Motion compensation using lighting superposition

By combining images captured under different illumination conditions and using motion compensation, the method addresses misalignment artifacts in three-dimensional reconstruction, achieving improved accuracy and reduced noise in shape-from-shading techniques.

JP2026514959APending Publication Date: 2026-05-13GELSIGHT INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GELSIGHT INC
Filing Date
2024-04-24
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing three-dimensional reconstruction techniques suffer from misalignment artifacts due to camera movement when capturing images under different illumination conditions, particularly in high-resolution reconstruction using handheld scanners, leading to errors in shape-from-shading methods.

Method used

The method employs a series of images captured under different illumination conditions, combined with a composite illumination image, and uses motion compensation and illumination multiplexing techniques to align and reconstruct a super-resolution or noise-reduced three-dimensional model by minimizing a cost function based on pixel differences.

Benefits of technology

This approach effectively reduces noise and enhances the accuracy of three-dimensional reconstruction by aligning images captured under varying illumination conditions, improving the precision of shape-from-shading techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514959000001_ABST
    Figure 2026514959000001_ABST
Patent Text Reader

Abstract

Using the superposition principle of linear systems, a series of images of a surface captured under different illumination conditions (e.g., different patterns or illumination directions) can be combined with an additional composite illumination image captured while illuminating the surface under all of the constituent illumination conditions (e.g., simultaneous directional illumination from all directions, or simultaneous illumination using multiple different illumination patterns). Additional images may also be captured under various combinations of illumination conditions and used in conjunction with illumination multiplexing techniques to obtain super-resolution or noise-reduced images of the surface. These individual super-resolution images can then be used to obtain a super-resolution or noise-reduced three-dimensional reconstruction based on the improved source image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the priority of U.S. Provisional Patent Application No. 63 / 461,463, filed on April 24, 2023, the entire content of which is incorporated herein by reference.

[0002] The present disclosure generally relates to three-dimensional reconstruction, and more specifically, to techniques for aligning a plurality of images of a target surface captured under different illumination conditions.

Background Art

[0003] Shape-from-shading techniques and similar three-dimensional reconstruction techniques enable the restoration of three-dimensional surface information from a number of images of a target surface captured under different illumination conditions. However, when images are captured sequentially over time, camera movement may cause misalignment between images, thus generating misalignment artifacts such as ghosts, blurring, and smearing of three-dimensional feature elements, thereby degrading the accuracy of the restored three-dimensional data. This problem can be particularly severe in imaging applications such as high-resolution reconstruction using images from a handheld scanner, where the time interval between images is sufficiently long (e.g., longer than 50 milliseconds) and the target resolution is sufficiently small (e.g., 20 microns or less), and camera shake becomes a major cause of three-dimensional reconstruction errors.

[0004] There is still a need for motion compensation and improved image alignment for shape-from-shading and similar three-dimensional reconstruction to address misalignment in a series of temporally separated source images used for three-dimensional reconstruction.

[0005] Summary Using the superposition principle of linear systems, a series of images of a surface captured under different illumination conditions (e.g., different illumination patterns or directions) can be combined with an additional composite illumination image captured while illuminating the surface under all of the constituent illumination conditions (e.g., simultaneous directional illumination from all directions, or simultaneous illumination using multiple different illumination patterns). Additionally, under various combinations of illumination conditions, additional images may be captured and used in conjunction with illumination multiplexing techniques to obtain super-resolution or noise-reduced images of the surface. These individual super-resolution images can then be used to obtain a super-resolution or noise-reduced three-dimensional reconstruction based on the improved source image.

[0006] The computer program products described herein include computer executable code embodied in a persistent computer-readable medium, which, when executed on one or more computing devices, causes one or more computing devices to perform the following steps: capturing a first image of a surface while illuminating the surface from a first direction; capturing a second image of the surface while illuminating the surface from a second direction; capturing a third image of the surface while illuminating the surface simultaneously from the first and second directions; and aligning the first image to the second image by applying a motion model for the third image to at least one of the first and second images, thereby aligning the first image to the second image according to the motion model, while minimizing a cost function representing the difference between the third image and the sum of the first and second images.

[0007] In some manifestations, the cost function is based on the difference in pixel values. In some manifestations, the cost function evaluates a subset of pixels. In some manifestations, the cost function is based on an image similarity index. In some manifestations, the cost function is based on a normalized correlation coefficient. In some manifestations, the step involves reconstructing the three-dimensional shape of the surface by shape-from-shading based on the first and second images, when aligned according to consistency. In some manifestations, minimizing the cost function involves minimizing the difference between the pixel values ​​at one or more pixel positions in the first pixel array for the third image and the sum of the first and second images at one or more corresponding positions.

[0008] In some embodiments, the motion model includes one or more rigid body motion models of rigid body translation and rigid body rotation. In another embodiment, the motion model may include a rigid body motion model for the motion of an image induced by six degrees of freedom in the orientation of an imaging device that captures a first image, a second image, and a third image. In some embodiments, the motion model uses independent motion tracking for one or more sub-regions of the first image, a second image, and a third image. In some embodiments, the motion model uses one or more visible criteria for tracking image differences. In some embodiments, matching includes matching downsampled instances of the first and second images and calculating motion parameters to match the first and second images by scaling up the motion parameters from the downsampled instances to the scale of the first and second images. In some embodiments, matching includes recursively downsampling, matching, and scaling motion parameters with respect to two or more downsampled versions of the first, second, and third images. In some embodiments, aligning involves dividing the respective pixel arrays of the first, second, and third images into multiple regions, and selecting one or more pixel locations from each of the multiple regions to evaluate the cost function. In some embodiments, aligning involves selecting a subset of pixel locations in the respective pixel arrays of the first, second, and third images to minimize the cost function. In some embodiments, selecting a subset of pixel locations involves selecting at least one subset of pixel locations based on the magnitude of the cost function in one of the subsets of pixel locations between the third image and the sum of the first and second images.

[0009] The method disclosed herein includes capturing two or more images of a surface, each of which is captured under two or more different lighting conditions; capturing a composite image of the surface while it is simultaneously illuminated under each of the two or more different lighting conditions; and aligning the two or more images by applying a motion model while minimizing the image difference between the composite image and the sum of the two or more images.

[0010] Two or more different lighting conditions may include two or more different lighting directions. Two or more different lighting conditions may include two or more different lighting wavelengths. Two or more different lighting conditions may include two or more different lighting patterns. Minimizing image difference may involve minimizing a cost function that represents the difference between the composite image and the sum of two or more images. Aligning two or more images may involve aligning two or more images using a multi-image alignment algorithm. Aligning two or more images may involve minimizing an optimization function. Two or more images may include three images. Two or more images may include six images. Aligning two or more images may involve aligning two or more images of a first group with each other in a first image alignment, aligning two or more images of a second group with each other in a second image alignment, and aligning the first image alignment with the second image alignment in a third image alignment. The method may involve calculating initial estimates of the displacement of the motion model based on input from an inertial measurement device. The method may include calculating an initial estimate of the motion model's displacement based on one or more visible criteria in each of two or more images. The method may also include calculating an initial estimate of the motion model's displacement based on evaluation by a machine learning model trained to associate one or more predetermined displacements with one or more visual artifacts in a combination of images illuminated under two or more different lighting conditions.

[0011] The system described herein includes a retrographic sensor comprising a deformable medium along with a detection surface, the deformable medium being formed from an optically transparent and deformable material, and the detection surface covering a portion of the deformable medium and providing a reflective surface visible through a second surface of the deformable medium. The system also includes a camera positioned to capture an image of the reflective surface through the second surface of the deformable medium. The system also includes an illumination system configured to provide directional illumination of the detection surface by independently illuminating the detection surface through the deformable medium from each of three or more directions around the optical axis of the camera. The system also includes a processing circuit, which controls the camera and lighting system to capture an image of the detection surface with a camera while the detection surface is individually illuminated from three or more directions, to provide three or more images of the detection surface; controls the camera and lighting system to capture a composite image of the detection surface while it is illuminated simultaneously from all three or more directions; aligns the three or more images by applying a motion model while minimizing the image difference between the composite image and the sum of the three or more images, thereby aligning the three or more images, and reconstructs the three-dimensional shape of the detection surface using shape-from-shading.

[0012] In another embodiment, the lighting system may be configured to illuminate a reflective surface under different lighting conditions, such as different lighting patterns or different wavelengths, which can be used in place of or in addition to directional lighting to extract three-dimensional data, along with a number of sequential images.

[0013] The processing circuit may include a controller for the imaging system, which may include a retrographic sensor, a camera, and an illumination system. The processing circuit may include cloud computing resources configured to receive three or more images and a composite image, align the three or more images, and reconstruct the three-dimensional shape of the detection surface using shape-from-shading and alignment of the three or more images. The system may include a substrate for the retrographic sensor, on which a deformable medium is placed, and the substrate is formed from a rigid, optically transparent material that mechanically supports the deformable medium, and the substrate is placed between the deformable medium and the camera. The processing circuit may be configured to acquire one or more additional images under different illumination conditions and to illuminate multiplex decoupling the three or more images, the composite image, and one or more additional images to obtain surface normal values ​​from the detection surface at a resolution higher than the camera's nominal resolution. In this regard, as understood, the term “illumination multiplex decoupling” is intended to mean processing that occurs after image acquisition. Generally, images may be multiplexed and multiplexed / decoupled. For example, in this case, the image may be multiplexed during image acquisition and then multiplexed / decoupled during image processing, and the system may use one or both of these techniques for three-dimensional reconstruction. Alternatively, the processing circuit may be configured to acquire one or more additional images under different lighting conditions and to multiplex the three or more images, the composite image, and one or more additional images under different lighting conditions to reduce pixel noise in the three or more images and the composite image.

[0014] Embodiments of the devices, systems, and methods described herein are shown in the following drawings. The drawings are not necessarily to scale and are more focused on illustrating the principles of the disclosure. [Brief explanation of the drawing]

[0015] [Figure 1] This is a diagram showing the imaging system.

[0016] [Figure 2] It is a perspective view of a tactile sensor.

[0017] [Figure 3] It is a side view of the tactile sensor of FIG. 2.

[0018] [Figure 4] It is a diagram showing a robot system using a tactile sensor.

[0019] [Figure 5] It is a flowchart of a method for processing an image.

[0020] [Figure 6] It is a diagram showing an imaging system having a retro-graphic sensor.

[0021] [Figure 7] It is a diagram showing an imaging system.

[0022] [Figure 8] It is a diagram showing a cross-sectional view of an imaging system.

[0023] [Figure 9] It is a diagram showing a motion compensation technique.

[0024] [Figure 10] It is a flowchart of a motion compensation method.

[0025] [Figure 11] It is a diagram showing the addition result (total) of the images before and after alignment.

[0026] [Figure 12] It is a diagram showing a comparison of the composite image with the added image before alignment and the added image after alignment.

[0027] [Figure 13]This figure shows a three-dimensional surface rendered with and without surface normal maps and motion compensation.

[0028] Detailed explanation All references herein are invoked as a whole by reference. Unless explicitly stated otherwise or evident from the context, references to singular items should be understood to include plural items, and vice versa. Grammatical conjunctions are intended to represent any and all disjunctive and conjunctive combinations, such as coordinated clauses, sentences, and words, unless otherwise specified. Thus, the term "or" should generally be understood to mean "and / or," etc.

[0029] In this specification, descriptions of ranges of values ​​are not intended to be limiting unless otherwise indicated, but rather to refer individually to any and all values ​​that fall within the expressed range, and each distinct value within such range is incorporated herein as if it were individually enumerated. The words “about,” “approximately,” or similar, when accompanied by numbers, should be interpreted as indicating tolerances that would be recognized by those skilled in the art to operate satisfactorily for the intended purpose. Ranges of values ​​and / or numbers are provided herein merely as examples and do not constitute a limitation on the scope of the described embodiments. Any and all examples or use of exemplary language (such as “for example,” “like,” or similar) provided herein is intended merely to better illustrate embodiments and does not limit the claims unless expressly stated. The absence of words herein should be interpreted as indicating any element not described in the claims that is essential to the implementation of the disclosed embodiments.

[0030] In the following explanation, terms such as "first," "second," "top (summit)," "bottom," "up," and "down" are for convenience only and should not be interpreted as restrictive unless otherwise specified.

[0031] The devices, systems, and methods described herein may include, or may be used in conjunction with, the methods, systems, and devices described in U.S. Patent No. 10,965,854 issued on March 30, 2021, and International Patent Application No. PCT / US2022 / 046129 published on April 13, 2023. The entire contents of each of the above documents are incorporated herein by reference. In certain embodiments, the devices, systems, and methods described herein may be used to improve the alignment of multiple images used in shape-from-shading three-dimensional reconstruction, particularly when camera movement may introduce artifacts into a series of images. For example, the systems described herein may be useful for aligning a series of images captured, for example, by a robot end-effector or a handheld imaging system.

[0032] Figure 1 shows an imaging system. Generally, the imaging system 100 may be any system for quantitative or qualitative surface shape measurement and / or visualization, such as any of those described in the literature revealed above. The imaging system 100 may be used to derive quantitative data from an image, such as a surface normal map or height map of a three-dimensional surface shape. Alternatively, the imaging system may acquire other contact data, such as a force map, an elastic map, or other measure of the softness / hardness of a target surface. As to be understood, the term “imaging system” is used to describe some of the intended embodiments, but the tactile sensor may also be in a system that does not generate an image, for example, when raw sensor data is provided to a neural network or other machine learning system for decision-making without converting it into any image or quantitative surface reconstruction. All such rearrangements, combinations, or variations described above are intended to fall within the scope of this description and the scope of imaging systems as described herein, unless otherwise specified.

[0033] In one embodiment, the imaging system 100 may include a tactile sensor 102, which may include a removable and replaceable cartridge for the imaging system 100. The imaging system 100 may also include a fixture 104 for removablely and replaceably holding the tactile sensor 102. The fixture 104 may have a predetermined geometric configuration with respect to the imaging system 100 and / or to the imaging device 106, such as a camera and an illumination source 108 (such as one or more light-emitting diodes or other light sources), such that the tactile sensor 102 has a known position and orientation (positional relationship) with respect to the imaging device 106 and the illumination source 108 when fixed to the fixture 104. This forced geometric arrangement is advantageous in that it allows for the reuse of calibration data for the tactile sensor 102 and facilitates the reliable and repeatable placement of the tactile sensor 102 within the optical train of the imaging system 100. Furthermore, the fixed spatial relationship between the illumination source 108 and the imaging device 106 also provides useful constraints for specific alignment techniques and / or multiple illumination. Generally, as long as the illumination source 108 is stationary relative to the imaging device 106, a series of images captured by the imaging device 106 under different illumination conditions can be aligned with each other using the linear superposition described herein.

[0034] The following description focuses on the use of a tactile sensor 102 that is removable and replaceable to an imaging system 100, for example, configured as a cartridge or similar for modular and reusable use. Thus, the terms tactile sensor, cartridge, and imaging cartridge are sometimes used interchangeably in this specification. However, as will be understood, or perhaps instead, the tactile sensor 102 or a part thereof may be integrated into the imaging system in a way that it is generally not removable. Thus, the advantages of the systems and methods described herein may also apply to an imaging system 100 that does not include a removable tactile sensor 102 (for example, by permanently or indistinctly incorporating some or all components of the tactile sensor 102 into the main body of the imaging system 100), particularly in relation to superimposed lighting conditions. Alternatively, the portion of the tactile sensor 102, such as a rigid substrate, may be integrated into the main body of the imaging system 100, while other portions, such as those that contact the target surface or contain a fluid imaging medium, may be detachable and replaceable to allow for reuse of the imaging system 100 after the contact surface becomes contaminated or damaged with use.

[0035] The tactile sensor 102 may include an optical element 110 formed of at least a portion of a rigid, optically transparent material such as glass, polycarbonate, acrylic, polystyrene, polyurethane, optically transparent epoxy resin, or any other material having suitable mechanical and optical properties for use in the systems described herein. In this regard, and more generally, as to be understood when the term is used herein, “optically transparent” may mean transparent within the visible light range, or, conversely, transparent within or within a range of wavelengths of interest. Thus, for example, if imaging is performed in the infrared region, a “transparent” material will transmit most of the incident light in the infrared region. As another example, imaging may be channeled or multiplexed using different ranges of wavelengths for useful purposes, and the optical element 110 may be transparent to these aggregated wavelength ranges, or may include multiple components that are transparent in one or more different wavelength ranges. Also to be understood, “transparent” in this context means sufficiently transparent to capture an image. This can generally be understood as a transmittance of 90% or more with less than 10% double reflection absorption over the wavelength of interest. However, optical elements 110 with a transmittance of less than 90% may also be used, for example, if the optical element 110 transmits enough light to support imaging in the imaging system 100 with a resolution that satisfies the intended use, due to specific material or cost constraints.

[0036] The imaging system 100 may include an illumination source 108, such as one or more LEDs or other light sources, arranged to direct illumination toward the optical elements 110 and / or the deformable material layer 116. This may include LEDs of different wavelengths and luminances (intensities), as well as lenses, filters, or other photoshaping mechanisms for directing illumination from the LEDs to suit three-dimensional reconstruction as described herein. While LEDs provide a low-cost, narrowband illumination source, or alternatively, the illumination source 108 may include other light sources, such as fluorescent light sources, incandescent light sources, laser light sources, or optical fibers that direct external light sources to one or more suitable locations within the imaging system 100.

[0037] Generally, the optical element 110 may form a substrate for the tactile sensor 102, or it may be a window or similar within a larger mechanical substrate for the tactile sensor 102, i.e., in this case, a window made of an optically transparent material is embedded in another structure for mounting to a fixture 104 or other component of the imaging system 100. In one embodiment, the optical element 110 may be formed from a silicone such as hard platinum-cured silicone, or any other polymer of optical quality. The optical element 110 may have a first surface 112 including an area with an optically transparent surface for capturing an image through the optical element 110 by, for example, an imaging device 106. The optical element 110 may also have a second surface 114 opposite the first surface 112, in which case the central axis 117 passes through (through, penetrates) the first surface 112 and the second surface 114. The central axis 117 may, for example, coincide with or be parallel to the optical axis of the imaging system 100.

[0038] Generally, the first surface 112 may have optical properties suitable for transmitting an image (or more generally, light rays) from the second surface 114 to the imaging device 106 via the optical element 110. To support this function, the first surface 112 may include, for example, a curved portion that provides a lens for optically magnifying, focusing, or otherwise modifying the image from the second surface 114. For example, the first surface 112 may include an aspherical surface shaped to address spherical aberration or other optical aberrations in the image captured from the second surface 114 via the optical element 110. Alternatively, the first surface 112 may include a free-form surface shaped to reduce or otherwise mitigate geometric distortion in the image captured via the optical element 110. For example, imaging through a thick medium may generally lead to spherical aberration of a magnitude that depends on the numerical aperture of the imaging system 100 (or more specifically, the imaging device 106). Thus, the first surface 112 of the optical element 110 may be curved or otherwise adapted to address spherical aberration (and other higher-order aberrations) caused by the propagation of a focused beam of light through a thick medium. More generally, the first surface 112 may include any shape or surface treatment suitable for focusing, shaping, or modifying an image to support the capture of optical data through the optical element 110. Alternatively, the second surface 114 may be modified to improve image capture. For example, the second surface 114 of the optical element 110 may include a convex surface extending from the optical element 110 (e.g., toward the target surface 130 being imaged) to enlarge or otherwise shape the image transmitted from the target surface 130 to the imaging device 106. More generally, the first surface 112 may include any photoshaping features suitable for facilitating imaging as described herein, such as filters, lenses, focusing curves, and diffusers.

[0039] The optical element 110 may generally serve a number of purposes in the imaging system 100, as intended herein. In one embodiment, the optical element 110 functions as a rigid body for transmitting relatively uniform pressure across the target surface 130 when capturing an image. In particular, the body of the optical element 110 may apply substantially uniform or continuous pressure to the imaging medium so that the reflective coating on the opposite side of the imaging medium matches the surface shape of the surface being measured. In one embodiment, the optical element 110 may provide glazing angle illumination or shallow angle illumination from, for example, one or more illumination sources 108 located along its edges. Alternatively, the optical element 110 may provide directional dark-field illumination. For this purpose, sufficiently thick optical material may be used and may function as a light guide to provide controlled, uniform illumination and collimated or nearly collimated dark-field or glazing illumination of the reflective surface from different directions (e.g., when a single LED segment of the illumination source 108 is turned on) or from all directions (e.g., when all LED segments of the illumination source 108 are turned on). This configuration for glazing illumination may be useful, for example, when the illumination source 108 includes LEDs of different colors used for multiple light channels for multispectral illuminance difference stereo, where each color is associated with a specific illumination direction.

[0040] A layer 116 made of an optically transparent elastomer or other transparent and deformable material is placed on the second surface 114 and can be attached to the second surface 114 by any suitable means, such as any of those described herein. Generally, the layer 116 may be formed from a gel or other relatively flexible material that can be deformed to conform to the surface shape of the target surface 130, so that the complementary shape formed on the layer 116 can be optically captured through the opposite side of the layer 116. For example, an elastomer with a Shore 00 durometer value of about 5 to 60 may serve usefully as the layer 116 intended herein. Alternatively, other materials, including fluids, gels, and elastic polymers, may be used if they are sufficiently optically transparent with respect to imaging and sufficiently soft to conform to the shape of the target surface. In one embodiment, the first side 118 of the layer 116 adjacent to the second surface 114 of the optical element 110 may have a refractive index that matches the refractive index of the second surface 114. As recognized, the term “matching” as used herein in reference to refractive index does not necessarily mean having the same refractive index. Rather, the term “matching” generally means having refractive indices close enough to transmit (transmit) an image across the corresponding interface between two materials for capture by the imaging device 106. Thus, for example, acrylic has a refractive index of about 1.49, while polydimethylsiloxane has a refractive index of about 1.41, and these materials match well enough to be placed adjacent to each other and used to transmit an image between them with sufficient intensity and focus for quantitative or qualitative surface shape measurements as intended herein.

[0041] The second side 120 of layer 116 may be configured to shape-match with the target surface 130 while providing a surface facing the imaging device 106 that facilitates surface shape imaging and surface shape measurement by the imaging system 100. The second side 120 may include, for example, an opaque coating or a reflective coating, or more generally, any optical coating having a predetermined reflectivity suitable for supporting surface shape imaging as intended herein. Generally, this coating can facilitate the capture of an image that is independent of the optical properties of the target surface 130, such that surface properties such as color, translucency, gloss, and specularity do not interfere with light intensity measurement. In one embodiment, the second side 120 may include a convex surface extending away from the optical element 110 (for example, toward the target surface 130). This geometric configuration can offer several advantages, such as facilitating imaging of surfaces with large collective concave shapes and reducing the accumulation of bubbles in the field of view when the tactile sensor 102 is first placed in contact with the target surface 130.

[0042] The sidewall 122 may be formed around the interior 124 of the optical element 110, extending from the first surface 112 to the second surface 114. Generally, the sidewall 122 may include one or more photoforming feature elements configured to control the illumination of the second surface 114 through the sidewall 122, for example, from an illumination source 108. The sidewall 122 may have various geometric shapes useful for photoforming, for example, to direct and pass light into the optical element 110 at a desired angle and uniformity. For example, the sidewall 122 may include a continuous surface forming a frustoconical shape between two circles formed on the first surface 112 and the second surface 114. Alternatively, the sidewall 122 may include a truncated hemisphere between part or all of the region between the first surface 112 and the second surface 114. In another embodiment, the side wall 122 may include two or more separate planes configured to form regular or irregular polygonal shapes, such as hexagons or octagons (e.g., forming truncated hexagonal or truncated octagonal pyramids), around the central axis 117. In this latter embodiment with planes, the illumination source 108 may be formed by a number of light-emitting diodes adjacent to each plane, or adjacent to two or more planes. This configuration can usefully support multi-directional side illumination via the optical element 110. It should be understood that, in this regard, the planes may act as photoshaping mechanisms that refract or filter light rays and / or otherwise control illumination in a desired manner within the imaging volume of the imaging system 100.

[0043] Alternatively, other photoforming features may be incorporated into the sidewall 122, for example, to focus or direct incident light from the illumination source 108, or to control the reflection of light within the optical element 110 and / or the optically transparent elastomer layer 116. For example, one or more photoforming mechanisms may include diffusing surfaces for diffusing a point source of incident light along the sidewall 122. This allows for the diffusion of light from individual light-emitting diode elements in the illumination source 108 and / or provides a more uniform illumination field from the plane of the sidewall 122. Alternatively, the sidewall 122 may include polished surfaces for refracting incident light toward the optical element 110. As recognized, diffusing and reflective surfaces may also be used in various combinations to shape the illumination within the optical element 110. Alternatively, the sidewall 122 may include curved surfaces that form lenses within the sidewall 122, for example, to focus or direct incident light toward the optical element 110 as needed.

[0044] In another embodiment, the sidewall 122 may include a light-reducing filter with stepped attenuation to compensate for the distance from the sidewall 122. For example, to avoid excessive illumination of areas of the second surface 114 closer to the sidewall 122 and / or under-illumination of areas of the second surface 114 further away from the sidewall 122 (and / or closer to the central axis 117), the sidewall 122 may include a light-reducing filter that provides greater attenuation in areas of the sidewall 122 closer to the second surface 114 and less attenuation in areas of the sidewall 122 closer to the first surface 112. Thus, rays directly illuminating the second surface 114 at a downward angle adjacent to the sidewall 122 may be attenuated more than other rays emanating from the illumination source 108 toward the center of the second surface 114. This attenuation may be continuous, individual, or gradually change to provide attenuation that is generally greater as the otherwise guided light approaches the sidewall 122, or otherwise balance the illumination in the field of view.

[0045] In another embodiment, the photoforming feature element(s) may include one or more color filters that can be usefully used, for example, to correlate a particular color to a particular direction of illumination within the optical element 110, or otherwise to control the use of colored illumination from the illumination source 108. Also, if the imaging system uses wavelength division multiplexing imaging, the color filters on the sidewall may reduce stray light in the cartridge by selectively reflecting or transmitting a frequency range of interest. In another embodiment, the photoforming feature element may include a non-perpendicular angle of the sidewall 122 with respect to the second surface 114. For example, as shown in Figure 1, the sidewall 122 is inclined away from the second surface 114 so as to form an obtuse angle with the second surface 114. This technique may favorably support indirect illumination of the second surface 114 by, for example, total internal reflection without light on the first surface 112 and to the optical element 110. In another embodiment, the sidewall 122 may be inclined toward the second surface so as to provide an acute angle with the second surface, for example, to support greater direct illumination of the second surface 114. These techniques can be used individually or in combination to direct light into and through the optical element 110 as desired.

[0046] Alternatively, the photoforming feature(s) may include geometric feature(s), such as focusing lenses, non-planar regions, or similar elements, to direct incident light as desired. Alternatively, other optical elements may be formed on or within the sidewall 122. For example, the photoforming feature(s) may include optical films, such as any of the various commercially available films for filtering, attenuating, polarizing, or otherwise shaping incident light. Alternatively, the photoforming feature(s) may include a microlens array or similar for steering or focusing incident light from the illumination source 108. Alternatively, the photoforming feature(s) may include multiple micro-replicated and / or diffractive optical feature(s), such as lenses, diffraction gratings, or similar elements. For example, the sidewall 122 may include a microstructured sidewall 122, which may include, for example, a microimaging lens, a lenticular lens, and a microprism, as photoforming features for directing light from the illumination source 108 to the optical element 110 to improve variations in surface shape imaging relative to the imaging surface of the tactile sensor 102 on the second side 120 of the layer 116 made of an optically transparent elastomer. For example, the microstructured features facilitate the shaping of the illumination pattern to provide a uniform light distribution over the area to be measured and to reduce reflections of light returning into or out of the optical element 110. The microstructure may be provided, for example, during injection molding of the optical element 110 or by adding an optical film having the desired microstructure to the side. For example, suitable commercially available optical films for applying surface microstructures include Vikuiti® (registered trademark) and Advanced Light Control Film (ALCF), both sold by 3M.

[0047] As should be understood, or instead, other surfaces of the optical element 110 (such as the top surface (facing the image sensor 106) or the bottom surface (facing the target surface 130)) may include any of the aforementioned optical processing and structures that can be used alone or in combination to control illumination to, from, or within the optical element 110 in order to support imaging as described herein.

[0048] A mechanical key 126 may be positioned on the outer surface of the optical element 110 (more commonly, the tactile sensor 102) to force it into a predetermined position within the fixture 104 of the imaging system 100. The mechanical key 126 may include, for example, at least one radially asymmetric feature element around the central axis 117 to force a unique rotational orientation of the optical element 110 within the fixture 104 of the imaging system 100. More commonly, the mechanical key 126 may include any number of mechanical elements or similar appropriate for holding the optical element 110 in a predetermined positional relationship and / or position within the imaging system 100. Alternatively, the mechanical key 126 may include a matched geometric shape between the optical element 110 and the fixture 104. For example, the mechanical key 126 may include a cylindrical structure extending from the optical element 110, and / or the mechanical key 126 may include an elliptical prism or similar, which may usefully enforce both position and rotational orientation.

[0049] In one embodiment, the mechanical key 126 may include one or more magnets 128 that may secure the optical element 110 to the fixture 104 of the imaging system 100. The one or more magnets 128 may be further encoded via positioning and / or polarity to ensure that the optical element 110 is inserted only in a specific rotational orientation around the central axis 117. Alternatively, the mechanical key 126 may include a plurality of protrusions, at least one of which has a different shape from the others of the plurality of protrusions in order to force a unique rotational orientation of the optical element 110 around the central axis 117 within the fixture 104 of the imaging system 100. Alternatively, the mechanical key 126 may include at least three protrusions (e.g., exactly three protrusions) that are shaped and sized to form a kinematic coupling with the fixture 104 of the imaging system 100. Alternatively, the mechanical key 126 may include feature elements such as flanges, dovetail joints, or any other mechanical shapes(s) or features(s) to firmly anchor the optical element 110 to the fixture 104 in a predetermined position and / or orientation. More generally, any combination and mechanism for mechanically anchoring feature elements may be used to provide a mechanical key 126 that physically positions the tactile sensor 102 in the correct orientation within the imaging system 100.

[0050] The surface of the tactile sensor 102 may be further treated as needed or to be helpful when acquiring an image with the imaging system 100. For example, the top, side, and bottom areas of the optical element 110 or other parts of the tactile sensor 102 may be covered with a light-absorbing layer, such as black paint, to contain light from the illumination source 108 or to reduce the intrusion of ambient light.

[0051] One challenge in fixing a flexible elastomer (within layer 116) to a rigid surface, such as an optical element 110, is delamination, which can occur due to shear forces and other edge effects after repeated compression, decompression, and shearing of layer 116 during image capture, particularly where the target surface 130 tends to adhere to the elastomer. To address this problem, the optical element 110 and the transparent elastomer layer 116 may be formed as a cartridge provided to the end user as an integrated, detachable, and replaceable device. In embodiments, this cartridge may be removed and replaced by the end user, for example, to change to a tactile sensor 102 with different optical properties for different imaging applications or resolutions, or to replace a worn or damaged tactile sensor 102. Simultaneously, replacing the optical element 110 and layer 116 together allows for a more robust attachment of the elastomer layer 116 to the optical element 110 compared to a configuration in which the user manually replaces only the elastomer layer 116.

[0052] Figure 2 is a perspective view of a tactile sensor. The tactile sensor 202 may, for example, have a generally rectangular configuration and may include one or more flanges 204 or similar so that the tactile sensor 202 can linearly slide-engage with a housing fixture. This type of engagement mechanism may be particularly advantageous for robotic applications or similar when the tactile sensor 202 can be removed from and replaced by an end effector of a robot handler. The tactile sensor 202 may be any of the tactile sensors described herein. A deformable layer 206, such as any of the fluid layer, elastomer layer or other deformable layer described herein, may provide an optically transparent medium, a film 208 for contact with a target surface, or a substrate 210 for mechanical support.

[0053] Figure 3 is a side view of the tactile sensor 202 shown in Figure 2.

[0054] Figure 4 shows a robotic system using a tactile sensor. Generally, the system 400 may include a robot handler 402 having a housing 404 on its end configured to detachably and replaceably receive a tactile sensor 406, such as one of the cartridges or other tactile sensors described herein or other optical devices. Generally, the robot handler 402 may include any robotic component or combination of components suitable for positioning and manipulating objects. For example, the robot handler 402 may include a robotic arm, gantry, SCARA robot, Cartesian robot, delta arm, or any combination thereof or other position controller, along with appropriate sensors, actuators and the like to control its movement. The robot handler 402 may also include any suitable manipulator, grasping device, end-effector, or similar for grasping (grasping) or otherwise handling and manipulating objects. The system 400 may also include a processor or other controller or similar for providing a programmable interface or user interface to control the operation of the robot handler 402.

[0055] The robot handler 402 may be configured to position the tactile sensor 406 in contact with the target surface 408 in order to capture a surface shape image of the target surface 408 using, for example, a camera or other imaging device located in the housing 404. As recognized, the components of such imaging device may generally be located within the housing 404, or remotely arranged and optically coupled to the tactile sensor 406 by, for example, an optical fiber or similar, or a combination thereof. In one embodiment, the system 400 may consist of, for example, computer executable code stored in the memory of the system 400, executed by the processor of the system 400, to automatically detach the tactile sensor 406 from the fixture of the system 400 (e.g., within the housing 404) and insert a second tactile sensor 410, which includes a replacement sensor, into the housing 404. The second tactile sensor 410 may be identical in structure and function to the tactile sensor 406, for example, to allow for replacement after normal wear and tear, or the second tactile sensor 410 may have a different optical configuration from the tactile sensor 406 to provide, for example, greater magnification, a larger field of view, better resolution of feature elements, deeper illumination of feature elements, and different aggregate surface shapes. The second tactile sensor 410 may be housed in a container or other container accessible to the robot handler 402 of the system 400. Generally, the system 400 may include one or more magnets, electromechanical latches, and actuators, etc., within the housing 404 or more generally within the system 400, to facilitate the removal and replacement of the tactile sensor 406 as described herein. More generally, the system 400 may include any gripping device, clamp, or other electromechanical end-effector, or similar, suitable for removing and replacing the tactile sensor 406 for use in the imaging process and for positioning the tactile sensor 406.

[0056] In one embodiment, the robot handler 402 may be operated manually by a human technician from a console or similar device. Alternatively, the robot handler 402 may be programmed to operate automatically, for example, in a test facility or manufacturing facility. In this regard, the robot handler 402 may, for example, use a sensing network, machine learning algorithms and other techniques to automatically position the tactile sensor 406 on the workpiece of interest, and may control parameters such as contact force, fluid pressure, temperature, or other parameters in preparation for measurement and / or during acquisition of imaging data. After appropriate positioning, the robot handler 402 may control an imaging system (either in or accessible from the housing 404) to acquire data for three-dimensional reconstruction of the target surface of the workpiece. This general technique may be used, for example, for part inspection and measurement.

[0057] In another embodiment, the robot handler 402 may use tactile feedback to guide decision-making. For example, the robot handler 402 may determine whether a workpiece meets certain physical requirements and then classify the workpiece as acceptable, unacceptable, and / or uncertain (e.g., requiring manual inspection). In another embodiment, the robot handler 402 may use tactile feedback from the tactile sensor 406 to control the gripping force of a robot hand, gripping device or other end-effector or similar, or to control the amount of instantaneous contact force, torque or similar applied to a workpiece being manipulated by the robot handler 402. In another embodiment, the tactile sensor 406, which may include an array of tactile sensors, may be used to produce a visualization of a contact force field, pressure field, or surface topology that can be presented to the operator on the display 410 to assist the human operator in controlling the robot handler 402's actions toward a workpiece. In another embodiment, the tactile sensor 406 may be used to quantify a contact force field, pressure field, or surface topology for use by the robot handler 402 in automatic adjustment of grip strength (grip force) and grip orientation, or in determining whether there is a sufficient grip to initiate the next step (e.g., moving the workpiece).

[0058] The system 400 may include a computing device 414 for processing data from the tactile sensor 406, controlling the operation of the robot handler 402, and providing a user interface to the robot handler 402. For example, the computing device 414 may consist of code stored in memory and executed on the processor of the computing device 414 for purposes such as identifying an object or surface contacted by the tactile sensor 406 of the robot handler 402, generating alerts to the user based on tactile feedback obtained from the tactile sensor 406, and determining actions for a workpiece in contact with the tactile sensor 406 (including user-recommended decisions and decisions automatically performed by the robot handler 402). In one embodiment, the code may utilize, for example, machine learning models or similar for identification, decision-making, and other intelligent sensing and / or data-driven actions.

[0059] Figure 5 shows a flowchart of a method for processing images. By using the superposition principle of linear systems, a series of images of a surface captured under different lighting conditions can be aligned (aligned) with each other based on an additional composite lighting image captured while the surface is illuminated simultaneously under all different lighting conditions. For example, when using directional illumination, this may involve illuminating the surface individually from each direction to obtain a series of directional lighting images, and then providing directional illumination from all directions simultaneously while simultaneously capturing a composite image of the surface. Alternatively, additional images may be acquired under various combinations of lighting conditions, such as various combinations of directional illumination, and used in conjunction with illumination multiplexing techniques to obtain a noise-reduced image of the surface. Or, instead, these same techniques may be used to obtain a super-resolution image based on the difference between the source images of the images to be aligned. The aligned images may then be used in conjunction with shape-from-shading, structured lighting, and illuminance difference stereo, etc., to obtain the corresponding noise reduction and / or super-resolution three-dimensional reconstruction.

[0060] Method 500 can generally be carried out using the apparatus and other methods described herein. For example, Method 500 can be carried out using an imaging system including a tactile sensor. In one embodiment, the imaging system may include a handheld imaging device as described herein, and motion compensation using the techniques described herein can be advantageously used to address camera shake associated with manually acquired images. Method 500 can be carried out, for example, by computer-executable code stored in a persistent computer-readable medium, which, when run on one or more computing devices, causes one or more computing devices to perform the steps described below.

[0061] As shown in step 502, method 500 may include capturing images of a surface under a number of different illumination conditions. For example, this may include capturing images during directional illumination from two or more different directions. For example, a first image of the surface may be captured while the surface is illuminated from a first direction, and a second image of the surface may be captured while the surface is illuminated from a second direction. More generally, a number of directional illumination images may be captured, for example, to facilitate de-obfuscation of surface normals (for shape-from-shading), to address occlusion, and otherwise to support accurate reconstruction of three-dimensional surface data. Thus, three or more images may be captured during illumination from different directions, in which case each image is captured during illumination from one of the different illumination directions. For example, this may include capturing six images illuminated from different side directions, which has been demonstrated to support high-resolution three-dimensional reconstruction based on surface normals. In some embodiments, a minimum of three images of the surface are captured while the surface is illuminated from different directions, and in some embodiments, seven or more images may be used.

[0062] In another embodiment, two or more different illumination conditions may include two or more different illumination wavelengths. That is, capturing two or more images of a surface under a number of different illumination conditions may include capturing an image of the surface under illumination at two or more different illumination wavelengths. As understood, illumination at different wavelengths may be used instead of, or in addition to, illumination from different directions. In another embodiment, two or more different illumination conditions may include two or more different illumination patterns. Thus, structured or patterned illumination may be created, for example, using lenses, filters, diffractive optical elements, or other techniques, and three-dimensional surface information may be extracted using these illumination patterns. In this case, a series of different illumination patterns may be used, and then all of the different illumination patterns may be projected at once to obtain a composite image. Structured lighting may include any suitable lighting pattern(s) such as dots (including random and / or ordered dots), lines, and sinusoidal patterns (e.g., a given spatial frequency, polygons, circles, textures), any of which may be applied from one or more different directions in various orders to facilitate three-dimensional reconstruction. As described herein, the principle of superposition can be appropriately utilized to align multiple images captured under these different lighting patterns.

[0063] As shown in step 504, method 500 may include capturing an image of a surface while simultaneously illuminating it under each of two or more different lighting conditions. For example, in the case of directional illumination, this may include capturing a composite image while illuminating the surface from each of the different directions used in the directional illumination image of step 502. More generally, the composite image may include images captured while simultaneously illuminating from different directions, in different patterns, and / or at different wavelengths. A single composite image may be used to accommodate movement between images, but additional composite images may also be used, for example, when each composite image is associated with a different combination of lighting conditions, as recognized. Thus, for example, if there are six directional light sources, a first composite image may be captured while illuminating with one group consisting of three of the six light sources, and a second composite image may be captured while illuminating with the other group consisting of three of the six light sources. Finally, a third composite image may be captured using illumination from all directions to allow alignment across the first and second composite images, for example. This common technique of capturing multiple composite images under different lighting conditions may be used more commonly, for example, to facilitate hierarchical alignment, to avoid local minimums during optimization, and to increase the robustness of the alignment process. As will be recognized, in this specification, a composite image is sometimes referred to as a “third” image, but this is not intended to suggest a particular order of image acquisition, and a composite image may be captured before, during, or after acquisition of other images under a single lighting condition, and / or combined with the acquisition of other image sets. Also, as will be understood, if the movement between images becomes too large, the superposition may no longer result in a meaningful combination of images.Therefore, the time between a single-condition image and a corresponding composite image can be constrained to facilitate the use of superposition for extracting an image with improved resolution, and / or physical motion can be measured and used as a threshold for image matching based on superposition.

[0064] As shown in step 506, method 500 may include obtaining an initial alignment estimate. The initial alignment estimate may be obtained, for example, using any suitable supplemental data source(s). In one embodiment, method 500 may include calculating an initial estimate of the displacement of the motion model based on input from an inertial measuring device (such as an inertial measuring device coupled to a handheld sensor) or other sensors, devices, or systems that enable spatial tracking of the device. Alternatively, image alignment may include calculating an initial estimate of the displacement of the motion model based on evaluation by a machine learning model trained to estimate displacement based on visual artifacts in a combination of images illuminated from two or more corresponding lighting directions. For example, this model may be trained by capturing a series of images having various given displacement patterns, and then training the machine learning model to classify or predict the nature of the displacement based on visual artifacts in a composite image formed from them. In general, providing initial estimates, especially high-quality estimates, can speed up the optimization process by helping to start from a state close to ideal alignment and avoid local minimums that may mislead and / or lead to false results in the optimization process. Another advantage is that measured initial estimates can help ensure that the images are sufficiently close in alignment to support the assumption of linear superposition. Also, as is understood, if measurements are not available for initial estimates, initial estimates can nevertheless be provided based on, for example, patterns or history of misalignment between images, relating to a particular device (apparatus), a particular user, and a particular image application.

[0065] Alternatively, a reference may be used to obtain an initial fit estimate. For example, Method 500 may include calculating an initial estimate of the displacement of a motion model based on one or more references visible in each of two or more images. In one embodiment, this may include a pattern of references visible in a particular spectral band, along with a camera having pixels that are sensitive in a particular spectral band (and possibly only in that spectral band). Thus, for example, this may include creating a pattern on the contact surface of a sensor using an infrared absorbing or infrared reflective pigment, and then capturing an image of the contact surface in the infrared spectrum. In this way, some or all of the images, including individually illuminated images and a composite image, can be decomposed into an infrared image containing references for initial fit to the other images, and another image (e.g., a red, green, or blue image) for use in three-dimensional reconstruction. In another embodiment, a separate optical channel (e.g., using wavelengths outside the range for reconstruction) may be used to perform a rough alignment of the references for the initial motion estimate.

[0066] As shown in step 508, method 500 may include aligning images. Generally, images captured during surface illumination under different conditions (e.g., a first image and a second image) are aligned with each other using the principles of composite image and superposition, so that the images are better aligned with each other before calculating surface normals at locations across the surface using luminance (intensity) data in the images, or otherwise before performing calculations for three-dimensional reconstruction based on the images. Aligning images may include applying an initial alignment based on estimates obtained in step 506.

[0067] To match directional illumination, the initially directionally illuminated images can be added together to form a single image. Assuming the imaging system has a linear response, the position-by-position (e.g., pixel-by-pixel) sum of the directionally illuminated images should be equal to the brightness (intensity) of the composite image at each corresponding position. In practice, the exposure or illumination time of one or more images may be shortened (or extended as needed) to maximize the pixel dynamic range of the single-illuminated image while avoiding saturation of the multiplexed image. In this case, a pixel intensity multiplier may be used to compensate for the difference in exposure time, and / or each image may be assigned a multiplier or weight to allow for a direct combination of pixel intensities between the images. Thus, generally, a simple sum of pixel intensities may be used, or a scaled sum may be used to account for differences such as exposure time between images. Generally, the difference between the composite image and the sum of the directionally illuminated images (or the scaled sum) provides a basis for identifying positional misalignments between the directionally illuminated images. By using a motion model to translate (translate) the images relative to each other or move them in other ways, correct matching can be achieved while optimizing the cost function that measures the similarities (or differences) with the composite image.

[0068] Although directional illumination is described herein, as can be understood, other types of illumination conditions, such as structured light or wavelength multiplexing, may be used to provide illumination under different conditions, and then provide simultaneous illumination for the acquisition of a composite image. Thus, for example, in the case of structured light, the images individually illuminated under each illumination pattern may be added together and compared to a composite image illuminated under all the illumination patterns used for the individual images. Assuming that the imaging system has a linear response, the sum of images specifically illuminated at each position (e.g., per pixel) should be equal to the luminance of the composite image at each corresponding position.

[0069] In one embodiment, the matching process may include preprocessing of image data to support more efficient optimization. For example, this may include reducing the resolution of the matching process by, for example, downsampling the image for initial matching and then upsampling the resulting motion parameters to provide more accurate initial estimates for matching the full-resolution image. In another embodiment, the preprocessing of image data may include selecting a subset of pixels for matching based on, for example, texture and position.

[0070] In general, optimization for selecting the correct alignment position of two or more images can be performed by moving the image images relative to each other using a given motion model. This may include, for example, a rigid motion model that applies rigid translation and / or rigid rotation to the image. Alternatively, this may include a deformable motion model that adapts to local deformations across the alignment region, etc. For example, when imaging with a retrographic sensor having a shape-matchable contact surface, the contact surface may stretch, slide, or otherwise move in a non-rigid manner while slowly conforming to the shape of the target surface. In this case, a deformable motion model can be usefully used to account for non-rigid changes along the contact surface between the sensor and the target. In another embodiment, the motion model may use any other technique suitable for independent motion (movement) tracking for one or more sub-regions of various captured images, or for tracking specific surface positions between images. More generally, any motion model suitable for characterizing motion that may occur during imaging, for example, due to camera shake or other undesirable motion sources, may be used for motion compensation as described herein.

[0071] In one embodiment, the selection of a motion model may be optimized to suit a specific measurement situation. For example, the contact-based tactile sensors described herein are generally applied perpendicular to the target surface. When this is in effect as a constraint, the dominant motion is generally constrained to three degrees of freedom, namely x-translation, y-translation, and z-rotation, although z-translation, x-rotation, and y-rotation tend to be physically constrained and limited. In this case, a appropriately simplified rigid motion model may be usefully deployed to simplify the matching calculation. However, in imaging techniques using perspective projection lenses where contact or perpendicularity is not enforced, appropriate compensation may require a rigid motion model that takes into account additional degrees of freedom, possibly three translational degrees of freedom and three rotational degrees of freedom of the imaging device. Thus, in one embodiment, motion compensation as described herein may use a rigid motion model having six degrees of freedom, such as a rigid motion model that includes motion (movement, action) induced by six degrees of freedom with respect to the orientation of the imaging device that captures individual and composite images.

[0072] As should be understood, while rigid motion models are emphasized in the previous paragraph, other motion models may be used, or may be used instead. For example, motion models may include non-rigid or deformable motion models to facilitate image alignment, for instance, when the target surface or the contact surface between the target surface and the sensor is flexible or deformable. While deformable motion models tend to be more computationally intensive, techniques such as finite element methods, mesh-based techniques, and statistical models can provide improved accuracy when a deformable model is appropriate, especially when the assumption of rigid motion is inaccurate or unreliable. This may be useful, for example, when imaging human skin or tissue, or when imaging non-rigid objects such as elastic foam and soft rubber. It may also be appropriate when the target surface has non-uniform surface friction that could cause local shear deformation or other deformation in the sensor. More generally, various motion models may be adapted for use according to the characteristics of the actual or expected target surface from which the image is being captured, as known in the technique.

[0073] Generally, alignment can be performed by iterating over motion parameters using a selected motion model while evaluating the superposition of different images. A cost function may be provided to evaluate the success (or failure) of a particular image displacement. The cost function may generally evaluate the similarity or difference between two images and may provide an objective measure of whether the images are better aligned (e.g., the composite image is more similar to the sum of the other images) or poorly aligned (e.g., the composite image is less similar to the sum of the other images). The cost function may be based on the difference in pixel values. In one embodiment, the cost function may evaluate a subset of pixels, which may favorably reduce the complexity of the process. Alternatively, the cost function may be based on an image similarity index (e.g., the closeness of values ​​between brightness (intensity) at the image locations). In another embodiment, the cost function may use the correlation coefficient of the two images to measure similarity, for example, based on normalized and / or windowed images. Thus, in one embodiment, the cost function may be based on the normalized cross-correlation of the two images. More generally, minimizing the cost function may involve minimizing the difference between the pixel values ​​at one or more pixel locations in the first pixel array for the composite image and the sum of the corresponding one or more pixel locations in the second pixel array for each of the many separate non-composite images captured in step 502.

[0074] In another embodiment, processing efficiency can be improved by selecting a subset of pixels for each alignment / optimization. Numerous methods can be used to select pixels. In one embodiment, this may involve subdividing the image into regions (e.g., on a grid) and then selecting specific or random pixels for use in alignment. However, this simple technique can be greatly improved by selecting specific pixels / regions of interest. For example, this may involve selecting one or more pixels that have a high cost function with respect to the sum of unaligned or initially aligned images for the composite image, e.g., those containing important information for misalignment. In another embodiment, the image may be divided into sub-regions, and for each sub-region, one or more high-cost pixel locations may be selected. Accordingly, alignment may involve dividing the pixel array of each captured image into multiple regions and selecting one or more pixel locations from each of the multiple regions for evaluating the cost function.

[0075] Alternatively, pixel selection may involve selecting pixels and / or the size or shape of the area in which pixels are selected, based on gradients (slope, inclination, slope), the strength of local textures, the presence of a reference, or any other suitable indicator (e.g., measuring the amount and type of change within an area or at a pixel location). Such local textures may be characterized by how clearly defined and identifiable the texture details are, and the presence of a strong local texture can improve the accuracy of the three-dimensional reconstruction from the underlying two-dimensional image, essentially containing more information about misalignment. Conversely, pixel values ​​in areas with weak or insufficient textures may not be sufficiently distinguishable, potentially impairing the reconstruction. Thus, selecting areas and / or pixels associated with a strong texture can improve alignment by providing a greater signal to the cost function used to align the image. Numerous metrics for measuring the intensity of a texture are known in this technique, such as entropy, contrast, homogeneity, correlation, and gradient, and any of these metrics may be used to measure the intensity of a texture for the purpose of selecting areas and / or pixels for use in alignment.

[0076] In accordance with the above, or alternatively, image matching may involve selecting a subset of pixel locations in the pixel array for each captured image in order to minimize the cost function, and selecting a subset of pixel locations may involve selecting at least one of the subsets of pixel locations based on the magnitude of the cost function in one of the subsets of pixel locations between the composite image and the sum of the images captured in step 502. In another embodiment, image matching may involve selecting a subset of pixel locations based on a metric (indicator) relating to the intensity of the texture.

[0077] Optimization may be performed at a given location using appropriate metrics (e.g., initial fit estimates, pixel selection, motion models, cost functions), which attempts to optimize the cost function by repeatedly moving (non-composite) images relative to each other, adding the images (and scaling them, if necessary or useful, to account for variations in exposure or illumination time between images, for example), and comparing the pixel values ​​of the added images with the corresponding pixel values ​​of the composite image according to the cost function. As recognized, optimization is generally considered a minimization problem herein, but the nature of the optimization depends on the nature of the underlying cost function used to perform the optimization. Thus, for example, if the cost function evaluates the difference between images (e.g., the result of adding a composite image versus an image), the optimization will be minimal. However, if the cost function measures similarity and larger quantitative values ​​are assigned to higher image similarity, the optimization will generally be a maximization of the cost function. Any such technique, or combination thereof, may be used to perform image fit optimization as described herein.

[0078] Accordingly, in one embodiment, the alignment may include applying a rigid body model to two or more images to transform the images relative to each other, while minimizing a cost function that represents the difference between the composite image and the sum of the individual directional illumination images from step 502, thereby providing alignment that aligns the images according to a motion model.

[0079] In one embodiment, the alignment process can also be broken down into numerous separate alignments of different groups of images. For example, in a three-dimensional reconstruction using six different directional illuminations, a first group consisting of three illumination sources may be activated, and a composite image may be captured for these three. Then, a second group of illumination sources may be activated, and a second composite image may be captured for these three. Each of these image sets (three light source images) can be aligned to a corresponding composite image. Finally, a third composite image may be captured and used to align the images of the first set (single illumination) with the images of the second set (single illumination images), for example, by using the third composite image for a pair of directional illuminations spanning the first and second groups. Naturally, this generally requires capturing additional composite images for different combinations of directional illuminations. However, if the computational complexity of the six-image alignment is high, an overall improvement in processing speed may be achieved even considering the additional image acquisition and multiple sequential alignments. Furthermore, this may increase the robustness of motion compensation by introducing redundancy between the individual images and the composite image used for alignment. Accordingly, image alignment may include aligning a first group of images to one another in a first image alignment, aligning a second group of images to one another in a second image alignment, and aligning the first image alignment to the second image alignment in a third image alignment based on an additional composite image captured under lighting conditions including the lighting used for at least one image from the first group and at least one image from the second group.

[0080] In another embodiment, the matching may involve performing an initial coarse matching using downsampled images, followed by a fine matching using parameters upscaled from the initial coarse matching. This common technique can be performed hierarchically using any number of downsampled images with progressively lower resolution. Generally, the most efficient number and scale of downsampled matching may depend on the image content, initial resolution, and degree of misalignment, among other things. In one embodiment, a predetermined hierarchical scheme may be used. In another embodiment, the number and scale of downsampled representations may be dynamically adjusted based on an initial assessment of the image content, for example, by a human user, a machine learning process, or other automated, semi-automatic, or manual technique. Thus, in one embodiment, the matching may involve matching downsampled instances of the image captured in step 502, and then calculating motion parameters for matching the image captured in step 502 by scaling up the motion parameters from the downsampled instances to the full resolution of the image captured in step 502. As mentioned above, this may be performed recursively or hierarchically, for example, alignment may involve recursively downsampling pixel data, aligning the image, and then upsampling motion parameters for alignment of the downsampled image.

[0081] As noted, Method 500 describes in this specification the alignment of each of the single illumination images with respect to one another. However, this alignment between images is generally indirect and is performed by a process of aligning the individual images to one of the composite images when they are added together using an optimization function or the like. Thus, or instead, Method 500 may be described as the alignment of individual images, or groups of images, to a composite image, which then provides a valid coordinate system for inferring the relative alignment of each of the individual images. For simplicity, this is called the relative alignment of images. This is a desired outcome for the purpose of three-dimensional reconstruction, even if this alignment is based on the alignment of the added group of images to one of the composite images or some other reference frame.

[0082] As shown in step 510, method 500 may include reconstructing the three-dimensional shape of a surface based on aligned or aligned images, such as any of images illuminated under various conditions as described herein. This may include reconstructing the three-dimensional shape using shape-from-shading, illuminance difference stereo, structured light, or any other suitable three-dimensional reconstruction technique.

[0083] In another embodiment, Method 500 may include reconstructing a super-resolution image from a aligned image. This may be performed independently of three-dimensional reconstruction, for example, simply to obtain a high-resolution image of the surface for display or analysis, or it may be performed as a preliminary step to three-dimensional reconstruction so that the three-dimensional reconstruction can resolve in more detail than the source image. Thus, in one embodiment, reconstructing the three-dimensional shape of the surface may include obtaining multiple super-resolution images of the surface from a aligned image and reconstructing the super-resolution three-dimensional shape of the surface from the super-resolution images using, for example, one of the three-dimensional reconstruction techniques described herein. Generally, the super-resolution process may include aligning a set of images (for example, using the techniques described herein) and then merging the images into a high-resolution image by utilizing the residual (unprocessed) data or the difference between the pixel values ​​of the aligned / aligned frames. This final step can be performed using, for example, an interpolation, back projection, a deep learning model such as a convolutional neural network trained on high-resolution and low-resolution models, or any other technique known in this art for merging data from two or more images into a single high-resolution image.

[0084] The steps of Method 500 may be repeated as necessary or insofar as they are useful for a particular imaging or control application. More generally, unless otherwise specified, the various features and techniques described herein may be used individually or in any suitable combination and may be applied simultaneously or sequentially to multiple image sets without departing from the scope of this disclosure. However, it should be noted that the superposition assumption may not apply when there is a large displacement between images. Thus, the acquisition rate of a series of such images may be usefully constrained in order to keep, or attempt to keep, the motion between images within a predetermined threshold suitable for a particular imaging system. Alternatively, other metrics relating to motion between images, such as initial inertia or image-based motion estimation, may be used as thresholds or gating criteria for image matching using superposition as described herein.

[0085] Furthermore, as can be understood, although Method 500 is described in relation to, for example, a handheld device for capturing source data for shape-from-shading, the technique may be used in other imaging situations. For example, in one embodiment, the technique described herein may be used for motion compensation in applications of high-speed robotic tactile sensing. In this regard, the imaging system may have a camera equipped with an RGB sensor, along with multiple RGB light sources for illumination. As a result, the illumination may be partially multiplexed in the spectral domain and in the temporal domain. For example, when using two sets of RGB light sources, Method 500 may include capturing an image using each light source while one light source is off, and then capturing a third (composite RGB) image with both sets of RGB light sources turned on. The resulting composite image allows for precise matching of the two images illuminated from each of the individual RGB light sources using the technique described herein.

[0086] Figure 6 shows an imaging system 600 comprising a retrographic sensor 602. While the retrographic sensor 602 is shown as a removable and replaceable cartridge, as can be understood, or alternatively, the retrographic sensor 602 may include any of the retrographic sensors described herein, or other elastomer or shape-matchable optical sensors. The imaging system 600 may also include an illumination system 604, an imaging device such as a camera 606, a processing circuit 608, and an imaging volume 610. An optical element 612 may be arranged to control the illumination of the imaging volume 610 by the illumination system 604. The imaging system 600 may include, for example, a handheld imaging device such as any of those described in International Application No. PCT / US2022 / 046129, issued on April 13, 2023, the entire contents of which are incorporated herein by reference.

[0087] The retrographic sensor 602 may be detachably and interchangeably coupled to the imaging system 600, and may be mechanically key-engaged to the imaging system 600, or otherwise coupled, so that the detection surface 614 of the retrographic sensor 602 is aligned with the imaging volume 610 of the imaging system 600. The retrographic sensor 602 may include, for example, an elastomer optical element having a flexible and optically transparent elastomer on a first side facing the imaging device 606, and a thin reflective coating on a second side facing away from the imaging device 606. The retrographic sensor may be configured to deform when positioned in contact with a target surface for measurement. More generally, the retrographic sensor 602 may include any deformable element having a reflective coating suitable for image capture as described herein, and may include sensors using a fluid medium or other deformable medium instead of elastomer. The imaging system 600 may have an axis 616 (such as an imaging axis or optical axis) passing through the imaging volume 610. When the retrographic sensor 602 is positioned for use within the imaging system 600, for example, the detection surface 614 of the retrographic sensor 602 intersects with the axis 616 of the imaging system 600 and is located within the imaging volume 610, so that the imaging device 606 can capture an image of the detection surface 614 of the retrographic sensor 602 within the imaging volume 610 of the imaging system 600.

[0088] The imaging system 600 may include a substrate for the retrographic sensor 602. The substrate may be formed from a rigid, optically transparent material that is placed between the deformable medium and the imaging device 606 and mechanically supports the deformable medium.

[0089] The illumination system 604 may include any light source or combination of light sources suitable for providing illumination into the imaging volume 610 via the optical element 612, for example, including directional and / or structured light sources. When the retrographic sensor 602 is positioned for use within the imaging system 600, the illumination system 604 may illuminate the detection surface 614 of the retrographic sensor 602, enabling image capture by the imaging device 606. These images may then be processed by the processing circuit 608 to resolve (interpret) three-dimensional surface information of an object in contact with the detection surface 614 of the retrographic sensor 602. In one embodiment, the illumination system 604 may include a laser or other device having a coherent fixed focus and / or providing collimated illumination. In this regard, as understood, the fixed focus may include light that is focused to infinity and collimated, or formed in parallel ray trajectories, as well as any other light having a fixed focus that can be used to create the illumination patterns described herein. In another embodiment, the illumination system 604 may provide non-focus illumination by making appropriate modifications to the optical elements 612 and other optical feature elements. Alternatively, the illumination system 604 may include one or more LEDs or other light sources (such as those shown in Figure 1) arranged to provide controllable lateral illumination of the detection surface 614 within the imaging volume 610, for example, to support shape-from-shading of a target surface or similar three-dimensional reconstruction.

[0090] The imaging device 606 may include a camera suitable for capturing an image of a reflective surface provided by a detection surface 614 of a deformable medium for use by a processing circuit 608 when interpreting (resolving) a three-dimensional image, or any other combination of optical devices, lenses, filters and other hardware. Generally, the imaging device 606 may have an imaging axis or optical axis, such as an axis 616 of the imaging system 600 that passes through the imaging volume 610 to capture its image.

[0091] The processing circuit 608 may include any processor, processing circuit, controller, microcontroller, or other circuit, or a combination thereof, suitable for controlling the operation of the imaging system 600 to acquire three-dimensional information as described herein. In particular, the processing circuit 608 may be configured to control the imaging device 606 and the illumination system 604 to capture an image of the deformable medium detection surface 614 by the imaging device 606 under different illumination conditions, thereby providing two or more images of the detection surface 614. Alternatively, the processing circuit 608 may also be configured to control the imaging device 606 and the illumination system 604 to capture a composite image of the detection surface 614 while it is simultaneously illuminated under all of the different illumination conditions used to capture two or more images of the detection surface 614. Or, instead, the processing circuit 608 may include computing resources to perform additional functions as described herein. This may include computing resources running locally on the imaging system 600, other local user resources such as a local desktop computer or workstation coupled to the imaging system 600, or cloud computing resources configured to support image processing using image data from the imaging system 600. As an example, the processing circuit 608 may include cloud computing resources configured to align multiple differentially illuminated images into a composite image using cost function optimization, and to perform three-dimensional reconstruction with the resulting aligned image using shape-from-shading, illuminance difference stereo, and structured light, etc.

[0092] The processing circuit 608 may be configured in a multi-image matching algorithm to match three or more images. In this case, the three or more images are matched simultaneously, for example, by applying a motion model while minimizing the image difference between the composite image and the sum of the three or more images, thereby providing matching that aligns the three or more images. The processing circuit 608 may also be configured to reconstruct the three-dimensional shape of the detection surface using shape-from-shading or other techniques based on the three or more images that are matched according to the matching.

[0093] In one embodiment, the processing circuit 608 may be configured to control the imaging system 600 to acquire images under different illumination conditions. Alternatively, the processing circuit 608 may be configured to multiplex and / or decouple groups of images and composite images captured under various combinations of directional illumination to acquire surface normal values ​​from the detection surface 614 at a resolution higher than the nominal resolution of the imaging device 606 (e.g., superresolution). In this context, superresolution imaging is understood to include any imaging resolution higher than the limitations imposed by the optical system (e.g., hardware pixels or other hardware or software limitations) that limit the resolution of the acquired image to a nominal resolution lower than the superresolution obtained from multiple images. For example, if the camera has a predetermined resolution, the superresolution image may include a greater number of pixels (either total or per unit of imaging area) than the predetermined or nominal resolution of the camera. Alternatively, the processing circuit 608 may multiplex and / or decouple a set of images to reduce pixel noise in three or more images.

[0094] In one embodiment, a processing circuit 608 physically coupled to the imaging system 600 may perform limited control over data acquisition, for example, to acquire data to be sent to a separate processor for processing, in which case the image data is sent to other processing resources for other preprocessing and three-dimensional reconstruction. In another embodiment, the processing circuit 608 may include one or more microprocessors, field-programmable gate arrays, graphics processing units, and / or other processors for locally processing the image and decomposing the image data into three-dimensional data of the surface in the imaging volume 610. In one embodiment, the processing circuit 608 includes a processor consisting of instructions stored in memory, which receives from the imaging device 606 an image of patterned light generated by the optical element 612 and reflected by a thin reflective coating on an elastomer optical element or other detection surface as it deforms against a target surface of an object in the imaging volume 610. This processor, or another processor integrated into or coupled to the imaging system 600 in a communication relationship with the imaging system 600, may be further configured by instructions stored in memory to calculate the quantitative surface shape of a surface based on an image captured by the imaging device 606.

[0095] The imaging surface in these reconstructions may include, for example, a deformable surface of an elastomer optical element configured to intersect the imaging volume 610 and shape-match with the target surface of the object to be measured. If the target surface intersects the imaging volume 610, an image of the deformable surface captured by the imaging device 606 (e.g., the detection surface 614 of the retrographic sensor 602) may be used to estimate the three-dimensional shape of the target surface. The processing circuit 608 may be configured to control the imaging device 606 and the illumination system 604 to illuminate the deformable surface under various conditions while capturing images. In one embodiment, the processing circuit 608 includes a cloud computing resource configured to receive three or more images and a composite image, align the three or more images based on the alignment of the sum of the images to the composite image, and reconstruct the three-dimensional shape of the detection surface using shape-from-shading or other techniques based on the aligned image.

[0096] The illumination system 604 may be configured to independently illuminate the detection surface 614 from multiple directions via the deformable medium of the retrographic sensor 602, or using differential illumination, for example, structured light or other forms based on different wavelengths. In the case of directional illumination, the detection surface 614 can be independently illuminated from each of three or more directions centered on the axis 616 of the imaging device 606 (which may include or be parallel to the optical axis), thereby providing directional illumination of the detection surface 614. In the case of structured light, the detection surface 614 can be independently illuminated in three or more different patterns.

[0097] The imaging volume 610 may generally define a three-dimensional field of view for the imaging device 606. As described above, the imaging device 606 may have an imaging axis, such as the axis 616 of the imaging system 600, which passes through the imaging volume 610. The plane may intersect the imaging volume 610 and be positioned substantially perpendicular to the imaging axis of the imaging device 606. Alternatively, this plane may be positioned substantially perpendicular to the plane in Figure 6 and shown as line 620, in which case the plane intersects the plane in Figure 6 and intersects the imaging volume 610 shown inside.

[0098] The optical elements 612 may include diffraction gratings, lenses, filters, microtextured surfaces, and metasurfaces suitable for creating desired illumination directions and / or patterns within the imaging volume 610. Generally, the pattern may include multiple feature elements, such as dots (including random and / or ordered dots), lines, sinusoidal patterns (e.g., having a predetermined spatial frequency), polygons, or similar, and combinations thereof, which can be projected onto the detection surface 614 of the imaging volume 610 imaged by the imaging device 606. In one embodiment, the pattern may include a first set of feature elements densely packed in a plane (represented by lines 620), and a second set of feature elements visually distinguishable from the first set of feature elements and scattered within the plane. In this pattern, the scattered feature elements may provide references or other high-texture markers within the imaging volume 610 to aid in alignment and other processing, while the densely packed feature elements assist in the extraction of high-resolution three-dimensional information that is more sensitive to surface shape. Alternatively, the pattern may include a first set of feature elements and a second set of feature elements that collectively form a regular geometric pattern in a plane, in which case the second set of feature elements form visually distinguishable anchor points or references within the pattern.

[0099] The anchor points or markers may be spaced far enough apart so that they are unlikely (or physically unable) to intersect in the imaging plane due to deflection along axis 616. In these embodiments, the pattern may generally include a first set of feature elements that are densely packed to provide high-resolution depth detection within the imaging volume, and a second set of feature elements that are spaced far enough apart in the plane passing through the imaging volume 610 to avoid intersection along the imaging axis (e.g., axis 616) within the imaging volume 610 during the maximum expected deformation of the detection surface 614 of the retrographic sensor 602 within the imaging volume 610. By arranging the second set of feature elements in this way, the feature elements can remain clearly associated directionally with their movement in the two-dimensional imaging field throughout the entire range of expected deformation. As understood, in this regard, the expected deformation may include Z-axis displacement, as well as any X-axis or Y-axis displacement resulting from slip and wrinkle of the elastomer optical elements as the imaging system 600 is positioned relative to the target surface and manipulated by the user.

[0100] In one embodiment, the optical element 612 may include a diffractive optical element that receives illumination from the illumination system 604 (e.g., a coherent light source such as a laser) on a first surface 612a (e.g., the surface facing the illumination system 604) and is arranged to create a three-dimensional illumination pattern within the imaging volume 610 from a second surface 612b opposite the first surface 612a. When a diffractive optical element is used, the diffractive optical element may include a micropatterning structure, optionally with additional lenses, on one or both of the first surface 612a and the second surface 612b, for example, to cooperate to create a desired illumination pattern when a suitable light source is directed toward the first surface 612a. Various types of diffractive optical elements are known in the art and can be used to create illumination patterns in which the brightness changes in the far-field plane and along the imaging axis the brightness and / or focus changes. As a significant advantage, these properties can be utilized to create a three-dimensional illumination pattern within the imaging volume 610 of the imaging system 600 in order to facilitate the resolution of three-dimensional information from the detection surface 614 of the retrographic sensor 602. More specifically, diffractive optical elements can be used to create illumination patterns with complex three-dimensional structures (e.g., not simple two-dimensional projections linearly proportional with distance). These patterns can usefully encode distances within the imaging volume so that shape reconstruction from a single image can be facilitated. Alternatively, any number of additional optical components may be included to create illumination patterns as described herein. For example, the optical system may incorporate photoshaping features such as lenses and filters, for example, to control refractive power and to compensate for distortion or wavefront errors.

[0101] Furthermore, while a suitable diffractive optical element (DOE) may be constructed of micropatterned and / or nanopatterned structures on various optical surfaces of a separate optical element, as shown in Figure 6, for example, the DOE may also be realized at other physical locations in the light path and / or other optical components, for example, by micropatterns on the side walls, top and / or bottom of the substrate of the retrographic sensor 602, and / or within other optical elements of the system.

[0102] In this regard, the three-dimensional illumination pattern may include any three-dimensional shape, pattern, or structure that changes with depth or distance from the optical element 612. For example, the three-dimensional illumination pattern may include diverging illumination projections such as a grid pattern, a dot array pattern, a conical pattern, or a pyramidal pattern that diverges (e.g., becomes larger in the imaging plane) as the distance from the optical element 612 increases, or more generally, a three-dimensional pattern that changes along the imaging axis (e.g., axis 616) within the imaging volume 610. In another embodiment, the three-dimensional illumination pattern may include a pattern having one or more feature elements that change along the line of projection from the optical element 612. For example, a circle, dot, or other image may change in brightness or focal point (with or without a change in size) as the distance of the projected image from the optical element 612 increases. These geometric characteristics of the three-dimensional illumination pattern can be usefully created by the diffractive optical element and used to improve the accuracy of the three-dimensional data based on the image of the detection surface 614 captured by the imaging device 606. In one embodiment, these lighting techniques and / or other lighting techniques described herein may be used, for example, to provide a basis for initial estimates of motion parameters of a motion model used to align images, in order to create a standard for aligning multiple images.

[0103] In one embodiment, the optical element 612 may be positioned to create a pattern from its surface into the imaging volume 610 at an oblique angle (e.g., at least 30 degrees, at least 45 degrees, at least 60 degrees, about 60 degrees, or 50 to 70 degrees) with respect to a plane intersecting the imaging volume 610. As understood, the ray trace from the optical element 612 may change angle multiple times as the light from the optical element 612 is optically coupled to the detection surface 614. For example, the light may travel through the surface of the quartz sheet 640, such as a quartz disc or similar used to protect / seal the inside of the imaging system 600 from the external environment when the retrographic sensor 602 is detachably coupled to the body of the retrographic sensor 602. In this regard, unless otherwise specified, the angle of interest is the angle at which these ray traces intersect the plane (specified by line 620) and / or detection surface 614 where the illumination strikes the deformable surface of the retrographic sensor 602 and the image data is captured for interpretation of the three-dimensional shape.

[0104] As recognized, a plane intersecting the imaging volume 610 provides a useful coordinate system (reference frame) for considering other features and structures of the imaging system 600. However, in one embodiment, for example, if the retrographic sensor 602 is pre-shaped to measure spherical, cylindrical, or other concave or convex surfaces, or more generally, any other target surface having a known and non-planar characteristic shape, the imaging volume 610 may be bounded by curved surfaces. In such cases, a single plane may omit a considerable portion of the imaging volume 610. Nevertheless, a plane of interest may be selected, such as a plane perpendicular to the optical axis of the imaging device used to capture an image of the imaging volume 610, or a plane perpendicular to the axis of the lens used to focus an image from the imaging volume 610, or a plane tangent to the contact area of ​​the target surface, or a plane oriented to provide a reference frame for representing angles such as illumination, imaging, and contact.

[0105] In many illumination patterns, a steep angle of incidence (e.g., a sharper angle relative to the plane) can provide greater sensitivity to three-dimensional displacement. Therefore, as shown in Figure 6, when side illumination is performed, it may be advantageous to include one or more additional light sources and / or optical elements in the illumination system 604 to provide illumination from different directions around the axis 616 of the imaging system 600, so that different regions of the imaging volume 610 can benefit from steep side illumination. In one embodiment, these additional light sources may use different spectral bands so that various patterns can be captured simultaneously, for example, in a single image frame, where visual features can be associated with specific light sources and DOEs (or other optical elements) based on wavelength. This technique can also advantageously improve the detection of steep or sharp surface features of closed (confined) regions and / or surfaces. Thus, in one embodiment, three-dimensional data of different parts of the detection surface 614 can be calculated using illumination from various light sources and / or optical elements. Any group of images can be combined as described herein if it can be added to and referenced in a composite image acquired under the corresponding combination of simultaneous illumination conditions.

[0106] In another embodiment, various illumination sources can be multiplexed, for example, by using light of different wavelength ranges (or different specific wavelengths) to illuminate the imaging volume 610 from different directions, and by processing these images from different wavelength ranges separately so that multiple images from multiple illumination directions can be captured and / or processed in parallel. As described above, the imaging system 600 may usefully include a second diffractive optical element positioned and configured to create a second pattern within the imaging volume 610 with respect to a different location from the first diffractive optical element, around the periphery of the imaging volume. These multiple images illuminated with different patterns can be aligned with each other using the techniques described herein, or otherwise combined using the techniques described herein. More generally, two or more additional light sources and / or optical elements can be incorporated into the imaging system 600 to improve imaging under various imaging conditions and / or imaging of various surface shapes.

[0107] In another embodiment, additional imaging techniques may be incorporated into the imaging system 600, for example, to improve the accuracy and robustness of the imaging system 600, to support high-speed, low-resolution processing for specific imaging circumstances (such as image preview or sparse three-dimensional processing), or for other reasons. Thus, in one embodiment, the imaging system 600 may include a multi-view imaging system using multiple imaging techniques (e.g., stereoscopic imaging system, illuminance difference stereo system, or similar) each configured to calculate the quantitative surface shape of a surface within the imaging volume 610 based on images of the surface from two or more different viewpoints and / or using different image and / or three-dimensional reconstruction processing techniques. In this regard, the multi-view imaging system may include stereoscopic imaging systems, illuminance difference stereo systems, and wavelength multiplexing imaging systems (e.g., based on visible and infrared, fluorescence, etc.). In another embodiment, a gradient-based system may use unfocused illumination from various directions to interpret three-dimensional surface information. Generally, these alternative imaging modes may be optically multiplexed for parallel operation with the systems described above, or otherwise combined. For example, these alternative systems may interpret the three-dimensional shape of the surface using light from a second light source, having a second spectral band with wavelengths that do not overlap with the first spectral band of the illumination system 604 and / or one or more other light sources used by the imaging system 600. Alternatively, the imaging system 600 may utilize confocal three-dimensional imaging to incrementally capture an image by eliminating out-of-focus light in two-dimensional slices passing through the imaging volume. These individual slices of the focused surface can then be combined for three-dimensional reconstruction.

[0108] More generally, various complementary imaging modes can be used to support imaging of various surface types, imaging at various speeds, imaging with various resolutions, and imaging over various depth ranges. For example, these complementary techniques can be used in combination to support improved measurement of low spatial frequency three-dimensional feature elements, such as macroscopic large-scale feature elements of a target surface that are preferably removed before measuring micron-scale surface features with gradient-based depth calculations or similar methods. Furthermore, these depth measurements can provide information on the compression of the elastomer within the imaging gel, provide real-time advice and user feedback regarding optimal compression, support faster rendering (e.g., using sparse data arrays), and support measurement of high-frequency forces (e.g., using finite element models of the elastomer).

[0109] In one embodiment, the imaging system 600 may include a complementary depth measurement mode for measuring the distance to a target surface in order to estimate the compression of an elastomer imaging medium, such as any of the elastomer optical elements described herein, and to provide the user with feedback to guide the user to an optimal range of contact force. This may include, for example, user feedback via a number of LEDs or similar, an auditory output device, or a display of the user interface of the handheld imaging device (e.g., a computer or similar coupled to the handheld device), which may be configured to guide the user to an optimal position, orientation, and / or contact force using visual or other feedback while positioning the handheld scanner.

[0110] In one embodiment, the imaging system 600 may include a lens 630 for variably focusing the imaging device 606 onto a surface within the imaging volume 610. For example, the lens 630 may be a liquid lens using a combination of optical fluid and polymer film to change focus by changing its shape, or any other adaptive lens or similar. Liquid lenses advantageously provide a compact mechanism for controlling focus without using mechanical moving parts and without physically moving the lens along the imaging axis to change the focal length. However, or alternatively, other lenses may be used to control the focus of illumination and / or imaging by the imaging device 606 through the imaging volume 610 and at various depths (depths) or z-axis positions along the imaging axis, and may be adapted for use in the imaging system 600 as described herein.

[0111] In one embodiment, lens 630 may include a lens system focused by a piezo focus drive, a voice coil motor, or any other electromechanical actuator(s) suitable for z-stack image acquisition. For example, lens 630 may include one or more high-resolution lenses with a narrow depth of field. To avoid low-pass filtering that may otherwise be caused by locally out-of-focus lenses, lens 630 may be variably focused to scan various depths (e.g., along the z-axis or imaging axis) to provide partially locally focused images at each desired depth. This stack of images may be assembled into a single image with a deeper depth of field for subsequent three-dimensional processing (e.g., using illuminance difference stereo) or to directly measure quantitative depth information by finding the optimal focus from various depths of focus for a local area within the imaging field of view. This single image with improved depth of field can also be easily restored, including textures, and, when combined with other imaging techniques (such as differential stereo), may provide more accurate, high-resolution surface measurements across the imaging field of view, free from distortion artifacts.

[0112] Figure 7 shows the imaging system 700. The imaging system 700 may be a handheld imaging system and may include a retrographic sensor 702 in a cartridge that can be removed from and replaced by the housing 704 of the imaging system 700. Generally, handheld imaging systems can be susceptible to user tremor, which can affect imaging resolution, especially when image capture times are long (e.g., 15-50 milliseconds per frame for 7 or more frames) and resolution is high (e.g., less than 20 μm, or less than 2 or 3 μm). In such cases, it may be useful to align multiple images with each other, as described herein, before performing three-dimensional reconstruction based on multiple images.

[0113] Figure 8 shows a broken diagram of the imaging system of Figure 7. Generally, the imaging system 700 may include any of the characteristic elements of the imaging systems described herein, such as the retrographic sensor 702, the camera 708, the illumination system 710, and the processing circuit 712.

[0114] The retrographic sensor 702 may include any of the retrographic sensors or other detection elements described herein. In one embodiment, the retrographic sensor 702 includes a deformable medium, such as an optically transparent deformable material, together with a detection surface that covers a portion of the deformable medium and provides a reflective surface visible through a second surface of the deformable medium. The retrographic sensor 702 may be configured as a cartridge for detachable and replaceable use with the imaging system 700. As further described herein, the retrographic sensor 702 has a substrate, in which case the deformable medium is placed on the substrate, and the substrate is placed between the deformable medium and the camera 708. The substrate may generally be formed from a rigid, optically transparent material suitable for mechanically supporting the deformable medium.

[0115] Camera 708 may be any camera, group of cameras, or other imaging system, optical sensor, or similar suitable for acquiring images for use in three-dimensional reconstruction as described herein. Generally, camera 708 may be positioned to capture an image of the reflective surface of retrographic sensor 702 through a second surface of a deformable medium.

[0116] The illumination system 710 may include any light source or combination of light sources suitable for differential illumination of the detection area of ​​the retrographic sensor 702, such as directional illumination, patterned illumination, and differential wavelength illumination. The illumination system 710 may include one or more light-emitting diodes, lasers, and electroluminescent light sources, along with associated lenses, filters, and other optical elements for focusing and directing illumination from the light sources for illumination, as outlined herein. Generally, the illumination system 710 may be configured to illuminate the detection surface independently from each of three or more directions around the optical axis of the camera via a deformable medium, or to provide differential illumination of the detection surface as otherwise described herein.

[0117] The processing circuit 712 may include local computing resources 714, such as a controller, microcontroller, processor, or other circuitry inside the imaging system 700, for example, to control the operation of the imaging system 700. Alternatively, the processing circuit 712 may include remote computing resources 716 outside the imaging system 700, such as a desktop computer locally coupled to the imaging system 700, or virtualized or cloud computing resources remotely coupled to the housing 704 and used to process and three-dimensionally reconstruct images from the imaging system 700.

[0118] The processing circuit 712, for example, is composed of computer executable code stored in the imaging system 700 and executable by local computing resources 714, and controls the camera 708 and the illumination system 710 to capture an image of the retrographic sensor's detection surface with the camera 708 during differential illumination of the detection surface, for example, from each of three or more directions individually, thereby providing three or more images of the detection surface. Alternatively, the processing circuit 712 may be configured to capture a composite image of the detection surface while it is illuminated simultaneously under all differential illumination conditions (e.g., three or more different directions, or any other differential illumination conditions).

[0119] As described above, the processing circuit 712 may also include cloud computing resources, which are composed of, for example, computer executable code stored in the remote computing resource 716, and can align the three or more images by applying a motion model while minimizing the image difference between the composite image and the sum of the three or more images, thereby aligning the three or more images. As understood, and or instead, two images may be used, but using three images can reduce or avoid spatial ambiguity in all directions within the imaging plane. The processing circuit 712 may also be configured to reconstruct the three-dimensional shape of the detection surface based on the aligned image using shape-from-shading or any other suitable three-dimensional reconstruction technique or combination of techniques. In one embodiment, the remote computing resource 716 may be configured to receive three or more images and a composite image from the imaging system 700, align the three or more images (for example, using any of the techniques described herein), and reconstruct the three-dimensional shape of the detection surface using shape-from-shading. Alternatively, if useful for additional processing, alignment motion parameters may be used to align three or more images with each other.

[0120] It should be noted that the terms alignment and alignment are frequently used herein. While these terms are closely related, they have slightly different meanings in the field of image processing. Alignment generally refers to the process of transforming different datasets into a single coordinate system. Alignment, on the other hand, generally means aligning features of two or more images into a common coordinate system by applying transformations such as translation or rotation to make local adjustments. Formally, alignment may be considered a subset of alignment, in which case the transformations are usually assumed to be minor and the overall complexity of the task is low. However, these terms are generally used interchangeably herein unless a different meaning is specifically given or it is clear from the context, and for example, they mean transforming two images into a common aligned coordinate system to facilitate direct pixel-by-pixel comparison.

[0121] In one embodiment, multiplexing may be used in connection with improving resolution and reducing pixel noise, etc. Generally, multiplexing is a technique used in various imaging and optical systems to improve the quality of captured images, increase data acquisition speed, or enable the extraction of additional information from a scene. In this context, time-multiplexing may be used with a series of time-separated images, for example, when a surface of interest is illuminated sequentially in different patterns or from different directions, and the camera captures the resulting image in synchronization with the illumination system. More specifically, by acquiring various combinations of composite and individual illumination, this technique favorably enables improved image resolution and noise reduction. Thus, for example, the processing circuit 712 may be configured to acquire one or more additional images under different illumination conditions, and to improve image matching by multiplexing three or more images, a composite image, and one or more additional images to improve image matching and to acquire surface normal values ​​from the detection surface at a resolution higher than the camera's nominal resolution. In another embodiment, the processing circuit 712 may be configured to acquire one or more additional images under different lighting conditions and to light-multiplex decouple the three or more images, the composite image, and one or more additional images to reduce pixel noise (e.g., caused by unwanted motion) in the three or more images and the composite image.

[0122] Figure 9 illustrates motion compensation techniques. As shown, a first image (image "A") and a second image (image "B") are captured using their own dedicated directional illumination or other differential illumination, respectively. If these two images are recorded at different times, a positional misalignment may occur between the images due to camera movement. To address this potential misalignment, an image motion model is selected along with corresponding motion parameters, which can be variably applied to the images to realign them to a common coordinate system. For example, a rigid body motion model with two translational degrees of freedom (x and y) and one rotational degree of freedom (ψ) can be represented in a similar motion model as follows:

[0123] TIFF2026514959000002.tif6161

[0124] TIFF2026514959000003.tif16161

[0125] Here, the motion parameters dx, dy, and ψ are used to characterize the motion resulting from the camera's movement. These three parameters, dx, dy, and ψ, usefully capture the expected motion of a handheld scanner, such as any of the handheld scanners described herein. In this case, the camera is primarily perpendicular to the imaging plane, and the camera's angular rotation and z-displacement generally contribute relatively little to the image motion, for example, when using a tactile sensor where a thin deformable medium (relative to the thickness of the deformable medium) is placed on a large surface area. These motion constraints allow for a reduction in the number of degrees of freedom required for an effective motion model, thereby reducing computational complexity. However, as is understood, or perhaps instead, various other motion models may be used. For example, this may include models that use different or additional rotational or translational motion parameters. In another embodiment, the motion model may include a deformable motion model for imaging applications where strong local motion may exist independently of the camera's movement.

[0126] Generally, motion parameters are iteratively estimated and applied to differential images such as the first image 904 ("IMAGE A") and the second image 906 ("IMAGE B") (step 902) to obtain a matched image 908 that attempts to compensate for camera motion in the images. The pixels may then be scaled to account for differences in exposure time and / or illumination time (step 910), and the corrected image is then summed (step 912) and compared to a third image 916 ("IMAGE A+B") captured simultaneously as a composite image under each of the differential illumination conditions (step 914).

[0127] Next, standard optimization techniques may be used to evaluate the success of the resizing based on the current motion parameters (step 918). For example, the current motion parameters may be evaluated using a cost function, and the resizing and evaluation may be repeated until a satisfactory result is obtained, in which case a matched image may be output along with the accompanying alignment data, as shown in step 920. Optimization may involve minimizing or maximizing a cost function, such as any of the cost functions described herein. Alternatively, optimization may use termination conditions, such as the convergence of the value of the cost function within a certain threshold over a certain number or range of adjustments to the motion parameters. More generally, optimization may use any condition or group of conditions suitable for evaluating whether a number of images have been successfully resized and / or whether further improvement cannot be achieved (or, alternatively, when alignment cannot be achieved) based on the source image, motion model, motion parameters, and cost function.

[0128] Figure 10 is a flowchart of a motion compensation method. In particular, a hierarchical technique for scaling the matching is shown. This may include scaling a set of differentially illuminated images to a low-resolution representation, matching the low-resolution images using the techniques described herein, and then scaling the resulting motion parameters to high resolution as an initial matching of the high-resolution images.

[0129] This technique can be applied hierarchically by downsampling an image into a number of progressively lower resolution image sets, then upsampling the alignment parameters into progressively higher resolution image sets, and finally providing a highly accurate initial estimate for the full-resolution image. Thus, Method 1000 may include reducing the range of motion of the image by formulating the estimation of the image motion parameters as a hierarchical problem (in this case, the initial parameters are estimated in the downsampled images). For example, motion parameters identified for the alignment of downsampled image sets using the optimization technique described herein may then be upsampled to provide an initial motion estimate for the alignment of images at the next higher resolution level.

[0130] As shown in step 1002, method 1000 may include filtering or otherwise preprocessing the scanned image (hereinafter referred to as the scanned image) 1004, such as any of the differentially illuminated images described herein. In one embodiment, this may include applying a Gaussian spatial filter to each image to smooth the image and reduce noise. The parameters of the Gaussian spatial filter may vary depending on the imaging conditions, but in one practical application, a Gaussian spatial filter with a standard deviation of 3 or about 3 has been demonstrated to adequately reduce the effects of random pixel noise without loss of three-dimensional accuracy. As is recognized, or or instead, other filters or techniques may be used to preprocess the image for hierarchical matching as described herein. Also, as is understood, filters may be usefully used in conjunction with motion compensation as described herein even when hierarchical matching is not used.

[0131] As shown in step 1006, method 1000 may include scaling the images by, for example, downsampling the image set to an image set of progressively decreasing resolution. For example, this may include downsampling the images to scales of 0.1 (or about 1 / 10) and 0.55 (or about 1 / 2). These image sets may be stored together with the preprocessed full-resolution images as scaled images 1008 for further processing.

[0132] As shown in step 1010, method 1000 may include aligning the lowest resolution set of scaled images 1008 (including low-resolution composite images) with each other using the techniques described herein. For the lowest resolution images, initial alignment estimates may be provided using any of the techniques described herein. This pre-alignment step may include, for example, a brute-force estimation of the similarity or difference of images across a pre-selected number of motion parameters, or the use of initial motion parameter estimates based on an independent source, such as a camera's inertial measurement device or other motion sensor, or other sources of initial estimates described herein. Aligning images may include generating a set of motion parameters that characterize the motion between two or more images so that two or more images can be aligned or aligned with each other, for example, when performing three-dimensional reconstruction or other processing.

[0133] As shown in step 1012, method 1000 may include scaling motion parameters from the lowest resolution image to correspond to the scale of the next highest resolution image set in the scaled image 1008. For example, in the case of a progressive scale of 0.1 and 0.55 (relative to full resolution), this may include multiplying the x and y displacement parameters by 5.5.

[0134] As shown in step 1014, method 1000 may include matching the scaled image 1008 at the next highest resolution. Generally, the scaled motion parameters from the lower resolution match can be used as initial estimates for the match at this next higher resolution. The match can then be performed using the scaled image and the scaled motion parameters, with the differentially illuminated image and the composite image, as described more generally herein.

[0135] As shown in step 1016, method 1000 may include determining whether an additional hierarchical resolution layer exists for processing. If the image currently being matched from the scaled image 1008 is a full-resolution image (e.g., relative to the scanned image 1004, or some other target resolution of interest based on the scanned image), method 1000 may proceed to step 1018, in which case the matched image may be stored, along with any motion parameters, residual data, and / or other information of interest as needed, for use in, for example, three-dimensional reconstruction, super-resolution imaging, or other processing as described herein. If the image currently being matched is not full resolution, method 1000 may return to step 1012, in which case the motion parameters may be scaled (upsampled) to the next higher resolution, and additional matching may be performed in the higher resolution space. This may be repeated until full resolution (e.g., scale = 1) or other target resolution is achieved.

[0136] Figure 11 shows the summation results of images before and after alignment. The image on the left shows the pixel-by-pixel difference between the summation result of (a) the composite image and (b) the single-illumination image before alignment. As expected, motion artifacts between the temporally separated single-illumination images result in a summation result with pixel values ​​that differ significantly from the composite illumination image, with particularly large residuals around areas of locally large height changes. The image on the right shows a second pixel-by-pixel difference. In particular, the image on the right shows the pixel-by-pixel difference between the summation result of (a) the composite image and (b) the single-illumination image after alignment using the technique described herein. Because motion artifacts have been reduced, the summation result of the single-illumination images closely matches the composite illumination image, and the difference across the field of view is significantly smaller.

[0137] In one embodiment, large local image differences in the left-hand image (indicated by dark shading) can provide a useful indicator of the magnitude and direction of the displacement. Motion compensation techniques described herein may use pixels in these high-cost or high-texture regions to provide initial estimates of motion parameters related to the displacement. Alternatively, these regions of large local differences may suggest the selection of favorable pixels or regions for performing iterative optimization. Thus, in one embodiment, the selection of pixels for performing optimization calculations may include selecting pixels for optimization based on local differences between the composite image and the pre-adjusted summation image, such local differences may be evaluated, for example, based on the value of the optimization cost function, the brightness (intensity) value, or the value of some other estimator or function at each pixel location in the pre-adjusted summation image.

[0138] Figure 12 shows a comparison of the composite image with the pre-alignment summation image and the post-alignment summation image. The composite image shows the target surface illuminated simultaneously under all differential illumination conditions (in this case, while illuminated from three different directions). This composite image provides a quantitative target for the alignment of the individual illumination images. The pre-alignment summation image is obtained by adding the pixel values ​​of each pixel position of the individual illumination images before any alignment or alignment is performed. Similar to Figure 11 above, this pre-alignment image contains artifacts based on large motion, particularly in areas of high texture or high local differences. The individual images can be aligned with each other as described herein by optimizing the cost function to identify motion parameters that minimize motion-based artifacts. The resulting post-alignment summation image is shown in Figure 12 and can be seen to clearly converge to the composite image that provides the alignment target.

[0139] Figure 13 shows surface normal maps and rendered three-dimensional surfaces with and without motion compensation. The images in Figure 13 were captured using a commercially available GelSight retrographic sensor. The first column shows a representative two-dimensional image under differential illumination (directional illumination in this case). The second column shows the surface normal map derived from a set of images obtained under multiple differentially illuminated images. The uncompensated calculated surface normal map may be observed to include reduced sharpness or focus as a result of the accumulation of motion artifacts across the imaging surface. As shown in the lower row, the compensated three-dimensional reconstruction includes noticeably improved detail compared to the uncompensated three-dimensional reconstruction. Both the surface normal map image and the three-dimensional rendered surface, when motion compensation is applied, include geometric shapes at a visibly smaller scale, and consequently, sharper images.

[0140] The systems, devices, methods, and processes described above may be implemented in hardware, software, or any combination thereof suited to a particular application. Hardware may include general-purpose computers and / or dedicated computing devices. This may include implementations in one or more microprocessors, microcontrollers, embedded microcontrollers, programmable digital signal processors, or other programmable devices or processing circuits, along with internal and / or external memory. Alternatively, this may include one or more application-specific integrated circuits, programmable gate arrays, programmable array logic components, or any other device(s) that can be configured to process electronic signals. As is further recognized, implementations of the processes or devices described above may include computer executable code written using structured programming languages ​​such as C, object-oriented programming languages ​​such as C++, or other high-level or low-level programming languages ​​(including assembly languages, hardware description languages, and database programming languages ​​and techniques), which can be stored, compiled, or interpreted for execution on any of the devices described above, as well as heterogeneous combinations of processors, processor architectures, or different hardware and software combinations. In another embodiment, the method may be embodied in a system that performs its steps and may be distributed across devices in numerous ways. Simultaneously, the processing may be distributed across devices such as the various systems described above, or all functions may be integrated into a dedicated standalone device or other hardware. In another embodiment, the means for performing the steps related to the process described above may include any of the hardware and / or software described above. All such substitutions and combinations are intended to fall within the scope of this disclosure.

[0141] Embodiments disclosed herein may include computer program products that include computer executable code or computer-available code that, when executed on one or more computing devices, performs any and / or all of the steps thereof. The code may be persistently stored in computer memory, which may be memory on which the program is executed (such as random access memory associated with a processor), or it may be a storage device such as a disk drive, flash memory, or any other optical device, electromagnetic device, magnetic device, infrared device, or any other device or combination of devices. In another embodiment, any of the systems and methods described above may be embodied in any suitable transmission or propagation medium that carries the computer executable code and / or any input or output from the computer executable code.

[0142] As is to be recognized, the devices, systems, and methods described above are provided as examples and are not intended as limitations. Unless expressly indicated otherwise, the disclosed steps may be modified, supplemented, omitted, and / or reordered without departing from the scope of this disclosure. Numerous variations, additions, omissions, and other modifications will be apparent to those skilled in the art. Furthermore, the order or presentation of the method steps in the above description and drawings is not intended to require this order to perform the described steps unless a specific order is expressly required or is evident from the context.

[0143] The method steps of the embodiment described herein are intended to include any suitable method of causing such method steps to be performed in a manner that conforms to the patentability of the following claims, unless otherwise explicitly stated or evident from the context. For example, the execution of step X may include any suitable method of causing another party, such as a remote user, a remote processing resource (e.g., a server or cloud computer), or a machine, to perform step X. Similarly, the execution of steps X, Y, and Z may include any method of commanding or controlling any combination of such other individuals or resources to perform steps X, Y, and Z in order to benefit from such steps. Thus, the method steps of the embodiment described herein are intended to include any suitable method of causing one or more other parties or entities to perform the steps in a manner that conforms to the patentability of the following claims, unless otherwise explicitly stated or evident from the context. Such parties or entities do not need to be under the command or control of any other party or entity, nor do they need to be located within a particular jurisdiction.

[0144] As will be recognized, the methods and systems described above are described as examples and not as limitations. Thus, although certain embodiments have been illustrated and described, as will be apparent to those skilled in the art, various modifications and variations may be made in form and detail without departing from the spirit and scope of this disclosure, and are intended to form part of the present invention as defined by the following claims.

Claims

1. A computer program product comprising computer executable code embodied in a persistent computer-readable medium, which, when executed on one or more computing devices, causes one or more computing devices to perform the following steps, wherein the steps are: A step of capturing three images, wherein the three images are A first image of the surface while the surface is illuminated from a first direction, A second image of the surface while illuminating the surface from a second direction, A capturing step including a third image of the surface while simultaneously illuminating the surface from the first and second directions, The steps include: aligning the first image with the second image by applying a motion model to at least one of the first and second images to the third image, while minimizing a cost function that represents the difference between the third image and the sum of the first and second images; A computer program product comprising the steps of: performing alignment to match the first image to the second image according to the motion model.

2. The cost function is based on the difference in pixel values, as described in claim 1.

3. The computer program product according to claim 1, wherein the cost function is evaluated for a subset of pixels.

4. The cost function is based on an image similarity index, as described in claim 1 for the computer program product.

5. The cost function is based on a normalized correlation coefficient, as described in claim 1 for the computer program product.

6. The computer program product according to claim 1, further comprising code causing one or more computing devices to perform the step of restoring the three-dimensional shape of the surface by shape-from-shading based on the first image and the second image when aligned according to the aforementioned alignment.

7. The computer program product according to claim 1, wherein minimizing the cost function includes minimizing the difference between the pixel values ​​at one or more pixel positions in the first pixel array for the third image and the sum of the first image and the second image at one or more corresponding positions.

8. The computer program product according to claim 1, wherein the motion model includes a rigid body motion model.

9. The computer program product according to claim 8, wherein the rigid body motion model includes one or more rigid body rotations and rigid body translations.

10. The computer program product according to claim 1, wherein the motion model includes a rigid body motion model for the motion of an image induced by six degrees of freedom in the orientation of an imaging device that captures the first image, the second image, and the third image.

11. The computer program product according to claim 1, wherein the motion model includes deformable motion.

12. The computer program product according to claim 1, wherein the motion model uses independent motion tracking for one or more sub-regions of the first image, the second image, and the third image.

13. The computer program product according to claim 1, wherein the motion model uses one or more visible criteria for tracking image differences.

14. The computer program product according to claim 1, wherein the matching comprises matching downsampled instances of the first image and the second image, and calculating the motion parameters in order to match the first image and the second image by scaling up the motion parameters from the downsampled instances to the scale of the first image and the second image.

15. The computer program product according to claim 1, wherein the matching includes recursively downsampling, matching, and scaling motion parameters with respect to two or more downsampled versions of the first image, the second image, and the third image.

16. The computer program product according to claim 1, wherein the matching includes dividing the pixel arrays of the first image, the second image, and the third image into a plurality of regions, and selecting one or more pixel positions from each of the plurality of regions for evaluating the cost function.

17. The computer program product according to claim 1, wherein the matching comprises selecting a subset of pixel positions in each of the pixel arrays of the first image, the second image, and the third image in order to minimize the cost function, and the selection of the subset of pixel positions comprises selecting at least one of the subsets of pixel positions based on the magnitude of the cost function in one of the subsets of pixel positions between the third image and the sum of the first image and the second image.

18. The computer program product according to claim 1, wherein the sum of the first image and the second image includes a scaled sum of pixel luminance values.

19. Capture at least three images of the surface, wherein the at least three images include two or more images captured under two or more different lighting conditions, and a composite image of the surface while simultaneously illuminated under each of the two or more different lighting conditions. A method comprising aligning two or more images by applying a motion model while minimizing the image difference between the composite image and the sum of the two or more images.

20. The method according to claim 19, wherein the two or more different lighting conditions include two or more different lighting directions.

21. The method according to claim 19, wherein the two or more different illumination conditions include two or more different illumination wavelengths.

22. The method according to claim 19, wherein the two or more different lighting conditions include two or more different lighting patterns.

23. The method according to claim 19, wherein minimizing the image difference includes minimizing a cost function that represents the difference between the composite image and the sum of the two or more images.

24. The method according to claim 19, wherein aligning the two or more images includes aligning the two or more images using a multi-image alignment algorithm.

25. The method according to claim 19, wherein aligning the two or more images includes minimizing an optimization function.

26. The method according to claim 19, wherein the two or more images include three images.

27. The method according to claim 19, wherein the two or more images include six images.

28. The method according to claim 19, wherein aligning the two or more images includes aligning the two or more images of a first group with each other in a first image aligning, aligning the two or more images of a second group with each other in a second image aligning, and aligning the first image aligning with the second image aligning in a third image aligning.

29. The method according to claim 19, further comprising calculating an initial estimate of the displacement of the motion model based on input from an inertial measuring device.

30. The method according to claim 19, further comprising calculating an initial estimate of the displacement of the motion model based on one or more visible criteria in each of the two or more images.

31. The method according to claim 19, further comprising calculating an initial estimate of the displacement of the motion model based on evaluation by a machine learning model trained to associate one or more predetermined displacements with one or more visual artifacts in a combination of images illuminated under two or more different lighting conditions.

32. It is a system, A retrographic sensor comprising a deformable medium together with a detection surface, wherein the deformable medium is formed from an optically transparent and deformable material, and the detection surface covers a portion of the deformable medium and provides a reflective surface visible through a second surface of the deformable medium, A camera positioned to capture an image of the reflective surface through a second surface of the deformable medium, An illumination system configured to provide directional illumination of the detection surface by independently illuminating the detection surface through the deformable medium from each of three or more directions centered on the optical axis of the camera, It includes a processing circuit, and the processing circuit is During individual illumination of the detection surface from each of the three or more directions, the camera captures an image of the detection surface and controls the camera and the illumination system to provide three or more images of the detection surface. The camera and the illumination system are controlled to capture a composite image of the detection surface while it is being illuminated simultaneously from all three or more directions. The three or more images are aligned by applying a motion model while minimizing the image difference between the composite image and the sum of the three or more images, thereby aligning the three or more images. A system configured to reconstruct the three-dimensional shape of the detected surface using shape-from-shading.

33. The system according to claim 32, wherein the processing circuit includes a controller for the camera and the lighting system.

34. The system according to claim 32, wherein the processing circuit includes a cloud computing resource configured to receive the three or more images and the composite image, align the three or more images, and restore the three-dimensional shape of the detected surface using shape-from-shading and alignment to align the three or more images.

35. The material further includes the substrate of the retrographic sensor, The deformable medium is placed on the substrate, The substrate is formed from a rigid, optically transparent material that mechanically supports the deformable medium. The system according to claim 32, wherein the substrate is placed between the deformable medium and the camera.

36. The system according to claim 32, wherein the processing circuit is configured to acquire one or more additional images under different lighting conditions, and to illuminate multiplex decoupling the three or more images, the composite image and the one or more additional images to acquire surface normal values ​​from the detection surface at a resolution higher than the nominal resolution of the camera.

37. The system according to claim 32, wherein the processing circuit is configured to acquire one or more additional images under different lighting conditions, and to light-multiplex decouple the three or more images, the composite image, and the one or more additional images to reduce pixel noise in the three or more images and the composite image.