Planar purge
A camera-based method using multiple lighting conditions purges 2D features from images, enhancing 3D features for accurate robotic grasping by subtracting and dividing images, addressing the challenges of noisy point cloud data and unreliable 2D detection.
Patent Information
- Application Number
- JP2025071037
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-04-23
- Publication Date
- 2025-11-07
AI Technical Summary
Existing image processing techniques struggle to accurately identify the shape and dimensions of objects on a pallet due to noise in point cloud data and unreliable 2D feature detection, particularly when boxes have text, graphics, and tape on their surfaces, leading to incorrect robotic grasping.
A method involving a single camera capturing three images of an object under different lighting conditions, using ambient light and two additional light sources, to calculate an output image that removes color and other 2D features by subtracting and dividing the images, emphasizing 3D features.
The method effectively purges 2D features, enhancing 3D features for reliable robotic grasping by eliminating color-based shading variations and reflections, improving the accuracy of object identification.
Smart Images

Figure 2025168301000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of image analysis, and more particularly, to a removal method and system for removing 2D features from flat surfaces in a 2D image and enhancing 3D features. Three images are captured by a fixed camera using different combinations of ambient and auxiliary light sources. Image operations are used to calculate an output image devoid of 2D features, including shading variations due to color, graphics, markings, and tape. [Background technology]
[0002] The use of camera images as input to machine control systems is well known, with applications ranging from automotive collision avoidance to industrial robot motion programs. In one common robotic application, a pallet of boxes is provided, and an industrial robot is used to remove one box at a time from the pallet and place each box on a destination, such as a conveyor, where the box will undergo further processing. In such depalletizing operations, one or more cameras are used to provide images of the boxes on the pallet, and the images are known to be analyzed to identify the corners, edges, and sides of the boxes. A depalletizing algorithm is then used to select a particular box for the next robotic picking operation, and the process is repeated until the pallet is empty. Summary of the Invention [Problem to be solved by the invention]
[0003] Accurately identifying the shape and dimensions of boxes using known image processing techniques, including both two-dimensional (2D) and three-dimensional (3D) methods, can be challenging depending on the nature of the boxes on a pallet. 3D methods typically use point cloud data, e.g., from optical methods such as stereoscopic imaging, structured light, or time-of-flight, to approximate the three-dimensional shape of the observed object's surface. However, point cloud data can be noisy, often have sparsely spaced data points, and contain areas of "dropouts," or missing surface points. These and other issues with point cloud data make it difficult to accurately determine the shape of observed objects using 3D image analysis.
[0004] Analyzing 2D images to identify objects is also problematic. Many boxes contain features on their surface, such as text, graphics, color variations, and tape, which makes detecting edges and corners using 2D camera images unreliable. If the lines or color features on a box are mistaken for the box's edges, the robot may attempt to grasp in the wrong place, potentially resulting in a failed grasp or dropping the box.
[0005] In view of the foregoing, there is a need for improved image analysis techniques that remove color shading variations and other 2D features from flat surfaces to improve image-based object identification. [Means for solving the problem]
[0006] According to the teachings of the present disclosure, a technique is provided for eliminating 2D features from flat surfaces in 2D images. The technique involves taking three digital images of an object, such as boxes stacked on a pallet. All images are taken by a single camera at a fixed position. In one image (I A ) is taken with ambient light only, another image (I1) is taken with ambient light and a first additional light source at a first location, and yet another image (I2) is taken with ambient light and a second additional light source at a second location. The output image Q is Q = (I1 - I A) / (I2-I A ) formula. Subtracting the ambient light image removes ambient diffuse and specular light. Division eliminates all variation in the output image caused by color. The only remaining variation is due to the angle between the normal to each surface point and the direction from the light to that point. The output image Q eliminates all color-based shading variations and reflections from marks, graphics, tape, etc., while retaining all 3D features such as gaps between boxes. The output image Q is particularly well-suited for computing the act of robotically grasping objects in the image.
[0007] Additional features of the presently disclosed apparatus and methods will become apparent from the following description and claims, taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] This figure shows an image of a collection of boxes on a palette, illustrating how the patterns and colors of the graphics on the boxes make it difficult to distinguish one box from another in 2D image analysis.
[0009] [Figure 2] An image of a collection of boxes on a pallet shows how the tape at the seams of the boxes and the resulting reflections make it difficult to distinguish one box from another in 2D image analysis.
[0010] [Figure 3] FIG. 1 illustrates the behavior of diffuse and specular reflections from a surface.
[0011] [Figure 4]FIG. 1 illustrates a system for acquiring an image of an object and processing the image to generate an output image that is free of shading variations due to 2D features such as color, graphics, tape, etc., according to one embodiment of the present disclosure.
[0012] [Figure 5] 1 is a flowchart of a method for acquiring an image of an object and processing the image to generate an output image that is free of shading variations due to 2D features such as color, graphics, tape, etc., according to one embodiment of the present disclosure.
[0013] [Figure 6A] FIG. 2 is a diagram showing an image of a collection of boxes in FIG. 1. [Figure 6B] 6 is a corresponding output image in which planar features have been purged using the system of FIG. 4 and the method of FIG. 5;
[0014] [Figure 7A] FIG. 3 is a diagram showing an image of a collection of boxes in FIG. 2. [Figure 7B] 6 is a corresponding output image in which planar features have been purged using the system of FIG. 4 and the method of FIG. 5;
[0015] [Figure 8A] FIG. 1 shows an image of a collection of flat packages. [Figure 8B] 6 is a corresponding output image in which planar features have been purged using the system of FIG. 4 and the method of FIG. 5; DETAILED DESCRIPTION OF THE INVENTION
[0016] The following description of embodiments of the present disclosure directed to removal methods and systems for removing 2D features from planar surfaces and enhancing 3D features in 2D images is merely exemplary and is not intended to limit the disclosed apparatus and techniques or their applications or uses.
[0017] It is known to use camera images and / or sensor data as input to a wide variety of machine control systems. One known application is box depalletizing, where a large number of boxes are present on a pallet and a robot has the task of picking up the boxes one at a time and moving each box to a secondary location. Both two-dimensional (2D) and three-dimensional (3D) image processing and data processing techniques exist that can be used to identify individual boxes within a stack as input to the box depalletizing operation. However, these existing techniques have difficulty accurately analyzing both 2D and 3D images or data.
[0018] In 3D methods using point cloud data, the point cloud may be sparse, resulting in "drop outs" where data points are missing in some areas of the point cloud. The resolution of the 3D data can also be an issue; coarse resolution can lead to incorrect recognition of the box shape, while fine resolution can make processing slow and computationally intensive. In 2D image analysis techniques, text, images, color patterns, and other 2D features on the face of a box can cause the box shape to be identified inaccurately, as shown in the figure below and discussed later.
[0019] FIG. 1 shows an image 100 of a collection of boxes on a pallet, illustrating how graphic patterns and colors on the boxes can make it difficult for 2D image analysis to distinguish one box from another. The object in image 100 is a pallet with two layers of boxes stacked on top of each other, including an upper layer 110 of smaller boxes and a lower layer 120 of larger boxes. The boxes in upper layer 110 contain many graphical designs, images, and text on their surfaces, making their edges difficult to identify using 2D image analysis. In particular, a dark bar 130 surrounded by a lighter area can be mistaken for a box edge, as can a linear transition from a light to a dark background, as shown at 140 and 150. Several other graphic features that could easily be mistaken for box edges in 2D image analysis are readily recognizable on the boxes in upper layer 110. Techniques for overcoming misidentification of box edges, such as using multiple cameras, can be complex and have proven only partially effective.
[0020] FIG. 2 shows an image 200 of a collection of boxes on a pallet, illustrating how tape at the seams of the boxes and the resulting reflections can make it difficult to distinguish one box from another during 2D image analysis. The object in image 200 is a pallet with two layers of boxes stacked on top of each other, including the top layer 210, which is the focus of this discussion. The boxes in top layer 210 are constructed of plain brown cardboard and do not display any of the graphical images of the boxes in FIG. 1. However, the boxes in top layer 210 have clear tape applied across their surfaces. A strip of tape 220 is placed across the approximate center of the cardboard panel of the box, away from the edges. A strip of tape 230 is placed across the edge of the flap in the top center of the box. Finally, a strip of tape 240 is placed along the upper edge of the box, holding down the edge of the flap. Even if the tape is transparent, 2D image analysis may be used to identify any of the strips of tape 220, 230 and 240 as the edge of the box, as the tape will clearly change the "color" (shading or pixel intensity) of the box and glare in the image will add bright reflections making it difficult to determine what the tape is covering.
[0021] Both Figures 1 and 2 illustrate how 2D features, such as graphics and tape on a flat surface (the top surface of a box), can interfere with identifying the edges of a box through 2D image analysis. This disclosure provides methods and systems for processing images to purge 2D features (intensity variations due to color, tape, etc.) and emphasize 3D features (such as the true edges of the box), thereby enabling the image analysis described below to more reliably and confidently identify the box being processed. Throughout this disclosure, when a technique is described as "removing color" from an output image, this does not mean simply converting color to grayscale, as is done by a black-and-white camera. Rather, "removing color" refers to completely purging all intensity variations from the output image, making it as if the color, markings, or 2D features were never present on the object to begin with.
[0022] 3 is an illustration of the behavior of diffuse and specular reflection from a surface 310. Specular reflection (e.g., glare) is visible only when the normal vector of the reflective surface bisects the vector from the surface to the light source and the vector from the surface to the observer. In other words, if incident light from light source 312 strikes surface 310 along vector 320, the specular reflection will only be visible to observer 314, who is viewing surface 310 along vector 322. Vectors 320 and 322 are symmetrically opposed to each other across the surface normal.
[0023] Diffuse reflection has the same brightness to an observer no matter what angle the observer views surface 310 from. For example, if a light source strikes surface 310 at a low angle (nearly parallel to the surface) along vector 330, a low brightness or low intensity diffuse reflection will be visible from any angle relative to surface 310, as indicated by short dashed arrow 332. If a light source strikes surface 310 at a high angle (nearly perpendicular to the surface) along vector 340, a high brightness or high intensity diffuse reflection will be visible from any angle relative to surface 310, as indicated by long dashed arrow 342.
[0024] Most objects exhibit a combination of both specular and diffuse reflections, which are problematic to sort out in 2D image analysis. However, this disclosure provides a technique for capturing multiple images of an object under specific, different lighting conditions and mathematically combining these images to remove undesirable reflection characteristics and emphasize desirable characteristics.
[0025] FIG. 4 is an illustration of a system 400 for acquiring images of an object and processing the images to generate an output image lacking intensity variations due to 2D features such as color, graphics, tape, etc., according to one embodiment of the present disclosure. The object 410 is one or more objects shown in the image or sensor data. To continue with the example used throughout this disclosure, the object 410 may be a collection of boxes on a pallet. However, the object 410 may also be a single object such as boxes, or any other multiple objects. The techniques of the present disclosure work with any embodiment of the object 410 and work particularly well when the object 410 has a relatively flat top surface 412 or any other relatively flat surface where 2D features such as color, graphics, tape, etc., are present that need to be removed from the output image.
[0026] FIG. 4 illustrates one embodiment of the disclosed technology in some detail. The exemplary embodiment of FIG. 4 shows a scenario in which an object 410 has a horizontal top surface 412 for which an image without intensity variations due to 2D features is desired. The light source and camera arrangement in FIG. 4 is suitable for a top surface imaging scenario. However, the scenario of FIG. 4 is merely an example for purposes of illustrating the disclosed technology. The same disclosed technology can be applied to many other imaging scenarios, for example, imaging a vertical side of an object, an imaging surface that is tilted relative to the vertical and horizontal directions, etc. In some scenarios, a combination of one or more surfaces to be imaged, two or more light sources, and output images is hybrid. These scenarios are discussed below. FIG. 4 again illustrates one specific, non-limiting example.
[0027] The first light source 420 and the second light source 422 are fixed at different positions, facing downward at an oblique angle toward the object 410. The light sources 420 / 422 should not be facing vertically downward toward the horizontal top surface 412, but rather should have an aiming angle between vertical and horizontal. In one embodiment, the light sources 420 / 422 have an elevation angle above horizontal of 25° to 45°, although higher or lower elevation angles are also suitable. In a preferred embodiment, the light sources 420 / 422 each have the same elevation angle above horizontal.
[0028] The light sources 420 / 422 need to be positioned at different locations relative to one another to illuminate the object 410 differently in different images. In one embodiment, when viewed from directly above the object 410, the light rays 420 / 422 have positions and aiming vectors that are approximately 90° apart in a top view. However, other relative position angles are also suitable, including light sources 420 / 422 that are 180° apart in a top view (directly opposite one another). In one embodiment, the light sources 420 / 422 are light-emitting diode (LED) lights, although other types of lights are also suitable.
[0029] The workspace in which object 410 is positioned typically has sources of ambient light in addition to light sources 420 / 422. Ambient light may be, for example, a combination of artificial lighting (such as fluorescent fixtures in a warehouse) and natural sunlight. The presence and uncontrollability of ambient light is typically unavoidable, and therefore, as described below, the techniques of the present disclosure have been developed to recognize and compensate for this fact.
[0030] The two-dimensional (2D) sensor 430 is fixed at another location, preferably a different location from the light sources 420 / 422, and is configured to capture detection data or images of the object 410. In one embodiment, the sensor 430 is a digital camera that captures black-and-white (grayscale) images of the object 410, with such images having a resolution ranging from 1 megapixel to 5 megapixels. Higher or lower resolutions can also be used. While color images may be used, grayscale images are preferred because the image processing techniques described below operate on pixel intensity values. The data provided by the 2D sensor 430 will hereinafter be referred to as an image, although it should be understood that other types of 2D sensors and data can also be used. The top surface 412, which will be compensated for in the image to eliminate 2D features, including color effects, must be illuminated by both the light sources 420 and 422, and of course, the surface 412 must be within the field of view of the sensor 430.
[0031] A computer 440 receives a set of three images from the sensor 430 for the disclosed image analysis techniques. The computer 440 can communicate with the sensor 430 wirelessly or via a hardwired connection. The computer 440 may also control the light sources 420 / 422. After the object 410 is positioned, the computer 440 can automatically control the acquisition of three images, including an image with ambient light only (referred to as I_A), an image in ambient light and the light source 420 with the light source 422 turned off (referred to as I_1), and an image in ambient light and the light source 422 with the light source 420 turned off (referred to as I_2).
[0032] Below is a description of how image operations are performed to combine the three images in a specific way to reduce 2D features and enhance 3D features in the object 410. Lambert's diffuse reflection law can be stated as follows:
number
[0033] Equation (1) can be rewritten to help explain the image processing concepts of this disclosure. The dot product of vectors L and N is defined as L·N=|L||N|cosθ, where θ is the angle between the surface normal vector and the vector from the surface to the light. Because L and N are unit vectors, both |L| and |N| are equal to 1. Therefore, L·N=cosθ. Substituting L·N into equation (1), we obtain D =CI L cosθ. This formula is known as Lambert's cosine law.
[0034] Using Lambert's cosine law, the contribution of all ambient light sources to the diffuse light in an image can be defined as:
number
[0035] Referring back to equation (1) and rewriting it as Lambert's cosine law, the contribution of the fill light source (light source 420 / 422) to the diffuse light in the image can be defined as:
number
number
[0036] When an image illuminated by one of the auxiliary light sources (light source 420 or 422) is captured using system 400 of Figure 4, the resulting image contains diffuse light contributions from both the auxiliary light source and the ambient light.
number
number
[0037] Because ambient light sources are inherently uncontrollable and may contain undesirable variations in light source, it is desirable to first subtract the ambient light contribution from each of the supplementally illuminated images. Equations (5) and (6) can be rewritten by rearranging the terminology to yield:
number
number
[0038] 1 and 2, where the edges of the palette boxes are difficult to discern due to color and other 2D features on their top surfaces, another objective of the techniques of this disclosure is to completely remove color-dependent intensity variations from the output image. The output image Q can be defined as:
number
[0039] Substituting equations (3) and (4) into equation (9), we get:
number
number
[0040] In equation (11), color component C does not appear because it is canceled out from the numerator and denominator. Therefore, the image segmentation technique of equations (9)-(11) eliminates color (i.e., pixel intensity variations associated with different colors) from the output image Q.
[0041] Light intensity value I L1 , I L2 Assuming that is constant over the image field of view, equation (11) can be further modified to:
number
[0042] Equation (12) clearly shows that brightness variations between points in the output image Q are caused only by variations in the relative angles (θ1 and θ2) from the normal vector to the two lights at each point on the surface. For flat surfaces, the normal vector is constant, so all of the pixels of a given flat surface have the same constant value. Pixels of different flat surfaces have different constant values depending on the direction of the normal vector of the surface.
[0043] It is understood that the above equations are all approximations, since light rays from a real physical source do not all emanate from the same distant point, and therefore are not all parallel, nor do they all have the same intensity. However, in practice, the approximations have been found to be good enough to make the Planar Purge method useful.
[0044] Substituting equations (7) and (8) into equation (9), we get:
number
[0045] As mentioned above, specular reflection I D1,A By ignoring I D1,A is the image I1 (ambient light + light source 420) captured using system 400, and I D2,A is the image I2 (ambient light + light source 422) captured using system 400, and I DA is an image I captured using system 400 A (ambient light only). Thus, the three images of object 410 captured by sensor 430 can be combined to produce an output image Q that is devoid of color and other 2D features as follows:
number
[0046] The above discussion explains how the diffuse reflection component is treated in the calculation. Peripheral specular reflections are also removed by the image subtraction process in equation (13). Specular reflections added by fill light sources are not removed, which is a good thing, because they can help reveal the top edges and other 3D features of the object 410. In the output image Q, the only variations of the top surface 412 that remain are due to the angle between the surface normal at each point and the light direction, which are due to variations in the 3D shape of the surface 412 (such as wrinkles, flap seams, and box edges).
[0047] Many of equations (1) through (14) involve image operations, specifically image subtraction and image division. Each of these operations is a pixel-by-pixel calculation using pixel intensity values for each pair of corresponding pixels in the two images. The pixel intensity value for each pixel is an integer value in a predefined range, such as 0-255, or 0-4095. Other pixel intensity value ranges may also be used.
[0048] For example, consider an example where each image is 1000 x 1000 pixels (1000 rows x 1000 columns = 1 million pixels). If pixel 1 (row 1, column 1) in the first image has an intensity of 247 and the corresponding pixel 1 in the second image has an intensity of 133, when the second image is subtracted from the first image, pixel 1 in the resulting image will have an intensity of 114 (=247-133). Such calculations are performed for each pixel in each pair of images being subtracted. The same pixel-by-pixel calculation concept applies to image division: the intensity value of a pixel in one image is divided by the intensity value of the corresponding pixel in another image.
[0049] Each image I A, I1, and I2 are imaged using equal exposure times, which are chosen to be low enough to avoid clipping the intensity values of light-colored pixels in any of the images (i.e., the maximum intensity value of all pixels in all images should be less than a maximum value such as 255 or 4095).
[0050] The same technique can be applied to color images to remove 2D features if the same process (Equation (14)) is performed separately on the red, green, and blue channels. However, since the 3D features remaining in the image do not have distinct colors, there is no particular advantage to using color.
[0051] Using the techniques described above, system 400 can be employed to capture an image of object 410, which can be rapidly processed to produce an output image that is cleared of color and other 2D features and is particularly suitable for subsequent processing, such as segmentation to identify the true edges and corners of a box and thereby determine the box's shape and dimensions required for robotic box depalletizing operations.
[0052] Figure 5 is a flowchart 500 of a method for acquiring an image of an object, processing the image, and generating an output image free of intensity variations due to 2D features, such as color, graphics, tape, etc., according to one embodiment of the present disclosure. In step 502, a workspace such as that shown in Figure 4 is provided, in which system 400 is placed, including light sources 420 / 422 and a 2D sensor 430. At least sensor 430 is in communication with computer 440. Flowchart 500 in Figure 5 illustrates a general image analysis methodology, and the specific orientation shown in Figure 4 (horizontal top surface, multiple lights, and sensor / camera down-pointing angle) should be construed as merely a non-limiting example.
[0053] In step 504, three images of the object 410 are captured using the sensor 430. The three images are a first image I in ambient light only; A , a second image I1 of ambient light and light source 420, and a third image I2 of ambient light and light source 422. These images may be acquired fully manually (by manually controlling light sources 420 / 422 and sensor 430 with switches / buttons), or fully automated (by controlling the operation of light sources 420 / 422 and sensor 430 with computer 440), or any combination thereof. A , I1, I2 are provided to a computer 440 for processing.
[0054] In step 506, image subtraction is performed as in equations (7) and (8) above, and similarly in the numerator and denominator of equation (14). That is, the first intermediate image (representing the diffuse reflection from the first auxiliary light source) is calculated as I1-I A The second intermediate image (representing the diffuse reflection from the second supplementary light source) is calculated as I2-I A Subtracting the ambient image removes the ambient diffuse and specular components from images I1 and I2.
[0055] In step 508, the images are divided as in equation (14) above. That is, one of the intermediate images calculated in step 506 is divided by the other. By dividing the first and second assist light source images (after subtraction of the periphery), color and other 2D surface features are removed from the output image Q. In the division, one of the first and second assist light source images may be used as the numerator and the other as the denominator. This is because which light of the light sources 420 / 422 corresponds to image I1 and which light corresponds to image I2 is ultimately a matter of arbitrary definition, as will be explained below.
[0056] In step 510, the output image Q is used for further processing or analysis. For example, the output image Q is particularly suitable for computing a "segmentation" of boxes on a palette, where individual boxes are identified by their edges, which are normally visible in the output image Q, but color and 2D features are purged from the output image Q.
[0057] It should be noted that Figure 5 illustrates the disclosed methodology in terms of very deliberate steps. In actual implementation, the planar purge technique is essentially performed in two steps: capturing the image and calculating the output image Q using Equation (14). Calculating the output image Q does not necessarily require the calculation of intermediate images. Instead, the output image Q may be calculated directly by applying Equation (14) (two subtractions followed by a division) to each pixel in one step.
[0058] The aforementioned technique has been developed and demonstrated in numerous experiments, showing that the removal of color and 2D features from flat surfaces is highly effective for a variety of objects, in particular the top surfaces of pallets containing different styles and arrangements of boxes. Using the disclosed technique, "removing color" from an output image does not simply mean converting color to grayscale, as is done by a black-and-white camera. Rather, "removing color" means completely purging all intensity variations due to color and 2D features from the output image, making it appear as if the color, markings, or 2D features were never present on the original object. Three examples are presented and described below.
[0059] Figure 6A is an image of the collection of boxes from Figure 1, and Figure 6B is the corresponding output image after flat features have been purged using the system of Figure 4 and the method of Figure 5. Figure 6A recreates image 100 from Figure 1 above, in which the top layer 110 of boxes has numerous graphic designs and other features on their surfaces, including dark bars 130 and light-to-dark transitions 140, 150.
[0060] FIG. 6B is an image 600 created using the planar purge technique described above. Specifically, image 600 is a composite of three input images I A , I1, and I2 using equation (14). It is noteworthy that none of the box's geometric features (130, 140, and 150) from image 100 appear in image 600. Thus, the problem of misidentifying the box's edges due to 2D features 130, 140, and 150 is resolved. In image 600, only the 3D features of the boxes are visible on the top surface of the box's top layer 110. Most of these 3D features are the true edges of the boxes (610, 620, and 630), whose visibility has been enhanced by the image processing techniques of the present disclosure. A small dent in one of the boxes is also visible, shown as 640. Image 600 demonstrates how the planar purging techniques of the present disclosure effectively eliminate 2D features while highlighting 3D features within an object. It is clear that image 600 is much preferable to image 100 for further analysis applications such as box segmentation.
[0061] Figure 7A is an image of the collection of boxes from Figure 2, and Figure 7B is the corresponding output image after planar features have been purged using the system of Figure 4 and the method of Figure 5. Figure 7A recreates image 200 from Figure 2 above, where the top layer 210 of the boxes has many 2D features (lightness and reflectivity) associated with transparent tape applied to its surface.
[0062] 7B is an image 700 created using the planar purge technique described above. Specifically, image 700 is a composite of three input images I A , I1, and I2 using equation (14), including the same top layer 210 of the box. It is notable that in image 700, the transparent tape, both for lightness / darkness effects and specular reflection, has disappeared. Where strips of tape 220 were located in FIG. 7A, they are completely invisible in FIG. 7B (just plain cardboard). Where strips of tape 230 cover the flap seams in FIG. 7A, only the flap seam 730 itself (with its 3D trough shape) is visible in FIG. 7B. Where strips of tape 240 were applied along the edges of the box in FIG. 7A, the edges of those boxes are visible in FIG. 7B as 740. Other box edges are also clearly visible in FIG. 7B, e.g., as shown at 750, 760, and 770. Image 700 again shows how the planar purging technique of the present disclosure effectively eliminates 2D features while highlighting 3D features within the object. It is clear that image 700 is far preferable to image 200 for further analysis applications such as box segmentation.
[0063] Figure 8A is an image 800 of a collection of flat packages, and Figure 8B is a corresponding output image 850 in which planar features have been purged using the system of Figure 4 and the method of Figure 5. Figures 8A and 8B show another clear and dramatic example of the effectiveness of the disclosed planar purging technique, in which light / white packages and dark / black packages appear the same in output image 850, and labels, marks, and other 2D features are absent in output image 850. It is clear that image 850 is far preferable to image 800 for further analysis applications such as package segmentation.
[0064] Many other stacks of boxes and other arrangements of packages have been processed using the planar purging techniques of the present disclosure with equally dramatic results, eliminating 2D features from the target surface of the box / package while highlighting 3D features such as, but not limited to, print, graphics, labels, colors, straps, and tape, etc.
[0065] As noted above, the division process in equation (14) is performed between one of the images of the first and second fill light sources in the numerator and the other of the images of the first and second fill light sources in the numerator. Reversing the numerator and denominator of equation (14) (or reversing the light sources associated with images I1 and I2) has the effect of changing the intensity levels at the pixels of output image Q; however, in either case, output image Q remains devoid of color effects and other 2D features; only 3D features are apparent in output image Q at the top surface 412 of object 410. FIG. 7B shows output image Q based on a particular demo workspace established by image I1 associated with one fill light source and image I2 associated with the other fill light source. If the light source or image division were reversed, the effect would be that the near / left end of the box (indicated by arrow 780) would become very dark instead of very bright, and the near / right side of the box (indicated by arrow 790) would become very bright instead of very dark. The top surface of the stack of boxes is essentially uniform in appearance, except for the highly visible seams at the edges and flaps of the boxes.
[0066] The dramatic results seen in Figures 6A-6B, 7A-7B, and 8A-8B—where planar features are suppressed and 3D features are enhanced—are, in part, the result of the considered placement of the fill light sources. For example, in Figures 7A-7B, two fill light sources are preferably positioned so that the aiming vector of one light source (e.g., light source 420) has a horizontal component perpendicular to one edge of the stack box, and the aiming vector of the other light source (e.g., light source 422) has a horizontal component perpendicular to the adjacent edge of the stack box (90° clockwise or counterclockwise as viewed from above). This configuration provides the most dramatic bright reflection lines and dark shadow lines at the edges of the boxes, as seen in Figure 7B. Other fill light source placements may be appropriate for other types of objects. For example, where the boxes have shapes or stacking arrangements which result in box edges which are on diagonal angles relative to the pallet itself, the lights may be aligned with an aiming vector perpendicular to those diagonal edges.
[0067] As previously mentioned, the system of Figure 4 and the before and after images shown in Figures 6A-8B are provided to clearly explain and demonstrate the effectiveness of the planar purge technique applied to a horizontal top surface. However, we emphasize that the planar purge technique is applicable to any relatively flat surface, not just horizontal top surfaces. That is, the disclosed technique can be applied equally effectively to remove color and 2D features from vertical surfaces (e.g., the side of a box) and beveled surfaces (the surfaces of objects that are oddly shaped or positioned so as to have a surface that is neither horizontal nor vertical).
[0068] Consider the example of a collection of boxes shown in Figures 7A and 7B. The disclosed planar purge technique may be applied to generate an image of the proximal / left end of the box (indicated by arrow 780) that is not the top surface but is devoid of color and 2D features. To generate a planar purged image of the proximal / left end of the box, two auxiliary light sources are provided, both of which are incident on the proximal / left end, and three images (I A , I1, I2) and use equation (14) to calculate the output image Q. The same concept can be applied to other vertical and sloping surfaces.
[0069] Continuing with the example of Figures 7A and 7B, it is easy to imagine how an output image of the entire stack of boxes could be created, with all surfaces purged of intensity variations due to color and 2D features. This could be achieved by providing multiple supplemental light sources at different locations around the stack of boxes, and generating an ambient light image (I) for each flat surface (e.g., top, proximal / left edge, and proximal / right side). A), and then capturing a pair of auxiliary light source images (I1, I2). The corresponding images can then be used to calculate an output image Q for each of the three planar faces (top, proximal / left end, and proximal / right side), and then a composite output image Q can be generated using the planar purged images for each of the three planar faces. It is easy to imagine how such fully planar purged composite images would be most useful for use in 3D box segmentation of the entire stack of boxes.
[0070] All of the above discussions are directed to removing color and 2D features from images of flat surfaces. Using more advanced lighting, image capture, and computational techniques, the same concepts can be applied to removing color and 2D features from non-flat surfaces. It may be impractical to completely remove coloration effects for all surface orientations on objects of arbitrary shape. However, the disclosed techniques may produce a similar effect of reducing shading variations due to surface color without affecting shading variations due to 3D shape. Thus, it can be summarized that the disclosed image processing techniques can be used to reduce shading variations in output images for objects of arbitrary shape, and that the techniques work particularly well for images of flat surfaces, where intensity variations in the image due to color and 2D features can be completely eliminated.
[0071] The disclosed planar purge technique describes an imaging method. While the above example illustrates the application of the method to box segmentation in an industrial environment, many other applications are envisioned. For example, the disclosed technique can be applied to other computer vision and non-computer vision applications, including, but not limited to: Visual surface inspection using computer vision Imaging device that enhances surface appearance for human inspection Artistic effects for photographic images Filmmaking special effects Applications involving cameras or lights that use non-visible wavelengths such as infrared or ultraviolet, so that the flash is not noticeable to a human observer Outdoor images for camouflage removal or reduction: Corporate intelligence gathering: Imaging prototype cars and trucks with camouflage paint and removing or reducing the camouflage to make their shapes clearer. Military Surveillance: Imaging camouflaged weapons and vehicles using fixed or aerial drones equipped with thermal cameras and light sources to create images with the camouflage removed or reduced to improve visibility.
[0072] Throughout the preceding discussion, various image computations are described and contemplated. It will be understood that these computations may be performed by software applications and modules on a computer, such as computer 440 in FIG. 4. Computer 440, equipped with one or more processors and memory and input / output ports, receives and processes input images from sensor 430, provides output images resulting from equation (14), and may also control sensor 430 and light sources 420 / 422. Computer 440 may also be a robot controller—i.e., the controller of a robot performing box depalletizing operations. In such an embodiment, sensor 430 provides its images / data directly to a robot controller that performs the image computations, eliminating the need for an additional computer.
[0073] As previously described, embodiments in which planar features from 2D images are purged offer significant advantages over existing image processing methods, providing an output image of an object that is essentially devoid of 2D features, including color, at the top surface, where the output image is suitable for direct human viewing and / or further computer processing, such as for box segmentation applications.
[0074] Having described a number of exemplary aspects and embodiments of methods and systems for removing 2D features from planar surfaces and enhancing 3D features in 2D images, those skilled in the art will recognize modifications, permutations, additions, and subcombinations thereof. Accordingly, the appended claims, and any claims hereafter introduced, are intended to be construed as including all such modifications, permutations, additions, and subcombinations as fall within their true spirit and scope. The following additional notes are provided regarding the above embodiments and modifications. (Appendix 1) A method for removing two-dimensional features (2D features) from an image, comprising: providing a workspace having a single 2D sensor at a fixed position and orientation, and a first auxiliary light source and a second auxiliary light source fixed at positions different from each other; capturing, with the 2D sensor, a first input image of the object in ambient light, a second input image of the object in the ambient light and the first auxiliary light source, and a third input image of the object in the ambient light and the second light source; 1. A method for removing 2D features in a computer having a processor and a memory, comprising: subtracting the first input image from the second input image to generate a first difference; subtracting the first input image from the third input image to generate a second difference; and dividing the first difference by the second difference to calculate an output image in which the 2D features have been removed. (Appendix 2) The removal method described in Appendix 1, characterized in that the 2D sensor is a 2D camera. (Appendix 3) 2. The removal method of claim 1, wherein the object has a flat surface from which the 2D features are removed in the output image. (Appendix 4) 4. The method of claim 3, wherein the 2D sensor is oriented vertically or at an angle inclined toward the flat surface. (Appendix 5) 4. The method of claim 3, wherein the auxiliary light source is oriented at an oblique angle relative to the flat surface. (Appendix 6) subtracting the first input image from the second input image and the first input image from the third input image includes subtracting pixel intensity values on a corresponding pixel-by-pixel basis; 2. The removal method of claim 1, wherein dividing the first difference by the second difference includes dividing pixel intensity values on a corresponding pixel basis. (Appendix 7) The removal method of Appendix 6, wherein calculating the output image includes subtracting a portion or all of the first input image from a corresponding portion or all of the second input image to calculate a first intermediate image, subtracting a portion or all of the first input image from a corresponding portion or all of the third input image to calculate a second intermediate image, and then calculating the output image by dividing the first intermediate image by the second intermediate image. (Appendix 8) 7. The removal method of claim 6, wherein calculating the output image includes calculating the first difference and the second difference for each pixel of the output image, and dividing the first difference by the second difference. (Appendix 9) the objects are a plurality of boxes arranged on a pallet, 2. The removal method of claim 1, further comprising using the output image in a box segmentation calculation, wherein edges of the boxes are identified in the output image and dimensions and shapes of individual boxes are determined from the edges. (Appendix 10) the object is a plurality of flat packages arranged on a surface; The removal method of claim 1, further comprising using the output image in a package qualification calculation, wherein edges of the packages are identified in the output image and dimensions and shapes of individual packages are determined from the edges. (Appendix 11) the object has one or more curved surfaces that exclude the 2D features in the output image; 2. The removal method of claim 1, wherein multiple auxiliary light sources are provided in the workspace and multiple input images are used to selectively remove the 2D features from local portions of the output image. (Appendix 12) the object has a plurality of flat surfaces; 2. The removal method of claim 1, wherein a plurality of auxiliary light sources are provided in the workspace, and a plurality of input images are used to selectively remove the 2D features from each of the planar surfaces in separate output images, and the separate output images are combined in a composite output image in which the 2D features have been removed from the planar surfaces. (Appendix 13) 2. The removal method of claim 1, wherein the 2D features removed from the output image include variations in intensity due to color, markings, graphics, and tape. (Appendix 14) A method for removing two-dimensional features (2D features) from an image, comprising: capturing, with a 2D sensor, a first input image of an object in ambient light, a second input image of the object in the ambient light and a first assist light source, and a third input image of the object in the ambient light and a second assist light source; 1. A method for removing 2D features in a computer having a processor and a memory, comprising: subtracting the first input image from the second input image to generate a first difference; subtracting the first input image from the third input image to generate a second difference; and dividing the first difference by the second difference to calculate an output image in which the 2D features have been removed. (Appendix 15) A removal system for removing two-dimensional features (2D features) from an image of an object, comprising: a 2D sensor at a fixed position and orientation directed towards the object; a first auxiliary light source and a second auxiliary light source at different fixed positions directed toward the object; a computer having a processor and a memory, the computer in communication with the 2D sensor; The computer receiving from the 2D sensor a first input image of the object in ambient light, a second input image of the object in the ambient light and the first assist light source, and a third input image of the object in the ambient light and the second assist light source; subtracting the first input image from the second input image to generate a first difference, subtracting the first input image from the third input image to generate a second difference, and dividing the first difference by the second difference to calculate an output image in which the 2D features have been removed. (Appendix 16) 16. The removal system of claim 15, wherein the 2D sensor is a 2D camera. (Appendix 17) 16. The removal system of claim 15, wherein the object has a flat surface from which the 2D features are removed in the output image. (Appendix 18) 18. The removal system of claim 17, wherein the 2D sensor is oriented vertically or at an angle inclined relative to the flat surface. (Appendix 19) 18. The removal system of claim 17, wherein the auxiliary light source is oriented at an oblique angle relative to the flat surface. (Appendix 20) subtracting the first input image from the second input image and the first input image from the third input image includes subtracting pixel intensity values on a corresponding pixel-by-pixel basis; 16. The removal system of claim 15, wherein dividing the first difference by the second difference includes dividing a pixel intensity value by a corresponding pixel unit. (Appendix 21) 21. The removal system of claim 20, wherein calculating the output image includes subtracting a portion or all of the first input image from a corresponding portion or all of the second input image to calculate a first intermediate image, subtracting a portion or all of the first input image from a corresponding portion or all of the third input image to calculate a second intermediate image, and then calculating the output image by dividing the first intermediate image by the second intermediate image. (Appendix 22) 21. The removal system of claim 20, wherein calculating the output image includes, for each pixel of the output image, calculating the first difference and the second difference, and dividing the first difference by the second difference. (Appendix 23) 16. The removal system of claim 15, wherein the computer controls the 2D sensor and the auxiliary light source to automatically capture the first input image, the second input image, and the third input image. (Appendix 24) the objects are a plurality of boxes arranged on a pallet, The removal method may further comprise the steps of: using the output image to calculate a box segmentation by the computer or a different computer; 16. The removal system of claim 15, wherein edges of the boxes are identified in the output image and the dimensions and shape of each of the boxes are determined from the edges. [Explanation of symbols]
[0075] 400 System 410 Object 412 Top surface 420 light source 422 Light source 430 Sensors 440 Computer
Claims
1. 1. A method for removing two-dimensional features (2D features) from an image, comprising: providing a workspace comprising a single 2D sensor at a fixed position and orientation, and a first auxiliary light source and a second auxiliary light source fixed at different positions from each other; capturing, with the 2D sensor, a first input image of the object in ambient light, a second input image of the object in the ambient light and the first auxiliary light source, and a third input image of the object in the ambient light and the second light source; 1. A method for removing 2D features in a computer having a processor and a memory, comprising: subtracting the first input image from the second input image to generate a first difference; subtracting the first input image from the third input image to generate a second difference; and dividing the first difference by the second difference to calculate an output image in which the 2D features have been removed.
2. The method of claim 1 , wherein the 2D sensor is a 2D camera.
3. The method of claim 1 , wherein the object has a planar surface from which the 2D features are removed in the output image.
4. The method of claim 3 , wherein the 2D sensor is oriented vertically or at an angle oblique to the flat surface.
5. The method of claim 3 , wherein the assist light source is oriented at an oblique angle relative to the flat surface.
6. subtracting the first input image from the second input image and the first input image from the third input image includes subtracting pixel intensity values on a corresponding pixel-by-pixel basis; The method of claim 1 , wherein dividing the first difference by the second difference comprises dividing pixel intensity values on a per-pixel basis.
7. 7. The removal method of claim 6, wherein calculating the output image includes calculating a first intermediate image by subtracting a portion or all of the first input image from a corresponding portion or all of the second input image, calculating a second intermediate image by subtracting a portion or all of the first input image from a corresponding portion or all of the third input image, and then calculating the output image by dividing the first intermediate image by the second intermediate image.
8. 7. The method of claim 6, wherein calculating the output image comprises: calculating, for each pixel of the output image, the first difference and the second difference; and dividing the first difference by the second difference.
9. the objects are a plurality of boxes arranged on a pallet, 2. The method of claim 1, further comprising using the output image in a box segmentation calculation, wherein edges of the boxes are identified in the output image and dimensions and shapes of individual boxes are determined from the edges.
10. the object is a plurality of flat packages arranged on a surface; 2. The method of claim 1, further comprising using the output image in a package qualification calculation, wherein edges of the packages are identified in the output image and dimensions and shapes of individual packages are determined from the edges.
11. the object has one or more curved surfaces that exclude the 2D features in the output image; The method of claim 1 , wherein a plurality of auxiliary light sources are provided in the workspace, and a plurality of input images are used to selectively eliminate the 2D features from localized portions of the output image.
12. the object has a plurality of flat surfaces; 10. The method of claim 1, wherein a plurality of auxiliary light sources are provided in the workspace, and a plurality of input images are used to selectively remove the 2D features from each of the planar surfaces in separate output images, and the separate output images are combined in a composite output image in which the 2D features have been removed from the planar surfaces.
13. The method of claim 1 , wherein the 2D features removed from the output image include variations in intensity due to color, markings, graphics, and tape.
14. 1. A method for removing two-dimensional features (2D features) from an image, comprising: capturing, with a 2D sensor, a first input image of an object in ambient light, a second input image of the object in the ambient light and a first assist light source, and a third input image of the object in the ambient light and a second assist light source; 1. A method for removing 2D features in a computer having a processor and a memory, comprising: subtracting the first input image from the second input image to generate a first difference; subtracting the first input image from the third input image to generate a second difference; and dividing the first difference by the second difference to calculate an output image in which the 2D features have been removed.
15. 1. A system for removing two-dimensional features (2D features) from an image of an object, comprising: a 2D sensor at a fixed position and orientation aimed at the object; a first auxiliary light source and a second auxiliary light source at different fixed positions directed toward the object; a computer having a processor and a memory, the computer in communication with the 2D sensor; The computer receiving from the 2D sensor a first input image of the object in ambient light, a second input image of the object in the ambient light and the first assist light source, and a third input image of the object in the ambient light and the second assist light source; subtracting the first input image from the second input image to generate a first difference, subtracting the first input image from the third input image to generate a second difference, and dividing the first difference by the second difference to calculate an output image in which the 2D features have been removed.
16. The removal system of claim 15 , wherein the 2D sensor is a 2D camera.
17. The removal system of claim 15 , wherein the object has a planar surface from which the 2D features are removed in the output image.
18. 20. The removal system of claim 17, wherein the 2D sensor is oriented vertically or at an oblique angle relative to the flat surface.
19. 20. The removal system of claim 17, wherein the auxiliary light source is oriented at an oblique angle relative to the planar surface.
20. subtracting the first input image from the second input image and the first input image from the third input image includes subtracting pixel intensity values on a corresponding pixel-by-pixel basis; The removal system of claim 15 , wherein dividing the first difference by the second difference comprises dividing pixel intensity values on a corresponding pixel-by-pixel basis.
21. 21. The removal system of claim 20, wherein calculating the output image includes subtracting a portion or all of the first input image from a corresponding portion or all of the second input image to calculate a first intermediate image, subtracting a portion or all of the first input image from a corresponding portion or all of the third input image to calculate a second intermediate image, and then calculating the output image by dividing the first intermediate image by the second intermediate image.
22. 21. The removal system of claim 20, wherein calculating the output image comprises, for each pixel of the output image, calculating the first difference and the second difference, and dividing the first difference by the second difference.
23. The removal system of claim 15 , wherein the computer controls the 2D sensor and the auxiliary light source to automatically capture the first input image, the second input image, and the third input image.
24. the objects are a plurality of boxes arranged on a pallet, The removal method may further comprise the steps of: using the output image to calculate a box segmentation by the computer or a different computer; The removal system of claim 15 , wherein edges of the boxes are identified in the output image and the size and shape of each of the boxes are determined from the edges.