System and method for stitching continuous images of an object

The system addresses the challenge of capturing high-resolution images of moving objects by using two-dimensional cameras and a controller to stitch together consecutive images, resulting in efficient and accurate image reconstruction.

JP7696692B2Active Publication Date: 2025-06-23COGNEX CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2019082792
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-04-25
Filing Date
2019-04-24
Publication Date
2025-06-23
Estimated Expiration
2039-04-24

AI Technical Summary

Technical Problem

Existing image stitching technologies face challenges in efficiently capturing and stitching high-resolution images of moving objects, particularly the bottom surface of objects on high-speed conveyor belts, due to physical constraints and the need for precise alignment and processing.

Method used

A system utilizing one or more two-dimensional cameras and a controller to capture consecutive images of an object's surface as it moves through a conveyor system. The controller stitches these images together using a stitching algorithm, which includes determining a two-dimensional coordinate transformation, warping, and mixing the images to generate a high-resolution stitched image.

Benefits of technology

The system effectively captures and stitches high-resolution images of moving objects, overcoming the limitations of traditional line scan cameras and achieving efficient image reconstruction with improved resolution and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696692000001
    Figure 0007696692000001
  • Figure 0007696692000002
    Figure 0007696692000002
  • Figure 0007696692000003
    Figure 0007696692000003
Patent Text Reader

Abstract

To provide systems and methods for stitching sequential images of an object.SOLUTION: A system comprises: a transport device for moving at least one object; at least one 2D digital optical sensor; and a controller operatively coupled to the 2D digital optical sensor. The controller performs the steps of: a) receiving a first digital image; b) receiving a second digital image; and c) stitching the first digital image and the second digital image using a stitching algorithm to generate a stitched image.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Field of the Technology) The present invention relates to the technical field of image stitching. In particular, the present invention relates to the fast stitching of two-dimensional images generated in a conveyor system with windows.

Background Art

[0002] (Background of the Invention) Vision systems for performing measurement, inspection, alignment of objects, and / or decoding of symbol systems such as one-dimensional and two-dimensional barcodes are used in a wide range of applications and industries. These systems process the acquired images by using an image sensor to acquire an image of the object and using a mounted or interconnected vision system processor. Images are typically captured as an array of pixels where each pixel has various colors and / or intensities.

[0003] A common use for such imaging systems is to track and classify objects moving along a conveyor in manufacturing and logistics operations. Typically, such imaging systems capture all sides of the object being tracked, which in many cases results in the capture of the six sides of a cubic object. In such a system, it is necessary to capture the bottom of the object, which is the side of the object in contact with the conveyor.

[0004] In some instances, a line scan camera can be employed to address the movement of objects and wide fields of view. However, such a solution is not applicable to certain object geometries and line arrangements. Additionally, while line scan image sensors tend to be less expensive than conventional format area scan sensors, the overall system cost of using line scan image sensors can be significantly higher than that of conventional format area scan sensors due to increased computational and processing requirements. For example, line scan image sensors require that a complete image be reconstructed in software line by line. Further, the alignment of the object's motion and the timing of image acquisition are important.

[0005] Area scan image sensors rapidly image a defined area, enabling easier setup and alignment and greater flexibility than line scan sensors. However, even when using area scan image sensors, partial views of the object may be captured depending on factors such as the object's size, the imager's resolution, and the geometry of the conveyor. Thus, some processing may be required to reconstruct a complete image.

[0006] Another consideration is which side of the object should be imaged, as the system configuration can vary depending on whether the bottom, top, or other side of the object is to be imaged. For example, in some applications, some or all sides of the object, including the bottom, are imaged. In other applications, imaging of only one side is required. For example, the bottom can be imaged instead of the top to avoid complications associated with imaging objects of different heights that may have different distances from the camera.

[0007] In some applications, one or more cameras may be used to image each side of an object, and in other applications, two or more cameras may be used to image one side of an object. In applications where one side of an object is imaged, a single standard-resolution camera may not provide an appropriate resolution, and an appropriate target area and two or more cameras may be used to image that side of the object.

Summary of the Invention

Means for Solving the Problems

[0008] Embodiments of the present invention may provide a high-resolution image of the surface of an object as the object of various shapes and sizes moves through a conveyor system. Such embodiments may relate to a solution to the problem of scanning the bottom surface of a box or other object moving on a high-speed conveyor belt by using one or more two-dimensional cameras. Current systems rely on line scan cameras to image objects such as boxes.

[0009] Scanning the bottom surface in a logistics tunnel application is particularly challenging because the physical constraints of the system limit the visibility of the bottom surface of the box to only a small slice at a time. To generate an image for the bottom surface, small image slices can be acquired (through the gap between two sections of the conveyor belt) and then stitched together.

[0010] In an embodiment, the system is a transport device for moving at least one object, wherein at least one substantially planar surface of the object is moved within a plane that is locally known in the vicinity of the viewing area, and the substantially planar surface of the object is blocked except when at least one substantially planar surface passes through the viewing area; a transport device; at least one two-dimensional digital optical sensor configured to capture at least two consecutive two-dimensional digital images of at least one substantially planar surface of at least one object that is moved within a plane that is known in the vicinity of the viewing area; and a controller operably coupled to the two-dimensional digital optical sensor, the controller performing: a) receiving a first digital image; b) receiving a second digital image; and c) stitching together the first digital image and the second digital image using a stitching algorithm to generate a stitched image.

[0011] In an embodiment, the controller may repeatedly perform steps a) to c). The first digital image may comprise at least one of a captured digital image or an image as a result of a previously stitched image. The surface moved through the known plane may be one of the bottom surface or the side surface of the object, the transport device may comprise a viewing area positioned on a corresponding surface of the transport device, and the two-dimensional digital image may be captured through the viewing area. The transport device may comprise a conveyor, and the viewing area may comprise one of an optical window in the conveyor, a gap in the surface of the conveyor, or a gap between two conveyor devices.

[0012] In an embodiment, the system may further include a light source configured to illuminate one or more objects in the viewing area, the controller may be coupled to the light source to control the illumination from the light source, and the controller intermittently blinks the light source. The sensor may be configured to detect the presence or absence of an object on the transport device and may control the acquisition of an image based on the presence or absence of an object on the transport device. The system may further include a mirror, and the two-dimensional digital optical sensor may capture a digital image through the mirror. The controller may decode a mark based on the stitched image.

[0013] In an embodiment, the stitching algorithm may include determining a two-dimensional coordinate transformation to be used to align a first digital image and a second digital image, warping the second digital image by using the found two-dimensional coordinate transformation, and mixing the warped second digital image with the first digital image. The transport device may operate according to a substantially linear motion, the controller may include a model of the substantially linear motion of the transport device, and the stitching algorithm may use the model of the substantially linear motion of the transport device. The system may further include a motion encoder, control the acquisition of an image based on the motion encoder, and determining the image transformation may be further based on an estimation of object translation generated by the motion encoder. The system may further include a light source configured to illuminate one or more objects in the viewing area, the controller may be coupled to the light source to control the illumination from the light source based on an estimation of object translation generated by the motion encoder, and the controller may intermittently blink the light source based on an estimation of object translation generated by the motion encoder.

[0014] In an embodiment, the two-dimensional optical sensor may be configured to capture a reduced field of view, the reduced field of view being determined manually or by analyzing the entire image acquired at a set time. The approximate two-dimensional coordinate transformation may be estimated through a training process that uses a plurality of digital image slices captured for a calibration moving object having features of a known pattern. The system may further comprise a plurality of two-dimensional digital optical sensors, each two-dimensional digital optical sensor being configured to capture a plurality of consecutive digital images associated with one or more objects through a field of view, and the plurality of two-dimensional digital optical sensors being configured to capture the digital images substantially simultaneously.

[0015] In an embodiment, at least one of each digital image from a camera may be stitched together with image slices captured substantially simultaneously by each other camera to form a combined image slice, and the combined image slices may be stitched together to form a complete image, or each digital image from a camera may be stitched together with digital images of continuously captured digital images from that camera, and the stitched images from each of the plurality of cameras may be stitched together to form a complete image, or each digital image from a camera may be directly stitched together with an image as the same result of previously stitched images from all cameras.

[0016] In an embodiment, a computer-implemented method is to move at least one object by a transport device, wherein at least one substantially planar surface of the object is moved in a locally known plane around the field of view, and the substantially planar surface of the object is blocked except when at least one substantially planar surface passes through the field of view, and capturing at least two consecutive two-dimensional digital images of the surface of the object that is being translated in a locally known plane around the field of view by at least one two-dimensional digital optical sensor, and stitching together a first digital image and a second digital image by using a stitching algorithm to generate a stitched image.

[0017] In an embodiment, the stitching may be repeatedly performed and may include determining a two-dimensional coordinate transformation to be used to align a first digital image and a second digital image, warping the second digital image by using the found two-dimensional coordinate transformation, and mixing the warped second digital image with the first digital image. The controller may decode a mark based on the stitched image. There may be a plurality of two-dimensional digital optical sensors that capture images, and at least one of each digital image from the two-dimensional digital optical sensors may be stitched together with an image slice captured substantially simultaneously by each other two-dimensional digital optical sensor to form a combined image slice, and the combined image slices may be stitched together to form a complete image, or each digital image from each two-dimensional digital optical sensor may be stitched together with a digital image of continuously captured digital images from that two-dimensional digital optical sensor, and the stitched images from each two-dimensional digital optical sensor among the plurality of two-dimensional digital optical sensors may be stitched together to form a complete image, or each digital image from each two-dimensional digital optical sensor may be directly stitched together with an image as the same result of previously stitched images from all two-dimensional digital optical sensors.

[0018] In an embodiment, the stitching may be performed by using parallel processing, or the controller may decode a mark based on the stitched image, and the stitching and decoding may be performed by using parallel processing.

[0019] In an embodiment, stitching may be performed repeatedly, and stitching may include determining a two-dimensional coordinate transformation to be used to align a first digital image and a second digital image, warping the second digital image by using the found two-dimensional coordinate transformation, and mixing the warped second digital image with the first digital image. The controller may decode a mark based on the stitched image. There may be a plurality of two-dimensional digital optical sensors that capture images, and each digital image from the two-dimensional digital optical sensors is stitched together with an image slice captured substantially simultaneously by each of the other two-dimensional digital optical sensors to form a combined image slice, and the combined image slices are stitched together to form a complete image, or each digital image from each two-dimensional digital optical sensor is stitched together with a digital image of a continuously captured digital image from that two-dimensional digital optical sensor, and the stitched images from each of the plurality of two-dimensional digital optical sensors are stitched together to form a complete image, or each digital image from each two-dimensional digital optical sensor is directly stitched together with an image as the same result of previously stitched images from all two-dimensional digital optical sensors. Stitching may be performed by using parallel processing, or the controller may decode a mark based on the stitched image, and stitching and decoding may be performed by using parallel processing.

[0020] The method may further include controlling illumination from a light source configured to illuminate an object in the field of view based on an estimation of object translation generated by a motion encoder. The method may further include intermittently flashing the light source based on an estimation of object translation generated by the motion encoder. The method may further include controlling illumination from a light source configured to illuminate an object in the field of view. The method may further include intermittently flashing the light source. The method may further include detecting the presence or absence of an object on a conveyor device by using a sensor. The method may further include controlling the acquisition of an image based on the presence or absence of an object on the conveyor device. The method may further include controlling illumination from the light source based on the presence or absence of an object on the conveyor device. The method may further include forming a single stitched image based on a fixed distance at which a two-dimensional digital camera is positioned from the field of view. The method may further include capturing a digital image by using a two-dimensional digital camera via a mirror.

[0021] The conveyor may operate according to a substantially linear motion. The method may further include modeling the substantially linear motion of the conveyor. The method may further include using a model of the substantially linear motion of the conveyor for stitching. The method may further include capturing a plurality of digital images substantially simultaneously by using a plurality of two-dimensional digital cameras, each two-dimensional camera being configured to capture a plurality of consecutive digital images associated with one or more objects through the field of view. The method may further include stitching each digital image from a camera with an image captured substantially simultaneously by each other camera to form a combined image, and the combined image slices stitched together will form a complete image. The method may further include stitching each digital image from each camera with a digital image of the continuously captured digital images from that camera, and stitching together the stitched images from each of the plurality of cameras to form a complete image.

[0022] Throughout this specification, many other embodiments are described. All of these embodiments are intended to be within the scope of the invention disclosed herein. It should be understood that while various embodiments are described herein, not all objectives, advantages, features, or concepts need to be achieved according to any particular embodiment. Thus, for example, one skilled in the art will recognize that the invention may be implemented or carried out without necessarily achieving other objectives or advantages as taught or suggested herein in a manner that achieves or optimizes one advantage or group of advantages as taught or suggested herein.

[0023] The methods and systems disclosed herein may be implemented in any means for achieving various aspects and may be executed in the form of a machine-readable medium that realizes a set of instructions for causing a machine to perform any of the operations disclosed herein when executed by the machine. These features, aspects, and advantages of the present invention, as well as other features, aspects, and advantages of the present invention, will be readily apparent to those skilled in the art and will be understood by reference to the following description, the appended claims, and the accompanying drawings, and the invention is not limited to any particular disclosed embodiment(s). The present invention provides, for example, the following items. (Item 1) A system, the system comprising A transport device for moving at least one object, at least one substantially planar surface of the object being moved in a locally known plane around the viewing area, the substantially planar surface of the object being blocked except when the at least one substantially planar surface passes through the viewing area; a transport device; At least one two-dimensional digital optical sensor, wherein the at least one two-dimensional digital optical sensor is configured to capture at least two consecutive two-dimensional digital images of at least one substantially planar surface of at least one object that is being moved within the known plane at the periphery of the viewing area. A controller operably connected to the two-dimensional digital optical sensor, the controller a) receiving a first digital image; b) receiving a second digital image; c) stitching together the first digital image and the second digital image using a stitching algorithm to generate a stitched image; and performing. A system comprising. (Item 2) The system according to the above item, wherein the controller repeatedly performs steps a) to c). (Item 3) The system according to any of the above items, wherein the first digital image comprises at least one of a captured digital image or an image as a result of a previously stitched image. (Item 4) The surface being moved through the known plane is one of the bottom surface of the object or a side surface of the object, The transport device comprises a viewing area positioned on a corresponding surface of the transport device, The system according to any of the above items, wherein the two-dimensional digital image is captured through the viewing area. (Item 5) The system according to any of the above items, wherein the transport device comprises a conveyor, and the viewing area comprises one of an optical window in the conveyor, a gap on the surface of the conveyor, or a gap between two conveyor devices. (Item 6) The system according to any of the above items, further comprising a light source configured to illuminate the one or more objects in the visual recognition area, wherein the controller is connected to the light source to control the illumination from the light source, and the controller intermittently blinks the light source. (Item 7) The system according to any of the above items, further comprising a sensor configured to detect the presence or absence of an object on the transport device and control the acquisition of an image based on the presence or absence of the object on the transport device. (Item 8) The system according to any of the above items, further comprising a mirror, wherein the two-dimensional digital optical sensor captures the digital image through the mirror. (Item 9) The system according to any of the above items, wherein the controller decodes a mark based on the combined image. (Item 10) The stitching algorithm determines a two-dimensional coordinate transformation to be used to align the first digital image and the second digital image, warps the second digital image by using the found two-dimensional coordinate transformation, and mixes the warped second digital image with the first digital image The system according to any of the above items, including. (Item 11) The transport device operates according to a substantially linear motion, the controller includes a model of the substantially linear motion of the transport device, and the stitching algorithm uses the model of the substantially linear motion of the transport device. The system according to any of the above items. (Item 12) The system according to any of the above items, further comprising a motion encoder, wherein the system controls the acquisition of an image based on the motion encoder, and determining the image transformation is further based on the estimation of the object translation generated by the motion encoder. (Item 13) The system according to any of the above items, further comprising a light source configured to illuminate the one or more objects in the viewing area, wherein the controller is coupled to the light source to control the illumination from the light source based on the estimation of the object translation generated by the motion encoder, and the controller intermittently blinks the light source based on the estimation of the object translation generated by the motion encoder. (Item 14) The system according to any of the above items, wherein the two-dimensional optical sensor is configured to capture a reduced field of view determined by one of manually or by analyzing an entire image acquired at a set time. (Item 15) The system according to any of the above items, wherein the approximate two-dimensional coordinate transformation is estimated through a training process using a plurality of digital image slices captured for a calibration moving object having known pattern features. (Item 16) The system according to any of the above items, further comprising a plurality of two-dimensional digital optical sensors, each two-dimensional digital optical sensor being configured to capture a plurality of consecutive digital images through the viewing area, the plurality of consecutive digital images being associated with the one or more objects, and the plurality of two-dimensional digital optical sensors being configured to capture digital images substantially simultaneously. (Item 17) Each digital image from a camera is spliced together with an image slice captured substantially simultaneously by each other camera to form a combined image slice, and the combined image slices are spliced together to form a complete image. Each digital image from each camera is spliced together with the digital images of the continuously captured digital images from that camera, and the spliced images from each of the plurality of cameras are spliced together to form a complete image, or Each digital image from each camera is directly joined to an image as the same result of the previously joined images from all cameras. The system according to any of the above items, which is at least one of them. (Item 18) A computer-implemented method, the computer-implemented method comprising: Moving at least one object by a transport device, wherein at least one substantially planar surface of the object is moved in a locally known plane at the periphery of the viewing area, and the substantially planar surface of the object is blocked except when the at least one substantially planar surface of the object passes through the viewing area; Capturing at least two consecutive two-dimensional digital images of the surface of the object that is being translated in the locally known plane at the periphery of the viewing area by at least one two-dimensional digital optical sensor; Using an algorithm to join a first digital image and a second digital image and joining them to generate a joined image; A computer-implemented method comprising the above. (Item 19) The joining is performed repeatedly, and the joining comprises: Determining a two-dimensional coordinate transformation to be used to align the first digital image and the second digital image; Warping the second digital image by using the found two-dimensional coordinate transformation; Mixing the warped second digital image with the first digital image; The method according to any of the above items, comprising the above. (Item 20) The method according to any of the above items, wherein the controller decodes a mark based on the joined image. (Item 21) There are a plurality of two-dimensional digital optical sensors for capturing images. Each digital image from the two-dimensional digital optical sensor is joined together with image slices captured substantially simultaneously by each other two-dimensional digital optical sensor to form combined image slices, and the combined image slices are joined together to form a complete image, or Each digital image from each two-dimensional digital optical sensor is joined together with the digital images of the continuously captured digital images from that two-dimensional digital optical sensor, and the joined images from each two-dimensional digital optical sensor among the plurality of two-dimensional digital optical sensors are joined together to form a complete image, or The method according to any of the above items, wherein each digital image from each two-dimensional digital optical sensor is directly joined to an image as the same result of the previously joined images from all two-dimensional digital optical sensors. (Item 22) The joining is performed by using parallel processing, or The controller decodes a mark based on the joined image, and the joining and the decoding are performed by using parallel processing, the method according to any of the above items. Abstract (Summary) The system is a transport device for moving at least one object, wherein at least one substantially planar surface of the object is moved in a locally known plane at the periphery of the viewing area, and the substantially planar surface of the object is blocked except when at least one substantially planar surface passes through the viewing area. The system includes a transport device, at least one two-dimensional digital optical sensor configured to capture at least two consecutive two-dimensional digital images of at least one substantially planar surface of at least one object that is moved in a known plane at the periphery of the viewing area, and a controller operably coupled to the two-dimensional digital optical sensor. The controller performs steps of: a) receiving a first digital image; b) receiving a second digital image; and c) stitching together the first digital image and the second digital image using a stitching algorithm to generate a stitched image.

Brief Description of the Drawings

[0024] A more specific description of the invention, briefly summarized above, may be had by reference to the embodiments, some of which are illustrated in the accompanying drawings, in such a way that the above features of the invention can be understood in detail. However, it should be noted that the accompanying drawings illustrate only typical embodiments of the invention, and the invention may admit of other equally effective embodiments.

[0025]

Figure 1

[0026]

Figure 2

[0027]

Figure 3

[0028]

Figure 4

[0029]

Figure 5

[0030]

Figure 6

[0031]

Figure 7a

Figure 7b

[0032]

Figure 8

[0033]

Figure 9

[0034]

Figure 10

[0035]

Figure 11

[0036]

Figure 12

[0037] Other features of this embodiment will become apparent from the following detailed description.

Mode for Carrying Out the Invention

[0038] (Detailed Description of the Embodiment) In the following detailed description of the preferred embodiments, reference is made to the accompanying drawings that form a part hereof, and which are shown by way of illustration of specific embodiments in which the invention may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the invention. Electrical, mechanical, logical, and structural changes may be made to the embodiments without departing from the spirit and scope of the present teachings. Thus, the following detailed description should not be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0039] FIG. 1 illustrates a schematic view from above of a system according to an embodiment. The system may include a transport device such as a conveyor 120 or a robotic arm (not shown), at least one camera 132, and a controller / processor 1200. The conveyor 120 may translate an object 110 such as a parcel across a belt, rollers, or other conveyor mechanism through the field of view of the camera / imager 132. Similarly, a robotic arm (not shown) may grip or otherwise hold an object 110 such as a parcel, for example, by a robotic hand or robotic claw, or by a suction device, etc. The robotic arm may move an object through the field of view of the camera / imager 132. The transport device (regardless of type) may translate the object 110 in a locally known plane at the periphery of a viewing area such as viewing area 121. However, in an embodiment, the translation of the object 110 when it is not in the viewing area may not be a translation in a known plane. Rather, outside the viewing area, the transport device may translate the object 110 in any direction or in any form of motion. In the specific examples described herein, a conveyor or robotic arm may be referred to, but it should be understood that in an embodiment, the transport device may include a conveyor, a robotic arm, or any other device capable of transporting an object.

[0040] The conveyor 120 may have a viewing area 121 on its surface. For example, the viewing area 121 may include a window such as a glass window or a scratch-resistant plastic window, or the viewing area 121 may include an opening. The viewing area 121 can be of any size, but typically, the width of the viewing area 121 (perpendicular to the direction of travel of the conveyor) is as large as possible to enable capture of the entire bottom of each object 110. For example, in an embodiment, the width of the viewing area 121 may be approximately equal to the width of the conveyor 120, or may be slightly smaller than the width of the conveyor 120 to adapt to the structure of the conveyor 120. In embodiments where the viewing area 121 includes a window, the length of the window (in the direction of travel of the conveyor) is not so large as to impede the travel of the objects on the conveyor, but can be relatively large. For example, a typical window can be approximately 1 / 4 inch to 2 inches in length. In embodiments where the viewing area 121 includes a gap, the gap cannot be so large as to impede the travel of the objects on the conveyor. For example, a relatively large gap can be large enough for an object to fall through the gap from the conveyor 120. Similarly, the gap can be large enough for the objects on the conveyor to have a non-uniform motion due to the tip that comes back after a slight fall. In some embodiments, the gap can be provided by an opening in a section of the conveyor 120, while in other embodiments, the gap can be provided by a separation between two sections of the conveyor 120. These are merely examples of viewing areas. In embodiments, the viewing area may have several physical implementations such as a window or a gap, while in other embodiments, the viewing area may be merely an area in a space where the camera / imager 132 can capture an image (of appropriate quality and resolution) of the surface of the object 110 to be processed as described below.

[0041] In the example shown in FIG. 1, the object can be translated along the conveyor 120 to image the bottom surface of the object. In this example, the bottom surface of the object is within the plane of the conveyor 120. Thus, the plane of the conveyor 120 can be considered to form a known plane. In an embodiment, the bottom surface of the object can be imaged in relation to gravity to ensure that the bottom surface of the object is within a known plane, the conveyor 120. And the known plane can form a plane on which the surface of the object can be imaged directly in any device in the known plane or through a window or gap.

[0042] In other embodiments, any surface of the object can be imaged in relation to a device to ensure that the surface of the object is within a known plane or parallel to a known plane. And this known plane can form a plane on which the surface of the object can be imaged directly in any device in the known plane or through a window or gap. Accordingly, embodiments can image the bottom surface, top surface, or any side surface of the object. It should be noted further that such a surface may not be completely planar or flat itself. For example, in some cases, the surface of a box can be planar, but in some cases, the surface of a box may be dented, curved, or folded. In some cases, the object may not be a box and can be an envelope, a pouch, or other irregularly shaped object. The important point is that the surface of the object to be captured is substantially planar, i.e., planar enough to capture an image of the surface of the object with sufficient clarity for the image to be processed as described herein by the camera.

[0043] Camera 132 can capture an image through the viewing area 121. Camera 132 can capture an image directly, or Camera 132 can capture an image through a series of mirrors (not shown). In this example, one camera is shown, but in embodiments, one or more cameras can be used. For example, if the width (perpendicular to the direction of travel of conveyor 120) is too large to be captured by one camera, or if it is too large to be captured with an appropriate image quality by one camera, two or more cameras can be used. Embodiments can be applied to any number of cameras. Examples of such embodiments are described in more detail below.

[0044] Camera 132 can continuously capture images when an object is being translated through the camera's field of view. The continuous images can be captured in a particular order. It is not necessary for the continuous images to be captured one after another without any intermediate images being captured. For example, in embodiments, fewer images than all possible images can be captured, but such images are still continuous. In embodiments, some of the images in the sequence can be skipped, but the resulting images are still continuous. For example, in embodiments, every nth image can be captured (or skipped) to form a sequence, or all images can be captured or skipped at irregular intervals. Similarly, in embodiments, each captured image can be processed, or every nth captured image can be processed (or skipped). The term continuous encompasses any and all such sequences.

[0045] The light source 131 may be associated with the camera to provide appropriate illumination and illuminate objects on dark surfaces. For example, in an embodiment, the light source 131 may include a strobe such as a strobe light or a strobe lamp, and the strobe may create a flash of light when controlled by the controller / processor 1200. Such a strobe light may include an intermittent light source such as a xenon flash lamp or a flash tube, or may include a more continuous light source that operates in an intermittent manner, such as a light-emitting diode (LED), an incandescent lamp, or a halogen lamp. The purpose of the strobe light is to provide a sufficient amount of light at a very short exposure time so that the acquired image does not have motion blur when the object is moving. Compared to a strobe light, a line scan camera may require a constantly lit light, which may not be energy efficient. The intensity and duration of each flash may be controlled by the controller / processor 1200 or other exposure control circuitry based on factors that may include the reflectivity of the object being captured, reflections from any windows that may be present in the viewing area 121, the speed of the conveyor, etc. In other embodiments, the light source 131 may include fixed or constant illumination. Similarly, the intensity of such illumination may be controlled by the controller / processor 1200 or other exposure control circuitry based on factors that may include the reflectivity of the object being captured, reflections from any windows that may be present in the viewing area 121, the speed of the conveyor, etc. The light source 131 may be at a different angle to the window than the camera to prevent oversaturation of a portion of the image due to preventing reflections (e.g., which may occur if the window present in the viewing area 121 is made of glass or another transparent material). In other embodiments, polarization may be used to prevent reflection problems.

[0046] In some embodiments, the system may include a speed or translation sensor or encoder 122 attached to the conveyor 120. The sensor / encoder 122 may be connected to the controller / processor 1200 to provide information about the speed of the conveyor 120 and / or the translation of the object 110 on the conveyor 120. Such speed and / or translation may be measured periodically or otherwise repeatedly to detect changes in speed and / or translation. In an embodiment, the sensor / encoder 122 may be a rotary encoder, shaft encoder, or any other electromechanical device that can convert the angular position or motion of the wheels, shafts, or axles of the conveyor 120, or the wheels, shafts, or axles attached to the conveyor 120, into an analog signal or digital code that can be received by the controller. The controller / processor may convert such an analog signal into a digital representation by using an analog-to-digital converter. In an embodiment, the encoder 122 may provide an absolute or relative speed or translation output.

[0047] In an embodiment, the encoder 122 may determine the translation of an object on the conveyor 120 and may output one or more pulses per unit distance of travel of the object 110 or the conveyor 120. For example, the encoder 122 may be arranged to provide one rotation of rotational motion per 12 inches of travel of the conveyor 120. The encoder 122 may be arranged to provide a number of pulses per rotation, such as 24 pulses per rotation. In this example, one pulse from the encoder 122 may occur for every half-inch of motion of the conveyor 120. In an embodiment, the number of rotations of rotational motion per travel of the conveyor 120 and the number of pulses per rotation may be fixed, or these parameters may be adjustable or programmable.

[0048] The controller / processor 1200 can be arranged to cause the camera 132 to capture an image based on each pulse or many pulses from the encoder 122. Each such captured image can be used to form image slices that are to be stitched together as described below. If the image resolution is known or can be determined, the image resolution for each motion unit for each pulse can be determined. For example, if the camera 132 is arranged or determined to capture an image that is 1 inch in length (in the direction of the conveyor's motion) and the camera 132 has a resolution of 200 dots per inch in that direction, since each pulse of the encoder 122 occurs every half inch, it can be determined that new image slices are captured every 100 pixels or every half inch. Additionally, if each image slice is 200 pixels or 1 inch in length and new image slices are captured every 100 pixels or every half inch, it can be understood that each new image slice overlaps the previous image slice by 100 pixels or half an inch. It should be noted that in many examples, the terms "dot" and "pixel" can be used interchangeably. It should further be noted that these numbers are for example only. In embodiments, any value of rotational motion per advancement of the conveyor 120, any value of number of pulses per rotation, any value of total camera 132 resolution, and any value of camera 132 resolution per unit distance, as well as any combination of rotational motion per advancement of the conveyor 120, number of pulses per rotation, total camera 132 resolution, and camera 132 resolution per unit distance can be utilized.

[0049] The controller / processor 1200 shown in the embodiment in FIG. 1 can be connected to and communicate with the camera 132. The controller / processor 1200 can receive image data from the camera 132 and perform image stitching to form a combined image. The received images can be stored in a buffer for later processing or processed immediately. In an embodiment, the controller / processor 1200 can receive a stream of images and use the captured images to determine whether an object is present. For example, the controller / processor 1200 can perform background detection on the captured images to determine that an object is present in the viewing area 121. If the presence of an object in the viewing area 121 is detected, the controller / processor 1200 can perform image stitching and / or additional image capture. In other embodiments, the controller / processor 1200 can receive a trigger from sensors 140, 141, such as electro-optical sensors like photoeyes, to start capturing images and intermittently flashing a light source. The sensors can include a photoeye that includes an infrared or visible light emitter 140 and a photocell, photodiode, or phototransistor to detect when the light from the emitter 140 is blocked (indicating the presence of an object) or not blocked (indicating the absence of any object). The controller / processor 1200 can also be connected to the optical encoder 122 to receive object translation data.

[0050] FIG. 2 illustrates a schematic view from below of a system according to an embodiment. In this example, the parcel 110 is partially visible in the viewing area 121. A single image captured by the camera 132 may not be able to capture the entire bottom surface of the parcel 110. By stitching several images together, a composite image of the entire bottom surface of the parcel 110 can be created.

[0051] FIG. 3 illustrates exemplary consecutive images 300 obtained by a two-camera continuous connection system according to an embodiment. Each strip represents a single image captured by a camera. The left column shows image capture from the first camera, and the right column shows images captured by the second camera. The viewing area 121 in this case is a wide rectangle. The captured images can overlap from image to image for each camera. Additionally, there can be an overlap in the target range between the two cameras. In an embodiment, at least an image and the immediately following image can be joined to form an output image. In an embodiment, each image in a series of images can be continuously joined to the previously joined images to form a composite image.

[0052] FIG. 4 is an illustration of an example of a joined image generated by a system according to an embodiment. In this image, the entire surface of the package can be seen. An image that appears as a single capture is presented, although multiple images were joined to create it.

[0053] FIG. 5 illustrates a flowchart of an embodiment of the image stitching operation. The operations in the flowchart can be performed, for example, by the controller / processor 1200. In an embodiment, the process 500 can continuously stitch image slices to form a complete resultant image of the target object. The process starts from the input process 501, and the image slices can be continuously acquired from each camera as described above and typically stored in a memory buffer located in the controller / processor 1200. At 502, the system reads the image slices from the buffer. Then, the system can perform an alignment process 510 to find a two-dimensional coordinate transformation that aligns the current image slice with the resultant image. And the process 522 can be used to warp the current image slice and mix the current image slice into the resultant image by using that two-dimensional coordinate transformation. And the processes 502-522 can be repeated for all the acquired target image slices to create the final resultant image 524.

[0054] Regarding two-dimensional coordinate transformation, a digital image is a two-dimensional array where each pixel has a position (coordinate) and an intensity value (color or grayscale level). The two-dimensional coordinate transformation relates the positions from one image to the other. The image warping and mixing processes apply that two-dimensional coordinate transformation to determine the intensity value at each pixel in the resultant image.

[0055] In process 512, it can be determined whether the image slice read from the buffer is the first slice in a series of acquired slices for one object. If so, process 500 continues to 514, and an identity two-dimensional coordinate transformation can be applied to the image slice. For example, the identity transformation can lead to simply copying the image slice as is to the resulting image in step 522. In other words, for each object, a first image slice can be acquired and copied as is to a resulting image that can be an empty resulting image. When the next image slice is acquired, image warping can be applied to warp the next image slice to the resulting image so as to enable proper stitching.

[0056] In process 512, if it is determined that the image slice read from the buffer is not the first slice in a series of acquired slices, process 500 continues to process 520, and it can be determined that a two-dimensional coordinate transformation should be used to align the two digital images currently being processed. Then, in process 522, the newly read digital image can be warped by using the found two-dimensional coordinate transformation, and the warped digital image can be blended with the other digital image being processed.

[0057] For example, in an embodiment, a two-dimensional coordinate transformation can be found by extracting features in the form of keypoints from a current image slice. And the keypoints can be used to search for matching keypoints in a second image, which can be represented in this process by a result image that includes all previously stitched-together image slices. In process 520, a corresponding set of matched keypoints can be used to fit a two-dimensional coordinate transformation that aligns the current image slice to the result image. For example, in some embodiments, a two-dimensional coordinate transformation can be determined by extracting a first set of features from a first digital image, extracting a second set of features from a second digital image, matching at least one extracted feature from the first set of features and the second set of features, and determining an image transformation based on the matched features.

[0058] The process of feature extraction can use, for example, Histogram of Oriented Gradients (HOG) features to find keypoints having rich texture. The HOG feature map is an array having n-dimensional feature vectors as entries. Each feature vector describes a local image patch called a tile. An image region of interest (ROI), which can be an overlapping region between two slices, is first divided into non-overlapping tiles of a fixed size (e.g., 8 pixels × 8 pixels). And for each tile, a one-dimensional histogram of gradient orientations is calculated over its pixels. The gradient magnitude and orientation are calculated at each pixel, for example, by using a finite difference filter. For a color image, the color channel having the maximum gradient magnitude is typically used. And the gradient orientation at each pixel is quantized into one of "n" orientation bins having a voting strength according to the gradient magnitude, in order to collectively build the histogram of gradient orientations in this tile as a vector of length "n".

[0059] Taking a brief look at FIG. 6, an example of extracting key points based on HOG features is illustrated. Image ROI 602 can be divided into non-overlapping tiles to calculate HOG features. And these tiles can be grouped into larger regions 603 (dotted lines). For each region 604, as illustrated, every 2 tiles × 2 tiles can be grouped into one patch having a candidate key point (the candidate key point is at the center of this patch). For example, 606 and 610 are candidate key points at the centers of 2 tiles × 2 tiles patches 608 and 612 respectively. These patches overlap, and each patch is represented by its center point and a score that is a function of the gradient orientation and magnitude determined in the HOG features of its tiles. For each region, the patch having the maximum score can be determined and selected as a key point if its score exceeds a predetermined threshold; otherwise, this region may not be represented by any key point, which indicates that this region does not have reliable key points that can be used for alignment. This results in the distribution of key points across the image ROI, which can be the output from the feature extraction process. The image ROI can be defined as the overlap between two images to be stitched together. The number and size of the regions can be predetermined based on factors such as the resolution of the image, the total size of the image, the number and size of the tiles, the number and size of the patches, etc.

[0060] Returning to FIG. 5, the process 520 may then use the keypoints extracted from the current image slice to find matching keypoints in the resulting image. Referring to FIG. 7a, each keypoint extracted from the image ROI of the new image slice 704 may be used to search for a corresponding keypoint in the resulting stitched image 702 that includes all of the previously stitched image slices. For example, the region represented by 708 in the new slice 704 had a patch 714 centered on the keypoint 716 as the selected keypoint having the maximum score in this region. The goal of the process is to find the correct matching point 724 that is at the center of the matching patch 726 in the region 706 in the resulting image 702. For example, template matching of the patch 714 within the region 722 in the region 706 may be used to find the matching patch 726.

[0061] In some embodiments, new image slices are acquired at specific predetermined physical distances, which map through the image resolution to a specific known transformation (typically a transformation such as in a two-dimensional coordinate transformation), and the specific known transformation associates all of the image slices with the previous image slice and thus with the resulting image.

[0062] In an embodiment, the object can be translated according to a substantially linear motion. A substantially linear motion is a motion at a substantially constant speed and in a substantially single direction. In embodiments where the object is being translated at a substantially constant speed, new image slices can be acquired at a particular predetermined physical distance by acquiring images at substantially fixed time intervals. If the changes in speed and interval are small enough such that the resulting change in the acquisition distance is small enough for the stitching and other processes described herein to still work well, the speed can be considered to be substantially constant and the time interval can be considered to be substantially fixed. For a point or region to which a line is oriented, if the line along which the object moves is straight enough for the stitching and other processes described herein to still work well, the motion can be considered to be in a substantially single direction. One example of linear motion can be an object that can be moved on a straight (linear) conveyor. However, this is just an example. The present technology is equally applicable to any embodiment where the motion occurs in a substantially single direction. In embodiments where an encoder is provided, new image slices can be acquired at a particular predetermined physical distance based on the distance information provided by the encoder.

[0063] The transformation (typically, the conversion) that maps to a specific known conversion through the image resolution is based on the ideal acquisition settings and may be referred to herein as an approximate two-dimensional coordinate transformation. In an actual setting, this two-dimensional coordinate transformation may map the keypoint 716 in the new image slice 704 to the point 718 in the resulting image 702, which may not be a perfect match. Such a mismatch may be due to, for example, the vibration of the object or a slight delay in camera acquisition. And the relatively small search space 722 may be centered on the point 718, and template matching may be performed only in this search space to find the correct match 724 for the keypoint 716. In embodiments, in order to enable the decoding of marks on the object or to enable further processing or analysis of the image, the mismatch may be small enough or the stitching accuracy may be high enough. For example, the marks on the object that can be decoded may include one-dimensional or two-dimensional barcodes, etc. Similarly, the text on the object can be recognized, or other features of the object or other features on the object can be recognized or analyzed. In addition, the approximate two-dimensional coordinate transformation can be used as is without refinement in the case of an object that does not have features for refinement, such as a plain cardboard box. In this case, since an image without distinct edges does not require perfect alignment, the approximate two-dimensional coordinate transformation can be good enough to be used for alignment.

[0064] It should be noted that the stitching and decoding processes can be performed by using continuous processing, parallel processing, or some combination of the two. For example, in embodiments, the stitching of the images can be performed in parallel. For example, each of a plurality of pairs of images can be stitched in parallel. Similarly, in embodiments, the stitching of the images can be performed in parallel with the decoding of the marks on the object or in parallel with other further processing of the image of the object. Such parallel stitching, decoding, and / or other processing can be performed by using any number of known parallel processing techniques.

[0065] Figure 7b illustrates an important case related to the extraction of key points. Region 712 in image 704 illustrates an example of a bad key point 728. Patch 730 centered on point 728 has a plurality of strong edges, but all of those edges have the same direction. When an approximate two-dimensional coordinate transformation is applied to key point 728, the transformation maps that key point to a point 732 that is very close to a correct match, but the search within search space 736 can result in a plurality of key points (such as 738 and 742) that perfectly match 728. This explains why the selection of key points is based on a score using HOG features to ensure that the patch has not only strong edges but also a plurality of different directions. This means that key point candidates such as 728 are not selected as key points (since the key point with the maximum score does not exceed the selection threshold), and further, regions such as 712 do not have key points extracted therefrom. In embodiments not shown in FIG. 7, features to be matched can be extracted from each region and then matched. For example, the matching method can be one of matching corresponding key points by using the least squares method, or finding the best corresponding position of the key points by using a known alignment method such as normalized correlation.

[0066] Returning to FIG. 5, process 520 may then use the corresponding set of key points calculated to fit a two-dimensional coordinate transformation that aligns the new image slice to the resulting image. For example, a Random Sample Consensus (RANSAC) process may be used to fit an accurate transformation that removes any outliers resulting from the matching problem. In process 522, the image slices to be stitched together may be warped by using the determined two-dimensional coordinate transformation. The two-dimensional coordinate transformation changes the spatial configuration of the image. In this specification, the two-dimensional coordinate transformation may be used to correct spatial mismatches in the images or image slices to be stitched together. For example, the two-dimensional coordinate transformation may be used to align two images so that they are ready to be blended. Preferably, all overlapping pixels can be aligned to exactly the same location in the two images. However, even if this is not possible, the two-dimensional coordinate transformation may provide sufficient alignment for a successful blend.

[0067] If at least a portion of the image slices to be stitched together overlaps the image being stitched, the overlapping portion may be blended to form the resulting image. In some embodiments, the slices are roughly aligned with the translation in the direction of movement. This divides all new slices into overlapping regions and new regions as illustrated in FIG. 7. Blending of the overlapping pixels may use a weighted average that allows for a seamless transition between the resulting image and the new image slice. For example, the upper portion of the overlap region may give a higher weight to the resulting image (which gradually decreases to the lower portion of the overlap region), which may give a higher weight to the new slice so that there is no seam line between the overlap region and the new region of each new slice at the edge of the overlap region. In embodiments, other types of blending may be utilized.

[0068] After completing process 520 for the current slice, at 524, process 500 branches back to 502 to obtain another slice to be processed. If there are no more slices to be processed, at 524, the stitched-together image can be output. It should be noted that in embodiments, the output image may contain only two stitched-together slices, or may contain all the slices in a series of image captures that are stitched together. In embodiments where multiple cameras are used, such as the example shown in FIG. 10, it should be noted that the system creates overlapping image slices that overlap between the two cameras 133a and 133b and also in the translation of the package along the conveyor 120. In such embodiments, the images captured by two or more cameras can likewise be stitched together. For example, the image slices from each camera can be stitched together with other image slices from that camera, and the resulting stitched-together images from each of the multiple cameras can be stitched together to form the final complete image. In some embodiments, each image slice from a camera can be stitched together with one or more image slices captured substantially simultaneously by one or more other cameras to form a combined image slice, and the combined image slices can be stitched together to form a complete image. In some embodiments, each image slice obtained from each camera can be directly stitched together with the same resulting image of the previously stitched-together images from all the cameras. The stitching process that can be used to stitch together images or image slices from different cameras is similar to the process described above for stitching together image slices from the same camera.

[0069] FIG. 8 illustrates an exemplary calibration plate 800 that can be used to automate calculations for different parameters of the system. The plate can include a checkerboard pattern including a plurality of alternating white and black squares or checkers 802. These alternating squares can have a fixed dimension such as 10 mm × 10 mm. These fixed dimensions can provide the ability to determine the image resolution at a particular fixed operating distance. Additionally, the calibration plate or pattern 800 can include a data matrix reference or pattern 804. These data matrix references or patterns 804 can encode physical details about the plate (such as the exact physical dimensions of the checkers and the coordinates of each point with respect to a fixed physical coordinate system defined on the plate). Additionally, it can also be used to determine the presence or absence of a mirror in the optical path between the camera and the object.

[0070] The setup process can be used to initialize the acquisition setup illustrated in FIG. 1. This process can be used to set up and calibrate the hardware and software used to perform the stitching process 500. The setup process begins by acquiring at least one image similar to the image illustrated in FIG. 9 for a stationary calibration plate placed over the viewing window. Region 121 in FIG. 9 corresponds to region 121 representing the viewing area in FIG. 1. The corners of the checkers are extracted from the image and the data matrix code is decoded, which provides an exact set of pairs of points in the image domain and the corresponding locations in the physical domain. This set can be used to automatically perform the following: - Finding the effective field of view of the camera that defines the portion of the acquired image 902 corresponding to the viewing area. The image sensor can be set to acquire only that portion, which in many sensors is an important factor in the acquisition speed. -Determining whether the view in the effective field of view is a perspective view or a non-perspective view. In many embodiments, perspective can be considered to be in the form of distortion of the captured image. Thus, for the purposes of configuration, it may be advantageous to confirm that the image does not contain such perspective distortion by physically correcting the configuration or by warping the image after acquisition to correct the perspective effect. -Calculating the resolution of the image in dpi at this working distance. And after the configuration process, the system can be locked in place, and the orientation of all components, including the orientation of the camera with respect to the viewing area that defines perspective, can be fixed.

[0071] In the stitching algorithm 500, particularly in the matching process which is part of process 520, an approximate two-dimensional coordinate transformation is referenced, which can facilitate the matching process by enabling a small search space for the matching of each keypoint. In some embodiments, this approximate two-dimensional coordinate transformation can be a simple transformation in the direction of movement that can be calculated by multiplying the physical distance between successive acquisitions and the image resolution calculated in the configuration process. Such a transformation can be determined based on a translation which is a substantially linear motion. As explained above, the motion can be considered to be a substantially linear motion when the motion is at a substantially constant speed and in a substantially single direction.

[0072] For example, if an encoder is used to cause the acquisition of image slices and is set to produce 24 pulses per resolution (one resolution being 12 inches), this results in the acquisition of a new image slice every half inch the object moves. If the image resolution is calculated as 200 dpi in the setup process, the approximate two-dimensional coordinate transformation is a simple 100-pixel translation in the direction of movement. In embodiments, any value of rotational movement per advancement of the conveyor 120, any value of pulses per resolution, and any value of resolution, as well as any combination of rotational movement per advancement of the conveyor 120, pulses per resolution, and resolution, may be utilized.

[0073] In some embodiments, the relationship of the individual image slices (for estimating the approximate two-dimensional coordinate transformation) can be established by using a training process that uses the image slices acquired for a calibration moving object having features of a known pattern such as the calibration plate or pattern 800 shown in FIG. 8. In this training process, the approximate two-dimensional coordinate transformation that associates consecutive image slices can be determined by analyzing the data matrix fiducials 804 and the checker 802 acquired across the consecutive image slices.

[0074] FIG. 10 illustrates a top - view schematic of a two - camera system 133 according to an exemplary embodiment. Two cameras are shown in this example, but the embodiments can include more than two cameras, and the techniques described in relation to two cameras can be adapted to three or more cameras. The system in this figure can include a conveyor 120, two cameras 133a, 133b, a light source 131, and a controller / processor 1200. The conveyor 120 translates an object 110, such as a packet, across the conveyor 120, and the conveyor 120 can include a viewing area 121, which can include an opening, a window, a gap, etc. The two cameras 133a, 133b can capture images through the viewing area 121 via mirrors 135. For example, if the width (perpendicular to the direction of travel of the conveyor 120) is too wide to be captured by one camera, or too wide to be captured with adequate image quality by one camera, two or more cameras can be used. The embodiments can be applicable to any number of cameras. Examples of such embodiments are described in more detail below. In an embodiment, the cameras 133a, 133b can be configured to capture images substantially simultaneously (e.g., within the timing accuracy of the electronic circuits for operating the cameras 133a, 133b simultaneously).

[0075] In this example, the system overlaps between the two cameras 133a and 133b and creates overlapping image slices in the translation of the packet along the conveyor 120. In such embodiments, the images captured by two or more cameras can similarly be stitched together. In an embodiment, the use of the mirrors 135 allows the cameras 133a, 133b to be positioned such that they do not need to face the viewing area 121 directly. The controller / processor 1200 can be arranged to cause the cameras 133a, 133b to capture images simultaneously or at different times.

[0076] An example of an embodiment in which the side surface of an object can be imaged is shown in FIG. 11. In this example, the object 110 can be translated on the conveyor 120 so as to image the side surface of the object. In this example, the side surface of the object is parallel to a known plane 1104 of the side surface. The known plane 1104 of the side surface can be implemented, for example, by using side walls attached to the conveyor 120. Such side walls can be made of a transparent material such as glass or transparent plastic as shown, or the side walls can be made of an opaque material such as metal. In an embodiment, such side walls can be tall enough to include a viewing area 121 such as a window, gap, or opening in the known plane 1104 of the side surface. In an embodiment, the camera 132 can be positioned or directed to image the object in or through the viewing area 121. In an embodiment, such side walls can be low enough such that the side surface of the object can be imaged directly without the use of a window, gap, or opening. In an embodiment, the alignment device 1102 can be used to ensure that the side surface of the object 110 is parallel to or aligned with the known plane 1104 of the side surface. In an embodiment, the alignment device 1102 can be implemented using, for example, a mechanical device such as a spring-loaded flap as shown in this example, and the mechanical device applies pressure to the object 110 to align the side surface of the object 110 parallel to the known plane 1104 of the side surface. In an embodiment, the alignment device 1102 can be implemented using, for example, an electromechanical device, and the electromechanical device applies pressure to the object 110 to align the side surface of the object 110 parallel to the known plane 1104 of the side surface.

[0077] FIG. 12 illustrates a schematic diagram of the components of a controller / processor 1200 according to an embodiment. The controller / processor 1200 includes an input / output interface 1204 for receiving images from a camera. The input / output interface 1204 can also be connected to an encoder 122 or a light source 131 to control the encoder 122 or the light source 131. Input / output devices (including but not limited to keyboards, displays, and pointing devices) can be directly connected to the system or connected to the system through an intervening input / output controller.

[0078] The controller / processor 1200 may have one or more CPUs 1202A. In an embodiment, the controller / processor 1200 has network capabilities provided by a network adapter 1206 connected to a communication network 1210. The network adapter can also be connected to other data processing systems or storage devices through an intervening private or public network. The network adapter 1206 enables software and data to be transmitted between the controller / processor 1200 and external devices. Examples of network adapters can include modems, network interfaces (such as Ethernet (registered trademark) cards), communication ports, or PCMCIA slots and cards. The software and data transferred through the network adapter can be in the form of signals, such as electrical signals, electromagnetic signals, optical signals, or other signals that can be received by the network adapter. These signals are provided to the network adapter through the network. This network can be implemented by using wires or cables, optical fibers, telephone lines, cellular phone links, RF links, and / or other communication channels to carry the signals.

[0079] The controller / processor 1200 may include one or more computer memory devices 1208 or one or more storage devices. The memory device 1208 can include a camera data capture routine 1212 and an image stitching routine 1214. An image data buffer 1216, such as an operating system 1218, is also included in the memory device.

[0080] The routines provided for the invention take the form of a computer program product accessible from a computer-usable medium or a computer-readable medium (providing program code for use by or in connection with a computer or any instruction execution system). For the purposes of this description, a computer-usable medium or a computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0081] The medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. Examples of computer-readable media include semiconductor or solid state memory, magnetic tape, removable computer diskettes, random access memory (RAM), read-only memory (ROM), rigid magnetic disks, or optical disks.

[0082] A data processing system suitable for storing and / or executing program code includes at least one processor directly or indirectly coupled to a memory device through a system bus. The memory device can include program code for temporarily storing at least some of the program code to reduce the number of times the code has to be read from mass storage during execution, mass storage, and local memory used during actual execution of the cache memory.

[0083] Various software embodiments are described in terms of this exemplary computer system. After reading this description, it will be apparent to those skilled in the relevant art how to implement the invention using other computer systems and / or computer architectures.

[0084] Controller / processor 1200 can include a display interface (not shown) that transfers graphics, text, and other data to a display on the display portion (from, for example, a frame buffer (not shown)).

[0085] In this disclosure, the terms "computer program medium," "computer usable medium," and "computer readable medium" are generally used to refer to media such as main memory and storage memory, removable storage devices, and hard disks installed in hard disk drives.

[0086] A computer program (also referred to as computer control logic) is stored in main memory and / or storage memory. When such a computer program is executed, it enables the computer system to implement the features of the invention as discussed herein. In particular, when the computer program is executed, it enables one or more processors to perform the operations described above.

[0087] From the above description, it can be understood that the present invention provides a system, a computer program product, and a method for efficient execution of image stitching. In the claims, a reference to an element in the singular is not intended to mean "only" unless explicitly stated otherwise, but rather is intended to mean "one or more." All structural and functional equivalents of the elements of the exemplary embodiments described above that are now known or later become known to those of ordinary skill in the art are intended to be encompassed by the claims of this patent. No claim element herein shall be construed under the provisions of 35 U.S.C. § 112, paragraph 6, unless the element is expressly recited using the phrase "means for" or "step for."

[0088] The foregoing description of the invention enables those of ordinary skill in the art to make and use what is currently considered to be the best mode, while those of ordinary skill in the art will recognize the existence of alternatives, modifications, variations, combinations, and equivalents to the specific embodiments, methods, and examples herein. Those of ordinary skill in the art will recognize that the disclosure herein is merely exemplary and that various modifications may be made within the scope of the invention. Additionally, while a particular feature of the teachings may be disclosed with respect to only one of several implementations, such feature may be combined with one or more other features of the other implementations as desired and advantageous for any given or particular function. Further, as long as the terms "comprising," "comprises," "having," "has," "include," or variations thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "including."

[0089] Other embodiments of the teachings will be apparent to those of ordinary skill in the art from a consideration of the specification and practice of the teachings disclosed herein. Accordingly, the invention should not be limited by the described embodiments, methods, and examples, but should be limited only by all embodiments and methods within the scope and spirit of the invention. Accordingly, the present invention is not limited to the specific embodiments as illustrated herein, but is limited only by the following claims.

Claims

1. A system, the system comprising: A transport device for moving at least one object, at least one substantially planar surface of the object being moved within a known plane, the known plane being a planar surface of the transport device including a viewing area, and the substantially planar surface of the object being blocked except when the at least one substantially planar surface passes through the viewing area; At least one two-dimensional digital optical sensor configured to capture at least two consecutive two-dimensional digital images of the at least one substantially planar surface of the at least one object being moved within the known plane; A controller operably coupled to the two-dimensional digital optical sensor; and the controller is configured to: a) receive a first digital image; b) receive a second digital image; c) generate a stitched image by stitching the first digital image and the second digital image using a stitching algorithm; and the stitching comprises: extracting a first set of features from the first digital image; extracting a second set of features from the second digital image; matching at least one feature from the first set of features with at least one feature from the second set of features; Determining a two-dimensional coordinate transformation to be used to align the first digital image and the second digital image, wherein the determined two-dimensional coordinate transformation is based on matching at least one feature from the first set of features and at least one feature from the second set of features, and Warping the second digital image using the determined two-dimensional coordinate transformation, and Mixing the warped second digital image with the first digital image is performed by An approximate two-dimensional coordinate transformation is estimated through a training process, and the training process uses a plurality of digital image slices captured for a calibration moving object having features of a known pattern, a system. **Claim 2** The controller repeatedly executes steps a) to c), the system according to claim 1. **Claim 3** The first digital image comprises at least one of a captured digital image or an image as a result of a previously stitched image, the system according to claim 1. **Claim 4** The surface moved through the known plane is one of the bottom surface or the side surface of the object, The transport device comprises a viewing area disposed on a corresponding surface of the transport device, The two-dimensional digital image is captured through the viewing area, the system according to claim 1. **Claim 5** The transport device comprises a conveyor, and the viewing area comprises one of an optical window in the conveyor, a gap on the surface of the conveyor, or a gap between two conveyor devices, the system according to claim 1. **Claim 6** The system further comprises a light source configured to illuminate the one or more objects in the viewing area, the controller being coupled to the light source to control the illumination from the light source, the controller intermittently flashing the light source, the system according to claim 1.

7. The system further comprises a sensor configured to detect the presence or absence of an object on the transport device and to control the acquisition of an image based on the presence or absence of an object on the transport device, the system according to claim 6.

8. The system further comprises a mirror, the two-dimensional digital optical sensor capturing the digital image via the mirror, the system according to claim 1.

9. The controller decodes a mark based on the stitched image, the system according to claim 1.

10. The transport device operates according to a substantially linear motion, the controller comprising a model of the substantially linear motion of the transport device, the stitching algorithm using the model of the substantially linear motion of the transport device, the system according to claim 1.

11. The system further comprises a motion encoder, the system controlling the acquisition of an image based on the motion encoder, determining image transformation being further based on an estimation of object translation generated by the motion encoder, the system according to claim 1.

12. The system further comprises a light source configured to illuminate the one or more objects in the viewing area, the controller being generated by the motion encoder coupled to the light source to control illumination from the light source based on an estimate of object translation, the controller intermittently flashing the light source based on an estimate of object translation generated by the motion encoder, the system of claim 11.

13. The two-dimensional digital optical sensor is configured to capture a reduced field of view, the reduced field of view being determined manually or by analyzing an entire image acquired at a set time, the system of claim 1.

14. The system further comprises a plurality of two-dimensional digital optical sensors, each two-dimensional digital optical sensor being configured to capture a plurality of sequential digital images associated with the one or more objects through the viewing area, the plurality of two-dimensional digital optical sensors being configured to capture digital images substantially simultaneously, the system of claim 1.

15. Each digital image from the two-dimensional digital optical sensor is joined together with an image slice captured substantially simultaneously by each other two-dimensional digital optical sensor to form a combined image slice, and then the combined image slices are joined together to form a complete image, or Each digital image from each two-dimensional digital optical sensor is joined together with the digital images of the sequentially captured digital images from that two-dimensional digital optical sensor, and then the joined images from each two-dimensional digital optical sensor of the plurality of two-dimensional digital optical sensors are joined together to form a complete image, or Each digital image from each two-dimensional digital optical sensor is joined directly to an image as the same result of the previously joined images from all two-dimensional digital optical sensors, the system of claim 14.

16. A computer-implemented method, wherein the computer-implemented method comprises: A transport device moving at least one object, wherein at least one substantially planar surface of the object is moved within a known plane, the known plane being a planar surface of the transport device including a viewing area, and the substantially planar surface of the object being blocked except when the at least one substantially planar surface of the object passes through the viewing area; At least one two-dimensional digital optical sensor capturing at least two consecutive two-dimensional digital images of the surface of the object being translated within the known plane; Generating a stitched image by stitching a first digital image and a second digital image using a stitching algorithm; comprising; said stitching comprising: extracting a first set of features from the first digital image; extracting a second set of features from the second digital image; matching at least one feature from the first set of features with at least one feature from the second set of features; determining a two-dimensional coordinate transformation to be used to align the first digital image and the second digital image, the determined two-dimensional coordinate transformation being based on matching at least one feature from the first set of features with at least one feature from the second set of features; warping the second digital image using the determined two-dimensional coordinate transformation; mixing the warped second digital image with the first digital image; and performed by; The approximate two-dimensional coordinate transformation is estimated through a training process, and the training process uses a plurality of digital image slices captured for a calibration moving object having characteristics of a known pattern, a computer-implemented method.

17. The controller decodes a mark based on the combined image, and the controller is operably coupled to the two-dimensional digital optical sensor, the method according to claim 16.

18. There are a plurality of two-dimensional digital optical sensors for capturing images, Each digital image from a two-dimensional digital optical sensor is combined with an image slice substantially simultaneously captured by each other two-dimensional digital optical sensor to form a combined image slice, and then the combined image slices are combined together to form a complete image, or, Each digital image from each two-dimensional digital optical sensor is combined with a digital image of the continuously captured digital images from that two-dimensional digital optical sensor, and then the combined images from each two-dimensional digital optical sensor among the plurality of two-dimensional digital optical sensors are combined together to form a complete image, or, Each digital image from each two-dimensional digital optical sensor is directly combined with an image as the same result of the previously combined images from all two-dimensional digital optical sensors, the method according to claim 16.

19. The combining is performed using parallel processing, or, The controller decodes a mark based on the combined image, and the combining and the decoding are performed using parallel processing, and the controller is operably coupled to the two-dimensional digital optical sensor, the method according to claim 16.

Citation Information

Patent Citations

  • Generating method of still image data by imaging multiple articles on transport line one after another, imaging apparatus, solid-state imaging-element camera and manufacturing method of the camera

    JP2002232755A

  • Imaging apparatus, method for processing image, and program

    JP2007208704A

  • Tunnel or portal scanner and method of scanning for automated checkout

    US20130020392A1

  • System and method for reading optical codes on bottom surface of items

    US20130292470A1

  • Item image stitching from multiple line-scan images for barcode scanning systems

    US20180005392A1