System and method for three-dimensional scanning of moving objects longer than the field of view

By combining a vision system with an encoder to track object features and acquire a series of image snapshots to generate aggregated feature data, the problem of image stitching difficulties when region scanning sensors capture long objects is solved, and efficient and accurate size determination is achieved.

CN115362473BActive Publication Date: 2025-12-12COGNEX CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180027055.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-18
Filing Date
2021-02-18
Publication Date
2025-12-12
Estimated Expiration
2041-02-18

AI Technical Summary

Technical Problem

Existing area scanning sensors require multiple snapshots and image stitching to capture objects longer than the field of view, resulting in computationally intensive and inaccurate results.

Method used

By using a vision system combined with an encoder, object features are tracked and a series of image snapshots are acquired as the object passes through the field of view. Aggregated feature data is generated to determine the overall size, avoiding the image stitching process.

Benefits of technology

It enables lightweight processing of long objects, improves the accuracy and efficiency of size determination, reduces computational burden, and can handle complex shapes and multiple objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115362473B_ABST
    Figure CN115362473B_ABST
Patent Text Reader

Abstract

A system and method for using area scan sensors of vision systems, in conjunction with encoders or other motion knowledge, to capture accurate measurements of objects larger than the individual field of view (FOV) of the sensor. It identifies features / edges of the object, tracks from image to image, providing a lightweight method to process the overall extent of the object for dimensioning. Logic automatically determines if the object is longer than the FOV, resulting in a series of image acquisition snapshots occurring while the moving / conveying object remains within the FOV and until the object is no longer within the FOV. At this point, acquisition stops and the individual images are combined into a segment in the overall image. These images can be processed to derive the overall dimensions of the object according to the application details entered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to machine vision systems that analyze objects in three-dimensional (3D) space, and more particularly, to systems and methods for analyzing objects transported through an inspection region on a conveyor. BACKGROUND

[0002] Machine vision systems (also referred to herein as “vision systems”) that perform measurements, inspections, alignments, and / or symbol (e.g., barcode— also referred to as “ID code”) decoding of objects are widely used in a variety of applications and industries. These systems are based on the use of image sensors that acquire images (typically grayscale or color, and one-, two-, or three-dimensional) of objects or objects, and use on-board or interconnected vision system processors to process these acquired images. The processors typically include processing hardware and non-transitory computer readable program instructions that perform one or more vision system processes to generate a desired output from the processed information of the images. This image information is typically provided within an image pixel array, with each image pixel having various color and / or intensity.

[0003] As noted above, one or more vision system cameras can be arranged to acquire two-dimensional (2D) or three-dimensional (3D) images of objects in an imaged scene. A 2D image is typically characterized as having x and y components of pixels within an overall N X M image array (typically defined by a pixel array of a camera image sensor). In the case of acquiring images in 3D, there is an additional height or z-axis component in addition to the x and y components. 3D image data can be acquired using a variety of mechanisms / techniques, including triangulation of stereo cameras, laser radar, time-of-flight sensors, and (for example) laser displacement analysis.

[0004] Typically, a 3D camera is arranged to capture 3D image information of objects falling within its field of view (FOV), which constitutes a volume of space that fans out in lateral x and y dimensions as a function of distance from the camera sensor in the orthogonal z dimension. A sensor that acquires an image of the entire volume of space simultaneously / synchronously (i.e., in a “snapshot”) is referred to as a “area scan sensor”. Such area scan sensors are distinguished from line scan sensors (e.g., profilometers) that capture 3D information in slices, and use motion (e.g., of a conveyor) and a measurement of that motion (e.g., through a motion encoder or stepper) to move objects through the inspection region / FOV.

[0005] Line scan sensors have the advantage that the object being inspected can be arbitrarily long - the object length is along the direction of motion of the conveyor. In contrast, area scan sensors take image snapshots of a volume space, do not require an encoder to capture the 3D scene, but cannot image the entire object in a single snapshot if the object is longer than the field of view. If only a portion of the object is acquired in a single snapshot, then another snapshot (or multiple snapshots) of the remaining length must be acquired as the trailing portion of the object (not yet imaged) enters the FOV. For multiple snapshots, the challenge is how to register (stitch together) the multiple 3D images in an efficient manner so that the overall 3D image accurately represents the features of the object. SUMMARY

[0006] The present invention overcomes the disadvantages of the prior art by providing a system and method for using an area scan sensor of a vision system, in combination with an encoder or other motion knowledge, to capture accurate measurements of objects that are larger than the single field of view (FOV) of the sensor. The system and method specifically address the disadvantage of snapshot area scan vision systems defining a limited field of view, which often requires the system to acquire multiple snapshots and other data needed to combine the snapshots. This avoids the task of combining raw image data and then post-processing such data, which can be computationally intensive. Instead, the example embodiments identify features / edges of the object (also referred to as "vertices" in relation to the polygonal shape identified) that are tracked between images, providing a lightweight way to process the overall extent of the object for dimensioning purposes. The system and method can employ logic that automatically determines whether the object is longer (in the direction of motion) than the FOV, resulting in a series of image acquisition snapshots occurring while the moving / conveyed object remains within the FOV until the object no longer appears in the FOV. At that point, acquisition stops, and optionally the individual images are combined into a segment in an overall image. The overall image data can be used for various downstream processing. Aggregated feature data from the discrete image snapshots, whether or not an actual overall image is generated, can be processed to obtain overall dimensions of the object according to input application specifics. Such aggregated feature data can be used to determine other attributes and features of the object(s), including but not limited to skew, length, width, and / or height out-of-tolerance, confidence score, liquid volume, classification, actual imaging data quantity (QTY) of the object versus expected imaging data quantity of the object, object location features, and / or damage detection relative to the object. The system and method can also accommodate complex / multi-level objects that would typically be separated by conventional imaging of area scan sensors due to loss of 3D data from shadows and the like.

[0007] In illustrative embodiments, a vision system and method of using the same are provided. The vision system and method can include a 3D camera assembly arranged as a region scanning sensor, and a vision system processor receiving 3D data from images of an object acquired within a field of view (FOV) of the 3D camera assembly. The object is conveyed through the FOV in a conveyance direction, and the object can define an overall length in the conveyance direction that is longer than the FOV. A dimensioning processor measures the overall length in conjunction with a plurality of 3D images of the object based on motion tracking information derived from the conveyance of the object through the FOV. The images can be acquired by the 3D camera assembly in a sequence having a predetermined amount of conveyance motion between 3D images. A presence detector associated with the FOV can provide a presence signal when the object is adjacent thereto. Responsive to the presence signal, the dimensioning processor can be arranged to determine whether the object appears in more than one image as the object moves in the conveyance direction. Responsive to information related to a feature on the object, the dimensioning processor can be arranged to determine whether the object is longer than the FOV as the object moves in the conveyance direction. An image processor can generate aggregate feature data in conjunction with information related to a feature on the object from successive image acquisitions by the 3D camera, thereby determining an overall dimension of the object without combining discrete individual images into an overall image. Illustratively, the image processor can be arranged to acquire a series of image acquisition snapshots while the object remains within the FOV and until the object exits the FOV in response to the overall length of the object being greater than the FOV. The image processor can be further arranged to derive overall attributes of the object using the aggregate feature data and based on input application data, and wherein the overall attributes include at least one of a confidence score, an object classification, an object dimension, a skew, and an object volume. Based on the overall attributes of the object, a process of the object can be performed that includes at least one of redirecting the object, rejecting the object, sounding an alarm, and correcting a skew in the object. Illustratively, the object can be conveyed by a mechanical conveyor or manually, and / or the tracking information can be generated by an encoder operably connected to the conveyor. A motion sensing device can be operably connected to the conveyor, an external feature sensing device, and / or a feature-based sensing device. The dimensioning processor can use the presence signal to determine a continuity of the object between each of the images as the object moves in the conveyance direction. A plurality of images can be acquired by the 3D camera with a predetermined overlap between the plurality of images, and a removal process removes an overlapping portion from the object dimension using the tracking information to determine the overall length. An image rejection process can reject a last one of the plurality of images acquired after the presence signal is asserted due to a previous one of the plurality of images containing a trailing edge of the object.The dimensioning processor can be further arranged to use information related to features on the object to determine continuity of the object between each of the images as the object moves along the conveyance direction. The dimensioning system can further define a minimum spacing between objects in an image, multiple objects below the minimum spacing being considered a single object lacking 3D image data. Illustratively, the image processor can be arranged to generate aggregate feature data about the object involving (a) data beyond a length limit, (b) data beyond a width limit, (c) data beyond a height limit, (d) data beyond a volume limit, (e) a confidence score, (f) a liquid volume, (g) a classification, (h) a quantity (QTY) of actual imaging data of the object versus an expected quantity of imaging data of the object, (i) a positional characteristic of the object, and / or (j) damage detection related to the object.

[0008] In illustrative embodiments, a vision system and associated method can include a 3D camera assembly arranged as a region scanning sensor having a field of view (FOV), the region scanning sensor can operate a capture process that captures one or more images of an object as the object passes through the FOV and determines (a) whether the object will occupy more than a single image, (b) determines when the object will no longer occupy a next image, and (c) calculates a size and relative angle of the object from the one or more images taken.

[0009] In an illustrative embodiment, a method for dimensioning an object can be provided using a vision system having a 3D camera assembly arranged as a region scanning sensor and a vision system processor that receives 3D data from images of an object acquired within a field of view (FOV) of the 3D camera assembly. The object can be conveyed in a conveyance direction through the FOV, and the object can define an overall length in the conveyance direction that is longer than the FOV. The method can further include the steps of measuring the overall length in conjunction with a plurality of 3D images of the object based on motion tracking information obtained from the object conveying through the FOV, the plurality of 3D images being obtained by the 3D camera assembly in a sequence having a predetermined amount of conveyance motion between 3D images. A presence signal is generated when the object is adjacent to the FOV, and in response to the presence signal, it can be determined whether the object appears in more than one image as the object moves in the conveyance direction. In response to information related to a feature on the object, it can be determined whether the object is longer than the FOV as the object moves in the conveyance direction. The information related to the feature on the object can be combined with successive image acquisition from the 3D camera to generate aggregate feature data to determine an overall dimension of the object without combining discrete individual images into an overall image. Illustratively, in response to the overall length of the object being greater than the FOV, a series of image acquisition snapshots can be acquired while the object remains within the FOV and until the object exits the FOV. The overall properties of the object can be derived using the aggregate feature data and based on inputted application data, and the overall properties can include at least one of a confidence score, an object classification, an object dimension, a skew, and an object volume. Illustratively, the method can perform a task with respect to the object based on the overall properties, including at least one of redirecting the object, rejecting the object, sounding an alarm, and correcting a skew in the object. Illustratively, the presence signal can be used to determine continuity of the object between each image as the object moves in the conveyance direction. In combining the information, the method can further generate (a) data that exceeds a length limit, (b) data that exceeds a width limit, (c) data that exceeds a height limit, (d) data that exceeds a volume limit, (e) a confidence score, (f) a liquid volume, (g) a classification, (h) a quantity (QTY) of actual imaging data of the object versus an expected quantity of imaging data of the object, (i) a location feature of the object, and / or (j) damage detection related to the object. A minimum spacing between objects in an image can be defined, and below the minimum spacing, multiple imaged objects are treated as a single object with missing 3D image data. BRIEF DESCRIPTION OF DRAWINGS

[0010] The application is herein described, by way of example only, with reference to the accompanying drawings, wherein:

[0011] Figure 1is an overview of a system for acquiring and processing 3D images of objects longer than the field of view using an area scan sensor, employing a sizing process (processor) to generate an overall image from multiple snapshots as the object moves through the field of view;

[0012] Figure 2 is Figure 1 a side view of the conveyor and area scan sensor arrangement, detailing the available region of interest for 3D imaging of an object trigger plane for image acquisition according to an exemplary FOV;

[0013] Figure 3 is a diagram showing the runtime operation of the arrangement of Figure 1 and Figure 2 an exemplary object having a length in the direction of conveyance greater than the available region of interest (ROI), reaching the trigger plane to trigger an initial snapshot;

[0014] Figure 4 is a diagram showing the runtime operation of the arrangement of Figure 3 an exemplary object partially extending out of the ROI and an encoder indicating the degree of motion sufficient to trigger a second snapshot;

[0015] Figure 5 is a diagram showing the runtime operation of the arrangement of Figure 1 and Figure 2 an exemplary object representing a complex top surface that can result in loss of 3D image data and / or indicate two discrete objects in the FOV;

[0016] Figure 6 is a flowchart showing the process for operating the sizing process according to an exemplary embodiment, including imaging multiple objects and / or complex objects;

[0017] Figure 7 is a diagram showing a top view of the imaging conveyor surface, with exemplary overlapping image segments representing an overall imaged object; and

[0018] Figure 8 is a flowchart of a general image acquisition process for the arrangement of Figure 1 and Figure 2 as applied to processing overlapping image segments as shown in Figure 7 to generate an overall length of the object, as well as its maximum width and height and relative angle with respect to the direction of conveyance (or another coordinate system). DETAILED DESCRIPTION

[0019] I. System Overview

[0020] Figure 1An overview of the arrangement 100 is shown, in which a vision system camera assembly (also simply referred to as "camera" or "sensor") 110 acquires 3D image data of an exemplary object 120 as it passes under its field of view (FOV) relative to a moving conveyor 130 moving in a downstream direction (arrow 132). As an example, a local / relative 3D coordinate system 138 is shown with x, y, and z orthogonal coordinates. Other coordinate systems, such as polar coordinates, can be used to represent 3D space. Note that in this example, the object 120 defines a length LO from upstream to downstream (generally along the y-axis / motion direction) that is longer than the upstream and downstream boundaries BF and BR of the FOV, which will be described further below.

[0021] The 3D camera / imaging assembly 110 contemplated herein can be any assembly that acquires 3D images of objects, including but not limited to stereo cameras, time-of-flight cameras, lidar, ultrasonic range-finding cameras, structured illumination systems, and laser displacement sensors (profilers), and thus the term 3D camera should be understood broadly to include these systems and any other system that generates height information associated with 2D images of objects. Further, a single camera or an array of multiple cameras can be provided, and the terms "camera" and / or "camera assembly" can refer to one or more cameras that acquire images in a manner that generates desired 3D image data of a scene. The depicted camera assembly 110 is shown mounted above the surface of the conveyor 130 in the manner of a checkpoint or inspection station that images the flowing object as it passes by. In this embodiment, the camera assembly 110 defines an optical axis OA that is approximately perpendicular (along the z-axis) relative to the surface of the conveyor 130. Other non-perpendicular orientations of the axis OA relative to the conveyor surface are expressly contemplated. The object 120 can remain in motion (generally) or be temporarily stopped for imaging, depending on the running speed of the conveyor and the acquisition time (partly dependent on frame rate and aperture setting) of the camera image sensor and associated electronics 110. The camera 110 acquires 3D images of the object 120 that are sufficiently within the FOV, which can be triggered by a photodetector or other trigger mechanism 136 and result in a trigger signal to the camera 110 and associated processor. Thus, the camera assembly 110 in this embodiment is arranged as a area scan sensor.

[0022] The camera assembly 110 includes an image sensor S adapted to generate 3D image data 134. The camera assembly also includes (optional) integral illumination assembly I, e.g., a ring illuminator of LEDs, which projects light in a predictable direction relative to the axis OA. External illumination (not shown) can be provided in alternative arrangements. An appropriate optical package O is shown in optical communication with the sensor S along the axis OA. The sensor S is in communication with an internal and / or external vision system process (processor) 140 that receives image data 134 from the camera 110 and performs various vision system tasks on the data in accordance with the systems and methods herein. The process (processor) 140 includes underlying processing / processors or functional modules, including a set of vision system tools 142, which can include various standard and custom tools that identify and analyze features in image data, including but not limited to, edge detectors, blob tools, pattern recognition tools, deep learning networks, ID (e.g., barcode) finders and decoders, etc. In accordance with the systems and methods, the vision system process (processor) 140 can also include a dimension determination process (processor) 144. The process (processor) 144 performs various analysis and measurement tasks on features identified in 3D image data in order to determine the presence of particular features for which further results can be computed. In accordance with exemplary embodiments, the process (processor) interfaces with various conventional and custom (e.g., 3D) vision system tools 142.

[0023] System setup and results display can be handled by a separate computing device 150, e.g., a server (e.g., cloud-based or local), a personal computer (PC), a laptop, a tablet, and / or a smartphone. The computing device 150 is depicted (as a non-limiting example) as having a conventional display or touchscreen 152, a keyboard 154, and a mouse 156 that collectively provide graphical user interface (GUI) functionality. Various interface devices and / or form factors can be provided in alternative implementations of the device 150. The GUI can be driven in part by a web browser application that resides on the device operating system and displays web pages from the process (processor) 140 with control and data information in accordance with exemplary arrangements herein.

[0024] Note that processing (processor) 140 can reside entirely or partially on the housing of camera assembly 110, and various process modules / tools 142 and 144 can be instantiated entirely or partially on board processing (processor) 140 or remote computing device 150 as appropriate. In exemplary embodiments, all vision system and interface functionality can be instantiated on board processing (processor 140), and computing device 150 can be used primarily for training, monitoring, and operations related to interface web pages (e.g. HTML) generated by board processing (processor) 140 and transmitted to the computing device over wired or wireless network links. Alternatively, all or part of processing (processor) 140 can reside in computing device 150. Analysis results from the processor can be transmitted to downstream utilization devices or processes 160. Such devices / processes can use the results 162 to process objects / packages— e.g. strobe the conveyor 130 to direct objects to different destinations based on analyzed features and / or reject defective objects.

[0025] The camera assembly includes on-board calibration data established by factory and / or in-field calibration procedures, and maps x, y, and z coordinates of imaged pixels to the coordinate space of the camera. This calibration data 170 is provided to the processor for use in analyzing image data. In addition, in exemplary embodiments, the conveyor and / or its drive mechanism (e.g. stepper motor) includes an encoder or other motion tracking mechanism 180 that reports relative motion data 182 to the processor 140. Motion data can be delivered in a variety of ways— e.g. distance-based pulses, each of which defines a predetermined increment of conveyor motion. By summing the pulses, the total motion over a given time period can be determined.

[0026] Note that the term "conveyor" as used herein should be understood broadly to include arrangements in which objects are passed through the FOV by another technique— e.g. manual motion. Thus, an "encoder" as defined herein can be any acceptable motion measurement / tracking device, including steppers, marker readers, and / or those devices that track features or fiducials on the conveyor or objects as they pass through the FOV, including various internal techniques that use knowledge applied by the underlying vision system (e.g. features in images) to determine the extent of motion between image snapshots.

[0027] II. Size Determination Processor

[0028] A. Setup

[0029] Figure 2The setup of the arrangement 200 is shown, which is adapted to image objects that are longer in the direction of motion 132 than the FOV. The camera assembly 110 is shown covering a portion of the conveyor 130, which includes a checkpoint for such objects. The conveyor 130 delivers an encoder count 210, which can be reset with each triggering event, where the conveyor 130 is depicted in the area of the 3D camera assembly. A trigger is issued whenever an object passes through the triggering plane 230 aligned with the presence detector 136 Figure 1 ) described above. The 3D FOV defines a usable region of interest 250, which extends from the surface of the conveyor 130 along the camera axis OA and on either side of the axis OA. In this example, the region of interest (ROI) defines a usable height HR, which can be defined by the user according to the maximum object height expected. The points 256 and 258 where the top 254 of the ROI 250 intersects the front BF and back BR boundaries of the FOV, respectively, effectively define the length LR of the usable ROI in the direction of motion 132 (y-axis). Thus, the higher the maximum height HR of the ROI, the shorter the length LR.

[0030] B. Object size inference and trigger / encoder logic

[0031] After the usable ROI is defined, the processor determines how many pulses the conveyor travels to reach the length LR. Referring to Figure 3 , the system is shown operating in a run-time mode to infer the size (length) of the object 120, which is depicted as the front end 310 of the object 120 reaching the triggering plane 230. The back end 320 is outside the back boundary 340 of the usable ROI 250. At this point, the 3D camera assembly is triggered to acquire a first image (snapshot) of the object, and the encoder count 330 is set to zero in the processor. The snapshot is analyzed to determine if the back end of the object appears within the ROI (using vision system tools, height changes, etc.). If the object 120 ends within a single ROI 250, the snapshot is reported as the complete 3D image of that object, and the system waits for the next object / trigger. Conversely, if the object 120 appears to end without being within the usable ROI (e.g., there is no substantial change in object geometry within the back boundary 340 of the ROI 250), as in Figure 3 , the system stores the first image and calculates the ROI length (LR) in encoder pulses (or a known proportion of that length LR). This count is referenced in Figure 4 , where the back end 320 of the object has now entered the usable ROI 250 (past the back boundary 320 of the ROI). At this point (as in Figure 4The pulse count 430 has counted a length LR of the ROI. In this example, the pulse count 430 is 300 mm. This length is equal to or less than the length LR. As described below, the pulse count can be less than the length LR where overlap between snapshot images is desired. This ensures that details of the image edges are not lost. As described below, the pulse count can be greater than the length LR where overlap between snapshot images is not desired. This ensures that the entire object is imaged. Figure 4 As shown, the trigger state 420 is also positive (indicated by green or other color / shading, and indicated by the diagonal shading) at this stage because the object is present at the trigger plane 230. This logical combination of events (positive trigger state and full encoder count) triggers the camera assembly to take a second snapshot of the object 120. The second snapshot is also stored in association with the object from the first snapshot.

[0032] The system analyzes the second snapshot to determine whether the back end 320 of the object 120 is now present downstream of the back ROI boundary 340. If so, the entire object length has been imaged, and the two snapshots can be combined and processed as described further below. If the back end of the object is not in the ROI of the second snapshot, the encoder count is reset, the trigger state 420 remains high due to the continued presence of the object, and the system counts until the next ROI length LR is reached. Another snapshot is taken and the above steps are repeated until the nth snapshot, where the back end of the object is finally detected within the ROI 250. At this point, the trigger goes low when the object is completely out of the FOV, and the image results are delivered. Due to the overlap between snapshot images, there are sometimes special cases that can be handled by evaluating the encoder count between discrete image captures. If a previous snapshot has imaged an edge of the object and reported a dimension, but the trigger still goes high (positive), and a new snapshot is taken, then in this case, if the encoder count of the current snapshot is equal to the length of the ROI, the new snapshot is discarded, and no dimensions are reported. This is because it is an extra / unused snapshot caused only by the overlap between snapshot images. Typically, as described below, all snapshots can be combined for processing in association with a single object image.

[0033] C. Complex and / or multiple objects

[0034] In some implementations of the runtime operation, the object can present a complex 3D shape. Referring to Figure 5 , an object 520 is shown with varying height along its top surface 522 and / or overhanging (occluding) features 525 reaches the trigger plane 230. At this point, the trigger state 540 goes positive (green, as indicated by the diagonal shading), and the encoder count 530 starts from 0. However, the complex shape of the object can cause loss of 3D image data or confuse the vision system processor as to the actual location of the back end 526 of the object - in this example, it is still outside the available ROI 250 when the camera assembly 110 takes the first snapshot. Also referring to Figure 6of the process 600, in step 610, the object initially triggers the camera assembly to take a 3D snapshot. In this initial image there can be a single (simple or complex shape) object or multiple objects. The vision system analyzes the image to determine if multiple objects are detected - typically by looking for a boundary that extends to the conveyor surface (baseline) and a gap between the separating boundaries. If two items are not detected, the vision system considers the object to be a single item (decision step 620). Conversely, if two or more objects are detected, the decision step 620 of the process 600 branches to a further decision step 630 in which the gap(s) are analyzed to determine if the gap(s) are less than a minimum distance (determined using calibration data in the camera and conventional measurement techniques). If the gap is greater than the minimum distance, set by the user or automatically, the system indicates (step 640) multiple items in the image. In the case of two items greater than the minimum gap distance (step 640) or a single item greater than the minimum gap distance (decision step 620), the process 600 branches to a decision step 650 in which the vision system determines if all of the object ends are in the image and the trigger state has gone negative. If so, the process 600 branches to step 660 in which the results of the image are output and the encoder is reset, waiting for the next object / trigger to occur. Conversely, if the trigger has not gone negative after the current (initial) image and the back end of the object(s) remains outside the FOV, the decision step 650 branches to step 670 in which the camera assembly counts the encoder pulses and takes another snapshot of the FOV. This continues until the encoder goes negative and the back of the object is detected.

[0035] Referring again to decision step 630 in which there is a gap, but it is less than the minimum distance, the process 600 assumes that the imaging contains a single object and the gap is a result of the absence or missing 3D data within the imaged object. Thus, the object is considered to be a single item and the decision step branches to a further decision step 650 (described above) in which the presence or absence of the back of the object in the image determines the next step. Note that one technique for predicting and providing missing or lost 3D image data is described in commonly-assigned co-pending U.S. Provisional Application Serial No. 62 / 972,114, filed February 10, 2020, entitled “COMPOSITE THREE-DIMENSIONAL BLOB TOOL AND METHOD FOR OPERATING THE SAME,” the teachings of which are incorporated by reference as useful background information. Such a blob tool can be used to generate image results to pass to further use processes.

[0036] D. Image Data Tracking and Result Overlay

[0037] ReferringFigure 7 which shows a continuous top-down (x-y plane) view 710 of a conveyor 720, composed of, for example, three consecutive image acquisitions in the presence of a long object. The three acquired images 730, 732 and 734 of the entire object are shown in sequence. Notably, the encoder count has been set so that a first predetermined overlap distance Dl is between the first pair of images 730 and 732, and a second predetermined overlap distance D2 is between the second pair of images 732 and 734. The overlap distance between object images can ensure that acquired object features are not inadvertently missed or blurred at the edges of the 3D ROI. The vision system can appropriately remove or blend the overlapping regions to compute the actual object dimensions and features, as described below with reference to Figure 8 Note that the term "feature" as used herein should be considered to include the term "vertex", or can be used interchangeably therewith, as the acquired images of the object typically define one or more polygons having associated vertices for use in deriving shape boundaries.

[0038] In operation, Figure 8 The process 800 begins with acquiring 3D images in the manner described above in step 810. That is, a trigger is generated and the encoder is set to count the movement of the object through the FOV. After the count reaches a distance value that allows overlap with a subsequent image, the count is reset in step 820, and the system determines whether this is the last image - the back of the object is in the ROI. If not (via decision step 830), the process 800 branches back to step 810 and acquires another overlapping image. If the image is the last image in the sequence, the process 800 branches (via decision step 830) to step 840, where the series of 3D images (segments of the entire object) are transmitted to a dimensioning process (processor) and associated tools to determine the object dimensions and (optionally) resolve features. Using the distances established by the encoder relative to the image pixel positions, the x-y positions and bounding box heights of each object segment are mapped in step 850. Then, based on this mapping, the overlap (constituting the given known encoder distance between segments) is removed to compute the actual object length. The resulting data can be defined as aggregate feature data, which effectively combines data from individual shapes to establish an overall picture or features of the object, without the need to (have) combine actual images of the object. This resulting aggregate feature data can also be used to determine, for example, the maximum width and maximum height of the entire object, and its relative angle with respect to the direction of conveyor movement. This information is particularly useful in various usage processes, such as logistics to ensure proper handling of objects (e.g. packages).

[0039] Note that in accordance with the systems and methods herein, it is expressly contemplated that generating an actual overall (composite or stitched together) image from the discrete image captures (snapshots) is optional. The aggregated feature data can be used independently of the creation of an overall image based on N x M pixels to provide suitable results for determining object dimensions and / or other processes described below. When desired, an overall image can be generated and used for further processing and / or to provide a visual record of all or part of the imaged object.

[0040] E. Application of Results

[0041] It is contemplated that the aggregated feature data (and / or data related to the vertices) resulting from the above operations can be applied to a variety of tasks and functions related to the imaged object, including but not limited to a stream of packages of various sizes and shapes. Notably, the processes herein can be used to determine a skew angle of the object / package relative to, for example, the direction of travel and / or the boundaries of the surrounding support surface. One potential issue that can be identified and measured using the aggregated feature data herein is the skew angle of the object relative to the direction of travel of the conveyor (and parallel side edges) or other support surface. In addition to potentially causing jams at narrow chutes, gates, or other transitions, skewed data can also cause the object to appear longer than it is and can cause the system to generate false defects. Skew angle data should be considered so that corrective action (i.e., ignoring false defects or straightening the object) can be taken. Notably, skew angle information (and / or other measured features) can be part of the metadata tags applied to the results (aggregated feature data) for use in various downstream operations. Other relevant feature data can include out-of-length, out-of-width, out-of-height, and / or out-of-volume values, which would indicate when an object is too long, too wide, too tall, or too voluminous for a parameter limit. Further relevant data can relate to, but is not limited to:

[0042] (a) a confidence score, which can inform on the shape of the object or the quality of the received data;

[0043] (b) a liquid volume, which can inform on the shape of the object or the true (non-minimal cuboid) volume;

[0044] (c) a classification, which can inform on whether the surface of the object is flat;

[0045] (d) a quantity (QTY) of viewing / imaging data versus an expected quantity of viewing / imaging data;

[0046] (e) location features (e.g., corners, center of mass, distance from a reference (e.g., conveyor edge)); and / or

[0047] (f) damage detection (e.g., based on actual imaged shape vs. expected shape, presence of protrusions / recesses wrapping).

[0048] The use of aggregated feature data generally avoids the need to include more detailed pixel-based image data of the combined object, allowing additional processing to be performed on the object. In processing and manipulating such aggregated feature data, the overall size of the data allows for faster processing and lower processor overhead. Some example tasks can include automation of processes in a warehouse - e.g., rejection and / or redirection of objects that are not dimensionally and / or shape-wise consistent. Thus, the use of such data to divert objects can result in such oversized objects getting stuck in the chutes or curves of a conveyor. Similarly, aggregated feature data derived by the system and method can assist in automating the labeling process to ensure that the position of the object is correct. Further, as noted above, skew information can be used to avoid false defect cases and / or allow the system or user to correct the object in the conveyor stream, thus avoiding soft jam cases. Other object / packaging processing tasks that rely on data can use the data in ways that should be clear to those skilled in the art.

[0049] III. CONCLUSION

[0050] It should be clear that the above-described system and method provide an effective, reliable, and robust technique for determining the length of oversized objects that can not fit entirely within the 3D area scan sensor FOV along the direction of movement of the conveyor. The system and method use conventional encoder data and detector triggers to generate an accurate set of object dimensions, and can operate in cases where 3D data is missing or not present due to complex shapes and / or the presence of multiple objects in the conveyor stream.

[0051] Illustrative embodiments of the application have been previously described in detail. Various modifications and additions can be made without departing from the spirit and scope of the application. Features of each of the individual embodiments described above can be combined with each other's features in accordance with the application. Furthermore, while the foregoing describes a number of separate embodiments of the apparatus and method of the present application, what described herein is merely illustrative of the application of the principles of the present application. For example, as used herein, the terms "process" and / or "processor" should be taken broadly to include various electronic hardware and / or software-based functions and components (and can alternatively be referred to as functional "modules" or "elements"). Moreover, the depicted processes or processors can be combined with or divided into various sub-processes or sub-processors. Such sub-processes and / or sub-processors can be variously combined in accordance with the embodiments herein. Likewise, it is expressly contemplated that any of the functions, processes and / or processors herein can be implemented using electronic hardware, software composed of program instructions of a non-transitory computer readable medium, or a combination of hardware and software. Furthermore, various directional and orientation terms, such as "vertical", "horizontal", "up", "down", "bottom", "top", "side", "front", "back", "left", "right", etc., as can be used herein, are merely used as relative conveniences and not as absolute directions / orientations with respect to a fixed coordinate space (e.g., the direction of force of gravity). Moreover, where the term "substantially" or "approximately" is used in reference to a given measurement, value or characteristic, it is meant to encompass amounts that are within normal operational ranges, but also includes amounts that are within a range of inaccuracies and errors that are inherent in systems that are permitted within a tolerance range. Accordingly, this description is by way of example only, and is not intended to limit the scope of the application.

Claims

1. A vision system having a 3D camera assembly arranged as an area scan sensor, and a vision system processor that receives 3D data from images of an object acquired within a field of view (FOV) of the 3D camera assembly, the object being conveyed in a direction of conveyance through the FOV, and the object defining an overall length in the direction of conveyance between opposite edges of the object that is longer than the FOV, the FOV defining an available region of interest (ROI), the vision system comprising: a dimensioning processor that measures the overall length based on motion tracking information derived from the object being conveyed through the FOV, and in conjunction with a plurality of 3D images of the object acquired by the 3D camera assembly in a sequence having a predetermined amount of conveyance motion between 3D images; a presence detector associated with the FOV that provides a presence signal when the object is located in proximity to the presence detector, wherein the dimensioning processor, in response to the presence signal, is arranged to determine whether the object appears in more than one image as the object moves in the direction of conveyance, wherein the dimensioning processor, in response to information related to a feature on the object, is arranged to determine whether the object is longer than the FOV as the object moves in the direction of conveyance; and wherein the dimensioning processor is arranged to determine a length of the available ROI (LR) based on the motion tracking information, wherein when the length of the available ROI (LR) is greater than the predetermined amount of conveyance motion between the 3D images, the dimensioning processor is arranged to determine an overlap of the 3D images such that the opposite edges are included in overlapping 3D images; and an image processor that generates aggregate feature data in conjunction with information related to a feature on the object from successive image captures from the 3D camera assembly, to determine an overall dimension of the object without combining discrete individual images into an overall image.

2. The vision system of claim 1, wherein, the image processor is arranged to acquire a series of image capture snapshots while the object remains within the FOV and until the object exits the FOV, in response to the overall length of the object being greater than the FOV.

3. The vision system of claim 2, wherein, the image processor is arranged to derive overall attributes of the object using the aggregate feature data and based on inputted application data, and wherein the overall attributes include at least one of a confidence score, an object classification, an object dimension, a skew, and an object volume.

4. The vision system of claim 3, further comprising an object handling process that performs a task with respect to the object based on the overall attributes, the task including at least one of redirecting the object, rejecting the object, sounding an alarm, and correcting a skew in the object.

5. The vision system of claim 1, wherein, the object is conveyed by a mechanical conveyor or manual operation.

6. The vision system of claim 5, wherein, The tracking information is generated by an encoder operably connected to the conveyor, a motion sensing device operably connected to the conveyor, an external feature sensing device, or a feature-based sensing device.

7. The vision system of claim 1, wherein, The dimensioning processor uses the presence signal to determine continuity of the object between each of the images as the object moves in the convey direction.

8. The vision system of claim 7, wherein, The plurality of images are acquired by the 3D camera assembly, have a predetermined overlap between the plurality of images, and further include a removal process that uses the tracking information to remove overlapping portions from the object dimensions to determine the overall length.

9. The vision system of claim 8, further comprising an image rejection process that rejects a last one of the plurality of images acquired due to assertion of the presence signal after a previous one of the plurality of images contains a trailing edge of the object.

10. The vision system of claim 1, wherein, The dimensioning processor is arranged to use information related to features on the object to determine continuity of the object between each of the images as the object moves in the convey direction.

11. The vision system of claim 1, wherein, The dimensioning processor defines a minimum spacing between objects in the images below which multiple objects are considered to be a single object with missing 3D image data.

12. The vision system of claim 1, wherein, The image processor is arranged to generate aggregate feature data about the object including: (a) data that exceeds a length limit, (b) data that exceeds a width limit, (c) data that exceeds a height limit, (d) data that exceeds a volume limit, (e) a confidence score, (f) a liquid volume, (g) a classification, (h) a quantity of actual imaging data of the object QTY versus an expected quantity of imaging data of the object, (i) a location feature of the object, or (j) damage detection related to the object.

13. A method of dimensioning an object with a vision system having a 3D camera assembly arranged as a zone scanning sensor, the vision system having a vision system processor that receives 3D data from object images acquired within a field of view FOV of the 3D camera assembly, the object being conveyed in a convey direction through the FOV, and the object defining an overall length in the convey direction between opposing edges of the object that is longer than the FOV, the FOV defining an available region of interest ROI, the method comprising: measuring the overall length in conjunction with a plurality of 3D images of the object based on motion tracking information derived from the object being conveyed through the FOV, the plurality of 3D images being acquired by the 3D camera assembly in a sequence having a predetermined amount of convey motion between the 3D images; generating a presence signal when the object is in proximity to the FOV, and in response to the presence signal, determining whether the object appears in more than one image as the object moves in the convey direction, and determining whether the object is longer than the FOV as the object moves in the conveyance direction in response to information related to features on the object; based on the motion tracking information, determining a length LR of the available ROI, wherein when the length LR of the available ROI is greater than a predetermined amount of transfer motion between the 3D images, the method comprises determining an overlap of the 3D images such that the relative edges are included in overlapping 3D images; and in combination with information related to features on the object from successive image acquisition from the 3D camera assembly, generating aggregated feature data to provide overall dimensions of the object without combining discrete individual images into an overall image.

14. The method of claim 13, further comprising: in response to the overall length of the object being longer than the FOV, acquiring a series of image acquisition snapshots while the object remains within the FOV and until the object exits the FOV.

15. The method of claim 14, further comprising: using the aggregated feature data and based on inputted application data to derive overall attributes of the object, and wherein the overall attributes include at least one of a confidence score, an object classification, an object size, a skew, and an object volume.

16. The method of claim 15, further comprising: performing a task with respect to the object based on the overall attributes, the task including at least one of redirecting the object, rejecting the object, sounding an alarm, and correcting a skew in the object.

17. The method of claim 16, using the presence signal to determine continuity of the object between each of the images as the object moves in the conveyance direction.

18. The method of claim 13, wherein, the step of combining information includes generating, (a) data that exceeds a length limit, (b) data that exceeds a width limit, (c) data that exceeds a height limit, (d) data that exceeds a volume limit, (e) a confidence score, (f) a liquid volume, (g) a classification, (h) a quantity QTY of actual imaged data of the object versus an expected quantity of imaged data of the object, (i) a location feature of the object, or (j) a damage detection related to the object.

19. The method of claim 13, further comprising: defining a minimum spacing between objects in the images, and treating multiple objects as a single object with missing 3D image data below the minimum spacing.

Citation Information

Patent Citations

  • Visual inspection method and system for assembly lines

    CN104796617A

  • System and method for reading optical codes on bottom surface of items

    US20130292470A1

  • Self-checkout with three dimensional scanning

    US20180189763A1