Methods and apparatus for sample tube tilt orientation determination and correction

The image processing apparatus with a convolutional neural network corrects sample tube orientations in automated diagnostic laboratories, addressing detection errors and enhancing handling accuracy.

WO2025178816A1PCT designated stage Publication Date: 2025-08-28SIEMENS HEALTHCARE DIAGNOSTICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/015752
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-13
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing image-based detection methods in automated diagnostic laboratories often erroneously detect sample tubes due to interference from objects like barcode tags, caps, or clots, leading to inaccurate handling and processing, which can cause errors and system contamination.

Method used

An image processing and control apparatus using a convolutional neural network to capture and process images of sample tubes, identifying key points and determining their tilt orientation to correct the orientation before handling, thereby improving the accuracy of sample tube handling and reducing errors.

Benefits of technology

The apparatus enhances the reliability of sample tube handling by correcting tilt orientations, reducing errors, and ensuring proper placement, thus improving the efficiency and accuracy of diagnostic processes in automated laboratories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025015752_28082025_PF_FP_ABST
    Figure US2025015752_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Methods for image-based detection of a tilt orientation of one or more sample tubes residing in a holder used in an imaging and control apparatus of an automated diagnostic system. An extent of a tilt angle of one or sample tubes in the holder is determined based on obtaining two more keypoints of the one or more sample tubes from one or more captured images. This may be accomplished by using a convolutional neural network to process the one or more images of the sample tube(s) that appear in the image. Image processing and control apparatus configured to carry out the methods are also described, as are other aspects.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS FOR SAMPLE TUBE TILT ORIENTATION DETERMINATION AND CORRECTIONCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims benefit under 35 USC § 119(e) of U.S. Provisional Patent Application No. 63 / 552,865, filed on February 13, 2024, the disclosure of which is hereby incorporated by reference herein in its entirety.FIELD

[0002] This disclosure relates to methods and apparatus for determining an orientation of one or more sample containers in a holder used in an automated diagnostic system.BACKGROUND

[0003] In vitro diagnostic analyzers are provided in automated medical diagnostic laboratories to assist in the diagnosis of disease based on conducting assays and / or chemical analyses on patient bio-fluid samples. These assays and / or chemical analyses are typically conducted on diagnostic analyzers (e.g., assay or chemical analysis apparatus) that are part of the automated diagnostic laboratory. Such automated diagnostic laboratories may include automated conveyance systems including carriers onto which sample containers containing patient bio-fluid samples (e.g., whole blood, blood serum, blood plasma, urine, saliva, sputum, cerebrospinal, and / or other bodily fluids) can be loaded. Because of the variety of assays and chemical analyses that may be performed, and the sheer volume of testing taking place in such automated diagnostic laboratories, multiple automated analyzers are often included in the automated diagnostic laboratory.

[0004] In some embodiments of automated diagnostic laboratories, an input / output device may be used to receive sample containers resident in trays. The sample containers may be moved from the tray to a carrier that is configured to circuit about a track, for example. The movement from the tray to the carrier (e.g., the pick and then place) may be carried out by any suitable robot. A tray is a holder including an array of receptacles (hereinafter “slots”) into or from which patient samples stored in sample containers in the form of sample collection tubes, vials, and / or the like (hereinafter any of which are referred to as “sample tubes”) are inserted or withdrawn. One or more sample tubes may be received in the holder in the form of a capped tube, uncapped tube, tube top sample cup, and the like, and may be of varying sizes (e.g., heights and widths).

[0005] In automated diagnostic laboratories including a track configured to transport sample tubes, once on the carrier, the sample tubes can be routed to one or more respective analyzers associated with (e.g., arranged about) the track. Typically, a single sample tube iscarried by each carrier and each carrier circuits about the track. In some automated diagnostic laboratories, one or more individual analyzers can accept a tray of samples tubes, while others can accept carriers. In such embodiments, a robot including an aspiration device may be present in the analyzer and may be used to aspirate and dispense samples from the sample tubes resident in the trays or carriers.

[0006] To facilitate handling and processing of numerous sample tubes in such automated diagnostic laboratories, image-based tube top detection methods and apparatus have been described that capture top-down images of the tops of the sample tubes in order to control a robot during picks from or placements to one or more slots of the tray, such as is described in US Pub. 2020 / 0167591 to Siemens Healthcare Diagnostics, Inc. Other methods for picking and placing sample containers are described in US Pub. 2019 / 0160666 and US Pub. 2019 / 0250180 to Siemens Healthcare Diagnostics, Inc. In other embodiments, side images captured of the sample tube at an imaging station can be used to characterize each sample tube (e.g., tube height, width, tube type, and / or cap type) as is described in US Pub. 2018 / 0365530.

[0007] However, some existing image-based detection methods and systems may erroneously detect the sample tube. For example, images may contain other objects (e.g., sample tube barcode tags, caps, tube tray slot circles, tube tray metal springs, clots and / or blood on the sample tube) that may be confused with the geometry of the sample tube. This may adversely affect sample tube handling and processing carried out by the robot. Accordingly, there is a need for improved image-based detection methods and apparatus that process images of sample tubes used in automated diagnostic laboratories and analyzers.SUMMARY

[0008] According to a first embodiment, an image processing and control apparatus is provided. The image processing and control apparatus comprises one or more image capture devices configured to capture one or more images of one or more sample tubes residing in a holder, and a controller comprising a processor and a memory, the controller configured via programming instructions stored in the memory to process the one or more images of the one or more sample tubes to: identify a location of at least two keypoints on each of the one or more sample tubes, and determine a tilt orientation of the one or more sample tubes from the at least two keypoints.

[0009] According to another embodiment, a method of imaging processing an image or images and controlling a robot based thereon is provided. The image processing and control method, comprises capturing one or more images of the one or more sample tubes residing in a holder, and processing the one or more images of the one or more sample tubes and information on one or more slots of the holder to: determine a location of at least twokeypoints on each of the one or more sample tubes, and determine a tilt angle orientation relative to the holder of the one or more sample tubes from the at least two keypoints.

[0010] According to another embodiment, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium comprises computer instructions of a convolutional neural network and parameters thereof capable of being executed in a processor and of applying the convolutional neural network and the parameters to an image of one or more sample tubes to generate one or more heat maps to be stored in the non- transitory computer-readable medium and accessible to a controller to control a robot based on the one or more heat maps, the convolutional neural network comprising convolution layers, max pooling layers, up-sampling layers, and activation layers.

[0011] Still other aspects, features, and advantages of this disclosure may be readily apparent from the following detailed description illustrating a number of example embodiments and implementations, including various modes contemplated for carrying out the present invention. The present disclosure may also be capable of other and different embodiments, and its several details may be modified in various respects.

[0012] Accordingly, the drawings and descriptions are to be regarded as illustrative in nature, and not as restrictive. This disclosure is to cover all modifications, equivalents, and alternatives falling within the scope of the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings, described below, are for illustrative purposes only, and are not necessarily drawn to scale. The drawings are not intended to limit the scope of the disclosure in any way.

[0014] FIG. 1 A illustrates a partially cross-sectioned side view of an image processing and control apparatus configured to capture and processing images of, determine tilt orientation of, and control movements of, one or more sample tubes residing in a holder according to embodiments.

[0015] FIG. 1 B illustrates a top view of an image processing and control apparatus configured to capture and process images of, determine tilt orientation of, and control movements of, one or more sample tubes residing in holders of an input / output device (module) comprising one or more sample trays and / or carriers according to embodiments.

[0016] FIG. 2A illustrates a top view diagram of example sample tray including slots configured to receive sample tubes according to embodiments, the sample tubes being shown in slots in a substantially “ideal” orientation, i.e. , substantially vertical orientation.

[0017] FIG. 2B illustrates a top view diagram of example sample tray according to embodiments, the sample tubes being shown in slots in both “ideal” and “non-ideal” (i.e., tilted above a threshold tilt angle) orientations.

[0018] FIGs. 3A and 3B illustrate a cross-sectioned side and partial top view, respectively, of a sample tube that is received in a holder in a substantially “ideal” orientation according to embodiments.

[0019] FIG. 3C illustrates a cross-sectioned side view of a sample tube that is received in a holder in a “non-ideal” orientation according to embodiments.

[0020] FIG. 3D illustrates a perspective side view of a sample tube that is received in a holder comprising a carrier in a substantially “ideal” orientation according to embodiments.

[0021] FIG. 3E illustrates a cross-sectioned side view of a sample tube that is received in a holder in a non-ideal orientation and is undergoing orientation correction according to embodiments.

[0022] FIG. 4A illustrates a schematic view of an example convolutional neural network configured to determine one or more heat maps from which the tilt orientation of one or more sample tubes may be determined according to embodiments.

[0023] FIG. 4B illustrates a schematic view of a structure of an example convolutional neural network (CNN) configured to determine one or more heat maps from which the tilt orientation of one or more sample tubes may be determined according to embodiments.

[0024] FIG. 4C illustrates an example of a heat map of centroid keypoints for multiple sample tubes residing in a holder (e.g., sample tray) according to embodiments.

[0025] FIG. 4D illustrates an example of a holder comprising a sample tray including multiple sample tubes and showing one or more calibration markers according to embodiments.

[0026] FIG. 4E illustrates a schematic view of an example training method configured to generate a full-frame keypoint detection model from which tilt orientation of one or more sample tubes in a holder may be determined according to embodiments.

[0027] FIG. 5 illustrates a flowchart of a method of capturing and processing one or more images of one or more sample tubes to determine a tilt angle orientation thereof according to embodiments.DETAILED DESCRIPTION

[0028] Independent of the grammatical term usage, individuals with male, female, or other gender identities are included within the term.

[0029] Automated diagnostic laboratories deal with a high volume of samples per day, and are expected to process these samples with a high level of accuracy. These diagnostic laboratories can include several diagnostic analyzers that sample tubes can be routed to in order to perform assays or clinical chemistry analysis on samples. Such analyzers can be connected to each other through a transportation mechanism (e.g., a track) on which carriers ride that hold the sample tubes. In some embodiments, the connected analyzers pick the sample tube off from the track, aspirate a sample, perform an analysis, and then place thesample tube back onto a carrier resident on the track so that it can be transported to other analyzer. In some embodiments, the sample tube may be placed on the carrier by a robot at an input / output device (hereinafter “I / O module”) and may be routed back to the I / O module for removal from the track by the robot after processing.

[0030] The robots of the analyzers and I / O module carrying out the pick operations make several assumptions about the position and orientation of the sample tube in order to pick it and then place it, but in many cases these assumptions may be too restrictive or otherwise flawed. For example, an analyzer or I / O module may expect the sample tube carried by a carrier on the track to stop at a certain position on the track, be substantially perpendicular to the track / carrier surface (i.e., substantially upright), and have a specific tube height or range of tube heights. The particular robot performing the pick operation may have a fixed approach and fixed grasp position with which it will attempt the pick operation on the sample tube. Combining these assumptions can create a very restrictive environment within which the sample tube can undergo the pick operation.

[0031] Thus, it should be recognized that sample tube orientation within a carrier in a real-world diagnostic system can vary due to variety of factors. For example, a sample tube may rattle around in its carrier or sample rack and thus may not be positioned in an ideal orientation, such as substantially upright (i.e., may be tilted in a non-ideal orientation) or simply may have been loaded in a disorientated position. In some instances, disorientation can happen if the carrier hits a side wall of the track or the track introduces sudden accelerations into the carrier, causing the sample tube to shift out of place within a slot of the carrier. A robot with rigid pick constraints may attempt to pick up a tilted sample tube and this may cause a crash into the sample tube, potentially breaking the sample tube and / or possibly spilling the sample. This may cause bio-fluid to contaminate the system so that movement of subsequent sample tubes is hindered. Furthermore, in some cases, it may cause the system to be shut down until an appropriate cleanup or other rectifying action can be completed.

[0032] In some cases of extreme misalignment of the sample tube in a slot of a holder, the robot may completely miss the grasp in the physical world but nonetheless register the grasp as being successful. This can put the individual analyzer or I / O module into an inconsistent state and cause it to generate an incorrect result for the current sample tube, and potentially any other samples it processes afterwards. Downstream components (e.g., analyzers) may also be affected if a later system component or analyzer expects a successful result from the previous one, causing errors to propagate throughout the entire automated diagnostic laboratory.

[0033] According to embodiments of the disclosure herein, the inventive method operates to identify an optimal grasping location for each sample tube that makes the pick operationeasy to execute. If it is determined that a tilt orientation of the sample is too far out of the desired or threshold orientation, then the robot can adjust the orientation of the sample tube to a more tolerable orientation prior to making the pick. In this manner, a perfect or substantially perfect pick of the sample tube can be accomplished. This has great benefits in that a proper pick can improve the propensity of a proper place operation when placing the sample tube in a destination slot of a carrier or sample tray, for example.

[0034] Accordingly, because the present disclosure operates to determine the current orientation, that orientation can be corrected so that the sample tube can be reoriented to a substantially ideal state (i.e. , substantially upright). Thus, embodiments described herein will allow a diagnostic laboratory system to handle a wider variety of types of sample tubes and reduce errors, thereby resulting in improved outcomes for the samples being analyzed. As should be recognized, the orientation adjustment may occur in the I / O module (e.g., sample handler) where sample tubes may be misplaced or otherwise disturbed from their expected position. Further, the orientation adjustment can also occur while on the track in a situation where the sample tube is not correctly sitting in a slot of a carrier.

[0035] Apparatus, systems, and methods of the disclosure will now be described with reference to FIG. 1 A through FIG. 5 herein. In some embodiments, as shown in FIGs. 1A- 1B, the configuration of the image processing and control apparatus 100 can include a holder 106T, such as the sample tray (hereinafter “tray”), for example. Holder 106T can be part of an I / O module as shown in FIG. 1 B or an analyzer as shown in FIG. 1 A, or otherwise as part of a handling system. Holder 106T can include a plurality of slots 106S (e.g., receptacles) that are configured to receive sample tubes 104 therein as shown in FIG. 2A and 2B, for example. In the tray embodiments of FIGs. 1A and 1 B, slots 106S may number from 2 to several hundred. For example, the tray may include 5 rows by 11 columns for a total of 55 slots, such as is shown in top view in FIG. 2A. Other configurations of rows and columns can be used, such as 6x12, 8x12, and the like.

[0036] In the embodiment of FIG. 1 B, the imaging and control apparatus 100 is shown as part of an I / O module. The imaging and control apparatus 100 may be configured to pick and / or place sample tubes 104 from and to holders 106C comprising carriers that reside on a track 107. T rack 107 may include one or more offshoots 1070 providing the opportunity for holders 106C comprising carriers 106C to branch off from a main channel of the track 107, as shown.

[0037] In FIGs. 1A and 1 B, the imaging and control apparatus 100 may include an end effector positioning apparatus 101 and an imaging and processing apparatus 102. The end effector positioning and control apparatus 101 is configured as a robot 110 to cause an end effector 114 (e.g., a gripper) to move about above the holder 106T (e.g., a tray) and access and grip and move sample tubes 104 in a work area, such as between the tray 106T and thecarrier 106C, as shown in FIG. 1 B, or between the tray 106T and a component of an instrument (e.g., cuvette ring -not shown) as in FIG. 1A. The imaging and processing apparatus 102 is configured to capture one or more images of the holder 106T and / or holder 106C. Imaging and processing apparatus 102 may include one or more image capture devices 112 and an image capture and processing module 122 coupled thereto.

[0038] In particular, the one or more image capture devices 112 (e.g., a digital camera or other digital image obtaining device) may be placed at any suitable location. In one or more embodiments, multiple images of the holder 106T (or holder 106C) and one or more sample tubes 104 in may be obtained from multiple perspectives (viewpoints or scenes) so that the tilt of the one or more sample tubes 104 can be obtained in three dimensions. For example, the one or more image capture devices 112 may be placed above a moveable loading drawer 125 of an analyzer or I / O module on which the holder 106T or holder 106C is resident. The moveable loading drawer 125 may be moveable relative to a frame 116 of the analyzer or I / O module. The holder 106T may be supported by the loading drawer 125 and moved into the testing instrument or I / O module to a position accessible by the robot 110. Thus, an operator may lay the holder 106T on the loading drawer 125 and slide the loading drawer 125 with resident holder 106T into a position generally below the robot 110. During that movement or even when received stationary in its final position underneath the robot 110, the one or more image capture devices 112 may take one or more digital images (frames) of holder 106T and the sample tubes 104 therein.

[0039] The sample tubes 104 may, in some embodiments, include caps 104C (see dotted outline of cap 104C in FIG. 3A) that close a top of the sample tube 104 for when the sample tube 104 is transported. The sample rack(s) can include sample tubes 104 with caps 104C when the diagnostic system includes a decapper apparatus for removing the cap 104C, such as at a location along the track 107. Each of the slots 106S may include a retaining device 106R, such as a spring (e.g., leaf spring shown or other spring device) that functions to retain the sample tube 104 within the slot 106S, such as by applying a retaining force to a side surface thereof. The retaining device 106R then may cause the sample tube 104 to be offset slightly from the central axis 306C of the slot 106S in some embodiments by a distance D as is shown in FIG. 3A. However, in some embodiments, the retaining device 106R may involve multiple retaining elements (e.g., springs) that attempt to center the sample tube 104 in the slot 106S, such that, in the ideal case, the central axis 306C and the long axis 304C are coincident when the sample tube 104 is substantially upright (vertically oriented in a substantially ideal orientation).

[0040] Again referring to FIGs. 1A and 1 B, the robot 110 includes a gripper 114 coupled to a moveable part of the robot 110, such as a moveable arm 110R or portion of a gantry, such as cross slide 110CS. In some embodiments, the robot 110 may be an R-theta robot asshown in FIG. 1A. Alternatively, as shown in FIG. 1 B, the robot 110 may comprise a gantry robot having a gantry assembly that facilitates movement of the gripper 114. Other types of robots, as well as robots with additional degrees of freedom may also be used. In each case, the robot 110 moves an end effector 114 (e.g., a gripper) in a robot coordinate system (e.g., in X, Y, and Z, where Y is into and out of the paper). The robot 110 shown in FIG. 1A may include a base 110B that may be coupled to a frame 116 of the analyzer, other equipment, or I / O module, and an upright portion 110U. Robot 110 may include a Q rotary motor 110© configured to move the arm 110R rotationally about the vertical axis 111Z (+O and -©). Robot 110 may also include a Z motor 110Z configured to move the arm 110R vertically along the vertical axis 111Z (+X and -X).

[0041] In some embodiments, the arm 110R can be configured to move radially (in the +R and -R directions). Robot 110 may include a radial motor 110RM configured to move the arm 110R axially in the +R and -R directions. “Gripper” as used herein means any end effector member coupled to a robot component (e.g., coupled to a robot arm 11 OR or gantry member such as cross slide 110CS) that is used in robotic operations to grasp and / or move an article (e.g., a sample tube 102) from one location to another. Thus, in one embodiment, an end effector 114 may be a gripper with two or more fingers 114A, 114B and can facilitate carrying out a pick and / or a place operation of a sample tube 104. For example, the robot 110 may be used to pick the sample tube 104 containing a sample 103 (FIGs. 1 A, 1 B, and 3A) from a slot 106S of the holder (e.g., sample tray 106T) or holder (e.g., carrier 106C). Further, the end effector 114 may place a sample tube 104 into a slot 106S in the sample tray 106T or carrier 106C. Some of the sample tubes 104 may be empty (devoid of sample 103); such as shown in row A of FIG. 1A. As shown in FIG. 3A, sample 103 may include multiple components, such red blood cells 303A and serum or plasma 303B. Sample 103 may be centrifuged.

[0042] In some embodiments, the end effector 114 may include two or more fingers 114A, 114B that are moveable relative to one another. For example, some may be generally opposed to one another, and are adapted to grasp articles, such as sides of sample tubes 104. The fingers 114A, 114B may be driven to open and close by any suitable actuation mechanism coupled to the one or more of the fingers 114A, 114B. Actuation mechanism may be any suitable mechanism that moves the fingers 114A, 114B to open and close relative to one another. The actuation mechanism may be linearly acting to move one or more of the fingers 114A, 114B in linear translation to accomplish the grip. In some embodiments, the actuation mechanism may optionally function to rotate the fingers 114A, 114B about a gripper axis 111G. The relative amount of movement of the fingers 114A, 114B may be the same (but in opposite directions) or a different amount or only one may move.

[0043] As will be apparent from this disclosure, the population and / or tilt data obtained by imaging can be used to determine an ideal gripping location for each sample tube 104 by the end effector 114. This may involve gripping at a desired depth or rotation orientation based on the determined tilt and size of the sample tube 104.

[0044] The actuation mechanism and the various motors 110©, 110Z, 110RM may be driven responsive to drive signals from a robot control module 116 of a controller 105. One or more linear position encoders and / or rotational encoders may be included to provide position feedback concerning the motion of the motors and the extent of opening and optionally rotational position of the fingers 114A, 114B. Furthermore, although two fingers 114A, 114B are shown, embodiments of the present disclosure are equally applicable to an end effector 114 having more than two fingers 114A, 114B. Other gripper types may be used, as well. The robot 110 may be any suitable robot type capable of moving the end effector 114 in space (e.g., three-dimensional space) to transport the sample tubes 104. Moreover, although an R-theta robot and a gantry robot are shown, other suitable robot types, robot motors, and mechanisms for imparting X, Y, Z, R, 0, and / or other combinations may be provided. Suitable position feedback mechanisms may be provided for each degree of motion capability (X, Y, Z, R, and / or 6).

[0045] In more detail, FIG. 1 B illustrates another embodiment of imaging and control apparatus 100 on which the orientation determination and control method may be practiced. The imaging and control apparatus 100 includes the end effector 114 mounted to a cross slide 110CS of a gantry robot, which can be moved back and forth on the cross slide 110CS to access any column of the sample tray 106T, for example. Likewise, the cross slide 110CS may move forward and backward along the slide rails 110G to allow access to any row of the sample rack 106. The end effector 114 may be moved vertically (into and out of the paper) to raise and lower the sample tubes 104. Thus, the end effector 114 may be moveable in an X, Y, Z robot coordinate system.

[0046] A pick operation may be made according to the method described above to pick from the sample rack 106 and transport the desired or target sample tube 104 to holder 106C comprising a carrier that resides on, and moves around, track 107. For example, the place may take place on the offshoot 1070, after which the sample tube 104 may proceed onto a main portion of the track 107. Likewise, the sample tubes 104 may be placed, using a place operation, back into the sample tray 106T upon returning from analysis and / or processing of the sample 103. T rack 107 may transport the sample tubes 104 to various pieces of equipment or analyzers to preform testing or otherwise process samples 103 contained in the sample tubes 104. Track 107 may include one or more offshoots 1070 from a main channel to allow loading and unloading from analyzers and / or I / O modules.

[0047] The imaging and control apparatus 100 may further include one or more image capture devices 112 (e.g., a camera, charge coupled device (CCD), active pixel device (e.g., CMOS sensor), or the like) that is configured to capture one or more images such as one or more greyscale or color digital images of the scene including the one or more sample tubes 104 and the holder 106T and / or 106C. In some embodiments, one or more image capture devices 112 may optionally include one or more lighting sources (not shown), such as an LED (light emitting diode) flash or other suitable lighting for the capture of the images.

[0048] In some embodiments, the robot 110 can include a robot component with a moving image capture device 112M mounted thereto to capture one or more images. For example, in some embodiments, a moving image capture device 112M may be mounted on a robot arm 110R or optionally on the end effector 114, such that the moving image capture device 112M is moveable with the motion of the robot 110. This allows capture of one or more images from one or multiple different viewpoints.

[0049] In another embodiment, one or more fixed image capture devices 112F may be mounted at one or more fixed locations that is / are able to capture images of the one or more sample tubes 104 within the workspace of the robot 110 proximate to the holder 106T or 106C, as well as images of the holder 106T and / or 106C. For example, the one or more fixed image capture devices 112F may be mounted to a frame 116 of the imaging and control apparatus 100 either directly or through a structural interconnecting member 1161 (shown dotted in FIG. 1A) or otherwise fixedly / stationarily mounted as shown in FIG. 1 B. Images captured may be processed by an image capture & processing module 122 of the controller 105, which may include a convolutional neural network (CNN) 122 configured to process the one or more images and therefrom determine a tilt orientation of the one or more sample tubes 104 within the holder 106T and / or 106C in three dimensions.

[0050] The controller 105 may include a suitable processor 105P (e.g., microprocessor), memory 105M, power supply, conditioning electronics, drivers and other circuitry adapted to carry out and control the motions of the robot 110 and to control positioning of the end effector 114 in the X,Y,Z robot coordinate system, as well as control an extent of finger 212A, 212B opening distance and optionally a rotational orientation thereof.

[0051] Memory 105M may be any type of non-transitory computer readable medium, such as, e.g., random access memory (RAM), hard, magnetic, or optical disk, flash memory, combinations thereof, and the like. Memory 105M may be configured to receive and store programs (i.e., computer-executable programming instructions) executable by the processor 105P. Note that processor 105P may also be configured to execute programming instructions embodied as firmware. The controller 105 may implement corrective actions to the determined tilt angle and in some cases correct tilt angle orientation of the one or more sample tubes 104 based on the processing of the images captured.

[0052] The one or more image capture devices 112 and image capture & processing module 122 are configured to capture and process one or more images of one or more sample tubes 104. This allows various features of the one or more sample tubes 104 to be automatically imaged and processed. The processing of the one or more images determines various geometrical points (keypoints 325) on or related to the one or more sample tubes 104. For example, two or more keypoints 325 (FIG. 3A), such as a tube top location 325T, tube cap top 325TC, tube cap centroid 325CC, tube centroid location 325C, tube center of mass 325CM, and / or tube bottom location 325B, may be determined. Other suitable keypoints 325 associated with the one or more sample tubes 104 may be determined. For example, keypoints 325 may be located anywhere on the side surface of the sample tube 104 or the space that it occupies. For example, surface keypoints 325S1 and 325S2 may be determined along one side. Any other two spaced apart points on the same side and at the same relative circumferential location may be used. Preferably they are spaced apart so as to provide good readings. The center location 325H of one or more slots 106S of the holder 106T or 106C that are configured to receive the one or more sample tubes 104 may also be imaged and processed.

[0053] Existing image-based tube top circle detection methods and systems may rely heavily on edge detection as the pre-processing step. However, in some cases, these existing edge detection image-based systems can be prone to errors and misidentification of the sample tubes. Moreover, prior systems attempt to pick the sample tubes in a disoriented condition (e.g., including a tilt). As such, if picked (grasped and withdrawn) in that condition, there is a likelihood, especially in situations with a high amount of sample tube tilt, of an improper insertion and possibly a crash or spillage of sample when inserted into another slot, such as a receptacle of a carrier or tray, for example.

[0054] According to the method carried out by the imaging and control apparatus 100, a first stage of the method 500 (FIG. 5) involves, in block 502, capturing, with an image capture device 112, one or more images of one or more sample tubes 104 residing in a holder (e.g., sample tray 106T or carrier 106C). In block 504, the method 500 comprises generating one or more heat maps for each selected one of the two or more keypoints 325 for each one of the sample tubes 104 in the one or more images. For example, FIG. 4C illustrates a heat map 435C of a centroid keypoint 325C (one labeled) on each a plurality of sample tubes 104 in the image for one viewpoint.

[0055] FIG. 4D illustrates a heat map 435TC of a cap top keypoints 325TC on a plurality of sample tubes 104 in the image for the same viewpoint. However, in some embodiments, two or more types of keypoints may be included in one heat map 435. For example, a tube top keypoint 325T, tube bottom keypoint 325B, and centroid keypoint 325C for each tube 104 may be included in one heat map 325.

[0056] The method 500 further include, in block 506, detecting a tilt orientation of the one or more sample tubes 104. In particular, the method operates by detecting an orientation (including ideal and non-ideal orientations) of the one or more sample tubes 104. In the case where the holder 106 is a tray (see tray 106T in FIGs. 1A and 1 B) with multiple slots 106S with multiple sample tubes 104 inserted therein as shown in FIGs. 1 and 2A-2B, the method detects the orientations of all of one or more sample tubes 104 in the tray. In the case of a holder 106 being a carrier 106C, such as shown in FIG. 3D, the method 500 detects the orientation of the sample tube 104 held in the slot 106S of the holder 106. Specifically, the method comprises, in each of the above cases, identifying an amount of a tilt angle eTof a long axis 304L of a sample tube 104 with respect to the central axis 306C of a slot 306S of the holder 106. Thus, a slot 106S may be a slot of: 1) a sample rack, such as at an I / O unit, or 2) a slot of a carrier configured to move sample tube 104 on a track 307.

[0057] Holder 106 comprising a carrier can include an suitable base 306B that is coupled to or is in contact with the track 307 as is shown in FIG. 3D. In the depicted embodiment, the track 307 may be stationary and the carrier (holder 106) may reside and move on a moving member 309, such as a linear motor. Control signals to the linear motor 309 cause linear motion of the holder 106 along the track 307, such as in the direction of the arrow shown. Any other suitable configuration of track 307 may be used, including a moving conveyor-type track.

[0058] Evaluation of the tilt angle (GT) between the central axis 306C of the slot 106S and the long axis 304L of the sample tube 104 can involve comparison of the tilt angle (GT) to two threshold values indicating the tilt state of the sample tube 104. A lower threshold (LT) as shown in FIG. 3C can be selected and preset to indicate a maximum allowable tilt angle GTof the tilt orientation of a normal example of sample tube 104. In an ideal orientation, the sample tube 104 will be perfectly aligned parallel to the central axis 306C of the slot 106 (i.e., perfectly vertically oriented, such as is shown in FIG.3D). However in real-world applications, there may be a deviation (i.e., a finite tilt angle GT) as measured between the long axis 304L and the central axis 306C as is shown in FIG. 3C. A lower threshold LT as shown in FIG. 3C is a tilt angle GT that defines the amount of deviation below or equal to which the imaging and control apparatus 100 considers the sample tube 104 to be in an acceptable state (acceptable degree of tilt angle GT when GT < LT). In this acceptable state, the sample tube 104 may be grasped by the end effector 114 and the pick may be accomplished by the robot 110 without any further adjustment in the orientation of the sample tube 104 within the slot 106S. This allows a high probability that a place operation following the pick operation of the sample tube 104 can occur without incident.

[0059] On the other hand, if the lower threshold LT, as shown in FIG. 3C, is exceeded (i.e., GT > LT) then the sample tube 104 may be in need of an orientation adjustment in theslot 106S prior to undergoing a pick operation. For example, if the tilt angle ©Tis greater than the lower threshold LT (i.e., ©T> LT), but less than or equal to an upper threshold UT (i.e., © < UT), then the deviation is such that the imaging and control apparatus 100 considers the sample tube 104 to be in a state that is non-ideal. Thus, some adjustment should be undertaken in the tilt orientation before any pick may take place on the sample tube 104.

[0060] In this non-ideal state, the sample tube 104 may be nudged or moved. For example, the nudging or movement may be accomplished by the end effector 114 of the robot 110 by way of a suitable control signal from the robot control module 116. For example, the end effector 114 of the robot 110 may be positioned on the side of the tilt as shown in FIG. 3E and moved in the direction of the arrow shown to adjust the orientation of the tilt angle ©Tto be such that the tilt angle eTis less than LT. However, it should be understood that a separate robot (not shown) whose sole function is to cause orientation adjustments may be used. In some embodiments, the goal is to right the orientation of the sample tube 104 to a substantially ideal orientation where the long axis 304L is substantially co-parallel with the central axis 306C. In this manner, the sample tube 104 can be reoriented to an acceptable and substantially ideal state where the tilt angle ©Tis less than or equal to LT (i.e., ©T< LT).

[0061] After the adjustment move, the sample tube 104 may be reimaged to determine the tilt angle GTand to confirm that an acceptable amount of adjustment has taken place. Such re-evaluation the state of the sample tube 104 thus can determine whether additional action is desired. This adjustment can continue until a desired adjustment threshold is met, such as ©TLT. In this acceptable state with the tilt angle ©Tis less than or equal to LT, the sample tube 104 may be properly grasped and the pick may be accomplished by the robot 110.

[0062] The upper threshold UT, as is shown in FIG. 3C, is also used to determine cases where the tilt orientation of the sample tube 104 deviates so much from the central axis 306C that it may not be recoverable to an ideal or near ideal orientation just by having the robot 110 nudge or move it into place. For example, severe cases of this disorientation would show a tilt angle ©Tabove the upper threshold UT (i.e., ©T > UT). The upper threshold UT can be preselected to be a tilt angle ©T that is greater than 60 degrees, or even greater than 80 degrees in some embodiments. Exceeding the upper threshold UT (i.e., ©T> UT) indicates that the sample tube 104 is too far out of the slot 106S of the holder 106 to be pushed back into place or that the sample tube 104 has fallen out of the slot 106S of the holder 106 and onto an adjacent surface. Deviations that lie in between these two thresholds (LT < ©TUT) and are situations in which the robot 110 or another robot may effectively be used to correct the tilt orientation of the sample tube 104 in the slot 106S. The exactmeasurement of the tilt angle (0T) relative to the thresholds LT and UT may be by any suitable convention, such as >, >, <, or <. In the case where the upper threshold UT is exceeded (i.e., 0T> UT) the sample tube 104 has essentially, fallen out of the holder 106 and then grasping procedure can be executed by the robot 110 that would pick up the tube 104 from wherever it fell and put it back into the slot 106S or move it to another location where additional manipulation or correction may take place. For angles above UT, the system won’t be able to push it back into place, so one option is to execute a grasp procedure, with the aid of the imaging devices 112M or 112F that restores the workspace into a consistent state.

[0063] Determining the tilt orientation eTof the sample tube 104 with respect to the central axis 306C of the slot 106S of the holder 106 that contains the sample tube 104 can comprise two phases, accomplished in any order. The first phase can comprise detecting the location of the holder 106 and, in particular, the location of the one or more slots 106S in a horizontal (X-Y plane) as shown in FIG. 1. The second phase comprises detecting the degree of tilt (detecting the tilt angle OT) of the long axis 304L of the one or more sample tubes 104 relative to the central axis 306C of the one or more corresponding slot 106S.

[0064] According to the method, it is assumed that the central axis 306C of the holder 106 (e.g., rack as shown in FIGs. 1-2B or carrier as shown in FIG. 3D) will always be oriented such that the central axis 306C of the slot 106S which holds the particular sample tube 104 is always facing directly upwards (i.e., is vertically oriented).

[0065] To determine the long axes 304L of the one or more sample tubes 104, embodiments of the method can make use of semantic keypoints 325 on the sample tube 104. These semantic keypoints 325 mark important features of the one or more sample tubes 104. The locations of the keypoints 325 can be used to derive the orientation of the respective long axes 304L of the one or more sample tubes 104. Semantic keypoints 325 can include, but are not limited to, at least two of the following: top of the sample tube 104, top of the tube cap 104C, centroid of the tube cap 104TC, centroid of the sample tube 104, bottom of the tube cap, bottom of the sample tube 104, and so on. The previously listed keypoints 325 lie on the long axis 304L of the sample tube 104 and at least two of the keypoints 325 can be used to compute the direction and extent (orientation in 3D space) of the tilt of the one or more sample tubes 104 in the one or more slots 106S of the holder 106.

[0066] Such keypoints 325 can be detected in one embodiment of the present method using an appropriately trained convolutional neural network (CNN 105C) to generate the keypoint predictions. Now referring to FIG. 5, the method 500 comprises, in block 502, capturing (e.g., obtaining), with an image capture device (e.g., image capture device 112), one or more images (e.g., two or more 2D images from different viewpoints or 1 3D image containing depth information) containing the holder 106 and the one or more sample tubes104 residing in the holder 106. In the case where the holder 106 is a sample rack, the image can contain two or more sample tubes 104 residing in the holder 106, which may be in various and different orientations therein, some of which may be non-ideal.

[0067] According to the method 500, the CNN 105C can be trained to generate heat map(s) that denote the region(s) in which the desired keypoints 325 are likely to be found. Thus, in block 504, the method 500 comprises generating one or more heat maps (e.g., heat map(s) 435) for each of the two or more keypoints 325. The two or more keypoints 325 may be selected so as to be able to determine the location of the long axis 304L, as shown in FIGs. 3A through 3C, and thus the tilt angle eTthereof as shown in FIG. 3C. The two or more keypoints 325 may be, but are not limited to, any two or more of the following geometrical points:• a center top 325T of the sample tube 104 (geometrical center at a plane of the top surface 326 of the sample tube 104,• a center top 325CT of the tube cap 104C of the sample tube 104 (geometrical center at a plane of the cap top surface 328 - if the sample tube 104 includes a cap 104C),• a cap centroid 325CC of the tube cap 104C of the sample tube 104 (geometrical center at a plane of the cap top surface 328 - if the sample tube 104 includes a cap 104C),• a centroid 325C of the sample tube 104,• a bottom 325B of the sample tube 104 (e.g., in holders 106 that include a side viewing capability such as holder 106 in FIG. ), and• any two points along a side surface of the sample tube.Any other definable keypoint 325 along a physical center or an extension of the physical center of the sample tube 104 may be used. The long axis 304L extends through and includes any two or more of the keypoints 325. The two or more keypoints are determined, in block 504, by processing the one or more images of the one or more sample tubes (e.g., one or more sample tubes 104) and information on one or more slots (e.g., slots 106s) of the holder (e.g., holder 106) in order to: determine the location of at least two keypoints (e.g., 325) on each of the one or more sample tubes, and based thereon determine a tilt angle orientation (e.g., tilt angle GT) relative to the holder of the one or more sample tubes from the at least two keypoints.

[0068] The one or more heat maps 435 (FIGs. 4A-4C) are generated as follows. The captured input image 401 is processed by being passed through the convolutional neural network 105C that performs convolution and pooling operations in the first phase, and then deconvolution in the second phase. The output of the CNN 105C can be an image with thesame dimensions as the captured image, but designed to have values in the range of 0 to 1. This output image is compared to the ground truth heat map, which is also in the range of 0 to 1 , in order to train the CNN 105C to predict the region where the one or more keypoints 325 are located. The ground truth heat map is generated by using the known 2D locations of a particular keypoint 325, and then generating a 2D probability distribution around that location. For such a distribution, the maximum value of 1 will occur at the exact location of the particular keypoint 325, and will decrease to zero over some pre-defined radius. Outside this radius, the rest of the output image has a value of 0.

[0069] In more detail, a heat map 435 can be generated for each desired keypoint 325 for each of the sample tubes 106, or a single heat map 435 can contain all of the keypoints 325 for each of the sample tubes 104 in the scene, or some subset of the keypoints 325.

[0070] Once two or more keypoints 325 have been detected in 2D, the apparatus 100 and method 500 then is configured to find their corresponding 3D points in a world coordinate system. This can be done with one of at least two methods. The first method involves using a distance calibrated 3D image capture device such as a distance calibrated 3D digital camera, and the second method involves acquiring multiple images of the sample tube 104 and holder 106 from different perspectives (locations) such that at least two keypoints 325 are visible in all the images captured.

[0071] Using the first method, a calibrated red, green, blue - depth (RGBD) image pair can capture a single image that will provide all the information necessary to determine the 3D locations of the two or more keypoints 325. The CNN 105C can be trained using image types of color (e.g., Red, Green, and Blue), depth, or both, to output 2D locations of the two or more keypoints 325, and the intrinsics of the image capture device 112 can be used in conjunction with the depth image. A depth image is an image with the same size as the color image, but each pixel contains the distance of the image capture device 112 to the location in the real-world captured by that pixel. The output 2D locations of the two or more keypoints 325, and the intrinsics can be used to reproject the predicted two or more keypoints 325 into the image capture device’s coordinate system. Reprojection is the process of converting the locations of the 2D keypoints 325 found in the color image into their real-world 3D counterparts.

[0072] When using the second method, several images are taken in which at least two (but preferably all the possible keypoints 325 are expected to be seen) are imaged and provides a set of correspondences between two different 2D viewpoints that can be used to compute the 3D locations of the at least two keypoints 325. This method relies on stereo vision techniques to calculate the relative motion of the image capture device 112 between images and infers the depth of the keypoints 325 using correspondences between the at least two captured images. 3D reconstruction via stereo vision is achieved by capturing twoor more images of a scene, and one or more sample tubes 104 in this case, and identifying keypoints 325 within those images. By identifying the same keypoint 325 in all the images, the method 500 can use the intrinsic properties of the image capture device 112 and the mathematics of projection in order to calculate the location of the keypoint 325. In this case, knowledge of the motion of the image capture device 112M between the different images is used in order to recover the pose of the robot 110 and the image capture device to gripper calibration. The two views are combined to generate the 3D information.

[0073] More succinctly, the method can take one image of the tube 104 and estimate the keypoints 325 then move the robot 110 to a different position and orientation and again observe the tube 104 and estimate the keypoints 325. The motion between the two views is retrieved by querying the robot 110 for its pose at both the first and second locations, and then applying an image capture device to gripper transform to get the pose of the image capture device 112M at each location.

[0074] Once found, the 3D locations of the two or more keypoints 325 for each sample tube 104 can be used to calculate the orientation of the long axis 304L of each of the sample tubes 104 in the world coordinate system (X-Y-Z coordinate system). The angle between the two vectors (long axis 304L and slot axis) can be calculated using a dot product thereof. Vectors are obtained from a unit vector that points along the central axis 306C, and a unit vector from the 3D keypoints 325 along the tube axis 304L. The dot product equation can be rearranged to solve for the angle between the two vectors. The dot product equation is: u ■ v — ||it|| || v\\ cos 9

[0075] The central axis 306C of the holder 106 can be found using the center bottom of the slot 106S (can be coincident with bottom keypoint 325B when long axis 304L and central axis 306C are coincident) and the assumption that this central axis 306C will point directly vertical from the holder 106. The central axis 306C of each slot 106S of the holder 106 can be found, in one method, by using a slot detector comprising a separate CNN model. This model can use a bounding box for each of the slot 106S, and the center of the slot 106S can be found by averaging the locations of the box corners. Using a tray placed on a flat surface, the method can transform between the image capture device 112 and the world. With this information, the central axis 306C of each slot 106S can be determined. Other suitable slot detection methods may be used to determine the slot locations. With the slot detector, the 2D locations of slots 106S in the tray 106, can be reprojected to get their 3D locations, and then use the cross product formed using the vectors centered on the slot 106S of interest in order to get the axis. At least three points are required for this method. For a single-slot carrier, a similar keypoint approach can be taken where keypoints on the same plane can be detected (four corners of the carrier 106C, for example), and the cross-product method can be applied.

[0076] By combining this information concerning the two or more keypoints 325 and the location of the central axis 306C, the tilt angle of deviation from ideal (tilt angle OT) can be found can be easily mathematically computed.

[0077] In the case of a biological fluid being the sample 103 contained in the sample container 104, there are many instances where the sample tube 104 can be accidentally broken and sample 103 can seep out, or sample 103 can escape during a de-capping procedure taking place. This sample 103 may flow onto the outside surface of the sample tube 104, which can affect the appearance of the sample tube 104 and potentially even the shape of the sample tube 104, such as when the sample 103 or a blood clot is formed on the outside surface thereof. Additionally, the sample 103 can leak onto the surrounding apparatus and possibly other sample tubes 104. Advantageously, the present method can be trained to be robust to these situations by detecting one or more keypoints 325 above or below that avoid these areas of the sample tube 104 prone to contamination. Due to the symmetry of sample tubes 104, detecting equivalent keypoints 325 elsewhere on the sample tube 104 can be easily achieved.

[0078] The present method is also able to take into account the condition of a barcode label 315 adhered to a sample tube 104. Barcode labels 315 can be encountered on the sample tube 104 in a variety of conditions. Some may be pristinely and accurately applied as per the recommended instructions, but some may be applied with the barcode label 315 being skewed off the long axis 304L of the sample tube 104, or may include an un-adhered portion, such as a label flap. During transportation and processing, the barcode label 315 can lose its adhesiveness and partially peel off the sample tube 104. Parts of the barcode label 315 may also tear during processing, leaving pieces of barcode label 315 that can alter the overall composite shape of the sample tube 104. The present method is able to handle these situations as well and avoid suggesting certain keypoints 325 in locations where there is greater variability in the shape of the sample tube 104 due to the partially adhered barcode label 315, for example. By considering the condition of the barcode label 315 together with the possible presence of sample 320, the present method can lead to more accurate handling of sample tubes 104 in challenging laboratory diagnostics settings. Bad labels can be discovered by using additional keypoints 325. For example, keypoints 325 could include the four corners of the label, and the four corners of the barcode within the label. These keypoints 325 have predictable geometric relationships with each other, and this structure can be exploited to detect erroneous barcode conditions. For example, if a label is torn and part of the barcode is also torn, that part of the label may hang off of the tube 104. The keypoint detector would be able to detect the 8 keypoints and then compare their 3D locations to a template configuration and see that keypoints close to the tear are out of place. This would be an indication of a non-standard condition. In the case of blood clots,the clot could cover an area of the barcode or label that may distort its appearance in the image. This may cause the model to predict the keypoint in a slightly different location due to the deformation. Again, this is a condition that could be detected by comparing the keypoint structure to a template as is done for label tears or label bunching.

[0079] Once the two or more keypoints 325 are determined for the one or more sample tubes 104, and the orientation of the long axis 304L is determined, orientation correction can be implemented, provided the upper threshold UT has not been exceeded. The orientation correction, as shown in block 504, involves controlling a robot (e.g., end effector 114 of a robot 110 as shown in FIG. 3E, for example) based on the determined tilt angle (e.g., eT) orientation of the one or more sample tubes (e.g., one or more sample tubes 104). To correct the tilt orientation of a particular one of the sample tubes 104, the robot 110 determines how to adjust the deviation to a substantially ideal orientation (e.g., substantially upright). By substantially upright, what is meant is that the sample tube 104 is adjusted to be within + / - 5 degrees of being co-parallel with the central axis 306C of the holder 106 and less than LT. The processing of the one or more images captured by the one or more image capture devices 112 is used to determine the geometry of both the sample tube 104 and the holder 106, and then determine an optimal trajectory along which the end effector 114 can push the sample tube 104 back into place. The 3D keypoints 325 detected by using the CNN 105C can be used for this purpose.

[0080] Keypoints 325, for example, that indicate the top (325T) and bottom (325B) of the sample tube 104 can be used to calculate the height of the sample tube 104, as well as the current location of the top 325T of the sample tube 104 (if determined that the sample tube 104 has no cap 104C). Using the keypoint 325H of the holder 106 and the height of the sample tube 104, a target location (final placement location) can be defined in the world coordinate system. Using these two keypoints 325, a robot trajectory can be defined in which some part of the end-effector 114, such as a gripper finger 114A or 114B, can be positioned to push the target push point 329 of the sample tube 104 from the starting position 330 (e.g., as shown in FIG.3E) to the target position (e.g., as shown in FIG. 3A). Trajectories can be generated using linear functions or even higher-order functions. The trajectories can be determined based on the current positions of the keypoints 325 and their desired endpoint. For example, if a tube 104 is tilted at 45 degrees, the top point is to the side of the desired position and slightly lower in height. The method can then design a trajectory curve that the top keypoint can take to move to the desired location and use that as information for the robot 110. The method can take evenly spaced 3D points along the curve, and calculate the inverse kinematics of the robot 110 for each point. These points will tell where the end effector 114 should be at each point in space to push the tube 104 back into place. Higher- order functions can be used to take into account the center of mass 325CM of the sampletube 104, given that a sample tube 104 can be filled with different amounts of sample 320. These functions can use the center of mass 325CM to safely restore the sample tube 104 to the target position 331 with less potential to disturb or spill the sample 320 in the sample tube 104. The center of mass of the sample tube 104 can be approximated as a cylinder of plastic and use geometry to get the center of mass of this when it is not filled with fluid. When there is fluid present, we can use a pre-determined value for the density of blood and then estimate the volume of sample in the tube 104 by observing the meniscus. From these pieces of information, the method can find the center of mass using standard calculations for composite bodies.

[0081] The CNN 105C may be trained in order to recognize the keypoints 325. The CNN 105C may also be trained to recognize the size (height and width of the sample tube 104), the center of mass 325CM of the sample tube 104 given the depth of the sample 320 therein, as well as whether the sample tube 104 includes a cap 104C.

[0082] The ground truth data for training the CNN 105C can be generated in several ways. In one example, the method uses both synthetic data and real data to train the CNN 105C. To generate the synthetic data, a 3D visualization tool can be used along with a generalized, parameterized shape model for the sample tube 104. According to the method, shape parameters may be modified on the shape model, such as the height (e.g., Height to the top of the cap HC, height to the top of the tube HT (if no cap) and width W of the sample tube 104. These simulated sample tubes 104 are inserted into a virtual environment that is similar to the real world setting. The keypoints 325 of interest are defined on the virtual sample tubes 104, and their location within the virtual world is recorded. The simulation environment provides the relationship between the world coordinate system and the coordinate system of the image capture device 112, as well as the ideal intrinsic parameters of the image capture device 112. The intrinsic parameters are used to map 3D points in the world into the 2D image plane. In a simulation environment, these parameters are configurable and can be set to match whatever real-world image capture device 112 (e.g., camera) that is being used.

[0083] Using this information, the method can use the simulation environment to generate synthetic images with 2D image coordinates of the keypoints 325 of the sample tube 104. This simulation process can be repeated by inserting different virtual sample tubes 104 with a variety of feasible shapes and sizes (heights HT, HC, and widths W, with and without caps 104C) in order to introduce many variations into the scene (image frame). Any number of these computer-generated scenes can be initiated and captured as training data, and the amount of data is limited only by time and computing capacity. For example, several hundred or more examples of different sample tubes 104 may be trained by the CNN 105C.

[0084] While simulated data has benefits in terms of rapid training of the CNN 105C to recognize different tube types and capped or capless sample tubes 104, real annotated data can be acquired by the method and the CNN 105C can be trained on it in order for the imaging and control system 100 to perform well in real-world scenarios. The real-world ground truth data is generated by making use of coordinate transformations within the workspace (the space including the one or more sample tubes 104).

[0085] According to the method, the keypoints 325 are defined in the tube coordinate system, which provides a base reference frame from which all the other transformations can be performed. The keypoints 325 in the tube coordinate system are defined by annotation. The method involves populating a holder 106 (e.g., a tray with several sample tubes 104 or a carrier with one sample tube 104) that is desired to be included in the data collection. The holder 106T (e.g., tray) or 106C (e.g., carrier) has a holder coordinate system associate with it, and this may be defined based on the tray or carrier geometry, as the case may be. For example, the tray geometry is used by the method to compute the location of the slots 106S in the holder 106 in the holder coordinate system. The method can then compute a transformation between the tube coordinate system and the holder coordinate system, which then can be used to get the 3D keypoint coordinates in the holder coordinate system. After that, the method computes a transformation between the holder coordinate system and the world coordinate system.

[0086] This transformation can be accomplished in several ways depending on the nature of the workspace scene, and what recognizable features that can be taken advantage of on the holder 106. One possible method is to define an origin of the workspace, and a grid 440 with different locations for the holder 106 can be created. This grid 440 and one or more calibration markers 440CM and locations of various features of the holder 106, such as the slots 106S, can be used to compute the location of the holder 106 and features thereof in the world coordinate system.

[0087] FIG. 4D illustrates a mock type of sample tray 106T arranged proximate to a grid 440, wherein the grid 440 and the sample tray 106T are capable of being imaged within the scene by an image capture device 112F or 112M. The grid 440 can include one or more grid features, such as calibration markers 440CM that allow the orientation of the grid 440, and thus the location of the sample tray 106T to be determined in the tray coordinate system. For example, two or more corners or other prominent features of the calibration marker 440CM may be used for grid orientation and for determining a grid origin, which may be any point on the grid within the scene, for example. The calibration markers 440CM provided in the scene enable quick determination of the orientation of the tray 106T within the scene.

[0088] The last 3D transformation carried out by the method 500 is the transformation from the world coordinate system to the particular image capture device 112F or 112M used.This transformation can be accomplished in several ways. One way is to use one or more calibration markers 440CM (e.g., locator patterns) in the workspace scene. The image capture device 112 can detect these one or more calibration markers 440CM (e.g., Hoffman markers) in any images or collection of images (e.g., videos) that it captures. The transformation between the image capture device 112 and the origin of the one or more calibration markers 440CM can be computed using known calibration techniques. By detecting the pattern and corners of the one or more calibration markers 440CM, for example, the pose of image capture device 112F with respect to the origin of the world coordinate system can be found, regardless of which one or more calibration markers 440CM are seen.

[0089] Alternatively, the image capture device 112M can be moving while attached to a component of the robot 110, such as a robot arm 110R or end effector 114 and the transformation between the robot 110 and the world coordinate system is known, then the position of the image capture device 112M can be computed using the world-to-robot and forward kinematics of the robot 110. Once all these 3D transformations are discovered, they can all be combined to create a general tube-to-image capture device transform. This, combined with the intrinsics of the image capture device 112, can directly compute the 2D image coordinates for every sample tube 104 in the workspace scene for any single image captured. With this method for annotating keypoints 325 in a single captured image of the workspace scene, we can extend the data collection to annotate data for images captured from more than one viewpoint (e.g., vantage point). For example, images from two generally orthogonal vantage points or from two or more images from a video can be used. For example, the robot 110 with attached image capture device 112M can be moved in the workspace and record two or more images (e.g., a video made up of multiple image frames from different viewpoints) of the holder 106T or 106C containing the one or more sample tubes 104. From this an annotation can be created for two or more image frames of the video and the CNN 105C trained thereon. Once the synthetic and real annotated data are obtained, they are combined by indicating the percentage (e.g., 40%) will be simulated images while the other percentage (e.g., 60%) are real-world images.

[0090] The architecture of the CNN 105C can be configured to receive images that can be RGB (Red-Green-Blue) images, depth images (images obtained from a depth sensor of the image capture device 112), or RGB-depth combination of images, and produces another image (or data). That produced image (or data) can be interpreted in a manner that can extract the predictions of the keypoints 325. One approach is to have the CNN 105C predict a heat map 435 (e.g., heat maps 435TC) as is representatively shown in FIG. 4C. A heat map 435 is an output from the CNN 105C that is a digital image or collection of data that has high values where the keypoints 325 are likely to be found, and lower values where they arenot likely to be found. Heat maps 435 can be predicted for each semantic keypoint 325. Moreover, multiple keypoints 325 can be predicted by the CNN 105C. Heat maps 435 are generated from the ground truth annotation by defining a probability distribution around the 2D image locations, with the maximum value of the distribution occurring at the location of the keypoint 325, and the values decreasing to zero further away from the particular keypoint 325.

[0091] To ensure that the predicted keypoints 325 are consistent with the shape of the sample tube 104, they can be matched with other keypoints 325 that may belong to the same sample tube 104, and the CNN 105C can be taught to recognize the sample tube 104 as a whole via finding two or more of its constituent keypoints 325. For example, the center of the cap top keypoint 325TC and centroid 325C can be jointly found. The loss function of the CNN 105C can be designed to reward predictions where all the keypoints 325 are found, and increasingly penalized if fewer keypoints 325 in the sample tube 104 are found.

[0092] The tilt angle ©Tof the sample tube 104 is determined based on the predicted two or more keypoints 325 relative to the orientation of the holder (e.g., sample tray 106T or carrier 106C). Using the predicted locations of the keypoints 325, a vector describing the long axis 304C of the sample tube 104 can be calculated. The long axis 304C of the sample tube 104 can be calculated by subtracting a first vector, such as representing the bottommost keypoint 325 on the tube 104, from a second vector, such as representing the top-most keypoint 325, to obtain a third vector representing the long axis 304C. Both keypoints 325 are 3D points. The origin of the first vector and second vector can be the origin of the world coordinate system, which can be a defined origin point of the calibration grid based on the location of a point on one or more calibration markers 440CM in the scene. Other suitable points in the scene may be used as the origin. Any two corresponding keypoints 325 along the center or side of the sample tube 104 may be used for determination of the tilt of the long axis.

[0093] Additionally, the method can use knowledge of the holder 106 to get a vector describing the central axis 306C of the slot 106S. Knowledge of the location of the one or more slots 106S may be determined via training the CNN 105C or by relating the location of one or more slots 106S to one or more calibration markers 440CM. It is assumed that the central axis 306C of the slot 106S is always perfectly vertical. Thus, a vector can be determined for each of the one or more slots 106S. Once these two vectors (vector one describing the long axis 304C and vector two describing the central axis 306C) are calculated, a number of different metrics, such as the dot product, can measure the amount of the tilt angle eTbetween them. The tilt angle orientation can then be transmitted to the robot 110 in order to correct the orientation of the sample tube 104, if needed.

[0094] During operation of the method, the correction can happen in two ways. First, the image capture device 112F or 112M can take one or more images of the scene from an appropriate viewpoint, pass it to the trained CNN 105C, and then acquire all the information about the tilt angle ©T. If one image is taken, then a depth image (from a depth sensor of the image capture device 112F or 112M) would be used to obtain the 3D tilt information. Two images can be used to obtain the 3D data when color (non-depth) images are obtained (such as from an RGB camera). Then, a trajectory can be planned for the robot 110 to correct the tilt angle, and the robot 110 can subsequently execute that plan without any additional information.

[0095] In the second method, an image capture device 112M can capture two or more images, or even capture images continually, as the robot 110 moves, and passes them through the CNN 105C. For each image, a prediction of keypoints 325 and tilt angle analysis is performed. This stream of information can be continually monitored for the locations of the keypoints 325, and the tilt angle correction plan can be adjusted as the image capture device 112M and robot 110 move closer to the sample tube 104 being adjusted. This provides advantages over the first method in that if the initial viewpoint generates a sub-optimal trajectory, then subsequent observations of the scene can correct those errors. By iteratively or continually monitoring the scene like this, the imaging and control apparatus 100 has a higher probability of successfully determining and then correcting the tilt angle GTof the sample tube 104.

[0096] In each case, image processing software stored in the image capture and processing module 122 of the controller 105 may receive and process the one or more digital images. From the images, data may be produced by the CNN 105C including tilt angle data. The tilt angle data may be accessed by the robot control module 116. Access may be either through a download of the data from the image capture and processing module 122 or by gaining access to a database resident in memory 105M. In some embodiments, the robot control module 116 and image capture and processing module 122 may reside in separate controllers, and communication there between may be by LAN or other suitable electronic communication between the controllers.

[0097] Further examples of the sample rack imaging systems may be found in U.S. Pat. Pub. No. US2016 / 0025757 filed March 14, 2014, to Pollack et al. entitled “Tube Tray Vision System”; PCT Application Pub. No. WO2015 / 191702 filed June 10, 2015, and entitled “Drawer Vision System”; PCT Application No. PCT / US2016 / 018100 filed February 16, 2016, and entitled “Locality-Based Detection Of Tray Slot Types And Tube Types In A Vision System”; PCT Application No. PCT / US2016 / 018112 filed February 16, 2016, and entitled “Locality-Based Detection Of Tray Slot Types And Tube Types In A Vision System”; andPCT Application No. PCT / US2016 / 018109 filed February 16, 2016, and entitled “Image- Based Tube Slot Circle Detection For A Vision System.”

[0098] In some embodiments, once the sample tubes 104 are in a substantially ideal orientation, pick operations may take place in any ordered sequence, such as row-by-row, column-by-column, or in any other ordered pattern. In some embodiments, after a first pick is made, the system and method can be used to analyze the rest of the sample rack 106 and create a pick strategy that takes into account the remaining sample tubes 104 and their tilt orientation. This pick order may be determined rapidly, and while the first sample tube 104 is being picked.

[0099] The methods and image capture and processing module 122 in accordance with one or more embodiments may be based on a convolutional neural network 105C, an example of which is shown in FIG. 4A, that “learns” how to turn an input image 401 into one or more heat maps 435. The convolutional neural network 105C includes a trained model 405TM configured to process the one or more input images 401 and generate the corresponding one or more heat maps 435. The one or more heat maps 435 may be stored in a non-transitory computer-readable storage medium, such as, e.g., memory 105M of FIG. 1 A. The trained model 405TM includes the network parameters used to process the images 401 input into the convolutional neural network 105C and may include kernel weights, kernel sizes, stride values, and padding values, each described in more detail below.

[0100] FIG. 4B illustrates one example configuration of a convolutional neural network (CNN 105C) in accordance with one or more embodiments. It should be understood, however, that other configurations of CNNs 105C may be used. The CNN 105C may be a fully convolutional network in some embodiments. In some embodiments, the convolutional network 105C may be optionally derived from a patch-based convolutional neural network by using a patch-based training procedure. The network configuration of the convolutional network 105C and the trained model 405TM may be stored in, e.g., memory 105M, and may be part of a method 500 performed by image capture & processing module 122 of the controller 105.

[0101] As shown in FIG. 4B, the convolutional neural network 105C may include a plurality of layers. Layers may include convolutional layers, pooling layers, activation layers (e.g., rectified linear activation (ReLu) function layers), and up-sampling layers. This CNN 105C may include a U-net architecture that is an encoder-decoder network. The CNN 105C is designed to take an input image 401 and produce an image as output. In particular, the network architecture is configured to generate one or more heat maps 435, each of which can be considered a one dimensional (1D) image. The contracting path contains convolutional layers (e.g., CONV 1 404), which extract features from the input, and maxpooling layers (e.g., Max Pool 406), which downsample the inputs. These two layers are successively applied in order to learn features at several different resolutions.

[0102] The decoder is the second part (e.g., including sections 403A, 403B, 403C) of the CNN 105C in which up-sampling (e.g., up-sample 1 410) and convolution layers (e.g., CONV4 412) are successively applied to recover the original resolution of the input image 401. After each up-sampling layer 410, features with the same resolution from earlier in the CNN 105C are concatenated (concatenations are labeled C1 through C3) to the newly up- sampled data to better learn the data representation. The output of the CNN 105C is an image with the same resolution as the input image 401 in which each pixel contains a probability indicating how likely it is that a keypoint 325 is detected there.

[0103] The activation (e.g., ReLu) layers in each case, operate to output the input directly if positive, and will output a zero if negative or zero. The max pooling layers operate to down-sample the input, which decreases the resolution of the input to the next layer by a scale factor. The max pooling layers are used aggregate information at different levels of resolution. Further, the convolutional neural network 105C may include a second section including up-sample layers (e.g., up-sample 1 410) that operate to attain the same resolution as the input image 401.

[0104] For example, the CNN 105C may comprise multiple sections, such as first section 402A and additional sections (e.g., section 2 402B, and section 3 402C). In this embodiment, the first section 402A comprises a first convolutional layer (e.g., CONV1 404), a first max pool layer 406, and a first activation layer (e.g., ReLui 408). The other one or more sections (e.g., Section 2 402B and Section 3402C) may be identical in structure to section 1 402A, and each may contain a convolutional layer, a max pool layer, and an activation (e.g., ReLu) layer. Other types of activation layers may be used.

[0105] The CNN 105C may comprise additional sections of a different type than Section 1 402A through Section 3 402C. For example, multiple additional sections (e.g., Section 4 403A, Section 5 403B, and section 6 403C) may be included. These sections (e.g., Section 4 403A, Section 5 403B, and Section 6 403C) may include, as best shown in forth section 403A, an up-sample layer (e.g., up-sample layer 410), a convolutional layer (e.g., convolutional layer 412), and a ReLu layer (e.g., ReLu layer 414). Section 5403B and Section 6 403C may be identical to Section 3403A. Note that the number of sections may be different in other embodiments.

[0106] The input to the first convolutional layer (e.g., CONV1 404) may be the one or more input images 401 , which may be one or more images of sample tubes 401 captured by, e.g., the one or more image capture devices 112M or 112F. For example, one or more images 401 may be images capable of displaying the tilt orientation of the one or more sample tubes 104 in three dimensions (3D). If the image capture device 112M is moving, theimages 401 can be made up of two or more image frames of different scenes where the tilt orientation can be observed in each scene or frame so as to be able to determine (i.e., through coordinate transformation) the tilt angle in the tray coordinate system. The one or more input images 401 may be, e.g., any one of the sample tube 104 images as shown respectively in FIG. 1A and FIGs. 3A-3E that contain X, Y, and Z information concerning the tilt angle and the positioning of the one or more sample tubes 104 in the holder 106T or 106C.

[0107] Input image 401 may be an arbitrarily-sized image, such as, e.g., 270 x 180 pixels, 540 x 360 pixels, 630 x 450 pixels, 900 x 720 pixels, or 1280 x 720 pixels, or the like. The one or more input images 401 may be received in any suitable file format, such as, e.g., JPEG, TIFF, or the like and, in some embodiments, may be color, grayscale, or black-and- white (monochrome). The one or more input images 401 may be received by the CNN 105C as an array of pixel values.

[0108] The first convolutional layer CONV1 402 may receive one or more input images401 as input and may generate one or more output activation maps (i.e., representations of input image 401). The first convolutional layer CONV1 402 may be considered a low level feature detector configured to detect, e.g., simple edges. First convolutional layer CONV1402 may generate one or more output activation maps based on kernel size, stride, and padding values. For example, a kernel size of 5, a stride of 1 , and a padding of 0 may be used. Other values may be used. The kernel, having its parameters stored in the trained model 405TM, may include an array of numbers (known as “weights”) representing a pixel structure configured to identify edges or contours of sample tube sides and cap 104C in the one or more input images 401. The kernel size may be thought of as the size of a filter applied to the array of pixel values of input image 401 . That is, the kernel size indicates the portion of the input image (known as the “receptive field”) in which the pixel values thereof are mathematically operated on (i.e., “filtered”) with the kernel’s array of numbers. The mathematical operation may include multiplying a pixel value of the input image 401 with a corresponding number in the kernel and then adding all the multiplied values together to arrive at a total. The total may indicate the presence in that receptive field of a desired feature (e.g., a portion of a side edge or contours of a cap 104C of a sample tube 104). The stride may be the number of pixels by which the kernel shifts in position to filter a next receptive field. The kernel continues to shift (“convolve”) around the input image 401 by the stride until the entire input image 401 has been filtered. Padding may refer to the number of rows and columns of zero pixel values to be added around the border of the output activation maps. In this first convolutional layer, the padding can be 0, meaning no rows and columns of zero pixel values are added to the border of the output activation maps. The output activation maps may thus include calculated pixel numbers representing pixel intensities inthe input image 401 based on the kernel size, weights, and original pixel values of the input image 401 .

[0109] The one or more activation maps generated by the first convolutional layer CONV1 402 may be applied to the first Max Pool 1 layer 406 that operate to decrease the resolution of the output by a scale factor and collect higher-level information from different regions of the image.

[0110] The activation map from the first Max Pool layer (e.g., Max Pool 1 406) may be applied to the first ReLu layer (e.g., ReLui 408). The first ReLu (e.g., ReLui 404) may generate output activation maps having maximum pixel values appearing in the one or more activation maps received from the first Max Pool layer (e.g., Max Pool 1 406), or zero if negative. That is, applying the ReLu, the maximum value of the calculated pixel values in each positive receptive field are included in output activation maps generated by the first ReLu 408 whereas the negative values are converted to zeros.

[0111] The output activation maps generated by the first section 402A may be input to a second section 402B, and so on. The number of sections may be 2 or more three or more, or even more. As more sections are added, the input image to each section becomes smaller and smaller, and that section can learn more general features about the image when compared to the finer details that are available at higher resolutions. As the network gets deeper, the accuracy will improve up to a point. The resolution of the input at each section will teach the model to learn different features. Thus, for example, the second section 402B may be configured to detect other features, e.g., more edge features (e.g., sides, tops and bottoms) than the first section 402A.

[0112] The activation maps generated by the first section 402A may be applied to a second section (e.g., second section 402B). The second section 402b may generate output activation maps that can be applied to a third section (e.g., third section 402C).

[0113] The output activation maps from the third section 402C may be applied to a fourth section 403A. In particular, the activation maps returned from the ReLu of section three 402C may be input to an up-sample layer (e.g., Up-Sample 1 410). The first Up-Sample layer (e.g., Up-Sample 1 410) operates to increase the resolution of the output by a scale factor. The activation maps returned from the up-sample layer (e.g., Up-Sample 1 410) may be then passed to a fourth convolution layer (e.g., CONV 4412) and then on to fourth ReLu layer (e.g., ReLu 4414). Additional Up-Sample layers as well as additional ReLu layers may be provided in Section 5403B and section 6 403C. The model structures of sections 5 403B and Section 6 403C may be identical to the structure shown in section four 403A.

[0114] Based on the various operations taking place, one or more heat maps 435 are generated, wherein each heat map 435 includes at least one determined location for one particular keypoint 325 (e.g., cap top 325TC and / or tube centroid 325C). Heat maps 435correspond to input image 401 and may be stored in a non-transitory computer-readable storage medium, such as, e.g., memory 105M of FIG. 1A. As should be recognized, some heat maps 435 may include more than one keypoints 325 per sample tube 104, such as both cap top 325TC and tube centroid 325C.

[0115] It should be understood that each heat map 435 may determine any one (or more) of the keypoints 325, as shown respectively in FIGs. 3A-3C. Each heat map 435 is representative of the location of at least one keypoint 325 on each of the one or more sample tubes 104 in the imaged scene.

[0116] The convolutional neural network 105C may, in some embodiments, be derived from a patch-based convolutional neural network in accordance with one or more embodiments. The patch-based convolutional neural network may include a plurality of layers like is shown in FIG. 4B. However, the input image 401 may be a fixed-sized image patch (which may be, e.g., 32 pixels x 32 pixels), which comprises a subset of the full input image. Thus, the fixed-sized image patch has fewer pixels than the full input image. The desired output may be a subpatch corresponding to a ground truth map. Ground truth map may be, e.g., 8 pixels x 8 pixels, derived from manual annotation, for example. The ground truth image may be a highly accurate image against which other images (e.g., from image processing methods) are compared.

[0117] To generate the consolidated heat map for a whole image as an input, a fixedsized (e.g., 32 pixels x 32 pixels) window may be used to raster scan the image and accumulate the response from each pixel thereof. In other embodiments, the convolutional network 105C can be derived from the patch-based convolutional neural network by using a CNN 105C similar to that shown in FIG. 4B. In this way, any arbitrarily-sized image may be input to the fully convolutional network 105C to directly generate a corresponding heat map 435 without using a scanning window.

[0118] FIG. 4E illustrates an example method 450 of training that may be performed, which may be optionally trained with the patch-based convolutional neural network. First, training images including keypoints 325 (e.g., centroid keypoint 325C, cap top keypoint 325TC, tube top keypoint 325T, bottom keypoint 325B, side keypoint 325S1 , 325S2, and / or the like) are input in training images and keypoint annotation block 452. Training images may be manually annotated to generate a tube image for each corresponding training image with the keypoints identified thereon. Positive and negative patch generation can optionally be provided in patch generation block 454. In patch generation block 454, patches, which may be of size 32x32 pixels, can be randomly extracted from each image where positive patches contain at least one corresponding keypoint 325 and negative patches contain no keypoints 325. The generated patches can be input into a patch-based convolutional neural network to train the patch-based keypoint detection model 456 that attempts to turn the 32 x32 pixel gray level image into an 8 x 8 pixel keypoint response heat map that matches the ground truth map as closely as possible. Then, fully-connected layers of the patch-based convolutional neural network and the trained patch-based keypoint detection model 456 are converted into the fully-convolutional layers of the fully convolutional network 105C via fullframe keypoint detection training in training block 558, as described above and the fully convolutional network 105C is refined with a full-size input gray-level image such that its output matches the keypoint heat map of the whole image as closely as possible. The refined convolutional neural network 105C with the trained full-frame keypoint detection model 560 (including the network parameters) can handle an arbitrary-sized image as the input and can estimate the corresponding keypoint response heat map 435 as the output.

[0119] For a patch-based convolutional neural network with the patch-based keypoint detection model 456, the input image may be uniformly divided into patches of size 32 x 32 pixels. Each patch may be input into the patch-based keypoint detection model 456 to estimate the keypoint response of the corresponding central patch of size 8 x 8 pixels. Then, using map fusion, the responses from multiple patches are fused by concatenating and normalizing the central patches to generate the estimated keypoint map for the whole test image. For a full-frame keypoint detection model 560, an arbitrary-sized test image can be provided as input, and the convolutional neural network 105C can directly estimate the corresponding keypoint response for the whole input image. Thus, using the full-frame keypoint detection model 560 may be much more efficient in terms of computation. The patch generation approach of blocks 454, 456 may be optional.

[0120] As should be recognized, the detected sample tube 104 may be used for tube characterization to estimate the tube height HT, Cap height HC, tube width W, and center offset dimension D from the slot 106S of the holder 106T or 106C, which are useful properties in an automated diagnostic analysis system.

[0121] The image processing and control method 500 will now be described with reference to FIG. 5. Image processing and control method 500 comprises, in block 502, capturing one or more images of the one or more sample tubes (e.g., sample tubes 104) residing in a holder (e.g., holder 106, such as tray holder 106T or carrier holder 106C). If a depth sensor is provided as part of the image capture device 112, then only one image is needed to obtain all of the orientation information, including the two or more keypoints 325 and the location of the one or more slots in the holder 106T or 106C. Imaging and processing may also include capturing one or more images of the one or more slots 106S in the holder 106T or 106C.

[0122] The image processing and control method 500 further comprises, in block 504, processing the one or more images of the one or more sample tubes 104 and information on one or more slots 106S of the holder 106 to: determine a location of at least two keypoints(e.g., keypoints 325) on each of the one or more sample tubes (e.g., one or more sample tubes 104), and determine a tilt angle orientation (e.g., tilt angle ©) relative to the holder (e.g., holder 106T or 106C) of the one or more sample tubes (e.g., one or more sample tubes 104) from the at least two keypoints (e.g., two or more keypoints 325). In particular, a tilt angle Q is determined based on the heat maps 435 generated by processing the one or more images through the CNN 105C. The method 500 may apply the image to a convolutional neural network 105C to determine the location of the at least two keypoints 325 in 3D space. The location of the one or more slots 106S in the holder 106T or 106C may also be determined from the convolutional neural network 105C. The location of each slots 106S may be otherwise determined, such as by imaging markers that have a known spatial relationship to the slots 106S.

[0123] The method 500 may further include, in block 506, controlling a robot (e.g., robot 110) based on the determined tilt angle orientation (e.g., tilt angle ©T) of the one or more sample tubes (e.g., one or multiple sample tubes 104). For example, the controlling of the robot 110 can involve movement of an end effector 114 based on the determined tilt angle orientation (e.g., tilt angle ©T). For example, based on determined tilt angle orientation (e.g., tilt angle ©T), the controller 105 (e.g., shown in FIG. 1A) may control the operation of robot 110 to make an adjustment the determined tilt angle orientation of the one or more sample tubes 104 in the holder 106T or 106C. The adjustment may be along a path directly opposing the tilt angle ©T. Thus, the adjustment in orientation can right the orientation to a substantially upright orientation in the holder 106T or 106C. By substantially upright, what is meant is anything that is adjusted to be in an acceptable state (acceptable degree of tilt angle ©Twhen ©T< LT). After the righting, the robot 110 and end effector 114 thereof may then grasp and move (e.g., pick) one of sample tubes 110 that is in an acceptable state and transfer and place the sample tube 104 in a destination slot 106S of a holder (tray 106T or carrier 106C). The degree of opening of the end effector 114 may be set based on the determined width W and the determined tilt angle ©T. Likewise, a corresponding offset of the end effector from the center of the slot 106S may be accomplished if the tube cap top keypoint 325TC or the tube top keypoint 325T (if no cap 104C) is offset in the holder 106T or 106C. As such, crashes between the sample tube 104 and the end effector 114 can be minimized or avoided. Moreover, placement may be improved as well by lowering the incidence of crashes during place operations occurring after the pick.

[0124] In some embodiments, a non-transitory computer-readable medium, such as, e.g., a removable storage disk or device, may include computer instructions capable of being executed in a processor, such as, e.g., processor 105P, and of performing the method 500. NON-LIMITING ILLUSTRATIVE EMBODIMENTS

[0125] The following is a list of non-limiting illustrative embodiments disclosed herein:

[0126] Illustrative embodiment 1 . An image processing and control apparatus, comprising: one or more image capture devices configured to capture one or more images of one or more sample tubes residing in a holder; and a controller comprising a processor and a memory, the controller configured via programming instructions stored in the memory to process the one or more images of the one or more sample tubes to: identify a location of at least two keypoints on each of the one or more sample tubes, and determine a tilt orientation of the one or more sample tubes from the at least two keypoints.

[0127] Illustrative embodiment 2. The image processing and control apparatus according to the preceding illustrative embodiment, wherein the at least two keypoints on each of the one or more sample tubes comprise two or more of: a center top of the one or more sample tubes appearing in the one or more images; a center top of a tube cap of the of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a tube cap; a location of a tube centroid of the one or more sample tubes appearing in the one or more images; a location of a cap centroid of a tube cap of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a cap; a location of at least two points along a side surface of the one or more sample tubes appearing in the one or more images; or a location of a tube bottom of the one or more sample tubes appearing in the one or more images.

[0128] Illustrative embodiment 3. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the image capture device is coupled to, and moveable with, a robot.

[0129] Illustrative embodiment 4. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the one or more image capture devices are configured to capture one or more side images of the one or more sample tubes.

[0130] Illustrative embodiment 5. The image processing and control apparatus according to one of the preceding illustrative embodiments, comprising a robot controller configured to control motion of an end effector of a robot, wherein the robot controller is configured to command the end effector to change the tilt orientation of the one or more sample tubes based on the tilt orientation determination.

[0131] Illustrative embodiment 6. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the robot controller is configured to cause the end effector to adjust the tilt orientation of all of the one or more sample tubes prior to a pick or place operation of any of the sample tubes.

[0132] Illustrative embodiment 7. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the robot controller is configured to cause the end effector to adjust the tilt orientation of the one or more sample tubes to a substantially vertical orientation.

[0133] Illustrative embodiment 8. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the robot controller is configured to cause a pick operation after the adjust of the tilt orientation.

[0134] Illustrative embodiment 9. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the determining of the tilt orientation of the one or more sample tubes is relative to a location of one or more slots of the holder.

[0135] Illustrative embodiment 10. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the holder comprises a sample tube tray comprising multiple slots configured to receive the one or more sample tubes therein.

[0136] Illustrative embodiment 11 . The image processing and control apparatus according to one of the preceding illustrative embodiments, comprising processing the one or more images of the one or more sample tubes by applying the one or more images to a convolutional neural network to determine the tilt orientation of the one or more sample tubes from an output of the convolutional neural network.

[0137] Illustrative embodiment 12. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the convolutional neural network is a fully convolutional network comprising a plurality of convolution layers.

[0138] Illustrative embodiment 13. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the fully convolutional network comprises one or more convolution layers and one or more max-pooling layers.

[0139] Illustrative embodiment 14. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the convolutional neural network is a patch-based convolutional network comprising a plurality of convolution layers followed by a fusion module that fuses keypoint responses of individual patches into one heat map representing an input image.

[0140] Illustrative embodiment 15. The image processing and control apparatus according to one of the preceding illustrative embodiments, wherein the patch-based convolutional network comprises one or more convolution layers and one or more maxpooling layers.

[0141] Illustrative embodiment 16. A non-transitory computer-readable medium, comprising: computer instructions of a convolutional neural network and parameters thereof capable of being executed in a processor and of applying the convolutional neural network and the parameters to one or more images of one or more sample tubes to generate one or more heat maps to be stored in the non-transitory computer-readable medium and accessible to a controller to control a robot based on the one or more heat maps, theconvolutional neural network comprising convolution layers, max pooling layers, up-sampling layers, and activation layers.

[0142] Illustrative embodiment 17. An image processing and control method, comprising: capturing one or more images of the one or more sample tubes residing in a holder; and processing the one or more images of the one or more sample tubes and information on one or more slots of the holder to: determine a location of at least two keypoints on each of the one or more sample tubes, and determine a tilt angle orientation relative to the holder of the one or more sample tubes from the at least two keypoints.

[0143] Illustrative embodiment 18. The image processing and control method according to the preceding illustrative embodiment, comprising capturing one or more images of the one or more slots of the holder.

[0144] Illustrative embodiment 19. The image processing and control method according to one of the preceding illustrative embodiments, wherein the at least two keypoints on each of the one or more sample tubes comprise: a location of a tube top of the one or more sample tubes appearing in the image; a location of a center top of the one or more sample tubes appearing in the one or more images; a location of a center top of a tube cap of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a tube cap; a location of a tube centroid of the one or more sample tubes appearing in the one or more images; a location of a cap centroid of a tube cap of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a cap; a location of at least two points along a side surface of the one or more sample tubes appearing in the one or more images; or a location of a tube bottom of the one or more sample tubes appearing in the one or more images.

[0145] Illustrative embodiment 20. The method according to one of the preceding illustrative embodiments, comprising processing the one or more images of the one or more sample tubes and one or more slots in the holder by applying the image to a convolutional neural network to determine a location of the at least two keypoints.

[0146] Illustrative embodiment 21 . The method according to one of the preceding illustrative embodiments, comprising controlling a robot based on the determined tilt angle orientation of the one or more sample tubes.

[0147] Illustrative embodiment 22. The method according to one of the preceding illustrative embodiments, wherein the controlling of the robot makes an adjustment of the determined tilt angle orientation of the one or more sample tubes in the holder.

[0148] Illustrative embodiment 23. The method according to one of the preceding illustrative embodiments, wherein the adjustment of the one or more sample tubes is to a substantially upright orientation in the holder.

[0149] Illustrative embodiment 24. The method according to one of the preceding illustrative embodiments, wherein the processing of the one or more images of the one or more sample tubes comprises generating one or more heat maps.

[0150] The foregoing description discloses only example embodiments of the invention; modifications of the above disclosed apparatus and methods that fall within the scope of the disclosure will be readily apparent to those of ordinary skill in the art. Accordingly, while the present invention has been disclosed in connection with the example embodiments thereof, it should be understood that other embodiments may fall within the scope of the disclosure, as defined by the claims.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. An image processing and control apparatus, comprising: one or more image capture devices configured to capture one or more images of one or more sample tubes residing in a holder; and a controller comprising a processor and a memory, the controller configured via programming instructions stored in the memory to process the one or more images of the one or more sample tubes to: identify a location of at least two keypoints on each of the one or more sample tubes, and determine a tilt orientation of the one or more sample tubes from the at least two keypoints.

2. The image processing and control apparatus of claim 1 , wherein the at least two keypoints on each of the one or more sample tubes comprise two or more of: a center top of the one or more sample tubes appearing in the one or more images; a center top of a tube cap of the of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a tube cap; a location of a tube centroid of the one or more sample tubes appearing in the one or more images; a location of a cap centroid of a tube cap of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a cap; a location of at least two points along a side surface of the one or more sample tubes appearing in the one or more images; or a location of a tube bottom of the one or more sample tubes appearing in the one or more images.

3. The image processing and control apparatus of claim 1 , wherein the image capture device is coupled to, and moveable with, a robot.

4. The image processing and control apparatus of claim 1 , wherein the one or more image capture devices are configured to capture one or more side images of the one or more sample tubes.

5. The image processing and control apparatus of claim 1 , comprising a robot controller configured to control motion of an end effector of a robot, wherein the robot controller is configured to command the end effector to change the tilt orientation of the one or more sample tubes based on the tilt orientation determination.

6. The image processing and control apparatus of claim 5, wherein the robot controller is configured to cause the end effector to adjust the tilt orientation of all of the one or more sample tubes prior to a pick or place operation of any of the sample tubes.

7. The image processing and control apparatus of claim 5, wherein the robot controller is configured to cause the end effector to adjust the tilt orientation of the one or more sample tubes to a substantially vertical orientation.

8. The image processing and control apparatus of claim 5, wherein the robot controller is configured to cause a pick operation after the adjust of the tilt orientation.

9. The image processing and control apparatus of claim 1 , wherein the determining of the tilt orientation of the one or more sample tubes is relative to a location of one or more slots of the holder.

10. The image processing and control apparatus of claim 1 , wherein the holder comprises a sample tube tray comprising multiple slots configured to receive the one or more sample tubes therein.11 . The image processing and control apparatus of claim 1 , comprising processing the one or more images of the one or more sample tubes by applying the one or more images to a convolutional neural network to determine the tilt orientation of the one or more sample tubes from an output of the convolutional neural network.

12. The image processing and control apparatus of claim 11 , wherein the convolutional neural network is a fully convolutional network comprising a plurality of convolution layers.

13. The image processing and control apparatus of claim 12, wherein the fully convolutional network comprises one or more convolution layers and one or more maxpooling layers.

14. The image processing and control apparatus of claim 1 , wherein the convolutional neural network is a patch-based convolutional network comprising a plurality of convolution layers followed by a fusion module that fuses keypoint responses of individual patches into one heat map representing an input image.

15. The image processing and control apparatus of claim 14, wherein the patch-based convolutional network comprises one or more convolution layers and one or more maxpooling layers.

16. A non-transitory computer-readable medium, comprising: computer instructions of a convolutional neural network and parameters thereof capable of being executed in a processor and of applying the convolutional neural network and the parameters to one or more images of one or more sample tubes to generate one or more heat maps to be stored in the non-transitory computer-readable medium and accessible to a controller to control a robot based on the one or more heat maps, the convolutional neural network comprising convolution layers, max pooling layers, up-sampling layers, and activation layers.

17. An image processing and control method, comprising: capturing one or more images of the one or more sample tubes residing in a holder; and processing the one or more images of the one or more sample tubes and information on one or more slots of the holder to: determine a location of at least two keypoints on each of the one or more sample tubes, and determine a tilt angle orientation relative to the holder of the one or more sample tubes from the at least two keypoints.

18. The image processing and control method of claim 17, comprising capturing one or more images of the one or more slots of the holder.

19. The image processing and control method of claim 17, wherein the at least two keypoints on each of the one or more sample tubes comprise: a location of a tube top of the one or more sample tubes appearing in the image; a location of a center top of the one or more sample tubes appearing in the one or more images;a location of a center top of a tube cap of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a tube cap; a location of a tube centroid of the one or more sample tubes appearing in the one or more images; a location of a cap centroid of a tube cap of the one or more sample tubes appearing in the one or more images, if the one or more sample tubes includes a cap; a location of at least two points along a side surface of the one or more sample tubes appearing in the one or more images; or a location of a tube bottom of the one or more sample tubes appearing in the one or more images.

20. The method of claim 17, comprising processing the one or more images of the one or more sample tubes and one or more slots in the holder by applying the image to a convolutional neural network to determine a location of the at least two keypoints.21 . The method of claim 17, comprising controlling a robot based on the determined tilt angle orientation of the one or more sample tubes.

22. The method of claim 21 , wherein the controlling of the robot makes an adjustment of the determined tilt angle orientation of the one or more sample tubes in the holder.

23. The method of claim 22, wherein the adjustment of the one or more sample tubes is to a substantially upright orientation in the holder.

24. The method of claim 17, wherein the processing of the one or more images of the one or more sample tubes comprises generating one or more heat maps.

Citation Information

Patent Citations

  • Methods and apparatus for dynamic position adjustments of a robot gripper based on sample rack imaging data

    US20190160666A1

  • Methods and systems for learning-based image edge enhancement of sample tube top circles

    US20200167591A1

  • Automated foam detection

    US20230357698A1

  • Sample handlers of diagnostic laboratory analyzers and methods of use

    WO2023225123A1