Multi-camera image processing
By generating and merging 3D views using a multi-camera system, the problem of a single camera being unable to fully capture objects in the workspace is solved, thus improving the operational efficiency and accuracy of the robot system.
Patent Information
- Application Number
- CN201980092389.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2019-12-03
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2039-12-03
AI Technical Summary
When robotic systems are dealing with complex workspaces, a single camera may struggle to generate complete image data for all objects, leading to operational failures or inefficiencies.
Multiple cameras are used to generate 3D views, and image data merging and calibration techniques are used to dynamically update the complete view of the workspace, ensuring that objects are fully captured and manipulated.
It enables comprehensive capture and manipulation of objects in the workspace, improving the grabbing and placement efficiency of the robot system and reducing the need for human intervention.
Smart Images

Figure CN113574563B_ABST
Abstract
Description
[0001] Cross-references to other applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 809,389, filed February 22, 2019, entitled ROBOTIC MULTI-ITEM TYPE PALLETIZING & DEPALLETIZING, which is incorporated herein by reference for all purposes.
[0003] This application is a continuation portion of co-pending U.S. Patent Application No. 16 / 380,859, filed April 10, 2019, entitled ROBOTIC MULTI-ITEM TYPE PALLETIZING & DEPALLETIZING, which is incorporated herein by reference for all purposes. This U.S. Patent Application claims priority to U.S. Provisional Patent Application No. 62 / 809,389, filed February 22, 2019, entitled ROBOTIC MULTI-ITEM TYPE PALLETIZING & DEPALLETIZING, which is incorporated herein by reference for all purposes. Background Technology
[0004] For example, robots are used in many environments to pick up, move, manipulate, and place objects. In order to perform tasks in a physical environment (sometimes referred to as the “workspace” in this paper), robotic systems typically use cameras and other sensors to detect objects to be manipulated by the robotic system, such as items to be picked up and placed using a robotic arm, and generate and execute a plan to manipulate the objects, such as grasping one or more objects in the environment and moving such objects(s) to a new location within the workspace.
[0005] Sensors may include multiple cameras, one or more of which may be three-dimensional (“3D”) cameras that generate conventional (e.g., red-blue-green or “RBG”) image data, as well as “depth pixels” indicating the distance to points in the image. However, due to factors such as occlusion of objects or parts thereof, a single camera may not be able to generate image data and / or complete 3D image data for all objects in the workspace. For successful operation, the robotic system must be able to respond to changing conditions and must be able to plan and execute operations within an operationally meaningful timeframe. Attached Figure Description
[0006] Various embodiments of the invention are disclosed in the following detailed description and accompanying drawings.
[0007] Figure 1 This is a diagram illustrating an embodiment of the robot system.
[0008] Figure 2 This is a flowchart illustrating an embodiment of a process for performing robot operations using image data from multiple cameras.
[0009] Figure 3 This is a flowchart illustrating an embodiment of a process for performing robot operations using segmented image data.
[0010] Figure 4 This is a flowchart illustrating an embodiment of a process for calibrating multiple cameras deployed in a workspace.
[0011] Figure 5 This is a flowchart illustrating an embodiment of a process for performing object instance segmentation processing on image data from a workspace.
[0012] Figure 6 This is a flowchart illustrating an embodiment of the process for maintaining camera calibration in a workspace.
[0013] Figure 7 This is a flowchart illustrating an embodiment of the process for recalibrating a camera in a workspace.
[0014] Figure 8 This is a flowchart illustrating an embodiment of a process for using image data from multiple cameras to perform robot operations in a workspace and / or provide visualization of the workspace.
[0015] Figure 9A This is a diagram illustrating an embodiment of a multi-camera image processing system for robot control.
[0016] Figure 9B This is a diagram illustrating an embodiment of a multi-camera image processing system for robot control.
[0017] Figure 9C This is a diagram illustrating an embodiment of a multi-camera image processing system for robot control.
[0018] Figure 10 This is a diagram illustrating an example of a visual display generated and provided in an embodiment of a multi-camera image processing system.
[0019] Figure 11 This is a flowchart illustrating an embodiment of a process for generating code for processing sensor data in a multi-camera image processing system for robot control. Detailed Implementation
[0020] This invention can be implemented in a variety of ways, including as a process; an apparatus; a system; a composition of matter; a computer program product implemented on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by memory coupled to the processor. In this specification, these implementations or any other form of the invention may be referred to as technology. Generally, the order of steps of the disclosed process can be varied within the scope of this invention. Unless otherwise stated, components such as processors or memory described as being configured to perform a task can be implemented as general-purpose components temporarily configured to perform a task at a given time or manufactured as specific components to perform a task. As used herein, the term 'processor' refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0021] The following provides a detailed description of one or more embodiments of the present invention, along with accompanying drawings illustrating the principles of the invention. The invention has been described in conjunction with such embodiments, but is not limited to any particular embodiment. The scope of the invention is limited only by the claims, and the invention includes many alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description to provide a thorough understanding of the invention. These details are provided for illustrative purposes, and the invention may be practiced according to the claims without requiring some or all of these specific details. For clarity, technical materials known in the art related to the invention have not been described in detail, so as not to unnecessarily obscure the invention.
[0022] Techniques are disclosed for generating a three-dimensional view of a workspace using a set of sensors, including multiple cameras or other image sensors. In some embodiments, the three-dimensional view is used to programmatically perform work in the workspace using a robotic system including one or more robots (e.g., a robotic arm with a suction, gripper, and / or other end effector at the end of operation), such as palletizing / depalletizing and / or otherwise packaging and / or unpacking any group of non-homogeneous items (e.g., different sizes, shapes, weights, weight distributions, rigidity, fragility, etc.).
[0023] In various embodiments, 3D cameras, force sensors, and other sensors are used to detect and determine the properties of items to be picked up and / or placed and / or programmatically generate a plan for grasping one or more items at an initial location and moving each of the items to a corresponding destination location within the workspace. Items whose type is determined (e.g., with sufficient confidence, as indicated by a programmatically determined confidence score) can be grasped and placed using strategies derived from an item type-specific model. Items that cannot be identified can be picked up and placed using strategies not specific to a given item type. For example, a model using size, shape, and weight information can be used.
[0024] In some embodiments, the techniques disclosed herein can be used to generate and display a visual representation of at least a portion of a workspace. In various embodiments, the visual representation can be displayed via a computer or other display device, including a workstation used by a human operator to monitor a robot operating in a fully or partially automated mode and / or to control a robotic arm or other robotic actuator via teleoperation.
[0025] For example, in some embodiments, if the robotic system gets stuck, such as being unable to perform or complete the next task or operation within configured parameters (e.g., timeout, confidence score, etc.), human intervention can be invoked. In some embodiments, displayed images and / or videos of the workspace can be used to perform teleoperation. A human operator can manually control the robot, using the displayed images or videos to view the workspace and control the robot. In some embodiments, the display can be incorporated into an interactive, partially automated system. For example, a human operator can indicate the point in the displayed image of the scene where the robot should grasp an object via the display.
[0026] Figure 1This is a diagram illustrating an embodiment of a robotic system. In the example shown, the robotic system 100 includes a robotic arm 102 and is configured to palletize and / or depalletize heterogeneous items. In this example, the robotic arm 102 is stationary, but in various alternative embodiments, the robotic arm 102 may be fully or partially mobile, for example, mounted on a track, fully mobile on a motorized chassis, etc. As shown, the robotic arm 102 is used to pick up arbitrary and / or different items from a conveyor belt (or other source) 104 and stack them on a pallet or other container 106. In the example shown, the container 106 includes a pallet or base having wheels at the four corners and being at least partially closed on three of the four sides, sometimes referred to as a three-sided “roller pallet,” “roller cage,” and / or “roller” or “cage” “trolley.” In other embodiments, roller or wheelless pallets with more, fewer, and / or no sides may be used. In some embodiments, Figure 1 Other robots, not shown, may be used to push container 106 into the location to be loaded / unloaded and / or into a truck or other destination to be transported.
[0027] In the example shown, the robotic arm 102 is equipped with a suction-type end effector 108. The end effector 108 has multiple suction cups 110. The robotic arm 102 is used to position the suction cups 110 of the end effector 108 above an item to be picked up, as shown, and a vacuum source provides suction to grasp the item, lift it from the conveyor 104, and place it at its destination location on the container 106.
[0028] In various embodiments, one or more of the cameras 112 mounted on the end effector 108 and cameras 114, 116 mounted in the space in which the robot system 100 is deployed are used to generate image data for identifying items on the conveyor 104 and / or determining a plan for gripping, picking up / placing and stacking items on the container 106. In various embodiments, additional sensors, not shown, such as weight or force sensors implemented in and / or adjacent to conveyor 104 and / or robot arm 102, force sensors in the xy plane and / or z direction (vertical direction) of suction cup 110, etc., can be used to identify, determine their attributes, grasp, pick up, move through a defined trajectory, and / or place items on conveyor 104 and / or other source and / or staging areas onto container 106 or at a destination location within container 106, where items can be positioned and / or repositioned, for example, by system 100, in conveyor 104 and / or other source and / or staging areas.
[0029] In the example shown, camera 112 is mounted on the side of the body of end effector 108. However, in some embodiments, camera 112 and / or additional cameras may be mounted in other locations, such as on the underside of the body of end effector 108, for example pointing downwards from between suction cups 110, or mounted on segments or other structures or locations of robotic arm 102. In various embodiments, cameras such as 112, 114, and 116 may be used to read text, logos, photographs, drawings, images, markings, barcodes, QR codes, or other coded and / or graphic information, or content visible on and / or including items on conveyor 104.
[0030] Further reference Figure 1 In the example shown, system 100 includes a control computer 118, which is configured to communicate wirelessly (but in various embodiments, with one or both wired and wireless communication) to elements such as a robotic arm 102, a conveyor 104, an actuator 108, and cameras such as cameras 112, 114, and 116 and / or weight, force, and / or pressure. Figure 1 Sensor communication with other sensors not shown. In various embodiments, control computer 118 is configured to use sensors such as cameras 112, 114 and 116 and / or weight, force and / or Figure 1 Input from other sensors (not shown) is used to view, identify, and determine one or more attributes of items to be loaded into and / or unloaded from container 106. In various embodiments, control computer 118 uses item model data stored on and / or in a library accessible to control computer 118 to identify items and / or their attributes, for example, based on images and / or other sensor data. Control computer 118 uses the model corresponding to the item to determine and implement a plan for stacking the item with other items in / on a destination (such as container 106). In various embodiments, item attributes and / or the model are used to determine strategies for grasping, moving, and placing items at the destination location (e.g., determining a specific location thereon as part of a planning / replanning process for stacking items in / on container 106).
[0031] In the example shown, control computer 118 is connected to “on-demand” teleoperation device 122. In some embodiments, if control computer 118 is unable to operate in fully automated mode, for example, if it cannot determine a strategy for grasping, moving, and placing items and / or fails in a way that prevents control computer 118 from having a strategy for completing the picking and placing of items in fully automated mode, control computer 118 prompts human user 124 to intervene, for example, by using teleoperation device 122 to operate robotic arm 102 and / or end effector 108 to grasp, move, and place items.
[0032] In various embodiments, control computer 118 is configured to receive and process image data (e.g., two-dimensional RGB or other image data, consecutive frames including video data, point cloud data generated by 3D sensors, consecutive point cloud datasets, each associated with a corresponding frame of 2D image data, etc.). In some embodiments, control computer 118 receives aggregated and / or merged image data generated by and received from separate computers, applications, services, etc., based on image data generated by and received from cameras 112, 114, and 116 and / or other sensors (such as laser sensors and other light, heat, radar, sonar, or other sensors), which use projection, reflection, radiation, and / or other electromagnetic radiation and / or signals received in other ways to detect and / or transmit information for creating an image or for use in creating an image. As used herein, an image includes a visual and / or computer- or other machine-perceptible representation, depiction, etc., of objects and / or features existing in a physical space or scene, such as in Figure 1 In the example shown, the robot system 100 is positioned within the workspace.
[0033] In various embodiments, image data generated and provided by cameras 112, 114, and / or 116 and / or other sensors is processed and used to generate a three-dimensional view of at least a portion of the workspace in which the robotic system 100 is positioned. In some embodiments, image data from multiple cameras (e.g., 112, 114, 116) are merged to generate a three-dimensional view of the workspace. The merged imaging data is segmented to determine the boundaries of objects of interest within the workspace. The segmented image data is used to perform tasks such as determining strategies or plans through automated processing to accomplish one or more tasks, including grasping objects within the workspace, moving objects through the workspace, and placing objects at destination locations.
[0034] In various embodiments, 3D point cloud data views generated by multiple cameras (e.g., cameras 112, 114, 116) are merged into a complete model or view of the workspace via a process called registration. The corresponding positions and orientations of objects and features in the workspace captured in separately acquired views are converted into a global 3D coordinate frame such that their intersecting regions overlap as perfectly as possible. For each set of point cloud datasets acquired from different cameras or other sensors (i.e., different views), in various embodiments, the system aligns them together into a single point cloud model as disclosed herein, allowing subsequent processing steps such as segmentation and object reconstruction to be applied.
[0035] In various embodiments, at least in part, a three-dimensional view of the workspace is generated using image data generated and provided by cameras 112, 114, and / or 116, and by merging data through cross-calibrating cameras such as 112, 114, and / or 116, to generate views of the workspace and the items / objects present within it from as many available angles and views as possible. For example, in Figure 1 In the example shown, cameras 112 and 116 may be positioned to view objects on conveyor 104, while camera 114, shown pointing at container 106 in the example, may (currently) not have any image data from the portion of the workspace in which conveyor 104 is positioned. Similarly, arm 102 may be moved to a position where camera 112 no longer has a view of conveyor 104. In various embodiments, image data (e.g., RGB pixels, depth pixels, etc.) from cameras in the workspace are combined to dynamically generate and continuously update a three-dimensional view of the workspace that is as complete and accurate as possible given the image data received from the cameras and / or other sensors at any given moment. If a camera obstructs its view of an object or area in the workspace, or if a camera is moved or pointed in a different direction, image data from those cameras that continue to have a line of sight to the affected object or area will continue to be used to generate the most complete and accurate view of the object or area possible.
[0036] In various embodiments, the techniques disclosed herein enable the use of multiple image data from cameras to generate and maintain a more complete view of the workspace and objects within it. For example, using multiple cameras at different locations and / or orientations within the workspace, a smaller object that may be occluded from one angle by a larger object can be visible via image data, with one or more cameras positioned to view the object from a vantage point where it is not occluded. Similarly, objects can be viewed from multiple angles, making it possible to identify all unoccluded sides and features of the object, thereby facilitating operations such as determining and implementing grasping strategies, determining to place items snugly adjacent to objects, and maintaining a view of the object as human workers or robotic actuators (e.g., robotic arms, conveyors, robot-controlled movable shelves, etc.) move through the workspace.
[0037] In some embodiments, segmented image (e.g., video) data is used to generate and display a visualization of the workspace. In some embodiments, objects of interest may be prominently displayed in the displayed visualization. For example, colored boundary shapes or outlines may be displayed. In some embodiments, a human-operable interface is provided to enable a human operator to correct, refine, or otherwise provide feedback on the automatically generated boundaries of the objects of interest. For example, an interface may be provided to enable a user to move or adjust the position of the automatically generated boundary shapes or outlines, or to indicate that the prominently displayed area actually includes two (or more) objects, rather than one. In some embodiments, the displayed visualization may be used to enable a human operator to control a robot in the workspace in a teleoperated mode. For example, a human operator may use the segmented video to move a robotic arm (or other actuator) into position, grasp a prominently displayed object (e.g., from conveyor 104), and move the prominently displayed object to a destination location (e.g., on container 106).
[0038] In various embodiments, in order to enable the merging of image data from multiple cameras to perform the tasks disclosed herein, at least the main camera or calibration reference camera is calibrated relative to a calibration pattern, object, or other reference having a stationary and / or otherwise known position, orientation, etc. Figure 1 In the examples shown, for instance, one or more of cameras 112, 114, and 116 can be calibrated based on one or more images of marker 130 mounted in a static location in the workspace (e.g., on a wall at a known location) and / or marker 132 mounted (printed, etc.) on the robotic arm 102. In various embodiments, images generated by the calibrated master camera are used to cross-calibrate one or more of the other cameras in the workspace.
[0039] In some embodiments, processing is performed to detect the need for recalibrating and / or cross-calibrating the camera, for example, due to camera error, camera collision, or intentional repositioning or reorientation; failure of operations based on image data to indicate camera error or misalignment; the system detecting, based on image data from one camera, that the position, orientation, etc., of another camera differs from expectations; etc. In various embodiments, system 100 (e.g., control computer 118) is configured to automatically detect the need for recalibrating one or more cameras and to automatically and dynamically recalibrate them as disclosed herein. For example, in various embodiments, recalibration is performed by one or more of the following: repositioning a reference marker (e.g., marker 130) in the workspace using a camera mounted on a robot actuator (e.g., camera 112); re-estimating the camera-to-workspace translation using the reference marker; and recalibrating a marker on the robot (e.g., marker 132).
[0040] Figure 2 This is a flowchart illustrating an embodiment of a process for performing robot operations using image data from multiple cameras. In various embodiments, Figure 2 Process 200 is performed by a computer or other processor (such as...) Figure 1 The control computer 118 performs the operation. In the example shown, image data (202) is received from multiple cameras positioned to capture video (e.g., 3D video including RGB and depth pixels) in the workspace. The received image data is processed and merged to generate a three-dimensional view of the workspace (also called a “scene”), which is segmented to distinguish objects in the workspace (204). The segmented video is used to perform operations (206) on one or more of the objects, such as grabbing the object and moving it to a new location in the workspace.
[0041] Figure 3 This is a flowchart illustrating an embodiment of a process for performing robotic operations using segmented image data. In various embodiments, Figure 3 The process was used to achieve Figure 2 Step 206 of the process. In the example shown, segmented video data is used to determine and implement strategies for grabbing, moving, and placing one or more objects (302). The segmented video and associated bounding boxes (or other shapes) are used to generate and display a visualization of the workspace (304). For example, in some embodiments, the segmented data is used to generate one or more mask layers to overlay a semi-transparent colored shape conforming to (as closely as possible to) the outline of the object on each of at least a subset of objects in the workspace.
[0042] In various embodiments, visualization can be used by a human operator to monitor the operation of the robotic system in autonomous mode and / or to operate the robotic arm or other robotic actuators via teleoperation.
[0043] Figure 4 This is a flowchart illustrating an embodiment of a process for calibrating multiple cameras deployed in a workspace. In various embodiments, Figure 4 The process 400 is performed by a computer (such as...) Figure 1 The control computer 118 is configured to control and process data from multiple cameras or other sensors (such as...) in the workspace. Figure 1 The data received by cameras 112, 114, and 116. In the example shown, a calibration reference (402) is obtained by generating images of reference markers or other references in the workspace using one or more cameras. For example, in Figure 1 In the example shown, one or more of cameras 112, 114, and 116 can be used to generate one or more images of marker 130 and / or marker 132. In some embodiments, such as by inserting a key or other object or accessory into a corresponding hole or other receiver and generating an image at a known location and orientation, a robotic arm or other actuator can be moved to a known fixed location.
[0044] Further reference Figure 4 The calibration reference is used to cross-calibrate all cameras in the workspace (404). At runtime, iterative closest point (ICP) processing is performed to merge point clouds from multiple cameras (406). Instance segmentation processing is performed to identify, label (e.g., by type, etc.) and tag objects in the workspace (408).
[0045] Figure 5 This is a flowchart illustrating an embodiment of a process for performing object instance segmentation processing on image data from a workspace. In various embodiments, Figure 5 The process was used to achieve Figure 4Step 408 of the process. In the example shown, segmentation processing (502) is performed on RGB (2D) image data from fewer than all cameras in the workspace. In some embodiments, RGB data from one camera is used to perform segmentation. RGB pixels identified as associated with object boundaries in the segmentation process are mapped to corresponding depth pixels (504). The segmented data and mapped depth pixel information are used to deproject to a point cloud with segmentation boxes (or other shapes) surrounding the points for the object (506). For image data from each camera, the point cloud for each object is labeled and the centroid is calculated (508). Nearest neighbor computation is performed between the centroids of the corresponding object point clouds from the corresponding cameras to segment the object (510).
[0046] Figure 6 This is a flowchart illustrating an embodiment of a process for maintaining camera calibration in a workspace. In various embodiments, Figure 6 The process 600 is performed by a computer or other processor (such as...) Figure 1 The control computer 118 performs the recalibration. In the example shown, the need to recalibrate one or more cameras in the workspace is detected (602). In various embodiments, one or more of the following may indicate the need for recalibration: a camera sees that the robot base position has moved (e.g., based on an image of an Aruco or other marker on the base); a camera on the robot arm or other actuator sees that a camera mounted in the workspace has moved (e.g., intentionally bumped, repositioned, etc. by a human or robot worker); and the system detects several (or more than a threshold number) missed picks in a row. The recalibration is performed dynamically, e.g., in real time, without interrupting pick-up and placement or other robot operations, and without human intervention (604). In various embodiments, recalibration may include one or more of the following: repositioning a reference marker (e.g., marker 130) in the workspace using a camera mounted on a robot actuator (e.g., camera 112); re-estimating the camera-to-workspace translation using the reference marker; and recalibrating to a marker on the robot (e.g., marker 132).
[0047] Figure 7 This is a flowchart illustrating an embodiment of a process for recalibrating a camera in a workspace. In various embodiments, Figure 7 The process was used to achieve Figure 6 Step 604 of the process. In the example shown, the operation of the robot system is paused—such as for picking up and placing a specified set of items to a desired destination or a set of destinations (such as...). Figure 1The operation (702) involves the tray or other container shown in the example. A reference camera is determined, and if not already specified or otherwise established, one or more reference images are generated (704). For example, a stationary camera, a camera mounted on a robotic arm, etc., can be designated as a reference camera. Alternatively, a camera included in a group that appears to be synchronized with one or more other cameras can be selected as a reference camera for recalibrating other cameras to it. The reference images can include images from the reference camera and one or more other cameras, including reference markers or other references in the workspace. The reference images are used for cross-calibrating the cameras (706).
[0048] Figure 8 This is a flowchart illustrating an embodiment of a process for using image data from multiple cameras to perform robotic operations in a workspace and / or provide visualization of the workspace. In various embodiments, Figure 8 The process 800 is performed by a computer or other processor (such as...) Figure 1The control computer 118 executes the operation. In the example shown, image data is received from multiple cameras in the workspace, and a workspace filter (802) is applied. In various embodiments, the workspace filter may remove image and / or point cloud data associated with portions of the workspace, features of the workspace, items in the workspace, etc., and may ignore the image and / or point cloud data for the purpose of the robot operation performed using the image / sensor data. Filtering out irrelevant information allows for a clearer and / or more focused view of the elements to be operated on in the workspace and / or about the workspace, such as objects to be grasped and placed, trays or other destinations for items to be grasped or placed therefrom, obstacles that may be encountered in moving items to their destination, etc. In some embodiments, statistical outliers may be removed by the workspace filter to clear noise from the sensors. Point cloud data from the respective cameras are merged (804). Segmentation (806) is performed using RGB image data from one or more of the cameras. For example, RGB segmentation may initially be performed on image data from only one camera. The point cloud data is subsampled and clustering is performed (808). Point cloud data from secondary sampling and clustering, along with RGB segmentation results, are used to perform "box (or other 3D geometric primitive) fitting" processing (810). Stable object matching is performed (812). In various embodiments, multiple 3D representations of the same object can be generated from 2D methods using different cameras. These representations may only partially overlap. In various embodiments, "stable object matching" includes reconcile which segments correspond to the same object and merging representations. In various embodiments, spatial, geometric, curvature properties, or features of the point cloud (including RGB data about each point) can be used to perform stable object matching. In some embodiments, stable object matching can be performed over time, across multiple frames from the same camera rather than two different cameras, etc. The processed image data is used to perform grasping synthesis (814). For example, object location and boundary information can be used to determine a strategy for grasping the object using a robotic gripper or other end effector. The processed image data is used to generate and display a visualization of the workspace (816). For example, the raw video of the workspace can be displayed using colors or other striking displays applied to the bounding boxes (or other geometric primitives) associated with different objects in the workspace.
[0049] Techniques for configuring a robotic system to process and integrate sensor data from multiple sensors, such as multiple 3D cameras, to perform robotic operations are disclosed. In various embodiments, management user interfaces, configuration files, application programming interfaces (APIs), or other interfaces can be used to identify sensors and define one or more processing pipelines to process and use sensor outputs to perform robotic operations. In various embodiments, a pipeline can be defined by identifying processing modules and how the corresponding inputs and outputs of such modules should be linked to form a processing pipeline. In some embodiments, this definition is used to generate binary code to receive, process, and use sensor inputs to perform robotic operations. In other embodiments, this definition is used by a single generic binary code that dynamically loads plugins to perform processing.
[0050] Figure 9A , 9B Examples of pipelines using the technical configurations and implementations disclosed herein are shown in various embodiments, respectively.
[0051] Figure 9A This is a diagram illustrating an embodiment of a multi-camera image processing system for robot control. In the example shown, system 900 includes multiple sensors, such as 3D or other cameras, which... Figure 9A The data is represented by sensor nodes 902 and 904. The sensor outputs of nodes 902 and 904 are processed by workspace filters 906 and 908, respectively. The resulting point cloud data (the 3D portion of the filtered sensor output) is merged 910. The merged point cloud data 910 is then resampled 912 and clustered 914, and in this example, the resulting object clustering data is used by the "target estimation" module / process 915 to estimate the position and orientation of the "target" object (such as an item present in the workspace that will be moved to its tray or plate).
[0052] exist Figure 9A In the example shown, segmentation processes 916 and 922 are performed on the RGB data from sensor nodes 902 and 904, and the corresponding segmentation results are used to perform "box fitting" processes 918 and 924 to determine the bounding boxes (or other polyhedra) of the object instances identified in the RGB segmentation. The box fitting results 918 and 924 are merged 925 and used to perform stable object matching 926, the results of which are used to perform grasp synthesis 927 and generate and display a visualization of the workspace 928.
[0053] Although Figure 9AIn the example shown, only RGB segmentation information is merged 925 and used to perform stable object matching 926, perform grasp synthesis 927, and generate visualization of objects in the workspace 928. However, in some alternative embodiments, the results of merging point cloud data 910 and secondary sampling 912, as well as object clustering based on the merged point cloud data 914, can be merged together with the results of RGB segmentation and box fitting (916, 918, 922, 924) 925 to perform stable object matching 926, perform grasp synthesis 927, and generate visualization of objects in the workspace 928.
[0054] In various embodiments, the pipeline can be predefined and / or dynamically adapted in real time based on conditions. For example, if the objects in the workspace are significantly cluttered, RGB segmentation results may be a better signal than a 3D clustering process, and in some embodiments, under such conditions, only bounding box (polygon) fitting can be applied to the RGB segmentation output. Under other conditions, when performing geometric primitive fitting, etc., two sources of segmentation (RGB, point cloud data) can be applied.
[0055] Although Figure 9A In the example shown, the “box fitting” results 918, 924 based on RGB segmentation 916, 922 of data from all sensors (e.g., cameras) 902, 904 are merged 925 and used to perform downstream tasks such as stable object matching 926, capture synthesis 927, and visualization 928. However, in various embodiments, RGB segmentation and / or box fitting of data from fewer than all sensors may be merged and used to perform one or more of the downstream tasks. In some embodiments, it may be dynamically determined that the quality and / or content of data from a given sensor is unreliable and / or less useful for performing a given downstream task than data from other sensors, and that data from that sensor may be omitted (i.e., discarded and not used) when performing that task.
[0056] In some embodiments, a pipeline can be defined as omitting a given sensor from a given pipeline path and / or task. For example, a user defining a pipeline as disclosed herein can determine, based on the sensor's capabilities, quality, reliability, and / or location, that a sensor may be useful for some tasks but not for others, and can define a pipeline to use the sensor's output only for those tasks for which the sensor is deemed appropriate and / or useful.
[0057] In various embodiments, such as Figure 9AThe pipeline 900 can be modified at any time, for example, to modify how certain sensors are used and / or to add or remove sensors to the group of sensors 902, 904. In various embodiments, new or updated pipelines are defined, and code for implementing the pipeline is generated and deployed as disclosed herein, so that new pipelines, added sensors, etc., can be deployed and used to perform subsequent robotic operations.
[0058] Figure 9B This is a diagram illustrating an embodiment of a multi-camera image processing system for robot control. Figure 9B In the example processing pipeline 940 shown, the outputs of sensors 902 and 904, such as camera frame data, are processed by workspace filters 906 and 908 to generate filtered sensor outputs, the 3D point cloud data portion of which is merged 910 to generate merged point cloud data for the workspace. In this example, the merged point cloud data is provided to three separate modules: a pre-visualization module 942, which performs processing to enhance the information present in the merged point cloud data; a "voxel" processing module 944, which identifies the 3D space in the workspace occupied by and / or not occupied by objects of interest; and a visualization module 946, which generates and provides visualizations based on the merged point cloud data and the outputs of the pre-visualization 942 and voxel 944 modules, for example, via a display device. In various embodiments, the pre-visualization module 942 reformatts the data from the rest of the processing pipeline to allow efficient rendering, enabling highly interactive visualizations where users can smoothly pan, zoom, rotate, etc.
[0059] Figure 9C This diagram illustrates an embodiment of a multi-camera image processing system for robot control. In the example shown, pipeline 960 includes a single sensor node 962 whose output (e.g., 3D or other camera frame data) is processed by a workspace filter 964 to provide filtered data to a clustering module or process 966, a pre-visualization module 970, and a visualization module 972. Furthermore, the output of clustering module 966 is provided to visualization module 972 and grasping synthesis module 968, whose output is subsequently provided to pre-visualization module 970 and visualization module 972. Figure 9C The example shown illustrates how modules can be easily chained together to define a processing pipeline, as disclosed herein, where the output of intermediate modules propagates along multiple paths to ensure that each processing module has the information needed to provide optimal information as the output of subsequent modules in the pipeline. For example, in Figure 9CIn the example shown, the visualization module generates a pre-visualization result 970 via sensor 962 and workspace filter 964, clustering information 966, grasping synthetic data 968, and workspace-filtered sensor outputs 962, 964, and grasping synthetic results 968. This pre-visualization result 970 has access to the original filtered sensor node outputs, thereby enabling information-rich, high-quality visualizations that may be very useful for human operators when monitoring and / or intervening in robot operations (e.g., via teleoperation).
[0060] Figure 10 This diagram illustrates an example of a visual display generated and provided in an embodiment of a multi-camera image processing system. In various embodiments, display 1000 can be generated and displayed based on image data from multiple cameras in a workspace, as disclosed herein, for example, via... Figure 8 The process and system illustrated in Figure 9 are used to generate and display this. In the example shown, Display 1000 shows a workspace (or a portion thereof) including a robotic arm 1002 with a suction cup end effector 1004, on which a camera is mounted (as shown to the right). The robotic system in this example can be configured to retrieve items from a workbench 1006 and move them to a destination location, such as a tray or... Figure 10 Other destinations not shown. Objects 1008, 1010, and 1012 shown on workbench 1006 have each been identified as object instances, and fill patterns as shown indicate different colors used to highlight the objects and distinguish them. In various embodiments, display 1000 may be provided to enable a human operator to detect errors in the segmentation of objects to be manipulated by the robotic system, such as whether two adjacent objects are shown within a single color and / or bounding box (or other shape), and / or to control the robotic arm 1002 via teleoperation.
[0061] Figure 11 This is a flowchart illustrating an embodiment of a process for generating code for processing sensor data in a multi-camera image processing system for robot control. In various embodiments, process 1100 is performed by a computer (such as...) Figure 1 The control computer 118) executes the process. In the example shown, the pipeline definition is received and parsed (1102). For example, the pipeline definition can be received via a user interface, configuration file, API, etc. As defined, instances of the processing components to be included in the processing pipeline are created and linked together as defined in the pipeline definition (1104). The binary code for implementing the components and pipeline is compiled (if necessary) and deployed (1106).
[0062] In various embodiments, the techniques disclosed herein can be used to perform robotic operations fully or partially autonomously and / or via full or partial teleoperation based on image data generated by multiple cameras in the workspace (in some embodiments including one or more cameras mounted on a robotic arm or other robotic actuator).
[0063] Although the foregoing embodiments have been described in considerable detail for clarity of understanding, the invention is not limited to the details provided. Many alternative ways of implementing the invention exist. The disclosed embodiments are illustrative and not restrictive.
Claims
1. A system comprising: A communication interface is configured to receive image data from each of a plurality of sensors associated with the workspace, wherein for each of the plurality of sensors, the image data includes two-dimensional visual image information and depth information, and the plurality of sensors include a plurality of cameras, wherein the plurality of cameras are automatically and dynamically recalibrated. as well as The processor, coupled to the communication interface, is configured to: Image data from multiple sensors is merged to generate merged point cloud data; Segmentation is performed based on a subset of visual image data from multiple sensors to generate segmentation results; as well as Use the merged point cloud data and segmentation results to generate a merged 3D and segmented view of the workspace. In the segmentation process, RGB pixels identified as associated with object boundaries are mapped to corresponding depth pixels, and for each camera's image data, point clouds for each object are labeled and centroids are calculated, and nearest neighbor calculations are performed between the centroids of the corresponding object point clouds for each camera to segment the objects.
2. The system as claimed in claim 1, wherein, The plurality of sensors includes one or more three-dimensional 3D cameras.
3. The system as described in claim 1, wherein, Visual image data includes RGB data.
4. The system as claimed in claim 1, wherein, The processor is configured to perform bounding box fitting of objects in the workspace using one or both of the merged point cloud data and the segmentation results.
5. The system as described in claim 4, wherein, The processor is further configured to use one or both of the merged point cloud data and the segmentation results to synthesize a strategy for using a robotic arm to grasp objects.
6. The system of claim 5, wherein, The processor is further configured to implement strategies for using a robotic arm to grasp objects.
7. The system of claim 6, wherein, The processor is configured to combine robot operations to grasp objects, pick them up from a starting position and place them at a destination location in the workspace.
8. The system of claim 1, wherein, The processor is further configured to display a visualization of the workspace using merged 3D and segmented views of the workspace.
9. The system of claim 8, wherein, The displayed visualizations prominently showcase the objects depicted within the workspace.
10. The system of claim 1, wherein, The processor is configured to use merged point cloud data and segmentation results to generate a merged 3D and segmented view of the workspace, at least in part, by projecting a set of points, including the segmentation results, onto the merged point cloud.
11. The system of claim 1, wherein, The processor is further configured to perform secondary sampling on the merged point cloud data.
12. The system of claim 11, wherein, The processor is further configured to perform clustering processing on the resampled point cloud data.
13. The system of claim 12, wherein, The processor is configured to use subsampled and clustered point cloud data, along with segmentation results, to generate bounding box fitting results for objects in the workspace.
14. The system of claim 1, wherein, The processor is further configured to verify the merged 3D and segmented views of the workspace based at least in part on visual image data associated with sensors not included in a subset of the sensors.
15. The system of claim 14, wherein, The processor is configured to generate a first bounding box fit result about objects in the workspace, based at least in part on visual image data associated with sensors not included in a subset of sensors, and at least in part on the combined 3D and segmented views of the workspace verified by using combined 3D and segmented views of the workspace. A second box fit is generated using visual image data associated with sensors not included in a subset of the sensors; and a box fit for verification of the object is determined using the first and second box fits.
16. The system of claim 1, wherein, The processor is configured to merge and process image data, at least in part, by implementing a user-defined processing pipeline.
17. The system of claim 16, wherein, The processor is configured to receive and parse the definition of a user-defined processing pipeline.
18. The system of claim 17, wherein, The processor is configured to use definitions to create instances of modules that include pipelines and automatically generate binary code for implementing the modules and pipelines.
19. A method comprising: Image data is received from each of a plurality of sensors associated with the workspace. For each of the plurality of sensors, the image data includes two-dimensional visual image information and depth information, and the plurality of sensors include a plurality of cameras, wherein the plurality of cameras are automatically and dynamically recalibrated. Image data from multiple sensors is merged to generate merged point cloud data; Segmentation is performed based on a subset of visual image data from multiple sensors to generate segmentation results; as well as Use the merged point cloud data and segmentation results to generate a merged 3D and segmented view of the workspace. In the segmentation process, RGB pixels identified as associated with object boundaries are mapped to corresponding depth pixels, and for each camera's image data, point clouds for each object are labeled and centroids are calculated, and nearest neighbor calculations are performed between the centroids of the corresponding object point clouds for each camera to segment the objects.
20. A computer program product implemented in a non-transitory computer-readable medium and comprising computer instructions for performing the following operations: Image data is received from each of a plurality of sensors associated with the workspace. For each of the plurality of sensors, the image data includes two-dimensional visual image information and depth information, and the plurality of sensors include a plurality of cameras, wherein the plurality of cameras are automatically and dynamically recalibrated. Image data from multiple sensors is merged to generate merged point cloud data; Segmentation is performed based on a subset of visual image data from multiple sensors to generate segmentation results; as well as Use the merged point cloud data and segmentation results to generate a merged 3D and segmented view of the workspace. In the segmentation process, RGB pixels identified as associated with object boundaries are mapped to corresponding depth pixels, and for each camera's image data, point clouds for each object are labeled and centroids are calculated, and nearest neighbor calculations are performed between the centroids of the corresponding object point clouds for each camera to segment the objects.
Citation Information
Patent Citations
A man-machine cooperation oriented real-time posture detection method for hand-held objects
CN109255813A
Systems and methods for extrinsic calibration of a plurality of sensors
US20180308254A1