Robotic manipulation in cluttered environments

CN116901034BActive Publication Date: 2026-08-11EBOTS INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但是,典型的组装流水线需要99.9995%甚至更高的成功率

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116901034B_ABST
    Figure CN116901034B_ABST
Patent Text Reader

Abstract

This disclosure relates at least to reliable robot manipulation in cluttered environments. One embodiment may provide a robot system. The robot system may include: a robot arm including an end effector; an illumination unit including multiple monochromatic light sources of different colors; a structured light projector for projecting an coded light pattern onto a scene; one or more cameras for capturing a pseudo-color image of the scene illuminated by the monochromatic light sources of different colors and a scene image with the projected coded light pattern; a pose determination unit for determining the pose of a component based on the pseudo-color image and the scene image with the projected coded light pattern; a path planning unit for generating a motion plan for the end effector based on the determined component pose and the current pose of the end effector; and a robot controller for controlling the movement of the end effector according to the motion plan to allow the end effector to grasp the component.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 332,922, filed April 20, 2022, entitled “Reliable Robotic Manipulation in a Clutered Environment”, the disclosure of which is incorporated herein by reference in its entirety for all purposes. Technical Field

[0003] This disclosure generally relates to the control of robotic systems. In particular, this disclosure relates to controlling robotic systems based on machine learning models and 3D computer vision. Background Technology

[0004] Advanced robotics has dramatically transformed how products are manufactured, sparking the Fourth Industrial Revolution (also known as Industry 4.0). This revolution improves upon the computing and automation technologies developed during the Third Industrial Revolution by allowing computers and robots to connect and communicate, ultimately enabling decision-making without human intervention. The convergence of cyber-physical systems, the Internet of Things (IoT), and the Internet of Systems (IoS) makes Industry 4.0 possible and smart factories a reality. Intelligent machines (e.g., robots) become smarter as they acquire more data and learn new skills, potentially leading to more efficient, productive, and less wasteful factories. Ultimately, a network of digitally connected intelligent machines capable of creating and sharing information will enable truly unsupervised "lights-out manufacturing."

[0005] Artificial intelligence (AI) plays a crucial role in enabling robots to autonomously perform complex tasks, learn from experience, and adapt to ever-changing environments. Over the past decade, various machine learning methods have been developed for controlling robot operation, such as supervised learning, unsupervised learning, and reinforcement learning. While current machine learning methods for robot control have demonstrated strong versatility, adaptability, and dexterity, they often fall short of the precision requirements of consumer electronics manufacturing, which often necessitates robots identifying components of varying shapes and sizes (some of which are minute) in cluttered environments, grasping components without dropping or damaging them, and manipulating components according to the needs of manufacturing tasks. The success rate of robot manipulation based on existing machine learning models ranges from 90% to 97%. However, typical assembly lines require success rates of 99.9995% or even higher. Summary of the Invention

[0006] One embodiment may provide a robotic system. The robotic system may include: a robotic arm including an end-effector; an illumination unit including multiple monochromatic light sources of different colors; a structured light projector for projecting an encoded light pattern onto a scene; one or more cameras for capturing pseudo-color images of the scene illuminated by the monochromatic light sources of different colors and scene images with the projected encoded light pattern; a pose determination unit for determining the pose of a component of interest based on the pseudo-color image of the scene and the scene image with the projected encoded light pattern; a path planning unit for generating a motion plan for the end-effector based on the determined component pose and the current pose of the end-effector; and a robot controller for controlling the movement of the end-effector according to the motion plan to allow the end-effector to grasp the component of interest.

[0007] In a variant of this embodiment, the robot system may further include an error compensation unit for compensating for errors in the movement of the end effector.

[0008] In other variants, the error compensation unit will apply machine learning techniques to determine the controller’s desired pose corresponding to the camera-indicated pose of the end effector, so that when the robot controller controls the movement of the end effector based on the controller’s desired pose, the end effector achieves the camera-indicated pose as observed by the camera.

[0009] In a variant of this embodiment, different monochromatic light sources of different colors can be turned on alternately, one color at a time.

[0010] In a variant of this embodiment, the monochromatic light source may include a light-emitting diode (LED), and the color of the monochromatic light source may be between ultraviolet and infrared.

[0011] In a variant of this embodiment, the robotic system may further include an image segmentation unit for generating a segmentation mask for the scene image based on a pseudo-color image.

[0012] In other variations, the robotic system may also include a point cloud generation unit for generating a three-dimensional (3D) point cloud of the component of interest by overlaying a segmentation mask onto a scene image with a projected coded light pattern.

[0013] In other variants, the pose determination unit can determine the pose of the component of interest based on the component's 3D point cloud and geometric model.

[0014] In other variants, the image segmentation unit can generate segmentation masks by implementing a machine learning model that includes a convolutional neural network based on mask regions (mask R-CNN).

[0015] In a variant of this embodiment, the encoded light pattern can be encoded based on maximum min-SWgray codes.

[0016] One embodiment provides a computer-implemented method for controlling a robotic arm. During operation, a robot controller can generate a set of initial instructions to control the robotic arm to move an end effector toward a component of interest in a work scene. In response to determining that the end effector is near the component of interest, the controller can configure multiple monochromatic light sources of different colors to illuminate the work scene, configure a structured light projector to project an coded light pattern onto the work scene, and configure one or more cameras to capture a pseudo-color image of the work scene illuminated by the different monochromatic light sources and an image of the work scene with the projected coded light pattern. The controller can determine the pose of the component of interest based on the pseudo-color image of the work scene and the image of the work scene with the projected coded light pattern, generate a set of refined instructions based on the determined component pose and the current pose of the end effector, and control the movement of the end effector according to the set of refined instructions to allow the end effector to grasp the component of interest. Attached Figure Description

[0017] Figure 1 An exemplary robot system according to one embodiment of this application is illustrated.

[0018] Figure 2 A block diagram of an exemplary refinement subsystem according to an embodiment of this application is illustrated.

[0019] Figure 3 A flowchart illustrating an exemplary operation process of a robot system according to an embodiment of this application is presented.

[0020] Figure 4 A flowchart illustrating an exemplary process for performing an assembly task according to an embodiment of this application is presented.

[0021] Figure 5 An exemplary computer system for facilitating the operation of a robot system according to one embodiment of this application is illustrated.

[0022] In each figure, similar reference numerals refer to the same figure elements. Detailed Implementation

[0023] Overview

[0024] The embodiments described herein address the technical problem of improving the operational accuracy of robotic systems. More specifically, the robotic system may include a refinement subsystem that can be combined with a machine learning model to control the operation of a robotic arm. The initial movement of the robotic arm can be guided by any known machine learning model, and the refinement subsystem can be activated when the end effector of the robotic arm is near a target component. Once activated, the refinement subsystem can determine the poses of the target component and the end effector with refined accuracy. The refinement subsystem may include a 3D computer vision system with a structured light projector and a multi-wavelength illumination unit to facilitate accurate segmentation of the work scene. The refinement subsystem can generate a refined motion plan based on the determined component pose and the end effector pose, and control the movement of the robotic arm based on the refined motion plan to perform an assembly task. The refinement subsystem may also include an error compensation unit, which can be used to reduce the pose error of the end effector during the movement of the robotic arm.

[0025] Robotic systems with 3D computer vision

[0026] Efficient robotic systems can mimic humans and may include an arm / hand, eyes, and a brain. Like a human arm, a robotic arm can use its hands and fingers (e.g., end effectors) to pick up or grasp components of interest, bring them to the desired mounting location, and then install them. Just as humans use their eyes to guide arm movements, robotic systems can use computer vision to guide the movement of the robotic arm. Just as the movement of a human arm is controlled by the brain, similarly, the movement of a robotic arm is controlled by a robot controller. The robot controller uses visual information provided by computer vision to determine the posture of one or more grippers in order to perform a specific task or a series of tasks.

[0027] As previously mentioned, robot controllers can implement various machine learning models developed for robot control applications to guide the movement of a robotic arm based on images of a work scene. These machine learning models can be trained using real-world or synthetic data. For example, a masked region-based convolutional neural network (masked R-CNN) model can be trained to segment images of a work scene to detect the pose of a target component. In another example, the robot controller can implement an autoregressive grasping planning model that maps sensor inputs to a scene (e.g., images captured by a camera) to a probability distribution of possible grasps. However, due to the uncertainty of predictions by existing machine learning models, the movement of the robotic arm cannot achieve the sub-millimeter accuracy required for consumer electronics manufacturing. In some embodiments of this application, the robot system can implement a refinement subsystem that can improve the accuracy of the robotic arm's motion or component manipulation by enabling more accurate image segmentation techniques, geometry-based pose determination techniques, and machine learning-based error compensation techniques.

[0028] Figure 1 An exemplary robot system according to one embodiment of this application is illustrated. The robot system 100 may include a robot arm 102, a three-dimensional (3D) computer vision system 104, and a refinement subsystem 106. In some embodiments, the robot arm 102 may include a base 108, multiple joints (e.g., joints 110 and 112), and a gripper 114. The combination of multiple joints enables the robot arm 102 to have a wide range of motion and six degrees of freedom (6DoF). The gripper 114 can grasp and manipulate a target component 116 (e.g., move it to a desired location, place it in a desired pose, etc.) to perform a desired task.

[0029] A 3D computer vision system 104 (which may include multiple cameras) can capture images of a work scene, including a gripper 114, components 116 grasped by the gripper 114, and other components that may be present in the work scene. In addition to cameras, the 3D computer vision system 104 may also include various mechanisms for “understanding” the work scene based on the captured images. For example, the 3D computer vision system 104 may include mechanisms for detecting / identifying components, mechanisms for measuring component dimensions, mechanisms for calculating gripper pose, etc.

[0030] Figure 1 Three different Cartesian coordinate systems (e.g., XYZ) are shown, including a coordinate system whose origin is at the robot's base (called the robot-base coordinate system), a coordinate system whose origin is centered on the tool / gripper (called the tool-center coordinate system), and a coordinate system whose origin is at the camera's center (called the camera coordinate system). Robot controller ( Figure 1(Not shown in the image) The movement of the robot arm 102 relative to the robot-base coordinate system is typically controlled. The camera typically refers to the camera coordinate system to observe the scene (including the observed pose of the gripper 114). The actual pose of the grasped component can be calculated in the tool-center coordinate system. Various mechanisms can be used to facilitate coordinate transformations between different coordinate systems. For example, calibration targets and machine learning techniques can be used to calibrate the transformation from the camera coordinate system to the robot-base coordinate system (this transformation is called eye-hand coordination).

[0031] The refinement subsystem 106 may include various functional blocks and units that can enhance the robot system 100’s perception of the work scene and improve the reliability and accuracy of the movement of the robot arm 102 and gripper 114.

[0032] In the initial phase of robot operation, the robot controller can control the movement of the robot arm 102 based on instructions generated by a conventional machine learning model (e.g., a convolutional neural network (CNN)). For example, the robot controller can implement a component classification network to identify various components within the work scene and a pose classification network to determine the pose of the component of interest based on images captured by the 3D computer vision system 104. The robot controller can then control the movement of the robot arm 102 and the gripper 114 to pick up the component and perform the desired task. Ideally, the gripper 114 can move to the exact location of the component of interest, with its pose aligned with the component, in order to pick it up. However, due to the predictive uncertainty of the machine learning model, the gripper 114 may not reach the exact location or have the exact pose. In this case, the refinement subsystem 106 needs to be activated by the robot controller to improve the movement accuracy of the robot arm 102.

[0033] The refinement subsystem 106 may include an illumination unit with multiple monochromatic light sources and a structured light projector. Once activated, the monochromatic light sources can alternately emit different colors of light, allowing a black-and-white (BW) camera within the 3D computer vision system 104 to capture pseudo-color images of the work scene. For example, an image captured under red illumination can be referred to as a pseudo-red image. The structured light projector can also project encoded images (e.g., spatially varying light patterns) onto the work scene. The refinement subsystem 106 may include a high-resolution image segmentation model (e.g., a masked region-based convolutional neural network (masked R-CNN)) that can segment images of the work scene based on these pseudo-color images. Image segmentation using pseudo-color images can more accurately depict target components even in cluttered environments compared to conventional image segmentation techniques. A 3D point cloud of the target component can be generated based on both the segmented image and the structured light pattern. According to some embodiments, the 3D point cloud can be generated based on an image of a scene illuminated by a light pattern encoded with Gray code (such as max-minimum SW (MMSW) Gray code). A more detailed description of image segmentation and 3D point cloud generation can be found in U.S. Patent Application No. 18 / 098,427 (Attorney’s File No. EBOT22-1001NP), filed January 18, 2023, by inventors Zheng Xu, John W. Wallerius, and Sabarish Kuduwa Sivanath, entitled “SYSTEM AND METHODFOR IMPROVING IMAGE SEGMENTATION”, the disclosure of which is incorporated herein by reference.

[0034] The refinement subsystem 106 may also include a model library storing geometric models (e.g., computer-aided design (CAD) models) of various components. The pose of the component of interest can be determined based on the corresponding CAD model. In some embodiments, template matching techniques can be used to match the CAD model with a 3D point cloud. For example, the 3D CAD model of the component can be manipulated (e.g., rotated) until the surface points of the CAD model match the surface points of the 3D point cloud of the component. In other examples, least squares matching (LSM) techniques can be used to match the model pose with the component pose. The pose of the gripper 114 can be determined similarly.

[0035] Once both the orientation of the component and the orientation of the gripper 114 are determined, the refinement subsystem 106 can generate a motion plan based on the determined orientation and environment. For example, the generated motion plan can specify a path for the gripper 114 to reach the component without colliding with other components or fixtures in the work environment. The robot controller can then generate motion commands to send to the various motors of the robot arm 102 to control the movement of the gripper 114. When generating motion commands, the refinement subsystem 106 can compensate for errors in robot hand-eye coordination (i.e., transformations between the robot-base coordinate system and the camera coordinate system). A more detailed description of the technique for reducing hand-eye coordination errors in robots can be found in U.S. Patent Application No. 17 / 751,228 (Attorney’s File No. EBOT21-1001NP), filed May 23, 2022, by Sabarish Kuduwa Sivanath and Zheng Xu, entitled “SYSTEM AND METHOD FOR ERRORCORRECTION AND COMMENSATION FOR 3D EYE-TO-HAND COORDINATION,” the disclosure of which is incorporated herein by reference.

[0036] Figure 2 A block diagram of an exemplary refinement subsystem according to an embodiment of this application is illustrated. The refinement subsystem 200 can facilitate the operation of a robot system. More specifically, the refinement subsystem 200 can interface with a robot controller to improve the accuracy and reliability of robot operation. The refinement subsystem 200 may include a control unit 202, a lighting unit 204, a camera unit 206, a structured light projector 208, a segmentation unit 210, a point cloud generation unit 212, a posture determination unit 214, a motion planning unit 216, an error compensation unit 218, and a component library 220.

[0037] The control unit 202 is responsible for controlling and coordinating the operation of various units within the refinement subsystem 200, such as the lighting unit 204, the camera unit 206, and the structured light projector 208. More specifically, the control unit 202 can synchronize the lighting of the work scene by the lighting unit 204 and / or the structured light projector 208 with the image capture operation of the camera unit 206.

[0038] The lighting unit 204 may include multiple monochromatic light sources. In some embodiments, the monochromatic light sources may include light-emitting diodes (LEDs) of various colors, with wavelengths ranging from the ultraviolet band (e.g., about 380 nm) to the infrared band (e.g., about 850 nm). In one example, the monochromatic light source may include infrared LEDs, yellow LEDs, green LEDs, and violet LEDs. The LEDs may be mounted above the work area, with LEDs of the same color arranged symmetrically to reduce shadows. When the refinement subsystem 200 is activated, or more specifically, when the lighting unit 202 is turned on, different colors of monochromatic light sources may be turned on alternately, one color at a time, so that the work area is illuminated by light of a specific color at a given moment. In alternative embodiments, more than one color (e.g., two or three colors) of LEDs may be turned on simultaneously to achieve a desired lighting effect. The control unit 202 may control the turning of the LEDs on and off.

[0039] Camera unit 206 may include multiple cameras mounted above the work scene. The cameras can be arranged such that they can capture images of the work scene from different angles. The cameras may be BW cameras to achieve high resolution. In some embodiments, the BW cameras may capture pseudo-color images of the work scene. A pseudo-color image refers to an image captured under illumination of a specific color of light. The operation of the cameras within camera unit 206 can be synchronized with the LEDs in illumination unit 204 via control unit 202.

[0040] The structured light projector 208 can be responsible for projecting coded light patterns onto a work scene. In some embodiments, the light source of the structured light projector 208 may include a laser source having an emission wavelength of approximately 455 nm. In an alternative embodiment, the light source of the structured light projector 208 may include an LED. The structured light projector 208 can project a series of coded light patterns onto the work scene. The projection of the light patterns can be synchronized with the image capture operation of the camera unit 206 via the control unit 202, enabling at least one image to be captured for each structured light pattern. In one embodiment, the structured light projector 208 can project light patterns encoded using binary code (e.g., max-min SW Gray code).

[0041] Segmentation unit 210 can be responsible for segmenting images of a work scene. In some embodiments, segmentation unit 210 can implement a machine learning model, such as Masked R-CNN, which can receive pseudo-color images of different colors as input and output image segmentation results. The machine learning model can have multiple input channels (one channel for each color), and the pseudo-color images of different colors can be concatenated along the channel dimension (e.g., in ascending wavelength order) before being sent to multiple input channels, with each pseudo-color image being sent to its corresponding input channel. The output of the machine learning model can include a segmentation mask. Segmentation unit 210 can also use depth information of the scene to enhance the accuracy of image segmentation. In one embodiment, the depth map of the scene can also be used as input to the machine learning model. In an alternative embodiment, the segmentation mask can be superimposed on an image of the scene captured under structured light (i.e., an image of a structured light pattern) to generate segmentation of the structured light pattern.

[0042] Point cloud generation unit 212 can be responsible for generating 3D point clouds of components of interest. In one embodiment, point cloud generation unit 212 can generate 3D point clouds of components by overlaying a segmentation mask of the component onto a structured light pattern to depict the component from its surroundings. In other embodiments, point cloud generation unit 212 can generate 3D point clouds of components by overlaying a segmentation mask of the component onto an image of a scene illuminated by a pattern using MMSW Gray code encoding.

[0043] The pose determination unit 214 can be responsible for determining the pose (e.g., position and orientation) of a component based on its 3D point cloud. In some embodiments, the pose determination unit 214 can determine the pose of the component based on its 3D point cloud and a geometric model (CAD model). For example, the pose determination unit 214 can apply template matching techniques to match the pose of the CAD model to the observed component pose (i.e., the pose of the 3D cloud). More specifically, the pose determination unit 214 can manipulate the pose of the CAD model until it matches the observed component pose. In one embodiment, the pose determination unit 214 can apply least-squares matching techniques to match the pose. In addition to the component of interest, the pose determination unit 214 can also determine the pose of a gripper. In one embodiment, the gripper's pose can be determined from a segmented image of the work scene (e.g., by segmenting the gripper from the background). In an alternative embodiment, the gripper's pose can be determined based on the on-the-fly settings of various motors within the robot arm (i.e., the robot controller always knows the pose of the end effector).

[0044] Motion planning unit 216 can be responsible for generating motion plans based on the poses of the gripper and components. Note that the pose of the component includes not only positional information but also orientation information. Similarly, the pose of the gripper includes not only positional information but also orientation information. When generating the motion plan, motion planning unit 216 needs to consider environmental factors, such as other components or equipment that may be located on the gripper's path. Motion planning unit 216 can calculate the gripper's path so that the gripper can reach the component of interest without interference from other components or equipment in the work environment. In some embodiments, motion planning unit 216 can calculate the path based on a 3D image of the work environment. In other embodiments, motion planning unit 216 can use machine learning techniques to calculate the path. For example, a trained deep learning neural network can be used to plan the path between two locations within the work environment.

[0045] Error compensation unit 218 is responsible for real-time compensation of movement errors when the gripper moves according to a motion plan. More specifically, due to transformation errors between the camera coordinate system (which can be used to represent the observed pose of the component and gripper) and the robot-base coordinate system (which can be used by the robot controller to generate motion commands / instructions to be sent to various motors in the robot arm) and mechanical defects of the motors, the movement of the robot arm planned by the controller may differ from the actual movement observed by the camera. In other words, there may be a pose error between the pose desired by the controller and the pose indicated by the camera. To compensate for such errors, the pose of the component of interest can first be determined by the camera in the camera coordinate system, and then this pose can be transformed to coordinates in the robot-base coordinate system (e.g., based on a predetermined transformation matrix), referred to as the camera-indicated pose. In some embodiments, error compensation unit 218 may apply machine learning techniques to infer the error matrix of the camera-indicated pose, and then calculate the controller-desired pose of the gripper based on the camera-indicated pose and the error matrix.

[0046] Applying machine learning techniques may include training a deep learning neural network using multiple test samples. The trained neural network can take the camera-indicated pose as input and output an error matrix. The error matrix correlates the camera-indicated pose with the controller's desired pose. In one embodiment, the controller's desired pose can be obtained by multiplying the camera-indicated pose by the error matrix. The gripper's controller's desired pose can then be sent to the robot controller to generate motion commands that will cause the gripper to reach the camera-indicated pose for successful component grasping. In one example, the motion plan may include multiple steps, and the error compensation unit 218 can compensate for pose errors at each step.

[0047] Component library 220 may include geometric models (e.g., CAD models) of various components in the work scene.

[0048] Figure 3 A flowchart illustrating an exemplary operation of a robot system according to an embodiment of this application is presented. During operation, the robot controller may generate a set of initial instructions or motion commands based on an image of the work scene (operation 302). In some embodiments, the robot controller may implement one or more conventional machine learning models (e.g., CNN) to locate the component of interest and estimate its pose based on 2D and / or 3D images of the work scene. The estimated pose may include errors, and the set of initial instructions may be generated based on the estimated pose. The robot controller may send initial instructions or motion commands to various motors of the robot arm to control the movement of the end effector (operation 304). The end effector may be instructed to move toward the component of interest.

[0049] Once the end effector stops moving, the robot controller can determine whether the end effector is near the component of interest based on the current image of the working scene (operation 306). For example, based on a scene image captured by a high-resolution camera with a small field of view (FOV), the robot controller can determine whether the end effector is within the camera's FOV. If the end effector is not near the component of interest, the robot controller can regenerate the initial instructions or motion commands based on the current image of the working scene (operation 302).

[0050] If the end effector is near the component of interest, the robot controller can activate the refinement subsystem (operation 308). The refinement subsystem can be similar to... Figure 2The refinement subsystem is shown in the diagram. Once activated, the various units within the refinement subsystem (e.g., control unit, illumination unit, camera unit, structured light projector, segmentation unit, point cloud generation unit, pose determination unit, motion planning unit, and error compensation unit) can interact to provide more accurate information about the pose of the component of interest. In one example, once activated, the refinement subsystem can be configured to alternately illuminate the work scene with monochromatic light sources of different colors, one color at a time, and to configure one or more cameras to capture pseudo-color images of the work scene. The refinement subsystem can also configure the structured light projector to project an coded light pattern onto the work scene and configure the cameras to capture images of the work scene using the projected coded light pattern. The refinement subsystem can further determine the refined pose of the component of interest based on the pseudo-color image of the work scene and the image of the work scene with the projected coded light pattern. In some embodiments, the refinement subsystem can generate a 3D point cloud of the component by generating a segmentation mask based on the pseudo-color image and overlaying the segmentation mask onto the image of the work scene with the projected coded light pattern. Then, the refined pose of the component of interest can be determined based on template matching between the component's 3D point cloud and 3D geometric model.

[0051] The robot controller can then generate and send a set of refined instructions or motion commands to the motors based on the refined pose of the component (operation 310). Note that generating refined instructions may include compensating for errors that may occur in the transformation between the camera coordinate system and the robot-base coordinate system. Upon receiving the refined instructions, the motors in the robot arm can operate accordingly, causing the end effector to grasp the component of interest (operation 312).

[0052] Figure 3 The process illustrated is an example of how a robotic arm can perform a simple operation of grasping a component of interest with the help of a refinement subsystem. In practice, a robotic arm can perform more complex assembly operations, such as attaching radio frequency (RF) pads or RF cable heads to corresponding receptacles. The refinement subsystem also needs to be activated to ensure that such tasks can be performed successfully (e.g., the end effector can identify pads or cable heads from various components scattered around the work area, grasp the pads or cable heads, move the pads or cable heads to the vicinity of the corresponding receptacle, and successfully attach the pads or cable heads to the receptacle).

[0053] Figure 4A flowchart illustrating an exemplary process for performing an assembly task according to an embodiment of this application is presented. The assembly task may include insertion operations, such as inserting RF pads or cable heads into corresponding sockets. During the operation, the robot system may determine the refined pose of the component to be assembled (operation 402) based on the output of the refined subsystem. The component to be assembled may include RF pads or cable heads. More specifically, determining the refined pose of the component may include image segmentation based on pseudo-color images of the work scene captured by a camera under monochromatic light source illumination of different colors. Determining the refined pose of the component may also include overlaying a segmentation mask onto an image with a structured light pattern having projections to separate the 3D image of the component from the background. A 3D point cloud corresponding to the component may be generated based on the segmented 3D image, and the refined pose of the component may be determined by template matching between the 3D point cloud of the component and the 3D geometric model.

[0054] Subsequently, the robot system can execute advanced motion planning (operation 404). For example, the robot system can calculate a path to bring the end effector to the vicinity of the component. In some embodiments, conventional machine learning-based motion planning techniques can be used to calculate the path without activating the refinement subsystem to ensure that the end effector can be brought to the vicinity of the component without interference from other components or equipment in the work environment. The robot controller can then generate motion commands and send them to the motors to bring the end effector to the vicinity of the component (operation 406). In some embodiments, multiple iterations may be required to ensure that both the end effector and the component are within the FOV of the high-resolution camera of the refinement subsystem.

[0055] The robotic system can then determine the refined pose of the end effector (operation 408). Note that the robotic system can use similar segmentation techniques to segment the 3D image of the end effector from the background. The refined pose of the end effector can also be determined by template matching between the 3D point cloud of the end effector and its geometric model. In some embodiments, the end effector may include one or more labels that can be identified as features for template matching. The location of each label can be used as a starting point for model-based point cloud registration.

[0056] By determining the refined pose of the end effector and components, the robot system can perform pose error compensation until the pose error is below a predetermined threshold (operation 410). In some embodiments, performing pose error compensation may include inferring an error matrix that can be used to transform the camera-indicated pose observed by the camera into a pose desired by the controller, such that when the robot controller controls the movement of the robot arm based on the controller-desired pose, the end effector achieves the camera-indicated pose observed by the camera.

[0057] The robot controller can then generate and send motion commands (operation 412) to enable the end effector to successfully grasp the component to be assembled. For example, the robot controller can generate motion commands based on the desired posture of the controller.

[0058] The robotic system can determine the refined orientation of the mating component (operation 414). For the insertion operation, the mating component can be a corresponding socket that accepts RF pads or cable heads. Similar techniques discussed earlier can be used to determine the refined orientation of the mating component.

[0059] Based on the refined pose of the mating components, the robot system can generate a motion plan to bring the end effector to the vicinity of the mating components (operation 416). The robot system can then determine a new refined pose of the component to be assembled (because its pose changes after being grasped by the end effector) and a new refined pose of the end effector (operation 418), and then generate and send motion commands that enable the end effector to successfully perform the insertion operation (operation 420). Generating the motion commands may include the robot system determining a transformation matrix (e.g., a component transformation matrix) that correlates the pose of the end effector with the pose of the grasped component. Note that generating the motion commands may also include operations for compensating for pose errors. A detailed description of the component transformation matrix and pose error compensation can be found in the aforementioned U.S. Patent Application No. 17 / 751,228.

[0060] Figure 5 An exemplary computer system for facilitating the operation of a robotic system according to one embodiment is illustrated. The computer system 500 includes a processor 502, a memory 504, and a storage device 506. Furthermore, the computer system 500 may be coupled to a peripheral input / output (I / O) user device 510, such as a display device 512, a keyboard 514, and a pointing device 516. The storage device 506 may store an operating system 520, a robot control system 522, and data 540.

[0061] The robot control system 522 may include instructions that, when executed by the computer system 500, cause the computer system 500 or processor 502 to perform the methods and / or processes described in this disclosure. Specifically, the robot control system 522 may include instructions for controlling the initial movement of the end effector to bring the end effector near the operating part (initial movement control instruction 524), instructions for controlling the lighting unit to synchronize the operation of the lighting unit and the camera (lighting control instruction 526), ​​instructions for controlling the structured light projector to synchronize the operation of the structured light projector and the camera (structured light control instruction 528), instructions for performing image segmentation based on a pseudo-color image (image segmentation instruction 530), instructions for generating a 3D point cloud of the component of interest (point cloud generation instruction 532), instructions for determining the refined pose of the component and the end effector (refined pose determination instruction 534), instructions for executing a high-level motion plan (motion planning instruction 536), and instructions for compensating for pose errors during end effector movement (error compensation instruction 538). Data 540 may include component models 542 and training samples 544 used by various machine learning models.

[0062] Generally, embodiments of the present invention can provide a system and method for refining the movement of an end effector of a robotic arm. The robotic system may include a refinement subsystem that can improve the accuracy and reliability of robot operation. Compared to conventional robot control systems that rely on existing machine learning techniques, the refinement subsystem can implement more advanced image segmentation and point cloud generation techniques to determine the refined poses of components and the end effector in the working scene. Advanced image segmentation can be performed based on pseudo-color images, and the refined poses can be determined based on 3D point clouds and geometric models of the components and the end effector. Furthermore, machine learning-based error compensation techniques can be used to compensate for pose errors caused by transformations between the camera coordinate system and the robot-base coordinate system.

[0063] The foregoing description is presented to enable any person skilled in the art to make and use the embodiments and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this disclosure. Therefore, the invention is not limited to the illustrated embodiments, but is accorded the widest scope consistent with the principles and features disclosed herein.

[0064] The methods and processes described in the detailed description section can be implemented as code and / or data, which can be stored in a computer-readable storage medium as described above. When a computer system reads and executes the code and / or data stored on the computer-readable storage medium, the computer system executes the methods and processes implemented as data structures and code and stored in the computer-readable storage medium.

[0065] Furthermore, the methods and processes described above may be incorporated into hardware devices or apparatuses. Hardware modules or apparatuses may include, but are not limited to, application-specific integrated circuit (ASIC) chips, field-programmable gate arrays (FPGAs), dedicated or shared processors that execute specific software units or code segments at specific times, and other programmable logic devices now known or developed hereafter. When a hardware device or apparatus is activated, it executes the methods and processes contained therein.

Claims

1. A robot system, comprising: Robotic arm, including end effector; The lighting unit includes multiple monochromatic light sources of different colors; Structured light projectors are used to project coded light patterns onto a scene; One or more black and white cameras; as well as A computer system includes a processor and a storage device, the storage device storing instructions that, when executed by the processor, cause the processor to perform a method, the method comprising: Estimate the initial pose of the components of interest; Without activating the lighting unit, structured light projector, and monochrome camera, the robot arm moves its end effector based on the estimated initial posture of the components. In response to determining that the end effector is near the component, the lighting unit, the structured light projector, and the monochrome camera are activated, such that the monochrome camera captures a pseudo-color image of the scene illuminated by the lighting unit and an image of the scene having an coded light pattern projected by the structured light projector, wherein the corresponding pseudo-color image is captured when the scene is illuminated by a monochromatic light source of the corresponding color. Multiple pseudo-color images of different colors are cascaded in ascending wavelength order; Multiple pseudo-color images of different colors are cascaded and input into a neural network with multiple input channels, where each pseudo-color image is sent to a corresponding input channel. The refined pose of the components of interest is determined based on the output of the neural network. A movement plan is generated for the end effector based on the refined pose of the defined components and the current pose of the end effector; and The movement of the end effector is controlled according to the motion plan to allow the end effector to grasp the component of interest.

2. The robot system of claim 1 further includes an error compensation unit for compensating for errors in the movement of the end effector.

3. The robot system of claim 2, wherein the error compensation unit applies machine learning techniques to determine the controller desired posture corresponding to the camera indicated posture of the end effector, such that when the robot controller controls the movement of the end effector based on the controller desired posture, the end effector achieves the camera indicated posture as observed by the camera.

4. The robot system of claim 1, wherein different monochromatic light sources of different colors are turned on alternately, one color at a time.

5. The robot system of claim 1, wherein the monochromatic light source comprises a light-emitting diode (LED), and wherein the color of the monochromatic light source is between ultraviolet and infrared.

6. The robot system of claim 1 further includes an image segmentation unit for generating a segmentation mask for the scene image based on a pseudo-color image.

7. The robot system of claim 6, further comprising: A point cloud generation unit is used to generate a 3D point cloud of the component of interest by overlaying a segmentation mask onto an image of a scene with a projected coded light pattern.

8. The robot system of claim 7, wherein the pose determination unit determines the pose of the component of interest based on the 3D point cloud and geometric model of the component.

9. The robotic system of claim 6, wherein the image segmentation unit generates a segmentation mask by implementing a machine learning model including a convolutional neural network based on mask regions (mask R-CNN).

10. The robot system of claim 1, wherein the encoded light pattern is based on max-min SW Gray code encoding.

11. A computer-implemented method for controlling a robotic arm, the method comprising: Without activating the refinement subsystem, the robot controller generates a set of initial instructions to control the robot arm to move the end effector toward the component of interest in the work scene. The refinement subsystem includes an illumination unit with multiple monochromatic light sources of different colors, a structured light projector, and one or more black and white cameras. In response to determining that the end effector is near the component of interest, the refinement subsystem is activated by the following operation: Alternately turn on the multiple monochromatic light sources of different colors to illuminate the work scene; Activate the structured light projector to project the coded light pattern onto the work scene; as well as Activate the one or more monochrome cameras to capture a pseudo-color image of the work scene and an image of the work scene having an coded light pattern projected by a structured light projector, wherein a corresponding pseudo-color image is captured when the scene is illuminated by a monochromatic light source of the corresponding color; Multiple pseudo-color images of different colors are cascaded in ascending wavelength order; Multiple pseudo-color images of different colors are cascaded and input into a neural network with multiple input channels, where each pseudo-color image is sent to a corresponding input channel. The refined pose of the components of interest is determined based on the output of the neural network. A set of refined instructions is generated based on the refined pose of the defined components and the current pose of the end effector; as well as The robot controller controls the movement of the end effector according to the set of refined instructions, allowing the end effector to grasp the component of interest.

12. The method of claim 11, further comprising compensating for errors in the movement of the end effector.

13. The method of claim 12, wherein compensating for errors in the movement of the end effector includes applying machine learning techniques to determine a controller desired pose corresponding to the camera indicated pose of the end effector, such that when the robot controller controls the movement of the end effector based on the controller desired pose, the end effector achieves the camera indicated pose as observed by the camera.

14. The method of claim 11, wherein the monochromatic light sources of different colors are configured to be turned on alternately, one color at a time.

15. The method of claim 11, wherein the monochromatic light source comprises a light-emitting diode (LED), and wherein the color of the monochromatic light source is between ultraviolet and infrared.

16. The method of claim 11, further comprising generating a segmentation mask for the image of the work scene based on a pseudo-color image.

17. The method of claim 16, further comprising generating a three-dimensional (3D) point cloud of the component of interest by overlaying a segmentation mask onto an image of a scene having a projected coded light pattern.

18. The method of claim 17, wherein the pose of the component of interest is determined based on the 3D point cloud and geometric model of the component.

19. The method of claim 16, wherein generating the segmentation mask includes implementing a machine learning model comprising a convolutional neural network based on the mask regions (masked R-CNN).

20. The method of claim 11, wherein the encoded light pattern is based on max-min SW Gray code encoding.

Citation Information

Patent Citations

  • System and method for error correction and compensation for 3D eye-to-hand coordination

    US12030184B2

  • System and method for improving image segmentation

    US12518391B2

  • Method for recovering real-time three-dimensional body posture based on multimodal fusion

    CN102800126A

  • Dispersedly stacked material pickup apparatus and method

    CN106934833A

  • Robot gravity pose decomposition joint error offline compensation method, system and terminal

    CN114131611A