Precise placement based on real-time object pose estimation
By installing image capture devices on the moving parts of the robot system, object pose information can be acquired in real time, and joint movement and image processing can be controlled in parallel. This solves the problem of pose uncertainty of objects in grippers, and achieves precise placement and improved production efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2026-03-31
AI Technical Summary
When robots use general gripping tools, the uncertainty of the object's pose in the gripper affects the placement accuracy, and the cycle time for precise placement affects production efficiency.
By installing image capture devices on the moving parts of the robot system, object pose information can be acquired in real time. The joint movement and image processing can be controlled in parallel by the computing system and the control system to achieve precise placement of the object.
It improves the accuracy of object placement and production efficiency, reduces extra movements or pauses, and enhances the task flexibility of the robot system.
Smart Images

Figure CN121773006A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an automated robot system. Background Technology
[0002] Robotic devices can be and / or include general-purpose grippers and / or end effectors capable of manipulating a variety of objects of different shapes and sizes. Such tools (e.g., general-purpose grippers and / or end effectors) can enhance the robot's task flexibility, enabling it to perform a wider range of tasks and adapt to changing environments. However, using general-purpose grippers can introduce uncertainty into the pose of the object in the gripper after each gripping, which may affect placement accuracy.
[0003] To address this challenge, accurate placement of objects in general-purpose grippers typically requires estimating the object's pose in the gripper after each pick-and-place maneuver. However, the cycle time for accurate placement is critical to productivity, as it impacts overall efficiency. Therefore, there remains a technological need to enhance the robot's ability to reduce additional movements or pauses when performing manipulations between pick-and-place operations. Summary of the Invention
[0004] A first aspect of this disclosure provides a method for controlling a robot comprising multiple joints, wherein the multiple joints include a first joint and one or more other joints, the first joint rotating the entire robot along a fixed base. The method includes: moving the robot from a first workspace to a second workspace based on controlling the first joint to rotate the robot along the fixed base, wherein the robot grips an object and the object corresponds to a first tool center point (TCP); acquiring one or more images of the object using an image capturing device while the robot is moving from the first workspace to the second workspace, wherein the image capturing device is mounted on a link connected to the first joint and moves with the first joint of the robot; determining a second TCP based on the one or more images; controlling one or more other joints of the robot to adjust the object gripped by the robot according to the second TCP; and placing the object at a target location in the second workspace according to the TCP refined based on the adjustment result converging to a predetermined reference.
[0005] According to the implementation of the first aspect, the method further includes: determining that the adjustment result has not converged to a predetermined reference; acquiring one or more other images of the object using an imaging device while the robot is moving from a first workspace to a second workspace; determining a third TCP based on the one or more other images; controlling one or more other joints of the robot to adjust the object held by the robot according to the third TCP; and placing the object at a target location in the second workspace according to a refined TCP after the new adjustment result has converged to the predetermined reference.
[0006] According to the implementation of the first aspect, placing an object at a target location in the second workspace based on a TCP refined by convergence of the adjustment result to a predetermined reference further includes: determining that the adjustment result converges to the predetermined reference; determining a third TCP of the object based on the second TCP and the adjustment of the object; acquiring one or more images of the object after adjustment using an image capture device; and generating a refined TCP of the object based on the third TCP and the one or more images captured after adjustment.
[0007] According to the implementation of the first aspect, the method further includes: processing one or more images to obtain a viewpoint image; and performing pose estimation based on the viewpoint image, wherein determining the second TCP based on one or more images is based on pose estimation.
[0008] According to the implementation of the first aspect, processing one or more images to obtain a viewpoint image further includes: cropping one or more images based on objects captured in one or more images; and removing background from one or more cropped images based on one or more images obtained by an imaging device.
[0009] According to the implementation of the first aspect, when the robot is moving from a first workspace to a second workspace, one or more images of an object are acquired using an image capture device, and the method further includes: controlling the movement of one or more other joints of the robot to allow the image capture device to capture images of the object from different viewpoints.
[0010] According to the implementation of the first aspect, the first TCP is associated with a relative coordinate system established based on the field of view of the image capture device.
[0011] According to the implementation of the first aspect, the adjustment result includes the updated TCP of the object. The predetermined reference includes the target TCP of the object, wherein the target TCP of the object corresponds to a predetermined point on the object and is defined by coordinates in a relative coordinate system corresponding to the image capturing device. The adjustment result converges to the predetermined reference based on the updated TCP of the object converging to the target TCP of the object.
[0012] According to the implementation of the first aspect, the adjustment result includes the updated pose of the object, the predetermined reference includes the desired pose of the object, and the adjustment result converges to the predetermined reference based on the updated pose of the object converging to the desired pose of the object.
[0013] According to the implementation of the first aspect, the convergence between the updated pose of the object and the desired pose of the object is determined based on the image containing the updated pose of the object and the image containing the desired pose of the object.
[0014] According to the implementation of the first aspect, determining the second TCP based on one or more images further includes: determining the pose of the object in a relative coordinate system based on one or more images; determining the offset between the first TCP and the pose of the object in the relative coordinate system; and determining the second TCP based on the first TCP and the offset.
[0015] According to the implementation of the first aspect, determining the second TCP based on one or more images further includes: determining the pose of the object in a relative coordinate system based on one or more images; obtaining one or more parameters of the object associated with the position and orientation of the object in the relative coordinate system; and determining the second TCP based on the one or more parameters of the object.
[0016] According to the implementation of the first aspect, one or more images obtained by the image capture device record the motion of an object caused by one or more other joints, the recorded motion of the object being independent of the motion of the first joint.
[0017] According to the implementation of the first aspect, the operation of (i) the first joint and (ii) the image capture device and the operation of one or more other joints are controlled in parallel.
[0018] A second aspect of this disclosure provides a robot system comprising: a robot including a plurality of joints, the plurality of joints including a first joint and one or more other joints, the first joint rotating the entire robot along a fixed base; an image capturing device mounted on a link connected to the first joint, the image capturing device being configured to move with the first joint of the robot and acquire one or more images of an object held by the robot; and a control system. The control system is configured to: move the robot from a first workspace to a second workspace based on controlling the first joint to rotate the robot along the fixed base, wherein the robot holds an object and the object corresponds to a first tool center point (TCP); acquire one or more images of the object using the image capturing device while the robot is moving from the first workspace to the second workspace; determine a second TCP based on the one or more images; control one or more other joints of the robot to adjust the object held by the robot according to the second TCP; and control the robot to place the object at a target location in the second workspace according to the TCP refined based on the adjustment result converging to a predetermined reference.
[0019] According to the implementation of the second aspect, the control system is further configured to: determine that the adjustment result has not converged to a predetermined reference; acquire one or more other images of the object using an imaging device while the robot is moving from a first workspace to a second workspace; determine a third TCP based on the one or more other images; control one or more other joints of the robot to adjust the object held by the robot according to the third TCP; and control the robot to place the object at a target location in the second workspace according to a refined TCP after the new adjustment result has converged to the predetermined reference.
[0020] According to the implementation of the second aspect, the control system is further configured to: determine that the adjustment result converges to a predetermined reference; determine a third TCP of the object based on the second TCP and the adjustment of the object; acquire one or more images of the object after adjustment using an image capture device; and generate a refined TCP of the object based on the third TCP and one or more images captured after adjustment.
[0021] According to the implementation of the second aspect, the control system is further configured to: process one or more images to obtain a viewpoint image; and perform pose estimation based on the viewpoint image; wherein the control system is configured to determine a second TCP based on the pose estimation.
[0022] According to the implementation of the second aspect, one or more images obtained by the image capture device record the motion of an object caused by one or more other joints, and the recorded motion of the object is independent of the motion of the first joint.
[0023] According to the implementation of the second aspect, the control device is also configured to control (i) the operation of the first joint and (ii) the operation of the image capture device and one or more other joints in parallel. Attached Figure Description
[0024] Embodiments of this disclosure will now be described in more detail with reference to the exemplary accompanying drawings. This disclosure is not limited to exemplary embodiments. All features described and / or illustrated herein may be used individually or in different combinations in the embodiments of this disclosure. The features and advantages of various embodiments of this disclosure will become apparent from reading the detailed description with reference to the accompanying drawings, which illustrate the following:
[0025] Figure 1 A simplified block diagram depicting a robot system according to one or more examples of this disclosure is shown;
[0026] Figure 2 This is a schematic diagram of an exemplary control system according to one or more examples of this disclosure;
[0027] Figure 3A An exemplary robotic arm according to one or more embodiments of the present disclosure is demonstrated;
[0028] Figure 3B Examples of robotic systems operating in a working environment according to one or more embodiments of the present disclosure are shown;
[0029] Figure 4 The illustration shows a process 400 for operating a robot system according to one or more embodiments of the present disclosure;
[0030] Figure 5 This is a flowchart illustrating an example process for operating a robot system according to one or more embodiments of the present disclosure. Detailed Implementation
[0031] This disclosure describes system integration strategies and algorithms that allow a robotic system to estimate the pose of a manipulated object during motion and make adjustments accordingly, thereby achieving precise object placement and improving efficiency.
[0032] The robotic system implements a vision system and / or a perception system (e.g., via an image capture device) that moves with the robotic system to acquire information about the pose of an object in real time (e.g., capture an image of the object). For example, the image capture device can be placed on a moving part of the robotic system (e.g., on link 1 of the robotic arm) and configured to move with that moving part. This configuration enables the image capture device to collect pose information of the object decoupled from the motion of the moving part. Based on the information collected by the image capture device, the robotic system can manipulate other parts of the robotic system (other joints in the robotic arm) to adjust the pose of the object according to the information provided by the image capture device. Furthermore, the robotic system can move the object along with the entire robot via this moving part and adjust the object's pose in parallel with other parts of the robotic system. It is important to note that the moving part on which the image capture device is placed can be associated with any joint in the robotic arm. In some variations, some or all joints in the robotic arm can have the image capture device placed on them. For example, as... Figure 3A As shown, a second image capture device can be placed on a link connected to joint 3 to capture images while the robot arm is moving in the vertical direction.
[0033] Traditional robotic systems estimate the pose of an object in the gripper of the robot system after each manipulation. In some cases, traditional robotic systems may rely on image capture devices deployed at specific locations in the environment, such as the pick-up point or destination. For example, a camera may be positioned next to the pick-up point and configured to face upwards from the ground to monitor the robot's object-picking process. In this setup, the robot needs to move the object toward the camera to obtain sufficient information to perform pose estimation on the object picked up by the robot. Some robotic systems may be equipped with image capture devices that move with the robot system. However, when using these image capture devices for pose estimation, traditional robotic systems alternately perform the following operations: (i) fixing the entire robot to capture images and estimate the object's pose; and (ii) driving the robot to change its pose and / or move. The techniques disclosed herein offer advantages over traditional robotic system integration strategies / algorithms by enabling the robotic system to perform motion and adjustment operations simultaneously, and in particular, allowing the robotic system to collect and process information independently of certain movements of the robotic system.
[0034] In some cases, multiple image capture devices (such as color cameras, depth cameras, or any suitable combination) can be placed on a moving part of a robotic system (e.g., on link 1 of a robotic arm in a robotic system) and configured to move with that moving part of the robotic system. Information collected by the multiple image capture devices can be combined to estimate the pose of an object based on a relative coordinate system that moves with the moving part and remains stationary relative to the image capture devices.
[0035] In some variations, the robotic system can implement one or more computer vision algorithms, including model-based and / or machine learning-based methods, to estimate the pose of an object. The robotic system can rely on the estimated pose of the object to determine adjustments to that pose. For example, the robotic system can be trained to manipulate an object to a desired pose independent of the robot system's motion relative to an image-capturing device via a specific moving part (e.g., a joint with an image-capturing device mounted on it), thereby allowing the robotic system to adjust the object's pose from one workspace to another as it moves via that specific moving part.
[0036] Specifically, exemplary aspects of the robot system and / or robot of this disclosure will be further illustrated below with reference to the exemplary embodiments depicted in the accompanying drawings. The exemplary embodiments illustrate some implementations of this disclosure, but are not intended to limit the scope of this disclosure.
[0037] In all the accompanying drawings, the same reference numerals indicate similar but not identical elements. The drawings are not drawn to scale, and the dimensions of some parts may be enlarged to illustrate the examples shown more clearly. Furthermore, the drawings provide examples and / or implementations consistent with the specification; however, the specification is not limited to the examples and / or implementations provided in the drawings.
[0038] Unless otherwise expressly stated, any term expressed in the singular form herein shall also include the plural form, and vice versa. Furthermore, as used herein, the terms “a” and / or “one” shall mean “one or more,” even if the phrase “one or more” is used herein. Additionally, when something described herein is “based on” another thing, that thing may also be based on one or more other things. In other words, unless otherwise expressly stated, as used herein, “based on” means “at least partially based on” or “at least partially based on.”
[0039] Figure 1 A simplified block diagram depicting a robot system 100 according to one or more embodiments of the present disclosure is shown.
[0040] See Figure 1 The robot system 100 includes a robot 110 and a computing system 130. The robot 110 includes various hardware and software components configured to perform specific tasks, such as sensing its surrounding environment, touching and / or detaching from objects. The computing system 130 is configured to process data from the robot 110 and generate signals / instructions based on the processed data. In some examples, the computing system 130 may be integrated into or communicatively coupled to a control device 118 within the robot 110. In some cases, the robot system 100 may be communicatively connected to sensors (such as image capture devices, depth sensors, etc.) positioned in the environment to obtain information about the environment in which the robot system 100 performs its tasks. Objects interacting with the robot 110 refer to items or entities that come into contact with or touch the robot 110 during a task or operation. These objects can be any type of item, object, device, product, etc., which the robot 110 can manipulate from a first location to a second location. For example, Figure 3B Box 388 is shown as an example of an object.
[0041] The robot 110 includes one or more image capture devices 112, multiple joints 114, multiple actuators / motors 116, and a control system 200.
[0042] Multiple image capture devices 112 are configured to capture images of objects touched by the robot 110. The image capture devices 112 may be color cameras, depth cameras, combinations thereof, or other suitable electronic image acquisition devices. The image capture devices 112 may be located and / or positioned on joints 114 of the robot 110, such as on links of a specific joint 114. The multiple image capture devices 112 may move with a specific joint 114 of the robot 110 to capture images in real time. This is explained below. Figure 3B It is shown and described in the text.
[0043] Multiple joints 114 are configured to enable the robot 110 to move along different axes of motion. Each joint 114 is driven by one or more actuators / motors 116, thereby allowing the robot 110 to move with high precision and high flexibility. The actuators / motors 116 include AC motors, DC motors, geared motors, linear motors, actuators, or any other electrically controlled devices used to implement the kinematics of the robot 110.
[0044] The control device 118 includes one or more controllers and / or control units, through which the control device 118 sends signals / instructions to control the operation of other components in the robot 110, such as causing multiple actuators / motors 116 to control the movement of corresponding joints 114, or causing (multiple) image capture devices 112 and / or sensors to collect information, or causing an end effector (e.g., extension equipment 120) to interact with the environment (e.g., touch / detach from an object).
[0045] Additional sensors 122 may be located and / or positioned at the robot 110, and may optionally be included within the control device 118. These additional sensors 122 may combine (or serve as a backup of) information provided to the control device 118 with information provided by the image capture device 112 (e.g., images). For example, these additional sensors 122 may include light sensors and / or flash camera sensor systems for providing light / illumination to images captured using the image capture device 112.
[0046] Robot 110 may optionally include an extension device 120, which may be a general-purpose gripping tool attached to robot 110 via a tool flange. Extension device 120 may be embodied as other types of end effectors attached to robot 110 via mounting mechanisms such as quick-change mechanisms, threaded couplings, or other suitable mounting mechanisms. In some cases, extension device 120 may be integrated with robot 110; in others, robot 110 may be detached from extension device 120 but may be accessible to support extension device 120. In some variations, extension device 120 may be electrically connected to control device 118, enabling control device 118 to send signals / commands to control the operation of extension device 120.
[0047] In some examples, control device 118 manipulates robot 110 by changing the physical position and / or orientation of joints 114, thereby aligning an object touched by robot 110 (e.g., via extension equipment 120) with a predefined location in the workspace. For example, control device 118 may move the base joints of robot 110 to move an object to the vicinity of a predefined location in the placement workspace, while simultaneously moving other joints to orient the object so that it is aligned with the predefined location in the workspace. In other cases, control device 118 may dynamically move multiple joints 114 in any order (including simultaneously) to provide smooth movement of the object and placement into a predefined location in the placement workspace, and to reduce cycle time.
[0048] Go back and see Figure 1 The computing system 130 may be part of or an extension of the control device 118 of the robot 110. The computing system 130 includes one or more processors 132, a communication interface 134, and a memory 136, which are communicatively coupled to a bus 138. Data can be transferred between the one or more processors 132, the communication interface 134, and the memory 136 via the bus 138.
[0049] One or more processors 132 are configured to perform operations according to instructions stored in memory 136. The processors 132 may be any suitable type of general-purpose or special-purpose microprocessor (e.g., CPU or GPU, respectively), digital signal processor, microcontroller, etc.
[0050] Memory 136 is configured to store computer-readable instructions that, when executed by processor 132, cause processor(s) 132 to perform the various operations disclosed herein. Memory 136 may be any non-transient mass storage medium, such as volatile or non-volatile, magnetic, semiconductor, magnetic tape, optical disc, removable, non-removable, or other types of storage devices or tangible computer-readable media, including but not limited to read-only memory (“ROM”), flash memory, dynamic random access memory (“RAM”), and / or static RAM.
[0051] Communication interface 134 is configured to transmit information between computing system 130 and robot 110, such as Figure 1 As shown in the diagram. As an example, communication interface 134 may include an Integrated Services Digital Network (“ISDN”) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity. For example, communication interface 134 may include a Local Area Network (LAN) card for providing data communication connectivity with a compatible LAN. As another example, communication interface 134 may include a high-speed network adapter, such as a fiber optic network adapter, a 10G Ethernet adapter, etc. Wireless links can also be implemented by communication interface 134. In this implementation, communication interface 134 can send and receive electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information via a network. This network can typically include cellular communication networks, wireless local area networks (WLANs), wide area networks (WANs), etc. In some variations, communication interface 134 may include various I / O devices, such as a keyboard, mouse, touchpad, touchscreen, microphone, camera, biosensors, etc.
[0052] In some examples, computing system 130 may implement a machine learning (ML) / artificial intelligence (AI) training system that trains ML and / or AI models, datasets, and / or algorithms (e.g., neural networks (NNs) and / or convolutional neural networks (CNNs)). Robotic system 100 can use this ML / AI model to assist in manipulating robot 110 to perform specific tasks and, where possible, reduce computational overhead. For example, computing system 130 may train or implement a trained ML / AI model, which can then be used to predict the next movement of robot 110, thereby precisely placing an object at a predefined location within the placement workspace.
[0053] In some cases, the computing system 130 may be implemented using one or more computing platforms, devices, servers, and / or apparatuses. In other cases, the computing system 130 may be implemented as an engine, software function, and / or application. In other words, the functionality of the computing system 130 may be implemented as software instructions, which are stored in a storage device (e.g., memory) and executed by one or more processors.
[0054] In some variations, robot system 100 uses one or more images, or a series of consecutive images and / or videos, captured by image capture devices 112 to manipulate robot 110 to perform specific tasks. For example, robot system 100 captures one or more images including an object touched by robot 110. Robot system 100 uses the one or more images to manipulate robot 110 to adjust the pose of the object so that robot 110 can precisely place the object at a predefined location in a placement workspace. In some cases, robot system 100 can use a trained neural network to determine regions of interest and / or points of interest (e.g., keypoints) within the images. Robot system 100 uses the determined regions of interest / keypoints to determine the pose of the object and / or to determine adjustments to the pose of the object. Based on the pose of the object and / or based on adjustments to the pose of the object, robot system 100 manipulates robot 110 to adjust the pose of the object.
[0055] It should be understood that Figure 1 The exemplary robot system 100 depicted herein is merely an example, and the principles discussed herein may also apply to other situations—for example, other types of robots 110.
[0056] Figure 2 This is a schematic diagram of an exemplary control system 200 according to one or more embodiments of the present disclosure. The control system 200 includes data from... Figure 1 The control device 118 and other suitable entities (e.g., image capture device 112, computing system 130, etc.). It should be understood that... Figure 2 The control system 200 shown is merely an example, and additional / alternative embodiments of the control system 200 are contemplated within the scope of this disclosure.
[0057] Control system 200 includes controller 210. Controller 210 is not limited to any particular hardware, and the configuration of the controller can be implemented through any type of programming (e.g., embedded Linux) or hardware design, or a combination of both. For example, controller 210 can be formed by a single processor (such as a general-purpose processor) and its corresponding software that implements the described control operations. On the other hand, controller 210 can also be implemented by dedicated hardware, such as ASIC (Application-Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), DSP (Digital Signal Processor), GPU (Graphics Processing Unit), NVIDIA Jetson devices, hardware accelerators, processors and / or other devices running TensorFlow, TensorFlowLite, PyTorch, and / or other ML software. In some cases, control system 200 and / or controller 210 can be edge computing hardware located on and / or included within robot 110.
[0058] Controller 210 communicates electrically with memory 230. Memory 230 may be and / or include computer-usable or computer-readable media, such as, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor computer-readable media. More specific examples of computer-readable media (e.g., a non-exhaustive list) may include: electrical connections having one or more wires; tangible media such as portable computer floppy disks, hard disks, time-varying memory (RAM), ROM, erasable programmable read-only memory (EPROM or flash memory), compact disc read-only memory (CD-ROM), or other tangible optical or magnetic storage devices. Memory 230 may store corresponding software, such as computer-readable instructions (code, scripts, etc.). The computer instructions cause controller 210 to control control system 200 when executed by controller 210, thereby providing operation of robot 110 as described herein.
[0059] The controller 210 is configured to provide and / or acquire information, such as one or more images from the image capture device 112. For example, the image capture device 112 may capture one or more images, or a continuous sequence of images, including objects that come into contact with the robot 110, and may provide these images to the controller 210. The controller 210 may use these images (either alone or in combination with other elements of the control system 200) to determine the pose and / or other appropriate state of the objects.
[0060] Additional sensors 122 may optionally be included within the control system 200. These additional sensors 122 may provide information to the control system 200 in combination with (or as a backup of) information provided by the image capture device 112 (such as images).
[0061] Additionally and / or alternatively, the additional sensor 122 may optionally include another image capture device (2D or 3D), a LiDAR sensor, a radio frequency identification (RFID) sensor, an ultrasonic sensor, a capacitive sensor, an inductive sensor, a magnetic sensor, and / or similar sensors to refine the trajectory of the robot end effector as it manipulates the object to a predefined pose, following visual recognition of the object's initial pose using the visual or video information described herein. Generally, any sensor capable of providing signals to enhance or improve the control system 200's ability to manipulate the robot 110 for efficient and precise object picking and / or placement can be included in the control system 200.
[0062] In some variations, the image capture device 112 and the additional sensor 122 form a flash-based photographic sensing system. This flash-based photographic system includes at least an image capture device (e.g., device 112) and a light source (e.g., used to emit a flash). In operation, the emitter flashes cyclically, or provides constant light by keeping the emitter continuously emitting light, and the resulting image is captured by the image capture device 112.
[0063] The control system 200 is configured to drive the actuators / motors 116 of the joints 114 of the robot 110. Each joint 114 may be driven by an actuator / motor 116. As used herein, the actuator / motor 116 includes an AC motor, a DC motor, a geared motor, a linear motor, an actuator, or any other electrically controlled device used to implement the kinematics of the robot 110. Therefore, the control system 200 is configured to automatically and continuously determine the physical state of the robot 110 and automatically control the individual actuators / motors 116 of the joints 114 to manipulate the robot 110 to move objects it touches and / or adjust its pose.
[0064] The control system 200 includes control devices 118 (e.g., motor control units or microcontrollers) for joint 114. These control devices 118 may be part of controller 210 or as stand-alone devices. For a specific joint (e.g., joint 1), the control devices 118 use feedback from one or more sensors 226 (e.g., encoders) to control the actuator 222 to provide real-time control of the actuator / motor 116. Therefore, the control devices 118 receive instructions for controlling the actuator / motor 116 (e.g., receiving motor / actuator control signals from controller 210) and interpret these instructions in conjunction with feedback signals from the sensors 226, thereby providing control signals to the actuator 222 for accurate and real-time control of the actuator / motor 116 (e.g., sending motor / actuator driver signals). The actuator 222 converts the control signals transmitted by the control devices 118 into drive signals for driving the actuator / motor 116 (e.g., sending individual operating signals to the motor / actuator). In another example, control device 118 is integrated with the circuit system to directly control actuator / motor 116.
[0065] Control device 118 may be included as part of controller 210 or as a stand-alone processing system (e.g., a microprocessor). Therefore, like controller 210, control device 118 is not limited to any particular hardware, and the configuration of control device can be achieved through any type of programming or hardware design or a combination of both.
[0066] The control system 200 may include input / output (I / O) terminals 240 for sending and receiving various input and output signals. For example, the control system 200 may send / receive external communications or data to / from users, servers (e.g., billing servers and / or enterprise computing systems), power supply units, etc., via the I / O terminals 240. The control system 200 may also control a user feedback interface via the I / O terminals 240 (or other means).
[0067] Figure 3A An exemplary robotic arm 300 according to one or more embodiments of the present disclosure is demonstrated. The robotic arm 300 may be embodied as Figure 1 The robot system 100 shown is a part of or included within the robot 110. The robot arm 300 can be composed of... Figure 2 The control system 200 shown is used for control. It should be understood that... Figure 3A The robotic arm 300 shown is merely an example, and additional / alternative embodiments of the robotic arm 300 are contemplated within the scope of this disclosure.
[0068] See Figure 3AThe robotic arm 300 has six joints and is referred to as a six-axis robot or articulated robot. In a six-axis robot, each joint has a specific range of motion and direction of movement. The joints of the robotic arm 300 are connected by rigid segments or sections called links. The links between adjacent joints determine the overall reach and flexibility of the robotic arm 300. Depending on the specific application, links can vary in length, material, and shape.
[0069] The first joint 310 (also referred to as joint 1) is configured for base rotation. For example, the first joint 310 allows the entire arm to rotate horizontally about a vertical axis (e.g., arrow 315 shows the rotation of joint 1), thus providing the robot with the ability to turn or swing. Joint 1 can be connected to the base 302 (referred to as link 0).
[0070] A second joint 320 (also referred to as joint 2) is configured for shoulder rotation. The second joint 320 allows the upper arm to rotate vertically about a horizontal axis (e.g., arrow 325 shows the rotation of joint 2), thus allowing the arm to be raised or lowered. Joints 1 and 2 can be connected by one or more links 312 (referred to as link 1).
[0071] The third joint 330 (also referred to as joint 3) is configured for elbow rotation. The third joint 330 allows the forearm to rotate vertically about a horizontal axis (e.g., arrow 335 shows the rotation of joint 3), thereby enabling the arm to bend or straighten. Joints 2 and 3 can be connected by one or more links 322 (referred to as link 2).
[0072] The fourth joint 340 (also referred to as joint 4) is configured for wrist yaw. The fourth joint 340 allows the wrist to rotate vertically about a horizontal axis (e.g., arrow 345 shows the rotation of joint 4), thereby enabling the arm to tilt or twist the end effector. Joints 3 and 4 can be connected by one or more links 332 (referred to as link 3).
[0073] The fifth joint 350 (also referred to as joint 5) is configured for wrist pitch. The fourth joint 350 allows the wrist to pivot about an axis that rotates together with joint 4 (e.g., arrow 355 shows the movement of joint 5), thereby controlling the pitch motion of the end effector. Joints 4 and 5 can be connected by one or more links 342 (referred to as link 4).
[0074] The sixth joint 360 (also referred to as joint 6) is configured for wrist roll. The sixth joint 360 enables the wrist to rotate vertically about a horizontal axis (e.g., arrow 365 shows the rotation of joint 6), thereby controlling the roll motion of the end effector. Joints 5 and 6 can be connected by one or more links 352 (referred to as link 5).
[0075] The robotic arm 330 may optionally include a tool flange 370 or a suitable type of mounting mechanism configured to connect to an end effector (e.g., extension assembly 120). In some cases, the tool flange 370 may be integrated into or formed by one or more links (referred to as links 6) between the joint 6 and the end effector.
[0076] Figure 3B An example of a robot system 100 operating in a work environment 380 according to one or more embodiments of the present disclosure is shown. The robot system includes... Figure 3A The robotic arm 300 is shown. The robotic system 100 can be composed of... Figure 2 The control system 200 shown is used for control. It should be understood that... Figure 3B The robot system 100 shown is merely an example, and additional / alternative embodiments of the robot 380 are contemplated within the scope of this disclosure.
[0077] See Figure 3B The first joint 310 of the robotic arm 300 is configured to move the entire robotic arm 300 to rotate horizontally about a vertical axis. A link 382 connects two opposite sides of the first joint 310. An image capture device 384 is mounted on the link 382. Thus, when the first joint 310 rotates, the image capture device 384 moves with the entire robotic arm 300. The camera angle of the image capture device 384 can be fixed or variable. Multiple image capture devices 384 can be mounted on the link 382. In some variations, one or more image capture devices 384 can be mounted on other joints in the robotic arm 300.
[0078] In addition, the robotic arm 300, through its mounting mechanism ( Figure 3B (Not shown in the image) is connected to the adsorption device 386. For example... Figure 2 As shown, the operation of the adsorption device 386 can be controlled by the control system 200. For example, the control system 200 can control the adsorption device 386 to grasp an object (e.g., box 388) by creating a vacuum seal with the object's surface. Once the adsorption device 386 has firmly grasped the object, the control system 200 can control the robotic arm 300 to move the object to a desired location (e.g., insertion slot 390). At the target location, the control system 200 can deactivate the adsorption device 386 to release the vacuum seal with the object's surface. In this way, the control system 200 can control the robotic system 100 to place the object at the target location.
[0079] Figure 4 The illustration depicts a process 400 for operating a robot system 100 according to one or more embodiments of the present disclosure. Process 400 may be executed by a control system 200, and in particular by... Figure 2 The control system 200 shown is executed. However, it should be recognized that any of the following blocks can be executed in any suitable order, and process 400 can be executed in any suitable environment by any suitable controller or processor.
[0080] Robot 110 in robot system 100 includes multiple joints. The multiple joints include a first joint and one or more other joints, which rotate the entire robot 110 along a fixed base.
[0081] At block 402, the control system 200 rotates the robot 110 along the fixed base based on controlling a first joint (e.g., first joint 310), moving the robot 110 from a first workspace to a second workspace. The robot 110 grips an object, and the object corresponds to a first tool center point (TCP). The first workspace may be a pick-up workspace, from which the control system 200 moves the robot 110 to pick up an object. The second workspace may be a placement workspace, into which the control system 200 moves the robot 110 to place an object.
[0082] The TCP of robot 110 can be a specific point on the end effector (e.g., a suction cup device) of robot 110, where actions such as grasping, manipulating, or interacting with an object are performed. The TCP of robot 110 indicates the location where the tool contacts or interacts with the environment (e.g., an object). The TCP is a critical reference point used for programming and controlling the movement of robot 110. The position and orientation of the TCP determine how robot 110 approaches, handles, and interacts with an object during a task. The accuracy and control of the TCP are crucial for ensuring precise movement and successful task execution. In some examples, the TCP can be defined and calibrated based on the design and geometry of the end effector or tool attached to the robot arm. In some cases, the TCP can be determined based on a coordinate system established by selecting a reference point (e.g., the base of robot 110) and defining three axes (X, Y, and Z) for position and orientation. This coordinate system can be a relative coordinate system or a universal coordinate system.
[0083] For a robot 110 with an object in its end effector (e.g., a suction device), the TCP of the robot 110 can be defined as a specific point on the object when the robot 110 comes into contact with the object (e.g., when the robot 110 grips the object using the suction device). The object's TCP reflects the pose of the object as seen from the object's field of view when the object comes into contact with the robot 110. For example, for a robot 110 with an object in its end effector, the TCP of the robot 110 can be defined as the geometric center of the top / bottom surface of the object; while for a robot 110 without an object in its end effector, the TCP of the robot 110 can be defined as the center point of the contact surface of the robot 110's end effector (e.g., a suction device). When the robot 110 successfully picks up an object, it needs to find a suitable TCP based on the object's actual pose in the end effector so that the object placement task can be handled correctly. To this end, the control system 200 can calculate the offset between the new robot's TCP and the TCP of the robot 110 without an object in its end effector based on the pose information obtained for the object.
[0084] In some examples, the control system 200 may determine the TCP of the robot with an object in its end effector as the first TCP corresponding to that object. The control system 200 may determine the first TCP based on information collected by the image capture device 112 and / or other sensors in the robot system 100 (e.g., captured images). In some cases, the control system 200 may determine the first TCP based on the field of view of the image capture device 112 mounted on the robot 110. The field of view of the image capture device 112 may be associated with a relative coordinate system, allowing the image capture device 112 to record changes in the pose of the object in that relative coordinate system, thereby allowing the control system 200 to calculate offsets to update the object's TCP.
[0085] In some variations, when robot 110 picks up an object from a first workspace, control system 200 may obtain an initial TCP for object placement manipulation. The first workspace may be an environment equipped with several sensors (e.g., image capture devices, distance sensors, etc.). Control system 200 may determine the initial TCP based on information collected by the sensors or receive such information from a computing system in communication with them. Alternatively, when control system 200 successfully controls robot 110 to pick up an object from the first workspace, control system 200 may set a default TCP for object placement manipulation.
[0086] At block 404, as robot 110 moves from a first workspace to a second workspace, control system 200 uses image capture device 112 to acquire one or more images of the object. Image capture device 112 is mounted on a first joint of robot 110 and moves with it. The one or more images may be individually captured images or frames included in a video stream. In some variations, the one or more images may include color images, depth images, or combinations thereof.
[0087] The control system 200 controls the first joint to move the entire robot 110 together with the image capture device 112. As the robot 110 moves (e.g., in parallel), the control system 200 may also control one or more other joints to adjust the pose of the object based on one or more images captured by the image capture device 112.
[0088] The captured image comprises a large number of pixels for capturing an object or a portion thereof within a specific field of view. The control system 200 can control the movement of one or more other joints of the robot 110 to allow the image capture device 112 to capture images of the object from different perspectives. Furthermore, the control system 200 can process the raw images from the image capture device 112 to suppress background and / or noise information in the captured images. For example, the control system 200 can crop the captured images to remove the background around the object. Additionally and / or alternatively, the control system 200 can combine multiple captured images to filter out background and / or noise in the images. For example, different images may contain different backgrounds as the robot 110 moves through its first joint. The control system 200 can identify the changing background from the combined images and then remove the background.
[0089] At block 406, the control system 200 determines the second TCP based on one or more images.
[0090] The control system 200 can identify the pose (e.g., position and / or orientation) of an object based on (multiple) captured images. In some examples, the control system 200 can implement a trained ML model (e.g., a trained CNN) and input the captured images into the trained ML model to determine one or more regions of interest. The control system 200 can also use one or more image processing algorithms or techniques (e.g., scale-invariant feature transform (SIFT) techniques) on the determined regions of interest to determine the object's pose. In some examples, the ML model can be trained to take one or more captured / processed images (e.g., color images, depth images, or combinations thereof) as input to output an estimated pose of the object. The object's pose can be defined by six degrees of freedom (DOF), including displacement along the x, y, and z axes in a specific coordinate system, and rotation about three axes (e.g., roll, pitch, and yaw).
[0091] The control system 200 can quantify the pose of an object based on a relative coordinate system corresponding to the field of view of the image capture device 112. For example, when the system is set up, the pose of the image capture device 112 relative to the coordinate system of the robot system (e.g., relative to the base of the robot 110) can be calibrated. In this way, the control system 200 can perform transformations between the relative coordinate system corresponding to the image capture device 112 and the coordinate system corresponding to the robot 110 based on the calibration results. In some variations, the control system 200 can describe the pose of the object based on a point of interest on the object (e.g., the geometric center at the bottom surface of the object) and other suitable parameters associated with shape or orientation.
[0092] The control system 200 can determine the second TCP based on the object's pose. For example, the control system 200 can calculate the offset between the object's first TCP and the object's pose described in a relative coordinate system corresponding to the image capture device 112, and then apply this offset to the object's first TCP to obtain the object's second TCP. Alternatively, the control system 200 can extract suitable parameters (e.g., the object's geometric center) based on the object's pose described in a relative coordinate system corresponding to the image capture device 112, and use the extracted parameters to determine the object's second TCP.
[0093] At block 408, control system 200 controls one or more other joints of robot 110 to adjust the object held by robot 110 according to the second TCP.
[0094] Based on the second TCP of the object, the control system 200 can determine the adjustments to be performed by one or more other joints, such as moving, pitching, yawing or rotating the object to a certain angle.
[0095] In some examples, the control system 200 can determine adjustments based on the difference between the object's second TCP and the target TCP. The target TCP can be a predefined TCP that indicates the desired pose of the object held by the robot 110, enabling the robot 110 to achieve precise placement of the object. The goal of precise robot placement is defined based on the target TCP. When the image capture device 112 is set to a fixed camera angle, the desired pose of the object can be reflected as a specific pose in the captured image.
[0096] The control system 200 can calculate the offset between the second TCP and the target TCP based on the relative coordinate system corresponding to the image capture device 112. Based on the calculated offset, the control system 200 can determine the adjustments to be performed by other joints and generate control signals / instructions to the other joints accordingly to reduce or even eliminate the offset.
[0097] In some cases, the control system 200 can compare the object's current pose with its desired pose to determine the adjustments to be performed. As mentioned above, when the image capture device 112 is set to a fixed (or predetermined) camera angle, the object's desired pose can be reflected as a specific pose in the captured image. During calibration, the robot 110 can be configured to hold the object in the desired pose, and the control system 200 can control the image capture device 112 to capture the desired pose from one or more predetermined camera angles. For this purpose, the control system 200 can store (multiple) images corresponding to the object's desired pose for later reference. Additionally and / or alternatively, the control system 200 can extract a set of parameters to represent the object's desired pose and store this set of parameters for later reference.
[0098] The control system 200 can apply a suitable visual detection algorithm (e.g., by applying a trained ML / AI model) to compare the object's current pose with its desired pose to determine the adjustments to be performed by other joints. After several iterations, the image capture device 112 can "see" the object in the field of view in the state where it is exactly in the ideal pickup position.
[0099] Based on the determined adjustments, the control system 200 can manipulate (e.g., move / orient) other joints to adjust the pose of the object. Specifically, the control system 200 can send control signals (which may include instructions) to operate actuators / motors 116 to properly move and position other joints 114, thereby manipulating the pose of the object held by the robot 110. More specifically, the control system 200 determines control signals for the actuators / motors 116, which are configured (when executed) to controllably manipulate the actuators / motors 116 to position and orient the object touched by the robot 110 (e.g., via the extension equipment 120). The control system 200 then sends these control signals to perform the specified movement. The actuators / motors 116 may include multiple actuators / motors that are collectively configured to (ultimately) place the object at a target location in a second workspace. In some cases, actuators may be specifically designed to fine-tune the orientation / position of the object, and actuator control signals are directed to control such actuators.
[0100] Control device 118 for a specific joint 114 receives motor / actuator control signals and may also receive feedback signals. The feedback signals are provided by sensors 226, which detect the state / position of various motors / actuators in other joints of the robot 110. Based on the feedback signals and the motor / actuator signals, control device 118 determines motor driver signals and actuator driver signals for the specific joint. Thus, the control signals can be high-level instructions for the operation (or final position) of the various elements of the robot 110, and control device 118 can interpret these high-level instructions (based on the information provided by the feedback signals) to provide low-level control signals for driving the respective motors / actuators. Control device 118 directly sends the motor driver signals and actuator driver signals to the actuators 222 for the specific joint. In some examples, control device 118 includes a circuit system that can operate with appropriate voltage and current to drive an actuator coupled to its processing system (e.g., microcontroller, FPGA, ASIC, etc.), and thus can send motor driver signals and actuator driver signals directly to actuator / motor 116.
[0101] After adjustment, the control system 200 can determine whether the adjustment result (e.g., the resulting TCP / pose of the object) converges to a predetermined reference.
[0102] In some examples, the control system 200 can obtain an estimated TCP indicating the pose of the object after adjustment. The estimated TCP can be determined based on the motion of the second TCP and other joints. The control system 200 can then determine whether the estimated TCP converges to the target TCP.
[0103] In some cases, the control system 200 can control the image capture device 112 to capture one or more images of the object after adjustment. By applying a visual detection algorithm to the captured images(s), the control system 200 can determine whether the updated pose of the object converges to a predetermined reference. The control system 200 can compare the updated pose of the object in the captured images with the desired pose shown in a pre-stored reference image. For example, the control system 200 can implement a trained ML / AI model to determine whether key features in two input images match. Additionally and / or alternatively, the control system 200 can extract a set of parameters to represent the updated pose of the object and compare the extracted set of parameters with a pre-stored set of parameters corresponding to the desired pose.
[0104] When the control system 200 determines that the adjustment result (e.g., the resulting TCP / pose of the object) converges to a predetermined reference, the control system 200 can control the robot system 100 to collect appropriate information to obtain a refined TCP of the object, and then proceed to the next operation (e.g., at block 410). For example, the control system 200 can use the TCP calculated after adjustment as the refined TCP of the object. Optionally, the control system 200 can fine-tune the refined TCP based on additional images captured by the image capture device 112. The control system 200 can also further adjust the pose of the object via other joints. In the example, the control system 200 can adjust the object (e.g., via other joints) and fine-tune the refined TCP until the difference between the pose estimates of two adjacent objects is less than 1 mm. When the control system 200 determines that the adjustment result has not converged to the predetermined reference, the control system 200 can perform the operations in blocks 404-408 over several iterations until the adjustment result converges to the predetermined reference.
[0105] At block 410, the control system 200 places the object at the target location in the second workspace according to the refined TCP. In some examples, the control system 200 may determine the refined TCP of the robot 110 to perform the object placement. The refined TCP of the robot 110 may be related to the actual pose of the object in the end effector by an offset calculated based on the expected pose of the object in the end effector.
[0106] When the control system 200 determines that the adjustment result converges to the predetermined reference, the control system 200 can terminate the pose (or object TCP) adjustment loop and continue to obtain refined TCP for the object or robot 110.
[0107] The control system 200 determines that the first joint has completed its movement from the first workspace to the second workspace, and that the object's fine-grained TCP is ready. Then, the control system 200 controls the robot 110 to place the object at the target location in the second workspace and release the object.
[0108] Figure 5 This is a flowchart illustrating an example process 500 for operating a robot system 100 according to one or more embodiments of the present disclosure. Process 500 can be executed by a control system 200, particularly as... Figure 2 As shown in the figure. However, it should be recognized that any of the following blocks can be executed in any suitable order, and process 500 can be executed in any suitable environment by any suitable controller or processor.
[0109] See Figure 5 The robot 110 in the robot system 100 includes a first joint (e.g., joint 1) and five other joints (e.g., joints 2 to 6). Joint 1 rotates the entire robot 110 along a fixed base, and an image capture device 112 is mounted on a link 1 connecting joint 1 and joint 2.
[0110] The control system 200 controls joints 1 and 2 through 6 to perform parallel tasks. The control system 200 controls joint 1 to move the object held by the robot 110 from the pick-up workspace 502 to the placement workspace 504, as illustrated in the exemplary embodiment. Figure 4 Block 402 in the middle.
[0111] Simultaneously, the control system 200 controls the image capture device 112 to acquire an image 506 of the object in real time as joint 1 rotates the entire robot 110, and controls joints 2-6 to adjust the pose of the object based on the captured image. The control system 200 continuously updates the TCP of the object based on the pose of the object under the grasp of the robot 110 obtained from the real-time captured image 506.
[0112] Specifically, the control system 200 performs image processing 508 on the captured image to enhance information about the object and / or suppress background / noise information. The control system 200 may perform image processing 508 in two steps. In the first step, the control system 200 may crop the captured image to remove the background around the object. The control system 200 may apply feature extraction techniques or other suitable image processing algorithms to achieve optimal cropping of the captured image. Next, the control system 200 may combine multiple captured images to filter out background and / or noise in the captured images. For example, when the robot 110 is moving through its first joint, different images may contain different backgrounds. The control system 200 can identify the changing background from the combined images and then remove the background. These image processing operations can reduce image size and filter out background noise, thereby improving the quality and efficiency of image processing.
[0113] Based on the processed image, the control system 200 performs pose estimation 510 for the object and updates the robot TCP corresponding to the object accordingly (in block 512), see exemplary embodiment. Figure 4 In block 406 of the diagram, as an example, the control system 200 can implement a computer vision algorithm to perform pose estimation. For example, the model can be trained to take one or more processed images as input and output an estimated pose of an object. The object's pose can be defined by six degrees of freedom (DOF), including displacement along the x, y, and z axes in a specific coordinate system, and rotation about three axes (e.g., roll, pitch, and yaw). In block 512, the control system 200 can determine an updated TCP corresponding to the object based on the estimated pose of the object and use the updated TCP as the current TCP corresponding to the object.
[0114] Based on the updated TCP, the control system 200 generates control signals / commands for joints 2-6 to adjust the pose of the object. See also the exemplary embodiment of controlling joints 2-6 to adjust the object according to the updated TCP. Figure 4 Block 408. The control system 200 can further update the object's TCP based on the adjustment. In block 520, the control system 200 can compare the adjustment result with a predetermined reference to determine whether to perform another iteration of pose adjustment (e.g., via blocks 506, 508, 510, 512, joints 2-6, and 520) or terminate the loop. Figure 4 Process 400 (e.g., in blocks 408 and 410) provides an exemplary embodiment for determining convergence based on the target TCP or desired pose of the object. In block 522, when the control system 200 determines that the adjustment result has converged to a predetermined reference, the control system 200 continues to generate a refined robot TCP corresponding to the object for final placement.
[0115] In block 530, when the control system 200 determines that the fine-grained TCP of the object (in block 522) is ready and joint 1 has completed its movement from the pick-up workspace 502 to the placement workspace 504, the control system 200 continues to perform precise placement of the object, see the exemplary embodiment of placing an object based on fine-grained TCP. Figure 4 Block 410 in the middle.
[0116] In some examples, the robotic system 100 can train an ML / AI model to perform... Figure 4 The process 400 and / or shown Figure 5 The process shown is 500.
[0117] During the training phase, the robot system 100 can start from the ideal pick-up point, the ideal TCP (Portable Interceptor) of the object, and / or the ideal placement point, which require only minor adjustments to the object (if any). These ideal values can be stored in the memory of the control system 200. In some cases, the object can be manually placed on the gripper to ensure that the object is picked up at the ideal pick-up point. Small errors (e.g., offset from the ideal pick-up point) can be introduced to allow the robot system 100 to learn to adjust the object.
[0118] During the training of the robot system 100, the control system 200 uses the image capture device 112 to capture one or more images of an object picked up at an ideal pickup point. In this way, the control system 200 knows the ideal location of the object and its appearance (e.g., how it looks in the captured images). The control system 200 can compare the updated TCP of the object with the target TCP (e.g., the ideal TCP) to determine convergence. Additionally and / or alternatively, the control system 200 can compare the updated pose of the object with the ideal pose (e.g., the desired pose of the object) to determine convergence. For example, the control system 200 can apply a visual detection algorithm to determine the degree of matching between the updated pose captured by the image capture device 112 and the ideal pose captured in the ideal image. The ideal image is an image captured by the image capture device 112 at a predetermined viewpoint, showing the object in the desired pose. Then, based on the updated TCP / pose corresponding to the object, the control system 200 can learn to control the robot 110 to place the object at the ideal placement point (e.g., ...). Figure 3B The insertion slot 390 shown is illustrated. The control system 200 can learn from the difference between the ideal placement point and the actual placement point to update the learnable parameters in the ML / AL model, where a suitable loss function can be used.
[0119] While embodiments of the invention have been illustrated and described in detail in the accompanying drawings and the foregoing description, such illustrations and descriptions should be considered exemplary rather than restrictive. It should be understood that changes and modifications can be made by those skilled in the art within the scope of the following claims. Specifically, the invention covers other embodiments having any combination of features of the different embodiments described above and below. For example, various embodiments of kinematic, control, electrical, mounting, and user interface subsystems can be used interchangeably without departing from the scope of the invention. Furthermore, the descriptions characterizing the invention herein refer only to embodiments of the invention, and not all embodiments.
[0120] The terms used in the claims should be interpreted as having the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the articles “a” or “the” when referring to an element should not be interpreted as excluding multiple elements. Similarly, the expression “or” should be interpreted as inclusive, so the expression “A or B” does not exclude “A and B” unless the context or the foregoing description clearly indicates that it refers to only one of A and B. Furthermore, the expression “at least one of A, B, and C” should be interpreted as one or more of a set of elements consisting of A, B, and C, and should not be interpreted as requiring at least one of each of the listed elements A, B, and C, regardless of whether A, B, and C belong to the same category or other categories. In addition, the expressions “A, B, and / or C” or “at least one of A, B, or C” should be interpreted as including any single entity of the listed elements, such as A, any subset of the listed elements, such as A and B, or the entire list of elements A, B, and C.
Claims
1. A method for controlling a robot comprising a plurality of joints, the plurality of joints comprising a first joint and one or more other joints, the first joint rotating the entire robot along a fixed base, the method comprising: moving the robot from a first workspace to a second workspace based on controlling the first joint to rotate the robot along the fixed base, wherein the robot is gripping an object and the object corresponds to a first tool center point (TCP); obtaining one or more images of the object using an image capture device while the robot is being moved from the first workspace to the second workspace, wherein the image capture device is mounted on a link connected to the first joint and moves with the first joint of the robot; determining a second TCP based on the one or more images; controlling the one or more other joints of the robot to adjust the object gripped by the robot according to the second TCP; and placing the object at a target location in the second workspace according to a refined TCP based on the adjustment converging to a predetermined reference.
2. The method of claim 1, further comprising: determining that the adjustment did not converge to the predetermined reference; obtaining one or more other images of the object using the imaging device while the robot is being moved from the first workspace to the second workspace; determining a third TCP based on the one or more other images; controlling the one or more other joints of the robot to adjust the object gripped by the robot according to the third TCP; and placing the object at the target location in the second workspace according to a refined TCP after a new adjustment converges to the predetermined reference.
3. The method of claim 1, wherein placing the object at the target location in the second workspace according to a refined TCP based on the adjustment converging to a predetermined reference, further comprises: determining that the adjustment converged to the predetermined reference; determining a third TCP of the object based on the second TCP and the adjustment to the object; obtaining one or more images of the object after the adjustment using the image capture device; and generating a refined TCP of the object based on the third TCP and the one or more images captured after the adjustment.
4. The method of claim 1, further comprising: processing the one or more images to obtain a viewpoint image; and performing pose estimation based on the viewpoint image, wherein determining the second TCP based on the one or more images is based on the pose estimation.
5. The method of claim 4, wherein processing the one or more images to obtain the viewpoint image further comprises: cropping the one or more images based on the object captured in the one or more images; and removing a background from the cropped one or more images based on the one or more images obtained by the imaging device.
6. The method of claim 1, wherein the one or more images of the object are obtained using the image capture device while the robot is moving from the first workspace to the second workspace, further comprising: controlling movement of the one or more other joints of the robot to allow the image capture device to capture images of the object from different viewpoints.
7. The method of claim 1, wherein the first TCP is associated with a relative coordinate system established based on a field of view of the image capture device.
8. The method of claim 7, wherein the adjustment result comprises an updated TCP of the object, wherein the predetermined reference comprises a target TCP of the object, wherein the target TCP of the object corresponds to a predetermined point on the object and is defined by coordinates in a relative coordinate system corresponding to the image capture device, and wherein the adjustment result converging to the predetermined reference is based on the updated TCP of the object converging to the target TCP of the object.
9. The method of claim 7, wherein the adjustment result comprises an updated pose of the object, wherein the predetermined reference comprises a desired pose of the object, and wherein the adjustment result converging to the predetermined reference is based on the updated pose of the object converging to the desired pose of the object.
10. The method of claim 9, wherein the convergence between the updated pose of the object and the desired pose of the object is determined based on an image containing the updated pose of the object and an image containing the desired pose of the object.
11. The method of claim 7, wherein determining the second TCP based on the one or more images further comprises: determining a pose of the object in the relative coordinate system based on the one or more images; determining an offset between the first TCP and the pose of the object in the relative coordinate system; and determining the second TCP based on the first TCP and the offset.
12. The method of claim 7, wherein determining the second TCP based on the one or more images further comprises: determining a pose of the object in the relative coordinate system based on the one or more images; obtaining one or more parameters of the object associated with a position and an orientation of the object in the relative coordinate system; and determining the second TCP based on the one or more parameters of the object.
13. The method of claim 1, wherein the one or more images obtained by the image capture device record motion of the object caused by the one or more other joints, the recorded motion of the object being independent of motion of the first joint.
14. The method of claim 1, wherein (i) operation of the first joint and (ii) operation of the image capture device and the one or more other joints are controlled in parallel. 15. A robotic system comprising: a robot comprising a plurality of joints, the plurality of joints comprising a first joint and one or more other joints, the first joint rotating the entire robot along a fixed base; an image capture device mounted on a link connected to the first joint, the image capture device configured to move with the first joint of the robot and obtain one or more images of an object held by the robot; and a control system configured to: move the robot from a first workspace to a second workspace based on controlling the first joint to rotate the robot along the fixed base, wherein the robot holds the object and the object corresponds to a first tool center point (TCP); obtain one or more images of the object using the image capture device while the robot is being moved from the first workspace to the second workspace; determine a second TCP based on the one or more images; control the one or more other joints of the robot to adjust the object held by the robot according to the second TCP; and control the robot to place the object at a target location in the second workspace according to a TCP refined based on a result of the adjustment converging to a predetermined reference.
16. The robotic system of claim 15, wherein the control system is further configured to: determine that the result of the adjustment does not converge to the predetermined reference; obtain one or more other images of the object using the imaging device while the robot is being moved from the first workspace to the second workspace; determine a third TCP based on the one or more other images; control the one or more other joints of the robot to adjust the object held by the robot according to the third TCP; control the robot to place the object at the target location in the second workspace according to a refined TCP after a new result of the adjustment converges to the predetermined reference.
17. The robotic system of claim 16, wherein the control system is further configured to: determine that the result of the adjustment converges to the predetermined reference; determine a third TCP of the object based on the second TCP and the adjustment to the object; obtain one or more images of the object after the adjustment using the image capture device; and generate a refined TCP of the object based on the third TCP and the one or more images captured after the adjustment.
18. The robotic system of claim 15, wherein the control system is further configured to: process the one or more images to obtain a view point image; and perform pose estimation based on the view point image; wherein the control system is configured to determine the second TCP based on the pose estimation. 19. The robotic system of claim 15, wherein the one or more images obtained by the image capture device record motion of the object caused by the one or more other joints, the recorded motion of the object being independent of motion of the first joint.
20. The robotic system of claim 15, wherein the control device is further configured to control (i) operation of the first joint and (ii) operation of the image capture device and the one or more other joints in parallel.