Method and device for demonstration-based robot programming with adaptive reference frames
By allowing operators to define reference frames based on object positions during execution, the method and device improve motion accuracy in demonstration-based robot programming, addressing environmental discrepancies and ensuring precise object manipulation.
Patent Information
- Application Number
- PCT/EP2024/066288
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-12-18
AI Technical Summary
Demonstration-based robot programming faces challenges in maintaining motion accuracy due to discrepancies between the demonstration environment and the execution environment, particularly when equipment is displaced or workpieces change shape or appearance.
A method and device that allow an operator to define a reference frame based on an object's position during execution, using video and operator input to generate a robot program with adaptive reference frames, enabling improved motion accuracy by aligning the robot's actions to the actual environment.
Enhances motion accuracy by allowing the robot to adapt its reference frames to the current environment, ensuring precise manipulation of objects despite changes, and providing operators with control over reference frame definition and duration.
Smart Images

Figure EP2024066288_18122025_PF_FP_ABST
Abstract
Description
METHOD AND DEVICE FOR DEMONSTRATION-BASED ROBOT PROGRAMMING WITH ADAPTIVE REFERENCE FRAMESTECHNICAL FIELD
[0001] The present disclosure generally relates to the field of robotic control, and specifically to demonstration-based programming of an industrial robot. More precisely, methods and devices are proposed herein which generate a robot program with motion commands that are expressed in a reference frame defined by the position of an object at the time of execution.BACKGROUND
[0002] Emerging markets for robotic applications, such as laboratory automation and food preparation, are characterized by the presence of less structured environments and users with no prior programming experience. To provide the necessary programs for these novel applications, the demand for programming interfaces that are relatively easy to use and / or to start using is expected to increase. This includes programming interfaces adapted for demonstration-based teaching.
[0003] Motion accuracy is a challenge in the demonstration-based programming paradigm. An important error source maybe discrepancies between the demonstration environment, in which the programming was carried out, and the environments where the robot program is later executed. The discrepancies may include that equipment in the environment has been displaced, that a workpiece has changed its shape or appearance, or the like.
[0004] US20230211499A1 discloses a method for teaching a welding robot. The robot movements are expressed relative to a predetermined reference plane, which is defined by the welding start point, the welding end point, and a configured angle value representing the spatial orientation of a tool. This is claimed to avoid the need for defining a coordinate system of the tool in advance.
[0005] US20230241773A1 discloses a system for generating a robot program for manipulating a novel object based on a user demonstration on a demonstration object. With an understanding that these objects are of the same category, the system represents the trajectory in a generalizable form, in so-called category space. In one example, the system can express the object poses in a receptacle’s coordinate frame,e.g., a gear’s pose relative to the shaft in the gear insertion task, which thus makes the process generalizable to arbitrary scene configurations regardless of their absolute poses in the world. To address the challenge of having to adopt one constant category-level canonical coordinate frame for different tasks while capturing the geometric variation across object instances, the system uses an attention mechanism to automatically and dynamically select an anchor point p that defines the origin of the category-level canonical coordinate system. It is reported that the attention mechanism allows the system to dynamically anchor the coordinate system to the local attended point, capturing the variation between demonstration and testing objects in scaling and local typology.
[0006] Still within the demonstration-based paradigm, the research paper F. Amadio et al., “Target-Referred DMPs for Learning Bimanual Tasks from Shared- Autonomy Telemanipulation,” 2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids), Ginowan, Japan, 2022, pp. 496-503, [doi: io.no9 / Humanoids53995.2022.10000233] discloses a novel implementation of Dynamic Movement Primitives (DMPs), called Target-Referred DMP (TR-DMP). The authors demonstrate that TR-DMPs improve generalization capacities and overcome a number of known limitations of DMPs, including difficulties to replicate learned manipulation tasks for a target object with a different orientation from the demonstration.
[0007] It would be desirable to provide a demonstration-based robot programming framework that meets the user community’s expectations of ease of use, yet produces robot programs that retain excellent motion accuracy despite changes in the working environment.SUMMARY
[0008] One objective of the present disclosure is to propose methods and devices that allow more efficient and / or more accurate demonstration-based robot programming than is currently possible according to the state of the art. A further objective is to provide a program causing the robot to manipulate a single object with improved motion accuracy. A further objective is to provide a program causing the robot to combine or separate two objects, or to otherwise manipulate the objects jointly, with improved motion accuracy. A further objective is to allow the operator tocontrol or influence which object shall define the reference frame of motion commands included in the robot program. A still further objective is to allow the operator to control or influence for how long each reference frame shall apply.
[0009] At least some of these objectives are achieved by the invention defined by the independent claims. The dependent claims relate to advantageous embodiments of the invention.
[0010] In a first aspect of the present disclosure, there is provided a method of programming an industrial robot, which includes a robot manipulator and robot controller. The method comprises: recording movements of the robot manipulator during a demonstration-based robot programming session, for thereby obtaining a robot trajectory; acquiring a video of the robot programming session; and generating, on the basis of the robot trajectory and the video, a robot program executable by the robot controller, wherein the robot program includes at least one motion command which is expressed in a first reference frame. According to the first aspect, the method further comprises capturing operator input data indicating a first object in an image in the video. Further, the generated robot program includes a command to identify a position of an object resembling the first object in a work area of a robot which executes the robot program, and the first reference frame is defined with respect to the identified position of the object resembling the first object.
[0011] Because the programming method according to the first aspect accepts operator input data indicating an object that maybe used as a basis for defining a reference frame at the time of execution, the method allows the operator to influence how the motion commands in the robot program will be expressed. Further, because the robot program causes the industrial robot to adapt the reference frame to the object’s actual position during execution, the motion accuracy relative to the object is improved. It is noted that the object resembling the first object can coincide with the first object, or it can be a visually similar object, e.g., any object that passes a preconfigured resemblance test.
[0012] Some embodiments of the programming method according to the first aspect generalize the above-outlined programming technique, in such manner that the robot program includes motions commands which are expressed in different reference frames. Each reference frame is defined by the positions (and optionallyorientations) assumed at the time of program execution by objects which resemble a respective first, second, third etc. object indicated by the operator.
[0013] In further embodiments, alternatively or additionally, the programming technique is generalized to include the reliance on multiple reference frames defined by a time sequence of positions (and optionally orientations) assumed by an object which resembles an object indicated by the operator. This is to say, after a first position of an object which resembles the operator-indicated object has been identified (at time ti during program execution) and used to define a first reference frame, a second position of the same object can be identified (at time t2 during program execution, where t2 > ti) and used similarly to define a second reference frame. Alternatively, after a first position of an object which resembles the operator- indicated object has been identified and used to define a first reference frame, a second position of a different object which resembles the operator-indicated object can be identified and used similarly to define a second reference frame.
[0014] In some embodiments, as already intimated above, the program execution may further include an identification of a pose of the object resembling the user- indicated first object. The reference frame, in which the motion commands are expressed, can then be oriented in accordance with the identified pose of the object, e.g., such that the axes of the reference frame are parallel with the pose or they differ by a predefined rotation from the pose.
[0015] In some embodiments, the operator input data may further include data representing a value of a robot parameter. The robot parameter may be, for example, a state of a tool carried by the robot manipulator, a movement execution parameter, a degree of compliance with the robot trajectory, a degree of movement precision, a reference frame, a drive system parameter. The operator input data may include speech data captured while recording the movements of the robot manipulator, or it may include character data and / or a menu selection through a graphical operator interface.
[0016] In some embodiments, the method according to the first aspect includes a concluding step of enabling the operator to review and improve the robot program. In particular, the step may include executing the robot program while capturing operator speech data including requests to modify the robot program.
[0017] In a second aspect of the present disclosure, there is provided a device for facilitating programming of an industrial robot with a robot manipulator and a robot controller. The device comprises memory and processing circuitry configured to: record movements of the robot manipulator during a demonstration-based robot programming session, for thereby obtaining a robot trajectory; acquire a video of the robot programming session; and generate, on the basis of the robot trajectory and the video, a robot program executable by the robot controller, wherein the robot program includes at least one motion command which is expressed in a first reference frame. According to the second aspect, the programming device further comprising an interface configured to capture operator input data indicating a first object in an image in the video. The robot program includes a command to identify a position of an object resembling the first object in a work area of a robot which executes the robot program, and the first reference frame is defined with respect to the identified position of the object resembling the first object.
[0018] The present disclosure further relates to a computer program containing instructions for causing a computer - or the device according to the second aspect in particular - to carry out the above method. The computer program may be stored or distributed on a data carrier. As used herein, a “data carrier” may be a transitory data carrier, such as modulated electromagnetic or optical waves, or a non-transitory data carrier. Non-transitory data carriers include volatile and non-volatile memories, such as permanent and non-permanent storage media of magnetic, optical or solid-state type. Still within the scope of “data carrier”, such memories may be fixedly mounted or portable.
[0019] In the terminology of the present disclosure, “industrial robot” is used in a broad sense, to cover in particular manufacturing robots, material-handling robots, assembly robots, cutting / welding robots, service robots, collaborative robots, hygiene robots, industrial robot tracks, industrial robot positioners. An industrial robot may be a stationary robot or a mobile robot, such as an automated guided vehicle (AGV), an autonomous mobile robot (AMR) or an autonomous mobile manipulator robot (AMMR). The term industrial robot covers the full range from lightweight robots designed to replace human manual work, over collaborative robots for supporting a human worker, all the way up to heavy-duty robots.
[0020] As used herein, “demonstration-based” robot programming includes a step where the operator performs a portion of a robot task or guides the robot manipulator such that it performs a portion of this robot task. The operator’s activities are imaged or otherwise observed by a programming device. Demonstration-based robot programming includes kinesthetic programming as a special case. The term kinesthetic programming - by allusion to the robot’s proprioceptive ability to sense its own position, pose and movements - is a programming approach in which the operator physically moves the robot manipulator to imitate the desired motions. The terms kinesthetic programming and lead-through programming are synonymous or at least partially overlapping in meaning. The state of the robot during a kinesthetic programming session is typically recorded by means of the robot’s onboard sensors, e.g., joint angles and torques. As used herein, kinesthetic programming also includes teleoperation as a special case, where the movement of the robot manipulator is controlled by an external input to the robot through a joystick, graphical user interface or other input means; kinesthetic programming by teleoperation does not require the operator to be present in the work area of the robot.
[0021] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to “a / an / the element, apparatus, component, means, step, etc.” are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order described, unless this is explicitly stated.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Aspects and embodiments are now described, by way of example, with reference to the accompanying drawings, on which: figure 1 shows a work area, in which a robot manipulator operates under the control of a robot controller, and a programming device supporting demonstration-based robot programming; figure 2 is a flowchart of a method of programming an industrial robot, according to embodiments herein;figure 3 shows an example appearance of a graphical user interface of the programming device; figure 4 illustrates a process of translating a robot trajectory into motion commands in a robot program, wherein a succession of moving and neutral reference frames is used; figure 5 contains four consecutive snapshots of a robot manipulator picking up a receptable and placing it at the top of a pile of similar receptacles; and figure 6 shows an alternative starting configuration for the pick-and-place task illustrated in figure 5.DETAILED DESCRIPTION
[0023] The aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, on which certain embodiments of the invention are shown. These aspects may, however, be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and to fully convey the scope of all aspects of the invention to those skilled in the art. Like numbers refer to like elements throughout the description.System overview
[0024] Figure 1 shows an industrial robot 100 made up of a robot manipulator no and a robot controller 120. The robot manipulator 110 and the robot controller 120 are joined by a wired or wireless bidirectional data connection, which conveys control signals, sensor data etc.
[0025] The robot manipulator 110 includes an arm, which extends from a base 112 and is made up of structural elements 114 and at least one linear or rotary joint 113. The arm may further carry tools 111 that allow it to interact with objects 151 in the form of various workpieces, which are present in a work area 150 of the robot manipulator no. The workpieces are subject to manufacture, processing or handling by the robot manipulator 110. The work area 150 may further include non-workpiece objects 152, such as containers, fixtures, robot positioners, separators, insulators, supports etc. These objects 152 maybe generic or maybe specifically adapted to the workpieces 151 handled by the robot manipulator 110. Apart from wear, staining etc.,they are generally in the same condition at the beginning and end of a work cycle. The arm of the robot manipulator no is movable by action of internal motors, drives or actuators (not shown), and it includes transducers, sensors and other measuring equipment (not shown), from which the manipulator’s no current position, pose, technical condition, load etc. can be derived, to some degree of accuracy. The position, pose etc. of the robot manipulator no may in particular refer to a point on the arm, particularly to a tool-center point (TCP).
[0026] The positions, poses etc. may be expressed with respect to one of multiple possible reference frames, including a first fixed reference frame O0defined with respect to a point in the work area 150, a second fixed reference frame (not shown) defined with respect to the base 112, a first reference frame O defined with respect to an initial position 151 of a workpiece, a second reference frame O2defined with respect to a position of a non-workpiece object 152 in the work area 150, or a third reference frame O3defined with respect to a later position 151.1 of the workpiece. Here, the first fixed reference frame O0is defined with respect to the point in the work area 150 in the sense that its origin is situated in that point. Alternatively, a reference frame maybe defined with respect to a reference point in the sense that its origin is situated a predefined translation from the reference point. The first and second fixed reference frames, as well as any further reference frames which are independent of the positions 151, 152 of objects in the work area 150, maybe collectively referred to as neutral reference frames. All reference frames shown in figure 1 have a common orientation, i.e., their respective first, second and third axes are parallel. Without departing from the scope of the present disclosure, a further option is to use reference frames with mutually different orientations, including a reference frame which is oriented in accordance with a pose of an object 151, 152 in the work area 150 or another suitable reference objects. In particular, one may use a reference frame which has its origin situated in the TCP and is oriented parallel to the tool 111 of the manipulator no at all times.
[0027] The robot controller 120 comprises processing circuitry 121, a memory 122 and a communication interface 123. Example content of the memory 122 during operation may includeS: an operating system, basic settings, software implementing generic movements, sensing, self-monitoring, generically useful functionalities and services (alltypically contributed by an original manufacturer), task- or role-specific configurations, configuration templates (typically contributed by a robot system integrator), and site-specific settings (typically contributed by an end user);C: robot programs or projects causing the robot manipulator no to perform useful or intended tasks in its work area 150.Processes that execute the robot programs may do so in accordance with the program-independent memory content S, e.g., by making calls to available functionalities, libraries, routines or parameter values therein. It is recalled that the memory 122 and processing circuitry 121 of the robot controller 120 maybe distributed and / or contain networked resources having a different physical localization than figure 1 suggests.
[0028] Each of the programs C may contain a plurality of movement instructions relating to locations such as points, poses, paths as well as modulated paths. A program may be a compiled executable (binary) or a script. A movement instruction relating to a modulated path may be expressed as - or may include - a process-on- path instruction. The programs C maybe created by an operator 190 with the aid of a robot programming device (programming station) or a general-purpose computer, or they may be created directly at the robot controller 120 if it has an operator interface. In the first two cases, versions of the programs C may be downloaded to the robot controller 120 over a wired or wireless connection or by being temporarily stored on a portable memory. As will be described below, a robot program may further be included by means of demonstration-based programming according to the teachings herein.
[0029] A dedicated programming device 160 is shown in the right-hand portion of figure 1. From the programming device 160, the created programs C can be transmitted to the communication interface 123 of the robot controller 120 and then stored in the memory 122 where they are available for execution. It is noted that the programming device 160 may be implemented as a set of collaborating components within the robot controller 120. In fact, the components can be shared with the robot controller 120, e.g., by using the processing circuitry 121 for the dual purposes of robot control and programming and / or using the memory 122 for the same dual purposes. In other words, the programming device 160 may constitute a portion of amulti-purpose device; it need not be a standalone device or a device with programming as its sole or main purpose.
[0030] The programming device 160 - which is shown as a standalone device in the non-limiting example of figure 1 - comprises processing circuitry 161, memory 162 with executable software 163, and at least one communication interface 164, 165. The programming device 160 may further comprise at least one operator interface 166, in particular a graphical operator interface or graphical user interface (GUI).
[0031] The programming device 160 further has access to an imaging device (e.g., camera, video camera, depth camera, lidar, radar) 167, by which images or video of the demonstration-based programming session can be captured. In particular, the imaging device may be a depth camera, lidar or radar, which in addition to - or instead of - the two-dimensional appearance of an object determines a depth coordinate of the object, e.g., by time-of-flight measurements, triangulation, reflection or other per se known techniques. An RGB-D camera is an example of a depth camera. The imaging device 167 may be a part of the programming device 160, or it maybe integrated in the industrial robot 100 and optionally be used for other tasks as well.
[0032] The programming device 160 is optionally enabled to capture speech data, which could represent utterances or narration by the operator 190 during the programming session. The programming device 160 may for this purpose utilize one or more acoustic transducers (microphones) 168 arranged in the vicinity of the operator’s 190 normal position. Alternatively, the speech data maybe captured by a portable microphone or a headset worn by the operator 190. To parse speech data, the programming device 160 may optionally utilize a natural-language model 170, which may be either stored internally or - as in the example configuration shown in figure 1 - may be available from an external server or host computer.
[0033] The natural-language model 170 may constitute a large language model (LLM), that is, a type of machine-learning algorithm which has been trained on very large datasets using deep learning techniques to be able to perform natural-language processing (NLP) tasks. Example NLP tasks are recognizing, summarizing, translating, predicting and generating plausible textual content. A very large dataset in this sense may include of the order of one million parameters, such as tens of millions of parameters, such as hundreds of millions of parameters. An LLM mayhave a transducer architecture, particularly a transducer architecture with four cascaded key links when it processes the input data, namely: word embedding, position encoding, self-attention mechanism, feedforward neural network. At the time of filing this disclosure, noteworthy example LLMs include Bidirectional Encoder Representations from Transformers (BERT), Bard, BLOOM, Claude 2, various versions of Generative Pre-trained Transformer (GPT), Llama, PaLM 2, RoBERTa, T5, LaMDA, Turing NLG. LLMs include as a special case multimodal models, such as Gemini. One benefit of an LLM is that the vocabulary is practically open-ended. The operator 190 can start using it without prior training. The demonstration-based programming can be carried out without requiring the operator to be extremely focused on using the right command words (or avoiding them) and / or syntax.
[0034] During a demonstration-based programming session, the programming device 160 has access to position data representing an actual position, a recorded position or recorded movements of the robot manipulator no, as well as images or video of the robot programming session. The position data may for example be obtained through the intermediary of the robot controller 120, which monitors position data in the normal course of its operation; alternatively, the programming device 160 is granted access to corresponding signals from the transducers, sensors or other measuring equipment in the robot manipulator no.Programming method
[0035] Turning to the flowchart in figure 2, embodiments of a method 200 of programming an industrial robot 100 by demonstration-based robot programming are illustrated. Not all steps shown in figure 2 are necessarily carried out in all embodiments. A possible output of the method 200 is a robot program (or robot routine) C with motion commands expressed adaptively in different reference frames. At least one of the reference frames is defined with respect to the identified position of an object at the time of execution, but there may as well be a fixed reference frame or another neutral reference frame. The robot program C maybe executable only by the robot controller no of the industrial robot 100 for which the programming method 200 was carried out. In some embodiments of the method 200, the robot program C can be executed also by further robot controllers 120 of the same model as- or compatible with - the one for which the programming method 200 was carried out.
[0036] The method 200 may be considered as a description of the behavior which the programming device 160 is configured to carry out. The method 200 may as well correspond to the behavior of the robot controller 120 in a special programming mode, which the robot controller 120 can be requested to enter. Instructions for causing a computer - or the programming device 160 in particular - to carry out the method may be provided in the form of a computer program 164 (see figure 1).
[0037] A first step 210 of the method 200, which is executed during a demonstration-based robot programming session, includes recording movements of the robot manipulator no. On the basis of the recorded movements, a robot trajectory can be obtained, e.g., by combining the recorded movements. The robot trajectory may indicate the robot manipulator’s 110 position and / or pose (e.g., in Cartesian or joint-space coordinates) as a function of time. The position or pose need not be indicated for all points in time. Alternatively, the trajectory may be expressed as a number of reference times at which the robot manipulator no is to assume corresponding setpoint positions or setpoint poses, wherein the robot manipulator 110 is free to have arbitrary positions or poses in the intervals between the reference times.
[0038] As explained initially, the demonstration-based programming session may for example be a kinesthetic programming session. Kinesthetic programming may include phases where the operator 190 moves the robot manipulator no into desired poses and positions by direct manual manipulation. Kinesthetic programming may as well include at least some phases of teleoperation, where the operator 190 causes the robot manipulator 110 to perform such movements by providing an external input through a joystick, graphical user interface or other input means connected to the robot controller 120 or the programming devices 160. The trajectory may correspond to a state-space representation, such as joint angles as a function of time. Alternatively, in other use cases, it may suffice to express the trajectory in terms of the position of a reference point on the robot manipulator no as a function of time.
[0039] It is important to note that the motion commands in the robot program C are not necessarily configured to make the industrial robot replicate the robot trajectory identically. For example, if the object resembling the first object at the time of execution occupies a different position than the first object during thedemonstration-based robot programming session, then clearly the industrial robot will need to visit different positions and poses (or different coordinates in joint space or in general state space) to complete the demonstrated task at runtime.
[0040] In a second step 211, a video (video recording, video sequence) of the robot programming session is acquired. The video may be acquired by means of the imaging device 167. Each frame of the video may be a visual image, a depth image or a combination of these. For the purposes of the present method 200, it is not essential that the video have a high frame rate (e.g., a frame rate which is capable of capturing naturally-looking body movements). Rather, the video maybe a sequence of sparse, even irregularly spaced, images of the work area 150.
[0041] In a third step 212, operator input data is captured. This input data indicates, according to the operator’s 190 intention, a first object 151 which appears in a frame of the video acquired in step 211. The operator input data may include speech data captured while recording the movements of the robot manipulator. The operator input data may include character data and / or a menu selection, which is optionally received via a graphical operator interface 166. For example, the operator 190 may provide said operator input data by indicating an approximate location or approximate extent of the first object 151 in a current frame of the video while the video is being acquired.
[0042] Optionally, the operator input data may include an approximate location of the first object 151, such as a right or left half of the work area 150, a sector or quadrant of the work area 150. This information maybe included in the robot program C to help the executing robot controller 120 distinguish between objects which have a similar appearance but are expected to be located in different parts of the work area 150 during execution.
[0043] In particular, the operator input data may include an object mask corresponding to the first object 151, wherein the object mask refers to a respective frame (image) in the video. As illustrated in figure 3, the operator may be shown a number of proposed masks 301, 302 in a graphical operator interface 166, and the operator is asked to select the desired one. In the example of figure 3, a first mask 301 corresponds to the first object 151, whereas a second mask corresponds to the tool 111 of the robot manipulator. As used herein, a “mask” or “object mask” of a visual object in an image may include a vector-based or bitmap-based representation from which itis derivable what regions of the image correspond to the object. In a depth image (e.g., depth-video frame), the mask maybe expressed in the plane of the image, that is, independently of the depth component. This to say, the mask generally does not specify the extent of the physical object in the depth direction.
[0044] The masks 301, 302 may be generated based on the video frame by means of a video object segmentation (VOS) algorithm. The VOS sub-algorithm may for example include elements of one or more of the XMem architecture (see H. K. Cheng et al., “XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model”, arXiv: 2207.07115 [cs.CV], retrieved from arXiv.org), the XMem++ architecture (see M. Bekuzarov et al., “XMem++: Production-level Video Segmentation From Few Annotated Frames”, arXiv:23O7.i5958 [cs.CV], retrieved from arXiv.org), the Track Anything architecture (see J. Yang et al., “Track Anything: Segment Anything Meets Videos”, arXiv: 2304.11968 [cs.CV], retrieved from arXiv.org), the Cutie architecture (see H. K. Cheng et al., “Putting the Object Back into Video Object Segmentation”, arXiv: 2310.12982 [cs.CV], retrieved from arXiv.org), or it maybe an algorithm using reference-guided mask propagation (see S. W. Oh et al., “Fast Video Object Segmentation by Reference-Guided Mask Propagation”, Proceedings of the IEEE conference on computer vision and pattern recognition 2018, pp. 7376-7385) or an algorithm using a space-time memory network (see S. W. Oh et al., “Video Object Segmentation using Space-Time Memory Networks”, Proceedings of the IEEE conference on computer vision and pattern recognition 2018, pp. 9226-9235; and H. K. Cheng et al., “Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation”, arXiv:2io6.05210 [cs.CV], retrieved from arXiv.org). Reference is made to the corresponding sections of the applicant’s prior disclosure PCT / EP2024 / 061921.
[0045] In a next step 214, a robot program C is generated on the basis of the robot trajectory and the video. The robot program C is executable by the robot controller 120. The robot program C includes at least:- a command to identify a position of an object resembling the first object (i.e., the object identified by the operator input data captured in step 212) in the work area 150 of the industrial robot 100 which executes the robot program; and- a motion command which is expressed in a first reference framewherein the first reference frame is defined with respect to the identified position of the object resembling the first object.
[0046] From these, the command to identify the position of an object resembling the first object can include or refer to information describing the appearance of the first object, such as one or more images from the same or different angles. In particular, the command may include or refer to a storable data item representing the first object, which has been obtained by applying a composite pose-estimation algorithm to the video, as described in detail in the applicant’s earlier disclosure PCT / EP2024 / 061921. The composite pose estimation algorithm maybe executed locally or on a processing resource outside the programming device 160, such as a networked (‘cloud’) processing resource. The pose estimation algorithm may be implemented as a combination of a video-object segmentation (VOS) sub-algorithm configured to determine a mask of a visual object in an image, and an object -pose tracking (OPT) sub-algorithm configured to track a pose of a visual object over multiple depth-video frames. Advantageous implementations of the VOS subalgorithm were reviewed above in relation to step 212. The OPT sub-algorithm, for its part, may include elements of the BundleTrack architecture (see B. Wen et al., “BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category- Level 3D Models”, Proceedings of 2021 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROSf) or BundleSDF architecture (see B. Wen et al., “BundleSDF: Neural 6-D0F Tracking and 3D Reconstruction of Unknown Objects”, arXiv: 2303.14158 [cs.CV], retrieved from arXiv.org). A generic composite pose-estimation algorithm has been described in the research paper M. Sun et al., “Instance-Agnostic Geometry and Contact Dynamics Learning”, arXiv:23O9.05832 [cs.CV], retrieved from arXiv.org.
[0047] It is not essential for the present invention what method or algorithm is used by the robot controller 120 to identify the position of the object resembling the first object at the time of executing the robot program C. Accordingly, the resemblance test maybe defined freely to suit the use case under consideration. In particular, if the command to identify the position of an object resembling the first object includes or refers to a storable data item of the type described inPCT / EP2024 / 061921, then the composite pose-estimation algorithm is the preferred option to use at the time of execution.
[0048] It is assumed that, before execution of this step 214, the trajectory recorded in step 211 - typically a geometric description or a joint-space description the video of the robot programming session acquired in step 211, and the operator input data captured in step 212 are available. The step 214 may include converting the trajectory into a sequence of robot commands that cause the robot manipulator no to realize the trajectory, while applying at least some of the objects identified by the operator input data as a reference frame of the objects. The robot commands maybe selected from a predefined set of robot commands executable by the robot controller 120. The robot commands may for example be compliant with the RAPID™ robot programming language. In the RAPID™ robot programming language, example motion commands are MoveL and MoveJ, and the reference frame may be included as a value of parameter wobjdata. Geometric input parameters to the motion commands (e.g., po, pi, p2 in the snippet in Table 3) may need to be converted in such manner that they are expressed in coordinates of the chosen reference frame. The step 214 may include a first substep of sampling the trajectory into a sequence of discrete points and a second substep of selecting robot commands that cause the robot manipulator no to move between each pair of consecutive discrete points. The sampling used in the first substep may be time-uniform sampling (constant step duration), space-uniform sampling (constant step length), or a non-uniform sampling algorithm with controlled deviation, such as Ramer-Douglas-Peucker. In a RAPID™ environment, the second substep may be performed so as to output instances of the command MoveL (cartesian linear motion) or MoveJ (joint-space linear motion) or a combination of these. Optionally, the step 214 may include a postprocessing substep applied to the sequence of generated robot commands, such as formatting the sequence into a predefined script format by appending a header, performing a consistency check, or the like.
[0049] As already pointed out, it is not essential for the sequence of robot commands to replicate the trajectory exactly. Rather, deviations are to be expected when if objects resembling the first, second etc. objects occupy different positions at the time of execution than during the demonstration-based programming session.
[0050] In other words, by clicking (operator input data) on a touchscreen, activating a “capture” button or using speech input during the demonstration-based programming session while video is being recorded, the operator informs the programming device 160 to use the first object 151 at its current position as anchor for the first reference frame. Conceptually, as a result of the operator’s 190 actions, the first reference frame is frozen. After the demonstration-based programming session, the generated RAPID™ code will make the robot go to the approximate position where the video capture was done at that time, look for the first object with the pose estimation algorithm, update the reference frame in the code so that it uses the current pose of the detected object, and replicate the motions done during the demonstration but now with respect to this updated reference frame.
[0051] The operator can as well express a preference or an instruction to use a fixed reference frame. Because a fixed reference point is by definition non-moving, it does not need to be imaged like movable objects do. When the operator prefers for some portion of the future robot program C to use a fixed (or another neutral) reference frame, the operator has the option of indicating this explicitly to the programming device 160 by speech input (e.g., consistently saying “with respect to” to describe motions) or a click on a corresponding button in an operator interface 166. The second option may be relatively more time consuming.
[0052] The choice of reference frames for the different motion commands in the robot program may be fully or completely automated. In particular, a large language model (LLM) or another natural language processing tool could be utilized to extract from speech input which reference frame to use. In an experiment described in the applicant’s prior disclosure PCT / EP2023 / 087949, GPT-3.5-turbo, an LLM, was prompted to segment certain timestamped text transcribed from the operator’s 190 utterances into phases, each phase annotated with a start and end time, to identify values of specified robot parameters, including reference^frame, and finally to format the output in accordance with the YAML format, which was fed to a program code generation substep. In the experiment, a prompt including the text in Table 1 was given to the LLM.To implement the teachings herein, a prompt including the text in Table 2 maybe used. It differs from Table 1 in that it instructs the LLM to take the operator input data into account in the decision-making concerning reference frames. See also Example 3.
[0053] In such implementations of the method 200 where the LLM is sole responsible for choosing the reference frames, the operator during the demonstration only has to worry about capturing images (e.g., video frames) of the objects before interacting with them, to make sure that the pose estimation algorithm can detect the pose of the object and that this reference frame is stored as part of the demonstration data. The LLM then decides, based on learned or explicitly prompted criteria, which reference frame to use for different parts of the robot program. Naturally, the LLM can use the knowledge that the operator captured pictures at different moments to steer the selection of reference frames. For example, if the operator captures a picture of an object 151 at time ti of the demonstration, it is most likely that the operator will interact with the object 151 in the next part of the demonstration, unless speech input or other input is explicitly saying otherwise.
[0054] It is not essential for the present invention what method or algorithm is used by the robot controller 120 to identify the position of the object resembling the first object at the time of executing the robot program C. Accordingly, the resemblance test may be defined freely to suit the use case under consideration. The resemblance test maybe, for example, whether a particular object recognition algorithm provides a positive outcome (e.g., ‘object recognized’, ‘position detected’, ‘pose detected’). In particular, if the command to identify the position of an object resembling the first object includes or refers to a storable data item of the type described in PCT / EP2024 / 061921, then the composite pose-estimation algorithm is the preferred option to use at the time of execution.
[0055] Some further developments of the robot programming method 200 may generate a robot program C which further includes- a command to identify a later position 151.1 of the object resembling the first object in the work area 150 of the robot which executes the robot program; and- at least one motion command which is expressed in a third reference frame O3defined with respect to the identified later position of the object resembling the first object.The command to identify the later position may include or refer to information describing the appearance of the first object or a storable data item of the type described in PCT / EP2024 / 061921. It maybe sufficient to refer back to preexisting information or a preexisting storable data item referring to the first object. However, in such embodiments where the robot program C includes an approximate location of the first object 151, to help the executing robot controller 120 distinguish between objects which have a similar appearance but are normally located in different parts of the work area 150 during execution, the first object 151 should be imaged anew at its new location.
[0056] In some further developments of such embodiments of the robot programming method 200 where the operator has full or partial control over the choice of reference frame, the captured operator input data further determines validity periods of the first and any further reference frames. For example, the operator input data may explicitly indicate a start time, an end time and / or a duration of validity. Such times maybe expressed in terms of program execution time, clock time, section of program, command line number etc. Further, the operator input data may implicitly indicate an end time if, say, the program generation step 214 applies a rule to the effect that a reference frame shall be used only up to a point in time where a next object (second object) in the video is indicated or a preexisting object (first object) is re-indicated. As another example of an implicit indication, the validity time of a reference frame maybe bounded above by a preconfigured value, such that it stops being applied in the robot program C after expiry of the preconfigured validity time counting from capturing the operator input data.
[0057] If a validity time has been defined in this way, the program generation step 214 may include a substep 214.1 of identifying any motion command which is expressed in a reference frame outside its validity period and a further substep 214.2 of transforming the motion command into an equivalent motion command which is expressed in a neutral reference frame O0.
[0058] In some further developments, the robot programming method 200 comprises a further step 213 of applying a pose-estimation algorithm to the acquiredvideo, for thereby obtaining a storable data item representing the first object. The storable data item of the type described in PCT / EP2024 / 061921, which is suitable for use in initializing an instance of the pose-estimation algorithm by which the position of the object resembling the first object is to be identified during execution of the robot program. Optionally, the pose-estimation algorithm is furthermore used to identify a pose of the object resembling the first object. Then, the robot program may include a command to identify a pose of the object resembling the first object, and the first reference frame O is oriented in accordance with the thus identified pose of the object. The first reference frame may for example be oriented such that the axes of the reference frame are parallel with the pose or they differ by a predefined rotation from the pose.
[0059] In some further developments of the robot programming method 200, further operator input data is captured which includes data indicating at least one robot parameter value of at least one robot parameter, and the robot trajectory is annotated with the at least one robot parameter value. Techniques for annotating different phases of a robot trajectory recorded during a demonstration-based robot programming session with robot-parameter values that are parsed from speech data are described in the applicant’s earlier disclosure PCT / EP2023 / 087949. The robot parameter may be one or more of: a state of a tool 111 carried by the robot manipulator, a movement execution parameter, a degree of compliance with the robot trajectory, a degree of movement precision, a reference frame, a drive system parameter. This robot-parameter values in the annotations may be included in the robot program C in the step 214.
[0060] In some further developments, the robot programming method 200 comprises a step 215 of executing the generated robot program C while capturing operator speech data (substep 215.1). The operator speech data may include one or more requests to modify the robot program (substep 215.2), e.g., to use a different reference frame in a certain portion of the robot program C, to change the speed of the robot manipulator no, to limit the acceleration of the robot manipulator 110 etc. In substep 215.1, the operator speech data maybe captured using the transducer 168, and it maybe processed by the LLM 170 discussed above. This step 215 gives the operator 190 an opportunity to devote full visual attention to the execution of robot program C visually and request adjustments in connection with actions or phasesthereof which are not satisfactory. To summarize, in some of the embodiments outlined above, particularly combinations thereof, the operator input may relate to all of the following:- robot parameters,- narration of the tasks performed,- (re)identification of first, second etc. objects, and it may take the form of speech, character data, menu selections and other GUI manipulations.Example 1
[0061] Figure 4, where the rightward horizontal direction represents time, illustrates how different sources of information are combined into a robot program C during an execution of the method 200 of figure 2. The top line represents the recorded movements of the robot manipulator no during the demonstration-based robot programming session. The lowest line represents operator input data items provided by the operator 190 which indicate objects 151, 152 in the industrial robot’s 100 work area 150. The objects 151, 152 are assumed to be visible in images acquired by the imaging device 167.
[0062] When operator input data is captured which causes the method 200 to use a reference frame related to a first object 151 (it is recalled that the operator 190 may control this choice independently, or the choice may be partially or fully automated, e.g., delegated to an LLM), information sufficient to identify a position of an object resembling the first object 151 during execution is extracted from the image, and a first reference frame O defined with respect to the identified (future) position of the object resembling the first object is used for expressing subsequent motion commands in the robot program C. The first reference frame O is used until the reference frame is to be changed. In the example illustrated in figure 4, the change occurs when the operator provides input data - or an LLM decides - to use a neutral reference frame O0. This may imply that the subsequent motion commands are expressed in the neutral reference frame O0. Generally, no information needs to be extracted from the current image of the work area 150. The next change of reference frame in figure 4 occurs as a result of operator input data that identifies a second object 152. As a result, information sufficient to identify a position of an objectresembling the second object 152 during execution is extracted from the image, and a first reference frame O2defined with respect to the identified (future) position of the object resembling the second object 152 is used for expressing subsequent motion commands in the robot program C. It is noted that the introduction of a second reference frame into the robot program C may correspond to a repetition of at least step 212.Example 2
[0063] Figures 5 and 6 will further illustrate the technical benefits of the described robot programming method 200 in the context of a pick-and-place task. Figure 5 contains four consecutive snapshots of a robot tool 111 picking up a receptable (e.g., a bowl) 153, lifting and lowering it, and placing it at the top of a pile 154 of similar receptacles. Because the receptacle 153 may assume different positions in the work area 150 during different executions, for various random reasons, it is preferable to express the motion in a reference frame defined by the actual position of the receptable 153. A fixed reference frame (e.g., defined with respect to the robot base 112) would be affected by any inaccuracies relative to the programming session.
[0064] The pick-and-place task may correspond to the RAPID™ program code in Table 3, where wobj_bowl refers to a reference frame anchored in the receptable 153, wobjo refers to the fixed reference frame of the robot base 112, and where po, ..., p2 are defined robtargets. The robtarget type in RAPID™ represents a position and orientation point specified for the TCP (tool center point) of the robot.It is notable that only two motion commands of this code use the workpiece (receptable 153) to redefine the first two motions, since they are the only two motions that need to be expressed relative to the position of the bowl to pick up on the table. Here, the dropping of the receptable 153 from point p2 and onto the pile 154 is not treated as a movement requiring additional accuracy. The motion of the bowl upwards to gain clearance could be done with respect to “wobjo” instead, but in this case the bowl will not be moved directly upwards from the position it was picked up, but rather to a fixed point in the robot workspace instead.
[0065] Figure 6 shows an alternative starting configuration for the pick-and-place task - corresponding to the first snapshot in the sequence of figure 5 - where the receptable 153 is slightly shifted to the right in the work area 150. Because the actual position of the receptable 153 will be identified during execution, the robot program C will still execute with a motion accuracy equivalent to that achieved with the starting configuration according to figure 5.
[0066] As mentioned above, according to some embodiments of the method 200, an approximate location of the first object 151 may be included in the robot program C to help an executing robot controller 120 distinguish between objects which have a similar appearance. This option may be relevant in the example of figures 5-6, where the receptable 153 and the pile of receptables 154 may have a somewhat similar visual appearance. Confusing the receptable 153 with the pile 154 during execution, so that the reference frame is instead defined with respect to the pile of receptables 154, would be an unwanted scenario. This scenario could be easily avoided by including an instruction that the pile 154 is normally located in the rear third of the work area 150, that the pile 154 is relatively closer to the rear than the receptable 153, or the like.ExampleIn an experiment, the prompt according to Table 4 was given to GPT4, an LLM. Step 3.8 of Table 4 is identical to the snippet presented above in Table 2.Here, the final paragraph (“{"words": [{"text": "I", "start": 1.7, "end": 1.84}, {"text": "will", "start": 1.84, "end": 1.98}, {"text": "teach", "start": 1.98, "end": 2.2}, ...”) is the beginning of a timestamped transcription of words uttered by the operator 190 during the demonstration-based programming session. While the rest of the transcription is intentionally omitted here for brevity, a full example is found in PCT / EP2023 / 087949. The prompt according to Table 4 led to the YAML output in Table 5.It is notable that the reference frame “bowlo” is only used after the reference frame was defined by the image capture at time 11.33 s, which can be seen by the identified start and end times.
[0067] To generate RAPID™ code from this YAML code, along the lines of the above discussion, the following substeps maybe executed:- For each step in the YAML file, the recorded reference frames are used to transform the recorded trajectory into relative motions with respect to these reference frames.- A work object is defined for each reference frame that was recorded during the demonstration (e.g., “wobj_bowlo”, “wobj_bowli”, “wobj_cubeo”). The initial values of these reference frames are not important, as they will be updated at runtime.- In the generated RAPID routine, a camera capture routine is included that specifies which object needs to be identified, and which work object needs to be updated as a result. The name of the object that needs to be identified can be retrieved from the label of the recorded captured objects by removing the numbers at the end (e.g., “bowlo” represents object “bowl”). In the RAPID code the call to the camera capture routine is included in such manner that it happens between movement commands according to the recorded timestamp from the demonstration. To illustrate, “bowlo” was identified at time 11.33 s, so that the call to the camera capture routine is inserted at aposition of the program code causing it to occur between motions happening before and after that timestamp.- After the camera capture routine, the work object can be used to describe relative motion according to the first substep.Optionally, the YAML code is initially preprocessed by ensuring that only reference frames are defined after they have been identified by the camera. Since the point in time at which the object defining a certain reference frame was identified during the demonstration-based programming session, the code can be searched for motion commands that refer to reference frame earlier than they were identified (substep 214.1 of method 200). It has been observed that such motion commands could sometimes be generated by LLMs, especially early releases of LLMs, which lack a perfect ability to enforce this type of consistency; it is recalled that the basis of an LLM is to create patterns mimicking how words occur in a language. Those steps in which the reference frame was set before identified, can be modified such that they refer to either a neutral reference frame (e.g., “robot_base”) or that they keep referring to the most recent previous reference frame (substep 214.2 of method 200).
[0068] During the experiment, these operations provided the output seen in Table 6.The routine pick_and_drop executeQ is called to replicate the motion, since it calls all the other routines in order to replicate the motion that was taught during the demonstration. Further, as shown by the boldface parts of the program code in Table 6, the image capturing routine “capture_wobj” provides data that allows the reference frame pick_and_drop wobj_bowlo to be defined. The reference frame is then used for expressing some of the subsequent motion commands. However, the motion command “MoveL pick_and_drop p2i, spd, zone, toolo, \WObj:=wobjo” uses a fixed reference frame, which is better adapted for lifting the bowl to a fixed point in the robot work area 150.
[0069] The aspects of the present disclosure have mainly been described above with reference to a few embodiments and examples thereof. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the invention, as defined by the appended patent claims.
Claims
CLAIMS1. A method (200) of programming an industrial robot (100), which comprises a robot manipulator (no) and a robot controller (120), comprising: recording (210) movements of the robot manipulator during a demonstration-based robot programming session, for thereby obtaining a robot trajectory; acquiring (211) a video of the robot programming session; and generating (214), on the basis of the robot trajectory and the video, a robot program (C) executable by the robot controller, wherein the robot program includes at least one motion command which is expressed in a first reference frame (Ox), the method further comprising: capturing (212) operator input data indicating a first object (151) in an image in the video, wherein the robot program includes a command to identify a position of an object resembling the first object in a work area (150) of a robot which executes the robot program; and wherein the first reference frame is defined with respect to the identified position of the object resembling the first object.
2. The method (200) of claim 1, wherein: the captured operator input data further indicates a second object (152) in an image in the video; and the robot program further includes- a command to identify a position of an object resembling the second object in the work area (150) of the robot which executes the robot program; and- at least one motion command which is expressed in a second reference frame (O2) defined with respect to the identified position of the object resembling the second object.
3. The method (200) of claim 1 or 2, wherein the robot program further includes- a command to identify a later position of the object resembling the first object in the work area (150) of the robot which executes the robot program; and- at least one motion command which is expressed in a third reference frame (O3) defined with respect to the identified later position of the object resembling the first object.
4. The method (200) of any of the preceding claims, wherein the captured operator input data further determines validity periods of the first and any further reference frames.
5. The method (200) of claim 4, wherein generating (214) the robot program comprises: identifying (214.1) any motion command which is expressed in a reference frame outside its validity period; and transforming (214.2) the motion command into an equivalent motion command which is expressed in a neutral reference frame (O0).
6. The method (200) of any of the preceding claims, wherein the captured (212) operator input data indicating the first and any second object includes object masks referring to the respective images in the video.
7. The method (200) of any of the preceding claims, further comprising: applying (213) a pose-estimation algorithm to the acquired video, for thereby obtaining a storable data item representing the first object (151), the storable data item being suitable for use in initializing an instance of the pose-estimation algorithm by which the position of the object resembling the first object is to be identified during execution of the robot program.
8. The method (200) of claim 7, wherein: the robot program includes a command to identify a pose of the object resembling the first object; and the first reference frame (Ox) is oriented in accordance with the identified pose of the object.
9. The method (200) of any of the preceding claims, wherein: the captured (212) operator input data further includes data indicating at least one robot parameter value of at least one robot parameter; and the robot trajectory is annotated with the at least one robot parameter value.
10. The method (200) of claim 9, wherein the at least one robot parameter includes one or more of: a state of a tool (111) carried by the robot manipulator, a movement execution parameter, a degree of compliance with the robot trajectory, a degree of movement precision, a reference frame, a drive system parameter.
11. The method (200) of any of the preceding claims, wherein the operator input data includes speech data captured (212) while recording the movements of the robot manipulator.
12. The method (200) of any of the preceding claims, wherein the operator input data includes character data and / or a menu selection, which is optionally received via a graphical operator interface (166).
13. The method (200) of any of the preceding claims, wherein the first object (151) is a workpiece to be handled or processed by the robot manipulator (no).
14. The method (200) of any of the preceding claims, wherein the first object (151) is a fixed or mobile object other than the workpiece in the work area (150).
15. The method (200) of any of the preceding claims, wherein the demonstrationbased robot programming session is a kinesthetic programming session.
16. The method (200) of any of the preceding claims, wherein generating (214) the robot program includes converting the robot trajectory into a sequence of robot commands selected from a predefined set of robot commands executable by the robot controller.
17. The method (200) of any of the preceding claims, further comprising: executing (215) the robot program (C) while capturing (215.1) operator speech data including requests to modify (215.2) the robot program.
18. A programming device (160) for facilitating programming of an industrial robot (100), which comprises a robot manipulator (no) and a robot controller (120), theprogramming device comprising memory (162) and processing circuitry (161) configured to: record movements of the robot manipulator during a demonstration-based robot programming session, for thereby obtaining a robot trajectory; acquire a video of the robot programming session; and generate, on the basis of the robot trajectory and the video, a robot program (C) executable by the robot controller, wherein the robot program includes at least one motion command which is expressed in a first reference frame (Ox), the programming device further comprising an interface (166) configured to capture operator input data indicating a first object (151) in an image in the video, wherein the robot program includes a command to identify a position of an object resembling the first object in a work area (150) of a robot which executes the robot program; and wherein the first reference frame is defined with respect to the identified position of the object resembling the first object.
19. A computer program (163) comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method (200) of any of claims 1 to 17.
Citation Information
Patent Citations
Robot system, robot control device, control method, and computer program
US20230211499A1
Method and device for speech-supplemented kinesthetic robot programming
WO2025140777A1
Model-free six-dimensional object pose estimation
WO2025228515A1
Control of robotic arms through hybrid inverse kinematics and intuitive reference frame
US20220288780A1
Robot program generation method from human demonstration
US20230120598A1
Cited By
Method of providing a specialized large language model for controlling a machine
WO2026046496A1
Method and device for demonstration-based robot programming supplemented by video
WO2026077521A1
Method and device for demonstration-based programming of a robot operable in multiple control modes
WO2026092842A1
Method and device for demonstration-based robot programming supplemented by spoken interaction
WO2026098774A1