Method and device for demonstration-based robot programming supplemented by spoken interaction
The method and device provide spoken guidance during robot programming, using natural-language models to address incomplete or unclear trajectories, enhancing learning efficiency and reducing errors without visual distraction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ABB (SCHWEIZ) AG
- Filing Date
- 2024-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing demonstration-based robot programming methods require significant experimentation and time for beginners to learn and correct mistakes, and there is a need for real-time feedback on program completeness and errors without visual distraction.
A method and device that provides spoken guidance during and after the programming session, using natural-language models to parse speech data into robot parameters and offer feedback on incompleteness or potential problems, suggesting solutions and summarizing inputs, all while maintaining operator focus on the task.
Enables faster learning for new programmers by providing hands-free, non-visual feedback on robot program completeness and errors, reducing the need for trial-and-error and minimizing cognitive load.
Smart Images

Figure EP2024081368_15052026_PF_FP_ABST
Abstract
Description
METHOD AND DEVICE FOR DEMONSTRATION-BASED ROBOT PROGRAMMING SUPPLEMENTED BY SPOKEN INTERACTIONTECHNICAL FIELD
[0001] The present disclosure generally relates to the field of robotic control, and specifically to demonstration-based programming of an industrial robot. More precisely, a programming device is proposed herein which an operator can utilize to generate a robot program from a recorded robot trajectory together with the operator’s spoken inputs. The operator receives spoken guidance from the programming device during the demonstration-based programming session and / or at later stages of adapting or finalizing the robot program.BACKGROUND
[0002] Lead-through robot programming, once an area exclusive to seasoned programmers of high technical expertise, has seen a broadening user base lately thanks to the advent of Large Language Models (LLMs), and Generative Pre-trained Transformers (GPTs) in particular. LLMs are increasingly being applied to robot control and robot programming. Reference is made to the applicant’s prior disclosure PCT / EP2023 / 087949, which proposed a speech-supplemented robot programming method based on the following workflow:• During a demonstration-based (in particular, a kinesthetic) programming session, movements of the robot manipulator are recorded and saved as a robot trajectory.• Speech data is recorded during the demonstration-based programming session.• By suitable prompting, an LLM is caused to parse the speech data into values of robot parameters, such as a tool state, a movement speed, a degree of compliance with the trajectory, a degree of movement precision, a choice of reference frame to express the movements in.• The robot trajectory is annotated with each robot parameter value at a point of the trajectory which corresponds to the time of utterance.• A robot program is generated as a sequence of commands which realizes the robot trajectory while applying the robot parameter values as modifiers.The programming method according to PCT / EP2023 / 087949 offers the operator a convenient, hands-free way of adding explanatory remarks to the recorded movements (purpose of a step, aspects to pay particular attention to), or of requesting certain robot parameter settings which cannot be seen or sensed during the demonstration (speed, precision) etc., ultimately to provide a usable robot program in shorter time. The programming workflow as such appears to be proper to the applicant.
[0003] Although this workflow, as well as other robot programming tools, have greatly simplified demonstration-based programming, the truth remains that practice makes perfect. Indeed, beginners in robot programming still have to make typical beginner’s mistakes to evolve as professionals: they need to execute their robot programs in test environments, discover mistakes and learn from them. A robot programmer in training not only learns to eliminate mistakes, but gradually acquires a feeling for which approaches are best suited to address which practical situations. Until now, this has required a certain amount of experimentation, which usually has a monetary cost and consumes time.
[0004] It would be desirable to help new robot programmers up to speed in shorter time. Another desirable aim would be to make the programmers - whether beginners or experts - aware of errors and room for improvement in a robot program without having to execute it.SUMMARY
[0005] One objective of the present disclosure is to propose methods and devices for speech-supplemented robot programming, which inform the operator about points of incompleteness or potential problems with the future robot program. A further objective is to propose such robot programming methods and devices which use a preexisting communication modality to inform the operator of incompletenesses or problems. A further objective is to propose such robot programming methods and devices which inform the operator of incompletenesses or problems without distracting the operator visually. A further objective is to proposesuch robot programming methods and devices that have a further ability of suggesting solutions to the incompletenesses or problems. A further objective is to propose such robot programming methods and devices that have a further ability of summarizing the received inputs by way of confirmation, and / or a further ability of answering questions from the operator, and / or a further ability of providing an instruction or recommendation related to the performing of a robot task. A still further objective is to provide a robot programming method with one or more of these characteristics.
[0006] At least some of these objectives are achieved by the invention as defined by the independent claims. The dependent claims relate to advantageous embodiments of the invention.
[0007] In a first aspect of the present disclosure, there is provided a method of programming an industrial robot, which comprises a robot manipulator and a robot controller. The method comprises: recording movements of the robot manipulator during a demonstration-based programming session, for thereby obtaining a robot trajectory; capturing speech data while recording the movements of the robot manipulator; using a natural-language model, parsing the speech data into at least one robot parameter value and annotating the robot trajectory with the robot parameter value; and, on the basis of the annotated robot trajectory, generating a robot program. According to the first aspect, the method further comprises outputting spoken guidance to an operator before or during the demonstration-based programming session, or between completion of the programming session and the generating of the robot program, or after an execution of the robot program by the industrial robot.
[0008] In relation to the robot programming method according to PCT / EP2023 / 087949, the method according to the first aspect of this disclosure is characterized by spoken communication not only from an operator to a programming device, but also from the programming device to the operator (spoken guidance). It is understood that the spoken guidance from the programming device to the operator relates to synthetic speech or combinations of one or more prerecorded speech segments. This way, the operator can be informed, for example, about points of incompleteness or potential problems with the future robot program. Because the method relies on speech - this is a preexisting communication modality because the operator is already providing speech data while the manipulator movements arebeing recorded - the spoken guidance does not add a significant cognitive load on the operator. Nor does the spoken guidance distract the operator visually, but the operator’s visual attention may remain focused on the demonstration-based programming session.
[0009] In some embodiments, the spoken guidance includes a request for a demonstration of a robot task, or a portion of a robot task. The movements of the robot manipulator are recorded during the requested demonstration, which occurs after outputting the spoken guidance to the operator. In particular, the demonstration requested by the spoken guidance may be a repetition of a past demonstration of (a portion of) a robot task, i.e., the spoken guidance is given during the demonstrationbased programming session. Further, the demonstration requested by the spoken guidance may be the only demonstration within the programming session. By outputting this type of spoken guidance, a robot programming device may ensure that all portions of the robot trajectory are sufficiently complete and clear to enable generating a robot program. Although such incomplete or unclear portions of the trajectory could be either genuinely unclear (e.g., the robot task has been poorly planned by the operator, or the robot task is infeasible), experience shows that incomplete or unclear portions are oftentimes the result of accidental omissions or oversight, which the operator can resolve once he or she becomes aware thereof.
[0010] In some embodiments, the spoken guidance includes a request for clarification or for additional information. The operator is expected to provide the requested clarification or additional information in the form of speech data, which is recorded. In these embodiments, the programming device can output the spoken guidance to the operator as soon as it has noticed the need for clarification or additional information. Accordingly, the spoken guidance could be output during the demonstration-based programming session, or between completion of the programming session and the generating of the robot program. By outputting this type of spoken guidance, the robot programming device may handle points of factual incompleteness or seemingly contradictory inputs from the operator.
[0011] In some embodiments, the spoken guidance includes a request to identify one or more objects involved in a robot task (e.g., workpieces, tools, consumables). The spoken guidance is output before the demonstration-based programming session ends, or even before it begins. The identified objects are then imaged, or they areidentified by being visible in an indicated image region. By outputting this type of spoken guidance, the robot programming device may proactively request the operator to focus the robot programming device’s attention on the intended objects. Furthermore, the objects may optionally be assigned human-intelligible labels, maybe used as reference frame in which to express robot movements, and the like.
[0012] In some embodiments, the spoken guidance includes a request to indicate whether the programming session is still active. If the operator confirms that this is the case, the programming session continues with further recording of manipulator movements and / or speech data. If not, the programming session ends. In the latter case, the robot programming device may enter a sleep mode, or it may proceed to subsequent steps of processing the trajectory and speech data into a robot program.
[0013] In some embodiments, the spoken guidance includes a confirmatory summary of the robot trajectory, the captured speech data, at least one robot parameter which has been parsed from the speech data, or the annotated robot trajectory, or a combination of one or more of these. The operator maybe expected to verify that the confirmatory summary is in accordance with the operator’s intention and indicate this to a robot programming device, in which case the robot programming device may proceed to generate the robot program on this basis.
[0014] In some embodiments, the spoken guidance includes an answer to a question by an operator, wherein the question has been detected in the recorded speech data. A robot programming device may determine the answer to the question by deriving it from a conversation context relating to the conversation between the operator and the robot programming device, or the robot programming device may retrieve the answer from a knowledge database. By outputting this type of spoken guidance, the robot programming device supports the operator in a handsfree manner by sharing reference information, as well as facts and historic details relating to the demonstration-based programming section at hand.
[0015] In some embodiments, the spoken guidance includes an instruction or recommendation related to the performing of a robot task. The instruction or recommendation may contain common general knowledge of practitioners in the field, such as a recognized best practice for addressing common difficulties.
[0016] In some embodiments, the spoken guidance includes an explanation of an error which a robot programming device has detected during the recording of robot manipulator movements or during the capturing of speech data. The explanation may include a likely cause of the error, as determined by the robot programming device or logic therein.
[0017] It is appreciated that the invention is not limited to using a single type of spoken guidance. Rather, two or more of the embodiments reviewed so far can be combined. Also, the method of the first aspect may be embodied such that spoken guidance is output on multiple occasions, e.g., both before and during the demonstration-based programming session.
[0018] In some embodiments, the method of the first aspect is performed by a programming device which is operable in at least a conversational mode, a learning mode and an inactive mode. Here, the inactive mode may correspond to a period during which the generated robot program has been transferred to a memory in the industrial robot and is being executed by the industrial robot. In these embodiments, the method further includes giving a visual indication of the current mode of the programming device. This visual indication informs the operator when the robot programming device is ready to accept speech data, so that the operator does not waste his or her efforts at other times. Alternatively or additionally, the visual indication may inform the operator whether the operator should listen for spoken guidance. An operator who is aware of such periods where no spoken guidance will be output (where the operator runs no risk of missing such guidance) may experience a reduced cognitive and sensory burden when using the robot programming device.
[0019] In a second aspect of the present disclosure, there is provided a programming device (in particular, a robot programming device) for facilitating programming of an industrial robot. The programming device comprises a bidirectional speech interface, a conversation engine configured to conduct a dialog with an operator through the bidirectional speech interface, as well as memory and processing circuity configured to carry out the robot programming method of the first aspect. The bidirectional speech interface may include a speech-to-text converter and a text-to-speech converter.
[0020] A programming device with these features generally shares the effects and advantages of the first aspect, and it can be implemented with an equivalent degree of technical variation.
[0021] The present disclosure further relates to a computer program containing instructions for causing a computer, or the programming device in particular, to carry out the method of the first aspect. The computer program may be stored or distributed on a data carrier. As used herein, a “data carrier” maybe a transitory data carrier, such as modulated electromagnetic or optical waves, or a non-transitory data carrier. Non-transitory data carriers include volatile and non-volatile memories, such as permanent and non-permanent storage media of magnetic, optical or solid-state type. Still within the scope of “data carrier”, such memories may be fixedly mounted or portable.
[0022] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to “a / an / the element, apparatus, component, means, step, etc.” are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order described, unless this is explicitly stated.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Aspects and embodiments are now described, by way of example, with reference to the accompanying drawings, on which: figure 1 shows a work area, in which a robot manipulator operates under the control of a robot controller, and a programming device supporting demonstration-based robot programming; figure 2 is a flowchart of a robot programming method according to embodiments herein; and figures 3 and 4 are illustrations of information flows during an execution of the method illustrated in figure 2.DETAILED DESCRIPTION
[0024] The aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, on which certain embodiments of the invention are shown. These aspects may, however, be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and to fully convey the scope of all aspects of the invention to those skilled in the art. Like numbers refer to like elements throughout the description.System overview
[0025] As explained initially, the present disclosure relates to the field of robotic control and specifically to demonstration-based programming of an industrial robot. In the terminology of the present disclosure, “industrial robot” is used in a broad sense, to cover in particular manufacturing robots, material-handling robots, assembly robots, cutting / welding robots, service robots, collaborative robots, hygiene robots, industrial robot tracks, industrial robot positioners. An industrial robot may be a stationary robot or a mobile robot, such as an automated guided vehicle (AGV), an autonomous mobile robot (AMR) or an autonomous mobile manipulator robot (AMMR). The term industrial robot covers the full range from lightweight robots designed to replace human manual work, over collaborative robots for supporting a human worker, all the way up to heavy-duty robots.
[0026] By way of illustration and not limitation, figure 1 shows an industrial robot 100 made up of a robot manipulator no and a robot controller 120. The robot manipulator no and the robot controller 120 are joined by a wired or wireless bidirectional data connection, which conveys control signals, sensor data etc.
[0027] The robot manipulator 110 includes an arm, which extends from a base 112 and is composed of structural elements 114 and at least one linear or rotary joint 113. The arm may further include an end-effector 111 allowing it to carry tools by which it may interact with objects in the form of various workpieces 151, which are present in a work area 150 of the robot manipulator no. The workpieces are subject to manufacture, processing or handling by the robot manipulator no. The work area 150 may further include non-workpiece objects 152, such as containers, fixtures, robot positioners, separators, insulators, supports etc. These objects 152 maybe generic orthey may be specifically adapted to the workpieces 151 handled by the robot manipulator no. Apart from wear, staining etc., they are generally in the same condition at the beginning and end of a work cycle. The non-workpiece objects 152 could also be consumables or containers for consumables.
[0028] The arm of the robot manipulator no is movable by action of internal motors, drives or other sorts of actuators (not shown). The arm of the robot manipulator 110 further includes transducers, encoders, sensors and other measuring equipment (not shown), from which the manipulator’s no current position, pose, technical condition, load etc. can be derived, to some degree of accuracy. The position, pose etc. of the robot manipulator 110 may in particular refer to a point on the arm, particularly to a tool-center point (TCP).
[0029] The actuators of the robot manipulator 110 can be controlled based on feedback from the measuring equipment. In particular, the different degrees of freedom (DOFs) of the robot manipulator 110 can be controlled in accordance with force control (also known as impedance control, stiffness control) or position control. In broad terms, position control signifies that the actuators are controlled in such manner that the robot manipulator no achieves a desired setpoint position (optionally subject to constraints on speed or acceleration), whereas force control may be said to establish a connection between the contact force and the position when the robot manipulator no makes contact with an object. The robot can conceptually ‘feel’ its surroundings. It is able to apply a constant force on a surface, even if the exact position of the surface is not known. Hybrid control may signify that some DOFs are position-controlled and others are force-controlled.
[0030] The positions, poses etc. of the robot manipulator 110 maybe expressed with respect to one of multiple possible reference frames, including a first fixed reference frame O0defined with respect to a point in the work area 150, a second fixed reference frame (not shown) defined with respect to the base 112, a first reference frame O defined with respect to an initial position of a workpiece 151, or a second reference frame O2defined with respect to a position of a non -workpiece object 152 in the work area 150. Here, the first fixed reference frame O0is defined with respect to the point in the work area 150 in the sense that its origin is situated in that point. Alternatively, a reference frame may be defined with respect to a reference point in the sense that its origin is situated a predefined translation from the reference point.The first and second fixed reference frames, as well as any further reference frames which are independent of the positions of objects 151, 152 in the work area 150, may be collectively referred to as neutral reference frames. All reference frames shown in figure 1 have a common orientation, i.e., their respective first, second and third axes are parallel. Without departing from the scope of the present disclosure, a further option is to use reference frames with mutually different orientations, including a reference frame which is oriented in accordance with a pose of an object 151, 152 in the work area 150 or another suitable reference objects. In particular, one may use a reference frame which has its origin situated in the TCP and is oriented parallel to the tool 111 of the manipulator no at all times.
[0031] The robot controller 120 comprises processing circuitry 121, a memory 122 and a communication interface 123. Example content of the memory 122 during operation may includeS: an operating system, basic settings, software implementing generic movements, sensing, self-monitoring, generically useful functionalities and services (all typically contributed by an original manufacturer), task- or role-specific configurations, configuration templates (typically contributed by a robot system integrator), and site-specific settings (typically contributed by an end user);C: robot programs or projects causing the robot manipulator no to perform useful or intended tasks in its work area 150.Processes that execute the robot programs may do so in accordance with the program-independent memory content S, e.g., by making calls to available functionalities, libraries, routines or parameter values therein. It is recalled that the memory 122 and processing circuitry 121 of the robot controller 120 maybe distributed and / or contain networked resources having a different physical localization than figure 1 suggests.
[0032] Each of the programs C may contain a plurality of movement instructions relating to locations such as points, poses, paths as well as modulated paths. A program maybe a compiled executable (binary) or a script. A movement instruction relating to a modulated path may be expressed as, or may include, a process-on-path instruction. The programs C maybe created by an operator 190 with the aid of a robot programming device (programming station) or a general-purpose computer, orthey may be created directly at the robot controller 120 if it has an operator interface. In the first two cases, versions of the programs C may be downloaded to the robot controller 120 over a wired or wireless connection or by being temporarily stored on a portable memory. A robot program maybe created by means of demonstration-based programming according to the teachings in the applicant’s prior disclosures PCT / EP2023 / 087949, PCT / EP2024 / 066288 and PCT / EP2024 / 078213.
[0033] As used herein, “demonstration-based” robot programming (or programming by demonstration [PbD], or programming from demonstration) includes a step where the operator performs a portion of a robot task or guides the robot manipulator such that it performs a portion of this robot task. The operator’s activities are imaged or, in other ways, observed by a programming device. Demonstration-based robot programming includes kinesthetic programming as a special case. The term kinesthetic programming - by allusion to the robot’s proprioceptive ability to sense its own position, pose and movements - is a programming approach in which the operator physically moves the robot manipulator to imitate the desired motions. The terms kinesthetic programming and lead-through programming are synonymous or at least partially overlapping in meaning. The state of the robot during a kinesthetic programming session is typically recorded by means of the robot’s onboard sensors, e.g., joint angles and torques. As used herein, kinesthetic programming also includes teleoperation as a special case, where the movement of the robot manipulator is controlled by an external input to the robot through a joystick, graphical user interface or other input means; kinesthetic programming by teleoperation does not require the operator to be present in the work area of the robot.
[0034] A dedicated programming device (robot programming device) 160 is shown in the right-hand portion of figure 1. From the programming device 160, the created programs C can be transmitted to the communication interface 123 of the robot controller 120 and then stored in the memory 122 where they are available for execution. It is noted that the programming device 160 maybe implemented as a set of collaborating components within the robot controller 120. In fact, the components can be shared with the robot controller 120, e.g., by using the processing circuitry 121 for the dual purposes of robot control and programming and / or using the memory 122 for the same dual purposes. In other words, the programming device 160 mayconstitute a portion of a multi-purpose device; it need not be a standalone device or a device with programming as its sole or main purpose.
[0035] During a demonstration-based programming session, the programming device 160 has access to position data representing an actual position, a recorded position or recorded movements of the robot manipulator no, as well as speech data captured while recording the movements of the robot manipulator. Optionally, according to some embodiments, the programming device 160 further has access to images or video of the robot programming session. The position data may for example be obtained through the intermediary of the robot controller 120, which monitors position data in the normal course of its operation; alternatively, the programming device 160 is granted access to corresponding signals from the transducers, sensors or other measuring equipment in the robot manipulator no.
[0036] The programming device 160 - which is shown as a standalone device in the non-limiting example of figure 1 - comprises processing circuitry 161, memory 162 with executable software 163, and a communication interface 164. The programming device 160 may further comprise at least one operator interface 166, in particular a graphical operator interface or graphical user interface (GUI).
[0037] The programming device 160 may further have an imaging device (e.g., camera, video camera, depth camera, lidar, radar) 167 at its disposal, by which images or video of the demonstration-based programming session can be captured. In particular, the imaging device 167 may be a depth camera, lidar or radar, which in addition to - or instead of - the two-dimensional appearance of an object determines a depth coordinate of the object, e.g., by time-of-flight measurements, triangulation, reflection or other per se known techniques. An RGB-D camera is an example of a depth camera. The imaging device 167 may be a part of the programming device 160, or it maybe integrated into the industrial robot 100 and optionally be used for other tasks as well. More generally, an equivalent function may be achieved by any imaging device which is in rigid relationship with the robot base 112.
[0038] The programming device 160 further comprises a bidirectional speech interface 165, a conversation engine 165.1, one or more acoustic transducers for input (e.g., microphones 168) and output (e.g., loudspeakers 169). The input could be speech data representing utterances or narration by the operator 190 or a bystander during the programming session. The output could be spoken guidance to theoperator 190. For this purpose, the acoustic transducers 168, 169 maybe arranged in the vicinity of the operator’s 190 normal position, or alternatively they are arranged in a headset to be worn by the operator 190.
[0039] The bidirectional speech interface 165 is associated with a speech-to-text converter 165.2 for converting captured speech data into a computer-readable form, and a text-to-speech converter 165.3 adapted for synthesis or playback of spoken guidance to the operator 190. The speech-to-text converter 165.2 may optionally be operable to detect starts and ends of clauses and sentences based on syntax and / or pausing. The software Azure Al Speech™, which is available from Microsoft Corporation, Redmond, Washington, United States, includes speech-to-text and text-to- speech functionalities. A bidirectional spoken communication with the operator 190 is coordinated and managed by a conversation engine 165.1. The conversation engine 165.1 may in turn be interfacing with a natural-language model 171. The naturallanguage model 171 may be either stored internally or - as in the example configuration shown in figure 1 - may be available from a host computer or an external memory 170. A further external memory stores a knowledge database 180, which maybe static or maybe susceptible of extensions and updates during operation. In detail, the conversation engine 165.1 maybe configured to process the text input through the natural-language model 171 to manage dialog, generate text responses, determine states (conversation, learning, execution etc.), and / or to control the robot's actions and the visual indications 118, 119.
[0040] The natural-language model 171 stored in the external memory 170 may constitute a large language model (LLM), that is, a type of machine-learning algorithm which has been trained on very large datasets using deep learning techniques to be able to perform natural-language processing (NLP) tasks. Example NLP tasks are recognizing, summarizing, translating, predicting and generating plausible textual content. A very large dataset in this sense may include of the order of one million parameters, such as tens of millions of parameters, such as hundreds of millions of parameters. An LLM may have a transducer architecture, particularly a transducer architecture with four cascaded key links when it processes the input data, namely: word embedding, position encoding, self-attention mechanism, feedforward neural network. At the time of filing this disclosure, noteworthy example LLMs include Bidirectional Encoder Representations from Transformers (BERT), Bard,BLOOM, Claude 2, various versions of Generative Pre-trained Transformer (GPT) including GPT-40, Llama, PaLM 2, RoBERTa, T5, LaMDA, Turing NLG, Gemini 1.0. LLMs include, as a special case, multimodal models which are also vision-enabled, such as Gemini 1.5. One benefit of an LLM is that the vocabulary is practically open- ended. The operator 190 can start using it without prior training. The demonstrationbased programming can be carried out without requiring the operator 190 to be extremely focused on using the right command words (or avoiding them) and / or syntax. Similarly, a well-tuned LLM generates naturally-sounding textual content with a high likelihood of being understood by the operator 190 without prior training; additionally, the LLM can provide a rephrased version of this textual content if the operator 190 indicates that he or she did not fully understand the initial version.
[0041] It is noted that the speech-to-text converter 165.2 and the text-to-speech converter 165.3 maybe implemented as integral parts of the natural-language model 171. Such an implementation of the natural-language model 171 maybe described as a speech-to-speech LLM.
[0042] It is furthermore noted that, for the purposes of implementing the aspects of the present disclosure, it is not necessary to use a natural -language model 171 that has been exposed to training data relating to the industrial robot 100 under consideration, or to industrial robots at all. To the contrary, the inventors have achieved very satisfying results while using a generic natural-language model, which appears to master the basics of robotic technology and its practical uses. The knowledge base of the natural-language model 171 may be further reinforced by implementing a suitable retrieval-augmented generation (RAG) technique; this may be particularly useful for the embodiments with a question-answering functionality.
[0043] Reference is made to figure 4. The conversation engine 165.1 in the robot programming setup disclosed herein may have the following characteristics and roles. The conversation engine 165.1 is stateful. It maintains a conversation context during operation, e.g., while the programming device is in the conversational mode. A conversation context maybe understood as the information that has been established up to the current time; the conversation engine 165.1 may rely upon this information to generate subsequent statements in the conversation (e.g., spoken guidance) and to interpret (contextualize) the operator’s 190 utterances. Maintaining the conversation context may include performing updates based on captured speech data 302, whichhas been transcribed into text input 303 by the speech-to-text converter 165.2, and based on spoken guidance 311, which has been synthesized by the text-to-speech converter 165.3 from text output 310 generated by the natural-language model 171. In some embodiments, the text output 310 is additionally displayed in the operator interface 166. The conversation engine 165.1 may further be configured to determine when to apply the conversational mode, the learning mode or neither of these.
[0044] The natural-language model 171 is initially prompted to learn a new robot task, and it may provide spoken guidance which is relevant and / or adequate in view of- the text input 303,- a trajectory 301 recorded during the demonstration-based programming session,- the current conversation context maintained by the conversation engine 165.1,- an operational state of the industrial robot 100, or combinations of one or more of these. The conversation engine 165.1 listens to sensor data 306 from the measuring equipment in the manipulator no of the industrial robot 100, which allows it to estimate said operational state. Further, it may feed action protocols 307 to the robot. As used herein, an action protocol command may define a certain action in the industrial robot 100 (e.g., turning on / off lead-through, turning on / off demonstration recording, capturing a picture with the camera 167) which is associated with a certain spoken guidance to the operator.Example 1. Spoken guidance to operator: “I will turn on lead-through mode. Robot command to execute: lead-through_on.Example 2. Spoken guidance to operator: “Please indicate the speed for this part of the motion.” Robot command to execute: None, will capture the expected speech data 302 using acoustic transducer 168.
[0045] The conversation engine 165.1 may be operable for a bidirectional exchange with the knowledge database 180; in particular, the conversation engine 165.1 may update or extend the knowledge database with logged operational data 308, and it may receive responses 309 to queries. The interface 308, 309 between the conversation engine 165.1 and the knowledge database 180 maybe configured tosupport retrieval-augmented generation (RAG). Further, the conversation engine 165.1 maybe configured to trigger visual indications 118, 119 informing the operator 190 whether the programming device 160 is in the conversational mode (triggered by signal 304) or in a learning mode (triggered by signal 305). In some implementations, the conversational mode and learning modes are mutually exclusive; in other implementations both maybe contemporaneously active, such that audio recording may continue during the physical interaction of the operator 190 and the robot manipulator no. The visual indications 118, 119 are preferably given by light indicators or visual displays located at or near the robot manipulator 110.Robot programming method
[0046] Turning to the flowchart in figure 2, embodiments of a method 200 of programming an industrial robot 100 of the type depicted in figure 1 will now be described. Not all steps shown in figure 2 are necessarily carried out in all embodiments. The output of the method 200 includes a robot program C to be executed by the robot controller 120 of the industrial robot 100 in figure 1. The robot program C maybe executable only by the robot controller 120 of the industrial robot 100 for which the programming method 200 was carried out. In some embodiments of the method 200, the robot program C can be executed also by further robot controllers of the same model as - or compatible with - the one for which the programming method 200 was carried out.
[0047] A robot program C may include a sequence of robot commands that cause the robot manipulator no to reproduce the trajectory. The robot commands maybe selected from a predefined set of robot commands executable by the robot controller 120, such as commands compliant with the RAPID™ robot programming language.
[0048] The method 200 in figure 2 may be performed by a general-purpose processor, and in particular by the programming device 160. As mentioned, the programming device 160 may constitute a portion of a multi-purpose device. The method 200 maybe considered to be a description of a behavior that the programming device 160 is configured to reproduce while active. The method 200 may as well correspond to the behavior of the robot controller 120 in a special programming mode, which the robot controller 120 enters on request. Instructions for causing a computer - or the programming device 160 in particular - to carry out the method 200 maybe provided in the form of a computer program 163 (see figure 1).
[0049] The functioning and other characteristics of the method 200 may be better understood from the illustrations in figures 3 and 4 of the information flows (shown as arrows) which occur during an execution of the method 200. Here, in addition to the reference numbers from figures 1 and 2, the following notation is used:301 trajectory recorded during demonstration-based robot programming session302 speech data303 text input transcribed from speech data304 conversational mode indication signal305 learning mode indication signal306 sensor data307 action protocols308 operational data309 responses to queries310 text output311 spoken guidance synthesized (or played back) from text output312 robot parameter value313 robot trajectory annotated with robot parameter values314 robot program.
[0050] A first step 212 of the method 200, which is executed during a demonstration-based robot programming session, includes recording movements of the robot manipulator no. On the basis of the recorded movements, a robot trajectory 301 can be obtained, e.g., by combining the recorded movements. The robot trajectory 301 thus obtained may indicate the robot manipulator’s 110 position and / or pose as a function of time, wherein the position or pose may be expressed in cartesian or joint-space coordinates. The position or pose need not be indicated for all points in time. Alternatively, the trajectory 301 maybe expressed as a number of reference times at which the robot manipulator no is to assume correspondingsetpoint positions or setpoint poses, wherein the robot manipulator no is free to have arbitrary positions or poses in the intervals between the reference times.
[0051] In an optional second step 213, a video of the programming session is captured. The video may be captured by means of the fixed imaging device 167 or by an imaging device (not shown) mounted on the robot manipulator 110. Each frame of the video may be a visual image, a depth image or a combination of these. The video may be relied upon to confirm the correctness of the generated robot program C, to resolve ambiguities in the trajectory 301 and / or in speech data 302. For example, poses and waypoints (in a fixed reference frame) of objects 150, 151 handled during the robot task may be extracted from the video.
[0052] In a next step 214, speech data 302 is captured while the movements of the robot manipulator no are being recorded. For this purpose, the acoustic transducer 168 maybe utilized. The speech data 302 maybe stored in any suitable audio format, and at a bitrate considered adequate for the application at hand.
[0053] The speech data 302 may include utterances or narration by the operator 190 or a bystander during the programming session. One the one hand, the speech data may act to confirm an activity which can in principle be inferred from the recorded trajectory 301. On the other hand, the speech data 302 may add new information, such as a purpose or desired result of a certain sequence of movements, or a property / quality which is not derivable from the recorded trajectory 301 and / or not visible in a video of the programming section.
[0054] For instance, the speech data 302 maybe used as a medium for indicating a setpoint speed of the robot manipulator no. This maybe helpful as it is, most often, practically impossible to perform the desired manipulator movements at the true speed without sacrificing accuracy and safety. The speed indications maybe provided in terms of everyday speed-related vocabulary like “fast”, “slowly”, “quite slowly” or relative speed indications such as “faster”. Further, the speech data 302 may indicate a required degree of compliance with the recorded robot trajectory 301, such as: “Move the robot close to this object”, “Reorient the box handle towards the top”, “Reorient the object so it is aligned with the place position”, “Carefully proceed with the insertion”, “Distribute the paint evenly”. Further, the speech data 302 may indicate a required degree of movement precision or an acceptable geometric deviation from the trajectory, such as: “Move carefully downwards to do theinsertion”, “I am moving the robot back to the home position”, “I will move the robot close to the object to grip”. Further, the speech data 302 may indicate, explicitly or implicitly, a reference frame in which the movements of the robot manipulator no are to be expressed, such as: “I am taking the robot back to the home position”, “I am approaching the red cube to pick up”, “I am inserting the charger in the socket”. Further still, the speech data 302 may indicate a handling force or a gripping force (or contact force), such as: “firmly”, “gently”, “hard”.
[0055] Optionally, in a step 216, the speech data 302 is decomposed into a plurality of phases. The phases may correspond to different time intervals of a captured audio track or different time-contiguous segments. They should preferably correspond to different robot activities, so that parsed parameter values (see step 217) can be applied naturally. Normally, a time-uniform decomposition is not very useful, as it does not account for differences in duration of the different activities. The robot activities maybe sub-tasks (e.g., picking a cube, rotating a handle), steps explicitly mentioned by the operator 190 (e.g., “Now I will begin the process of moving backwards”) or robot operating modes (e.g., linear motion, motion parallel to a table, opening gripper, closing gripper), or various combinations of these. The phases may be annotated with respective human-intelligible labels. Optionally, the content of these labels is derived from the speech data, e.g., based on utterances by the operator 190 such as “The preheating is now complete, and I move the workpiece from the oven to the mold to begin the shaping.” Such a statement maybe interpreted as marking the end of a preheating phase and the start of a shaping or molding phase. The decomposing of the speech data 302 into phases in step 216 maybe carried out using the natural-language model 171.
[0056] Corresponding phases of the trajectory 301 may be identified. The phases may be corresponding with respect to the activity or subject-matter occurring therein. Alternatively, the corresponding phases in the trajectory 301 maybe contemporaneous with those of the speech data 302, in which case they can be identified by matching start and end times of the phase of the speech data 302 with time indications forming part of the recorded trajectory 301, including the mentioned reference times.
[0057] In a next step 217, the speech data 302 is parsed into at least one robot parameter value 312 using the natural-language model 171 and the robot trajectory isannotated with the robot parameter value. The robot parameter value shall apply locally, i.e., it applies at or around the time the speech was captured but not - at least not by default - during the full robot trajectory. The annotation may have a granularity of the phases into which the speech data 302 was optionally decomposed in step 216. More precisely, a robot parameter value parsed from a phase of the speech data 302 shall be used for annotating a corresponding phase of the trajectory 301.
[0058] The parsing in step 217 may include transcribing the speech 302 data into text 303, which preferably carries timestamps (cf. speech-to-text converter 165.2 in figure 4). After this, the natural-language model 171 is used for extracting robot parameter values 312 from the timestamped text. Step 217 may return robotparameter values 312 pertaining to one or more robot parameters; further, step 217 may return one or more parsed robot-parameter values for each of said robot parameter or robot parameters. The robot parameter maybe, for example, a state of a tool carried by the robot manipulator, a movement execution parameter, a degree of compliance with the robot trajectory (e.g., to what extent smoothing, straightening or removal of sharp bends is permissible), a degree of movement precision, a reference frame, a drive system parameter, a motion template to be used for realizing the robot trajectory, or the like. As used herein, a motion template is a preconfigured combination of robot commands (or a program snippet), which represents a sequence of movements and / or status changes of the robot manipulator no, which can be used as a building block when carrying out the robot task. For example, a motion template may be suitable for realizing linear motion, spiral search, hole insertion, polishing, welding, an act of opening or releasing a gripping tool. So-called reference movements in RAPID™ are motion templates in this sense.
[0059] The annotated trajectory 313 which forms the output of step 217 is to be provided in a form that can be handed over safely to the downstream processing steps without losing or corrupting the information. The output of step 217 may include the extracted robot parameter values associated with respective points in time or respective phases of the robot trajectory 301. In particular, the extracted robot parameter values may be formatted in accordance with a data serialization format segmented into phases. Examples data serialization formats are JSON (specified in The JSON Data Interchange Syntax, Standard ECMA-404, 2ndedition (2017-12)), YAML (specified in YAML Ain’t Markup Language (YAML™), version 1.2, revision 1.2.2(2021-10-01)) and XML (specified in XML Signature Syntax and Processing, version 1.1, W3C Recommendation (201304-11)). If phase granularity is used (cf. step 216), the output is preferably organized phase by phase, that is, the set of robotparameter assignments for each new phase are contained in a new item in the output.
[0060] In a next step 220, a robot program C is generated on the basis of the annotated robot trajectory 313, in accordance with the robot parameter values 312. As mentioned, the robot program C may be expressed as a sequence of robot commands which, when executed by the robot controller 120, cause the robot manipulator no to realize the trajectory 301 while applying the robot parameter values as modifiers to these robot commands. Figure 3 shows an embodiment where step 220 includes generating the robot program 314 based on a combination of the annotated robot trajectory 313 and an image / video 315a or an indicated image region 315b that depicts an object 151, 152 (cf. figure 1) which is involved in the robot task.
[0061] Step 220 may include a first substep of sampling the trajectory 301 into a sequence of discrete points and a second substep of selecting robot commands that cause the robot manipulator no to move between each pair of consecutive discrete points. The sampling used in the first substep may be time-uniform sampling (constant step duration), space-uniform sampling (constant step length), or a non- uniform sampling algorithm with controlled deviation, such as Ramer-Douglas- Peucker. In a RAPID™ environment, the second substep may be performed so as to output instances of the command MoveL (cartesian linear motion) or MoveJ (jointspace linear motion) or a combination of these. Optionally, the step 220 may include a postprocessing substep applied to the sequence of generated robot commands, such as formatting the sequence into a predefined script format by appending a header, performing a syntactic consistency check, or the like.
[0062] According to the present disclosure, embodiments of the robot programming method 200 includes at least one step of outputting spoken guidance 311 to the operator 190. The spoken guidance 311 maybe provided during the demonstrationbased programming session, or between completion of the programming session and the generating of the robot program, or after the robot program has been executed (step 221) by the industrial robot 100, or in combinations of these time periods.
[0063] In a first embodiment of the method 200, the spoken guidance 311 includes a request for a demonstration of a robot task or a portion thereof. Themovements of the robot manipulator are recorded (step 212.1) during the requested demonstration, which occurs after the spoken guidance has been output to the operator (step 210 or 215). In particular, the spoken guidance be a request to repeat a past demonstration of the robot task (or the portion thereof), i.e., the spoken guidance is given during the demonstration-based programming session. The conversation engine 165.1 may provide this type of spoken guidance 311 in circumstances when it is determined (e.g., by the natural-language model 171) that the so far recorded robot trajectory 301 is incomplete or a portion of it is of inferior quality.
[0064] For illustration, an example conversation with the operator (“User”) 190 is presented in Table 1. The running example in this and the following tables relates to a task of placing a receptable with a blood sample in a sample rack. In the example conversations, “Robot” is used as an umbrella term referring to any entity which carries out the present method 200, such as the robot controller 120 or a robot programming device 160. The operator’s 190 perception maybe that the spoken guidance 311 emanates from the industrial robot 100 if it is output by a loudspeaker 169 in the vicinity of a visible element of the industrial robot 100. It is further noted that the italicized text in the tables does not form part of the spoken communication.In Table i, “auto mode” - as opposed to “manual mode” - refers to an operational mode of the industrial robot 100, which does not necessarily have an impact on the robot programming device 160. In the auto mode, the robot controller 120 is authorized to move the robot manipulator no autonomously, e.g., in accordance with a robot program or a previously given instruction. In the manual mode, the robot controller 120 normally carries out any received instructions without delay. Further, “lead-through mode” corresponds the act of recording the movements of the robot manipulator 110. In some implementations, such recording is possible in auto mode as well as in manual mode of the industrial robot, but it maybe restricted to the learning mode of the robot programming device 160. Accordingly, the robot programming device 160 may enter the conversational mode before, between or after the periods of recording manipulator movements. In other implementations, the programming device 160 may record manipulator movements while it is in the conversational mode.
[0065] Although Robot’s output in the example of Table 1 conveys some robotspecific knowledge (e.g., suitability of the auto mode), this example was generated using a generic LLM as the natural-language model 171.
[0066] In a second embodiment, the spoken guidance 311 includes a request for clarification or for additional information. The operator is expected to provide the requested clarification or additional information in the form of speech data, which is recorded (substep 214.1). In these embodiments, the conversation engine 165.1 can output the spoken guidance to the operator as soon as it notices (e.g., with the assistance of the natural-language model 171) the need for clarification or additionalinformation. Accordingly, the spoken guidance 311 could be output during the demonstration-based programming session, or between completion of the programming session and the generating of the robot program.
[0067] An example conversation with the operator 190 is found in Table 2.
[0068] In a third embodiment, the spoken guidance 311 includes a request to identify one or more objects involved in a robot task (e.g., workpieces, tools, consumables), wherein the spoken guidance is output before the demonstration-based programming session ends, or even before it begins (step 210). The objects are identified by being depicted in an image or video 315a (step 211a), or they are identified (step 211b) as being visible in an indicated region 315b of an image or video. The acquired or indicated image data maybe used to recognize the object during the demonstration-based programming session, and / or to focus the programming device’s 160 attention accordingly. Optionally, some of the objects are depicted in multiple poses in order to facilitate subsequent pose estimation, including pose estimation by the composite algorithm described in the applicant’s prior disclosure PCT / EP2024 / 061921.
[0069] An example conversation with the operator 190 is found in Table 3.
[0070] In a fourth embodiment, the spoken guidance 311 includes a request to indicate whether the programming session is still active. If the operator confirms that this is the case, the programming session continues with further recording of manipulator movements and / or speech data. If not, the programming session ends. In the latter case, the robot programming device may enter a sleep mode (e.g., leave conversational mode), or it may proceed to subsequent steps 217, 220 of processing the trajectory and speech data into a robot program.
[0071] A request to confirm that the programming session is still active (“It seems we have paused ...”) is found in the example conversation in Table 3 above.
[0072] In a fifth embodiment, the spoken guidance 311 includes a confirmatory summary of the robot trajectory 301, the captured speech data 302, at least one robot parameter 312 which has been parsed from the speech data 302, or the annotated robot trajectory 313, or a combination of one or more of these. The confirmatory summary is provided in a step 218. The summary is confirmatory in the sense that it normally does not go beyond the information contained in the robot trajectory, speech data etc., other than possibly contextualizing it in a non-creative way. The operator 190 may be expected to verify that the confirmatory summary is in accordance with the operator’s 190 intention and indicate this to a robot programmingdevice 160 (e.g, by saying “Correct”, “OK, go ahead”; step 219), in which case the robot programming device may proceed to generate the robot program on this basis. In other words, the present fifth embodiment makes the generating 220 of the robot program C conditional upon receiving 219 the operator’s 190 approval of the confirmatory summary. If no approval is received (N branch from step 219), it is decided that the information collected so far is insufficient for generating a robot program C. In this event, the execution flow of the method 200 may stop or loop back to one of the initial steps.
[0073] An example conversation according to the fifth embodiment with the operator 190 is found in Table 4.As Table 4 illustrates, the method 200 may include using post-teaching spoken guidance 311 which is output (step 222) after the robot program has been executed in full or in part (step 221). More precisely, the operator 190 is asked whether he wishes to adjust any behavior laid down in the program, and the operator may react accordingly.
[0074] In a sixth embodiment, the spoken guidance 311 includes an answer to a question by an operator, wherein the question has been detected (step 214.3) in the recorded speech data 302. In a robot programming device 160, where a conversation engine 165.1 maintains a conversation context relating to the conversation with the operator 190, the answer to the detected question maybe determined 214.4 by deriving it from the conversation context. At the time of this disclosure, a question answering model is (implicitly) part of many available LLMs, which can be used as the natural-language model 171. Alternatively, the answer to the detected question may be retrieved from the knowledge database 180 discussed above.Implementations of the sixth embodiments may include meaningful reactions to incomprehensible or unanswerable questions. In particular, the natural-language model 171 maybe trained or tuned to provide a clear indication to this effect (e.g., “I’m not sure”) rather than giving an erroneous output that the operator 190 could mistake for a successfully retrieved answer. Similarly, the operator 190 may at alltimes request a clarification of a spoken guidance 311 that he or she does not readily understand.
[0075] An example exchange of questions and answers is found in Table 5.
[0076] Table 6 illustrates how a situation is handled where the robot is unable to determine an answer to a question.
[0077] Table 7 illustrates how a situation is handled where the robot is unable to understand the question and requests clarification.
[0078] In a seventh embodiment, the spoken guidance 311 includes an instruction or recommendation related to the performing of a robot task. The instruction or recommendation may contain common general knowledge of practitioners in thefield, such as a recognized best practice for addressing a known difficulty. Common general knowledge of this kind may be incorporated in the natural -language model 171 (and even in a generic natural-language model 171, as experiments suggest) or in the knowledge database 180, or both.
[0079] An example conversation where the operator 190 is given a recommendation on the carrying out of a task is found in Table 8.
[0080] In an eighth embodiment, the spoken guidance 311 includes an explanation of an error which the robot programming device 160 has detected (step 212.2) during the recording of robot manipulator movements 301 or which the robot programming device 160 has detected (step 214.2) during the capturing ofspeech data 302. The explanation may include a likely cause of the error, as determined by the robot programming device 160 or logic therein. In particular, the natural-language model 171 or the knowledge database 180 maybe relied upon.
[0081] The example conversation in Table 8 contains an explanation of error related to excessive contact force (“I noticed you applied excessive force ...”). Another error may be, for example, that a motor in the robot manipulator no is out of operation.
[0082] In a ninth embodiment, an exchange of speech data 302 and spoken guidance 311 is used to initialize the demonstration-based programming session, announce its purpose to the industrial robot 100, and to transition from a conversational mode to a learning mode. The exchange may have an appearance similar to the example conversation in Table 1 above.
[0083] As already mentioned, technical features from the first, second, third etc. embodiments can be advantageously combined.
[0084] The aspects of the present disclosure have mainly been described above with reference to a few embodiments and examples thereof. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the invention, as defined by the appended patent claims.
Claims
CLAIMS1. A method (200) of programming an industrial robot (100), which comprises a robot manipulator (no) and a robot controller (120), comprising: recording (212) movements of the robot manipulator during a demonstration-based programming session, to obtain a robot trajectory (301); capturing (214) speech data (302) while recording the movements of the robot manipulator; using a natural-language model (171), parsing (217) the speech data into at least one robot parameter (312) value and annotating the robot trajectory with the robot parameter value; and, on the basis of the annotated robot trajectory (313), generating (220) a robot program (314), characterized by further comprising: outputting (210, 215, 222) spoken guidance (311) to an operator (190) before or during the demonstration-based programming session, or between completion of the programming session and the generating of the robot program, or after an execution (221) of the robot program by the industrial robot.
2. The method (200) of claim 1, wherein the spoken guidance (311) includes a request for a demonstration of a portion of a robot task, the method further comprising: recording movements (212.1) of the robot manipulator during the requested demonstration.
3. The method (200) of claim 1 or 2, wherein the spoken guidance (311) includes a request for clarification or for additional information, the method further comprising: recording (214.1) speech data including the requested clarification or additional information.
4. The method (200) of any of the preceding claims, wherein the spoken guidance (311) includes a request to identify, before an end of the demonstration-based programming session, one or more objects (151, 152) involved in a robot task,the method further comprising: recording (211a) an image or video (315a) of said one or more objects, and / or receiving (211b) an indication (315b) of an image region in which said one or more objects are visible.
5. The method (200) of any of the preceding claims, wherein the spoken guidance (311) includes a request to indicate whether the programming session is still active.
6. The method (200) of any of the preceding claims, further comprising: providing (218) a confirmatory summary relating to one or more of: the robot trajectory (301), the captured speech data (302), at least one robot parameter (312), the annotated robot trajectory (313), wherein the spoken guidance (311) includes said confirmatory summary.
7. The method (200) of claim 6, wherein the generating (220) of the robot program is conditional upon receiving (219) the operator’s (190) approval of the confirmatory summary.
8. The method (200) of any of the preceding claims, further comprising: detecting (214.3) a question in the recorded speech data; and determining (214.4) an answer to the question, wherein the spoken guidance (311) includes said answer to the question.
9. The method (200) of claim 8, wherein the answer to the question is determined (212.4) by deriving it from a conversation context or from a knowledge database (180).
10. The method (200) of any of the preceding claims, wherein the spoken guidance (311) includes an instruction or recommendation related to the performing of a robot task.
11. The method (200) of any of the preceding claims, further comprising detecting (212.2, 214.2) an error during said recording (212) of robot manipulator movements or during the capturing (214) of speech data, wherein the spoken guidance (311) includes an explanation of the error.
12. The method (200) of any of the preceding claims, which is performed by a programming device (160) operable in at least a conversational mode, a learning mode and an inactive mode, the method further comprising:at or near the robot manipulator (no), giving a visual indication (118, 119) of a current mode of the programming device.
13. A programming device (160) for facilitating programming of an industrial robot (100), which comprises a robot manipulator (no) and a robot controller (120), the programming device comprising: a bidirectional speech interface (165), associated with a speech-to-text converter (165.2) and a text-to-speech converter (165.3); a conversation engine (165.1) configured to conduct a dialog with an operator (190) through the bidirectional speech interface; memory (162) and processing circuitry (161) configured to carry out the method (200) of any of claims 1 to 12.
14. A computer program (163) comprising instructions to cause the programming device (160) of claim 13 to carry out the method (200) of any of claims 1 to 12.