Method of providing a specialized large language model for controlling a machine
The method enhances LLMs for industrial control by augmenting operator instructions with spoken features and fine-tuning, addressing token constraints and data dependence, enabling efficient machine control on consumer hardware.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2026-03-05
AI Technical Summary
Existing large language models (LLMs) for robot control and industrial machinery are limited by token constraints and require extensive domain-specific training data, making them costly and inefficient.
A method to generate a specialized LLM by augmenting operator instructions with randomly sampled features of spoken natural language and fine-tuning a baseline LLM using example operator instructions, prompts, and machine parameter values, reducing dependence on domain-specific data.
The method enables effective control of industrial machines with reduced reliance on specific training data, allowing operation on consumer-level hardware and improving parsing of operator instructions into machine parameter values.
Smart Images

Figure EP2024073896_05032026_PF_FP_ABST
Abstract
Description
METHOD OF PROVIDING A SPECIALIZED LARGE LANGUAGE MODEL FOR CONTROLLING A MACHINETECHNICAL FIELD
[0001] The present disclosure relates to the field of industrial control. In particular, it presents methods and devices configured to use fine-tuning in order to provide a specialized large language model for controlling a machine such that the machine performs a task in the field of an industry branch.BACKGROUND
[0002] Large Language Models (LLMs), including Generative Pre-trained Transformers (GPTs), are increasingly being applied to robot control and robot programming. See for instance the applicant’s prior disclosure PCT / EP2023 / 087949. LLMs are also being used more generally in the control of industrial machinery and equipment.
[0003] While the LLM-oriented development is very fruitful on the whole, it is being held back by several factors. One is the token limitations of the available LLMs, e.g., limitations on the combined size of input (prompt) to the LLM and its output (completion). Another limiting factor is the scarcity, or cost, of good training data. Fine-tuning is a technique for specializing a generic, pre-trained baseline LLM for a particular task or particular domain. A fine-tuning scheme may include a phase of continued training of the baseline LLM on a domain-specific dataset. The continued training embeds extra information into the LLM and / or it conditions the type of outputs from the LLM based on the domain-specific dataset. Therefore, the need for good training data from the relevant domain is also felt by the fine-tuning technology in its generic form.
[0004] To promote the ongoing efforts of applying LLMs to machinery control, and to robot control in particular, it would be desirable to develop a new fine-tuning scheme that is less dependent on domain-specific training data than generic fine- tuning schemes.SUMMARY
[0005] One objective of the present disclosure is to make available a method of generating a specialized LLM for controlling a machine in order for the machine toperform a task in the field of an industry branch. The aimed-for method should be less dependent on training data from that industry branch than fine-tuning schemes according to the state of the art. The aimed-for method should be less dependent on machine-specific training data. The aimed-for method should be possible to execute without the availability of a baseline LLM that has been trained on the machine. A further objective is to generate a specialized LLM adapted to parse an operator instruction into a set of time-stamped machine parameter values, similarly to the processing of speech data captured while recording movements of a robot manipulator in the prior disclosure PCT / EP2023 / 087949. A still further objective is to make available a device and a computer program with these characteristics.
[0006] At least some of these objectives are achieved by the invention as defined by the independent claims. The dependent claims relate to advantageous embodiments of the invention.
[0007] In a first aspect of the present disclosure, there is provided a method of generating a specialized LLM for controlling a machine to perform a task in the field of an industry branch. The method comprises: generating a plurality of example operator instructions in written natural language using a source LLM, each example operator instruction relating to a task in the field of the industry branch; generating a plurality of example prompts in written natural language using the source LLM, wherein each example prompt is suitable for requesting an LLM to parse any of the example operator instructions into a set of time-stamped machine parameter values; augmenting the example operator instructions with randomly sampled features of spoken natural language (e.g., speed, prosody, intonation, volume, word stress) for a plurality of random combinations of an example prompt and an example operator instruction; causing the source LLM to parse the example operator instructions into a set of machine parameter values; and providing the specialized LLM by fine-tuning a baseline LLM based on the thus obtained example operator instructions, example prompts and sets of machine parameter values.
[0008] As used herein, an “industry branch” refers to an official or unofficial industry classification or industry taxonomy, which classifies companies, organizations and traders into industrial groupings based on them having similar production processes, similar raw materials and / or similar products. An industry branch is one of the industrial groupings in such a classification or taxonomy.Example industry branches include: agriculture, hunting and forestry; fishing; mining and quarrying; manufacture of food products, beverages, tobacco; manufacture of textiles; manufacture of leather; manufacture of wood products; manufacture of pulp, paper, paper products, printed products; manufacture of coke, refined petroleum products, nuclear fuel; manufacture of chemicals, chemical products; manufacture of rubber and plastics; manufacture of non-metallic mineral products; manufacture of metals and metal products; manufacture of machinery; manufacture of electric and optic equipment; manufacture of transport equipment; electricity gas and water supply; construction; wholesale and retail trade; repair of motor vehicles; transport, storage and communication; real estate and building activities; education; healthcare; domestic work. A task in the field of an industry branch is a task which occurs in the industry branch (e.g., such that known building blocks and other common knowledge of practitioners in the industry branch can be relied upon) and / or which shall be solved in a manner consistent with the practices in the industry branch (as regards dimensional accuracy, chemical purity, minimum hygiene requirements, safety, etc.).
[0009] As used herein, a set of “operator instructions” in written natural language is human-intelligible instructions for carrying out a task. The operator instructions may in particular be intelligible to a worker who is familiar with the industry branch in question. An operator instruction in this sense may contain an implicit or explicit reference (“as you see me doing”, “please pay attention to the movement”) to a demonstration of the task which is intended to be available while reading or listening to the operator instruction, so as to make the operator instruction easier to understand and / or more unambiguous. The demonstration maybe a demonstration in kinesthetic robot programming, wherein the robot poses are recorded by proprioceptive sensors, or in the form of an image or a video of the demonstration, etc. The demonstration may also be a robot-less demonstration, e.g., a so-called passive observation. For example, the demonstration maybe in the form of a video of a human who carriers out a pick-and-place task without any assistance from a robot; from such a passive observation, the robot learns skills from the video and reproduces the motions by which the task is completed. The combination of the demonstration and the operator instruction may be considered to be a multi-modal operator instruction. Further, the operator instructions maybe a task explanation, i.e., human-intelligible reasoning that describes the task to be solved and / or how theactions indicated by the operator instruction solve the task and / or a desired result of solving the task.
[0010] In this disclosure, therefore, the inventors propose an implementation of fine-tuning which performs particularly well in the context of industrial machinery control, including robot control, in a desired industry branch. The method according to the first aspect reduces the need for training data from the relevant industry branch - or eliminates it altogether - by instead extracting the necessary domainspecific knowledge from the source LLM. The source LLM need not be familiar with the machine or robot to be controlled either; this information maybe fed into an automated code-generation process downstream of the processing carried out by the specialized LLM.[oon] A wanted effect of the step of augmenting the example operator instructions with randomly sampled features of spoken natural language is to stop the LLM from learning to find patterns in a non-random way of adding timestamps to the words of the operator instruction. If the step of augmenting the example operator instructions with randomly sampled features of spoken natural language was replaced with deterministic timestamps with an increment of, say, 0.2 seconds per word (placeholder timestamps), then the LLM during training might learn only to process operator instructions with timestamps that obeyed this 0.2 second / word duration, while it would have a limited capability of handling operator instructions with differently spaced words or operator instructions with irregular timing. Thanks to the step of augmenting the example operator instructions with randomly sampled features of spoken natural language, the resulting specialized LLM is less likely to perceive word timing as semantically significant.
[0012] In the interest of computational feasibility, the baseline LLM is preferably smaller than the source LLM with respect to the number of parameters and / or with respect to the amount of training data. In some embodiments, the baseline LLM is a local LLM. A “local LLM”, as this term is used in the present disclosure, may refer an LLM that is small enough in size that it can run on consumer-level hardware, such as a processor with a performance similar to NVIDIA® GeForce RTX™ 3080, NVIDIA® GeForce RTX™ 4090 or a comparable graphics processing unit (GPU). Local LLMs differ in this sense from regular LLMs, which at the time of filing thepresent disclosure are conventionally run on large or very large GPU clusters and which consumer-grade local clients will usually access through a cloud service.
[0013] The further embodiments that will be presented in the below detailed description have proven to be advantageous in experiments conducted by the inventors. Supporting data will be presented.
[0014] In a second aspect of the present disclosure, there is provided a method of controlling a machine to perform a task in the field of an industry branch, comprising: performing the method according to the first aspect, whereby a specialized LLM is provided; capturing speech data containing an operator instruction in spoken natural language; parsing the operator instruction into at least one machine parameter value using the specialized LLM; and operating the machine in accordance with said at least one machine parameter value.
[0015] In a third aspect of the present disclosure, there is provided a control device for facilitating control of a machine adapted for an industry branch. The device comprises memory and processing circuitry configured to perform the method according to the first aspect.
[0016] The second and third aspects presented herein generally share the effects and advantages of the first aspect, and they can be embodied with a corresponding degree of technical variation.
[0017] The present disclosure further relates to a computer program containing instructions for causing a computer, or the control device in particular, to carry out the method according to the first aspect. The computer program may be stored or distributed on a data carrier. As used herein, a “data carrier” maybe a transitory data carrier, such as modulated electromagnetic or optical waves, or a non-transitory data carrier. Non-transitory data carriers include volatile and non-volatile memories, such as permanent and non-permanent storage media of magnetic, optical or solid-state type. Still within the scope of “data carrier”, such memories may be fixedly mounted or portable.
[0018] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to “a / an / the element, apparatus, component, means, step, etc.” are to be interpreted openly as referring to at least one instance of the element,apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order the steps are described, unless this is explicitly stated.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Aspects and embodiments are now described, by way of example, with reference to the accompanying drawings, on which: figure 1 shows a work area, in which a robot manipulator operates under the control of a robot controller, and a programming device supporting demonstration-based robot programming; figure 2 is a flowchart of a method of generating a specialized LLM for controlling a machine to perform a task in the field of an industry branch; figure 3 is a flowchart of a method of controlling a machine to perform a task in the field of an industry branch, wherein the method includes the method of figure 2; figure 4 illustrates information flows during an execution of the method of figure 2; figure 5 illustrates information flows during the use of a specialized LLM generated by the method of figure 2; figures 6 and 7 present experimental data recorded during experiments conducted by the inventors which support the advantageous performance of some embodiments disclosed herein.DETAILED DESCRIPTION
[0020] The aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, on which certain embodiments of the invention are shown. These aspects may, however, be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and to fully convey the scope of all aspects of the invention to those skilled in the art. Like numbers refer to like elements throughout the description.System overview
[0021] As explained initially, the present disclosure relates to the field of industrial control and particularly to technologies for controlling industrial machines.Industrial machines include machines which are conventionally used in the example industry branches enumerated above, ranging from large / complex machines to small / simple machines, such as assembly machines, packaging / labelling machines, chemical machinery, papermaking machinery, agricultural equipment, energy generators and converters, and industrial robots. In the terminology of the present disclosure, “industrial robot” is used in a broad sense, to cover in particular manufacturing robots, material-handling robots, assembly robots, cutting / welding robots, service robots, collaborative robots, hygiene robots, industrial robot tracks, industrial robot positioners. An industrial robot may be a stationary robot or a mobile robot, such as an automated guided vehicle (AGV), an autonomous mobile robot (AMR) or an autonomous mobile manipulator robot (AMMR). The term industrial robot covers the full range from lightweight robots designed to replace human manual work, over collaborative robots for supporting a human worker, all the way up to heavy-duty robots.
[0022] By way of illustration and not limitation, figure 1 shows an industrial robot 100 made up of a robot manipulator no and a robot controller 120. The robot manipulator no and the robot controller 120 are joined by a wired or wireless bidirectional data connection, which conveys control signals, sensor data etc.
[0023] The robot manipulator 110 includes an arm, which extends from a base 112 and is made up of structural elements 114 and at least one linear or rotary joint 113. The arm may further carry tools 111 that allow it to interact with objects 151 in the form of various workpieces, which are present in a work area 150 of the robot manipulator no. The workpieces are subject to manufacture, processing or handling by the robot manipulator 110. The work area 150 may further include non-workpiece objects 152, such as containers, fixtures, robot positioners, separators, insulators, supports etc. These objects 152 maybe generic or they maybe specifically adapted to the workpieces 151 handled by the robot manipulator 110. Apart from wear, staining etc., they are generally in the same condition at the beginning and end of a work cycle. The arm of the robot manipulator no is movable by action of internal motors, drives or actuators (not shown), and it includes transducers, sensors and other measuring equipment (not shown), from which the manipulator’s no current position, pose, technical condition, load etc. can be derived, to some degree ofaccuracy. The position, pose etc. of the robot manipulator no may in particular refer to a point on the arm, particularly to a tool-center point (TCP).
[0024] The positions, poses etc. may be expressed with respect to one of multiple possible reference frames, including a first fixed reference frame O0defined with respect to a point in the work area 150, a second fixed reference frame (not shown) defined with respect to the base 112, a first reference frame O defined with respect to an initial position 151 of a workpiece, or a second reference frame O2defined with respect to a position of a non-workpiece object 152 in the work area 150. Here, the first fixed reference frame O0is defined with respect to the point in the work area 150 in the sense that its origin is situated in that point. Alternatively, a reference frame may be defined with respect to a reference point in the sense that its origin is situated a predefined translation from the reference point. The first and second fixed reference frames, as well as any further reference frames which are independent of the positions 151, 152 of objects in the work area 150, maybe collectively referred to as neutral reference frames. All reference frames shown in figure 1 have a common orientation, i.e., their respective first, second and third axes are parallel. Without departing from the scope of the present disclosure, a further option is to use reference frames with mutually different orientations, including a reference frame which is oriented in accordance with a pose of an object 151, 152 in the work area 150 or another suitable reference objects. In particular, one may use a reference frame which has its origin situated in the TCP and is oriented parallel to the tool 111 of the manipulator no at all times.
[0025] The robot controller 120 comprises processing circuitry 121, a memory 122 and a communication interface 123. Example content of the memory 122 during operation may includeS: an operating system, basic settings, software implementing generic movements, sensing, self-monitoring, generically useful functionalities and services (all typically contributed by an original manufacturer), task- or role-specific configurations, configuration templates (typically contributed by a robot system integrator), and site-specific settings (typically contributed by an end user);C: robot programs or projects causing the robot manipulator no to perform useful or intended tasks in its work area 150.Processes that execute the robot programs may do so in accordance with the program-independent memory content S, e.g., by making calls to available functionalities, libraries, routines or parameter values therein. It is recalled that the memory 122 and processing circuitry 121 of the robot controller 120 maybe distributed and / or contain networked resources having a different physical localization than figure 1 suggests.
[0026] Each of the programs C may contain a plurality of movement instructions relating to locations such as points, poses, paths as well as modulated paths. A program maybe a compiled executable (binary) or a script. A movement instruction relating to a modulated path may be expressed as - or may include - a process-on- path instruction. The programs C maybe created by an operator 190 with the aid of a robot programming device (programming station) or a general-purpose computer, or they maybe created directly at the robot controller 120 if it has an operator interface. In the first two cases, versions of the programs C may be downloaded to the robot controller 120 over a wired or wireless connection or by being temporarily stored on a portable memory. A robot program may be created by means of demonstration-based programming according to the teachings in the applicant’s prior disclosures PCT / EP2023 / 087949 and PCT / EP2024 / 066288.
[0027] As used herein, “demonstration-based” robot programming (or programming by demonstration [PbD], or programming from demonstration) includes a step where the operator performs a portion of a robot task or guides the robot manipulator such that it performs a portion of this robot task. The operator’s activities are imaged or, in other ways, observed by a programming device. Demonstration-based robot programming includes kinesthetic programming as a special case. The term kinesthetic programming - by allusion to the robot’s proprioceptive ability to sense its own position, pose and movements - is a programming approach in which the operator physically moves the robot manipulator to imitate the desired motions. The terms kinesthetic programming and lead-through programming are synonymous or at least partially overlapping in meaning. The state of the robot during a kinesthetic programming session is typically recorded by means of the robot’s onboard sensors, e.g., joint angles and torques. As used herein, kinesthetic programming also includes teleoperation as a special case, where the movement of the robot manipulator is controlled by an external input to the robot through a joystick, graphical user interface orother input means; kinesthetic programming by teleoperation does not require the operator to be present in the work area of the robot.
[0028] A dedicated programming device 160 is shown in the right-hand portion of figure 1. From the programming device 160, the created programs C can be transmitted to the communication interface 123 of the robot controller 120 and then stored in the memory 122 where they are available for execution. It is noted that the programming device 160 may be implemented as a set of collaborating components within the robot controller 120. In fact, the components can be shared with the robot controller 120, e.g., by using the processing circuitry 121 for the dual purposes of robot control and programming and / or using the memory 122 for the same dual purposes. In other words, the programming device 160 may constitute a portion of a multi-purpose device; it need not be a standalone device or a device with programming as its sole or main purpose.
[0029] The programming device 160 - which is shown as a standalone device in the non-limiting example of figure 1 - comprises processing circuitry 161, memory 162 with executable software 163, and at least one communication interface 164, 165. The programming device 160 may further comprise at least one operator interface 166, in particular a graphical operator interface or graphical user interface (GUI).
[0030] The programming device 160 further has access to an imaging device (e.g., camera, video camera, depth camera, lidar, radar) 167, by which images or video of the demonstration-based programming session can be captured. In particular, the imaging device may be a depth camera, lidar or radar, which in addition to - or instead of - the two-dimensional appearance of an object determines a depth coordinate of the object, e.g., by time-of-flight measurements, triangulation, reflection or other per se known techniques. An RGB-D camera is an example of a depth camera. The imaging device 167 maybe a part of the programming device 160, or it maybe integrated into the industrial robot 100 and optionally be used for other tasks as well.
[0031] The programming device 160 is optionally enabled to capture speech data, which could represent utterances or narration by the operator 190 during the programming session. The programming device 160 may for this purpose utilize one or more acoustic transducers (microphones) 168 arranged in the vicinity of the operator’s 190 normal position. Alternatively, the speech data maybe captured by a portable microphone or a headset worn by the operator 190. To parse speech data,the programming device 160 may utilize a natural-language model, which may be either stored internally or - as in the example configuration shown in figure 1 - may be available from a host computer or an external memory 170.
[0032] The natural-language models stored in the external memory 170 may constitute a large language model (LLM), that is, a type of machine-learning algorithm which has been trained on very large datasets using deep learning techniques to be able to perform natural-language processing (NLP) tasks. Example NLP tasks are recognizing, summarizing, translating, predicting and generating plausible textual content. A very large dataset in this sense may include of the order of one million parameters, such as tens of millions of parameters, such as hundreds of millions of parameters. An LLM may have a transducer architecture, particularly a transducer architecture with four cascaded key links when it processes the input data, namely: word embedding, position encoding, self-attention mechanism, feedforward neural network. At the time of filing this disclosure, noteworthy example LLMs include Bidirectional Encoder Representations from Transformers (BERT), Bard, BLOOM, Claude 2, various versions of Generative Pre-trained Transformer (GPT), Llama, PaLM 2, RoBERTa, T5, LaMDA, Turing NLG. LLMs include, as a special case, multimodal models, such as Gemini. One benefit of an LLM is that the vocabulary is practically open-ended. The operator 190 can start using it without prior training. The demonstration-based programming can be carried out without requiring the operator to be extremely focused on using the right command words (or avoiding them) and / or syntax.
[0033] During a demonstration-based programming session, the programming device 160 has access to position data representing an actual position, a recorded position or recorded movements of the robot manipulator no, as well as images or video of the robot programming session. The position data may for example be obtained through the intermediary of the robot controller 120, which monitors position data in the normal course of its operation; alternatively, the programming device 160 is granted access to corresponding signals from the transducers, sensors or other measuring equipment in the robot manipulator no.LLM fine-tuning method
[0034] Turning to the flowchart in figure 2, embodiments of a method 200 of generating a specialized LLM for controlling a machine to perform a task in the fieldof an industry branch are illustrated. Not all steps shown in figure 2 are necessarily carried out in all embodiments. The output of the method 200 may include, for example, a full or partial definition of the specialized LLM, such as set of weights of an artificial neural network. The specialized LLM may be used in a process of generating a program to be executed on the machine, such as a robot program C to be executed by the robot controller no of the industrial robot 100 in figure 1.
[0035] The method 200 in figure 2 may be executed on a general-purpose processor and on the programming device 160 in particular. As mentioned, the programming device 160 may constitute a portion of a multi-purpose device. The method 200 may be considered to be a description of the behavior which the programming device 160 is configured to carry out. The method 200 may as well correspond to the behavior of the robot controller 120 in a special programming mode, which the robot controller 120 can be requested to enter. Instructions for causing a computer - or the programming device 160 in particular - to carry out the method 200 maybe provided in the form of a computer program 164 (see figure 1).
[0036] The functioning and other characteristics of the method 200 may be better understood from the illustration in figure 4 of the data flows which occur during an execution of the method 200. Here, the following reference numbers are used:401* example operator instructions,401** example operator instructions with features of spoken natural language,402* example prompts,403* sets of machine parameter values provided from example input data.
[0037] A first step 210 of the method 200 is carried out using a source LLM 171.In this step 210, the source LLM 171 is requested to generate a plurality of example operator instructions in written natural language. The request to the source LLM 171 specifies that each example operator instruction shall relate to a task in the field of the industry branch. A task in the field of the industry branch, in this sense, is one which occurs in the industry branch and / or which shall be solved in a manner consistent with the practices in the industry branch.
[0038] As already explained, operator instructions in written natural language are human-intelligible instructions for carrying out a task. The operator instructions may in particular be intelligible to a worker who is familiar with the industry branch inquestion. An operator instruction may contain an implicit or explicit reference to a demonstration of the task (e.g., “look here, you gotta inspect the belt from end to end” in task_explanation_3 in Table 2), which demonstration maybe a kinesthetic robot programming demonstration. The combination of the demonstration and the operator instruction may be considered to be a multi-modal operator instruction. The operator instructions may in particular be a task explanation.
[0039] The operator instructions to be generated in step 210 are example operator instructions in the sense that they are generated at random, they are fictitious and / or unrelated to any request formulated by a user or a machine operator. Their main purpose is to provide knowledge expected to be useful to help the specialized LLM process actual (i.e., non-example) operator instructions. (Below, in the description of step 211, the term example prompt is used in the same or a similar sense.)
[0040] The source LLM 171 is a general-purpose LLM which embodies world knowledge from a plurality of different domains, such as a plurality of industry branches. It is not essential for the present teachings to confirm or ascertain that the source LLM 171 embodies world knowledge from the industry branch to which the generated example operator instructions relate. The source LLM 171 may even lack prior exposure to the industry branch, but it might be able to fill any gaps by interpolation and deductions. Further, it is not necessary for the source LLM 171 to be knowledgeable about the machine (e.g., an industrial robot 100) to be controlled. The source LLM 171 does not need to have been exposed to training data relating to the machine. The source LLM 171 may be machine-agnostic.
[0041] The request may be formulated as a prompt to the source LLM 171. Examples of prompts which are suitable for carrying out the first step 210 are disclosed in the Examples section.
[0042] The request may in particular instruct the source LLM 171 to generate each example operator instruction with a length of at least 200 words, preferably at least 400 words. The number of words refer to an English-language version of an example operator instruction, whether this is the original version or a translation. Although requests (prompts) with this substantive content have proved to perform well in the inventors’ experiments, it is not essential to the present invention that the actual outputs of the source LLM 171 have the instructed number of words. In fact, the relatively high number of word-counting problems, which GPT users report aroundthe time of filing the present disclosure, suggests that some discrepancies may be expected.
[0043] In particular, the number of example operator instructions generated in step 210 maybe of the order of thousands. In one example, at least 2,000 example operator instructions maybe generated in step 210. In another example, at least 5,000 example operator instructions are generated. In another example, at least 10,000 example operator instructions are generated.
[0044] Also the second step 211 of the method 200 is carried out using the source LLM 171. The source LLM 171 is requested to generate a plurality of example prompts in written natural language. The request includes an instruction that each example prompt shall be suitable for requesting an LLM to parse any of the example operator instructions into a set of time-stamped machine parameter values. The request may be formulated as a prompt to the source LLM 171. Accordingly, the example prompt is suitable for requesting the source LLM 171 to parse the example operator instructions into machine parameter values. Example prompts suitable for carrying out the second step 211 are disclosed in the Examples section.
[0045] The request may in particular instruct the source LLM 171 to generate said plurality of example prompts such that the maximum length in the plurality is at least twice the minimum length in the plurality. Preferably, the maximum length shall be at least three or four times the minimum length in the plurality of example prompts. Including this instruction in the request in step 211 is optional; it may have the advantage of preventing the source LLM 171 from generating a population of example prompts which are all of very low complexity, which may be detrimental to the processing in step 214.
[0046] In a third step 212 of the fine-tuning method 200, the example operator instructions are augmented with randomly sampled features of spoken natural language. The features of spoken natural language may include one or more of speed (speech rate), prosody, intonation, volume and word stress. In particular, the features of spoken natural language may include timing of single words. The processing in step 212 may include adding metadata to the example operator instructions which indicates or quantifies the realization of these features. Preferably each metadata entry is associated with one word or one syllable in an example operator instruction. The metadata maybe added inline or in a different layer / channel.
[0047] In one example, an operator instruction containing the following text fragment is processed:“Alright, so the first thing we've gotta do”As an effect of the processing in step 212, start and end times (in seconds, counted from a reference point) are inserted after each word. The text fragment changes into the following annotated text:“Alright, 8.18 9.08 so 10.11 10.39 the 10.43 10.64 first 10.69 n.oi thing 11.0711.84 we've 11.9 12.4 gotta 12.48 12.91 do 12.9713.13”
[0048] The features to be adduced to the example operator instructions may be randomly sampled from a natural language model. In particular, a model of speech rate in spoken English can be constructed by applying the data reported in Huang, Lan-fen & Graf, Tomas, “Speech Rate and Pausing in English: Comparing learners at different levels of proficiency with native speakers”, Taiwan Journal ofTESOL, vol. 17 (2020), no. 1, pp. 57-86, DOI: 1O.3O397 / TJTESOL.2O2OO4_17(1).OOO3. Such a model of speech rate may for example relate the duration of a word to various phonetic or semantic characteristics of the word, so that the random sampling process may include inputting said phonetic or semantic characteristics of a word to the model and obtaining a (non-deterministic) word duration as output.
[0049] As explained in an earlier section of this disclosure, an effect of augmenting the example operator instructions in this way is to prevent an unintended learning that the word timing (or another feature of spoken language) would be information-carrying or somehow necessary for parsing the operator instruction into machine parameter values. In a further development, the augmentation in step 212 purposefully uses different timestamp formats, so as to prevent an unintended learning that the timestamp format would be significant.
[0050] In a fourth step 213 of the method 200, the source LLM 171 is caused to parse the example operator instructions (augmented with the sampled features of spoken natural language) into a set of machine parameter values, and indeed for a plurality of random combinations of an example prompt and an example operator instruction. Each example prompt maybe used once or multiple times. Each example operator instruction maybe used once or multiple times. Within the method 200, according to some embodiments, the number of example prompts is equal to thenumber of example operator instructions (e.g., 1,000 or 10,000), which is in turn equal to the number of combinations. In other embodiments, the number of example operator instructions is greater (or significantly greater) than the number of example prompts, and the number of combinations is greater (or significantly greater) than the number of example operator instructions.
[0051] The parsing in step 213 may be initiated by the entity executing the method 200, namely, by inputting a request (prompt) to the source LLM 171. Example prompts suitable for carrying out the fourth step 213 are disclosed in the Examples section below.
[0052] The machine parameters to which step 213 refers is a generalization of the concept of robot parameters discussed in detail in the applicant’s prior disclosures PCT / EP2023 / 087949 and PCT / EP2024 / 066288. For example, the machine parameters may be a state of a tool or subsystem of the machine, a movement execution parameter, a degree of compliance with a predefined machine behavior (e.g., motion), a degree of movement or processing precision, a reference frame for the machine’s movement, a parameter relating to a drive system of the machine. In the special case of an industrial robot, the machine parameter maybe a state of a tool carried by the robot manipulator, a movement execution parameter, a degree of compliance with the robot trajectory, a degree of movement precision, a reference frame, a drive system parameter. Further possible machine parameters are found in the Examples sections below.
[0053] Next, in a fifth step 214 of the method 200, a specialized LLM 173 is provided by fine-tuning a baseline LLM 172 based on the thus obtained (augmented) example operator instructions 401**, example prompts 402* and sets of machine parameter values 403*. In figure 4, the inputs to the fine-tuning process are indicated by dashed arrows. The example operator instructions, example prompts and sets of machine parameter values are provided to the fine-tuning as triplets. More precisely, an aim of step 214 is to provide the specialized LLM 173 such that it returns a particular set of machine parameter when it is prompted by the associated example prompt to parse the associated example operator instruction.
[0054] The paradigm of fine-tuning includes techniques for specializing a generic, pre-trained baseline LLM for a particular task or particular domain by means of continued training. Fine-tuning may include so-called transfer learning as a specialcase; wherein the continued training is restricted to a subset of the baseline LLM, such as one or more layers of an artificial neural network. In contrast to available fine-tuning methods, where the continued training uses a domain-specific dataset, the fine-tuning proposed in the present disclosure uses example data (example operator instructions, example prompts and sets of machine parameter values) generated by the source LLM 171.
[0055] The fine-tuning in step 214 may include performing one or more of- full fine-tuning;- a low-rank adapter method (LoRA);- a quantized LoRA method (QLoRA);- weight-decomposed low-rank adaptation method (DoRA; see the recent preprint Liu et al., “DoRA: Weight-Decomposed Low- Rank Adaptation”, arXiv:24O2.09353 [cs.CL], retrieved from arXiv.org); and- a supervised fine-tuning method.It is not essential to the present invention which option or options are selected from the above. At the time of filing, the inventors have been able to confirm the efficiency of the LoRA and QLoRA methods as far as time and memory usage is concerned.
[0056] The baseline LLM 172 is preferably smaller than the source LLM 171 with respect to the number of parameters and / or with respect to the amount of training data used.
[0057] In some embodiments of the method 200, the baseline LLM 172 is a local LLM in the sense explained and exemplified above. Briefly put, a local LLM is an LLM which allows execution on consumer-level processing circuitry. Examples of local LLMs include the Phi models from Microsoft Corp. (Phi 1, Phi 2, Phi 3), the Llama models offered by Facebook Inc. (Llama 1, Llama 2, Llama 3, including the LLaMa 2 7B model), and the Gemma models offered by Google Inc. At the time of filing the present disclosure, it was reasonable to expect that LLMs with less than 13 billion parameters would be able to run on consumer-level hardware, or would do so with some adaptation. Because the fine-tuning 214 does not appreciably increase the data volume occupied by the LLM, the fact that the baseline LLM 172 is a local LLM generally implies that the requirements of the specialized LLM 173 are comparable tothose of a local LLM too. Accordingly, the specialized LLM 173 can run on a programming device 160 with non-special hardware, such as a consumer-level processor or consumer-level computer.
[0058] After completion of the fifth step 214, the example operator instructions, example prompts and sets of machine parameter values may be deleted or marked as free to be deleted / overwritten.Machine control method
[0059] One possible use of the specialized LLM 173 obtained by executing the fine-tuning method 200 is to assist automated programming of a machine or an industrial robot. In particular, the specialized LLM 173 may be used to parse an operator instruction into a set of time-stamped machine parameter values, similar to the processing of speech data in the prior disclosure PCT / EP2023 / 087949.
[0060] This intended use is illustrated by the machine-control method 300 shown in figure 3, in which the fine-tuning method 200 constitutes an early stage.
[0061] The machine-control method 300 further comprises a step 310 of capturing speech data containing an operator instruction in spoken natural language. The speech data may include utterances by an operator and it may be captured using the transducer 168. With the consumer-level processing power available at the time of filing the present disclosure, due to token limitations of the LLM used in step 311, it is possible to process a captured speech data track with a duration of up to 2-3 minutes. The speech data may be transcribed using any suitable transcription software, such as Whisper, which is available from OpenAI, Inc. or its subsidiaries.
[0062] Then, in a subsequent step 311, the specialized LLM 173, which has been provided by the execution of the fine-tuning method 200, is requested to parse the operator instruction into at least one machine parameter value. The parsing operation is illustrated in figure 5, where the following reference numbers are used:401 (actual) operator instruction,402 (actual) prompt for the specialized LLM 173 to perform parsing,403 set of machine parameter values parsed from the operator instruction.
[0063] The request may be a prompt with a similar form and content as the example prompts used in the fine-tuning method 200. A benefit of the parsingprocess is more compact representation (i.e., with fewer tokens) of a transcribed speech data track.
[0064] The machine is then operated, in a further step 312, in accordance with said at least one machine parameter value.
[0065] In some embodiments of the machine-control method 300, the speech data has been captured (step 310) during a machine programming session, in particular a demonstration-based machine programming session. In such embodiments, the machine-control method 300 includes generating a program C to be executed by the machine, wherein the program C incorporates said at least one machine parameter value. The program C is generated such that it can be executed by the machine or by a controller associated with the machine. This may include, for example, using commands selected from a predefined set of machine commands executable by the machine or machine controller. As such, because machine-specific knowledge is introduced in the program-generation step, the fine-tuning method 200 may rely on a source LLM 171 which is machine-agnostic.
[0066] An example of automatically generating a program C in the particular context of an industrial robot is described in detail in PCT / EP2024 / 066288. According to that disclosure, a combination of a recorded robot trajectory (e.g., a geometric description or a joint-space description) and robot parameter values indicating desired reference frames to be used to express the robot movements. The combination is converted into a sequence of robot commands that cause the robot manipulator no to reproduce the trajectory, while applying the desired reference frames. The robot commands may be selected from a predefined set of robot commands executable by the robot controller 120, such as commands compliant with the RAPID™ robot programming language. The conversion may include sampling the trajectory into a sequence of discrete points and a second substep of selecting robot commands that cause the robot manipulator 110 to move between each pair of consecutive discrete points.Example 1
[0067] To validate the technologies disclosed herein, the inventors used GPT-4 as source LLM 171 and further, as baseline LLM 172, LLaMa 2 7B and LLaMa 3. GPT-4 is available from OpenAI, Inc. or its subsidiaries, and LLaMa 2 7B is available fromMeta Platforms, Inc. Approximately 13,000 example operator instructions were generated and used for fine-tuning by means of QLoRA (4 bit).
[0068] In detail, the prompt shown in Table 1 was used to execute step 210 in the inventors’ implementation of the fine-tuning method 200. The result is presented in Table 2.The expression “List of 10 tasks:” causes the source LLM 171 to output the tasks directly, i.e., without an introductory phrase such as “Sure, I will now give you the 10 tasks that you requested”, which could make parsing more failure prone.It is seen in Table 2 that the output from the source LLM 171 begins with a table of contents, which is not necessary for the subsequent processing steps.
[0069] The prompt shown in Table 3 was used to execute step 211 in the inventors’ implementation of the fine-tuning method 200. The operator instruction is referred to as a “task explanation”. The result is presented in Table 4.For the avoidance of doubt, it is remarked that the ten “requirements” in Table 3 are not in any kind or relationship - let alone a one-to-one relationship - with the ten outputs in Table 4.
[0070] All words of the operator instructions were augmented with randomly sampled start and end times. The example inline metadata format presented above in connection with step 212 was used.
[0071] Combinations were formed based on the operator instructions in Table 2 and the prompts in Table 4 and used to execute step 213 in the inventors’ implementation of the fine-tuning method 200. An example combination is shown in Table 5, where the operator instruction has been abridged (“[...]”). The result is presented in Table 6.The string “\u2013”, which represents a non-ASCII character generated during the transcription of the speech data, can be fed to the fine-tuning process without detriment. Alternatively, it may be removed.
[0072] The performance of the baseline LLM 172 (i.e., LLaMa 2 7B) and the specialized LLM 173 (i.e., LLaMa 2 7B after the fine-tuning in step 214) were very different. For the example prompt in Table 7, in which the operator instruction (“task description”) has been abridged (“[...]”), LLaMa 2 7B did not yield any output at all. After fine-tuning it provided the output in Table 8, which maybe considered usable.
[0073] As mentioned, the present disclosure is not limited to the control of industrial robots but relates generally to technologies for controlling industrial machines. It is evident from Table 2 that the tasks for which the successful performance of the fine-tuning method 200 was demonstrated are not restricted to the robotics field.Example 2
[0074] Figures 6 and 7 show results from systematic performance evaluations, in which the inventors compared the performance of different LLMs available at the time of filing the present disclosure, and various improvements thereof. In general terms, the evaluation makes reference to ground-truth data (a YAML file) generated using GPT-4 for the same type of prompts. The inventors compared this ground truth with outputs from other LLMs. Most of the performance metrics used take values in the range [o, 1], where 1 represents the highest similarity between the output and the ground truth. One of the performance metrics was BERTscore, which is adapted for evaluating similarity of generated text with ground truth text. For details see the preprint Zhang et al., “BERTScore: Evaluating Text Generation with BERT”, arXiv:i9O4.O9675 [cs.CL], retrieved from arXiv.org. Separate evaluations were performed on text relating to the robot-motion parameters step similarity, step description, start time, end time, speed, precision, exact following and reference frame. A mean similarity value was computed for each individual parameter throughout the YAML file.
[0075] Figure 6 reports on results from the following models:- baseline GPT-3.5: single prompt with full audio data;- test GPT-3.5: two-step prompt with simplified audio data;- improved GPT-3.5: two-step prompt with simplified audio data and new information for reference frame (for expressing robot motion);- improved GPT-4: two-step prompt with simplified audio data and new information for reference frame; and- final GPT-3.5: single prompt with simplified audio data.The reported values are mean values over thirteen demonstrations.
[0076] Figure 7 presents results from the following models:- baseline LLaMa 2: LLaMa 2 7B chat (llama-2-7b-chat-hf);- baseline LLaMa 3: LLaMa 3 8B Instruct (llama-3-8b-instruct);- fine-tuned LLaMa 2: merged_llama-2-7b-chat-chulengo-o.i.i: ik dataset of simple YAML output (few fields), 1 epoch;- fine-tuned LLaMa 2: merged_llama-2-7b-chat-chulengo-i.o.i: 13k dataset of complex YAML output (multiple fields), 1 epoch;- fine-tuned LLaMa 3: merged-llama-3-8b-instruct-chulengo-o.i: ik dataset of simple YAML output (few fields), 1 epoch;- fine-tuned LLaMa 3: merged-llama-3-8b-instruct-chulengo-i.o: 13k dataset of complex YAML output (multiple fields), 1 epoch;- fine-tuned LLaMa 3: merged-llama-3-8b-instruct-chulengo-i.i: 13k dataset of complex YAML output (multiple fields), 2 epochs;- fine-tuned LLaMa 3: merged-llama-3-8b-instruct-chulengo-i.2: 13k dataset of complex YAML output (multiple fields), 3 epochs;- fine-tuned LLaMa 3: merged-llama-3-8b-instruct-chulengo-i.2 (see previous item), in quantized mode (QM).In the QM, the inventors applied 4 -bit quantization, for which BitsAndBytes is an example implementation (see github.com / bitsandbytes-foundation / bitsandbytes). Each value is an average over twelve demonstrations.
[0077] As can be seen from a comparison of reported values in figures 6 and 7, the fine-tuned Llama 3 model achieves an overall similarity of 0.898 and a step similarityof 0.742, which are better than GPT-3.5 and comparable to GPT-4. Generally, the results reported in figures 6 and 7 are very promising and they attest to the potential usefulness of the technologies disclosed herein.
[0078] The aspects of the present disclosure have mainly been described above with reference to a few embodiments. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the invention, as defined by the appended patent claims.
Claims
CLAIMS1. A method (200) of generating a specialized large language model, LLM, for controlling a machine to perform a task in the field of an industry branch, the method comprising: using a source LLM (171), generating (210) a plurality of example operator instructions in written natural language, each example operator instruction relating to a task in the field of the industry branch; using the source LLM, generating (211) a plurality of example prompts in written natural language, each example prompt being suitable for requesting an LLM to parse any of the example operator instructions into a set of time-stamped machine parameter values; augmenting (212) the example operator instructions with randomly sampled features of spoken natural language; for a plurality of random combinations of an example prompt and an example operator instruction, causing (213) the source LLM to parse the example operator instructions into a set of machine parameter values; and providing (214) the specialized LLM (173) by fine-tuning a baseline LLM (172) based on the thus obtained example operator instructions, example prompts and sets of machine parameter values.
2. The method (200) of claim 1, wherein the features of spoken natural language include timing of single words.
3. The method (200) of claim 1 or 2, wherein the baseline LLM is smaller than the source LLM with respect to the number of parameters and / or with respect to the amount of training data.
4. The method (200) of any of the preceding claims, wherein the source LLM (171) is machine-agnostic.
5. The method (200) of any of the preceding claims, wherein the baseline LLM (172) is a local LLM.
6. The method (200) of any of the preceding claims, wherein the source LLM (171) is instructed to generate (210) each example operator instruction with a length of at least 200 words, preferably at least 400 words.
7. The method (200) of any of the preceding claims, wherein the plurality of example operator instructions is generated (210) such that at least one of the example operator instructions contains a reference to a demonstration of the task.
8. The method (200) of any of the preceding claims, comprising generating (210) at least 2,000 example operator instructions, preferably at least 5,000 example operator instructions, preferably at least 10,000 example operator instructions.
9. The method (200) of any of the preceding claims, wherein said plurality of example prompts is generated (211) such that the maximum length is at least twice the minimum length, preferably at least three or four times the minimum length.
10. The method (200) of any of the preceding claims, wherein fine-tuning (214) the baseline LLM (172) includes performing one or more of: full fine-tuning; a low-rank adapter, LoRA, method; a quantized LoRA, QLoRA, method; weight-decomposed low-rank adaptation, DoRA, method; a supervised fine-tuning method.
11. The method (200) of any of the preceding claims, wherein the machine is an industrial robot (100).
12. A method (300) of controlling a machine to perform a task in the field of an industry branch, comprising: performing the method (200) of any of the preceding claims, whereby a specialized LLM (173) is provided; capturing (310) speech data containing an operator instruction in spoken natural language; parsing (311) the operator instruction into at least one machine parameter value using the specialized LLM; andoperating (312) the machine in accordance with said at least one machine parameter value.
13. The method (300) of claim 12, wherein the speech data is captured (310) during a machine programming session, the method further comprising generating a program (C) to be executed by the machine, wherein the program incorporates said at least one machine parameter value.
14. A control device (160) for facilitating control of a machine (100) adapted for an industry branch, the device comprising memory (161) and processing circuitry (162) configured to: using a source LLM (171), generate a plurality of example operator instructions in written natural language, each operator instruction relating to a task in the field of the industry branch; using the source LLM, generate a plurality of example prompts in written natural language, each example prompt suitable for requesting an LLM to parse any of the example operator instructions into a set of time-stamped machine parameter values; augment the example operator instructions with randomly sampled features of spoken natural language; for a plurality of random combinations of an example prompt and an example operator instruction, cause the source LLM to parse the example operator instructions into a set of machine parameter values; and provide a specialized LLM (173) by fine-tuning a baseline LLM (172) based on the thus obtained example operator instructions, example prompts and sets of machine parameter values, wherein the specialized LLM is provided for use in controlling the machine to perform a task in the field of the industry branch.
15. A computer program (163) comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 13.
Citation Information
Patent Citations
Method and device for speech-supplemented kinesthetic robot programming
WO2025140777A1
Method and device for demonstration-based robot programming with adaptive reference frames
WO2025256740A1
Robot task analysis method and device based on large language model and readable medium
CN116861921A
Voice interaction method, server and computer readable storage medium
CN117373456A
Training method and device of action plan generation model based on large language model
CN118152528A