Method, apparatus, device, medium and product for providing an interaction

WO2026188396A1PCT designated stage Publication Date: 2026-09-17ABB (SCHWEIZ) AG +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081818
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2026-09-17

Smart Images

  • Figure CN2025081818_17092026_PF_FP_ABST
    Figure CN2025081818_17092026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of present disclosure provide a method, an apparatus, an electronic device, a computer-readable storage device, and a computer program product for providing an interaction. The method comprises receiving, from a user, a user input comprising one or more instructions for controlling the robot. The method further comprises determining, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on an electronic device and associated with the controlling of the robot. The method further comprises presenting, to the user, at least one operation to be performed on the electronic device. In this way, user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, APPARATUS, DEVICE, MEDIUM AND PRODUCT FOR PROVIDING AN INTERACTIONFIELD

[0001] Embodiments of the present disclosure generally relate to the field of computer technology and in particular, to a method, an apparatus, an electronic device, a computer-readable medium and a computer program product for providing an interaction with an electronic device.BACKGROUND

[0002] A robot is a machine that can perform tasks automatically. The robot is usually controlled by a program, and has perception, decision-making and execution capabilities. Robots can imitate human behavior and complete complex or repetitive tasks. With the development of artificial intelligence and sensor technology, robots are becoming more intelligent and autonomous, and are able to adapt to diverse environments and collaborate with people.SUMMARY

[0003] In general, various example embodiments of the present disclosure provide a method, an apparatus, an electronic device, a computer-readable storage device, and a computer program product for providing an interaction with an electronic device.

[0004] In a first aspect, it is provided a method for providing an interaction with an electronic device. The method comprises receiving, from a user, a user input comprising one or more instructions for controlling the robot. The method further comprises determining, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on the electronic device and associated with the controlling of the robot. The method further comprises presenting, to the user, at least one operation to be performed on the electronic device.

[0005] In a second aspect, it is provided an apparatus for providing an interaction with an electronic device. The apparatus comprises a receiving module configured to receive, from a user, a user input comprising one or more instructions for controlling the robot. The apparatus further comprises a determining module configured to determine, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on the electronic device and associated with the controlling of the robot. The apparatus further comprises a presenting module configured to present, to the user, at least one operation to be performed on the electronic device.

[0006] In a third aspect, it is provided an electronics device. The electronics device comprises a processor; and a memory coupled to the processor, wherein the memory has instructions stored therein, and the instructions, when executed by the processor, cause the device to execute actions of the first aspect.

[0007] In a forth aspect, it is provided a computer-readable medium. The computer-readable medium comprises instructions stored therein, which when executed by a processor, cause the processor to perform methods of the first aspect.

[0008] In a fifth aspect, it is provided a computer program product. The computer program product comprises instructions stored therein, which when executed by a processor, cause the processor to perform methods of the first aspect.

[0009] It is to be understood that the Summary is not intended to identify key or essential features of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become readily comprehensible through the description below.DESCRIPTION OF DRAWINGS

[0010] Through the following detailed descriptions with reference to the accompanying drawings, the above and other objectives, features and advantages of the example embodiments disclosed herein will become more comprehensible. In the drawings, several example embodiments disclosed herein will be illustrated in an example and in a non-limiting manner, wherein:

[0011] FIG. 1 illustrates a schematic diagram of an example environment in which a plurality of embodiments of the present disclosure can be implemented;

[0012] FIG. 2 illustrates a schematic diagram of another example environment in which a plurality of embodiments of the present disclosure can be implemented;

[0013] FIG. 3 illustrates an example scenario of providing an interaction with an electronic device in accordance with some embodiments of the present disclosure;

[0014] FIG. 4 illustrates an example process of providing an interaction with an electronic device in accordance with some embodiments of the present disclosure;

[0015] FIG. 5 illustrates a diagram of an example retrieval-augmented generation (RAG) in accordance with some embodiments of the present disclosure;

[0016] FIG. 6 illustrates a flowchart of an example method for providing an interaction with an electronic device in accordance with some embodiments of the present disclosure;

[0017] FIG. 7 illustrates a block diagram of an example apparatus for providing an interaction with an electronic device in accordance with some embodiments of the present disclosure; and

[0018] FIG. 8 illustrates a block diagram illustrating an electronic device in accordance with some embodiments of the present disclosure.

[0019] Throughout all the drawings, the same or similar reference numerals represent the same or similar elements.DETAILED DESCRIPTION OF EMBODIMENTS

[0020] Principles of the present disclosure will now be described with reference to several example embodiments shown in the drawings. Though example embodiments of the present disclosure are illustrated in the drawings, it is to be understood that the embodiments are described only to facilitate those skilled in the art in better understanding and thereby achieving the present disclosure, rather than to limit the scope of the disclosure in any manner.

[0021] The term comprises "or" includes "and" its variants are to be read as open terms that mean "includes, but is not limited to" . The term "or" is to be read as "and / or" unless the context clearly indicates otherwise. The term "based on" is to be read as "based at least in part on" . The term "being operable to" is to mean a function, an action, a motion or a state can be achieved by an operation induced by a user or an external mechanism. The term "one embodiment" and "an embodiment" are to be read as "at least one embodiment" . The term "another embodiment" is to be read as "at least one other embodiment" . The terms "first" , "second" , and the like may refer to different or same objects. Other definitions, explicit and implicit, may be included below. A definition of a term is consistent throughout the description unless the context clearly indicates otherwise.

[0022] The functions or algorithms described herein may be implemented in software in one embodiment. The software may consist of computer executable instructions stored on computer readable media or computer readable storage device such as one or more non-transitory memories or other type of hardware-based storage devices, either local or networked. Further, such functions correspond to modules, which may be software, hardware, firmware or any combination thereof. Multiple functions may be performed in one or more modules as desired, and the embodiments described are merely examples. The software may be executed on a digital signal processor, ASIC, microprocessor, or other type of processor operating on a computer system, such as a personal computer, server or other computer system, turning such computer system into a specifically programmed machine.

[0023] The functionality can be configured to perform an operation using, for instance, software, hardware, firmware, or the like. For example, the phrase "configured to" can refer to a logic circuit structure of a hardware element that is to implement the associated functionality. The phrase "configured to" can also refer to a logic circuit structure of a hardware element that is to implement the coding design of associated functionality of firmware or software. The term "module" refers to a structural element that can be implemented using any suitable hardware (e.g., a processor, among others) , software (e.g., an application, among others) , firmware, or any combination of hardware, software, and firmware. The term "logic" encompasses any functionality for performing a task. For instance, each operation illustrated in the flowcharts corresponds to logic for performing that operation. An operation can be performed using, software, hardware, firmware, or the like. The terms, "component" , "system" , and the like may refer to computer-related entities, hardware, and software in execution, firmware, or combination thereof. A component may be a process running on a processor, an object, an executable, a program, a function, a subroutine, a computer, or a combination of software and hardware. The term, "processor" may refer to a hardware component, such as a processing unit of a computer system.

[0024] The terms "a" or "an" as used herein, are defined as one or more than one. Also, the use of introductory phrases such as "at least one" and "one or more" in the claims should not be construed to imply that the introduction of another claim element by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim element to disclosures containing only one such element, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an" . The same holds true for the use of definite articles.

[0025] Furthermore, the claimed subject matter may be implemented as a method, apparatus, or article of manufacture using standard programming and engineering techniques to produce software, firmware, hardware, or any combination thereof to control a computing device to implement the disclosed subject matter. Computer-readable storage media can include, but are not limited to, magnetic storage devices, e.g., hard disk, floppy disk, magnetic strips, optical disk, compact disk (CD) , digital versatile disk (DVD) , smart cards, flash memory devices, among others. In contrast, computer-readable media, i.e., not storage media, may additionally include communication media such as transmission media for wireless signals and the like.

[0026] A robot teach pendant is an electronic device used to interact with an industrial robot. It may be equipped with a touch screen, buttons, and / or joystick for programming and controlling one or more robots. Through the robot teach pendant, an operator can manually guide the robot to complete specific actions, record the path and position information, and save it as a reproducible program. The teach pendant may also provide real-time monitoring, parameter adjustment, and fault diagnosis functions. This can make the operation more intuitive and flexible and provide a good user experience.

[0027] A robot can interact with an electronic device which can provide interaction between a user and the robot. For example, the electronic device may be a robot teach pendant. Operating a robot teach pendant typically demands specialized training or substantial experience. This is great challenges for new users. Critical functions and essential information are often nested within multi-layered directories, necessitating complex navigation through multiple operational steps, which consequently diminishes work efficiency. Furthermore, the evolution of user interfaces during product upgrades from older to newer versions frequently creates adaptation difficulties for experienced users, potentially disrupting their workflow and productivity. There is a need for providing an interaction with an electronic device, where the interaction mode user friendly.

[0028] The present disclosure proposed a solution for providing an interaction with an electronic device. By implementing the proposed solution, it can provide an easy and user friendly user interface with the electronic device. For example, it is more efficient for a beginner because the beginner can operate the electronic device without learning a manual of how to use the electronic device.

[0029] Reference is made to FIG. 1, which illustrates a schematic diagram of an example environment 100 in which a plurality of embodiments of the present disclosure can be implemented. The example environment 100 is only illustrated and is not intended to suggest any limitations as to scope of use or functionality of embodiments of the disclosure described herein.

[0030] As shown in FIG. 1, the example environment 100 comprises an electronic device 104. An example of the computing device 104 may be a robot teach pendant. Other examples of the computing device 104 may be a tablet, a smart phone, a computer or other terminal device.

[0031] The electronic device 104 may incorporate a language model, for example, a large language model (LLM) 108. The LLM 108 may be a multimodal generative model and / or a sophisticated natural language processing (NLP) model capable of handling diverse multimodal inputs. Its functionality is to enable computational systems to understand, analyze, and produce human-like language content. The LLM 108 may integrate a comprehensive set of computer algorithms specifically engineered to process and interpret vast quantities of natural language data. Furthermore, the LLM 108 may receive user inputs (e.g., user input 102) , accurately interpret user requirements, and generate an appropriate skill list 110 in response to the analyzed user input 102.

[0032] The user input 102 may encompass various formats, including textual paragraphs, audio recordings, speech segments, or visual images. The LLM 108 may be designed to accommodate direct input of these diverse natural language carriers, supporting text, speech, and image data either individually or in combination. When the user input 102 consists of text, it undergoes immediate semantic analysis without requiring additional preprocessing. For speech-based input, the LLM 108 may utilize automatic speech recognition (ASR) technology to transcribe the audio content into textual format for subsequent semantic processing. In cases where the user input 102 is an image, advanced image processing techniques may be employed to extract and convert visual information into textual data suitable for semantic analysis. This multimodal capability can enable the LLM 108 to effectively process and integrate multiple input formats simultaneously, enhancing its versatility in natural language understanding tasks.

[0033] The LLM 108 may be a statistical language model which is built on deep neural networks, with huge parameter scale and strong context-awareness. The LLM 108 may use a large amount of unlabeled text for training through self-supervised learning methods. The LLM 108 may understand and generate human language, and handle a variety of natural language tasks such as text classification, question and answer, dialogue, article writing, code generation, etc. The LLM 108 may have multi-language and multi-modality support and may be applied to cross-cultural and cross-media scenarios. For example, the LLM 108 may receive the user input 102 and understand what the user would like to control or configure for the robot. The LLM 108 may then output the skill list 110 for controlling or configuring for the robot.

[0034] The skill list 110 in the electronic device 104 (e.g., the robot teach pendant) may be a collection of predefined or user-customized tasks or functions that a robot can perform. It may include basic motion skills (e.g., point-to-point movement, linear motion, circular motion) , process skills (e.g., welding, painting, assembly, grinding) , interaction skills (e.g., communication with external devices, force control, vision guidance) , and advanced functions (e.g., path optimization, lead-through, adaptive control, multi-robot collaboration) .

[0035] After the LLM 108 determines the skill list 110, the electronic device 104 may perform the skill list 110. For example, the electronic device 104 may determine what functions can be called from a skill library to perform the skill list 110. The electronic device 104 may call the needed one or more function (s) from the skill library. The electronic device 104 may execute the called one or more function (s) . The result of executing the function (s) is the one or more operations 106. The electronic device 104 may present the one or more operations 106 to the user.

[0036] For example, the user input 102 may be a speech of a user. The speech is: please let us know what is the IP address of the robot. The LLM 108 may receive the user input 102 and output the skill list 110. The skill list 110 may comprise at least a function for retrieving the IP address of the robot, and a function for switching to a page of displaying the IP address of the robot. The electronic device 104 may execute the function for retrieving the IP address of the robot to obtain the IP address of the robot. The electronic device 104 then may execute the function for switching to a page of displaying the IP address of the robot. the result of switching to a page of displaying the IP address of the robot may be the operation 106. The user may see the IP address of the robot on the page of displaying the IP address of the robot.

[0037] Reference is made to FIG. 2, which illustrates a schematic diagram of an example environment 200 in which a plurality of embodiments of the present disclosure can be implemented. The elements in FIG. 2 are similar to the elements in FIG. 1. For example, the user input 202 in FIG. 2 may correspond to the user input 102 in FIG. 1. The electronic device 204 in FIG. 2 may correspond to the electronic device 104 in FIG. 1. The skill list 210 in FIG. 2 may correspond to the skill list 110 in FIG. 1. The operation 206 in FIG. 2 may correspond to the operation 106 in FIG. 1.

[0038] The only difference is that the LLM 208 is deployed outside the electronic device 204. The LLM 208 may be deployed on a cloud server, edge computing devices or a personal computer and so on, and the specific choice of the computing devices depends on the model scale and application requirements. For some examples, cloud servers and high-performance computing clusters are suitable for processing large-scale models and high-concurrency requests. Personal computers are convenient for local development and testing. Edge computing devices and embedded systems provide low latency and data privacy protection. Mobile devices are suitable for lightweight models and are easy to carry. Dedicated artificial intelligence (AI) accelerators (such as GPUs and TPUs) provide high performance and low power consumption, and are suitable for deep learning tasks.

[0039] The user scenarios of FIG. 2 is also similar to FIG. 1. For example, the user input 202 may be a speech of a user. The speech is: please control the robot in the lead-through mode. The LLM 208 may receive the user input 202 and output the skill list 210 as well. The skill list 210 may at least comprise a function for checking if the robot meets the requirement of the lead-through mode, a function for configuring the robot with the lead-through mode, and a function for switching to a page of displaying user interactions for the paths to lead the robot. Each of these functions may comprises a lot of sub-functions. For the purpose of simplification, they will not be discussed herein.

[0040] The electronic device 204 may execute these functions and complete the configuration for the robot with the lead-through mode. The electronic device 204 may present a user interface for inputting the path. The user may input the path via the user interface. The electronic device 204 may receive the path from user via the user interface. The electronic device 204 may control the robot to follow the path.

[0041] Reference is made to FIG. 3, which illustrates an example scenario 300 of providing an interaction with an electronic device in accordance with some embodiments of the present disclosure. FIG. 3 is a logic diagram. This means that although the LLM platform 304 is shown separate to the robot teach pendant 306, yet the LLM platform 304 can be inside or outside the robot teach pendant 306.

[0042] The user voice 302 may be captured from a voice input device. A command parsing module may be in the LLM platform 304. User commands may be captured via the voice input device and subsequently transmitted to the command parsing module in the LLM platform 304. The LLM platform 304, which may be implemented as a cloud server with command parsing capabilities, a local computer, or an edge computing device, may process the input commands. The robot teach pendant 306 may then execute the corresponding operations based on the parsed output from the command parsing module.

[0043] As an example, user voice 302 may be input to the LLM platform 304. The LLM platform 304 (may be or may represent the command parsing module) may determine corresponding skills and the robot teach pendant 306 may execute the skills. After executing the skills, the resulted user interactions may be determined can present to the user. For example, the following use cases can be applied without limitations.

[0044] The robot teach pendant 306 may deliver the robot status feedback in response to voice queries, such as providing the robots IP address or identifying the currently active tool center point (TCP) configuration upon receiving user commands like "what is the robot's IP address? " or "which TCP is currently in use? " , with responses dynamically generated based on the robot's operational status.

[0045] The robot teach pendant 306 may enable voice-activated interface navigation by automatically switching screen content in response to verbal instructions, such as instantly displaying the 'jog' page upon receiving the command "switch to the 'jog' page" . This can facilitate seamless teach pendant operation through voice-controlled menu transitions.

[0046] The robot teach pendant 306 may implement a voice-guided operational workflow system that generates and displays step-by-step instructions for completing complex tasks. When receiving voice commands such as "create a new TCP, " the robot teach pendant 306 may dynamically present a sequential guide on the interface sidebar, detailing each necessary step for TCP creation. The robot teach pendant 306 may provide contextual prompts and visual indicators for the current operation phase. This can enable users to efficiently complete technical procedures through interactive, voice-initiated guidance.

[0047] The robot teach pendant 306 may enable automated execution of operational tasks through voice command recognition, facilitating seamless control of key robotic functions including TCP switching, operation mode configuration, robot target creation, and lead-through function activation and / or deactivation. This can enhance operational efficiency through voice-driven automation.

[0048] The robot teach pendant 306 may enable voice-activated automation of operational tasks, including TCP switching, operation mode transitions, robot target generation, and lead-through function management. This can allow for hands-free control and enhanced operational efficiency through intelligent voice command execution.

[0049] When encountering ambiguous voice commands, the robot teach pendant 306 may intelligently identify and display potential matching functions to prompt the user to clarify their intent through a selection interface. This can ensure an accurate command interpretation and execution.

[0050] The voice command interface may support multiple activation methods. For example, keyword wake-up triggers, voiceprint authentication, or manual button activation. This can provide flexible and secure access to voice control functionality.

[0051] Reference is made to FIG. 4, which illustrates an example process 400 of providing an interaction with an electronic device in accordance with some embodiments of the present disclosure. When the command parsing function is activated, the voice command comprising the instruction 402 can be captured and transmitted to the LLM 404 in either voice or text format.

[0052] Th LLM 404 may be developed through fine-tuning of existing pre-trained models or by implementing customized prompt engineering techniques, for example, using custom prompts. The robot teach pendant 410 may generate an optimized skill execution sequence for the user's command by integrating the robot teach pendant knowledge 412. Knowledge 412 can be incorporated through retrieval-augmented generation (RAG) technology or utilized as training data for LLM refinement.

[0053] The knowledge 412 may include basic knowledge of robots, user manuals for robot teach pendants, explanations of the skill library that can communicate with and control the teach pendant, and descriptions of all skills. The knowledge 412 may include other information or knowledge of robots or robot teach pendant. Based on this knowledge, the LLM 404 may provide a more reasonable skill list 406 and a more reasonable skill execution sequence 408. The skill execution module may call the skills from the teach pendant control skill library 414 according to the skill execution sequence 408 and may execute them. The skill library 414 may include all skills related to communicating with and controlling the teach pendant, such as obtaining robot information, page switching, and state switching functions. The execution of skill list 406 may be present to the user by the robot teach pendant 410.

[0054] Reference is made to FIG. 5, which a diagram of an example retrieval-augmented generation (RAG) 500 in accordance with some embodiments of the present disclosure. Regarding RAG 500 in the field of robot teach pendants, its architecture may comprise four parts: user input 502, retriever 504, generator 506 and result 508. The user input 502 may comprises one or more of multi-modality inputs comprising: voice 510, text 512, image 514, video 516, and robot knowledge 518 in any form, and other information 520.

[0055] In the data preparation and index construction phase, the RAG 500 may collect, organize and index information such as operation manuals, programming language documents, troubleshooting guides related to robot teach pendants. The RAG 500 may build a comprehensive knowledge base, and the RAG 500 may clean, organize and encode the above data 522 to build an efficient index so that the retriever can quickly retrieve relevant information and send the encoded data 522 to the retrieval module 504.

[0056] Then, the retrieval module 504 may receive the user's input query, preprocess and analyze the user's input query. The retrieval module 504 may use vector retrieval technology to retrieve the most relevant documents or paragraphs from the knowledge base to the user's query, and may reorder the retrieved documents to ensure that the most relevant documents are ranked first. For example, data 522 may be embedded into vectors, and the retrieval module 504 may retrieve the needed relevant information by the sparse retrieval approach 524, the dense retrieval approach 526 and / or other retrieval approach 528. The retrieval module 504 may send the retrieve information to the generator 506.

[0057] Then, the generator 506 may combine the retrieved relevant information with the user's original query to form an enhanced input, and uses an LLM to generate the final answer based on the enhanced input to ensure that the answer is both accurate and meets the user's expectations. The LLM may comprise one or more of a transformer 530, long short-term memory (LSTM) 532, diffusion 534, generative adversarial network (GAN) 536 and / or other architecture 536. The generator 506 may output the final answer to the user query as the result 508.

[0058] When training (for example, refine-tuning) the RAG 500, a feedback and optimization process can be used. It is responsible for collecting user feedback on the generated answers, evaluating the quality and relevance of the answers, and optimizing and adjusting the retrieval module and generation module based on user feedback and performance in actual applications to continuously improve the performance of the system. Through this architecture, the RAG 500 can improve the performance of the robot teaching pendant in processing complex queries and generating tasks, and can provide more accurate, timely and useful answers and / or user interactions between the user and the robot teaching pendant.

[0059] In some example, embodiments, the design of prompt for LLM may be used independently or used in combination with RAG. The prompts may explicitly include robotics field terms (such as TCP, joint angles, path planning) , clarifying the task type (code generation, fault diagnosis, parameter optimization) , and providing structured input (such as robot model, controller type, sensor configuration) and constraints (such as speed limit, safety specifications) . Through multi-modality input integration and error handling guidance, a good quality of the output (such as Script codes, operation steps) can be generated. It can lower the programming threshold and improve task execution efficiency while meeting the strict specifications of industrial scenarios.

[0060] In some example, embodiments, the preparation of prompts for LLM may be needed and may be combined with the RAG architecture to clarify the task objectives and user needs, such as operation guidance, troubleshooting or programming support, and design targeted prompts accordingly. The prompts may be extracted form user inputs and relevant information retrieved from the robot knowledge base, such as operation steps, troubleshooting solutions, programming examples, etc. The prompts may be provided as context to the LLM.

[0061] The prompts may clearly tell the LLM how to use the context to answer questions. Few-sample learning can also be used to provide simple question-answer pairs to help the LLM understand the task. In addition, the output format may be specified according to needs, such as step lists, tables, etc. As the robot knowledge base is updated and user feedback accumulates, the prompts may be dynamically updated and optimized. Through these methods, prompts can be effectively prepared to improve the application effect of LLM in the field of robot teach pendants.

[0062] Reference is made to FIG. 6, which illustrates a flowchart of an example method 600 for providing an interaction with an electronic device in accordance with some embodiments of the present disclosure. FIG. 6 can be implemented in the environment of FIG. 1 and / or FIG. 2.

[0063] At 602, the electronic device receives, from a user, a user input comprising one or more instructions for controlling the robot. The user input may comprise a user voice representing the one or more instructions for controlling the robot; an image representing the one or more instructions for controlling the robot; and / or a text representing the one or more instructions for controlling the robot.

[0064] In some example embodiments, the one or more instructions for controlling the robot may comprise creating a TCP for the robot; switching an IP address for the robot; changing an operation mode of the robot; enabling a function of leading the robot; disabling a function of leading the robot; obtaining information associated with the robot; and / or switching to a second interface from a first interface which is currently presented.

[0065] In some example embodiments, if the user input is the user voice, a voice recognition function of the language model may be activated by a keyword wake-up; a voiceprint recognition; or a physical button press.

[0066] At 604, the electronic device determines, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on the electronic device and associated with the controlling of the robot. At 606, the electronic device presents, to the user, at least one operation to be performed on the electronic device.

[0067] In some example embodiments, the language model may be deployed on at least one of a local server, a cloud server, an edge computing device or the electronic device. In some example embodiments, the language model may comprise a multi-modality generative model. In some example embodiments, the electronic device may comprise a robot teach pendant.

[0068] In some example embodiments, the robot knowledge may comprise basic knowledge of at least one robot; at least one user manual for at least one robot teach pendant; at least one description of at least one skill of at least one robot teach pendant; and / or at least one explanation of a skill library. In some example embodiments, the skill library may comprise at least one skill related to at least one of communicating with the at least one robot teach pendant or controlling the at least one robot teach pendant.

[0069] By implementing the embodiments of the method 600, it can enhance language-based human-computer interaction capabilities. By applying LLM technology to an electronic device such as the robot teach pendant, usability and efficiency can be improved. It can provide an easy and user friendly user interface with the electronic device. For example, it is more efficient for a beginner because the beginner can operate the electronic device without learning a manual of how to use the electronic device.

[0070] Reference is made to FIG. 7, which illustrates a block diagram of an example apparatus 700 for providing an interaction with an electronic device in accordance with some embodiments of the present disclosure. The apparatus 700 comprises a receiving module 702 configured to receive, from a user, a user input comprising one or more instructions for controlling the robot. The apparatus 700 further comprises a determining module 704 configured to determine, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on the electronic device and associated with the controlling of the robot. The apparatus 700 further comprises a presenting module 706 configured to present, to the user, at least one operation to be performed on the electronic device.

[0071] In some example embodiments, if the at least one operation to be performed comprises two or more operations to be performed, the apparatus 700 may further comprise a second determining module configured to determine, by the language model, a performing sequence of the two or more operations to be performed. Further, the presenting module 706 may be configured to present, to the user, the two or more operations to be performed based on the performing sequence.

[0072] In some example embodiments, the apparatus 700 may further comprise a first generating module configured to generate, by the language model, the information associated with the robot; and a second presenting module configured to present, to the user, the information associated with the robot.

[0073] In some example embodiments, the apparatus 700 may further comprise a switching module configured to switch from first interface to the second interface if the one or more instructions for controlling the robot comprises switching to the second interface from the first interface which is currently presented to the user.

[0074] In some example embodiments, the determining module 704 may further be configured to determine a list of operations associated with the controlling of the robot. In some example embodiments, the determining module 704 may further comprises a third presenting module configured to present, to the user, a list of functions associated with the one or more instructions; a second receiving module configured to receive, from the user, a function selected by the user from the list of functions; and a third determining module configured to determine the list of operations based on the selected function and the one or more instructions.

[0075] In some example embodiments, the presenting module 706 may further comprise a fourth presenting module configured to present, to the user, a first operation in the list of operations; a third receiving module configured to receive, from the user, a user response to the first operation; a performing module configured to perform the user response for the robot to determine a second operation in response to the user response to the first operation; and a fifth presenting module configured to present, to the user, the second operation.

[0076] In some example embodiments, the determining module 704 may further comprise an extracting module configured to extract, by the language model, one or more prompts based on the one or more instructions, wherein the one or more prompts are pre-determined by fine-tuning the language model based on a skill library comprising skills of the robot teach pendant; a fourth determining module configured to determine, by the language model, a list of skills to be called from the skill library based on the one or more prompts and the robot knowledge; and a fifth determining configured to determine, by the language model, the at least one operation to be performed on the electronic device based on performing the list of skills.

[0077] In some example embodiments, the apparatus 700 may further comprise a fourth receiving module configured to receive, from the user, the interaction in response to the at least one operation to be performed on the electronic device; and a controlling module configured to control the robot based on the interaction.

[0078] By implementing the example embodiments of FIG. 7, the integration of LLM technology significantly can enhance natural language-based human-computer interaction capabilities. When implemented in electronic devices such as robot teach pendants, it can improve usability and operational efficiency. This can provide an intuitive and user-friendly interface. New users can benefit from enabling device operation without requiring extensive manual study or prior technical knowledge.

[0079] FIG. 8 illustrates a block diagram illustrating an electronic device 800 in accordance with some embodiments of the present disclosure. As indicated, the device 800 includes a central processing unit (CPU) 801, which can execute various appropriate actions and processing based on the computer program instructions stored in a read-only memory (ROM) 802 or the computer program instructions loaded into a random-access memory (RAM) 803 from a storage unit 808. The RAM 803 also stores all kinds of programs and data required by operating the electronic device 800. CPU 801, ROM 802 and RAM 803 are connected to each other via a bus 804, to which an input / output (I / O) interface 805 is also connected.

[0080] A plurality of components in the device 800 are connected to the I / O interface 805, comprising: an input unit 806, such as a keyboard, a mouse and the like; an output unit 807, such as various types of displays, loudspeakers and the like; a storage unit 808, such as a storage disk, an optical disk and the like; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver and the like. The communication unit 809 allows the device 800 to exchange information / data with other devices through computer networks such as Internet and / or various telecommunication networks.

[0081] Each procedure and processing described above, such as the method 600, can be executed by a processing unit 801. For example, in some embodiments, the method 800 can be implemented as computer software programs, which are tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, the computer program can be partially or completely loaded and / or installed to the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded to the RAM 803 and executed by the CPU 801, one or more steps of the above described method 600 are implemented. Alternatively, in other embodiments, the CPU 801 may also be configured in any proper manner to implement the above process / method.

[0082] The present disclosure may be a method, a device, a system and / or a computer program product. The computer program product can include a computer-readable storage medium loaded with computer-readable program instructions thereon for executing various aspects of the present disclosure.

[0083] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (anon-exhaustive list) of the computer readable storage medium would include: a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , a static random access memory (SRAM) , a portable compact disc read-only memory (CD-ROM) , a digital versatile disk (DVD) , a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination thereof. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable) , or electrical signals transmitted through a wire.

[0084] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium, or downloaded to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0085] Computer readable program instructions for carrying out operations of the present disclosure may be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN) , or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider) . In some embodiments, by means of state information of the computer readable program instructions, an electronic circuitry including, for example, programmable logic circuitry (PLC) , field-programmable gate arrays (FPGA) , or programmable logic arrays (PLA) can be personalized to execute the computer readable program instructions, thereby implementing various aspects of the present disclosure.

[0086] Aspects of the present disclosure are described herein with reference to flowchart and / or block diagrams of methods, apparatus (systems) , and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0087] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which are executed via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0088] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which are executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0089] The flowchart and block diagrams illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, snippet, or portion of codes, which comprises one or more executable instructions for implementing the specified logical function (s) . In some alternative implementations, the functions noted in the block may be implemented in an order different from those illustrated in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.

[0090] Further, while operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0091] A person of ordinary skill in the art may be aware that, in combination with the examples described in the embodiments disclosed in this specification, units and algorithm steps can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed by hardware or software depends on particular applications and design constraints of the technical solutions. A person skilled in the art may use different methods to implement the described functions for each particular application, but it should not be considered that the implementation goes beyond the scope of this application.

[0092] It may be clearly understood by a person skilled in the art that, for the purpose of convenient and brief description, for a detailed working process of the foregoing system, apparatus, and unit, refer to a corresponding process in the foregoing method embodiment. Details are not described herein again.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiment is merely an example. For example, the unit division is merely logical function division and may be other division in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.

[0094] The units described as separate parts may be or may not be physically separate, and parts displayed as units may be or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of the embodiments.

[0095] In addition, functional units in the embodiments of this application may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

[0096] When the functions are implemented in a form of a software functional unit and sold or used as an independent product, the functions may be stored in a computer readable storage medium. Based on such an understanding, the technical solutions in this application essentially, or the part contributing to the prior art, or some of the technical solutions may be implemented in a form of a software product. The computer software product is stored in a storage medium, and includes several instructions for instructing a computer device (which may be a personal computer, a server, a network device, or the like) to perform all or some of the steps of the methods described in the embodiments of this application. The foregoing storage medium includes: any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (Read-Only Memory, ROM) , a random access memory (Random Access Memory, RAM) , a magnetic disk, or an optical disc.

[0097] The foregoing descriptions are merely specific implementations of this application, but are not intended to limit the protection scope of this application. Any variation or replacement readily figured out by a person skilled in the art within the technical scope disclosed in this application shall fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1.A method for providing an interaction with an electronic device associated with a robot, comprising:receiving, from a user, a user input comprising one or more instructions for controlling the robot;determining, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on the electronic device and associated with the controlling of the robot; andpresenting, to the user, at least one operation to be performed on the electronic device.2.The method of claim 1, further comprising:in response to the at least one operation to be performed comprising two or more operations to be performed, determining, by the language model, a performing sequence of the two or more operations to be performed,and wherein presenting, to the user, the at least one operation to be performed on the electronic device comprises:presenting, to the user, the two or more operations to be performed based on the performing sequence.3.The method of claim 1, wherein receiving the user input comprising the one or more instructions for controlling the robot comprises at least one of the following:a user voice representing the one or more instructions for controlling the robot;an image representing the one or more instructions for controlling the robot; ora text representing the one or more instructions for controlling the robot.4.The method of claim 1, wherein in response to the user input being the user voice, a voice recognition function of the language model is activated by at least one of the following:a keyword wake-up;a voiceprint recognition; ora physical button press.5.The method of claim 1, wherein the one or more instructions for controlling the robot comprises at least one of the following:creating a tool-center point (TCP) for the robot;switching an Internet protocol (IP) address for the robot;changing an operation mode of the robot;creating a target for the robot;enabling a function of leading the robot;disabling a function of leading the robot;obtaining information associated with the robot; orswitching to a second interface from a first interface which is currently presented.6.The method of claim 5, further comprising:in response to the one or more instructions for controlling the robot comprising obtaining the information associated with the robot, generating, by the language model, the information associated with the robot; and presenting, to the user, the information associated with the robot; andin response to the one or more instructions for controlling the robot comprising switching to the second interface from the first interface which is currently presented to the user, switching from first interface to the second interface.7.The method of claim 1, wherein determining the least one operation comprises determining a list of operations associated with the controlling of the robot, and wherein determining the list of operations associated with the controlling of the robot comprising:in response to the one or more instructions for controlling the robot being determined by the language model being ambiguous, presenting, to the user, a list of functions associated with the one or more instructions;receiving, from the user, a function selected by the user from the list of functions; anddetermining the list of operations based on the selected function and the one or more instructions.8.The method of claim 7, wherein presenting the two or more operations to be performed based on the performing sequence comprises:presenting, to the user, a first operation in the list of operations;receiving, from the user, a user response to the first operation;performing the user response for the robot to determine a second operation in response to the user response to the first operation; andpresenting, to the user, the second operation.9.The method of claim 8, wherein at least one of the following:the language model is deployed on at least one of a local server, a cloud server, an edge computing device or the electronic device;the language model comprises a multi-modality generative model; orthe electronic device comprises a robot teach pendant.10.The method of claim 1, wherein determining, by the language model, the at least one operation associated with the controlling of the robot based on the one or more instructions comprises:extracting, by the language model, one or more prompts based on the one or more instructions, wherein the one or more prompts are pre-determined by fine-tuning the language model based on a skill library comprising skills of the robot teach pendant;determining, by the language model, a list of skills to be called from the skill library based on the one or more prompts and the robot knowledge; anddetermining, by the language model, the at least one operation to be performed on the electronic device based on performing the list of skills.11.The method of claim 1, wherein the robot knowledge comprises at least one of the following:basic knowledge of at least one robot;at least one user manual for at least one robot teach pendant;at least one description of at least one skill of at least one robot teach pendant; orat least one explanation of a skill library,and wherein the skill library comprises at least one skill related to at least one of communicating with the at least one robot teach pendant or controlling the at least one robot teach pendant.12.The method of claim 1, further comprising:receiving, from the user, the interaction in response to the at least one operation to be performed on the electronic device; andcontrolling the robot based on the interaction.13.An apparatus for providing an interaction with an electronic device, comprising:a receiving module configured to receive, from a user, a user input comprising one or more instructions for controlling the robot;a determining module configured to determine, by a language model trained with robot knowledge and based on the user input, at least one operation to be performed on the electronic device and associated with the controlling of the robot; anda presenting module configured to present, to the user, at least one operation to be performed on the electronic device.14.An electronic device, comprising:a processor; anda memory coupled to the processor, wherein the memory has instructions stored therein, and the instructions, when executed by the processor, cause the device to execute operations of any of claims 1-12.15.A computer program product having instructions stored therein, which when executed by a processor, cause the processor to perform a method of any of claims 1-12.