Robot control system, robot control device, and robot control method

The robot control system uses a display/input device and combined language models to accurately interpret user instructions, addressing the inefficiencies of existing systems by enabling precise and intuitive robot operation.

JP2026002459APending Publication Date: 2026-01-08PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024100463
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing robot control systems struggle to accurately interpret intuitive and precise verbal instructions from users, particularly for tasks requiring fine movements, leading to user burden and inefficiency.

Method used

A robot control system that integrates a display/input device for environmental image display and editing, combined with a control device using both large-scale and visual language models to infer user instructions, allowing for accurate robot operations based on natural language and image inputs.

Benefits of technology

Enables intuitive and accurate robot operation by combining large-scale and visual language models to interpret user instructions, reducing the burden on users and improving operational precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026002459000001_ABST
    Figure 2026002459000001_ABST
Patent Text Reader

Abstract

To accurately operate a robot on the basis of an instruction by receiving the intuitive instruction from a user on the operation of the robot.SOLUTION: The robot control system includes a display input device that displays an environmental image including an object to be processed by the robot and receives an instruction input by an editing operation on a natural language and / or the environmental image from a user, and a control device that infers, based on the instruction input, at least one of a predetermined object to be processed by the robot by an operation of the user and a predetermined operation of the robot for performing processing on the predetermined object, and outputs an inference result to the display input device.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a robot control system, a robot control device, and a robot control method. [Background technology]

[0002] Conventionally, systems that use working robots (hereinafter simply referred to as "robots") to perform various tasks in place of humans have become widespread. In recent years, there has been a high demand for robot control systems that operate robots in response to verbal instructions from users. This type of robot control system must accurately respond to the user's verbal instructions and operate the robot. For example, Patent Document 1 discloses a conversational artificial intelligence (AI) assistant as a method for accurately responding to the user's verbal instructions. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication No. 2023-525173 Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure has been devised in view of the above-described conventional situation, and aims to accept intuitive instructions from a user regarding the operation of a robot and to make the robot operate accurately based on the instructions. [Means for solving the problem]

[0005] The present disclosure provides a robot control system comprising: a display input device that displays an environmental image including an object to be processed by a robot and accepts instruction input from a user using natural language and / or editing operations on the environmental image; and a control device that, based on the instruction input, infers a predetermined object to be processed by the robot through operation by the user and at least one predetermined action of the robot to perform the processing on the predetermined object, and outputs the inference result to the display input device.

[0006] The present disclosure also provides a robot control device comprising: an input unit that displays an environmental image including an object to be processed by a robot, and receives instruction input from a display input device that receives instruction input from a user using natural language and / or editing operations on the environmental image; and a control unit that, based on the instruction input, infers at least one of a predetermined object to be processed by the robot through operation by the user and a predetermined action of the robot for performing the processing on the predetermined object, and outputs the inference result to the display input device.

[0007] The present disclosure also provides a robot control method that displays an environmental image including an object to be processed by a robot, and receives instruction input from a display input device that receives instruction input from a user using natural language and / or editing operations on the environmental image, and based on the instruction input, infers at least one of a predetermined object to be processed by the robot through operation by the user and a predetermined action of the robot for performing the processing on the predetermined object, and outputs the inference result to the display input device.

[0008] Any combination of the above components, and conversion of the expression of the present disclosure into a method, device, system, storage medium, computer program, etc., are also valid aspects of the present disclosure. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to accept intuitive instructions from a user regarding the operation of a robot and to make the robot operate accurately based on the instructions. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram illustrating an example of a system configuration according to an embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a display input device according to an embodiment of the present disclosure; [Figure 3] 1 is a flowchart of a control process of a system according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram illustrating a first operation example of a system according to an embodiment of the present disclosure. [Figure 5] FIG. 10 is a schematic diagram illustrating a second operation example of the system according to the embodiment of the present disclosure. [Figure 6] FIG. 10 is a schematic diagram illustrating a third operation example of the system according to the embodiment of the present disclosure. [Figure 7] FIG. 10 is a schematic diagram for explaining a fourth operation example of the system according to the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] (Background to this disclosure) Conventionally, robots that perform various tasks on behalf of humans have been known, including robots with a work tool at the tip that can grasp and transport objects using the work tool. For example, examples of objects that can be grasped include packages to be delivered or sorted at logistics sites. In recent years, there has been a trend toward robots of this type being required to operate in response to instructions, such as linguistic instructions in natural language, from users. However, it is difficult to flexibly understand the intent of a user's instructions, making it difficult to intuitively and accurately instruct the robot's operations. Patent Document 1 discloses a technology in which an AI assistant responds to a user's voice, video, and / or text input. Patent Document 1 discloses an example in which the AI ​​assistant's operation is determined using gestures, etc. acquired from video and the analysis results of the voice or text input. However, since Patent Document 1 is a technology primarily intended for dialogue with users, it is not suitable for situations requiring precise instructions, such as instructions to a robot. Patent Document 1 acquires a user's gestures from video and analyzes the user's natural language linguistic instructions from the voice or text input. Therefore, for example, when instructing a robot to move an item from one compartment to another on a tray measuring several tens of centimeters square, the user must either perform extremely fine gestures with precision down to a few centimeters, or give a lengthy instruction such as, "Please move item A stored in the top left compartment to the second compartment from the right and the third compartment from the top." Such actions place a heavy burden on the user, making them unsuitable for frequently issuing instructions to a robot.

[0012] Hereinafter, with reference to the drawings as appropriate, detailed descriptions of each embodiment specifically disclosing the system according to the present disclosure will be provided. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters or redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter recited in the claims.

[0013] <First Embodiment> [System Configuration] FIG. 1 is a block diagram showing an overall overview of a robot control system 1 according to this embodiment. Note that the system configuration shown in FIG. 1 is an example, and one device may be divided into multiple parts, or multiple devices may be integrated into one. Also, multiple devices may be provided that are identical to the one shown in FIG. 1. Furthermore, the processing entities shown below are examples, and part of the function of one device may be realized as the function of another device. Also, FIG. 2 is a schematic diagram for explaining a display input device according to this embodiment.

[0014] The robot control system 1 includes a robot 100, a control device 200, and a display / input device 300. The control device 200 and the display / input device 300 may be physically configured as a robot control device composed of one or more devices. In this example, the robot 100 is a picking robot equipped with a robot arm, a gripper provided at the tip of the robot arm as a work tool, and the gripper grips a package 400, which is an object to be grasped. That is, in this example, the action of the robot 100 grasping an object corresponds to "processing by the robot." Note that, hereinafter, an object grasped by the robot 100 is also referred to as a "grasped object." In addition, although the robot control system 1 in this example includes the robot 100, when the control device 200 and the display / input device 300 of this example are introduced into an environment where the robot 100 is already installed, a system not including the robot 100 may be referred to as the robot control system 1.

[0015] Here, the object to be grasped corresponds to the baggage 400 (401 to 406) shown in Fig. 2 etc. The baggage 400 as the object to be grasped includes a wide variety of baggage in terms of size, shape, weight, contents (contents), internal arrangement, etc.

[0016] The control device 200 functions as a control device for the robot 100. The control device 200 has an inference unit 210 and an operation control unit 220. The inference unit 210 and the operation control unit 220 are physically one or more processors, and may be configured using, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), a GPU (Graphical Processing Unit), or an FPGA (Field Programmable Gate Array).

[0017] The inference unit 210 performs inference using language and images. For example, the inference unit 210 performs inference using language and images using a large-scale language model (LLM) and a vision-language model (VLM), and creates an inference result using at least one of language and image to be output to the display / input device 300 as a response to the user.

[0018] A large-scale language model is a model built by machine learning the associations of words, etc. contained in a large amount of natural language text. A large-scale language model can construct natural language text from words, etc. related to questions, etc. received from a user, and infer answers to the questions. In this case, a large-scale language model can infer answers that reflect a context that reflects a history of questions and answers, thereby enabling inference specialized for past interactions with the user. On the other hand, a large-scale language model may output a standard answer to a question that is not included in the context. A visual language model is a model built by machine learning a large amount of natural language text and images. Because a visual language model learns the associations between words, etc. contained in text and images, it can output the content contained in an image as text, answer text questions about the content of the image, and generate an image corresponding to the text. On the other hand, a visual language model is a model that outputs corresponding to an input image or an image and text, and therefore cannot provide answers to events that are not contained in the input image, even if they are included in past interactions with the user.

[0019] To accurately instruct a robot, it is desirable to consider both contextual information and information contained in the image. For example, consider a case where a user gives an instruction such as "Please move the previous item to that location" while marking the destination on the image. In this case, the identity of the "previous item" is inferred from past interactions with the user, while the "location" is inferred from the image. Conversely, a visual language model that cannot use context cannot infer the identity of the "previous item." Furthermore, a large-scale language model cannot identify the location of the "location" unless the location has been specified in past interactions.

[0020] As in the above example, instructions to a robot often include instructions to move or process items, so it is necessary to handle both information that has continuity with past interactions, such as which item is to be processed, and information that has little continuity with past interactions, such as the destination, etc. Therefore, in this embodiment, accurate instructions to a robot are realized by using a large-scale language model and a visual language model together.

[0021] More specifically, for example, by having the visual language model generate text describing the current environment and including it in the context of the large-scale language model, the large-scale language model can perform inferences that include instructions obtained from images. This allows inferences to be made that include information that could not be obtained in past interactions. Note that the structure for using a visual language model and a large-scale language model together is not limited to this. For example, if a user's input contains words that are difficult for the large-scale language model to infer, it is possible to generate text that queries the visual language model about the meaning of the words and then query the visual language model. The above example describes an example in which a large-scale language model and a visual language model exchange information via text. However, since the large-scale language model and the visual language model are models that have been trained by converting text into a format that is easy for each model to handle, information equivalent to the text to be exchanged may be exchanged using data generated in that format. Note that, because it is difficult for the visual language model to perform inferences that involve time factors, such as the next processing of an item, the inference unit 210 uses the output of the large-scale language model as the inference result. That is, the inference unit 210 adopts the result of inference made by the large-scale language model based on the information obtained by the visual language model as the inference result. Note that there is no limit to the number of times that the information obtained by the visual language model can be provided to the large-scale language model. For example, by performing multiple exchanges between the visual language model and the large-scale language model, the information obtained by the visual language model can be made more specific and included in the context, thereby improving the accuracy of the inference result made by the large-scale language model.

[0022] In this way, based on an instruction (instruction input) from the user, the inference unit 210 infers the predetermined grasp object (in this example, one of the luggage items 401 to 406) that the user has instructed the robot 100 to acquire and grasp from the grasp object by the user's operation, and the predetermined action of the robot 100 for acquiring and grasping the predetermined grasp object, and generates an inference result. Note that the inference unit 210 may use a language model other than the large-scale language model and the visual language model. Also, the inference result may be output as text in a natural language or as an image. When outputting as text, the inference result is input to the large-scale language model to generate the text to be output. When outputting as an image, the inference result is input to the visual language model to generate the image to be output. Also, the inference result may be output as both text and an image. In this case, the inference result is input to both the large-scale language model and the visual language model to generate the text and the image, respectively.

[0023] The movement control unit 220 generates a control signal for the robot 100 based on the inference result. The movement control unit 220 is also capable of transmitting a control signal related to the predetermined movement to the robot 100, and determines whether the predetermined movement is possible before transmitting the control signal related to the predetermined movement to the robot 100. For example, when a linguistic instruction related to the predetermined movement is input from the inference unit 210, the movement control unit 220 generates a control signal related to the predetermined movement to the robot 100, and determines whether the predetermined movement is possible before transmitting the control signal related to the predetermined movement to the robot 100.

[0024] Although not shown in FIG. 1, the control device 200 includes, for example, a memory, an input / output device, a robot IF, a display input device IF, and a communication device, and each part is connected to enable communication.

[0025] The memory is a storage area for storing and holding various data, and may be composed of, for example, a non-volatile storage area such as a ROM (Read Only Memory) or an HDD (Hard Disk Drive), or a volatile storage area such as a RAM (Random Access Memory). For example, a processor such as the inference unit 210 and the operation control unit 220 reads and executes various data (for example, a conversation log, etc., which will be described later) and programs stored in the memory, thereby realizing various functions, which will be described later.

[0026] The input / output device receives instructions from a user via, for example, a mouse or keyboard (not shown) and outputs various types of information data via a display (not shown). The robot IF is an interface for connecting to the robot 100 and transmits and receives various control signals to the robot 100 based on, for example, instructions from the operation control unit 220. The display / input device IF is an interface for connecting to the display / input device 300 and transmits and receives various control signals to the display / input device 300 based on, for example, instructions from the inference unit 210. The communication device communicates with an external device (not shown) via a wired / wireless network and transmits and receives various types of data and signals. The communication method used by the communication device is not particularly limited and may be compatible with multiple communication methods. For example, a wide area network (WAN), a local area network (LAN), power line communication, short-range wireless communication (e.g., Bluetooth (registered trademark)), etc. may be used.

[0027] The display input device 300 has a function of displaying the environmental image P, a conversation log, etc., and also of receiving instruction input using language and the environmental image through user operation. The display input device 300 is, for example, a touch panel.

[0028] Here, the environmental image P is an image of the surrounding environment of the gripping unit of the robot 100, including objects to be grasped (packages 401 to 406 in this example) that are arranged around the robot 100, which is the main operator, and can be grasped by the robot 100. In this example, the environmental image P shows an appearance in which packages 401 to 406 are arranged in each compartment of a tray formed into six compartments. The environmental image P is captured, for example, by a camera (not shown) mounted on the gripping unit of the robot 100, and is displayed on the display input device 300. The environmental image P displayed on the display input device 300 is updated as appropriate.

[0029] The display input device 300 has an image editing section 320, a language input section 330, and a conversation log display section 340 on its display screen 310.

[0030] The image editing unit 320 has a function of accepting editing operations by the user on the environmental image P displayed on the display screen 310. This editing operation may be performed using, for example, a touch pen. The image editing unit 320 may also accept a language input operation on the environmental image P displayed on the display screen 310.

[0031] The language input unit 330 has a function of accepting a language input operation by the user. This language input operation may be performed, for example, by inputting characters displayed in keyboard format on the display screen 310, by handwriting on a freeboard displayed on the display screen 310, or by voice input. However, when the language input operation is performed by voice input, the input voice is converted into text.

[0032] The conversation log display unit 340 displays a conversation log including user instruction inputs and responses from the control device 200. This conversation log includes text generated using a large-scale language model in response to the user instruction inputs and the control device 200 responses, and images (language and images) generated using a visual language model, and is displayed in, for example, a chatbot format. This conversation log may be stored in a memory or the like of the control device 200, and inference may be performed using the stored conversation log, for example. Note that the image displayed in the conversation log may be an image obtained by combining an image generated using a visual language model with an environmental image at the time of the conversation log. For example, it may be an image in which a mark or the like generated by the visual language model is superimposed on an environmental image corresponding to the time of the conversation log.

[0033] [Robot control system operation] The operation of the robot control system 1 according to this embodiment will be described below with reference to FIGS.

[0034] 3 is a flowchart for explaining the flow of control processing by the control device 200 in the robot control system 1 according to this embodiment. To simplify the explanation, the processing is collectively described as the control device 200, but the components constituting the control device 200 (particularly the inference unit 210 and the operation control unit 220) work together to realize the following processing.

[0035] At the start of this control process, it is assumed that the luggage to be grasped is placed at a predetermined position, that the luggage is photographed to identify the area of ​​each luggage, and that an environmental image P including the photographed luggage is displayed on the display input device 300 (see Figure 2).

[0036] When a user performs an instruction input operation on the display / input device 300, the control device 200 (e.g., the inference unit 210) acquires the instruction input (input information) from the display / input device 300 (step S501). This instruction input operation may be an operation intuitively performed by the user, such as a linguistic input such as "pick up the square toy," or an linguistic input and image editing such as adding a mark to a package or the like included in the environmental image P and saying "move this to the box on the right." This instruction input is displayed on the conversation log display unit 340 of the display / input device 300 (see the upper part of FIG. 2).

[0037] Based on the acquired instruction input, the control device 200 (e.g., the inference unit 210) infers the predetermined grasp object that the user is instructed to acquire and grasp from the grasp object by the user's operation, i.e., the predetermined baggage (step S502). In step S502, as described above, the large-scale language model and the visual language model are used in combination to perform inference based on the acquired instruction input. For example, the control device 200 uses the visual language model to acquire information contained in the environmental image P and the editing operation on the environmental image P, and uses the large-scale language model to infer the predetermined baggage based on the information acquired by the visual language model and the conversation log with the user including the instruction input and the inference result.

[0038] The control device 200 (e.g., the inference unit 210) determines whether or not the specified baggage has been identified by inference (step S503). If the specified baggage has been identified (step S503: Yes), the processing of the control device 200 proceeds to step S505. On the other hand, if the specified baggage cannot be identified (step S503: No), the processing of the control device 200 proceeds to step S510. Note that cases in which the specified baggage cannot be identified include cases in which the specified baggage cannot be found, cases in which baggage that may be the specified baggage is found but the reliability of the inference result is low, and cases in which the number of specified baggage found is too many or too few in comparison with the user's instructions.

[0039] The control device 200 (e.g., the inference unit 210) infers a predetermined operation of the robot 100 for acquiring and gripping the predetermined load based on the acquired instruction input (step S504). In step S504, for example, the inference is performed in the same manner as in step S502.

[0040] The control device 200 (e.g., the operation control unit 220) determines whether the inferred predetermined operation is possible (step S505). If it is determined that the predetermined operation is possible (step S505: Yes), the processing of the control device 200 proceeds to step S506. On the other hand, if it is determined that the predetermined operation is impossible (step S505: No), the processing of the control device 200 proceeds to step S510.

[0041] The control device 200 (e.g., the inference unit 210) outputs two inference results, each of which is the inferred predetermined baggage and predetermined action, to the display input device 300 (step S507). These two inference results use at least one of text generated using a large-scale language model and images (language and images) generated using a visual language model so that the user can intuitively understand the response from the control device 200, and are displayed on the conversation log display unit 340 of the display input device 300 (see the middle part of Figure 2). At this time, the control device 200 also outputs an instruction request to the display input device 300 to confirm whether or not to accept the two inference results.

[0042] After step S507, when the user performs an instruction input operation for the two inference results on the display input device 300, the control device 200 (e.g., the inference unit 210) acquires the instruction input for the two inference results from the display input device 300 (step S508). The instruction input for the two inference results is displayed on the conversation log display unit 340 of the display input device 300 (see the lower part of FIG. 2).

[0043] The control device 200 (e.g., the inference unit 210) determines whether the two inference results have been accepted based on the instruction input for the two acquired inference results (step S508). That is, in step S509, it is determined whether the two inference results match the user's instruction. If the two inference results have been accepted (step S508: Yes), the processing of the control device 200 proceeds to step S509. On the other hand, if the two inference results have not been accepted (step S508: No), the processing of the control device 200 proceeds to step S510.

[0044] The control device 200 (for example, the operation control unit 220) generates a control signal related to the inferred and approved predetermined operation, and transmits the generated control signal related to the predetermined operation to the robot 100 (step S509). Then, the robot 100 operates based on the control signal. Then, this processing flow ends.

[0045] If the specified baggage cannot be identified (step S503: No), if it is determined that the specified operation is impossible (step S505: No), or if the two inference results are not accepted (step S508: No), the control device 200 (e.g., the inference unit 210) outputs an instruction request to the display input device 300 to prompt the user to provide a necessary instruction (step S510). This instruction request uses at least one of language and images so that the user can intuitively understand the response from the control device 200, and is displayed on the conversation log display unit 340 of the display input device 300 (see the second row of FIG. 5, etc.). Then, the processing of the control device 200 proceeds to step S501.

[0046] (First example of robot control system operation) A first operation example of the robot control system 1 will be described with reference to Fig. 4. Fig. 4 is a schematic diagram for explaining the first operation example of the robot control system 1 according to the present embodiment, and shows the conversation log display unit 340. In this example, a case will be taken as an example in which two types of inference results match the user's instructions.

[0047] First, the user inputs an instruction to the display / input device 300. In this example, the user marks one of the packages 406 (see FIG. 2 for the reference numerals) (corresponding to the "predetermined package") shown in the environmental image P and verbally inputs "please move this to the bottom middle box" (see the upper part of FIG. 4). Then, the control device 200 acquires the instruction input from the display / input device 300 (step S501).

[0048] Next, the control device 200 performs an inference regarding the specified predetermined cargo (step S502). Here, it is inferred which is the specified cargo.

[0049] Next, the control device 200 determines whether the predetermined baggage has been identified by inference (step S503). In this example, since one of the bags 406 (see FIG. 2 for the reference numerals) is marked in the environmental image P included in the instruction input, it is determined that the predetermined baggage has been identified (step S503: Yes).

[0050] Next, the control device 200 infers a predetermined operation of the robot 100 for acquiring and grasping the predetermined load (step S504), where it infers how to move the predetermined load.

[0051] Next, the control device 200 determines whether the inferred predetermined operation is possible (step S505). In this example, since the predetermined item can be moved to the bottom middle box, it is determined that the predetermined operation is possible (step S505: Yes).

[0052] Next, the control device 200 outputs two types of inference results, namely, the inferred predetermined baggage and the predetermined action, to the display / input device 300 (step S506). In this example, the inference result is output using an image in which the predetermined action is represented by an arrow indicating the movement of the predetermined baggage shown in the environmental image P, and a phrase such as "Is this movement correct?" to prompt an instruction from the user (see the middle part of FIG. 4). At this time, the control device 200 may output an instruction request, such as "Do you want to perform the action?", to the display / input device 300 to confirm whether or not to accept the two types of inference results, or the instruction request may be output to the display / input device 300 between step S507 and step S508.

[0053] Next, the user performs an instruction input operation for the two inference results to the display / input device 300. In this example, since the two inference results match the user's instruction, it is assumed that the user verbally inputs "yes" (see the lower part of FIG. 4). Then, the control device 200 acquires an instruction input for the two inference results from the display / input device 300 (step S507).

[0054] Next, the control device 200 determines whether the two inference results have been accepted based on the instruction input for the two acquired inference results (step S508). In this example, it is assumed that the two inference results have been accepted (step S508: Yes).

[0055] Next, the control device 200 generates a control signal related to the approved predetermined action, and transmits the generated control signal related to the predetermined action to the robot 100 (step S509). As a result, the robot 100 operates based on the user's instruction.

[0056] (Second example of robot control system operation) A second operation example of the robot control system 1 will be described with reference to Fig. 5. Fig. 5 is a schematic diagram for explaining the second operation example of the robot control system 1 according to the present embodiment, and shows the conversation log display unit 340. In this example, a case where the specified baggage cannot be identified will be taken as an example.

[0057] First, the user inputs an instruction to the display / input device 300. In this example, it is assumed that the user has not edited the environmental image P and has input the verbal instruction "Please take the plastic bottle" (see the first row in FIG. 5). Then, the control device 200 acquires the instruction input from the display / input device 300 (step S501).

[0058] Next, the control device 200 performs inference on the specified baggage (step S502).

[0059] Next, the control device 200 determines whether or not the predetermined baggage has been identified by inference (step S503). In this example, it is assumed that it has been determined that the predetermined baggage has not been identified (step S503: No).

[0060] Next, the control device 200 outputs an instruction request to the display / input device 300 to prompt the user to give a necessary instruction (step S510). Here, the user is prompted to give an instruction to identify the specified package. In this example, the instruction request is output using language such as "Where is the plastic bottle?" (see the second row of FIG. 5).

[0061] Next, the user performs an instruction input operation in response to the instruction request to the display input device 300. In this example, the user marks the box of the luggage 401 (see FIG. 2 for the symbol) shown in the environmental image P and verbally inputs "This is the transparent bottle inside this box" (see the third row in FIG. 5). Then, the control device 200 acquires the instruction input in response to the instruction request from the display input device 300 (step S501).

[0062] Next, the control device 200 performs inference on the specified predetermined item (step S502), and determines whether the specified item can be identified by inference (step S503). In this example, it is determined that the specified item cannot be identified (step S503: No).

[0063] Next, the control device 200 outputs an instruction request to the display / input device 300 to prompt the user to give a necessary instruction (step S510). In this example, the instruction request is output using language such as "There are multiple options. Which one is it?" (see the fourth row of FIG. 5).

[0064] Next, when the user inputs an instruction in response to the instruction request to the display / input device 300, the control device 200 performs the process of step S501. Then, the control device 200 performs each of the above-mentioned processes until it finally operates the robot 100 based on the user's instruction.

[0065] (Third example of robot control system operation) A third operation example of the robot control system 1 will be described with reference to Fig. 6. Fig. 6 is a schematic diagram for explaining the third operation example of the robot control system 1 according to the present embodiment, and shows the conversation log display unit 340. In this example, a case where a predetermined operation is not possible will be taken as an example.

[0066] First, the user inputs an instruction to the display / input device 300. In this example, a mark is placed on one of the packages 406 (see FIG. 2 for the reference numerals) shown in the environmental image P (corresponding to the "predetermined package"), and the user verbally inputs "please pick up this toy" (see the upper part of FIG. 6). Then, the control device 200 acquires the instruction input from the display / input device 300 (step S501).

[0067] Next, the control device 200 performs inference on the specified baggage (step S502).

[0068] Next, the control device 200 determines whether the predetermined baggage has been identified by inference (step S503). In this example, since one of the bags 406 (see FIG. 2 for the reference numerals) is marked in the environmental image P included in the instruction input, it is determined that the predetermined baggage has been identified (step S503: Yes).

[0069] Next, the control device 200 performs inference on a predetermined operation of the robot 100 for acquiring and gripping the predetermined load (step S504).

[0070] Next, the control device 200 determines whether the inferred predetermined operation is possible (step S505). In this example, it is determined that the predetermined operation is impossible because another load (hereinafter also referred to as an "obstacle") is placed on the predetermined load (step S505: No). Other cases in which it is determined that the predetermined operation is impossible include, for example, when another robot or equipment is present on the trajectory of the arm for the predetermined operation, or when the robot's performance is insufficient to grasp the predetermined load due to its size, weight, shape, material, etc. In other words, even if the predetermined load itself can be identified, if there is a factor that prevents the predetermined operation on the predetermined load, it is determined that the predetermined operation is impossible.

[0071] Next, the control device 200 outputs an instruction request to the display / input device 300 to prompt the user to give a necessary instruction (step S510). Here, the user is prompted to give an instruction that will enable the predetermined operation. In this example, the instruction request is output using language such as "Another object is in the way and I can't get it" (see the middle part of FIG. 6).

[0072] Next, the user performs an instruction input operation in response to the instruction request to the display / input device 300. In this example, the user indicates with an arrow that an obstacle shown in the environmental image P should be removed, and verbally inputs, "Please move this in the direction of the arrow first" (see the lower part of FIG. 6). Then, the control device 200 acquires the instruction input in response to the instruction request from the display / input device 300 (step S501).

[0073] Next, the process of step S502 is performed by the control device 200. Then, the control device 200 performs the above-described processes until the robot 100 is finally operated based on the user's instruction.

[0074] (Fourth example of robot control system operation) A fourth operation example of the robot control system 1 will be described with reference to Fig. 7. Fig. 7 is a schematic diagram for explaining the fourth operation example of the robot control system 1 according to the present embodiment, and shows the conversation log display unit 340. In this example, a case where the inference is incorrect will be taken as an example.

[0075] First, the user inputs an instruction to the display / input device 300. In this example, it is assumed that the environmental image P has not been edited by the user and that the user verbally inputs "Please move the square toy to the box below" (see the first row in FIG. 5). Then, the control device 200 acquires the instruction input from the display / input device 300 (step S501). In this example, it is assumed that one of the packages 403 (see FIG. 2 for the reference numerals) is the specified package.

[0076] Next, the control device 200 performs inference on the specified baggage (step S502).

[0077] Next, the control device 200 determines whether or not the predetermined package has been identified by inference (step S503). In this example, it is assumed that the package 405 (see FIG. 2 for the reference numeral) has been mistakenly identified as the predetermined package (step S503: Yes).

[0078] Next, the control device 200 performs inference on a predetermined operation of the robot 100 for acquiring and gripping the predetermined load (step S504).

[0079] Next, the control device 200 determines whether the inferred predetermined action is possible (step S505). In this example, it is assumed that it is determined that the predetermined action is possible (step S505: Yes).

[0080] Next, the control device 200 outputs two types of inference results, namely, the inferred predetermined baggage and the predetermined action, to the display / input device 300 (step S506). In this example, the inference result is output using an image in which the predetermined action is represented by an arrow indicating the movement of the predetermined baggage shown in the environmental image P, and a message such as "Is this movement correct?" to prompt an instruction from the user (see the middle part of FIG. 7). At this time, the control device 200 may output an instruction request, such as "Do you want to perform the action?", to the display / input device 300 to confirm whether or not to accept the two types of inference results, or the instruction request may be output to the display / input device 300 between step S507 and step S508.

[0081] Next, the user performs an instruction input operation for the two inference results to the display / input device 300. In this example, the two inference results are incorrect, that is, they do not match the user's instructions, so the user verbally inputs "That's wrong. Please move the toy in the box on the top right to the box on the bottom right" (see the bottom part of FIG. 7). Then, the control device 200 acquires an instruction input for the two inference results from the display / input device 300 (step S507).

[0082] Next, the control device 200 determines whether the two inference results have been accepted based on the instruction input for the two acquired inference results (step S508). In this example, since an instruction input indicating that the two inference results are incorrect (answer "That's wrong...") has been received, it is determined that the two inference results have not been accepted (step S508: No).

[0083] Next, when the user inputs an instruction in response to the instruction request to the display / input device 300, the control device 200 performs the process of step S501. Then, the control device 200 performs each of the above-mentioned processes until it finally operates the robot 100 based on the user's instruction.

[0084] In the above explanations of each operation example of the robot control system, a final goal (for example, "move the square toy to the adjacent box") is given as an instruction, but in the robot control system 1, it is also possible to give step-by-step instructions to achieve the final goal (for example, "first, move the round toy that is on top of the square toy") Furthermore, in the robot control system 1, it is possible to issue operation instructions for the robot 100 from a location remote from the robot 100, so it is possible to respond to unexpected situations without going to the site.

[0085] As described above, a robot control system (e.g., 1) according to this embodiment includes a display / input device (e.g., 300) that displays an environmental image (e.g., P) including an object to be processed by a robot (e.g., 100) and accepts instruction input from a user using natural language and / or editing operations on the environmental image, and a control device (e.g., 200) that infers, based on the instruction input, at least one of a predetermined object (e.g., 401-406) to be processed by the robot through user operation and a predetermined robot action for processing the predetermined object, and outputs the inference result to the display / input device. This configuration makes it possible to accept intuitive instructions from a user regarding the robot's actions and to accurately operate the robot based on the instructions. In addition, it is possible to operate the robot with fewer instructions.

[0086] The control device also generates a control signal for the robot based on the inference result. This configuration allows for a high level of compatibility between various inferences and robot control based on user instructions.

[0087] The control device can transmit a control signal related to a predetermined action to the robot and determine whether the predetermined action can be performed before transmitting the control signal related to the predetermined action to the robot. This configuration allows the robot to operate more accurately based on user instructions and prevents the robot from operating unnecessarily. It also prevents accidents and breakdowns caused by causing the robot to perform actions that are difficult to achieve given the robot's performance or environment.

[0088] Furthermore, if the control device determines that the predetermined movement is impossible, it outputs an instruction request to the display / input device to prompt the robot to provide the necessary instructions to make the predetermined movement possible. This configuration allows the user's instructions to be accurately understood, allowing the robot to operate more accurately.

[0089] Furthermore, if the control device determines that the predetermined action is possible, it outputs an instruction request to the display / input device to confirm whether or not the user accepts the inference result, and if the inference result is accepted by the user, it transmits a control signal related to the predetermined action to the robot. With this configuration, the user's instructions can be accurately understood, and the robot can be operated more accurately.

[0090] Furthermore, if the control device is unable to identify the target object, the control device outputs to the display / input device an instruction request prompting the robot to provide instructions necessary to identify the target object. This configuration allows the user's instructions to be accurately understood, allowing the robot to operate more accurately.

[0091] The control device also uses the visual language model to acquire information contained in the environmental image and editing operations on the environmental image, and uses the large-scale language model to infer a predetermined object or a predetermined action based on the information acquired by the visual language model and a conversation log with the user including the instruction input and the inference result. This configuration makes it possible to accurately accept intuitive instructions from the user regarding the robot's actions.

[0092] The control device also outputs, to the display / input device, at least one of an image generated from the inferred predetermined object or predetermined action using a visual language model and a text generated from a large-scale language model as an inference result. This configuration makes it possible to more accurately receive intuitive instructions from a user regarding the robot's actions.

[0093] The display / input device also includes an image editing unit (e.g., 320) that accepts editing operations by the user on the displayed environmental image, a language input unit (e.g., 330) that accepts language input operations by the user, and a conversation log display unit (e.g., 340) that displays a conversation log including input instructions and inference results. This configuration makes it possible to more accurately accept intuitive instructions from the user regarding the robot's operations.

[0094] The image editing unit also receives user input of language into the displayed environmental image. This configuration allows the robot to more accurately receive intuitive instructions from the user regarding the robot's operation.

[0095] Furthermore, a robot control device (e.g., ) according to this embodiment includes an input unit that displays an environmental image (e.g., P) including an object to be processed by a robot (e.g., 100) and receives instruction input from a display / input device (e.g., 300) that receives instruction input from a user using natural language and / or editing operations on the environmental image, and a control unit (e.g., 200) that infers, based on the instruction input, at least one of a predetermined object to be processed by the robot through user operation and a predetermined robot action for processing the predetermined object, and outputs the inference result to the display / input device. This configuration makes it possible to receive intuitive instructions from a user regarding the robot's actions and to accurately operate the robot based on the instructions. In addition, it is possible to operate the robot with fewer instructions.

[0096] Furthermore, a robot control method according to this embodiment displays an environmental image (e.g., P) including an object to be processed by a robot (e.g., 100), and receives instruction input from a display / input device (e.g., 300) that receives instruction input from a user using natural language and / or editing operations on the environmental image. Based on the instruction input, the method infers at least one of a predetermined object to be processed by the robot through user operation and a predetermined robot action for processing the predetermined object, and outputs the inference result to the display / input device. This configuration makes it possible to receive intuitive instructions from a user regarding the robot's actions and to operate the robot accurately based on the instructions. In addition, it is possible to operate the robot with fewer instructions.

[0097] (Other variations) In the above-described embodiment, an example has been described in which two types of inference results, a predetermined graspable object and a predetermined action, are output. However, the object of inference may be either the object or the action. For example, when control is performed such that the object is first identified and then the action is identified, only the inference result for the object may be output at the stage of identifying the object. Furthermore, when the object has already been identified or there is only one object, it is not necessary to infer the object again and it is sufficient to infer only the action, so only the inference result for the action may be output. In the above-described embodiment, an example has been described in which instructions are mainly given to a robot. However, the concept of the present embodiment may be applied to things other than giving instructions to a robot. For example, application to dialogue with an AI assistant as in Patent Document 1 is conceivable. In the above-described embodiment, input by a user and output to the user are performed using the display input device 300. However, input and output may be performed using different devices. Also, some or all of the output may be performed on an output device for a user other than the display input device 300 of the user who input the instruction. For example, the inference result may be displayed on a terminal owned by the superior of the user who input the instruction, for approval. Also, output is not limited to display on a screen, and may be output by voice or the like. Furthermore, the above-described embodiment has been described with reference to an example of a robot that grasps and transports a load, which is a graspable object. However, the concept of the present disclosure can also be applied to the control of a robot that grasps an object other than load and performs a task other than transportation. For example, the present disclosure can be applied to the control of a robot that paints a predetermined area of ​​an object or glues multiple objects together. The present disclosure also applies to programs and storage media that realize the functions of the devices of the above-mentioned embodiments, which are supplied to the device via a network or various storage media and read and executed by a computer within the device.

[0098] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to these examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents may be made within the scope of the claims, and it is understood that these also fall within the technical scope of the present disclosure. Furthermore, the components of the various embodiments described above may be combined in any manner without departing from the spirit of the invention.

[0099] (Addendum) The above-described embodiments disclose the following techniques. (Technology 1) a display / input device that displays an environmental image including an object to be processed by the robot and accepts instruction input from a user using natural language and / or editing operations for the environmental image; a control device that infers at least one of a predetermined object to be processed by the robot in response to an operation by the user and a predetermined action of the robot for performing the processing on the predetermined object based on the instruction input, and outputs the inference result to the display / input device; Equipped with Robot control system. This configuration makes it possible to receive intuitive instructions from the user regarding the robot's operation and to make the robot operate accurately based on the instructions.

[0100] (Technology 2) The control device generating a control signal for the robot based on the inference result; The robot control system according to claim 1. This configuration enables both inference and robot control to be performed at a high level.

[0101] (Technology 3) The control device a control signal relating to the predetermined operation can be transmitted to the robot, and a determination is made as to whether the predetermined operation is possible or not before transmitting the control signal relating to the predetermined operation to the robot; The robot control system according to Technology 1 or Technology 2. This configuration allows the robot to operate more accurately based on user instructions, and prevents the robot from operating unnecessarily.

[0102] (Technology 4) The control device When it is determined that the predetermined action is impossible, an instruction request is output to the display / input device to prompt the user to provide an instruction required to make the predetermined action possible. The robot control system described in technique 3. This configuration allows the user's instructions to be accurately understood, allowing the robot to operate more accurately.

[0103] (Technology 5) The control device further When it is determined that the predetermined operation is possible, an instruction request for confirming whether or not to accept the inference result is output to the display / input device; transmitting a control signal relating to the predetermined action to the robot when the inference result is accepted by the user; The robot control system according to Technology 3 or Technology 4. This configuration allows the user's instructions to be accurately understood, allowing the robot to operate more accurately.

[0104] (Technology 6) The control device further When the predetermined object cannot be identified, an instruction request is output to the display input device to prompt an instruction required to identify the object. A robot control system according to any one of techniques 1 to 5. This configuration allows the user's instructions to be accurately understood, allowing the robot to operate more accurately.

[0105] (Technology 7) The control device using a visual language model to obtain information contained in the environmental image and editing operations on the environmental image; using a large-scale language model, inferring the predetermined object or the predetermined action based on the information acquired by the visual language model and a conversation log with the user including the instruction input and the inference result; A robot control system according to any one of techniques 1 to 6. This configuration makes it possible to receive intuitive instructions from the user regarding the robot's operation with high accuracy.

[0106] (Technology 8) The control device outputting at least one of an image generated from the inferred predetermined object or predetermined action using the visual language model and a text generated from the inferred predetermined object or predetermined action to the display input device as the inference result; The robot control system described in technique 7. This configuration makes it possible to receive intuitive instructions from the user regarding the robot's operation with high accuracy.

[0107] (Technology 9) the display input device, an image editing unit that accepts an editing operation by the user on the displayed environmental image; a language input unit that accepts a language input operation by the user; a conversation log display unit that displays a conversation log including the instruction input and the inference result, The robot control system described in technique 7. This configuration makes it possible to receive intuitive instructions from the user regarding the robot's operation with high accuracy.

[0108] (Technology 10) The image editing unit accepting a language input operation by the user into the displayed environmental image; The robot control system described in technique 9. This configuration makes it possible to more accurately receive intuitive instructions from the user regarding the robot's operation.

[0109] (Technology 11) an input unit that displays an environmental image including an object to be processed by the robot and receives an instruction input from a display input device that receives an instruction input from a user using natural language and / or an editing operation for the environmental image; a control unit that infers at least one of a predetermined object to be processed by the robot in response to an operation by the user and a predetermined action of the robot for performing the processing on the predetermined object based on the instruction input, and outputs the inference result to the display / input device; A robot control device comprising: This configuration makes it possible to receive intuitive instructions from the user regarding the robot's operation and to make the robot operate accurately based on the instructions.

[0110] (Technology 12) receiving an instruction input from a display input device that displays an environmental image including an object to be processed by the robot and receives an instruction input from a user using natural language and / or an editing operation for the environmental image; based on the instruction input, inferring at least one of a predetermined object to be processed by the robot in response to an operation by the user and a predetermined action of the robot for performing the processing on the predetermined object, and outputting the inference result to the display / input device; Robot control method. This configuration makes it possible to receive intuitive instructions from the user regarding the robot's operation and to make the robot operate accurately based on the instructions. [Industrial Applicability]

[0111] The present disclosure is useful as a robot control system, a robot control device, and a robot control method. [Explanation of symbols]

[0112] 1. Robot Control System 100 robots 200 control device 210 Reasoning Department 220 Motion control section 300 Display input device 310 display screen 320 Image Editorial Department 330 Language input section 340 Conversation log display section 400 luggage P Environmental Image

Claims

1. a display / input device that displays an environmental image including an object to be processed by the robot and accepts instruction input from a user using natural language and / or editing operations for the environmental image; a control device that infers at least one of a predetermined object to be processed by the robot in response to an operation by the user and a predetermined action of the robot for performing the processing on the predetermined object based on the instruction input, and outputs the inference result to the display / input device; Equipped with Robot control system.

2. The control device generating a control signal for the robot based on the inference result; The robot control system of claim 1 .

3. The control device a control signal relating to the predetermined operation can be transmitted to the robot, and a determination is made as to whether the predetermined operation is possible or not before transmitting the control signal relating to the predetermined operation to the robot; The robot control system of claim 2 .

4. The control device When it is determined that the predetermined action is impossible, an instruction request is output to the display / input device to prompt the user to provide an instruction required to make the predetermined action possible. The robot control system of claim 3 .

5. The control device further When it is determined that the predetermined operation is possible, an instruction request for confirming whether or not to accept the inference result is output to the display / input device; transmitting a control signal relating to the predetermined action to the robot when the inference result is accepted by the user; The robot control system of claim 4.

6. The control device further When the predetermined object cannot be identified, an instruction request is output to the display input device to prompt an instruction required to identify the object. The robot control system of claim 1 .

7. The control device using a visual language model to obtain information contained in the environmental image and editing operations on the environmental image; using a large-scale language model, inferring the predetermined object or the predetermined action based on the information acquired by the visual language model and a conversation log with the user including the instruction input and the inference result; The robot control system of claim 1 .

8. The control device outputting at least one of an image generated from the inferred predetermined object or predetermined action using the visual language model and a text generated from the inferred predetermined object or predetermined action to the display input device as the inference result; The robot control system of claim 7.

9. the display input device, an image editing unit that accepts an editing operation by the user on the displayed environmental image; a language input unit that accepts a language input operation by the user; a conversation log display unit that displays a conversation log including the instruction input and the inference result, The robot control system of claim 7.

10. The image editing unit accepting a language input operation by the user into the displayed environmental image; The robot control system of claim 9.

11. an input unit that displays an environmental image including an object to be processed by the robot and receives an instruction input from a display input device that receives an instruction input from a user using natural language and / or an editing operation for the environmental image; a control unit that infers at least one of a predetermined object to be processed by the robot in response to an operation by the user and a predetermined action of the robot for performing the processing on the predetermined object based on the instruction input, and outputs the inference result to the display / input device; A robot control device comprising:

12. receiving an instruction input from a display input device that displays an environmental image including an object to be processed by the robot and receives an instruction input from a user using natural language and / or an editing operation for the environmental image; based on the instruction input, inferring at least one of a predetermined object to be processed by the robot in response to an operation by the user and a predetermined action of the robot for performing the processing on the predetermined object, and outputting the inference result to the display / input device; Robot control method.

Citation Information

Patent Citations

  • A conversational AI platform that utilizes rendered graphical output

    JP2023525173A