Autonomous running robot operation system, autonomous running robot operation method, and program

The autonomous mobile robot operation system simplifies robot control through voice and handwriting inputs, leveraging AI for intuitive operation and efficient tasks like grasping and manipulation.

JP2025136027APending Publication Date: 2025-09-19NAGOYA ELECTRICAL EDUCATIONAL FOUNDATION +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024034184
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing methods for controlling autonomous mobile robots require complex text chat input, complicating the operation process.

Method used

An autonomous mobile robot operation system that utilizes voice input and handwriting input to simplify the operation of autonomous mobile robots, incorporating machine-learned voice recognition and object estimation using artificial intelligence for intuitive control.

Benefits of technology

Enables easier and more intuitive operation of autonomous mobile robots by allowing voice and handwritten inputs, facilitating operations such as grasping, cutting, moving, and welding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025136027000001_ABST
    Figure 2025136027000001_ABST
Patent Text Reader

Abstract

To provide an autonomous running robot operation system that operates an autonomous running robot more easily, using sound input and handwriting input.SOLUTION: The autonomous running robot operation system comprises: an autonomous running robot that photographs an image of a circumferential environment and operates an object to be operated; a handwriting input interface that displays the image photographed by the autonomous running robot and receives handwriting input to the displayed image; and a sound input interface that receives sound input to the object to be operated, which operates the autonomous running robot so that the robot operates the object to be operated, in accordance with an instructions for the handwriting input and the sound input.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an autonomously moving robot operation system, an autonomously moving robot operation method, and a program. [Background technology]

[0002] Patent Document 1 discloses a remote control system that controls an end effector through handwriting input and text chat. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-94604 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the method of Patent Document 1, text chat must be done by handwriting, and input is complicated because handwriting input and operation overlap. Therefore, an object of the present disclosure is to provide an autonomous mobile robot operation system that uses voice input and handwriting input to more easily operate an autonomous mobile robot. [Means for solving the problem]

[0005] The robot manipulation system of the present disclosure comprises: an autonomous robot that captures images of its surroundings and operates an object to be operated; a handwriting input interface that displays an image captured by the autonomous robot and accepts handwriting input on the displayed image; a voice input interface that accepts voice input for the target to be operated, The autonomous mobile robot operation system operates the autonomous mobile robot so as to operate the operated object in accordance with the handwritten input and voice input instructions.

[0006] The above configuration provides an autonomous mobile robot operation system that allows for easier operation of an autonomous mobile robot using voice input and handwritten input.

[0007] The robot manipulation system of the present disclosure comprises: The voice input interface is characterized in that it uses a machine-learned voice recognition unit that inputs voice, recognizes the voice, and outputs an action, and inputs the actions of the autonomous robot through dialogue.

[0008] The above configuration is an example of speech recognition using artificial intelligence (AI).

[0009] The robot manipulation system of the present disclosure comprises: The voice input interface is used to give instructions to the target object that is not displayed in the image.

[0010] With the above configuration, it is possible to give instructions to an object to be operated that is not directly displayed on the image.

[0011] The robot manipulation system of the present disclosure comprises: The handwriting input interface is characterized in that the image is input, and the object to be operated is input using an object estimation unit that has been trained by machine learning to estimate and output an object in the image.

[0012] The above configuration is an example of object estimation using artificial intelligence (AI).

[0013] The robot manipulation system of the present disclosure comprises: The handwriting input interface is used to input a trajectory of the autonomous mobile robot.

[0014] With the above configuration, the trajectory of the autonomous mobile robot can be easily input.

[0015] The robot manipulation system of the present disclosure comprises: The handwriting input interface is used to input adverbial operations of actions.

[0016] With the above configuration, the movement of the autonomous mobile robot can be intuitively controlled.

[0017] The robot manipulation system of the present disclosure comprises: The operation of the autonomous mobile robot is characterized by being grasping, cutting, moving, screwing, or welding.

[0018] The above configuration is an example of the operation of an autonomous mobile robot.

[0019] The robot operation method of the present disclosure includes: The method for operating an autonomously moving robot is to operate the autonomously moving robot in the same way as operating an object to be operated according to instructions input by handwriting and voice.

[0020] The above configuration provides a method for operating an autonomous mobile robot that allows for easier operation of the autonomous mobile robot using voice input and handwritten input.

[0021] The program of the present disclosure is This is a program that causes an information processing device to operate an autonomous robot in the same way as operating an object to be operated according to handwritten and voice input instructions.

[0022] The above configuration provides a program for an information processing device that executes an autonomously moving robot operation that allows the autonomously moving robot to be operated more easily using voice input and handwritten input. [Effects of the Invention]

[0023] The present disclosure provides an autonomous mobile robot operation system that uses voice input and handwriting input to more easily operate an autonomous mobile robot. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a diagram illustrating an overview of an autonomous traveling robot operation system according to an embodiment. [Figure 2] FIG. 10 is a diagram illustrating an example of a display screen displayed on a display panel of a remote terminal according to an embodiment. [Figure 3] 1 is an external perspective view showing an example of the external configuration of an autonomous traveling robot according to an embodiment; [Figure 4] 1 is a block diagram showing a configuration of an autonomous mobile robot according to an embodiment; [Figure 5] 1A and 1B are diagrams illustrating examples of captured images acquired by an autonomous mobile robot according to an embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of an operable region output by a trained model according to an embodiment. [Figure 7] FIG. 2 is a block diagram illustrating a configuration of a remote terminal according to the embodiment. [Figure 8] 1 is a flowchart of a method for operating an autonomous mobile robot according to an embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of a display screen displayed on a display panel of a remote terminal according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0025] Embodiment Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the invention according to the claims is not limited to the following embodiments. Furthermore, not all of the configurations described in the embodiments are necessarily essential means for solving the problems. For clarity of explanation, the following description and drawings have been omitted and simplified as appropriate. In each drawing, the same elements are given the same reference numerals, and duplicate explanations are omitted as necessary. Furthermore, in the following embodiments, a robot having a hand at the end of an arm as an end effector will be described as an example of an object to be operated, but the object to be operated is not limited to this.

[0026] (Description of an Autonomous Mobile Robot Operation System According to an Embodiment) 1 is a diagram showing an overview of an autonomous mobile robot operation system according to an embodiment. An autonomous mobile robot 100 that performs various actions in a first environment is remotely controlled by a user who serves as a remote operator in a second environment away from the first environment, by operating a remote terminal 300 (operation terminal) via a system server 500 connected to the Internet 600. The remote terminal 300 is also referred to as a handwriting input interface 300.

[0027] In a first environment, the autonomous mobile robot 100 is connected to the Internet 600 via a wireless router 700. In a second environment, the remote terminal 300 is connected to the Internet 600 via the wireless router 700. The system server 500 is connected to the Internet 600. The autonomous mobile robot 100 performs a grasping action or the like using the hand 124 in accordance with the operation of the remote terminal 300.

[0028] In this embodiment, the operation by the hand 124 is not limited to the operation of simply gripping (grasping) an object to be operated, but also includes, for example, the following operations. - Grab and lift the object being operated If the object being operated is the knob of a door or drawer of a dresser, etc., the action of grabbing the knob and opening or closing the door or drawer. If the object being operated is a doorknob, the action of grabbing the doorknob and opening or closing the door - Disconnecting the controlled object -Moving the controlled object - Screw fixing of the operated object - Welding the object being operated

[0029] The autonomous mobile robot 100 captures an image of the first environment in which the autonomous mobile robot 100 exists using the stereo camera 131 (imaging unit), and transmits the captured image to the remote terminal 300 via the Internet 600. The example in Fig. 1 shows the autonomous mobile robot 100 capturing an image of a table 400 that exists in the first environment.

[0030] The remote terminal 300 is, for example, a tablet terminal and includes a display panel 341 on which a touch panel is superimposed. The display panel 341 displays a captured image received from the autonomous mobile robot 100, allowing the user to indirectly view the first environment in which the autonomous mobile robot 100 exists. The user can also input handwritten input information by hand on the captured image displayed on the display panel 341. The handwritten input information is, for example, information instructing an operation target that requires operation by the hand 124 or an adverbial manner of operating the operation target. Methods for inputting handwritten input information include, but are not limited to, touching a touch panel superimposed on the display panel 341 with the user's finger or a touch pen. The handwritten input information input by the user on the captured image is transmitted to the autonomous mobile robot 100 via the Internet 600.

[0031] The remote terminal 300 also includes a voice input interface 360 ​​for the user to have a voice conversation with the autonomous mobile robot 100. The voice input interface 360 ​​uses a machine-learned voice recognition unit that inputs voice, recognizes the voice, and outputs an action to interact and input the action of the autonomous mobile robot 100. The utterance input by the user is transmitted to the autonomous mobile robot 100 via the Internet 600. The response utterance generated by the autonomous mobile robot 100 in response to the user's utterance is received from the autonomous mobile robot 100 via the Internet 600.

[0032] 2 is a diagram illustrating an example of a display screen displayed on a display panel of a remote terminal according to an embodiment. In the example of Fig. 2, a captured image 311 captured by the autonomous mobile robot 100 is displayed on a display screen 310.

[0033] The captured image 311 shows a table 400, a cup 401 placed on the table 400, a calculator 402, a smartphone 403, and a piece of paper 404. The cup 401, the calculator 402, the smartphone 403, and the piece of paper 404 are operable objects that can be operated by the hand 124. Therefore, the captured image 311 is processed to display the names of the operable objects in speech bubbles so that the user can visually recognize the operable objects. Handwritten input information 931 for the captured image 311 has been handwritten by the user.

[0034] Based on the handwritten input information entered by the user into the captured image and the dialogue history of the voice input, the autonomous mobile robot 100 estimates the object to be operated that is being requested to be operated by the hand 124, and also estimates the manner in which the hand 124 is requested to operate on the estimated object to be operated.

[0035] In the example of FIG. 2, handwritten input information 931 is input at the position of the smartphone 403 on the captured image 311. Assume also that a voice input requests a grasping action to grab and lift the target to be operated. Based on the handwritten input information 931 and the voice input, the autonomous mobile robot 100 can estimate that the target to be operated is the smartphone 403 placed on the table 400 and that the manner of movement is to grab and lift the smartphone 403. Note that, in the example of FIG. 2, the handwritten input information 931 is an image simulating grabbing the smartphone 403 from above, but this is not limiting. The handwritten input information 931 may simply be an image indicating that the smartphone 403 is the target to be operated, and the manner of movement may be instructed by a user in a dialogue via voice input. The image of the handwritten input information 931 indicating that the smartphone 403 is the target to be operated may be, for example, an image in which an arrow points at the smartphone 403 or an image in which the smartphone 403 is surrounded by an arbitrary shape (for example, a circle).

[0036] Furthermore, the autonomously moving robot 100 may determine, based on the dialogue history of the text chat, whether or not there is any additional action that is being requested of the autonomously moving robot 100, and if there is any additional action that is being requested of the autonomously moving robot 100, estimate how that action will be performed.

[0037] For example, a voice input is used to request that the smartphone 403 be transported to the living room. Therefore, the autonomous mobile robot 100 can infer, based on the voice input, that an additional request has been made to the autonomous mobile robot 100 to transport the smartphone 403 that has been grasped by the grasping motion to the living room.

[0038] Therefore, in the example of FIG. 2, the autonomous mobile robot 100 can infer that the overall action required of the autonomous mobile robot 100 is to grab the smartphone 403 and carry it to the living room.

[0039] FIG. 3 is an external perspective view showing an example of the external configuration of an autonomous mobile robot according to an embodiment. The autonomous mobile robot 100 is broadly composed of a cart unit 110 and a main body unit 120. The cart unit 110 is configured in a cylindrical housing and supports two drive wheels 111 and one caster 112, each of which comes into contact with the traveling surface. The two drive wheels 111 are arranged so that their rotation axes coincide with each other. Each drive wheel 111 is independently driven and rotated by a motor (not shown). The caster 112 is a driven wheel, and is provided such that a pivot extending vertically from the cart unit 110 is spaced apart from the rotation axis of the wheel to pivotally support the wheel, and follows the direction of movement of the cart unit 110.

[0040] The cart unit 110 is equipped with a laser scanner 133 on the periphery of its upper surface. The laser scanner 133 scans a certain range in the horizontal plane at each step angle and outputs whether or not an obstacle is present in each direction. Furthermore, if an obstacle is present, the laser scanner 133 outputs the distance to the obstacle.

[0041] The main body 120 mainly includes a trunk 121 mounted on the upper surface of the cart 110, a head 122 placed on the upper surface of the trunk 121, an arm 123 supported on the side of the trunk 121, and a hand 124 attached to the tip of the arm 123. The arm 123 and the hand 124 are driven via a motor (not shown) to grasp an object to be operated. The trunk 121 can rotate around a vertical axis relative to the cart 110 by the driving force of the motor (not shown). Depending on the application, a welding tool or a cutting tool may be attached instead of the hand 124.

[0042] Head 122 mainly includes stereo camera 131 and display panel 141. Stereo camera 131 has a configuration in which two camera units having the same angle of view are arranged apart from each other, and outputs image signals captured by each camera unit.

[0043] Display panel 141 is, for example, a liquid crystal panel, and displays the face of a set character as an animation, or displays information about autonomous mobile robot 100 as text or icons. Displaying a character's face on display panel 141 gives people around the display panel the impression that display panel 141 is a pseudo-face.

[0044] Head 122 can rotate around a vertical axis relative to body 121 by the driving force of a motor (not shown). Therefore, stereo camera 131 can capture images in any direction, and display panel 141 can present display content in any direction.

[0045] 4 is a block diagram showing the configuration of the autonomous mobile robot according to the embodiment. Here, the main elements related to the estimation of the operated object and the manner of movement will be described, but the configuration of the autonomous mobile robot 100 may also include other elements, and other elements that contribute to the estimation of the operated object and the manner of movement may also be added.

[0046] The autonomous mobile robot 100 is equipped with an information processing device including a control unit 150 with a processor that executes programs and performs processing, and a memory 180 that stores the programs. The control unit 150 is, for example, a CPU, and is stored in a control unit provided in, for example, the body 121. The carriage drive unit 145 includes drive wheels 111 and a drive circuit and motor for driving the drive wheels 111. The control unit 150 controls the rotation of the drive wheels by sending drive signals to the carriage drive unit 145. The control unit 150 also receives feedback signals from an encoder or the like from the carriage drive unit 145 to determine the direction and speed of movement of the carriage unit 110.

[0047] Upper body drive unit 146 includes arm 123, hand 124, torso 121, and head 122, as well as drive circuits and motors for driving these. Control unit 150 realizes grasping actions and gestures by sending drive signals to upper body drive unit 146. Control unit 150 also receives feedback signals from an encoder or the like from upper body drive unit 146 to determine the positions and movement speeds of arm 123 and hand 124, and the orientations and rotation speeds of torso 121 and head 122.

[0048] The display panel 141 receives and displays the image signal generated by the control unit 150. The control unit 150 also generates image signals for characters and the like and causes the display panel 141 to display them, as described above.

[0049] In response to a request from the control unit 150, the stereo camera 131 captures an image of the first environment in which the autonomous mobile robot 100 exists and passes the captured image signal to the control unit 150. The control unit 150 performs image processing using the captured image signal and converts the captured image signal into a captured image in accordance with a predetermined format. In response to a request from the control unit 150, the laser scanner 133 detects whether or not an obstacle exists in the direction of movement and passes a detection signal representing the detection result to the control unit 150.

[0050] The hand camera 135 is, for example, a distance image sensor, and is used to recognize the distance, shape, direction, etc., of the operated object. The hand camera 135 includes an image sensor in which pixels that photoelectrically convert an optical image incident from the target space are arranged two-dimensionally, and outputs the distance to the object for each pixel to the control unit 150. Specifically, the hand camera 135 includes an irradiation unit that irradiates the target space with pattern light, receives the reflected light with the image sensor, and outputs the distance to the object captured by each pixel based on the distortion and size of the pattern in the image. The control unit 150 grasps the state of the wider surrounding environment using the stereo camera 131, and grasps the state near the operated object using the hand camera 135.

[0051] The memory 180 is a non-volatile storage medium, such as a solid state drive. The memory 180 stores various parameter values, functions, lookup tables, and the like used for control and calculation, in addition to a control program for controlling the autonomous mobile robot 100. In particular, the memory 180 stores a trained model 181, an utterance DB 182, and a map DB 183.

[0052] Trained model 181 is a trained model that takes a captured image as an input image and outputs an operable object that appears in the captured image. Trained model 181 is also a trained model that inputs voice, recognizes the voice, and outputs an action. Utterance DB 182 is configured by a recording medium such as a hard disk drive, and is a database in which individual terms systematized as a corpus are stored together with reproducible utterance data.

[0053] The map DB 183 is configured by a recording medium such as a hard disk drive, and is a database that stores map information that describes the space in the first environment in which the autonomous mobile robot 100 exists.

[0054] The communication unit 190 is, for example, a wireless LAN unit, and performs wireless communication with the wireless router 700. The communication unit 190 receives handwritten input information for the captured image and the user's voice input sent from the remote terminal 300, and passes them on to the control unit 150. The communication unit 190 also transmits the captured image captured by the stereo camera 131 to the remote terminal 300 under the control of the control unit 150. The communication unit 190 also transmits to the voice input interface 360 ​​the voice output of a response utterance sentence to the user's utterance sentence, which is generated by the control unit 150.

[0055] The control unit 150 executes a control program read from the memory 180 to control the entire autonomous mobile robot 100 and perform various arithmetic processing. The control unit 150 also serves as a function execution unit that executes various calculations and controls related to control. As such function execution units, the control unit 150 includes a recognition unit 151 and an estimation unit 152. The estimation unit 152 is also referred to as an object estimation unit.

[0056] The recognition unit 151 uses an image captured by one of the camera units of the stereo camera 131 as an input image, obtains an operable area that can be operated by the hand 124 and is shown in the captured image from the trained model 181 read out from the memory 180, and recognizes the operable part.

[0057] Fig. 5 is a diagram illustrating an example of a captured image acquired by the autonomous mobile robot according to the embodiment. For example, Fig. 5 illustrates an example of a captured image 311 of a first environment acquired by the autonomous mobile robot 100 using the stereo camera 131. The captured image 311 in Fig. 5 includes a table 400, a cup 401 placed on the table 400, a calculator 402, a smartphone 403, and a piece of paper 404. The recognition unit 151 provides the captured image 311 to the trained model 181 as an input image.

[0058] 6 is a diagram illustrating an example of an operable area output by the trained model according to the embodiment. For example, FIG. 6 is a diagram illustrating an example of an operable area output by the trained model 181 when the captured image 311 in FIG. 5 is used as an input image. Specifically, the area surrounding the cup 401 is detected as an operable area 801, the area surrounding the calculator 402 as an operable area 802, the area surrounding the smartphone 403 as an operable area 803, and the area surrounding the paper 404 as an operable area 804. Therefore, the recognition unit 151 recognizes the cup 401, the calculator 402, the smartphone 403, and the paper 404, which are respectively surrounded by the operable areas 801 to 804, as operable parts.

[0059] The trained model 181 is a neural network trained using training data that is a combination of an image showing an operable part that can be operated by the hand 124 and a correct answer value indicating which area of ​​the image is the operable part. In this case, by using the training data as training data that further indicates the name, distance, and direction of the operable part in the image, the trained model 181 can be a trained model that not only outputs the operable part when a captured image is used as an input image, but also outputs the name, distance, and direction of the operable part. The trained model 181 may be a neural network trained by deep learning. Furthermore, the trained model 181 may be trained by adding training data as needed.

[0060] Furthermore, when the recognition unit 151 recognizes an operable part, the recognition unit 151 may process the captured image so that the user can visually recognize the operable object. As a method for processing the captured image, there is a method of displaying the name of the operable object in a speech bubble as in the example of FIG. 2, but the method is not limited to this.

[0061] Furthermore, the recognition unit 151 receives a voice from the user as input and recognizes the voice from the trained model 181.

[0062] The estimation unit 152 has a function of having a voice conversation with the user of the voice input interface 360. Specifically, the estimation unit 152 refers to the utterance DB 182 and generates a voice output of a response utterance sentence appropriate for the utterance sentence input by the user to the voice input interface 360. At this time, if the user has also input handwritten input information for the captured image to the remote terminal 300, the estimation unit 152 also refers to the handwritten input information to generate the voice output of the response utterance sentence.

[0063] The estimation unit 152 estimates an operated object that is requested to be operated by the hand 124, based on the handwritten input information input by the user to the captured image and the dialogue history of the voice input, and estimates the manner in which the hand 124 is requested to move the estimated operated object. The estimation unit 152 may also determine, based on the dialogue history of the voice input, whether or not there is any additional movement that is requested of the autonomous mobile robot 100, and, if there is any additional movement that is requested of the autonomous mobile robot 100, estimate the manner of that movement. In this case, it is preferable that the estimation unit 152 analyzes the content of the handwritten input information and the content of the voice input, and performs the above estimation while confirming the analyzed content to the voice input interface 360 ​​using voice output.

[0064] (Description of a method for estimating an object to be operated and how it behaves according to an embodiment) Hereinafter, an estimation method for estimating the operated object, the manner of movement, etc., in the estimation unit 152 of the autonomous mobile robot 100 will be described in detail using FIG. 2 as an example. 2, the autonomous mobile robot 100 first receives a voice input of the user's utterance sentence "Take this" from the voice input interface 360. At this time, the graspable objects shown in the captured image 311 captured by the autonomous mobile robot 100 are a cup 401, a calculator 402, a smartphone 403, and a piece of paper 404, which have been recognized by the recognition unit 151. The autonomous mobile robot 100 also receives handwritten input information 931 input from the remote terminal 300 at the position of the smartphone 403 on the captured image 311.

[0065] Therefore, based on the voice input of "take this," the estimation unit 152 analyzes that the manner of the action is a movement of grabbing and lifting the operated object. Furthermore, based on the handwritten input information 931, the estimation unit 152 analyzes that the operated object is the smartphone 403, which is at the input position of the handwritten input information 931, among the operable objects recognized by the recognition unit 151. The estimation unit 152 can recognize the input position of the handwritten input information 931 on the captured image 311 using any method. For example, if the remote terminal 300 transmits the handwritten input information 931 together with position information indicating the input position of the handwritten input information 931, the estimation unit 152 can recognize the input position of the handwritten input information 931 based on the position information. Alternatively, if the remote terminal 300 transmits the captured image 311 processed so that the handwritten input information 931 has been input, the estimation unit 152 can recognize the input position of the handwritten input information 931 based on the captured image 311.

[0066] Then, in order to confirm to the user that the object to be operated is the smartphone 403, the estimation unit 152 generates a voice output of a response utterance sentence, "Got it. Is it a smartphone?", and transmits the generated voice output to the voice input interface 360.

[0067] Next, the autonomous mobile robot 100 receives a voice input of the user's utterance, "Yes, bring it to me," from the remote terminal 300. Therefore, the estimation unit 152 estimates that the operated object that is being requested to be grasped by the hand 124 is the smartphone 403, and that the manner of movement is to grasp and lift up the smartphone 403.

[0068] Furthermore, since the estimation unit 152 was able to estimate the operated object and the manner of operation, it generates a voice output of a response utterance sentence of “Got it,” and transmits the generated voice output to the voice input interface 360.

[0069] Furthermore, based on the voice input of "bring it to me," the estimation unit 152 analyzes that the autonomous mobile robot 100 is additionally requested to transport the smartphone 403 that it has grasped through the grasping action to "me."

[0070] Then, the estimation unit 152 generates a voice output of the response utterance sentence "Are you in the living room?" to confirm where "my place" is, and sends the generated voice output to the voice input interface 360.

[0071] Next, the autonomous mobile robot 100 receives a voice input of the user's utterance, "That's right. Thank you," from the voice input interface 360. Therefore, the estimation unit 152 estimates that an additional operation of transporting the smartphone 403 to the living room has been requested of the autonomous mobile robot 100. As a result, the estimation unit 152 will estimate that the overall action required of the autonomous mobile robot 100 is to grab the smartphone 403 and carry it to the living room.

[0072] In this way, the estimation unit 152 can estimate the operated object that is requested to be grasped by the hand 124 and the manner in which the hand 124 is requested to perform on the operated object. Furthermore, if there is an additional action that is requested of the autonomous mobile robot 100, the estimation unit 152 can also estimate the manner of that action.

[0073] When the estimation unit 152 completes the above estimation, the control unit 150 prepares to start the requested operation of the hand 124 on the operated object. Specifically, the control unit 150 first drives the arm 123 to a position where the hand camera 135 can observe the operated object. Next, the control unit 150 causes the hand camera 135 to capture an image of the operated object and recognizes the state of the operated object.

[0074] The control unit 150 then generates a trajectory for the hand 124 to achieve the requested action on the operated object based on the state of the operated object and the manner in which the hand 124 is requested to act on the operated object. At this time, the control unit 150 generates the trajectory for the hand 124 so as to satisfy predetermined gripping conditions. The trajectory may be indicated by a straight line, a curve, an arrow, or the like using a handwriting input interface. The predetermined operation conditions include conditions for the hand 124 to grip the operated object and conditions for the trajectory until the hand 124 grasps the operated object. An example of a condition for the hand 124 to operate the operated object is that the arm 123 is not overextended when the hand 124 operates the operated object. An example of a condition for the trajectory until the hand 124 operates the operated object is that the hand 124 follows a straight trajectory when the operated object is a drawer knob.

[0075] When the control unit 150 generates the trajectory of the hand 124, it transmits a drive signal corresponding to the generated trajectory to the upper body drive unit 146. The hand 124 performs an action on the operated object in accordance with the drive signal.

[0076] When the estimation unit 152 estimates the manner of an additional operation requested of the autonomous mobile robot 100, the control unit 150 causes the autonomous mobile robot 100 to execute the additional operation requested before or after the trajectory generation and grasping operation of the hand 124. At this time, depending on the additional operation requested of the autonomous mobile robot 100, an operation to move the autonomous mobile robot 100 may be required. For example, as in the example of FIG. 2, when an additional operation to grab and transport the operated object is requested, the autonomous mobile robot 100 needs to be moved to the destination. Furthermore, when there is a distance from the current position of the autonomous mobile robot 100 to the operated object, the autonomous mobile robot 100 needs to be moved to a position near the operated object.

[0077] When an operation to move the autonomous mobile robot 100 is required, the control unit 150 acquires map information describing the space in the first environment in which the autonomous mobile robot 100 exists from the map DB 183 in order to generate a path for moving the autonomous mobile robot 100. The map information may describe, for example, the location and layout of each room in the first environment. The map information may also describe obstacles, such as chests of drawers and tables, present in each room. However, regarding obstacles, the presence or absence of obstacles in the direction of movement of the autonomous mobile robot 100 can also be detected based on a detection signal from the laser scanner 133. Furthermore, when there is a distance from the current position of the autonomous mobile robot 100 to an operated object, the distance and direction of the operated object can be obtained from the captured image acquired by the stereo camera 131 depending on the trained model 181. The distance and direction of the operated object may be obtained by image analysis of the captured image of the first environment or from information from other sensors.

[0078] Therefore, when moving the autonomous mobile robot 100 to the vicinity of the operated object, the control unit 150 generates a route for moving the autonomous mobile robot 100 from its current position to the vicinity of the operated object while avoiding obstacles, based on map information, the distance and direction of the operated object, the presence or absence of obstacles, etc. Furthermore, when moving the autonomous mobile robot 100 to a destination, the control unit 150 generates a route for moving the autonomous mobile robot 100 from its current position to the destination while avoiding obstacles, based on map information, the presence or absence of obstacles, etc. Then, the control unit 150 transmits a drive signal according to the generated route to the cart drive unit 145. The cart drive unit 145 moves the autonomous mobile robot 100 in accordance with the drive signal. Note that, if there is, for example, a door on the route to the destination, the control unit 150 needs to generate a trajectory for the hand 124 to grasp the doorknob near the door and open or close the door, and also control the hand 124 in accordance with the generated trajectory. In this case, the trajectory can be generated and the hand 124 can be controlled using, for example, the same method as described above.

[0079] 7 is a block diagram showing the configuration of a remote terminal according to an embodiment. Here, the main elements related to the process of the user inputting handwritten input information to the captured image received from the autonomous mobile robot 100 and the process of the user interacting by voice input will be described, but the configuration of the remote terminal 300 may also include other elements, and other elements that contribute to the process of the user inputting handwritten input information may also be added.

[0080] The calculation unit 350 is, for example, a CPU, and executes a control program read from the memory 380 to control the entire remote terminal 300 and perform various calculation processes. The display panel 341 is, for example, a liquid crystal panel, and displays, for example, captured images sent from the autonomous mobile robot 100. The display panel 341 may also display, on a chat screen, voice input of a speech sentence entered by the user and voice output of a response speech sentence sent from the autonomous mobile robot 100.

[0081] The input unit 342 includes a touch panel superimposed on the display panel 341, push buttons provided on the periphery of the display panel, etc. The input unit 342 transfers handwritten input information input by the user by touching the touch panel and voice input of a spoken sentence to the calculation unit 350. Examples of the handwritten input information and voice input are as shown in, for example, FIG.

[0082] Memory 380 is a non-volatile storage medium, such as a solid state drive, that stores a control program for controlling remote terminal 300, as well as various parameter values, functions, lookup tables, and the like used for control and calculation.

[0083] The communication unit 390 is, for example, a wireless LAN unit, and performs wireless communication with the wireless router 700. The communication unit 390 receives captured images sent from the autonomous mobile robot 100 and passes them to the calculation unit 350. The communication unit 390 also cooperates with the calculation unit 350 to send handwritten input information to the autonomous mobile robot 100.

[0084] (Description of an operation method for an autonomous mobile robot according to an embodiment) Next, the overall processing of the autonomous mobile robot operation system 10 according to this embodiment will be described. Fig. 8 is a flowchart of the autonomous mobile robot operation method according to this embodiment. The flow on the left represents the processing flow of the autonomous mobile robot 100, and the flow on the right represents the processing flow of the remote terminal 300 and the voice input interface 360. In addition, the exchange of handwritten input information, captured images, and voice input via the system server 500 is indicated by dotted arrows.

[0085] The control unit 150 of the autonomous mobile robot 100 causes the stereo camera 131 to capture an image of the first environment in which the autonomous mobile robot 100 exists (step S11), and transmits the captured image to the remote terminal 300 via the communication unit 190 (step S12).

[0086] When the calculation unit 350 of the remote terminal 300 receives the captured image from the autonomous mobile robot 100 via the communication unit 390 , the calculation unit 350 displays the received captured image on the display panel 341 . Thereafter, the user uses the voice input interface 360 ​​to have a conversation with the autonomous mobile robot 100 (step S21). Specifically, when the user inputs a voice of an utterance sentence, the voice input interface 360 ​​transmits it to the autonomous mobile robot 100 via the communication unit. Furthermore, when the voice input interface 360 ​​receives a voice output of a response utterance sentence from the autonomous mobile robot 100 via the communication unit, it outputs the voice output to the headphones.

[0087] Furthermore, the calculation unit 350 of the remote terminal 300 transitions to a state in which it accepts input of handwritten input information for the captured image (step S31). When the user inputs handwritten input information for the captured image via the input unit 342, which is a touch panel (Yes in step S31), the calculation unit 350 transmits the handwritten input information to the autonomous mobile robot 100 via the communication unit 390 (step S32).

[0088] When the estimation unit 152 of the autonomous mobile robot 100 receives handwritten input information input by the user to the captured image from the remote terminal 300, it estimates an operated object that is requested to be operated by the hand 124 based on the handwritten input information and the dialogue history of the voice input, and estimates the manner of action that the hand 124 is requested to take on the estimated operated object (step S13). At this time, with regard to the operated object, the estimation unit 152 acquires information on the operable parts shown in the captured image to which the handwritten input information was input from the recognition unit 151, and estimates the operated object from among the operable parts based on the handwritten input information and the dialogue history of the voice input. In addition, the estimation unit 152 analyzes the content of the handwritten input information and the dialogue history of the voice input, and performs the above estimation while confirming the analyzed content to the voice input interface 360 ​​using voice output.

[0089] Thereafter, the control unit 150 of the autonomous mobile robot 100 generates a trajectory for the hand 124 to realize the requested action on the operated object (step S14). After generating the trajectory for the hand 124, the control unit 150 controls the upper body drive unit 146 according to the generated trajectory, and the hand 124 performs an action on the operated object (step S15).

[0090] In step S13, the estimation unit 152 may determine whether or not there is an additional action that the autonomous mobile robot 100 is requested to perform, based on the dialogue history of the voice input, and if there is an additional action that the autonomous mobile robot 100 is requested to perform, estimate how that action will be performed. This estimation may be performed by analyzing the content of the dialogue history of the voice input, and confirming the analyzed content to the voice input interface 360 ​​using voice output.

[0091] If the estimation unit 152 estimates the manner of an additional operation that is required of the autonomous mobile robot 100, the control unit 150 causes the autonomous mobile robot 100 to execute the additional operation that is required before or after steps S14 and S15. If the execution of such an operation requires the autonomous mobile robot 100 to move, the control unit 150 generates a path along which the autonomous mobile robot 100 should move. The control unit 150 then transmits a drive signal according to the generated path to the cart drive unit 145. The cart drive unit 145 moves the autonomous mobile robot 100 according to the drive signal.

[0092] As described above, according to this embodiment, the estimation unit 152 estimates an object to be operated that is requested to be operated by the hand 124, based on handwritten input information entered by the user into an image of the environment in which the autonomous mobile robot 100 exists, and the dialogue history of voice input, and also estimates the manner in which the hand 124 is requested to operate the estimated object to be operated.

[0093] This allows the user to remotely control the autonomous mobile robot 100 to execute an action without having to remember a preset instruction figure and input it by hand. Therefore, it is possible to realize an autonomous mobile robot operation system 10 that allows for more intuitive operation.

[0094] Furthermore, according to this embodiment, the estimation unit 152 may analyze the contents of the handwritten input information input to the captured image and the contents of the dialogue history of the voice input, and may confirm the analyzed contents to the voice input interface 360 ​​(user) using voice output.

[0095] This allows communication with the user regarding operation of the robot while confirming the user's intentions through voice, thereby realizing an autonomous mobile robot operation system 10 that allows intuitive operation that better reflects the user's intentions.

[0096] Furthermore, some or all of the processes of the autonomous mobile robot 100 and the system server 500 described above can be implemented as a computer program. Such a program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible recording media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be provided to a computer by various types of temporary computer-readable media. Examples of temporary computer-readable media include electrical signals, optical signals, and electromagnetic waves. The temporary computer-readable media can provide the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0097] The present invention is not limited to the above-described embodiment, and can be modified as appropriate within the scope of the invention. For example, in the above embodiment, the display screen 310 displayed on the display panel 341 of the remote terminal 300 is a screen on which a captured image 311 is arranged, as shown in FIG. 2, but is not limited to this. The display screen 310 may be a screen on which a chat screen is arranged superimposed on a captured image, for example. FIG. 9 is a diagram showing an example of a display screen displayed on the display panel of the remote terminal according to the embodiment. For example, FIG. 9 is a diagram showing an example of the display screen 310 on which a chat screen 312 is arranged superimposed on a captured image 311.

[0098] In the above embodiment, the estimation unit 152 confirms the analysis of the handwritten input information input to the captured image to the voice input interface 360 ​​(user) by using voice output. At this time, the operation target analyzed from the handwritten input information may be confirmed to the remote terminal 300 (user) by cutting out an image of the operation target from the captured image and displaying it on the chat screen. To confirm to the user that the operation target analyzed from the handwritten input information 931 is the smartphone 403, the estimation unit 152 transmits an image of the smartphone 403 cut out from the captured image 311 together with a voice output of a response utterance such as "Got it. Is it this smartphone?" to the remote terminal 300, and causes them to be displayed on the display panel 341.

[0099] In the above embodiment, an example in which one piece of handwritten input information is input to a captured image has been described, but the present invention is not limited to this. A plurality of pieces of handwritten input information may be input to a captured image. When a plurality of pieces of handwritten input information are input to a captured image, the estimation unit 152 analyzes each of the plurality of pieces of handwritten input information and estimates the operation target and the manner of operation while confirming the analyzed content with the voice input interface 360 ​​(user) using voice output. In this case, the estimation unit 152 may estimate the order of the operations as the order in which the handwritten input information corresponding to the operations was input. Alternatively, the estimation unit 152 may estimate the order of the operations while confirming with the voice input interface 360 ​​(user) using voice output.

[0100] In the above embodiment, the recognition unit 151 and the estimation unit 152 are provided in the autonomous mobile robot 100, but this is not limiting. The functions of the recognition unit 151 and the estimation unit 152, excluding the function of interacting with the user of the voice input interface 360, may be provided in the remote terminal 300 or the system server 500.

[0101] In the above embodiment, the autonomous mobile robot 100 and the remote terminal 300 exchange captured images, handwritten input information, and voice input via the Internet 600 and the system server 500, but this is not limiting. The autonomous mobile robot 100 and the remote terminal 300 may exchange captured images, handwritten input information, and voice input via direct communication.

[0102] Furthermore, in the above embodiment, an imaging unit (stereo camera 131) provided in the autonomous mobile robot 100 is used, but this is not limited to this. The imaging unit may be any imaging unit provided at any location in the first environment in which the autonomous mobile robot 100 exists. Furthermore, the imaging unit is not limited to a stereo camera, and may be a monocular camera or the like.

[0103] In the above embodiment, the example has been described in which the operated object is the autonomous mobile robot 100 equipped with the hand 124 at the tip of the arm 123 as an end effector, but the present invention is not limited to this. The operated object may be anything that is equipped with an end effector and performs an action using the end effector. The end effector may also be a gripping unit other than a hand (for example, a suction unit, etc.).

[0104] When issuing instructions to an autonomous robot by combining handwritten input with voice input, it can be difficult to determine which combination of instructions corresponds to a single action instruction, or whether the instructions have already ended and it is okay to start moving. Therefore, an input corresponding to a division between handwritten and voice instructions may be further received. For example, after receiving multiple handwritten and voice inputs, such as gripping method and trajectory, an input corresponding to an instruction division, such as "start," is received, and the robot will begin to move. The instructions received up to the time of receiving the division can be integrated and judged, and a clarified action can be instructed.

[0105] For example, when inputting by hand, you can fill in the areas where you want to avoid collisions, and then supplement that information with voice input so that the autonomous robot can understand the meaning.

[0106] For example, when you speak a shape, a sketch of that shape will be displayed, and you can manipulate the sketch to specify the robot's movements.

[0107] For example, adverbial manipulation of the robot's movements can be input using voice input such as "move slowly" or handwriting input such as drawing a line slowly.

[0108] For example, after drawing a trajectory and making the autonomous robot move, the movements of the autonomous robot, such as starting, stopping, changing speed, changing direction, and changing trajectory, can be controlled in real time by voice input and handwriting input. Starting and stopping movements can also be done by handwriting input instead of voice input.

[0109] For example, location information specified by handwriting and voice input may be converted into text and saved in chat format. This allows instructions for an autonomous robot to be generated. Missing information can be supplemented by further voice and handwriting input.

[0110] For example, you can specify something that is not displayed in the captured image by voice input. Also, you can use handwriting input to supplement something that is displayed in the captured image but cannot be clearly specified. For example, you can use handwriting input to supplement something that does not have a target, such as the right side of the center of a desk.

[0111] For example, the autonomous robot can provide visual and auditory feedback regarding the instructions given, and the user can confirm this to ensure clear instructions.

[0112] For example, when approaching a fork in the road such as an intersection or a dead end, you can select left or right by handwriting input or button operation and enter the direction you want to go.

[0113] For example, it is possible to provide the user with the part of the content that has already been instructed that the user understands, and to present further necessary instructions to the user. [Explanation of symbols]

[0114] 10 Autonomous mobile robot operation system, 100 Autonomous mobile robot, 110 Cart unit, 111 Drive wheel, 112 Caster, 120 Main body unit, 121 Torso unit, 122 Head unit, 123 Arm, 124 Hand, 131 Stereo camera, 133 Laser scanner, 135 Hand camera, 141 Display panel, 145 Cart drive unit, 146 Upper body drive unit, 150 Control unit, 151 Recognition unit, 152 Estimation unit, 180 Memory, 181 Trained model, 182 Speech DB, 183 Map DB, 190 Communication unit, 300 Remote terminal, 310 Display screen, 311 Captured image, 312 Chat screen, 341 Display panel, 342 Input unit, 350 Calculation unit, 360 Voice input interface, 380 Memory, 390 Communication unit, 400 Table, 401 Cup, 402 Calculator, 403 Smartphone, 404 Paper, 500 System Server, 600 Internet, 700 Wireless Router, 801-804 Operational Area, 901 Image (User), 902 Image (Robot), 911-913, 931 Handwritten Input Information

Claims

1. an autonomous robot that captures images of its surroundings and operates an object to be operated; a handwriting input interface that displays an image captured by the autonomous robot and accepts handwriting input on the displayed image; a voice input interface that accepts voice input for the target to be operated, An autonomous mobile robot operation system that operates the autonomous mobile robot so as to operate the operated object in accordance with instructions from the handwritten input and the voice input.

2. The autonomous mobile robot operation system according to claim 1, wherein the voice input interface inputs voice, uses a machine-learned voice recognition unit that inputs voice, recognizes the voice, and outputs an action, and inputs the action of the autonomous mobile robot through dialogue.

3. The autonomous mobile robot operation system according to claim 1 , wherein the voice input interface is used to give instructions to the operated object that is not displayed in the image.

4. 2. The autonomous mobile robot operation system according to claim 1, wherein the handwriting input interface inputs the image and inputs the operated object using an object estimation unit that has been trained by machine learning to estimate and output an object in the image.

5. The autonomous mobile robot operation system according to claim 1 , wherein the trajectory of the autonomous mobile robot is input using the handwriting input interface.

6. The autonomous mobile robot operation system according to claim 1 , wherein the handwriting input interface is used to input adverbial operations of actions.

7. The autonomous mobile robot operation system according to claim 1 , wherein the operation of the autonomous mobile robot is grasping, cutting, moving, screwing, or welding.

8. An autonomous robot operation method for operating an autonomous robot in such a way that the autonomous robot operates an object to be operated according to handwritten and voice input instructions.

9. A program that causes an information processing device to operate an autonomous robot in the same way as operating an object to be operated according to handwritten and voice input instructions.

Citation Information

Patent Citations

  • Remote operation system and remote operation method

    JP2021094604A