Teleoperation system, teleoperation method, and computer-readable storage medium
By photographing and recognizing graspable objects and movable areas, the system automatically identifies the type of transport path, solving the problem of insufficient path differentiation in the robot's remote operating system and improving operational accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2023-08-31
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, when the robot remote operating system contains delivery path information in the user's handwritten information, it cannot accurately distinguish between the robot's movement path and the end effector's movement path, resulting in insufficient operational accuracy.
The camera unit captures images of the environment, identifies graspable objects and movable areas, uses the recognition unit to identify and infer the graspable object and the content of the action, and the inference unit automatically distinguishes whether the transport path is the robot's movement path or the end effector's movement path, and displays the detour path information in the operation terminal.
This technology improves the accuracy of remote robot operation while allowing for intuitive user control, enabling accurate execution of grasping and transport tasks.
Smart Images

Figure CN117644528B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a remote operating system, a remote operation method, and a control program. Background Technology
[0002] Technology is known for remotely operating and causing an object equipped with an end effector (e.g., a robot with a gripping part (e.g., a gripper or suction part) at the tip of its arm) to perform gripping actions. For example, Japanese Patent Application Publication No. 2021-094605 discloses a system that determines the robot's control method based on handwritten information received by a user on an operating terminal displaying images taken of the robot's surroundings, and then remotely operates the robot. This system enables remote operation of the robot based on the user's intuitive handwritten instructions. Summary of the Invention
[0003] However, in the system disclosed in Japanese Patent Application Publication No. 2021-094605, when the handwritten information contains information about the transport path of the object being grasped, there is a problem that the robot (the object being manipulated) cannot perform high-precision processing because it does not determine whether the transport path represents the robot's movement path (driving path) or the movement path of the end effector implemented by the robot's arm operation.
[0004] This disclosure was made to solve such a problem by providing a remote operating system, remote operation method, and control program that enables the operated object to perform high-precision processing while still allowing the user to perform intuitive operations.
[0005] The remote operating system disclosed herein is a remote operating system for remotely operating an object equipped with an end effector. The remote operating system includes: a camera unit that captures images of the environment in which the object is located; a recognition unit that identifies, based on the captured image of the environment, a graspable object that the end effector can grasp, a movable area of the object, and a non-movable area of the object; an operation terminal that displays the captured image and accepts input of handwritten information related to the captured image; and an inference unit that, based on the graspable object and the handwritten information, infers from the graspable object an object for which the end effector requests grasping, and infers the action content requested from the object for grasping. If the handwritten information includes information about the transport path of the object, the inference unit infers the transport path located within the movable area of the object as the movement path of the object, and infers the transport path located within the non-movable area of the object as the movement path of the end effector that grasped the object. This remote operating system allows users to perform handwriting input without having to recall pre-set instruction graphics. Instead, it enables more intuitive handwriting input to execute desired processes, such as transporting the grasped object. Furthermore, since the handwritten information includes information about the transport path of the grasped object, the system automatically distinguishes whether the transport path represents the movement path (driving path) of the grasped object or the movement path of the end effector. Therefore, the remote operating system can enable the grasped object to perform highly accurate processing while still allowing for intuitive user operation.
[0006] Alternatively, when the handwritten information includes information about the transport path of the object being grasped, if there is an immovable area sandwiched between two movable areas on the transport path, the inference unit infers a detour path for the object being grasped, where there is no immovable area sandwiched between two movable areas. The remote operating system also includes an output unit that outputs information about the detour path inferred by the inference unit.
[0007] Alternatively, the output unit can cause the operating terminal to display information about the detour path.
[0008] Alternatively, the identification unit may identify objects other than the object to be grasped as immovable areas of the object to be manipulated.
[0009] Alternatively, the object being operated on can be a robot capable of autonomous movement.
[0010] Alternatively, if the handwritten information includes information about the transport path of the grasped object, the inference unit infers the transport path located within the movable area of the manipulated object as the movement path of the manipulated object, i.e., the robot, and infers the transport path located within the immovable area of the manipulated object as the movement path achieved by the end effector installed at the tip of the robot's arm through the operation of the arm.
[0011] Alternatively, the camera unit can be mounted on the object being operated on.
[0012] Alternatively, the camera unit can be positioned in a location different from the object being operated on within the environment.
[0013] Alternatively, the input of the handwritten information can be completed by the terminal through a predetermined handwriting end operation performed by the user.
[0014] Alternatively, the handwritten information may include a first image simulating the actions of the manipulated object toward the grasped object.
[0015] Alternatively, the inference unit may use a learned model, such as a neural network that has been trained through deep learning, to infer the action content requested for the grasped object based on the first image of the handwritten information.
[0016] The remote operation method disclosed herein is a remote operation method implemented by a remote operating system for remotely operating an object equipped with an end effector. In this remote operation method, the environment in which the object is located is photographed; based on the photographed image of the environment, the graspable objects that the end effector can hold, the movable areas of the object, and the immovable areas of the object are identified; in an operating terminal displaying the photographed image, input of handwritten information related to the displayed photographed image is accepted; and based on the graspable objects and the handwritten information, a request is inferred from the graspable objects... The end effector grasps the object and infers the requested action content for the grasped object. In inferring the requested action content for the grasped object, if the handwritten information includes information about the object's transport path, the transport path within the movable area of the object is inferred as the object's movement path, and the transport path within the immovable area of the object is inferred as the end effector's movement path. This remote operation method allows users to perform desired processes, such as transporting the grasped object, without recalling pre-set instruction graphics during handwriting input, through more intuitive handwriting input. Furthermore, since this remote operation method automatically distinguishes whether the transport path represents the object's movement path (driving path) or the end effector's movement path when the handwritten information includes the object's transport path information, it enables the object to perform highly accurate processing while still allowing for intuitive user operation.
[0017] The control program disclosed herein is a control program that enables a computer to perform remote operation processing. This remote operation processing is implemented by a remote operating system that remotely operates an object equipped with an end effector. The control program enables the computer to perform the following processes: capturing an image of the environment in which the object is located; identifying, based on the captured image, a graspable object that the end effector can hold, the movable area of the object, and the immovable area of the object; accepting handwritten information input based on the captured image in an operating terminal displaying the captured image; and processing based on the graspable object and the handwritten information. The system infers the object to be grasped by the end effector from the graspable object, and infers the processing of the action content requested by the manipulated object for the grasped object. In inferring the action content requested by the manipulated object for the grasped object, if the handwritten information includes information about the transport path of the grasped object, the transport path within the movable area of the manipulated object is inferred as the movement path of the manipulated object, and the transport path within the immovable area of the manipulated object is inferred as the movement path of the end effector that grasped the grasped object. This control program allows the user to perform desired processing, such as transporting the grasped object, by performing more intuitive handwriting input instead of recalling pre-set instruction graphics while performing handwriting input. Furthermore, in this control program, since the handwritten information contains information about the transport path of the object being grasped, the program automatically identifies whether the transport path represents the movement path (driving path) of the object being manipulated or the movement path of the end effector. Therefore, the control program can enable the object being manipulated to perform high-precision processing while still allowing the user to perform intuitive operations.
[0018] According to this disclosure, a remote operating system, remote operation method, and control program can be provided that enable the operated object to perform high-precision processing while still allowing the user to perform intuitive operations.
[0019] The above and other objects, features and advantages of this disclosure will become more fully understood from the detailed description and accompanying drawings given below, which are given by way of illustration only and should not be considered as limiting the disclosure. Attached Figure Description
[0020] Figure 1 This is a conceptual diagram illustrating an example of the overall environment using the remote operating system described in Implementation 1.
[0021] Figure 2 A diagram illustrating an example of the first environment in which the robot is located.
[0022] Figure 3 A diagram illustrating an example of handwritten information.
[0023] Figure 4 A diagram illustrating an example of handwritten information.
[0024] Figure 5 A diagram illustrating an example of handwritten information.
[0025] Figure 6 A diagram illustrating an example of handwritten information.
[0026] Figure 7 A three-dimensional view showing the external structure of a robot.
[0027] Figure 8 A block diagram illustrating an example of the functional block structure of a robot.
[0028] Figure 9 A diagram illustrating an example of an image captured by a robot.
[0029] Figure 10 A diagram illustrating an example of the graspable region output by the first learned model.
[0030] Figure 11 A diagram illustrating an example of handwritten information.
[0031] Figure 12 A diagram illustrating an example of teacher data used in the second completed learning model.
[0032] Figure 13 A diagram illustrating an example of teacher data used in the second completed learning model.
[0033] Figure 14 A block diagram illustrating an example of the functional block structure of a remote terminal.
[0034] Figure 15 This is a flowchart illustrating an example of the overall processing flow of the remote operating system involved in Implementation 1.
[0035] Figure 16 A diagram illustrating an example of the first environment in which the robot is located.
[0036] Figure 17A This diagram illustrates the process of inputting handwritten information into an image that is displayed on a remote terminal.
[0037] Figure 17B This diagram illustrates the process of inputting handwritten information into an image that is displayed on a remote terminal.
[0038] Figure 17CThis diagram illustrates the process of inputting handwritten information into an image that is displayed on a remote terminal.
[0039] Figure 17D This diagram illustrates the process of inputting handwritten information into an image that is displayed on a remote terminal.
[0040] Figure 18 A diagram illustrating an example of handwritten information ultimately entered in response to an image captured and displayed on a remote terminal.
[0041] Figure 19 To indicate Figure 15 A flowchart illustrating the detailed process of step S15.
[0042] Figure 20 This diagram illustrates an example of the robot's movable and immovable areas as identified by the recognition unit.
[0043] Figure 21 This diagram illustrates an example of the robot's movable and immovable areas as identified by the recognition unit.
[0044] Figure 22 This diagram illustrates a method for inferring the robot's actions using the movable and immovable areas identified by the recognition unit.
[0045] Figure 23 A diagram illustrating the circuitous path of a grasped object transported by a robot.
[0046] Figure 24 A diagram illustrating an example of handwritten information.
[0047] Figure 25 A diagram illustrating an example of inputting multiple handwritten messages. Detailed Implementation
[0048] The present invention will be described below through embodiments, but the invention involved in the technical solution is not limited to the following embodiments. Furthermore, the structures described in the embodiments are not all necessary technical means for solving the problem. In addition, in the following embodiments, a robotic arm with an arm tip is used as an example of a robot as an end effector, but the object of operation is not limited to this.
[0049] Implementation Method 1
[0050] Figure 1This is a conceptual diagram illustrating an example of the overall environment of the remote operating system 10 according to Embodiment 1. A robot 100 performing various actions in a first environment is remotely operated via a system server 500 connected to a network 600 by a remote operator, i.e., a user, residing in a second environment separate from the first environment, operating the remote terminal 300 (operation terminal).
[0051] In the first environment, robot 100 is connected to network 600 via wireless router 700. Furthermore, in the second environment, remote terminal 300 is connected to network 600 via wireless router 700. System server 500 is connected to network 600. Based on the operations performed by the user via remote terminal 300, robot 100 performs actions such as grasping objects and transporting the grasped objects from their current location to their destination, all controlled by robotic arm 124.
[0052] Furthermore, in this embodiment, the grasping action performed by the robot arm 124 is not limited to grasping (holding) the object, but also includes the following actions.
[0053] • The action of grasping and lifting the object being grasped.
[0054] • When grasping the handle of a cabinet door or drawer, the action of grasping the handle and opening or closing the door or drawer.
[0055] • When the object being grasped is a door handle, the action of grasping the door handle and opening or closing the door.
[0056] Furthermore, in this embodiment, the transport operation from the current location of the grasped object to the destination includes the operation of transporting the grasped object by moving the robot 100, and the operation of transporting the grasped object held by the robotic arm 124 by arm operation, for example, without moving the robot 100.
[0057] The robot 100 uses a stereo camera 131 (image capture unit) to capture images of the first environment in which the robot 100 exists, and transmits the captured images to a remote terminal 300 via a network 600. In addition, the robot 100 identifies graspable objects that the robotic arm 124 can grasp based on the captured images.
[0058] Figure 2 This is a diagram illustrating an example of the first environment in which robot 100 exists. Figure 2In the example, in the first environment, there are a table 400, cabinets 410, 420, 430, and a door 440. In addition, the graspable objects in the first environment are objects 401 and 402 placed on the table 400, handles 411 of cabinet 410, handles 421 and 422 of cabinet 420, handles 431 and 432 of cabinet 430, and door handle 441 of door 440.
[0059] The remote terminal 300 is, for example, a tablet terminal, having a display panel 341 with an overlaid touch panel. On the display panel 341, images received from the robot 100 are displayed, allowing the user to indirectly visually confirm the first environment in which the robot 100 is located. Furthermore, the user can input handwritten information (a first image) simulating the robot 100's actions towards a grasped object, based on the images displayed on the display panel 341.
[0060] As a method for inputting handwritten information, there are methods such as using a user's finger or stylus to touch a corresponding part of a captured image on a touch panel that is superimposed on the display panel 341, but it is not limited to this. Figures 3-6 A diagram illustrating an example of handwritten information input for a captured image 310. Figure 3 The example shows a simulation of grabbing a cuboid object 401 placed on a table 400 from above and writing 901. Figure 4 The example shows a simulation of handwritten information 902 being grasped laterally from a cylindrical object 402 placed on a table 400. Figure 5 The example shows a handwritten message 903 simulating grasping the handle 411 of cabinet 410 and opening the door. Figure 6 The example illustrates handwritten information 904 simulating grasping a cylindrical object 402 placed on a table 400 laterally and moving it to another location on the table 400. Furthermore, handwritten information 904 consists of handwritten information 904a simulating grasping the cylindrical object 402 laterally, handwritten information 904b simulating the transport path of the object 402, and handwritten information 904c simulating the destination of the object 402. For example... Figures 3-6 As shown, the image of the handwritten information can be an image composed only of lines or other shapes, or an image composed of a combination of lines or other shapes and text. The handwritten information entered by the user in response to the captured image is sent to the robot 100 via network 600.
[0061] Furthermore, the remote terminal 300 is not limited to a tablet terminal, and the handwritten information input to the remote terminal 300 is not limited to input through touching a touch panel using a user's finger or stylus. For example, the remote terminal 300 could be configured as VR (Virtual Reality) goggles, and the handwritten information input to the remote terminal 300 could be input by a user wearing VR goggles operating a VR controller to input images displayed in the VR goggles.
[0062] Based on the graspable objects identified from the captured image and the handwritten information input by the user in relation to the captured image, the robot 100 infers the graspable object requested to be grasped by the robotic arm 124 from the graspable objects, and infers the action content requested by the robot 100 for the graspable object. Furthermore, the action content requested by the robot 100 for the graspable object includes, for example, the grasping action performed by the robotic arm 124, and the transport action from the current location of the graspable object to its destination.
[0063] Figure 7 This is a perspective view showing the external structure of robot 100. Robot 100 is generally divided into a carriage section 110 and a main body section 120. The carriage section 110, within a cylindrical housing, supports two drive wheels 111, each grounded to the travel surface, and a small caster 112. The two drive wheels 111 are arranged such that their rotation axes coincide. Each drive wheel 111 is independently driven to rotate by a motor (not shown). The small caster 112 is a driven wheel, and its steering axis, extending vertically from the carriage section 110, is separate from the wheel's rotation axis, providing axle support for the wheel, and it follows the direction of movement of the carriage section 110.
[0064] The trolley section 110 has a laser scanner 133 at the periphery of its upper surface. The laser scanner 133 scans a fixed range in the horizontal plane at each step angle and outputs whether there are obstacles in each direction. Furthermore, if there are obstacles, the laser scanner 133 outputs the distance to the obstacle.
[0065] The main body 120 mainly includes a torso 121 mounted on the upper surface of the carriage 110, a head 122 mounted on the upper surface of the torso 121, an arm 123 supported on the side of the torso 121, and a robotic hand 124 provided at the top of the arm 123. The arm 123 and the robotic hand 124 are driven by a motor (not shown) to grasp an object. The torso 121 can rotate about a vertical axis relative to the carriage 110 by the driving force of a motor (not shown).
[0066] The head 122 mainly includes a stereo camera 131 and a display panel 141. The stereo camera 131 has a structure in which two camera units with the same field of view are arranged separately from each other, and outputs the shooting signals captured by each camera unit.
[0067] Display panel 141 is, for example, an LCD panel, and displays the face of a set character through animation, or displays information related to robot 100 through text and icons. If the face of a character is displayed on display panel 141, it can leave people around with the impression that display panel 141 is a simulated face.
[0068] The head 122 can rotate about a vertical axis relative to the torso 121 by the driving force of a motor (not shown). Therefore, the stereo camera 131 can capture images from any direction, and the display panel 141 can display content facing any direction.
[0069] Figure 8 This is a block diagram illustrating an example of the functional block structure of robot 100. Here, while the main elements related to the inference of the object to be grasped and the inference of the action content performed by robot 100 on the object to be grasped are described, other elements are also present in the structure of robot 100. Furthermore, other elements that contribute to the inference of the object to be grasped and the inference of the action content performed by robot 100 on the object to be grasped can be added.
[0070] The control unit 150 is, for example, a CPU, and is stored in, for example, a controller unit included in the body 121. The trolley drive unit 145 includes a drive wheel 111 and a drive circuit and motor for driving the drive wheel 111. The control unit 150 executes rotation control of the drive wheel by sending drive signals to the trolley drive unit 145. In addition, the control unit 150 receives feedback signals from the trolley drive unit 145 such as encoders to determine the direction and speed of movement of the trolley section 110.
[0071] The upper body drive unit 146 includes an arm 123 and a robotic hand 124, a torso 121 and a head 122, and drive circuits and motors for driving them. The control unit 150 sends drive signals to the upper body drive unit 146 to realize grasping actions and postures. In addition, the control unit 150 receives feedback signals from encoders and the like from the upper body drive unit 146 to monitor the position and movement speed of the arm 123 and the robotic hand 124, as well as the orientation and rotation speed of the torso 121 and the head 122.
[0072] The display panel 141 receives and displays the image signals generated by the control unit 150. Furthermore, as described above, the control unit 150 generates image signals of characters, etc., and displays them on the display panel 141.
[0073] Stereo camera 131, upon request from control unit 150, captures images of the first environment in which robot 100 is located and submits the captured images to control unit 150. Control unit 150 uses the captured images to perform image processing or converts the captured images into captured images according to a predetermined format. Laser scanner 133, upon request from control unit 150, detects whether obstacles exist in the direction of movement and submits a detection signal as its detection result to control unit 150.
[0074] The robotic arm camera 135, for example, is a distance image sensor used to identify the distance, shape, and orientation of the object being grasped. The robotic arm camera 135 includes an imaging element composed of pixels arranged in a two-dimensional shape that photoelectrically converts an optical image incident from the object space, thereby outputting the distance to the subject to each pixel to the control unit 150. Specifically, the robotic arm camera 135 includes an illumination unit that illuminates patterned light into the object space, receives the reflected light through the imaging element, and outputs the distance to the subject captured by each pixel based on the deformation and size of the pattern in the image. Furthermore, the control unit 150 uses a stereo camera 131 to grasp a wider range of surrounding conditions and the robotic arm camera 135 to grasp the situation near the object being grasped.
[0075] The memory 180 is a non-volatile storage medium, such as a solid-state drive. Besides storing the control program for controlling the robot 100, the memory 180 also stores various parameter values, functions, lookup tables, etc., used for control and calculation. Specifically, the memory 180 stores a first learned model 181, a second learned model 182, and a third learned model 183. The first learned model 181 takes a captured image as input and outputs a model of the graspable object captured in the image. The second learned model 182 takes an image of handwritten information as input and outputs a model representing the meaning of the robot 100's actions simulated by the handwritten information. The third learned model 183 takes a captured image as input and outputs a model of the robot 100's movable and immovable areas captured in the image.
[0076] The communication unit 190 is, for example, a wireless LAN unit, which enables wireless communication with the wireless router 700. The communication unit 190 receives handwritten information sent from the remote terminal 300 and submits it to the control unit 150. In addition, the communication unit 190, under the control of the control unit 150, sends images captured by the stereo camera 131 to the remote terminal 300.
[0077] The control unit 150 executes the control program read from the memory 180, thereby performing overall control and various calculations of the robot 100. Furthermore, the control unit 150 also functions as a functional execution unit for performing various calculations and controls related to the robot. As such a functional execution unit, the control unit 150 includes a recognition unit 151, an inference unit 152, and an image adjustment unit 153.
[0078] (Details of the Identification Unit 151)
[0079] The identification unit 151 mainly extracts the graspable area that the robotic arm 124 can grasp from the image captured by the stereo camera 131 and identifies it as a graspable object. This will be explained in detail below.
[0080] First, the recognition unit 151 uses the first learned model 181 read from the memory 180 to extract the graspable area that the robotic arm 124 can grasp in the captured image taken by the stereo camera 131, and identifies it as a graspable object.
[0081] Figure 9 This diagram illustrates an example of a first environmental image 310 captured by robot 100 using stereo camera 131. Figure 9 The captured image 310 shows a cabinet 410 with handle 411 and a cabinet 420 with handles 421 and 422. The recognition unit 151 provides this captured image 310 as an input image to the first learned model 181.
[0082] Figure 10 To indicate in the Figure 9 This diagram illustrates an example of the grippable regions output by the first learned model 181 when the captured image 310 is used as the input image. Specifically, the region surrounding handle 411 is identified as grippable region 801, the region surrounding handle 421 as grippable region 802, and the region surrounding handle 422 as grippable region 803, respectively. Therefore, the recognition unit 151 identifies handles 411, 421, and 422, which are surrounded by grippable regions 801, 802, and 803, as grippable objects.
[0083] The first learned model 181 is a neural network that learns using teacher data, which is a combination of images of objects that the robotic arm 124 can grasp and the correct value of which region in the image is the object. In this case, the first learned model 181 learns using teacher data including information related to the distance and orientation of the object in the image, thus becoming a learned model capable not only of extracting the object from the captured image but also of extracting information related to the distance and orientation of the object. Preferably, the first learned model 181 is a neural network that learns using deep learning. Furthermore, the first learned model 181 can be supplemented with teacher data at any time for further learning.
[0084] Furthermore, the identification unit 151 not only identifies graspable objects, but also identifies the movable and immovable areas of the robot 100. The process of identifying the movable and immovable areas of the robot 100 by the identification unit 151 is basically the same as that of identifying graspable objects.
[0085] In other words, the recognition unit 151 uses the third learned model 183 read from the memory 180 to extract and recognize the movable and immovable regions of the robot 100 captured in the image taken by the stereo camera 131. Here, the third learned model 183 is a neural network that has been learned using teacher data that combines images of the movable regions of the robot 100 with the correct values of which regions in the images are the movable regions of the robot 100, and also uses teacher data that combines images of the immovable regions of the robot 100 with the correct values of which regions in the images are the immovable regions of the robot 100.
[0086] (Details of Inference Section 152)
[0087] The inference unit 152 infers the object to be grasped by the robot arm 124 from the graspable objects identified by the recognition unit 151, and infers the action content requested by the robot 100 for the object to be grasped. This will be explained in detail below.
[0088] First, the inference unit 152 infers the object to be grasped by the robot arm 124 from the graspable objects identified by the recognition unit 151 from the captured image and the handwritten information input by the user for the captured image.
[0089] Figure 11 To indicate that the user is targeting in remote terminal 300 Figure 9An example of handwritten information entered from captured image 310. Figure 11 In the example, handwritten information 905 is input at the position of handle 411 on the captured image 310. Therefore, the inference unit 152 infers that, among the handles 411, 421, and 422 identified by the recognition unit 151 as graspable objects, the object for which the robot arm 124 is requesting grasping is handle 411. Furthermore, the inference unit 152 can identify the input position of handwritten information 905 on the captured image 310 in any way. For example, if the remote terminal 300 sends position information indicating the input position of handwritten information 905 on the captured image 310 in the handwritten information 905, the inference unit 152 can identify the input position of handwritten information 905 based on this position information. Alternatively, if the remote terminal 300 sends the captured image 310 processed to have handwritten information 905 input, the inference unit 152 can identify the input position of handwritten information 905 based on the captured image 310.
[0090] Furthermore, the inference unit 152 uses the second learned model 182 read from the memory 180 to obtain the meaning of the action content of the robot 100 simulated by the handwritten information input by the user in response to the captured image, thereby inferring the action content of the robot 100 requested for grasping the object.
[0091] The second completed model 182 is a neural network that has learned by using teacher data that combines images as handwritten information with the meaning of the action content of the robot 100 simulated by the handwritten information. Figure 12 A diagram illustrating an example of teacher data used in the second completed learning model 182. Figure 12 An example is teacher data used to enable the second learned model 182 to learn from three images representing the grasping action of "grasping," four images representing the grasping action of "opening," and four images representing the transporting action of "transporting." Furthermore, with... Figure 12 Compared to the teacher data, the second learning model 182 can learn more detailed teacher data after completion. Figure 13 A diagram illustrating an example of teacher data used to enable the second learned model 182 to learn the grasping action of "grabbing" in greater detail. Figure 13 Examples include teacher data used to enable the second learned model 182 to learn from images representing grasping actions such as "grabbing from above," "grabbing horizontally," and "grabbing diagonally above." Alternatively, the second learned model 182 can be a neural network learned through deep learning. Furthermore, the second learned model 182 can be supplemented with teacher data at any time for further learning.
[0092] exist Figure 11 In the example, the inference unit 152 recognizes from the second learned model 182 that the handwritten information 904 means a grasping action such as "open". Therefore, the inference unit 152 infers that the action requested by the robot 100 for the grasping object is to grasp the handle 411 of the grasping object and open the cabinet door.
[0093] Furthermore, a more detailed inference method for the inference unit 152, which uses the movable and immovable areas of the robot 100 identified by the recognition unit 151, to infer the action content requested by the robot 100 for the object to be grasped, will be described below.
[0094] (Overview of image adjustment unit 153)
[0095] The image adjustment unit 153 follows the handwritten information input by the user regarding the captured image displayed on the display panel 341 of the remote terminal 300, changes the shooting direction and shooting range of the stereo camera 131, and displays the captured image taken by the stereo camera 131 on the display panel 341 of the remote terminal 300. Details regarding the image adjustment unit 153 will be described below.
[0096] As described above, in the control unit 150, the inference unit 152 can infer the object to be grasped by the manipulator 124 and the action requested by the robot 100 for the object to be grasped. Furthermore, the control unit 150 can obtain the distance and direction of the object to be grasped from the images captured by the stereo camera 131 based on the learning content of the first learned model 181. Additionally, the control unit 150 can obtain the distance and direction of the object to be grasped by image analysis of the captured images of the first environment, or by obtaining the distance and direction of the object to be grasped based on the detection results of other sensors. Furthermore, the control unit 150 can also detect whether there are obstacles in the direction of movement of the robot 100 based on the detection signal from the laser scanner 133.
[0097] Therefore, the control unit 150 generates a path for the robot 100 to move from its current position to the vicinity of the object to be grasped, based on the distance and direction of the object to be grasped, as well as the presence or absence of obstacles. The control unit 150 then sends a drive signal corresponding to the generated path to the trolley drive unit 145. The trolley drive unit 145 then moves the robot 100 to the vicinity of the object to be grasped based on this drive signal.
[0098] When the robot 100 moves near the object to be grasped, the control unit 150 prepares to initiate actions requested by the robot 100 for grasping the object. Specifically, first, the control unit 150 drives the arm 123 to position the object to be grasped so that the robotic arm camera 135 can observe its location. Next, the control unit 150 causes the robotic arm camera 135 to photograph the object to be grasped, thereby recognizing the state of the object to be grasped.
[0099] Then, based on the state of the object to be grasped and the action content requested by the robot 100 for the object to be grasped, the control unit 150 determines the detailed content of the actions of each part of the robot 100 to implement the action content requested by the robot 100 for the object to be grasped.
[0100] For example, the control unit 150 generates a track for the robot arm 124 to perform the grasping action of the robot arm 124 on the object being grasped. At this time, the control unit 150 generates the track of the robot arm 124 in a manner that satisfies predetermined grasping conditions. The predetermined grasping conditions include conditions when the robot arm 124 grasps the object, and track conditions before the robot arm 124 grasps the object. For example, the conditions when the robot arm 124 grasps the object include, for example, preventing the arm 123 from excessively extending when grasping the object. Furthermore, the track conditions before the robot arm 124 grasps the object include, for example, in the case where the object is a drawer handle, conditions such as the robot arm 124 taking a straight track.
[0101] When the track of the robotic arm 124 is generated, the control unit 150 sends a drive signal corresponding to the generated track to the upper body drive unit 146. The robotic arm 124 then performs a grasping action on the object to be grasped based on the drive signal.
[0102] Alternatively, the control unit 150 generates the track of the manipulator 124 and the movement path (travel path) of the robot 100 for carrying out the object-grabbing action performed by the robot 100. In this case, the control unit 150 sends a drive signal corresponding to the generated track of the manipulator 124 to the upper body drive unit 146 and a drive signal corresponding to the generated movement path of the robot 100 to the trolley drive unit 145. The manipulator 124 and the robot 100 perform the object-grabbing action according to these drive signals.
[0103] Figure 14This is a block diagram illustrating a functional block structure example of the remote terminal 300. Here, the main elements related to the processing of inputting handwritten information in response to a captured image received from the robot 100 are described. However, other elements also exist as part of the configuration of the remote terminal 300, and additional elements that facilitate the processing of inputting handwritten information can be added.
[0104] The arithmetic unit 350, for example, is a CPU, which executes a control program read from the memory 380 to perform overall control and various arithmetic operations on the remote terminal 300. The display panel 341, for example, is an LCD panel, which displays images captured, for example, from the robot 100.
[0105] The input unit 342 includes a touch panel superimposed on the display panel 141, buttons on the periphery of the display panel 141, etc. The input unit 342 submits handwritten information—images simulating the actions of the robot 100 towards an object being grasped—inputted by the user using a finger or stylus to touch the touch panel, to the processing unit 350. Examples of handwritten information include... Figures 3-6 As shown.
[0106] The memory 380 is a non-volatile storage medium, such as a solid-state drive. In addition to storing the control program for controlling the remote terminal 300, the memory 380 also stores various parameter values, functions, lookup tables, etc., used for control and calculation.
[0107] The communication unit 390 is, for example, a wireless LAN unit, which enables wireless communication with the wireless router 700. The communication unit 390 receives images captured from the robot 100 and submits them to the computing unit 350. In addition, the communication unit 390 cooperates with the computing unit 350 to send handwritten information to the robot 100.
[0108] Next, use Figures 15-18 The overall processing of the remote operating system 10 involved in this embodiment will be explained.
[0109] Figure 15 This is a flowchart illustrating an example of the overall processing flow of the remote operating system 10 involved in this embodiment. Figure 15 In the diagram, the left side represents the processing flow of robot 100, and the right side represents the processing flow of remote terminal 300. Additionally, the dotted arrows indicate the handover of handwritten information and captured images between robot 100 and remote terminal 300 via system server 500.
[0110] Figure 16 This is a diagram illustrating an example of the first environment in which robot 100 is located. Additionally, in... Figure 16In the example, it is assumed that robot 100 (not shown) is configured in the first environment in an area opposite to the paper. Therefore, in Figure 16 In the example, robot 100 can use stereo camera 131 to photograph a first predetermined area A1 or a second predetermined area A2. The following describes a scenario where a user remotely operates robot 100 to transport an object 402 positioned on table 400 to cabinet 420.
[0111] also, Figures 17A to 17D This diagram illustrates the process of inputting handwritten information onto an image displayed on a remote terminal 300. Figure 18 A diagram illustrating an example of handwritten information ultimately entered in response to an image captured and displayed on a remote terminal 300.
[0112] In robot 100, control unit 150 uses stereo camera 131 to capture a first predetermined area A1 of the first environment in which robot 100 is located (step S11), and sends the captured image to remote terminal 300 via communication unit 190 (step S12).
[0113] In the remote terminal 300, when the arithmetic unit 350 receives a captured image of the first predetermined area A1 sent from the robot 100, it displays the received captured image on the display panel 341. Afterward, the arithmetic unit 350 switches to a state that accepts handwritten information input for the captured image (step S31). Then, the user begins to input handwritten information for the captured image via the input unit 342, which functions as a touch panel (step S31, "Yes").
[0114] Here, if a portion of the captured image displayed on the display panel 341 is designated as a predetermined outer periphery 311 of the captured image displayed on the display panel 341 ("Yes" in step S32), the arithmetic unit 350 sends the handwritten information to the robot 100 via the communication unit 390 (step S33).
[0115] exist Figure 17A In the example, in the captured image of the predetermined area A1 displayed on the display panel 341, a portion P1 of the transport path of the grasped object 402, as indicated in the handwritten information 906, is the predetermined outer peripheral area 311 of the captured image displayed on the display panel 341 ("Yes" in step S32). In this case, the arithmetic unit 350 sends the handwritten information 906 to the robot 100 via the communication unit 390 (step S33).
[0116] At this time, in robot 100, image adjustment unit 153 uses stereo camera 131 to change the shooting area from first predetermined area A1 to second predetermined area A2 and takes a picture (step S13). The captured image is then sent to remote terminal 300 via communication unit 190 (step S14). Here, image adjustment unit 153 adjusts the shooting direction and shooting range of stereo camera 131 so that the second predetermined area A2 includes at least the portion of the area specified in the handwritten information and the area adjacent to the first predetermined area A1. The second predetermined area A2 may include all or part of the first predetermined area A1.
[0117] In the remote terminal 300, when the computing unit 350 receives an image captured from the second predetermined area A2 sent from the robot 100, it displays the received image on the display panel 341. That is, the computing unit 350 switches the image displayed on the display panel 341 from the image captured from the first predetermined area A1 to the image captured from the second predetermined area A2.
[0118] exist Figure 17B In the example, the captured image displayed on the display panel 341 is switched from a captured image of the first predetermined area A1 to a captured image of the second predetermined area A2. Furthermore, the captured image of the second predetermined area A2 is a captured image of an area that includes at least a portion of the transport path of the grasped object 402 as indicated in the handwritten information 906, such as area P1, and an area adjacent to the first predetermined area A1.
[0119] In this way, even in areas with low visual clarity, such as the area at the end of the captured image specified in the handwritten information, the captured image displayed on the display panel 341 switches according to the area specified in the handwritten information, thus maintaining high visual clarity. As a result, the user can input a wide range of handwritten information, for example, spanning from the first predetermined area A1 to the second predetermined area A2.
[0120] exist Figure 17C In the example, handwritten information 906 is continuously input from the first predetermined area A1 to the second predetermined area A2.
[0121] Additionally, if the area specified in the handwritten information (e.g., the transport path of the grasped object indicated in the handwritten information) in the captured image displayed on the display panel 341 is not a predetermined outer periphery area 311 of the captured image displayed on the display panel 341 (No in step S32), the captured image displayed on the display panel 341 will not be switched.
[0122] Then, when the user wants to end the input of handwritten information on the captured image, a predetermined handwriting end operation is performed. To end the input of handwritten information on the captured image, the user can tap any part of the touch panel twice consecutively using a double-tap motion. Alternatively, the user can continuously tap the same part of the touch panel for a predetermined time or more. Alternatively, if an operation button is displayed on the display panel 341, the user can tap the area of the touch panel on the button that ends the handwriting input.
[0123] exist Figure 17D In the example, as a pre-defined handwriting end operation, the user double-tap any part of the touch panel twice to input handwritten information 907 for the captured image.
[0124] Therefore, the arithmetic unit 350 ends the acceptance of handwritten information input for the captured image ("Yes" in step S34), and... Figure 18 The handwritten information 906 shown is sent to the robot 100 via the communication unit 390 (step S35).
[0125] In robot 100, if the recognition unit 151 receives handwritten information input for a captured image sent from remote terminal 300, it recognizes a graspable object based on the captured image. The inference unit 152 infers the graspable object that the robot arm 124 requests to grasp from the graspable objects recognized by the recognition unit 151 and the handwritten information input for the captured image, and infers the action content requested by robot 100 for the graspable object (step S15).
[0126] exist Figure 18 In the example, the inference unit 152 infers from the multiple graspable objects identified by the recognition unit 151 that object 402 is the graspable object, and infers that the action requested by the robot 100 for the graspable object is to transport the graspable object, i.e., object 402, which is arranged on the table 400, to the cabinet 420.
[0127] Subsequently, in robot 100, control unit 150 controls trolley drive unit 145 to move robot 100 to the vicinity of the object to be grasped, and at the time point when robot 100 moves to the vicinity of the object to be grasped, determines the details of the actions of various parts of robot 100 to implement the actions requested by robot 100 for grasping the object.
[0128] For example, the control unit 150 generates a track for the robot arm 124 to perform a grasping action on the object being grasped (step S16). When the track of the robot arm 124 is generated, the control unit 150 controls the upper body drive unit 146 based on the generated track. Thus, the grasping action of the robot arm 124 on the object being grasped is performed (step S17).
[0129] Alternatively, the control unit 150 generates the track of the manipulator 124 and the movement path (travel path) of the robot 100 for carrying out the grasping and transporting action of the robot 100 (step S16). When the control unit 150 generates the track of the manipulator 124 and the movement path of the robot 100, it controls the upper body drive unit 146 and the trolley drive unit 145 based on this information. Thus, the manipulator 124 and the robot 100 carry out the grasping and transporting action of the object (step S17). Figure 18 In the example, the robotic arm 124 and the robot 100 also perform the action of transporting the grasped object.
[0130] Next, use Figure 19 ,right Figure 15 The details of step S15 will be explained. Figure 19 To represent the actions performed by robot 100 Figure 15 A flowchart illustrating the detailed process of step S15.
[0131] First, if the recognition unit 151 receives handwritten information input for the captured image sent from the remote terminal 300, it uses the first learned model 181 read from the memory 180 to extract the graspable area captured in the captured image and recognize it as a graspable object (step S151).
[0132] Next, the inference unit 152 infers the object to be grasped by the robot arm 124 from the graspable objects identified by the recognition unit 151 based on the input position of the handwritten information on the captured image (step S152). Furthermore, the input position of the handwritten information on the captured image is identified, for example, using the method already described.
[0133] Next, the recognition unit 151 uses the third learning completion model 183 read from the memory 180 to extract and recognize the movable and immovable regions of the robot 100 captured in the image with input handwritten information (step S153). Furthermore, the recognition processing (segmentation processing) of the movable and immovable regions of the robot 100 performed by the recognition unit 151 can also be performed simultaneously with the recognition processing of the graspable object.
[0134] Figure 20 This diagram illustrates an example of the movable and immovable areas of the robot 100 identified by the identification unit 151. Figure 20 The captured image 310 shown is for... Figure 6 The image shown is an example of the identification processing performed by the recognition unit 151 on the captured image 310 to identify the movable and immovable areas of the robot 100. Figure 20 In the captured image 310 shown, the floor 450 is identified as a movable area of the robot 100, and the table 400 is identified as a non-movable area of the robot 100. Additionally, in Figure 20 In the example, it is inferred that the object being grasped is object 402.
[0135] also, Figure 21 A diagram illustrating other examples of the movable and immovable areas of the robot 100 identified by the identification unit 151. Figure 21 The image shown is for... Figure 18 The image shown is an example of the identification processing performed by the identification unit 151 to identify the movable and immovable areas of the robot 100 based on the captured image. Figure 21 In the captured images shown, the floor 450 is identified as a movable area of the robot 100, while the table 400 and cabinet 420 are identified as non-movable areas of the robot 100. Furthermore, graspable objects other than object 401 (such as object 402) are also identified as non-movable areas. Additionally, in Figure 21 In the example, it is inferred that the object being grasped is object 402.
[0136] Next, the inference unit 152 uses the second learned model 182 read from the memory 180 to obtain the meaning of the action content of the robot 100 simulated by the handwritten information from the image of the handwritten information, and infers the action content requested by the robot 100 for the grasping object (step S154).
[0137] exist Figure 20 In the example, the inference unit 152 infers from the image of the handwritten information 904 that the action requested by the robot 100 is to transport the object 402, represented by the handwritten information 904a, to the destination represented by the handwritten information 904c along the transport path represented by the handwritten information 904b. Figure 20In the example, the transport path of the object 402 represented by the handwritten information 904 is entirely located within the immovable area of the robot 100. In this case, the inference unit 152 infers that the transport path of the object 402 represented by the handwritten information 904 is not the travel path of the robot 100, but rather a movement path achieved by the manipulator 124 mounted at the tip of the arm of the robot 100 through arm operation.
[0138] In addition, Figure 21 In the example, the inference unit 152 infers from the image of the handwritten information 906 that the action requested by the robot 100 is to transport the object 402, represented by the handwritten information 906a, to the destination represented by the handwritten information 906c along the transport path represented by the handwritten information 906b. Figure 21 In the example, part of the transport path of object 402 represented by handwritten information 906 exists within the movable area of robot 100, and another part exists within the immovable area of robot 100. Specifically, as... Figure 22 As shown, in the transport path represented by handwritten information 906b, transport path R2 exists within the movable area of robot 100, while transport paths R1 and R3 exist within the immovable area of robot 100. In this case, the inference unit 152 infers transport path R2, which is located within the movable area of robot 100, as the movement path (driving path) of robot 100, and infers transport paths R1 and R3, which are located within the immovable area of robot 100, as the movement paths of the manipulator 124 mounted on the tip of the arm of robot 100, achieved through arm operation.
[0139] Furthermore, the inference unit 152 can also infer a detour path for the grasped object where there is no immovable area sandwiched between two movable areas on the transport path represented by the handwritten information. The robot 100 has an output unit that outputs information about this detour path; for example, it can be displayed on the display panel 341 of the remote terminal 300. Figure 23 In the example, on the transport path represented by handwritten information 906, there is an immovable area (obstacle 451) sandwiched between two movable areas. In this case, the inference unit 152 can, for example, infer a detour path 916 to avoid obstacle 451 and make a suggestion to the user.
[0140] As described above, in the remote operating system 10 of this embodiment, in the robot 100, the recognition unit 151 identifies the graspable objects that the robotic arm 124 can grasp based on the captured images obtained by taking pictures of the environment where the robot 100 is located. Furthermore, the inference unit 152 infers the graspable object that the robotic arm 124 requests to grasp from the graspable objects based on the graspable objects identified by the recognition unit 151 from the captured images and the handwritten information input by the user in relation to the captured images, and infers the action content (grasping action, transporting action) requested by the robot 100 for the graspable object.
[0141] Therefore, in the remote operating system 10 of this embodiment, the user can perform handwriting input without having to recall a pre-set instruction graphic, and instead perform more intuitive handwriting input to enable the robot 100 to perform desired processes such as transporting the grasped object.
[0142] Furthermore, in the remote operating system 10 of this embodiment, since the captured image displayed on the display panel 341 of the remote terminal 300 switches according to the area specified in the handwritten information, a high degree of visual confirmation can be maintained, for example, enabling the input of handwritten information across a wide range. In other words, the remote operating system 10 improves the ease of operation.
[0143] Furthermore, in the remote operating system 10 of this embodiment, when the handwritten information includes information about the transport path of the grasped object, it automatically distinguishes whether the transport path represents the movement path (driving path) of the robot 100 or the movement path of the manipulator 124. Therefore, the remote operating system 10 can enable the robot 100 to perform high-precision processing while maintaining the user's intuitive operation capabilities.
[0144] Furthermore, the present invention is not limited to the above-described embodiments and can be appropriately modified without departing from the spirit of the invention.
[0145] In the above embodiment, the case where the handwritten information is an image simulating the grasping action requested by the robot 124 for a grasping object has been described as an example, but it is not limited to this. The handwritten information may also include an image indicating the degree of the grasping action. In this case, the inference unit 152 may further infer the degree of the grasping action performed by the robot 124 for the grasping object based on the handwritten information. Figure 24 A diagram illustrating an example of handwritten information containing an image showing the degree of grasping action. Figure 24 The example shows that in relation to Figure 5The handwritten information 903 was appended to the same image as handwritten information 908, with the handwritten information 908 indicating the degree of grasping action as "30°". Figure 24 In the example, the inference unit 152 infers that the gripping action requested by the robotic arm 124 for the handle 411 is an action such as gripping the handle 411 and opening the door 30°. As a result, the user can perform more detailed and intuitive operations.
[0146] Furthermore, although the above embodiment describes an example of inputting one handwritten message for a captured image, it is not limited to this. Multiple handwritten messages can also be input for a captured image. When multiple handwritten messages are input for a captured image, the inference unit 152 infers the grasping object and the robot's actions based on each handwritten message. In this case, the inference unit 152 may also infer that the robot's actions on the grasping object inferred from the previously input handwritten messages will be prioritized. Alternatively, the multiple handwritten messages may also include images representing the order of the grasping actions. In this case, the inference unit 152 further infers the order of the grasping actions based on the handwritten messages. Figure 25 A diagram illustrating an example of inputting handwritten information from multiple images containing the sequence of grasping actions. Figure 25 The example shows two examples of handwritten information input for the captured image 310: handwritten information 909 for handle 422 and handwritten information 910 for handle 411. Here, handwritten information 909 includes an image such as "1" indicating the sequence of gripping actions, and handwritten information 910 includes an image such as "2" indicating the sequence of gripping actions. Therefore, in Figure 25 In the example, the inference unit 152 infers that the robot arm 124 first performs a gripping action on the handle 422 (such as gripping the handle 422 and opening the drawer), and then performs a gripping action on the handle 411 (such as gripping the handle 411 and opening the cabinet door).
[0147] Furthermore, while the above embodiments use three learning completion models—first learning completion model 181, second learning completion model 182, and third learning completion model 183—these are not limiting. For example, instead of the first learning completion model 181 and the second learning completion model 182, a transfer learning completion model can be used that applies the output of the first learning completion model 181 to the second learning completion model 182. This transfer learning completion model, for example, is a model that takes a photographed image containing handwritten information as input and outputs the meaning of the graspable object captured in the photographed image, the grasping object within the graspable object, and the action content of the robot 100 simulated by the handwritten information towards the grasping object.
[0148] Furthermore, although the recognition unit 151, the inference unit 152, and the image adjustment unit 153 are provided in the robot 100 in the above embodiment, it is not limited thereto. All or part of the recognition unit 151, the inference unit 152, and the image adjustment unit 153 may be provided in the remote terminal 300 or in the system server 500.
[0149] Furthermore, although in the above embodiment, the robot 100 and the remote terminal 300 exchange captured images and handwritten information via the network 600 and the system server 500, this is not a limitation. The robot 100 and the remote terminal 300 can also exchange captured images and handwritten information through direct communication.
[0150] Furthermore, although the imaging unit (stereo camera 131) of the robot 100 was used in the above embodiment, it is not limited to this. The imaging unit can be any imaging unit provided at any location in the first environment where the robot 100 is located. In addition, the imaging unit is not limited to a stereo camera, and can also be a single-lens reflex camera or the like.
[0151] Furthermore, although the above embodiment describes a robot 100 with a robotic arm 124 at the tip of arm 123 as the end effector, the robot is not limited to this. The object being manipulated can be any object equipped with an end effector, and the grasping action can be performed using the end effector. Additionally, the end effector can be any grasping component other than a robotic arm (e.g., a suction unit).
[0152] Furthermore, this disclosure enables the remote operating system 10 to perform some or all of its processing by having the CPU (Central Processing Unit) execute computer programs.
[0153] When the above-described program is read by a computer, it includes a set of instructions (or software code) as described in the embodiments for causing the computer to perform one or more functions. The program may be stored in a non-transitory computer-readable medium or physical storage medium. By way of example, and not limitation, computer-readable media or physical storage media include RAM (Random-Access Memory), ROM (Read-Only Memory), flash memory, SSD (Solid-State Drive) or other storage technologies, CD-ROM, DVD (Digital Versatile Disc), Blu-ray disc or other optical disc storage, magnetic tape, including magnetic tape, disk storage, or other magnetic storage devices. The program may be transmitted on a temporary computer-readable medium or communication medium. By way of example, and not limitation, temporary computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagation signals.
[0154] It is apparent from the disclosure described herein that embodiments of this disclosure can be varied in many ways. Such variations should not be considered a departure from the technical spirit and scope of this disclosure, and it will be clear to those skilled in the art that all such modifications are included within the technical solutions described above.
Claims
1. A remote operating system for remotely operating an object having an end effector, the remote operating system comprising: The camera unit takes pictures of the environment in which the object being operated is located; The recognition unit identifies, based on images captured of the environment, the graspable objects that the end effector can grasp, the movable areas of the manipulated object, and the immovable areas. The operating terminal displays the captured image and accepts input of handwritten information related to the captured image; The inference unit, based on the graspable object and the handwritten information, infers from the graspable object the object to be grasped by the end effector, and infers the action content requested for the graspable object from the object to be manipulated; When the handwritten information includes information about the transport path of the grasped object, the inference unit infers the transport path located within the movable area of the manipulated object as the movement path of the manipulated object, and infers the transport path located within the immovable area of the manipulated object as the movement path of the end effector that grasped the grasped object.
2. The remote operating system as described in claim 1, wherein, When the handwritten information includes information about the transport path of the grasped object, if there is an immovable area sandwiched between two movable areas on that transport path, the inference unit infers a detour path for the grasped object. This detour path is one where there is no immovable area sandwiched between the two movable areas. The remote operating system also includes an output unit that outputs information about the detour path inferred by the inference unit.
3. The remote operating system as described in claim 2, wherein, The output unit enables the operating terminal to display information about the detour path.
4. The remote operating system as described in claim 1, wherein, The identification unit identifies objects other than the object to be grasped as immovable areas of the object being manipulated.
5. The remote operating system as described in claim 1, wherein, The object being manipulated is a robot capable of autonomous movement.
6. The remote operating system as described in claim 5, wherein, When the handwritten information includes information about the transport path of the grasped object, the inference unit infers the transport path located within the movable area of the manipulated object as the movement path of the manipulated object, i.e., the robot, and infers the transport path located within the immovable area of the manipulated object as the movement path achieved by the end effector installed at the tip of the robot's arm through the operation of the arm.
7. The remote operating system as described in claim 1, wherein, The camera is mounted on the object being operated on.
8. The remote operating system as described in claim 1, wherein, The camera is positioned in a location different from the object being operated on within the environment in which the object is located.
9. The remote operating system as described in claim 1, wherein, The operation terminal completes the input of handwritten information by performing a predetermined handwriting end operation based on the user's instructions.
10. The remote operating system as described in claim 1, wherein, The handwritten information includes a first image simulating the actions of the manipulated object in relation to the grasped object.
11. The remote operating system as described in claim 10, wherein, The inference unit uses the learned model and the first image of the handwritten information to infer the action content requested for the grasped object.
12. A remote operation method, which is a remote operation method implemented by a remote operating system for remotely operating an object equipped with an end effector, wherein, The environment in which the object being operated is located is photographed. Based on images captured of the environment, the system identifies the objects that the end effector can grasp, the movable areas of the manipulated object, and the immovable areas. The operating terminal that displays the captured image accepts input of handwritten information related to the displayed captured image. Based on the graspable object and the handwritten information, the grasping object requesting grasping by the end effector is inferred from the graspable object, and the action content requested for the grasping object is inferred from the manipulated object. In inferring the action content requested by the manipulated object for the grasped object, if the handwritten information contains information about the transport path of the grasped object, the transport path located within the movable area of the manipulated object is inferred as the movement path of the manipulated object, and the transport path located within the immovable area of the manipulated object is inferred as the movement path of the end effector that grasped the grasped object.
13. A computer-readable storage medium storing a control program that causes a computer to perform remote operation processing, said remote operation processing being implemented by a remote operating system that remotely operates an operable object having an end effector, said control program causing the computer to perform the following processing: The process of photographing the environment in which the object being operated is located; Based on the images captured in the environment, the process identifies the graspable objects that the end effector can grasp, the movable areas of the manipulated object, and the immovable areas. In the operating terminal that displays the captured image, the input of handwritten information for the displayed captured image is processed. Based on the graspable object and the handwritten information, the system infers from the graspable object the object to be grasped by the end effector, and infers the processing of the action content requested by the manipulated object for the graspable object. In the process of inferring the action content requested for the manipulated object regarding the grasped object, if the handwritten information contains information about the transport path of the grasped object, the transport path located within the movable area of the manipulated object is inferred as the movement path of the manipulated object, and the transport path located within the immovable area of the manipulated object is inferred as the movement path of the end effector that grasped the grasped object.