Remote control method, program, remote control device, and remote control system

JP7915634B2Active Publication Date: 2026-09-04HONDA MOTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022149426
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2026-09-04
Estimated Expiration
2042-09-20

AI Technical Summary

Benefits of technology

【0016】 (1)~(10)によれば、遅延が生じる環境であっても、操作性を損なわない画像情報を提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007915634000001
    Figure 0007915634000001
  • Figure 0007915634000002
    Figure 0007915634000002
  • Figure 0007915634000003
    Figure 0007915634000003
Patent Text Reader

Abstract

To provide a remote control operation method, a program, a remote control operation device, and a remote control operation system which can provide image information to prevent damage of the operability even in an environment where delay occurs.SOLUTION: A remote control operation method of a robot comprises: a conversion process which converts, when an operator operates, an operation input into a joint angle command value of the robot; an acquisition process which acquires real visual information of a space where the robot exists; a creation process which creates virtual visual information of arbitrary viewpoint in the space where the robot exists by three dimensional expression for the operator; and a projection process which projects, when the real visual information cannot be newly obtained, the virtual visual information onto a screen the operator visually recognizes, and which projects, when the real visual information can be newly obtained, a mixed image where an image based on the real visual information newly obtained is synthesized with the virtual visual information onto the screen the operator visually recognizes.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote control method, a program, a remote control device, and a remote control system.

Background Art

[0002] When performing remote control of a robot, an operator is required to grasp the surrounding environment of the robot located in a remote place. Since visual perception accounts for most of the perception proportion of human five senses, presentation of visual information is effective for environment grasping. Systems that synthesize and provide image data from remote locations have been proposed (see, for example, Patent Document 1).

Prior Art Literature

Patent Literature

[0003]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0004] However, visual information has a large data capacity, and the communication bandwidth is limited during remote control, so communication delay occurs. For example, a delay of approximately 0.75 seconds occurs in a round trip between the International Space Station and the ground, and a delay of approximately 3 seconds occurs in a round trip between the Moon and the ground. Furthermore, this delay time may vary depending on the condition of the transmission path. Operability is impaired in an environment where communication delay varies. As described above, in the prior art, operability is impaired due to transmission delay of image information during remote control.

[0005] The present invention has been made in view of the above problems, and an object of the present invention is to provide a remote control method, a program, a remote control device, and a remote control system that can provide image information without impairing operability even in an environment where delay occurs.

Means for Solving the Problem

[0006] (1) In order to achieve the above objective, a remote control method according to one aspect of the present invention is: The computer executes A method for remotely controlling a robot, comprising a conversion step of converting the operator's control input to the robot into joint angle command values ​​of the robot, communication The robot via the network to Actual visual information of existing space , and the posture information of the robot The acquisition process to obtain, An estimation step which estimates the posture of the robot and the object to be worked at a future time after the acquisition time, which is the time when the actual visual information was acquired, based on the posture information acquired via the communication network in the acquisition step, the joint angle command value, and a delay amount estimated based on the time when the posture information was acquired on the robot side and the time when the posture information was acquired via the communication network in the acquisition step, and based on the joint angle command value, the posture of the robot and the object to be worked at the future time estimated in the estimation step, and the actual visual information acquired in the acquisition step, 3D representation At the aforementioned future time A generation process for generating virtual visual information from an arbitrary viewpoint in the space in which the robot exists, and, if the real visual information has not been newly acquired, a virtual image based on the virtual visual information. perspective The image is projected onto a screen visible to the operator, and if the actual visual information has been newly acquired, the actual visual image is an image based on the newly acquired actual visual information, and the virtual perspective A remote control method for a robot, comprising a projection step of projecting a mixed image, which is a composite image formed by combining two images, onto a screen visible to the operator.

[0007] (2) In addition, in the remote control method according to one aspect of the present invention, the virtual visual information is obtained by taking the visual information held at the time of acquisition of the real visual information and the visual information held at the time the real visual information arrives at the operator as input, and by the inverse transformation of the three-dimensional representation which predicts parameters such as motion and shape from the visual information by regression, and by interpolating the parameters between time periods using a time series network.

[0008] (3) In addition, in the remote control method according to one aspect of the present invention, the generation step generates the virtual visual information by inputting the joint angle command value, the orientation of the screen viewed by the operator, and the acquired real visual information into a trained model.

[0009] (4) In addition, in any one of the remote control methods (1) to (3) according to one aspect of the present invention, the virtual visual information is generated using posture information which is estimated to be the posture of the robot and the object to be worked on at the time the actual visual information was acquired.

[0010] (5) In addition, in any one of the remote control methods (1) to (4) according to one aspect of the present invention, the actual visual information includes metadata including time information of the image being captured, position information of the work target object recognized by recognition processing on the captured image, and posture information of the work target object, and in the generation step, the virtual visual information is generated using the posture of the robot and estimated information obtained by estimating the posture of the work target object using the metadata.

[0011] (6) In addition, in any one of the remote control methods (1) to (5) according to one aspect of the present invention, in the projection step, the mixing ratio of the actual visual information and the virtual visual information is determined according to the rate of change of the actual visual information over time.

[0012] (7) In addition, in any one of the remote control methods (1) to (6) according to one aspect of the present invention, the actual visual information includes RGB data and depth data.

[0013] (8) To achieve the above objective, a program according to one aspect of the present invention causes a remotely controlled computer to convert the operator's input to the robot into joint angle command values ​​of the robot, communication The robot via the network to Actual visual information of existing space , and the posture information of the robot Let it obtain, Based on the posture information acquired via the communication network, the joint angle command values, and a delay amount estimated based on the time difference between the time the posture information was acquired on the robot side and the time the posture information was acquired via the communication network, the posture of the robot and the work object at a future time after the acquisition time, which is the time the actual visual information was acquired, and based on the joint angle command values, the estimated posture of the robot and the work object at the future time, and the acquired actual visual information, 3D representation At the aforementioned future time The robot generates virtual visual information for an arbitrary viewpoint in the space in which it exists, and if the real visual information has not been newly acquired, it generates a virtual image based on the virtual visual information. perspective The image is projected onto a screen visible to the operator, and if the actual visual information is newly acquired, the actual visual image is an image based on the newly acquired actual visual information, and the virtual perspective The system will perform the action of projecting a composite image, created by combining two images, onto a screen visible to the operator. for It is a program.

[0014] (9) To achieve the above object, a remote control device according to one aspect of the present invention comprises: a conversion unit that converts an operation input for a robot by an operator into a joint angle command value of the robot; communication via a network, the robot to real visual information of the space in which said robot exists , and the posture information of the robot an acquisition unit that acquires An estimation unit estimates the posture of the robot and the work object at a future time after the acquisition time, which is the time when the actual visual information was acquired, based on the posture information acquired via the communication network, the joint angle command value, and a delay amount estimated based on the time difference between the time when the posture information was acquired on the robot side and the time when the posture information was acquired via the communication network; and based on the joint angle command value, the estimated posture of the robot and the work object at the future time, and the acquired actual visual information, based on three-dimensional representation At the aforementioned future time a generation unit that generates virtual visual information from an arbitrary viewpoint in the space where the robot exists; and when the real visual information has not been newly acquired, a virtual image that is an image based on the virtual visual information perspective is projected onto a screen visually recognized by the operator, and when the real visual information has been newly acquired, a real visual image which is an image based on the newly acquired real visual information and the virtual perspective a projection unit that projects a mixed image obtained by combining the image onto a screen visually recognized by the operator.

[0015] (10) To achieve the above object, a remote control system according to one aspect of the present invention comprises a remote control device and a remote site device, wherein the remote site device comprises a robot and a visual sensor that detects real visual information of a space where the robot exists; A posture sensor for detecting the posture information of the robot, wherein the remote control device comprises: a conversion unit that converts an operation input for the robot by an operator into a joint angle command value of the robot; communication via a network, from the remote site device, the real visual information and the aforementioned posture information an acquisition unit that acquires An estimation unit estimates the posture of the robot and the object to be worked at a future time after the acquisition time, which is the time when the actual visual information was acquired, based on the posture information acquired via the communication network, the joint angle command value, and a delay amount estimated based on the time difference between the time when the posture information was acquired on the remote device side and the time when the posture information was acquired via the communication network, and based on the joint angle command value, the estimated posture of the robot and the object to be worked at the future time, and the acquired actual visual information, based on three-dimensional representation At the aforementioned future time a generation unit that generates virtual visual information from an arbitrary viewpoint in the space where the robot exists; and when the real visual information has not been newly acquired, a virtual image that is an image based on the virtual visual information perspective is projected onto a screen visually recognized by the operator, and when the real visual information has been newly acquired, a real visual image which is an image based on the newly acquired real visual information and the virtual perspective a projection unit that projects a mixed image obtained by combining the image onto a screen visually recognized by the operator. Effects of the Invention

[0016] According to (1) to (10), image information that does not impair operability can be provided even in an environment where delay occurs. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] [Figure 1] It is a diagram showing a configuration example of a remote control system according to an embodiment. [Figure 2] It is a diagram showing a configuration example of a remote site apparatus according to an embodiment. [Figure 3] It is a diagram showing a configuration example of an operation side apparatus according to an embodiment. [Figure 4] It is a diagram for explaining data acquisition timing, reception timing, extrapolation of estimated information, and the like according to an embodiment. [Figure 5] It is a diagram showing a configuration example in a case where processing for a plurality of viewpoints is performed. [Figure 6] It is a flowchart of a processing procedure example of the remote control system according to an embodiment. [Figure 7] It is a flowchart of a processing procedure example for detection data of a visual sensor of a remote site apparatus according to an embodiment. [Figure 8] It is a flowchart of a processing procedure example for a detection result of an internal sensor of a remote site apparatus according to an embodiment. [Figure 9] It is a flowchart of a processing procedure example for integrated data of a remote site apparatus according to an embodiment. [Figure 10] It is a flowchart of a processing procedure example for compressed data of a remote site apparatus according to an embodiment. [Figure 11] It is a flowchart of a processing procedure example such as generation of a virtual viewpoint image and mixing of images according to an embodiment. [Figure 12] It is a diagram showing an example of an image captured by a visual sensor. DESCRIPTION OF EMBODIMENTS

[0018] Embodiments of the present invention will be described below with reference to the drawings. Note that in the drawings used in the following description, the scale of each component has been appropriately changed to ensure that each component is recognizable. In all the figures used to illustrate the embodiments, components with the same function are given the same reference numerals, and repeated explanations are omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on another element in addition to XX. Also, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on something that has been calculated or processed from XX. "XX" is any element (for example, any information).

[0019] [Example of a remote control system configuration] Figure 1 shows an example of the configuration of a remote control system according to this embodiment. As shown in Figure 1, the remote control system 1 comprises an operating device 2 and a remote location device 3. The operating device 2 includes, for example, a first communication device 21, a remote control device 22, a display unit 23, a sensor 24, and a sensor 25. The remote device 3 includes, for example, a robot 31 and a second communication device 32. The operating device 2 (remote control device) and the remote location device 3 are connected via a network NW.

[0020] The first communication device 21 receives information transmitted by the remote device 3 and outputs the received information to the remote control device 22. The first communication device 21 transmits the transmission information output by the remote control device 22 to the remote device 3.

[0021] The remote control device 22 generates an image to be displayed on the display unit 23 based on the received information and displays the generated image on the display unit 23. Based on the results detected by the gaze detection unit of the display unit 23, the remote control device 22 generates a control command for the robot 31 based on the angle and movement of the operator's hands and fingers detected by the sensor 24, the angles of the robot's joints etc. included in the received information, and the environmental information included in the received information. The remote control device 22 outputs the generated control command as transmission information to the first communication device 21. The remote control device 22 estimates the operator's head posture based on the detection results of the sensor 25.

[0022] The display unit 23 is, for example, an HMD (Head Mounted Display) and includes a gaze detection unit that detects the operator's line of sight. The display unit 23 displays images generated by the remote control device 22.

[0023] Sensor 24, for example, the operator's hand to This is an operation instruction detection unit that is attached to the device. Sensor 24 detects the position, angle, and movement of the operator's hands and fingers. Sensor 25 detects the operator's viewpoint and line of sight.

[0024] The robot 31 includes at least a manipulator, a vision sensor (camera), and sensors for detecting the angle, position, and movement of the manipulator. The robot 31 outputs the image captured by the vision sensor and the detection results detected by the sensors to the second communication device 32. The robot 31 operates according to the control instructions included in the transmission information output by the second communication device 32.

[0025] The second communication device 32 receives the transmission information sent by the operating device 2 and outputs the control instructions included in the transmission information to the robot 31. The second communication device 32 also transmits the image output by the robot 31 and the detection results to the operating device 2.

[0026] [Example of configuration for remote-side equipment] Figure 2 shows an example of the configuration of the remote-side device according to this embodiment. As shown in Figure 2, the remote-side device 3 includes, for example, a visual sensor 301 (301-1, 301-2), an internal sensor 302, a detection and estimation unit 303 (303-1, 303-2), an encoder 304, a data integration unit 305, a data compression unit 306, a communication unit 307, a control unit 308, an actuator 309, and a processing unit 310. Note that the configuration example in Figure 2 is an example of a configuration in which the operating-side device 2 is made to present a stereo image.

[0027] The visual sensor 301 (301-1, 301-2) is, for example, a photographic device equipped with a fisheye lens. The visual sensor 301 is attached to the head of, for example, the robot 31. The image captured by the visual sensor 301 also includes depth information. The number of visual sensors 301 may be one or three or more.

[0028] The internal sensor 302 is a sensor attached to the joints of the end effector of the robot 31, and is, for example, a joint encoder, a tension sensor, a torque sensor, etc.

[0029] The detection and estimation unit 303 (303-1, 303-2) performs a well-known instance segmentation process on the captured RGBD data to separate the region of interest (ROI) from the background, etc. If multiple objects are captured in the image, object recognition is performed for each object. The detection and estimation unit 303 performs a well-known image processing on the captured RGBD data to estimate the pose of the object. The detection and estimation unit 303 (303-1, 303-2) is metadata that includes, for example, a timestamp, pose, and instance (ROI of interest, class information (ID)).

[0030] The encoder 304 encodes the RGBD (W (width) × H (height) × 4 (four RGBD)) data output by the vision sensor 301-1, the metadata output by the detection and estimation unit 303-1, the RGBD (W × H × 4) data output by the vision sensor 301-2, and the metadata output by the detection and estimation unit 303-2 into streamable data using a predetermined method.

[0031] The data integration unit 305 integrates the encoded data.

[0032] The data compression unit 306 compresses the detection results detected by the internal sensor 302 using a predetermined method.

[0033] The communication unit 307 transmits the integrated data output by the data integration unit 305 and the compressed detection results from the data compression unit 306 to the operating device 2 via the network NW. The communication unit 307 receives transmission information, including control instructions, from the operating device 2 and outputs the received control instructions to the control unit 308. The transmitted image data has either a reduced frame rate or a reduced resolution to suit the communication environment. The communication unit 307 includes a first communication unit that transmits the integrated data and a second communication unit that transmits the compressed data.

[0034] The control unit 308 drives the actuator 309 using control instructions.

[0035] The actuator 309 is mounted on the end effector. The actuator 309 is driven in accordance with the control of the control unit 308.

[0036] The processing unit 310 controls data acquisition from the vision sensor 301, data acquisition from the internal environment sensor 302, the encoder 304, data compression, data integration, transmission, etc.

[0037] [Example of configuration for the operating device] Figure 3 shows an example of the configuration of the operator-side device according to this embodiment. As shown in Figure 3, the operator-side device 2 includes, for example, an HMD 201 (projection unit), an operation instruction detection unit 204 (sensor 24), a command value generation unit 205 (conversion unit), a communication unit 206 (acquisition unit), a decoder 207, a data expansion unit 208, an attitude estimation unit 209, a model 210 (generation unit), an output unit 211, an output unit 212, an output unit 213, a resolution restoration unit 214, a resolution restoration unit 215, a first mixed data generation unit 216 (generation unit, projection unit), a second mixed data generation unit 217 (generation unit, projection unit), and a processing unit 218 (generation unit, projection unit). The HMD201 includes a gaze detection unit 202 (sensor 25), a sensor 203 (sensor 25) for detecting the position and orientation of the HMD201, and a display unit 23 (Figure 1). Note that the configuration example in Figure 3 is an example where the operating device 2 presents a stereo image.

[0038] The gaze detection unit 202 (sensor 25) detects the operator's viewpoint and gaze direction. The gaze detection unit 202 is provided, for example, on both the left and right sides to accommodate both eyes. The gaze detection unit 202 inputs the detected left and right viewpoints and gaze directions to the model 210.

[0039] The operation instruction detection unit 204 (sensor 24) detects the position and movement of the operator's fingers and hands. The operation instruction detection unit 204 is the operation input interface for remote control. Operation instruction detection unit Unit 204 outputs the detected result to the command value generation unit 205. The operation instruction detection unit 204 is, for example, a data glove, an exoskeleton-type device, etc.

[0040] The command value generation unit 205 generates control commands for the hands and arms of the robot 31 using the detection results output by the operation instruction detection unit 204, and outputs the generated control commands to the communication unit 206. In other words, when an operator performs an operation, the command value generation unit 205 converts that operation input into joint angle command values ​​for the robot 31. The command value generation unit 205 generates joint angles for the hands and arms of the robot 31 using the detection results output by the operation instruction detection unit 204, and inputs information indicating the generated joint angles into the model 210.

[0041] The communication unit 206 transmits the control command output by the command value generation unit 205 to the remote device 3 via the network NW. The communication unit 206 outputs the integrated data contained in the information received from the remote device 3 to the decoder 207. The communication unit 206 outputs the compressed data contained in the information received from the remote device 3 to the data decompression unit 208. The communication unit 206 comprises a first communication unit that receives the integrated data and a second communication unit that receives the compressed data.

[0042] Decoder 207 performs a decoding process on the decoded and integrated data using a predetermined method. Through this process, Decoder 207 extracts RGBD data and metadata. Decoder 207 outputs the metadata to the attitude estimation unit 209. Decoder 207 outputs the RGBD data to the resolution restoration unit 214 and the resolution restoration unit 215.

[0043] The data decompression unit 208 decompresses and extracts the compressed data. Through this process, the data decompression unit 208 extracts the detection results from the internal sensor 302 and outputs them to the attitude estimation unit 209 and the output unit 213. to Output.

[0044] The posture estimation unit 209 is an inference unit. Based on metadata input from the decoder 207, detection results from the internal sensor 302 input from the data expansion unit 208, and virtual viewpoint (virtual vision) image data input from the model 210, the posture estimation unit 209 estimates the posture of the object to be worked on and the posture of the robot 31's end effector (arm, hand) without delay using a well-known method. For this reason, the posture of the object to be worked on and the posture of the robot 31's end effector (arm, hand) input from the posture estimation unit 209 to the model 210 have a delay time. It doesn't seem to exist.This is the estimated posture data. The posture estimation unit 209 estimates the posture based on the time difference (delay) between the time the internal sensor 302 acquired the data and the time the operating device acquired the data, using the timestamp included in the acquired data. Alternatively, the posture estimation unit 209 may pre-store a preferred posture that the robot 31 will take at the time after the delay during operation, based on the command value generated by the command value generation unit 205, and then estimate the posture by referring to the stored information.

[0045] Model 210 is a model that has learned virtual viewpoint image data (RGBD data) of the working state using, for example, the operator's viewpoint and gaze direction, the operator's joint angles, the object's pose estimated by the pose estimation unit 209, and the pose of the robot 31's end effector. Model 210 can be represented, for example, in NeRF (Neural Radiance Fields) representation (neural 3D representation). NeRF can generate a 3D model from images of multiple viewpoints and render video from any viewpoint. Furthermore, NeRF can represent a 3D model as a machine learning model without using polygons (see, for example, Reference 1). Based on the input data, Model 210 can, for example, capture images from the left and right. done An RGBD (NW (width) × NH (height) × 4 (four RGBDs)) with a resolution higher than the image data is output to output units 211 and 212.

[0046] Reference 1; Ben Mildenhall Pratul P. Srinivasan , et al, “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, ECCV 2020 (oral), 2020

[0047] The output unit 211 includes, for example, a decoder. The output unit 211 outputs the RGBD (NW (width) × NH (height) × 4) data output by the model 210 to the first mixed data generation unit 216. The output unit 211 also includes a buffer.

[0048] The output unit 212 includes, for example, a decoder. The output unit 212 outputs the RGBD (NW (width) × NH (height) × 4) data output by the model 210 to the second mixed data generation unit 217. The output unit 212 also includes a buffer. Furthermore, output units 211 and 212 are also used to provide stereoscopic vision to the HMD201. Therefore, if the provided image is not stereoscopic, the operating device 2 only needs to be equipped with one of the output units 211 and 212.

[0049] The output unit 213 outputs the detection results of the deployed internal sensor 302 to the first mixed data generation unit 216 and the second mixed data generation unit 217. The output unit 213 is equipped with a buffer. The output unit 213 also encodes the values ​​estimated by the internal sensor 302 to obtain feature quantities.

[0050] The resolution restoration unit 214 restores the resolution of the decoded image data using a well-known method. The resolution restoration unit 214 outputs the restored super-resolution image data, which is RGBD(NW (width) × NH (height) × 4 (four RGBD)) data, to the first mixed data generation unit 216. The resolution restoration unit 214 is equipped with a buffer.

[0051] The resolution restoration unit 215 restores the resolution of the decoded image data using a well-known method. The resolution restoration unit 214 outputs the restored super-resolution image data, which is RGBD (NW (width) × NH (height) × 4 (four RGBD)) data, to the second mixed data generation unit 217. The resolution restoration unit 215 is equipped with a buffer.

[0052] The first mixed data generation unit 216 includes, for example, an encoder and a decoder. The first mixed data generation unit 216 mixes virtual images and real images to generate data, and displays the generated mixed data on the display unit 23 of the HMD 201. Virtual visual information is calculated, for example, by taking visual information held when real visual information detected by the visual sensor 301 is acquired and visual information held at the time the real visual information arrives at the operator as input, predicting parameters such as motion and shape from the visual information using regression in an inverse transformation of a 3D representation, and interpolating parameters between time periods using a time series network. The first mixed data generation unit 216 may also mix virtual images and real images using the feature quantities of the internal sensor 302 obtained by the output unit 213. In this embodiment, mixing means providing real visual information when real visual information is available, and when real visual information is not available, as will be described later. Virtual visual information The objective is to provide.

[0053] The second mixed data generation unit 217 includes, for example, an encoder and a decoder. The second mixed data generation unit 217 generates mixed data of a virtual image and a real image, and displays the generated mixed data on the display unit 23 of the HMD 201. The second mixed data generation unit 217 mixes the images by, for example, determining the mixing ratio according to the rate of change. The second mixed data generation unit 217 may also mix the virtual image and the real image using the feature quantities of the internal sensor 302 obtained by the output unit 213. The operating device only needs to include one of the first mixed data generation unit 216 and the second mixed data generation unit 217. Furthermore, the first mixed data generation unit 216 and the second mixed data generation unit 217 are also used to provide stereoscopic vision to the HMD 201. Therefore, if the provided image is not stereoscopic, the operating device 2 only needs to be equipped with one of the first mixed data generation unit 216 and the second mixed data generation unit 217.

[0054] The processing unit 218 controls processes such as data transfer and parameter setting.

[0055] The display unit 23 displays either the mixed data mixed by the first mixed data generation unit 216 or the mixed data mixed by the second mixed data generation unit 217.

[0056] [Data acquisition timing, reception timing, information estimation] Next, we will explain the data acquisition timing, reception timing, and extrapolation of estimated information. Figure 4 is a diagram illustrating the data acquisition timing, reception timing, and extrapolation of estimated information according to this embodiment. In Figure 4, the horizontal axis represents frames or time.

[0057] Code g101 is an example of the data acquisition timing for the first visual sensor 301-1 (RGBD) (e.g., left side). Code g102 is an example of the data acquisition timing for the second visual sensor 301-2 (RGBD) (e.g., right side). Note that because the transmission bandwidth is often narrow, it is not possible to increase the FPS (frame rate) or resolution as shown by the double arrow g111. Therefore, in this embodiment, the operating device 2 performs a resolution restoration process to obtain image data with a higher resolution than the captured image. Note that in this embodiment, the data from the visual sensor 301 is transmitted and received via streaming.

[0058] Code g103 represents the data acquisition timing of the internal sensor 302 (joint angle, etc.). For example, the communication protocol in outer space uses CCSDS (Consultative Committee for Space Data Systems), and packets with errors are discarded. Therefore, interactive remote control is difficult in systems with delays in the second. Monitoring can be made interactive by mixing in virtual images, so in this embodiment, virtual images are mixed in.

[0059] Codes g104 and g105 represent examples of timing for receiving past information due to communication delays.

[0060] In the example in Figure 4, at time t1, the remote device 3 transmits data from the internal sensor 302. However, due to a delay in the communication system, at time t3, the operating device 2 receives the transmitted data. Similarly, at time t3, data from the internal sensor 302 is transmitted, but the operating device 2 receives the data at time t 4 The data sent at that time is received.

[0061] Furthermore, although the data from the visual sensor 301 is transmitted at time t2, due to a delay in the communication system, the operating device 2 receives the transmitted data at time t5. Although the data from the visual sensor 301 is transmitted at time t6, the operating device 2 receives the transmitted data at time t7. Note that the delay time of transmission and reception of the internal sensor 302 is used for the visual sensor 301 Send The reason for the long reception delay is that the data from the visual sensor 301 is image data and is larger than the data from the internal sensor 302.

[0062] For example, in the case of code g106, the signal received at time t4 is past data, for example, from time t3. Therefore, in this embodiment, as with code g106, current information is estimated from the received past information, and the information is extrapolated until more information is received. For example, the attitude estimation unit 209 uses the data received at time t3 to estimate the data that will be received at time t4.

[0063] In the case of image data, for example, time t 5 The image data received is from time t2, and the current image data is not available. Therefore, in this embodiment, Model 210 estimates the current image and first mixture The data generator 216 mixes the received image data with the estimated image data and presents it.

[0064] Furthermore, in this embodiment, as shown by code g107, a viewpoint image with high resolution and high frame rate is generated using a virtual image through a resolution restoration process. For example, such an image is continuously provided during the period from time t6 to t7 when actual image data cannot be received.

[0065] Furthermore, any viewpoint image can be acquired without delay because the operator is using the HMD201. However, there is a delay in acquiring the actual visual information of the space in which the robot 31 exists. Therefore, in this embodiment, the posture of the robot 31 estimated by the posture estimation unit 209 and the delayed actual visual information are input to the model 210 to obtain a virtual viewpoint image with no delay at the current time. For example, since the joint angles of the robot 31 are estimated, a virtual viewpoint image with no delay at the current time, which is the viewpoint image as seen from the robot 31, is obtained. In order to perform this processing, in this embodiment, since estimated information is used, a virtual viewpoint image can be appropriately generated even if the timing of obtaining information from the visual sensor 301 (external sensor) and the internal sensor 302 is different. Furthermore, when the posture estimation unit 209 estimates the posture of the robot 31, the processing unit 218 can obtain command values ​​based on the information acquired from the operation instruction detection unit 204 without delay. Therefore, it is possible to extrapolate during periods when information from the internal sensor 302 is not available, for example, by using a Kalman filter. Thus, in this embodiment, virtual viewpoint images can be provided without delay, even during periods when data cannot be received.

[0066] [Example of processing from multiple perspectives] Here, we will explain an example configuration for processing from multiple perspectives. Figure 5 shows an example configuration for processing multiple viewpoints. When processing multiple viewpoints, the streaming signal may use, for example, the HEVC standard of MPEG (Moving Picture Experts Group). In this case, as shown in Figure 5, for example, for the first viewpoint, RGB image data 351-1 and depth data 352-1 are input to encoder 353-1. For the Nth viewpoint, RGB image data 351-N and depth data 352-N are input to encoders 353-N. Then, the outputs of each encoder 353-1 to N are input to mixer 354. The signal format and configuration described above are merely examples and are not limited to them. The signal format may be any other format, and the configuration may be any configuration appropriate to the signal format.

[0067] [Example of processing procedure for a remote control system] Next, the processing procedure of the remote control system 1 will be described. Figure 6 is a flowchart of an example of the processing procedure of the remote control system according to this embodiment.

[0068] (Step S1) When an operator performs an operation, the processing unit 218 of the operating device 2 converts that operation input into joint angle command values ​​for the robot 31.

[0069] (Step S2) The processing unit 218 acquires real visual information of the space in which the robot 31 is located, as detected by the visual sensor 301.

[0070] (Step S3) The processing unit 218 generates virtual visual information for an arbitrary viewpoint in the space where the robot is located, for example, using the model 210, through neural 3D representation.

[0071] (Step S4) The processing unit 218 determines whether or not new real visual information has been acquired. If the processing unit 218 determines that new real visual information has been acquired (frame has been updated) (Step S4; YES), it proceeds to the process in Step S5. If the processing unit 218 determines that new real visual information has not been acquired (frame has not been updated) (Step S4; NO), it proceeds to the process in Step S6.

[0072] (Step S5) The processing unit 218 projects a composite image onto the HMD 201, which is created by combining an image based on newly received real visual information with the virtual viewpoint image. Because there is a delay in the newly received real visual information, the processing unit 218 predicts the delay and combines the newly received image with the image generated in step S3. At this time, the processing unit 218 uses the pixel values ​​of the real image for pixels that do not change even when considering the delay, or for new regions. At this time, the pixels of the real image and the composite image are not projected as they are from the real image, but rather a composite image that takes the delay time into account is created and projected. At this time, the processing unit 218 corrects extrapolated parameters such as the movement of objects from the real image and real data. After processing, the processing unit 218 terminates the process.

[0073] (Step S6) The processing unit 218 uses the previously received real visual information in step S3 The generated virtual viewpoint image is projected onto the HMD201. After processing, the processing unit 218 terminates its operation.

[0074] Furthermore, the processing unit 218 repeats the processes in steps S1 to S6 while the operation is ongoing.

[0075] [Example of processing procedure for remote device] Next, we will explain an example of the processing procedure for the remote device 3. First, an example of a processing procedure for the visual sensor 301 will be described. Figure 7 is a flowchart of an example of a processing procedure for detection data from the visual sensor of the remote device according to this embodiment.

[0076] (Step S101) The processing unit 310 determines whether or not processing (data acquisition) has been performed for all visual sensors 301. If the processing unit 310 determines that processing has been performed for all visual sensors 301 (Step S101; YES), it proceeds to the process in Step S106. If there are visual sensors 301 that have not been processed, the processing unit 310 determines whether or not processing has been performed for all visual sensors 301. and If a determination is made (step S101; NO), the process proceeds to step S102.

[0077] (Step S102) The processing unit 310 acquires RGBD data from the vision sensor 301.

[0078] (Step S103) The processing unit 310 obtains the detection result of the internal sensor 302 at the timing of RGBD data acquisition.

[0079] (Step S104) The detection and estimation unit 303 performs object recognition using the acquired RGBD data.

[0080] (Step S105) The detection and estimation unit 303 uses the acquired RGBD data to recognize the position and orientation of the object. After processing, the detection and estimation unit 303 returns to the process of step S101.

[0081] (Step S106) The encoder 304 performs encoding on the RGBD data and the recognition result.

[0082] (Step S107) The data integration unit 305 performs data integration processing on the encoded data from multiple viewpoints. The data integration processing is performed on the data from multiple viewpoints, as shown in the mixer 354 in Figure 5.

[0083] (Step S108) The communication unit 307 transmits the integrated data to the operating device 2.

[0084] Next, an example of a processing procedure for the internal sensor will be described. Figure 8 is a flowchart of an example of a processing procedure for the detection result of the internal sensor of the remote device according to this embodiment.

[0085] (Step S151) The processing unit 310 obtains the status of the robot 31 by acquiring detection results from the internal sensor 302.

[0086] (Step S152) The data compression unit 306 performs data compression on the detection results, for example, using an irreversible compression method.

[0087] (Step S153) The communication unit 307 transmits the compressed data to the operating device 2.

[0088] [Example of processing procedure for the operating device] Next, an example of the processing procedure for the operating device 2 will be explained. First, we will describe an example of a processing procedure for integrated data. Figure 9 is a flowchart of an example of a processing procedure for integrated data from the remote device according to this embodiment.

[0089] (Step S201) The communication unit 206 receives the integrated data transmitted by the remote device 3.

[0090] (Step S202) The communication unit 206 determines whether or not there are error packets in the received data. If the communication unit 206 determines that there are no error packets in the received data (Step S202; YES), it proceeds to step S204. If the communication unit 206 determines that there are error packets in the received data (Step S202; NO), it proceeds to step S203.

[0091] (Step S203) The communication unit 206 deletes the error packet and proceeds to the process in step S204.

[0092] (Step S204) The decoder 207 decodes and expands the integrated data.

[0093] (Step S205) The resolution restoration units 214 and 215 perform a resolution restoration process on the RGBD data among the decoded data.

[0094] (Step S206) Resolution restoration units 214, 215 restore the resolution of The restored data is stored in a buffer.

[0095] Next, an example of a processing procedure for compressed data will be described. Figure 10 is a flowchart of an example of a processing procedure for compressed data by the remote device according to this embodiment.

[0096] (Step S251) The communication unit 206 receives the compressed data transmitted by the remote device 3.

[0097] (Step S252) The communication unit 206 determines whether or not there are error packets in the received data. If the communication unit 206 determines that there are no error packets in the received data (Step S252; YES), it proceeds to step S254. If the communication unit 206 determines that there are error packets in the received data (Step S252; NO), it proceeds to step S253.

[0098] (Step S253) The communication unit 206 deletes the error packet, and then in step S2 5 Proceed to step 4.

[0099] (Step S254) The data decompression unit 208 decompresses and extracts the compressed data.

[0100] (Step S255) The output unit 213 holds the decompressed data in a buffer.

[0101] Next, virtual perspective This section explains examples of processing procedures such as image generation and image mixing. Figure 11 is a flowchart showing an example of the processing procedure for generating a virtual viewpoint image and mixing images according to this embodiment.

[0102] (Step S301) The processing unit 218 obtains the latest data from the buffer.

[0103] (Step S302) The processing unit 218 determines whether or not the operation device 2 is being started for the first time (initial startup). If the processing unit 218 determines that the operation device 2 is being started for the first time (Step S302; YES), it proceeds to the process in step S303. If the processing unit 218 determines that the operation device 2 is not being started for the first time (Step S302; NO), it proceeds to the process in step S305.

[0104] (Step S303) The processing unit 218 updates (initializes) the posture estimation parameters of the robot 31 to the posture estimation unit 209.

[0105] (Step S304) The processing unit 218 updates (initializes) the parameters for estimating the object's attitude to the attitude estimation unit 209.

[0106] (Step S305) The processing unit 218 acquires information (position, orientation) of the HMD201.

[0107] (Step S306) The processing unit 218 determines whether or not there is new posture information for the robot 31 from the internal sensor 302. If the processing unit 218 determines that there is new posture information (Step S306; YES), it proceeds to the process in Step S307. If the processing unit 218 determines that there is no new posture information (Step S306; NO), it proceeds to the process in Step S308.

[0108] (Step S307) The processing unit 218 updates the parameters of the attitude estimation unit 209.

[0109] (Step S308) The processing unit 218 acquires the estimation results from the posture estimation unit 209 (the posture of the robot 31 and the posture of the object).

[0110] (Step S309) The processing unit 218 receives the data acquired from the model 210 (attitude estimation unit) 209 The estimated results, the position and orientation of the HMD201, and the detection results of the internal sensor 302 are input to generate a virtual viewpoint image.

[0111] (Step S310) The processing unit 218 determines whether the frame has been updated (i.e., whether the actual image has been obtained). If the processing unit 218 determines that the frame has been updated (Step S310; YES), it proceeds to the process in step S311. If the processing unit 218 determines that the frame has not been updated (Step S310; NO), it proceeds to the process in step S312.

[0112] (Step S311) At least one of the first mixed data generation unit 216 and the second mixed data generation unit 217 generates a mixed image by combining an image based on newly received real visual information with a virtual viewpoint image.

[0113] (Step S312) At least one of the first mixed data generation unit 216 and the second mixed data generation unit 217 displays a virtual viewpoint image on the HMD201 when a real image is not available, and mixes the image when a real image is available. image Display this on the HMD201.

[0114] [Example of an image displayed on the HMD] Here, we will explain an example of an image displayed on the HMD201. Figure 12 shows an example of an image captured by a visual sensor. Image g101 is an example of an image captured by the first visual sensor 301, showing the working state including the robot 31. Image g111 is the first 2 The image g101 and g111 show examples of images from the perspective of the robot 31, as seen by the visual sensor 301. There is a delay in the real-vision images acquired remotely for the robot 31. Therefore, in this embodiment, a virtual viewpoint image is generated using previously acquired real-vision images, or a mixed image is generated by combining an image based on newly received real-vision information with the virtual viewpoint image, and this is presented to the HMD 201. The virtual viewpoint image is, for example, an image in which a 360-degree viewpoint is estimated in a spherical shape.

[0115] According to this embodiment, since a virtual viewpoint image is generated, the operator can observe the environment in which the robot 31 exists from a free viewpoint, regardless of the robot 31's viewpoint. Furthermore, in this embodiment, since the current virtual viewpoint image is estimated using real visual information obtained with a delay, the operator can be continuously provided with images of the robot 31's workspace by providing a virtual viewpoint image during periods when real visual information is unavailable.

[0116] According to this embodiment, a photorealistic image can be presented to the operator that shows no delay even in a delayed environment. Furthermore, according to this embodiment, a virtual image can be created even in areas invisible to the robot.

[0117] Furthermore, the method of this embodiment described above can be applied to any location where the robot 31 exists, not just outer space, as delays will occur during data transmission and reception if the location is far from the operator.

[0118] Furthermore, the HMD201 may be binocular or monocular. In addition, the virtual viewpoint images provided may be monochrome or color. The images provided may be a series of still images or a video of a series of images over a short period of time.

[0119] Furthermore, a program to implement all or part of the functions of the operating device 2 and the remote device 3 in this invention may be recorded on a computer-readable recording medium, and all or part of the processing performed by the operating device 2 and the remote device 3 may be performed by loading the program recorded on this recording medium into a computer system and executing it. Here, "computer system" includes hardware such as an OS and peripheral devices. Also, "computer system" includes a WWW system equipped with a homepage provisioning environment (or display environment). Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer system. Moreover, "computer-readable recording medium" also includes volatile memory (RAM) inside a computer system that acts as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, which holds the program for a certain period of time.

[0120] Furthermore, the above program may be transmitted from a computer system that stores the program in a memory device or the like to another computer system via a transmission medium or by transmission waves within the transmission medium. Here, the "transmission medium" for transmitting the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. In addition, the above program may be for the purpose of realizing a part of the functions described above. Furthermore, it may be a so-called differential file (differential program) that can realize the functions described above in combination with a program already recorded in the computer system.

[0121] Although embodiments for carrying out the present invention have been described above using examples, the present invention is not limited in any way to these embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention. [Explanation of Symbols]

[0122] 1...Remote control system, 2...Operating device, 3...Remote location device, 21...First communication device, 22...Remote control device, 23...Display unit, 24...Sensor, 25...Sensor, 31...Robot, 32...Second communication device, NW...Network, 201...HMD, 204...Operation instruction detection unit, 205...Command value generation unit, 206...Communication unit, 207...Decoder, 208...Data expansion unit, 209...Attitude estimation unit, 210...Model, 211...Output unit, 212...Output unit, 213…Output unit, 214…Resolution restoration unit, 215…Resolution restoration unit, 216…First mixed data generation unit, 217…Second mixed data generation unit, 218…Processing unit, 301, 301-1, 301-2…Vision sensor, 302…Internal environment sensor, 303, 303-1, 303-2…Detection estimation unit, 304…Encoder, 305…Data integration unit, 306…Data compression unit, 307…Communication unit, 308…Control unit, 309…Actuator, 310…Processing unit

Claims

1. A method for remotely controlling a robot, performed by a computer, A conversion step that converts the operator's input to the robot into joint angle command values ​​for the robot, An acquisition step of acquiring real visual information of the space in which the robot is located and posture information of the robot via a communication network, An estimation step, based on the posture information acquired via the communication network in the acquisition step, the joint angle command values, and a delay amount estimated based on the time difference between the time the posture information was acquired on the robot side and the time the posture information was acquired via the communication network in the acquisition step, estimates the posture of the robot and the work object at a future time after the acquisition time, which is the time the actual visual information was acquired. A generation step generates virtual visual information of an arbitrary viewpoint in the space where the robot is located at the future time, using a three-dimensional representation, based on the joint angle command value, the posture of the robot and the work object at the future time estimated in the estimation step, and the actual visual information acquired in the acquisition step. If the aforementioned real visual information has not been newly acquired, a projection step is performed to project a virtual viewpoint image, which is an image based on the virtual visual information, onto the screen viewed by the operator; and if the aforementioned real visual information has been newly acquired, a projection step is performed to project a mixed image, which is a composite image obtained by combining the newly acquired real visual image (which is an image based on the aforementioned real visual information) and the virtual viewpoint image, onto the screen viewed by the operator. A method for remotely controlling a robot having [a specific feature / ability].

2. In the generation step, the virtual visual information is generated by inputting the joint angle command value, the orientation of the screen viewed by the operator, and the acquired real visual information into a trained model. A method for remotely controlling a robot according to claim 1.

3. In the estimation step, the posture of the work object at the time of acquisition is estimated using the actual visual information acquired in the acquisition step, and the posture of the robot and the work object at the future time is estimated using the estimated posture of the work object at the time of acquisition. In the generation step, the virtual visual information is generated using the postures of the robot and the work object at the future time estimated in the estimation step. A method for remotely controlling a robot according to claim 1 or claim 2.

4. In the projection step, the mixing ratio of the real visual image and the virtual viewpoint image is determined according to the rate of change of the real visual information at each acquisition time. A method for remotely controlling a robot according to claim 1 or claim 2.

5. The aforementioned real-world visual information includes RGB data and depth data. A method for remotely controlling a robot according to claim 1 or claim 2.

6. To a remotely controlled computer, The operator's input to the robot is converted into joint angle command values ​​for the robot. The system acquires real-world visual information of the space in which the robot is located, and posture information of the robot, via a communication network. Based on the posture information acquired via the communication network, the joint angle command values, and a delay amount estimated based on the time difference between the time the posture information was acquired on the robot side and the time the posture information was acquired via the communication network, the posture of the robot and the work object at a future time after the acquisition time, which is the time the actual visual information was acquired, is estimated. Based on the joint angle command values, the estimated postures of the robot and the work object at the future time, and the acquired real visual information, virtual visual information of an arbitrary viewpoint in the space where the robot exists at the future time is generated in a three-dimensional representation. If the aforementioned real visual information has not been newly acquired, a virtual viewpoint image, which is an image based on the aforementioned virtual visual information, is projected onto the screen viewed by the operator. If the aforementioned real visual information has been newly acquired, a composite image, which is a combination of the newly acquired real visual image (based on the aforementioned real visual information) and the aforementioned virtual viewpoint image, is projected onto the screen viewed by the operator. A program to perform a task.

7. A conversion unit that converts the operator's input to the robot into joint angle command values ​​for the robot, An acquisition unit that acquires real-world visual information of the space in which the robot is located and posture information of the robot via a communication network, An estimation unit estimates the posture of the robot and the work object at a future time after the acquisition time, which is the time when the actual visual information was acquired, based on the posture information acquired via the communication network, the joint angle command value, and a delay amount estimated based on the time when the posture information was acquired on the robot side and the time when the posture information was acquired via the communication network. A generation unit generates virtual visual information of an arbitrary viewpoint in the space where the robot is located at the future time, using a three-dimensional representation, based on the joint angle command value, the estimated posture of the robot and the work object at the future time, and the acquired real visual information. A projection unit that, when the aforementioned real visual information has not been newly acquired, projects a virtual viewpoint image, which is an image based on the aforementioned virtual visual information, onto the screen viewed by the operator, and when the aforementioned real visual information has been newly acquired, projects a mixed image, which is a composite image obtained by combining the newly acquired real visual image (which is an image based on the aforementioned real visual information) and the aforementioned virtual viewpoint image, onto the screen viewed by the operator. A remote control device equipped with the following features.

8. It comprises a remote control device and a remote location device, The remote-side device is, Robots and, A visual sensor that detects real visual information of the space in which the robot exists, The robot comprises a posture sensor for detecting posture information, The remote control device is A conversion unit that converts the operator's input to the robot into joint angle command values ​​for the robot, An acquisition unit that acquires the actual visual information and the attitude information from the remote device via a communication network, An estimation unit estimates the posture of the robot and the work object at a future time after the acquisition time, which is the time when the actual visual information was acquired, based on the posture information acquired via the communication network, the joint angle command value, and a delay amount estimated based on the time difference between the time when the posture information was acquired on the remote device side and the time when the posture information was acquired via the communication network. A generation unit generates virtual visual information of an arbitrary viewpoint in the space where the robot is located at the future time, using a three-dimensional representation, based on the joint angle command value, the estimated posture of the robot and the work object at the future time, and the acquired real visual information. The system includes a projection unit that, when new real visual information cannot be acquired, projects a virtual viewpoint image, which is an image based on the virtual visual information, onto the screen viewed by the operator, and when new real visual information can be acquired, projects a mixed image, which is a composite of the newly acquired real visual image (based on the real visual information) and the virtual viewpoint image, onto the screen viewed by the operator. Remote control system.

Citation Information

Patent Citations

  • Using augmented reality for controlling intelligent devices

    WO2019046559A1

  • View synthesis robust to unconstrained image data

    WO2022026692A1