Remote operation system and remote operation device
The remote control system addresses communication delays by predicting future images and positions of moving objects using a three-dimensional scene construction and operation information, enhancing accuracy and operability.
Patent Information
- Application Number
- PCT/JP2024/012683
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-10-02
AI Technical Summary
Communication delays in remote control systems for moving objects reduce the accuracy of movement prediction and operability due to the time lag in obtaining speed information from cameras, leading to decreased operability.
A remote control system that predicts future images of a moving obstacle by constructing a three-dimensional scene based on ambient images and operation information, using a remote computer to generate predicted images without communication delays, thereby improving accuracy and operability.
The system enhances the accuracy of movement prediction and image prediction by predicting the position and orientation of the moving object based on operation information and a motion model, improving the overall operability of the remote control device.
Smart Images

Figure JP2024012683_02102025_PF_FP_ABST
Abstract
Description
Remote control system and remote control device
[0001] The present disclosure relates to a remote control system for a moving object, and more particularly to a remote control system that suppresses degradation of operability due to communication delays.
[0002] In remote control systems in which an operator controls a mobile object from a remote location based on images captured by a camera mounted on the object, communication delays can reduce operability, but technology has been developed to improve operability by predicting images.
[0003] For example, Patent Document 1 discloses a technology that uses an environmental information sensor such as a laser range finder to recognize a moving obstacle from a camera image of a moving object, separates the moving obstacle image from a background image, predicts a future image of the moving obstacle based on the speed information of the moving obstacle, and presents an image in which the prediction result is combined with the camera image, thereby suppressing a decrease in operability due to communication delays.
[0004] Patent No. 6054513
[0005] In the technology disclosed in Patent Document 1, the moving body predicts future images of a moving obstacle based on the speed information of the moving obstacle acquired from a camera image, so there is a delay time until the speed information is obtained, which reduces the accuracy of the movement prediction, reduces the accuracy of the image prediction, and therefore reduces operability.
[0006] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a remote control system that suppresses deterioration in operability of a remote control device.
[0007] The remote control system according to the present disclosure is a remote control system for remotely controlling a moving body from a remote control device, wherein the moving body acquires and outputs a first ambient image of the area around the moving body, and the remote control device comprises an operation device for operating the moving body and a remote computer that calculates a predicted image after a predetermined time based on the first ambient image output by the moving body, and the remote computer comprises a three-dimensional scene construction unit that constructs a three-dimensional scene based on the first ambient image and a first depth image generated based on the first ambient image, a moving body movement prediction unit that predicts the position and orientation of the moving body after the predetermined time based on operation information to the operation device and a motion model of the moving body, and a predicted image generation unit that generates the predicted image based on the three-dimensional scene and the position and orientation after the predetermined time.
[0008] According to the remote control system of the present disclosure, since the operation information does not include a delay, by predicting the movement of the moving object based on the operation information and the behavior model of the moving object, the accuracy of predicting the position after a specified time is improved, the accuracy of image prediction is also improved, and the operability of the remote control device is also improved.
[0009] FIG. 1 is a functional block diagram showing a configuration of a remote control system according to a first embodiment of the present disclosure. FIG. 2 is a system configuration diagram showing an example of a schematic configuration when a moving body is an automobile. FIG. 3 is a diagram schematically showing a host vehicle coordinate system. FIG. 4 is a diagram showing the relationship of communication delay between a remote control device and a moving body. FIG. 5 is a schematic diagram showing an example of a procedure for generating a predicted image in a remote computer. FIG. 6 is a functional block diagram showing a configuration of a remote control system according to a second embodiment of the present disclosure. FIG. 7 is a functional block diagram showing a configuration of a remote control system according to a third embodiment of the present disclosure. FIG. 8 is a functional block diagram showing a configuration of a remote control system of a modified example of the third embodiment of the present disclosure. FIG. 9 is a functional block diagram showing a configuration of a remote control system according to a fourth embodiment of the present disclosure. FIG. 10 is a functional block diagram showing a configuration of a remote control system according to a fifth embodiment of the present disclosure. FIG. 11 is a functional block diagram showing a configuration of a remote control system according to a sixth embodiment of the present disclosure. FIG. 12 is a functional block diagram showing a configuration of a remote control system according to a seventh embodiment of the present disclosure. FIG. 13 is a diagram showing a hardware configuration for realizing the remote control systems according to the first to seventh embodiments of the present disclosure.
[0010] 1 is a functional block diagram showing a configuration of a remote control system 1000 according to a first embodiment of the present disclosure. As shown in FIG. 1, the remote control system 1000 includes a moving object 20 and a remote control device 30.
[0011] The mobile body 20 is equipped with a camera 21 that acquires an image of the surroundings of the mobile body 20 (first surrounding image), an internal sensor 22 that acquires the state quantity of the mobile body 20, a wireless device 26 for communicating with the remote control device 30, an actuator 24 for moving the mobile body 20, and a control device 25 for controlling the mobile body 20.
[0012] <Mobile Body> The mobile body 20 is a movable object, and examples thereof include automobiles, personal mobility, mobile robots, drones, ships, airplanes, and helicopters. In the following, an automobile will be used as an example.
[0013] The camera 21 includes a forward-facing camera, and any of a monocular camera, a stereo camera, and a 360-degree camera can be used. When using a monocular camera, a depth image can be generated using known techniques such as Visual SLAM (Simultaneous Localization and Mapping) and neural networks (NN). Compared to a stereo camera, Visual SLAM requires images from multiple viewpoints and also requires estimation of the amount of camera movement between viewpoints, but it reduces the number of cameras, thereby reducing costs. In addition, Visual SLAM using a camera is cheaper than LiDAR (Light Detection and Ranging), and its accuracy has improved due to technological advances, making it rapidly popular.
[0014] The method using a neural network is a method of inputting a monocular video and estimating the depth of the image using a neural network. For example, the method described in "Casser, Vincent, et al. "Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos." Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 33. 2019." can be used.
[0015] When using a stereo camera, a depth image is generated using a known method such as stereo matching. Although the number of cameras increases compared to a monocular camera, the depth can be estimated from an image taken from a single viewpoint, eliminating the need to estimate the amount of camera movement between viewpoints, improving the accuracy of depth image estimation.
[0016] The internal sensor 22 includes a vehicle speed sensor, a steering angle sensor, an acceleration sensor, a yaw rate sensor, etc., and outputs the vehicle speed V, steering angle δ, and longitudinal acceleration a as the sensor value z. x , acceleration in the left and right direction a y , yaw rate γ, and other vehicle state quantities are acquired.
[0017] The control device 25 controls the mobile object 20 based on operation information u from the remote control device 30. The actuators 24 include an EPS motor, an engine, a brake, an EV motor, etc. The wireless device 26 is connected to the camera 21, the internal sensor 22, and the control device 25, and can transmit sensor information acquired by the camera 21 and the internal sensor 22 to the remote control device 30.
[0018] <Remote Control Device> The remote control device 30 includes an operation device 31 for operating the mobile object 20, a display 32, a remote computer 33, and a wireless device 34 for communicating with the mobile object 20. The remote operator operates the operation device 31 to remotely control the mobile object 20 while looking at the display 32 on which the first surrounding image is displayed. However, the communication delay (transmission delay) from the remote control device 30 to the mobile object 20 is L. send , the communication delay (reception delay) from the mobile object 20 to the remote control device 30 is L recv Let's say.
[0019] The display 32 displays the predicted image at a constant interval. The display interval depends on the generation interval of the predicted image and can be shorter than the reception interval of the first surrounding image. This effectively improves the frame rate of the first surrounding image, thereby improving operability. For example, if the first surrounding image is received at 30 frames per second (fps) and the predicted image is generated at 60 fps, the scenery around the moving object appears to the operator to change at 60 fps, making it appear as if the first surrounding image is being received at 60 fps.
[0020] The operation device 31 includes a steering wheel, an accelerator pedal, a brake pedal, a joystick, etc. From the operation information u of these by the operator, a target value to be transmitted to the moving body 20 is calculated by a known method. For example, if the moving body 20 is a vehicle, the steering angle δ of the operation device 30 is calculated by a known method. s , accelerator pedal depression amount p g , brake pedal depression amount p b Based on the above, the target steering angle δ t , target vehicle speed V t Alternatively, the target yaw rate γ t, target acceleration a t Calculate the following:
[0021] If the operation device 31 is a joystick, the target value is calculated from the tilt of the joystick, etc. Although the target value can be calculated from the operation information u by either the remote operation device 30 or the moving body 20, hereinafter, it is assumed that the calculation is performed by the remote operation device 30. When the calculation is performed by the remote operation device 30, the calculated target value is also referred to as operation information u in a broad sense, or simply as operation information u, and operation information u that does not include the target value is referred to as operation information u in a narrow sense. An important point here is that the operation information u does not include communication delays. Therefore, when the moving body 20 calculates the target value from the operation information u in the narrow sense, it is necessary to predict the movement of the moving body 20 using operation information u in the narrow sense that does not include delays.
[0022] In the following description, the operation device 31 is assumed to be a steering wheel, an accelerator pedal, and a brake pedal, and the moving body 20 is set to a target steering angle δ t , target vehicle speed V t The steering angle δ of the operation device 31 is transmitted. s , accelerator pedal depression amount p g , brake pedal depression amount p b From the above, the target steering angle δ t and target vehicle speed V t is calculated, for example, by the following formula (1).
[0023]
[0024] The upper part of the equation (1) is the target steering angle δ t is the steering angle δ s This shows that the steering angle δ of the operation device 31 can be obtained by substituting s The target steering angle δ t This indicates that it can be used as a
[0025] The middle equation of equation (1) is the target acceleration a t is a mathematical formula for calculating k g and k b is the control gain.
[0026] In the lower part of Equation (1), T s is the control period, and the previous target vehicle speed V t and a t T s The current target vehicle speed V t is calculated, and the operation information u is u=[δ t ,V t ]
[0027] The remote computer 33 includes a depth image generation unit 331 , a three-dimensional scene composition unit 332 , a moving object movement prediction unit 333 , a predicted image generation unit 334 , and an output unit 335 .
[0028] The depth image generator 331 generates a first depth image based on the first surrounding image, and the three-dimensional scene constructor 332 constructs a three-dimensional scene based on the first surrounding image and the first depth image.
[0029] The moving body movement prediction unit 333 predicts the position and posture of the moving body 20 after a predetermined time based on operation information to the operation device 31 and a motion model of the moving body 20, the predicted image generation unit 334 generates a predicted image after a predetermined time based on the three-dimensional scene and the position and posture after the predetermined time, and the output unit 335 outputs the predicted image to the display 32.
[0030] The function of the remote computer 33 is to 0 Operation information u, which is information that can be used in 0:t0 , sensor value z0:t0-Lrecv and image I 0:t0-Lrecv From time t 0 -L recv than the predetermined time L pred The goal is to predict future images and present them to the remote operator.
[0031] The depth image generating unit 331 calculates the communication delay L from the past (t=0) to the present. recv the past (t = t 0 -L recv ) Time series data I 0:t0-Lrecv Using this, a depth image D is generated by a known method such as Visual SLAM or NN. 0:t0-Lrecv Generate.
[0032] The three-dimensional scene construction unit 332 converts a point (u, v) on the first ambient image I into a point (X img ,Y img ,Z img ) to construct a three-dimensional scene. Here, the definition of three-dimensional space is arbitrary, but from now on, we will use the time t = t 0 -L recv The coordinate system is considered with the center of gravity of the moving body 20 at the origin, the front being the positive X-axis, the left being the positive Y-axis, and the top being the positive Z-axis.
[0033] It is also possible to store a 3D map of the area around the moving object as a database on the remote control device 30 or on the cloud, and to construct a 3D scene based on that information. In particular, for static targets such as roads and buildings, using a 3D map from a database can improve accuracy and reduce blind spots.
[0034] The moving object movement prediction unit 333 calculates the time from the past (t=0) to the present (t=t 0 ) Time series data of operation information up to 0:t0 and the communication delay L from the past (t=0) to the present. recv the past (t = t 0 -L recv ) and the motion model dx / dt=f(x, u) of the moving object 20, 0 -L recv From a predetermined time L pred After t = t 0 -L recv +L pred The state quantity xt0-Lrecv+Lpred including the position and orientation at
[0035] A known model is used as the motion model, and in the case of a vehicle, a geometric model or a two-wheel model can be used, but the motion model can also be expressed using a neural network.
[0036] When using a neural network, if part of the state quantity x is estimated from a sensor value, operation information u and the state quantity x can be collected and an operation model can be trained offline or online based on the time-series data. This improves the accuracy of the operation model and also improves the accuracy of movement prediction. In particular, if the mobile object 20 is equipped with a global navigation satellite system (GNSS) and LiDAR, and the sensor value z includes position and orientation information, the accuracy of the position and orientation of the state quantity x is improved. Therefore, if this is used as training data to train the neural network for the operation model, the accuracy of movement prediction is greatly improved.
[0037] The operation information u is a target value to be transmitted to the moving body 20. In the case of a vehicle, for example, a target steering angle δ t and target vehicle speed V t Alternatively, the target yaw rate γ t and target acceleration a t etc.
[0038] The state quantity x is the horizontal position (X g ,Y g ), and the yaw angle θ. Here, the origin of the XY coordinate system is at time t = t 0 -L recv The center of gravity of the moving body 20 is the position at the X axis, and the front is the positive X axis and the left is the positive Y axis. g,t0-Lrecv ,Yg, t0-Lrecv ,θ t0-Lrecv is set to 0. When considering three-dimensional movement such as drones, the three-dimensional position and roll, pitch, and yaw are included.
[0039] The state quantity x can use a sensor value acquired by the internal sensor 22, but can also be estimated by a known estimation method such as a Kalman filter. The Kalman filter is a method for identifying the position of a moving object with a smaller error by combining a sensor value and an equation of motion.
[0040] The target steering angle δ is set to the moving body 20. t and target vehicle speed V tWhen transmitting the target acceleration a to the moving body 20, if the steering angle tracking ability and the speed tracking ability of the moving body 20 are known, the steering angle δ and the vehicle speed V can be estimated with high accuracy only from the operation information u without using the sensor values acquired by the internal sensor 22. t When transmitting the vehicle speed V, if an attempt is made to estimate the vehicle speed V from only the operation information u, an integral error will be accumulated, so it is desirable to also use the vehicle speed V acquired by the internal sensor 22.
[0041] Predetermined time L pred For example, L pred =L recv +L send As will be explained later with reference to FIG. t The time when this is reflected is t = t 0 +L send and time t=t 0 +L send This is because the operability is improved by presenting the image in u to the remote operator, and the information available at the current time is 0:t0 , x0:t0-Lrecv, and x0:t0-Lrecv to x t0+Lsend This takes into consideration the need to predict the
[0042] The key point here is the operation information 0:t0 The prediction is made using the operation information u 0:t0 does not include communication delays, so x t0+Lsend However, the predetermined time L pred can be set to other values, and the predetermined time L pred Since there is a trade-off between the length and accuracy of , it can be set smaller.
[0043] Communication delay L recv and L send If is constant, it can be set to a constant value, and if it fluctuates, the communication delay L recv and L send The delay in receiving the sensor value L can be a variable to be estimated. recv,sens and image reception delay L recv,img If it is different from L, the purpose is to compensate for the delay of the image. pred≦L recv,img +L send and L recv,img Set the predicted time based on the
[0044] When a geometric model is used as the motion model, for example, the state quantity is expressed as [X g ,Y g ,θ,V,δ]. In this case, u=[δ t ,V t ], the state equation (operation model) is given by the following equation (2).
[0045]
[0046] In formula (2), T d,V and T d,δ is the time constant when the speed tracking ability and the steering angle tracking ability are modeled as a first-order delay system. β is the sideslip angle, and γ is the yaw rate, which are given by the following equations (3) and (4), respectively.
[0047]
[0048]
[0049] However, l f and l r are the distances from the center of gravity of the vehicle to the front axle and rear axle, respectively, and l is the wheelbase, then l = l f +l r is.
[0050] The prediction of xt0-Lrecv+Lpred is to set the initial value of the state quantity as x t0-Lrecv and calculates it by integrating the state equation using the time series of operation information ut0-(Lrecv+Lsend):t0-(Lrecv+Lsend)+Lpred. That is, it is calculated by the following formula (5).
[0051]
[0052] However, L pred ≦L recv +L send In addition, in the formula (5), x t0-Lrecv Although X is used t0-Lrecv ,Y t0-Lrecv ,θ t0-Lrecvis set to 0 in each control cycle, and the steering angle δ and the vehicle speed V can be estimated from the operation information u, so the state quantity xt0−Lrecv+Lpred can be estimated without using the sensor value z.
[0053] The predicted image generation unit 334 moves the viewpoint p of the camera 21 based on the predicted state quantity xt0-Lrecv+Lpred of the moving body 20, and generates a predicted image lt0-Lrecv+Lpred when the three-dimensional scene S is viewed from the viewpoint pt0-Lrecv+Lpred after the movement using a known method such as perspective projection.
[0054] As described above, the generation cycle of the predicted image can be made shorter than the reception cycle of the first surrounding image. In other words, the operation cycle of the predicted image generation unit 334 can be upsampled. This allows the frame rate of the image presented to the remote operator to be upsampled, improving operability. In this case, the operation cycle of the mobile object movement prediction unit 333 can also be upsampled, and the prediction time can be adjusted to be consistent with the upsampling. The predicted image lt0-Lrecv+Lpred generated by the predicted image generation unit 334 is displayed on the display 32 via the output unit 335.
[0055] 2 is a system configuration diagram showing an example of a schematic configuration of a vehicle 1 when the moving body 20 is an automobile. As shown in FIG. 2, the vehicle 1 includes a steering wheel 2, a steering shaft 3, a steering unit 4, an EPS (Electric Power Steering) motor 5, a power train unit 6, and a brake unit 7 as a drive system.
[0056] The sensor system also includes a forward camera 111, a radar sensor 112, a GNSS sensor 121, a steering angle sensor 131, a steering torque sensor 132, a yaw rate sensor 133, a speed sensor 134, and an acceleration sensor 135.
[0057] In addition to these, the vehicle is equipped with a navigation device 122, a vehicle control unit 200, an EPS controller 311, a powertrain controller 312, and a brake controller 313.
[0058] The steering wheel 2, which is installed so that the driver can operate the vehicle 1, is coupled to a steering shaft 3. A steering unit 4 is connected to the steering shaft 3. The steering unit 4 rotatably supports two front tires as steered wheels and is steerably supported on the vehicle frame. Therefore, torque generated by the driver's operation of the steering wheel 2 rotates the steering shaft 3, and the steering unit 4 steers the front wheels left and right. This allows the driver to control the amount of lateral movement of the vehicle 1 when moving forward and backward. The steering shaft 3 can also be rotated by an EPS motor 5, and by controlling the current flowing through the EPS motor 5 with an EPS controller 311, the front wheels can be steered freely independent of the driver's operation of the steering wheel 2.
[0059] The vehicle control unit 200 is an integrated circuit such as a microprocessor, and includes an A / D (Analog / Digital) conversion circuit, a D / A (Digital / Analog) conversion circuit, a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc.
[0060] The vehicle control unit 200 is connected to a forward camera 111, a radar sensor 112, a GNSS sensor 113, a navigation device 122, a steering angle sensor 131 that detects the steering angle, a steering torque sensor 132 that detects the steering torque, a yaw rate sensor 133 that detects the yaw rate, a speed sensor 134 that detects the speed of the vehicle 1, an acceleration sensor 135 that detects the acceleration of the vehicle 1, an EPS controller 311, a powertrain controller 312, and a brake controller 313.
[0061] The vehicle control unit 200 processes the information input from the connected sensors according to a program stored in ROM, transmits the target steering angle to the EPS controller 311, and transmits the target acceleration to the powertrain controller 312 and the brake controller 313.
[0062] The front camera 111 is installed in a position where it can detect the lane markings ahead of the vehicle as an image, and based on the image information, it detects the environment ahead of the vehicle 1, such as lane information and the location of obstacles. Note that, although the present embodiment has exemplified only a camera that detects the environment ahead of the vehicle 1, cameras that detect the environment behind and to the sides may also be installed.
[0063] The radar sensor 112 emits radar and detects the reflected waves, thereby outputting the relative distance and relative speed to an obstacle present around the vehicle 1. As this distance measurement sensor, a well-known type of sensor such as a millimeter wave radar, LiDAR, laser range finder, or ultrasonic radar can be used.
[0064] The GNSS sensor 121 receives radio waves from positioning satellites with an antenna, performs positioning calculations, and outputs the absolute position and absolute direction of the vehicle 1 .
[0065] The navigation device 122 has a function of calculating the optimum driving route for the destination set by the driver, and stores road information on the driving route. The road information is map node data that represents the road alignment, and each map node data incorporates information such as latitude, longitude, and altitude that indicate the absolute position at each node, as well as lane width, cant angle, and inclination angle information.
[0066] The EPS controller 311 controls the EPS motor 5 based on the target steering angle transmitted from the vehicle control unit 200 , thereby controlling the traveling trajectory of the vehicle 1 .
[0067] The powertrain controller 312 controls the powertrain unit 6 so as to achieve the target acceleration transmitted from the vehicle control unit 200 .
[0068] In this embodiment, a vehicle that uses only an engine as a driving force source is given as an example, but this embodiment can also be applied to vehicles that use only an electric motor as a driving force source, vehicles that use both an engine and an electric motor as driving force sources, etc.
[0069] The brake controller 313 controls the brake unit 7 so as to achieve the target acceleration transmitted from the vehicle control unit 200 .
[0070] 3 is a diagram showing a schematic representation of the vehicle coordinate system used in this embodiment. That is, the X-axis and Y-axis in FIG. 3 represent the inertial coordinate system. g ,Y g , θ represent the center of gravity position and vehicle body direction of the vehicle 1 in the inertial coordinate system. The x-axis and y-axis in FIG. 3 are the vehicle coordinate system with the center of gravity of the vehicle 1 as the origin, the x-axis pointing forward, and the y-axis pointing leftward. The angle θ is set with the positive direction of the x-axis as the reference direction, and the counterclockwise direction as positive. Here, the center of gravity position X of the vehicle 1 is calculated for each execution cycle. g ,Y g and the vehicle body orientation θ are initialized to 0. In other words, the inertial coordinate system and the host vehicle coordinate system are made to coincide with each other for each execution cycle.
[0071] 4 is a schematic diagram showing the relationship of communication delay between the remote control device 30 and the mobile object 20, with the horizontal axis representing time. As shown in FIG. 4, the remote control device 30 transmits operation information u to the mobile object 20, and receives a sensor value z and an image I from the mobile object 20.
[0072] The communication delay from the remote control device 30 to the mobile object 20 is L send The communication delay in receiving the signal at the remote control device 30 is L recv is.
[0073] There is a communication delay L during transmission between the remote control device 30 and the mobile object 20. send and communication delay L when receiving recv By L send +L recv delays will occur.
[0074] Current time t=t 0 The information available in the operation information u 0:t0 , sensor value z0: t0-Lrecv, state quantity x0: t0-Lrecv, and image I 0:t0-Lrecv is.
[0075] Latest operation information t0 is later t=t 0 +L send Since the image is reflected on the moving object 20 at t0-Lrecv From I t0-Lrecv+Lsend It is desirable to predict the time Lpred is at most L send +L recv Let's say.
[0076] 5 is a schematic diagram showing an example of a procedure for generating a predicted image in the remote computer 33 (FIG. 1). A forward image, which is a first surrounding image acquired by the camera 21 of the moving body 20, is input to the depth image generating unit 331 from the wireless device 26 of the moving body 20 via the wireless device 34 of the remote control device 30. As explained using FIG. 4, the communication delay in reception at the remote control device 30 is L. recv Therefore, the forward image is the time series data I 0:t0-Lrecv The depth image generation unit 331 generates a depth image D as a first depth image by a known method such as Visual SLAM or NN, as described above. t0-Lrecv Generate and output.
[0077] Output depth image D t0-Lrecv is input to the three-dimensional scene construction unit 332, and the three-dimensional scene S is constructed using the original front image and depth image. t0-Lrecv The three-dimensional scene construction unit 332 constructs (reconstructs) and outputs the three-dimensional scene S using the corresponding depth image data and front image data, using a known three-dimensional scene generation method. t0-Lrecv Configure.
[0078] Output 3D scene S t0-Lrecv is input to the predicted image generating unit 334, and the viewpoint of the camera 21 is moved based on the state quantities including the predicted position and orientation of the moving object 20, and the three-dimensional scene S is generated by a known method such as perspective projection. t0-Lrecv A predicted image is generated as seen from the viewpoint after the movement.
[0079] 5, in order to generate a predicted image for a three-dimensional scene around a moving object represented by a cube, a viewpoint represented by a triangular pyramid is moved from left to right according to a movement prediction. This movement of the moving object is predicted by the moving object movement prediction unit 333 based on the state quantity x0:t0-Lrecv and the operation information u 0:t0The state quantity obtained by the moving body movement prediction unit 333 is input to the predicted image generation unit 334.
[0080] As described above, in the remote operation system 1000 of the first embodiment, state quantities including the position and attitude of the moving object 20 at a predetermined time are predicted based on the motion model of the moving object 20 using operation information of the operator via the operation device 31 of the remote operation device 30 and state quantities past the communication delay. Because the operation information does not include communication delay, movement prediction based on the motion model improves the prediction accuracy of the movement position and also improves the accuracy of image prediction. This also improves the operability of the remote operation device 30.
[0081] Second Embodiment Fig. 6 is a functional block diagram showing a configuration of a remote control system 2000 according to a second embodiment of the present disclosure. As shown in Fig. 6, the remote control system 2000 includes a moving body 20A and a remote control device 30A. In Fig. 6, the same components as those in the remote control system 1000 described using Fig. 1 are denoted by the same reference numerals, and redundant description will be omitted.
[0082] The mobile body 20A further includes an external sensor 23 and a mobile body computer 27 in addition to the configuration of the mobile body 20 of the first embodiment shown in FIG.
[0083] The external sensor 23 is a sensor that acquires distance information (first distance information) to targets around the moving body 20A, and is a sensor that can measure the distance to targets, such as LiDAR and millimeter wave radar.
[0084] The mobile computer 27 has a depth image generation unit 271 that generates a first depth image based on the first surrounding image input from the camera 21 and first distance information input from the external sensor 23. The depth image generation unit 271 uses the first surrounding image and the first distance information to generate the first depth image by a known method such as coordinate transformation and perspective projection.
[0085] The accuracy of the depth image is improved by using the first distance information acquired by the external sensor 23. However, since the point cloud information obtained from LiDAR or the like has a larger data volume than the depth image, by generating the depth image on the mobile body side and sending it to the remote control device side, the amount of communication can be reduced compared to when the first distance information acquired by the external sensor 23 is sent to the remote control device side.
[0086] The remote control device 30A has a remote computer 33A, but the remote computer 33A does not have a depth image generation unit 331 like the remote control device 30 of the first embodiment shown in Fig. 1. The first depth image generated by the moving body 20A is input to the three-dimensional scene construction unit 332 of the remote computer 33A from the wireless device 26 of the moving body 20A via the wireless device 34 of the remote control device 30A, and a three-dimensional scene is constructed in the three-dimensional scene construction unit 332.
[0087] As described above, in the remote control system 2000 of the second embodiment, the operability of the remote control device 30A is improved as in the remote control system 1000 of the first embodiment, and the accuracy of the depth image is improved by using the first distance information acquired by the external sensor 23. Furthermore, by generating a depth image on the mobile object side and sending it to the remote control device side, the amount of communication can be reduced compared to when the first distance information is sent to the remote control device side.
[0088] Third Embodiment Fig. 7 is a functional block diagram showing a configuration of a remote control system 3000 according to a third embodiment of the present disclosure. As shown in Fig. 7, the remote control system 3000 includes a mobile object 20, a remote control device 30A, and a roadside unit 40. In Fig. 7, the same components as those in the remote control system 1000 described using Fig. 1 are denoted by the same reference numerals, and redundant description will be omitted.
[0089] The roadside unit 40 includes a roadside computer 41, a roadside sensor 42 that acquires an image of the surroundings of the roadside unit 40 (a second surrounding image) and distance information (second distance information) to surrounding landmarks of the roadside unit 40, and a wireless device 43 for communicating with the mobile body 20 and the remote control device 30A.
[0090] The roadside computer 41 is equipped with a depth image generation unit 411 that generates a first depth image and a second depth image corresponding to the first surrounding image and the second surrounding image, respectively, based on the first surrounding image, the second surrounding image, and the second distance information input from the wireless device 26 of the mobile body 20 via the wireless device 43.
[0091] The roadside sensor 42 is a sensor that acquires images of the surroundings of the roadside unit 40 and information on the distance to targets in the surroundings of the roadside unit 40, and includes cameras such as a monocular camera, a stereo camera, and a 360-degree camera, and sensors that measure the distance to targets such as LiDAR and millimeter-wave radar.
[0092] The roadside unit 40 is installed at the side of a road (roadside) on which vehicles such as the mobile object 20 travel, and is an advanced device that links vehicles with camera information and the like using V2I (Vehicle to Infrastructure) technology, and is linked not only with the vehicle but also with the remote control device 30A via a wireless device 43. When the mobile object 20 enters the detection range of the roadside sensor 42 of the roadside unit 40, the roadside unit 40 acquires an image of the surroundings of the roadside unit 40 and information on distances to objects in the surroundings of the roadside unit 40, and generates a first depth image and a second depth image.
[0093] The remote operation device 30A has a remote computer 33A, but the remote computer 33A does not have a depth image generation unit 331 like the remote operation device 30 of embodiment 1 shown in Fig. 1. The first depth image and the second depth image generated by the depth image generation unit 411 of the roadside unit 40 are input to the 3D scene construction unit 332 of the remote computer 33A from the wireless device 43 of the roadside unit 40 via the wireless device 34 of the remote operation device 30A, and a 3D scene is constructed in the remote computer 33A.
[0094] As described above, in the remote control system 3000 of the third embodiment, the operability of the remote control device 30A is improved, as in the remote control system 1000 of the first embodiment. In addition, by using the surrounding images and distance information of the roadside unit 40, it is possible to construct a three-dimensional scene of an area that is a blind spot of the camera 21 of the moving object 20, thereby improving the accuracy of image prediction and the operability of the remote control device 30A.
[0095] <Modification> Figure 8 is a functional block diagram showing the configuration of a remote control system 3001 according to a modification of the third embodiment of the present disclosure. As shown in Figure 8, the remote control system 3001 includes a mobile object 20, a remote control device 30, and a roadside unit 40A. In Figure 8, the same components as those in the remote control system 1000 described using Figure 1 are denoted by the same reference numerals, and duplicated descriptions will be omitted.
[0096] The roadside unit 40A of the remote control system 3001 shown in FIG. 8 includes a roadside sensor 42 that acquires a surrounding image (second surrounding image) of the roadside unit 40A and distance information (second distance information) to surrounding landmarks of the roadside unit 40A, and a wireless device 43 that communicates with the mobile object 20 and the remote control device 30.
[0097] The roadside sensor 42 is a sensor that acquires images of the surroundings of the roadside unit 40A and information on the distance to targets in the surroundings of the roadside unit 40A, and includes cameras such as a monocular camera, a stereo camera, and a 360-degree camera, and sensors that measure the distance to targets such as LiDAR and millimeter-wave radar.
[0098] The roadside unit 40 outputs the second surrounding image and second distance information acquired by the roadside sensor 42 from the wireless device 43 to the depth image generator 331 via the wireless device 34 of the remote control device 30. The depth image generator 331 generates a first depth image and a second depth image corresponding to the first surrounding image and the second surrounding image, respectively, based on the first surrounding image input from the wireless device 26 of the mobile object 20 via the wireless device 34, and the second surrounding image and second distance information input from the roadside sensor 42 via the wireless device 34.
[0099] The three-dimensional scene construction unit 332 of the remote computer 33 receives the first depth image and the second depth image from the depth image generation unit 331 and constructs a three-dimensional scene.
[0100] As described above, in the remote operation system 3001 according to the modification of the third embodiment, the roadside unit 40A simply outputs the second surrounding image and the second distance information acquired by the roadside sensor 42 to the remote operation device 30. Unlike the roadside unit 40 of the remote operation system 3000 according to the third embodiment, the roadside unit 40A does not have the roadside computer 41. Therefore, it is possible to use a known roadside unit, thereby reducing the cost of building the system.
[0101] <Fourth embodiment> Fig. 9 is a functional block diagram showing the configuration of a remote control system 4000 according to a fourth embodiment of the present disclosure. As shown in Fig. 9, the remote control system 4000 includes a moving object 20A, a remote control device 30A, and a roadside unit 40. In Fig. 9, the same components as those in the remote control system 2000 described using Fig. 6 are denoted by the same reference numerals, and redundant description will be omitted.
[0102] The roadside unit 40 includes a roadside computer 41, a roadside sensor 42 that acquires a surrounding image of the roadside unit 40 (second surrounding image) and distance information (second distance information) to surrounding landmarks of the roadside unit 40, and a wireless device 43 that communicates with the mobile body 20A and the remote control device 30A.
[0103] The roadside computer 41 includes a depth image generator 411 that generates a second depth image corresponding to the second surrounding image based on the second surrounding image and the second distance information.
[0104] The roadside sensor 42 is a sensor that acquires images of the surroundings of the roadside unit 40 and information on the distance to targets in the surroundings of the roadside unit 40, and includes cameras such as a monocular camera, a stereo camera, and a 360-degree camera, and sensors that measure the distance to targets such as LiDAR and millimeter-wave radar.
[0105] The depth image generating unit 271 of the moving body 20A generates a first depth image based on the first surrounding image input from the camera 21 and the first distance information input from the external sensor 23 .
[0106] The remote control device 30A has a remote computer 33A, but the remote computer 33A does not have a depth image generation unit 331 like the remote control device 30 of the first embodiment shown in Fig. 1. The first depth image generated by the moving body 20A is input to the three-dimensional scene construction unit 332 of the remote computer 33A from the wireless device 26 of the moving body 20A via the wireless device 34 of the remote control device 30A, and a three-dimensional scene is constructed in the three-dimensional scene construction unit 332.
[0107] In addition, the second depth image generated by the roadside unit 40 is input to the three-dimensional scene construction unit 332 from the wireless device 43 of the roadside unit 40 via the wireless device 34 of the remote control device 30A, and a three-dimensional scene is constructed in the three-dimensional scene construction unit 332.
[0108] As described above, in the remote control system 4000 of the fourth embodiment, the operability of the remote control device 30A is improved as in the remote control system 1000 of the first embodiment, and the accuracy of the first depth image is improved by using the first distance information acquired by the external sensor 23. Furthermore, by generating a depth image on the mobile object side and sending it to the remote control device side, the amount of communication can be reduced compared to when the first distance information is sent to the remote control device side.
[0109] Furthermore, by utilizing the surrounding image and distance information from the roadside unit 40, a three-dimensional scene of the blind spot of the camera 21 of the moving object 20A can be constructed, improving the accuracy of image prediction and the operability of the remote control device 30A.
[0110] Fifth Embodiment Fig. 10 is a functional block diagram showing a configuration of a remote control system 5000 according to a fifth embodiment of the present disclosure. As shown in Fig. 10, the remote control system 5000 includes a moving object 20 and a remote control device 30B. In Fig. 10, the same components as those in the remote control system 1000 described using Fig. 1 are denoted by the same reference numerals, and redundant description will be omitted.
[0111] The remote control device 30B includes an operation device 31, a display 32, a remote computer 33B, and a wireless device 34. The remote computer 33B further includes a dynamic target movement prediction unit 336 in addition to the configuration of the remote computer 33 of the remote control device 30 shown in FIG.
[0112] The dynamic target movement prediction unit 336 recognizes a dynamic target using the three-dimensional scene output from the three-dimensional scene construction unit 332, estimates the speed of the dynamic target, predicts the dynamic target movement prediction result, which is the position and attitude of the dynamic target after a predetermined time, and inputs it to the predicted image generation unit 334.
[0113] In addition, instead of a three-dimensional scene, a dynamic target can also be recognized using the first depth image output from the depth image generation unit 331, or a dynamic target can also be recognized using the first surrounding image.
[0114] Dynamic targets are objects that may move, and include, for example, vehicles, people, animals, etc. Static targets include roads, buildings, etc. that do not move.
[0115] For the recognition of dynamic targets, known techniques such as image processing and neural networks can be used. For example, the position and type of targets in an image can be estimated by YOLO (You Only Look Once) or the like.
[0116] YOLO is a representative image analysis method that utilizes artificial intelligence (AI), and uses a convolutional neural network (CNN) to simultaneously detect and identify multiple targets in an image with a single image scan.
[0117] The position and attitude of the dynamic target can be predicted by a known method in the dynamic target movement prediction unit 336, such as predicting that the recognized dynamic target is moving at a constant velocity while maintaining its current speed. In addition, when the dynamic target moves on a road, it can be predicted that the dynamic target is moving along the road, or that the dynamic target is turning from the change in the azimuth θ when the vehicle turns.
[0118] Furthermore, if the moving body 20 is equipped with a millimeter wave radar or the like, the speed of the dynamic target can be estimated from the millimeter wave radar, and the estimated result can be input from the wireless device 26 of the moving body 20 to the dynamic target movement prediction unit 336 via the wireless device 34 of the remote control device 30B, thereby determining the speed of the dynamic target.
[0119] The predicted image generating unit 334 generates a predicted image for a predetermined time later based on the dynamic target movement prediction result input from the dynamic target movement predicting unit 336 .
[0120] As described above, in the remote operation system 5000 of the fifth embodiment, the predicted image generation unit 334 generates a predicted image for a predetermined time period after the predetermined time based on the input dynamic target movement prediction result. That is, the predicted image is generated by moving the dynamic target in the three-dimensional scene based on the dynamic target movement prediction result. Moving the dynamic target improves the accuracy of image prediction, and improves the operability of the remote operation device 30B.
[0121] Sixth Embodiment Fig. 11 is a functional block diagram showing a configuration of a remote control system 6000 according to a sixth embodiment of the present disclosure. As shown in Fig. 11, the remote control system 6000 includes a moving body 20B and a remote control device 30. In Fig. 11, the same components as those in the remote control system 1000 described using Fig. 1 are denoted by the same reference numerals, and redundant description will be omitted.
[0122] The moving body 20B further includes a moving body computer 28 in addition to the configuration of the moving body 20 of the first embodiment shown in Fig. 1. The moving body computer 28 has a target state estimation unit 281 that estimates the state of surrounding targets based on information of the first surrounding image input from the camera 21.
[0123] The state of surrounding targets includes the position, type, color, etc. of the target, and the state of surrounding targets can be estimated using known techniques such as image processing and neural networks. For example, the position, type, and color of targets in an image can be estimated using YOLO, as described above. Furthermore, by using YOLO, multiple targets in an image can be simultaneously detected and estimated with a single image scan.
[0124] The data (second data) including the states of the surrounding targets estimated by the target state estimation unit 281 is transmitted to the remote control device 30 via the wireless device 26 so as to reduce communication delay compared to the data (first data) including the first surrounding image acquired by the camera 21 of the mobile object 20. In other words, the second data including the states of the surrounding targets is sent quickly.
[0125] A method for shortening the communication delay of the second data is to separate packets and lines, and transmit the first data and the second data in separate packets or over separate lines.
[0126] The second data including the states of the surrounding targets estimated by the target state estimation unit 281 is input to a predicted image generation unit 334 of the remote computer 33 via the wireless device 34, and the predicted image generation unit 334 generates a predicted image after a predetermined time based on the states of the surrounding targets as well.
[0127] As described above, in the remote control system 6000 of the sixth embodiment, the predicted image generating unit 334 generates a predicted image after a predetermined time based on the states of surrounding targets estimated by the target state estimating unit 281 of the moving body 20B. Therefore, targets not shown in the surrounding image can be reflected in the predicted image, and discontinuously changing information such as traffic signal displays can be reflected in the predicted image.
[0128] Furthermore, by quickly transmitting only information on the estimated state of surrounding targets to the remote control device 30, discontinuous changes in targets not shown in the first surrounding image, such as the appearance of a new obstacle or a change in the color of a traffic light, can be reflected in the predicted image, thereby improving safety.
[0129] Note that transmitting the second data not including the first surrounding image via the wireless device 26 so that the communication delay is shorter than that of the first data including the first surrounding image is not limited to the remote control system 6000 of the sixth embodiment, but can be applied to any of the remote control systems of the first to fifth embodiments. Furthermore, the second data is not limited to data including the state of surrounding targets.
[0130] Seventh Embodiment Figure 12 is a functional block diagram showing a configuration of a remote control system 7000 according to a seventh embodiment of the present disclosure. As shown in Figure 12, the remote control system 7000 includes a moving object 20 and a remote control device 30C. In Figure 12, the same components as those in the remote control system 1000 described using Figure 1 are denoted by the same reference numerals, and redundant description will be omitted.
[0131] The remote control device 30C includes an operation device 31, a display 32, a remote computer 33C, a wireless device 34, and a database 35. The remote computer 33C further includes a target type recognition unit 337 in addition to the configuration of the remote computer 33 of the remote control device 30 shown in FIG.
[0132] The database 35 stores data related to the texture of targets in general. Textures include CG, polygons, images, etc. Images can also be referenced from sites such as Google (registered trademark) Street View.
[0133] The target type recognition unit 337 recognizes the type of target in the surrounding image or the image present in the three-dimensional scene. The type can be roughly classified as "vehicle," "pedestrian," or "building," or can be classified as proper nouns, such as the type of vehicle or the name of a building.
[0134] The type of target in the image can be recognized by known techniques such as image processing and neural networks. For example, the type of target in the image can be recognized by the YOLO algorithm described above. Furthermore, by using YOLO, multiple targets in the image can be simultaneously detected and estimated with a single image scan.
[0135] The predicted image generation unit 334 assigns textures stored in the database 35 to the three-dimensional scene constructed by the three-dimensional scene construction unit 332 based on the type of target input from the target type recognition unit 337, thereby complementing blind spots in the three-dimensional scene and generating a predicted image.
[0136] For example, textures such as CG are applied to vehicles, roads, guardrails, buildings, etc. classified by the target type recognition unit 337. Also, instead of CG, a published texture of the entire target is applied.
[0137] When the target classified by the target type recognition unit 337 is a vehicle and only the rear part of the vehicle is visible in the three-dimensional scene, CG is assigned to the other parts of the vehicle that are blind spots, thereby complementing the scene so that the entire vehicle can be seen.
[0138] Furthermore, if the target classified by the target type recognition unit 337 is a vehicle and there are parts of the three-dimensional scene that are hidden by the vehicle and cannot be seen, the hidden parts are supplemented with images from Google (registered trademark) Street View.
[0139] The display of the completed areas can be changed to show them in grayscale or hatched to indicate low accuracy.
[0140] Furthermore, when the predicted image includes a blind spot, the predicted image generation unit 334 can input the predicted image including the blind spot into a trained neural network model without using a database or a target type recognition unit, thereby complementing the blind spot, thereby simplifying the configuration of the remote control device 30.
[0141] As described above, in the remote operation system 7000 of the seventh embodiment, targets are classified by the target type recognition unit 337, and textures stored in the database 35 are assigned to the three-dimensional scene constructed by the three-dimensional scene construction unit 332. This makes it possible to generate a predicted image for areas that would be blind spots if only image information were used, thereby improving prediction accuracy and the operability of the remote operation device 30C.
[0142] <Hardware Configuration> Each of the components of the remote computers 33 to 33C, the mobile computer 27, and the roadside computer 41 in the first to seventh embodiments described above can be configured using a computer, and is realized by the computer executing a program. That is, the remote computers 33 to 33C, the mobile computer 27, and the roadside computer 41 are realized, for example, by a processing circuit 500 shown in Fig. 13. A processor such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor) is applied to the processing circuit 500, and the functions of each part are realized by executing a program stored in a storage device.
[0143] Dedicated hardware may be applied to the processing circuit 500. When the processing circuit 500 is dedicated hardware, the processing circuit 500 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a combination thereof.
[0144] The functions of the remote computers 33 to 33C, the mobile computers 27 and 28, and the roadside computer 41 may be realized by individual processing circuits, or the functions may be realized together by a single processing circuit.
[0145] 14 also shows a hardware configuration in the case where the processing circuit 500 is configured using a processor. In this case, the functions of each of the remote computers 33 to 33C, the mobile computers 27 and 28, and the roadside computer 41 are realized by a combination of software, etc. (software, firmware, or software and firmware). The software, etc. is written as a program and stored in memory 520. The processor 510 functioning as the processing circuit 500 realizes the functions of each of the components by reading and executing the program stored in memory 520 (storage device). In other words, it can be said that this program causes a computer to execute the procedures and methods of operation of the components of the remote computers 33 to 33C, the mobile computers 27 and 28, and the roadside computer 41.
[0146] Here, the memory 520 may be, for example, a non-volatile or volatile semiconductor memory such as RAM, ROM, flash memory, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), HDD (Hard Disk Drive), magnetic disk, flexible disk, optical disk, compact disk, mini disk, DVD (Digital Versatile Disc) and its drive device, or any storage medium that will be used in the future.
[0147] The above describes a configuration in which the functions of the components of the remote computers 33 to 33C, the mobile computers 27 and 28, and the roadside computer 41 are realized either by hardware or software, etc. However, this is not a limitation, and a configuration in which some of the components of the remote computers 33 to 33C, the mobile computers 27 and 28, and the roadside computer 41 are realized by dedicated hardware and other components are realized by software, etc. For example, it is possible to realize the functions of some of the components by the processing circuit 500 as dedicated hardware, and to realize the functions of other components by the processing circuit 500 as the processor 510 reading and executing a program stored in the memory 520.
[0148] As described above, the remote computers 33 to 33C, the mobile computers 27 and 28, and the roadside computer 41 can realize the above-mentioned functions by hardware, software, or a combination of these.
[0149] Although the present disclosure has been described in detail, the above description is illustrative in all respects and does not limit the disclosure thereto. It is understood that countless variations not illustrated can be envisioned without departing from the scope of the present disclosure.
[0150] It should be noted that, within the scope of the present disclosure, the embodiments can be freely combined, modified, or omitted as appropriate.
Claims
1. A remote control system for remotely controlling a moving object from a remote control device, wherein the moving object acquires and outputs a first surrounding image around the moving object, the remote control device comprises: a control device for controlling the moving object; and a remote computer for calculating a predicted image after a predetermined time based on the first surrounding image output by the moving object, the remote computer comprising: a three-dimensional scene construction unit for constructing a three-dimensional scene based on the first surrounding image and a first depth image generated based on the first surrounding image; a moving object movement prediction unit for predicting the position and orientation of the moving object after the predetermined time based on operation information to the control device and a motion model of the moving object; and a predicted image generation unit for generating the predicted image based on the three-dimensional scene and the position and orientation after the predetermined time.
2. The remote control system of claim 1, wherein the remote computer has a depth image generator that generates the first depth image based on the first ambient image, and the first depth image is input to the three-dimensional scene construction unit.
3. The remote control system of claim 1, wherein the mobile body is equipped with a mobile body calculator, acquires first distance information to targets in the surroundings of the mobile body and inputs it to the mobile body calculator, the mobile body calculator has a depth image generation unit that generates the first depth image based on the first surrounding image and the first distance information, and the first depth image is transmitted to the remote control device.
4. The remote control system according to claim 1, wherein the remote control system comprises a roadside unit, the roadside unit acquires a second surrounding image of the surroundings of the roadside unit, and the three-dimensional scene construction unit constructs the three-dimensional scene based on the second surrounding image and also on a second depth image generated based on the second surrounding image.
5. The remote control system according to claim 1, wherein the remote control system comprises a roadside unit, the roadside unit having a roadside computer, which acquires a second surrounding image of the roadside unit and inputs it to the roadside computer, the roadside computer having a depth image generation unit which generates the first depth image and the second depth image corresponding to the first surrounding image and the second surrounding image, respectively, based on the first surrounding image and the second surrounding image output by the mobile object, the first depth image and the second depth image are transmitted to the remote control device, and the three-dimensional scene construction unit constructs the three-dimensional scene also based on the second depth image.
6. A remote control system as described in claim 4, wherein the mobile body is equipped with a mobile body calculator, acquires first distance information to targets in the surroundings of the mobile body and inputs it to the mobile body calculator, the mobile body calculator has a depth image generation unit that generates the first depth image based on the first surrounding image and the first distance information, and the roadside unit has a roadside calculator, and the roadside computer has a depth image generation unit that generates the second depth image based on the second surrounding image.
7. A remote control system as described in claim 5 or claim 6, wherein the roadside unit also acquires second distance information to surrounding objects of the roadside unit, and the depth image generation unit generates the second depth image based also on the second distance information.
8. A remote control system as described in any one of claims 4 to 6, wherein the remote computer further has a dynamic target movement prediction unit that recognizes a dynamic target from the first depth image or the three-dimensional scene, estimates the speed of the dynamic target, and predicts a dynamic target movement prediction result including the position and attitude of the dynamic target after the predetermined time, and the predicted image generation unit generates the predicted image also based on the dynamic target movement prediction result.
9. The remote control system according to claim 8, wherein the dynamic target movement prediction unit recognizes the dynamic target by image processing or a neural network.
10. A remote control system according to claim 8, wherein the dynamic target movement prediction unit predicts the dynamic target movement prediction result based also on the second surrounding image or the second depth image.
11. The remote control system of claim 1, wherein the mobile body transmits second data that does not include the first surrounding image so that the communication delay is shorter than that of first data that includes the first surrounding image.
12. The remote control system according to claim 11, wherein the mobile unit transmits the first data and the second data in separate packets or over separate lines so that the communication delay for the second data is shorter.
13. A remote control system as described in claim 11, wherein the mobile body is equipped with a target state estimation unit that estimates the state of surrounding targets of the mobile body based on the first surrounding image, the mobile body transmits the state of the surrounding targets as the second data, and the predicted image generation unit generates the predicted image also based on the state of the surrounding targets.
14. The remote control system according to claim 13, wherein the target state estimation unit estimates the state of the surrounding targets by image processing or a neural network.
15. A remote control system as described in claim 1, wherein the remote control device further comprises a database relating to target textures, the remote computer further comprises a target type recognition unit that recognizes the type of target in the first surrounding image or an image present in the three-dimensional scene, and the predicted image generation unit assigns the texture to a blind spot of the target in the image in the three-dimensional scene based on the type of target in the image and the database.
16. A remote control system according to claim 15, wherein the target type recognition unit recognizes the type of target in the image by image processing or a neural network.
17. The remote control system of claim 1, wherein, if the predicted image includes a blind spot, the predicted image generation unit complements the blind spot using a trained neural network model.
18. A remote control system as described in claim 1, wherein the remote control device further comprises a database of a three-dimensional map of the surroundings of the mobile object, and the three-dimensional scene construction unit constructs the three-dimensional scene also based on the database of the three-dimensional map.
19. The remote control system according to claim 1, wherein the predicted image generating unit generates the predicted image at a cycle shorter than the reception cycle of the first surrounding image.
20. The remote control system according to claim 1, wherein the moving object movement prediction unit expresses the movement model using a neural network.
21. The remote control system according to claim 20, wherein the moving object movement prediction unit trains the neural network based on the operation information and the state quantity of the moving object.
22. A remote control device for remotely controlling a moving object, wherein the moving object acquires and outputs a first surrounding image around the moving object, the remote control device comprising: an operation device for operating the moving object; and a remote computer for calculating a predicted image after a predetermined time based on the first surrounding image output by the moving object, the remote computer comprising: a three-dimensional scene construction unit for constructing a three-dimensional scene based on the first surrounding image and a first depth image generated based on the first surrounding image; a moving object movement prediction unit for predicting the position and orientation of the moving object after the predetermined time based on operation information to the operation device and a motion model of the moving object; and a predicted image generation unit for generating the predicted image based on the three-dimensional scene and the position and orientation after the predetermined time.
Citation Information
Patent Citations
Remote operation device for moving body
JP1998275015A
Image generation device for vehicles and image generation method for vehicles
JP2013025528A
Route generation device, route control system, and route generation method
JP2018160228A
Remote control device, remote control method, and remote control program
JP2023078020A
Information processing device, information processing method, and program
WO2021002116A1