Method for training a robotic device controller, method for controlling a robotic device, robotic device control system, computer program elements, and computer-readable media
A trained robotic device controller using encoder and decoder networks enables autonomous navigation with high-level commands, addressing the inefficiency of constant human input by accurately navigating and avoiding obstacles.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DCONSTRUCT TECH PTE LTD
- Filing Date
- 2021-09-17
- Publication Date
- 2026-05-11
AI Technical Summary
Existing robotic control systems require constant human user input to navigate obstacles and follow paths, necessitating continuous attention and involvement, which is inefficient and limits concurrent tasks.
A robotic device controller is trained using an encoder network, decoder network, and policy model to autonomously interpret high-level commands and navigate environments without prior knowledge of the path, using supervised learning to determine traversability and distance from obstacles, enabling hands-free operation.
The system allows for autonomous navigation with simple commands, reducing user burden and enabling concurrent tasks by accurately avoiding obstacles and following paths without prior environmental knowledge or path recording.
Smart Images

Figure 0007856754000001 
Figure 0007856754000002 
Figure 0007856754000003
Abstract
Description
Technical Field
[0001] Various aspects of the present disclosure relate to an apparatus and method for controlling a robot device, and an apparatus and method for training a robot device controller.
Background Art
[0002] A robot device such as a mobile robot can be controlled using remote control by a human user. For this, a human user can be supplied with an image from the perspective of the robot, for example, and react accordingly, for example, can steer the robot around obstacles. However, this requires accurate input by the user at the correct time, and thus requires constant attention from the human user.
[0003] Therefore, a technique that enables a robot to move more autonomously according to high-level commands of a human user, such as "move forward" (along a path such as a corridor), "turn right", or "turn left", is desirable.
Summary of the Invention
[0004] According to various embodiments, a method for training a robot device controller includes, for each of a plurality of digital training input images, an encoder network encoding the digital training input image into features in a latent space, a decoder network determining, for each of a plurality of regions shown in the digital training input image from the features, whether the region is traversable, and determining information regarding the distance between the perspective and the region of the digital training input image, and a policy model determining control information for controlling the movement of the robot device from the features, including training a neural network including an encoder network, a decoder network, and a policy network, wherein at least the policy model is trained in a supervised manner using control information ground truth data of the digital training input image.
[0005] According to one embodiment, training an encoder network and a decoder network includes training an autoencoder comprising an encoder network and a decoder network.
[0006] According to one embodiment, the method includes training an encoder network together with a decoder network.
[0007] According to one embodiment, the method includes training an encoder network together with a decoder network and a policy network.
[0008] According to one embodiment, the decoder network includes a semantic decoder and a depth decoder, and the neural network is trained such that for each digital training input image, the semantic decoder determines from features whether each of the multiple regions shown in the digital training input image is traversable, and the depth decoder determines from one or more features information regarding the distance between the viewpoint of the digital training input image and the region for each of the multiple regions shown in the digital training input image.
[0009] According to one embodiment, the semantic decoder is trained in a supervised manner.
[0010] According to one embodiment, the depth decoder is trained in a supervised manner, or the depth decoder is trained in an unsupervised manner.
[0011] According to one embodiment, one or more of the encoder network, decoder network, and policy network are convolutional neural networks.
[0012] According to one embodiment, the control information includes control information for each of the multiple robot device movement commands.
[0013] In one embodiment, the neural network is trained so that the policy model determines control information from features encoded by the encoder from multiple training input images.
[0014] According to one embodiment, a method is provided for controlling a robotic device, comprising: training a robotic device controller according to a method according to any one of the embodiments described above; acquiring one or more digital images showing the surroundings of the robotic device; encoding one or more digital images into one or more features using an encoder network; supplying one or more features to a policy network; and controlling the robot in response to one or more features according to the control information output of a policy model.
[0015] According to one embodiment, the method includes receiving one or more digital images from one or more cameras of a robotic device.
[0016] According to one embodiment, the control information includes control information for each of a plurality of robot device movement commands, and the method includes receiving instructions for a robot device movement command and controlling the robot according to the control information for the instructed robot device movement command.
[0017] According to one embodiment, a policy model is a neural network trained to determine control information from features encoded by an encoder of multiple training input images, and the method includes acquiring multiple digital images showing the surroundings of a robotic device, encoding the multiple digital images into multiple features using an encoder network, supplying the multiple features to a policy network, and controlling the robot in response to the multiple features according to the control information output of the policy model.
[0018] According to one embodiment, the multiple digital images include images received from different cameras.
[0019] According to one embodiment, the plurality of digital images include images taken from different viewpoints.
[0020] According to one embodiment, the plurality of digital images include images taken at different times.
[0021] According to one embodiment, a robot device control system configured to implement any one of the methods of the above embodiments is provided.
[0022] According to one embodiment, a computer program element including program instructions that cause one or more processors to implement any one of the methods of the above embodiments when executed by the one or more processors is provided.
[0023] According to one embodiment, a computer-readable medium including program instructions that cause one or more processors to implement any one of the methods of the above embodiments when executed by the one or more processors is provided.
Brief Description of the Drawings
[0024] The present invention will be better understood by reference to the detailed description when considered in conjunction with non-limiting examples and the accompanying drawings. [Figure 1] Showing a robot. [Figure 2] Showing a control system according to one embodiment. [Figure 3] Showing a machine learning model according to one embodiment. [Figure 4] Showing a machine learning model for processing a plurality of input images according to one embodiment. [Figure 5] Showing a method for training a robot device controller according to one embodiment.
Modes for Carrying Out the Invention
[0025] The following detailed description refers to the accompanying drawings that illustrate specific details and embodiments by which the present disclosure can be implemented. These embodiments are described in sufficient detail to enable those skilled in the art to implement the present disclosure. Without departing from the scope of the present disclosure, other embodiments can be utilized and structural and logical changes can be made. Since some embodiments can combine with one or more other embodiments to form new embodiments, various embodiments are not necessarily mutually exclusive.
[0026] Embodiments described in the context of one apparatus or method are equally valid for other apparatuses or methods. Similarly, embodiments described in the context of an apparatus are equally valid for a vehicle or a method, and vice versa.
[0027] Features described in the context of one embodiment may be applicable corresponding to the same or similar features of other embodiments. Features described in the context of one embodiment may be applicable corresponding to other embodiments even if not explicitly described in these other embodiments. Further, the additional and / or combinations and / or alternatives described for features in the context of one embodiment may be applicable corresponding to the same or similar features of other embodiments.
[0028] In the context of various embodiments, the articles “a,” “an,” and “the” used with respect to a feature or element include reference to one or more of the feature or element.
[0029] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0030] Hereinafter, embodiments will be described in detail.
[0031] FIG. 1 shows a robot 100.
[0032] Robot 100 is a mobile device. In the example in Figure 1, it is a quadruped robot having four legs 101 for walking on the ground 102 and a camera 103 (or more cameras) for observing its environment (i.e., its surroundings), in particular the ground 102 and obstacles 104 such as objects or people.
[0033] Camera 103, for example, acquires an RGB image 105 (red, green, blue, i.e., a color image) of the robot's environment.
[0034] Image 105 can be used to control the path taken by the robot 100. This can be done, for example, by remote control. This means that a remote control device 106 is provided, which is operated by a human user 107. The human user 107 sends back to the robot 100, specifically to the robot's controller 108, which generates control commands for the robot 100 that control the robot's movements accordingly. For example, the legs include actuators 109 configured to be controlled by the controller 108 according to the transmitted commands.
[0035] To generate control commands, the robot 100 can transmit image 105 to a control device 106 that presents the image 105 to a human user 107 (on a screen). The human user 107 can then generate control commands for the robot (for example, by a control device including a joystick and / or console).
[0036] However, such a control method requires a certain level of involvement from a human user, because a human user must constantly look at the RGB images delivered by the robot 100 and select the corresponding control commands in order to avoid obstacles 104 and follow the appropriate path on the ground 102, for example.
[0037] In view of the above, according to various embodiments, instead of operating the robot with a control device that requires a certain level of involvement from a human user, a human user 107 can operate the robot with simple (high-level) commands ("move left", "move right", "move forward").
[0038] Therefore, control systems in various embodiments allow a human user (i.e., an operator, such as a driver) to direct the moving device using simple commands such as move forward, turn left, or turn right. This reduces the burden of operating the device and allows the operator to perform other tasks concurrently.
[0039] According to various embodiments, the control system deploys a QR code (Registered trademark) This provides operators with more convenient control (especially a hands-free control experience) without requiring enhancements to the environment in which the robot moves, or prior knowledge of the robot's movement path, such as point cloud maps, which need to be prepared in advance and consumed during operation. In particular, according to various embodiments, the control system does not need to record the robot's control over the path for replay after control.
[0040] Furthermore, the embodiments go beyond intervention when the operator (human user 107) makes a mistake, such as stopping the robot 100 when an obstacle 104, such as a pedestrian, is too close. While this only helps to avoid collisions, various embodiments allow the human user 107 to maneuver the robot 100 to reach a destination from a starting point with fewer simple control commands. For example, according to various embodiments, a machine learning model can be trained (by appropriately labeling the training data for a policy model, as described later) to stop and bypass a collision before it occurs.
[0041] Therefore, the control systems provided according to various embodiments can operate in any unconventional environment without prior knowledge of the environment or path the robot should take, and do not require the placement of reference markers in the environment to guide the system, nor do they require prior recording of the path.
[0042] Figure 2 shows a control system 200 according to one embodiment.
[0043] The control system 200 plays a role in controlling, for example, robot 201, which corresponds to robot 100.
[0044] The control system 200 includes a first processing unit (or computing unit) 202 and a second processing unit (or computing unit) 203, as well as a camera 204 (or multiple cameras).
[0045] Camera 204 and the first processing unit 202 are part of the payload 205 of robot 201, which is mounted on robot 201. Therefore, they may also be considered part of robot 201, for example, corresponding to camera 103 and controller 108, respectively.
[0046] The second processing unit 203 corresponds, for example, to the remote control device 106.
[0047] As described above, the control system 200 allows a human operator 206 to direct the movement of the robot 201 (generally speaking, the mobile and / or movable (robot) device) using simple commands (i.e., high-level control commands) such as forward, left turn, or right turn.
[0048] From these high-level control commands entered by user 206, the control system 200 automatically infers velocity and angular velocity control signals 207 (for example, for actuator 109) and maneuvers the robot 201 accordingly.
[0049] To this end, the first processing unit 202 implements a machine learning model 208. The first processing unit 202 uses the machine learning model 208 to determine control signals 207 according to high-level control commands 210 input by the user 206. For example, if there is a curve in the path (e.g., a corridor or passage), and the human user 206 simply inputs a forward command, the first processing unit 202 uses the machine learning model 208 to determine the appropriate speed and angular velocity, as well as the corresponding control signals 207, to keep the robot 201 on the path (for each of the series of control time stages, i.e., control time).
[0050] Similarly, when user 206 inputs a “turn left” or “turn right” command, the first processing unit 202 generates a control signal 207 to adapt to the available path, for example, so that the robot 201 turns at the correct time to avoid colliding with obstacles (in particular, corridors or building walls) or falling off the path.
[0051] Camera 204 (or multiple cameras) is calibrated, for example, to have a good field of view of the environment.
[0052] The first processing unit 202 communicates with the second processing unit 203 to send the image 209 generated by the camera 204 to the second processing unit 203 and receives a high-level command 210 input to the second processing unit 203 by the user 206.
[0053] For this communication, the first processing unit 202 and the second processing unit 203 communicate between processing units 202 and 203 (e.g., 5G network, WiFi, Ethernet). (Registered trademark) Bluetooth (Registered trademark) This includes communication devices that implement a corresponding wireless or wired communication interface (using a cellular mobile wireless network such as a cellular mobile wireless network).
[0054] The camera 204 generates the image 209 in the form of a message stream, for example, which is provided to the first processing unit 202.
[0055] The first processing unit 202 transfers the image 209 to the second processing unit 203, and the second processing unit 203 The image 209 can be displayed to the human operator 206 to show the operator the environment in which the robot is currently located. The human operator 206 issues a high-level command 210 using the second processing unit 203. The second processing unit 203 sends the high-level command 210 to the first processing unit 202.
[0056] The first processing unit 202 hosts (implements) a machine learning model 208, is connected to the camera 204 and components of the controlled robot 201 (e.g., actuator 109), and receives high-level commands 210 from the second processing unit 203. The first processing unit 202 generates control signals 207 by processing the image 209 and the high-level commands 210. This includes processing the image 209 using the machine learning model 208. The first processing unit 202 supplies the control signals 207 to the components of the controlled robot 201.
[0057] Camera 204 is positioned on the robot 201 in such a way as to provide a first-person perspective image to a machine learning model 208 for processing. For example, camera 204 provides a color image. Multiple cameras are used to achieve a sufficient field of view. 209 We can provide this.
[0058] Robot 201 provides mechanical means for operating according to control signals. First processing unit 202 provides computing resources for running a machine learning model 208 at a speed fast enough for real-time inference (of control signals 207 and high-level commands from image 209). Any number and type of cameras can be used depending on the form factor of robot 201. First processing unit 202 processes images 209Stitching and calibration can be performed (for example, to compensate for discrepancies between camera angles and positions).
[0059] To achieve better control performance, other types of sensors besides RGB cameras, particularly thermal cameras, motion sensors, and acoustic transducers, can be added.
[0060] The first processing unit 202 determines the control signal 207 using a control algorithm that includes processing by the machine learning model 208.
[0061] It should be noted that in one embodiment, the machine learning model 208 may also be hosted on a second processing unit 203 instead of the first processing unit 202. In that case, the determination of the control signal 207 is performed on the second processing unit 203. The control signal 207 is then sent by the second processing unit 203 to the first processing unit 202 (instead of a high-level command 210), and the first processing unit 202 The control signal 207 is transferred to the robot 201.
[0062] The machine learning model 208 may also be hosted on a third processing unit located between the first processing unit 202 and the second processing unit 203. In this case, the determination of the control signals 207 is carried out on the third processing unit, which may be located remotely and exchange data with the first processing unit 202 and the second processing unit 203. The control system remains complete in such an arrangement, as long as the second processing unit 203 receives images and transmits high-level user commands in real time. Similarly, the first processing unit 202 can transmit images and receive (low-level) control signals 207 in real time.
[0063] According to various embodiments, the machine learning model 208 is a deep learning model that processes images (i.e., frames) 209 provided by the camera 204 (or multiple cameras) to control information of the robot 201 at each control time stage. According to the embodiments described below, the machine learning model 208 makes predictions of control information for all possible intentions (i.e., all possible high-level commands) at each control time stage. Next, the first processing unit 202 determines a control signal 207 from the predicted control information according to the high-level commands provided by the second processing unit 203.
[0064] In this embodiment, the robot 201 is assumed to have low inertia so as to respond to changes in the control signal 207 at each time step.
[0065] Figure 3 shows machine learning model 300.
[0066] In the example in Figure 3, we assume that the machine learning model 300 receives a single RGB (i.e., color) input image 301, for example, an image 301 from a single camera 204, over one control time step.
[0067] The machine learning model includes an (image) encoder 302 for transforming an input image 301 into features 303 (i.e., feature values, or feature vectors containing multiple feature values) in a feature space (i.e., latent space). The policy model 304 generates control information predictions as the output 305 of the machine learning model 300.
[0068] The encoder 302 and policy model 304 are trained (i.e., optimized) during training time and deployed to process images during operation (i.e., during inference).
[0069] For training purposes, the machine learning model 300 includes a depth decoder 306 and a semantic decoder 307 (neither of which are deployed or used for inference).
[0070] The depth decoder 306 is trained to provide depth predictions of positions on the input image 301 (which is the training input image 301 at training time). This means predicting the distance of parts of the robot's environment (particularly objects) shown in the input image 301 from the robot. The output may be a high-density depth prediction and may be in the form of relative depth values or absolute (scale-consistent) depth values.
[0071] The depth decoder 306 is trained to provide semantic predictions of positions on the input image 301 (which is the training input image 301 at training time). This means predicting whether a portion of the robot's environment shown in the input image 301 is traversable.
[0072] The encoder 302 can use any standard convolutional neural network (CNN). The depth decoder 306 and semantic decoder 307 can also use any standard CNN (to the extent that they can be optimized for their respective use cases).
[0073] The policy model 304 infers control information (such as velocity and direction, which may include one or more angles) from the features 303. The quality of the features 303 is important to the policy model 304 so that the encoder 302 can be trained together with the policy model 304. Similarly, the encoder 302 can be trained together with the decoders 306, 307 so that the features 303 reliably represent depth information and semantic information.
[0074] The policy model 304 is trained in a supervised manner using control information ground truth (e.g., included in the labels of the training input images). For example, the policy model 304 is trained to slow down (the robot 201 decelerates) when an obstacle is near the robot. For forward intentions (i.e., high-level commands to move forward), it can also be trained to slow down when a human operator 206 needs to input an explicit command, i.e., in the case of a symmetric Y-junction where operator 206 needs to specify where to move forward.
[0075] With respect to angles, forward intent is defined as path following. Therefore, on curved paths, policy model 304 is trained to predict control information and take over robots to ensure the robot stays on the path.
[0076] For left or right intentions (i.e., high-level commands "turn left" and "turn right"), policy model 304 is trained to predict only control information that causes the robot to turn where possible, i.e., to continue moving forward until a path for turning becomes clear, rather than turning the robot towards an obstacle.
[0077] As described above, the policy model 304 is trained in a supervised manner, that is, by providing a training dataset containing training input images, and for each training input image, a label is provided specifying the target control information for each high-level command (i.e., ground truth control information). The mean squared error (MSE) can be used as the loss for training the policy model 304.
[0078] The depth decoder 306 is trained to ensure that depth predictions are geometrically accurate, for example, by not predicting a triangular space as a dome-shaped space. The depth decoder can be trained using supervised or unsupervised methods.
[0079] In supervised training, the label of each training input image further specifies the target (ground truth) depth information that the depth decoder 306 should output. The mean squared error (MSE) can be used as the loss for training the depth decoder 306.
[0080] In unsupervised training, for example, two cameras 204 can be used to generate images simultaneously. The depth decoder 306 can then be trained to minimize the loss between the image generated by the first camera and the image reconstructed from the depth prediction of the viewpoint of the second camera. Reconstruction is performed by a network trained to generate an image from the viewpoint of the second camera from the image and depth information captured by the first camera. The depth decoder can also be trained on sampled sequences within a video.
[0081] In one embodiment, the semantic decoder 307 performs traversable path segmentation rather than identifying the category of each pixel in the scene (which is a standard formulation for semantic segmentation). This means it is trained to understand the geometric shapes of non-convex objects such as people and chairs. In an image of a person standing, a standard semantic segmentation model would predict the space between the person's feet as "floor" or "ground." Instead, the semantic decoder 307 is trained to predict that it is not traversable because it is undesirable for the robot 201 to bump into the person. This also applies to many pieces of furniture, such as chairs.
[0082] The semantic decoder 307 is trained in a supervised manner. For this purpose, the label of each training input image further specifies whether the portion shown in the training image is traversable or not. The cross-entropy loss can be used as the loss for training the semantic decoder 307 (e.g., having "traversable" and "non-traversable" categories).
[0083] Encoder 302 is trained one or more times with other models. Encoder 302, policy model 304, depth decoder 306, and semantic decoder 307 can all be trained together by summing the output losses of policy model 304, depth decoder 306, and semantic decoder 307.
[0084] Figure 4 shows a machine learning model 400 for processing multiple input images 401.
[0085] The machine learning model 400, for example, processes the payload 205, each controlling the image at each time step. 209 This can be applied when including multiple cameras 204 that provide control information. The machine learning model 400 also predicts control information from multiple subsequent images 209 Please note that this may be used to take that into consideration.
[0086] All input images are supplied to the same encoder 402 (similar to encoder 302). This allows a feature 403 to be obtained for each input image.
[0087] The features 403 generated by encoder 402 are concatenated together before being consumed by policy model 404 to produce control information output 405. For training, the same set of decoders ( Depth Decoder 406 and Semantic Decoder 407) operates with each feature 403.
[0088] Training data can be selected according to the use case. For example, in the case of pedestrian navigation rather than car navigation, following car traffic rules is not the goal, and there is no need to clearly demarcate lanes.
[0089] In summary, according to various embodiments, the method shown in Figure 5 is provided.
[0090] Figure 5 shows a method for training the robot device controller.
[0091] A neural network 500, including an encoder network 501, a decoder network 502, and a policy network 503, is trained. As a result, for each of the multiple digital training input images 504, the encoder network 501 encodes the digital training input image into features in latent space; the decoder network 502 determines from the features whether each of the multiple regions shown in the digital training input image is traversable and determines information regarding the distance between the viewpoint of the digital training input image and the region; and the policy model 503 determines from the features control information for controlling the movement of the robotic device.
[0092] At least the policy model 503 is trained in a supervised manner using control information ground truth data 505 of the digital training input image 504.
[0093] According to various embodiments, in other words, the robotic device is controlled based on features representing information about the distance of each of one or more regions from the robot and whether the region is traversable to the robotic device. This is achieved by training an encoder / decoder architecture in which a decoder unit reconstructs distance (i.e., depth) information and semantic information (i.e., whether the region is traversable) from features generated by the encoder, and trains a supervised policy model to generate control information for controlling the robotic device from the features.
[0094] According to various embodiments, in other words, a method is provided for training a robotic device controller, comprising: training a neural encoder network to encode one or more digital training input images into one or more features in a latent space; training a neural decoder network to determine, for each of a plurality of regions shown in one or more digital training input images, from one or more features whether the region is traversable by the robot and determine information regarding the distance between the viewpoint from which the one or more digital training input images were taken and that region; and training a policy model to determine control information for controlling the movement of the robotic device from one or more features, wherein at least the policy model is trained in a supervised manner using control information ground truth data of the digital training input images.
[0095] The method shown in Figure 5 is performed by a robotic device control system that includes, for example, a communication interface, one or more processing units, and memory (for example, for storing a trained neural network).
[0096] The methods described above can be applied to the control of any device that is movable and / or has movable parts. This means that it can be used not only to control the movement of mobile devices such as walking robots (and thus, those in Figure 1), flying drones, and autonomous vehicles (e.g., for logistics), but also to control the movement of movable limbs of devices such as robotic arms (e.g., industrial robots that should avoid collisions with obstacles such as passing workers, as in a mobile robot) or access control systems (and thus, monitoring).
[0097] Therefore, the methods described above can be used to control the movement of any physical system, such as robots, vehicles, home appliances, tools, or computer-controlled machinery such as manufacturing equipment. The term “robot apparatus” is understood to include all mobile and / or movable devices of these kinds (i.e., stationary devices with movable components in particular).
[0098] The methods described herein can be implemented, and the various processing or computing units and devices and computing entities described herein can be implemented by one or more circuits. In one embodiment, “circuit” may be understood as any kind of logic implementation entity, which may be hardware, software, firmware, or any combination thereof. Thus, in one embodiment, “circuit” may be a wired logic circuit or a programmable processor, such as a microprocessor. “Circuit” may also be software implemented or executed by a processor, such as any kind of computer program, such as a computer program using virtual machine code. Any other kind of implementation of each function described herein may also be understood as a “circuit” in an alternative embodiment.
[0099] While this disclosure has been specifically shown and described with reference to certain embodiments, it should be understood by those skilled in the art that various modifications in form and detail can be made without departing from the spirit and scope of the invention as defined by the appended claims. Accordingly, the scope of the invention is indicated by the appended claims and is therefore intended to include all modifications that fall within the same meaning and scope as the claims.
Claims
1. A method for training a robotic device controller, For each of the multiple digital training input images, the encoder network encodes the digital training input image into features in the latent space. The decoder network determines, from the features, whether each of the multiple regions shown in the digital training input image is traversable, and determines information regarding the distance between the viewpoint of the digital training input image and the region in terms of relative depth. The policy model determines control information for controlling the movement of the robot device based on the aforementioned features. This includes training a neural network that includes the encoder network, the decoder network, and the policy model, At least the policy model is trained in a supervised manner using control information ground truth data of the digital training input images. method.
2. The method according to claim 1, wherein training the encoder network and the decoder network includes training an autoencoder comprising the encoder network and the decoder network.
3. The method according to claim 1 or 2, comprising training the encoder network together with the decoder network.
4. The method according to any one of claims 1 to 3, comprising training the encoder network together with the decoder network and the policy model.
5. The decoder network includes a semantic decoder and a depth decoder, and for each digital training input image, The semantic decoder determines, based on the features, whether each of the multiple regions shown in the digital training input image is traversable. Based on the characteristics, the depth decoder determines that for each of the multiple regions shown in the digital training input image, The method according to any one of claims 1 to 4, wherein the neural network is trained to determine information regarding the distance between the viewpoint of the digital training input image and the region.
6. The method according to claim 5, wherein the semantic decoder is trained in a supervised manner.
7. The method according to claim 5, wherein the depth decoder is trained in a supervised manner or in an unsupervised manner.
8. The method according to any one of claims 1 to 7, wherein one or more of the encoder network, the decoder network, and the policy model are convolutional neural networks.
9. The method according to any one of claims 1 to 8, wherein the control information includes control information for each of a plurality of robot device movement commands.
10. The method according to any one of claims 1 to 9, wherein the policy model is trained such that the neural network determines the control information from features encoded by the encoder network of multiple training input images.
11. A method for controlling a robotic device, Training a robot device controller according to the method described in any one of claims 1 to 10, Acquiring one or more digital images showing the surroundings of the robot device, Encoding one or more features of the one or more digital images using the encoder network, To supply one or more of the above features to the policy model, A method comprising controlling the robot device in accordance with the control information output of the policy model in response to one or more of the aforementioned features.
12. The method according to claim 11, comprising receiving one or more digital images from one or more cameras of the robot device.
13. The method according to claim 11 or 12, wherein the control information includes control information for each of a plurality of robot device movement commands, and the method includes receiving an instruction for a robot device movement command and controlling the robot device in accordance with the control information for the instructed robot device movement command.
14. The policy model is such that the neural network is trained to determine the control information from features encoded by the encoder network from multiple training input images, and the method is The acquisition of multiple digital images showing the surroundings of the robot device, The encoder network is used to encode the multiple digital images into multiple features, The above-mentioned multiple features are supplied to the policy model, The method according to any one of claims 11 to 13, further comprising controlling the robot device in accordance with the control information output of the policy model in response to the plurality of features.
15. The method according to claim 14, wherein the plurality of digital images include images received from different cameras.
16. The method according to claim 14 or 15, wherein the plurality of digital images include images taken from different viewpoints.
17. The method according to any one of claims 14 to 16, wherein the plurality of digital images include images taken at different times.
18. A robot device control system configured to carry out the method described in any one of claims 1 to 17.
19. A computer program element that, when executed by one or more processors, includes a program instruction causing the one or more processors to carry out the method described in any one of claims 1 to 17.
20. A computer-readable medium that, when executed by one or more processors, includes program instructions causing the one or more processors to carry out the method according to any one of claims 1 to 17.