MOBILE DEVICE FOR AUTOMATED TRANSPORT OF LOAD CARRIERS

DE502023001063D1Active Publication Date: 2025-06-18STILL GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502023001063
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-07
Filing Date
2023-01-11
Publication Date
2025-06-18
Estimated Expiration
2043-01-11

AI Technical Summary

Technical Problem

Existing automated industrial trucks and robots face challenges in accurately recognizing and orienting themselves relative to load carriers like pallets or wire mesh boxes for efficient handling and transport, particularly in dynamic warehouse environments.

Method used

A mobile device equipped with movement actuators, image capture devices, LiDAR sensors, and a processor implementing reinforcement learning neural networks to autonomously navigate and orient itself relative to load carriers, using a combination of image and LiDAR data to determine optimal control actions for movement.

Benefits of technology

Enables precise and efficient automated transport of load carriers by improving the mobile device's ability to recognize and align with load carriers, reducing collisions and enhancing operational efficiency in warehouse environments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a mobile device, in particular a mobile robot for the automated transport of load carriers, in particular a pallet or a wire mesh box, for example in a warehouse.

[0002] Load carriers with standardized dimensions, such as wire mesh boxes or pallets, and in particular Euro pallets, are often used to transport and store products, goods, and materials. Partially automated industrial trucks, such as forklift trucks, industrial robots, and the like, are often used to handle such load carriers, for example in intralogistics - i.e., the internal flow of materials, e.g. in a warehouse. For this purpose, industrial trucks generally have a load-handling device, such as two forks, which can be inserted into corresponding openings in the load carrier to pick it up. If this is to be done using automated industrial trucks or robots, the industrial truck or robot must be...the robot recognizes the load carrier, its pose and the insertion openings of the load carrier and moves and orients itself relative to the load carrier in such a way that the load handling device can be inserted into the insertion openings of the load carrier and the load carrier can thus be transported.

[0003] The present invention is based on the object of providing a mobile device for the automated transport of a load carrier, in particular a pallet or a wire mesh box.

[0004] US 2019 / 101917 A1 concerns a prediction of a state of an object in an environment using an action model of a neural network.

[0005] This object is achieved according to a first aspect of the invention in that the mobile device for the automated conveying of a load carrier comprises at least one movement actuator which is designed to be controlled by means of control data in order to move the mobile device, wherein the mobile device is designed to move and orient itself relative to the load carrier to be conveyed by means of the at least one movement actuator, an image capture device for capturing a plurality of images of a section of an environment of the mobile device at different times, at least one LiDAR sensor for capturing LiDAR data in the environment of the mobile device and a processor which is designed to implement a neural network trained by means of reinforcement learning, RL-trained neural network, wherein the RL-trained neural network is designed, on the basis of the plurality of images, the LiDAR data,a plurality of past control data (ie, control data determined in the past) and a plurality of past rewards (ie, rewards determined in the past) determined for the plurality of past control data, to determine the control data for controlling the at least one motion actuator for the current time to move the mobile device.,

[0006] According to one embodiment, the processor is further configured to implement another RL-trained neural network, wherein the another RL-trained neural network is configured to determine a Q-value based on the plurality of images, the LiDAR data, the plurality of past control data, and the plurality of past rewards.

[0007] In one embodiment, the RL-trained neural network comprises a subnetwork and the further RL-trained neural network also comprises the subnetwork configured to generate a feature vector based on the plurality of images.

[0008] According to one embodiment, the processor is designed to pre-train the subnetwork in a first training section and to train the complete RL-trained neural network and the complete further RL-trained neural network in a subsequent second training section.

[0009] In one embodiment, in the second training section, the processor is configured to train the RL-trained neural network using tasks (also referred to as "tasks") that are selected by the processor based on the further RL-trained neural network.

[0010] According to one embodiment, the processor is further configured to implement a further neural network, wherein the further neural network is configured to determine a probability of success of the control data for controlling the at least one motion actuator on the basis of the Q value.

[0011] In one embodiment, the further neural network is configured to determine the probability of success of the control data for controlling the at least one motion actuator based on the Q value and one or more geometric parameters of the task.

[0012] In one embodiment, the processor is configured to control the at least one motion actuator with the control data if the probability of success of the control data is greater than a predefined threshold.

[0013] According to one embodiment, the processor is configured to bring the mobile device to a standstill if the probability of success of the control data is less than a predefined threshold.

[0014] In one embodiment, the processor is configured to train the RL-trained neural network based on a soft actor critic algorithm.

[0015] Further advantages and details of the invention are explained in more detail with reference to the exemplary embodiments shown in the schematic figures. <h2 style=";text-align:left;direction:ltr">Figure 1 a schematic representation of a mobile device for the automated transport of a load carrier according to an embodiment; <h2 style=";text-align:left;direction:ltr"> Figure 2a a schematic representation of the architecture of a neural network of a mobile device for the automated transport of a load carrier according to one embodiment; <h2 style=";text-align:left;direction:ltr"> Figure 2ba schematic representation of the architecture of another neural network of a mobile device for the automated transport of a load carrier according to an embodiment; <h2 style=";text-align:left;direction:ltr"> Figur 2c a schematic representation of the architecture of a subnetwork of the neural networks of <h2 style=";text-align:left;direction:ltr"> Figure 2a and 2b ; and <h2 style=";text-align:left;direction:ltr"> Figure 3 a schematic representation of the architecture of another neural network of a mobile device for the automated transport of a load carrier according to an embodiment.

[0016] <h2 style=";text-align:left;direction:ltr"> Figure 1shows a schematic representation of a mobile device 110 for the automated transport of a load carrier 170 according to one embodiment. A load carrier 170 can be understood as aids that serve to combine a load, in particular comprising one or more packages, in particular packaging units, into a load unit. For example, a carton, a box, a pallet, a wire mesh box, a container, an exchangeable swap body, a shelf element or the like can be a load carrier 170. The mobile device 110 can be, in particular, an industrial robot 110 or an autonomous or automated industrial truck 110, for example, an industrial truck designed for driverless and / or autonomous operation. The industrial robot can be, in particular, an order-picking robot.

[0017] As in <h2 style=";text-align:left;direction:ltr"> Figure 1As indicated, the mobile device 110 is designed to move autonomously, for example in a warehouse 100 with a plurality of load carriers 170, which can be arranged with different orientations at different positions in the warehouse 100. In one embodiment, the mobile device 110 comprises one or more movement actuators 160 and corresponding drives for moving the mobile device 110. In one embodiment, the movement actuators 160 of the mobile device 110 comprise, for example, a plurality of wheels that can be driven by a motor by means of a control unit of the mobile device 110. According to the invention, the mobile device 110 is designed to perform, for example, a linear movement and a rotational movement by means of the movement actuators 160 in order to be able to move and orient itself relative to a load carrier 170 to be transported.

[0018] The mobile device 110 further comprises an image capture device 130, in particular a 2D camera 130, for capturing a plurality of image data of the surroundings of the mobile device 110 that change due to the movement, and at least one LiDAR sensor 140 for capturing LiDAR data in the surroundings of the mobile device 110. In one embodiment, the at least one LiDAR sensor 140 is configured to capture LiDAR data in the horizontal plane, i.e., the plane of movement of the mobile device 110. In one embodiment, the at least one LiDAR sensor 140 is configured to capture LiDAR data in an angular range of 360° or a section thereof around the mobile device 110.

[0019] The mobile device 110 further comprises at least one processor 120, which is configured to control the movement of the mobile device 110 by means of the movement actuators 160 based on the image data acquired by the camera 130 and the LiDAR data acquired by the at least one LiDAR sensor 140, as described below with further reference to the <h2 style=";text-align:left;direction:ltr"> Figure 2a-c is described in detail.

[0020] In one embodiment, the mobile device 110 for automated transport of a load carrier 170 may further include a non-volatile memory 150, for example, a flash memory 150. The non-volatile memory 150 is configured to store data and executable program code that, when executed by the processor 120 of the mobile device 110, causes the processor 120 to perform the functions, operations, and methods described below.

[0021] <h2 style=";text-align:left;direction:ltr"> Figure 2ashows the architecture of an artificial neural network (also referred to herein as a "neural network") 200 implemented by the at least one processor 120 according to one embodiment. <h2 style=";text-align:left;direction:ltr"> Figure 2a The neural network 200 shown is designed to determine, in an application phase of the mobile device 110, on the basis of a plurality of continuously acquired input data 210a-d, control data 290 for controlling the movement actuators 160, which enable the mobile device 110 to move and align itself relative to a load carrier 170 to be transported and to pick up the load carrier.

[0022] As described in more detail below, the architecture of the neural network 200 is based on <h2 style=";text-align:left;direction:ltr"> Figure 2abased on concepts, algorithms, and methods of machine learning known as reinforcement learning, in which an agent (in this case, the mobile device 110) independently learns a strategy (also referred to as a "policy") to maximize received rewards. The agent is not shown which action is best in which situation (also referred to as a "task"); instead, it receives a reward, which can also be negative, at specific times through its interaction with its environment.

[0023] Reinforcement learning considers the interaction of a learning agent (in this case, the mobile device 110) with its environment, which is defined, for example, by the plurality of load carriers 170. The latter is formulated as a Markov decision problem. Thus, the environment defines a set of states. The agent, i.e., the mobile device 110, can select an action from a set of available actions depending on the situation, for example, a linear and / or rotational movement defined by control data 290, thereby reaching a subsequent state and receiving a reward.

[0024] In reinforcement learning, the goal of the agent, i.e., the mobile device 110, is to maximize the future expected profit, which is composed of the rewards for future time steps. The expected profit is therefore something like the expected total reward. A discount factor can be used to weight future rewards. For episodic problems, i.e., the overall system transitions to a final state after a finite number of steps, the discount factor can be omitted. In this case, the reward is valued equally at each time step. For continuous problems, a suitable discount factor can be chosen so that the expected profit converges. To this end, the agent, i.e., the mobile device 110, pursues a strategy (also referred to as a "policy") that the agent continuously improves.Typically, the strategy is considered a function that assigns an action to each state. However, nondeterministic strategies (or mixed strategies) are also possible, so that an action is selected with a certain probability. Further details on reinforcement learning are described in the book "Reinforcement Learning: An Introduction," by Richard S. Sutton and Andrew G. Barto, which is incorporated herein by reference.

[0025] At the <h2 style=";text-align:left;direction:ltr"> Figure 2aIn the embodiment shown, the input data of the neural network 200 comprise the four last, i.e. most recent, images or image data 210a of the surroundings of the mobile device 110 captured by the camera 130, as well as the current LiDAR data 210b of the surroundings of the mobile device 110 captured by the at least one LiDAR sensor 140. In one embodiment, the camera 130 is designed to capture a respective image of the surroundings of the mobile device 110 at temporally constant intervals. For example, in one embodiment, the interval between two temporally successive images can be approximately 180 ms. In this exemplary embodiment, a new state can therefore be observed every 180 ms. The image data 210a used for a state according to one embodiment are then the current image as well as the images captured 180 ms, 360 ms and 540 ms ago.

[0026] As in <h2 style=";text-align:left;direction:ltr"> Figure 2aAs shown, the four last, ie most recent, images or image data 210a captured by the camera 130 can be stacked to generate a multi-dimensional image data array (or an image data tensor) 210a, which in the embodiment of <h2 style=";text-align:left;direction:ltr"> Figure 2a consists of the four most recent images, each with three color channels and an exemplary image resolution of 80x80 pixels. As will be explained in more detail below in connection with <h2 style=";text-align:left;direction:ltr"> Figur 2c As described, the neural network 200 comprises a subnetwork 220 (referred to as "Custom ResNet" in <h2 style=";text-align:left;direction:ltr"> Figure 2a(denoted by [labeled]), which is configured to generate a first feature vector with, for example, 3200 components on the basis of the multidimensional image data array 210a. For this first feature vector of the image data 210a with 3200 components, a second feature vector of the image data 210a with 256 components is created in a substantially known manner by a subnetwork 230 with a "fully connected" layer and a ReLU layer.

[0027] The LiDAR data 210b acquired by the at least one LiDAR sensor 140 can, as in <h2 style=";text-align:left;direction:ltr"> Figure 2ashown by way of example, are fed into the neural network 200 as a data vector with 256 components, wherein each of the components of the data vector corresponds, for example, to a direction, i.e., a directional angle in the horizontal plane. In one embodiment, the at least one LiDAR sensor 140 is configured to capture respective LiDAR data 210b of the surroundings of the mobile device 110 at temporally constant intervals. For example, in one embodiment, the interval between temporally successive LiDAR data 210b can be approximately 180 ms (consistent with the sampling interval of the camera 130).For the current LiDAR data 210b in the form of the LiDAR data vector with 256 components, a first feature vector of the LiDAR data 210b with 64 components and then a second feature vector of the LiDAR data 210b with 32 components is created in a substantially known manner by a first subnetwork 240a with a "fully connected" layer and a ReLU layer and a second subnetwork 240b with a "fully connected" layer and a ReLU layer.

[0028] In addition to the image data 210a acquired by the camera 130 and the LiDAR data 210b acquired by the at least one LiDAR sensor 140, the input data of the neural network 200 includes <h2 style=";text-align:left;direction:ltr"> Figure 2afurthermore, the four most recent actions 210c, i.e., the control data 210c determined at four previous points in time for controlling the motion actuators 160, as well as the rewards 210d determined for these four most recent actions 210c (how these rewards can be determined is described in detail below). In one embodiment, the control data 210c for controlling the motion actuators 160 can each include, for example, a linear velocity and a rotational velocity.

[0029] As in <h2 style=";text-align:left;direction:ltr"> Figure 2a As shown, the neural network 200 is designed to combine the control data 210c determined at the four previous points in time for controlling the movement actuators 160 and the corresponding rewards 210d into a data vector which, in the embodiment of <h2 style=";text-align:left;direction:ltr"> Figure 2aFor example, it has twelve components, i.e., it has dimensions of 1x12. For this data vector with 12 components, a feature vector of the actions, i.e., the control data and the rewards, with eight components is created in a substantially known manner by a subnetwork 250 with a "fully connected" layer and a ReLU layer.

[0030] In one embodiment, the processor 120 of the mobile device 110 is configured to determine a reward at each time step t based on the following equation, which represents a sum of multiple partial rewards: r t = r _ S t + r _ CD t + r _ C t + r _ G t + r _ F t .

[0031] In one embodiment, the partial reward r_{S}(t) has the constant value 0.1 and thus remains the same for each time step t. In one embodiment, the partial reward r_{CD}(t) has the value -0.1 if a collision with the load carrier 170 to be transported is detected for time step t, otherwise the value is 0. In one embodiment, the partial reward r_{C}(t) has the value -10 if a collision of the mobile device 110 with any object other than the load carrier 170 to be transported is detected in time step t, otherwise the value is 0. In one embodiment, the partial reward r_{G}(t) has the value 10 if the task is successfully completed in time step t, i.e., the load carrier 170 to be transported has been picked up, otherwise the value is 0. In one embodiment, the partial reward r_{F}(t) has the value -0.05 if the mobile device 110 does not move in a forward direction defined by the camera 130, otherwise the value is 0.

[0032] At the <h2 style=";text-align:left;direction:ltr"> Figure 2aIn the embodiment shown, the neural network 200 further comprises a device 260 which is designed to concatenate the second feature vector of the image data 210a with 256 components, the second feature vector of the LiDAR data 210b with 32 components and the feature vector of the actions and rewards with 8 components to create an overall feature vector with 296 components which, as one skilled in the art will recognize, contains information about the image data 210a, the LiDAR data 210b, the actions 210c and rewards 210d, at least in part for several points in time in the past.For this overall feature vector with 296 components, a first subnetwork 270 with a fully connected layer and a ReLU layer first creates a compressed overall feature vector with 64 components, and then a second subnetwork 280 with a fully connected layer and a ReLU layer creates the control data 290 for controlling the motion actuators 160, i.e., the actions for the next time step. As already described above, in the case of the system shown in . <h2 style=";text-align:left;direction:ltr"> Figure 2a In the illustrated embodiment, the control data 290 for controlling the movement actuators 160 comprises, for example, two components, namely a component for a linear movement and a component for a rotary movement of the mobile device 110.

[0033] In one embodiment, the processor 120 of the mobile device 110 is configured to implement a further neural network 200' for determining a Q-value 290', the architecture of which is described in <h2 style=";text-align:left;direction:ltr"> Figure 2b As is known to those familiar with reinforcement learning, the Q-value 290' represents an estimate of the cumulative reward that will be achieved due to the neural network 200 of <h2 style=";text-align:left;direction:ltr"> Figure 2a determine control data 290 for controlling the motion actuators 160 in the future. A mathematically exact definition of the Q-value 290' is described in the book "Reinforcement Learning: An Introduction", Richard S. Sutton and Andrew G. Barto, cited above, which is hereby incorporated by reference.

[0034] At the <h2 style=";text-align:left;direction:ltr"> Figure 2bIn the illustrated embodiment of the neural network 200' for determining the Q-value 290', the input data and the architecture of the neural network 200' are identical to the input data and the architecture of the neural network 200 of <h2 style=";text-align:left;direction:ltr"> Figure 2a to determine the control data 290 for controlling the motion actuators 160. In this embodiment, the two neural networks 200 and 200' differ only in that the weights of the individual sub-networks can be different and that the sub-network 280 generates two output values, namely the control data 290 for controlling the motion actuators 160, and the sub-network 280' provides one output value, namely the Q-value 290'. In one embodiment, the neural network 200 of <h2 style=";text-align:left;direction:ltr"> Figure 2a and the neural network 200' of <h2 style=";text-align:left;direction:ltr"> Figure 2buse the same subnetwork ("Custom ResNet") 220, i.e., with the same architecture and the same weights. As described in detail below, the subnetwork ("Custom ResNet") 220 may be pre-trained in one embodiment.

[0035] <h2 style=";text-align:left;direction:ltr"> Figur 2c shows the architecture of the subnetwork ("Custom ResNet") 220 of the <h2 style=";text-align:left;direction:ltr"> Figure 2a and 2b shown neural networks 200, 200' according to a preferred embodiment. As already described above, the subnetwork ("Custom ResNet") 220 is designed to generate the first feature vector of the image data 210a with, for example, 3200 components based on the multidimensional image data array 210a, which contains environmental information of the mobile device 110 at different times in the past. For this purpose, the subnetwork ("Custom ResNet") 220 according to the embodiment of <h2 style=";text-align:left;direction:ltr"> Figur 2cseveral consecutive residual blocks 222, 224, 226 and 228 with different dimensions, wherein a flatten function of the subnetwork ("Custom ResNet") 220 converts the multidimensional output of the last residual block 228 into the first feature vector of the image data 210a with, for example, 3200 components.

[0036] The general structure of the residual blocks 222, 224, 226, 228 used in the subnetwork ("Custom ResNet") 220 is known from the prior art. A preferred embodiment is described in <h2 style=";text-align:left;direction:ltr"> Figur 2c shown, specifically for the residual block 224 as an example. As already mentioned above, according to one embodiment, the other residual blocks 222, 226 and 228 may have substantially the same architecture as the residual block 224 with correspondingly adapted dimensions. As in <h2 style=";text-align:left;direction:ltr"> Figur 2cAs shown, the residual block 224 comprises, along a first partial path, a first sub-block which comprises a convolutional layer 224a (with a 3x3 kernel and padding), a ReLU layer 224b and a further convolutional layer 224c (with a 3x3 kernel and padding) and is designed by means of these layers to generate an output data array with the dimensions 128x40x40 from the input data array with dimensions 64x40x40. Along a parallel second partial path, the residual block 224 comprises a convolutional layer 224a (with a 1x1 kernel), which is designed to generate an output data array with the dimensions 128x40x40 from the input data array with dimensions 64x40x40.The residual block 224 further comprises an adder 342d, which is configured to combine, in particular to sum, the output data array having dimensions 128x40x40 generated by the first partial path with the output data array having dimensions 128x40x40 generated by the second partial path. The residual block 224 further comprises a second subblock, which comprises a ReLU layer 224e, a convolutional layer 224f (with a 2x2 kernel and without padding), and a further ReLU layer 224g, and is configured, using these layers, to generate an output data array having dimensions 128x20x20 from the data array having dimensions 128x40x40 provided by the adder 324d.

[0037] In one embodiment, the <h2 style=";text-align:left;direction:ltr"> Figur 2cThe subnetwork ("Custom ResNet") 220 shown here can be created and trained using an autoencoder known to those skilled in the art. Such an autoencoder is used to convert input data into compressed or coded data using a function (the so-called encoder) (which can be considered the extraction of features) and then to process the compressed data using another function (the so-called decoder) to reconstruct the original input data. If the error occurring during this reconstruction is small, the compressed or coded data represents a good representation of the input data (in other words, the compressed or coded data contains meaningful features of the input data).

[0038] In the <h2 style=";text-align:left;direction:ltr"> Figur 2cThe input data for the subnetwork ("Custom ResNet") 220 shown is the image data 210a, and the subnetwork ("Custom ResNet") 220 corresponds to the encoder, which generates the feature vector with, for example, 3200 components as compressed data. The decoder, which is not used in the application phase, can be, for example, a convolutional neural network.

[0039] During the training phase, the autoencoder, consisting of the encoder and decoder, can be pre-trained with the goal of minimizing the reconstruction error. In one embodiment, a database of images captured during a plurality of random movements of the mobile device 110 in a simulated environment at different times (e.g., every 180 ms) can be used to train the autoencoder and thus the subnetwork ("Custom ResNet") 220. After the autoencoder has been pre-trained in this way, according to one embodiment, only the now pre-trained encoder is used as the subnetwork ("Custom ResNet") 220 for the <h2 style=";text-align:left;direction:ltr"> Figure 2a and 2b shown neural networks 200, 200' are used. With the thus pre-trained subnetwork ("Custom ResNet") 220, the other subnetworks of the <h2 style=";text-align:left;direction:ltr"> Figure 2a and 2bshown neural networks 200, 200' for determining the control data 290 and the Q-value 290' can be trained faster and more efficiently.

[0040] <h2 style=";text-align:left;direction:ltr"> Figure 3 shows the architecture of another neural network 300, which according to one embodiment is implemented by the processor 120 of the mobile device 110. The neural network 300 is configured to perform a <h2 style=";text-align:left;direction:ltr"> Figure 2bcertain Q value 290' and one or more further geometric parameters (also referred to as "geometrical properties") 310 of a task given at a given time, i.e., task, to determine a probability of success 360 of the task. In one embodiment, the geometric parameters 310 can be, for example, a distance between the mobile device 110 and the load carrier 170 to be transported, a distance between the mobile device 110 and a nearest obstacle, a distance between the load carrier 170 to be transported and a nearest obstacle, and / or a relative angle between a reference direction and the direction from the mobile device 110 to the load carrier 170 to be transported.In one embodiment, the geometric parameters 310 may be one or more of the geometric parameters described as “geometric properties” in the article by Morad et al. “Embodied Visual Navigation with Automatic Curriculum Learning in Real Environments,” IEEE Robotics and Automation Letters, 2020.

[0041] As in <h2 style=";text-align:left;direction:ltr"> Figure 3As shown, the neural network 300 comprises an input layer 320, a first hidden layer 330, a second hidden layer 340, and an output layer 350. According to one embodiment, the input layer 320, the first hidden layer 330, and the second hidden layer 340 can each have one or more fully connected layers and one or more ReLU layers, and the output layer 350 can comprise one or more fully connected layers and one or more sigmoid layers. Further details of a possible architecture of the neural network 300 are described in the article by Morad et al., "Embodied Visual Navigation with Automatic Curriculum Learning in Real Environments," IEEE Robotics and Automation Letters, 2020, which is hereby incorporated by reference.

[0042] As one skilled in the art will recognize, the weights of both the neural network 200' for determining the Q-value 290' and the neural network 300 for determining the probability of success 360 are generally randomly distributed prior to training. During training, however, the neural network 300 learns to interpret its input variables, i.e., both the geometric parameters 310 and the Q-value 290', and thus to determine the probability of success 360. In one embodiment, a "soft actor-critic" algorithm can be used to train the neural network 200' for determining the Q-value 290'. As one skilled in the art knows, a "soft actor-critic" algorithm is a "deep RL" algorithm in which, on the one hand, a network for determining a policy, such as the network 200 of <h2 style=";text-align:left;direction:ltr"> Figure 2a , and secondly a network for determining a Q-value, such as the network 200' of <h2 style=";text-align:left;direction:ltr"> Figure 2b, is used. Further details on the "Soft Actor-Critic" algorithm are described in the book "Reinforcement Learning: An Introduction", Richard S. Sutton and Andrew G. Barto, cited above, which is hereby incorporated by reference. <h2 style=";text-align:left;direction:ltr"> Figure 2a and 2b However, the neural networks 200, 200' shown can also be trained using other "Deep RL" algorithms.

[0043] For example, if according to one embodiment a "Soft Actor-Critic" algorithm is used to train the neural networks 200, 200', in parallel the <h2 style=";text-align:left;direction:ltr"> Figure 3The variant of the NavACL algorithm underlying the neural network 300 shown can be executed in parallel to determine which task, i.e., task, is best suited to train the neural networks 200, 200' at a given time. Further details of such algorithms for automated "curriculum learning" (ACL) are described in the article by Morad et al., "Embodied Visual Navigation with Automatic Curriculum Learning in Real Environments," IEEE Robotics and Automation Letters, 2020, cited above.

[0044] In one embodiment, during the application phase, the weights of both the neural networks 200, 200' for determining the control data 290 and the Q value 290' and of the neural network 300 for determining the probability of success 360 are constant, i.e., according to one embodiment, no retraining takes place after training during the application phase of the mobile device 110. In an alternative embodiment, however, it is conceivable that training also takes place temporarily or continuously during the application phase of the mobile device, for example, in order to be able to readjust the weights of one or more of these networks 200, 200' and 300.

[0045] As already described above, during the application phase of the mobile device 110, the probability of success 360 can be determined for each task by means of the two neural networks 200' and 300. In one embodiment, the mobile device 110 is configured to perform the actions at a given time, ie, to use the control data 290 to control the movement actuators 160 for the movement of the mobile device 110, only if the probability of success 360 determined by means of the two neural networks 200' and 300 exceeds a threshold value. Otherwise, ieIf the probability of success 360 is too low, for example due to probable collisions of the mobile device 110 with load carriers 170 or other obstacles, the mobile device 110 can be configured not to perform the corresponding action and instead to send a warning signal, for example to a monitoring system installed in the warehouse 100.

Claims

1. Mobile apparatus (110) for the automated transport of a load carrier (170), wherein the mobile apparatus (110) comprises: at least one movement actuator (160) configured to be controlled by means of control data (290) to move the mobile apparatus (110), wherein the mobile apparatus (110) is configured to move and orient itself by means of the at least one movement actuator (160) relative to the load carrier (170) to be transported; an image capture device (130) for capturing a multiplicity of images (210a) of a section of an environment of the mobile apparatus (110) at different times; at least one LiDAR sensor (140) for capturing LiDAR data (210b) in the environment of the mobile apparatus (110); and a processor (120) configured to implement a neural network (200) trained by means of reinforcement learning, RL, wherein the RL-trained neural network (200) is configured to determine, on the basis of the multiplicity of images (210a), the LiDAR data (210b), a multiplicity of past control data items (210c) and a multiplicity of past rewards (210d) determined for the multiplicity of past control data items (210c), the control data (290) for controlling the at least one movement actuator (160) to move the mobile apparatus (110).

2. Mobile apparatus (110) according to Claim 1, wherein the processor (120) is also configured to implement a further RL-trained neural network (200'), wherein the further RL-trained neural network (200') is configured to determine a Q-value (290') on the basis of the multiplicity of images (210a), the LiDAR data (210b), the multiplicity of past control data items (210c) and the multiplicity of past rewards (210d).

3. Mobile apparatus (110) according to Claim 2, wherein the RL-trained neural network (200) comprises a subnetwork (220) and the further RL-trained neural network (200') comprises the subnetwork (220), which is configured to generate a feature vector on the basis of the multiplicity of images (210a).

4. Mobile apparatus (110) according to Claim 3, wherein the processor (120) is configured to train the subnetwork (220) in a first training section and to train the RL-trained neural network (200) and the further RL-trained neural network (200') in a second training section.

5. Mobile apparatus (110) according to Claim 4, wherein, in the second training section, the processor (120) is configured to train the RL-trained neural network (200) by means of tasks selected on the basis of the further RL-trained neural network (200').

6. Mobile apparatus (110) according to one of Claims 2 to 5, wherein the processor (120) is also configured to implement a further neural network (300), wherein the further neural network (300) is configured to determine, on the basis of the Q-value (290'), a probability of success (360) of the control data (290) for controlling the at least one movement actuator (160).

7. Mobile apparatus (110) according to Claim 6, wherein the further neural network (300) is configured to determine, on the basis of the Q-value (290') and one or more geometric parameters (310), the probability of success (360) of the control data (290) for controlling the at least one movement actuator (160).

8. Mobile apparatus (110) according to Claim 6 or 7, wherein the processor (120) is configured to control the at least one movement actuator (160) with the control data (290) if the probability of success (360) of the control data (290) is greater than a predefined threshold value.

9. Mobile apparatus (110) according to one of Claims 6 to 8, wherein the processor (120) is configured to bring the mobile apparatus (110) to a standstill if the probability of success (360) of the control data (290) is less than a predefined threshold value.

10. Mobile apparatus (110) according to one of the preceding claims, wherein the processor (120) is configured to train the RL-trained neural network (200) on the basis of a "soft actor-critic" algorithm.