Device and method for controlling a robot device
The method addresses the challenge of merging multiple sensor data types in robot control by using an encoding and fusion model with fusion layers, resulting in improved performance and efficient deep reinforcement learning for robotic manipulation tasks.
Patent Information
- Application Number
- DE102023201140
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Existing robot control methods struggle to efficiently merge and utilize multiple types of sensor data, such as vision, haptics, and self-perception data, which can lead to delayed or degraded training processes due to redundant or confusing information in image data.
A method for controlling a robotic device involves receiving sensor data from multiple types and processing it using an encoding and fusion model. This model consists of a sequence of encoding stages with fusion layers that combine features across different sensor data types, allowing for efficient extraction and merging of relevant information across multiple layers.
The proposed method enhances the performance of robot control tasks by effectively merging feature information across multiple layers, enabling efficient deep reinforcement learning for robot manipulation tasks and improving the overall control of robotic devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
State of the art
[0001] The present disclosure relates to apparatus and methods for controlling a robotic device.
[0002] The controller of a robot device can consider environmental state information (e.g., sensor data about the robot device's environment) as well as state information obtained from the robot's self-perception. Both types of state information (i.e., sensor data types) can be important for training an autonomous robot device. For example, two-dimensional images, along with the state information obtained from self-perception or haptic data, can also provide information for training a robot well.
[0003] However, images often contain a lot of redundant or even confusing information, which can delay or degrade the training process. Therefore, approaches for effective learning using image data together with state information from other sensor data types are desirable, e.g., approaches that enable the improvement of pixel-based deep reinforcement learning using additional sensor data. This means that such training should enable efficient fusion of different data modalities (i.e., sensor data types), e.g., vision, haptic, and self-perception data. In this context, a successful fusion technique would be able to efficiently extract relevant information from input sensor data (comprising sensor data corresponding to multiple sensor data types) regarding a state of the robot device and its environment to perform a given task (such as manipulating an object).
[0004] It is therefore an object of the invention to provide an improved control of a robot device. This object is achieved according to the claims.
[0005] Ashish Vaswani et al., “Attention is all you need.” in Proceedings of NeurIPS, pages 5998–6008, 2017, referred to below as Reference 1, describe a transformer network architecture.
[0006] DE 10 2021 118 885 B4 discloses machine learning for controlling object transfers.
[0007] DE 10 2020 128 653 A1 discloses a grasping determination for an object in disorder.
[0008] DE 11 2020 005 156 T5 reveals reinforcement learning of tactile grasping strategies.
[0009] DE 11 2020 005 799 T5 discloses an efficient utilization of a processing element array.
[0010] DE 10 2019134 794 B4 discloses a handheld device for training at least one movement and at least one activity of a machine.
[0011] DE 10 2018 215 826 A1 discloses a robot system and a workpiece gripping method. Disclosure of the invention
[0012] According to various embodiments, a method for controlling a robotic device is provided, comprising receiving sensor data for each of a plurality of sensor data types, processing the sensor data of the plurality of sensor data types by an encoding and fusion model comprising a sequence of encoding stages, each encoding stage comprising a coding layer for each of the sensor data types that generates features for the sensor data of the sensor data type, a plurality of fusion layers, each fusion layer combining features of the plurality of sensor data types generated by a corresponding one of the encoding stages and generating an input for the corresponding subsequent encoding stage in the sequence of encoding stages, and an output stage generating an output from an output of a last encoding stage of the sequence of encoding stages, selecting an action,to be performed by the robot device, using the generated output, and controlling the robot device to perform the selected action.
[0013] The method described above enables improving performance in a control task by fusing feature information across multiple layers under encoding processes of all multimodalities. This can be used, for example, in the context of deep reinforcement learning for a robot manipulation task.
[0014] Thus, fusion is performed at multiple coding stages (possibly after each coding stage in the sequence). Therefore, there is a high flow of information among the encoders (i.e., the coding paths, i.e., coding layer sequences, for the different sensor data types), enabling meaningful coding in terms of control. In particular, there may be trainable fusion layers after at least two (possibly all) of the "intermediate" coding layers, i.e., the coding layers that are not the last coding layers (i.e., the coding layers of the last stage in the sequence). Therefore, for at least one intermediate coding layer, there may be a fusion layer before the intermediate coding layer (i.e., a fusion layer that provides an output on which an input of the intermediate coding layer depends).The intermediate coding layers refer to trainable neural network layers, i.e., layers containing neurons, such as convolutional layers or fully connected layers. The final coding layer for each sensor data type can be a coding layer that performs pooling and flattening and may or may not be trainable.
[0015] Various examples are described below.
[0016] Example 1 is the method for controlling a robot device as described above.
[0017] Example 2 is the method of Example 1, wherein each fusion layer comprises at least one cross-attention layer.
[0018] This allows for an efficient combination of features related to the given task, i.e., control. For example, each fusion layer comprises a transformer-encoder (as described, for example, in Reference 1; the fusion layers can thus be implemented in a plug-and-play manner).
[0019] Example 3 is the method of example 1 or 2, wherein, for at least one of the fusion layers, the features of the multiple sensor data types comprise multiple components and the fusion layers mask some of the components before combining the features.
[0020] This allows for reducing the size of the cross-attention layer of the fusion layer, especially for early coding layers where feature dimensions are typically still very high (assuming that, as is typically the case, feature dimensions decrease over the sequence of coding stages). This can reduce computational costs. Furthermore, masking improves generalization and has a regularization effect.
[0021] Example 4 is the method of any one of Examples 1 to 3, wherein at least the coding layers of the coding stages, except for the last coding stage of the sequence of coding stages, comprise (or are) multilayer perceptrons or convolutional layers.
[0022] The coding layers of the final coding stage may or may not be multilayer perceptrons or a convolutional layer (they may also simply be a pooling and flattening layer). For example, image data may be encoded by convolutional layers, while haptic or self-perception data may be encoded by multilayer perceptrons. This enables efficient coding.
[0023] Example 5 is the method of any one of Examples 1 to 4, wherein the output stage comprises an additional fusion layer that combines features of the multiple sensor data types generated by the last encoding stage of the sequence of encoding stages.
[0024] In other words, a late fusion layer may be provided in addition to the fusion layers to perform the final fusion of the encoded sensor data (e.g., after concatenating the output of the final coding layers of the final coding stage of the sequence of coding stages).
[0025] Example 6 is the method of any one of Examples 1 to 5, comprising training the encoding and fusion model.
[0026] The encoding and fusion model (e.g., comprising one or more neural networks) may be trained (before being used for control or while being used for control), for example, using reinforcement learning, together with a control strategy according to which the action to be performed is selected using the generated output (which may, for example, be another neural network). In particular, the encoding and fusion model and other involved models (such as a model implementing the control strategy) may be trained together in an end-to-end manner.
[0027] Example 7 is a controller configured to perform a method of any one of Examples 1 to 6.
[0028] In particular, the controller is designed to implement the coding and fusion model and a control strategy for selecting actions to be performed by the robot device using outputs of the coding and fusion model.
[0029] Example 8 is a computer program comprising instructions that, when executed by a computer, cause the computer to perform a method of any of Examples 1 to 6.
[0030] Example 9 is a computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform a method of any of Examples 1 to 6.
[0031] In the drawings, similar reference characters generally designate the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being generally placed upon illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings, in which: Fig. 1 shows a control scenario according to one embodiment. Fig. 2 represents a late merger. Fig. 3 represents a transformer-based fusion. Fig. 4 shows an architecture of a transformer. Fig. Figure 5 illustrates the functionality of a transformer encoder. Fig. 6 illustrates transformer-based fusion according to one embodiment. Fig. 7 shows an architecture of a masked multimodal fusion transformer according to one embodiment. Fig. 8 shows a flowchart illustrating a method for controlling a robot device according to an embodiment.
[0032] The following detailed description refers to the accompanying drawings, which show, by way of illustration, specific details and aspects of this disclosure in which the invention may be practiced. Other aspects may be utilized, and structural, logical, and electrical changes may be made without departing from the scope of the present disclosure. The various aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form new aspects.
[0033] Various examples are described in more detail below.
[0034] Fig. 1 shows a robot 100.
[0035] The robot 100 comprises a robot arm 101, for example, an industrial robot arm for handling or assembling a workpiece (or one or more other objects). The robot arm 101 comprises manipulators 102, 103, 104 and a base (or support) 105 by which the manipulators 102, 103, 104 are supported. The term "manipulator" refers to movable elements of the robot arm 101, the actuation of which enables physical interaction with the environment, e.g., to perform a task. For control, the robot 100 comprises a (robot) controller 106, which is designed to implement the interaction with the environment according to a control program. The last element 104 (furthest from the support 105) of the manipulators 102, 103, 104 is also referred to as the end effector 104 and may comprise one or more tools, such as a welding torch, a gripping instrument, painting equipment or the like.
[0036] The other manipulators 102, 103 (closer to the support 105) can form a positioning device, so that, together with the end effector 104, the robot arm 101 is provided with the end effector 104 at its end. The robot arm 101 is a mechanical arm that can provide similar functions to a human arm (possibly with a tool at its end).
[0037] The robot arm 101 may include joint elements 107, 108, 109 that connect the manipulators 102, 103, 104 to each other and to the support 105. A joint element 107, 108, 109 may have one or more joint connections, each of which may provide rotational movement (i.e., a rotary movement) and / or translational movement (i.e., a translation) for associated manipulators relative to each other. The movement of the manipulators 102, 103, 104 may be initiated by means of actuators controlled by the controller 106.
[0038] The term "actuator" can be understood as a component adapted to influence a mechanism or process in response to being driven. The actuator can translate instructions issued by the controller 106 (called activation) into mechanical movements. The actuator, e.g., an electromechanical transducer, can be configured to convert electrical energy into mechanical energy in response to being driven.
[0039] The term "controller" can be understood as any type of logic implementation entity, which may include, for example, a circuit and / or a processor capable of executing software, firmware, or a combination thereof stored on a storage medium and capable of issuing instructions, in the present example, e.g., to an actuator. The controller may be configured, for example, through program code (e.g., software), to control the operation of a system, in the present example, a robot.
[0040] In the present example, the controller 106 includes one or more processors 110 and a memory 111 that stores code and data according to which the processor 110 controls the robot arm 101. According to various embodiments, the controller 106 controls the robot arm 101 based on a machine learning model 112 stored in the memory 111.
[0041] According to various embodiments, the machine learning model 112 is configured and trained to enable the robot 100 to perform a certain task, such as an insertion task, for example, inserting a plug 113 into a corresponding socket 114. To do so, the controller 106 captures images of its surroundings, here the plug 113 and the socket 114, using cameras 117, 119. The robot 100 (in particular, its controller 106) receives a visual observation of its surroundings.
[0042] Furthermore, the robot 100 may have self-perception information, ie, a self-perception state, as well as a haptic state (such as from a sensor in the end effector through which it can detect that it has grasped an object).
[0043] The controller 106 has data from multiple modalities at its disposal, in this example images, self-perception data, and other sensor data (e.g., haptic data).
[0044] A modality generally refers to a specific way of doing or experiencing something, and multimodality means a combination of two or more modalities. A modality refers to a source or form of information in the context of machine learning. Each modality has different information and perspectives on the surrounding environment (in the example from Fig. 1 the robot's workspace, comprising the plug 113 and the socket 114). Some modalities share redundant information (crossover), some provide missing information from a single modality (complementarity), and some even have a variety of different information interactions between them. Therefore, if the multimodal information can be appropriately fused, rich information (e.g., in the form of one or more features determined by a neural network) of the environment can be obtained.
[0045] There are three categories of fusion techniques: early, late, and intermediate fusion. • Early fusion: This naive approach is often used for tasks where the input modalities are RGB images and depth. Fusing these early would result in an RGB-D input with four channels. • Late fusion: Methods in this family often encode each modality separately and then fuse all the encoded features in a final latent layer. • Intermediate fusion: Information can also be fused at intermediate layers. A correct implementation of intermediate fusion can be superior to other fusion techniques because it allows for communication and exchange of information between modalities at (possibly all) intermediate layers, which can help learn a better latent representation of input data. According to various embodiments, this approach is applied to robot learning applications where, for at least one intermediate layer, fusion is performed before the intermediate layer.
[0046] In the above-mentioned Fig. 1, each element of the input data comprises, as mentioned, a visual observation, e.g., an RGB-D image I ∈ ℝ 4×H×W , a self-perception state x ∈ ℝ n and a haptic state y ∈ ℝ m. The robot 100 should learn a control strategy, e.g., using reinforcement learning, i.e., an RL-based strategy operating on such input data, i.e., the controller 106 optimizes an RL-based strategy using multimodal inputs, each of which has the form (l, x, y) in this example. Therefore, the goal of the RL problem is to find a strategy that maps from (l, x, y) to an action α∈A (i.e., in an action space) that maximizes task performance, i.e., maximizes a total return of rewards in RL settings (such as successfully inserting the plug 113 into the socket 114 within a certain time limit).
[0047] A fusion operation can be formulated in an abstract way as follows: assuming that a fusion function f maps an input data item (l, x, y) including all modalities to a D-dimensional fused feature θ ∈ ℝ D the control strategy can be defined as a mapping π=h∘f where h is implemented, for example, by one or more non-linear projection layers, e.g., MLPs (multilayer perceptrons), which are distributed over ℝ D into the action space.
[0048] It should be noted that the fusion function here is intended to include encoding input data. It is therefore also referred to as an encoding and fusion function (or model, pipeline, module, or (neural) network). According to various embodiments, two MLPs are used to encode the self-perception input x and the haptic input y, respectively, while a neural network consisting of convolutional layers encodes the visual input l. These neural networks (MLPs and convolutional neural network) can all be part of the machine learning model 112. The result of the encoding includes visual features θ v ∈ ℝ D×N×N (i.e., D-dimensional features of an N × N feature map), self-perception features θ p ∈ ℝ D and haptic features θh ∈ ℝ DIn addition to encoding, the fusion function f includes a fusion operation of these three modalities, called vision-self-perception-haptic fusion.
[0049] Fig. 2 represents a late merger.
[0050] As explained above, an input image 201 (here an RGB image) is encoded into visual features 205 by a convolutional neural network 202 (image or visual encoder) comprising convolutional layers 203 as encoding (intermediate) layers and an average pooling and flattening layer 204 (which can be considered as the final encoding layer).
[0051] A self-perception state 206 (e.g., comprising end effector position and gripper width) is encoded into self-perception features 209 by an MLP 207 (self-perception encoder), comprising a sequence of MLP layers (in particular, hidden layers) 208 as coding (intermediate) layers and a flattening layer 212 (which can be considered the final coding layer). The encoding of the haptic state information into haptic features is omitted for simplicity, but can be performed analogously to the self-perception state 206 by a haptic encoder.
[0052] A late fusion operation 210 (after the last coding layer) fuses the visual features 205 and the self-perception features 209 (and may similarly fuse haptic features). The results of the fusion are one or more fused features 211, which are the input to the function h of equation (1).
[0053] The late fusion operation 210 may require that all features (also referred to as latent features) be fused, including flattened features of the same length. Therefore, the average pooling and flattening layer 204 flattens the features output by the preceding convolutional layer 203. Alternatively, an additional MLP may be used to map D×N×N features to D, e.g., using pooling. This mapping generates the (flattened) visual features 205, for example, as g(θv)=θv*∈ℝD.
[0054] Examples of the late fusion operation 210 are • Mean and peak pooling: these two types of fusion are similar operations. The fused feature is called θmean=mean(θv*,θp,θh) or θmax=max(θv*,θp,θh) calculated. • Concatenation: this fusion operation is simply calculated as θcat=concat(θv*,θp,θh)∈ℝ3D • Bayesian fusion: The fusion technique uses Bayes' theorem to aggregate information from different inputs, i.e., using standard Gaussian conditioning. Essentially, it computes a posterior of the fused feature conditioning on input features. (θv*,θp,θh) Assume that each input feature is a Gaussian random variable with mean and variance (a diagonal variance), µ and σ 2 (assuming that each feature is divided into two parts, each with a length of D / 2). Specifically, the Gaussian distribution for each feature is (μv,σv2),(μp,σp2),(μh,σh2) represents, and assuming that a Gaussian prior of the fused distribution (μ0,σ02) , fusing them using Bayes' theorem results in a fused distribution (μout,σout2), where 1σout2=1σ02+∑i∈{v,p,h}1σi2 μout=μ0+σout2∑i∈{v, p, h}(μi−μ0)σi2.
[0055] Fig. 3 represents a transformer-based late fusion.
[0056] As referring to Fig. 2, an input image 301 is encoded into visual features 305 by a convolutional neural network 302 comprising convolutional layers 303 and an average pooling and flattening layer 304, and an eigenperception state 306 is encoded into eigenperception features 309 by an MLP 307 comprising a sequence of MLP layers (specifically, hidden layers) 308 and a flattening layer 312. Again, the encoding of the haptic state information into haptic features is omitted for simplicity.
[0057] The architecture from Fig. 3 differs from that of Fig. 2 in that a transformer 313 is included between the average pooling and flattening layer 304 and the preceding convolutional layer 303 as well as the flattening layer 312 and the preceding MLP layer 308.
[0058] Furthermore, as in Fig. 2, which processes features 305, 309 through a late fusion operation 310 that produces one or more fused features 311.
[0059] Fig. 4 shows an architecture of the transformer 313.
[0060] The transformer receives visual features (i.e., a feature map) 401 and self-perception features 402 and arranges them into a vector of tokens 403, which is then processed by a transformer encoder 404 into a result vector 405, which is separated into a visual feature result vector (output to the average pooling and flattening layer 304) and an self-perception feature result (output to the flattening layer 312).
[0061] In the case of visual, self-perceptual and haptic features (θ v , θ p1 θ h ) output by the last convolutional layer or MLP layer, they are rearranged to obtain N × N + 2 tokens and written as the vector of tokens Θ in ∈ ℝ (N2+2)×D which is processed by the transformer encoder 404.
[0062] The transformer 313, in particular the transformer encoder 404, may be configured, for example, as described in Reference 1.
[0063] Fig. 5 illustrates the functionality of the transformer encoder 404.
[0064] The multimodality input embedding 501, also known as F in , corresponding to the vector of tokens Θ in , is provided with a positional embedding 502 (to reflect the position of the tokens within, for example, the image feature map) and normalized 503. It is then fed into a multi-headed attention module 504, whose output is added to its input and normalized 505, processed by an MLP 506, whose output is added to its input and normalized 507.
[0065] The multi-headed attention is (as on the right side of Fig. 5) is determined by a set of queries, keys and values, denoted as (Q, K, V) and weighted by M q ∈ ℝ D × Dq , M k ∈ ℝ D × Dk , M v ∈ ℝ D × Dv be parameterized as Q=ΘinMq, K=ΘinMk, V=ΘinMv.
[0066] A scaled dot product attention module 509 determines the attention weights as A=softmax(QKTDk)V. The attention weights are concatenated 510. Linear projections 508 and 511 can be performed before the scaled dot product attention module 509 and after concatenation.
[0067] The output of the transformer-decoder is finally determined by the MLP 506 with residual connections and layer norms (LN) as Fout=LN(MLP(A)+Fin)
[0068] Multiple projection heads or multi-layer transformer layers can also be used. The output Fout has a similar shape to the input ℝ (N2+2) × D and is rearranged to again have three embeddings (θ v , θ p; θ h ). For example, average pooling is used on the visual embedding to receive a D-dimensional feature, then concatenate it with the other two embeddings to obtain a 3*D-dimensional feature, which is then used as input to the function h of the strategy of equation (1).
[0069] According to various embodiments, in order to fully exploit information communication across modality encoders, information communication across all (or at least several) coding layers is possible.
[0070] Fig. 6 illustrates transformer-based fusion according to one embodiment.
[0071] As referring to Fig. 2, an input image 601 is encoded into visual features 605 by a convolutional neural network 602 comprising convolutional layers 603 and an average pooling and flattening layer 604, and an eigenperception state 606 is encoded into eigenperception features 609 by an MLP 607 comprising a sequence of MLP layers (specifically, hidden layers) 608 and a flattening layer 612. Again, the encoding of the haptic state information into haptic features is omitted for simplicity.
[0072] The architecture from Fig. 6 differs from that of Fig. 2 in that, for multiple coding layers, a transformer 613 is included after the corresponding coding layer (of all encoders), which combines the features output by the coding layers. This means that the outputs of the coding layers of a certain coding (intermediate) stage, e.g., the i-th coding layer of each encoder, are fed into a transformer 613 and fused.
[0073] A transformer 603 may in particular be arranged before an intermediate coding stage (as in Fig. 6 before the third coding stage). For such a coding stage, for each encoder, the result of the fusion by the transformer 603 is added to the output of the previous coding layer, and the result of the addition is fed to the coding layer of the encoder.
[0074] The layers of the modality encoder are assumed to have the following forms • visual encoder: (D 1 × H 1 × W 1 ),..., (D L × H L × W L ) • Self-perception encoder: (D 1 ),...,(D L ) • Haptic encoder: (D 1 ),...,(D L )
[0075] The features output by the coding layers of, e.g., the i-th stage are fused by the corresponding transformer 613 using a cross-attention operation, as described with reference to Fig. 5. The output after fusion after coding level i has the form Fouti∈ℝ(Hi*Wi+2), which can be rearranged to include three corresponding intermediate features (θvi,out, θpi,out, θhi,out) as input to the next coding stage i + 1 (possibly added to the output of coding stage i, as mentioned above). The features are further coded by the coders (visual, self-perception, and haptic) to produce features (θvi+1,in, θpi+1,in, θhi+1,in) as output of the i+1-th coding stage (which can be fed into a transformer 613 from (i.e., to) the i+1-th coding stage, hence the superscript "in") and so on.
[0076] The transformers 613 can be configured like the transformer 313 described with reference to Fig. 5. However, according to various embodiments, the transformers are masked multimodal fusion transformers (MMFTs).
[0077] Fig. Figure 7 shows an architecture of an MMFT 613.
[0078] Like the transformer 313, which refers to Fig. 3, the MMFT 613 receives visual features (i.e., feature map) as well as self-perception features 702. However, before these are arranged into a vector of tokens 703, which is then processed by the transformer-encoder 704 into a result vector 705, which is separated into a visual feature result vector (output to the average pooling and flattening layer 704) and an self-perception feature result (output to the flattening layer 612), the visual features are partially masked by a masking operation 706 to partially masked features 701 (i.e., a partially masked feature map where features for certain patches are set to zero; e.g., there is a D i -dimensional feature for each patch for the MMFT 613 after the i-th coding level).
[0079] This means that random patches of the visual feature map are masked by the MMFT 613 (for each stage i where there is an MMFT). Therefore, fusion at each stage only works on the visible patches. Masking with a high proportion of features can both provide higher performance due to strong regularization and reduce computational costs due to attention operations on high-dimensional inputs, especially at early layers.
[0080] Each MMFT 613 can mask the visual features with a predefined ratio. The visual features remaining after masking (i.e., the features of unmasked patches) are flattened and concatenated with the self-perception features 602 as input to the transformer-encoder 604. Afterward, the output of the transformer-encoder 604 is rearranged so that results for unmasked patches go into their corresponding patch location, while the masked patches are zero in the result vector 605.
[0081] The transformer encoder 604 may be configured as described with reference to Fig. 5 is described.
[0082] All fusion techniques described above refer to the function f of Equation (1). Accordingly, any actor-critical algorithm can be used to learn the strategy π from end to end, as defined in Equation (1). For example, soft actor criticism (SAC), which is an out-of-strategy algorithm, can be used to optimize π.
[0083] In summary, according to various embodiments, a method is provided as described in Fig. 8 shown.
[0084] Fig. 8 shows a flowchart 800 illustrating a method for controlling a robotic device according to an embodiment.
[0085] In 801, sensor data is received for each of the multiple sensor data types.
[0086] In 802, the sensor data of the multiple sensor data types are processed by a coding and fusion model. The coding and fusion model includes • a sequence of coding levels, each coding level comprising a coding layer for each of the sensor data types that generates features for the sensor data of the sensor data type; • a plurality of fusion layers, each fusion layer combining features of the plurality of sensor data types generated by a corresponding one of the coding stages and generating an input for a corresponding subsequent coding stage in the sequence of coding stages; and • an output stage that generates an (output stage) output (ie, encoded sensor data or, in other words, a latent representation of the sensor data) from an (encoding stage) output of a last encoding stage of the sequence of encoding stages.
[0087] In 803, an action to be performed by the robot device is selected using the generated (output stage) output (ie, the output of the output stage).
[0088] In 804, the robot device is controlled to perform the selected action.
[0089] The approach from Fig. 8 can be used to calculate a control signal for controlling a technical system, such as a computer-controlled machine, such as a robot, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system. According to various embodiments, a strategy for controlling the technical system can be learned, and then the technical system can be operated accordingly.
[0090] Various embodiments may receive and use image data (i.e., digital image data) from various visual sensors (cameras), such as video, radar, LiDAR, ultrasound, thermal imaging, sonar, etc., as well as other sensor data types, such as pressure, force, etc.
[0091] According to one embodiment, the method is computer-implemented.
Claims
[1] A method (800) for controlling a robotic device (106), comprising: Receiving (801) sensor data for each of a plurality of sensor data types; Processing (802) the sensor data of the plurality of sensor data types by a coding and fusion model comprising a sequence of coding stages, each coding stage comprising a coding layer for each of the sensor data types that generates features for the sensor data of the sensor data type; a plurality of fusion layers, each fusion layer combining features of the plurality of sensor data types generated by a corresponding one of the coding stages and generating an input for a corresponding subsequent coding stage in the sequence of coding stages; and an output stage generating an output from an output of a last coding stage of the sequence of coding stages; selecting (803) an action to be performed by the robot device (106) using the generated output; and Controlling (804) the robot device (106) to perform the selected action. [2] The method (800) of claim 1, wherein each fusion layer comprises at least one cross-attention layer. [3] The method (800) of claim 1 or 2, wherein, for at least one of the fusion layers, the features of the multiple sensor data types comprise multiple components and the fusion layers mask some of the components before combining the features. [4] The method (800) of any one of claims 1 to 3, wherein at least the coding layers of the coding stages, with the exception of the last coding stage of the sequence of coding stages, comprise multilayer perceptrons or convolutional layers. [5] The method (800) of any one of claims 1 to 4, wherein the output stage comprises an additional fusion layer that combines features of the multiple sensor data types generated by the last coding stage of the sequence of coding stages. [6] The method (800) of any one of claims 1 to 5, comprising training the coding and fusion model. [7] Controller designed to carry out a method according to one of claims 1 to 6. [8] A computer program comprising instructions which, when executed by a computer, cause the computer to perform a method according to any one of claims 1 to 6. [9] A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform a method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Robot system and workpiece gripping method
DE102018215826A1
Generative modeling of neural networks for transforming speech utterances and extending training data
DE102019107928A1
Handheld device for training at least one movement and at least one activity of a machine, system and procedure.
DE102019134794B4
Grasping point for an object in disorder
DE102020128653A1
ESTIMATION OF VEHICLE DAMAGE
DE102021101082A1