Method for adapting a machine learning model to a varied control condition

By detecting sensor data elements under changed control situations and generating their enhancements, combining the output of the machine learning model to determine the target output, calculate the loss and adapt the model, the machine learning model adaptation problem is solved, and efficient adaptation and grabbing performance improvements without using annotated data.

CN120065713APending Publication Date: 2025-05-30ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411738364.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to adapt the machine learning model to the changed control situation without using annotated training data, resulting in a degradation of the grab performance.

Method used

By detecting sensor data elements, multiple enhancements are generated under changed control situations, the first instance of the machine learning model generates output for each enhancement, the output is combined to determine the target output, the loss between the second instance output and the target output, and the machine learning model is adapted to reduce the overall loss.

Benefits of technology

This enables the machine learning model to be adapted to the changed control situation without the need for supervised training and annotated training data, thereby improving the grab performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065713A_ABST
    Figure CN120065713A_ABST
Patent Text Reader

Abstract

According to various embodiments, a method for adapting a machine learning model to a changed control situation is described, the method comprising: detecting a sensor data element under a changed control situation; generating, for each determined sensor data element, a plurality of enhancements of the sensor data element; generating a respective output for each enhancement using a first instance of the machine learning model; determining a target output for the sensor data element by combining the generated outputs; and determining a loss between an output for a second instance of the sensor data element and the determined target output; and adapting the second instance of the machine learning model to reduce overall loss, wherein the overall loss includes the determined loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for adapting a machine learning model to a changed control situation. Background Art

[0002] Picking (i.e., grasping) an object is an important problem in robotics. Newer methods use machine learning to achieve model-free grasping of a large number of unseen objects. In practical applications, for example, when removing an object from a container, the performance of these methods, i.e., the performance of a correspondingly trained machine learning model usually depends on the conditions in the corresponding control situation, and every change in the camera, the object, or the surroundings (compared to the situation for which the machine learning model was trained) can have a negative impact on the grasping performance. To achieve reliable control (i.e., high grasping performance) even under changed conditions, supervised learning can be used to retrain (nachtrainieren) the machine learning model with the corresponding training data to adapt it to the corresponding situation. However, this requires a great deal of effort because additional training data (e.g., images) with relevant annotations (i.e., "labels") must be generated. Summary of the Invention

[0003] Therefore, it is desirable to have a method that can adapt a machine learning model to a changed control situation without much effort.

[0004] According to various embodiments, there is provided a method for adapting a machine learning model to a changed control situation (compared to the control situation for which the machine learning model was trained), the method comprising:

[0005] · Detecting sensor data elements in the changed control situation;

[0006] · For each determined sensor data element

[0007] o Generating a plurality of augmentations of the sensor data element;

[0008] o Using a first instance of the machine learning model to generate a corresponding output for each augmentation;

[0009] o Determining a target output for the sensor data element by combining the generated outputs; and

[0010] o Determining the loss between the output of a second instance for the sensor data element and the determined target output; and

[0011] ·Adapt a second instance of the machine learning model to reduce the overall loss, where the overall loss includes the determined loss.

[0012] The above method enables self-supervised test-time adaptation, i.e., adapting a machine learning model to inference conditions that have changed relative to training, e.g., to improve the performance of a changing neural grasping prediction network (Greifvorhersagenetz) of a camera that provides input images for the grasping prediction network (e.g., when recording input images using a new camera type or changing the installation situation), without the need for supervised training of the machine learning model using an annotated training dataset.

[0013] Different embodiments are described below.

[0014] Embodiment 1 is the method for adapting a machine learning model to a changed control situation as described above.

[0015] Embodiment 2 is the method according to Embodiment 1, which includes adapting a first instance of the machine learning model in the direction of an adapted second instance of the machine learning model.

[0016] Thus, the first instance (also referred to as the teacher model or specifically as the teacher network in the following example) can follow the second instance (also referred to as the student model or specifically as the student network in the following example) at a specific batch interval, as described in the following example.

[0017] Embodiment 3 is the method according to Embodiment 1, which for each batch in the batch sequence includes:

[0018] ·Detect the corresponding sensor data element in the changed control situation;

[0019] ·For each sensor data element determined for the batch:

[0020] o Generate multiple augmentations of the sensor data element;

[0021] o Generate a corresponding output for each augmentation by feeding the generated augmentations into the corresponding first instance of the machine learning model (for the batch);

[0022] o Determine the target output for the sensor data element by combining the generated outputs; and

[0023] o Determine the loss between the output of the corresponding second instance of the machine learning model (for the batch) for the sensor data element and the determined target output; and

[0024] ·Adapt the corresponding second instance of the machine learning model to reduce the overall loss, where the overall loss includes the determined loss,

[0025] For each batch in the sequence except the last batch, the corresponding adapted second instance of the machine learning model is used as the second instance of the machine learning model for subsequent batches in the sequence.

[0026] Thus, the adaptation described in the above method can relate to batches (i.e., the detected sensor data elements are sensor data elements of a batch) and the adaptation can be repeated accordingly for additional batches, where the second instance is continuously adapted during the course of the sequence, so that the accuracy of the machine learning model increases over time (e.g., during continuous operation). As described above, the first instance can follow the second instance:

[0027] Example 4 is a method according to Example 3, the method comprising adapting a first instance of a machine learning model in the direction of a second instance of the machine learning model after a predetermined number of batches.

[0028] The number of batches can also be one, i.e., the first instance can follow the second instance immediately, but e.g. in a weighted manner (see example below), such that the first instance still follows the second instance slowly. This ensures stability during the adaptation process. For the first batch of the sequence, the first instance and the second instance can be set as the machine learning model to be adapted.

[0029] Example 5 is a method according to one of Examples 1 to 4, wherein the sensor data element is an image data element.

[0030] In this case, the image data element is understood as a data element in the form of a matrix having one or more channels (i.e., each position in the matrix, i.e., each "pixel" has one or more values). Such a sensor data element can effectively represent the scene to be controlled. The machine learning model is (or includes) e.g. a convolutional neural network.

[0031] Example 6 is a method according to Example 5, wherein the change in the control situation to which it is adapted is: a change in the camera, and / or a change in one or more conditions of the image recording using the camera (114) for recording these image data elements.

[0032] Thus, the above method can be used to adapt a machine learning model to changes in the camera or image recording conditions (e.g., changes in illumination, color shift, etc.) with less training effort and without the need for explicit annotation of the sensor data elements (but using generated pseudo-labels, i.e., target outputs).

[0033] Example 7 is a method according to Example 5 or 6, wherein the image data elements have a plurality of channels, and generating a corresponding output for each enhancement and generating an output of a second instance for each sensor data element comprises: a trainable scaling of the respective values of these channels, wherein the scalings are adapted together to reduce the overall loss.

[0034] Example 8 is a method for controlling a robotic device, the method comprising:

[0035] · adapting a machine learning model to a control situation in which the robotic device is to be controlled by a method according to one of Examples 1 to 7;

[0036] · detecting one or more other sensor data elements in the control situation;

[0037] · processing the one or more other sensor data elements by an adapted second instance of the machine learning model or by a first instance of the machine learning model adapted in the direction of the adapted second instance; and

[0038] · generating a control signal for the robotic device based on the processing result.

[0039] Example 9 is a data processing device (in particular a control device for a robotic device) arranged to execute a method according to one of Examples 1 to 8.

[0040] Example 10 is a computer program having instructions which, when executed by a processor, cause the processor to execute the method according to one of Examples 1 to 8.

[0041] Example 11 is a computer-readable medium storing instructions which, when executed by a processor, cause the processor to execute the method according to one of Examples 1 to 8. Description of the Drawings

[0042] In the drawings, like reference numerals generally refer to the same parts in different views. The drawings are not necessarily to scale, but generally focus on illustrating the principles of the invention. In the following description, different aspects will be described with reference to the following drawings.

[0043] Figure 1 A robot is shown.

[0044] Figure 2 A test-time adaptation method for a machine learning model according to one embodiment is shown.

[0045] Figure 3A flowchart showing a method for adapting a machine learning model to an altered control situation according to an embodiment is shown.

[0046] The following detailed description refers to the accompanying drawings which, for purposes of explanation, show specific details and aspects of the present disclosure in which the invention can be implemented. Other aspects can be used and structural, logical, and electrical changes can be made without departing from the scope of the invention. Different aspects of the present disclosure are not necessarily mutually exclusive, as some aspects of the present disclosure can be combined with one or more other aspects of the present disclosure to form new aspects. Detailed Description

[0047] Different examples are described in more detail below.

[0048] Figure 1 A robot 100 is shown.

[0049] The robot 100 includes a robot arm 101, such as an industrial robot arm, for manipulating or assembling workpieces (or one or more other objects). The robot arm 101 includes movable arm elements 102, 103, 104 and a base (or support (Stütze)) 105 for supporting the arm elements 102, 103, 104. The term "movable arm element" refers to the movable components of the robot arm 101, the manipulation of which enables physical interaction with the surrounding environment (Umgebung) in order to, for example, perform tasks. For control, the robot 100 includes a (robot) control device 106 which is designed to implement the interaction with the surrounding environment according to a control program. The last arm element 104 of the arm elements 102, 103, 104 (which is the furthest from the support 105) is also referred to as the end effector (Endeffektor) 104 and can include one or more tools, such as a welding torch, a gripping tool, a painting device, etc.

[0050] The other arm elements 102, 103 (which are closer to the support 105) can form a positioning device such that, together with the end effector 104, the robot arm 101 with the end effector 104 at its own end is provided. The robot arm 101 is a robotic arm (possibly with a tool at its end).

[0051] The robotic arm 101 may include joint elements 107, 108, 109 that connect arm elements 102, 103, 104 to each other and to a support 105. The joint elements 107, 108, 109 may include one or more joints that can respectively provide rotational movement (i.e., rotational motion) and / or translational movement (i.e., displacement (Verlagerung)) of the associated arm elements relative to each other. The movement of the arm elements 102, 103, 104 may be initiated by actuators controlled by a control device 106.

[0052] The term "actuator (Aktuator)" may be understood as a component configured to affect a mechanical mechanism (Mechanismus) or process in response to being driven. The actuator may implement instructions (so-called activation) created by the control device 106 as mechanical motion. An actuator, such as an electromechanical converter, may be designed to convert electrical energy into mechanical energy in response to being activated.

[0053] The term "control device" may be understood as any type of entity implementing logic (including one or more computers), which may for example include circuits and / or processors, firmware, or a combination thereof capable of executing software stored in a storage medium, and capable of outputting instructions to, for example, the actuators in this example. For example, the control device may be configured by program code (such as software) to control the operation of a system that is a robot in this example.

[0054] In this example, the control device 106 includes one or more processors 110 and a memory 111 that stores the code and data based on which the processor 110 controls the robotic arm 101. According to various embodiments, the control device 106 controls the robotic arm 101 based on a machine learning model 112 stored in the memory 111.

[0055] According to various embodiments, the machine learning model 112 is designed and trained to enable the robot 100 to recognize manipulation poses at one or more objects 113, at which the robot 100 can pick up the one or more objects 113 (or otherwise interact with them, such as painting).

[0056] For example, the robot 100 may be equipped with one or more cameras 114 that enable the robot to record images of its workspace. For example, the camera 114 is fixed to the robotic arm 101 such that the robot can take images of the object 113 from different perspectives by moving (herumbewegen) its robotic arm 101 around. As Figure 1 shown, the camera 114 may also be fixedly assembled in the robot unit to detect the object to be grasped.

[0057] According to various embodiments, the machine learning model 112 is a neural network and the control device 106 feeds input data based on the one or more digital images of the object 113 (depth images with optional color images or point clouds with optional color images or also including other per-pixel information, such as information about surface normals) to the neural network 112, and the neural network (specifically a neural "grasp prediction network" in this example) determines, for example, for each of the multiple (on the surface) parts of the object, a quality that indicates how well the object can be grasped at the corresponding part. Instead of continuous values (i.e., instead of regression), the machine learning model 112 (e.g., a neural network) can also perform classification, for example, classify as "good parts for grasping" and "bad parts for grasping". It can also output other continuous values, for example, for each part, output the orientation for the end effector 104 (exemplarily assumed to be a gripper hereinafter), and then for each orientation output the (manipulation or hereinafter grasping) quality that indicates how well the object can be manipulated (exemplarily grasped) at this part with this orientation.

[0058] It may happen that the machine learning model 112 has been trained for a specific (control) situation (i.e., under specific conditions), but then should be used in other situations (referred to as "test time", and this also includes inference during use). For example, the camera 114 can be replaced, such that the way the object 113 is represented in the input data of the machine learning model 112 changes (e.g., lens properties, noise behavior, etc. change). To compensate for this, so-called test-time adaptation can be performed. Further examples of this are: compensating for changes in object properties or ambient conditions (such as lighting conditions), the composition and number of the corresponding object, or the positioning of the camera relative to the corresponding object. Generally speaking, whenever there is a domain shift (or domain gap) between the training data and the test-time data, test-time adaptation can always be applied.

[0059] According to various embodiments, test-time adaptation is provided, the goal of which is to adapt a pre-trained ML (machine learning) model for inference, for which no training data annotated (with ground truth, i.e., usually labels) is required. For example, this test-time adaptation is performed to adapt the neural grasp prediction network to a camera change (such as due to the replacement of the camera 114). According to various embodiments, for example, the test-time adaptation is used to enable a neural network for per-pixel grasp prediction to adapt to a camera change (or "domain shift" ") between training and test time without supervised (re)-training and thus without additional annotation effort.

[0060] Specifically, according to various embodiments, the mean teacher method is used for self-supervised test-time adaptation of a machine learning model (wherein input image channel scaling may be incorporated). The "mean teacher" (i.e., the machine teacher model) provides pseudo-labels (or soft pseudo-labels), which are used to adapt, for example, network weights and batch normalization statistics (Batch-Normalization-Statistiken) of a convolutional network (convolutional neural network CNN) serving as a grasping prediction network. This method can be used for real-time adaptation or during an initial adaptation phase to update network weights and batch normalization statistics according to, for example, new input images from a new camera type. The method is not limited to a specific CNN network architecture.

[0061] Thus, according to different embodiments, the mean teacher framework with test-time augmentation and image channel scaling is used to achieve robust network predictions for new input images with unknown domain shifts (e.g., from a new unknown camera type) even during model runtime. From this, the following aspects are derived:

[0062] (1) Given that a grasping prediction network trained with image data from one or more known cameras can be adapted to images with domain shifts (e.g., images from an unknown camera type or unknown camera mounting location or unknown properties of the grasped object, such as surface reflection, color shift, etc.) through this self-supervised retraining method without retraining the labeled image data.

[0063] (2) The method can be used in an offline or online setting to either adapt to a new domain shift using a fixed set of input images during an initial adaptation phase or continuously adapt during application runtime.

[0064] (3) Integrating test-time augmentation into the mean teacher prediction (i.e., pseudo-labels) enables the generation of robust pseudo-labels by exploiting the phenomenon that even after training with augmented training samples, the CNN is not completely invariant with respect to the symmetry in the data distribution of new input data.

[0065] (4) Integrating learnable channel scaling into the test-time adaptation method enables automatic adaptation of the weight factors for each image input channel (e.g., RGB, depth, etc.). This results in more robust predictions in cases where there is an imbalance in domain shift between input channels, such as when a new RGB-D camera type provides RGB images of similar quality but lower-precision depth images.

[0066] Figure 2Shows a test-time adaptation method for a machine learning model according to one embodiment.

[0067] The test-time adaptation method can be applied to, for example, neural networks, such as various CNN network architectures. For example, the machine learning model maps an input image (e.g., RGB or RGB-D (RGB plus depth information)) to an output having the same resolution as the input image. As described above, the output can be a classification (e.g., marking good or bad grasping positions in the input image or also object classification in autonomous driving) or a continuous value (e.g., probability of stable grasping). In the test-time adaptation method, two versions of the machine learning model are used: the "student network" 202 and the "teacher network" 203. Both are initialized to the machine learning model to be adapted (i.e., initially both are identical to the machine learning model to be adapted). The weights of the machine learning model before its adaptation (and thus the initial weights of the student network 202 and the teacher network 203) are denoted by denoted.

[0068] The input to the test-time adaptation method is a batch 201 of B input images which are, for example, recorded by a camera different from the (ones) used to record the images for training the machine learning model 201 (assumed to be a neural network hereinafter).

[0069] For the teacher network 203, for each input image, N different augmented input images are generated from the input image by using N different augmentation transforms. In this case, depending on the type of image channels (e.g., RGB, depth), different image augmentation techniques (e.g., size change, mirroring, adding noise, etc.) can be applied separately for each image channel for each augmentation. For each augmented input image (i.e., each augmentation), different image channels are then scaled by a learnable vector γ using a function sc (representing "scale") so that the test-time adaptation method can adapt the scaling for each input channel (i.e., the image channels in this example) at test time:

[0070] where C is the number of different input channels (e.g., RGB, depth, etc.).

[0071]

[0072] where,

[0073]

[0074] where C is the number of different input channels (e.g., RGB, depth, etc.).

[0075] This enables consideration of different domain transfers between input channels. For example, compared to the camera used to record images for training the network, there are different domain transfers between the RGB and depth channels in the case of a new camera.

[0076] The enhanced individual images are guided through the teacher network 203 after being scaled, which results in N different teacher outputs (output images in this embodiment):

[0077]

[0078] where θ T are the weights of the teacher network 203, which are initialized as described above by the weights and are initialized.

[0079] Then, the various outputs are combined into a single output by averaging 204

[0080] where avg represents various possible averaging techniques for calculating the average at each pixel coordinate (n, m), such as the arithmetic mean:

[0081]

[0082] or the geometric mean:

[0083]

[0084] However, weighted averaging can also be performed (e.g., different weights are applied to the augmentations depending on the augmentation transformation used (e.g., the type and / or intensity of the transformation)).

[0085] To achieve the merging of the output images, the inverse augmentation transformation can be applied to each output image (corresponding to the augmentation transformation used to generate the augmented input image that has been processed by the teacher network 203 into the output image).

[0086] Augmented average (which is used as the pseudo-label for training the machine learning model 202) The basic idea can be regarded as: achieving high-quality pseudo-labels by merging the output images calculated using the augmented versions of the same input image. Especially for convolutional networks, this "Ensembling" takes advantage of the phenomenon that even after training with augmented training samples, the convolutional network (CNN) is not completely invariant to augmentations.

[0087] In the case of the student network 202, the generation of the augmented input image is optional and can be obtained from one of the augmentations of the teacher network 203.

[0088] It is also possible to use multiple augmentations of the student network 202 and average the output images of the student network 203, as described for the teacher network 203, i.e.,

[0089] The different image channels of the student network 202 are then scaled by the learnable vector γS so that the test adaptation method can adapt the scaling of each input channel at test time:

[0090]

[0091] where

[0092]

[0093] and C are as described above.

[0094] Then, the augmented and thus scaled input images are passed through the student network 202 with network weights θS, which are initialized as described above by the pre-trained source model weights . If the input images are generated using the augmentation transformation, the corresponding inverse augmentation transformation is applied to o before calculating the loss 205 j .

[0095] Finally, the output of the teacher path and the output o j of the student path are used to calculate the (overall) loss 205 for the batch. The loss 205 is a consistency loss and is given, for example, by:

[0096]

[0097] where B is the number of input images in the batch, H is the height, W is the width of the output image, and l is the per-pixel loss between the j-th output of the student path and the teacher path. This value can be regarded as the loss (or "single loss (Einzelverlust)") for the j-th input image (usually a sensor data element).

[0098] The per-pixel loss l can be represented by different loss functions, such as I1 loss, I2 loss, or cross-entropy loss (depending on the type of pixel values).

[0099] The student network 202 is updated by backpropagating the loss 202. This loss enforces the consistency between the student network output and the pseudo-label.

[0100] According to one embodiment, the teacher network 203 is updated with an Exponential Moving Average (EMA): At each training step of the teacher network (index t, where, for example, the training step of the teacher network is performed after a specific number of batches), the weights of the student model 202 are used to update the teacher network 203, which results in: The teacher network 203 that is continuously trained and averaged over time. Since it can be assumed that the predictions of the (average) teacher network 203 are more accurate than the outputs of the student network 202, they can (as described above) be used as pseudo-labels for the self-supervised training of the student network 202. For example, the weight update of the teacher model 203 is as follows

[0101]

[0102] where 0 < α < 1 defines the mixing ratio of the student and teacher weights and causes smoothing. In this case, a higher value results in a slower moving average, which is useful, for example, when adapting to a large number of adaptation steps.

[0103] Furthermore, the batch normalization layers of the student network and the teacher network are in training mode during test-time adaptation and are re-estimated separately based on the test-time data.

[0104] Various methods are described below to adapt the grasping prediction network (e.g., to a new camera) using the test-time adaptation method described with reference Figure 2 to adapt the grasping prediction network (e.g., to a new camera) using the test-time adaptation method described.

[0105] There are three main variants for using the results of the test-time adaptation method for grasping prediction in a new control situation:

[0106] 1. Use the grasping prediction network with the batch normalization statistics and weights of the student network

[0107] 2. Use the grasping prediction network with the batch normalization statistics and weights of the teacher network and

[0108] 3. During the inference process, use the grasping prediction network with the batch normalization statistics and weights of the teacher network and (as described for the teacher network) the enhanced average. The batch normalization statistics are the parameters of the batch normalization (BN) layer of the corresponding neural network. Strictly speaking, these are not trained weights but are calculated for the training data during the training process. These BN statistics are specific to the training data. If there is a domain shift compared to the test-time data, then (in addition to the weights) the BN statistics can also be adapted in the same way (e.g., in a manner similar to the weights described above).

[0109] In Case 1, after self-supervised training, the student network 202 should result in more robust predictions than the original machine learning model. In Case 2, the network weights are calculated by averaging over time, which may lead to even more robust results. Finally, in Case 3, the prediction results are determined by multiple augmentations and averaging, which can reduce the error that occurs during a single forward computation. However, this incurs the computational cost of multiple forward computations during grasp prediction.

[0110] In addition, there are two main ways to implement the proposed test-time adaptation method in the application of grasp prediction:

[0111] 1. Adapt using a fixed set of input images before applying to additional input images during the adaptation phase.

[0112] 2. Online test-time adaptation as a continuous process during grasp prediction (i.e., typically using a machine learning model for inference).

[0113] In Case 1, the test-time adaptation method generates adapted network weights and BN statistics for a fixed set of input images during the adaptation phase. After the adaptation phase, the network weights and BN statistics are set and used for new images.

[0114] In Case 2, the test-time adaptation method is integrated into the pipeline for grasp prediction, and the network weights and BN statistics are continuously updated for each new input image. In this case, the adaptation can account for new domain shifts over time. However, effects such as error accumulation and catastrophic forgetting should be considered here.

[0115] In summary, according to different embodiments, a method as Figure 3 shown is provided.

[0116] Figure 3 FIG. 300 is a flowchart that represents a method for (self-monitoringly) adapting a machine learning model to an altered control situation (relative to the control situation for which it has been trained).

[0117] In 301, sensor data elements are detected in the altered control situation.

[0118] In 302, for each determined sensor data element,

[0119] · In 303, multiple augmentations (i.e., altered (transformed) versions) of the sensor data element are generated, such as by changing size, mirroring, adding noise, shifting, rotating, changing color, etc.

[0120] · In 304, a first instance of a machine learning model generates a corresponding output for each augmentation.

[0121] · In 305, a target output for a sensor data element is determined by combining (e.g., averaging) the generated outputs.

[0122] · In 306, a loss between an output of a second instance of the sensor data element (which is generated by processing the sensor data element or its augmentation using the second instance) and the determined target output is determined (the determination of the loss may include back-augmentation (Rückaugmentierung) to enable comparison between the outputs).

[0123] In 307, a second instance of the machine learning model is adapted to reduce an overall loss that includes the determined loss (e.g., by backpropagating the loss and adapting parameters (e.g., weights or batch normalization statistics) of the machine learning model in a direction of reducing the overall loss).

[0124] Figure 3 The method enables adaptation of a pre-trained ML model (e.g., a grasping prediction model) such that it takes into account changes (or "shifts") in the input image domain (e.g., due to a new camera type) without re-training and with access to labeled training data. The method can be used, for example, in combination with grasping by a robot when picking an object out of a container in order to achieve robust prediction performance even when using a new camera or a new camera mounting position, or when object properties (e.g., surface reflection) change.

[0125] Figure 3 The result of the method is (for an example of a neural network as a machine learning model) adapted network weights (and thus its statistical properties), which the grasping prediction network can directly use to predict a grasp based on an image with domain shift.

[0126] Figure 3The method can be executed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity capable of processing data or signals. For example, data or signals can be processed according to at least one (i.e., one or more than one) specific function executed by the data processing unit. The data processing unit can include analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), integrated circuits of programmable gate arrays (FPGAs), or any combination thereof, or be formed by them. Any other means for implementing the corresponding functions described in more detail herein can also be understood as a data processing unit or a logic circuit device. One or more of the method steps described in detail herein can be executed (e.g., implemented) by the data processing unit through one or more specific functions executed by the data processing unit.

[0127] The method is thus implemented, in particular, by a computer according to different embodiments.

[0128] After training (i.e., its test-time adaptation), the machine learning model can be applied to sensor data determined by at least one sensor. For example, after training, the machine learning model is used to generate control signals for a robotic device by feeding sensor data about the robotic device and / or its surrounding environment to the machine learning model. The term "robotic device" can be understood to refer to any technical system (having mechanical components with controlled self-movement), such as a computer-controlled machine, vehicle, household appliance, power tool, manufacturing machine, personal assistant, or access control system.

[0129] In addition to images having grayscale, color channels, or depth channels, various embodiments can also receive and use sensor data from various other sensors, such as video, radar, lidar, ultrasound, motion, thermal imaging, etc.

Claims

1. A method for adapting a machine learning model (112) to a changed control situation, the method comprising: detecting (301) a sensor data element (201) under a changed control condition; For each determined sensor data element generating (303) a plurality of enhancements of the sensor data element; generating (304) a corresponding output for each enhancement using the first instance (203) of the machine learning model; determining (305) a target output for the sensor data element by combining (204) the generated outputs; determining (306) a loss between an output for a second instance (202) of the sensor data element and the determined target output; and Adapting (307) the second instance (202) of the machine learning model to reduce an overall loss (205), wherein the overall loss includes the determined loss.

2. The method according to claim 1, comprising: Adapt the first instance (203) of the machine learning model in the direction of an adapted second instance (202) of the machine learning model.

3. The method according to claim 1, comprising: For each batch in the batch sequence, detecting corresponding sensor data elements of the changed control condition; For each sensor data element determined for the batch generating a plurality of enhancements of the sensor data element; Generating a corresponding output for each enhancement by feeding the generated enhancement to a corresponding first instance (203) of the machine learning model; determining a target output for the sensor data element by combining (204) the generated outputs; determining a loss between an output for a corresponding second instance (202) of the sensor data element and the determined target output; and adapting a corresponding second instance of the machine learning model (202) to reduce an overall loss (205), wherein the overall loss includes the determined loss, Wherein, for each batch in the sequence except the last batch, the corresponding adapted second instance (202) of the machine learning model is used as the second instance (202) of the machine learning model for subsequent batches in the sequence.

4. The method according to claim 3 includes adapting the first instance (203) of the machine learning model in the direction of the second instance (202) of the machine learning model after a predetermined number of batches.

5. The method according to any one of claims 1 to 4, wherein the sensor data elements are image data elements.

6. The method according to claim 5, wherein the change in the control situation to be adapted is: a change in the camera (114) and / or a change in one or more conditions of image recording using the camera (114) for recording the image data elements.

7. The method of claim 5 or 6, wherein the image data element has a plurality of channels, and generating a corresponding output for each enhancement and generating an output of a second instance (202) for each sensor data element comprises: A trainable scaling of the respective values ​​of the channels, wherein the scalings are adapted together to reduce the overall loss.

8. A method for controlling a robotic device (101), the method comprising: Adapting a machine learning model (112) to a control situation in which the robotic device (101) is to be controlled by means of a method according to any one of claims 1 to 7; detecting one or more other sensor data elements in said control situation; processing the one or more further sensor data elements by means of the adapted second instance (202) of the machine learning model or by means of the first instance (203) of the machine learning model adapted towards the adapted second instance (202); and A control signal for the robot device (101) is generated according to the result of the processing.

9. A data processing device (106) configured to execute the method according to any one of claims 1 to 8.

10. A computer program having instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.

11. A computer readable medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.