Method for adapting machine learning model to changed control situation
The method allows machine learning models for robotic grasping to adapt to changed control situations by detecting sensor data elements, generating extensions, and adapting the model outputs, effectively addressing the challenge of adapting to dynamic conditions without extensive retraining or annotated data.
Patent Information
- Application Number
- JP2024207085
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-10
AI Technical Summary
Existing machine learning models for robotic grasping face challenges in adapting to changed control situations, such as changes in camera settings or environmental conditions, without the need for extensive retraining or annotated data.
A method that involves detecting sensor data elements in a changed control situation, generating extensions of these data elements, using a first instance of a machine learning model to produce outputs, combining these outputs to identify a target output, calculating a loss between the output of a second instance of the model and the target output, and adapting the second instance to reduce this loss.
Enables self-supervised test-time adaptation of machine learning models, allowing them to adapt to changes in camera settings or environmental conditions without the need for new annotated training data, thereby improving grasping performance in dynamic situations.
Smart Images

Figure 2025087644000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for adapting a machine learning model to changed control situations.
Background Art
[0002] Picking up an object (i.e., grasping it) is an important issue in robotics. Recent approaches utilize machine learning to enable model-free grasping of multiple objects that have not been seen before. In actual applications, for example, when taking an object out of a container, the performance of these approaches, i.e., the performance of a correspondingly trained machine learning model, typically depends on the conditions in each control situation, and each change in the camera, object, or environment (relative to the situation when the machine learning model was trained) can have an adverse effect on the grasping ability. In order to achieve reliable control (i.e., high grasping ability) even when the conditions change, the machine learning model can be retrained by supervised learning using corresponding training data, whereby the machine learning model can be adapted to each situation. However, for this purpose, additional training data (e.g., images) having corresponding annotations (i.e., "labels") must be generated, which requires a great deal of cost.
Summary of the Invention
Problems to be Solved by the Invention
[0003] Therefore, an approach that enables adapting a machine learning model to changed control situations at low cost is desired.
Means for Solving the Problems
[0004] According to various embodiments, a method for adapting a machine learning model to a changed control situation (relative to the control situation when the machine learning model was trained), ·Detecting sensor data elements in a changed control situation; ·For each identified sensor data element, ○Generating a plurality of extensions of the sensor data element; ○For each extension, generating an output using a first instance of a machine learning model; ○Identifying a target output for the sensor data element by combining the generated outputs; ○Identifying a loss between the output of a second instance related to the sensor data element and the identified target output; ·Adapting a second instance of the machine learning model to reduce the total loss including the identified loss; A method is provided that includes the above.
[0005] According to the above method, for example, when the camera supplying the input image for the grasping prediction network changes (for example, when a new camera type for taking the input image is used, or when the mounting situation changes), self-supervised test-time adaptation becomes possible without the need for supervised training of the machine learning model using an annotated training data set, that is, during inference, it is possible to adapt the machine learning model to the changed conditions for training.
[0006] Various embodiments are described below.
[0007] Example 1 is a method for adapting a machine learning model to a changed control situation as described above.
[0008] Example 2 is the method described in Example 1, including adapting a first instance of the machine learning model in the direction of the adapted second instance of the machine learning model.
[0009] That is, the first instance (also referred to as the teacher model or specifically the teacher network in the following examples) can follow the second instance (also referred to as the student model or specifically the student network in the following examples) at a predetermined interval of batches, as described in the following examples.
[0010] Example 3 involves, for each batch of a sequence of batches, · detecting each sensor data element in a varying control situation, and · for each sensor data element identified for a batch, ○ generating a plurality of expansions of the sensor data element, and ○ generating respective outputs by supplying the plurality of generated expansions to respective first instances (for the batch) of a machine learning model, and ○ identifying a target output for the sensor data element by combining the generated outputs, and ○ identifying the loss between the output of each second instance (for the batch) related to the sensor data element and the identified target output, and · adapting each second instance of the machine learning model to reduce the total loss including the identified loss, and including The method according to Example 1, wherein for each batch of the sequence excluding the last batch, each adapted second instance of the machine learning model is used as the second instance of the machine learning model for the subsequent batch in the sequence.
[0011] That is, the adaptation described in the above method can be related to one batch (i.e., the detected sensor data elements are sensor data elements of one batch), and the adaptation described in the above method can be appropriately repeated for further batches. In this case, a second instance is adaptively adjusted so that the accuracy of the machine learning model improves over time (e.g., in an ongoing operation). As described above, the first instance can follow the second instance.
[0012] Example 4 is a method according to Example 3, including adapting a first instance of a machine learning model in the direction of a second instance of the machine learning model after a predetermined number of batches.
[0013] The number of batches may be one, that is, the first instance can directly follow the second instance, but in a weighted form (see the following example), for example. Thus, the first instance still slowly follows the second instance. This ensures stability during adaptation. For the first batch of the sequence, the first instance and the second instance can be set to the machine learning model to be adapted.
[0014] Example 5 is a method according to any one of Examples 1 to 4, where the sensor data element is an image data element.
[0015] In this specification, an image data element is understood as a data element in the form of a matrix having one or more channels (i.e., one or more values per position in the matrix, i.e., per "pixel"). Such sensor data elements can effectively represent the scene to be controlled. The machine learning model is, for example, a neural convolutional network (or includes a neural convolutional network).
[0016] Example 6 is the method according to Example 5, wherein the change in the control situation to be adapted is a change in the camera and / or a change in one or more conditions of the image capture using the camera that captures the image data elements.
[0017] Thereby, using the above method, with a small training cost and without explicit annotation of the sensor data elements (instead, using the generated pseudo-labels, i.e., the target output), it is possible to adapt the machine learning model to changes in the camera or image capture conditions (such as changes in lighting, color shift, etc.).
[0018] Example 7 is the method according to Example 5 or 6, wherein the image data elements have a plurality of channels, generating respective outputs for each expansion, and generating an output of a second instance for each sensor data element, including trainably scaling the respective values of the channels, and the scaling is adapted together to reduce the total loss.
[0019] Example 8 is a method for controlling a robot device, · adapting a machine learning model to the control situation in which the robot device is to be controlled by the method according to any one of Examples 1 to 7; · detecting one or more additional sensor data elements in the control situation; · processing one or more additional sensor data elements using the adapted second instance of the machine learning model or using the first instance of the machine learning model adapted in the direction of the adapted second instance; · generating a control signal for the robot device according to the result of the processing. The method includes the above steps.
[0020] Example 9 is a data processing device (in particular, a control device for a robot device) configured to implement the method according to any one of Examples 1 to 8.
[0021] Example 10 is a computer program comprising instructions for causing a processor to execute the method according to any one of Examples 1 to 8 when executed by the processor.
[0022] Example 11 is a computer-readable medium storing instructions for causing a processor to execute the method according to any one of Examples 1 to 8 when executed by the processor.
[0023] In the drawings, like reference numerals generally refer to the same parts throughout the several different drawings. The drawings are not necessarily to scale, and instead emphasis is generally placed on illustrating the principles of the invention. In the following description, various aspects will be described with reference to the following drawings.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Modes for Carrying Out the Invention
[0025] The following detailed description refers to the accompanying drawings, which show specific details and aspects of the present disclosure in which the invention can be practiced. Other aspects can be used and structural, logical, or electrical changes can be made without departing from the scope of the invention. Since some aspects of the present disclosure can be combined with one or more other aspects of the present disclosure to form new aspects, the various aspects of the present disclosure are not necessarily mutually exclusive.
[0026] Hereinafter, various examples will be described in more detail.
[0027] Figure 1 shows robot 100.
[0028] Robot 100 includes a robot arm 101, for example, an industrial robot arm for processing or assembling a workpiece (or one or more other objects). The robot arm 101 includes movable arm elements 102, 103, 104 and a base (or support) 105 that supports these arm elements 102, 103, 104. The term "movable arm element" refers to a movable member of the robot arm 101, and by operating this movable member, physical interaction with the environment becomes possible, for example, to perform a certain task. For control, the robot 100 includes a (robot) control device 106, and this control device 106 is configured to implement interaction with the environment according to a control program. The last arm element 104 (the one farthest from the support 105) among the arm elements 102, 103, 104 is also referred to as an end effector 104 and may include one or more tools such as a welding torch, a gripping device, painting equipment, etc.
[0029] The other arm elements 102, 103 (closer to the support 105) can constitute a positioning device, whereby, together with the end effector 104, a robot arm 101 having this end effector 104 at its end is provided. The robot arm 101 is a mechanical arm (which may have a tool at its end in some cases).
[0030] The robotic arm 101 may include joint elements 107, 108, 109, and these joint elements 107, 108, 109 connect the arm elements 102, 103, 104 to each other and connect the arm elements 102, 103, 104 to the support 105. The joint elements 107, 108, 109 may have one or more joint mechanisms, and each of these joint mechanisms can bring about a rotatable movement (i.e., rotational movement) and / or a translational movement (i.e., displacement) relative to each other for the associated arm element. The movement of the arm elements 102, 103, 104 can be initiated using actuators controlled by the control device 106.
[0031] The term "actuator" can be understood as a component configured to act on a mechanism or process in response to being driven. The actuator can execute commands (so-called activations) issued by the control device 106 to generate mechanical movement. The actuator, for example, an electromechanical transducer, can be configured to convert electrical energy into mechanical energy in response to an activation.
[0032] The term "control device" can be understood as any type of entity (including one or more computers) that implements logic, and this entity can include, for example, a circuit and / or a processor capable of executing software, firmware, or a combination thereof stored in a storage medium and, for example, in this example, capable of issuing commands to the actuator. For example, the control device can be configured to control the operation of the system, in this example, the operation of the robotic device, by program code (for example, software).
[0033] According to this example, the control device 106 includes one or more processors 110 and a memory 111 that stores code and data, and the processor 110 controls the robotic arm 101 based on this code and data. According to various embodiments, the control device 106 controls the robotic arm 101 based on a machine learning model 112 stored in the memory 111.
[0034] According to various embodiments, the machine learning model 112 is configured and trained to enable the robot 100 to recognize an operating posture in which the robot 100 can pick up one or more objects 113 (or interact with the objects 113 by other means, such as painting).
[0035] For example, one or more cameras 114 that enable the robot 100 to take images of its own work space can be provided on the robot 100. The camera 114 is, for example, assembled to the robotic arm 101, and thus the robot can take images of the object 113 from various viewpoints by rotating its own robotic arm 101. However, as shown in FIG. 1, the camera 114 can also be fixedly attached to the robot cell to detect the object to be grasped.
[0036] According to various embodiments, the machine learning model 112 is a neural network, and the control device 106 supplies input data to the neural network based on one or more digital images of the object 113 (a depth image with an optional color image, or a point cloud with an optional color image, or additional per-pixel information such as information regarding surface normals), and the neural network (in this example, specifically a neural "grasp prediction network") identifies, for example, for each of a plurality of locations (on the surface) of the object, a quality indicating how well the object can be grasped at each location. Instead of continuous values (i.e., instead of regression), the machine learning model 112 (e.g., a neural network) can also perform a classification, for example, into "good locations for grasping" and "bad locations for grasping". It can also output yet another continuous value, for example, for each location, it can output an orientation with respect to the end effector 104 (assumed to be a gripper, by way of example hereinafter), and then output a quality (of manipulation, or of grasping hereinafter) indicating how well the object can be manipulated in that orientation at that location.
[0037] There may be cases where a machine learning model 112 trained for a specific (control) situation (i.e., under specific conditions) is then desired to be used in other situations (also referred to as "test time", which includes inferences during use). For example, the camera 114 may be replaced, which will change the way the object 113 is represented in the input data of the machine learning model 112 (e.g., lens characteristics, noise behavior, etc. change). To compensate for this, so-called test-time adaptation can be implemented. Further examples in this regard are the compensation for changes in object characteristics or environmental conditions (e.g., light conditions), changes in the structure and number of each object, or changes in the camera positioning relative to each object. Generally, test-time adaptation can always be applied when there is a domain shift (or domain gap) between the training data and the test-time data.
[0038] According to various embodiments, test-time adaptation is provided for the purpose of adapting a pre-trained ML (machine learning) model for inference without the need for training data annotated for that purpose (using ground truth, i.e., typically using labels). For example, this test-time adaptation is implemented to adapt a neural grasping prediction network to changes in the camera (e.g., due to replacement of the camera 114). That is, according to various embodiments, for example, to adapt a neural network for per-pixel grasping prediction to changes in the camera (or "domain shift") between training time and test time without (new) supervised training and thus without additional annotation cost, test-time adaptation is utilized.
[0039] According to various embodiments, in particular, the mean teacher concept is used for self-supervised test-time adaptation of a machine learning model (in this case, input image channel scaling can be incorporated). The "mean teacher" (i.e., the machine teacher model) provides pseudo-labels (or soft pseudo-labels), which are used, for example, to adapt the network weights and batch normalization statistics of a convolutional neural network (CNN) used as a grasping prediction network. This approach, which is used to update the network weights and batch normalization statistics in response to new input images, for example, in response to a new camera type, can be used as real-time adaptation or as used during an initial adaptation phase. This approach is not limited to a specific CNN network architecture.
[0040] That is, according to various embodiments, a mean teacher framework with test-time augmentation and image channel scaling is used to enable robust network predictions for new input images with unknown domain shifts, for example, of a new unknown camera type, even during the runtime of the model. The results are as follows. (1) A grasping prediction network trained with image data from one or more known cameras can be adapted, by this approach of self-supervised post-training, to images with domain shifts (e.g., images of an unknown camera type, or images of an unknown camera mounting position, or images of unknown characteristics of a grasping object such as surface reflection, color shift, etc.) without the need for new training with labeled image data. (2) This approach can be used in an offline setting or an online setting to adapt to a new domain shift using a fixed set of input images during an initial adaptation phase or to adapt continuously during the runtime of the application. (3) By incorporating test time augmentation into the average teacher prediction (i.e., pseudo-labels), even after training with the augmented training samples, it is possible to generate robust pseudo-labels by exploiting the phenomenon that the CNN is not completely invariant to the symmetry of the data distribution of new input data. (4) By incorporating learnable channel scaling into the test time adaptation method, it is possible to automatically adapt the weighting coefficients for each image input channel (e.g., RGB, depth, etc.). This results in more robust predictions when the domain shift between input channels is unbalanced, for example, when a new RGB-D camera type supplies similar quality RGB images but less accurate depth images.
[0041] Figure 2 shows a test time adaptation method for a machine learning model according to an embodiment.
[0042] The test time adaptation method can be applied to, for example, neural networks, such as various CNN network architectures. The machine learning model maps, for example, an input image (e.g., RGB or RGB-D (depth information in addition to RGB)) to an output having the same resolution as the input image. As described above, the output may be a classification (e.g., an identifier of a good or bad grasping position in the input image, or a classification of an object in autonomous driving), or may be a continuous value (e.g., the probability of a stable grasp). In the test time adaptation method, two versions of the machine learning model, namely, the "student network" 202 and the "teacher network" 203 are used. Both are initialized based on the machine learning model to be adapted (i.e., initially both are identical to the machine learning model to be adapted). The weights of the machine learning model before adaptation (and thus the initial weights of the student network 202 and the teacher network 203) are
Number
[0043] The input of the test time adaptation method is, for example, B input images captured by a camera different from the camera used to capture images for training a machine learning model 201 (assumed to be a neural network hereinafter).
Number
[0044] In the teacher network 203, input images that are differently expanded for each of the input images N are generated by using N different expansion transformations from these input images.
Number
Number
[0045] Next, for each of the expanded input images (i.e., each expansion), different image channels are scaled by a function sc (representing "scale") using a learnable vector γ so that the scaling of each input channel (i.e., the image channel in this example) can be adapted to the test time by the test time adaptation method.
Number
Number
[0046] This allows, for example, considering different domain shifts between input channels when a new camera has different domain shifts between the RGB channel and the depth channel compared to the camera that captured the images for training the network.
[0047] The extended individual images are led through the teacher network 203 after these individual images are scaled, thereby resulting in N different teacher outputs (in this example, output images)
Number
Number
[0048] Subsequently, the different outputs are integrated into just one output by the average 204, where avg represents various possible averaging techniques for calculating the average at each pixel coordinate (n, m), such as the arithmetic mean
Number
Number
Number
[0049] However, weighted averaging can also be performed (e.g., the augmentations are weighted differently depending on the augmentation transformation used (e.g., type and / or strength of the transformation)).
[0050] To enable integration of multiple output images, an inverse augmentation transformation can be applied to each output image (corresponding to the augmentation transformation that generated the augmented input image processed into the output image by the teacher network 203).
[0051] (Augmentation average) used as a pseudo-label for training the machine learning model 202 [Number] The underlying idea can be seen as achieving a high-quality pseudo-label by integrating multiple output images calculated using augmented versions of the same input image. Particularly in a convolutional neural network, this "ensembling" takes advantage of the phenomenon that the convolutional neural network (CNN) is not completely invariant to augmentations even after training with augmented training samples.
[0052] In the case of the student network 202, the augmentation of the input image [Number] is optional and can be taken from one of the augmentations of the teacher network 203.
[0053] As described for the teacher network 203, it is also possible to use multiple augmentations of the student network 202 and average the output images of the student network 202, i.e., [Number] where.
[0054] Next, in order to be able to adapt the scaling of each input channel to the test time by means of the test time adaptation method, different image channels of the student network 202 are scaled by a learnable vector γ S by
Number
Number
[0055] Next, the input image that has been extended and scaled in this way is passed through the student network 202 with network weights θ S which network weights θ S as described above, are initialized by the pre-trained source model weights
Number
Number
[0056] Finally, to calculate the (total) loss 205 for the batch the output of the student path
Number
Number
Number
Number
[0057] The per-pixel loss l may be represented by various loss functions, for example, l1 loss, l2 loss, or cross-entropy loss (depending on the type of pixel values).
[0058] The student network 202 is updated by the feedback of the loss 205. This loss enforces the consistency between the output of the student network and the pseudo label.
[0059] According to one embodiment, the teacher network 203 is updated by exponential moving average EMA. That is, in each training step of the teacher network (index t, for example, after a predetermined number of batches, the training step of the teacher network is performed), the weights of the student model 202 are used to update the teacher network 203, thereby resulting in a continuously trained and time-averaged teacher network 203. Since the prediction of the (averaged) teacher network 203 can be assumed to be more accurate than the output of the student network 202, the prediction of the (averaged) teacher network 203 can be used as the pseudo label for the self-supervised training of the student network 202 (as described above). The weights of the teacher model 203 are, for example,
Number
[0060] In addition to this, the batch normalization layers of the student network and the teacher network are in training mode during test-time adaptation and are newly estimated respectively according to the test-time data.
[0061] In the following, various approaches for using the test-time adaptation method described with reference to FIG. 2 for adapting the grasping prediction network (e.g., for a new camera) will be described.
[0062] To use the results of the test-time adaptation method for grasping prediction in a new control situation, there are three main variations, namely, 1. Use of the grasping prediction network with the weights of the student network and the batch normalization statistics, 2. Use of the grasping prediction network with the weights of the teacher network and the batch normalization statistics, 3. Use of the grasping prediction network with the weights of the teacher network and the batch normalization statistics and the exponential average (as described for the teacher network) during inference exist. The batch normalization statistics are parameters related to the batch normalization (BN) layer of each neural network. The batch normalization statistics are, in a narrow sense, not trained weights but are calculated during training for the training data. These BN statistics are specific to the training data. Therefore, when a domain shift occurs for the test-time data, the BN statistics can also be adapted (like the weights as described above).
[0063] In Case 1, the student network 202 should yield more robust predictions than the original machine learning model after self-supervised training. In Case 2, the network weights are calculated by time averaging, which can result in more robust results. Finally, in Case 3, the prediction results are determined by multiple evaluations and mean value formation, which can reduce the error that occurs in the case of a single forward calculation. However, this is associated with the computational cost for multiple pre-calculations during the grasping prediction.
[0064] Furthermore, to implement the proposed test-time adaptation method in the application for grasping prediction, there are two main approaches, namely, 1. Adaptation using a fixed set of input images before application to additional input images in the adaptation phase 2. Online test-time adaptation as a continuous process during grasping prediction (i.e., generally, inference using a machine learning model) exist.
[0065] In Case 1, the test-time adaptation method generates network weights and BN statistics adapted to a fixed amount of input images during the adaptation phase. After the adaptation phase, the network weights and BN statistics are defined and used for new images.
[0066] In Case 2, the test-time adaptation method is incorporated into the pipeline for grasping prediction, and the network weights and BN statistics are continuously updated for each new input image. In this case, the adaptation can take into account new domain shifts over time. However, in this case, effects such as error accumulation and catastrophic forgetting should be considered.
[0067] In summary, according to various embodiments, the method is provided as shown in FIG. 3.
[0068] Figure 3 shows a flowchart 300 of a method for adapting a machine learning model (with self - teaching) to a changed control situation (with respect to the control situation when the machine learning model was trained).
[0069] At 301, sensor data elements in the changed control situation are detected.
[0070] At 302, the following is done for each identified sensor data element, namely · At 303, multiple augmentations of the sensor data element (i.e., modified (transformed) versions, e.g., by resizing, mirroring, adding noise, shifting, rotating, color - changing, etc.) are generated. · At 304, for each augmentation, an output is generated using a first instance of the machine learning model. · At 305, a target output for the sensor data element is identified by combining (e.g., averaging) the generated outputs. · At 306, a loss between the output of a second instance (generated by processing the sensor data element or its augmentation using the second instance) related to the sensor data element and the identified target output is identified (the identification of the loss may include inverse augmentation to enable comparison between the outputs).
[0071] At 307, a second instance of the machine learning model is adapted (e.g., by backpropagating the loss to adapt the parameters of the machine learning model (e.g., weights or batch normalization statistics too) in a direction where the total loss decreases) to reduce the total loss that includes the identified loss.
[0072] According to the approach of FIG. 3, it is possible to adapt a pre-trained ML model (e.g., a grasping prediction model) to account for changes (or "shifts") in the input image domain (e.g., due to a new camera type) without the need for access to new training and marked training data. This method can be used in the context of grasping with a robot, for example, when removing an object from a container, enabling robust prediction performance even when a new camera or new camera mounting position is used, or when object characteristics such as surface reflectance change.
[0073] The result of the method of FIG. 3 is the adapted network weights (and thus their statistical characteristics) (for the example of a neural network as a machine learning model), and these network weights can be directly used by a grasping prediction network to predict grasping for images with domain shift.
[0074] The method of FIG. 3 can be implemented by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any kind of entity that enables the processing of data or signals. The data or signals can be processed according to at least one (i.e., one or more) specific functions performed by the data processing unit, for example. The data processing unit can include analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), integrated circuits of programmable gate arrays (FPGAs), or any combination thereof, or can be composed of these. Any other method for implementing each function described in more detail herein can also be understood as a data processing unit or logic circuit device. One or more of the method steps described in detail herein can be implemented (e.g., implemented) by the data processing unit via one or more specific functions performed by the data processing unit.
[0075] That is, according to various embodiments, the method is particularly computer-implemented.
[0076] The machine learning model may be applied to sensor data identified by at least one sensor after training (i.e., test-time adaptation of the machine learning model). For example, by supplying sensor data regarding a robotic device and / or the environment of the robotic device to the machine learning model, the machine learning model can be used for generating control signals for the robotic device after training. The term "robotic device" can be understood as relating to any technical system having machine parts whose operation is controlled, such as a computer-controlled machine, vehicle, household appliance, power tool, manufacturing machine, personal assistant, or access control system.
[0077] In addition to images having grayscale, color channels, or depth channels, various embodiments can also receive and use sensor data from various other sensors such as, for example, video, radar, LiDAR, ultrasonic, motion, thermal imaging, etc.
Claims
1. 1. A method for adapting a machine learning model (112) to a changed control situation, comprising: Detecting (301) a sensor data element (201) in the changed control condition; For each identified sensor data element: generating (303) a plurality of extensions of the sensor data elements; For each extension, generating (304) a respective output using the first instance (203) of the machine learning model; identifying (305) target outputs for the sensor data elements by combining (204) the generated outputs; determining (306) a loss between an output of the second instance (202) of the sensor data element and the determined target output; adapting (307) the second instance (202) of the machine learning model to reduce a total loss (205) that includes the identified losses; and The method includes:
2. adapting the first instance of the machine learning model to the orientation of the adapted second instance of the machine learning model. The method of claim 1.
3. For each batch in the sequence of batches, detecting respective sensor data elements in said changed control condition; For each sensor data element identified for the batch: generating a plurality of extensions of the sensor data elements; providing the generated augmentations to a first instance (203) of each of the machine learning models to generate respective outputs; combining (204) the generated outputs to identify target outputs for the sensor data elements; determining a loss between an output of each second instance (202) of the sensor data element and the determined target output; adapting the second instances (202) of each of the machine learning models to reduce a total loss (205) that includes the identified losses; and Including, for each batch in the sequence except the last batch, the respective adapted second instance (202) of the machine learning model is used as the second instance (202) of the machine learning model for a subsequent batch in the sequence. The method of claim 1.
4. adapting the first instance of the machine learning model to a direction of the second instance of the machine learning model after a predetermined number of batches. The method according to claim 3.
5. the sensor data elements are image data elements; 5. The method according to any one of claims 1 to 4.
6. the change in the control state to be adapted is a change in a camera (114) and / or a change in one or more conditions of image capture using a camera (114) capturing the image data elements; The method according to claim 5.
7. the image data elements have a plurality of channels; generating respective outputs for each augmentation and generating outputs of the second instances (202) for each sensor data element includes trainably scaling the values of each of the channels; The scaling is adapted together to reduce the total loss. The method according to claim 5 or 6.
8. A method for controlling a robotic device (101), comprising: Adapting a machine learning model (112) to a control situation in which the robotic device (101) is to be controlled according to the method of any one of claims 1 to 7, Detecting one or more further sensor data elements in the control situation; and processing the one or more further sensor data elements with an adapted second instance (202) of the machine learning model or with a first instance (203) of the machine learning model adapted in the direction of the adapted second instance (202); generating a control signal for said robotic device (101) according to a result of said processing; The method includes:
9. A data processing apparatus (106) configured to perform the method according to any one of claims 1 to 8.
10. A computer program comprising instructions which, when executed by a processor, cause the processor to carry out a method according to any one of claims 1 to 8.
11. A computer readable medium storing instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 8.