Method and apparatus for subsequent training of machine learning system

By setting a pre-given spatial image structure in the machine learning system and adapting the parameters of the diffusion model in subsequent training, the problem of generating sufficient training data is solved by using distribution-based evolutionary algorithms and loss functions, and the similarity between training efficiency and image generation is improved.

CN119992246APending Publication Date: 2025-05-13ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411581863.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-11-07
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the training of machine learning systems, especially in safety-critical systems, it is difficult for prior art to effectively generate a sufficient amount of training data, especially to map rare but safety-critical situations.

Method used

By setting up a machine learning system to generate digital images under a pre-given spatial image structure and adapting the parameters of the diffusion model in subsequent training, the parameters are optimized using distribution-based evolutionary algorithms and loss functions to improve the similarity of image generation.

Benefits of technology

This method reduces the number of parameters that need to be adapted, improves training efficiency and speed, avoids the problems of gradient vanishing and insufficient memory, and supports non-continuous differentiable loss functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992246A_ABST
    Figure CN119992246A_ABST
Patent Text Reader

Abstract

A computer-implemented method for subsequent training of a machine learning system configured to generate a digital image with a predefined spatial image structure, the method comprising the steps of: receiving training data by the machine learning system, each training data comprises a specification of a spatial image structure, additional parameters are added to a machine learning system, an image is generated by the machine learning system for each training data, a spatial image structure of the image generated by the machine learning system is determined by a second machine learning system, and the spatial image structure of the image generated by the machine learning system is determined by the second machine learning system. The further parameters are adapted using a loss function that measures the similarity between the determined spatial image structure of the generated image and a spatial image structure predefined by the associated training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for post-training a machine learning system, wherein the machine learning system is configured to generate digital images given a predetermined spatial image structure. The invention also relates to a device configured to carry out the method, a computer program for implementing the method and a machine-readable data carrier having such a computer program. Background Art

[0002] In the training of machine learning systems in the field of computer vision, a training data set comprising real images can be expanded or supplemented by synthetically generated images. Alternatively, the training data set can also only include synthetically generated images. In particular, machine learning systems that assume safety-critical tasks in reasoning - for example, when corresponding machine learning systems are used in at least partially autonomous driving, automatic optical inspection, surveillance or vehicle interior space monitoring - require a sufficient amount of training data during training, which in particular must also map rare but possibly particularly safety-critical situations / situations. Therefore, in the case of insufficient availability of corresponding real image data, it may be necessary to generate the corresponding images synthetically with the help of a generative machine learning system.

[0003] Generative machine learning systems for generating images are given, for example, by diffusion models (DMs), which add random noise to the input image data at each step in a Markov chain of diffusion steps, and then learn the inverse of the diffusion process during training so that during inference, from the input of the noisy image and a textual prompt, the "desired" image data can be obtained based on the textual prompt.

[0004] In arXiv:2302.05543v1[cs.CV], a neural network architecture ControlNet is proposed, which is designed to check the diffusion model by adding additional conditions, that is, in other words, the neural network structure is designed to be able to control the structure of the image generated by the diffusion model more specifically in terms of the image structure and / or the arrangement and / or position of the objects shown in the image. As a result, specific characteristics of the image to be generated by the diffusion model can be controlled or influenced in a targeted manner through ControlNet. As described in arXiv:2302.05543v1[cs.CV], ControlNet is based on a stable diffusion model (Stable Diffusion Modell), see arXiv:2112.10752[cs.CV].

[0005] In https: / / doi.org / 10.1016 / j.neunet.2009.12.004, a method for adapting the parameters of a machine learning system based on an evolutionary algorithm is proposed. Unlike the back-propagation method for determining the minimum (or maximum) of the loss function under consideration, this method first involves initializing a current set of parameter values ​​of the machine learning system and extracting N additional sets of parameter values ​​from a distribution centered at the current initialized set of parameter values. For each of the N sets of extracted parameter values, the value of the loss function selected for the target task is calculated and the gradient of the loss function is estimated accordingly. This estimated gradient is used to update the set of parameter values. These steps are iterated accordingly until the minimum or maximum of the loss function under consideration is approached with the desired accuracy. Summary of the invention

[0006] In a first aspect, the invention relates to a computer-implemented method for subsequent training of a machine learning system. The machine learning system is configured to generate digital images under a predetermined spatial image structure. The images generated by the machine learning system can, for example, show an environment, in particular a driving environment of an at least partially autonomous robot, a workpiece or a part of a workpiece in the automated optical inspection of workpieces, or at least one person in an area to be monitored by a surveillance camera (for example inside or outside a building). The spatial image structure can be determined or predetermined, for example, by setting the arrangement, position, size, relative size of multiple objects to each other, description of foreground and background, position and / or orientation of individual objects in the image. The spatial image structure can be predetermined, for example, by assigning image regions (i.e., for example, continuous sets of pixels) to content to be represented in the generated image in an image mask, which has the same parameters as the image to be generated by the machine learning system in terms of the number and arrangement of image pixels. The image mask can, for example, be a semantic segmentation mask, by which different image regions can be separated from each other in terms of their semantic content, and the semantic content of these image regions can be predetermined in the form of labels. The separation of image regions in terms of their semantic content can be achieved, for example, in the following manner, that is, the pixel values ​​in the image regions of the same semantic content can be made to take the same value respectively, while the image regions of different contents are different from each other in the image mask in terms of the values ​​assigned to the pixels. The image mask can alternatively be an image mask with key point postures of the human body. In the prototype of the image to be generated, the prototype has the same size as the image to be generated itself, and the key point posture can represent the position and posture of human limbs (such as limbs, head, eyes, chin, neck, hands and / or torso) in the form of "stickman (part)" through interconnected image points. Preferably, the machine learning system is a generative probabilistic text-to-image diffusion model. Here, the blocks of the above-mentioned diffusion model and their parameters are copied as locked copies and trainable copies. Here, the method includes at least the steps described below. In a method step, the machine learning system receives training data. Each training data includes a pre-given spatial image structure here. Preferably, the spatial image structure is given by an image mask, which is a semantic segmentation mask or an image mask with key point postures of the human body. In a further method step, additional parameters are added to the trainable copy of at least one replicated block of the diffusion model. This can be achieved by inserting at least one additional layer parameterized with the additional parameters in the trainable copy of the replicated block. Alternatively or additionally, the additional parameters can be added by decomposing at least one weight matrix in the trainable copy of the replicated block to be adapted in subsequent training into the sum of a pretrained weight matrix and additional addends added in subsequent training. In this case, the additional addends are given by the matrix product of two additional matrices.The two further matrices are parameterized here with further parameters, and the parameters of the pre-trained weight matrix are adapted in the pre-training of the diffusion model and remain unchanged in the subsequent training. The ranks of the two further matrices are lower than the rank of the pre-trained weight matrix. In a further method step, the machine learning system generates an image for each received training data. In a next step, the spatial image structure of the image generated by the machine learning system is determined in the following manner, i.e., the corresponding generated image is fed to a second machine learning system and the second machine learning system determines the spatial image structure of the corresponding image, and the second machine learning system is configured to determine the spatial image structure of the digital image. In a further step, the further parameters are adapted using a loss function, wherein the loss function measures the similarity between the determined spatial image structure of the image generated by the machine learning system and the spatial image structure predetermined by the training data belonging to the generated image.

[0007] An advantageous aspect of the methods described above and below is that not all parameters in the trainable copy need to be adapted in subsequent training. This can in particular reduce the required training time and / or the required training data, thereby allowing a faster and / or more cost-effective adaptation of the machine learning system.

[0008] According to a preferred specific embodiment, the further parameters are adapted by means of a distribution-based evolutionary algorithm.

[0009] Advantageously, non-continuously differentiable learning systems or loss functions can thereby be used. This is useful, for example, when a metric such as the mean intersection over union (IoU) should be optimized directly. Another advantage compared to using gradient-based optimization is that, due to the iterative diffusion process in the diffusion model, the activations of all intermediate steps must be kept in memory during backpropagation. This can quickly lead to out-of-memory problems, which can be avoided in the use of the evolutionary algorithm proposed here. In addition, an advantage of using gradient-free optimization compared to gradient-based optimization is that "vanishing gradient" problems may occur during backpropagation through the "unrolled" diffusion model, which is similar to problems that may occur, for example, in recursive networks. This problem can make gradient-based optimization more difficult, which can be avoided again by using gradient-free methods.

[0010] In particular, distribution-based evolutionary algorithms can be derivative-free policy gradient estimation algorithms.

[0011] According to a preferred embodiment, the loss function is given by a non-differentiable metric. In the case where the semantic image structure is predetermined by taking the semantic segmentation mask as a predetermined value for the spatial image structure, the non-differentiable metric can be given, for example, by the mean IoU metric. According to another preferred embodiment, the pixel-wise cross entropy or the mean square error of each key point can be used as the loss function instead, both of which are differentiable compared to the mean IoU metric. The pixel-wise cross entropy can also provide a relatively good uncertainty calibration.

[0012] According to a preferred embodiment, two adapter layers are respectively inserted into at least the replicas of the Transformer block, and the loss function is used in subsequent training to adapt the parameters added by inserting the adapter layers.

[0013] The number of parameters to be adapted in the subsequent training is thereby advantageously reduced. This is particularly advantageous with regard to the duration of the subsequent training: Firstly, the method described here may be slower than subsequent training by means of backpropagation. However, this effect of a longer training duration may be partially compensated by reducing the number of parameters to be adapted. (Subsequent) training by means of an evolution-based estimation algorithm may be slower than (subsequent) training by means of backpropagation, since in each training cycle the losses must be determined for a plurality of sampled parameter sets in order to estimate the gradients, rather than calculating the gradients by means of backpropagation based on only one parameter set.

[0014] According to a preferred embodiment, a prefix of length K is respectively added in front of at least one trainable copy of the Transformer block. Here, the key vector and the value vector in the self-attention layer of the Transformer block can be modified by the prefix of the associated Transformer block. In addition, an additional gating mechanism with scalar parameters can be introduced, wherein the parameters related to the prefix and the scalar parameters of the gating mechanism are added parameters, which will be adapted in the subsequent training of the machine learning system.

[0015] In the above-described embodiment, the number of parameters to be adapted in the subsequent training is also advantageously reduced, so that the above-mentioned advantages associated with the reduction in the number of parameters to be adapted with respect to the duration of the subsequent training and thus possibly also with respect to the costs of the subsequent training can be achieved.

[0016] According to a preferred embodiment, the additional parameters to be adapted in the subsequent training can be obtained by: i Represented as the pre-trained weight matrix Wi,0 The sum of products of two other matrices W i =W i,0 +W i,A ·W i,B , the ranks of the two other matrices are respectively lower than the weight matrix W i,0 The elements of these two additional matrices are the added parameters to be adapted in subsequent training. i =W i,0 +W i,A ·W i,B The first term includes the AxB weight matrix W corresponding to the corresponding layer i,0 , whose entries have been adapted in the pre-training of the machine learning system and do not change in subsequent training. The second term is represented by the matrix W i,A and W i,B Here, W i,A represents the Axr matrix, and W i,B Represents the rxB matrix, whose entries are the added parameters, which will be adapted in subsequent training. Here, r represents a freely selectable hyperparameter that determines the matrix W i,A , W i,B The hyperparameter r can be smaller than the rank of the matrix W. i,0 The value of the rank value of . Preferably, r can take a value of r≤16, for example. For example, r can take a value of r=4. For example, the matrix W i,A The added parameters, i.e., entries, of can be initialized randomly (e.g., in a Gaussian distribution). i,B The parameters of , i.e., the matrix entries, can be set to zero initially, i.e., at the beginning of subsequent training.

[0017] In the above-described embodiments, the number of parameters to be adapted in the subsequent training is also advantageously reduced, so that the above-mentioned advantages associated with the reduction in the number of parameters to be adapted with respect to the duration of the subsequent training and thus possibly also with respect to the costs of the subsequent training can be achieved again.

[0018] Preferably, the specification of the spatial image structure can be respectively given by a semantically segmented image, ie a semantic segmentation mask, or respectively given by an image or an image mask having at least one key point pose of a human body.

[0019] According to a preferred embodiment, the training data also includes a text consisting of a plurality of words, wherein the text describes the quality and / or content of the image to be generated by the machine learning system. By including text in the training data in addition to the pre-determination of the spatial image structure, another possibility of checking or expanding the content of the image to be generated can be provided. Through the text, for example, the resolution or style of the image can be set, and / or whether the image should be a night / dark / hazy or unclear or "blurred" image due to fog or other weather conditions can be set. This provides another possibility of generating diversified images, such as for safety-critical situations in autonomous driving, automated optical inspection, or indoor and object monitoring, which represent a predetermined spatial structure but for example different times, seasons, weather, etc. and / or different types of objects in the image. Here, different types can be given, for example, in the background area in the semantic segmentation mask, in the case of the driving environment, by bushes, trees, agricultural land, urban environment, etc. In the case of key point gestures, the text can also be used to specify, in particular, what kind of body shape, what kind of clothing, and / or what kind of object should be held in the hands of the person represented. For example, the text can be used to specify that a person holding a gun or other dangerous object should be represented in the image to be generated. As a result, images with safety-critical situations can be generated by a machine learning system, which images can be used to train the machine learning system, wherein the latter machine learning system is to be used within the scope of autonomous driving, automated optical inspection or object monitoring.

[0020] According to a preferred embodiment, the diffusion model has a U-Net structure including multiple encoder blocks, multiple decoder blocks and an intermediate block. Here, each of the above blocks can include multiple layers of the machine learning system, which together constitute a functional unit within the structure of the machine learning system. For example, the block can be a ResNet block, a Transformer block, a multi-head attention block, a conv-bn-relu block, etc.

[0021] According to a preferred embodiment, the encoder block and the intermediate block of the diffusion model with a U-Net structure are copied as trainable copies and locked copies, respectively. Here, each trainable copy of the block can be connected to the associated locked copy via a convolutional layer. Here, the convolutional layer can be a zero convolutional layer in particular. The zero convolutional layer can be a convolutional layer in which the weights and biases are initialized with the value zero. The function of the zero convolutional layer can be to input a predetermined initial input of the spatial image structure into the representation space, or to input the representation determined by the trainable encoder block or the trainable intermediate block into the decoder block of the machine learning system by addition at a subsequent position in the architecture of the machine learning system.

[0022] According to a preferred embodiment, the digital images generated by the subsequently trained machine learning system are training data for an image classifier for image classification. The image classification performed by the above-mentioned image classifier can be based on low-level features. Low-level features can be, for example, edges or pixel properties of the image.

[0023] In particular, images generated by a subsequently trained machine learning system may show at least part of the driving environment of an autonomous robot, or at least a portion of a workpiece to be inspected for defects and / or functionality, or the environment as seen from the perspective of a surveillance camera.

[0024] According to another aspect, the invention relates to a device, such as a computer, comprising means for executing the above method.

[0025] Furthermore, the present invention also relates to a computer program comprising machine-readable instructions which, when executed on one or more computers, cause the one or more computers to perform one of the methods according to the present invention described above and below. The present invention also includes a machine-readable data carrier on which the above-mentioned computer program is stored, as well as a computer equipped with the above-mentioned computer program and / or the above-mentioned machine-readable data carrier. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The embodiments of the present invention will be explained in detail below with reference to the accompanying drawings. In the accompanying drawings:

[0027] Figure 1 The training system is schematically shown;

[0028] Figure 2 An overview of the information flow of the method described here is schematically shown. DETAILED DESCRIPTION

[0029] Figure 1 An exemplary embodiment of a training system 140 is shown, which is used for subsequent training of a machine learning system 60 with the aid of a training data set T, wherein the machine learning system is configured to generate digital images given a predetermined spatial image structure. The training data set T comprises a plurality of input signals {m i , t i}, these input signals are used to train the machine learning system 60. The training data T can respectively include the spatial image structure m for each digital image to be generated. i These pre-determined information about the spatial image structure can be given, for example, in the form of a semantic segmentation mask m i In the form of key point pose m i In general, the training data is given in the form of iThe pre-set may also include a text consisting of multiple words. i , which describes the quality and / or content of the image to be produced by the machine learning system.

[0030] The machine learning system 60 may be a generative probabilistic text-to-image diffusion model, where blocks of the diffusion model and their parameters are replicated as locked copies and trainable copies.

[0031] Additional parameters θ to be adapted in the subsequent training are added to the parameters of the machine learning system 60 that have already been adapted in the pre-training. This can be achieved by inserting at least one additional layer in the trainable copy of the replica block of the learning system 60. Alternatively or additionally, the additional parameters can be obtained by representing the weight matrix to be adapted in the subsequent training in the layer of the trainable copy of the replica block of the machine learning system as the sum of matrix products of the pre-trained weight matrix and two further matrices of respectively lower rank, and in this case the elements of the two further matrices describe the added additional parameters θ.

[0032] For training, the training data unit 150 accesses a computer-implemented database St2, wherein the database St2 provides a training data set T. The training data unit 150 preferably randomly determines at least one input signal {m i , t i}, the input signal preferably has a spatial image structure m i at least one text which predetermines and contains the content and / or quality of the image to be generated i The training data unit 150 converts the input signal {m i , t i} is transmitted to the machine learning system 60. The machine learning system 60 determines the output signal yi based on the input signal, that is, the learning system generates a digital image y i As the output signal. The generated image y i is transmitted by the machine learning system 60 to the second machine learning system 61. The second machine learning system is configured to determine, from the input of the digital image, the spatial image structure of the digital image received as input as an output. That is, the second machine learning system 61 can determine, based on the digital input image, the semantic segmentation of the image or the key point poses of one or more persons represented in the image. The second machine learning system 61 now determines the image y generated by the machine learning system 60 i The spatial image structure m i The spatial image structure m is predetermined by the associated training data i and the image y generated for the associated training data determined by the second machine learning system 61 i The spatial image structure mi is transmitted to the changing unit 180 .

[0033] Then based on the spatial image structure m given by the training data i and the determined spatial image structure m i , the modification unit 180 determines a new parameter θ for the machine learning system 60. To this end, the modification unit 180 transforms the predetermined semantic segmentation mask m into a new parameter θ by means of a loss function, for example. i Or a predetermined keypoint pose (one or more) m i and the determined segmentation mask m i or the determined keypoint pose(s) m i For comparison, the loss function measures the similarity between the determined spatial image structure of the image generated by the machine learning system and the spatial image structure predetermined by the training data belonging to the generated image. The loss function determines a first loss value, which characterizes how large the deviation between the determined spatial image structure and the predetermined spatial image structure is, or how large the similarity between the spatial image structures is. In this embodiment, it is preferred to select a non-differentiable metric as the loss function, such as selecting the average intersection over union (IoU). In alternative embodiments, other loss functions may also be considered.

[0034] The changing unit 180 determines a new parameter θ based on the first loss value. In this embodiment, this is done with the aid of a distribution-based evolutionary algorithm. Preferably, this can be a derivative-free policy gradient estimation algorithm.

[0035] The determined new parameter θ is stored in the model parameter memory St1. Preferably, the determined new parameter θ is provided to the machine learning system 60 as a parameter Φ.

[0036] In a further preferred embodiment, the described training is repeated in an iterative manner for a predefined number of iteration steps, or iteratively until the first loss value is below a predefined threshold. Alternatively or additionally, it can also be provided that the training is terminated when the average first loss value involving the test or validation data set is below a predefined threshold. In at least one iteration, the new parameter θ determined in the previous iteration is used as the parameter θ of the machine learning system 60.

[0037] Furthermore, the training system 140 may include at least one processor 145 and at least one machine-readable storage medium 146 , wherein the machine-readable storage medium contains instructions, which, when executed by the processor 145 , cause the training system 140 to perform the training method of one of the aspects of the present invention.

[0038] Figure 2A schematic information flow overview of an embodiment of a computer-implemented method for subsequent training of a machine learning system described herein is shown. The machine learning system is preferably configured to generate a digital image under a predetermined spatial image structure. Preferably, the machine learning system is a generative probabilistic text-to-image diffusion model, wherein blocks of the diffusion model and their parameters are copied as locked copies and trainable copies. In a first method step S1, training data are received by the machine learning system, each training data including a predetermined spatial image structure. In a further method step S2, additional parameters are added to the trainable copy of at least one replicated block of the diffusion model. Here, additional parameters can be added by inserting at least one additional layer parameterized with the additional parameters in the trainable copy of the replicated block. Alternatively, additional parameters can be added by decomposing at least one weight matrix to be adapted in the subsequent training in the trainable copy of the replicated block into the sum of a pre-trained weight matrix and additional addends added in the subsequent training. The additional addends are given by the matrix product of two additional matrices, wherein the two additional matrices are parameterized with the additional parameters. The parameters of the pre-trained weight matrix are preferably adapted in the pre-training and remain unchanged in the subsequent training. The ranks of the two other matrices are preferably lower than the rank of the pre-trained weight matrix. In method step S3, an image is generated for each training data by a machine learning system. In step S4, the spatial image structure of the image generated by the machine learning system is determined by a second machine learning system, which is used to determine the spatial image structure of the digital image. Then, in step S5, other parameters are adapted and a loss function is used, which measures the similarity between the determined spatial image structure of the image generated by the machine learning system and the spatial image structure predetermined by the training data belonging to the generated image. Here, a distribution-based evolutionary algorithm is preferably used.

[0039] The term "computer" includes any device for processing predefined calculation rules. These calculation rules can exist in the form of software, hardware, or a combination of software and hardware.

[0040] In general, a plurality may be understood as being indexed, i.e., each element in the plurality is assigned an explicit index, preferably by assigning consecutive integers to the elements contained in the plurality. Preferably, when the plurality comprises N elements, where N is the number of elements in the plurality, integers 1 to N are assigned to these elements.

Claims

1. A computer-implemented method for subsequent training of a machine learning system (60), wherein the machine learning system (60) is configured to train a given spatial image structure (m i ), wherein the machine learning system (60) is a generative probabilistic text-to-image diffusion model, wherein blocks of the diffusion model and their parameters are copied as locked copies and trainable copies, the method comprising the following steps: (S1) receiving training data (T) through the machine learning system (60), each training data including a spatial image structure (m i ), (S2) adding additional parameters (θ) in the trainable copy of at least one replicated block of the diffusion model, wherein the additional parameter (θ) is added by inserting at least one additional layer parameterized by the additional parameter in a trainable copy of the replicated block, and / or wherein the further parameters (θ) are added by decomposing at least one weight matrix to be adapted in the subsequent training in the trainable copies of the replicated block into a sum of a pretrained weight matrix and a further addend added in the subsequent training, wherein the further addend is given by a matrix product of two further matrices, wherein the two further matrices are parameterized with the further parameters, wherein the parameters of the pretrained weight matrix are adapted in the pretraining and remain unchanged in the subsequent training, wherein the ranks of the two further matrices are both lower than the rank of the pretrained weight matrix, (S3) Generate an image (y) for each training data (T) by the machine learning system i ), (S4) determining, by a second machine learning system (61), an image (y i ) spatial image structure (m i ), the second machine learning system is used to determine the spatial image structure of the digital image, (S5) adapting the further parameter (θ) using a loss function that measures the image (y i ) of the determined spatial image structure (m i ) and the spatial image structure (m) predetermined by the training data (T) associated with the generated image i ) between them. 2 . The method according to claim 1 , wherein the further parameter (θ) is adapted by means of a distribution-based evolutionary algorithm.

3. A method according to any one of the preceding claims, wherein the loss function is given by a non-differentiable metric.

4. The method according to claim 1 , wherein at least two adapter layers are respectively inserted in the replica of the Transformer block, The loss function is used in subsequent training to adapt the parameters (θ) added by inserting the adapter layer.

5. The method according to any one of claims 1 to 3, wherein a prefix of length K is respectively added in front of at least one trainable copy of the Transformer block, The key vector and value vector in the self-attention layer of the Transformer block can be modified by the prefix of the associated Transformer block, respectively. An additional gating mechanism with a scalar parameter is introduced, The parameters related to the prefix and the scalar parameters of the gating mechanism are added parameters (θ), and the added parameters will be adapted in subsequent training.

6. A method according to any one of claims 1 to 3, wherein the additional parameters are added by: a weight matrix W in a layer of a trainable copy of the replicated block to be adapted in subsequent training i Decomposed into a pre-trained weight matrix W i,0 The sum of products of two other matrices W i =W i,0 +W i,A ·W i,B , the ranks of the two other matrices are respectively lower than the weight matrix W i,0 rank, The elements of the two additional matrices are respectively the added parameters (θ) to be adapted in subsequent training, in -W i,0 represents the AxB weight matrix of the machine learning system corresponding to the layer, whose entries have been adapted in pre-training and will not change, -W i,A represents the Axr matrix, W i,B represents the rxB matrix, whose entries are the added parameters that will be adapted to the target task in subsequent training, -r represents a freely selectable hyperparameter that determines the matrix W i,A ,W i,B rank.

7. A method according to any one of the preceding claims, wherein the training data (T) also includes text (t i ), wherein the text describes the quality and / or content of the image to be generated by the machine learning system.

8. The method according to any one of the preceding claims, wherein the diffusion model has a U-Net structure including a plurality of encoder blocks, a plurality of decoder blocks and an intermediate block.

9. The method of claim 7, wherein the encoder block and the intermediate block of the diffusion model are replicated as trainable copies and locked copies, respectively, and wherein each trainable copy of a block is connected to the locked copy through a convolutional layer.

10. A method according to any one of the preceding claims, wherein the digital image (y) generated by the subsequently trained machine learning system (60) i ) is training data for image classifiers used for image classification, especially image classification based on low-level features.

11. The method according to claim 10, wherein the image generated by the subsequently trained machine learning system (60) displays at least part of the driving environment of the autonomous robot, or displays at least a portion of a workpiece to be inspected for defects and / or functionality, or displays the environment seen from the perspective of a surveillance camera.

12. An apparatus (140) comprising means for performing the method of any one of claims 1 to 11.

13. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 11.

14. A machine-readable storage medium having stored thereon a computer program according to claim 13.