Improved training of slot-based machine learning models

By decomposing the feature map into N components and initializing the features of K slots, the long calculation time and optimization difficulties in training machine learning models in the prior art are solved, and a more efficient and deterministic training process and better task result quality are achieved.

CN119948491APending Publication Date: 2025-05-06ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380067582.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-21
Filing Date
2023-09-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When training machine learning models, it is difficult to effectively allocate and process image features, resulting in long calculation time, difficulty in achieving global optimal solutions, and high uncertainty in training results.

Method used

By decomposing the feature map into N components and distributing features according to a predetermined number of slots K, the characteristics of the slots are initialized, thereby optimizing the training process, reducing the number of iterations, improving the possibility of global optimal solutions, and increasing the certainty and repeatability of training.

Benefits of technology

The number of iterations during the training process is reduced, the possibility of the optimized global optimal solution is improved, the determinism and repeatability of training is enhanced, the calculation time is reduced, and the quality of task results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948491A_ABST
    Figure CN119948491A_ABST
Patent Text Reader

Abstract

A method (100) for training a machine learning model (1) configured to process measurement data (2) regarding a given task, where the model (1) comprises a slot building portion (11) configured to assign logged features (3) of input measurement data (2) to a plurality of slots (4a-4d), and a task portion (12) configured to assign logged features (3) of the input measurement data (2) to the plurality of slots (4a-4d). The task part (12) is configured to process a feature (3) of a record of measurement data (2) respectively assigned to each slot (4a-4d), the method (100) comprising the steps of: providing (110) a training sample (2a) for the record of measurement data (2); generating (120) at least one feature map (3 #) from each training sample (2a) by at least one feature extraction layer (11a) in the slot construction section (11); decomposing (130) each feature map (3 #) into N components (3a-3f) according to a predetermined decomposition criterion and / or method; initializing (140) a predetermined number of K slots (4a-4d) by assigning features (3) from the feature map (3 #) to the slots (4a-4d) based at least in part on the N components (3a-3f); and training (150) the slot building portion (11) and the task portion (12) towards at least one given target.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the training of a specific class of machine learning models that, with respect to a given task, assign features of input data to these machine learning models to multiple slots before processing them slot by slot. Background Art

[0002] Many image processing tasks, such as discovering and classifying objects visible in an image, are dual in nature. First, object instances need to be identified, and second, some evaluation of the object instances with respect to a given task needs to be performed. It has proven promising to separate these two tasks and partition machine learning models for image processing accordingly. For example, the slot building part of the model can decompose the objects of an image using multiple slots, each of which is assigned only a single object. The contents of different slots can then be processed independently of each other with respect to a given task.

[0003] (F. Locatello et al., "Object-Centric Learning with Slot Attention", arXiv:2006.15055) discloses a machine learning model that uses a slot attention module to assign features of an image to slots and then processes these slots with respect to different tasks. Starting from an initialization randomly drawn from a normal distribution, the slot attention module is trained in an iterative manner using a recursive function. Summary of the invention

[0004] The present invention provides a method for training a machine learning model configured to process measurement data about a given task. In particular, the term "machine learning model" includes any function with trainable parameters that exhibits the ability to generalize from training samples to unseen data. For example, a machine learning model can be a neural network, or it can include one or more neural networks.

[0005] The machine learning model comprises a slot construction part and a task part, wherein the slot construction part is configured to assign features of input measurement data records to multiple slots, and the task part is configured to process the features of the measurement data records respectively assigned to each slot. In this article, "respectively" may in particular include: processing the features assigned to one slot is performed independently of processing the features assigned to all other slots. That is, if the features assigned to one slot are processed with respect to a given task to produce a certain task result, then changing the features assigned only to other slots will not change the task result.

[0006] In the process of the method, training samples of measurement data records are provided. At least one feature extraction layer in the slot construction part generates at least one feature map based on each training sample. For example, the feature map can have three dimensions, where two dimensions represent coordinates in the feature map and the third dimension represents different channels for each feature.

[0007] According to a predetermined decomposition criterion and / or method, each feature map is decomposed into N components. In particular, the value of N determines the level of detail of the original data to be retained, and the value of N can be selected according to the desired level of detail.

[0008] Initialize a predetermined number of K slots by assigning features from the feature map to the slots based at least in part on the N components. That is, each slot may be assigned one or more features from one or more components, or a mixture or superposition of one or more features from one or more components. In particular, the value of K may be motivated by the task at hand. For example, the value of K may represent the number of object instances sought in the image.

[0009] Then, the slot construction part and the task construction part can be trained towards at least one given goal. The two parts can be trained simultaneously, but it is also possible to first train the slot construction part and then start training the task part with or without further training of the slot construction part.

[0010] The advantage of initializing K slots based on features from the N components into which the feature map has already been decomposed is that the optimization process of training only needs to take a shorter path in the direction of the optimal assignment of features to slots. This is because the optimization does not need to be performed from scratch; instead, it starts with information that is already apparent from the feature map and has already been extracted from it. This starting point is already somewhat close to the optimal assignment in feature space, whereas a random initialization could be anywhere in feature space. This has a triple advantage.

[0011] First, the number of iterations required to reach an optimal assignment of features to slots is reduced, and hence the computation time is reduced, since there is no need to rediscover information that is already at hand.

[0012] Second, it increases the probability that the optimization will reach a global optimum for the assignment of features to slots. In addition to one global optimum, there are likely to be several local optima. The random initialization may be so close to one such local optimum that the optimization locks on that local optimum instead of striving to reach the global optimum.

[0013] Third, training becomes more deterministic and repeatable. The decomposition of each feature map into its N components can be performed using deterministic operations. This means that each time you perform training, you get the same assignment of features to slots. Given the specific application at hand, this is more reasonable than an assignment with some random variation. For example, if the slots represent object instances, there is only one objective fact about which image feature belongs to which object instance. When optimizing the slots again based on the same training data, there is no motivation why the contents of the slots should vary.

[0014] In particular, if both the slot construction part and the task part of training are performed simultaneously, it becomes easier to diagnose any problems if the final result of the training is not satisfactory. If both the slot construction part of training and the task part of training start from some random initialization, it is more difficult to pinpoint the problem.

[0015] In a particularly advantageous embodiment, the measurement data comprises at least one image. In this case, at least one feature extraction layer is a convolutional layer, which applies one or more filter kernels to the measurement data in a sliding manner with a predetermined stride. The image may encode any suitable observable quantity in pixel values, such as the intensity of light received by a detector. The pixels may be arranged in any suitable regular grid, such as a two-dimensional or three-dimensional grid. In particular, a stack of multiple convolutional layers can be used to first detect basic features in a first layer, and then detect more complex features in later layers.

[0016] In a further particularly advantageous embodiment, the slot represents an object type and / or an object instance whose presence the measurement data indicates. This facilitates the processing or manipulation of individual object instances or individual types of objects in the scene described by the measurement data.

[0017] In another advantageous embodiment, both the slot construction part and the task part are trained towards a single goal, which is measured based on the results produced by the task part. That is, the training is performed "end-to-end", and the performance of the machine learning model can be measured by comparing the final results with the ground truth, which is used to label the corresponding training samples. In this way, a separate ground truth is not necessary for the training of the slot construction part.

[0018] In a further particularly advantageous embodiment, during initialization of the slots, each feature in at least one feature map is assigned to at most one slot. In this way, the processing and / or manipulation of each entity (such as an object instance or object type) corresponding to a slot becomes more independent from the processing of other entities corresponding to other slots. In essence, similar to image editing software, the slots correspond to layers, which are stacked on top of each other, and their blending together produces the final image. The information on each layer can be manipulated without affecting the information on any other layer.

[0019] In a further particularly advantageous embodiment, the features of the feature map are clustered into a plurality of clustering results during the predetermined decomposition method. This allows different components of the feature map to be identified in an unsupervised and statistically motivated manner.

[0020] An exemplary clustering method is k-means clustering. This algorithm quickly finds the centroid of clustering results in feature space. It requires the number of clustering results N sought as input.

[0021] Another exemplary clustering method is mean shift clustering, which does not require presetting the number N of clustering results sought, but at the expense of making convergence more difficult.

[0022] In another advantageous embodiment, during the predetermined composition method, each feature is mapped to a numerical value by the convolutional layer of the slot construction part. The features are then assigned to components based at least in part on the numerical value. The numerical values ​​provide a simple way to sort the features, and numerical ranges and / or thresholds can be used to group the features. In particular, the differences between the numerical values ​​to which different features are mapped introduce the concept of distances between these features.

[0023] In a further particularly advantageous embodiment, the initialization of the slots comprises transforming features characterizing or belonging to the components into features characterizing or belonging to the slots using at least one trainable transformation layer of the slot building part. In this way, there is better flexibility to adapt the number N of components into which the feature map is decomposed to the desired number K of slots. As mentioned before, N determines the level of detail in the feature map that is retained, while K is motivated or given by the specific application (task) at hand.

[0024] In particular, at least one trainable transformation layer may be a fully connected layer with a rectified linear unit (ReLU) activation function. This provides a linear transformation between features belonging to a component and features belonging to a slot. That is, features may be transferred from a component to a slot as they are, but a feature in a slot may also be a linear combination of two or more features in one component or in different components.

[0025] Examples of given tasks for which this method can be used to train machine learning models include:

[0026] ● Reconstruct the original measurement data according to the characteristics assigned to the slots;

[0027] ● classify the type of object the measurement data indicates exists;

[0028] • creating a new record of the measurement data, the new record being modified relative to the particular object whose existence the measurement data indicates; and

[0029] ●Combining multiple records of measurement data into a new record of measurement data.

[0030] All of these tasks benefit from decoupling the identification of object types and object instances from the work performed with those object types and object instances. In particular, once the slot construction part has been trained, that training can be at least partially reused when switching to another task.

[0031] The goal of training is to enable the machine learning model to produce good task results. Therefore, in another particularly advantageous embodiment, once training is completed, the record of measurement data captured by at least one sensor is provided to the trained slot construction part. The features of the record of measurement data are assigned to the slot by the trained slot construction part. The features assigned to each slot are processed by the trained task part, which produces a task result for a given task. Due to the improvement in training, the quality of the task results is improved. That is, when the accuracy of the task is measured by comparing the task results with a benchmark truth value in the form of test or verification data, this accuracy will be improved.

[0032] In a further particularly advantageous embodiment, an actuation signal is calculated based on the task result. The vehicle, the driver assistance system, the monitoring system, the quality inspection system and / or the medical imaging system are actuated with the actuation signal. In this way, the probability is increased that the action performed by the actuation system in response to the actuation is appropriate for the situation encoded in the measurement data.

[0033] The method may be implemented in whole or in part by a computer and may therefore be embodied in software. Therefore, the present invention also relates to a computer program having machine-readable instructions which, when executed by one or more computers and / or computing instances, cause one or more computers and / or computing instances to perform the above method. In this document, control units for vehicles and other embedded systems in technical equipment capable of executing machine-readable instructions are also considered to be computers. Examples of computing instances are virtual machines, containers or serverless execution environments, in which machine-readable instructions can be executed in the cloud. The present invention also relates to a machine-readable data carrier and / or a download product having a computer program. A download product is a digital product having a computer program, which may be sold, for example, in an online store for immediate fulfillment and download to one or more computers. The present invention also relates to one or more computing instances having a computer program and / or a machine-readable data carrier and / or a download product.

[0034] In the following, further advantageous embodiments will be explained using the figures, without any intention of limiting the scope of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The figures show:

[0036] Figure 1 is an exemplary embodiment of a method 100 for training a machine learning model 1;

[0037] Figure 2 is a graphical representation of the effect of improved initialization on finding the optimal assignment of features 3 to slots 4a-4d;

[0038] Figure 3 is an exemplary process of initial assignment of training sample 2a toward features 3 to slots 4a-4d.

[0039] Figure 1 1 is a schematic flow chart of an embodiment of a method 100 for training a machine learning model 1. The machine learning model 1 includes a slot construction part 11 and a task part 12, the slot construction part 11 is configured to assign features 3 of records of input measurement data 2 to multiple slots 4a-4d, and the task part 12 is configured to process the features 3 of records of measurement data 2 assigned to each slot 4a-4d respectively.

[0040] In step 110 , a recorded training sample 2 a of measurement data 2 is provided.

[0041] In step 120, at least one feature extraction layer 11a in the slot construction part 11 extracts at least one feature from each training sample 2a. Figure 3 #.

[0042] In step 130, each feature is decomposed according to a predetermined decomposition standard and / or method. Figure 3 # is decomposed into N components 3a-3f.

[0043] According to box 131, the feature Figure 3 #Feature 3 can be clustered into multiple cluster results. According to box 132, in response to feature 3 belonging to a cluster result, the feature 3 can be assigned to the component corresponding to the cluster result. Such a cluster result can then be represented by a specific feature 3, such as a feature 3 corresponding to or closest to the center of the cluster result.

[0044] According to box 133, the feature Figure 3 #'s feature 3 may be mapped to a numerical value by the convolutional layer 11b of the slot construction portion 11. According to block 134, the feature 3 may then be assigned to components 3a-3f based at least in part on the numerical value.

[0045] In step 140, the N components 3a-3f are combined with the feature Figure 3# features 3 are assigned to these slots 4a-4d, initializing a predetermined number K of slots 4a-4d. In a simple case, K is equal to N, and components 3a-3f can be immediately reused as slots. However, as previously described, a higher number N of components 3a-3f that retain more details can be processed into a lower number K of slots 4a-3d in any suitable manner.

[0046] For example, according to block 141, features 3 representing or belonging to components 3a-3f may be transformed into features 3' representing or belonging to slots 4a-4d using at least one trainable transformation layer 11c of the slot construction portion 11. For example, features 3' in slots 4a-4d may be linear combinations of features 3 in components 3a-3f.

[0047] In step 150, the slot construction part 11 and the task part 12 are trained towards at least one given goal. Herein, the training of the slot construction part 11 can utilize the assignment of features 3 to slots 4a-4d that have been obtained for the training sample 2a in any suitable manner. The resulting training states of the slot construction part 11, the task part 12, and the machine learning model 1 as a whole are labeled with reference marks 11*, 12*, and 1*, respectively.

[0048] In step 160, a record of measurement data 2 captured by at least one sensor is transferred to the trained slot construction part 11*. Then, in step 170, features 3 of the record of measurement data 2 are assigned to slots 4a-4d by the trained slot construction part 11*. Then, the trained task part 12* processes the features assigned to each slot 4a-4d, which in step 180 produces a task result 5 for a given task.

[0049] In step 190, an actuation signal 190b is calculated based on the task result 5. In step 200, the vehicle 50, the driver assistance system 60, the monitoring system 70, the quality inspection system 80 and / or the medical imaging system 90 are actuated using the actuation signal 190a.

[0050] Figure 2 The diagram shows how the improved initialization of slots 4a-4d increases the probability of finding a global optimal solution G in the space of all possible assignments of features 3 to slots 4a-4d. In addition to the global optimal solution G, there are many local optimal solutions L in the space. The previous random initialization IR could place the initial assignment of features 3 to slots 4a-4d anywhere in the space. Therefore, this initial assignment may have been close to the global optimal solution G, but it may also be closer to one of the local optimal solutions L. In the latter case, the optimization may strive to reach such a local optimal solution, which is illustrated in a dashed line in one example.

[0051] In contrast, the current deterministic initialization according to steps 110 to 140 of method 100 has already allocated the initial allocation I D is placed very close to the global optimal solution. Therefore, when the slot construction part 11 is trained in step 150, the optimization is very likely to find the global optimal solution G.

[0052] Figure 3 The diagram illustrates how a training sample 2a is processed into an assignment of its features 3 to slots 4a-4d. The training sample 2a is an image that can have any practical number of pixels. With the aid of one or more feature extraction layers 11a, the features of the training sample 2a are generated. Figure 3 #.exist Figure 3 In the example shown in Figure 3 #Includes 64×64 features, and each feature includes 64 channels. Therefore, the feature Figure 3 # is a tensor of dimension 64×64×64.

[0053] According to blocks 131 and 132, by means of k-means clustering, the features Figure 3 The feature 3 of # is initially assigned to N=6 different components 3a-3f. Each such component 3a-3f can be characterized by a specific representative feature 3, such as the center of the clustering result corresponding to the corresponding component 3a-3f. Each such representative feature 3 has a dimension of 64.

[0054] According to block 141, features 3 that characterize or belong to components 3a-3f are transformed into features 3' in K=4 slots 4a-4d. For example, different slots 4a-4d may contain features 3' that are different linear combinations of features 3 in components 3a-3f. In this process, each feature 3' maintains a dimension of 64.

[0055] If the assignment of features 3 to slots 4 a - 4 d is calculated iteratively during inference, this initialization can also be reused for the actual measured data 2 .

Claims

1. A method (100) for training a machine learning model (1), the machine learning model (1) being configured to process measurement data (2) about a given task, wherein the model (1) comprises a slot construction part (11) and a task part (12), the slot construction part (11) being configured to assign features (3) of a record of input measurement data (2) to a plurality of slots (4a-4d), the task part (12) being configured to process the features (3) of the record of measurement data (2) assigned to each slot (4a-4d), the method (100) comprising the following steps: ● providing (110) training samples (2a) for the records of measurement data (2); ● generating (120) at least one feature map (3#) according to each training sample (2a) by at least one feature extraction layer (11a) in the slot construction part (11); ● Decomposing (130) each feature map (3#) into N components (3a-3f) according to a predetermined decomposition criterion and / or method; Initializing (140) a predetermined number K of slots (4a-4d) by assigning features (3) from the feature map (3#) to the slots (4a-4d) based at least in part on the N components (3a-3f); and - Training (150) the slot building part (11) and the task part (12) towards at least one given goal.

2. The method (100) according to claim 1, wherein the measurement data (2) comprises at least one image, and wherein at least one feature extraction layer (11a) is a convolutional layer which applies one or more filter kernels to the measurement data (2) in a sliding manner with a predetermined stride.

3. The method (100) according to any one of claims 1 to 2, wherein the slot (4a-4d) represents an object type and / or an object instance whose presence the measurement data (2) indicates.

4. A method (100) according to any one of claims 1 to claim 3, wherein both the slot building part (11) and the task part (12) are trained (151) towards a single goal, and the single goal (151) is measured based on the results produced by the task part.

5. A method (100) according to any one of claims 1 to 4, wherein during initialization (140) of the slots (4a-4d), each feature (3) of the at least one feature map (3#) is assigned to at most one slot (4a-4d).

6. The method (100) according to any one of claims 1 to 5, wherein the predetermined decomposition method comprises: ● clustering (131) the features (3) in the feature graph (3#) into a plurality of clustering results; as well as In response to a feature (3) belonging to a clustering result, the feature (3) is assigned (132) to a component corresponding to the clustering result.

7. The method (100) according to claim 6, wherein the clustering (131a) is performed by k-means clustering or mean shift clustering.

8. The method (100) according to any one of claims 1 to 7, wherein the predetermined decomposition method comprises: The convolutional layer (11b) of the slot construction part (11) maps (133) each feature (3) to a numerical value; and - assigning (134) said features (3) to components (3a-3f) based at least in part on said values.

9. A method (100) according to any one of claims 1 to 8, wherein the initialization (140) of the slot includes transforming (141) features (3) representing or belonging to components (3a-3f) into features (3') representing or belonging to slots (4a-4d) using at least one trainable transformation layer (11c) of the slot building part (11).

10. The method (100) of claim 9, wherein at least one trainable transformation layer (11c) is a fully connected layer with a rectified linear unit (ReLU) activation function.

11. The method (100) according to any one of claims 1 to 10, wherein the given task comprises: • reconstructing the original measurement data (2) according to the features (3) assigned to the slots (4a-4d); and / or ● classifying the type of object the measurement data (2) indicates is present; and / or ● creating a new record of measurement data (2'), said new record being modified with respect to the specific object whose existence said measurement data (2) indicates; and / or ●Combining multiple records of measurement data (2) into a new record of measurement data (2').

12. The method (100) according to any one of claims 1 to 11, further comprising: • providing (160) a record of measurement data (2) captured by at least one sensor to a trained slot construction part (11*); - assigning (170) features (3) of said record (2) of measurement data to slots (4a-4d) by said trained slot building part (11*); as well as The features (3) assigned to each slot (4a-4d) are processed (180) by the trained task part (12*) to obtain a task result (5) for the given task.

13. The method (100) of claim 12, further comprising: Calculating (190) an actuation signal (190a) based on the task result (5); and actuating (200) a vehicle (50), a driver assistance system (60), a monitoring system (70), a quality inspection system (80) and / or a medical imaging system (90) using the actuation signal (190a).

14. A computer program comprising machine-readable instructions which, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) according to any one of claims 1 to 13.

15. A non-transitory data carrier and / or download product having a computer program according to claim 14.

16. One or more computing instances having a computer program according to claim 14 and / or a non-transitory data carrier and / or a download product according to claim 15.