Method for training machine learning model
Patent Information
- Application Number
- JP2024113645
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-03
- Filing Date
- 2024-07-16
- Publication Date
- 2025-08-01
AI Technical Summary
【0020】 本発明の更なる利点、特徴、および詳細は、本発明の例示的な実施形態が図面を参照して詳細に説明される、以下の説明から得られる。特許請求の範囲および明細書において言及される特徴は、個々にまたは任意の所与の組み合わせにおいて、本発明にとり必須であり得る。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for training a machine learning model, and further to a machine learning model, a computer program, an apparatus and a memory medium for this purpose. [Background technology]
[0002] When multiple tasks (e.g., semantic segmentation, object recognition, classification, etc.) are processed simultaneously in a neural network, this is called multitasking. In most cases, all tasks share a part of the network (called the "backbone"). Each task gets its own layer (called the "head").
[0003] To train a multitasking network, the various losses of the individual tasks must be combined into an overall loss. Correct weighting of the task-specific losses is crucial to ensure the desired distribution of capacity within the neural network. Summary of the Invention [Means for solving the problem]
[0004] The subject matter of the present invention relates to a method having the features of claim 1, a machine learning model having the features of claim 7, a computer program having the features of claim 8, an apparatus having the features of claim 9 and a computer-readable memory medium having the features of claim 10. Further features and details of the invention emerge from the respective subclaims, the description and the drawings. Of course, the features and details explained in relation to the method according to the invention also apply in relation to the machine learning model according to the invention, the computer program according to the invention, the apparatus according to the invention and the computer-readable memory medium according to the invention. In each case, vice versa. As a result, in relation to the disclosure, cross-reference is always made or can always be made to the individual aspects of the invention.
[0005] The subject of the present invention relates to a method for training a machine learning model, in particular for machine applications, such as for object recognition in autonomous driving, said method comprising at least some of the following training steps: Specific to the application of a machine learning model, is the step of providing training data. Beginning processing of the training data where multiple tasks are processed simultaneously by the machine learning model. Determining a loss for each task, where the individual loss is based on the difference between the output generated by the machine learning model and a default (which may be pre-defined), providing a reference for application of the machine learning model, e.g. in the form of ground truth. A step of weighting the determined losses, which can be performed based on analytical calculations, preferably optimal calculations, in particular analytical calculations of optimal values, using at least one task-specific uncertainty, in order to compensate in particular for tasks of different scales. Updating weights of the machine learning model based on the weighted losses for training the machine learning model.
[0006] In the method according to the invention, various losses such as classification loss and regression loss can be simultaneously learned using multitask loss. This serves the particular purpose of allowing different quantities and units to be used. At least one task-specific, i.e. specifically task-dependent, uncertainty can be used. This uncertainty reflects the relative reliability between tasks and can be a function of the task representation or unit. The determination of the individual losses and / or the weighting of the determined losses can be performed based on at least one loss function. The loss function, preferably the multitask loss function used for the multitask loss, can be based on maximizing the Gaussian probability with homoscedastic uncertainty.
[0007] It is possible to exploit task-specific uncertainties to compensate for tasks of different scales. For example, one task may measure in millimeters and another in meters. This can result in dramatically different loss weights that must be compensated for. Using conventional methods, the initialization of task weights is not always sufficient. Known approaches focus mainly on calculating task weights during training. On the other hand, according to the present invention, the main focus can be placed especially on their initialization. Complex multitask systems include many different tasks. It turns out that in complex multitask systems, it is very important to initialize task weights in a well-distributed way. Neural networks are often trained using backpropagation based on gradient descent, which generally only adapts slowly. Thus, if the starting values are too far apart, large differences in loss weights are often not compensated for. This can result in divergence and slow down the learning process. This problem is becoming more important for multitask learning environments with multiple tasks (e.g., more than three) or very different tasks (e.g., object recognition and segmentation as well as different classification tasks). Thus, for example, in environments such as autonomous driving, various tasks such as pedestrian recognition, vehicle recognition, and road recognition are often required in various formats such as classification, bounding boxes, segmentation, and / or spline fitting.
[0008] Also advantageously, within the scope of the present invention, the step of updating the weights of the machine learning model is performed initially at the beginning of training and / or (subsequently) iteratively during training, based on analytical optimum calculations in each case. Alternatively or additionally, the losses can be weighted differently (i.e. not constant) at the beginning. The weights of the loss functions for the losses for various tasks, in particular the task weights, can be initialized differently. The present invention can have the advantage that large differences between tasks can be compensated within the first few iterations of training. It has been shown that this can result in more stable training (relying only on previous methods can lead to divergence), faster convergence (thus reducing computational costs), and improved overall performance across all tasks. Correct initialization of the uncertainty weights can be a key component in the training pipeline.
[0009] Optionally, the loss for each task can be determined based on a task-specific loss function. The loss function is preferably based on the distribution of a location scale family and / or on the approximation and / or maximization of a Gaussian probability with homoscedastic uncertainty. The location scale family can include a family of probability distributions parameterized by a location parameter and a non-negative scale parameter. Furthermore, the task-specific uncertainty can be determined based on the loss determined for the individual task. Furthermore, the representation or units of the tasks can be different, and the task-specific uncertainty can be a function of the representation or units. Furthermore, it is conceivable that the step of weighting the determined loss is performed based on an analytical optimum calculation and / or on a loss function of the current and / or previous iterations of the training step. For regression tasks, the probability can be defined as a Gaussian distribution with a mean value provided by the model output. For classification tasks, the model output can be processed by a softmax function. From the resulting probability vector, samples can be obtained. In this way, the uncertainty terms can be reduced during training. This improves the optimization process.
[0010] The analytical optimum calculation, in particular the loss weighting of the current iteration, can be additionally stabilized with the loss weighting of several previous iterations of the training step, preferably using a moving average. Furthermore, it is conceivable that the analytical optimum calculation includes the calculation of a moving average over the loss weighting of the current iteration and several previous iterations of the training step. Thus, a more reliable training of the machine learning model can be enabled.
[0011] Further advantages within the scope of the present invention are achievable when the machine learning model is designed as an artificial neural multitasking network to handle multiple tasks simultaneously to assist the operation of a machine, preferably to provide autonomous driving, and thus more complex tasks related to the operation of the machine can be handled.
[0012] Further, optionally, the training data is specific to sensor data collected from at least one sensor of the machine, preferably a vehicle or robot, preferably a camera sensor. The machine learning model can be trained for autonomous driving and / or autonomous navigation applications, in particular for pedestrian recognition and / or vehicle recognition and / or road recognition. In many cases, a single neural network may be required to solve multiple tasks simultaneously, share knowledge between tasks, and increase computational efficiency. Thus, in the case of autonomous driving, for example, lanes, pedestrians, vehicles, drivable surfaces, etc. may be evaluated.
[0013] Furthermore, the tasks may include semantic segmentation and / or object recognition and / or classification, which are preferably represented differently and / or involve different units, which may result in different scales of the loss function.
[0014] Furthermore, machines can be designed as tools, etc. For example, an intelligent drill may need to simultaneously evaluate the type of material and the depth of the drill bit. It is also possible to design machines as surveillance cameras, where recognition of pedestrians, potentially dangerous situations, and other tasks of interest can be derived simultaneously.
[0015] The subject of the invention further relates to a machine learning model, in particular trained by the method according to the invention, which assists the operation of the machine through the processing of multiple tasks for the machine, and thus offers the same advantages as those detailed with respect to the method according to the invention.
[0016] The subject matter of the invention further relates to a computer program, in particular a computer program product, which comprises commands that prompt the computer to carry out the method according to the invention when the computer program is executed by the computer. The computer program according to the invention therefore offers the same advantages as those detailed above with respect to the method according to the invention.
[0017] The subject of the invention further relates to a data processing device arranged to carry out the method according to the invention. For example, a computer running a computer program according to the invention can be provided as the device. The computer can have at least one processor for executing the computer program. Furthermore, a non-volatile data memory can be provided in which the computer program is stored and from which the processor can read the computer program for execution.
[0018] The subject of the present invention further relates to a computer-readable memory medium. The computer-readable memory medium comprises a computer program and / or a command, which, when executed by a computer, prompts the computer to execute the method according to the invention. The memory medium is designed, for example, as a data memory, such as a hard disk and / or a non-volatile memory and / or a memory card. The memory medium can, for example, be integrated into the computer.
[0019] Furthermore, the method according to the present invention can also be designed as a computer-implemented method.
[0020] Further advantages, features and details of the invention emerge from the following description, in which exemplary embodiments of the invention are explained in detail with reference to the drawings. The features mentioned in the claims and in the specification may be essential to the invention either individually or in any given combination. [Brief description of the drawings]
[0021] [Figure 1] 1 is a schematic diagram of a method, an apparatus, a memory medium, and a computer program product according to an exemplary embodiment of the present invention. [Diagram 2] 13 shows an example of the resulting equation for compensating for multiple tasks. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] A method 100, an apparatus 10, a memory medium 15, and a computer program 20 according to an exemplary embodiment of the invention are illustrated diagrammatically in Figure 1. Also illustrated is a machine learning model 50 according to an exemplary embodiment of the invention.
[0023] Fig. 1 shows a method 100 that can be used to train a machine learning model 50 for application to a machine. First, according to a first training step 101, training data is provided. The training data can be specific to the application of the machine learning model 50. This can mean that the training data includes inputs of a type that will also be processed by the machine learning model in a subsequent application. The training data can also include defaults, i.e. criteria that indicate in particular which outputs of the machine learning model 50 are expected from the processing.
[0024] Processing of the training data, in which multiple tasks are processed simultaneously by the machine learning model, can then begin according to a second training step 102. Losses can then be determined according to a third training step 103. The losses are determined specifically for each task. The individual losses can be based on the difference between the output generated by the machine learning model 50 and a default. The determined losses can then be weighted according to a fourth training step 104. The weighting can be performed using task-specific uncertainties based on analytical optimal calculations. In addition, as a fifth training step 105, the weights of the machine learning model 50 can be updated for training of the machine learning model 50 based on the weighted losses.
[0025] The present invention can build on known approaches, such as those described in "Multi-task learning using uncertainty to weigh losses for scene geometry and semantics" by Kendall, Alex, Yarin Gal, and Roberto Cipolla, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. This allows the use of task-specific uncertainty to compensate for tasks of different scales. For example, one task can be measured in millimeters and another in meters. This can result in dramatically different weightings of losses. These weightings can be compensated according to a variant of an embodiment of the present invention. However, the initialization for all task weights is traditionally a fixed constant, e.g. 1. This initialization can be improved according to a variant of an embodiment of the present invention, thus leading to significantly improved results in large and diverse multitask environments.
[0026] During training, Kendall et al. (see above) use the following equation to compensate the T tasks:
number
[0027] The invention can be advantageously split into several steps: estimating task weights for the first sample, using a moving average to weight the tasks, and seamlessly transitioning to the uncertainty weighting method, e.g., by Kendall et al. (see above).
[0028] According to the first step of estimating the task weights for the first sample, the following equation can be used to compensate the T tasks during training. For simplicity, the following discussion focuses on deriving the formula for regression tasks. For classification, the 2 in the denominator would be changed to, for example, 1.
number
[0029] To estimate the optimal initialization for a set of training samples and loss L(x, y), we can analytically compute the initialization based on the equations above. Since summations are involved, this can be done independently for each σ_i. The result is:
number
[0030] σ t Solving for gives the following result:
number
[0031] Thus, in the first iteration (e.g., based on the first image input to the neural network and the computation of the loss), which contains the loss for this particular task, the weights are
number
[0032] In Figure 2, Equation 1 above is expressed as the loss weighting σ t are plotted for each loss L0=1. The optimal solution is
number
[0033] This differs significantly from the previous solution by Kendall et al., which uses a constant to initialize σ_t and backpropagation (gradient descent) for updates.
[0034] According to the further step of using a moving average to weight the tasks, for a subsequent training iteration i, a moving average can be calculated over the loss weights of the current iteration and the previous iteration.
number
[0035] This can be performed for the first few iterations. The exact number N of the first iterations for which the analytical solution and the moving average are used is a hyperparameter. In general, the method has been empirically proven to be very robust with respect to this hyperparameter. As a general rule, the following applies: the higher the variance of the task-specific loss across samples, the higher the number of iterations that should be used to estimate the initialization. From a technical point of view, any value between 1 and the total number of training steps is possible. The actual network training can also be performed starting from the 0th iteration, since there is an estimation of the weighting of the loss at the 0th iteration.
[0036] Then, in the next step, we can seamlessly move to the uncertainty weighting method, e.g., by Kendall et al. After a predefined number of iterations used for initialization, we can change to a gradient descent based update of the uncertainty weights, similar to Kendall et al.
[0037] In the above description of the embodiments, the invention has been described exclusively by way of example. Of course, the individual features of the embodiments can be freely combined with one another, where technically feasible, without departing from the scope of the invention.
Claims
Claim 1 A method (100) for training a machine learning model (50) for application to a machine, the method comprising the following training steps, namely: providing (101) training data specific to the application of the machine learning model (50); starting (102) the processing of the training data in which a plurality of tasks are processed simultaneously by the machine learning model (50); determining (103) a loss for each individual task, the individual loss being based on the difference between the output generated by the machine learning model (50) and a default; weighting (104) the determined loss, the weighting step being performed using at least one task-specific uncertainty based on an analytical calculation; updating (105) the weights of the machine learning model (50) for training of the machine learning model (50) based on the weighted loss. The method (100) includes these steps. Claim 2 In the method (100) according to claim 1, the step (105) of updating the weights of the machine learning model (50) is performed initially at the start of training and / or subsequently iteratively during training, in each case based on an analytical optimal calculation, the losses being initially weighted differently, and the weighting of the loss functions for the losses for the various tasks, in particular the task weighting, being initialized differently. The method (100) is characterized by this. Claim 3 In the method (100) according to claim 1 or 2, the loss for each task is determined based on a task-specific loss function, the loss function preferably being based on a location-scale family of distributions and / or an approximation of a Gaussian probability with homoscedastic uncertainty and / or maximization, the task-specific uncertainty being determined based on the loss determined for the individual task, the task representations or units being different, the task-specific uncertainty being a function of the representation or unit, and the step (104) of weighting the determined loss being performed based on an analytical optimal calculation and / or on the loss functions of the current and / or previous iterations of the training step. The method (100) is characterized by this. Claim 4 In the method (100) according to claim 1 or 2, the method (100) is characterized in that an analytical optimal calculation, in particular the weighting of the loss of the current iteration, is additionally stabilized, in particular using a moving average, by the weighting of the losses of a plurality of previous iterations of the training step.
5. In the method (100) according to claim 1 or 2, the method (100) is characterized in that the machine learning model (50) is designed as an artificial neural multi-tasking network for simultaneously processing a plurality of tasks, preferably for providing autonomous driving, in order to assist the operation of a machine.
6. In the method (100) according to claim 1 or 2, the training data is specific to sensor data collected from at least one sensor of a machine, preferably a vehicle, preferably a camera sensor, and the machine learning model (50) is trained for application to autonomous driving, in particular for pedestrian recognition and / or vehicle recognition and / or road recognition, and the tasks include semantic segmentation and / or object recognition and / or classification.
7. A machine learning model (50) for assisting the operation of a machine through the processing of a plurality of tasks of the machine, trained by the method according to claim 1 or 2.
8. A computer program (20) comprising commands that, when executed by a computer (10), cause the computer to execute the method (100) according to claim 1 or 2.
9. A data processing device (10) configured to execute the method (100) according to claim 1 or 2.
10. A computer-readable memory medium (15) comprising commands that, when executed by a computer (10), cause the computer to execute the steps of the method (100) according to claim 1 or 2.