Computer-implemented method for balancing non-uniform distribution in training data during training of machine learning algorithms

By defining auxiliary categories and combining classification tasks and regression tasks, a weighted total loss function is formed, which solves the problem of uneven data distribution in machine learning algorithm training and improves the model's prediction performance for rare categories.

CN120030341APending Publication Date: 2025-05-23ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411678623.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

During the training process of machine learning algorithms, unevenly distributed training data are encountered, which makes the model difficult for appropriate detection and prediction of rare categories or events.

Method used

By defining auxiliary categories, creating classification tasks and regression tasks and combining them, calculating the classification probability of each auxiliary category, determining the weighted classification loss function, and combining them with the regression loss function to form a total loss function to train machine learning algorithms.

Benefits of technology

Effectively balance the uneven distribution in the training data, reduce the impact on the regression task, and improve the model's performance on rare category prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030341A_ABST
    Figure CN120030341A_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method for balancing non-uniform distribution in training data during training of a machine learning algorithm, where the training data comprises a plurality of data sets, where the machine learning algorithm solves regression tasks, where the training data has a non-uniform distribution in terms of their labels, where the machine learning algorithm comprises a plurality of data sets. Wherein the method comprises the steps of:-defining an assistance category for the training data (S10); creating a classification task (S12); -determining a classification probability for each auxiliary category (S14); determining a classification loss function for the classification task (S16); -weighting (S18) said classification loss function; determining a total loss function (S20); -training the machine learning algorithm (S22); and-providing the trained machine learning algorithm (S24).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of training machine learning algorithms, in particular training with incomplete or unevenly distributed training data. Background Art

[0002] Unevenly distributed training data refers to a situation where the distribution of values, categories, or classes in the training data is uneven. This means that there is a significant difference between the number of examples of different values ​​or classes, with some values ​​and classes being significantly more common than others. This uneven distribution can occur in a variety of applications and datasets, such as in image recognition, medical diagnosis, fraud detection, or other machine learning scenarios.

[0003] When training data is unevenly distributed, this can cause machine learning models, especially neural networks, to have difficulty properly detecting and predicting rare categories or events because they tend to focus on more common categories due to the uneven distribution. This means that models tend to achieve better results for common categories, while rare categories are neglected. In machine learning practice, managing and coping with unevenly distributed training data are important challenges and require special strategies to balance what the model learns, especially what is easy to learn due to the uneven distribution.

[0004] There are some methods that try to deal with imbalanced classification tasks. However, these methods cannot be used for regression tasks.

[0005] A known method for balancing the uneven distribution in training data involves determining a focal loss for identifying wrong predictions with high accuracy. The focal loss was developed to address the uneven distribution with the help of background classes. When solving regression tasks, such as determining the speed of objects in road traffic from camera images, most objects will belong to the background class. In this case, high-security wrong predictions for one class are penalized by assigning a higher weight to the corresponding loss.

[0006] Another approach is to use neural networks for ordered regression. Here, an attempt is made to solve a downstream classification task, in which the categories should be ordered and the names should be encoded according to the importance of the category. Here, a sigmoid cross entropy loss function can be used to compare the label encoding with the prediction. However, this has disadvantages. The larger the category, the greater its "weight" in the total loss. In the case of classic downstream regression tasks, this order of categories does not always exist. When converting a regression problem to an ordered classification, it is assumed that the order of the categories corresponds to the importance of the categories, for example 5>4>3>2>1>0.

[0007] When it comes to regression tasks such as road traffic, the training data usually has unevenly distributed regression data. This may be because in driving scenarios, most objects have low speeds or are located within a range of 100 meters. However, the advantage of camera sensors over other modalities such as lidar sensors or radar sensors is early object recognition from a long distance. In the case of training data with high uneven distribution, the model is motivated to learn the unevenly distributed data distribution, that is, the model is better cut in the representative data distribution and worse cut in the less representative data distribution. Therefore, the model may have to learn to make use of the imbalanced training data to predict the regression task.

[0008] According to the methods proposed in the prior art, each category name is encoded with a gradually increasing weight. However, this raises two problems. Not all regression categories are strictly ordered. For example, for traffic classification, the speed categories 0m / s, 1m / s, 2m / s, ...., 50m / s should have the same encoding as 0m / s, -1m / s, -2m / s, ...., -50m / s, because -50m / s and 50m / s are equally important. However, the sign will mean that another separate subtask is needed to predict the sign. In addition, it can be assumed that 50m / s is more important than 4m / s, but this is not applicable to most regression tasks. Summary of the invention

[0009] Therefore, the task of the present invention is to propose a method for training a machine learning algorithm, which can efficiently balance the uneven distribution in the training data so as to minimize the impact of the uneven distribution on the regression task.

[0010] This object is achieved by the subject matter of the independent claims.

[0011] According to a first aspect of the invention, the task is solved by a computer-implemented method for balancing uneven distributions in training data during training of a machine learning algorithm. The training data comprises a plurality of data sets for training the machine learning algorithm. The machine learning algorithm solves a regression task by determining an output value for each data set. A regression loss function is determined for the regression task. The regression loss function quantifies the quality of solving the regression task.

[0012] The training data has a non-uniform distribution in terms of its labels. This means that the distribution of the values ​​to be predicted via the regression task is non-uniform in the training data. Thus, there is a significant difference between the number of examples of different values, with some values ​​appearing more frequently than others.

[0013] The method comprises the following steps:

[0014] -Define auxiliary categories for training data;

[0015] -Create classification tasks to assign datasets into auxiliary categories and perform classification and regression tasks;

[0016] -Determine the classification probability of each auxiliary category, where the classification probability indicates whether the data set is correctly assigned to the corresponding category;

[0017] -Determine the classification loss function for the classification task;

[0018] -The classification loss function is weighted by the classification probability of each category to form a weighted classification loss function;

[0019] -Determine the total loss function from the regression loss function and the weighted classification loss function for training the machine learning algorithm;

[0020] - train the machine learning algorithm using the total loss function; and

[0021] -Provide trained machine learning algorithms.

[0022] A machine learning algorithm is an algorithm that has been developed to automatically identify patterns and relationships in data and make predictions or decisions. The machine learning algorithm is created by training with existing data and can then be applied to new, unknown data to generate predictions or classifications.

[0023] The data set is divided into auxiliary categories before training. The auxiliary categories are selected in particular so that they divide the data set into groups with respect to uneven distribution. Through these additional separations, another factor is provided for performing training for the model on which the machine learning algorithm is based.

[0024] Creating a classification task may include, among other things, converting the labels of a dataset into a scheme that conforms to auxiliary classes.

[0025] For example, the labels of a dataset may include the speed or distance of an object identified in a video. Transformation may then mean equipping the dataset with additional labels, which are derived from existing labels. For example, objects of different speeds may be divided into speed categories, such as "slow" 0km / s to 25km / s, "medium" 25km / s to 50km / h, and "fast" 50+km / s. In another example, a dataset may be divided by distance into "near" up to 10m, "medium" 10m to 50m, and "far" 50+m.

[0026] However, auxiliary classes can also be defined in the abstract. The parameters by which the auxiliary classes are defined do not necessarily correspond to known or commonly used variables. Thus, the division into auxiliary classes can be based on a quotient representing the ratio of two physical variables, for example, but has no practical significance outside of a machine learning algorithm.

[0027] Classification probability is used to determine the likelihood that a data set is assigned to a specific class.

[0028] As in each subtask of the model, a loss function is also determined for the classification of the data set with respect to the auxiliary classes. This classification loss function is multiplied by the classification probability of each class in order to take into account the influence of the frequency of the data set being assigned to the auxiliary classes on the calculation of the loss function.

[0029] From the regression loss function and the weighted classification loss function, a total loss function is determined, which quantifies the quality of the model. The total loss function can then be used to train the model. Since the weighted classification loss function is included in the total loss function, the uneven distribution in the training data can be balanced or at least partially compensated when training the model. This solves the task of the present invention.

[0030] In the last step, the trained machine learning algorithm can be provided so that it can be used for the preset regression task. The classification of the processed data set and the auxiliary categories associated with it are no longer required in the trained machine learning algorithm.

[0031] In one embodiment, the machine learning algorithm includes a linear model, a decision tree, a support vector machine, or a neural network.

[0032] Machine learning algorithms can take many forms such as linear models, decision trees, support vector machines, neural networks, etc. The algorithm optimizes by learning from the training data in such a way that it recognizes patterns and rules in order to make the best possible predictions or classifications for new data.

[0033] The effectiveness of a machine learning algorithm depends on various factors, including the quality and quantity of training data, the choice of algorithm, model configuration, and the evaluation of the model based on evaluation metrics. Models can be continuously improved and optimized to maximize accuracy and performance.

[0034] Linear regression models assume a linear relationship between a dependent variable and one or more independent variables and are used to make predictions about continuous values.

[0035] Support Vector Machines (SVM) is a model used for classification or regression and identifying patterns in data. The support vector machine finds the best separation between different classes, or tries to fit a continuous function to the data.

[0036] A decision tree is a model that creates decision rules in the form of a tree diagram. It divides data based on features and enables prediction or classification.

[0037] Naive Bayes is a probabilistic model that is based on Bayes' theorem and is used for classification. It assumes that features are independent of each other and calculates the probability of a specific class based on the given features.

[0038] Neural networks refer to models that are preferably used for processing images or other grid-based data. The architecture of a neural network includes a number of nodes, neurons or nodes, which are arranged in different layers.

[0039] In one embodiment, the machine learning algorithm includes a convolutional neural network.

[0040] A convolutional neural network (CNN) is a special type of artificial neural network that is usually designed for tasks in the field of image processing and pattern recognition, although it also has applications in other fields such as language processing. It is characterized by the use of convolution operations, which are used to extract features from input data.

[0041] The main components of CNN are convolutional layers. These layers are responsible for applying convolution operations to the input data. Convolution enables the network to recognize features in the data such as edges, textures, and other local patterns. In addition, convolutional neural networks can also include pooling layers and fully connected layers.

[0042] Pooling Layers are used to reduce the spatial dimension of features and improve computational efficiency. In this case, maximum pooling or average pooling is typically used. Fully connected layers are located at the end of the network and are used for classification or regression. The fully connected layers combine the extracted features to generate the final output.

[0043] Convolutional neural networks have demonstrated extremely high performance for tasks such as image classification, object recognition, face recognition, and segmentation. They have the ability to learn features hierarchically, meaning they can learn from simple features like edges to more complex concepts in the data such as shapes and objects. This makes them particularly useful for problems where the spatial structure of the data is important.

[0044] In one embodiment, a machine learning algorithm is trained to determine the parameters of each data set.

[0045] Machine learning algorithms that only have to determine one parameter for each dataset have several advantages over models that must perform multiple tasks simultaneously or sequentially. So-called single-task models are usually simpler in structure and have fewer parameters. This results in higher computational efficiency and lower computational resource requirements. They are easier to train and optimize. Single-task models usually converge faster during training because they target a specific task. The model can focus on features and patterns that are relevant to that task.

[0046] In addition, single-task models are often easier to interpret because they have clear connections between input and output. This is particularly important in applications such as medical diagnosis. In multi-task models that perform multiple tasks simultaneously or sequentially, different tasks may interfere with or influence each other. In contrast, single-task networks are less susceptible to such interference.

[0047] In one embodiment, the machine learning algorithm is applied in a vehicle or an autonomous robotic unit and the parameter is the velocity of an object in the surroundings of the vehicle or autonomous robotic unit.

[0048] The recognition of the speed of objects in the surroundings of a vehicle or a robotic unit is an important contribution to controlling them. Objects, in particular other vehicles and / or robotic units, may also move in the surroundings. Without central position and collision control, vehicles and / or robotic units rely on their perception to recognize all objects that could potentially be a danger to them or that could cross their path and thus constitute a collision risk.

[0049] By predicting the speed of an object, a motion vector of the object can be determined, which is used to predict the future movement of the object. The motion prediction can be performed in a subsequent step or in a dedicated machine learning algorithm for each detected object.

[0050] In one embodiment, the classification loss function is determined by using a cross entropy loss function.

[0051] The cross-entropy loss function, also known as the "Cross-Entropy Loss Function", is a mathematical function commonly used in machine learning theory, especially in the field of supervised learning. This function is usually used to measure the error or difference between the predicted probability distribution and the actual probability distribution, especially in classification tasks.

[0052] In classification tasks, the task is to train a model to classify input data into different categories or classes. The cross entropy loss function here helps determine how well the model's predictions agree with the actual classes.

[0053] In one embodiment, the cross entropy loss function is a Softmax function.

[0054] The Softmax function, also known as the Softmax activation function, is a mathematical function that is widely used in machine learning algorithms and especially in multi-class classification. The Softmax function is used to produce a probability distribution over multiple categories or classes based on a series of raw values ​​or activations.

[0055] The Softmax function takes a real vector as input and converts it into a probability distribution. The Softmax function produces a probability distribution in which all probabilities are between 0 and 1 and sum to 1. This makes it particularly suitable for classification tasks, in which the model calculates probabilities for each possible class. Typically, the class with the highest probability is used as the model's prediction. The Softmax function advantageously normalizes the output of a machine learning algorithm.

[0056] In one embodiment, the auxiliary categories divide the possible outcomes of the regression task into groups.

[0057] In this implementation, both the regression task and the classification task aim to predict the same attribute. In this case, it can be assumed that the learned weight parameters are uniform in the loss landscape, that is, the gradient that helps the classification task will also help the main regression task, because both learn to predict the same attribute.

[0058] In multi-task learning, gradient conflicts between multiple tasks are often observed. An auxiliary classification task that bears a small portion of the total loss can help the main regression task to balance these gradient conflicts. Under this assumption, the auxiliary classification learns to balance the unbalanced regression tasks with little or no observed gradient conflicts.

[0059] In one embodiment, the class loss function is normalized.

[0060] Advantageously, the normalized class loss function is easier to interpret and easier to further process.

[0061] In one embodiment, the class loss function and the regression loss function are weighted differently when determining the overall loss function.

[0062] Depending on the degree of imbalanced distribution in the training data, the classification loss function should have more or less influence on the total loss function. Preferably, the weights of the regression loss function and the classification loss function are trainable hyperparameters, so that the weights are adjusted according to the imbalanced distribution in the training data when training the machine learning algorithm.

[0063] In one embodiment, the machine learning algorithm is trained to solve other tasks and uses the other loss functions for those tasks when determining the overall loss function.

[0064] So-called multi-task models can solve several tasks in parallel or in sequence. This can be used, for example, to predict not only the speed of an object, but also its movement in the future. Thus, a multi-task model can have a superordinate task and several subtasks, of which the regression task is such a subtask.

[0065] In such a multi-task model, other tasks should be considered in the total loss function, especially when the results of the subtasks affect or are even correlated with each other. By considering the loss functions of other tasks in the total loss function, machine learning algorithms can be advantageously used to solve tasks beyond simple parameter prediction.

[0066] In another aspect, the invention relates to a computer program having a program code for executing the method as described above when the computer program is executed on a computer.

[0067] In another aspect, the invention relates to a computer-readable data carrier having a program code of a computer program in order to perform a method as described above when the computer program is executed on a computer.

[0068] In another aspect, the present invention relates to a system for training a machine learning algorithm, wherein the system is configured to perform the method as described above.

[0069] In summary, it should be pointed out that the present invention provides a method for balancing uneven distributions in training data during the training of a machine learning algorithm, a computer program for performing the method, a computer-readable data carrier having a program code of the computer program, and a system for performing the method.

[0070] The described embodiments and developments can be combined with one another as desired.

[0071] Further possible embodiments, refinements and implementations of the invention also include combinations of features of the invention which are not explicitly mentioned and which are described above or below with reference to the exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The accompanying drawings shall facilitate a further understanding of the embodiments of the present invention. The accompanying drawings illustrate the embodiments and, together with the description, serve to explain the principles and concepts of the present invention.

[0073] Further embodiments and many of the advantages described are apparent from the drawings, in which the elements shown are not necessarily shown to scale with respect to one another.

[0074] Figure 1 The flow of a method according to one embodiment is schematically shown.

[0075] In the figures of the drawings, identical reference numerals denote identical or functionally identical elements, components or parts, unless otherwise indicated. DETAILED DESCRIPTION

[0076] Figure 1 The flow of the method according to the embodiment of the present invention is schematically shown.

[0077] In a first step S10, auxiliary categories are created for the training data. Auxiliary categories support the solution of the regression task in the following way: the auxiliary categories realize a kind of pre-classification of the data set. If the regression task of the machine learning algorithm is, for example, to determine the speed of an object from a camera image, the auxiliary categories can divide the object into different speed classes. These classes can, for example, include "fast" for more than 50 km / h, "moderate" for 25 km / h to 50 km / h and "slow" for less than 25 km / h. Other divisions are also possible. The choice of auxiliary categories and their boundaries should be adapted to the regression problem to be solved.

[0078] The auxiliary categories are preferably selected so that they can divide the training data in terms of uneven distribution. Returning to the example where the regression task is to predict speed, the auxiliary categories are preferably defined so that they can divide the data set into a range that forms an over-represented group and a range that forms an under-represented group. In embodiments, there may also be ranges in between or outside of the two.

[0079] In step S12, a classification task is created. Creating the classification task may in particular include supplementing the training data with corresponding labels for auxiliary categories. In addition, independent segments, such as subnetworks as part of a neural network, may be created, which solve the classification task. Training is then performed so that the classification task is solved together with the regression task.

[0080] In step S14, the classification probability is determined. The classification probability indicates how often a data set is assigned to the correct class. From this, the probability of the correct classification for each auxiliary class is derived.

[0081] In step S16, a classification loss function is determined. The classification loss function is used to optimize the classification task. In the following step S18, the classification loss function is weighted so that the weights correspond to the frequency of the data sets in the corresponding auxiliary categories. This can balance one or more categories that contain significantly more data sets than other auxiliary categories.

[0082] In step S20, a total loss function is determined for the machine learning algorithm. The total loss function is determined by the regression loss function and the classification loss function. If the machine learning algorithm also solves other tasks, the loss functions used to solve these tasks are also used to determine the total loss function.

[0083] In step S 22, the machine learning algorithm is trained with the total loss function. This step includes, among other things, back-propagation or adjusting the hyperparameters and weights used by the machine learning algorithm to solve the regression task and possibly other tasks.

[0084] In some embodiments, steps S14 to S22 are repeated until the machine learning algorithm correctly solves the regression task for the data set in the training data and the uneven distribution in the training data is balanced.

Claims

1. A computer-implemented method for balancing uneven distributions in training data during training of a machine learning algorithm, The training data includes multiple data sets used to train machine learning algorithms. wherein the machine learning algorithm solves the regression task by determining output values ​​for each of the data sets, Wherein a regression loss function is determined for the regression task, where the training data has an uneven distribution in terms of its labels, The method comprises the following steps: - defining auxiliary categories for the training data (S10); - creating a classification task (S12) for assigning the data set to auxiliary categories and performing the classification task and regression task; - determining a classification probability for each auxiliary class (S14), wherein the classification probability indicates whether the data set is correctly assigned to the corresponding class; -determining a classification loss function for the classification task (S16); - weighting the classification loss function with the classification probability of each category (S18) to form a weighted classification loss function; - determining a total loss function (S20) from the regression loss function and the weighted classification loss function for training the machine learning algorithm; - training the machine learning algorithm using the total loss function (S22); and -Providing a trained machine learning algorithm (S24).

2. The computer-implemented method of claim 1, wherein the machine learning algorithm comprises a linear model, a decision tree, a support vector machine, or a neural network.

3. The computer-implemented method of claim 1 , wherein the machine learning algorithm comprises a convolutional neural network.

4. A computer-implemented method according to any of the above claims, wherein the machine learning algorithm is trained to determine the parameters of each data set.

5. The computer-implemented method of claim 4, wherein the machine learning algorithm is applied in a vehicle or an autonomous robotic unit, and wherein the parameter is the velocity of an object in the surrounding environment of the vehicle or autonomous robotic unit.

6. A computer-implemented method according to any one of the above claims, wherein the classification loss function is determined by using a cross entropy loss function.

7. The computer-implemented method of claim 6, wherein the cross entropy loss function is a Softmax function.

8. The computer-implemented method of any of the preceding claims, wherein the auxiliary categories divide possible outcomes of the regression task into groups.

9. A computer-implemented method according to any one of the above claims, wherein the class loss function is normalized.

10. The computer-implemented method of any one of the preceding claims, wherein the class loss function and the regression loss function are weighted differently in determining the total loss function.

11. A computer-implemented method according to any of the above claims, wherein the machine learning algorithm is trained to solve other tasks, and wherein the other loss functions of these tasks are used in determining the total loss function. 12 . A computer program having a program code for carrying out the method according to claim 1 , when the computer program is executed on a computer. 13 . A computer-readable data carrier having a program code of a computer program for executing the method according to claim 1 , when the computer program is executed on a computer.

14. A system for training a machine learning algorithm, wherein the system is configured to perform the method according to any one of claims 1 to 11.