Annotating data using an artificial neural network
The method addresses the challenge of inconsistent data annotation by using neural networks for similar data and alternative methods for uncertain data, enhancing reliability and efficiency through loss functions, resulting in cost-effective and accurate labeling.
Patent Information
- Application Number
- DE102023212587
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for annotating datasets using artificial neural networks face challenges in accurately and efficiently handling both similar and uncertain data, leading to inconsistent and costly annotation processes.
A method that combines the use of an artificial neural network for annotating similar data and an alternative method for uncertain data, utilizing similarity and uncertainty criteria, along with loss functions like triplet and virtual adversarial losses, to enhance annotation reliability and efficiency.
This approach improves the reliability and reduces the cost of data annotation by leveraging the strengths of both neural networks and human annotation, ensuring high-quality labeling across varying data types.
Abstract
Description
FIELD OF THE INVENTIONThe present invention relates to a method for annotation data of a data set by means of an artificial neural network comprising annotated data and unnotated data.SUMMARY OF THE INVENTIONAccordingly, the following is provided:a method for annotation of data of a data set by means of an artificial neural network comprising annotated data and unnotated data, comprising the following steps: selecting unnotated data which meet predetermined similarity criteria relating to a similarity to the annotated data; annotation of the selected unnotated data which meet the similarity criteria by means of the artificial neural network; selecting unnotated data which meet predetermined uncertainty criteria relating to an annotation based on knowledge of the artificial neural network; annotation of the selected unnotated data which meet the uncertainty criteria by means of an alternative method and / or provision of the selected unnotated data which meet the uncertainty criteria for the purposes of annotation thereof.A sensor, also referred to as a detector, (measurement variable or measurement) sensor or (measurement) sensor, is a technical component that can record specific physical, chemical properties or states, e.g. temperature, humidity, pressure, speed, brightness, acceleration, pH value, ionic strength, electrochemical potential and / or the material nature of its environment qualitatively or quantitatively as a measurement variable. These variables are detected by means of physical or chemical effects and transformed as sensor data into an electrical signal that can be processed further.A central system is, for example, a cloud, that is to say a storage space which is provided as a service via the Internet. Central systems comprise central elements, such as central data memories or central evaluation units, which are provided on the Internet.An artificial neural network (ANN) is in particular a network of networked artificial neurons simulated in a computer program. The artificial neurons are typically arranged on different layers (layers). Usually, the artificial neural network comprises an input layer and an output layer (output layer), the neuron output of which becomes visible as the only one of the artificial neural network. Layers lying between the input layer and the output layer are typically referred to as hidden layers (hidden layers). Typically, an architecture or topology of an artificial neural network is initially initiated and then trained in a training phase for a specific task or for a plurality of tasks in a training phase.The training of a neural network and the minimization of a loss function are closely linked to each other.The loss function, also called a cost function or an error function, is a mathematical function that quantitates the error between the values predicted by the neural network and the actual values (the target values). The loss function measures how well or poorly the model fits to the given data. It may have various forms depending on the type of problem, such as mean squared error (mean squared error) for regression problems or cross entropy (cross entropy) for classification problems.Training a neural network is to adjust the model parameters, such as weights and bias terms in the neurons, to minimize loss function. The network is iteratively fed the training data and after each forward and backward pass (forward and backpropagation) the parameters are adjusted stepwise to reduce the error. This process is often controlled by an optimization algorithm such as the gradient descent method.Annotation or labeling of data refers to the process of assigning categories or labels to data points in a dataset. These labels serve to prepare and identify the data for machine learning and other analytical tasks.An annotated or unnotated data record contains data with or without labels.A dataset is a structured collection of data organized and stored in a particular manner to represent information about a specific domain or problem. Data sets consist of individual data points, also called data, which contain information about one or more attributes or features.Computer program products typically comprise a sequence of instructions which, when the program is loaded, cause the hardware to perform a specific method which leads to a specific result.The basic idea of the invention is to annotation a data set comprising both annotated data and unnotated data by means of different means. In this case, it is provided that a part of the unnotated data is annotated by means of an artificial neural network, provided that the artificial neural network recognizes a similarity of the unnotated data to the annotated data. This conditioning is based on the consideration that unnotated data, which resemble annotated data in their structure, i.e. data whose annotation is known to the artificial neural network, can be annotated by the artificial neural network with a high reliability.Data whose annotation by means of the artificial neural network is subject to a great uncertainty are annotated by means of an alternative method.An alternative method is understood to mean a method which differs from the artificial neural network with which the data which meet the similarity criterion have been annotated.The alternative method may be less available or more expensive than the artificial neural network. It is expedient if the alternative method achieves better results on the data which meet the uncertainty criterion than the artificial neural network.It is understood that instead of carrying out the alternative method, it is also conceivable to provide the selected unnotated data that meet the uncertainty criteria for the purposes of their annotation and to receive the provided data after their annotation.Accordingly, it is conceivable that the artificial neural network is provided by a first unit and the alternative method is carried out by a second unit.Advantageous embodiments and refinements emerge from the further dependent claims and from the description with reference to the figures of the drawing.According to a preferred development of the invention, a first loss of the artificial neural network for the annotated data is determined by means of a first loss function, and a second loss of the artificial neural network for the unnotated data is determined by means of a second loss function before annotated data are annotated by means of the artificial neural network and / or before unnotated data which meet the predetermined similarity criteria or meet the predetermined uncertainty criteria are selected.Thus, the knowledge of the artificial neural network can be quantitatively evaluated via a data set with annotated and unlabeled data. The quantitative assessment of the knowledge of the artificial neural network with respect to a data set can be used for the specification of a similarity criterion and / or an uncertainty criterion.According to a preferred development of the invention, the alternative method comprises presenting data to a person for the purposes of annotation. Accordingly, it is conceivable to pass corresponding data directly to a person or to provide corresponding data for the purposes of presentation if the alternative method is carried out by a unit that has not carried out the preceding steps.Human annotation is more reliable and cost-intensive than artificial neural network annotation.Alternatively, it is also conceivable to use a further neural network as an alternative method.According to a preferred development of the invention, the similarity is determined on the basis of a distance from unnotated data to annotated data.The definition of the distance of the unnotated data from the annotated data can be based on the definition of a metric, wherein the metric can be determined for a data pair and determines the distance of a data pair.It is also expedient here if a similarity criterion forms the dropping below a distance threshold value and / or the similarity criterion comprises a list of those N data which have the smallest distance from the annotated data. Here, N is a natural number. The phrase "distance to the annotated data" is to be understood here in such a way that this means the smallest distance to an annotated data entry of the data set, i.e. the next annotated data entry is used for the distance calculation.Thus, the reliability of the method for annotation of data can be specified by means of a parameter that is easy to interpret, namely the similarity criterion, which can be interpreted as a number.It will be appreciated that high desired reliability is associated with a low distance threshold or value for N, and this results in the method requiring more passes to note a predetermined number of data than a method with a higher value for N.According to a preferred development of the invention, the uncertainty criterion is based on the result of a decision tree, in particular a random tree.Accordingly, it is conceivable to check by means of a tree structure of a decision tree whether data have a structure which is already known to the artificial neural network. Experiments by the applicant have shown that this analysis can be implemented favorably by a decision tree.It is particularly expedient here if the unnotated data are ordered with respect to their uncertainty by means of a decision tree and the first M data are selected, wherein M is a natural number.Alternatively, it is conceivable to select data which are processed by the decision tree up to a specific hierarchy level.Thus, the frequency of applying the alternative method for annotation of data can be specified by means of an easily interpreted parameter, namely the uncertainty criterion, which can be interpreted as a number.It is thus possible to set how frequently data should be annotated by means of the alternative method.If the alternative method is a more reliable and more complex method compared to annotation using the artificial neural network, it follows that annotation of the unnotated data becomes more reliable and more complex as the M or the hierarchy level of the tree increases.A decision tree is a diagram or tree structure used to represent decision rules. In a decision tree, decisions are made stepwise by making a query at each node of the tree, and based on the results of the query, the algorithm is continued along the branches of the tree until a final decision is made. Each node in the tree represents a decision rule for a particular feature or attribute, and the leaves of the tree represent the classification or regression value.Decision trees are easily interpretable and can be used to model complex decision processes.A random tree or random forest is a decision tree in the field of machine learning. Rather than using only a single decision tree, a random forest creates a set (ensemble) of decision trees. These trees are trained in different ways using examples and features randomly selected from the dataset.When predictions are made, the results of the individual trees are combined to make a final prediction. This approach results in the random forest being less prone to overfitting and having better generalization properties.According to a preferred development of the invention, the first loss function is designed as a triple loss function.A triple loss function, also a triplet loss function, aims to learn a so-called "embedding" function that enables embedding data in a multi-dimensional space such that similar data points are close to each other and dissimilar data points are far from each other.The triplet loss function uses three data points (triplets) at a time:Anchor (anchor): the data point for which a representation is to be created. Positive example: a data point which is similar to the anchor, i.e. belongs to the same class or category.Negative example: a data point which differs from the anchor, i.e. belongs to a different class or category.The triplet loss function evaluates the distances between these three data points in the embedding space. It minimizes the distance between the armature and the positive example while maximizing the distance between the armature and the negative example. The aim is to ensure that similar data points are close together and dissimilar data points are far from each other.Training with a triplet loss function helps to learn the embeds such that similar data points are close together and dissimilar data points are far from each other.The mathematical form of a triplet loss function may be as follows:Here, ||x|| is the Euclidean norm (L2normal) of x, and alpha is a hyperparameter that controls the margin (distance) between the positive and negative examples. The goal is that the difference between the distances is greater than alpha to minimize performance.According to a preferred development of the invention, the second loss function is designed as a virtual adventrial loss function.A virtual adventitial loss function or virtual adventitial loss is a technique for improving robustness and stability of neural networks.In minimizing a virtual adventarial loss, i.e. in minimizing it, the aim is to make a neural network robust with respect to small, imperceptible disturbances in the input data. This is accomplished by training the model to make consistent predictions on input data that has been easily changed in an environment around a given input. A "consistent prediction" is a prediction that closely matches in similar or repeated situations or on similar data points.Thus, an artificial neural network becomes more resistant to small disturbances in the data, especially in situations where the model might be faced with unintentional or adversary disturbances in the input data.According to a preferred development of the invention, a loss of the artificial neural network is minimized by means of a total loss function on the data record before the first and second loss is determined.Minimizing the total loss corresponds to training of the artificial neural network on the entire dataset. Thus, the results of annotation of the data by the artificial neural network can be improved.It is also conceivable that the total loss function is of the form: first loss function+C*second loss function. C can be understood as a so-called hyperparameter, which can be specified by a user for weighting the second loss. C is a real number.A hyperparameter is a parameter that controls the configuration of a model instead of being learned from the training data by learning itself. In other words, hyperparameters are settings that are set before the actual training process and that influence the performance and behavior of the model.According to a preferred development of the invention, the method is repeated until all unnotated data have been annotated. Accordingly, the method is designed as an iterative method which produces new annotations with each iteration step.According to a preferred development of the invention, the data record is variable during the execution of the method and the artificial neural network is configured to determine the first loss for annotated data of a changed data record and to determine the second loss for unlabeled data of a changed data record as soon as the method is called up again after a change of the data record.It is conceivable, for example, that further data are added to the data set. Added data can already be annotated or unnotated and are processed accordingly by the method described.An active artificial neural network can thus be provided, which automatically selects or changes a data set to be processed.An active neural network is an artificial neural network that selectively selects data to improve its performance, particularly in minimizing a loss function.It is understood that a computer program product configured to perform the method as described above is advantageous.Furthermore, a central system is also stored on which a variable data record is stored, which is to be annotated by means of a method as described above, wherein the central system can be connected by means of one or more data memories which are configured to transmit data to the central system in order to change the data record. Accordingly, it is conceivable that a data memory of the central system can be reached by a plurality of users or terminals in order to change a database of the data record. For example, it is conceivable that one or more vehicles continuously transmit data to the central system for the purposes of annotation. In order to annotate the data by means of an artificial neural network, a computer program product as described above is executed on the central system or the central system is configured to transmit the variable data record to a computing unit configured to execute a computer program product as described above. Accordingly, it is conceivable for the computer program product to be executed on the central system and / or on a local computing unit. If the computer program product is executed on the central system, it is conceivable that the alternative method or a program for supporting the alternative method is executed on the central system or is executed on a local computing unit.A computer program product according to a method of an embodiment of the invention executes the steps of a method according to the preceding description when the computer program product runs on a computer, in particular an in-vehicle computer. When the program in question is used on a computer, the computer program product causes an effect, namely the recognition of contents in data, in order to annotate them.CONTENT SPECIFICATION OF THE DRAWINGSThe present invention is explained in more detail below with reference to the exemplary embodiments indicated in the schematic figures of the drawings. The following are shown: FIG. 1 is a schematic block diagram of an embodiment of the invention.The accompanying drawings are intended to provide a further understanding of the embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention. Other embodiments and many of the advantages mentioned are evident with reference to the drawings. The elements of the drawings are not necessarily shown to scale with respect to each other.In the figures of the drawings, identical, functionally identical and identically functioning elements, features and components are provided with the same reference numerals, unless stated otherwise.DESCRIPTION OF EMBODIMENTSFIG. 1 shows a schematic block diagram of a method for annotation data of a data set by means of an artificial neural network, wherein the data set comprises annotated data and unlabeled data.In step S 1, unnotated data that satisfies predetermined similarity criteria regarding similarity to the annotated data is selected. In step S 2, the selected unnotched data that satisfies the similarity criteria is annotated using the artificial neural network. In step S 3, unanswered data that satisfies predetermined uncertainty criteria regarding an annotation based on knowledge of the artificial neural network is selected. In step S 4, the selected unnotated data that meet the uncertainty criteria are annotated using an alternative method or provided using an alternative method for the purposes of annotation.Reference numerals denote reference numeralsS1 -S4 Method steps
Claims
Method for annotation of data of a data set by means of an artificial neural network comprising annotated data and unnotated data, having the following steps: - selecting (S1) unnotated data which meet predetermined similarity criteria relating to a similarity to the annotated data; - annotation (S2) the selected unnotated data which meet the similarity criteria by means of the artificial neural network; - selecting (S3) unnotated data which meet predetermined uncertainty criteria relating to an annotation based on knowledge of the artificial neural network; annotation (S4) of the selected unnotated data that meet the uncertainty criteria using an alternative method and / or providing the selected unnotated data that meet the uncertainty criteria for the purposes of annotation thereof using an alternative method.Method according to claim 1, wherein the following steps are carried out before unnotated data are selected: - determining a first loss of the artificial neural network for the annotated data by means of a first loss function; - determining a second loss of the artificial neural network for the unnotated data by means of a second loss function.The method of any preceding claim, wherein the alternative method comprises presenting data to a human for annotation purposes.The method of any preceding claim, wherein the similarity is determined based on a distance from unnotated data to annotated data.Method according to Claim 4, wherein a similarity criterion comprises dropping below a distance threshold value and / or the similarity criterion comprises a list from those N unnotated data which have the smallest distance from the annotated data.Method according to one of the preceding claims, wherein the uncertainty criterion is based on the result of a decision tree, in particular a random tree.The method of claim 6, wherein the unanswered data is ordered by the decision tree for its uncertainty, and the most uncertain unanswered data is selected.Method according to one of the preceding claims, wherein the first loss function is designed as a triple loss function.Method according to one of the preceding claims, wherein the second loss function is designed as a virtual adventrial loss.The method of any preceding claim, wherein loss of the artificial neural network is minimized by minimizing an overall loss function on the dataset before the first and second losses are determined.The method of any preceding claim, wherein the total loss function is of the form: first loss function + C * second loss function.The method of any preceding claim, wherein the dataset is variable while the method is being performed, and the artificial neural network is configured to determine the first loss for all annotated data of a changed dataset and determine the second loss for unnotated data of a changed dataset once the method is invoked again after a change of the dataset.The method of any preceding claim, wherein the method is repeated until all unnotated data has been annotated.Computer program product for annotation of data by means of an artificial neural network and for preparing annotation of data by means of an alternative method, which is configured to carry out a method according to one of the preceding claims.Central system on which a variable data record is stored which is to be annotated by means of a method according to one of the preceding claims, wherein the central system is connectable by means of one or more data memories which are configured to transmit data to the central system in order to change the data record, wherein a computer program product according to claim 14 is executed on the central system, or the central system is configured to transmit the variable data record to a computing unit which is configured to execute a computer program product according to claim 14 in order to annotation data of the data record by means of an artificial neural network.