WORKFLOW FOR TRAINING A CLASSIFIER FOR QUALITY CONTROL IN METROLOGY
The method optimizes machine learning training dataset annotation by using 2D sectional planes and machine learning models to identify and supplement datasets based on predefined thresholds, addressing time and cost issues while enhancing classification accuracy and adaptability.
Patent Information
- Application Number
- DE102019110721
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2019-04-25
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2039-04-25
AI Technical Summary
Existing manual annotation processes for machine learning training datasets are time-consuming, costly, and prone to errors, particularly in handling 3D data and classifying anomalies, with issues like redundant data, incomplete defect classes, and difficulty in incremental adaptation.
A computer-implemented method and system for expanding training datasets using 2D sectional planes from 3D data, employing machine learning models to identify anomalies and supplement the dataset based on difference values above predefined thresholds, optimizing annotation with targeted data selection and manual intervention.
Reduces annotation costs and time, increases classification accuracy, and allows continuous dataset supplementation, effectively handling difficult anomalies with reduced errors.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the InventionThe invention relates to a workflow for annotations for training datasets and, in particular, to a computer-implemented method for expanding a training dataset for machine learning, and to a corresponding workflow system.Background ArtIndustrial production processes require continuous quality control to continuously meet customer (industrial companies as well as private persons) requirements. Even slight weaknesses with respect to a constantly high quality are not tolerated on the market. In order to meet the requirement for a constantly high quality of the manufactured products, automated quality assurance systems are frequently used. Some of these are based on purely optical methods, while others also rely on a mechanical / mechanical-electronic measurement of product parameters. Another category of quality measurement methods is based on data evaluation of electron microscopes, fluorescence microscopes, light microscopes, optical coherence tomography systems, interferometers, spectrometers or else computer tomography systems.Artificial intelligence techniques, in particular machine learning (machine learning) systems, can be used to assess the measured values. In order for these systems to function as desired, a large amount of training data is typically required to detect anomalies in the workpieces being produced by the machine learning systems then trained. All training data must be annotated manually, which requires a great deal of time for specialists. Annotation is to be understood here as meaning the process of assigning each anomaly to be classified to a desired, predetermined anomaly category. This can be a pore, a core fracture, a wall displacement, etc.This manual procedure is problematic in several respects: the collection of data is very time consuming as is the manual annotation process. With regard to data collection, it is problematic that the frequency distributions of the defect classes are often very different. In order to ensure that there are sufficient examples for training even of less frequently occurring defects, an enormous number of data must generally be recorded. However, this is technically very complicated and expensive. In particular, it is often not possible to store all anomalies.With regard to data annotation, it is also problematic that under certain circumstances very similar or even redundant data are recorded and stored. If these redundant data are annotated without suitable selection, it is hardly possible to expect an improvement in the classification performance, since the redundant annotated data are more or less irrelevant. Particularly within the scope of (feasibility) studies, large amounts of data can frequently be recorded; however, with unfavorable sampling of the data, many similar data points can be collected and annotated unnecessarily in a complicated manner.Often, the set of defect impressions is not complete. Instead, over time, there is often a need to include new defect classes in the existing catalog. Possible reasons are changing parameters of the monitored process or the determination that the original training data set has been chosen to be too small and thus not completely representative of the overall task of quality assurance. In this context, it must therefore be prevented that previously unknown defects during categorization are assigned the class (classification) of a known defect, or that characteristics of a known class are incorrectly classified as a non-defect or defect of another class. In selecting new examples, it should be avoided to select irrelevant outliers for annotation. Furthermore, in the selection and annotation process, the possibility should be given that an operator / vision tester influences the annotation process, for example because the selected example has been produced by a faulty recording and is not assigned to any relevant defect type.Third, it should be avoided that if the selected defects are not annotated by an application specialist, but rather by trained non-specialists, inconsistencies or errors occur in the annotations.Moreover, in practical applications, the incremental adaptation of a classifier after an augmented annotation is often impractical, since application experts want to annotation a plurality of data (e.g. 30-300) on the piece, for example, as well, because the training takes a relatively long time compared to the test time of a test object. Thus, only the extended training of the classifier by several new data points would be practicable. When selecting a plurality of examples for annotation, it is therefore absolutely necessary to ensure that not only is each individual example maximally informative, but also the combination of the examples (i.e. the training data) is not redundant.In addition, annotation of 3-dimensional data is difficult, since the data is difficult to visualize on a 2D display. The classification of the data is also very complicated because of the large amount of data. With a large number of subvolumes to be classified, a correspondingly large computing power is required.In this context, there is already the known document "Active Learning using conformal predictors: application to image classification" by MAKIL, Larazo; SANCHEZ, Jes us A. Vega; DOMIDO-CANTO, Sebastiaan; published in: Fusion Science and Technology, Vol. 62, 2012, pp. 347-355. - issn 1536-1055. This document addresses the problem of active learning, in particular an active selection of training data instead of randomly picked data samples, in order to make a set of training data that is as small as possible usable. Reliable prediction results can thereby be achieved. A 5-class classification problem serves as the basis.Despite the progress already made, the underlying object of the concepts presented here is to elegantly optimize the annotation process for training data sets, to relieve responsible personnel, to enable continuous supplementation of the training data set and to enable an appropriate possibility of reacting to anomalies which can be classified only imprecisely.Overview of the InventionThis object is achieved by the computer-implemented method proposed here for expanding a training dataset for machine learning and a corresponding workflow system for expanding a training dataset for machine learning according to the independent claims. Further embodiments are described by the respective dependent claims.According to a first aspect of the present invention, there is provided a computer-implemented method (100) for expanding a training dataset for machine learning. The method in this case comprises in particular the following: (i) providing a dataset of an object to be examined, wherein the dataset has coordinate values and measured values per coordinate, wherein the dataset is derived from an image recording method of a computer tomography apparatus, (ii) identifying an anomaly in a section of the dataset which corresponds to a partial region of the object to be examined, (iii) classifying the anomaly by means of a first machine learning model trained with a first training dataset into a number of predefined classification classes, (iv) determining a difference value of the anomaly from the trained first machine learning model on the basis of a combination of novelty measures by means of a second machine learning model trained with a second training dataset, wherein the first training dataset and the second training dataset are produced by a respective random selection of datasets from an overall training dataset, and (v) supplementing the first training dataset with data relating to the identified anomaly if the difference value is above a first predefined threshold value.In accordance with an aspect not explicitly claimed, a computer-implemented method for expanding a training dataset for machine learning is also discussed. The method comprises providing a data record of an object to be examined, wherein the data record has coordinate values and measured values per coordinate.The method further comprises identifying an anomaly in a section of the data set that corresponds to a partial region of the object to be examined, and classifying the anomaly into a number of predefined classification classes by means of a first machine learning model trained with a first training data set.In addition, the method comprises determining a difference value of the anomaly from the trained first machine learning model, and supplementing the first training dataset with data relating to the identified anomaly if the difference value is above a first predefined threshold value.According to another aspect of the present invention, a workflow system according to the inventive method is presented.Properties are also discussed in a workflow system that comprises:a receiving unit which is adapted to receive a data record of an object to be examined, wherein the data record has coordinate values and measured values per coordinate,a detection module adapted to identify an anomaly in a portion of the dataset corresponding to a portion of the object to be examined,a classification system adapted to classify the anomaly into a number of predefined classification classes by means of a first machine learning model trained with a first training data set,a difference value determination module adapted to determine a difference value of the anomaly from the trained first machine learning model; anda complementing unit adapted to supplement the first training dataset with data concerning the identified anomaly if the difference value is above a first predefined threshold value.Additionally, embodiments may be implemented in the form of a corresponding computer program product accessible by a computer usable or computer readable medium having program code for use by or in connection with a computer or instruction execution system. For the context of this specification, a computer-usable or computer-readable medium may be any device having means for storing, communicating, forwarding, or transporting the program for use by or in connection with an instruction execution system, apparatus, or apparatus.The computer-implemented method for expanding a training dataset for machine learning has several advantages and technical effects:The proposed concept initially allows a significant reduction in annotation costs and costs. In order to achieve a reduction in the amount of data and thus the complexity, instead of the complete 3D volume, the two-dimensional XY, YZ, XZ sectional planes through the center of the anomaly can be used for annotation and analysis / classification. The relevant sectional planes can be obtained from the probability map of the abnormality detection. As experiments have shown, the slice images contain sufficient information to distinguish fault classes, but are much easier to process in terms of handling than full 3D volume information. Alternatively, further 2D projections of the 3D volume can also be used (for example local minimum / maximum projections).In addition, reduced training times can be realized. Since overall fewer data result in the same information content due to the targeted selection of data, overall less time has to be provided for the classifier training. This is directly reflected in a faster turn around, which also allows interactive switching between training, data selection and annotation, and can thereby result in a faster implementation of solutions.Furthermore, the concept proposed here results in lower classification errors. Since defects that are difficult to detect are preferably selected for annotation, the defect classification can also reliably detect difficult anomalies in a quick manner. The classification accuracy is thus constantly increased, which improves the results of the quality check and thus of the associated products.Furthermore, it is possible to intervene manually in the automated or semi-automated annotation process in order to prevent incorrect annotation parameters in this way and to adapt them accordingly.Overall, the annotation process is therefore elegantly optimized, responsible personnel can be relieved, continuous supplementation of the training dataset is made possible, and appropriate possibilities of reacting to anomalies which can be classified only imprecisely are available.In the following, further embodiments of the inventive concept for the method, which can be used equally for the corresponding workflow system, are presented:According to an embodiment of the method, the anomaly may be a plurality of anomalies and the supplementing may comprise supplementing the training dataset with data of selected identified anomalies if the difference value is above a second predefined threshold value. In this case, the first threshold value can inform an operator that an anomaly cannot be unambiguously assigned to a class, and manual intervention in the annotation process seems to be required.If, however, the difference value determined is above a second difference value, for example a higher difference value, it can be assumed that a significantly different anomaly is present, which can then also be used automatically for annotation of training data sets.In addition, it should be noted that the data about the anomalies may not only consist of one scan data set, but may contain a plurality of data from a plurality of scans that are composed. In this way, the recognition probability can be significantly increased.According to another embodiment of the method, the difference value may be determined by a novelty measure method for a selected anomaly. This can be selected from the group consisting of: Novel Detection, Estimated Error Reduction, Estimated Entropy Minimization, Expected Model Output Changes (EMOC), MC Dropout, OpenMAX - the last two for neural networks - SVM Margin, neural network, ratio of the highest to the second highest classification probability, entropy of classification probabilities, GP variance, variance of individual classification probabilities of aggregated classifiers (e.g. Bagging), variance of the classification probability under perturbation of the input signal, i.e. here of the data set which relates to the object to be examined.Thus, a rich arsenal of methods for the difference value is available from which workflow responsibilities can be used without being subject to artificial restrictions. The skilled person can further supplement the aforementioned methods for detecting a difference value.According to a further possible embodiment of the method, a combination of the novelty measure methods can be determined by a method selected from the group consisting of: multi-armed bandit formulation, success driven selection among multiple criteria, a linear combination based on a Forward function, reinforcement learning. By means of this variant, further, more complex novelty measure methods can be taken into account in order to thus satisfy the combination of the identification of the anomaly and the classification used.According to an additional embodiment, when determining a difference value, the method can also comprise determining the difference value or combination of difference values by means of a second machine learning model trained with a second training data set. This can be, in particular, one of the methods forest regression, linear regression, Gaussian process, neural network or the like. This quasi-safety check allows an even more accurate detection of whether it is actually an additional class of anomalies that would have to be taken into account.An advantageous embodiment of the method can comprise creating the first training dataset and the second training dataset by a respective random selection of datasets from an overall training dataset, e.g. by means of bridging. In this way, too, a tendency to take into account or not taking into account specific anomalies can be limited and the method can thus be made more robust with respect to misinterpretations.A further advantageous embodiment of the method can also comprise retraining the first machine learning model with the supplemented training data set to generate a third machine learning model. In this case, the parameters of the first machine learning model can be used as starting values for the training. Alternatively, a complete restart of the machine learning model can also take place in order to create the machine learning model without a tendency.According to a specific embodiment of the method, the classification can be carried out by means of a classifier selected from the group consisting of: neural network-also deep neural network-random forest, logistics regression (somewhat atypical), support vector machine (SVM) and Gaussian regression. In principle, all known classifiers can be used, of which the person skilled in the art selects a suitable classifier or a classifier system depending on anomalies and objects to be examined.According to further useful embodiments of the method, the data record can be derived from an image recording method. The image recording method can be based on the following: an electron microscope-in particular other charged particles can also be used-a fluorescence microscope, an light microscope, an optical coherence tomography apparatus, an interferometer, a spectrometer, an operating microscope and a computer tomography apparatus. The skilled person could carry out the method using comparable image recording methods.The method according to a practicable embodiment could also comprise receiving - for example from a vision tester - a selection signal for an anomaly, wherein the selection signal increases the difference value such that it is above the first predefined threshold value or above the second predefined novelty threshold value. In this way, based on overwriting the determined difference value, an "uncertain candidate" could be changed to a "safe candidate" and thus the annotation of experts could be supported.Overview of the FiguresAspects already described above and additional aspects of the present invention are evident, inter alia, from the described exemplary embodiments and from the additional further concrete embodiments described by reference to the figures.Preferred embodiments of the present invention will be described by way of example and with reference to the following figures: FIG. 1 illustrates a block diagram of an embodiment of the inventive computer-implemented method for expanding a training dataset for machine learning. FIG. 2 illustrates a block diagram for a sequence of a usual training sequence of a classifier. FIG. 3 illustrates a block diagram of a more implementation-approximate flow chart of the proposed method. FIG. 4 illustrates a block diagram of an embodiment of the workflow system. FIG. 5 illustrates a block diagram of a computer system that additionally includes the workflow system.Detailed description of the FiguresIn the context of this specification, conventions, terms and / or terms are intended to be understood as follows:The term machine learning (or machine learning) is a basic term or function of artificial intelligence, wherein statistical methods are used, for example, to give computer systems the capability of learning. For example, certain behavior patterns are optimized within a specific task area. The methods used enable machine learning systems to analyze data without requiring explicit procedural programming. Typically, e.g. CNN (Convolutional Neural Network) is an example of a machine learning system, a network of nodes acting as artificial neurons, and artificial connections between the artificial neurons - so-called links - wherein parameters - e.g. weight parameters for the connection - can be assigned to the artificial connections. During neural network training, the weight parameters of the connections automatically adapt to generate a desired result based on input signals. During supervised learning, the images supplied as input values (training data) are supplemented-generally (input) data-by metadata (annotations) in order to display a desired output value. In unsupervised learning, such annotations are not required. Generally speaking, a mapping of input data to output data is learned.In this context, recursive neural networks (RNNs) are also to be mentioned, which also represent a type of deep neural networks in which the weight adaptation is effected recursively, so that a structured prediction can be generated via input data of variable size. Typically, such RNNs are used for sequential input data. In this case, just as in the case of CNNs, back-propagation functions are also used in addition to the forward-directed weight adaptation. RNNs can also be used in image analysis.In the context of this document, the term "training data set" essentially describes image data, which can be obtained by the aforementioned methods, of one or more anomalies in an object to be examined (e.g. workpiece), for which annotations are already present.The term "expanding a training dataset" describes the process of expanding an existing training dataset with new image data with corresponding classifications, which can then be usable for a renewed or expanded training of the classifier.The term "object to be examined" generally describes here a workpiece which is to be subjected to a quality check.The term "anomaly" basically describes a non-agitated abnormality in the volume-or also in a surface-of the object to be examined. This can be, for example, a pore, an inclusion, a core fracture, a wall displacement or other salient features in or on the workpiece.The term "classifying" describes the process of associating a detected anomaly with an anomaly class by the classifier operating on the principles of machine learning.The term "classifier"-also referred to in the context of machine learning as a classifier or classifier system-describes a machine learning system which is enabled by training with training data to assign input data-here in particular image data of anomalies of the object to be examined-to a specific class (in particular of anomalies).It should also be noted that a classifier typically classifies into a predefined number of classes. This is normally done by determining a classification value of the input data for each class and selecting a WTA (inner takes it all) filter to the class having the highest classification value as the classified class. In classifiers, the deviation from a 100% classification value is frequently used as a quality parameter of the classification or as a probability for the correctness of the classification.In the context of this document, the term "difference value" describes a value which is determined by the method presented here or the corresponding workflow system, which specifies how great the deviation of a recognized, i.e. identified, anomaly from a 100% classification is. The highest classification value (see above) can be used as an indicator; however, the other classification values of the other classes can also be additionally used. In the simplest case, the classification probability parameter of the classifier will thus be able to be used in a simple manner. This difference value is used in the decision as to whether or not the identified anomaly is usefully assigned to a specific, already known class-and thus the corresponding pixel data are subjected to a corresponding auto annotation for the purpose of supplementing the training dataset. These annotation values of the pixels are then used to supplement the data set consisting of coordinate values and measured values by the parameter "annotation value".If a unique classification is not possible, a specialist can release a new anomaly class or specify the classifier to be trained for further or new training rounds.The term "training the classifier" means that, for example, a machine learning system (machine learning system) is adjusted by a plurality of example images parameters in a neural network, for example, by partially repeatedly evaluating the example images in such a way that, after the training phase, unknown images are also associated with one or more categories with which the learning system has been trained. The example images are typically annotated-i.e. provided with metadata-in order to generate desired results based on the input images, such as, for example, indications about maintenance measures to be carried out.The term "convolutional neural network", as an example of a classifier, describes a class of artificial neural networks based on feed-forward techniques. They are frequently used for image analyses with images as input data. The main constituent of convolutional neural networks are convolutional layers (therefore the name), which enables efficient evaluation by parameter sharing.It should also be mentioned that deep neural networks consist of several layers of different functions-for example an input layer, an output layer and one or more intermediate layers for example for convolution operations, application of non-linear functions, dimension reduction, normalization functions, etc.:. The functions can be "executed in software" or special hardware modules can take over the calculation of the respective function values. Combinations of hardware and software elements are also known.The term "novel detection" describes a mechanism by which an "intelligent" system (for example an intelligent organism) can classify an incoming sensor pattern as unknown up to now. This principle can also be applied to artificial neural networks-or other classifiers. If a sensor pattern arriving at the artificial neural network does not generate an output signal in which the recognition probability is above a predefined (or dynamically adapted) threshold value, or generates a plurality of output signals in which the recognition probabilities are approximately the same size, the arriving sensor pattern-for example a digital image-can be classified as one with a new content (Novelty).If, on the other hand, a new example image is supplied to a neural network trained by known example images (configured as an autoencoder, for example) and passes through the autoencoder, then the latter should be able to reconstruct the new example image again at the output. This is possible because, when passing through the autoencoder, the new example image is strongly compressed in order to subsequently expand / reconstruct it again by the neural network. If the new input image (largely) matches the expanded / reconstructed image, the new example image corresponded to a known pattern or has a high similarity to the contents of the training image database. If a distinct difference is found in the comparison of the new example image with respect to the expanded / reconstructed image, i.e. if the reconstruction errors are significant, then the example image is one with a previously unknown distribution (image content). That is, the autoencoder cannot map the data distribution, it is an anomaly compared to the known database assumed to be normal.A detailed description of the figures is given below. It is understood that all details and instructions are schematically depicted in the figures. First, a block diagram of an embodiment of the computer-implemented method for expanding a training dataset for machine learning according to the invention is presented. Further exemplary embodiments or exemplary embodiments for the corresponding system are described below:FIG. 1 illustrates a block diagram of an embodiment of the inventive computer-implemented method 100 for expanding a training dataset for machine learning. The method 100 comprises providing 102 a data record of an object to be examined. The data record has coordinate values and measured values for each coordinate. This can be a 2D (2-dimensional) or 3D data set which has either gray values per pixel (generally also for voxels) as measured values or also color values. In addition, there may be further measured values per pixel or for groups of pixels. The data of the data set were recorded by the above-mentioned methods and optionally preprocessed, e.g. normalized.The method 100 further comprises identifying 104 an anomaly in a section of the data set-also referred to in the literature as "region of interest"-which corresponds to a partial region of the object to be examined. In one variant, therefore, this is a section; 3D image data is a partial volume.The method 100 also comprises classifying 106 the identified anomaly into a number of predefined classification classes by means of a first machine learning model trained with a first training data set. Anomalies in the object to be examined can thus be classified. The classifier used for the classification was trained with the first training data set before use-i.e. before the anomaly classification-in order to form the first machine learning model typically within the classifier.Furthermore, the method 100 comprises determining 108 a difference value of the anomaly with respect to the trained first machine learning model. The difference value is not necessarily the classification probability value of the classifier, but can be derived from it-i.e. a function thereof.In addition, the method 100 comprises supplementing 110 the first training dataset with data relating to the identified anomaly if the difference value is above a first predefined threshold value. The threshold value for the difference value can be set here on individually predefined quality standards, production processes and anomalies to be expected. Annotation is typically per pixel.FIG. 2 illustrates a block diagram for a sequence 200 of a usual training sequence of a classifier. In this case, first of all, a determination 202 of potential faults, i.e. anomalies, follows. These are annotated 204 according to defect classes. In this case, in particular those pixels which belong to the respective anomaly are provided manually and thus time-consumingly with a specific annotation flag for the respective anomaly class. This process is interactive by a user. In this way, the training data set is created. As is known, a large number of example anomalies are required for this purpose.In a next step, the classifier is trained 206 with a completed training data set. The trained classifier is then used 208 in a testing process of a quality assurance system. This procedure known from the prior art is time-consuming and requires special know-how in the assessment of the anomaly and thus the creation of the training data with corresponding annotations.The method presented here can use this procedure for creating a basic training dataset in order to have a basic stock of training data-in particular of the first training dataset-available.FIG. 3 illustrates a block diagram of a more implementation-approximate flow chart 300 of the proposed method. According to the workflow of the presented method, here too, first a determination 302 of potential errors-i.e. anomalies-is carried out. The corresponding data are typically manually annotated 304 according to defect classes in order in particular to form the first training data set for training the first machine learning model, in order then to be trained 306 as a result of an increasing training data set. This classifier trained in this way can furthermore be used in a normal test process 308, 308.However, if it is found during the ongoing test process that a difference value-based on the probability of the correctness of the classification or the respective quality parameters of the classification-of a classification is too large compared to a first predefined threshold value (i.e. simply greater than the threshold value), then a supplement to the training dataset can be made, in which a new anomaly class is then taken into account. At the same time, the pixels associated with the identified anomaly that cannot be classified unambiguously (i.e., difference value less than the predefined threshold value) can be automatically annotated with a new annotation flag associated with a new class of anomalies.Retraining the classifier with the extended training data set then also comprises learning the additional classification class, i.e. that the machine learning model of the classifier is also trained by the supplemented training data for the additional classification class.As already mentioned above, the determination 310 of the safety measure, i.e. the threshold value, is carried out in accordance with predefined parameters of the input data used (typically image data), of the material from which the object to be examined consists, and / or expected anomalies or defects. The additional annotation 312 by means of a security measure of selected additional data and the retraining 306 with the now increasing data volume have also already been explained above.The measure can then be used differently for the sequence in the test process: (1) In order to make the data collection more targeted, only data on partial volumes are stored, for which a high value of newly obtained information is estimated. This saves storage space and a large number of relevant subvolumes can be stored in a memory of predetermined size. Only those partial volumes are then taken into account for the annotations which are particularly informative, whereby an improvement of the classifier in the subsequent training is expected.A high estimated information content, i.e. a high difference value (for example due to a high uncertainty with respect to the classification), simultaneously signals that the associated subvolume can contain a defect that was unknown to date. Therefore, it should be taken into account in an extension of the training dataset. If the fault image appears clustered, the set of fault categories in the fault catalog can be extended by the new class. That is, a new class does not need to be introduced for each first time abnormality having a difference value larger than the threshold value occurs. In this case, model parameters of an already trained model can be used to initialize the new model. This can significantly reduce the training time and the required number of additional training data.When inspecting new data, the measure can be evaluated immediately for each localized effect in order to decide directly in the checking process whether it should be presented to an expert for follow-up checking (stream-based setting). Alternatively, a sufficiently large amount of localized defects (i.e. anomalies) can first be collected, from which a selection is subsequently made for annotation (pool-based setting) by evaluating the information measure of the most relevant examples.In order to prevent the selection of non-relevant data points, the information measure can additionally be supplemented by a factor which is used, for example, to estimate whether visually similar examples occur in the collection of new data (a novel cluster) or whether it is an outlier (for example due to recording errors which have occurred in the short term and lead to an image appearing completely differently).To avoid increased rejection by the vision tester, the information measure should adapt to estimate similar examples later with a lower information value.The effects of erroneous annotations by non-application experts can be reduced most easily by annotation the selected examples of multiple persons (multiple annotation). Alternatively, the error of a vision tester can already be estimated in the information measure, so that different vision testers are presented with different effects for annotation. This can be done independently of the reliability which is assigned to the respective vision tester.To ensure that a combination of several examples is also informative, it is possible to add a diversity measure for selection. This can either make a direct assessment of the similarity of the examples or be achieved by assessment based on the variance of the resulting classifier.For completeness, FIG. 4 illustrates a block diagram of an embodiment of the workflow system 400. The workflow system 400 comprises a receiving unit 402, which is adapted to receive a data set of an object to be examined. In this case, the data record has coordinate values and measured values per coordinate.The workflow system 400 also has a recognition module 404 which is adapted to identify an anomaly in a section of the dataset which corresponds to a partial region of the object to be examined, and a classification system 406 which is adapted to classify the anomaly into a number of predefined classification classes by means of a first machine learning model trained with a first training dataset.A difference value determination module 408 as part of the workflow system 400 is adapted to determine a difference value of the anomaly from the trained first machine learning model. Finally, the workflow system 400 also has a supplementing unit 410 which is adapted to supplement the training dataset with data relating to the identified anomaly if the difference value is above a first predefined threshold value.This allows both an autonotation and an expansion of the training dataset and the introduction of potentially new anomaly classes to be effected elegantly.In summary, it can thus be stated that, in order to realize the solution approach presented here, measures for the information content can be used which are known from the area active learning. This fundamentally comprises two classes of methods: explorative methods and explorative methods.Exploratory methods ignore the model of an already trained classifier and instead attempt to cover the entire feature space as well as possible with as few data points to be selected, i.e. to generate a representative, high-variance data collection as quickly as possible. High information is thus assumed in the case of those new data which are maximum and similar to already known data. For the problem underlying here, these methods would therefore select defects for annotation which have not yet been known in terms of their visual form. The advantages of these methods lie in the possibility of rapidly selecting defect categories (anomaly classes) for annotation which have hitherto been unknown. There are disadvantages with regard to susceptibility to outliers, caused for example by disturbing processes during data recording. Directly applicable are measures from the field of "Novel Detection", which explicitly address the finding of new, unknown data.With regard to exploitive methods, it must be assumed in the task addressed here that a classifier already achieves a performance relevant in practice (warm start) and the machine learning model does not have to be learned from zero (cold start). Therefore, information measures which assume an already trained classifier can be used successfully.Exploitative methods include the previously trained classifier in the evaluation of the information content of new examples, i.e. they "exploit" it. The aim is then to achieve a rapid reduction in the classification error. Depending on the quantity of available training data and the available computing time, the expected error reduction can either be evaluated directly (e.g. estimated error reduction or estimated entropy minimization), be approximated slightly (e.g. by EMOC), or be replaced by heuristically instantiated methods.In heuristics, the uncretainy-based sampling can be used to particular advantage, which assumes a high information content for that example, in which the current classifier is maximally uncertain (e.g. SVM margin, k-nearest neighbor, 1-vs-2 and entropy, GP mean).A selection of examples which will have a strong influence on the previous classifier during the new training is likewise well usable, which is known as ExpectedModelChange (e.g. for SVMs or for random forests).In the present application, an application of uncretainty sampling is particularly relevant, since the measures can be calculated quickly and thus permit a delay-free interaction of data recording system and data annotation experts.The information measures MC dropout and OpenMAX are particularly well suited for use with neural networks or deep learning.FIG. 5 illustrates a block diagram of a computer system that may include at least portions of the maintenance monitoring system. Embodiments of the concept proposed here can be used in principle with virtually any type of computer, independently of the platform used therein for storing and / or executing program codes. FIG. 5 illustrates, by way of example, a computer system 500 which is suitable for executing program code in accordance with the method presented here. A computer system already present in the microscope system can also serve--possibly with corresponding extensions--as a computer system for executing the concept presented here.The computer system 500 has a plurality of generally purpose functions (general purpose functions). The computer system can be a tablet computer, a laptop / notebook computer, another portable or mobile electronic device, a microprocessor system, a microprocessor-based system, a smartphone or computer system with specially configured special functions. The computer system 500 may be configured to execute computer system executable instructions, such as program modules, that may be executed to implement functions of the concepts proposed herein. To do so, the program modules may include routines, programs, objects, components, logic, data structures, etc., to implement particular tasks or particular abstract data types.The components of the computer system may include one or more processors or processing units 502, a memory system 504, and a bus system 506 that connects various system components including the memory system 504 to the processor 502. Typically, the computer system 500 includes a plurality of volatile or non-volatile storage media accessible by the computer system 500. In the storage system 504, the data and / or instructions (instructions) of the storage media may be stored in volatile form, such as in random access memory (RAM) 508, for execution by the processor 502. These data and instructions implement individual or multiple functions or steps of the concept presented here. Other components of the storage system 504 may be a permanent memory (ROM) 510 and a long-term memory 512 in which the program modules and data (reference numeral 516) may be stored.The computer system has a number of dedicated devices (keyboard 518, mouse / pointing device (not shown), display 520, etc.) for communication. These dedicated devices may also be combined into a touch-sensitive display. A separately provided I / O controller 514 provides smooth data exchange to external devices. A network adapter 522 is available for communication over a local or global network (LAN, WAN, for example, over the Internet). The network adapter may be accessed by other components of computer system 500 via bus system 506. It is understood that although not shown, other devices may be connected to computer system 500.In addition, at least parts of the workflow system 400 (cf. FIG. 4 ) can be connected to the bus system 506. The digital image data of an image sensor (not shown) can also be processed by a separate preprocessing system (not shown). This can support the provision of the data record of an object to be examined.The description of the various exemplary embodiments of the present invention has been presented for better understanding, but does not serve to directly restrict the inventive concept to these exemplary embodiments. Further modifications and variations will occur to those skilled in the art. The terminology used herein has been chosen to best describe the basic principles of the embodiments and to make them readily apparent to those skilled in the art.The principle presented here can be embodied both as a system, as a method, combinations thereof and / or as a computer program product. In this case, the computer program product may comprise one (or more) computer-readable / s storage medium / media comprising computer-readable program instructions for causing a processor or a control system to execute various aspects of the present invention.Media used are electronic, magnetic, optical, electromagnetic, infrared media or semiconductor systems as forwarding medium; for example SSDs (solid state device / drive as solid state memory), RAM (random access memory) and / or ROM (read-only memory), EEPROM (electrically erasable ROM) or any combination thereof. Also suitable as transmission media are propagating electromagnetic waves, electromagnetic waves in waveguides or other transmission media (e.g. light pulses in optical cables) or electrical signals which are transmitted in wires.The computer readable storage medium may be an embodying device that stores instructions for use by an instruction execution device. The computer-readable program instructions described here can also be downloaded to a corresponding computer system, for example as a (smartphone) app from a service provider via a cable-based connection or a mobile radio network.The computer readable program instructions for carrying out operations of the invention described herein may be machine dependent or machine independent instructions, microcode, firmware, status defining data, or any source code or object code written, for example, in C++, Java, or the like, or in conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may be executed completely by a computer system. In some embodiments, it may also be electronic circuits such as programmable logic circuits, field programmable gate arrays (FPGA), or programmable logic arrays (PLA), which execute the computer readable program instructions by utilizing status information of the computer readable program instructions to configure or individualize the electronic circuits according to aspects of the present invention.Moreover, the invention presented herein is presented with reference to flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the invention. It should be noted that practically any block of the flowcharts and / or block diagrams may be configured as computer readable program instructions.The computer readable program instructions may be provided to a general purpose computer, special purpose computer, or other programmable data processing system to produce a machine, such that the instructions executed by the processor or computer or other programmable data processing apparatus generate means for implementing the functions or acts illustrated in the flowchart and / or block diagrams. These computer-readable program instructions can accordingly also be stored on a computer-readable storage medium.In this sense, each block in the illustrated flowchart or block diagrams may represent a module, segment, or portion of instructions that represents multiple executable instructions for implementing the specific logic function. In some exemplary embodiments, the functions which are represented in the individual blocks can be executed in a different order-possibly also in parallel.REFERENCE NUMERALS100 Method 102 Method step 100 104 Method step 100 106 Method step 100 108 Method step 100 110 Method step 100 200 Method for annotating training data and using a classification 202,... 208 method steps to 200 300 more detailed methods compared to method 100 302,... 312 method steps to 300 400 workflow system 402 receiving unit 404 recognition module 406 classification system, classifier 408 difference value determination module 410 supplementing unit 500 computer system 502 processor 504 storage system 506 bus system 508 RAM 510 ROM 512 long-term memory 514 I / O controller 516 program modules or and data 518 keyboard 520 screen 522 network adapter
Claims
A computer-implemented method (100) for expanding a training dataset for machine learning, the method comprising: - providing a dataset of an object to be examined, wherein the dataset comprises coordinate values and measured values per coordinate, wherein the dataset is derived from an image recording method of a computer tomography apparatus, - identifying an anomaly in a section of the dataset corresponding to a partial region of the object to be examined, - classifying the anomaly into a number of predefined classification classes by means of a first machine learning model trained with a first training dataset, determining a difference value of the anomaly from the trained first machine learning model on the basis of a combination of novelty measures by means of a second machine learning model trained with a second training data set, wherein the first training data set and the second training data set are created by a respective random selection of data sets from an overall training data set, and supplementing the first training data set with data relating to the identified anomaly if the difference value is above a first predefined threshold value.The method of claim 1, wherein the anomaly is a plurality of anomalies, and the supplementing comprises: supplementing the training dataset with data from selected identified anomalies if the difference value is above a second predefined threshold.The method of claim 1 or 2, wherein the difference value is determined by a novelty measure method for a selected anomaly selected from the group consisting of: Novel Detection, Estimated Error Reduction, Estimated Entropy Minimization, Expected Model Output Changes (EMOC), MC Dropout, OpenMAX, SVM Margin, Neural Network, ratio of highest to second highest classification probability, entropy of classification probability, GP variance, variance of individual classification probabilities of aggregated classifiers, variance of the classification probability with perturbation of the input signal.The method of claim 3, wherein a combination of the novelty measure methods is determined by a method selected from the group consisting of: multi-armed bandit formulation, success driven selection among multiple criteria, a linear combination based on a Forward function, reinforcement learning.The method according to one of the preceding claims, also comprising - retraining the first machine learning model with the supplemented training data set to generate a third machine learning model, wherein the parameters of the first machine learning model are used as starting values.The method according to any one of the preceding claims, wherein the classifying is performed by means of a classifier selected from the group consisting of: neural network, random forest, logistics regression, support vector machine, Gaussian regression.The method according to any of the preceding claims, also comprising - receiving a selection signal for an anomaly, wherein the selection signal increases the novelty measure to be above the first predefined novelty threshold or above the second predefined novelty threshold.A workflow system (400) for expanding a training dataset for machine learning, the workflow system (400) comprising - a receiving unit (402) adapted to receive a dataset of an object to be examined, wherein the dataset comprises coordinate values and measured values per coordinate, wherein the dataset is derived from an image recording method of a computer tomography apparatus, - a recognition module (404) adapted to identify an anomaly in a section of the dataset corresponding to a partial region of the object to be examined, - a classification system (406) adapted to classify the anomaly into a number of predefined classification classes by means of a first machine learning model trained with a first training dataset, a difference value determination module (408) adapted to determine a difference value of the anomaly compared to the trained first machine learning model on the basis of a combination of novelty measures by means of a second machine learning model trained with a second training data set, wherein the first training data set and the second training data set are created by a respective random selection of data sets from an overall training data set, and a supplement unit (410) adapted to supplement the first training data set with data relating to the identified anomaly if the difference value is above a first predefined threshold value.