System and method for synthesizing data quality metrics for anomaly assessment

By combining anomaly detectors and data density estimators to evaluate the quality scores of synthetic anomalies, the problem of inability to effectively evaluate the quality of synthetic anomalies in existing technologies is solved, and the performance and robustness of the anomaly detection model are improved.

CN120611294APending Publication Date: 2025-09-09ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510260655.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2025-03-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing data quality metrics cannot effectively evaluate the quality of synthetic anomalies, resulting in poor performance of machine learning models in anomaly detection, especially in cases with limited or biased anomaly sets.

Method used

By leveraging the anomaly detector and data density estimator of the machine learning model, combined with class conditional probability and rarity score functions, we calculate the quality score of the synthetic dataset and evaluate the quality of the synthetic anomalies.

Benefits of technology

Improves the performance of machine learning models in anomaly detection, can effectively incorporate high-quality synthetic anomalies to improve the training set, and improve the model's predictive accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611294A_ABST
    Figure CN120611294A_ABST
Patent Text Reader

Abstract

The invention relates to a system and method for synthesizing data quality metrics for anomaly assessment. A method for machine learning model tagging data, where the method comprises: generating one or more synthetic datasets, the one or more synthetic datasets comprising one or more anomalies and one or more tags associated with a training dataset; determining, with the anomaly detector, an estimate associated with a class conditional probability of the one or more synthetic data sets in response to the mapping anomaly score assigned by the anomaly detector; determining, using a data density estimator, a data density associated with the one or more synthetic data sets in response to a rare score or kernel density estimation function associated with the training data set; determining a quality score associated with the one or more synthetic data sets using at least the data density and the estimate associated with the class conditional probability; and outputting a quality score associated with the anomalies of the one or more synthetic data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to machine learning networks, including machine learning networks utilized to create synthetic data or datasets. Background Art

[0002] Traditional data quality metrics evaluate an example by quantifying its impact on the performance of the model when it is included for training. For example, several methods such as DATASHAP compute the marginal contribution of an example (x,y)∈D by bootstrapping D and measuring the impact of using (x,y) when (x,y) is included in the bootstrapping training set. Alternatively, DATAOOB is an out-of-bag (OOB) estimator that measures the contribution of each example to the out-of-bag accuracy, i.e., the prediction accuracy of the unselected examples in the bootstrapping. Finally, methods such as LAVAEVALUATOR and influence function (INF) quantify the rate at which the utility value changes when a particular example is weighted more. Unfortunately, all existing data quality metrics aim to evaluate the impact of training examples, while none are specifically designed to estimate the quality of external anomalies.

[0003] Anomaly detection aims to identify examples that behave unexpectedly / unusually. Because anomalies are rare and unexpected, collecting realistic anomaly examples is often challenging in several applications. Furthermore, learning anomaly detectors with limited (or no) anomalies often leads to poor prediction performance. One option is to employ auxiliary synthetic anomalies to improve model training. However, synthetic anomalies can be of poor quality: anomalies that are unrealistic or difficult to distinguish from normal samples can deteriorate detector performance. Unfortunately, no existing metrics quantify the quality of auxiliary anomalies. Summary of the Invention

[0004] The first embodiment discloses a method for labeling data for a machine learning model, wherein the method includes: generating one or more synthetic data sets using a training data set, wherein the one or more synthetic data sets include one or more anomalies and one or more labels associated with the training data set; determining, using an anomaly detector of the machine learning model, an estimate associated with a class conditional probability of the one or more synthetic data sets in response to a mapped anomaly score assigned by the anomaly detector; determining, using a data density estimator, a data density associated with the one or more synthetic data sets in response to a rarity score function or a kernel density estimation function associated with the training data set; determining a quality score associated with the one or more synthetic data sets using at least the data density and the estimate associated with the class conditional probability; and outputting the quality score associated with the anomaly of the one or more synthetic data sets.

[0005] A second embodiment discloses a system configured to label data using a machine learning (ML) model, the system comprising an input interface configured to receive input data from a sensor, wherein the sensor comprises a video, radar, lidar, sound, sonar, ultrasonic, motion, or thermal imaging sensor. The system comprises a processor in communication with the input interface, wherein the processor is programmed to: generate one or more synthetic datasets using a training dataset, wherein the one or more synthetic datasets include one or more anomalies associated with the training dataset and one or more labels; determine, using an anomaly detector of the ML model, an estimate associated with a class conditional probability of the one or more synthetic datasets in response to a mapped anomaly score assigned by the anomaly detector; determine, using a data density estimator, a data density associated with the one or more synthetic datasets in response to a rarity score function or a kernel density estimation function associated with the training dataset; determine a quality score associated with the one or more synthetic datasets using at least the data density and the estimate associated with the class conditional probability; and output a quality score associated with one of the one or more anomalies of the one or more synthetic datasets.

[0006] A third embodiment discloses a computer program product storing instructions, which, when executed by a computer, causes the computer to: generate one or more synthetic data sets using a training data set, wherein the one or more synthetic data sets include one or more anomalies associated with the training data set; determine, using an anomaly detector of an ML model, an estimate associated with a class conditional probability of the one or more synthetic data sets in response to a mapped anomaly score assigned by the anomaly detector; determine, using a data density estimator, a data density associated with the one or more synthetic data sets in response to a rarity score function or a kernel density estimation function associated with the training data set; determine a quality score associated with the one or more synthetic data sets using at least the data density and the estimate associated with the class conditional probability; and output a quality score associated with one of the one or more anomalies of the one or more synthetic data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 A system for training a neural network according to an embodiment is shown.

[0008] Figure 2 A computer-implemented method for training and utilizing a neural network according to an embodiment is shown.

[0009] Figure 3 A flow chart associated with a method of evaluating anomalies according to one embodiment is shown.

[0010] Figure 4 Shown are graphs associated with anomalies for various methods of evaluating anomalies.

[0011] Figure 5 Depicted is a schematic diagram of the interaction between a computer-controlled machine and a control system according to an embodiment.

[0012] Figure 6 Depicted is a diagram according to an embodiment of the present invention. Figure 5 Schematic diagram of a control system configured to control a vehicle, which can be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot.

[0013] Figure 7 Depicts Figure 5 Schematic diagram of a control system configured to control a manufacturing machine, such as a punch, cutter, or gun drill, of a manufacturing system, such as part of a production line.

[0014] Figure 8 Depicts Figure 5 Schematic diagram of a control system configured to control a power tool, such as a drill or driver, having an at least partially autonomous mode.

[0015] Figure 9 Depicts Figure 5 Schematic diagram of a control system configured to control an automated personal assistant.

[0016] Figure 10 Depicts Figure 5 Schematic diagram of a control system configured to control a monitoring system, such as a controlled access system or a surveillance system.

[0017] Figure 11 Depicts Figure 5 Schematic diagram of a control system configured to control an imaging system, such as an MRI apparatus, an X-ray imaging apparatus, or an ultrasound apparatus. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure are described herein. However, it will be understood that the disclosed embodiments are merely examples, and other embodiments may take various forms and alternative forms. The figures are not necessarily drawn to scale; some features may be enlarged or minimized to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be interpreted as limiting, but merely as a representative basis for teaching those skilled in the art to adopt the embodiments in various ways. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any one of the figures may be combined with the features illustrated in one or more other figures to produce embodiments that are not explicitly illustrated or described. The combination of the illustrated features provides representative embodiments for typical applications. However, for specific applications or implementations, various combinations and modifications of features consistent with the teachings of the present disclosure may be desired.

[0019] Unless the context clearly dictates otherwise, as used herein, "a," "an," and "the" refer to both singular and plural references. By way of example, "a processor" programmed to perform various functions refers to one processor programmed to perform each function, or a plurality of processors collectively programmed to perform each of the various functions.

[0020] The systems and embodiments disclosed below disclose a novel metric that estimates the quality of synthetic anomalies by capturing the uncertainty of anomaly detectors. Intuitively, if an anomaly detector has high uncertainty in its predictions, the example is likely too close to (e.g., overlapping) a normal training example, or is located in an unexplored region of the example space (e.g., out of distribution). Specifically, an auxiliary anomaly that falls into a normal region may indicate that it is indistinguishable from its normal counterpart, while synthetic anomalies that fall into an unexplored region are likely unrealistic to appear in practice.

[0021] Anomalies are often associated with adverse events, such as production line defects, excessive water use, oil extraction failures, or wind turbine malfunctions. Collecting examples of anomalies is often a difficult task because anomalies are caused by rare, infrequent, and unexpected events. For example, in an industrial production line, a defect may only appear after months of successful production. Furthermore, the limited number of anomalies observed may be a biased sample, meaning it does not fully represent all potential anomalies. For example, due to changes in the production line, defects may initially be similar to each other (e.g., they are caused by the same event) and evolve over time.

[0022] Therefore, traditionally, machine learning models for anomaly detection suffer from limited access and biased anomalies. One option is to collect auxiliary synthetic anomalies (e.g., by leveraging generative models) and employ them to enhance the dataset. However, synthetic anomalies can be of poor quality: images of defective products may have imperceptible defects and therefore be difficult to distinguish from their normal counterparts, or conversely, the defects may be too severe to be practical. Including poor-quality anomalies can have the effect of worsening the model's performance. Unfortunately, existing methods cannot quantify the quality of synthetic anomalies, so that only high-quality examples are enriched in the training set.

[0023] In the present disclosure, the system and method can measure the quality of the synthesized auxiliary anomalies using a per-example metric. According to the embodiments disclosed herein, the performance of the model can be improved by incorporating high-quality auxiliary anomalies into the training set.

[0024] Existing data quality metrics may be targeted at evaluating training samples, but may not be specifically designed for estimating the quality of auxiliary anomalies. The methods proposed in the disclosed embodiments improve upon existing work by addressing many challenges.

[0025] Small and potentially biased anomaly sets make evaluating the performance of ML-based anomaly detectors unreliable. As a result, variations in the utility of a sample may not be indicative of its contribution to test performance. To address this issue, the system and method can consider both model predictions and data density, making data evaluation metrics more robust to unreliable test sets.

[0026] While existing data quality metrics retrain the model for each training example, this can be prohibitively expensive for external anomaly sets, as researchers may want to evaluate a large number of auxiliary anomaly sets. Consequently, the number of training steps does not scale with the size of the external set. The proposed method does not require retraining and is significantly more efficient than existing work.

[0027] The present invention can be used to evaluate the quality of synthetic anomalies or auxiliary anomalies in anomaly detection. Based on the data evaluation metric, high-quality auxiliary anomalies can be included in the training set to improve the training of the anomaly detection model, or included in the validation set to help hyperparameter tuning and model selection.

[0028] Reference is now made to the embodiments illustrated in the figures, which may apply these teachings to machine learning models or neural networks. Figure 1 A system 100 for training a neural network (e.g., a deep neural network) is shown. The system 100 may include an input interface for accessing training data 102 for the neural network. For example, Figure 1 As shown, the input interface can be constituted by a data storage interface 104, which can access the training data 102 from a data storage device 106. For example, the data storage interface 104 can be a memory interface or a persistent storage interface, such as a hard disk or SSD interface, or a personal area network, local area network, or wide area network interface, such as a Bluetooth, Zigbee, or Wi-Fi interface, or an Ethernet or fiber optic interface. The data storage device 106 can be an internal data storage device of the system 100 (such as a hard disk or SSD), but can also be an external data storage device, such as a network-accessible data storage device.

[0029] In some embodiments, data storage 106 may also include a data representation 108 of an untrained version of the neural network, which may be accessed by system 100 from data storage 106. However, it will be appreciated that training data 102 and data representation 108 of the untrained neural network may each be accessed from different data storages (e.g., via different subsystems of data storage interface 104). Each subsystem may be of the type described above with respect to data storage interface 104. In other embodiments, data representation 108 of the untrained neural network may be generated internally by system 100 based on design parameters for the neural network and, therefore, may not be explicitly stored on data storage 106. System 100 may also include a processor subsystem 110 that may be configured to provide an iterative function to replace a layer stack of the neural network to be trained during operation of system 100. Here, the corresponding layers of the replaced layer stack may have mutually shared weights and may receive as input the output of the previous layer, or, for the first layer of the layer stack, receive the initial activation and a portion of the input of the layer stack. Processor subsystem 110 may also be configured to iteratively train the neural network using training data 102. Here, the training iterations by the processor subsystem 110 may include a forward propagation portion and a backward propagation portion. The processor subsystem 110 may be configured to perform the forward propagation portion by determining an equilibrium point of the iterative function (at which equilibrium point the iterative function converges to a fixed point) and by providing the equilibrium point as a substitute for the output of the layer stack in the neural network, in addition to defining other operations that may be performed on the forward propagation portion, wherein determining the equilibrium point includes using a numerical root-finding algorithm to find a root solution of the iterative function minus its input. The system 100 may also include an output interface for outputting a data representation 112 of the trained neural network, which data may also be referred to as trained model data 112. For example, as also Figure 1 As illustrated, the output interface may be constituted by a data storage interface 104, which in these embodiments is an input / output ("IO") interface, via which trained model data 112 may be stored in a data storage device 106. For example, during or after training, data representations 108 defining an "untrained" neural network may be at least partially replaced by data representations 112 of a trained neural network, as parameters of the neural network (such as weights, hyperparameters, and other types of parameters of the neural network) may be adapted to reflect the training of the training data 102. This is also illustrated by reference numerals 108, 112. Figure 1, reference numerals 108 and 112 refer to the same data record on the data storage device 106. In other embodiments, the data representation 112 may be stored separately from the data representation 108 defining the "untrained" neural network. In some embodiments, the output interface may be separate from the data storage interface 104, but may generally be of the type described above for the data storage interface 104.

[0030] The structure of system 100 is one example of a system that may be utilized to train the machine learning models described herein. Figure 2 Additional structures for operating and training machine learning models are shown.

[0031] Figure 2 A system 200 is depicted for implementing the machine learning models described herein. The system 200 can be implemented to perform synthetic data generation and scoring as described herein. The system 200 can include at least one computing system 202. The computing system 202 can include at least one processor 204, which is operably connected to a memory unit 208. The processor 204 can include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. The CPU 206 can be a commercially available processing unit that implements an instruction set, such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, the CPU 206 can execute stored program instructions retrieved from the memory unit 208. The stored program instructions can include software that controls the operation of the CPU 206 to perform the operations described herein. In some examples, the processor 204 can be a system on a chip (SoC) that integrates the functionality of the CPU 206, the memory unit 208, a network interface, and an input / output interface into a single integrated device. The computing system 202 can implement an operating system for managing various aspects of operation. Although Figure 2 One processor 204, one CPU 206, and one memory 208 are shown in FIG, but of course multiple ones of each of processor 204, CPU 206, and memory 208 may be utilized throughout the system.

[0032] The memory unit 208 may include volatile memory and non-volatile memory for storing instructions and data. Non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is deactivated or powered off. Volatile memory may include static and dynamic random access memory (RAM) that stores program instructions and data. For example, the memory unit 208 may store a machine learning model 210 or algorithm, a training data set 212 for the machine learning model 210, and an original source data set 216.

[0033] The computing system 202 may include a network interface device 222 configured to provide communications with external systems and devices. For example, the network interface device 222 may include a wired and / or wireless Ethernet interface as defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface to an external network 224 or the cloud.

[0034] External network 224 may be referred to as the World Wide Web or the Internet. External network 224 may establish standard communication protocols between computing devices. External network 224 may allow information and data to be easily exchanged between computing devices and the network. One or more servers 230 may communicate with external network 224.

[0035] The computing system 202 may include an input / output (I / O) interface 220, which may be configured to provide digital and / or analog input and output. The I / O interface 220 is used to transfer information between internal storage devices and external input and / or output devices (e.g., HMI devices). The I / O 220 interface may include associated circuitry or a BUS network to transfer information to or between the processor(s) and the memory. For example, the I / O interface 220 may include digital I / O logic lines that can be read or set by the processor(s), handshake lines for supervising data transfer via the I / O lines, timing and counting facilities, and other structures known to provide such functionality. Examples of input devices include a keyboard, a mouse, sensors, etc. Examples of output devices include a monitor, a printer, a speaker, etc. The I / O interface 220 may include an additional serial interface (e.g., a universal serial bus (USB) interface) for communicating with external devices.

[0036] The computing system 202 may include a human-machine interface (HMI) device 218, which may include any device that enables the system 200 to receive control inputs. Examples of input devices may include human-machine interface inputs such as a keyboard, mouse, touch screen, voice input device, and other similar devices. The computing system 202 may include a display device 232. The computing system 202 may include hardware and software for outputting graphics and text information to the display device 232. The display device 232 may include an electronic display screen, a projector, a printer, or other suitable device for displaying information to a user or operator. The computing system 202 may also be configured to allow interaction with a remote HMI and a remote display device via a network interface device 222.

[0037] System 200 can be implemented using one or more computing systems. Although the example depicts a single computing system 202 that implements all of the described features, it is intended that the various features and functions can be separate and implemented by multiple computing units communicating with each other. The specific system architecture selected may depend on a variety of factors.

[0038] The system 200 may implement a machine learning algorithm 210 configured to analyze a raw source data set 216. The raw source data set 216 may include raw or unprocessed sensor data, which may represent an input data set for a machine learning system. The raw source data set 216 may include video, video clips, images, text-based information, audio or human speech, time series data (e.g., pressure sensor signals over time), and raw or partially processed sensor data (e.g., a radar map of an object). Figure 5-Figure 11 , several different input examples are shown and described. In some examples, the machine learning algorithm 210 can be a neural network algorithm (e.g., a deep neural network) designed to perform a predetermined function. For example, a neural network algorithm can be configured in an automotive application to identify street signs or pedestrians in an image. The machine learning algorithm(s) 210 can include algorithms configured to operate the models described herein.

[0039] The computer system 200 can store a training data set 212 for the machine learning algorithm 210. The training data set 212 can represent a set of previously constructed data for training the machine learning algorithm 210. The machine learning algorithm 210 can use the training data set 212 to learn weighting factors associated with the neural network algorithm. The training data set 212 can include a set of source data with corresponding outputs or results that the machine learning algorithm 210 attempts to replicate through the learning process. In this example, the training data set 212 can include input images that include objects (e.g., street signs). The input images can include various scenes in which objects are identified.

[0040] The machine learning algorithm 210 can operate in a learning mode using the training dataset 212 as input. The machine learning algorithm 210 can be executed in multiple iterations using data from the training dataset 212. In each iteration, the machine learning algorithm 210 can update internal weighting factors based on the achieved results. For example, the machine learning algorithm 210 can compare the output results (e.g., a reconstructed or supplemented image in the case where image data is input) with the results included in the training dataset 212. Because the training dataset 212 includes expected results, the machine learning algorithm 210 can determine when the performance is acceptable. After the machine learning algorithm 210 achieves a predetermined performance level (e.g., 100% consistency with the results associated with the training dataset 212) or converges, the machine learning algorithm 210 can be executed using data not in the training dataset 212. It should be understood that in the present disclosure, "convergence" can mean that a set (e.g., predetermined) number of iterations have occurred, or that the residual is sufficiently small (e.g., the change in the approximation probability during the iteration is less than a threshold), or other convergence conditions. The trained machine learning algorithm 210 can be applied to a new dataset to generate annotated data.

[0041] The machine learning algorithm 210 can be configured to identify specific features in the raw source data 216. The raw source data 216 can include multiple examples or input data sets for which supplementary results are desired. For example, the machine learning algorithm 210 can be configured to identify the presence of road signs in a video image and annotate the presence. The machine learning algorithm 210 can be programmed to process the raw source data 216 to identify the presence of specific features. The machine learning algorithm 210 can be configured to identify features in the raw source data 216 as predetermined features (e.g., road signs). The raw source data 216 can be obtained from a variety of sources. For example, the raw source data 216 can be actual input data collected by the machine learning system. The raw source data 216 can be machine-generated and used to test the system. As an example, the raw source data 216 can include raw video images from a camera.

[0042] In an example, the original source data 216 may include image data representing an image. Applying the machine learning algorithm described herein, the output may be a score that can be utilized to assess synthesis quality.

[0043] Figure 3 A flowchart for evaluating anomalies associated with a synthetic dataset is illustrated. Flowchart 300 may reflect an illustrative embodiment. A key concept is to model the probability px of each example being an anomaly. The proposed quality score is the expected posterior of this parameter.

[0044] At step 301, the system can generate a synthetic dataset. The synthetic dataset can include one or more anomalies and one or more labels associated with the training dataset. In one example, the label can be associated with data that classifies the training dataset.

[0045] To quantify the uncertainty of the detector, the system can assume a Bayesian perspective and model the anomaly detection problem as

[0046] Y|X=x~BERNOULLI(px),px~BETA(α0,β0),

[0047] where px is the class probability of anomaly, which has a Beta distribution with parameters is the prior. This parameter can be interpreted as the probability that the example is an anomaly.

[0048] Given an example x, assume the system can observe N pseudo observations with labels y1,...,yn from P(Y|X=x). The system can then use traditional Bayesian updating to derive the posterior distribution px.

[0049] px|y1,...,yn~BETA(α0+α1,β0+n-α1)(Equation 1)

[0050] in, is the number of labels for the abnormal class.

[0051] However, it is unrealistic to assume that we have access to m labels for the same instance, which makes it difficult to calculate α1. We can estimate α1 by simulating n examples where nP(X=x)≈m, where only P(Y=1|X=x) will belong to the anomaly class. Therefore, we can use data density and conditional probability to estimate α1 as

[0052]

[0053] At step 303, the system determines the class conditional probability. Specifically, embodiments of the system and method consider both: (i) the class conditional probability P(Y=1|X=x) of example x, which generally reflects the underlying aleatoric uncertainty of the data, and (ii) the density P(X=x) of example x, which reflects epistemic uncertainty. The posterior expectation in equation (1) reflects the quality of the auxiliary anomaly x: if the expected posterior is high, then the evidence is sufficient to rely on the expected conditional probability to evaluate the auxiliary anomaly, while if the expected posterior is low, then the quality reflects the prior belief.

[0054] Computing P(Y=1|X=x) in anomaly detection is a difficult task because (1) class probabilities are generally unreliable for imbalanced classification tasks and (2) the available anomalies may not be representative of the entire anomaly class (i.e., the system and method may have access to a biased set). This makes traditional calibration techniques generally impractical for anomaly detection. However, the system may be primarily concerned with having probabilities that satisfy two properties. First, they must be consistent with the detector's predictions, i.e., predicting an anomaly (normal) requires a probability greater than (less than) 0.5. Second, the system may want the proportion of predicted anomalies to match the expected proportion of true anomalies. This ensures that if the detector's ranking is accurate, the class predictions are optimally computed.

[0055] For this task, the system can utilize and employ a compression scaler to map the anomaly score f(x) assigned by the detector f to a [0,1] probability value:

[0056]

[0057] where λ is the decision threshold of the detector.

[0058] At step 305, the system can determine an estimate associated with the data density. The system can estimate P(X=x) using a rarity score, which is fast to compute and less affected by the curse of dimensionality. Roughly speaking, the rarity score (1) creates a k-NN sphere centered at each training example and (2) assigns a minimum radius to the sphere that contains a given synthetic example. If the synthetic example falls outside all spheres, it is considered too uncommon and its rarity equals 0. Formally, given a sample x, the estimated data density is equal to the rarity score, which is the function rk: Make

[0059]

[0060] where NNk(xi) is the distance between xi and its kth nearest neighbor in D, and Bk(xi) = {x|d(xi,x)≤NNk(xi)} is the k-NN sphere with xi as the center and NNk(xi) as the radius.

[0061] At step 307, the system can calculate the quality score. Using the estimates of the data density and conditional probability, the system can calculate the quality score by taking the expected posterior distribution of px and the estimated To calculate the mass of the synthetic auxiliary anomaly x:

[0062]

[0063] Thus, the system may output a score at step 309. The score may determine the quality of the synthesized auxiliary anomaly. It may be useful to determine how the system behaves when subjected to: (1) a large training set (n→+∞); (2) a small training set or zero density examples; and (3) high-level conditional probabilities. Given a real-world anomaly x R , indistinguishable anomaly x1 and unrealistic anomaly x U , determine whether the system and method The following systems and methods may have three related properties: (1) convergence to class conditional probabilities; (2) convergence to the prior mean; and (3) the quality of distinguishable anomalies increases with their density.

[0064] Regarding convergence to class conditional probabilities, the number of training examples indicates how strong the empirical evidence is. That is, does the detector f have enough evidence to properly estimate the class conditional probabilities. Therefore, increasing n makes EAP converge to class conditional probabilities:

[0065]

[0066] Regarding convergence to the prior mean, given an example x, if there is no empirical evidence to update px, then the posterior distribution is still equal to the prior. This lack of evidence could be due to n being relatively small or having low density.

[0067]

[0068] Regarding the quality of obvious anomalies, as their density increases, given that the class conditional probability of an example x is equal to 1 (i.e., it is easily distinguished from normal data by f), its quality depends only on its density. In fact, the closer it is to the training examples, the higher the density, and the higher the quality:

[0069]

[0070] Figure 4 A graph illustrating scores associated with anomalies for different methods is shown. The graph illustrates the average "area under the receiver operator curve" (AUCQLT) obtained by each method on a per-dataset basis (image on the left, table on the right). The embodiment utilized is labeled EAP and achieves the highest (i.e., best) performance on most datasets, beating runners-up RARITY and LAVA on 30 and 31 of the 40 datasets, respectively.

[0071] The machine learning model described in this article can be used for many different applications, not just in the context of road sign image processing. Figures 6-11Additional applications in which anomaly detection or classification may be used are shown in . Figure 5 The architecture used to train and use machine learning models for these applications (and others) is illustrated in

[15] . Figure 5 A schematic diagram depicting the interaction between a computer-controlled machine 500 and a control system 502 is shown. The computer-controlled machine 500 includes an actuator 504 and a sensor 506. The actuator 504 may include one or more actuators, and the sensor 506 may include one or more sensors. The sensor 506 is configured to sense a condition of the computer-controlled machine 500. The sensor 506 may be configured to encode the sensed condition into a sensor signal 508 and transmit the sensor signal 508 to the control system 502. Non-limiting examples of the sensor 506 include video, radar, lidar, ultrasonic, and motion sensors. In one embodiment, the sensor 506 is an optical sensor configured to sense an optical image of the environment near the computer-controlled machine 500.

[0072] The control system 502 is configured to receive sensor signals 508 from the computer-controlled machine 500. As explained below, the control system 502 may also be configured to calculate actuator control commands 510 depending on the sensor signals and transmit the actuator control commands 510 to the actuators 504 of the computer-controlled machine 500.

[0073] like Figure 5 As shown in FIG, control system 502 includes a receiving unit 512. Receiving unit 512 can be configured to receive sensor signals 508 from sensor 506 and convert sensor signals 508 into input signals x. In an alternative embodiment, sensor signals 508 are received directly as input signals x without receiving unit 512. Each input signal x can be a portion of each sensor signal 508. Receiving unit 512 can be configured to process each sensor signal 508 to generate each input signal x. Input signal x can include data corresponding to an image recorded by sensor 506.

[0074] The control system 502 includes a classifier 514. The classifier 514 can be configured to classify an input signal x into one or more labels using a machine learning (ML) algorithm, such as the neural network described above. The classifier 514 is parameterized by parameters, such as those described above (e.g., the parameter θ). The parameter θ can be stored in and provided by a non-volatile storage device 516. The classifier 514 is configured to determine an output signal y based on the input signal x. Each output signal y includes information assigning one or more labels to each input signal x. The classifier 514 can transmit the output signal y to a conversion unit 518. The conversion unit 518 is configured to convert the output signal y into an actuator control command 510. The control system 502 is configured to transmit the actuator control command 510 to the actuator 504, which is configured to actuate the computer-controlled machine 500 in response to the actuator control command 510. In another embodiment, the actuator 504 is configured to actuate the computer-controlled machine 500 directly based on the output signal y.

[0075] When the actuator control command 510 is received by the actuator 504, the actuator 504 is configured to perform an action corresponding to the associated actuator control command 510. The actuator 504 may include control logic configured to transform the actuator control command 510 into a second actuator control command that is utilized to control the actuator 504. In one or more embodiments, the actuator control command 510 may be utilized to control a display instead of or in addition to the actuator.

[0076] In another embodiment, the control system 502 includes the sensor 506 instead of or in addition to the computer-controlled machine 500 including the sensor 506. The control system 502 may also include the actuator 504 instead of or in addition to the computer-controlled machine 500 including the actuator 504.

[0077] like Figure 5 As shown in , the control system 502 also includes a processor 520 and a memory 522. The processor 520 may include one or more processors. The memory 522 may include one or more memory devices. The classifier 514 of one or more embodiments (e.g., a machine learning algorithm, such as the algorithm described above with respect to the pre-trained classifier 306) can be implemented by the control system 502, which includes the non-volatile storage device 516, the processor 520, and the memory 522.

[0078] The non-volatile storage 516 may include one or more persistent data storage devices, such as a hard drive, an optical drive, a tape drive, a non-volatile solid-state device, a cloud storage device, or any other device capable of persistently storing information. The processor 520 may include one or more devices selected from a high-performance computing (HPC) system, including a high-performance core, a microprocessor, a microcontroller, a digital signal processor, a microcomputer, a central processing unit, a field programmable gate array, a programmable logic device, a state machine, a logic circuit, an analog circuit, a digital circuit, or any other device that manipulates (analog or digital) signals based on computer-executable instructions residing in the memory 522. The memory 522 may include a single memory device or multiple memory devices, including but not limited to random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information.

[0079] The processor 520 can be configured to read the memory 522 and execute computer-executable instructions that reside in the non-volatile storage 516 and embody one or more ML algorithms and / or methods of one or more embodiments. The non-volatile storage 516 can include one or more operating systems and applications. The non-volatile storage 516 can store compiled and / or interpreted computer programs created using a variety of programming languages ​​and / or technologies, including but not limited to Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL, whether used alone or in combination.

[0080] When executed by the processor 520, the computer-executable instructions of the non-volatile storage device 516 may cause the control system 502 to implement one or more ML algorithms and / or methods disclosed herein. The non-volatile storage device 516 may also include ML data (including data parameters) that supports the functions, features, and processes of one or more embodiments described herein.

[0081] The program code embodying the algorithms and / or methods described herein can be distributed as a program product in a variety of different forms, either individually or collectively. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of one or more embodiments. Computer-readable storage media are non-transitory in nature and can include volatile and non-volatile, as well as removable and non-removable tangible media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media can also include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technology, portable compact disc read-only memory (CD-ROM) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or any other medium that can be used to store desired information and can be read by a computer. The computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing device or other device, or downloaded to an external computer or external storage device via a network.

[0082] The computer-readable program instructions stored in a computer-readable medium can be used to instruct a computer, other types of programmable data processing devices, or other devices to operate in a particular manner so that the instructions stored in the computer-readable medium (including instructions for implementing the functions, actions, and / or operations specified in the flowchart or diagram) produce an article of manufacture. Consistent with one or more embodiments, in certain alternative embodiments, the functions, actions, and / or operations specified in the flowchart and diagram can be reordered, processed serially, and / or processed simultaneously. In addition, any flowchart and / or diagram may include more or fewer nodes or blocks than those shown in accordance with one or more embodiments.

[0083] The process, method or algorithm may be embodied in whole or in part using suitable hardware components, such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), state machines, controllers or other hardware components or devices, or a combination of hardware, software and firmware components.

[0084] Figure 6A schematic diagram of a control system 502 is depicted, which is configured to control a vehicle 600, which can be an at least partially autonomous vehicle or an at least partially autonomous robot. Vehicle 600 includes an actuator 504 and a sensor 506. Sensor 506 can include one or more video sensors, cameras, radar sensors, ultrasonic sensors, lidar sensors, and / or location sensors (e.g., GPS). One or more of the one or more specific sensors can be integrated into vehicle 600. In the context of landmark recognition and processing described herein, sensor 506 is a camera mounted to or integrated into vehicle 600. As an alternative to or in addition to one or more specific sensors identified above, sensor 506 can include a software module that is configured to determine the state of actuator 504 when executed. A non-limiting example of a software module includes a weather information software module that is configured to determine the current or future state of the weather near vehicle 600 or other location.

[0085] The classifier 514 of the control system 502 of the vehicle 600 can be configured to detect an object in the vicinity of the vehicle 600 based on the input signal x. In such an embodiment, the output signal y can include information characterizing the proximity of the object relative to the vehicle 600. The actuator control command 510 can be determined based on this information. The actuator control command 510 can be used to avoid a collision with the detected object.

[0086] In embodiments where the vehicle 600 is at least partially autonomous, the actuator 504 may be embodied in the brakes, propulsion system, engine, transmission, or steering of the vehicle 600. Actuator control commands 510 may be determined such that the actuator 504 is controlled such that the vehicle 600 avoids a collision with a detected object. The detected object may also be classified based on what the classifier 514 deems most likely to be, such as a pedestrian or a tree. The actuator control commands 510 may be determined based on the classification. In the event of a potential adversarial attack, the system described above may also be trained to better detect objects or identify changes in lighting conditions or angles for sensors or cameras on the vehicle 600.

[0087] In other embodiments where the vehicle 600 is an at least partially autonomous robot, the vehicle 600 can be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and stepping. The mobile robot can be an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. In such embodiments, the actuator control commands 510 can be determined such that a propulsion unit, a steering unit, and / or a braking unit of the mobile robot can be controlled so that the mobile robot can avoid a collision with the identified object.

[0088] In another embodiment, vehicle 600 is an at least partially autonomous robot in the form of a gardening robot. In such an embodiment, vehicle 600 can use an optical sensor as sensor 506 to determine the state of plants in the environment near vehicle 600. Actuator 504 can be a nozzle configured to spray a chemical. Depending on the identified plant species and / or the identified plant state, actuator control commands 510 can be determined to cause actuator 504 to spray an appropriate amount of a suitable chemical onto the plant.

[0089] Vehicle 600 can be an at least partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances include a washing machine, stove, oven, microwave, or dishwasher. In such a vehicle 600, sensor 506 can be an optical sensor configured to detect the state of an object to be processed by the household appliance. For example, if the household appliance is a washing machine, sensor 506 can detect the state of the laundry inside the washing machine. Actuator control commands 510 can be determined based on the detected laundry state.

[0090] Figure 7 A schematic diagram of a control system 502 is depicted, which is configured to control a system 700 (e.g., a manufacturing machine), such as a punch, cutter, or gun drill, that is part of the manufacturing system 702 (e.g., a production line). The control system 502 can be configured to control an actuator 504 that is configured to control the system 700 (e.g., a manufacturing machine).

[0091] The sensor 506 of the system 700 (e.g., a manufacturing machine) can be an optical sensor configured to capture one or more properties of the manufactured product 704. The classifier 514 can be configured to determine a state of the manufactured product 704 based on the captured one or more properties. The actuator 504 can be configured to control the system 700 (e.g., a manufacturing machine) based on the determined state of the manufactured product 704 to perform subsequent manufacturing steps of the manufactured product 704. The actuator 504 can be configured to control the system 700 (e.g., a manufacturing machine) based on the determined state of the manufactured product 704 to control the function of the system 700 (e.g., a manufacturing machine) for the subsequent manufactured product 106 of the system 700 (e.g., a manufacturing machine).

[0092] Figure 8 A schematic diagram of a control system 502 configured to control a power tool 800 (such as a drill or driver) having at least a partially autonomous mode is depicted. The control system 502 may be configured to control an actuator 504 configured to control the power tool 800.

[0093] The sensor 506 of the power tool 800 can be an optical sensor that is configured to capture one or more properties of the work surface 802 and / or the fastener 804 driven into the work surface 802. The classifier 514 can be configured to determine a state of the work surface 802 and / or the fastener 804 relative to the work surface 802 based on the captured one or more properties. The state can be that the fastener 804 is flush with the work surface 802. Alternatively, the state can be the hardness of the work surface 802. The actuator 504 can be configured to control the power tool 800 so that the driving function of the power tool 800 is adjusted depending on the determined state of the fastener 804 relative to the work surface 802 or the one or more properties of the captured work surface 802. For example, if the state of the fastener 804 is flush with the work surface 802, the actuator 504 can stop the driving function. As another non-limiting example, the actuator 504 can apply additional or less torque depending on the hardness of the work surface 802.

[0094] Figure 9 Depicted is a schematic diagram of a control system 502 configured to control an automated personal assistant 900. The control system 502 may be configured to control an actuator 504, which may be configured to control the automated personal assistant 900. The automated personal assistant 900 may be configured to control a household appliance, such as a washing machine, stove, oven, microwave, or dishwasher.

[0095] Sensor 506 may be an optical sensor and / or an audio sensor. The optical sensor may be configured to receive a video image of gesture 904 of user 902. The audio sensor may be configured to receive a voice command of user 902.

[0096] Control system 502 of automated personal assistant 900 can be configured to determine actuator control commands 510, which are configured to control system 502. Control system 502 can be configured to determine actuator control commands 510 based on sensor signals 508 from sensor 506. Automated personal assistant 900 is configured to transmit sensor signals 508 to control system 502. Classifier 514 of control system 502 can be configured to execute a gesture recognition algorithm to identify gesture 904 performed by user 902 in order to determine actuator control commands 510 and transmit actuator control commands 510 to actuator 504. Classifier 514 can be configured to retrieve information from non-volatile storage in response to gesture 904 and output the retrieved information in a form suitable for receipt by user 902.

[0097] Figure 10A schematic diagram of a control system 502 configured to control a monitoring system 1000 is depicted. The monitoring system 1000 can be configured to physically control access through a door 1002. A sensor 506 can be configured to detect a scene relevant to determining whether to grant access. The sensor 506 can be an optical sensor configured to generate and transmit image and / or video data. The control system 502 can use such data to detect a person's face.

[0098] The classifier 514 of the control system 502 of the monitoring system 1000 can be configured to interpret the image and / or video data by matching the identities of known persons stored in the non-volatile storage device 516 to determine the identity of the person. The classifier 514 can be configured to generate an actuator control command 510 in response to the interpretation of the image and / or video data. The control system 502 can be configured to transmit the actuator control command 510 to the actuator 504. In this embodiment, the actuator 504 can be configured to lock or unlock the door 1002 in response to the actuator control command 510. In other embodiments, non-physical logical access control is also possible.

[0099] Monitoring system 1000 may also be a surveillance system. In such an embodiment, sensor 506 may be an optical sensor configured to detect a monitored scene, and control system 502 may be configured to control display 1004. Classifier 514 may be configured to determine a classification of the scene, e.g., whether the scene detected by sensor 506 is suspicious. Control system 502 may be configured to transmit actuator control commands 510 to display 1004 in response to the classification. Display 1004 may be configured to adjust displayed content in response to actuator control commands 510. For example, display 1004 may highlight an object deemed suspicious by classifier 514. Using embodiments of the disclosed system, a surveillance system may predict the appearance of an object at a certain time in the future.

[0100] Figure 11 A schematic diagram of a control system 502 is depicted, which is configured to control an imaging system 1100, such as an MRI apparatus, an X-ray imaging apparatus, or an ultrasound apparatus. Sensor 506 may be, for example, an imaging sensor. Classifier 514 may be configured to determine a classification of all or a portion of a sensed image. Classifier 514 may be configured to determine or select actuator control commands 510 in response to the classification obtained by the trained neural network. For example, classifier 514 may interpret a region of the sensed image as potentially abnormal. In this case, actuator control commands 510 may be determined or selected to cause display 1102 to display the image and highlight the region of potential abnormality.

[0101] While exemplary embodiments have been described above, these embodiments are not intended to describe all possible forms encompassed by the claims. The terms used in the specification are descriptive rather than restrictive, and it will be understood that various changes may be made without departing from the spirit and scope of the present disclosure. As previously described, features of the various embodiments may be combined to form additional embodiments of the present invention that may not be explicitly described or illustrated. While various embodiments may have been described as providing advantages or being superior to other embodiments or prior art implementations with respect to one or more desired features, those skilled in the art recognize that one or more features or characteristics may be compromised to achieve desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, lifecycle cost, marketability, appearance, packaging, size, applicability, weight, manufacturability, ease of assembly, and the like. Thus, to the extent that any embodiment is described as less desirable than another embodiment or prior art implementation with respect to one or more characteristics, such embodiments do not exceed the scope of the present disclosure and may be desirable for a particular application.

Claims

1. A method for labeling data for a machine learning (ML) model, the method comprising: generating one or more synthetic datasets using the training dataset, wherein the one or more synthetic datasets include one or more anomalies and one or more labels associated with the training dataset; determining, using an anomaly detector of the ML model, an estimate associated with a class conditional probability for the one or more synthetic data sets in response to the mapped anomaly scores assigned by the anomaly detector; determining, using a data density estimator, a data density associated with the one or more synthetic data sets in response to a rarity score function or a kernel density estimation function associated with the training data sets; determining a quality score associated with the one or more synthetic data sets using at least the data density and the estimate associated with the class conditional probability; and Quality scores associated with anomalies of the one or more synthetic data sets are output.

2. The method according to claim 1, wherein The anomalies are synthetic anomalies generated by data augmentation or generative models and are not present in the training dataset.

3. The method according to claim 1, wherein The dataset includes one or more images.

4. The method according to claim 1, wherein The method includes utilizing a compression scaler to map one or more anomaly scores to one or more probability values.

5. The method according to claim 1, wherein The training dataset includes a set of one or more images.

6. The method according to claim 1, wherein The rarity scores create one or more spheres using a k-nearest neighbor model. 7 . The method of claim 1 , further determining an estimate associated with a class conditional probability for the one or more synthetic data sets in response to the mapping score between 0 and 1. 8 .

8. The method according to claim 1, wherein The method includes taking the inverse of the rarity score and normalizing the inverse.

9. A system configured to label data using a machine learning (ML) model, comprising: an input interface configured to receive input data from a sensor, wherein the sensor comprises a video, radar, lidar, sound, sonar, ultrasonic, motion, or thermal imaging sensor; a processor in communication with the input interface, wherein the processor is programmed to: generating one or more synthetic datasets using the training dataset, wherein the one or more synthetic datasets include one or more anomalies and one or more labels associated with the training dataset; determining, using an anomaly detector of the ML model, an estimate associated with a class conditional probability for the one or more synthetic data sets in response to the mapped anomaly scores assigned by the anomaly detector; determining, using a data density estimator, a data density associated with the one or more synthetic data sets in response to a rarity score function or a kernel density estimation function associated with the training data sets; determining a quality score associated with the one or more synthetic data sets using at least the data density and the estimate associated with the class conditional probability; and A quality score associated with one of the one or more anomalies of the one or more synthetic data sets is output.

10. The system according to claim 9, wherein: The anomalies are synthetic anomalies generated by data augmentation or generative models and are not present in the training dataset.

11. The system according to claim 9, wherein: The dataset includes one or more images.

12. The system according to claim 9, wherein: The anomaly detector includes an anomaly detection model.

13. The system according to claim 9, wherein: The data density estimator includes a data density model.

14. The system according to claim 9, wherein: The rarity scores create one or more spheres using a k-nearest neighbor model.

15. The system according to claim 9, wherein: The system includes a compression scaler to map one or more anomaly scores to one or more probability values.

16. A computer program product storing instructions which, when executed by a computer, cause the computer to: Using the training dataset, generate one or more synthetic datasets, where The one or more synthetic data sets include one or more anomalies associated with the training data set; determining, using an anomaly detector of the ML model, an estimate associated with a class conditional probability for the one or more synthetic data sets in response to the mapped anomaly scores assigned by the anomaly detector; determining, using a data density estimator, a data density associated with the one or more synthetic data sets in response to a rarity score function or a kernel density estimation function associated with the training data sets; determining a quality score associated with the one or more synthetic data sets using at least the data density and an estimate associated with a class conditional probability; as well as A quality score associated with one of the one or more anomalies of the one or more synthetic data sets is output.

17. The computer program product of claim 16, wherein: The one or more synthetic data sets relate to data associated with video, camera, radar, lidar, sound, sonar, ultrasonic, motion, or thermal imaging sensors.

18. The computer program product of claim 16, wherein: The synthetic dataset includes one or more labels.

19. The computer program product of claim 16, wherein: The rarity scores create one or more spheres using a k-nearest neighbor model.

20. The computer program product of claim 16, wherein: The instructions include utilizing a compression scaler to map one or more anomaly scores to one or more probability values.