Anomaly detection using local neural transformations
Local neural transformations with CPC and DDCL enhance anomaly detection in time series by learning semantic and diverse transformations, enabling efficient and reliable detection of anomalous regions in time series data, surpassing deep learning baselines.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2026-03-17
AI Technical Summary
Existing anomaly detection methods in machine learning systems, particularly for time series data, struggle to effectively identify anomalous regions within sequences due to the lack of labeled data and reliance on domain-specific features, leading to inefficiencies in detecting and responding to anomalous behavior.
The method employs local neural transformations (LNT) that combine contrast predictive coding (CPC) with dynamic deterministic contrast loss (DDCL) to generate diverse vector representations, scoring each time step for anomaly detection, using a self-supervised approach that learns semantic and diverse transformations while ensuring locality, and applies hidden Markov models for continuous anomaly region detection.
The LNT method achieves robust and scalable anomaly detection in time series data, outperforming deep learning baselines and providing reliable, continuous anomaly region detection without requiring domain-specific features, as demonstrated by experiments on the LibriSpeech dataset.
Smart Images

Figure 0007832064000051 
Figure 0007832064000052 
Figure 0007832064000053
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to anomaly detection in machine learning systems. More specifically, the present application relates to improvements in anomaly detection of time series in machine learning systems via local neural transformation.
Background Art
[0002] Background Art In data analysis, anomaly detection (also referred to as outlier detection) is the identification or observation of specific data or events that arouse suspicion by being significantly different from most of the data. Typically, anomaly items are converted into some problem such as a defect in structure, malfunction, misoperation, medical problem, error, etc.
Summary of the Invention
Means for Solving the Problems
[0003] Summary The anomaly detection method includes steps of receiving time series data grouped into patches, encoding the data via parameters of an encoder to obtain a local latent representation for each patch, determining a representation loss from the local latent representation for each patch, converting the local latent representation associated with each patch into a series of diversely transformed vector representations via at least two local neural transformations, determining a dynamic deterministic contrastive loss (DDCL) from the series of diversely transformed vector representations, combining the representation loss and the DDCL to obtain updated parameters, updating the parameters of the encoder using the updated parameters, scoring each of the series of diversely transformed vector representations via the DDCL to obtain a diverse semantic requirement score associated with each patch, smoothing the diverse semantic requirement scores to obtain a loss region, masking data associated with the loss region to obtain verified data, and outputting the verified data.
[0004] The anomaly region detection system includes a controller, which receives time-series data grouped into patches from a first sensor, encodes the data via encoder parameters to obtain a local latent representation for each patch, calculates a contrast predictive coding (CPC) loss from the local latent representation for each patch, transforms the local latent representation associated with each patch into a series of diversely transformed vector representations via at least two local neural transformations, calculates a dynamic deterministic contrast loss (DDCL) from the series of diversely transformed vector representations, combines the CPC loss and DDCL to obtain updated parameters, updates the encoding parameters using the updated parameters, scores each of the series of diversely transformed vector representations via DDCL to obtain diverse semantic demand scores associated with each patch, smooths the diverse semantic demand scores to obtain a loss region, masks the data associated with the loss region to obtain verified data, and operates the machine based on the verified data.
[0005] The telecommunication system includes a controller, which receives time-series data grouped into patches, encodes the data via encoder parameters to obtain a local latent representation for each patch, determines a contrast predictive coding (CPC) loss from the local latent representation for each patch, transforms the local latent representation associated with each patch into a series of diversely transformed vector representations via at least two local neural transformations, calculates a dynamic deterministic contrast loss (DDCL) from the series of diversely transformed vector representations, combines the CPC loss and DDCL to obtain updated parameters, updates the encoding parameters using the updated parameters, scores each of the series of diversely transformed vector representations via DDCL to obtain diverse semantic demand scores associated with each patch, smooths the diverse semantic demand scores to obtain a loss region, masks the data associated with the loss region to obtain verified data, and operates the telecommunication system based on the verified data. [Brief explanation of the drawing]
[0006] [Figure 1] This is a flowchart of an anomaly detection system using local neural transformation (LNT). [Figure 2] This is a flowchart of local neural transformations related to latent representations. [Figure 3] This is a flowchart of a dynamic deterministic contrast loss (DDCL) with push / pull in latent space. [Figure 4] This is a flowchart for forward scoring and backward scoring. [Figure 5] This is a block diagram of an electronic computing system. [Figure 6] This is a graphical representation of the receiver operating characteristic (ROC) curves of the local neural transform (LNT), showing the true positive rate (TPR) relative to the false positive rate (FPR) for different design choices. [Figure 7] This is a graphical representation of an example signal with anomaly injection, score, and error detection signal. [Figure 8] This is a graphical representation of an example signal with anomaly injection, score, and error detection signal. [Figure 9] This is a flowchart of an anomaly detection system using comparative predictive coding (CPC). [Figure 10] This is an explanatory diagram of the contrast predictive coding (CPC) loss value, as well as the image corresponding to the superposition of both the image and the loss value. [Figure 11] This is a schematic diagram of a control system configured to control a vehicle. [Figure 12] This is a schematic diagram of a control system configured to control manufacturing machinery. [Figure 13] This is a schematic diagram of a control system configured to control power tools. [Figure 14] This is a schematic diagram of a control system configured to control an automated personal assistant. [Figure 15] This is a schematic diagram of a control system configured to control a monitoring system. [Figure 16] This is a schematic diagram of a control system configured to control a medical imaging system. [Modes for carrying out the invention]
[0007] Detailed explanation Where necessary, detailed embodiments of the present invention are disclosed herein, but it should be understood that these embodiments are merely illustrative examples of the present invention, which can be embodied in various alternative forms. The drawings are not necessarily to scale, and some features may be exaggerated or minimized to illustrate the details of certain components. Accordingly, certain structural and functional details disclosed herein should not be constrained, but rather should be interpreted simply as representative grounds to teach those skilled in the art how to employ the present invention in various ways.
[0008] In this specification, the term “substantially” may be used to describe the disclosed or requested embodiments. The term “substantially” may modify values or relative characteristics disclosed or requested in this disclosure. In such cases, “substantially” may mean values or relative characteristics that are modified within ±0%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, or 10%.
[0009] The term "sensor" refers to a device that detects or measures physical properties, records them, displays them, or responds to them by other means. This includes optical sensors, light sensors, imaging sensors, or photon sensors (e.g., charge-coupled devices (CCD), CMOS active pixel sensors (APS), infrared sensors (IR), CMOS sensors, etc.), acoustic sensors, sound sensors, or vibration sensors (e.g., microphones, geophones, hydrophones, etc.), automotive sensors (e.g., wheel speed sensors, parking sensors, radar, oxygen sensors, blind spot sensors, torque sensors, LIDAR, etc.), chemical sensors (e.g., ion-sensitive field-effect transistors (ISFETs), oxygen sensors, carbon dioxide sensors, chemistristors, holographic sensors, etc.), current sensors, potential sensors, magnetic sensors, or high-frequency sensors (e.g., Hall effect sensors, magnetometers, magnetoresistive meters, Faraday cups, galvanometers, etc.), environmental sensors, weather sensors, moisture sensors, or humidity sensors (e.g., weather radar, etc.). This includes devices such as cutinometers, flow sensors or velocity sensors (e.g., mass airflow sensors, anemometers, etc.), ionizing radiation sensors or particle sensors (e.g., ionization chambers, Geiger counters, neutron detectors, etc.), navigation sensors (e.g., GPS sensors, magnetohydrodynamic (MHD) sensors, etc.), position sensors, angle sensors, displacement sensors, distance sensors, velocity sensors or acceleration sensors (e.g., LiDAR, accelerometers, ultra-wideband radar, piezoelectric sensors, etc.), force sensors, density sensors or level sensors (e.g., strain gauges, nuclear densimeters, etc.), thermal sensors, heating sensors or temperature sensors (e.g., infrared thermometers, pyrometers, thermocouples, thermistors, microwave radiometers, etc.), or other devices, modules, machines or subsystems intended to detect or measure physical properties, record them, display them, or respond to them by other means.
[0010] Specifically, the sensor can measure the characteristics of a time-series signal, which may include spatial or spatiotemporal aspects such as location in space. The signal may include electromechanical, sound, light, electromagnetic, RF, or other time-series data. The technology disclosed in this application can be applied, for example, to time-series imaging using other sensors such as radio electromagnetic antennas and sound-collecting microphones.
[0011] The term "image" refers to a representation or artifact, such as a photograph or other two-dimensional image, that depicts the perception of physical properties (e.g., audible sound, visible light, infrared light, ultrasound, underwater acoustics, etc.), which are analogous to a subject (e.g., a physical object, scene, or property) and thus provide a depiction of it. Images may be multidimensional in that they may contain components of time, space, intensity, density, or other properties. For example, an image may include a time-series image. This technique can also be extended to image 3D sound sources or objects.
[0012] Anomaly detection is applicable to a wide range of systems, including medical devices, security systems, consumer systems, automotive systems, autonomous driving systems, wireless communication systems, sensors, and image defect detection using machine vision. It is often used in preprocessing to remove anomalous data from datasets. In supervised learning, removing anomalous data from a dataset often results in statistically significant improvements in accuracy.
[0013] The detection of anomalous regions within time series is critical in many applications, including consumer, medical, industrial, automotive, and aerospace. In this disclosure, a data-enhancement-based auxiliary task can be used to train a deep architecture, the performance of which may then be used as an anomaly score. This disclosure provides solutions for domains such as time-series data or tabular data. These domains benefit when augmentation is learned from the data. This disclosure provides an approach that extends these methods to the task of detecting anomalous regions within time series. This disclosure presents a Local Neural Transform (LNT), an end-to-end pipeline for detecting anomalous regions. This LNT learns to locally embed and augment the time series, generating an anomaly score for each time step. Through experiments, it has been shown that the LNT can detect synthetic noise in speech segments from the LibriSpeech dataset. While this disclosure focuses on time series, the concept can also be applied to other types of data, such as spatial or spatiotemporal data.
[0014] Detecting anomalous behavior in time series is a critical task in many machine learning applications. One aspect of this disclosure involves training a deep learning-based approach that can scan new time series for anomalies using an unlabeled dataset of (potentially multivariate) time series. Specifically, the focus is on the task of anomalous region detection, as anomalies should be judged at the subsequence level rather than at the entire sequence. In many applications, this setup can be critical for responding quickly to anomalous behavior.
[0015] In this disclosure, data augmentation is used to define tasks for training deep feature extractors, and performance on these tasks is used as an anomaly score. Specifically, this includes improvements to adapt self-supervised anomaly detection to domains other than images, and deep methods for time-series anomaly detection that utilize the advantages of data augmentation.
[0016] In this disclosure, a transformation learning approach of local neural transformations (LNTs) may be used to detect anomalies in time series data. The anomaly detection pipeline then combines feature extractors (such as contrast predictive coding CPCs) with the neural transformation learning. Finally, a scoring method based on a hidden Markov model (HMM) combines the scores to detect anomalous regions. This disclosure presents (1) a novel self-supervised method for detecting anomalous regions in time series by combining representation learning and transformation learning, and (2) a scoring method that combines different loss contributions into anomaly scores for each time step.
[0017] Here, Local Neural Transform (LNT) can be applied as an anomaly detection method for time series.
number
number
[0018] Figure 1 is a flowchart of the anomaly region detection system 100 via a local neural transform (LNT). The encoder 102 creates a local latent representation of the data via internal parameters to the encoder. This local latent representation is transformed into a series of diversely transformed vector representations via a local neural transform 104. The encoder also outputs representation losses 106, such as the CPC loss. Furthermore, the local neural transform 104 also outputs a dynamic deterministic contrast loss (DDCL) 108.
[0019] Local time series representation: The first step is to encode the time series to obtain the local latent representation z t This also generates representation losses, which are losses that evaluate the quality of the representation, such as CPC loss or autoencoder loss. The time series patch x t:t+τ (with window size τ) is encoded by strided convolution to obtain the local latent representation z t Then, they are processed by a recurrent model (GRU) to form the context representation c t Both are trained using contrastive predictive coding (CPC), which contrasts the linear k-step future prediction W j for the negative samples z k c t drawn from the proposed distribution Z, so that the context embedding c t promotes the prediction of neighboring patches. The resulting loss (Equation 1) is the cross-entropy that classifies the positive samples from a fixed-size set Z~Z of negative samples using the exponentiated similarity measure
Number
Number
Number
[0020] Local Neural Transformations: Next, the time-series representation is processed by local neural transformations to generate different views of each embedding. These local transformations enable the detection of anomalous regions.
[0021] Figure 2 is a flowchart of the local neural transformation 200 for the latent representation. The encoder 202 creates a local latent representation 204 of the data via internal parameters for the encoder. This local latent representation 204 is transformed via the local neural transformation 206 into a series of differently transformed vector representations 208 (208a, 208b, 208c, 208d, 208e).
[0022] For example, while the semantic transformation of an image may be rotation or color distortion, in one embodiment using other data, NeuTraL extracts a parameter θ from the data by comparing the original sample to different views of the same sample. l Neural transformation T using l The set of (·) was used to train the model. This did not rely on negative samples drawn from the noise distribution, as in Noise Contrast Estimation (NCE), and therefore resulted in a deterministic contrast loss.
[0023] This disclosure reveals different potential views. l (z t To obtain the learned latent representation z, this approach is used as shown in Figure 2. t This disclosure proposes its application to the following: Furthermore, this disclosure provides a future prediction W for different situations k, as given by Equation 2. k c t By incorporating it as a positive sample, we extend the time-dependent loss to a dynamic deterministic contrast loss (DDCL).
number
number
number
[0024] This, as a result, produces the sum (Equation 3) of different categorical cross-entropy losses (Equation 2), and these are all,
number
number
number
number
number
number
[0025] Ultimately, the two objectives of Equations 1 and 3 are met by an integrated loss L = L with hyperparameter λ that balances representation learning and transformation learning. CPC +λL DDCL We will use this to train collaboratively.
[0026] Transformation Learning Principle: Rather than assuming that every function learned from data creates a meaningful transformation (or view), this disclosure presents the motivation for the architecture in Figure 1, which has three key principles that must be held for the learned transformation.
[0027] The two principles are (1) semantics and (2) diversity, which help to eliminate common harmful contour cases in transformation learning and support robust transformation learning for tabular data. By strengthening the first two principles and a third principle, (3) locality, downstream tasks of anomaly region detection are assumed and provided.
[0028] Semantics is the view generated by learned transformations that are supposed to share important semantic information with the original sample.
[0029] Diversity is a learned transformation that should generate diverse views of each sample, resulting in diverse and challenging self-monitoring tasks that require strong semantic features to be resolved.
[0030] Locality refers to transformations that, while respecting the general context of a sequence, should only affect data within its local neighborhood.
[0031] Regarding locality, the performance of the induced self-monitoring task is required to be locally sensitive to anomalous behavior within some general contexts of the sequence. This is expected to detect "outlier windows" that are anomalous across the entire dataset, but it should be emphasized that this differs from a simple sliding window approach that would not detect anomalous behavior only within the context of a specific sequence.
[0032] A key insight from CPC is the local representation of time series z t and general expression c t This involves breaking it down into z. t It may also be utilized by applying neural transformations only to the transformations that satisfy the principle of locality.
[0033] From a different perspective, L DDCL In the latent space, this can be interpreted as different representations involving push and pull, as seen in Figure 3. The numerator of Equation 2 is the learned transformation with pull.
number
number
[0034] Anomaly scoring: To score anomalies at a specific point in time, L DDCL Loss reuse is considered. This has the advantage of being deterministic and therefore does not require the extraction of negative samples from proposed or noise distributions, as other comparative self-monitoring tasks may require.
[0035] As a result, the score against time t is as depicted in Figure 4. t To derive this, there are two possible methods for integrating time dependence into this approach.
[0036] Figure 3 is a flowchart of a dynamic deterministic contrast loss (DDCL) 300 with push / pull in the latent space. The encoder 302 creates a local latent representation 304 of the data via internal parameters to the encoder. Note that this flowchart shows the flow at different points in time, e.g., t, t-1, t-2, t+1, etc. This local latent representation 304 is transformed into a series of diversely transformed vector representations 308 via local neural transformations 306,310. These local transformations 306,310 are shown as recurrent neural networks (RNNs), but can also be implemented as convolutional neural networks (CNNs) or other neural transformations.
[0037] Figure 4 is a flowchart of forward scoring 400 and backward scoring 450.
[0038] Forward scoring is the transformed representation at time t.
number
number
[0039] Post-scoring is scoring in hindsight. t To update the future expression
number
number
[0040] It should be noted that during training, all loss contributions are summed up, making these considerations meaningless. Based on experiments in several embodiments, backward scoring smoothed and distorted predictions compared to forward scoring. Thus, although both forward and backward scoring provided acceptable results, forward scoring is used in the following example.
[0041] Hidden Markov Models: To derive a binary decision about anomalies, one method is to use a threshold and a score at each time point. t This would involve a step of comparing them separately. Another method, which leverages the sequential nature of the data, is to use a downstream hidden Markov model (HMM) with binary states and extract the maximum likelihood state trajectory using Viterbi decoding. This can smooth the output and help detect entire regions that appear anomalous, as shown in Figures 7 and 8.
[0042] Examples of telecommunication systems, machine architectures, and machine-readable media. Figure 5 is a block diagram of a computer system suitable for implementing the systems disclosed herein or for performing the methods disclosed herein. The machine in Figure 5 is shown as a standalone device suitable for implementing the above concepts. For the server embodiments described above, multiple such machines can be used operating in a data center, as part of a cloud architecture, etc. Not all of the illustrated functional units and devices are utilized in the server embodiments. For example, systems, devices, etc., that a user uses to interact with the server and / or cloud architecture may have screens, touchscreen inputs, etc., but servers often do not have screens, touchscreens, cameras, etc., and typically interact with the user through a connection system with appropriate input / output modes. Therefore, the following architectures should be understood as encompassing multiple types of devices and machines, and various embodiments may or may not be present in any particular device or machine depending on their form factor and purpose (for example, servers rarely have cameras, and wearables rarely contain magnetic disks). However, the illustrative description in Figure 5 is suitable for those skilled in the art to determine how the illustrated embodiments can be appropriately modified for specific devices, machines, etc., used, and how the embodiments described above can be implemented using appropriate combinations of hardware and software.
[0043] Although only a single machine is illustrated, the term “machine” should be interpreted to include any set of machines, individually or collectively, that perform one or more sets of instructions for carrying out any of the methodologies discussed herein.
[0044] An example of machine 500 includes at least one processor 502 (e.g., a controller, microcontroller, central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), advanced processing unit (APU), or a combination thereof), one or more memories such as main memory 504, static memory 506, or other types of memory communicating with each other via link 508. Link 508 may be a bus or other type of connection channel. Machine 500 may include any further embodiments such as a graphics display unit 510 including any type of display. The machine 500 may also include any other aspects such as an alphanumeric input device 512 (e.g., a keyboard, touchscreen, etc.), a user interface (UI) navigation device 514 (e.g., a mouse, trackball, touch device, etc.), a storage unit 516 (e.g., a disk drive or other storage device), a signal generating device 518 (e.g., a speaker), a sensor 522 (e.g., a global positioning sensor, accelerometer, microphone, camera, etc.), an output controller 528 (e.g., a wired or wireless connection for connecting and / or communicating with one or more other devices such as a universal serial bus (USB), near-field communication (NFC), infrared (IR), serial / parallel bus, etc.), and a network interface device 520 for connecting and / or communicating via one or more networks 526 (e.g., wired and / or wireless).
[0045] Various memories (i.e., 504, 506, and / or the memory of processor 502) and / or storage units 516 can store one or more sets of instructions and data structures (e.g., software) 524 that are implemented or utilized by any one or more methodologies or functional units described herein. When these instructions are executed by processor 502, they trigger various operations for implementing the disclosed embodiments.
[0046] Exemplary experiment: The LibriSpeech dataset was also used, along with artificial anomalies, to demonstrate the argument of concept and compare the merits of given design choices.
[0047] In the test data, additive pure sine wave tones of various frequencies and lengths were randomly placed within the dataset, resulting in a continuous anomalous region that constituted approximately 10% of the data.
[0048] The CPC hyperparameter is c t ∈R 256 , z t ∈R 512 And K=12 was used. Additionally, a separate learned transformation T with L=12 was used. l (z t ) are trained. Each consists of an MLP of size 64 with three hidden layers, with ReLU activation and no bias term, and these are trained on input z t This is applied as a residual multiplication mask with sigmoid activation. For co-training, pre-training is performed for 30 isolated epochs, and then λ=0.1 is selected. During isolated training, gradient flow from neural transformation to representation is prevented, i.e., both parts are trained separately.
[0049] LSTM and THOC were used as deep learning baselines, while OC-SVM and LOF were used as classical baselines. The latter were not specifically designed for time series, and therefore features were extracted from a fixed-size sliding window and applied to detection of these features. These features should be invariant under transformation. For speech data, it has been shown that Mel-scaled spectrograms should be strong domain-specific features.
[0050] Results: The anomaly scores predicted by the algorithm for each time step were compared individually. Considering the subseries of length 20480, the result was many scores and decisions per sample, and overall for the entire test set, approximately 10 8 This is how it works. For baseline algorithms that depend on smaller window sizes, the results from multiple subwindows are concatenated.
[0051] Using the receiver operating characteristic (ROC) curve reported in Figure 6, the LNT presented in this disclosure is compared to several baselines. The LNT outperformed all deep learning methods that depended on the learned representation and performed comparably to OCSVM, which depends on domain-specific features. While these features are particularly well suited to this particular anomaly that depends on pure frequency, the LNT presented in this disclosure can be applied to other tasks where strong domain-specific features are not present.
[0052] Figure 6 shows the effect of two different design choices on anomaly detection performance, while keeping other hyperparameters, specifically the isolated training and bias terms, fixed. Isolated single training results in poor anomaly detection. Therefore, co-training has a positive effect on the learned representation, resulting in not only less post-training loss but also better anomaly detection. Even a stronger effect can be explained by the presence of the bias term in LNT. The bias term makes the learned transformations somewhat independent of the input by design. This makes the performance on the self-monitoring task invariant under various inputs that would otherwise disrupt the semantic principles of this anomaly detection method.
[0053] Figure 6 is a graphical representation of the receiver operating characteristics (ROC) curve 600 for local neural transformation (LNT), showing the true positive rate (TPR) versus false positive rate (FPR) for different design choices.
[0054] This disclosure presents the LNT method, a novel anomalous region detection method for time series that combines representation learning and transformation learning. Furthermore, this disclosure provides results showing that this system and method can detect anomalous regions within a time series and also outperforms general deep learning baselines that acquire data representations rather than relying on domain-specific features.
[0055] Although not explicitly discussed in this disclosure, the disclosed systems and methods can be applied to other time-series datasets having anomalies as annotated, in particular to datasets with context-dependent anomalies, spatial or spatial-temporal datasets.
[0056] Neural transformation learning: Extends self-supervised anomaly detection methods to domains other than images. While promising results have recently been shown for tabular data, it lacks the scalability to detect anomalous regions within time series. The following provides a brief commentary on key considerations.
[0057] A trained transformation x has parameters trained by deterministic contrast loss (DCL) (see Equation 4). k :=T k Consider augmenting data D with (x). This loss promotes that the transformed samples are similar to their original samples with respect to cosine similarity h(·), while being dissimilar to other views of the same sample. This is motivated by semantics and diversity.
number
[0058] The ability to contrast these different views of the data is expected to reduce anomalous data and, through L reuse, lead to a deterministic anomaly score. One significant advantage over previous anomaly detection methods that rely on data augmentation is that the learned transformations are applicable even in domains where the techniques for manually designing semantic augmentation are unknown.
[0059] Viterbi decoding in Hidden Markov Models (HMMs): To derive binary decisions from the evaluation of a self-monitoring task, we consider a Hidden Markov Model (HMM) with two states. Therefore, the state s of the HMM t This corresponds to the condition whether the current time step t is part of the anomalous region. Then the score ` t This is considered as the emission of the generation probability model. The emission of probabilities is L after training. DDCL The Gaussian density is selected using the average of the values, which will be higher in the case of abnormal conditions. This design selection is characterized by bimodality. t It is motivated by the distribution of [something].
[0060] To further leverage the sequential nature of the data, state transition probabilities are selected to favor continuity. One method is to extract the maximum likelihood state trajectory from the network with the assistance of Viterbi decoding.
[0061] The results for samples randomly selected from the test set are reported in Figures 7 and 8. These plots show the L produced by LNT. DDCL The top row, below the loss, shows samples from a test set containing artificial anomalies, located in the yellow shaded area. The baseline, shown with the reference loss, is the output when the same sample without corruption is supplied. Note that this baseline is for visualization purposes only and is not an input to any method.
[0062] Viterbi decoding can overcome the effect of varying performance in the comparative self-monitoring task with respect to anomalous regions by extracting a continuous anomaly detection sequence with few leaps in change, as shown in the bottom row. For the selected samples, anomalous regions are detected almost completely.
[0063] Figure 7 is a graphical representation of an exemplary signal with injected anomalies 700. The corrupted region is the region where the anomaly was injected into the data stream. A graphical representation of the score 730 is shown.
number
[0064] Figure 8 is a graphical representation of an exemplary signal with injected anomalies 800. The corrupted region is the region where the anomaly was injected into the data stream. A graphical representation of the score 830 is shown.
number
[0065] When deploying machine learning models in practice, reliable anomaly detection is crucial, but remains challenging due to the lack of labeled data. The use of contrast learning approaches may also be common in settings of self-supervised representation learning. Here, a contrast anomaly detection approach is applied to images and configured to provide results that can be interpreted in the form of an anomaly segmentation mask. In this section of the disclosure, the use of a contrast predictive coding model is presented. The disclosure presents a patchy contrast loss that can be directly interpreted as anomaly scores and used in the creation of an anomaly segmentation mask. The resulting model has been tested for both anomaly detection and segmentation on the challenging MVTec-AD dataset.
[0066] An anomaly (or outlier, novelty, or out-of-distribution sample) is an observation that differs significantly from the majority of the data. Anomaly detection (AD) attempts to distinguish anomaly samples from those considered "normal" in the data. Detecting these anomalies is becoming increasingly important to improve the reliability of machine learning methods and enhance their applicability in real-world scenarios, such as automated industrial inspection, medical diagnosis, or autonomous driving. Typically, anomaly detection is treated as an unsupervised learning problem because labeled data is not generally available, and the goal is to develop methods that can detect anomalies that have not been seen before.
[0067] In this disclosure, contrast predictive coding (CPC) methods and systems can be applied to detect and segment anomalies in images. Furthermore, the system applies a representation loss (e.g., InfoNCE loss) that can be directly interpreted as anomaly scores. In this loss, patches from within the image are contrasted with each other and can be used to create an accurate anomaly segmentation mask.
[0068] Figure 9 is a flowchart of the anomaly region detection system 900 via contrast predictive coding (CPC). Image 902 is grouped into patches 904 and encoded by the encoder 910 for the creation of local latent representations along with the data within each patch 906. The encoder may include negative samples 908 when computing the local latent representations. The local latent representations pass through a local neural transform 912 to create a series of diversely transformed vector representations and dynamic deterministic contrast losses (DDCLs), both of which are used to create diverse semantic requirement scores associated with each patch.
[0069] In the schematic diagram of contrast predictive coding for anomaly detection and segmentation in images, after extracting (sub)patches from the input image, the encoded representation (z) from within the same image is obtained. t ,z t+k ) is a randomly coordinated representation (z t ,z j This is compared to subpatch x. The resulting InfoNCE loss is subpatch x t+k It is used to determine whether or not something is abnormal.
[0070] Figure 10 is an explanatory diagram of Image 1000, which shows the localization of anomalous regions for different classes in the MVTec-AD dataset. It includes the original input image 1002 (1002a, 1002b, 1002c, 1002d, 1002e, 1002f, 1002g), the corresponding InfoNCE loss values 1004 (1004a, 1004b, 1004c, 1004d, 1004e, 1004f, 1004g) (brighter shadows represent higher loss values), and a superimposed image of both 1006 (1006a, 1006b, 1006c, 1006d, 1006e, 1006f, 1006g). This demonstrates that the model presented in this disclosure consistently highlights anomalous regions across many classes.
[0071] To improve the performance of the CPC model for anomaly detection, this disclosure includes two adjustments. First, the negative sample setup during testing is adapted so that anomalous patches can only appear within positive samples. Second, the autoregressive portion of the CPC model is omitted. With these adjustments, the presented method achieves promising performance on real-world data such as the challenging MVTec-AD dataset.
[0072] Comparative learning: Self-supervised methods based on contrast learning work by having the model determine whether two (randomly or pseudo-randomly) transformed inputs originate from the same input sample or from two samples randomly selected from the entire dataset. Different transformations can be chosen depending on the domain and the downstream task. For example, in image data, random data augmentation such as random cropping or color jittering may be considered. In this disclosure, the contrast predictive coding model uses a temporal transformation. These approaches were evaluated by training a linear classifier on the resulting representation and measuring the performance achieved by this linear classifier on the downstream task.
[0073] Anomaly detection: Anomaly detection methods can be broadly categorized into three types: density-based, reconstruction-based, and discrimination-based methods. Density-based methods predict anomalies by estimating the probability distribution of data (e.g., GANs, VAEs, or flow-based models). Reconstruction-based methods are based on models trained for reconstruction purposes (e.g., autoencoders). Discrimination-based methods learn decision boundaries between anomalous and normal data (e.g., SVMs, one-class classification). The methods proposed in this disclosure include, but are not limited to, density-based methods for discriminative one-class purposes.
[0074] Contrast Predictive Coding: Contrast predictive coding is a self-supervised representation learning approach that leverages the data structure to force temporally close inputs to be coded similarly in latent space. This is achieved by having the model determine whether pairs of samples are made up of temporally close samples or randomly assigned samples. This approach can also be applied to static image data by dividing the image into patches and interpreting each row of the patch as a separate time step.
[0075] The CPC model uses a contrast loss function, which is based on noise contrast estimation and the latent representation (z) of the patch. t ) and their surrounding patches (c t+k This includes one known as InfoNCE, which is designed to optimize mutual information between ) and ).
number
[0076] CPC for anomaly detection: The CPC model is applied for anomaly detection and segmentation. To improve the performance of the CPC model in this setting, two adjustments to its architecture are considered (see Figure 9). First, the autoregressive model g arThis is omitted. As a result, the loss function changes as follows.
number
number
[0077] This adjustment results in a simpler model while still allowing for the learning of useful latent representations. Secondly, the setup of negative samples during testing is changed. In one implementation of the CPC model, random patches from within the same test batch are used as negative samples. However, this can result in negative samples containing anomalous patches, which can make it difficult for the model to detect anomalous patches within the positive samples. To avoid this, a novel sampling approach using a subset of non-anomalous images from the training set is considered.
[0078] During the testing phase, the loss function in equation (6) is equal to the image patch x t+k It can be used to determine whether something can be classified as an anomaly.
number
[0079] The threshold τ remains in implicit function form, and the lower region of the receiver operating characteristic curve (AUROC) can be used as a performance indicator. One solution is patch x t+kBy using each anomaly score, an anomaly segmentation mask is created. The other solution is to determine whether a sample is anomalous by averaging over the scores of patches within an image.
[0080] Experiment: The proposed contrastive prediction coding model for anomaly detection and segmentation was evaluated on the MVTec-AD dataset. This dataset contains high-resolution images of 10 objects and 5 textures with pixel-accuracy annotations, providing between 60 and 391 training images per class. Next, all images were randomly trimmed to 768×768 pixels and then split into patches of size 256×256, where each patch has a 50% overlap with adjacent patches. These patches were further split into sub-patches of size 64×64, also with a 50% overlap. These sub-patches were used with the InfoNCE loss (Figure 1) for anomaly detection. The trimmed images were flipped horizontally with a 50% probability during training.
[0081] Next, ResNet-18 v2 was used as the encoder g enc up to the third residual block. Separate models were trained from scratch for each class for 150 epochs with a batch size of 16 using the Adam optimizer with a learning rate of 1.5e-4. This model was trained and evaluated with grayscale images. To increase the accuracy of the InfoNCE loss as an indicator of anomalous patches, the model was applied in two directions. A shared encoder was used in both the direction from the top row to the bottom row and the direction from the bottom row to the top row of the image, but a separate W k was used for each direction.
[0082] Anomaly Detection The performance of the proposed model for anomaly detection was evaluated by averaging the top 5% InfoNCE loss values across all subpatches in a trimmed image and using this value to calculate the AUROC score. Table 1 shows an exemplary comparison, including systems without pre-trained feature extractors. The proposed CPC-AD model significantly improves upon kernel density estimation (KDE) and auto-coding (Auto) models. Although the performance of the model in this disclosure is inferior to the Cut-Paste model, CPC-AD provides a more generally applicable approach for anomaly detection. The Cut-Paste model relies heavily on randomly sampled artificial anomalies designed to resemble anomalies encountered in the dataset. Consequently, it cannot be applied to k-classes-out tasks where anomalies are semantically different from normal data.
[0083] [Table 1]
[0084] Table 1 shows the category-specific anomaly detection AUROC scores in the MVTec-AD test set. The proposed CPC-AD approach significantly outperforms kernel density estimation (KDE) and auto-coding (Auto) models, and its superior performance is due to the CutPaste model, which relies heavily on dataset-specific augmentation for training.
[0085] Abnormal segmentation: To evaluate the anomalous segmentation performance of the proposed CPC-AD model, the InfoNCE loss values at the subpatch level are upsampled for consistency with the ground truth level annotations at the pixel level. The InfoNCE losses of overlapping subpatches are averaged, and the resulting value is assigned to all affected pixels. An anomalous segmentation mask is created at the resolution of half the subpatch (32x32 pixels) of the same dimension as the cropped image (768x768).
[0086] Table 2 shows an illustrative comparison of the anomaly segmentation performance of the proposed CPC-AD methods. The best results on the MVTec-AD dataset are achieved by a wide range of models pre-trained on ImageNet, such as FCDD and PaDiM, or using additional artificial anomalies and ensemble methods such as CutPaste. The models presented in this disclosure are less complex and more general because they are trained from scratch and simply use the provided training data. The proposed CPC-AD approaches perform even better with two auto-encoding approaches (AE-SSIM, AE-L2) and a GAN-based approach (AnoGAN).
[0087] The model presented in this disclosure successfully generates accurate segmentation masks for large amounts of images across most classes (Figure 10). Even for classes with low pixel-level AUROC scores, such as transistors, the generated segmentation masks can be seen to adequately highlight anomalous input regions. This corresponds to the relatively high detection performance achieved by the CPC-AD method for this class (Table 1). These results suggest that some of the low segmentation scores (compared to the detection scores) may be due to small spatial deviations from the ground truth level. This effect may be exacerbated by the relatively low resolution of the segmentation masks created by this patch-level approach.
[0088] [Table 2]
[0089] Table 2 shows the anomalous segmentation, i.e., the pixel-level AUROC scores in the MVTec-AD test set for each category, for the presented CPC-AD models without using AE-SSIM, AE-L2, AnoGAN, CutPaste, or pre-trained feature extractors. (*) FCDD and PaDiM use pre-training to improve these results.
[0090] Overall, the CPC-AD model demonstrates that comparative learning can be applied not only to anomaly detection but also to anomaly segmentation. The proposed method performs well in anomaly detection tasks, with results that are second to none for most of the data. Furthermore, despite the fact that this model still outperforms modern segmentation methods, the generated segmentation masks offer a promising first step towards anomaly segmentation methods based on comparative learning.
[0091] Figures 11 to 16 illustrate exemplary embodiments, but the concepts of this disclosure may also apply to additional embodiments. Some exemplary embodiments include industrial applications in which embodiments may include video, weight, IR, 3D camera, and sound; power tool or home appliance applications in which embodiments may include torque, pressure, temperature, distance, or sound; medical applications in which embodiments may include ultrasound, video, CAT scan, MRI, or sound; robotic applications in which embodiments may include video, ultrasound, LIDAR, IR, or sound; and security applications in which embodiments may include video, sound, IR, or LIDAR. These embodiments may have diverse datasets; for example, a video dataset may include images, a LIDAR dataset may include point clouds, and a microphone dataset may include time series.
[0092] Figure 11 is a schematic diagram of a control system 1102 configured to control a vehicle or robot that may be at least partially autonomous. The vehicle includes sensors 1104 and actuators 1106. Sensors 1104 may include one or more wave energy-based sensors (e.g., charge-coupled devices CCD or video), radar, LiDAR, microphone arrays, ultrasound, infrared, thermal imaging, acoustic imaging, or other technologies (e.g., positioning sensors such as GPS). One or more of the one or more specific sensors may be integrated into the vehicle. Alternatively or in addition to the one or more specific sensors identified above, the control module 1102 may include a software module configured to determine the state of actuators 1104 at runtime.
[0093] In embodiments where the vehicle is at least partially autonomous, the actuator 1106 may be implemented in the vehicle's braking system, propulsion system, engine, drivetrain, or steering system. Actuator control commands may be determined to control the actuator 1106 so that the vehicle avoids collisions with detected objects. Detected objects may be classified according to what a classifier considers most likely to be, such as pedestrians or trees. Actuator control commands may be determined in relation to the classification. For example, the control system 1102 may classify (e.g., optical, acoustic, thermal) images or other inputs from the sensor 1104 into one or more background classes and one or more object classes (e.g., pedestrians, bicycles, vehicles, trees, traffic signs, traffic signals, road debris, or barrels / cones for a construction site) in order to avoid collisions with objects, and send control commands to the actuator 1106, in this implementation, to the braking system or propulsion system. In other examples, the control system 1102 can divide an image into one or more background classes and one or more sign classes (e.g., lane signs, guardrails, road edges, vehicle paths, etc.) and send control commands to actuators 1106, which are embodied in the steering system, to ensure that the vehicle avoids crossing signs and remains within its lane. In scenarios where hostile attacks may occur, the system described above may be further trained to better detect objects or to identify changes in lighting conditions or angles for sensors or cameras on the vehicle.
[0094] In other embodiments where the vehicle 1100 is at least partially autonomous, the vehicle 1100 may be a mobile robot configured to perform one or more functions such as flying, swimming, diving, and stepping. This mobile robot may be at least partially autonomous lawnmower or at least partially autonomous cleaning robot. In such embodiments, the actuator control command 1106 may be determined so that the propulsion unit, steering unit and / or brake unit of the mobile robot can be controlled so that the mobile robot can avoid collisions with identified objects.
[0095] In other embodiments, the vehicle 1100 is at least partially autonomous robot in the form of a horticultural robot. In such embodiments, the vehicle 1100 may use optical sensors as sensors 1104 to determine the state of plants in the environment adjacent to the vehicle 1100. The actuator 1106 may be a nozzle configured to spray chemicals. Depending on the identified species and / or identified state of the plant, the actuator control command 1102 may be determined to cause the actuator 1106 to spray an appropriate amount of appropriate chemicals onto the plant.
[0096] The vehicle 1100 may be at least partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances include washing machines, stoves, ovens, microwave ovens, or dishwashers. In such a vehicle 1100, the sensor 1104 may be an optical or acoustic sensor configured to detect the state of an object being processed by the household appliance. For example, if the household appliance is a washing machine, the sensor 1104 can detect the state of the laundry inside the washing machine. Actuator control commands may be determined based on the detected state of the laundry.
[0097] In this embodiment, the control system 1102 receives image and annotation information (optically or acoustically) from the sensor 1104. This information, along with a predetermined number of classes k and similarity scale stored in the system, is used to obtain the necessary information.
number
[0098] Figure 12 shows a schematic diagram of a control system 1202 configured to control a system 1200 (e.g., a manufacturing machine) such as a punch cutter, cutter, or gun drill in a manufacturing system 102, such as part of a production line. The control system 1202 may also be configured to control an actuator 1206 configured to control the system 1200 (e.g., a manufacturing machine).
[0099] Sensor 1204 of system 1200 (e.g., manufacturing machine) may be an optical sensor, acoustic sensor, or wave energy sensor such as a sensor array, configured to capture one or more characteristics of the manufactured product. Control system 1202 may be configured to determine the state of the manufactured product from one or more of the captured characteristics. Actuator 1206 may be configured to control system 1202 (e.g., manufacturing machine) for subsequent manufacturing steps of the manufactured product, depending on the determined state of the manufactured product 104. Actuator 1206 may be configured to control the functional part of Figure 12 (e.g., manufacturing machine) with respect to subsequent manufactured products of the system (e.g., manufacturing machine) depending on the determined state of the previously manufactured product.
[0100] In this embodiment, the control system 1202 receives image and annotation information from the sensor 1204 (for example, optically or acoustically). This information, along with a predetermined number of classes k and similarity scale stored in the system, is used to obtain the image and annotation information.
number
[0101] Figure 13 shows a schematic diagram of a control system 1302 configured to control a power tool 1300, such as an electric drill or screwdriver, which has at least a partially autonomous mode. The control system 1302 may also be configured to control an actuator 1306 configured to control the power tool 1300.
[0102] The sensor 1304 of the power tool 1300 may be a wave energy sensor, such as an optical or acoustic sensor, configured to capture one or more characteristics of the work surface and / or fasteners driven into the work surface. The control system 1302 may be configured to determine the state of the work surface and / or the state of the fasteners relative to the work surface from one or more of the captured characteristics.
[0103] In this embodiment, the control system 1302 receives image and annotation information from the sensor 1304 (for example, optically or acoustically). This information, along with a predetermined number of classes k and similarity scale stored in the system, is used to obtain the image and annotation information.
number
[0104] Figure 14 shows a schematic diagram of a control system 1402 configured to control an automated personal assistant 1401. The control system 1402 may also be configured to control an actuator 1406 configured to control the automated personal assistant 1401. The automated personal assistant 1401 may be configured to control household appliances such as a washing machine, stove, oven, microwave oven, or dishwasher.
[0105] In this embodiment, the control system 1402 receives image and annotation information from the sensor 1404 (for example, optically or acoustically). This information, along with a predetermined number of classes k and similarity scale stored in the system, is used to obtain the image and annotation information.
number
[0106] Figure 15 shows a schematic diagram of a control system 1502 configured to control a monitoring system 1500. The monitoring system 1500 may be configured to physically control access through door 252. Sensor 1504 may be configured to detect relevant scenes in determining whether access is permitted or not. Sensor 1504 may be an optical or acoustic sensor or sensor array configured to generate and transmit image data and / or video data. Such data may be used by the control system 1502 to detect human faces.
[0107] The monitoring system 1500 may be a surveillance system. In such an embodiment, the sensor 1504 may be a wave energy sensor such as an optical sensor, infrared sensor, or acoustic sensor configured to detect a scene under surveillance. The control system 1502 is configured to control the display 1508. The control system 1502 is configured to classify the scene, for example, to determine whether a scene detected by the sensor 1504 is suspicious or not. Perturbation objects may be used for a predetermined type of object detection to enable the system to identify such objects under suboptimal conditions (e.g., night, fog, rain, interference from background noise, etc.). The control system 1502 is configured to transmit actuator control commands to the display 1508 according to the classification. The display 1508 may be configured to adjust the displayed content in response to the actuator control commands. For example, the display 1508 may highlight objects deemed suspicious by the controller 1502.
[0108] In this embodiment, the control system 1502 receives image and annotation information (optically or acoustically) from the sensor 1504. This information, along with a predetermined number of classes k and similarity scale stored in the system, is used to obtain the necessary information.
number
[0109] Figure 16 shows a schematic diagram of a control system 1602 configured to control an imaging system 1600, such as an MRI machine, an X-ray imaging machine, or an ultrasound machine. Sensor 1604 may be, for example, an imaging sensor or an acoustic sensor array. The control system 1602 may be configured to determine the classification of all or part of the sensed image. The control system 1602 may be configured to determine or select actuator control commands according to the classification obtained by a trained neural network. For example, the control system 1602 may interpret a region of the image (optically or acoustically) sensed as potentially anomalous. In this case, the actuator control command may be determined or selected to display the image on the display 1606 and highlight the potentially anomalous region.
[0110] In this embodiment, the control system 1602 receives image and annotation information from the sensor 1604. This information, along with a predetermined number of classes k and similarity scale stored in the system, is used to obtain the necessary information.
number
[0111] Program code embodying the algorithms and / or methodologies described herein can be distributed individually or as a collection of various different forms of program products. This program code may be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of one or more embodiments. A computer-readable storage medium that is essentially non-temporary may include volatile and non-volatile, removable and non-removable tangible media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. A computer-readable storage medium may further include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, portable compact disc read-only memory (CD-ROM) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is readable by a computer. Computer-readable program instructions may be downloaded from a computer-readable storage medium to a computer, other types of programmable data processing equipment, or other devices, or they may be downloaded via a network to an external computer or external storage device.
[0112] Computer-readable program instructions stored on a computer-readable medium may be used to instruct a computer, other types of programmable data processing devices, or other devices to function in a particular manner, thereby causing the instructions stored on the computer-readable medium to generate a manufactured article containing instructions that perform functions, actions, and / or operations specified in a flowchart or diagram. In a given alternative embodiment, the functions, actions, and / or operations specified in the flowchart and diagram may be rearranged, processed sequentially, and / or simultaneously, in accordance with one or more embodiments. Furthermore, any flowchart and / or diagram may contain more or fewer nodes or blocks than those shown in accordance with one or more embodiments.
[0113] While the entirety of the present invention is illustrated by the description of various embodiments, and these embodiments have been described in great detail, it is not the applicant's intention to limit or restrict the scope of the appended claims in any way to such details. Additional advantages and modifications will be readily apparent to those skilled in the art. Accordingly, the present invention in its broader aspects is not limited to the specific details, representative apparatus and methods, and exemplary embodiments illustrated and described. Thus, such details can be used as a starting point without departing from the spirit or scope of the overarching concept of the invention.
Claims
1. A remote communication system including a controller, The aforementioned controller, Receive time-series data that will be grouped into patches, To obtain a local latent representation for each of the aforementioned patches, the data is encoded via the encoder parameters, For each of the aforementioned patches, the contrast predictive coding (CPC) loss is determined from the local latent representation. The local latent representation associated with each patch is transformed into a series of diversely transformed vector representations via at least two local neural transformations. From the aforementioned series of diversely transformed vector representations, the dynamic deterministic contrast loss (DDCL) is calculated. To obtain the updated parameters, the CPC loss and the DDCL are combined, The encoding parameters are updated using the updated parameters. To obtain diverse semantic requirement scores associated with each of the aforementioned patches, each of the series of diversely transformed vector representations is scored via the DDCL, In order to obtain the loss region, the various semantic requirement scores are smoothed, To obtain verified data, the data related to the loss region is masked, The remote communication system is operated based on the verified data. A remote communication system configured in such a way.
2. The remote communication system according to claim 1, wherein the controller is configured to smooth the diverse semantic request scores via Viterbi decoding.
3. The controller, in order to acquire the loss region, The remote communication system according to claim 1, configured to create a history of the diverse semantic request scores and to smooth the diverse semantic request scores by filtering out data having scores exceeding a threshold.
4. The controller calculates the CPC loss from the vector representation using the following equation [Math 1] It is configured to calculate according to the above, where [Math 2] is the expectation function, and the aforementioned z t is a local latent representation of the encoded data, where t is a time index, and W is a time index. k is a matrix of parameters, where k is the width of the time index, and f is the width of the time index. k (a, b)=f(a, W k b) is the aforementioned W k A comparison function between two arguments parameterized by L CPC This is the contrast predictive coding (CPC) loss, and the c t The remote communication system according to claim 1, wherein the contextual representation is [this].
5. The controller processes each of the series of diversely transformed vector representations using the following equation [Math 3] It is configured to score according to the above, where [Math 4] is the score, and the z t is the local latent representation of the encoded data, the t is the time index, and the W k is the matrix of parameters, the k is the width of the time index, and the [Math 5] This is a cosine similarity function, and the aforementioned c t The remote communication system according to claim 1, wherein the contextual representation is [this].
6. Steps include receiving time-series data to be grouped into patches, To obtain a local latent representation for each of the aforementioned patches, the data is encoded via the encoder parameters, For each of the aforementioned patches, the step of determining the representation loss from the local latent representation, The steps include transforming the local latent representation associated with each patch into a series of diversely transformed vector representations via at least two local neural transformations, The steps include determining the dynamic deterministic contrast loss (DDCL) from the series of diversely transformed vector representations, The steps include combining the representation loss and the DDCL to obtain updated parameters, The steps include updating the encoder parameters using the updated parameters, To obtain diverse semantic requirement scores associated with each of the aforementioned patches, the steps include scoring each of the series of diversely transformed vector representations via the DDCL, To obtain the loss region, the steps include smoothing the various semantic requirement scores, To obtain verified data, the steps include masking the data related to the loss region, The steps include outputting the verified data, An abnormal region detection method including
7. The anomaly detection method according to claim 6, wherein the step of encoding the data is performed via a recurrent neural network (RNN) or a convolutional neural network (CNN).
8. The anomaly region detection method according to claim 6, wherein the step of smoothing the diverse semantic requirement scores is performed via Viterbi decoding.
9. The step of smoothing the diverse semantic requirement scores is performed in order to obtain the loss region. The steps include creating a history of the various semantic requirement scores, A step of filtering out data with scores exceeding a threshold, The abnormal region detection method according to claim 6, including the method described in claim 6.
10. The anomaly region detection method according to claim 6, wherein the step of calculating the representation loss is a step of calculating the contrast predictive coding (CPC) loss.
11. The step of calculating the CPC loss from the vector representation is given by the following equation [Math 6] It is calculated according to the above, where [Number 7] is the expectation function, and the aforementioned z t is a local latent representation of the encoded data, where t is a time index, and W is a time index. k is a matrix of parameters, where k is the width of the time index, and f is the width of the time index. k (a, b)=f(a, W k b) is the aforementioned W k A comparison function between two arguments parameterized by L CPC This is the contrast predictive coding (CPC) loss, and the c t The abnormal region detection method according to claim 10, wherein the contextual representation is...
12. The step of calculating the DDCL loss from the various transformed vector representations is given by the following equation [Number 8] It is calculated according to the above, where [Number 9] is the expectation function, where t is the time index, k is the width of the time index, l is the index of the local neural transformation, and [Number 10] This is the score, and L DDCL The abnormal region detection method according to claim 6, wherein the loss is a dynamic deterministic contrast loss.
13. The step of scoring each of the aforementioned series of diversely transformed vector representations is given by the following equation [Math 11] It is calculated according to the above, where [Math 12] is the score, and the aforementioned z t is a local latent representation of the encoded data, where t is a time index, and W is a time index. k is a matrix of parameters, where k is the width of the time index, and [Number 13] This is a cosine similarity function, and the aforementioned c t The abnormal region detection method according to claim 6, wherein the contextual representation is...
14. The method for detecting an abnormal area according to claim 6, further comprising the step of operating a machine based on the verified data, wherein the time-series data is received from a sensor.
15. An anomaly detection system including a controller, The aforementioned controller, The first sensor receives time-series data that is grouped into patches. To obtain a local latent representation for each of the aforementioned patches, the data is encoded via the encoder parameters, For each of the aforementioned patches, the contrast predictive coding (CPC) loss is calculated from the local latent representation. The local latent representation associated with each of the aforementioned patches is transformed into a series of diversely transformed vector representations via at least two local neural transformations. From the aforementioned series of diversely transformed vector representations, the dynamic deterministic contrast loss (DDCL) is calculated. To obtain the updated parameters, the CPC loss and the DDCL are combined, The encoding parameters are updated using the updated parameters. To obtain diverse semantic requirement scores associated with each of the aforementioned patches, each of the series of diversely transformed vector representations is scored via the DDCL, In order to obtain the loss region, the various semantic requirement scores are smoothed, To obtain verified data, the data related to the loss region is masked, The machine is operated based on the verified data. An anomaly detection system configured as follows.
16. The anomaly region detection system according to claim 15, wherein the controller is configured to smooth the diverse semantic request scores via Viterbi decoding.
17. The controller, in order to acquire the loss region, An anomaly region detection system according to claim 15, configured to create a history of the diverse semantic request scores and to smooth the diverse semantic request scores by filtering out data having scores exceeding a threshold.
18. The controller calculates the CPC loss from the vector representation using the following equation [Number 14] It is configured to calculate according to the above, where [Number 15] is the expectation function, and the aforementioned z t is a local latent representation of the encoded data, where t is a time index, and W is a time index. k is a matrix of parameters, where k is the width of the time index, and f is the width of the time index. k (a, b)=f(a, W k b) is the aforementioned W k A comparison function between two arguments parameterized by L CPC This is the contrast predictive coding (CPC) loss, and the c t The abnormal region detection system according to claim 15, wherein the contextual representation is...
19. The abnormal area detection system according to claim 15, wherein the first sensor is an optical sensor, an in-vehicle sensor, or an acoustic sensor, and the machine is an autonomous vehicle.
20. The abnormal area detection system according to claim 15, wherein the first sensor is an RF sensor and the machine is a remote communication machine.
Citation Information
Patent Citations
Detector and detection method
JP2019215757A
Action selection neural network training using imitation learning in latent space
US20200104680A1