Image set anomaly detection with transducer encoder
By using a combination of a transformer encoder without position encoding and a convolutional neural network, the arrangement sensitivity problem of the set input depth network in image set abnormal detection is solved, and the arrangement unchanged abnormal detection is achieved, which improves the accuracy and applicability of the detection.
Patent Information
- Application Number
- CN202280102629.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-22
AI Technical Summary
The existing set input depth network has arrangement sensitivity problems in image set abnormal detection, resulting in different arrangement methods leading to different state representations and policy outputs, affecting the accuracy of autonomous driving decisions.
Using a combination of a transformer encoder and a convolutional neural network that does not perform position encoding, anomaly detection of the image set is achieved by extracting low-dimensional features in the image set and generating an estimate of the arrangement invariant.
It improves the accuracy of image set abnormal detection and is suitable for multi-instance learning, three-dimensional shape recognition, small sample image classification, intelligent manufacturing, medical image analysis, security and abnormal event detection in automatic transportation tools.
Smart Images

Figure CN120359548A_ABST
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure generally relate to anomaly detection in image sets. Background Art
[0002] An artificial neural network may include interconnected groups of artificial neurons (e.g., neuron models). An artificial neural network may be a computing device or represented as a method to be executed by a computing device. A convolutional neural network is a feedforward artificial neural network. A convolutional neural network may include a collection of neurons, where each neuron has a receptive field and together they tile the input space. Convolutional neural networks (CNNs), such as deep convolutional neural networks (DCNs), have numerous applications. Specifically, these neural network architectures are used in various technologies, such as image recognition, speech recognition, acoustic scene classification, keyword spotting, autonomous driving, and other classification tasks.
[0003] Many machine learning tasks, such as multi-instance learning, 3D shape recognition, and few-shot image classification, are defined over a collection of instances. Today's image data is often multi-view (e.g., individual objects are described from several perspectives or views), where each view highlights different characteristics of the object: multi-images from multiple cameras / sensors (signal-to-image, same / different times, localization) or multi-images from the same camera / sensor (signal-to-image, same / different times, localization).
[0004] A set-input deep network is a deep neural network that can learn to aggregate information across members of an input set. Set-input deep networks have recently attracted interest in computer vision and machine learning. This may be partly due to the increasing number of tasks defined over set inputs, such as meta-learning, clustering, and anomaly detection. These networks take any number of input samples and produce an output that is invariant to permutations of the input set. Permutation-invariant outputs are beneficial, for example, for autonomous driving decision-making. Autonomous driving decision-making uses real-valued representations, such as speed and localization. It cascades the perceptual information of the ego vehicle, surrounding vehicles, and the road into a state vector and then performs policy learning based on the vectorized state space. However, the information of surrounding vehicles must be permuted according to manually designed sorting rules because different permutations may lead to different state representations and different policy outputs. These set-input deep networks suffer from the permutation sensitivity problem and indicate that the information of surrounding vehicles must be permuted according to manually designed sorting rules because different permutations may lead to different state representations and policy outputs. Summary of the Invention
[0005] The present disclosure is set forth in the independent claims respectively. Some aspects of the present disclosure are described in the dependent claims.
[0006] In one aspect of the present disclosure, a processor-implemented method includes: receiving, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images. The method further includes: extracting low-dimensional features of each image in the set of images. The method also further includes: generating, by a transformer encoder, an estimate of each image in the set of images.
[0007] Another aspect of the present disclosure relates to an apparatus that includes components for receiving, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images. The apparatus further includes components for extracting low-dimensional features of each image in the set of images. The apparatus also further includes components for generating, by a transformer encoder, an estimate of each image in the set of images.
[0008] In another aspect of the present disclosure, a non-transitory computer-readable medium having non-transitory program code recorded thereon is disclosed. The program code is executed by a processor and includes program code for receiving, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images. The program code further includes program code for extracting low-dimensional features of each image in the set of images. The program code also further includes program code for generating, by a transformer encoder, an estimate of each image in the set of images.
[0009] Another aspect of the present disclosure relates to an apparatus that has a memory and one or more processors coupled to the memory. The processors are configured to receive, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images. The processors are further configured to extract low-dimensional features of each image in the set of images. The processors are also further configured to generate, by a transformer encoder, an estimate of each image in the set of images.
[0010] Additional features and advantages of the present disclosure will be described below. Those skilled in the art should understand that the present disclosure can be easily used as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. Those skilled in the art should also recognize that such equivalent structures do not depart from the teachings of the present disclosure set forth in the appended claims. The novel features that are considered to be characteristics of the present disclosure, both in terms of its organization and method of operation, together with further objectives and advantages, will be better understood when considered in conjunction with the following description taken in connection with the accompanying drawings. However, it should be clearly understood that each of the accompanying drawings is provided for illustrative and descriptive purposes only and is not intended as a definition of the limitations of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The features, nature, and advantages of the present disclosure will become more apparent when considered in conjunction with the detailed description set forth below, in which like reference numerals are always correspondingly identified in the accompanying drawings.
[0012] Figure 1Illustrate an example implementation of a neural network using a system - on - chip (SOC) (including a general - purpose processor) according to certain aspects of the present disclosure.
[0013] Figure 2A 、 Figure 2B and Figure 2C are diagrams illustrating neural networks according to aspects of the present disclosure.
[0014] Figure 2D is a diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.
[0015] Figure 3 is a block diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.
[0016] Figure 4 is a block diagram illustrating an exemplary software architecture that enables modularization of artificial intelligence (AI) functions.
[0017] Figure 5 is a block diagram illustrating an example architecture for image - set anomaly detection according to aspects of the present disclosure.
[0018] Figure 6 Illustrates a processor - implemented method for operating a neural network according to aspects of the present disclosure. Detailed Description
[0019] The following detailed description presented in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the described concepts may be practiced. To provide a thorough understanding of the various concepts, the detailed description includes specific details. However, it will be apparent to those skilled in the art that the concepts may be practiced without these specific details. In some instances, well - known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0020] Based on the teachings, those skilled in the art should recognize that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether implemented independently of any other aspect of the present disclosure or in combination with any other aspect. For example, a device may be implemented or a method may be practiced using any number of the aspects described. Additionally, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functionality, or a combination of structures and functionality that supplement or are different from the various aspects of the present disclosure as described. It should be understood that any aspect of the present disclosure disclosed may be embodied by one or more elements of a claim.
[0021] The term "exemplary" is used to mean "serving as an example, instance, or illustration". Any aspect described as "exemplary" need not be construed as superior or better than other aspects.
[0022] Although specific aspects are described, numerous variations and permutations of these aspects fall within the scope of the present disclosure. Although some benefits and advantages of the preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or purposes. Instead, the aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the accompanying drawings and the following description of the preferred aspects. The detailed description and the drawings are merely illustrative of the present disclosure and not limiting, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0023] Aspects of the present disclosure generally relate to anomaly detection. An anomaly is a data pattern having data characteristics different from normal instances. Anomaly detection generally aims to identify anomalies in a given dataset. That is, anomaly detection is the process of identifying unexpected event items in a dataset that are different from normal values. Anomaly detection can generally be applied to unlabeled data. Thus, anomaly detection can be used in a wide variety of applications. Visual anomaly detection has broad application prospects. For example, in the field of intelligent manufacturing, visual anomaly detection can be applied to defect detection. In the field of medical image analysis, it can be used to detect lesions in medical images; in the field of intelligent security, it can be used to detect abnormal events in videos. The emergence of autonomous vehicles provides an opportunity to apply visual anomaly detection in more dynamic applications.
[0024] A common application of image set anomaly detection is performed on a collection of images, where N - 1 images belong to the same class / have the same high-level features, and one image belongs to another class. Note that the class does not necessarily have to be related to the classes in a standard classification problem, but can be a combination of multiple features.
[0025] Set-input deep networks have recently drawn interest in computer vision and machine learning, partly due to the increasing number of tasks (such as clustering and anomaly detection) defined on set inputs (e.g., autonomous driving). Conventional set-input deep networks have been applied to tasks such as autonomous driving decision-making that use real-valued representations (such as speed and localization). However, conventional set-input deep networks suffer from the permutation sensitivity problem, and the information indicating surrounding vehicles must be permuted according to manually designed sorting rules because different permutations may result in different state representations and policy outputs.
[0026] Similarly, many conventional feedforward neural networks are also permutation variants. For example, recurrent neural networks (RNNs) are sensitive to the input order and are applied to sets only by assuming an order in the data.
[0027] To address these and other challenges, aspects of the present disclosure relate to a permutation-invariant anomaly detection model for an image set. According to aspects of the present disclosure, the anomaly detection model includes a transformer encoder that does not perform positional encoding.
[0028] Certain aspects of the subject matter described in the present disclosure can be implemented to realize one or more of the following potential advantages. In some examples, the described techniques can improve the accuracy of anomaly detection. Accordingly, aspects of the present disclosure can be beneficially applied to the fields of multi-instance learning, three-dimensional (3D) shape recognition, and few-shot image classification, as well as intelligent manufacturing, medical image analysis, security, detection of abnormal events in video, and autonomous vehicles.
[0029] Figure 1 An example embodiment of a system-on-chip (SOC) 100 is illustrated. The SOC may include a central processing unit (CPU) 102 or a multi-core CPU configured to perform anomaly detection in an image set of multiple images. Variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), latencies, frequency slot information, and task information may be stored in a storage block associated with a neural processing unit (NPU) 108, a storage block associated with the CPU 102, a storage block associated with a graphics processing unit (GPU) 104, a storage block associated with a digital signal processor (DSP) 106, storage block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or from memory block 118.
[0030] The SOC 100 may further include additional processing blocks customized for specific functions, such as the GPU 104, the DSP 106, a connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that can detect and recognize poses, for example. In one embodiment, the NPU 108 is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 may further include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which may include a global positioning system.
[0031] The SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general-purpose processor 102 may include code for receiving a set of images through an artificial neural network. The set of images includes a plurality of images. The general-purpose processor 102 may further include code for extracting low-dimensional features of each image in the set of images. The general-purpose processor 102 may additionally include code for generating an estimate of each image in the set of images through a transformer encoder.
[0032] Deep learning architectures can perform object recognition tasks by learning to represent the input at successively higher levels of abstraction in each layer, thus constructing a useful feature representation of the input data. In this way, deep learning addresses the main bottleneck of traditional machine learning. Before the emergence of deep learning, machine learning methods for object recognition problems might rely heavily on features designed by humans and might be combined with shallow classifiers. The shallow classifier can be a two-class linear classifier, for example, where the weighted sum of the feature vector components can be compared with a threshold to determine which class the input belongs to. The features designed by humans can be templates or kernels customized by engineers with domain expertise for a specific problem domain. In contrast, although deep learning architectures can learn to represent features similar to those that human engineers might design, they need to be trained. In addition, deep networks can learn to represent and recognize new types of features that humans might not have considered.
[0033] Deep learning architectures can learn a hierarchy of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power in specific frequencies. The second layer takes the output of the first layer as input and can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.
[0034] When applied to problems with a natural hierarchy, deep learning architectures can perform particularly well. For example, the classification of motorized vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher layers to recognize cars, trucks, and airplanes.
[0035] Neural networks can be designed to have a variety of connectivity patterns. In a feedforward network, information passes from lower layers to higher layers, where each neuron in a given layer communicates with neurons in a higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one block of input data in a sequence of input data blocks fed to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the identification of high-level concepts can assist in discerning specific low-level features of the input.
[0036] The connections between the layers of a neural network can be fully connected or locally connected. Figure 2A An example illustrating a fully connected neural network 202. In a fully connected neural network 202, neurons in the first layer can communicate their outputs to every neuron in the second layer, such that each neuron in the second layer will receive inputs from every neuron in the first layer. Figure 2B An example illustrating a locally connected neural network 204. In a locally connected neural network 204, neurons in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layers of a locally connected neural network 204 can be configured such that each neuron in the layer will have the same or a similar connectivity pattern, but the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can produce spatially distinct receptive fields in higher layers, because neurons in a given region of a higher layer can receive inputs that are tuned through training to the characteristics of a restricted portion of the total input to the network.
[0037] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example illustrating a convolutional neural network 206. A convolutional neural network 206 can be configured such that the connection strengths associated with the inputs to each neuron in the second layer are shared (e.g., 208). Convolutional neural networks can be well-suited to problems where the spatial location of the input is meaningful.
[0038] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2D A detailed example illustrating a DCN 200 designed to identify visual features from an image 226 input from an image capture device 230 (such as an on-vehicle camera). The DCN 200 of the current example can be trained to identify traffic signs and the numbers provided on the traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic signals.
[0039] The DCN 200 can be trained using supervised learning. During training, the DCN 200 can be presented with an image such as an image 226 of a speed limit sign, and then a forward pass can be computed to produce an output 222. The DCN 200 can include a feature extraction part and a classification part. When receiving the image 226, the convolutional layer 232 can apply a convolutional kernel (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 can be a 5x5 kernel that generates 28x28 feature maps. In this example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to the image 226 at the convolutional layer 232. The convolutional kernel can also be referred to as a filter or a convolutional filter.
[0040] The first set of feature maps 218 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can also be convolved by one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).
[0041] In Figure 2D the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Additionally, the first feature vector 224 is also convolved to generate a second feature vector 228. Each feature of the second feature vector 228 can include numbers corresponding to possible features of the image 226, such as "sign", "60", and "100". A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.
[0042] In this example, the probabilities of "sign" and "60" in the output 222 are higher than the probabilities of other numbers (such as "30", "40", "50", "70", "80", "90", and "100") in the output 222. Before training, the output 222 produced by the DCN 200 may be incorrect. Therefore, the error between the output 222 and the target output can be computed. The target output is the ground truth of the image 226 (e.g., "sign" and "60"). Then the weights of the DCN 200 can be adjusted such that the output 222 of the DCN 200 is closer to the target output.
[0043] To adjust the weights, the learning algorithm can compute the gradient vector of the weights. The gradient can indicate the amount by which the error will increase or decrease when the weights are adjusted. At the top layer, the gradient can directly correspond to the value of the weights connecting the activated neurons in the penultimate layer and the neurons in the output layer. In the lower layers, the gradient can depend on the value of the weights and the error gradients computed in the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be referred to as "backpropagation" because it involves a "backward pass" through the neural network.
[0044] In practice, the error gradient of the weights can be computed over a small number of examples such that the computed gradient is close to the true error gradient. This approximation method can be referred to as stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, new images can be presented to the DCN, and the forward pass through the network can produce an output 222 that can be considered an inference or estimate of the DCN.
[0045] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training data set. A DBN can be obtained by stacking layers of restricted Boltzmann machines (RBMs). An RBM is a type of artificial neural network that can learn a probability distribution from a set of inputs. Since an RBM can learn a probability distribution without information about the class to which each input should be classified, an RBM is typically used for unsupervised learning. Using a hybrid paradigm of unsupervised and supervised, the bottom RBM of a DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the inputs from the previous layer and the target classes) and can be used as a classifier.
[0046] A deep convolutional network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. A DCN has achieved state-of-the-art performance on many tasks. A DCN can be trained using supervised learning, where both the input targets and the output targets are known for many paradigms and are used to modify the weights of the network by using gradient descent methods.
[0047] A DCN can be a feedforward network. Additionally, as described above, the connections from the neurons in the first layer of a DCN to a set of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational burden of a DCN may be much smaller than that of a similarly sized neural network that includes recurrent or feedback connections.
[0048] The processing of each layer of a convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then the convolutional network trained on this input can be considered three-dimensional, where two spatial dimensions are along the axes of the image, and the third dimension captures color information. The output of the convolutional connection can be considered to form a feature map in subsequent layers, where each element in this feature map (e.g., 220) receives inputs from a certain range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature map can be further processed with non-linearities such as rectification, max(0,x). Values from adjacent neurons can be further pooled, which corresponds to downsampling, and can provide additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied through lateral inhibition between neurons in the feature map.
[0049] The performance of deep learning architectures can increase as more labeled data points become available or as computing power increases. Modern deep neural networks are typically trained with computing resources that are thousands of times the computing resources available to a typical researcher only fifteen years ago. New architectures and training paradigms can further enhance the performance of deep learning. Rectified linear units can reduce the training problem known as vanishing gradients. New training techniques can reduce overfitting and thus enable larger models to achieve better generalization. Encapsulation techniques can extract data within a given receptive field and further improve overall performance.
[0050] Figure 3 is a block diagram illustrating a deep convolutional network 350. Based on connectivity and weight sharing, the deep convolutional network 350 can include multiple different types of layers. As Figure 3 shown, the deep convolutional network 350 includes convolutional blocks 354A, 354B. Each of the convolutional blocks 354A, 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAXPOOL) 360.
[0051] The convolutional layer 356 can include one or more convolutional filters that can be applied to input data to generate a feature map. Although only two convolutional blocks 354A, 354B are shown, the present disclosure is not limited thereto, but instead, any number of convolutional blocks 354A, 354B can be included in the deep convolutional network 350 according to design preferences. The normalization layer 358 can normalize the output of the convolutional filters. For example, the normalization layer 358 can provide whitening or lateral inhibition. The max pooling layer 360 can provide spatially downsampled aggregation to achieve local invariance and dimensionality reduction.
[0052] For example, the parallel filter bank of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the deep convolutional network 350 can access other processing blocks that may be present on the SOC 100, such as the sensor processor 114 and the navigation module 120 dedicated to sensors and navigation, respectively.
[0053] The deep convolutional network 350 may also include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may also include a logistic regression (LR) layer 364. There are weights (not shown) to be updated between each layer 356, 358, 360, 362, 364 of the deep convolutional network 350. The output of each of these layers (e.g., 356, 358, 360, 362, 364) can be used as the input to the subsequent layer among these layers (e.g., 356, 358, 360, 362, 364) in the deep convolutional network 350 to learn a hierarchical feature representation from the input data 352 (e.g., image, audio, video, sensor data, and / or other input data) supplied at the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 can be a set of probabilities, where each probability is the probability that the input data includes a feature from a set of features.
[0054] Figure 4 is a block diagram illustrating an exemplary software architecture 400 that enables modularization of artificial intelligence (AI) functionality. According to aspects of the present disclosure, by using the architecture 400, various processing blocks (e.g., CPU 422, DSP 424, GPU 426, and / or NPU 428) of a system-on-chip (SoC) 420 (which may be similar to Figure 1 the SOC 100) can be supported to apply adaptive rounding for post-training quantization as disclosed for AI applications 402.
[0055] The AI application 402 can be configured to call functions defined in the user space 404, which can, for example, provide detection and recognition of a scene indicating the current operating location of the device. For example, the AI application 402 can configure the microphone and camera differently depending on whether the recognized scene is an office, a lecture hall, a restaurant, or an outdoor environment such as a lake. The AI application 402 can make a request for compiled program code associated with a library defined in an AI functional application programming interface (API) such as the SceneDetect API 406 to provide an estimate of the current scene. This request can ultimately rely on the output of a deep neural network configured to provide an inference response based on, for example, video and location data. The deep neural network can be a differential neural network configured to provide a scene estimate based on, for example, video and location data.
[0056] The runtime engine 408 (which can be the compiled code of a runtime framework) can further be accessible by the AI application 402. For example, the AI application 402 can cause the runtime engine 408 to request an inference, such as a scene estimate, triggered by a specific time interval or an event detected by the user interface of the application 402. When causing the runtime engine 408 to provide an inference response (e.g., to estimate a scene), the runtime engine can in turn send a signal to an operating system in the operating system (OS) space (such as the Linux kernel 412) running on the SoC 420. The operating system can then cause continuous quantization relaxation to be performed on the CPU 422, DSP 424, GPU 426, NPU 428, or some combination thereof. The CPU 422 can be directly accessed by the operating system, while the other processing blocks can be accessed through drivers (such as drivers 414, 416, or 418 for the DSP 424, GPU 426, or NPU 428, respectively). In an exemplary example, the deep neural network can be configured to run on a combination of processing blocks (such as the CPU 422, DSP 424, and GPU 426), or can run on the NPU 428.
[0057] As described, aspects of the present disclosure relate to anomaly detection of an image set using a transformer neural network.
[0058] Figure 5 is a block diagram illustrating an example architecture 500 for image set anomaly detection in accordance with aspects of the present disclosure. Refer to Figure 5, the example architecture 500 includes a Convolutional Neural Network (CNN) 504 and a Transformer encoder 506. Different from conventional Transformer encoders, the Transformer encoder 506 does not include a positional encoding module. The CNN 504 can be pre-trained to extract lower-dimensional features of an image. For example, the CNN 504 can receive an input with P features for N observations, where P >> N. Further, the CNN 504 can generate an output with O features for N observations, where O > N and O << P. The lower-dimensional features can be translation-invariant. In some aspects, the CNN 504 can be pre-trained on a dataset with many classes and multiple resolutions.
[0059] The example architecture can receive an image set 502. The image set 502 can include raw data from multiple cameras or sensors. The image set 502 can include an unordered collection of multiple images. For example, the image set 502 can include multiple images from multiple cameras (as shown in input set 502a), multiple images from sensors to image elements (as shown in input set 502b), or multi-view images from multiple cameras around a vehicle (as shown in input set 502c). The image set 502 can be received by the CNN 504. The CNN 504 can extract lower-dimensional features of each image in the image set 502. A multi-dimensional tensor including the extracted lower-dimensional features corresponding to each image in the image set 502 can be supplied to the Transformer encoder 506.
[0060] The Transformer encoder 506 can be configured to process data without positional encoding. The Transformer neural network is a deep learning model that uses multi-head attention and provides context information for any element within the input set. In some aspects, the Transformer encoder 506 can be an Efficiently Learning Encoder for Accurate Token Replacement (ELECTRA) small model. Of course, this is just an example, and other architectures can also be adopted, such as Bidirectional Encoder Representations from Transformers (BERT), Robustly Optimized BERT Approach (RoBERTa), XLNet, Transformer-XL, and Generative Pretrained Transformer (GPT) Transformer series.
[0061] The transformer encoder 506 vectorizes the multi-dimensional tensors of the lower-dimensional features of each image in the image set 502, and processes the vectors using a multi-head attention layer to generate a set 508 of logits. The multi-head attention layer of the transformer encoder 506 uses scaled dot products between queries and keys to find correlations and similarities between the images of the image set 502. For example, the example image set 502a shows multiple images of animals. The first four elements (e.g., 1 - 4) of the example image set 502a show images of foxes, while the element 5 image shows a different animal. The image in the first element of the example image set 502a can be used as a query, and can be compared with each of the other images in the other elements of the image set by taking the dot product of the vector corresponding to the first element and the vectors corresponding to each of the other elements to determine a set of keys. The keys can be used to provide attention weights. Then, the attention weights can be multiplied by each of the images in the example image set 502a to generate a set of values or scores for each of the elements in the example image set 502a. The scores for each element (e.g., 1 - 5) of the example image set 502a can be supplied to a multi-layer perceptron (MLP) of the transformer encoder 506 and processed to generate the logits for each element of the example image set 502a.
[0062] Logits represent probability values from 0 to 1. The set 508 of logits includes one logit for each image of the image set 502 to provide permutation-equivariant estimates. That is, if two elements in the elements of an image set (e.g., 502a) are switched in order, the output will be the same as when the elements are not switched in order.
[0063] A classifier can be implemented in addition to the transformer encoder 506 that does not perform positional encoding. Thus, the transformer encoder 506 can generate an estimate for each image set element.
[0064] An activation function can be applied to the set 508 of logits. For example, as Figure 5 shown, a softmax function 510 can be applied to the set 508 of logits. In some aspects, the classifier can be trained such that an anomalous image has the highest score / probability in the set 508 of logits. This is somewhat different from a conventional classification layer in that the softmax function 510 is applied to the images rather than the output classes in the classical sense. However, if two images swap their positions, their positions are swapped in the output softmax. Thus, the estimate is equivariant with respect to the input.
[0065] Figure 6 Illustrates a processor-implemented method 600 for image set anomaly detection according to aspects of the present disclosure. As Figure 6As shown, the processor-implemented method 600 includes, at block 602, receiving, by an artificial neural network, a set of images. The set of images includes a plurality of images. For example, as described with reference to Figure 5 Example architecture 500 may receive a set of images 502. The set of images 502 may include an unordered collection of a plurality of images. For example, the set of images 502 may include a plurality of images from multiple cameras (as shown in example set of images 502a), a plurality of images from sensors to image elements (as shown in example set of images 502b), or multi-view images from multiple cameras around a vehicle (as shown in example set of images 502c). The set of images 502 may be received by CNN 504.
[0066] The processor-implemented method 600 includes, at block 604, extracting low-dimensional features of each image in the set of images. For example, as described with reference to Figure 5 Example architecture 500 includes a convolutional neural network (CNN) 504 and a transformer encoder 506. The set of images 502 may be received by CNN 504. CNN 504 may extract high-level, low-dimensional features of each image in the set of images 502.
[0067] The processor-implemented method 600 includes: at block 606, generating, by a transformer encoder, an estimate for each image in the set of images. As described with reference to Figure 5 Transformer encoder 506 vectorizes the low-dimensional features of each image in the set of images 502 and uses a multi-head attention layer to process the vectors to generate a set of logits 508. The multi-head attention layer of transformer encoder 506 uses scaled dot product between queries and keys to find correlations and similarities between the images of the set of images. The set of logits 508 includes one logit for each image of the set of images 502 to provide permutation-equivariant estimates. That is, if two elements among the elements in the set of images (e.g., 502a) are switched in order, the output will be the same as when the elements are not switched in order.
[0068] A classifier may be implemented in addition to transformer encoder 506 that does not perform positional encoding. Thus, transformer encoder 506 may generate an estimate for each element of the set of images.
[0069] Specific implementation examples are provided in the following numbered clauses.
[0070] 1. A processor-implemented method, the processor-implemented method including:
[0071] Receiving, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images;
[0072] Extracting low-dimensional features of each image in the set of images; and generating, by a transformer encoder, an estimate for each image in the set of images.
[0073] 2. The method implemented by a processor according to Clause 1, the method implemented by the processor further comprising: detecting an anomaly in the image set based on the estimation of each image in the image set.
[0074] 3. The method implemented by a processor according to Clause 1 or 2, wherein the estimation includes a score for each image in the image set, and the anomaly in the image set comprises that the image corresponds to the maximum score in the image set.
[0075] 4. The method implemented by a processor according to any one of Clauses 1 to 3, the method implemented by the processor further comprising:
[0076] generating the log odds of each image in the image set by the transformer encoder; and
[0077] using an activation function to calculate the estimation of each image in the image set.
[0078] 5. The method implemented by a processor according to any one of Clauses 1 to 4, wherein the activation function is a softmax function, and the softmax function is applied to each image in the image set.
[0079] 6. The method implemented by a processor according to any one of Clauses 1 to 5, wherein the image set comprises an unordered set of multiple images.
[0080] 7. The method implemented by a processor according to any one of Clauses 1 to 6, wherein the ANN comprises a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.
[0081] 8. An apparatus, the apparatus comprising:
[0082] a memory; and
[0083] at least one processor coupled to the memory, the at least one processor
[0084] being configured to:
[0085] receive an image set by an artificial neural network (ANN), the image set comprising multiple images;
[0086] extract the low-dimensional features of each image in the image set; and generate an estimation of each image in the image set by the transformer encoder. 9. The apparatus according to Clause 8, wherein the at least one processor is further configured to be based on
[0087] Detecting an anomaly in the image set based on the estimation of each image in the image set. 10. The apparatus according to clause 8 or 9, wherein the estimation includes a score for each image in the image set, and the anomaly in the image set includes that the image corresponds to the maximum score in the image set.
[0088] 11. The apparatus according to any one of clauses 8 to 10, wherein the at least one processor is further configured to:
[0089] Generate the log odds of each image in the image set by the transformer encoder; and
[0090] Use an activation function to calculate the estimation of each image in the image set.
[0091] 12. The apparatus according to any one of clauses 8 to 11, wherein the activation function is a softmax function, and the at least one processor is further configured to apply the softmax function to each image in the image set.
[0092] 13. The apparatus according to any one of clauses 8 to 12, wherein the image set includes an unordered set of multiple images.
[0093] 14. The apparatus according to any one of clauses 8 to 13, wherein the ANN includes a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.
[0094] 15. A non-transitory computer-readable medium, on which program code is recorded, the program code being executed by a processor and including:
[0095] Program code for receiving an image set by an artificial neural network (ANN), the image set including multiple images;
[0096] Program code for extracting the low-dimensional features of each image in the image set;
[0097] And
[0098] Program code for generating an estimation of each image in the image set by a transformer encoder.
[0099] 16. The non-transitory computer-readable medium according to clause 15, wherein the program code further includes program code for detecting an anomaly in the image set based on the estimation of each image in the image set.
[0100] 17. The non-transitory computer-readable medium according to clause 15 or 16, wherein the estimation includes a score for each image in the image set, and the anomaly of the image set includes an image corresponding to the maximum score in the image set.
[0101] 18. The non-transitory computer-readable medium according to any one of clauses 15 to 17, wherein the program code further includes:
[0102] Program code for generating the log-odds of each image in the image set by the transformer encoder; and
[0103] Program code for calculating the estimation for each image in the image set using an activation function.
[0104] 19. The non-transitory computer-readable medium according to any one of clauses 15 to 18, wherein the activation function is a softmax function, and the softmax function is applied to each image in the image set.
[0105] 20. The non-transitory computer-readable medium according to any one of clauses 15 to 19, wherein the image set includes an unordered set of multiple images.
[0106] 21. The non-transitory computer-readable medium according to any one of clauses 15 to 20, wherein the ANN includes a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.
[0107] 22. An apparatus, the apparatus includes:
[0108] Components for receiving an image set by an artificial neural network (ANN), the image set including multiple images;
[0109] Components for extracting the low-dimensional features of each image in the image set; and components for generating an estimation for each image in the image set by a transformer encoder.
[0110] 23. The apparatus according to clause 22, the apparatus further includes: components for detecting an anomaly of the image set based on the estimation for each image in the image set.
[0111] 24. The apparatus according to clause 22 or 23, wherein the estimation includes a score for each image in the image set, and the anomaly of the image set includes an image corresponding to the maximum score in the image set.
[0112] 25. The apparatus according to any one of clauses 22 to 24, the apparatus further includes:
[0113] A component for generating the log odds of each image in the image set by the transducer encoder; and
[0114] A component for calculating the estimate for each image in the image set using an activation function.
[0115] 26. The apparatus according to any one of clauses 22 to 25, wherein the activation function is a softmax function, and the softmax function is applied to each image in the image set.
[0116] 27. The apparatus according to any one of clauses 22 to 26, wherein the image set includes an unordered set of multiple images.
[0117] 28. The apparatus according to any one of clauses 22 to 27, wherein the ANN includes a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.
[0118] In one aspect, the receiving component, the extracting component, and / or the generating component can be the CPU 102, the program memory associated with the CPU 102, the dedicated memory block 118, the fully connected layer 362, and / or the routing connection processing unit 216 configured to perform the functions. In another configuration, the foregoing components can be any module or any apparatus configured to perform the functions described by the foregoing components.
[0119] The various operations of the foregoing method can be performed by any suitable component capable of performing the corresponding functions. These components can include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, in the case where an operation is illustrated in the drawings, these operations can have corresponding paired components plus functional components with similar numbers.
[0120] As used, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, and so on. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Further, "determine" can include parsing, selecting, choosing, establishing, and so on.
[0121] As used, the phrase referring to "at least one of" a list of items means any combination of these items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c.
[0122] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the described functions. While a general-purpose processor may be a microprocessor, in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0123] The steps or algorithms of the methods described in connection with the present disclosure may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software modules may reside in any form of storage medium known in the art. Some examples of storage media that may be used include random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, hard disk, removable disk, CD-ROM, and the like. The software modules may include a single instruction, or many instructions, and may be distributed over several different code segments, among different programs, and across multiple storage media. The storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor.
[0124] The disclosed methods include one or more steps or acts for implementing the described methods. The method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or acts is specified, the order and / or use of specific steps and / or acts may be modified without departing from the scope of the claims.
[0125] The described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnected buses and bridges. The bus may link together various circuits, including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, via the bus, a network adapter, for example. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as a timing source, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and will not be described further herein.
[0126] The processor may be responsible for managing the bus and general processing, including executing software stored on the machine-readable medium. The processor may be implemented using one or more general-purpose processors and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuits capable of executing software. Software should be broadly interpreted to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. By way of example, the machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard disk drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging material.
[0127] In a hardware implementation, the machine-readable medium may be part of a processing system separate from the processor. However, as will be readily understood by those skilled in the art, the machine-readable medium or any part thereof may be external to the processing system. By way of example, the machine-readable medium may include a transmission line, a carrier modulated by data, and / or a computer product separate from the device, all of which may be accessed by the processor via the bus interface. Alternatively or in addition, the machine-readable medium or any part thereof may be integrated into the processor, such as in the case of having a cache and / or a general register file. Although the various components discussed may be described as having a specific location, such as local components, they may also be configured in various ways, such as some components being configured as part of a distributed computing system.
[0128] The processing system can be configured as a general-purpose processing system having one or more microprocessors providing processor functionality and an external memory providing at least a portion of the machine-readable medium, all of these components being linked together via an external bus architecture to other support circuitry. Alternatively, the processing system can include one or more neuromorphic processors for implementing the described neuron models and nervous system models. As yet another alternative, the processing system can be implemented with an application-specific integrated circuit (ASIC) having a processor, bus interface, user interface, support circuitry, and at least a portion of the machine-readable medium integrated on a single chip, or with one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic components, discrete hardware components, or any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality of the processing system depending on the particular application and overall design constraints imposed on the overall system.
[0129] The machine-readable medium can include a plurality of software modules. These software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. The software modules can include a sending module and a receiving module. Each software module can reside in a single storage device or be distributed across multiple storage devices. By way of example, when a triggering event occurs, the software module can be loaded from a hard disk drive into RAM. During the execution of the software module, the processor can load some of the instructions into a cache to improve access speed. One or more cache lines can then be loaded into the general register file for execution by the processor. When reference is made hereinafter to the functionality of a software module, it will be understood that such functionality is implemented by the processor when executing instructions from that software module. In addition, it should be understood that aspects of the present disclosure result in improvements to the functionality of a processor, computer, machine, or other system implementing such aspects.
[0130] If implemented in software, the functions can be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and that is accessible by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and (Blu- ) disc, where disks typically reproduce data magnetically, while discs reproduce data optically with a laser. Thus, in some aspects, the computer-readable medium may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, the computer-readable medium may include transitory computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0131] Accordingly, some aspects may include a computer program product for performing the operations of rendering. For example, such a computer program product may include a computer-readable medium having (and / or encoded with) instructions that can be executed by one or more processors to perform the described operations. For some aspects, the computer program product may include packaging materials.
[0132] Furthermore, it should be understood that modules and / or other suitable components for performing the described methods and techniques can be downloaded and / or otherwise obtained by a user terminal and / or base station, where applicable. For example, such devices can be coupled to a server to facilitate the transfer of components for performing the described methods. Alternatively, the various methods are provided via a storage component (e.g., RAM, ROM, physical storage media such as a compact disc (CD) or floppy disk) such that once the storage component is coupled to or provided to the device, the user terminal and / or base station can obtain the various methods. Additionally, any other suitable technology can be utilized that is adapted to provide the described methods and techniques to the device.
[0133] It should be understood that the claims are not limited to the exact configurations and components illustrated above. Various modifications, variations, and alterations can be made to the arrangements, operations, and details of the methods and apparatuses described above without departing from the scope of the claims.
Claims
1. A processor-implemented method, the processor-implemented method comprising: Receiving, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images; Extracting low-dimensional features of each image in the set of images; And Generating, by a transformer encoder, an estimate for each image in the set of images.
2. The processor-implemented method according to claim 1, the processor-implemented method further comprising: Detecting an anomaly in the set of images based on the estimate for each image in the set of images.
3. The processor-implemented method according to claim 2, wherein the estimate includes a score for each image in the set of images, and the anomaly in the set of images includes an image corresponding to the maximum score in the set of images.
4. The processor-implemented method according to claim 1, the processor-implemented method further comprising: Generating, by the transformer encoder, a logit for each image in the set of images; And Using an activation function to calculate the estimate for each image in the set of images.
5. The processor-implemented method according to claim 4, wherein the activation function is a softmax function, and the softmax function is applied to each image in the set of images.
6. The processor-implemented method according to claim 1, wherein the set of images includes an unordered set of a plurality of images.
7. The processor-implemented method according to claim 1, wherein the ANN includes a convolutional neural network (CNN), the CNN being pre-trained to extract the low-dimensional features of each image in the set of images.
8. An apparatus, the apparatus comprising: A memory; And At least one processor coupled to the memory, the at least one processor being configured to: Receive, by an artificial neural network (ANN), a set of images, the set of images including a plurality of images; Extract low-dimensional features of each image in the set of images; And Generate, by a transformer encoder, an estimate for each image in the set of images.
9. The apparatus according to claim 8, wherein the at least one processor is further configured to detect an anomaly in the set of images based on the estimate for each image in the set of images.
10. The apparatus according to claim 9, wherein the estimate includes a score for each image in the set of images, and the anomaly in the set of images includes an image corresponding to the maximum score in the set of images.
11. The apparatus according to claim 8, wherein the at least one processor is further configured to: Generate, by the transformer encoder, a logit for each image in the set of images; and Use an activation function to calculate the estimate for each image in the set of images.
12. The apparatus according to claim 11, wherein the activation function is a softmax function, and the at least one processor is further configured to apply the softmax function to each image in the set of images.
13. The apparatus according to claim 8, wherein the set of images includes an unordered set of a plurality of images.
14. The apparatus according to claim 8, wherein the ANN includes a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.
15. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by a processor and including: program code for receiving an image set by an artificial neural network (ANN), the image set including a plurality of images; program code for extracting low-dimensional features of each image in the image set; and program code for generating an estimate of each image in the image set by a transformer encoder.
16. The non-transitory computer-readable medium according to claim 15, wherein the program code further includes program code for detecting an anomaly in the image set based on the estimate of each image in the image set.
17. The non-transitory computer-readable medium according to claim 16, wherein the estimate includes a score for each image in the image set, and the anomaly in the image set includes an image corresponding to the maximum score in the image set.
18. The non-transitory computer-readable medium according to claim 15, wherein the program code further includes: program code for generating the log odds of each image in the image set by the transformer encoder; and program code for calculating the estimate of each image in the image set using an activation function.
19. The non-transitory computer-readable medium according to claim 18, wherein the activation function is a softmax function, and the softmax function is applied to each image in the image set.
20. The non-transitory computer-readable medium according to claim 15, wherein the image set includes an unordered set of a plurality of images.
21. The non-transitory computer-readable medium according to claim 15, wherein the ANN includes a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.
22. An apparatus, the apparatus including: means for receiving an image set by an artificial neural network (ANN), the image set including a plurality of images; means for extracting low-dimensional features of each image in the image set; and means for generating an estimate of each image in the image set by a transformer encoder.
23. The device according to claim 22, wherein the device further comprises: means for detecting an anomaly in the image set based on the estimate of each image in the image set.
24. The apparatus according to claim 23, wherein the estimate includes a score for each image in the image set, and the anomaly in the image set includes an image corresponding to the maximum score in the image set.
25. The apparatus according to claim 22, the apparatus further including: means for generating the log odds of each image in the image set by the transformer encoder; and means for calculating the estimate of each image in the image set using an activation function.
26. The apparatus according to claim 25, wherein the activation function is a softmax function, and the softmax function is applied to each image in the image set.
27. The apparatus according to claim 22, wherein the image set includes an unordered set of multiple images.
28. The apparatus according to claim 22, wherein the ANN includes a convolutional neural network (CNN), and the CNN is pre-trained to extract the low-dimensional features of each image in the image set.