SYSTEM AND METHOD FOR SCENE ANNOMALY DETECTION - Patent application

Video anomaly detection systems address the inefficiencies of traditional sensor-based failure detection in factory automation by using video processing and feature vector analysis to identify deviations from normal behavior, enhancing detection efficiency and reducing costs.

JP7682394B2Active Publication Date: 2025-05-23MITSUBISHI ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024538514
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-08
Filing Date
2022-07-08
Publication Date
2025-05-23
Estimated Expiration
2042-07-08

AI Technical Summary

Technical Problem

Existing technologies for detecting failures in factory automation systems are inefficient and costly, as they rely on customized sensors that may not detect unexpected failures and require extensive installation and configuration.

Method used

The implementation of video anomaly detection systems that use normal video models to identify anomalies in factory automation scenarios by processing video frames and generating motion and appearance feature vectors to detect deviations from expected behavior.

Benefits of technology

This approach allows for the efficient and cost-effective detection of both expected and unexpected failures in factory automation systems, reducing downtime and operational costs compared to traditional sensor-based methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682394000006
    Figure 0007682394000006
  • Figure 0007682394000007
    Figure 0007682394000007
  • Figure 0007682394000008
    Figure 0007682394000008
Patent Text Reader

Abstract

A system for detecting anomalies in videos of factory automation scenes is disclosed, the system can receive the video, receive a set of training feature vectors derived from spatio-temporal regions of the training video associated with one or more training feature vectors, segment the video into a plurality of sequences of video volumes, generate a sequence of binary difference images for each of the video volumes, generate input feature vectors including an input motion feature vector defining a temporal variation in counts of the predetermined patterns for each of the video volumes by counting occurrences of each of the predetermined patterns of pixels in each binary difference image for each of the video volumes, generate a set of distances based on the generated input feature vectors and the set of training feature vectors, and detect the anomaly based on the generated set of distances.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to image processing, and more specifically to video anomaly detection in a scene. [Background technology]

[0002] In recent years, automation in the workplace has been used in a variety of applications to reduce the cost of processes (such as manufacturing processes) for developing final products. As an example, factory automation can be a factory assembly line that includes robots, conveyors, and other machines that can automatically pick and assemble raw materials into more complex devices and products. In some cases, the factory assembly line can have a problem (e.g., a breakdown) that requires human intervention to fix. If this breakdown is not identified in a timely manner, it can lead to a larger problem and ultimately to a longer downtime.

[0003] Currently, there are various technologies that aim to detect failures at a fixed time. These available technologies use customized sensors to detect failures associated with a factory assembly line. As an example, the customized sensors may be manually installed in specific locations where failures are known to occur. However, if an unexpected failure occurs, these available technologies may not detect the unexpected failure because the customized sensors were not installed in the location where the unexpected failure occurred and / or the customized sensors are not configured to detect the unexpected failure. Furthermore, installing multiple customized sensors to detect both expected and unexpected failures can be a time-consuming and costly process.

[0004] Therefore, there is a need for a system that detects expected and unexpected failures associated with such automation in an efficient and feasible manner. Summary of the Invention

[0005] To solve the above problem, an objective of some embodiments is to adapt anomaly detection to video anomaly detection from video cameras supervising automation scenarios such as factory assembly lines. An "anomaly" as used herein may correspond to a fault associated with an automation scenario. As an example, faults associated with a factory assembly line may include an abnormal orientation of a robot arm, an unexpected stop of a conveyor, etc. In video anomaly detection, some embodiments aim to automatically detect as an anomaly an activity (e.g., a machine operation) in a portion of a video when the former activity differs from an activity seen in a normal video of the same scene. Thus, the detected anomalies include both expected and unexpected faults associated with factory assembly, since all activities that differ from the activity of the normal video are detected as an anomaly. Furthermore, video anomaly detection can reduce the cost of detecting anomalies in factory automation compared to techniques that aim to detect anomalies by installing customized sensors. As an example, video anomaly detection can be less costly than these techniques because it does not use customized sensors to detect anomalies.

[0006] To detect anomalies, some embodiments aim to build a model using normal videos. Hereinafter, "normal video" and "training video" may be used interchangeably. As used herein, a "training video" may correspond to a video that includes a set of video frames that correspond to normal operation of a machine performing a task in an automation scenario. In an example embodiment, a model may be built for a training video by dividing the training video into multiple spatiotemporal regions and learning a separate model for each spatial region of the video. By way of example, each spatiotemporal region may be defined by a video bounding box. For example, the video bounding box includes spatial and temporal dimensions to divide the training video into multiple spatiotemporal regions. Furthermore, feature vectors may be calculated for short sequences of the training video, and all "unique" feature vectors occurring in each spatial region may be stored as a model. A short sequence of videos in a spatial region (i.e., a spatiotemporal region) may be referred to as a sequence of training image patches. The unique feature vectors occurring in each spatial region may be referred to as "exemplars."

[0007] It is also an objective of some embodiments to compute a feature vector for each sequence of training image patches in such a way that the computed feature vector is simple yet accurate enough to represent anomalous patterns of time-series motion data in an automation scenario. To this end, some embodiments use a motion feature vector that defines the temporal variation of counts of predefined patterns. As an example, the predefined patterns may indicate different types of motion captured in the video. Using counts of predefined patterns of motion instead of these patterns themselves simplifies the motion feature vector while preserving some motion information. Some embodiments use binary difference images of consecutive frames of the training video to compute the motion feature vector. The binary difference image shows a threshold difference of these two frames, which indicates the relative motion captured by the two frames. Patterns formed by pixels of the binary difference image that are greater than a threshold "1" or smaller than a threshold "0" are the predefined patterns that are counted by the motion feature vector. Furthermore, using the temporal variation of the counts allows for taking into account motion over time, which is advantageous for factory automation. Furthermore, considering only predefined patterns and counts of predefined patterns allows for having a fixed size of the motion feature vector, which is favorable for distance-based anomaly detection.

[0008] During the control of the machine performing the task, an input test video, for example from the same fixed camera used to acquire the training video, is processed in the same manner to generate input motion feature vectors that are compared with the motion feature vectors obtained from the training video for the corresponding spatial region to detect anomalies. In an example embodiment, the shortest distance (e.g., Euclidean distance) between each input motion vector and the training motion vector of the same spatial region may be calculated. Furthermore, the calculated shortest distance may be compared with an anomaly detection threshold to detect anomalies. As an example, an anomaly in the input video may be detected if at least one calculated shortest distance is greater than the anomaly detection threshold. Using the simplified feature vectors (i.e., the motion feature vectors from the training video and the test video) makes it possible to detect anomalies in the input test video in a feasible manner.

[0009] Some embodiments use an appearance feature vector representing appearance information in the video in addition to the motion feature vector. The appearance information is optional but convenient to complement the motion feature vector to take into account the situation of motion changes while detecting anomalies without requiring additional sensors. In this way, the hardware requirements for anomaly detection can be reduced. In one embodiment, histogram of oriented gradients (HoG) features calculated for an image patch of a video volume may be used as appearance information. In another embodiment, a binary difference image calculated for two consecutive image patches of a video volume may be used as appearance information.

[0010] Accordingly, one embodiment discloses a system for detecting anomalies in a video of a factory automation scene, the system comprising a processor and a memory having stored instructions that, when executed by the processor, cause the system to: receive an input video of a scene including a machine performing a task; receive a set of training feature vectors derived from spatio-temporal domains of a training video of normal operation of the machine performing the task, the spatio-temporal domains being associated with one or more training feature vectors, each training feature vector including a motion feature vector defining a temporal variation in counts of a predetermined pattern; and further, convert the input video into a plurality of sequences of video volumes corresponding to spatial and temporal dimensions of the spatio-temporal domains of the training video, the sequences of image patches defined by the spatial and temporal dimensions of the corresponding spatio-temporal domains of the video volumes. generating a sequence of binary difference images for each of the video volumes by determining a binary difference image for each successive pair of image patches in the sequence of image patches for each of the video volumes; generating input feature vectors including an input motion feature vector defining a temporal variation in the counts of the predetermined patterns for each of the video volumes by counting occurrences of each of the predetermined patterns in each binary difference image for each of the video volumes; generating a set of distances by calculating a shortest distance between the input feature vector for each of the video volumes and a training feature vector associated with a corresponding spatial region in the scene; and comparing each distance in the set of distances to an anomaly detection threshold to detect anomalies in the input video of the scene.

[0011] Another embodiment discloses a method for detecting anomalies in videos of factory automation scenes, the method including receiving an input video of a scene including a machine performing a task, and receiving a set of training feature vectors derived from spatio-temporal domains of a training video of normal operation of the machine performing the task, the spatio-temporal domains being associated with one or more training feature vectors, each training feature vector including a motion feature vector defining a temporal variation in counts of a predetermined pattern, the method further including partitioning the input video into a plurality of sequences of video volumes corresponding to spatial and temporal dimensions of the spatio-temporal domains of the training video, such that the video volumes include a sequence of image patches defined by the spatial and temporal dimensions of the corresponding spatio-temporal domain, and partitioning each of the video volumes into a plurality of sequences of video volumes corresponding to spatial and temporal dimensions of the corresponding spatio-temporal domain. the input feature vectors include an input motion feature vector defining a temporal variation in the counts of the predetermined patterns for each of the video volumes by counting occurrences of each of the predetermined patterns in each of the binary difference images for each of the video volumes; generating a set of distances by calculating a shortest distance between the input feature vectors for each of the video volumes and a training feature vector associated with a corresponding spatiotemporal region in the scene; and comparing each distance in the set of distances to an anomaly detection threshold to detect an anomaly in the input video of the scene.

[0012] Another embodiment discloses a system for detecting anomalies in videos of factory automation scenes, the system comprising a processor and a memory having instructions stored thereon, which when executed by the processor cause the system to: receive an input video of a scene including a machine performing a task; and receive a set of training feature vectors derived from spatio-temporal domains of a training video of normal operation of the machine performing the task, the spatio-temporal domains being associated with one or more training feature vectors, each training feature vector consisting of an appearance feature vector and a motion feature vector; partitioning the input video into a plurality of sequences of video volumes corresponding to spatial and temporal dimensions of the spatio-temporal domains of the training video, such that the video volumes include a sequence of image patches defined by the spatial and temporal dimensions of the corresponding spatio-temporal domains; and generating a sequence of binary difference images for each of the video volumes by determining a binary difference image for each successive pair of image patches in the sequence of image patches of each of the video volumes; the input video of the scene is generated by counting occurrences of each of the predetermined patterns of pixels in each binary difference image for each of the video volumes, thereby generating an input motion feature vector defining a temporal variation in the counts of the predetermined patterns for each of the video volumes; calculating an input appearance feature vector for each of the video volumes, wherein the input appearance feature vector represents a pattern of pixels occurring in the video volume; generating a set of motion distances by calculating a shortest distance between the input motion feature vector of each of the video volumes and a motion feature vector of the training feature vectors associated with a corresponding spatiotemporal region in the scene; generating a set of appearance distances by calculating a shortest distance between the input appearance feature vector of each of the video volumes and an appearance feature vector of the training feature vectors associated with a corresponding spatial region in the scene; and comparing each motion and appearance distance of the set of motion and appearance distances to at least one anomaly detection threshold to detect anomalies in the input video of the scene. [Brief description of the drawings]

[0013] [Figure 1] FIG. 1 illustrates an overview of a system for detecting anomalies in a video, according to some embodiments of the present disclosure. [Figure 2A] FIG. 2 illustrates a flowchart for detecting anomalies in an input video of a factory automation scene according to some embodiments of the present disclosure. [Figure 2B] FIG. 2 illustrates a pipeline for splitting an input video into multiple sequences of video volumes according to some embodiments of the present disclosure. [Figure 2C] FIG. 2 shows a schematic diagram for generating a sequence of binary difference images of a video volume according to some embodiments of the present disclosure. [Figure 2D] FIG. 2 illustrates a pipeline for generating an input feature vector for a particular video volume according to some embodiments of the present disclosure. [Diagram 3] FIG. 2 illustrates a flowchart for generating a set of training feature vectors for training a system to detect anomalies in a video, according to some embodiments of the present disclosure. [Figure 4] FIG. 2 shows a schematic diagram for computing an appearance feature vector of a video patch according to some embodiments of the present disclosure. [Figure 5A] FIG. 13 illustrates a flowchart for detecting anomalies in a video of a factory automation scene according to some other embodiments of the present disclosure. [Figure 5B] FIG. 13 illustrates a flowchart for detecting anomalies in a video of a factory automation scene according to some other embodiments of the present disclosure. [Figure 6] FIG. 1 illustrates a working environment of a system for detecting anomalies in a factory automation process according to some embodiments of the present disclosure. [Figure 7]FIG. 1 shows an overall block diagram of a system for detecting anomalies in videos of factory automation scenes according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] Description of the embodiments

[0015] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, the apparatus and methods are shown in block diagram form solely to avoid obscuring the present disclosure.

[0016] As used in the specification and claims, the terms "for example," "for instance," and "such as," as well as "comprising," "having," "including," and other forms of these verbs, when used in conjunction with a list of one or more components or other items, should be construed as open-ended, meaning that the list should not be considered to exclude further components or items. The term "based on" means based at least in part on. Furthermore, it should be understood that the style and terminology used herein are for purposes of illustration and should not be considered to be limiting. Any headings used herein are for convenience only and should not be considered to have any legal or limiting effect.

[0017] FIG. 1 illustrates an overview of a system 100 for detecting anomalies in a video 102, according to some embodiments of the present disclosure. According to an embodiment, the video 102 may include a set of video frames corresponding to an automation scene, such as a factory automation scene. As used herein, "factory automation" may be a process of automatically manufacturing a complex system using one or more machines. For example, factory automation may be realized as a factory assembly line, where various raw materials of a complex system are automatically assembled using one or more machines to manufacture the complex system. As an example, the complex system may be a vehicle, etc. The one or more machines may include, but are not limited to, a robotic arm, a conveyor, etc. As used herein, a "factory automation scene" may correspond to a scene showing one or more machines performing tasks to realize the factory automation.

[0018] During a factory automation process, in some scenarios, one or more machines may encounter a fault. For example, a fault associated with one or more machines may include, but is not limited to, an abnormal orientation of a robot arm, an unexpected stop of one or more machines during a factory automation process, etc. Hereinafter, "failure of one or more machines" and "fault" may be used interchangeably.

[0019] According to an example embodiment, the system 100 may be configured to detect anomalies in a factory automation process. In such a case, the system 100 may detect anomalies in the factory automation process using a video 102 of a factory automation scene. To that end, the system 100 may acquire the video 102 from an imaging device monitoring one or more machines performing a task on a factory floor. Thus, if one or more machines experience a fault, the fault is reflected in the video 102 acquired from the imaging device. By way of example, the imaging device may be a camera, a video player, etc. The system 100 may process the video 102 to detect anomalies in the factory automation process.

[0020] Additionally, the system 100 may provide an output 104 in response to detecting the anomaly. In one embodiment, the output 104 may be a control signal for controlling one or more machines to stop the abnormal activity. In another embodiment, the output 104 may be a notification to a user to stop the abnormal activity. The system 100 may detect anomalies in the video 102 of the factory automation scene as further described with reference to FIG. 2A.

[0021] 2A illustrates a flowchart 200 for detecting anomalies in an input video of a factory automation scene, according to some embodiments of the present disclosure. FIG. 2A is described in conjunction with FIG. 1. Flowchart 200 may be executed by system 100. Flowchart 200 may correspond to a testing phase of system 100.

[0022] In step S1, the system 100 may receive an input video 202. The input video 202 may correspond to the factory automation scene video 102. By way of example, the input video 202 may include a set of video frames corresponding to a scene including one or more machines performing a task.

[0023] In step S2, the system 100 may receive a set of training feature vectors 204. Vector The set of training feature vectors 204 may be obtained from a training video. The training video may include a set of video frames corresponding to normal operation of one or more machines performing a task. As used herein, "normal operation of one or more machines" may correspond to activity of one or more machines without any anomalies. As used herein, "training features" may correspond to values ​​or information extracted from video frames of a training video. In an example embodiment, the set of training feature vectors 204 may be obtained from spatiotemporal regions of the training video. For example, one or more training feature vectors may be obtained for each spatial region of the training video. In some cases, the set of training feature vectors 204 may be a matrix E (shown in FIG. 2A), where each element of the matrix E includes one or more training feature vectors for a corresponding spatial region of the training video.

[0024] The spatiotemporal regions of the training video may be defined by a video bounding box. The video bounding box may include spatial and temporal dimensions for dividing (or separating) the training video into multiple spatiotemporal regions. The spatial dimensions may include the size (width, height) of an image patch (e.g., a portion of an image frame). The temporal dimensions may include the number of video frames, which may be less than the number of image frames in the training video. In an example embodiment, each training feature vector associated with one particular spatiotemporal region includes a motion feature vector. As used herein, a "motion feature vector" may be a value or information that defines motion information associated with one or more machines in that particular spatiotemporal region. As an example, the motion feature vector may be obtained from the training video as further described with reference to FIG. 3.

[0025] In step S3, the system 100 may divide the input video 202 into multiple sequences of video volumes. For example, the system 100 may divide the input video 202 into multiple sequences of video volumes as described in the detailed description of FIG. 2B.

[0026] FIG. 2B illustrates a pipeline for partitioning an input video 202 into a sequence of video volumes 208, according to some embodiments of the present disclosure. In an example embodiment, the system 100 may partition the input video 202 into a sequence of spatiotemporal regions using a video bounding box. In doing so, the system 100 generates a partitioned input video 206. By way of example, the video bounding box may be spatially shifted in both horizontal and vertical directions with a fixed step size to generate a set of (possibly overlapping) spatial regions. In one embodiment, the video bounding box may be temporally shifted one frame at a time for each of the spatial regions to generate a sequence of overlapping spatiotemporal regions within each of the spatial regions. By way of example, each overlapping spatiotemporal region of the sequence of overlapping spatiotemporal regions may correspond to a video volume. The same set of spatial regions is used for both the training video and the input test video.

[0027] 2A, in step S4, the system 100 may determine a set 210 of binary difference images of a plurality of sequences of video volumes. To determine the set 210 of binary difference images, the system 100 may generate a sequence of binary difference images for each video volume. For example, the system 100 may generate a sequence of binary difference images for a particular video volume as further described with reference to FIG. 2C.

[0028] 2C shows a schematic diagram for generating a sequence of binary difference images 210a for a video volume 208a, according to some embodiments of the present disclosure. FIG. 2C is described in conjunction with FIG. 2A. The video volume 208a may correspond to one particular video volume 208 of a plurality of sequences of video volumes. By way of example, the video volume 208a may include n+1 image patches, such as image patches 208a-0, 208a-1, 208a-2, 208a-3...208a-n.

[0029] To generate the sequence of binary difference images 210a, the system 100 may determine a binary difference image for each successive pair of image patches in the sequence of image patches 208a-0, 208a-1, 208a-2, 208a-3...208a-n. For example, the system 100 may determine the binary difference image 210a-0 for the successive pair of image patches 208a-0 and 208a-1. To determine the binary difference image 210a-0, the system 100 may generate the difference image by determining a pixel difference value between the image patch 208a-0 and the image patch 208a-1. A "pixel difference value" as used herein may be the absolute value of the difference between (i) an intensity value of a first pixel in the image patch 208a-0 and (ii) an intensity value of a second pixel in the image patch 208a-1 that corresponds to the first pixel in the image patch 208a-0. A difference image as used herein may be an image whose pixel values ​​are pixel difference values.

[0030] Furthermore, the system 100 thresholds the pixel values ​​of the difference image to generate a binary difference image 210a. -0 For example, the system 100 may check whether each of the pixel difference values ​​is greater than a threshold pixel difference value. In one embodiment, if a particular pixel difference value is greater than the threshold pixel difference value, the system 100 may assign a value of "1" to the pixel corresponding to the particular pixel difference value. Otherwise, the system 100 may assign a value of "0" to the pixel corresponding to the particular pixel difference value. -0 is image patch 208 It is a binary image showing which pixels change significantly from a-0 to image patch 208a-1.

[0031] Similarly, the system 100 may determine a binary difference image 210a-1 for successive pairs of image patches 208a-1 and 208a-2. In this manner, the system 100 may generate a sequence of binary difference images 210a-0, 210-1...210a-m by iteratively determining a binary difference image from each successive pair of image patches in the sequence of image patches 208a-0, 208a-1, 208a-2, 208a-3...208a-n.

[0032] 2A , similarly, the system 100 may determine a set of binary difference images 210 by generating a sequence of binary difference images for each video volume of the input video 202. Once the set of binary difference images 210 has been determined, the system 100 may proceed to step S5.

[0033] In step S5, the system 100 may generate a set of input feature vectors 212 based on the determined set of binary difference images 210. As an example, the set of input feature vectors 212 may be a matrix F (shown in FIG. 2A), where each element of the matrix F includes one or more input feature vectors for a corresponding spatiotemporal region of the input video 202. Thus, to generate the set of input feature vectors 212, the system 100 may generate an input feature vector for each video volume of a plurality of sequences of video volumes. In an example embodiment, the system 100 may generate an input feature vector for a particular video volume based on a sequence of binary difference images determined for this particular video volume. For example, the system 100 may generate an input feature vector for one particular video volume as described in the detailed description of FIG. 2D.

[0034] FIG. 2D illustrates a pipeline for generating an input feature vector 212a for a particular video volume, according to some embodiments of the present disclosure. FIG. 2D is described in conjunction with FIG. 2A and FIG. 2C. As an example, the particular video volume may be a video volume 206a of the input video 202. As an example, a sequence of binary difference images associated with the particular video volume may be a sequence of binary difference images 210a. In an example embodiment, the system 100 may generate the input feature vector 212a for the video volume 206a based on the sequence of binary difference images 210a.

[0035] To generate the input feature vector 212a, the system 100 may identify a predetermined pattern for each pixel of the binary difference image in the sequence of binary difference images 210a in step S5-1. As an example, the system 100 may identify a predetermined pattern for each pixel of the binary difference image 210a-0. In an example embodiment, to identify a predetermined pattern for one particular pixel of the binary difference image 210a-0, the system 100 may apply a window 214 to the particular pixel. According to an embodiment, the size associated with the window 214 may be smaller than the size of the binary difference image 210a-0. As an example, the size of the window 214 is 3 pixels wide and 3 pixels long, covering 9 pixels. When the window 214 is applied to the particular pixel, the window 214 defines a 3 pixel by 3 pixel neighborhood 216 occurring in the binary difference image 210a-0 for the particular pixel. A "predetermined pattern" as used herein may be a particular number of bright ("1" value) or dark ("0" value) pixels within the window 216. In other words, a "predetermined pattern" may be a count of the number of pixels within the window 216 that are above a threshold. Because the window 214 covers nine pixels, there are ten possible predetermined patterns 218, e.g., zero pixels above the threshold, one pixel above the threshold, ..., and nine pixels above the threshold. As an example, where the pixels above the threshold correspond to bright pixels, the system 101 may identify the number "2" as a predetermined pattern for a particular pixel corresponding to the 3 pixel by 3 pixel neighborhood 216. In this manner, the system 100 may identify a predetermined pattern for each pixel of the binary difference image 210a-0 by repeatedly applying the window for each pixel of the binary difference image 210a-0.

[0036] In step S5-2, the system 100 can create a histogram 220 by counting the occurrence of each of the predefined patterns 218 of pixels in the binary difference image 210a-0. By way of example, the histogram 220 can include ten bins, such that each bin of the histogram 220 is associated with a corresponding one of the predefined patterns 218. For example, in step S5-2, the system 100 can formulate the histogram 220 by increasing the value of one particular bin by "1" if the predefined pattern corresponding to this particular bin was identified in step S5-1. The formulated histogram 220 is then a count of the number of pixels above a threshold in each 3 pixel by 3 pixel neighborhood that occur in the binary difference image 210a-0. The formulated histogram 220 thus encodes motion information associated with one or more machines in one consecutive pair of image patches (e.g., image patches 206a-0 and 206a-1). Once the histogram 220 for the binary difference image 210a-0 has been formulated, the system 100 can again proceed to step S5-1 and repeatedly perform steps S5-1 and S5-2 to formulate a histogram for each binary difference image in the sequence of binary difference images 210a.

[0037] In step S5-3, the system 100 may generate an input feature vector 212a by concatenating formulaic histograms associated with the binary difference images of the sequence of binary difference images 210a. Since the input feature vector 212a is generated by concatenating histograms encoding motion information associated with one or more machines, the input feature vector 212a may hereinafter be referred to as an input motion feature vector. By way of example, bin 0 of the input motion feature vector 212a is generated by concatenating values ​​of bin 0 of the formulaic histogram over time. Similarly, bin 1 ... bin 9 of the motion feature vector 212a is generated by concatenating values ​​of respective bins 1s ... bin 9s of the formulaic histogram over time. In this manner, the generated input motion feature vector 212a defines the temporal variation of counts of the predetermined pattern 218. Furthermore, because the generated input motion feature vectors 212a are the time variations in counts of a predetermined pattern 218 rather than a pattern representing an arrangement of pixels, the generated input motion feature vectors 212a are easy to calculate and compare.

[0038] For purposes of illustration, the predetermined patterns 218 are considered to be 10 patterns in FIG. 2D. However, the predetermined patterns 218 may include any finite number of patterns. As an example, the predetermined patterns 218 may include 5 patterns, 17 patterns, etc. Thus, if the predetermined patterns 218 correspond to 5 patterns, a window with a size of 2 pixels wide and 2 pixels long may be used to generate a motion feature vector having 5 bins. Alternatively, if the predetermined patterns 218 correspond to 17 patterns, a window with a size of 4 pixels wide and 4 pixels long may be used to generate a motion feature vector having 17 bins.

[0039] In one implementation, the system 100 may receive predefined patterns as inputs during the testing and / or training phases. As an example, a designer may select predefined patterns from a library of patterns based on a task to be performed by one or more machines, such that the selected predefined patterns provide accurate results for the performed task among other patterns in the library of patterns. For example, the library of patterns may include 5 patterns, 10 patterns, 17 patterns, etc. In another implementation, the system 100 may select predefined patterns from the library of patterns using a trained machine learning model. As an example, the machine learning model may be trained to select predefined patterns from a library of patterns based on a task to be performed by one or more machines, such that the selected predefined patterns provide accurate results for the performed task among other patterns in the library of patterns.

[0040] Referring again to FIG. 2A, similarly, the system 100 may generate an input motion feature vector 212a defining the temporal variation in counts of the predetermined pattern 218 for each of the video volumes to generate a set of input feature vectors (or a set of input motion feature vectors) 212.

[0041] In step S6, the system 100 may generate a set of distances. The set of distances may also be referred to as a set of motion distances. In an example embodiment, to generate the set of distances, the system 100 may calculate the shortest distance between each input feature vector of the video volume and a training feature vector associated with a corresponding spatial region of the training video. As an example, the system 100 may calculate the shortest distance between each element of the matrix F representing one or more input feature vectors (e.g., one or more motion feature vectors) and a corresponding element of the matrix E representing one or more training features of the same spatial region. In an example embodiment, this shortest distance may correspond to a Euclidean distance between each input feature vector of the video volume and a training feature vector associated with a corresponding spatial region of the training video.

[0042] In step S7, the system 100 may detect anomalies in the input video 202 based on the generated set of distances. According to an embodiment, the system 100 may compare each distance in the set of distances to an anomaly detection threshold to detect anomalies in the input video 202 of the factory automation scene. By way of example, the anomaly detection threshold may be a threshold that may be predefined based on experiments, etc. In an example embodiment, the system 100 may detect an anomaly in the input video 202 of the factory automation scene when at least one distance in the set of distances is greater than the anomaly detection threshold. Furthermore, the system 100 may execute a control action in response to the detection of anomalies. In one embodiment, the control action may be executed to control one or more machines to stop the abnormal activity. In another embodiment, the control action may be executed to generate a notification to a user to stop the abnormal activity.

[0043] In this manner, the system 100 may detect anomalies in a factory automation process using the input video 202. Because the anomalies in the factory automation process are detected using the input video 202, the cost of detecting anomalies in the factory automation process may be significantly reduced compared to a conventional method of monitoring one or more machines performing a task using customized sensors to detect anomalies. As a result, the system 100 efficiently detects anomalies in the factory automation process. Furthermore, to detect anomalies in the input video 202, the system 100 generates an input motion feature vector that is simple to calculate and compare. As a result, the system 100 detects anomalies in the factory automation process in a feasible manner. Furthermore, the system 100 may generate a motion feature vector for a training video, as will be further described with reference to FIG. 3.

[0044] FIG. 3 illustrates a flowchart 300 for generating a set of training feature vectors for training the system 100 to detect anomalies in a video 102, according to some embodiments of the present disclosure. FIG. 3 will be described in conjunction with FIGS. 1, 2A, 2B, 2C, and 2D. The flowchart 300 may be executed by the system 100. The flowchart 300 may correspond to a training phase of the system 100. In step 302, the system 300 may receive a training video. By way of example, the training video may include a set of video frames corresponding to normal operation of one or more machines performing a task.

[0045] In step 304, the system 100 can generate a corresponding sequence of training video volumes by dividing the training video into spatio-temporal regions. As an example, the system 100 can divide the training video into spatial regions, as described in the detailed description of FIG. 2B. Furthermore, the system 100 can divide each spatial region into a sequence of training video volumes using video bounding boxes, as described in the detailed description of FIG. 2B. In this manner, multiple sequences of training video volumes can be generated for the spatial regions of the training video. The sequence of training spatial regions in the training phase corresponds to the spatial regions in the test phase. The training video volumes represent the normal operation of one or more machines in the video.

[0046] In step 306, the system 100 may determine a binary difference image for each pair of training patches in each of the sequences of training patches. By way of example, the system 101 may determine a binary difference image for each pair of training patches as described in the detailed description of FIG. 2C. In this manner, a set of binary difference images may be determined during the training phase, which may be similar to the set of binary difference images 210.

[0047] In step 308, the system 100 may generate one or more training feature vectors for each of the video volumes by counting occurrences of a predetermined pattern of pixels in each binary difference image for each of the video volumes. As an example, the system 100 may generate one or more training feature vectors for each of the video volumes of the training video as described in the detailed description of FIG. 2A and FIG. 2D. As a result, each training feature vector of the one or more training feature vectors includes a motion feature vector that defines the temporal variation of the counts of the predetermined pattern. Thus, a set of training feature vectors may be generated by the system 100. Some embodiments are based on the recognition that the set of training feature vectors includes a plurality of similar training feature vectors. As an example, if the distance between at least two training feature vectors is close to zero (or is less than a distance threshold), the at least two training feature vectors may be referred to as similar training feature vectors. Some embodiments are based on the recognition that these multiple similar training feature vectors may add additional computational load during the testing phase (i.e., while comparing the input feature vector with the training feature vectors of the same spatial region). Therefore, an objective of some embodiments is to select one training feature vector (sometimes called a unique training feature vector) from multiple similar training feature vectors to avoid additional computational burden.

[0048] To select the training feature vectors, the system 100 may generate a set of distances between the training feature vectors by calculating the distance between each training feature vector that corresponds to the same spatial region in the scene in step 310. As an example, the system 100 may generate the set of distances between the training feature vectors by calculating the distance between each training feature vector and each other training feature vector of the same spatial region.

[0049] In step 312, the system 100 may select a training feature vector in the set of training feature vectors if all distances between this training feature vector and the corresponding feature vectors in the set of training feature vectors are above a distance threshold that is stored in memory and defines the shortest distance between training feature vectors corresponding to the same spatial region. As an example, for one particular spatial region of the training video, the system may select a training feature vector if the distances between this training feature vector and all other training feature vectors of the particular spatial region are above the distance threshold. In an example embodiment, the distance threshold may be determined by the system 100. As an example, the system 100 may calculate the average of the distances between all training feature vectors and the training feature vectors in the set of feature vectors and increase this average by a standard deviation to generate the distance threshold. In another embodiment, the distance threshold may be the shortest distance between the training feature vectors of the particular spatial region. In one embodiment, this shortest distance may be a function of an anomaly detection threshold. In this embodiment, system 100 may select a training feature vector (or multiple training feature vectors) for the particular spatial region if all distances between the training feature vector and all other training feature vectors for the particular spatial region are above an anomaly detection threshold. In another embodiment, the shortest distance may be the median distance between all possible pairs of training feature vectors for the particular spatial region. In this embodiment, system 100 may select a training feature vector (or multiple training feature vectors) for the particular spatial region if all distances between the training feature vector and all other training feature vectors for the particular spatial region are above the median distance.

[0050] In step 314, the system 100 may generate an updated set of training feature vectors.

[0051] In an example embodiment, the updated set of training feature vectors may be used in a testing phase to detect anomalies in an input test video. By way of example, the updated set of training feature vectors may correspond to the set of training feature vectors 204.

[0052] Thus, during the training phase, the system 100 may select the training features so that additional computational burden during the testing phase is avoided. Vector In some scenarios, it may be important to consider the context of motion variation. It is an objective of some embodiments to use an appearance feature vector in addition to the motion feature vector to consider the context of motion variation. In an example embodiment, each of the training feature vectors and each of the input feature vectors may further include a corresponding appearance feature vector obtained from the content of the video volume. As an example, the system 100 may calculate the appearance feature vector of the video volume as further described with reference to FIG. 4.

[0053] FIG. 4 shows a schematic diagram for computing an appearance feature vector of a video volume 400 according to some embodiments of the present disclosure. FIG. 4 is described in conjunction with FIG. 1, FIG. 2A, FIG. 2C, and FIG. 3. During a training phase of the system 100, the video volume 400 may correspond to a sequence of training patches of a training video. According to an embodiment, the system 100 may compute an appearance feature vector of the video volume 400 such that the computed appearance feature vector represents a pattern (e.g., a spatial arrangement) of pixels occurring in the video volume 400. In an embodiment, to compute an appearance feature vector representing a pattern of pixels occurring in the video volume 400, the system 100 may compute a binary difference image 402 from two consecutive video frames of the video volume 400. As an example, the system 100 may compute the binary difference image 402 as described in the detailed description of FIG. 2C. In this embodiment, the determined binary difference image 402 may be an appearance feature vector.

[0054]

number

[0055] In a testing phase of the system 100, the video volume 400 may correspond to one particular video volume (e.g., video volume 206a) of the input video 202. Furthermore, the appearance feature vector computed from the video volume 400 may be referred to as an input appearance feature vector in the testing phase of the system 100. In one embodiment, the system 100 may compute a binary difference image 402 of the video volume 400 and use the computed binary difference image 402 as the input appearance feature vector. In another embodiment, the system 100 may compute a HoG representation 404 of the video volume 400 and use the computed HoG representation 404 as the input appearance feature vector.

[0056] In some embodiments, during a testing phase, system 100 may use the calculated input appearance feature vector together with the input motion feature vector to detect anomalies in a factory automation scene. By way of example, a testing phase of system 100 using input appearance feature vectors and input motion feature vectors to detect anomalies in a factory automation scene is as further described with reference to Figures 5A and 5B.

[0057] 5A and 5B show a flowchart 500 for detecting anomalies in a video of a factory automation scene, according to some other embodiments of the present disclosure. FIGs. 5A and 5B are described in conjunction with FIGs. 1, 2A, 2B, 2C, 2D, 3, and 4. The flowchart 500 may be executed by the system 100. The flowchart 500 may correspond to a testing phase of the system 100. In step 502, the system 100 may receive an input video. By way of example, the input video may correspond to the input video 202 of a factory automation scene.

[0058]

number

[0059] At step 506, the system 100 may partition the input video into a plurality of sequences of video volumes. As an example, the system 100 may partition the input video into a plurality of sequences of video volumes using video bounding boxes, as described in the detailed description of FIG. 2B. To do so, the system 100 may generate a plurality of sequences of video volumes for the input video, such that each video volume includes a sequence of image patches defined by the spatial and temporal dimensions of a corresponding spatiotemporal domain.

[0060] In step 508, system 100 may generate a sequence of binary difference images for each of the video volumes by determining a binary difference image for each successive pair of image patches in the sequence of image patches for each of the video volumes. By way of example, system 100 may determine a binary difference image for each pair of image patches in the sequence of image patches for each of the video volumes as described in the detailed description of FIG. 2C.

[0061] In step 510, the system 100 may generate an input motion feature vector by counting the occurrence of each of the predefined patterns of pixels in each binary difference image for each of the video volumes. By way of example, the system 100 may generate an input motion feature vector for each of the video volumes as described in the detailed description of Figures 2A and 2D. The input motion feature vector may define the temporal variation of the counts of the predefined patterns.

[0062] In step 512, the system 100 may compute an input appearance feature vector for each of the video volumes. By way of example, the system 100 may compute an input appearance vector for each video volume as described in the detailed description of Figure 4. The computed input appearance feature vector may represent a pattern of pixels occurring in the video volume.

[0063]

number

[0064]

number

[0065] In step 518, the system 100 may compare each motion distance and each appearance distance from the set of motion distances and appearance distances to an anomaly detection threshold to detect an anomaly in the input video. For example, in one embodiment, the at least one anomaly detection threshold may include a motion anomaly detection threshold and an appearance anomaly detection threshold. In this embodiment, each motion distance of the set of motion distances is compared to the motion anomaly detection threshold to detect an anomaly in the input video. As an example, the system 100 may detect an anomaly in the input video if at least one motion distance of the set of motion distances is greater than the motion anomaly detection threshold. Furthermore, each appearance distance of the set of appearance distances is compared to the appearance anomaly detection threshold to detect an anomaly in the input video. As an example, the system 100 may detect an anomaly in the input video if at least one appearance distance of the set of appearance distances is greater than the appearance anomaly detection threshold.

[0066]

number

[0067] FIG. 6 illustrates a working environment 600 of a system 602 for detecting anomalies in a factory automation process, according to some embodiments of the present disclosure. FIG. 6 is described in conjunction with FIG. 1 and FIG. 2A. The system 602 may correspond to the system 100. The working environment 600 may correspond to a monitoring system 604 at a location 606. By way of example, the location 606 may be an area of ​​a factory where a factory automation process is performed. The location 606 may include one or more imaging devices 608a and 608b. An "imaging device" as used herein may correspond to a camera, a video player, and the like. The one or more imaging devices 608a and 608b may be positioned (or arranged) such that the one or more imaging devices 608a and 608b monitor one or more machines (such as machines 610a and 610b) performing a task. By way of example, the imaging device 608a may be positioned such that the imaging device 608a monitors a robotic arm 610a. As an example, imaging device 608b may be positioned such that imaging device 608b monitors conveyor 610b. For example, robotic arm 610a may pick up an object from a first level and place it on conveyor 610b at a second level different from the first level to accomplish a factory automation process. Furthermore, conveyor 610b may move the object from the first location to the second location different from the first location to accomplish a factory automation process.

[0068] One or more imaging devices 608a and 608b may separately capture videos including a scene of a factory automation process. As an example, imaging device 608a may capture a video of a scene including a robotic arm 610a picking up and placing an object. As an example, imaging device 608b may capture a video of a scene including a conveyor 610b moving an object. Furthermore, one or more imaging devices 608a and 608b may separately transmit the captured videos to system 602. System 602 may receive the captured videos from each of one or more imaging devices 608a and 608b as input videos. Furthermore, system 602 may execute flowchart 200 to detect anomalies in each of the input videos. As an example, an anomaly in the video captured by imaging device 608a may correspond to an abnormal orientation of robotic arm 610a, etc. As an example, an anomaly in the video captured by imaging device 608b may correspond to an unexpected stopping of conveyor 610b, etc. Additionally, the system 602 may execute control actions to control one or more of the machines 610a and 610b to stop the anomalous activity. Alternatively, the system 602 may generate a notification to an operator associated with the monitoring system 604 to stop the anomalous activity.

[0069] In this manner, the system 602 may detect anomalies in a factory automation scene using video captured by one or more imaging devices 608. As a result, the cost of detecting anomalies in a factory automation process may be significantly reduced as compared to traditional methods of using customized sensors to monitor one or more machines performing a task to detect anomalies.

[0070] In another implementation, the location 606 may include a single imaging device 608. In this implementation, the single imaging device 608 may be positioned such that the single imaging device 608 monitors tasks performed by each of one or more machines 610a and 610b. The single imaging device 608 may thereby capture a video including multiple interdependent processes of a factory automation scene. As an example, the multiple interdependent processes may be a robotic arm 610a picking up and placing an object and a conveyor moving the object. Furthermore, the single imaging device 608 may transmit the captured video to the system 602. The system 602 may receive the captured video as an input video. Furthermore, the system 602 may execute the flowchart 200 to detect anomalies in the input video. Thus, in this implementation, the system 602 detects anomalies in multiple interdependent processes from a single video without the expense of programming logic for anomaly detection.

[0071] FIG. 7 illustrates an overall block diagram of a system 700 for detecting anomalies in a video 702 of a factory automation scene, according to some embodiments of the present disclosure. FIG. 7 is described in conjunction with FIG. 1 and FIG. 2A. The system 700 may correspond to the system 100. The system 700 may have several interfaces connecting the system 700 to one or more imaging devices 704. For example, a network interface controller (NIC) 706 is adapted to connect the system 700 to a network 710 via a bus 708. The system 700 may receive an input video 702 of a factory automation scene through the network 710, either wirelessly or wired. In addition, additional information associated with the input video 702 may be received via an input interface 712. As an example, the additional information associated with the input video 702 may correspond to a number of predefined patterns. The input interface 712 may connect the system 700 to a keyboard 722 and / or a pointing device 724. By way of example, the pointing device 724 may include a mouse, a trackball, a touchpad, a joystick, a pointing stick, a stylus, or a touch screen, among others.

[0072] The system 700 includes a processor 714 configured to execute stored instructions and a memory 716 that stores instructions executable by the processor 714. The processor 714 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 716 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. Additionally, the system 700 includes a storage device 718 adapted to store various modules that store executable instructions for the processor 714. The storage device 718 may be implemented using a hard drive, an optical drive, a thumb drive, an array of drives, or any combination thereof.

[0073] The storage device 718 is configured to store the anomaly detection model 720. Additionally, the storage device 718 may store a set of training feature vectors. As an example, the set of training feature vectors may correspond to the set of training feature vectors 204. In some embodiments, the processor 714 may be configured to execute the anomaly detection model 720 to perform the steps of the flowchart 200 described in the detailed description of FIGS. 2A-2D. As an example, the system 700 may receive an input video 702 of a factory automation scene. Additionally, the system 700 may receive a set of training feature vectors obtained from a spatiotemporal region of the training video. The training video may include a set of video frames of normal operation of a machine performing a task. Each spatiotemporal region is associated with one or more training feature vectors, each training feature vector including a motion feature vector that defines the temporal variation of counts of a predetermined pattern.

[0074] Additionally, system 700 may partition the input video 702 into a plurality of sequences of video volumes corresponding to the spatial and temporal dimensions of the spatio-temporal domains of the training video, such that the video volume includes a sequence of image patches defined by the spatial and temporal dimensions of the corresponding spatio-temporal domain. Additionally, system 700 may determine a binary difference image for each successive pair of image patches in the sequence of image patches of each of the video volumes to generate a sequence of binary difference images for each of the video volumes.

[0075] Additionally, the system 700 may count occurrences of each of the predetermined patterns of pixels in each binary difference image for each of the video volumes to generate input feature vectors including an input motion feature vector that defines the temporal variation of the counts of the predetermined patterns for each of the video volumes. Additionally, the system 700 may generate a set of distances by calculating the shortest distance between the input feature vectors of each of the video volumes and a training feature vector associated with a corresponding spatial region in the scene. Additionally, the system 700 may compare each distance from the set of distances to an anomaly detection threshold to detect anomalies in the input video of the factory automation scene.

[0076] In addition, system 700 may include an imaging interface 726 and an application interface 728. Imaging interface 726 connects system 700 to a display device 730. By way of example, display device 730 may include a computer monitor, a television, a projector, or a mobile device, among others. Application interface 728 connects system 700 to an application device 732. By way of example, application device 732 may include a surveillance system, etc. In example embodiments, system 700 outputs results of the video anomaly detection via imaging interface 726 and / or application interface 728.

[0077] The above description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the above description of exemplary embodiments provides one of ordinary skill in the art with an enabling description for implementing one or more exemplary embodiments. What is intended is various changes that may be made in the function and arrangement of elements without departing from the spirit and scope of the disclosed subject matter, as set forth in the appended claims.

[0078] Specific details are provided in the above description to provide a thorough understanding of the embodiments. However, those skilled in the art will appreciate that the embodiments may be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments. Additionally, like reference numbers and names in the various drawings refer to like elements.

[0079] Also, the individual embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process may be terminated when its operations are completed, but may have other steps not discussed or included in the diagram. Moreover, not all operations in any process specifically described may occur in all embodiments. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the end of the function may correspond to a return of the function to the calling function or to the main function.

[0080] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, either manually or automatically. The manual or automated implementation may be performed or at least assisted through the use of machines, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments to perform the necessary tasks may be stored on a machine-readable medium. A processor may perform the necessary tasks.

[0081] The various methods or processes outlined herein may be coded as software executable on one or more processors employing any one of a variety of operating systems or platforms. In addition, such software may be written using any of a number of suitable programming languages ​​and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.

[0082] The embodiments of the present disclosure may be implemented as a method, an example of which is provided. The order of operations performed as part of the method may be determined in any suitable manner. Thus, the embodiments may be configured to perform operations in an order different from that illustrated, which may include performing some operations simultaneously, even though in the illustrated embodiment they are shown as a series of operations. Although the present disclosure has been described with reference to certain preferred embodiments, it should be understood that various other adaptations and modifications may be made within the spirit and scope of the present disclosure. Accordingly, the scope of the appended claims covers all such variations and modifications that fall within the true spirit and scope of the present disclosure.

Claims

1. 1. A system for detecting anomalies in a video of a factory automation scene, the system comprising: a processor; and a memory having instructions stored thereon, the instructions, when executed by the processor, causing the system to: receiving an input video of a scene including a machine performing a task; receiving a set of training feature vectors derived from spatio-temporal regions of a training video of normal operation of the machine performing the task, each of the spatio-temporal regions being associated with one or more training feature vectors, each training feature vector including a motion feature vector defining a temporal variation in counts of a predetermined pattern, the spatio-temporal regions being defined by a video bounding box; and using the video bounding box to partition the input video into a plurality of sequences of video volumes corresponding to spatial and temporal dimensions of the spatio-temporal regions of the training video, such that a video volume comprises a sequence of image patches defined by the spatial and temporal dimensions of the corresponding spatio-temporal regions, wherein partitioning the input video into a plurality of sequences of video volumes using the video bounding box is performed by: spatially shifting the video bounding box horizontally and vertically with a fixed step size to generate a set of spatial regions, and temporally shifting the video bounding box one frame at a time for each of the spatial regions to generate a sequence of overlapping spatio-temporal regions within each of the spatial regions, each overlapping spatio-temporal region of the sequence of overlapping spatio-temporal regions corresponding to a video volume; and generating a sequence of binary difference images for each of the video volumes by determining a binary difference image for each pair of consecutive image patches in the sequence of image patches for each of the video volumes; and generating input feature vectors including an input motion feature vector defining a temporal variation of counts of the predetermined patterns for each of the video volumes by counting occurrences of each of the predetermined patterns of pixels in each binary difference image for each of the video volumes, the predetermined pattern for a particular pixel being a count of a number of pixels above a threshold within a window applied to the particular pixel, the window being a count of a number of pixels above a threshold within a window applied to the particular pixel in a corresponding binary difference image for the particular pixel. A 3 pixel by 3 pixel neighborhood is defined, and further, generating a set of distances by calculating the shortest distance between the input feature vector of each of the video volumes and the training feature vectors associated with corresponding spatial regions in the scene; comparing each distance in the set of distances to an anomaly-detection threshold to detect anomalies in the input video of the scene; To determine the training feature vector, the processor: the video bounding box is configured to partition the training video into the spatio-temporal regions of the training video using the video bounding box to generate a sequence of corresponding training patches, wherein partitioning the training video into spatio-temporal regions using the video bounding box is performed by spatially shifting the video bounding box horizontally and vertically by the fixed step size to generate a set of further spatial regions, and temporally shifting the video bounding box one frame at a time for each of the further spatial regions to generate a sequence of overlapping further spatio-temporal regions within each of the further spatial regions, each overlapping spatio-temporal region of the sequence of overlapping further spatio-temporal regions corresponding to a spatio-temporal region of the training video; and to determine the training feature vector, the processor further comprises: determining a binary difference image for each pair of consecutive patches in each of said sequences of training patches; and generating a plurality of training feature vectors for each of the spatio-temporal regions by counting the occurrence of each of the predetermined patterns of pixels in each binary difference image for each of the spatio-temporal regions, each training feature vector including the motion feature vector defining the temporal variation of counts of the predetermined patterns, and for determining the training feature vectors, the processor further comprises: generating a set of distances between training feature vectors by calculating a distance between each training feature vector corresponding to the same spatial region in the scene; configured to select a training feature vector in the set of training feature vectors, the training feature vector being selected if all distances between the training feature vector and corresponding feature vectors in the set of training feature vectors exceed a distance threshold defining the shortest distance between the training feature vectors stored in the memory and corresponding to the same spatial region; the shortest distance corresponds to a Euclidean distance between the input feature vector of each of the video volumes and the training feature vector associated with a corresponding spatial region of the training video; the processor is further configured to formulate a histogram based on the count of occurrences of each of the predetermined patterns in the binary difference image, the histogram being a count of a number of pixels above a threshold for each of the 3 pixel by 3 pixel neighborhoods occurring in the binary difference image; the input feature vector is generated by concatenating the formulated histograms; The input feature vector generated by concatenating the formulated histograms is the input motion feature vector.

2. The processor, detecting an anomaly in the input video of the scene if at least one metric in the set of metric is greater than the anomaly-detection threshold; The system of claim 1 , configured to perform a control action in response to detecting the anomaly.

3. The system of claim 1 , wherein the training feature vectors for the spatio-temporal regions are selected such that a shortest distance between them is greater than or equal to the anomaly detection threshold.

4. The system of claim 1 , wherein the training feature vectors for the spatio-temporal region are selected such that a shortest distance between them is greater than a median distance between all possible training feature vector pairs for that spatial region.

5. To determine the binary difference image, the processor generating a difference image by determining pixel difference values ​​between successive image patches of the video volume; The system of claim 1 , configured to generate the binary difference image by thresholding pixel values ​​of the difference image.

6. The system of claim 1 , wherein the time dimension of the spatio-temporal domain defines a portion of the normal operation of the machine performing the task, or defines the entirety of the normal operation of the machine performing the task.

7. The system of claim 1 , wherein the processor is configured to receive the predetermined pattern selected for the task performed.

8. The system of claim 1 , wherein the processor is configured to select the predetermined pattern from a library of patterns.

9. The system of claim 1 , wherein each of the training feature vectors and each of the input feature vectors comprises a corresponding appearance feature vector derived from content of a video patch.

10. The system of claim 9 , wherein the appearance feature vector is a gradient orientation histogram computed for an image patch of the video patch.

11. The system of claim 9 , wherein the appearance feature vector is a binary difference image computed for two consecutive image patches of the video patch.

12. Each training feature vector is composed of an appearance feature vector and the motion feature vector, and to detect anomalies in the input video of the scene, the instructions include: calculating an input appearance feature vector for each of said video volumes, said input appearance feature vector representing a pattern of pixels occurring in the video volume; and generating a set of motion distances by calculating the shortest distance between the input motion feature vector of each of the video volumes and the motion feature vectors of the training feature vectors associated with corresponding spatial regions in the scene; generating a set of appearance distances by calculating a shortest distance between the input appearance feature vector of each of the video volumes and the appearance feature vectors of the training feature vectors associated with corresponding spatial regions in the scene; 2. The system of claim 1 , further comprising: comparing each motion distance in the set of motion distances to a motion anomaly detection threshold; and comparing each appearance distance in the set of appearance distances to an appearance anomaly detection threshold, to detect anomalies in the input video of the scene.

13. 1. A method for detecting anomalies in a video of a factory automation scene, the method comprising: receiving an input video of a scene including a machine performing a task; receiving a set of training feature vectors derived from spatio-temporal regions of a training video of normal operation of the machine performing the task, each of the spatio-temporal regions being associated with one or more training feature vectors, each training feature vector including a motion feature vector defining temporal variations in counts of a predetermined pattern, the spatio-temporal regions being defined by a video bounding box, the method further comprising: using the video bounding box to partition the input video into a plurality of sequences of video volumes corresponding to spatial and temporal dimensions of the spatio-temporal regions of the training video, such that a video volume comprises a sequence of image patches defined by the spatial and temporal dimensions of the corresponding spatio-temporal regions, wherein partitioning the input video into a plurality of sequences of video volumes using the video bounding box is performed by: spatially shifting the video bounding box horizontally and vertically with a fixed step size to generate a set of spatial regions, and temporally shifting the video bounding box one frame at a time for each of the spatial regions to generate a sequence of overlapping spatio-temporal regions within each of the spatial regions, each overlapping spatio-temporal region of the sequence of overlapping spatio-temporal regions corresponding to a video volume; and generating a sequence of binary difference images for each of the video volumes by determining a binary difference image for each pair of consecutive image patches in the sequence of image patches for each of the video volumes; generating input feature vectors including an input motion feature vector defining a temporal variation of counts of the predetermined patterns for each of the video volumes by counting occurrences of each of the predetermined patterns of pixels in each binary difference image for each of the video volumes, wherein the predetermined pattern for a particular pixel is a count of a number of pixels above a threshold within a window applied to the particular pixel, the window defining a 3 pixel by 3 pixel neighborhood occurring in a corresponding binary difference image for the particular pixel; and generating a set of distances by calculating the shortest distance between the input feature vector of each of the video volumes and the training feature vectors associated with corresponding spatial regions in the scene; comparing each distance in the set of distances to an anomaly-detection threshold to detect anomalies in the input video of the scene; To determine the training feature vectors, the method further comprises: using the video bounding box to partition the training video into the spatio-temporal regions of the training video to generate a sequence of corresponding training patches, wherein partitioning the training video into spatio-temporal regions using the video bounding box is performed by spatially shifting the video bounding box horizontally and vertically by the fixed step size to generate a set of further spatial regions, and temporally shifting the video bounding box one frame at a time for each of the further spatial regions to generate a sequence of overlapping further spatio-temporal regions within each of the further spatial regions, each overlapping spatio-temporal region of the sequence of overlapping further spatio-temporal regions corresponding to a spatio-temporal region of the training video, and to determine the training feature vector, the method further comprises: determining a binary difference image for each pair of consecutive patches in each of said sequences of training patches; and generating a plurality of training feature vectors for each of the spatio-temporal regions by counting occurrences of each of the predetermined patterns of pixels in each binary difference image for each of the spatio-temporal regions, each training feature vector including a motion feature vector defining the temporal variation in counts of the predetermined patterns, and to determine the training feature vectors, the method further comprises: generating a set of distances between training feature vectors by calculating a distance between each training feature vector corresponding to the same spatial region in the scene; selecting a training feature vector in the set of training feature vectors, the training feature vector being selected if all distances between the training feature vector and corresponding feature vectors in the set of training feature vectors exceed a distance threshold defining a shortest distance between the training feature vectors corresponding to the same spatial region; the shortest distance corresponds to a Euclidean distance between the input feature vector of each of the video volumes and the training feature vector associated with a corresponding spatial region of the training video; The method further includes formulating a histogram based on the count of occurrences of each of the predetermined patterns in the binary difference image, the histogram being a count of the number of pixels above a threshold for each of the 3 pixel by 3 pixel neighborhoods occurring in the binary difference image; the input feature vector is generated by concatenating the formulated histograms; The input feature vector generated by concatenating the formulated histograms is the input motion feature vector.

Citation Information

Patent Citations

  • Operation evaluation device, method, and program

    JP2015011664A

  • Image processing device and image processing method

    JP2017173098A

  • Sub-Pixel and Sub-Resolution Localization of Defects on Patterned Wafers

    US20160292840A1

  • Welded state monitoring device and method

    WO2009057830A1