Anomaly detection by ranking from algorithm
By transforming data from a high-dimensional space to a low-dimensional space and combining K-means clustering and deep learning, the problems of computational complexity and resource requirements in high-dimensional anomaly detection are solved, and efficient anomaly detection is achieved.
Patent Information
- Application Number
- CN202111353087.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-16
- Filing Date
- 2021-11-16
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-11-16
AI Technical Summary
When performing anomaly detection in high-dimensional space, existing technologies face challenges such as high computational complexity and sparse data processing. Furthermore, traditional convolutional learning-based methods require a large amount of computational resources and are difficult to adapt to unlabeled anomaly samples.
A distance-based vector classification method is adopted to transform the data from a high-dimensional space to a low-dimensional space. Using the K-means clustering algorithm and deep learning technology, anomaly detection is performed by generating score vectors, thereby reducing the computational resource requirements.
It improves the processing efficiency of the computing system, adapts to unlabeled anomaly samples, reduces computational complexity and resource requirements, and achieves efficient anomaly detection.
Smart Images

Figure CN114511723B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of machine learning, and more specifically to anomaly detection. Background Technology
[0002] Anomaly detection (e.g., outlier detection) is the identification of rare items, events, or observations that raise suspicion by being significantly different from most data. Anomaly detection is used in a variety of fields, such as, but not limited to, statistics, signal processing, finance, econometrics, manufacturing, networking, and data mining. Typically, anomalous items will be translated into problems such as bank fraud, structural defects, medical problems, or errors in text. Anomalies are also referred to as outliers, singularities, noise, biases, and exceptions.
[0003] Cluster analysis, or clustering, is the task of grouping a collection of objects in such a way that objects in the same group (e.g., cluster) are more similar to each other than objects in other groups (clusters). Clustering is a primary task in data mining and a common technique for statistical data analysis, used in many fields, including pattern recognition, image analysis, information retrieval, bioinformatics, data compression, computer graphics, and machine learning.
[0004] K-means clustering is a vector quantization method originally from signal processing. Its purpose is to divide "n" observations into "K" clusters, where each observation belongs to the cluster with the nearest mean (e.g., cluster center or centroid), serving as the prototype for the clusters. This results in the data space being partitioned into Voronoi cells. Cluster analysis is widely used in data mining. K-means clustering has a loose relationship with the K-nearest neighbor classifier, a popular machine learning technique for classification that is often confused with K-means clustering due to its name. Summary of the Invention
[0005] This invention discloses a method, computer program product, and system for distance-based vector classification in anomaly detection. The method includes one or more processors identifying one or more audio communications from a first user to a second user, wherein the one or more audio communications are transmitted using a first computing device. The method also includes one or more processors determining a target of the first user based at least in part on the audio communications of the first user. The method further includes one or more processors determining a set of conditions corresponding to the one or more audio communications and the target, wherein the set of conditions indicates a vulnerability in the first user's personal data. The method also includes one or more processors preventing the first computing device from transmitting audio data including the first user's personal data. Attached Figure Description
[0006] Figure 1 This is a functional block diagram of a data processing environment according to an embodiment of the present invention.
[0007] Figure 2 It is a description of an embodiment of the present invention. Figure 1 The flowchart illustrates the operational steps of a procedure in a data processing environment that uses distance vectors based on image features to classify images as anomalies.
[0008] Figure 3 According to an embodiment of the present invention Figure 1 A block diagram of the client device and server components. Detailed Implementation
[0009] Embodiments of the present invention allow distance-based vector classification in anomaly detection. Embodiments of the present invention identify visual features of one or more images. Embodiments of the present invention determine the distance vector of a test image. Additional embodiments of the present invention generate a score using the distance vector of the test image. Some embodiments of the present invention distinguish the test image from the set of images based on the generated score of the test image.
[0010] Some embodiments of the present invention recognize that challenges exist in anomaly detection algorithms regarding accuracy due to the generation of vectors in high-dimensional spaces. For example, working in high-dimensional spaces may be undesirable for many reasons, such as the original data often being sparse due to dimensionality reduction, and the analytical data often being computationally difficult. Embodiments of the present invention allow data to be transformed from a high-dimensional space to a low-dimensional space, such that the low-dimensional representation preserves some meaningful properties of the original data. Furthermore, embodiments of the present invention can be adapted to unlabeled anomaly samples as input.
[0011] Embodiments of the present invention can be operated to improve computing systems by utilizing less computational power (e.g., processing resources) than conventional convolutional learning-based techniques. Furthermore, various embodiments of the present invention improve the efficiency of computing system processing resources by utilizing clustering algorithms that require fewer processing resources than conventional convolutional learning-based techniques.
[0012] The embodiments of the present invention can be implemented in various forms, and exemplary implementation details will be discussed below with reference to the accompanying drawings.
[0013] The invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a functional block diagram illustrating a distributed data processing environment according to an embodiment of the present invention, typically designated as 100. Figure 1 This is merely an illustration of an implementation and does not imply any limitation on the environments in which different embodiments may be implemented. Many modifications can be made to the described environment by those skilled in the art without departing from the scope of the invention as set forth in the claims.
[0014] This invention may include various accessible data sources, such as database 144, which may include personal data, content, or information that the user wishes not to be processed. Personal data includes personally identifiable information or sensitive personal information, as well as user information such as tracking or geolocation information. Processing refers to any automated or non-automated operation or collection of operations, such as collecting, recording, organizing, structuring, storing, adapting, altering, retrieving, consulting, using, distributing publicly, or otherwise making a combination, restriction, erasure, or destruction of personal data available. Detection procedure 200 allows for authorized and secure processing of personal data. Detection procedure 200 provides informed consent, informing the user of the collection of personal data and allowing the user to opt in or out of processing personal data. Consent may take several forms. Opting in consent compels the user to take an affirmative action before personal data is processed. Alternatively, opting out consent compels the user to take an affirmative action to prevent personal data from being processed before it is processed. Detection procedure 200 provides information about the nature of the personal data and the processing (e.g., type, scope, purpose, duration, etc.). The detection program 200 provides the user with a copy of the stored personal data. The detection program 200 allows for the correction or completion of incorrect or incomplete personal data. The detection program 200 allows for the immediate deletion of personal data.
[0015] The distributed data processing environment 100 includes a server 140 and client devices 120, all interconnected via a network 110. The network 110 can be, for example, a telecommunications network, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN) such as the Internet, or a combination of these, and can include wired, wireless, or fiber optic connections. The network 110 can include one or more wired and / or wireless networks capable of receiving and transmitting data, voice, and / or video signals, including multimedia signals comprising voice, data, and video information. Typically, the network 110 can be any combination of connections and protocols supporting communication between the server 140 and client devices 120, as well as other computing devices (not shown) within the distributed data processing environment 100.
[0016] Client device 120 may be one or more of a laptop computer, tablet computer, smartphone, smartwatch, smart speaker, virtual assistant, or any programmable electronic device capable of communicating with various components and devices within the distributed data processing environment 100 via network 110. Typically, client device 120 represents one or more programmable electronic devices or a combination of programmable electronic devices capable of executing machine-readable program instructions and communicating with other computing devices (not shown) within the distributed data processing environment 100 via a network such as network 110. According to embodiments of the invention, client device 120 may include information regarding… Figure 3 The components are described and illustrated in further detail.
[0017] Client device 120 includes user interface 122 and application 124. In various embodiments of the invention, the user interface is a program that provides an interface between the user of the device and multiple applications residing on the client device. A user interface such as user interface 122 involves information presented to the user by the program (such as graphics, text, and sound), and control sequences used by the user to control the program. Various types of user interfaces exist. In one embodiment, user interface 122 is a graphical user interface (GUI). A graphical user interface (GUI) is a user interface that allows a user to interact with an electronic device (such as a computer keyboard and mouse) through graphical icons and visual indicators (such as auxiliary symbols), as opposed to text-based interfaces, typed command labels, or text navigation. In computing, GUIs were introduced as a response to the steep learning curve of a command-line interface that requires typing commands on a keyboard. Actions in a GUI are typically performed by directly manipulating graphical elements. In another embodiment, user interface 122 is a scripting or application programming interface (API).
[0018] Application 124 is a computer program designed to run on client device 120. Applications are often used to provide users with similar services accessed on a personal computer (e.g., web browsers, music players, email programs, or other media). In one embodiment, application 124 is mobile application software. For example, mobile application software or an "app" is a computer program designed to run on smartphones, tablets, and other mobile devices. In another embodiment, application 124 is a web user interface (WUI) and may display text, documents, web browser windows, user options, application interfaces, and instructions for operation, and includes information presented to the user by the program (such as graphics, text, and sound) and control sequences used by the user to control the program. In another embodiment, application 124 is a client-side application for detecting program 200.
[0019] In various embodiments of the present invention, server 140 may be a desktop computer, a computer server, or any other computer system known in the art. Generally, server 140 represents any electronic device or combination of electronic devices capable of executing computer-readable program instructions. According to embodiments of the present invention, server 140 may include... Figure 3 The components are described and illustrated in further detail.
[0020] Server 140 may be a standalone computing device, management server, web server, mobile computing device, or any other electronic device or computing system capable of receiving, sending, and processing data. In one embodiment, server 140 may represent a server computing system utilizing multiple computers as server systems, such as in a cloud computing environment. In another embodiment, server 140 may be a laptop computer, tablet computer, netbook computer, personal computer (PC), desktop computer, personal digital assistant (PDA), smartphone, or any programmable electronic device capable of communicating via network 110 with client devices 120 and other computing devices (not shown) within the distributed data processing environment 100. In yet another embodiment, server 140 represents a computing system utilizing cluster computers and components (e.g., database server computers, application server computers, etc.) that act as a single, seamless pool of resources when accessed within the distributed data processing environment 100.
[0021] Server 140 includes storage device 142, database 144, and detection program 200. Storage device 142 can be implemented using any type of storage device, such as a persistent storage device 305 capable of storing data accessible and utilized by client device 120 and server 140, such as a database server, hard disk drive, or flash memory. In one embodiment, storage device 142 may represent multiple storage devices within server 140. In various embodiments of the invention, storage device 142 stores various types of data, including database 144. Database 144 may represent one or more organized collections of data stored and accessed from server 140. For example, database 144 includes training images, test images, visual features, centroid information, etc. In one embodiment, data processing environment 100 may include an additional server (not shown) hosting additional information accessible via network 110.
[0022] Typically, detection procedure 200 can classify images as anomalies using distance vectors based on image features. In one embodiment, detection procedure 200 clusters multiple normal training images using a set of features from multiple normal training images. Additionally, detection procedure 200 obtains a set of centroid information for the set of training images. Furthermore, detection procedure 200 computes a distance vector for a test image using the set of centroid information and the set of features. Furthermore, detection procedure 200 ranks and selects the top "k" elements (e.g., from k-means clustering) to obtain a reduced distance vector. Furthermore, detection procedure 200 generates a score vector using the reduced distance vector and performs classification-based anomaly detection on the test image using the score vector.
[0023] Figure 2 This is a flowchart illustrating the operational steps of a detection program 200 according to an embodiment of the present invention, which classifies images as anomalies using distance vectors based on image features. In one embodiment, the detection program 200 is initiated in response to a user connecting a client device 120 to the detection program 200 via network 110. For example, the detection program 200 is initiated in response to a user registering (e.g., opting in) a laptop computer (e.g., client device 120) with the detection program 200 via a WLAN (e.g., network 110). In another embodiment, the detection program 200 is a background application that continuously monitors the client device 120. For example, the detection program 200 is initiated and monitors a client-side application (e.g., application 124) on the laptop computer used to input test images when the user's laptop computer (e.g., client device 120) boots.
[0024] In step 202, detection procedure 200 identifies features of the set of training images. In one embodiment, detection procedure 200 utilizes deep learning techniques to identify features of the set of training images. For example, detection procedure 200 utilizes a pre-trained machine learning algorithm (e.g., a deep neural network (DNN), a convolutional deep neural network (CNN), etc.) to identify features (e.g., derived values from an initial set of measurement data) of each image in the set of training images, which may include higher-level features from each image. In this example, detection procedure 200 extracts a set of features (e.g., a high-dimensional dataset) of each image in the set of training images, which includes relevant information from the input data (e.g., input images), enabling the performance of the desired task (e.g., anomaly detection) by using a reduced representation instead of the complete initial data (e.g., the set of training images).
[0025] In step 204, detection procedure 200 generates a set of images from the set of training images. In one embodiment, detection procedure 200 utilizes a clustering algorithm to generate two or more sets of images. For example, detection procedure 200 utilizes machine learning techniques (e.g., a K-means clustering model) that identify the cluster centroids that minimize the distance between data points (e.g., training images) and their nearest centroids to group the set of training images based on extracted high-dimensional features (as discussed in step 202). In this example, detection procedure 200 utilizes clustering to generate classes (e.g., a set of images) based on the set of training images, wherein when classifying the training images, the distance between the training images and each group center (e.g., centroid) (e.g., the centroid of the cluster closest to the data point) is utilized. Furthermore, detection procedure 200 uses the distances of the training images in k-means clustering to classify each image in the set of training images into groups with similar attributes and / or features, and data points in different groups should have highly dissimilar attributes and / or features. Alternatively, the detection program 200 generates a dictionary of "k" vectors such that the data vectors can be mapped to code vectors (e.g., the input image) that minimize the error in the reconstruction (i.e., vector quantization).
[0026] In step 206, detection procedure 200 determines the centroid information for each of the images in the set. In one embodiment, detection procedure 200 determines the centroid information for a set of two or more images for a clustering algorithm. For example, detection procedure 200 uses the location information (e.g., vector data) of each point in a single cluster from a clustering table (e.g., database 144) to determine the average of the location information for a single cluster (i.e., determine the centroid of the cluster). In this example, the centroid is a vector containing a number for each variable, where each number is the average of the observed variables in that cluster (i.e., the centroid can be considered as the multidimensional average of the cluster). Additionally, detection procedure 200 determines the average of the location information for each cluster in the set of training images.
[0027] In step 208, detection procedure 200 determines the distance vector of the test image. In one embodiment, detection procedure 200 determines the distance vector corresponding to the features of the test image provided by the user via client device 120. For example, detection procedure 200 utilizes a pre-trained machine learning algorithm (e.g., deep neural network (DNN), convolutional deep neural network (CNN), etc.) to identify the set of features of the test images in the set of test images. In this example, detection procedure 200 calculates the distance vector of each feature of the set of features of the test images relative to the centroid of each cluster determined in step 206.
[0028] In an example embodiment, the detection procedure 200 can use an equation to calculate the distance vector "d" of the features of the test image relative to the obtained centroid information. i,x ",in:
[0029]
[0030] Among them, "f" x "" refers to the image features of a given test image obtained from a pre-trained machine learning algorithm, and "c i " is the centroid information from the "k" centroids of the k-means cluster discussed in step 206 (e.g., 2 clusters of 2D data).
[0031] In step 210, detection procedure 200 determines the score of the distance vector of the test image. In one embodiment, detection procedure 200 performs dimensionality reduction of the distance vector of the test image. In one scenario, detection procedure 200 uses a k-means clustering model to cluster the high-dimensional feature vectors of the set of training images into "k" clusters, resulting in a set of "k" cluster centers (e.g., centroids). Detection procedure 200 can represent each of the original data points (e.g., features of the training images) based on how far the data points are from each of these cluster centers (i.e., the distance from the data point to each cluster center can be calculated), resulting in a set of "k" distances for each data point. Additionally, detection procedure 200 can use the set of "k" distances for each data point to form a new vector dimension "k", which represents the original data points as new vectors with a lower dimension relative to the original feature dimensions. For example, detection procedure 200 can use the calculated distance vector "d" i,x The corresponding values are used to sort each feature in the set of features of the test image in ascending order (e.g., the shorter the distance value, the higher the rank in the ranking). In this example, the detection procedure 200 selects features above a defined threshold (e.g., the first rank) to generate a reduced representation (i.e., a reduced distance vector) of the set of features of the test image.
[0032] In another embodiment, detection program 200 generates a score vector of distance vectors corresponding to the features of the test image. For example, detection program 200 uses a reduced distance vector of the set of features of the test image to generate a score vector of the set of features of the test image. In an example embodiment, detection program 200 may use an equation to convert the distance vector of the features of the test image into a score vector "s". i,x ",in:
[0033] S i,x =1 / (d i,x +∈) (2)
[0034] Among them, “d”i,x " is a reduced distance vector of the set of features of the test image, and "∈" is a small value so as to be divided by zero.
[0035] In step 212, detection procedure 200 classifies the test images. In one embodiment, detection procedure 200 uses a score vector of features from the test images to detect anomalies. For example, detection procedure 200 uses a machine learning algorithm (e.g., one-class support vector machine (OCSVM), artificial neural network, etc.) to classify the features using the generated score vector of features from the test images. In this example, detection procedure 200 uses one or more sets of training data to train the machine learning algorithm to classify the test images (i.e., identify images containing anomalies), the set of training data may include a set of training images without anomalies (as discussed in step 202) and / or a set of training images with anomalies. Furthermore, detection procedure 200 uses the score vector of generated features based on distance vectors relative to clusters (e.g., groups, classes, etc.) to detect anomalies (e.g., outliers) in the unlabeled test image dataset by identifying features in the set of images that appear to best fit the corresponding cluster (e.g., centroids, classes, etc.).
[0036] In one scenario, detection procedure 200 inputs the generated score vector of the image data input (e.g., features from a set of test image features) into OCSVM to determine whether the generated score vector of the image data input is within the distance of each centroid. Furthermore, if detection procedure 200 determines that the generated score vector is within the distance, it assigns a positive classification to the feature. Conversely, if detection procedure 200 determines that the generated score vector is not within the distance, it assigns a negative classification to the feature, indicating that the feature is an outlier.
[0037] Figure 3 A block diagram depicts the components of a client device 120 and a server 140 according to an illustrative embodiment of the present invention. It should be understood that... Figure 3 This is merely an illustration of an implementation and does not imply any limitation on the environments in which different embodiments may be implemented. Many modifications can be made to the described environment.
[0038] Figure 3The system includes a processor 301, a cache 303, a memory 302, a permanent storage device 305, a communication unit 307, an input / output (I / O) interface 306, and a communication structure 304. The communication structure 304 provides communication between the cache 303, memory 302, permanent storage device 305, communication unit 307, and input / output (I / O) interface 306. The communication structure 304 can be implemented using any architecture designed to transfer data and / or control information between processors (such as microprocessors, communication and network processors, etc.), system memory, peripheral devices, and any other hardware components within the system. For example, the communication structure 304 can be implemented using one or more buses or crossbar switches.
[0039] Memory 302 and persistent storage device 305 are computer-readable storage media. In this embodiment, memory 302 includes random access memory (RAM). Typically, memory 302 may include any suitable volatile or non-volatile computer-readable storage medium. Cache 303 is a fast memory that enhances the performance of processor 301 by storing recently accessed data and data near recently accessed data from memory 302.
[0040] Program instructions and data (e.g., software and data 310) used to implement embodiments of the present invention can be stored in persistent storage device 305 and memory 302 for execution by one or more corresponding processors 301 via cache 303. In embodiments, persistent storage device 305 includes a magnetic hard disk drive. As an alternative to or supplement to a magnetic hard disk drive, persistent storage device 305 may include a solid-state drive, semiconductor storage device, read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.
[0041] The media used in persistent storage device 305 can also be removable. For example, a removable hard disk drive can be used in persistent storage device 305. Other examples include optical discs and disks, thumb drives, and smart cards, which are inserted into the drive for transfer to another computer-readable storage medium that is also part of persistent storage device 305. Software and data 310 can be stored in persistent storage device 305 for access and / or execution by one or more corresponding processors 301 via cache 303. Regarding client device 120, software and data 310 includes data for user interface 122 and application 124. Regarding server 140, software and data 310 includes data for storage device 142 and detection program 200.
[0042] In these examples, communication unit 307 provides communication with other data processing systems or devices. In these examples, communication unit 307 includes one or more network interface cards. Communication unit 307 can provide communication by using one or both of physical and wireless communication links. Program instructions and data (e.g., software and data 310) for implementing embodiments of the invention can be downloaded to persistent storage device 305 via communication unit 307.
[0043] I / O interface 306 allows data input and output to other devices that can be connected to each computer system. For example, I / O interface 306 can provide connectivity to external device 308 (e.g., keyboard, keypad, touchscreen, and / or other suitable input devices). External device 308 may also include portable computer-readable storage media, such as thumb drives, portable optical discs or disks, and memory cards. Program instructions and data (e.g., software and data 310) used to implement embodiments of the invention can be stored on such portable computer-readable storage media and can be loaded onto permanent storage device 305 via I / O interface 306. I / O interface 306 is also connected to display 309.
[0044] The display 309 provides a mechanism for displaying data to the user and may be, for example, a computer monitor.
[0045] The programs described herein are identified based on applications that implement them in specific embodiments of the invention. However, it should be understood that any specific procedural terminology used herein is for convenience only, and therefore the invention should not be limited to use only in any specific application identified and / or implied by such terminology.
[0046] This invention can be a system, method, and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.
[0047] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0048] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0049] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages (including object-oriented programming languages such as Smalltalk, C++, etc.) and procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing the status information of the computer-readable program instructions.
[0050] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0051] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0052] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0053] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order as shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0054] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The terminology used herein is chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method comprising: generating, by one or more processors, one or more clusters of images in a set of training images using a clustering algorithm based at least in part on one or more sets of features of the set of training images to determine a respective centroid of each of the one or more clusters; determining, by one or more processors, one or more distance vectors for a test image based at least in part on a set of features of the test image and the respective centroids of the one or more clusters; generating, by one or more processors, a reduced distance vector for each of the one or more distance vectors of the test image, wherein generating the reduced distance vector further comprises: ordering, by one or more processors, the respective set of features in ascending order with values identified for each distance vector corresponding to the respective set of features of the test image; and selecting, by one or more processors, from the ascending order the set of features of the test image above a defined threshold to generate a reduced representation of the set of features of the test image as the reduced distance vector; generating, by one or more processors, a score vector based on the one or more distance vectors of the test image; and assigning, by one or more processors, an anomaly classification to the test image based on the score vector.
2. The method of claim 1, further comprising: identifying, by one or more processors, features corresponding to the test image with a deep learning cognitive model; and extracting, by one or more processors, a set of features corresponding to the test image, wherein the set of features comprises a high-dimensional dataset.
3. The method of claim 1, wherein assigning the anomaly classification to the test image based on the score vector further comprises: generating, by one or more processors, a class corresponding to each of the one or more clusters of the set of training images; and determining, by one or more processors, whether the score vector of the test image is within a range of distances of the classes of the one or more clusters.
4. The method of claim 1, wherein the clustering algorithm utilized is a k-means clustering model.
5. A computer program product comprising program instructions for implementing the steps in one of claims 1-4.
6. A computer system comprising: one or more computer processors; one or more computer-readable storage media; and program instructions stored on the computer-readable storage media for execution by at least one of the one or more processors, the program instructions comprising program instructions for implementing the steps in one of claims 1-4.
Citation Information
Patent Citations
Face image recognition method and device, electronic equipment and storage medium
CN109829433A
Method for automatic facial impression transformation, recording medium and device for performing the method
US20180268207A1