Procedure, program product and system for data identification
By recording data with time information and analyzing clustering probabilities, the method improves the detection of maliciously generated erroneously accepted data in supervised learning processes, enhancing accuracy and reliability in credit checks and insurance evaluations.
Patent Information
- Application Number
- DE112012003110
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2012-04-26
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2032-04-26
AI Technical Summary
Existing supervised machine learning methods struggle to accurately detect and prevent erroneously accepted data generated maliciously, leading to potential damage and unnoticed errors in processes like credit checks and insurance evaluations.
Record data with time information and perform clustering on both learning and test data, summarizing probability densities for each time interval, and analyze relative frequencies to identify statistically significant abnormalities indicative of malicious attacks.
Enhances the detection of malicious data with high accuracy, reducing reliance on data homogeneity and abnormality characteristics, and providing early warnings for potential attacks.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical area
[0001] The present invention relates to data identification using supervised machine learning and, more particularly, to a method for handling attacks in which data is maliciously manipulated. State of the art
[0002] Until now, for example, assessing damage claims at insurance companies and reviewing and granting loans and credit cards at financial institutions were essential and important tasks, and experienced experts were responsible for these tasks. However, as the number of tasks to be processed increases today, such tasks can no longer be completed manually by experts.
[0003] Accordingly, a process was recently implemented to assess claims for damages and to examine and grant loans and credit cards using a machine learning process using a computer.
[0004] Data used for loan evaluations and screenings is submitted by applicants and includes yes / no answers to questions, information such as age and annual income, and other descriptive text data. When such data is submitted on paper, designated clerks enter the data using computer keyboards or optical character recognition (OCR) to convert the data into electronic formats. In contrast, when applicants submit the data to a server using web browser operations, the data does not need to be converted into electronic formats.
[0005] If electronic applications are collected in this way, the experts first examine each application's details and determine for each application whether it is accepted or rejected, and then label the application accordingly with an electronic label. A supervised (training) dataset containing pairs, each of which consists of a feature vector x i (i = 1, ..., n) and a decision result (class label) y i (i = 1, ..., n) for each piece of application data, and which represents the above-described prior decision made by the experts, is defined as follows. DTraining={(x1,y1),…,(xn,yn)} Here y i ∈ C, where C represents a set of class labels. For example, C = {0, 1}, where 1 represents acceptance and 0 represents rejection.
[0006] An example of such a training set is shown in Fig. 1. That is, monitored data includes accepted (label 1) data 102, 104, 106, and 108, and rejected (label 0) data 110, 112, and 114. These data correspond to the individual application.
[0007] A supervised machine learning system configures a classifier using this training data. The classifier corresponds to a function h such as h:x→y where x represents a feature vector for the application and y represents a label for the application.
[0008] After the classification element is configured as described above, Fig. 2 Applications when test data is classified using the classification element. That is, data 202, 204, 206, and 208 are classified as accepted data, whereas data 210, 212, 214, and 216 are classified as rejected data. The main focus here is on data 208 and 210. If the data were properly classified, they should have been classified as rejected data; however, data 208 was classified as accepted data by the classification element and was incorrectly labeled as accepted data (FP = false positive). If data 210 had been properly classified, it should have been classified as accepted data; however, the data was classified as rejected data by the classification element and is labeled as incorrectly accepted data (FP = false positive).If the data 210 had been correctly classified, it would have been classified as accepted data; however, the data 210 was classified as rejected data by the classification element and is referred to as falsely rejected data (FN = false negative).
[0009] The classification element is configured based on probability. Accordingly, even using any machine learning program, it is difficult to completely eliminate falsely accepted data and falsely rejected data.
[0010] The classification element classifies test data of a sample, and the classification result is as in Fig. Figure 3 illustrates that data 302, 304, 306, 308, 310, and 312 are classified as accepted data, whereas data 314, 316, 318, 320, and 322 are classified as rejected data. Considering the classification result, assume that a malicious person accidentally encounters data 312 that was mistakenly accepted. The malicious person can analyze the content described in data 312 and obtain knowledge that they maliciously use, such as that elements of it need to be rewritten and how these elements need to be rewritten to make data that should be rejected into accepted data. They can create a manual using this knowledge. For example, this manual may be titled "How to Make an Insurance Claim That is Very Unlikely to Be Accepted, Very Easily Accepted."The malicious person could sell this manual, and people who have read the manual could create and submit a series of damage cases that falsely become accepted data, as indicated by reference numeral 324 in . Fig. 3 is illustrated.
[0011] The following documents describe known technologies for detecting such a malicious attack.
[0012] In the document Shohei Hido, Yuta Tsuboi, Hisashi Kashima, Masashi Sugiyama, Takafumi Kanamori, “Inlier-based Outlier Detection via Direct Density Ratio Estimation”, ICDM 2008 http: / / sugiyama-www.cs.titech.ac.jp / ~sugi / 2008 / ICDM2008.pdf, a method is disclosed in which an anomaly is detected by obtaining a density ratio between training data and test data.
[0013] In the paper by Daniel Lowd and Christopher Meek, "Adversarial Learning," KDD 2005 http: / / portal.acm.org / citation.cfm?id=1081950, an algorithm in the field of spam filtering is disclosed that aims to continuously handle a situation in which a single attacker conducts an attack using various techniques. The algorithm defines a distance from an ideal sample that the attacker wishes to forward as the adverse cost, and detects a sample with the least adverse cost (the first sample the attacker wishes to forward among the samples they can forward) and a sample with an adverse cost that is k times the minimum adverse cost, given a polynomial function number of attacks.
[0014] The document by Adam J. Oliner, Ashutosh V. Kulkarni, and Alex Aiken, "Community Epidemic Detection using Time-Correlated Anomalies," RAID 2010 (http: / / dx.doi.org / 10.1007 / 978-3-642-15512-3_19), describes a method for detecting a malicious attack when a computer is subject to a malicious attack. When multiple clients are grouped together under the same condition, the difference in behavior between the environments is calculated as an anomaly level. A situation in which the anomaly level temporarily increases for a single client may occur even in a normal situation, whereas a situation in which the anomaly levels increase for a certain number of abnormal clients at the same time indicates the occurrence of an attack. This is called a time-correlated anomaly, and a monitoring method for detecting a time-correlated anomaly is proposed.
[0015] Masashi Sugiyama's paper, "Kyouhenryoushifutokadeno kyoushitsuki gakushu" ("Supervised Learning under Covariate Shift"), Nihon Shinkei Kairo Gakkaishi (The Brain & Neural Networks), Vol. 13, No. 3, 2006, describes a discussion on how to correct a prediction model in supervised learning when training data and test data have different probability distributions. Specifically, this paper describes a method for increasing the importance of training data examples that are present in a range where test data frequently occurs, so that test data can be successfully classified.
[0016] According to the state of the art described above, a malicious attack can be detected in a specific situation. However, the state of the art has the limiting problem that properties specific to data, such as data homogeneity and degrees of anomaly, are assumed for individual data items. Another problem is that while a degree of vulnerability can be assessed, the fact that a "saturation attack" is being carried out using data that was falsely assumed cannot be detected. List of references
[0017] Document EP 1 376 420 A1 describes a system and method for processing an electronic document in a system that handles electronic documents in accordance with their classification into electronic document classes.
[0018] The document WO 2010 / 111 748 A1 describes methods and / or systems for processing, detecting and / or reporting the presence of anomalies or rare events in data. Non-patent literature Shohei Hido, Yuta Tsuboi, Hisashi Kashima, Masashi Sugiyama, Takafumi Kanamori, “Inlier-based Outlier Detection via Direct Density Ratio Estimation,” ICDM 2008 Daniel Lowd, Christopher Meek, “Adversarial Learning,” KDD 2005 http: / / portal.acm.org / citation.cfm?id=1081950 Adam J. Oliner, Ashutosh V. Kulkarni, Alex Aiken, Community Epidemic Detection using Time-Correlated Anomalies, RAID 2010 http: / / dx.doi.org / 10.1007 / 978-3-642-15512-3_19 Masashi Sugiyama, “Kyouhenryoushifutokadeno kyoushitsuki gakushu” (“Supervised Learning under Covariate Shift”) Nihon Shinkei Kairo Gakkaishi (The Brain & Neural Networks), Vol. 13, No. 3, 2006
[0019] The document "Strategies for detecting fraudulent claims in the automobile insurance industry" by Stijn Viaene, Mercedes Ayuso, Montserrat Guillen, Dirk Van Gheel, Guido Dedene, published in "European Journal of Operational Research, Vol. 176, 2007, pp. 565-583. - doi:10.1016 / j.ejor.2005.08.005" describes automatic detection systems that are used to decide whether or not to investigate claims suspected of fraud. Brief description of the invention
[0020] The invention is described by the features of the independent claims. Embodiments are specified in the dependent claims. Technical problems
[0021] It is accordingly an object of the present invention to provide a method by which, by means of supervised machine learning, falsely accepted data that was maliciously generated can be detected with a high degree of accuracy in a process in which examinations and evaluations of application documents are carried out.
[0022] It is a further object of the present invention to prevent an increase in damage by using an indication of an unavoidable erroneous decision in the process of performing reviews and evaluations of application documents by means of supervised machine learning.
[0023] It is yet another object of the present invention to prevent, by means of supervised machine learning, a situation in which damage occurs but is not noticed in the process of performing checks and evaluations of application documents. Solution to the problems
[0024] The present invention was created to solve the above problems. According to the present invention, both when creating supervised (learning) data and when creating test data, the data is recorded with time information appended to the data. This time is, for example, the time at which the data was input.
[0025] Subsequently, the system according to the present invention performs clustering on the training data in a target class (usually an acceptance class). Similarly, the system performs clustering on the test data in the target class (usually the acceptance class).
[0026] Subsequently, the system according to the present invention aggregates an identification probability density for each of the subclasses obtained by clustering. The aggregate is performed on the training data for each of the time intervals having different time points and durations (widths), and it is performed on the test data for each of the time intervals in the last period having different durations.
[0027] Subsequently, the system according to the present invention obtains, as a relative frequency in each of the time intervals for each of the subclasses, a ratio between a probability density obtained when learning is performed and a probability density obtained when testing is performed. The system recognizes an input having a relative frequency that statistically and noticeably increases as an anomaly and issues a warning, thus specifically examining whether it is an anomaly caused by an attack. In other words, according to the results of the present invention, such a case potentially indicates a high possibility that a malicious person may circulate learning knowledge acquired from the learning data. Advantageous effects of the invention
[0028] According to the present invention, in a process of performing examinations and evaluations of application documents using supervised machine learning, in both the case where learning data is created and the case where test data is created, the data is recorded with time information attached to the data. Furthermore, a frequency for each of the time intervals after clustering for the learning data is compared with that for the test data, thereby enabling the detection of malicious data. Accordingly, malicious data can be detected with high accuracy without having to assume data-specific properties such as data homogeneity and anomaly levels for each individual piece of data, resulting in increased reliability of the examinations. Furthermore, even a social connection between attackers can be considered. Brief description of the drawings Fig. Figure 1 is a diagram illustrating a supervised machine learning process. Fig. 2 is a diagram illustrating a classification process using a classification element configured through a supervised machine learning process. Fig. Figure 3 is a diagram illustrating a condition in which a classifier configured through a supervised machine learning process is attacked using falsely accepted data. Fig. 4 is a block diagram of a hardware configuration for implementing the present invention. Fig. 5 is a block diagram showing a functional configuration for carrying out the present invention. Fig. Figure 6 is a diagram illustrating a flowchart of an analysis process of a training input. Fig. Figure 7 is a diagram illustrating a flowchart of a process for generating sub-classification elements. Fig. Figure 8 is a diagram illustrating a flowchart of an analysis process of test input data. Fig. Figure 9 is a diagram illustrating a flowchart of a frequency analysis process for each of the time windows. Fig. Figure 10 is a graph illustrating individual frequencies in subclasses for training data and test data. Fig. Figure 11 is a graph illustrating frequencies of data that may be abnormal data. Description of the embodiment
[0029] An embodiment of the present invention will be described below based on the drawings. The same reference numerals denote the same objects throughout the drawings unless otherwise specified. It should be noted that one embodiment of the present invention will be described, and the present invention is not intended to be limited to the explanation of this embodiment.
[0030] In relation to Fig. 4 is a block diagram illustrating computer hardware for implementing a system configuration and process according to an embodiment of the present invention. Fig. 4, a CPU 404, random access memory (RAM) 406, a hard disk drive (HDD) 408, a keyboard 410, a mouse 412, and a display 414 are connected to a system bus 402. The CPU 404 is preferably based on 32-bit or 64-bit architecture. For example, a Pentium (trademark) 4, Core (trademark) 2 Duo, and Xeon (trademark) from Intel Corp., and an Athlon (trademark) from AMD Inc. can be used for the CPU 404. The random access memory 406 preferably has a capacity of 4 GB or more. The hard disk drive 408 desirably has, for example, a capacity of 500 GB or more for storing training data and test data for a large amount of application data such as claim evaluations in an insurance company and credit and credit card approvals in a financial company.
[0031] Hard disk drive 408 prestores an operating system, not specifically illustrated. The operating system may be any system compatible with CPU 404, such as Linux (trademark), Windows XP (trademark), or Windows (trademark) 2000 from Microsoft Corp., or Mac OS (trademark) from Apple Computer, Inc.
[0032] Hard disk drive 408 can store programming language processors such as C, C++, C#, and Java (trademark). These programming language processors are used to create and maintain routines or tools for the processes according to the present invention in the following manner. Hard disk drive 408 also includes development environments such as text editors for writing source code to be compiled using the programming language processors, and Eclipse (trademark).
[0033] The keyboard 410 and mouse 412 are used to activate the operating system or programs (not illustrated) loaded from the hard disk drive 408 into the memory 406 and displayed on the display 414, and to enter characters.
[0034] The display 414 is preferably a liquid crystal display. For example, a display of any resolution, such as XGA (1024 × 768 resolution) or UXGA (1600 × 1200 resolution), may be used for the display 414. The display 414 is used to display clusters containing erroneously accepted data that may have been maliciously generated (not illustrated).
[0035] Fig. Figure 5 is a functional block diagram illustrating processing routines, training data 502, and test data 504 according to the present invention. These routines are written using existing programming languages such as C, C++, C#, and Java (trademark) and stored in the hard disk drive 408 in executable binary format. The routines are called into memory 406 in response to operations from the mouse 412 or keyboard 410 and by means of functions of the operating system (not illustrated) so that they can be executed.
[0036] The training data 502 is stored on the hard disk drive 408 and has the following data structure. D(Trainig)={(x1(Training),y1(Training),t1(Training)),…,(xn(Training),yn(Training),tn(Training))}
[0037] In this data structure, x i (Training) represents a feature vector for the i-th training data, yi (Training) represents a class label for the i-th training data, and t i (Training) represents a timestamp for the i-th training data. The feature vector x i (Training) (i = 1, ..., n) is generated from elements in the electronic application data, preferably automatically by a computer process. When generating the feature vector, a technology such as text mining may be used. The class label y i (Training) (i = 1, ..., n) is set according to the result determined by a responsible competent expert who has previously reviewed the application data. At timestamp t i (Training) It is preferably the date and time when the application data was entered and, for example, it has a date and time format.
[0038] A routine for generating a classifier 506 has the function of generating a classification parameter 508, which uses a classifier 510 to classify the test data 504 based on the training data 502.
[0039] The test data 504 is stored on the hard disk drive 408 and has the following data structure. D'(Test)={(x1(Test),t1(Test)),…,(xm(Test),tm(Test))}
[0040] In this data structure, x i (Test) represents a feature vector for the i-th test data, and t i (Test) represents a timestamp of the i-th test data. The feature vector x i (test) (i = 1, ..., m) is generated automatically based on elements in the electronic application data, preferably by means of a computer process. The timestamp t i (Test)is preferably the date and time the application data was entered and has, for example, a date and time format.
[0041] The classification element 510 adds each individual test data (x i (Test) , t i (Test) ) a class label y i (Test) via a known supervised machine learning process. The function of the classifier 510 can be described as a function h(), and the equation y i (Test) = h(x i (Test) ) be used.
[0042] Known supervised machine learning is roughly classified into classification analysis and regression analysis. The supervised machine learning that can be used for the purpose of the present invention originates from the field of classification analysis. Methods known as classification analysis include linear classifiers such as the linear Fisher discriminant function, logistic regression, Naive Bayes classification, and perception. Apart from the above, the methods include a quadratic classifier, the k-nearest neighbor algorithm, boosting, a decision tree, a neural network, a Bayesian network, a support vector machine, and the hidden Markov model. For the present invention, any of these methods can be selected. However, according to this embodiment, a support vector machine is specifically used.For a more detailed description, see, for example, Christopher M. Bishop, “Pattern Recognition And Machine Learning,” 2006, Springer Verlag.
[0043] The classifier 510 reads the test data 504 and adds a class label to the test data 504 to generate classified data 512 expressed in the following equation. D'(Test)={(x1(Test),t1(Test),t1(Test)),…,(xm(Test),tm(Test))}
[0044] A cluster analysis routine 524 defines a distance, such as the Euclidean distance and the Manhattan distance, between the feature vectors of the data in the training data 502 and performs clustering by a known method, such as the k-means algorithm, using this distance to generate partitioning data 516, which is the result of the clustering. The partitioning data 516 is preferably stored on the hard disk drive 408. Since the partitioning data 516 specifies position data, such as boundaries or centers of clusters, a decision can be made as to which of the data should belong to which cluster by referring to the partitioning data 516. In short, the partitioning data 516 serves as a sub-classifier.It should be noted that the clustering method that can be used for the present invention is not limited to k-means, and any clustering method compatible with the present invention can be applied, such as a Gaussian mixture model, agglomerative clustering, branch clustering, and self-organizing maps. Alternatively, partitioned data sets can be obtained using grid division.
[0045] The cluster analysis routine 514 writes the partitioning data 516, which represents the result of the clustering, to the hard disk drive 408.
[0046] A time series analysis routine 518 reads the training data 502, calculates a data frequency and other statistics for each predetermined time window for each of the clusters (subclasses) corresponding to the partitioning data 516, and stores the result as time series data 520, preferably on the hard disk drive 408.
[0047] A time series analysis routine 522 reads the classified data 512, calculates a data frequency and other statistical data for each predetermined time window for each of the clusters (subclasses) corresponding to the partitioning data 516, and stores the result as time series data 524, preferably on the hard disk drive 408.
[0048] An anomaly detection routine 526 calculates data with respect to a time window for a cluster for the time series data 520 and with respect to a corresponding cluster for the time series data 524. The anomaly detection routine 526 has a function of activating a warning routine 528 when the result value is greater than a predetermined threshold.
[0049] The warning routine 528 has a function of displaying, for example, the cluster and the time window in which the anomaly is detected on the display 414 to inform an operator of the anomaly.
[0050] Regarding the schedules in the Fig. 6 to 9, the processes carried out are described individually below. Fig. Figure 6 is a diagram illustrating a flowchart of a training data analysis process.
[0051] In step 602 in Fig. 6, the classifier generation routine 506 generates the classification parameter 508 to generate the classifier 510.
[0052] In step 604, the cluster analysis routine 514 generates a subclassification element, ie, the partitioning data 516 for clustering.
[0053] In step 606, the time series analysis routine 518 calculates an input frequency statistic for each of the time windows for each of the subclasses to generate the time series data 520.
[0054] Fig. 7 is a diagram illustrating a flowchart describing the process, particularly in step 604. In other words, in this process, the cluster analysis routine 514 loops from step 702 to step 706 on each of the classes and, in step 704, generates a subclassifier for the data in the class.
[0055] It should be noted that in the process of the schedule in Fig. 7 Not all classes need to be subjected to the process. For example, if an attack needs to be detected for a specific class, only that class needs to be subjected to the process.
[0056] Fig. Figure 8 is a diagram illustrating a flowchart of an analysis process on the test data. In a loop from step 802 to step 810, each individual piece of data contained in the test data 504 is subjected to the process.
[0057] In step 804, the classifier 510 classifies each piece of data in the test data 504. Then, in step 806, the time series data analysis routine 522 classifies the classified data into a subclass (i.e., clustering) based on the partitioning data 516. In step 808, the time series data analysis routine 522 increases the input frequency for the subclass in the current time window while transitioning from a time window of a predetermined width.
[0058] Once the process loop from step 802 to step 810 is completed for all individual data items contained in the test data 504, the time series data analysis routine 522 writes the time series data 524 to the hard disk drive 408.
[0059] Fig. 9 is a diagram illustrating a flowchart of a process in which the anomaly detection routine 526 detects a possible occurrence of an anomaly within a given time window. In step 902, the anomaly detection routine 526 calculates a ratio of a test input frequency with respect to a training data frequency within the time window.
[0060] In step 904, the anomaly detection routine 526 calculates an incremental score for a statistically significant frequency for each of the subclasses. Here, statistical significance means that a sufficient number of samples are generated. An incremental score for a significant frequency can be obtained through a simple ratio calculation. However, according to the embodiment, the following equation is used to more accurately calculate an incremental score.
[0061] The width of a time window is represented by W. A function g() represents a function of subclassing. In the time window, a set of input feature vectors, labeled as j at time t, is expressed in the following equation: Xt(mode)(j)={xi(mode)|g(xi(mode))=j,t−W≤ti(mode)≤t}
[0062] Here, "mode" represents either "Training," i.e., the training data, or "Test," i.e., the test data. A probability of occurrence for input data with a label j is defined in the following way. Pt(mode)(j)=P(Xt(mode)(j))
[0063] Then, the increasing score of the anomaly is represented as the following equation. q(xk(test))≡Ps(test)(j)E(Pt(training)(j))(σ(Pt(training)(j)))+1)
[0064] Here, s=tk(test) and j=g(xk(test)).
[0065] In this equation, E() represents an expectation and σ() represents a deviation.
[0066] This equation essentially uses a moving average of frequencies and a moving average variance. Alternatively, a frequency transformation such as a wavelet transform can be applied to account for the periodic fluctuation of a relative frequency.
[0067] In step 906, the anomaly detection routine 526 determines whether the increasing score of the anomaly value exceeds a threshold. If the value exceeds a threshold, the warning routine 528 is activated in step 908, and information about a possibility that the subclass may be irregular is displayed on the display 414.
[0068] In making this decision, weighting may be added according to the magnitude of the cost for each of the samples, or natural variation may be identified using manipulation features that may cause an attack.
[0069] The process of the schedule in Fig. 9 is performed for each of the time slots.
[0070] Fig. 10 includes graphs for the training data and for the test data, illustrating data distributions over time for each of the subclasses A1, A2, ..., and An of a class A. In the process of the present invention, the possibility of occurrence of an anomaly is detected using a frequency ratio between the training data and the test data in a predetermined time window for the same subclass of the same class.
[0071] Fig.Figure 11 illustrates an example in which such a possibility of an anomaly occurrence is detected. In other words, in a certain time interval, as indicated by reference numeral 1104, the anomaly detection routine 526 detects a state in which a frequency of the test data is substantially similar to a frequency of the training data in the fourth cluster (subclass) and notifies the warning routine 528 that irregular data may be present.
[0072] By activating the warning routine 528, an operator is notified that the data in the cluster may have a problem during the time window, and they can narrow down the amount of data from which the problem needs to be identified. The data analysis result must be used to identify a detected misclassification that caused the attack, temporarily modifying the label and moving the data to a rejection set, providing an opportunity to modify the discriminatory model in the future.
[0073] Furthermore, while the input is subjected to detection, by limiting the detection to a case where subclasses with a property of frequent occurrence and causing a large deviation of statistics can be identified, a report can be generated, and only if it is assumed that, for example, a manual used to propagate automatic detection exists.
[0074] As described above, the present invention has been described based on the specific embodiment. It should be noted that the present invention is not limited to the specific embodiment, and various configurations and methods, such as modifications and substitutions, which are readily apparent to those skilled in the art, can be applied to the present invention.
[0075] For example, according to the embodiment, the application example was described in which the present invention is applied to the review of application documents for assessing claims in an insurance company and for reviewing and granting loans and credit cards in a financial company. However, the present invention can be applied to any documents that need to be reviewed in which the described content can be converted into feature vectors.
Claims
[1] A computer-implemented method for data identification for detecting an attack carried out using irregular data against a classifier (510) configured by means of supervised machine learning, the method comprising the steps of: - creating a plurality of individual training data (502), each comprising a feature vector, a class label and a timestamp for each individual training data (502); - generating (506, 602) the classification element (510) using the feature vector, the class label and the timestamp of the plurality of individual training data (502); - clustering the plurality of individual training data (502) into a plurality of subclasses based on a distance between the feature vectors of the plurality of individual training data (502); - creating a plurality of test data (504), each comprising a feature vector and a timestamp for each individual test data (504); - classifying the plurality of individual test data (504) using the classification element (510), wherein the classification element (510) adds a class label to each individual test data (504); - clustering the plurality of individual test data (504) that have been classified using the classification element (510) into the plurality of subclasses; - for each subclass, calculating (606, 902) statistical data representing a ratio of a frequency of the plurality of individual training data (502) clustered into the corresponding subclass and a frequency of the plurality of individual test data (504) clustered into the corresponding subclasses, wherein the ratio captures the frequency of the plurality of individual training data (502) and the frequency of the plurality of individual test data (504) in a predetermined time window for the same subclass of the same class, and - warning (528) of a possibility of the attack being carried out using the irregular data in one or more subclasses of the plurality of subclasses in response to a value of the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding a predetermined threshold (906). [2] The data identification method according to claim 1, wherein feature data is represented by the feature vector, the feature vector being obtained by converting a response to a question element in a financial application document into an electronic form, and the classes comprise an acceptance class and a rejection class. [3] The data identification method of claim 1, wherein the classifying element (510) is configured with a support vector machine. [4] The data identification method of claim 1, wherein the clustering uses a k-means algorithm. [5] A data identification method according to claim 2, wherein the irregular data is erroneously accepted data. [6] The data identification method according to claim 1, wherein the statistical data is calculated using a moving average of the frequency and a variance of the moving average. [7] A computer-executed data identification program product for detecting an attack carried out using irregular data against a classification element (510) configured by means of supervised machine learning, the program product causing a computer to perform the steps of: Creating a plurality of individual training data (502) each comprising a feature vector, a class label and a timestamp for each individual training data (502); Generating (506, 602) the classification element (510) using the feature vector, the class label and the timestamp of the plurality of individual training data (502); Clustering the plurality of individual training data (502) into a plurality of subclasses based on a distance between the feature vectors of the plurality of individual training data; Creating a plurality of test data (504), each comprising feature data, a feature vector, and a timestamp for each individual test data (504); Classifying the plurality of individual test data (504) using the classification element (510), wherein the classification element (510) adds a class label to each individual test data (504); Clustering the plurality of individual test data (504) that have been classified using the classification element (510) into the plurality of subclasses; Calculating (606, 902) statistical data for each subclass representing a ratio of a frequency of the plurality of individual training data items (502) clustered into the corresponding subclass and a frequency of the plurality of individual test data items (504) clustered into the corresponding subclasses, the ratio capturing the frequency of the plurality of individual training data items (502) and the frequency of the plurality of individual test data items (504) in a predetermined time window for the same subclass of the same class; and warning (528) of a possibility of the attack being carried out using the irregular data in one or more subclasses of the plurality of subclasses in response to a value of the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding a predetermined threshold (906). [8] A data identification program product according to claim 7, wherein feature data is represented by the feature vector, the feature vector being obtained by converting a response to a question element in a financial application document into an electronic form, and the classes comprise an acceptance class and a rejection class. [9] A program product for data identification according to claim 7, wherein the classification element (510) is configured by means of a support vector machine. [10] A data identification program product according to claim 7, wherein the clustering uses a k-means algorithm. [11] A data identification program product according to claim 8, wherein the irregular data is erroneously accepted data. [12] A data identification program product according to claim 7, wherein the statistical data is calculated using a moving average of the frequency and a variance of the moving average. [13] A computer-implemented data identification system for detecting an attack carried out using irregular data against a classifier (510) configured by means of supervised machine learning, the data identification system comprising: storage means; a plurality of individual training data (502), each comprising a feature vector, a class label and a timestamp, stored in the storage means; a classifier (510) generated (506) using the feature vector, the class label, and the timestamp of the plurality of individual training data (502); a cluster analysis routine (514) used to cluster the plurality of individual training data (502) into a plurality of subclasses based on a distance between the feature vectors of the plurality of individual training data (502); Data in the subclasses of the multitude of individual training data (502), wherein the data is generated by applying the clustering to the plurality of individual training data (502) and is stored in the storage means; a plurality of individual test data (504), each comprising a feature vector, a class label and a timestamp, stored in the storage means; Data in subclasses of the plurality of individual test data (504), the data being generated by applying the clustering to the plurality of individual test data (504) and being stored in the storage means; Calculation means for calculating (606, 902) statistical data for each subclass representing a ratio of a frequency of the plurality of individual training data items (502) clustered into the corresponding subclass and a frequency of the plurality of individual test data items (504) clustered into the corresponding subclasses, the ratio capturing the frequency of the plurality of individual training data items (502) and the frequency of the plurality of individual test data items (504) in a predetermined time window for the same subclass of the same class; and Warning means for warning (528) of a possibility of occurrence of the attack carried out using the irregular data in one or several subclasses of the plurality of subclasses, in response to a value of the statistical data calculated for the one or more subclasses of the plurality of subclasses exceeding a predetermined threshold (906). [14] The data identification system according to claim 13, wherein feature data is represented by the feature vector, the feature vector being obtained by converting a response to a question element in a financial application document into an electronic form, and the classes comprise an acceptance class and a rejection class. [15] The data identification system according to claim 13, wherein the classification element (510) is configured with a support vector machine. [16] The data identification system of claim 13, wherein the clustering uses a k-means algorithm. [17] A data identification system according to claim 14, wherein the irregular data is erroneously accepted data. [18] The data identification system according to claim 13, wherein the statistical data is calculated using a moving average of the frequency and a variance of the moving average.
Citation Information
Patent Citations
Method and system for classifying electronic documents
EP1376420A1
Systems and methods for detecting anomalies from data
WO2010111748A1
Cited By
Automatic inheritance of similar alert properties
US20240106694A1