An industrial data outlier detection system and method based on the PCA algorithm

Data processing through PCA algorithm dimensionality reduction and GrubbMAD algorithm, calculate exception scores to mark exception points, solving the problem of low accuracy of traditional algorithms in high-dimensional space, and significantly improving the accuracy of abnormal point detection in industrial high-dimensional data.

CN115828089BActive Publication Date: 2025-06-24JIANGSU ZHIHENG INFORMATION TECH SERVICES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211292098.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-06-24
Estimated Expiration
2042-10-21

AI Technical Summary

Technical Problem

Traditional anomaly point detection algorithms are not very accurate in high-dimensional space, making it difficult to effectively detect anomaly points in industrial high-dimensional data.

Method used

The PCA algorithm is used to reduce the dimensionality of the data set, and the GrubbMAD algorithm is used to calculate the anomaly degree value of each data point. The exception score of each data point is obtained by weight summing, thereby marking the exception points.

Benefits of technology

The accuracy of industrial high-dimensional data anomaly point detection is significantly improved, so that abnormal data points can be significantly different from normal data points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115828089B_ABST
    Figure CN115828089B_ABST
Patent Text Reader

Abstract

The present invention discloses an industrial data outlier detection system and method based on the PCA algorithm in the field of data mining technology, including: obtaining a data set to be detected; using the PCA algorithm to reduce the dimension of the data set to obtain the dimension-reduced data set and each eigenvalue; using the GrubbMAD algorithm to calculate the GM value of each data point in each dimension of the dimension-reduced data set; performing weighted summation on each eigenvalue and the GM value of each dimension to obtain the outlier score of each data point; and marking the outlier points in the data set according to the outlier score of each data point. The present invention uses the PCA algorithm and the GrubbMAD algorithm to process the data set, making the outlier data points significantly different from the normal data points, thereby effectively improving the accuracy of industrial high-dimensional data outlier detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an industrial data outlier detection system and method based on the PCA algorithm, belonging to the technical field of data mining. Background Art

[0002] In the field of industrial manufacturing, a large amount of operation data is generated during the production of industrial assembly lines and the operation of industrial equipment. These data are related to the quality of products produced by the assembly line and the stability of the operation of industrial equipment. However, almost all production processes may encounter situations such as machine failures and equipment anomalies, and the data generated at this time is also abnormal data. If the abnormality of the data can be detected in time, machine failures and equipment anomalies can be detected as early as possible. Therefore, using a good data outlier detection method to monitor the production of industrial assembly lines and the operation of industrial equipment, and detecting machine failures, equipment anomalies and other situations as early as possible will help enterprises take timely countermeasures and minimize the impact of abnormal failures on enterprises.

[0003] Traditional outlier detection algorithms are mostly based on distance metrics. For example, the LOF algorithm determines whether a point is an outlier by comparing the density of each point with its neighborhood. The density is defined and calculated based on the distance between points. The COF algorithm is similar to the LOF algorithm and determines outliers based on the connectivity between data points. The connectivity is also based on the distance between points. Such algorithms are more suitable for outlier detection tasks of small and low-dimensional data sets because in high-dimensional spaces, distance metrics will gradually lose their meaning. However, in the field of industrial manufacturing, the data generated by industrial production and equipment operation are mostly high-dimensional data, so the accuracy of traditional outlier detection algorithms is not high. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an industrial data outlier detection system and method based on the PCA algorithm. The PCA algorithm and the GrubbMAD algorithm are used to process the data set, so that the abnormal data points are significantly different from the normal data points, thereby effectively improving the accuracy of industrial high-dimensional data outlier detection.

[0005] To achieve the above purpose, the present invention is implemented by the following technical solutions:

[0006] In the first aspect, the present invention provides an industrial data outlier detection method based on the PCA algorithm, including:

[0007] Obtain the data set to be detected;

[0008] Use the PCA algorithm to reduce the dimension of the data set to obtain the reduced data set and the eigenvalue of each dimension;

[0009] Calculate the GM value of each data point in each dimension of the dataset after dimensionality reduction using the GrubbMAD algorithm;

[0010] Perform a weighted sum of each eigenvalue and the GM value of each dimension to obtain the anomaly score of each data point;

[0011] Mark the outliers in the dataset according to the anomaly scores of each data point.

[0012] Furthermore, use the PCA algorithm to perform dimensionality reduction on the dataset to obtain the dataset after dimensionality reduction and the eigenvalues of each dimension, including:

[0013] S1. Calculate the mean vector μ of the samples in the original dataset X, and the calculation formula is as follows:

[0014]

[0015] where: μ is the mean vector, n is the total number of data points in the dataset, i is the count, and x i represents the i-th data;

[0016] S2. Remove the mean from each sample, that is, centralize the sample data, and the calculation formula is as follows:

[0017]

[0018] where: is the data matrix after removing the mean, X is the original dataset, and μ is the mean vector;

[0019] S3. Construct the covariance matrix V of the data matrix The calculation formula is as follows:

[0020]

[0021] where: V is the covariance matrix, n is the total number of data points in the dataset, is the data matrix after removing the mean, is the transpose of;

[0022] S4. Perform eigenvalue decomposition on the covariance matrix V to obtain the eigenvalues λ j and the corresponding eigenvectors w j , and sort the eigenvalues λ j in descending order, where j is the count, representing the j-th dimension;

[0023] S5. Select the first k eigenvalues Λ = [λ1, λ2,..., λ k and the corresponding eigenvectors W = [w1, w2,..., w k as the basis of the subspace, then the k-dimensional data after dimensionality reduction is:

[0024]

[0025] Where D is the data after dimensionality reduction, is the data matrix after mean removal, and W T is the transpose of W.

[0026] Furthermore, the GrubbMAD algorithm is used to calculate the GM value of each data point in each dimension of the dataset after dimensionality reduction, including:

[0027] For the k-dimensional data D after dimensionality reduction, all data in the j-th dimension is D j , and the data of the i-th data point in the j-th dimension is d ij , then the calculation method of its GM value is as follows:

[0028]

[0029] Where: GM ij is the GM value of the i-th data in the j-th dimension, and median() represents the median.

[0030] Furthermore, a weighted sum is performed on each eigenvalue and each GM value of each dimension, and the calculation formula is:

[0031]

[0032] Where: i and j are the counting quantities, representing the data point and the dimension respectively, k is the total number of dimensions, and OS i is the anomaly score of the i-th data point, and λ j is the eigenvalue of the j-th dimension, and GM ij is the GM value of the i-th data in the j-th dimension.

[0033] Furthermore, according to the anomaly score of each data point, the dataset is marked for anomaly points, including:

[0034] Traverse each data point in the dataset;

[0035] In response to a signal that the anomaly score of a certain data point is greater than the threshold, mark that data point as an anomaly point;

[0036] In response to a signal that the anomaly score of a certain data point is not greater than the threshold, mark that data point as a normal point;

[0037] In response to marking all anomaly points, the process ends.

[0038] Furthermore, the method for determining the threshold value is: arrange the data points in the dataset to be detected in descending order according to their anomaly scores, and take the anomaly score of the int(m*0.1)-th point as the threshold, where m is the total number of data points and int represents rounding down.

[0039] In a second aspect, the present invention provides an industrial data outlier detection system based on the PCA algorithm, comprising:

[0040] A dataset acquisition module: used to acquire the dataset to be detected;

[0041] A dimensionality reduction module: used to reduce the dimensionality of the dataset using the PCA algorithm to obtain the reduced-dimensional dataset and the eigenvalue of each dimension;

[0042] A GM value calculation module: used to calculate the GM value of each data point in each dimension of the reduced-dimensional dataset using the GrubbMAD algorithm;

[0043] An outlier score calculation module: used to perform weighted summation on the eigenvalue of each dimension and the GM value of each dimension to obtain the outlier score of each data point;

[0044] An outlier marking module: used to mark the outliers in the dataset according to the outlier score of each data point.

[0045] In a third aspect, the present invention provides an industrial data outlier detection device based on the PCA algorithm, comprising a processor and a storage medium;

[0046] The storage medium is used to store instructions;

[0047] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the above.

[0048] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the above are implemented.

[0049] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0050] The present invention innovatively designs a PGOF outlier detection algorithm improved based on the PCA algorithm. After reducing the dimensionality of industrial high-dimensional data using the PCA algorithm, the eigenvalue of each dimension of the data is weighted and summed with each dimension of the data processed using the GrubbMAD algorithm, so that the outlier data points are significantly different from the normal data points, thereby effectively improving the accuracy of industrial high-dimensional data outlier detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of an industrial data outlier detection method based on the PCA algorithm provided in Embodiment 1 of the present invention;

[0052] Figure 2 is an overall schematic diagram of the PGOF algorithm provided in Embodiment 1 of the present invention. Detailed implementation mode

[0053] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0054] Embodiment 1:

[0055] An industrial data outlier detection method based on the PCA algorithm, the detailed steps are as follows:

[0056] Step 01: Obtain the industrial data set Data to be detected n (m n-dimensional data points), and proceed to Step 02.

[0057] Step 02: Use the PCA algorithm to reduce the dimension of the data set Data n to obtain the reduced data set Data k (m k-dimensional data points), and the eigenvalue λ of each dimension j , and proceed to Step 03.

[0058] Step 03: Use the GrubbMAD algorithm to process each dimension of the data in the data set Data k to obtain the outlier degree value GM of each data point in each dimension ij , and proceed to Step 04.

[0059] Step 04: Perform weighted summation on the eigenvalue λ of each dimension j and GM ij to obtain the outlier score OS of each data point i (Outlier_Score), and proceed to Step 05.

[0060] Step 05: Traverse the data set Data k , for each data point: if the outlier score of the data point > the threshold threshold, mark the data point as an outlier; if the outlier score of the data point <= the threshold threshold, mark the data point as a normal point. Until all outliers are marked, the process ends.

[0061] The PCA algorithm in Step 02, that is, the principal component analysis method, is a widely used data dimension reduction algorithm. Its main idea is to map the n-dimensional features of the original data to new abstract k-dimensional features. Generally speaking, n >= k, that is, to find a low-dimensional representation of high-dimensional data, so as to achieve the purpose of dimension reduction. Specifically:

[0062] (1) Calculate the mean vector μ of the samples in the original data set X, and the calculation formula is as follows:

[0063]

[0064] where μ is the mean vector, n is the total number of data points in the data set, i is the count, and x i represents the i-th data.

[0065] (2) De-mean each sample, that is, center the sample data. The calculation formula is as follows:

[0066]

[0067] where is the data matrix after de-meaning, X is the original data set, and μ is the mean vector.

[0068] (3) Construct the covariance matrix V of the data matrix The calculation formula is as follows:

[0069]

[0070] where V is the covariance matrix, n is the total number of data points in the data set, is the data matrix after de-meaning, is the transpose of.

[0071] (4) Perform eigenvalue decomposition on the covariance matrix V to obtain the eigenvalues λ j and the corresponding eigenvectors w j , and sort the eigenvalues λ in descending order j , where j is the count, representing the j-th dimension.

[0072] (5) Select the first k eigenvalues Λ = [λ1, λ2,..., λ k and the corresponding eigenvectors W = [w1, w2,..., w k as the basis of the subspace. Then the k-dimensional data after dimensionality reduction is:

[0073]

[0074] where D is the data after dimensionality reduction, is the data matrix after de-meaning, and W T is the transpose of W.

[0075] The specific details of the GrubbMAD algorithm in step 03 are as follows:

[0076] For the k-dimensional data D after dimensionality reduction, all the data in the j-th dimension is D j , and the data in the j-th dimension of the i-th data point is d ij . Then the calculation method of its GM value is as follows:

[0077]

[0078] Where: GM ij is the GM value of the i-th data in the j-th dimension, and median() represents the median. The denominator uses the median of the median and the median absolute deviation to repair the influence deviation of the outliers in the j-th dimensional data on the overall mean and standard deviation of the j-th dimensional data, so as to make the outliers as prominent as possible.

[0079] The weighted summation in step 04 is calculated as follows:

[0080]

[0081] Where: i and j are counting numbers, representing data points and dimensions respectively, k is the total number of dimensions, and OS i is the outlier score of the i-th data point, and λ j is the eigenvalue of the j-th dimension, and GM ij is the GM value of the i-th data in the j-th dimension.

[0082] The specific value of the threshold threshold in step 05 needs to be given according to specific data. The setting method is: sort the data points in Data n (m n-dimensional data points) in descending order according to their outlier scores OS i and take the outlier score of the int(m*0.1)-th point as the threshold, where m is the total number of data points and int represents rounding down.

[0083] Embodiment 2:

[0084] An industrial data outlier detection system based on the PCA algorithm can implement the industrial data outlier detection method described in Embodiment 1, including:

[0085] Dataset acquisition module: used to acquire the dataset to be detected;

[0086] Dimensionality reduction module: used to reduce the dimensionality of the dataset using the PCA algorithm to obtain the reduced dataset and the eigenvalue of each dimension;

[0087] GM value calculation module: used to calculate the GM value of each data point in each dimension of the reduced dataset using the GrubbMAD algorithm;

[0088] Outlier score calculation module: used to perform weighted summation on the eigenvalue of each dimension and the GM value of each dimension to obtain the outlier score of each data point;

[0089] Outlier marking module: used to mark the outliers in the dataset according to the outlier score of each data point.

[0090] Embodiment 3:

[0091] An embodiment of the present invention further provides an industrial data outlier detection device based on the PCA algorithm, which can implement the industrial data outlier detection method based on the PCA algorithm described in Embodiment 1, including a processor and a storage medium;

[0092] The storage medium is used to store instructions;

[0093] The processor is used to operate according to the instructions to execute the steps of the following method:

[0094] Obtain the data set to be detected;

[0095] Use the PCA algorithm to reduce the dimension of the data set to obtain the reduced data set and the eigenvalue of each dimension;

[0096] Use the GrubbMAD algorithm to calculate the GM value of each data point in each dimension of the reduced data set;

[0097] Perform weighted summation on the eigenvalue of each dimension and the GM value of each dimension to obtain the outlier score of each data point;

[0098] Mark the outliers in the data set according to the outlier score of each data point.

[0099] Embodiment 4:

[0100] An embodiment of the present invention further provides a computer-readable storage medium, which can implement the industrial data outlier detection method based on the PCA algorithm described in Embodiment 1. A computer program is stored thereon, and when the program is executed by a processor, it implements the steps of the following method:

[0101] Obtain the data set to be detected;

[0102] Use the PCA algorithm to reduce the dimension of the data set to obtain the reduced data set and the eigenvalue of each dimension;

[0103] Use the GrubbMAD algorithm to calculate the GM value of each data point in each dimension of the reduced data set;

[0104] Perform weighted summation on the eigenvalue of each dimension and the GM value of each dimension to obtain the outlier score of each data point;

[0105] Mark the outliers in the data set according to the outlier score of each data point.

[0106] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0107] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0108] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0110] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. An industrial data outlier detection method based on the PCA algorithm, characterized in that, Including: Obtain the dataset to be detected; Use the PCA algorithm to reduce the dimension of the dataset, obtaining the reduced-dimension dataset and the eigenvalue of each dimension; Use the GrubbMAD algorithm to calculate the GM value of each data point in each dimension in the reduced-dimension dataset; Perform weighted summation on the eigenvalue of each dimension and the GM value of each dimension to obtain the anomaly score of each data point; Mark the anomaly points in the dataset according to the anomaly score of each data point; Among them, using the GrubbMAD algorithm to calculate the GM value of each data point in each dimension in the reduced-dimension dataset includes: For the k-dimensional data D after dimensionality reduction, all the data in the j-th dimension is D j , and the data in the j-th dimension of the i-th data point is d ij . Then the calculation method of its GM value is as follows: Where: GM ij is the GM value of the i-th data in the j-th dimension, and median() represents the median.

2. The industrial data outlier detection method based on the PCA algorithm according to claim 1, characterized in that, Using the PCA algorithm to reduce the dimension of the dataset, obtaining the reduced-dimension dataset and the eigenvalue of each dimension includes: S1. Calculate the mean vector μ of the samples in the original dataset X, and the calculation formula is as follows: where: μ is the mean vector, n is the total number of data points in the data set, i is the count, and x i represents the i-th data; S2. Remove the mean value from each sample, that is, centralize the sample data, and the calculation formula is as follows: Wherein: is the data matrix after mean removal, X is the original data set, and μ is the mean vector; S3. Construct a data matrix The covariance matrix V is calculated as follows: Where: V is the covariance matrix, n is the total number of data points in the data set, is the data matrix after mean removal, is the transpose of; S4. Perform eigenvalue decomposition on the covariance matrix V to obtain the eigenvalues λ for each dimension j and the corresponding eigenvectors w j , and sort the eigenvalues λ in descending order j , where j is the count, representing the j-th dimension; S5. Select the first k eigenvalues Λ = [λ1, λ2,..., λ k and the corresponding eigenvectors W = [w1, w2,..., w k as the basis of the subspace. Then the k-dimensional data after dimensionality reduction is: where D is the data after dimensionality reduction, is the data matrix after mean removal, and W T is the transpose of W.

3. The industrial data outlier detection method based on the PCA algorithm according to claim 1, characterized in that, The formula for performing weighted summation on the eigenvalue of each dimension and the GM value of each dimension is: where: i and j are counting numbers, representing data points and dimensions respectively, k is the total number of dimensions, and OS i is the anomaly score of the i-th data point, and λ j is the eigenvalue of the j-th dimension, and GM ij is the GM value of the i-th data in the j-th dimension.

4. The industrial data outlier detection method based on the PCA algorithm according to claim 1, characterized in that, Marking the anomaly points in the dataset according to the anomaly score of each data point includes: Traverse each data point in the dataset; In response to a signal that the anomaly score of a certain data point is greater than the threshold, mark the data point as an anomaly point; In response to a signal that the anomaly score of a certain data point is not greater than the threshold, mark the data point as a normal point; In response to marking all the anomaly points, the process ends.

5. The industrial data outlier detection method based on the PCA algorithm according to claim 4, characterized in that, The way to determine the threshold value is: sort the data points in the dataset to be detected in descending order according to their anomaly scores, and take the anomaly score of the int(m*0.1)th point as the threshold, where m is the total number of data points, and int represents rounding down.

6. An industrial data outlier detection system based on the PCA algorithm, characterized in that, Including: Dataset acquisition module: used to obtain the dataset to be detected; Dimension reduction module: used to use the PCA algorithm to reduce the dimension of the dataset, obtaining the reduced-dimension dataset and the eigenvalue of each dimension; GM value calculation module: used to use the GrubbMAD algorithm to calculate the GM value of each data point in each dimension in the reduced-dimension dataset; Anomaly score calculation module: used to perform weighted summation on the eigenvalue of each dimension and the GM value of each dimension to obtain the anomaly score of each data point; Anomaly point marking module: used to mark the anomaly points in the dataset according to the anomaly score of each data point; Among them, using the GrubbMAD algorithm to calculate the GM value of each data point in each dimension in the reduced-dimension dataset includes: For the k-dimensional data D after dimensionality reduction, all the data in the j-th dimension is D j , and the data in the j-th dimension of the i-th data point is d ij , then the calculation method of its GM value is as follows: where: GM ij is the GM value of the i-th data in the j-th dimension, and median() represents the median.

7. An industrial data outlier detection device based on the PCA algorithm, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • An abnormality detection method of injection molding machine blocking based on ensemble learning

    CN109145948A

  • Mobile phone key detection method and system based on data classification matching

    CN114594865A