Anomaly detection method for mining scenarios based on graph-regularized incremental non-negative matrix factorization
Through the method of regular incremental non-negative matrix decomposition based on graph, the existing anomaly detection methods are solved in timeliness and network resource utilization, and dynamic training and comprehensive analysis of multi-sensor data are realized. It can track and accurately detect abnormalities in mining scenarios in real time, analyze the causes of abnormalities, and effectively utilize network resources.
Patent Information
- Application Number
- CN202111509423.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-12-10
AI Technical Summary
Existing anomaly detection methods are difficult to take into account both timeliness and network resource utilization, and it is difficult to effectively analyze multi-sensor data to determine the root cause of scene abnormalities.
The method based on regular incremental non-negative matrix decomposition of graph is adopted to realize dynamic data training and comprehensive analysis of multi-sensor data. Through regular incremental non-negative matrix decomposition of graph, the data content can be represented online, the geometric structure of the data is maintained, and combined with incremental learning, reducing repeated calculations and reducing operation time.
Dynamic training and real-time tracking are realized, timeliness is ensured, and multiple factors can be considered comprehensively to ensure the accuracy of detection, and the causes of abnormalities can be analyzed, network resources can be used effectively, and efficiency can be improved.
Smart Images

Figure CN114238854B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of anomaly detection and diagnosis technology and intelligent security, and in particular to a mining scene anomaly detection method based on graph regularized incremental non-negative matrix decomposition. Background Art
[0002] Mining safety has always been a hot topic of concern. At present, most mines, especially metal mines, use underground mining. However, during underground construction, toxic gases, mine collapse, etc. pose a huge threat to the safety of construction workers; therefore, it is of great significance to conduct risk detection and early warning for underground construction environments. At present, there are many anomaly detection methods, but it is a challenging task to ensure timeliness, ensure the effective use of network resources, and comprehensively analyze multi-sensor (multi-factor) data to ensure accuracy and determine the root cause of scene anomalies.
[0003] A mining scene anomaly detection method based on graph regularized incremental non-negative matrix factorization is proposed. Graph regularized incremental non-negative matrix factorization can overcome the difficulties of traditional non-negative matrix factorization in online processing of large data sets. It can not only represent data content online and maintain the geometric structure of data, but also make full use of the decomposition results of the previous step in combination with incremental learning to avoid repeated calculations, thereby reducing the operation time; at the same time, it significantly reduces the dimension and has better clustering accuracy; based on the point-to-point element operation of the matrix, the decomposition results of the graph regularized incremental non-negative matrix method are more likely to represent the local characteristics of the data; in addition, the non-negative basis vectors obtained by the graph regularized incremental non-negative matrix factorization have certain linear independence and sparsity, which can better describe the large-scale high-dimensional data generated in the industrial process; finally, the graph regularized incremental non-negative matrix method does not make any special ideal assumptions on the data distribution during the decomposition process, so it can also process non-Gaussian data, and only needs to design appropriate monitoring statistics in combination with the density estimation method. Summary of the invention
[0004] In view of this, the present invention provides a mining scene anomaly detection method based on graph regular incremental non-negative matrix decomposition, which realizes dynamic data training, comprehensive analysis of multi-sensor data, saves time and space costs, and can analyze the cause of the anomaly if an anomaly occurs. It includes the following steps:
[0005] Step 1: Data collection; use two sets of equipment to sample the environment at multiple normal working times; the data obtained by the two sets of equipment are processed in the same way as follows. The operations performed on the data of one set of equipment will be described below, and they will be described separately when necessary, and no repetition will be made later.
[0006] Step 2: Data preprocessing: The measured data in the industrial process may not necessarily meet the non-negative condition. For example, the readings of sensors such as temperature and pressure may be negative. The values can be made non-negative by adjusting the units. Therefore, the data collected in step 1 are preprocessed by grayscale, vectorization, and normalization to obtain the training set X'. The initial training sample matrix (There are k samples in this matrix) is composed of data samples of a certain period in the training set X', and the data samples at that time in the training set X' are taken as the next newly added samples.
[0007] Initial training matrix Where t0≤t i ≤…≤t i+k-1 ≤t1, indicating that the data is collected during the initial period t0 to t1; represents the data collected by the i-th device in the j-th sample, which is collected at t i collected at any time; for example, i=1 means the device is a camera, i=2 means the device is a gas sensor, Represents the data of the gas sensor in the third sample, which is obtained at t i+2 Collected at all times; Indicates the first data sample, which is at t i Collected at all times; It means that the k+1th data sample is collected at time t2 and is also the next new sample. The same applies to the new samples thereafter.
[0008] Step 3: Training phase: The initial training sample matrix obtained in step 2 And the newly added samples perform graph regular incremental non-negative matrix decomposition to obtain the optimal basis matrix W new And the coefficient matrix H new This not only preserves the geometric structure information of the sample in low-dimensional space, but also makes full use of the decomposition results of the previous step in combination with incremental learning to avoid repeated calculations, thereby reducing the calculation time. The specific steps are as follows:
[0009] Step 3.1: First, the initial training sample matrix Perform SVD decomposition to obtain the singular value matrix ∑ and singular vector matrices U and V. Use the singular value matrix and singular vector matrix to Initialize the basis matrix and coefficient matrix in the graph regular non-negative matrix decomposition, and then update and iterate the initialized basis matrix and coefficient matrix until the objective function tends to be stable, and get W k ,H kIn this way, a better global optimal solution can be obtained, and no data structure change is required for the input matrix, the data structure of the original data will not be destroyed, and more detailed information can be retained, thereby improving the decomposition effect.
[0010] SVD decomposition formula:
[0011] W k ,H k initialization:
[0012] Where |U| means taking the absolute value of the matrix U, V Τ Represents the transpose of the matrix V.
[0013] Objective function:
[0014] Iteration rules:
[0015] Where R is the weight matrix and D is the diagonal matrix L k is the Laplace matrix (L k =DR), λ is the regularization parameter.
[0016] Step 3.2: When a new sample is added at the next moment, the graph regular incremental non-negative matrix factorization is used to obtain the optimal basis matrix W new And the coefficient matrix H new ;
[0017] For example, when When added, the objective function is:
[0018]
[0019] Iteration formula
[0020] Step 3.3: Repeat the above operation for all new samples, and get the optimal basis matrix W after the update. new And the coefficient matrix H new .
[0021] Step 4: Calculate the control limits; W obtained from step 3 new , H new Calculate monitoring statistics N 2 and SPE, N 2 To monitor the changes in the feature space, SPE is used to monitor the changes in the residual space; the kernel density estimation (KED) method is then used to estimate the probability density of the process data and extract the actual distribution information of the data, thereby determining the statistical control limits corresponding to the training samples of the two sets of equipment. SPE1'; SPE'2.
[0022] Step 5: Testing phase; re-collect data (as test set X″) for testing, perform the same processing as S2 and S3 on the test set X″, calculate the corresponding statistics, and compare the monitoring statistics with the two sets of control limits. The situations can be divided into the following three types:
[0023] Case 1: When both statistics are within the two sets of control limits, the scenario is normal and mining can be carried out.
[0024] Case 2: When any one or two statistics are outside the two sets of control limits, it means that the environment must be abnormal, and a level 1 alarm is immediately issued, and mining cannot be carried out; the contribution values are calculated and sorted, and the largest or larger contribution values are uploaded to the control interface as the cause of the abnormality for display.
[0025] Case 3: When any one or two statistics are outside only one set of control limits, further confirm whether it is an equipment failure. If not, immediately issue a secondary alarm, calculate and sort the contribution values, and upload the largest or larger contribution values as the cause of the abnormality to the control interface for display.
[0026] The beneficial effects of the present invention are as follows: the present invention avoids the drawbacks of traditional detection, can perform dynamic training, real-time tracking and prediction, and ensure timeliness; can comprehensively consider various factors such as toxic gases, water gushing, mine collapse, etc., and conduct comprehensive analysis of multiple sensor data to ensure accuracy; if an abnormality occurs, the root cause of the scene abnormality can be found; and can also ensure the effective use of network resources and improve efficiency.
[0027] The beneficial effects of the present invention are as follows: using the graph regular incremental non-negative matrix decomposition method for anomaly detection can overcome the difficulties of traditional non-negative matrix decomposition in online processing of large data sets, and can not only represent data content online and maintain the geometric structure of data, but also significantly reduce the dimension, reduce the operation time, and have better clustering accuracy; based on the point-to-point element operation of the matrix, the decomposition result of the graph regular incremental non-negative matrix method is more likely to characterize the local characteristics of the data; at the same time, the non-negative basis vectors obtained by the graph regular incremental non-negative matrix decomposition have certain linear independence and sparsity, which can better describe the large-scale high-dimensional data generated in the industrial process; in addition, the graph regular incremental non-negative matrix method does not make any special ideal assumptions on the data distribution during the decomposition process, so it can also process non-Gaussian data, and it only needs to design appropriate monitoring statistics in combination with the density estimation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A general flow chart of a mining scene anomaly detection method based on graph-regular incremental non-negative matrix factorization provided by an embodiment of the present invention;
[0029] Figure 2 A schematic diagram of collecting mining environment data provided by an embodiment of the present invention;
[0030] Figure 3 A schematic diagram of data composition provided by an embodiment of the present invention;
[0031] Figure 4 A specific process of graph regularized incremental non-negative matrix decomposition provided by an embodiment of the present invention; DETAILED DESCRIPTION
[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0033] The purpose of the present invention is to provide a mining scene anomaly detection method based on graph regularized incremental non-negative matrix decomposition, which can track and predict the safety of mining scenes in real time, reduce the safety risks of miners, and feedback the cause of the anomaly when the environment is abnormal, and ensure the effective use of network resources to improve efficiency.
[0034] like Figure 1 As shown, it is a general flow chart of the mining scene anomaly detection method based on graph regularized incremental non-negative matrix decomposition provided by an embodiment of the present invention. It includes the following steps:
[0035] Step 1: Data collection: Use two sets of equipment (each set includes a camera and multiple sensors (gas sensor, CO2 sensor, etc.), the quantity is one) to sample the environment under multiple normal working hours.
[0036] like Figure 2 The figure shows a schematic diagram of collecting mining environment data. The data collected by the two sets of equipment are transmitted to the node and then uploaded to the cloud server for processing. The equipment is not limited to the above three types and can be determined according to the specific mining environment.
[0037] The data obtained by the two sets of equipment are processed in the same manner as follows. The operations performed on the data collected by one set of equipment will be described below. If necessary, they will be described separately and will not be repeated hereafter.
[0038] Step 2: Data preprocessing: grayscale, vectorize, and normalize the data collected in step 1 to obtain the training set X' and the initial training sample matrix (There are k samples in this matrix) is composed of data samples of a certain period in the training set X', and the data samples at that time in the training set X' are taken as the next newly added samples.
[0039] like Figure 3The data is shown in Figure 1. The collected data is preprocessed first, and the image data collected by the camera is grayed to obtain the image matrix V∈a×b.
[0040] Vectorize the obtained image matrix V by extracting each column and recombining it into a column vector such as Figure 3 middle Then, the data collected by other sensors in the corresponding set of equipment are combined into a column vector such as Figure 3 middle Finally, the equipment was i The data collected at each moment is normalized to between 0 and 1 and combined into a column vector Indicates the first sample data, which is at t i Collected at all times, including Represents the data of the gas sensor of the second device in the first sample, also at t i Collected at all times.
[0041] Continue to collect data at multiple times, repeat the above process, and finally get the initial training sample matrix The data collected at different times later are used as additional samples in turn.
[0042] In this embodiment, 200 groups of normal data samples are collected within 5 hours as samples in the initial training sample matrix, 40 groups are collected every hour, and 10 groups are collected every hour after 5 hours as additional samples, and a total of 100 groups are added as new samples.
[0043] Step 3: Training phase: The initial training sample matrix obtained in step 2 And the newly added samples perform graph regular incremental non-negative matrix decomposition to obtain the optimal basis matrix W new And the coefficient matrix H new .
[0044] like Figure 4 As shown in FIG. 1 , the specific process of the regular incremental non-negative matrix decomposition provided by the embodiment of the present invention is shown in FIG. First, the initial training sample matrix Perform SVD decomposition to obtain the singular value matrix ∑ and singular vector matrices U and V.
[0045] SVD decomposition:
[0046] Use singular value matrix and singular vector matrix to Initialize the basis matrix and coefficient matrix in graph regularized non-negative matrix factorization.
[0047] W k ,H k initialization:
[0048] Where |U| means taking the absolute value of the matrix U, V Τ is the transpose of the matrix V.
[0049] The initialized basis matrix and coefficient matrix are updated and iterated again until the objective function becomes stable, and W is obtained. k ,H k .
[0050] Objective function:
[0051] Iteration rules:
[0052] Where R is the weight matrix and D is the diagonal matrix L k is the Laplace matrix (L k =DR), λ is the regularization parameter.
[0053] Add samples one by one Each time a sample is added, a graph regular incremental non-negative matrix factorization is performed, and it is iteratively updated until the objective function meets the convergence conditions to obtain the optimal basis matrix and coefficient matrix W new ,H new .
[0054] For example, when When added, the objective function is:
[0055]
[0056] W k+1 and H k+1 Respectively represent the sample set X k+1 The basis matrix and coefficient matrix obtained after graph regular non-negative matrix decomposition, L k+1 is the sample set X k+1 When the number of training samples is large enough, the impact of adding a new training sample on the basis matrix and the coefficient matrix is relatively small. Therefore, it is assumed that when a new sample is added, the coefficient matrix H k+1 The first k column vectors are approximately equal to H k The column vector of H k+1 =[H k ,h k+1 ], then the objective function F k+1 Can be rewritten as follows:
[0057]
[0058] Get the objective function F k+1After obtaining the incremental expression, the corresponding iterative update formula can be derived using the gradient descent method:
[0059]
[0060] Repeat the above operation for all new samples, and the optimal basis matrix W is obtained after the update is completed. new And the coefficient matrix H new .
[0061] Step 4: Calculate the control limits; get the corresponding W from the two training sets new , H new Calculate the corresponding monitoring statistic N 2 and SPE.
[0062] Statistics N used to monitor changes in feature space 2 :N 2 (i) = X Τ (i)WW Τ X(i).
[0063] The statistic SPE used to reflect the degree of data deviation:
[0064] in Represents the reconstructed value of the i-th sample vector, which is calculated as follows:
[0065] Calculate the control limits of the two statistics: Use the kernel density estimation method to estimate the probability density of the two statistics, extract the actual distribution of the statistics, and calculate the control limits of the statistics of the two sets of equipment training samples by setting the significance level α. SPE1'; SPE2'.
[0066] Step 5: Testing phase; re-collect data (as test set X″) for testing, perform the same processing as steps 2 and 3 on the test set X″, and calculate the corresponding statistic N 2 and SPE, comparing the monitoring statistic to two sets of control limits.
[0067] when When both statistics are within the control limits trained by the two equipment training sets, it means that the scene is normal and mining can be carried out.
[0068] when When any one or two statistics are outside the control limits trained by the two sets of equipment training sets, it means that the environment must be abnormal, and a level 1 alarm is immediately issued, and mining cannot be carried out; the contribution values are calculated and sorted, and the largest or larger contribution values are uploaded to the control interface as the cause of the abnormality for display.
[0069] Calculate contribution
[0070] The subscript j represents the label of the variable, abs means to find the absolute value; and δ j is the jth column of the n×n unit matrix; assuming there are four types of equipment, for the second type of equipment (variable) gas sensor, δ2=[0 1 0 0] Τ .
[0071] When any one or two statistics are only greater than the control limit trained by one of the equipment training sets, it is further confirmed whether it is an equipment failure. If not, a secondary alarm is immediately issued, the contribution value is calculated and sorted, and the largest or larger contribution values are uploaded to the control interface as the cause of the abnormality for display.
Claims
1. A mining scene anomaly detection method based on graph regularized incremental non-negative matrix factorization, characterized by: The following steps are involved: S1: Data collection: Two sets of equipment are used, each set of equipment includes a camera, a gas sensor, and a CO2 sensor, and the number of each sensor is one, to sample the environment under multiple normal working hours; the data obtained by the two sets of equipment are processed in the same way as follows; S2: Data preprocessing: The data collected in S1 is grayed, vectorized, and normalized to obtain the training set X' and the initial training sample matrix There are k samples in this matrix, which are composed of data samples of a certain period in the training set X', and the data samples at the subsequent time in the training set X' are taken as the next new samples in turn; S3: training phase; the initial training sample matrix obtained in S2 is And the newly added samples perform graph regular incremental non-negative matrix decomposition to obtain the optimal basis matrix W new And the coefficient matrix H new ; where used The singular value matrix and singular vector matrix obtained by SVD decomposition are respectively Initialize the basis matrix and coefficient matrix in graph regularized non-negative matrix factorization; This will not destroy the data structure of the original data, can maintain the geometric structure information of the sample in low-dimensional space, and can also combine incremental learning to make full use of the decomposition results of the previous step to avoid repeated calculations, thereby reducing the calculation time; S4: Calculate control limits; W obtained from S3 new , H new Calculate monitoring statistics N 2 and SPE, N 2 To monitor the changes in feature space, SPE is to monitor the changes in residual space; then the kernel density estimation (KED) method is used to estimate the probability density of process data to extract the actual distribution information of the data, so as to determine the statistical control limits corresponding to each set of equipment training samples. S5: Testing phase; re-collect the test set data X” for testing, perform the same processing as S2 and S3 on the test set X”, calculate the corresponding statistics, and compare the monitoring statistics with the two sets of control limits; if the statistics are all within the control limits of the two sets of equipment training sets, it indicates a normal state; if any one or both statistics are outside the control limits of the two sets of equipment training sets, it indicates an abnormal state, and a level one alarm is immediately issued; if any one or both statistics are only outside the control limits of one set of equipment training sets, further check whether it is an equipment abnormality, if not, it is an environmental abnormality, and a level two alarm is issued; when the scene is abnormal, calculate the contribution value and sort it, and upload the largest contribution values as the cause of the abnormality to the control interface for display.
2. The mining scene anomaly detection method based on graph regularized incremental non-negative matrix factorization according to claim 1 is characterized by: The specific parameters of the data sample in step S2 are as follows: Initial training sample matrix Where t0≤t i ≤…≤t i+k-1 ≤t1, indicating that the data is collected during the initial period t0 to t1; represents the data collected by the i-th device in the j-th sample, which is collected at t i i=1 means the device is a camera, i=2 means the device is a gas sensor, Represents the data of the gas sensor in the third sample, which is obtained at t i+2 Collected at all times; Indicates the first data sample, which is at t i Collected at all times; It means that the k+1th data sample is collected at time t2 and is also the next new sample. The same applies to the new samples thereafter.
3. The mining scene anomaly detection method based on graph regularized incremental non-negative matrix factorization according to claim 1 is characterized by: The specific steps of the training phase in step S3 are as follows: S3.1: First, the initial training sample matrix Perform SVD decomposition to obtain the singular value matrix ∑ and singular vector matrices U and V. Use the singular value matrix and singular vector matrix to Initialize the basis matrix and coefficient matrix in the graph regular non-negative matrix decomposition, and then update and iterate the initialized basis matrix and coefficient matrix until the objective function tends to be stable, and get W k ,H k ; In this way, a better global optimal solution can be obtained, and no data structure change processing is required for the input matrix, the data structure of the original data will not be destroyed, and more detailed information can be retained, thereby improving the decomposition effect; SVD decomposition formula: W k ,H k initialization: Objective function: Iteration rules: Where R is the weight matrix and D is the diagonal matrix L k is the Laplace matrix L k =DR, λ is the regularization parameter; S3.2: When the new samples are added at the next moment, the graph regularized incremental non-negative matrix factorization is used to obtain the optimal basis matrix and coefficient matrix; when When added, the objective function is: Iteration Rules S3.3: Repeat the above operation for all newly added samples, and the optimal basis matrix W is obtained after the update is completed. new And the coefficient matrix H new .
4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the mining scene anomaly detection method as described in any one of claims 1-3.
Citation Information
Patent Citations
Image decomposition method based on NMF
CN108985356A
Supervised image recognition method based on orthogonality constraint increment non-negative matrix factorization
CN110334761A