Granular network threat detection method based on subspace learning
Through the particle network threat detection method based on subspace learning, the problems of high-dimensional complex network data and fuzzy and scarce threat data in the new power system are solved, and efficient and accurate network threat detection is achieved to adapt to the complex and diverse power system network environment.
Patent Information
- Application Number
- CN202510460250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-01
AI Technical Summary
The network data in the new power system is highly dimensional and complex, traditional algorithms are difficult to process efficiently, threat data is fuzzy and scarce, supervision models rely on a large number of labeled samples to adapt, data sources are diverse, structures are heterogeneous, and the system is highly dynamic, which increases the difficulty of network threat analysis and response.
The particle network threat detection method based on subspace learning is adopted, and the representation matrix is calculated through the orthogonal matching pursuit method, and the subspace representation abnormal score is obtained, and the particle abnormal score is calculated through the fuzzy similarity relationship matrix. Finally, the weighted fusion obtains the fusion abnormal score, realizing the detection of network threat data.
Effectively process high-dimensional network data, improve the accuracy and real-time nature of network threat detection, without the need for large numbers of labeling samples, and adapt to complex and diverse power system network environments.
Smart Images

Figure CN120238358A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network data analysis, and particularly relates to a granular network threat detection method based on subspace learning. Background Art
[0002] With the continuous acceleration of the construction of a new power system with new energy as the main body in China, the rapid development of renewable energy such as wind energy and photovoltaic energy has promoted the transformation of the power system from centralized to distributed and diversified directions. In this context, the network security problem of the power system has become increasingly prominent, and network threats may have a significant impact on the stable operation of the system and national energy security. Therefore, it is of great practical significance and application value to carry out research on network threat detection in the new power system. At present, intelligent analysis-based methods have been preliminarily applied in the field of power network security, but still face many challenges: such as high-dimensional and complex network data, which are difficult to process efficiently by traditional algorithms; fuzzy and scarce threat data, and supervised models are difficult to adapt due to the dependence on a large number of labeled samples; diverse data sources, heterogeneous structures, and strong system dynamics, which increase the difficulty of threat analysis and response. Therefore, it is urgent to introduce more efficient and robust intelligent algorithms to improve the accuracy and real-time performance of network threat detection in the new power system. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a granular network threat detection method based on subspace learning.
[0004] The technical solution adopted by the present invention is as follows: a granular network threat detection method based on subspace learning, comprising the following steps:
[0005] S1. Obtain network data, perform normalization processing on the network data to obtain the data after normalization processing;
[0006] S2. According to the normalized data, use the orthogonal matching pursuit method to calculate the representation matrix;
[0007] S3. Obtain the reconstruction points of each sample point according to the representation matrix, and obtain the subspace representation anomaly score through the reconstruction error between the reconstruction point and the original sample point;
[0008] S4. According to the normalized data, calculate the fuzzy similarity relation matrix;
[0009] S5. Calculate the sample similarity according to the fuzzy similarity relation matrix, and further calculate the reachable similarity between samples;
[0010] S6. Calculate the fuzzy neighborhood density according to the reachable similarity, and then calculate the granular anomaly score;
[0011] S7. Weightedly add the subspace representation anomaly score and the granular anomaly score to obtain the fusion anomaly score;
[0012] S8. Compare the anomaly scores of each sample fusion with the threshold. Samples with scores greater than the threshold are determined as network threat data.
[0013] The beneficial effects of the present invention are as follows: The present invention makes full use of the advantages of subspace learning in processing high-dimensional data and the advantages of granular computing in processing uncertain information, so that even if the network data of the new power system has high-dimensionality and ambiguity, the detection method of the present invention can still effectively detect network threat data, providing strong support for the network security protection of the new power system.
[0014] Further, the expression for the normalization process in step S1 is:
[0015]
[0016] where f m (·) is the normalization process of numerical attributes, is the value of the sample f m (·) on the numerical attribute m; M m is the set composed of the values of all samples in the power consumption data on the numerical attribute m; x is the sample in the network data; m is the numerical attribute; min is the minimum value function; max is the maximum value function; for all the data mentioned in the specification, its attributes are all numerical attributes, that is, there are no discrete attributes in the data.
[0017] The beneficial effect of the above further solution is to reduce the data calculation amount by normalizing the numerical attributes.
[0018] Further, the calculation method of the matrix representation in step S2 is as follows:
[0019] S201. According to the normalized data matrix X n×d , where n is the number of samples, d is the number of attributes, and the iteration threshold ∈, the maximum number of iterations k max , for each sample point x i (i = 1, 2,..., n), perform the processing of S202 to S204;
[0020] S202. For the sample point x i , let the residual q0 = x i , initialize the support set the iteration number k = 0, and the matrix to be operated X -i ;
[0021] where X -i represents the matrix obtained by removing the i-th row from X (i = 1, 2,..., n), and the residual represents the distance to fully represent x iThe vector of the difference, the support set T is a set of natural numbers from 1 to n - 1, reflecting the non-zero column indices of the i-th row of the representation matrix, and the iteration number k reflects the current iteration number;
[0022] S203. When k < k max and ||q k ||2 > ∈, perform the following operations:
[0023] 1) T k+1 = T k ∪{j *}, where
[0024] 2) where is the projection matrix onto the span of the vector group {x j T , j ∈ T k+1};
[0025] 3) k = k + 1;
[0026] S204. Obtain
[0027] Then insert a 0 into the i-th column of , then represents the self-representation vector of the sample point x i ;
[0028] S205. Obtain the representation matrix
[0029]
[0030] Thus, the calculation of the representation matrix is completed.
[0031] The beneficial effect of the above further solution is: calculating the representation matrix provides a calculation basis for the subspace representation anomaly score.
[0032] Furthermore, the calculation method of the reconstruction point in step S3 is as follows:
[0033] S301. For each sample point x i , perform the processing of S302 to S303;
[0034] S302. Let where represents the reconstruction error of the sample point x i ;
[0035] S303. Define Then where L maxDenote the maximum reconstruction error, Denote the sample point x i The subspace representation of the outlier score.
[0036] The beneficial effect of the above further scheme is: calculating the outlier score of the subspace representation to prepare for the subsequent fusion
[0037] Further, the calculation method of the fuzzy similarity relation matrix in step S4 is as follows:
[0038] S401. Let
[0039] where o and p are two sample points randomly selected from X; is the attribute set, representing the set composed of all attributes, and d(o, p) represents the Euclidean distance between o and p in the sense of the attribute set;
[0040] S402. Let where represents the fuzzy similarity relation of the sample points; represents the cardinality of the attribute set, that is, the number of elements in;
[0041] S404. Let where, is the required fuzzy similarity relation matrix.
[0042] The beneficial effect of the above further scheme is: using the fuzzy similarity relation matrix to describe the similarity relationship between sample points, which is conducive to the definition of similarity in the following text.
[0043] Further, the calculation methods of sample similarity and reachable similarity in step S5 are as follows:
[0044] S501. Randomly select a sample point o in X, and define its fuzzy k-nearest neighbor in the following way with a given number of neighbors k
[0045]
[0046] where p is also a sample point, represents the k-nearest neighbor similarity value of sample o, and its value is determined by the following conditions:
[0047] 1) At least k sample points p ∈ X - {o} satisfy the condition
[0048] 2) At most k - 1 sample points p ∈ X - {o} satisfy the condition
[0049] S502. Randomly select two sample points o and p in X, and define the reachable similarity from p to o:
[0050]
[0051] In the above definition, if the distance between p and o is very far, then their similarity will be very small, and their fuzzy similarity relationship will be used as the reachable similarity from p to o; conversely, if their similarity is high, then the k-nearest neighbor similarity value of o will be used as the reachable similarity from p to o.
[0052] The beneficial effect of the above further solution is: Defining the reachable similarity according to the distance reduces the adverse impact of statistical fluctuations on the results and improves the stability of the calculation of the granular anomaly score.
[0053] Furthermore, the calculation method of the sample granular anomaly score in step S6 is as follows:
[0054] S601. On the premise that the attribute set is and the number of nearest neighbors is k, the fuzzy neighborhood density is defined as:
[0055]
[0056] The value of the fuzzy neighborhood density is the ratio of the sum of the reachable similarities of the fuzzy k-nearest neighbors of the sample point o to the cardinality of its fuzzy k-nearest neighbors, where Obviously, the smaller the fuzzy neighborhood density, the more likely the sample point o is an outlier;
[0057] S602. The calculation of the granular anomaly score is as follows:
[0058]
[0059] Among them, the definition of the f function is as follows:
[0060]
[0061] The beneficial effect of the above further solution is: Calculating and normalizing the granular anomaly score using the neighborhood deviation degree effectively balances the fusion weight of the granular anomaly score and the subspace representation anomaly score.
[0062] Furthermore, the calculation method of the sample fusion anomaly score in step S7 is as follows:
[0063] Given the subspace representation weight α (0 < α < 1), the fusion anomaly score of the sample point o is defined as follows:
[0064] MS o = α·SSCS o +(1 - α)·FNOS o
[0065] The beneficial effects of the above further solution are as follows: By calculating the fusion anomaly score obtained from the particle anomaly score and the subspace representation anomaly score, and comparing it with the threshold, the threat detection problem of high-dimensional fuzzy network data is solved, and the accuracy of threat data screening is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0067] Embodiment
[0068] As Figure 1 shown, in an embodiment of the present invention, a particle network threat detection method based on subspace learning includes the following steps:
[0069] S1. Obtain network data, perform normalization processing on the network data to obtain the normalized data;
[0070] S2. According to the normalized data, use the orthogonal matching pursuit method to calculate the representation matrix;
[0071] S3. Obtain the reconstruction points of each sample point according to the representation matrix, and obtain the subspace representation anomaly score through the reconstruction error between the reconstruction point and the original sample point;
[0072] S4. According to the normalized data, calculate the fuzzy similarity relation matrix;
[0073] S5. Calculate the sample similarity according to the fuzzy similarity relation matrix, and further calculate the reachable similarity between samples;
[0074] S6. Calculate the fuzzy neighborhood density according to the reachable similarity, and then calculate the particle anomaly score;
[0075] S7. Add the subspace representation anomaly score and the particle anomaly score by weighting to obtain the fusion anomaly score;
[0076] S8. Compare the size of the fusion anomaly score of each sample with the threshold, and the sample with a score greater than the threshold is determined as network threat data.
[0077] In this embodiment, the subspace learning theory and the fuzzy neighborhood theory provide an effective tool to overcome the problem that it is difficult to detect the network threat data of the new power system due to its high-dimensional ambiguity, and can be directly applied to the analysis model of numerical attributes. In the specific application process, the network data of the new power system is imported into an information system (or called information table), where each row represents an object (or called sample), and each column represents an attribute (feature) of the object. The values of the attributes are numerical. An information system is represented as (U, A), where U represents the set of all objects, and A represents the set of all attributes.
[0078] In this embodiment, first, a normalization function is used to normalize the data, effectively reducing the amount of calculation; the self-representation matrix is used to characterize the reconstructibility of the samples, and the subspace representation anomaly score is obtained; the neighborhood deviation degree of the samples is calculated through the fuzzy neighborhood theory and normalized to obtain the granular anomaly score; this method does not require labeled data for model training and can effectively realize the unsupervised threat detection of the network data of the new power system.
[0079] The expression of the normalization process in step S1 is as follows:
[0080]
[0081] where f m (·) is the normalization process of the numerical attribute, is the value of the sample f m (·) on the numerical attribute m; M m is the set composed of the values of all samples in the electricity consumption data on the numerical attribute m; x is the sample in the network data; m is the numerical attribute; min is the minimum value function; max is the maximum value function; for all the data targeted in the specification, its attributes are all numerical attributes, that is, there are no discrete attributes in the data.
[0082] In this embodiment, through the min-max normalization operation, the value range of all numerical attributes is adjusted to the real number interval from 0 to 1. Suppose the data matrix obtained after normalization is
[0083]
[0084] The calculation method of the representation matrix in step S2 is as follows:
[0085] S201. According to the data matrix X n×d , where n is the number of samples, d is the number of attributes, and the iteration threshold ∈, the maximum number of iterations k max , for each sample point x i (i = 1, 2,..., n), the processes of S202 to S204 are performed;
[0086] S202. For the sample point x i , let the residual q0 = x i , and initialize the support set The iteration number k = 0, and the matrix X to be operated on -i ;
[0087] where X -i represents the matrix obtained by removing the i-th row from X (i = 1, 2,..., n), the residual represents the vector of the difference from completely representing x i , the support set T is a set of natural numbers from 1 to n - 1, reflecting the non-zero column indices of the i-th row of the representation matrix, and the iteration number k reflects the current iteration number;
[0088] S203. When k < k max and ||q k ||2 > ∈, perform the following operations:
[0089] 1) T k+1 = T k ∈{j *}}, where
[0090] 2) where is the projection matrix of onto the span of the vector group;
[0091] 3) k = k + 1;
[0092] S204. Obtain
[0093] Then insert a 0 into the i-th column of , then represents the self-representation vector of the sample point x i ;
[0094] S205. Obtain the representation matrix
[0095]
[0096] Thus, the calculation of the representation matrix is completed.
[0097] In this embodiment, the representation matrix C is the coefficient matrix for the self-representation reconstruction of the data matrix X.
[0098] The specific steps of step S3 are as follows:
[0099] S301. For each sample point x i , perform the processing of S302 to S303;
[0100] S302. Let where represents the reconstruction error of the sample point x i ;
[0101] S303. Define Then where L max represents the maximum reconstruction error, represents the subspace representation anomaly score of the sample point x i .
[0102] In this embodiment, the subspace representation anomaly score is calculated through the self-representation reconstruction error, which specifically includes the following steps:
[0103] Step 1: Calculate the reconstructed sample
[0104] Calculate the reconstructed data matrix X using the representation matrix C and the data matrix X *
[0105] X * = CX
[0106] X * represents the reconstructed data matrix. Calculate the subspace representation anomaly score using the difference between the reconstructed data matrix and the original sample point. For any sample point x i , denote its reconstructed point as
[0107] Step 2: Calculate the reconstruction error
[0108]
[0109]
[0110] represents the reconstruction error of the sample point x i . In this process, the anomaly degree of the sample point is reflected by calculating the distance error between the reconstructed point and the original sample point. Sample points with larger reconstruction errors often have a higher anomaly degree. L max is the maximum reconstruction error value.
[0111] Step 3: Calculate the subspace representation anomaly score
[0112] For each sample point o in X, calculate its subspace representation score using the reconstruction error
[0113]
[0114] This step adjusts the degree of anomaly to the real number interval from 0 to 1, improving comparability and preventing the imbalance of fusion weights in the fusion process caused by the dominance of large numbers. A high anomaly score represents a high degree of anomaly, and points with high subspace representation of anomaly scores tend to have a higher degree of anomaly.
[0115] The calculation method of the fuzzy similarity relation matrix in step S4 is as follows:
[0116] S401. Let
[0117] where o and p are any two sample points in X; is the attribute set, representing the set composed of all attributes, and d(o, p) represents the Euclidean distance between o and p in the sense of the attribute set;
[0118] S402. Let where, represents the fuzzy similarity relation of the sample point; represents the cardinality of the attribute set, that is, the number of elements in;
[0119] S404. Let where, is the required fuzzy similarity relation matrix.
[0120] In this embodiment, represents the distance metric between sample point o and sample point p. Obviously, the farther the distance between two points, the smaller their similarity, represents the fuzzy similarity between two sample points. The closer the two points, the smaller the fuzzy similarity. The fuzzy similarity matrix is composed of fuzzy similarity relations, providing a basis for characterizing the similarity between sample points below.
[0121] Step S5 is specifically as follows:
[0122] S501. Arbitrarily take a sample point o in X, and define its fuzzy k-nearest neighbor in the following way given the number of neighbors k
[0123]
[0124] where p is also a sample point, represents the k-nearest neighbor similarity value of sample o, and its value is determined by the following conditions:
[0125] 1) At least k sample points p ∈ X - {o} satisfy the condition
[0126] 2) At most k - 1 sample points p ∈ X - {o} satisfy the condition
[0127] S502. Randomly select two sample points o and p in X, and define the reachable similarity from p to o.
[0128]
[0129] In the above definition, if the distance between p and o is very far, then their similarity will be very small, and their fuzzy similarity relationship will be used as the reachable similarity from p to o; conversely, if their similarity is relatively high, then the k-nearest neighbor similarity value of o will be used as the reachable similarity from p to o; adopting this definition method can reduce the adverse impact of statistical fluctuations on the results.
[0130] In this embodiment, the reachable similarities between all sample points are calculated, and the reachable similarities can be used to calculate the fuzzy neighborhood density and the fuzzy neighborhood degree.
[0131] The specific steps for calculating the granular anomaly score in step S6 are as follows:
[0132] S601. On the premise that the attribute set is and the number of nearest neighbors is k, the fuzzy neighborhood density is defined as:
[0133]
[0134] The value of the fuzzy neighborhood density is the ratio of the sum of the reachable similarities of the fuzzy k-nearest neighbors of the sample point o to the cardinality of its fuzzy k-nearest neighbors, where Obviously, the smaller the fuzzy neighborhood density, the more likely the sample point o is an outlier;
[0135] S602. The calculation of the granular anomaly score is as follows:
[0136]
[0137] Among them, the definition of the f function is as follows:
[0138]
[0139] This step adjusts the degree of anomaly to the real number interval from 0 to 1, improves the comparability, and prevents the imbalance of the fusion weights caused by the dominance of large numbers in the fusion process.
[0140] The specific steps of step S7 are as follows:
[0141] Given the subspace representation weight α (0 < α < 1), the fusion anomaly score of the sample point o is defined as follows:
[0142] MS o = α · SSCS o + (1 - α) · FNOSo
[0143] In this embodiment, the fusion anomaly score of the sample points is used to measure the likelihood of a sample belonging to a threat point, and this score is obtained by fusing the subspace representation anomaly score and the granular anomaly score.
[0144] In this embodiment, let the outlier threshold of the anomaly point be ξ ∈ (0, 1). If the sample x i has a fusion anomaly score then x i is determined to be an anomaly point. By comparing the outlier degrees of all samples with the threshold ξ one by one, the threat data in the information system can be calculated.
[0145] Finally, it should be noted that the above are only preferred examples of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A granular network threat detection method based on subspace learning, characterized in that: The following steps are involved: S1. Acquire network data, perform normalization processing on the network data, and obtain normalized data; S2. Calculate the representation matrix based on the normalized data using the orthogonal matching pursuit method; S3, obtain the reconstruction point of each sample point according to the representation matrix, and obtain the subspace representation anomaly score through the reconstruction error between the reconstructed point and the original sample point; S4. Calculate the fuzzy similarity relationship matrix based on the normalized data; S5, calculating sample similarity according to the fuzzy similarity relationship matrix, and further calculating the achievable similarity between samples; S6, calculating the fuzzy neighborhood density according to the reachable similarity, and then calculating the particle anomaly score; S7, weighted addition of the subspace representation anomaly score and the granular anomaly score to obtain a fused anomaly score; S8. Compare the fusion anomaly scores of each sample with the threshold, and determine the samples with scores greater than the threshold as network threat data.
2. According to the granular network threat detection method based on subspace learning in claim 1, it is characterized in that: The expression for the normalization process in step S1 is: Among them, f m (·) is the normalization processing of data, is a sample f in the data m (·) The value of the numeric attribute m; M m is a set of values of all samples in the electricity consumption data on the numerical attribute m; x is a sample in the network data; m is a numerical attribute; min is the minimum function; max is the maximum function; for all data targeted in the claims, their attributes are all numerical attributes, that is, there are no discrete attributes in the data.
3. The granular network threat detection technology based on subspace learning according to claim 1 is characterized in that: The calculation method of the representation matrix in step S2 is as follows: S201, according to the normalized data matrix X n×d , where n is the number of samples, d is the number of attributes, and the iteration threshold ∈, the maximum number of iterations k max , for each sample point x i (i=1, 2, ..., n) perform the processing from S202 to S204; S202, for sample point x i , let the residual q0 = x i , initialize the support set Iteration number k = 0, matrix to be operated X -i ; Among them, X -i represents the matrix after removing the i-th row of X (i=1,2,…,n), and the residual represents the distance completely representing x i The support set T is a set of natural numbers from 1 to n-1, reflecting the non-zero column index of the i-th row of the matrix, and the number of iterations k reflects the current number of iterations; S203, when k <k max And ||q k ||2>∈, perform the following operations: 1)T k+1 =T k ∪{j * },in, 2) in, yes The vector group {x j T ,j∈T k+1 }The projection matrix of the spanned space; 3) k = k + 1; S204, get Then in Insert a 0 into the i-th column of Represents the sample point x i The self-representing vector of ; S205, get the representation matrix This completes the calculation of the representation matrix.
4. According to the granular network threat detection method based on subspace learning in claim 1, it is characterized in that: The calculation method of the reconstruction points in step S3 is as follows: S301, for each sample point x i , execute the processing from S302 to S303; S302, Order in Represents the sample point x i The reconstruction error of S303, Definition but Where L max represents the maximum reconstruction error, Represents the sample point x i The subspace of represents the anomaly score.
5. According to the granular network threat detection method based on subspace learning in claim 1, it is characterized in that: The calculation method of the fuzzy similarity relationship matrix in step S4 is as follows: S401, HDM β (o, p) = d(o, p), Among them, o and p are two sample points randomly selected from X; is an attribute set, which means the set consisting of all attributes. d(o, p) means o and p in Euclidean distance in the sense of attribute set; S402, Order in, Represents the fuzzy similarity relationship of sample points; Represents the cardinality of the attribute set, i.e. The number of elements in ; S404, Order in, This is the desired fuzzy similarity relationship matrix.
6. The granular network threat detection method based on subspace learning according to claim 1 is characterized in that: The calculation method of the sample similarity and the reachable similarity in step S5 is as follows: S501, randomly select a sample point o in X, and define its fuzzy k nearest neighbors in the following way given the number of neighbors k. : Among them, p is also a sample point, It represents the k-nearest neighbor similarity value of sample o, and its value is determined by the following conditions: 1) There are at least k sample points p∈X-{o} that satisfy the condition 2) There are at most k-1 sample points p∈X-{o} that satisfy the condition S502: randomly select two sample points o and p in X, and define the reachable similarity from p to o. In the above definition, if the distance between p and o is very far, then their similarity will be small, and their fuzzy similarity relationship will be used as the reachable similarity from p to o; conversely, if their similarity is high, then the k-nearest neighbor similarity value of o will be used as the reachable similarity from p to o; the reason for adopting this definition is that it can reduce The adverse effect of statistical fluctuations on the results.
7. The granular network threat detection method based on subspace learning according to claim 1 is characterized in that: The calculation method of the sample particle abnormality score in step S6 is as follows: S601, in the attribute set Under the premise that the number of neighbors is k, the fuzzy neighborhood density is defined as: The value of the fuzzy neighborhood density is the ratio of the sum of the reachable similarities of the fuzzy k-nearest neighbors of the sample point o to the cardinality of its fuzzy k-nearest neighbors, where: Obviously, the smaller the fuzzy neighborhood density is, the more likely the sample point o is an outlier; S602: The calculation of the particle abnormality score is as follows: The definition of the f function is as follows: Represents the normalization of the results. This is because the value before normalization may be greater than 1. Normalization can significantly reduce the weight imbalance caused by excessive particle anomaly scores.
8. The granular network threat detection method based on subspace learning according to claim 1 is characterized in that: The calculation method of the sample fusion anomaly score in step S7 is as follows: Given a subspace representation weight α (0<α<1), the fusion anomaly score of sample point o is defined as follows: MS o =a·SSCS o +(1-α)·FNOS o 。