Internet of Things equipment intrusion detection method and system based on anchor graph learning
Building high-quality anchor points through anchor diagram learning method solves the problem of the computing power limitation of IoT devices, and realizes efficient and transparent intrusion detection, which is suitable for complex IoT systems.
Patent Information
- Application Number
- CN202510429324.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
AI Technical Summary
The computing power and storage limitations of IoT devices make it difficult to deploy standard encryption technologies and common intrusion detection systems, and existing methods lack transparency and network resilience, making it difficult to effectively detect cyber threats.
Using an anchor graph learning method, by constructing the objective function and optimization method of anchor graph learning, high-quality anchor points are dynamically generated to achieve data reduction and calculation complexity reduction, and abnormal detection is performed in combination with unsupervised clustering method.
It improves the transparency and detection performance of IoT device intrusion detection, reduces computing costs, can effectively detect attacks in the network, and is suitable for complex IoT systems.
Smart Images

Figure CN120301640A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to an Internet of Things device intrusion detection method and system based on anchor graph learning. Background Art
[0002] Intelligent Internet of Things transforms traditional linear manufacturing into dynamically interconnected intelligent manufacturing, bringing a brand-new operation mode to enterprises and the like. Although the Internet of Things brings many advantages to our modern life, security vulnerabilities and cyber threats hinder the significant progress of the Internet of Things. If these problems are not solved, it may have a negative impact on the deployment and operation of Internet of Things systems.
[0003] The limitations of the computing power and storage of Internet of Things devices make it difficult to deploy standard encryption technologies and common intrusion detection systems. Although some research focuses on using artificial intelligence technologies to detect abnormal events on the Internet of Things, some proposed methods lack transparency and network resilience, which concerns many cybersecurity experts. Therefore, customized intrusion detection as a second line of defense has become an urgent task for Internet of Things systems.
[0004] Recent work on intrusion detection systems has mainly focused on the use of machine learning models because of their ability to automatically learn from heterogeneous data. However, the black-box mode of deep learning is one of the reasons for the cautious adoption of deep learning in many critical Internet of Things systems. Therefore, understanding the reasoning behind intrusion detection decisions helps to cultivate trust and enables experts to verify whether the system can solve problems safely and reliably. In addition, due to the complexity and heterogeneity of Internet of Things networks, intrusion detection systems with interpretable functions can provide transparency for the way of predicting cyberattacks.
[0005] In recent years, graph-based methods have been deeply studied due to their ability to cluster data of arbitrary shapes. Their powerful data analysis capabilities enable graph learning to learn normal user behavior patterns from massive training data and determine anomalies by the deviation of subsequent test samples, becoming a new idea for intrusion detection.
[0006] Therefore, one aspect of this patent focuses on improving the transparency of the intrusion detection system for the Internet of Things. Another aspect is to detect hidden attacks in the network as accurately as possible through the established intrusion detection system, enhance the detection performance of the intrusion detection system, and reduce the overall computing cost. To solve the above problems existing in the Internet of Things network security, this patent proposes an Internet of Things device intrusion detection method and system based on one-step anchor graph learning. This method uses high-quality anchor graphs to obtain the final interpretable classification results, thereby achieving competitive anomaly detection performance. Specifically, considering data redundancy and corruption as well as the computing power of Internet of Things devices, this patent maps the original data to the embedding space to learn more representative and compatible anchor points, and then reduces the time overhead and computational complexity through the established anchor points. At the same time, in order to explore more discriminative information, this patent proposes dynamically guided anchor graph learning to learn high-quality anchor graphs. Finally, this patent unifies anchor graph learning, anchor graph construction, and anchor graph decomposition into a framework and designs an effective algorithm to solve the resulting optimization problem. In addition, this patent enforces the K-connectivity constraint on high-quality anchor graphs to directly generate detection results without additional subsequent processing methods. Summary of the Invention
[0007] Aiming at the shortcomings of the existing technology and comprehensively solving the problem of network traffic anomaly detection, the present invention proposes an Internet of Things device intrusion detection method and system based on anchor graph learning.
[0008] To achieve the above object, the present invention adopts the following technical solutions, including the following steps:
[0009] Step 1: At the control center, collect network traffic, represent it as a feature matrix, and use it as training data.
[0010] Step 2: At the control center, build an Internet of Things device intrusion detection method based on anchor graph learning.
[0011] Step 3: At the control center, use the collected network traffic to train the detection model established in Step 2.
[0012] Step 4: End.
[0013] The setting of the feature matrix described in Step 1 specifically includes the following steps:
[0014] Step A: At the control center, measure the network traffic and represent it as a traffic feature matrix X.
[0015] End-to-end network traffic, called origin-destination flow (OD), describes the traffic from an origin node to a destination node. For all possible OD pairs in a network with N nodes, we can describe it using a traffic matrix, and at the same time use the network security dataset traffic feature extraction tool (Cicflowmeter) to extract the traffic matrix to form a traffic feature matrix. The traffic feature matrix can be represented as X, and X is a d×n matrix, where d represents the dimension of the traffic data (indicating the number of extracted features), and n represents the total number of samples.
[0016] Step B: At the control center, perform standard normalization on the original traffic feature matrix data X and use it as the subsequent training data.
[0017] The setup of the intrusion detection method based on anchor graph learning described in Step 2 specifically includes the following steps:
[0018] Step A: At the control center, load the dataset X collected in Step 1 and read the traffic feature matrix.
[0019] Step B: At the control center, establish the objective function of anchor graph learning.
[0020] Given n samples in the dataset with dimension d, the construction of anchor graph learning is as follows:
[0021]
[0022] where represents the sample projection matrix, m (m << d) is the dimension of the embedding space. It should be noted here that the orthogonal constraint (i.e., W T W = I m , I m is the identity matrix with dimension m) can make the space bases uncorrelated, thereby improving the discrimination ability and avoiding trivial solutions at the same time. is the anchor point matrix in the embedding space, and a is the number of anchor points. A T A = I a restricts the anchor points to make them more discriminative, where I a represents the identity matrix with dimension a. Z is the anchor graph, and Z is used to represent the relationship between the anchor point matrix and the original data, ||·|| F represents the Frobenius norm, and k represents the number of labeled data types. Therefore, using formula (1), the anchor point matrix A can be dynamically obtained through learning instead of traditional static sampling. represents the rank constraint, where is the normalized Laplacian matrix of O, and the degree matrix D O is a diagonal matrix associated with O, and its diagonal elements are calculated by and O is the enhanced graph of B, defined as:
[0023]
[0024] Step C: Establish an optimization method for anchor graph learning in the control center.
[0025] In the control center, we solve Equation (1) according to the alternating optimization iterative update mechanism. Specifically, we can divide the detailed process into several steps by updating each variable while keeping the other variables fixed.
[0026] When A and Z are fixed, Equation (1) can be rewritten with respect to W as:
[0027]
[0028] where P W = XZ T A T , assuming that the matrix in Formula (3) has an m-order singular value decomposition form (SVD) as where To maximize the value of Equation (3), correspondingly, we get
[0029] When W and Z are fixed, Equation (1) can be rewritten with respect to A as:
[0030]
[0031] where P A = W T XZ T , similar to the solution of Formula (3), assuming that P A has a singular value decomposition form (SVD) as Correspondingly, we get where U A represents the left singular vector matrix of P A , and where V A represents the right singular vector matrix of P A .
[0032] When W and A are fixed, Equation (1) can be rewritten with respect to Z as:
[0033]
[0034] By solving Equation (5) through Lagrange, we get:
[0035]
[0036] where P :,i represents the matrix P Z = A T (W TThe i-th column element of X), Z :,i represents the i-th column element of matrix Z, γ represents the regularization term, T i,: represents the row vector composed of t ij t ij is expressed as follows:
[0037]
[0038] Among them, The optimal solution of F should be composed of the k eigenvectors with the smallest eigenvalues of, k represents the number of labeled data types. Considering that the entire similarity matrix O is composed of high-quality anchor graph Z, the first n rows of F actually represent the indicators of data points, and the remaining a rows are the indicators of anchor points. F (n) (i, :) and F (a) (j, :) respectively represent the i-th row and the j-th row of F (n) and F (a) , D (n) (i, i) and D (a) (j, j) respectively represent the i-th diagonal element and the j-th diagonal element of D (n) and D (a) . Finally, stack by column and take the mean of all results to generate the optimal Z.
[0039] Step D: In the control center, establish an alternating optimization update mechanism.
[0040] Solve equation (1) according to the iterative update mechanism. Specifically, solve by updating each variable while keeping other variables fixed, that is, solve the optimal solution of the optimal equation (1) through Step C.
[0041] The detection model described in Step 3 specifically includes the following steps:
[0042] Step A: In the control center, repeat Step D in Step 2, and the stop criterion is defined as (obj i-1 -obj i ) / obj i ≤10 -4 and obj i ≤10 -10 , where obj is the objective value of the iterator, and obj i-1 and obj i are the values at the previous moment and the current moment respectively.
[0043] Step B: In the control center, perform anomaly detection.
[0044] The method designed in this patent is an unsupervised clustering method. Its main anomaly detection method is to regard the points far from the clustering center as anomalies. For the optimal anchor point A obtained after iterative update, calculate the difference between the real-time network data points and the anchor point, regard those with too large difference from the anchor point as anomalies, and set those within the threshold range as data of the same type as the anchor point.
[0045] Due to the data reduction achieved by the embedding learning and anchor graph learning methods, the anomaly detection method proposed in this patent can effectively handle the big data environment. Given the n data with d dimensions collected in the subsequent network. For the subsequent data point x t on the embedding feature c j ∈M (j = 1, 2,..., m), t ∈W T X t and the adaptive fuzzy relationship with the anchor point x a ∈A is expressed as follows:
[0046]
[0047] where represents the adaptive neighborhood in the feature c j , represents the standard deviation on c j , and the adaptive neighborhood can be adjusted through multi-party verification according to different scenarios. represents the value of x j in the feature c t . By comparing the adaptive fuzzy relationships between each data point and a anchor points, anomaly detection discrimination can be carried out:
[0048]
[0049] where x i ∈A (i = 1,..., a). Among them, P(x t ) ∈ [0, m] represents the final judgment result. When the maximum value P < m / 2, it is considered that the detection result has no corresponding anchor point of the same type and is regarded as a newly emerging anomaly.
[0050] Advantages of the present invention:
[0051] Although there are already various methods for researching the optimization of traffic intrusion detection, the method proposed in this invention takes into account the low computing power of Internet of Things (IoT) devices and the complexity of the network, and the proposed method is more comprehensive and practical. This invention proposes an IoT device intrusion detection method and system based on anchor graph learning. This method considers the characteristics of traffic and uses an architecture based on anchor graph learning to capture network traffic characteristics and perform intrusion detection. Network devices interact with each other and are in large numbers, making the scale of network traffic data larger and difficult to process. The method of this invention can detect new types of attacks by extracting anchor point data and through prior knowledge. For complex systems such as IoT networks, the IoT device intrusion detection method and system based on anchor graph learning proposed in this invention has more efficient and accurate performance compared with many other current optimization methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flow chart of the present invention;
[0053] Figure 2 is a detection diagram for drones according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0054] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0055] As Figure 1 shown, the first step is to collect the original IoT device network traffic data and real-time network traffic data; the second step is to model an intrusion detection model based on anchor graph learning; the third step is to train the established intrusion detection model; the fourth step is to perform real-time detection on the real-time data based on the established detection model.
[0056] Applying the IoT device intrusion detection method and system based on anchor graph learning to the drone networking (such as Figure 2 ) specifically includes the following steps:
[0057] Step 1: At the control center, collect network traffic, represent it as a feature matrix, and use it as training data.
[0058] Step 2: At the control center, build an IoT device intrusion detection method based on anchor graph learning.
[0059] Step 3: At the control center, use the collected network traffic to train the detection model established in Step 2.
[0060] Step 4: End.
[0061] The setting of the feature matrix described in Step 1 specifically includes the following steps:
[0062] Step A: At the control center, measure the network traffic and represent it as a traffic feature matrix X.
[0063] The end-to-end network traffic, called the origin-destination flow (OD), describes the traffic from the origin node to the destination node. For all possible OD pairs in the unmanned network with N unmanned aerial vehicle nodes, we can describe it with a traffic matrix. Meanwhile, use the traffic feature extraction tool of the network security dataset to extract the traffic matrix and form a traffic feature matrix. The traffic feature matrix can be represented as X, and X is a d×n matrix, where d represents the dimension of the traffic data (indicating the number of extracted features, such as 76 features like traffic transmission speed, packet size, etc., that is, d = 76), and n represents the total number of samples.
[0064] Step B: At the control center, perform standard normalization on the original traffic feature matrix data X and use it as the subsequent training data.
[0065] The setting of the intrusion detection method based on anchor graph learning described in Step 2 specifically includes the following steps:
[0066] Step A: At the control center, load the dataset X collected in Step 1 and read the traffic feature matrix.
[0067] Step B: At the control center, establish the objective function of anchor graph learning.
[0068] Given n samples in the dataset with dimension d, the construction of anchor graph learning is as follows:
[0069]
[0070] where represents the sample projection matrix, m (m << d) is the dimension of the embedding space. It should be noted here that the orthogonal constraint (i.e., W T W = I m , I m is the identity matrix, with dimension m) can make the space bases uncorrelated, thus improving the discrimination ability and also avoiding the trivial solution. is the anchor point matrix in the embedding space, and a is the number of anchor points. A T A = I a restricts the anchor points to make them more discriminative, where I a represents the identity matrix with dimension a. Z is the anchor graph, and Z is used to represent the relationship between the anchor point matrix and the original data. ||·|| F represents the Frobenius norm, and k = 12 represents the number of labeled data types. Therefore, using formula (10), the anchor point matrix A can be obtained dynamically through learning instead of the traditional static sampling. represents the rank constraint, where is the normalized Laplacian matrix of O, and the degree matrix D O is a diagonal matrix associated with O, and its diagonal elements are calculated by where O is the enhanced graph of B, defined as:
[0071]
[0072] Step C: At the control center, establish an optimization method for anchor graph learning.
[0073] At the control center, we solve Equation (10) according to the alternating optimization iteration update mechanism. Specifically, we can divide the detailed process into several steps by updating each variable while keeping other variables fixed.
[0074] When A and Z are fixed, Equation (10) can be rewritten with respect to W as:
[0075]
[0076] where P W = XZ T A T , assuming that the matrix in Equation (12) has the m-order singular value decomposition form (SVD) as where To maximize the value of Equation (12), we can correspondingly obtain
[0077] When W and Z are fixed, Equation (10) can be rewritten with respect to A as:
[0078]
[0079] where P A = W T XZ T , similar to the solution of Equation (12), assuming that P A has the singular value decomposition form (SVD) as We can correspondingly obtain where U A represents the left singular vector matrix of P A , and V A represents the right singular vector matrix of P A .
[0080] When W and A are fixed, Equation (10) can be rewritten with respect to Z as:
[0081]
[0082] By solving Equation (14) through Lagrangian, we can obtain:
[0083]
[0084] Among them, P :,i represents matrix P Z = A T (W T X)'s i-th column element, Z :,i represents the i-th column element of matrix Z, γ represents the regularization term, T i,: represents the row vector composed of t ij . t ij is expressed as follows:
[0085]
[0086] Among them, The optimal solution of F should be composed of the k eigenvectors with the smallest eigenvalues of , where k represents the number of labeled data types. Considering that the entire similarity matrix O is composed of high-quality anchor graph Z, the first n rows of F actually represent the indicators of data points, and the remaining a rows are the indicators of anchor points. F (n) (i, :) and F (a) (j, :) respectively represent the i-th row and the j-th row of F (n) and F (a) , D (n) (i, i) and D (a) (j, j) respectively represent the i-th diagonal element and the j-th diagonal element of D (n) and D (a) . Finally, stack by column and take the mean of all results to generate the optimal Z.
[0087] Step D: At the control center, establish an iterative update mechanism.
[0088] Solve equation (10) according to the alternating optimization update mechanism. Specifically, solve it by updating each variable while keeping other variables fixed, that is, solve the optimal solution of the optimal equation (10) through Step C.
[0089] The detection model described in Step 3 specifically includes the following steps:
[0090] Step A: At the control center, repeat Step D in Step 2, and the stop criterion is defined as (obj i-1 - obj i ) / obj i ≤ 10 -4 and obj i ≤ 10 -10 , where obj is the objective value of the iterator, and obj i-1 and obj i are the values at the previous moment and the current moment respectively.
[0091] Step B: Perform anomaly detection in the control center.
[0092] The method designed in this patent is an unsupervised clustering method. Its main anomaly detection method is to regard the points far from the clustering center as anomalies. For the optimal anchor point A obtained after iterative update, calculate the difference between the real-time network data points and the anchor points, regard those with too large a difference from the anchor points as anomalies, and set those within the threshold range as data of the same type as the anchor points.
[0093] Since the embedding learning and anchor graph learning methods achieve data reduction, the anomaly detection method proposed in this patent can effectively process the big data environment. Given the n data with d dimensions collected in the subsequent network. For the subsequent data points x t on the embedding feature c j ∈M (j = 1, 2,..., m), t ∈W T X t The adaptive fuzzy relationship with the anchor point x a ∈A is expressed as follows:
[0094]
[0095] where represents the adaptive domain in the feature c j , represents the standard deviation on c j . The adaptive neighborhood can be adjusted through multi-party verification according to different scenarios. represents the value of x j in the feature c t . By comparing the adaptive fuzzy relationships between each data point and a anchor points, anomaly detection and discrimination can be carried out:
[0096]
[0097] where x i ∈A (i = 1,..., a). Among them, P(x t ) ∈ [0, m] represents the final judgment result. When the maximum value P < m / 2, it is considered that the detection result has no corresponding anchor point of the same type and is regarded as a newly emerging anomaly.
Claims
1. An intrusion detection method and system for Internet of Things devices based on anchor graph learning, characterized in that It is established by an anchor graph learning model and consists of an anchor-based intrusion detection method. It includes the following steps: Step 1: At the control center, collect network traffic, represent it as a feature matrix, and use it as training data. Step 2: At the control center, build an intrusion detection method for IoT devices based on anchor graph learning. Step 3: At the control center, use the collected network traffic to train the detection model established in Step 2. Step 4: End.
2. The establishment of the anchor graph learning model according to claim 1, wherein, The setting of the intrusion detection method for IoT devices based on anchor graph learning described in Step 2 specifically includes the following steps: Step A: At the control center, load the dataset X collected in Step 1 and read the traffic feature matrix. Step B: At the control center, establish the objective function of anchor graph learning. Given dataset For n samples of d dimensions in the dataset, the construction of anchor graph learning is as follows: Among them represents the sample projection matrix, and m (m << d) is the dimension of the embedding space. It should be noted here that the orthogonal constraint (i.e., W T W = I m , where I m is the identity matrix) can make the spatial bases uncorrelated, thereby improving the discrimination ability and avoiding the trivial solution at the same time. is the anchor matrix in the embedding space, and a is the number of anchors. A T A = I a restricts the anchors to make them more discriminative. Z is the anchor graph, and Z is used to represent the relationship between the anchor matrix and the original data. ||·|| F represents the Frobenius norm. Therefore, using formula (1), the anchor matrix A can be dynamically obtained through learning instead of the traditional static sampling. represents the rank constraint, where is the normalized Laplacian matrix of O, and the degree matrix D O is a diagonal matrix associated with O, and its diagonal elements are calculated by . O is the enhanced graph of B, which is defined as: Step C: At the control center, establish the optimization method of anchor graph learning. At the control center, we solve Equation (1) according to the alternating optimization iterative update mechanism. Specifically, we can divide the detailed process into several steps by updating each variable while keeping other variables fixed. When A and Z are fixed, Equation (1) can be rewritten with respect to W as: where P W = XZ T A T , assuming that the matrix in formula (3) has an m-order singular value decomposition form (SVD) as where To maximize the value of equation (3), correspondingly, we can obtain When W and Z are fixed, equation (2) can be rewritten with respect to A as: where P A = W T XZ T , similar to the solution of formula (3), assuming P A has a singular value decomposition form (SVD) of correspondingly, we can obtain where U A represents the left singular vector matrix of P A , and where V A represents the right singular vector matrix of P A . When W and A are fixed, Equation (1) can be rewritten with respect to Z as: By solving Equation (5) through Lagrange, we can get: where, P :,i represents matrix P Z = A T (W T X)'s i-th column element, Z :,i represents the i-th column element of matrix Z, γ represents the regularization term, T i,: represents the row vector composed of t ij . t ij is expressed as follows: Among them, The optimal solution of F should consist of the k eigenvectors with the smallest eigenvalues, where k represents the number of labeled data types. Considering that the entire similarity matrix O is composed of the high-quality anchor graph Z, the first n rows of F actually represent the indicators of data points, while the remaining a rows are the indicators of anchor points. F (n) (i, :) and F (a) (j, :) respectively represent the (n) i-th row and the (a) j-th row of F (n) and D (a) (i, i) and D (n) and D (a) the i-th and j-th diagonal elements of D Finally, stack by columns and take the mean of all results to generate the optimal Z. Step D: At the control center, establish an iterative update mechanism. Solve Equation (1) according to the iterative update mechanism. Specifically, solve it by updating each variable while keeping other variables fixed, that is, solve the optimal solution of the optimal Equation (1) through Step C.
3. The anchor-based intrusion detection method according to claim 1, wherein The setting of the detection model described in Step 3 specifically includes the following steps: Step A: In the control center, repeat step D in step two. The stop criterion is defined as (obj i-1 -obj i ) / obj i ≤10 -4 and obj i ≤10 -10 , where obj is the target value of the iterator, and obj i-1 and obj i are the values at the previous moment and the current moment respectively. Step B: At the control center, perform anomaly detection. The method designed in this patent is an unsupervised clustering method, and its main anomaly detection method is to regard points far from the clustering center as anomalies. For the optimal anchor point A obtained after iterative update, calculate the difference between real-time network data points and the anchor point, regard those with too large a difference from the anchor point as anomalies, and set those within the threshold range as data of the same type as the anchor point. Since the embedding learning and anchor graph learning methods achieve data reduction, the anomaly detection method proposed in this patent can effectively handle the big data environment. Given the n data points of dimension d in t . For the subsequent data points x j ∈ M (j = 1, 2,..., m) on the embedding feature c t ∈ W T X t and the adaptive fuzzy relationship with the anchor point x a ∈ A is represented as follows: Among them represents the adaptive domain in feature c j and represents the standard deviation on c j The adaptive neighborhood can be adjusted through multi - party verification according to different scenarios. represents the value of x j in feature c t . By comparing the adaptive fuzzy relationship between each data point and a anchor points, anomaly detection and discrimination can be carried out: where x i ∈ A (i = 1,..., a). Where P(x t ) ∈ [0, m] represents the final judgment result. When the maximum value P < m / 2, it is considered that there is no corresponding anchor point of the same type in the detection result, and it is regarded as a newly emerged anomaly.