Self-learning algorithm and device for network configuration auditing rules
Through the self-learning algorithm and decision tree model, the configuration data of network equipment is automatically processed and adaptive audit rules are generated, which solves the problems of time-consuming, labor-intensive and poor adaptability of traditional audit methods, and achieves efficient and accurate network configuration audits.
Patent Information
- Application Number
- CN202510330707.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Traditional network configuration auditing methods rely on manual auditing, which is time-consuming and labor-intensive and error-prone, and cannot meet the needs of fast and accurate auditing of large-scale networks. Preset auditing rules are difficult to adapt to changes in the network environment.
The self-learning algorithm is adopted to automatically collect and process the configuration files and operation logs of network devices through data mining and decision tree models, extract key features, generate adaptive audit rules, and use the powerful computing resources of the CPU and GPU for real-time auditing.
It realizes the intelligence and automation of network configuration audits, improves audit efficiency and accuracy, can continuously learn to adapt to changes in the network environment, reduces manual participation, and has high scalability and flexibility.
Smart Images

Figure CN119854127B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and particularly to a self-learning algorithm and device for network configuration audit rules. Background Art
[0002] In today's highly information-based society, the network has become an indispensable infrastructure for all walks of life. With the continuous expansion of the network scale and the increase in complexity, the management and audit of network configuration have become an important issue to be solved urgently. The traditional network configuration audit method mainly relies on manual review, which is not only time-consuming and laborious, but also prone to errors, and cannot meet the requirements of fast and accurate audit of large-scale networks.
[0003] In the prior art, although there are some automated network configuration audit tools, most of these tools use preset audit rules for auditing. However, with the continuous development of network technology and the continuous change of network environment, the preset audit rules are often difficult to adapt to new network configuration scenarios and requirements, resulting in limitations on the accuracy and reliability of audit results.
[0004] To solve the above problems, the present invention proposes a self-learning algorithm and device for network configuration audit rules. Summary of the Invention
[0005] The object of the present invention is to solve the problems existing in the background art, and to propose a self-learning algorithm and device for network configuration audit rules.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] In a first aspect, the present invention provides a self-learning algorithm for network configuration audit rules, including the following steps:
[0008] Step 1: Data collection and preprocessing;
[0009] Obtain the network device configuration file, historical network traffic data, and device operation logs to obtain the original configuration data. Extract abnormal configuration data through data mining of the original network configuration data, and locate potential network configuration problems.
[0010] Number the network devices including routers, switches, and firewalls, and the numbering symbol is i; i = 1, 2,..., n; n is the total number of network devices.
[0011] Collect the configuration files of all network devices i, and extract the configuration parameters therein, including the device name ID_i of the network device, the device type Device_i, the device enable status symbol State_i, the IP address IPAddress_i, the subnet mask Subnet_i, the destination network Destination_i, the gateway Gateway_i, and the interface Interface_i. Among them, when the value of the device enable status symbol State_i is 1, it represents that the network device is in the enabled state; when the value of the device enable status symbol State_i is 0, it represents that the network device is in the disabled state.
[0012] Collect the fault records, device restart records, and performance warnings in the device operation logs of all network devices i, and obtain the timestamp Timestamp_i of the most recent fault, the fault type FaultID_i, the affected device number DeviceID_i, the fault description code Description_i, and the restart time RestartTime_i. If no fault records are extracted from the device operation log of network device i, then set the values of the timestamp Timestamp_i of the most recent fault, the fault type FaultID_i, the affected device number DeviceID_i, the fault description code Description_i, and the restart time RestartTime_i to 0.
[0013] Generate the original configuration data feature matrix of each network device at the current moment t . At preset time intervals, re-obtain the original configuration data of each network device i, and dynamically update its corresponding original configuration data feature matrix.
[0014] Extract statistical features from the original configuration data. The specific process is as follows:
[0015] At preset statistical periods, count all the original configuration data feature matrices obtained by dynamic update for the same network device i, estimate the probability distribution through the frequency of element distribution in the matrix, and use a preset formula to calculate the joint probability distribution of each element in the original configuration data feature matrix and the marginal probability distribution and ; where is the number of original configuration data feature matrices that satisfy the condition: x1 = a and x2 = b; where is the number of original configuration data feature matrices that satisfy the condition: x1 = a; where The number of original configuration data feature matrices that satisfy the condition: x2 = b; where x1 and x2 are two elements in the original configuration data feature matrix; where T is the number of original configuration data feature matrices generated by the same network device within a statistical period. Where a and b are the specific values of the original configuration data x1 and x2 respectively.
[0016] Substitute the calculated joint probability distribution and marginal probability distribution into the preset formula Calculate the statistical influence value of the original configuration data when x1 = a and x2 = b . Where A and B are the sets composed of all specific values of the original configuration data x1 and x2 respectively.
[0017] As a preferred embodiment of the present invention, sort all the calculated statistical influence values in descending order. Its specific arrangement order reflects the strength of the statistical dependence between any two configuration parameters. According to the sorted statistical influence values, select the top k to form the influence feature vector Z(i) = {I1, I2,..., Ik}. Where k is a preset threshold; where I1, I2,..., Ik are the sorted statistical influence values.
[0018] Define the fault label: When it is recognized that in the previous statistical period, faults and configuration problems are recorded in the device operation log of network device i, then set the corresponding fault label Y(i) value to 1; otherwise, set the corresponding fault label Y(i) value to 0.
[0019] Step 2: Self-learning model training and parameter setting;
[0020] Create a configuration diagnosis self-learning model based on a decision tree and perform supervised learning based on the extracted key feature data.
[0021] Use the influence feature vector Z(i) = {I1, I2,..., Ik} of each network device i generated in Step 1 as the input feature of the model. And based on the historical network configuration problems and fault records, automatically label the configuration status of each network device to prepare label data for supervised learning.
[0022] As a preferred embodiment of the present invention, use the extracted feature vector and label data to train the decision tree model. The training process of the decision tree includes the following steps:
[0023] As a preferred embodiment of the present invention, select the optimal feature for splitting at each node in the decision tree and use the information gain as the criterion for feature selection. The information gain formula is:
[0024] ; where, is the entropy of the fault label Y(i); is a subset when the value of the influence feature vector Z(i) is v.
[0025] As a preferred embodiment of the present invention, the decision tree is recursively split. Starting from the root node of the decision tree, the influence feature vectors Z(i) = {I1, I2,..., Ik} of each network device i are used as the initial data to input into the decision tree. At the current node, the information gain of all influence feature vectors is calculated , and the element with the largest information gain in the influence feature vector is selected as the splitting feature.
[0026] According to the selected feature, the data set is divided into subsets, and each subset is recursively split until the stopping condition is met: the value of the information gain is less than the preset threshold or the recursive splitting times of the decision tree are greater than the preset threshold. Finally, a decision tree is generated, and each leaf node corresponds to a category, that is, the network device i is normal or abnormal. When it is determined that the network device i is normal, the leaf node outputs the judgment parameter U(i) = 1; when it is determined that the network device i is abnormal, the leaf node outputs the judgment parameter U(i) = 0.
[0027] As a preferred embodiment of the present invention, after the decision tree is generated, a pruning operation is performed to obtain the final configuration diagnosis self-learning model to prevent overfitting.
[0028] Step Three: Self-learning model evaluation;
[0029] Generate a validation data set, including validation influence feature vectors and validation fault labels .
[0030] After obtaining the final configuration diagnosis self-learning model. Repeat Step One to generate the influence feature vectors of each network device i again, denoted as validation influence feature vectors ={I1, I2,..., Ik}; generate the fault labels of each network device i again, denoted as validation fault labels .
[0031] Use the validation influence feature vectors and validation fault labels to evaluate the trained decision tree model. Input the influence feature vectors into the final configuration diagnosis self-learning model to obtain the output result, and calculate the following metrics:
[0032] For the same validation influence feature vector , if the output result of the configuration diagnosis self-learning model is abnormal, that is, the output Ui = 0; and the value of the validation fault label is 0; then it is recorded as a true positive;
[0033] For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is abnormal, that is, the output Ui = 1; and the value of the verification fault label is 1; then it is recorded as a true negative example;
[0034] For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is abnormal, that is, the output Ui = 0; and the value of the verification fault label is 1; then it is recorded as a false positive example;
[0035] For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is normal, that is, the output Ui = 1; and the value of the verification fault label is 0; then it is recorded as a false negative example;
[0036] Traverse all network devices i = 1, 2,..., n, and count the total number of true positive examples TP, the total number of true negative examples TN, the number of false positive examples FP, and the number of false negative examples FN.
[0037] Through the preset formula Calculate the accuracy Accuracy, precision Precision, recall Recall, and F1 score of the final configured diagnostic self-learning model.
[0038] When the conditions are met: the accuracy Accuracy is greater than the preset threshold and the F1 score is greater than the preset threshold, it is determined that the evaluation of the configured diagnostic self-learning model is qualified, and proceed to step four; otherwise, return to step three and re-perform the decision tree training for the configured diagnostic self-learning model.
[0039] Step four: Generate auditing rules;
[0040] Based on the trained decision tree-based configured diagnostic self-learning model, extract its splitting rules and leaf node judgment conditions to generate network configuration auditing rules. The network configuration auditing rules are expressed as a series of logical conditions for judging whether the configuration of network devices is normal. Specifically:
[0041] Whenever the impact feature vector of a certain network device is obtained, it is input into the root node of the configured diagnostic self-learning model to obtain the output value U(i) of the leaf node, and the path from the root node to the leaf node is recorded as a rule.
[0042] For example, assume that a path of the decision tree is: root node: I1 > 0.8. Internal node: I2 < 0.5; leaf node: U(i) = 0. Then generate the auditing rule:
[0043] If I1 > 0.8 AND I2 < 0.5, then the configuration of network device i is abnormal.
[0044] Then generate an audit rule:
[0045] If I1 > 0.8 AND I2 < 0.5, then the configuration of device i is abnormal.
[0046] Traverse each internal node of the decision tree and record its splitting feature and splitting condition.
[0047] As a preferred embodiment of the present invention, record all the generated audit rules for optimization, remove redundant conditions, and merge similar rules by using a rule induction algorithm.
[0048] In a second aspect, the present invention provides a self-learning device for network configuration audit rules, including a number of network nodes, and each network node includes a central processing unit (CPU), a graphics processing unit (GPU), a storage device, and a network interface.
[0049] Among them, the central processing unit is responsible for overall task scheduling, data preprocessing, rule generation, and logical control. The graphics processing unit is used to accelerate the model training and inference processes, especially the splitting calculation and feature importance analysis of the decision tree model, and supports a deep learning acceleration library to optimize the training efficiency of the decision tree model.
[0050] All network nodes are integrated through network connections and are respectively equipped with preset modules, including: a data collection and preprocessing module, a model training module, a rule generation module, and a real-time audit module.
[0051] The data collection and preprocessing module collects configuration data and operation logs from network devices and performs preprocessing. Built-in data cleaning and feature extraction algorithms are used to automatically generate impact feature vectors.
[0052] The model training module trains a configuration diagnosis self-learning model based on the decision tree algorithm, and built-in information gain calculation and recursive splitting algorithms are used to generate a decision tree model.
[0053] The rule generation module extracts audit rules from the trained decision tree model and removes redundant conditions and merges similar rules through a built-in rule optimization algorithm.
[0054] The real-time audit module applies the generated audit rules to real-time network configuration detection. Built-in anomaly detection and alarm mechanisms are used to trigger an alarm and send an anomaly notification signal to the administrator when a network configuration anomaly is identified.
[0055] Compared with the prior art, the beneficial effects of the present invention are:
[0056] 1. The present invention realizes the intelligent auditing of network device configurations through a self - learning algorithm for network configuration auditing rules. This algorithm can automatically collect and process the configuration files, historical traffic data, and operation logs of network devices, extract key features, and train a decision tree model based on these features. This process greatly reduces manual participation and improves the efficiency and accuracy of auditing. At the same time, as the network environment constantly changes, the algorithm can continuously learn and update to adapt to new configuration problems and fault patterns, ensuring the timeliness and effectiveness of auditing rules;
[0057] 2. The self - learning device for network configuration auditing rules of the present invention has high scalability and flexibility. Each network node in the device is equipped with powerful computing resources (including CPU and GPU) and can process large - scale network configuration data. In addition, the algorithm supports a variety of feature extraction and model training methods and can be customized and optimized according to different network environments and requirements. This scalability and flexibility enable the present invention to be widely applied to various scales and types of network environments to meet the auditing needs of different users;
[0058] 3. The present invention significantly improves the speed and accuracy of configuration problem diagnosis. By deeply mining the statistical dependencies between configuration data, the algorithm can accurately identify potential network configuration problems and generate corresponding auditing rules. These rules are represented in the form of logical conditions, which are easy to understand and apply, helping network administrators quickly locate and solve configuration problems and reducing the risk of network failures. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the drawings:
[0060] Figure 1 is the flowchart of the method of the present invention;
[0061] Figure 2 is the schematic diagram of the decision tree of the present invention;
[0062] Figure 3 is the schematic diagram of the network node of the self - learning device for a network configuration auditing rule of the present invention;
[0063] Figure 4 is the system block diagram of the self - learning device for a network configuration auditing rule of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0064] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] Please refer to Figure 1 as shown, a self-learning algorithm for network configuration auditing rules includes the following steps:
[0066] Step 1: Data collection and preprocessing;
[0067] Obtain the network device configuration files, historical network traffic data, and device operation logs to obtain the original configuration data. Extract abnormal configuration data through data mining of the original network configuration data to locate potential network configuration problems.
[0068] Number the network devices including routers, switches, and firewalls, and the numbering symbol is i; i = 1, 2,..., n; n is the total number of network devices.
[0069] Collect the configuration files of all network devices i, and extract the configuration parameters therein, including the device name ID_i, device type Device_i, device enabled status symbol State_i, IP address IPAddress_i, subnet mask Subnet_i, target network Destination_i, gateway Gateway_i, and interface Interface_i of the network device. Among them, when the value of the device enabled status symbol State_i is 1, it represents that the network device is in the enabled state; when the value of the device enabled status symbol State_i is 0, it represents that the network device is in the disabled state.
[0070] Collect the fault records, device restart records, and performance warnings in the device operation logs of all network devices i, and obtain the timestamp Timestamp_i of the most recent fault, fault type FaultID_i, affected device number DeviceID_i, fault description code Description_i, and restart duration RestartTime_i therein. If no fault record is extracted from the device operation log of network device i, then set the values of the timestamp Timestamp_i of the most recent fault, fault type FaultID_i, affected device number DeviceID_i, fault description code Description_i, and restart duration RestartTime_i to 0.
[0071] Generate the original configuration data feature matrix of each network device at the current moment t At every preset time interval, re-obtain the original configuration data of each network device i, and dynamically update the corresponding original configuration data feature matrix.
[0072] Furthermore, perform statistical feature extraction on the original configuration data. The specific process is as follows:
[0073] At every preset statistical period, count all the original configuration data feature matrices obtained by dynamic update for the same network device i, estimate the probability distribution through the frequency of element distribution in the matrix, and use a preset formula to calculate the joint probability distribution of each element in the original configuration data feature matrix and the marginal probability distribution and ; where is the number of original configuration data feature matrices that satisfy the condition: x1 = a and x2 = b; where is the number of original configuration data feature matrices that satisfy the condition: x1 = a; where is the number of original configuration data feature matrices that satisfy the condition: x2 = b; where x1 and x2 are two elements in the original configuration data feature matrix; where T is the number of original configuration data feature matrices generated by the same network device within a statistical period. Where a and b are the specific values of the original configuration data x1 and x2 respectively.
[0074] According to the calculated joint probability distribution and marginal probability distribution, substitute them into the preset formula to calculate the statistical influence value of the original configuration data when x1 = a and x2 = b . Where A and B are the sets composed of all specific values of the original configuration data x1 and x2 respectively.
[0075] It should be noted that the statistical influence value is an index to measure the non-linear relationship between two variables, which can capture the statistical dependence between variables. The statistical influence value measures the shared information amount between two variables. The larger the value of the statistical influence value, the stronger the correlation between the two variables. For example, when the subnet mask Subnet_i of xi is 255.255.255.0 and the IP address IPAddress_i is 192.168.0.1, the calculated statistical influence value has a value of 0.75, which indicates that there is a strong correlation between the corresponding IP address and subnet mask, that is, there is a certain statistical dependence between their values.
[0076] Further, sort all the calculated statistical impact values in descending order. The specific arrangement order reflects the strength of the statistical dependence between any two configuration parameters. According to the sorted statistical impact values, select the top k to form the impact feature vector Z(i) = {I1, I2,..., Ik}. Where k is a preset threshold; where I1, I2,..., Ik are the sorted statistical impact values.
[0077] It should be noted that the impact feature vector reflects the occurrence probability or type of network configuration problems and can be used for subsequent machine learning model training or diagnosis of network configuration problems.
[0078] Define the fault label: When it is recognized that in the device operation log of network device i in the previous statistical period, faults and configuration problems are recorded, then set the corresponding fault label Y(i) value to 1; otherwise, set the corresponding fault label Y(i) value to 0.
[0079] Step 2: Self-learning model training and parameter setting;
[0080] Create a configuration diagnosis self-learning model based on decision trees and perform supervised learning based on the extracted key feature data.
[0081] Take the impact feature vector Z(i) = {I1, I2,..., Ik} of each network device i generated in Step 1 as the input feature of the model. And based on the historical network configuration problems and fault records, automatically label the configuration status of each network device to prepare label data for supervised learning.
[0082] Further, use the extracted feature vector and label data to train the decision tree model. The training process of the decision tree includes the following steps:
[0083] Please refer to Figure 2 As shown, at each node in the decision tree, select the optimal feature for splitting and use information gain as the criterion for feature selection. The information gain formula is:
[0084] ; where is the entropy of the fault label Y(i); is the subset when the impact feature vector Z(i) takes the value v.
[0085] Further, perform recursive splitting on the decision tree. Starting from the root node of the decision tree, take the impact feature vector Z(i) = {I1, I2,..., Ik} of each network device i as the initial data input to the decision tree. At the current node, calculate the information gain of all impact feature vectors and select the element with the largest information gain in the impact feature vector as the splitting feature.
[0086] The data set is divided into subsets according to the selected features, and each subset is recursively split until the stopping condition is met: the value of the information gain is less than the preset threshold or the number of recursive splits of the decision tree is greater than the preset threshold. Finally, a decision tree is generated, and each leaf node corresponds to a category, that is, the network device i is normal or abnormal. When it is determined that the network device i is normal, the leaf node outputs the judgment parameter U(i)=1; when it is determined that the network device i is abnormal, the leaf node outputs the judgment parameter U(i)=0.
[0087] Furthermore, after the decision tree is generated, a pruning operation is performed to obtain the final configuration diagnosis self-learning model to prevent overfitting.
[0088] Step 3: Self-learning model evaluation;
[0089] Generate a validation data set, including validation impact feature vectors and validation fault labels .
[0090] After obtaining the final configuration diagnosis self-learning model. Repeat Step 1 to generate the impact feature vectors of each network device i again, denoted as validation impact feature vectors ={I1, I2,..., Ik}; generate the fault labels of each network device i again, denoted as validation fault labels .
[0091] Use the validation impact feature vectors and validation fault labels to evaluate the trained decision tree model. Input the impact feature vectors into the final configuration diagnosis self-learning model to obtain the output result, and calculate the following metrics:
[0092] For the same validation impact feature vector , if the output result of the configuration diagnosis self-learning model is abnormal, that is, the output Ui=0; and the value of the validation fault label is 0; then it is recorded as a true positive case;
[0093] For the same validation impact feature vector , if the output result of the configuration diagnosis self-learning model is abnormal, that is, the output Ui=1; and the value of the validation fault label is 1; then it is recorded as a true negative case;
[0094] For the same validation impact feature vector , if the output result of the configuration diagnosis self-learning model is abnormal, that is, the output Ui=0; and the value of the validation fault label is 1; then it is recorded as a false positive case;
[0095] For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is normal, that is, the output Ui = 1; while the value of the verification fault label is 0; then it is recorded as a false negative example;
[0096] Traverse all network devices i = 1, 2,..., n, and count the total number of true positive examples TP, the total number of true negative examples TN, the number of false positive examples FP, and the number of false negative examples FN.
[0097] Through the preset formula Calculate the accuracy, precision, recall, and F1 score of the final configured diagnostic self-learning model.
[0098] When the conditions are met: the accuracy is greater than the preset threshold and the F1 score is greater than the preset threshold, it is determined that the configured diagnostic self-learning model evaluation is qualified, and proceed to step four; otherwise, return to step three and re-perform the decision tree training for the configured diagnostic self-learning model.
[0099] It should be noted that the accuracy measures the overall prediction accuracy of the model; the precision measures the reliability of the model's prediction of positive examples; the recall measures the ability of the model to discover positive examples; the F1 score is a comprehensive index of precision and recall, suitable for scenarios where a balance between the two is required. By calculating these metrics, the performance of the decision tree model can be comprehensively evaluated, and the model can be optimized based on the evaluation results, thereby improving the accuracy and efficiency of network configuration auditing.
[0100] Step four: Generate auditing rules;
[0101] Based on the trained configuration diagnostic self-learning model based on the decision tree, extract its splitting rules and leaf node judgment conditions to generate network configuration auditing rules. The network configuration auditing rules are expressed as a series of logical conditions for judging whether the configuration of network devices is normal, specifically:
[0102] Whenever the impact feature vector of a certain network device is obtained, it is input into the root node of the configuration diagnostic self-learning model to obtain the output value U(i) of the leaf node, and the path from the root node to the leaf node is recorded as a rule.
[0103] For example, assume that a path of the decision tree is: root node: I1 > 0.8. Internal node: I2 < 0.5; leaf node: U(i) = 0. Then generate the auditing rule:
[0104] IF I1 > 0.8 AND I2 < 0.5 THEN the configuration of network device i is abnormal.
[0105] Then generate the auditing rule:
[0106] If I1 > 0.8 AND I2 < 0.5, then the configuration of device i is abnormal.
[0107] Traverse each internal node of the decision tree and record its splitting feature and splitting condition.
[0108] Furthermore, record all generated audit rules for optimization, remove redundant conditions, and merge similar rules by using a rule induction algorithm.
[0109] It should be noted that the significance of generating audit rules is to improve the speed of network configuration diagnosis by recording fixed mapping relationships. There is no need to call the configuration diagnosis self-learning model based on the decision tree in each network diagnosis, reducing unnecessary waste of computing power.
[0110] Please refer to Figure 3 and Figure 4 As shown in
[0111] A self-learning device for network configuration audit rules includes several network nodes, and each network node includes a central processing unit (CPU), a graphics processing unit (GPU), a storage device, and a network interface.
[0112] All network nodes are integrated through network connections and are respectively equipped with preset modules, including: a data collection and preprocessing module, a model training module, a rule generation module, and a real-time audit module.
[0113] The data collection and preprocessing module collects configuration data and operation logs from network devices and performs preprocessing. Built-in data cleaning and feature extraction algorithms are used to automatically generate impact feature vectors.
[0114] The model training module trains a configuration diagnosis self-learning model based on the decision tree algorithm. Built-in information gain calculation and recursive splitting algorithms are used to generate a decision tree model.
[0115] The rule generation module extracts audit rules from the trained decision tree model and removes redundant conditions and merges similar rules through a built-in rule optimization algorithm.
[0116] The real-time audit module applies the generated audit rules to real-time network configuration detection. Built-in anomaly detection and alarm mechanisms are used to trigger an alarm and send an anomaly notification signal to the administrator when a network configuration anomaly is identified.
[0117] It should be understood that the terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0118] It should also be understood that the terms used in this disclosure specification are for the purpose of describing particular embodiments only and are not intended to limit this disclosure. As used in this disclosure specification and claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should further be understood that the term "and / or" as used in this disclosure specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations;
[0119] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A self-learning algorithm for network configuration auditing rules, characterized in that, It includes the following steps: Step 1, data collection and preprocessing; Obtain the network device configuration files, historical network traffic data, and device operation logs to obtain the original configuration data; extract abnormal configuration data through data mining of the original network configuration data, locate potential network configuration problems, generate the corresponding impact feature vectors for each network device, and define the fault labels of the network devices; Step 2, self-learning model training and parameter setting; Create a configuration diagnosis self-learning model based on decision trees and perform supervised learning based on the extracted key feature data; Step 3, self-learning model evaluation; Evaluate the configuration diagnosis self-learning model to generate a validation data set, including validation impact feature vectors and validation fault labels; use the trained configuration diagnosis self-learning model to process the validation data set to obtain the output results; according to the comparison between the output results of the model and the validation fault labels, count the numbers of true positives, true negatives, false positives, and false negatives; Calculate the accuracy, precision, recall rate, and F1 score of the model through a preset formula; if both the accuracy and F1 score of the model are greater than the preset threshold, it is determined that the model evaluation is qualified, and proceed to Step 4; Otherwise, return to Step 3 to retrain the model; Step 4, audit rule generation; Extract audit rules. Based on the trained decision tree model, extract its splitting rules and leaf node judgment conditions; represent the extracted rules as logical conditions for judging whether the configuration of the network device is normal; specifically, input the impact feature vector into the root node of the decision tree, trace the path to the leaf node, and record the splitting features and conditions on the path to form the audit rules.
2. The self-learning algorithm for network configuration auditing rules according to claim 1, wherein The specific process of obtaining the original configuration data is as follows: Number the network devices including routers, switches, and firewalls, and the numbering symbol is i; i = 1, 2,..., n; n is the total number of network devices; Collect the configuration files of all network devices i and extract the configuration parameters, including the device name ID_i, device type Device_i, device enabled status symbol State_i, IP address IPAddress_i, subnet mask Subnet_i, destination network Destination_i, gateway Gateway_i, and interface Interface_i of the network device; among them, when the value of the device enabled status symbol State_i is 1, it represents that the network device is in the enabled state; when the value of the device enabled status symbol State_i is 0, it represents that the network device is in the disabled state; Collect the fault records, device restart records, and performance warnings in the device operation logs of all network devices i, and obtain the timestamp Timestamp_i of the most recent fault, the fault type FaultID_i, the affected device number DeviceID_i, the fault description code Description_i, and the restart time RestartTime_i among them; if no fault record is extracted from the device operation log of network device i, then set the values of the timestamp Timestamp_i of the most recent fault, the fault type FaultID_i, the affected device number DeviceID_i, the fault description code Description_i, and the restart time RestartTime_i to 0; Generate the original configuration data feature matrix of each network device at the current moment t ; At every preset time interval, re-obtain the original configuration data of each network device i, dynamically update its corresponding original configuration data feature matrix, and extract statistical features from the original configuration data.
3. The self-learning algorithm for network configuration auditing rules according to claim 2, wherein The specific process of extracting statistical features from the original configuration data is as follows: At every preset statistical period, all the original configuration data feature matrices obtained by a same network device i through dynamic update are statistically analyzed, and the probability distribution is estimated by the frequency of the element distribution in the matrix. Through a preset formula calculate the joint probability distribution of each element in the original configuration data feature matrix as well as the marginal probability distribution and ; where is the number of original configuration data feature matrices that satisfy the condition: x1 = a and x2 = b; where is the number of original configuration data feature matrices that satisfy the condition: x1 = a; where is the number of original configuration data feature matrices that satisfy the condition: x2 = b; where x1 and x2 are two elements in the original configuration data feature matrix; where T is the number of original configuration data feature matrices generated by the same network device within a statistical period; where a and b are the specific values of the original configuration data x1 and x2 respectively; Substitute the calculated joint probability distribution and marginal probability distribution into the preset formula Calculate the statistical influence value of the original configuration data when x1 = a and x2 = b ; where A and B are the sets composed of all specific values of the original configuration data x1 and x2 respectively; Sort all the calculated statistical influence values in descending order; its specific arrangement order reflects the strength of the statistical dependence between any two configuration parameters; according to the sorted statistical influence values, select the top k to form the influence feature vector Z(i) = {I1, I2,..., Ik}; where k is a preset threshold; where I1, I2,..., Ik are the sorted statistical influence values.
4. The self-learning algorithm for network configuration auditing rules according to claim 1, characterized in that The specific process of defining the fault label is as follows: when it is recognized that in the previous statistical period, faults and configuration problems are recorded in the device operation log of network device i, then set the corresponding fault label Y(i) value to 1; otherwise, set the corresponding fault label Y(i) value to 0.
5. The self - learning algorithm for network configuration auditing rules according to claim 1, characterized in that, The specific process of creating a self-learning configuration diagnosis model based on a decision tree is as follows: Take the influence feature vector Z(i) = {I1, I2,..., Ik} of each network device i generated in step one as the input feature of the model; And according to the historical network configuration problems and fault records, automatically label the configuration status of each network device to prepare label data for supervised learning; Use the extracted feature vectors and label data to train the decision tree model; The training process of the decision tree includes the following steps: Select the optimal feature for splitting at each node in the decision tree, and use the information gain as the criterion for feature selection. The information gain formula is: ; among them, is the entropy of the fault label Y(i); is the subset when the value of the influence feature vector Z(i) is v, and the decision tree is recursively split.
6. The self-learning algorithm for network configuration auditing rules according to claim 1, characterized in that: The specific process of recursively splitting the decision tree is as follows: Starting from the root node of the decision tree, take the influence feature vector Z(i) = {I1, I2,..., Ik} of each network device i as the initial data and input it into the decision tree. At the current node, calculate the information gain of all influence feature vectors , and select the element with the largest information gain in the influence feature vector as the splitting feature; Divide the data set into subsets according to the selected feature, and recursively split each subset until the stopping condition is met: the value of the information gain is less than the preset threshold or the number of recursive splits of the decision tree is greater than the preset threshold, and finally generate a decision tree. Each leaf node corresponds to a category, that is, network device i is normal or abnormal; when it is judged that network device i is normal, the leaf node outputs the judgment parameter U(i) = 1; when it is judged that network device i is abnormal, the leaf node outputs the judgment parameter U(i) = 0.
7. The self-learning algorithm for network configuration auditing rules according to claim 1, characterized in that, The specific process of evaluating the self-learning configuration diagnosis model is as follows: Generate a validation dataset, including validation impact feature vectors and validation fault labels ; After obtaining the final configuration diagnosis self-learning model; repeat step one to generate the impact feature vectors of each network device i again, denoted as verification impact feature vectors ={I1, I2,..., Ik}; generate the fault labels of each network device i again, denoted as verification fault labels ; Using verification impact feature vectors and verification fault labels evaluate the trained decision tree model, input the impact feature vectors into the final configuration diagnosis self-learning model, obtain the output results, and calculate the following metrics: For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is abnormal, that is, the output Ui = 0; and the verification fault label has a value of 0; then it is recorded as a true positive example; For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is abnormal, that is, the output Ui = 1; and the verification fault label has a value of 1; then it is recorded as a true negative example; For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is abnormal, that is, the output Ui = 0; and the verification fault label has a value of 1; then it is recorded as a false positive example; For the same verification impact feature vector , if the output result of the configured diagnostic self-learning model is normal, that is, the output Ui = 1; while the value of the verification fault label is 0; then it is recorded as a false negative example; Traverse all network devices i = 1, 2,..., n, and count the total number of true positives TP, the total number of true negatives TN, the number of false positives FP, and the number of false negatives FN; Through a preset formula Calculate the accuracy, precision, recall, and F1-score of the finally configured diagnostic self-learning model; When the conditions are met: the accuracy is greater than the preset threshold and the F1 score is greater than the preset threshold, it is determined that the configuration diagnosis self-learning model evaluation is qualified, and step four is entered. Otherwise, return to step three and re-perform the decision tree training for the configuration diagnosis self-learning model.
8. The self - learning algorithm for auditing rules of network configuration according to claim 1, characterized in that, The specific process of forming the audit rules is as follows: Based on the trained configuration diagnosis self-learning model based on the decision tree, extract its splitting rules and leaf node judgment conditions to generate network configuration audit rules; the network configuration audit rules are expressed as a series of logical conditions for judging whether the configuration of network devices is normal, specifically: Whenever the impact feature vector of a certain network device is obtained, it is input into the root node of the configuration diagnosis self-learning model to obtain the output value U(i) of the leaf node, and the path from the root node to the leaf node is recorded as a rule; traverse each internal node of the decision tree and record its splitting feature and splitting condition; Record all the generated audit rules for optimization, remove redundant conditions, and merge similar rules by using the rule induction algorithm.
9. A self-learning device for network configuration audit rules, which is used to implement the self-learning algorithm for network configuration audit rules described in any one of claims 1-8, characterized in that: It includes several network nodes, and each network node includes a central processing unit (CPU), a graphics processing unit (GPU), a storage device, and a network interface; The central processing unit is responsible for overall task scheduling, data preprocessing, rule generation, and logic control; the graphics processing unit is used to accelerate the model training and inference process, especially the splitting calculation and feature importance analysis of the decision tree model, and supports the deep learning acceleration library to optimize the training efficiency of the decision tree model; All network nodes are integrated through network connections and are respectively equipped with preset modules, including: data collection and preprocessing module, model training module, rule generation module, and real-time audit module; The data collection and preprocessing module collects configuration data and operation logs from network devices and performs preprocessing; built-in data cleaning and feature extraction algorithms are used to automatically generate impact feature vectors; The model training module trains the configuration diagnosis self-learning model based on the decision tree algorithm, and built-in information gain calculation and recursive splitting algorithms are used to generate a decision tree model; The rule generation module extracts audit rules from the trained decision tree model and removes redundant conditions and merges similar rules through the built-in rule optimization algorithm; The real-time audit module applies the generated audit rules to real-time network configuration detection; built-in anomaly detection and alarm mechanisms are used to trigger an alarm and send an anomaly notification signal to the administrator when network configuration anomalies are identified.
Citation Information
Patent Citations
Abnormal data auditing rule distinguishing method and device, electronic equipment and storage medium
CN118014617A
Cited By
Heterogeneous configuration text rapid matching method based on skeleton index and frozen semantic anchor point
CN121980284A
Cross-platform network configuration anomaly detection method based on heterogeneous evidence collaborative reasoning
CN122802379A