An AI-based data transmission anomaly supervision system and method
By analyzing abnormal data of communication networks based on artificial intelligence, building a collection of target inspection nodes, solving the problem of low efficiency in network node abnormal data inspection, and achieving efficient abnormal data source location and maintenance.
Patent Information
- Application Number
- CN202310575898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-22
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-05-22
AI Technical Summary
In a communication network, after abnormal data appears on the network node, it is difficult for managers to efficiently check the reasons for abnormal data. The existing technology methods are inefficient and may cause secondary damage to network operation.
Using an artificial intelligence-based method, by analyzing historical operation records and current abnormal data, calculating the association relationship and information entropy of nodes, building a collection of target inspection nodes, setting a risk value evaluation function, and giving inspection suggestions.
It improves the efficiency of managers in operating and maintaining the communication network, accurately locates abnormal data source nodes, and reduces secondary damage to the network.
Smart Images

Figure CN116614395B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network data detection, and particularly to an artificial intelligence-based data transmission anomaly supervision system and method. Background Technique
[0002] Data transmission is to transmit data from a data source to a data terminal through one or more data links according to certain regulations. Its main function is to realize information transmission and exchange between points. A network node refers to a network connected with an independent address and the function of transmitting or receiving data. Network nodes can be workstations, clients, network users or personal computers, and can also be servers, printers and other network-connected devices. With the rapid development of network technologies such as Internet technology and Internet of Things technology, the number of network nodes is not only increasing continuously, but also the services carried are becoming more diversified. In a communication network, when an abnormal data is detected by a certain network node, the source node that generates the abnormal data may not be the current node. The generation of this abnormal data has a certain degree of concealment, and it requires managers to conduct in-depth investigation of the nodes. In the actual operation process of a communication network, the reasons for the same abnormal data generated by network nodes may also be different. Therefore, after abnormal data appears in network nodes, it is becoming more and more difficult for managers to investigate the reasons for the generation of abnormal data.
[0003] Chinese Patent with publication number CN104993960A and name of a method for locating network node faults provides a method for locating network node faults. This method uses the method of traversing all alarm-generating nodes in a computer network to screen out faulty nodes. If a large number of alarm messages are generated simultaneously in a computer network, it will cause an operation burden on the supervision system, resulting in a decrease in the efficiency of troubleshooting faulty nodes. A large amount of data used for traversal calculation entering the computer network will also reduce the operation efficiency of the computer network and cause secondary damage to the computer network. Summary of the Invention
[0004] The purpose of the present invention is to provide an artificial intelligence-based data transmission anomaly supervision system and method to solve the problems proposed in the above background technique.
[0005] To solve the above technical problems, the present invention provides the following technical solution: An artificial intelligence-based data transmission anomaly supervision method, the method includes:
[0006] Step S100: Collect historical operation records of each data transmission node from the historical operation logs of the communication network, and record the historical operation records with data transmission anomalies as abnormal transmission records. Each abnormal transmission record includes several operation parameter items. After normalizing the values of each operation parameter item in the abnormal transmission record, record them in the abnormal transmission record set B;
[0007] Step S200: Classify the abnormal transmission records in the abnormal transmission record set B to obtain several abnormal data transmission types, extract the characteristics of the influence scope of each abnormal data transmission type, and obtain the feature vector of the influence scope of each abnormal data transmission type;
[0008] Step S300: Detect the abnormal data transmission information in the current communication network, set the data transmission node with abnormal transmission data as the target node, collect the operation data of the target node, analyze the operation data, obtain the correspondence between the operation data of the target node and the influence scope of the abnormal data transmission type, further obtain the first associated node of the target node, and set the first weight value of each first associated node;
[0009] Step S400: Extract the connection content data stream information within the T time period before the abnormal transmission data appears at the target node, calculate the information entropy of the connection content data stream, set the node influence threshold, calculate the separation information entropy, set the node corresponding to the separation information entropy as the second associated node, and set the second weight value of each second associated node according to the information entropy size of communicating with the target node;
[0010] Step S500: Assemble all the nodes in the first associated node and the second associated node to form the target troubleshooting node set, set the risk value evaluation function, calculate the node risk value, and give suggestions for the management to troubleshoot. If the management fails to detect the node that generates abnormal data in the target troubleshooting node set, set the node with the longest logical distance from the target node as the new target node, return to step S300 for cycling, and give new troubleshooting suggestions again.
[0011] Further, step S200 includes:
[0012] Step S201: Project the abnormal transmission records into points in space, where the number of operation parameter items in the abnormal transmission records corresponds to the space dimension, and calculate the Euclidean distance between the points corresponding to each abnormal transmission record;
[0013] Step S202: Select the core data in the abnormal transmission records, set the core data evaluation threshold α and the core data evaluation radius r1. In the space projected by the abnormal transmission records, use x to represent an abnormal transmission record. When the number of abnormal transmission records with a Euclidean distance less than r1 from x is greater than α, set x as the core data, and all the abnormal transmission records within the range with x as the center and r1 as the radius form an abnormal data transmission type;
[0014] Within a certain range, several abnormal transmission records appear, indicating that this range is where the distribution probability of node abnormal transmission records is relatively high when abnormal data transmission occurs. Compared with the traditional supervision method that corresponds one by one to the operation parameter values, by summarizing the distribution law of historical data and using the data range to replace the value of the data itself, data that conforms to the distribution law of historical data but is not completely consistent with the historical data can be compared, further improving the supervision effect;
[0015] Step S203: Set the influence radius r2 of the core data. The range centered on the core data with r2 as the radius is set as the core data influence range;
[0016] By setting different values of the influence radius of the core data, different core data influence ranges can be delimited, different core data management strategies can be deployed, and the management method for abnormal data transmission is more flexible;
[0017] Step S204: From the operation and maintenance records of the communication network, extract the nodes that generate abnormal data detected after abnormal data transmission occurs. The nodes that generate abnormal data correspond to the historical operation logs, extract the corresponding relationship between the nodes that generate abnormal data and the abnormal transmission records, and extract a combination of one or more nodes that generate abnormal data corresponding to the core data;
[0018] Step S205: Set a combination form of nodes of an abnormal data in the historical operation log as an abnormal event, where one abnormal event includes k core data;
[0019] The abnormal data generated by one or more nodes that generate abnormal data in the communication network will affect several data transmission nodes. The several core data extracted represent the impact of one or more nodes that generate abnormal data on the communication network;
[0020] Step S206: Calculate the abnormal event influence range D of the abnormal event, where D = C1 ∪ C2 ∪ C3 ∪ … ∪ C k , and use C1, C2, C3, … C k to represent the 1st, 2nd, 3rd, …, kth core data influence ranges included in the event influence range D respectively;
[0021] Step S207: Aggregate the abnormal influence ranges of all abnormal events, denoted as the first abnormal influence range, and calculate the feature vectors of the abnormal event influence ranges in the first abnormal influence range;
[0022] The feature vector is a vector representing the characteristics of each abnormal event influence range.
[0023] Furthermore, step S300 includes:
[0024] Step S301: Set the data transmission node that currently detects abnormal transmission data as the target node, and collect the operation data of the target node when abnormal data transmission occurs. The collected operation data includes at least one operation parameter item;
[0025] Step S302: After normalizing the collected alarm operation data, set it as the abnormal transmission state vector;
[0026] Step S303: Project the abnormal transmission state vector and the feature vector of the first abnormal influence range into the same dimensional space, and calculate the included angle θ between the abnormal transmission state vector and the feature vector of the first abnormal influence range;
[0027] Step S304: Calculate the cosine function value g of the included angle between the abnormal transmission state vector and the feature vector of the first abnormal influence range cos ;
[0028] Step S305: Extract the values greater than 0 in g cos Let the number of values greater than 0 in g cos be u. According to the correspondence between abnormal events and a group of nodes that generate abnormal data, extract all u types of nodes that generate abnormal data corresponding to abnormal events, and record them in the first associated node V1;
[0029] Step S306: Set the first weight value P of each node in the first associated node V1. The first weight value is the cosine value of the included angle between the feature vector of the first abnormal influence range where the node is located and the operation data vector;
[0030] By calculating the cosine value of the included angle between the alarm operation data vector and the feature vector of the first abnormal influence range, the closeness between the operation data vector and the operation parameters when data is abnormally transmitted can be obtained. The smaller the vector included angle, the closer the operation data vector is to the operation state corresponding to the abnormal event, and the greater the possibility of this abnormal cause in the communication network;
[0031] Within the range of the vector included angle [0, 2π], the cosine function value is monotonically decreasing. The calculation result is within the range of [-1, 1]. Extract the part greater than 0 in the calculation result, that is, extract the calculation result within the range of [0, 1]. These calculation results have normalization, which is conducive to the operations in subsequent steps. The cosine value of the vector included angle greater than 0 indicates the consistency between the operation data vector and the corresponding feature vector of the first abnormal influence range.
[0032] Further, step S400 includes:
[0033] Step S401: Extract the connection content data stream information of the target node within the T time period before the abnormal transmission data alarm is generated. The connection content data stream information includes the transmission path information of the target node connection content, and the connection content data stream information includes m parameter items, where m≥1;
[0034] Step S402: Set the nodes that send data to the target node within the T time period as the data sending nodes. Select one of the parameter items in the connection content data stream information parameter items, calculate the joint entropy of the target node connection content data stream information of this item and the information entropy of each data sending node. Use H(x) to represent the joint entropy of the parameter item in the target node connection content data stream;
[0035] The continuous entropy of the random variable X is defined as H(X) = -∫ x f(x)lnf(x)dx, where f(x) is the probability density function of the random variable X. The joint entropy of the random variables X and Y is defined as H(X,Y) = -∫∫ x,y f(x,y)lnf(x,y)dxdy, where f(x,y) is the joint probability density of the random variables X and Y. Take the nodes that have communication records with the target node within the T time period as random variable nodes. For the sake of simple calculation, it is considered that the communication behaviors of each random variable node with the target node within the T time period are independent of each other. Then the joint entropy is not less than the entropy of any of the variables, H(X1,X2,X3,…,X n )≥max(H(X1),H(X2),H(X3),……H(X n ))), where X1,X2,X3,…,X n , respectively represent the 1st, 2nd, 3rd, ……, nth random variable nodes, and H(X1),H(X2),H(X3),……H(X n )) respectively represent the information entropy of the 1st, 2nd, 3rd, ……, nth random variable nodes within the T time period;
[0036] Step S403: Set the node influence threshold ω, and separate the n separated information entropies greater than the threshold ω. Use H(X1),H(X2),H(X3),……H(X n ) to represent the separated information entropy of the 1st, 2nd, 3rd, ……, nth data sending nodes sending the parameter item respectively. Each separated information entropy corresponds to a separated data transmission node;
[0037] Step S404: Calculate the second weight influence value F of the connection content data stream information parameter item, where where F s represents the second weight influence value of the sth node among the n data sending nodes, H(Xs ) represents the s-th separate information entropy, minH(X) represents the minimum value among n separate information entropies, and maxH(X) represents the maximum value among n separate information entropies;
[0038] Step S405: Calculate the second weight influence value of the data flow information parameter item with connection content, gather the second weight influence values of each separate data transmission node, incorporate all the nodes with second weight influence values into the second associated node V2, and normalize the second weight influence values of each node in the second associated node to obtain the second weight value Q of each second associated node.
[0039] Further, step S500 includes:
[0040] Step S501: Gather each node in the first associated node V1 and the second associated node V2 to form a target troubleshooting node set W, where W = V1 ∪ V2, and each target node in the target troubleshooting node set corresponds to a first weight value, a second weight value, or a combination of the first weight value and the second weight value;
[0041] Step S502: Set the first target troubleshooting coefficient λ and the second target troubleshooting coefficient μ, and calculate the risk value γ of each node in the target troubleshooting node set, where γ = λP w + μQ w , P w is the first weight value corresponding to a target troubleshooting node w in W, and Q w is the second weight value corresponding to w;
[0042] When there is a node in the target troubleshooting node set W without a first weight value, substitute Pw = 0 into the formula = λPw + μQw for calculation. When there is a node in the target troubleshooting node set W without a second weight value, substitute Qw = 0 into the formula = λPw + μQw for calculation;
[0043] Step S503: Set the node that generates abnormal data in the communication network as the abnormal data source node, arrange the risk values γ of each node in W from high to low, and give troubleshooting suggestions to the node management personnel according to the order from high to low of the risk values. After inspection by the corresponding node management personnel, feedback the abnormal data source node;
[0044] Step S504: Set the logical distance DL between each node in the communication network, where DL = Lmin + 1, Lmin represents the minimum number of nodes on the path between two nodes, and the logical distance DL between a node and itself is 0;
[0045] Step S505: If the management personnel do not detect a source node with an anomaly in the target troubleshooting node set, set the node with the longest logical distance from the target node as the new target node, enter Step S301 to start the loop, and calculate the target troubleshooting node set again.
[0046] To better implement the above method, a data anomaly transmission system based on artificial intelligence is also proposed, including: a historical operation data preprocessing module, a weight value calculation module, and a data transmission node management module. The historical operation data preprocessing module is used to collect the historical operation data of each data transmission node from the historical operation logs and calculate the feature vectors corresponding to each anomaly influence range. The weight value calculation module is used to calculate the first weight value and the second weight value. The data transmission node management module is used to evaluate the risk value of the nodes and troubleshoot and feedback the anomaly data source nodes.
[0047] Further, the historical operation data preprocessing module includes: an abnormal transmission record extraction unit, a node extraction unit for abnormal data, an abnormal transmission record classification unit, a core data extraction unit, a core data influence range calculation unit, an anomaly influence range calculation unit, and an anomaly influence range feature vector calculation unit. The abnormal transmission record extraction unit is used to extract abnormal transmission records. The node extraction unit for abnormal data is used to extract the nodes that generate abnormal data. The abnormal transmission record classification unit is used to classify the abnormal transmission records according to different abnormal events. The core data extraction unit is used to extract core data. The anomaly influence range calculation unit is used to calculate the anomaly influence range of the core data. The anomaly influence range feature vector calculation unit is used to calculate the feature vectors of the anomaly influence range.
[0048] Further, the weight value calculation module includes: an abnormal transmission state vector extraction unit, a vector cosine value calculation unit, a first associated node extraction unit, a first weight value calculation unit, a connection content data stream information extraction unit, an information joint entropy calculation unit, an information entropy separation unit, a source node influence threshold determination unit, and a second weight value calculation unit. The abnormal transmission state vector extraction unit is used to generate an abnormal transmission state vector. The vector cosine value calculation unit is used to calculate the cosine value of the angle between the abnormal transmission state vector and the feature vector of the first anomaly influence range. The first associated node extraction unit is used to extract the first associated nodes. The first weight value calculation unit is used to calculate the first weight value. The connection content data stream information extraction unit is used to extract the connection content data stream information of the target node within the T time period before the abnormal transmission data warning is generated. The information joint entropy calculation unit is used to calculate the information joint entropy of the connection content data stream information parameter items. The information entropy separation unit is used to calculate the separation information entropy of the sending parameter items of the sending data node. The source node influence threshold determination unit is used to determine the second associated node. The second weight value calculation unit is used to calculate the second weight value.
[0049] Further, the data transmission node management module includes: a target troubleshooting node extraction unit, a node risk value calculation unit, a node risk value sorting unit, a logical distance calculation unit, and an abnormal data source node output unit. The target troubleshooting node extraction unit is used to extract target troubleshooting nodes, the node risk value calculation unit is used to calculate the risk values of the target troubleshooting nodes, the node risk value sorting unit is used to sort the risk values of the target troubleshooting nodes, the logical distance calculation unit is used to calculate the logical distances between data transmission nodes, and the abnormal data source node output unit is used to output information about abnormal data source nodes.
[0050] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: By analyzing the impact of abnormal data on each node of the communication network and combining the information entropy component of the data flow information of the communication connection content of the target node before the abnormal data transmission, the present invention locates the nodes generating abnormal data directionally, provides reference information for managers to troubleshoot the nodes generating abnormal data, and improves the efficiency of managers in operating and maintaining the communication network. Brief Description of the Drawings
[0051] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0052] Figure 1 is a schematic structural diagram of a data transmission anomaly supervision system based on artificial intelligence according to the present invention;
[0053] Figure 2 is a schematic flow diagram of a data transmission anomaly supervision method based on artificial intelligence according to the present invention;
[0054] Figure 3 is a schematic diagram of the core data evaluation method of a data transmission anomaly supervision method based on artificial intelligence according to the present invention;
[0055] Figure 4 is a schematic diagram of a method for generating the influence range of an abnormal event in a data transmission anomaly supervision method based on artificial intelligence according to the present invention;
[0056] Figure 5 is a schematic diagram of another method for generating the influence range of an abnormal event in a data transmission anomaly supervision method based on artificial intelligence according to the present invention. Detailed Embodiments
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Please refer to Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 , the technical solutions provided by the present invention are as follows:
[0059] Step S100: Collect the historical operation records of each data transmission node from the historical operation logs of the communication network, record the historical operation records with abnormal data transmission as abnormal transmission records. Each abnormal transmission record includes several operation parameter items. After normalizing the values of each operation parameter item in the abnormal transmission record, record them in the abnormal transmission record set B.
[0060] Step S200: Classify the abnormal transmission records in the abnormal transmission record set B to obtain several abnormal data transmission types, and extract the characteristics of the influence range of each abnormal data transmission type to obtain the feature vector of the influence range of each abnormal data transmission type;
[0061] Among them, step S200 includes:
[0062] Step S201: Project the abnormal transmission record into a point in space. The number of operation parameter items in the abnormal transmission record corresponds to the space dimension, and calculate the Euclidean distance between the points corresponding to each abnormal transmission record.
[0063] Step S202: Select the core data in the abnormal transmission record, set the core data evaluation threshold α and the core data evaluation radius r1. In the space projected by the abnormal transmission record, use x to represent a certain abnormal transmission record. When the number of abnormal transmission records with a Euclidean distance less than r1 from x is greater than α, set x as the core data. All abnormal transmission records within the range with x as the center and r1 as the radius form an abnormal data transmission type.
[0064] Step S203: Set the core data influence radius r2, and set the range with the core data as the center and r2 as the radius as the core data influence range.
[0065] Refer to Figure 3 , a method for selecting the core data in the extracted abnormal transmission record is provided;
[0066] By setting different values for the influence radius of core data, different core data influence ranges can be delimited, and different core data management strategies can be deployed, making the management method for abnormal data transmission more flexible;
[0067] Step S204: From the operation and maintenance records of the communication network, extract the nodes that generate abnormal data detected after abnormal data transmission occurs. The nodes that generate abnormal data correspond to the historical operation logs. Extract the correspondence between the nodes that generate abnormal data and the abnormal transmission records, and extract one node corresponding to the core data or a combination of multiple nodes that generate abnormal data;
[0068] Step S205: Set a combination form of nodes of an abnormal data type in the historical operation log as an abnormal event, where one abnormal event includes k core data;
[0069] In the communication network, the abnormal data generated by one node or multiple nodes that generate abnormal data will affect several data transmission nodes. The several types of core data extracted represent the impact of one node or multiple nodes that generate abnormal data on the communication network;
[0070] Step S206: Calculate the abnormal event influence range D of the abnormal event, where D = C1 ∪ C2 ∪ C3 ∪ … ∪ C k , and use C1, C2, C3, … C k to represent the 1st, 2nd, 3rd, …, kth core data influence ranges included in the event influence range D respectively;
[0071] Reference Figure 4 , provides a method for generating an abnormal event influence range where r2 > r1;
[0072] Reference Figure 5 , provides a method for generating an abnormal event influence range where r2 < r1;
[0073] Step S207: Aggregate the abnormal influence ranges of all abnormal events, denoted as the first abnormal influence range, and calculate the eigenvectors of the abnormal event influence ranges in the first abnormal influence range.
[0074] Step S300: Detect the abnormal data transmission information in the current communication network, set the data transmission node where the abnormal transmission data is detected as the target node, collect the operation data of the target node, analyze the operation data, obtain the correspondence between the operation data of the target node and the influence range of the abnormal data transmission type, further obtain the first associated nodes of the target node, and set the first weight values of each first associated node;
[0075] Among them, step S300 includes:
[0076] Step S301: Set the data transmission node that currently detects abnormal transmission data as the target node, collect the operation data when abnormal data transmission occurs at the target node, and the collected operation data includes at least one operation parameter item;
[0077] Step S302: After normalizing the collected alarm operation data, set it as the abnormal transmission state vector;
[0078] Step S303: Project the abnormal transmission state vector and the feature vector of the first abnormal influence range into the same dimensional space, and calculate the included angle θ between the abnormal transmission state vector and the feature vector of the first abnormal influence range;
[0079] Step S304: Calculate the cosine function value g of the included angle between the abnormal transmission state vector and the feature vector of the first abnormal influence range cos ;
[0080] Step S305: Extract the values greater than 0 in g cos Let the number of values greater than 0 in g cos be u. According to the correspondence between abnormal events and a group of nodes that generate abnormal data, extract all u types of nodes that generate abnormal data corresponding to the abnormal events, and record them in the first associated node V1;
[0081] Step S306: Set the first weight value P of each node in the first associated node V1. The first weight value is the cosine value of the included angle between the feature vector of the first abnormal influence range where the node is located and the operation data vector.
[0082] Step S400: Extract the connection content data stream information within the T time period before the target node has abnormal data transmission, calculate the information entropy of the connection content data stream, set the node influence threshold, calculate the separation information entropy, and set the node corresponding to the separation information entropy as the second associated node. According to the information entropy size of communicating with the target node, set the second weight value of each second associated node:
[0083] Among them, Step S400 includes:
[0084] Step S401: Extract the connection content data stream information within the T time period before the target node generates an alarm for abnormal data transmission. The connection content data stream information includes the transmission path information of the connection content of the target node, and the connection content data stream information includes m parameter items, where m≥1;
[0085] Step S402: Set the nodes that send data to the target node within the T time period as the data sending nodes. Select one of the parameter items in the connection content data stream information parameter items, calculate the joint entropy of the connection content data stream information of the target node for this item and the information entropy of each data sending node, and use H(x) to represent the joint entropy of the information of the parameter item in the connection content data stream of the target node;
[0086] Step S403: Set the node influence threshold ω, and separate n separated information entropies whose information entropies are greater than the threshold ω, denoted as H(X1), H(X2), H(X3), ……, H(X n ) represents the separated information entropy of the first, second, third, ……, nth sending data node for sending the parameter item, and each separated information entropy corresponds to a separated data transmission node;
[0087] As a common algorithm for separated information entropy, Fast Independent Component Analysis (FastICA) estimates the number of blind sources, iteratively increases the number of random variable estimations. As the number increases, the separated signals become purer, and thus the average information entropy of the random variables becomes lower. The average information entropy is
[0088] When using the FastICA analysis algorithm, the method for setting the node influence threshold ω is: where represents the average information entropy of m random variables;
[0089] Step S404: Calculate the second weight influence value F of the data stream information parameter item of this connection content, where where F s represents the second weight influence value of the sth node among the n sending data nodes, H(X s ) represents the sth separated information entropy, minH(X) represents the minimum value among the n separated information entropies, and maxH(X) represents the maximum value among the n separated information entropies;
[0090] Step S405: Calculate the second weight influence value of the data stream information parameter item with connection content, collect the second weight influence values of each separated data transmission node, incorporate all nodes with second weight influence values into the second associated node V2, and normalize the second weight influence values of each node in the second associated node to obtain the second weight value Q of each second associated node.
[0091] Step S500: Collect all nodes in the first associated node and the second associated node to form a target troubleshooting node set, set a risk value evaluation function, calculate the node risk value, and give troubleshooting suggestions to the management staff. If the management staff does not detect a node that generates abnormal data in the target troubleshooting node set, then set the node with the longest logical distance from the target node as the new target node, return to step S300 for looping, and give new troubleshooting suggestions again;
[0092] Among them, step S500 includes:
[0093] Step S501: Assemble each node in the first associated node V1 and the second associated node V2 to form a target troubleshooting node set W, where W = V1 ∪ V2, and each target node in the target troubleshooting node set corresponds to a first weight value, a second weight value, or a combination of the first weight value and the second weight value;
[0094] Step S502: Set a first target troubleshooting coefficient λ and a second target troubleshooting coefficient μ, and calculate the risk value γ of each node in the target troubleshooting node set, where γ = λP w + μQ w , P w is the first weight value corresponding to a target troubleshooting node w in W, and Q w is the second weight value corresponding to w;
[0095] Step S503: Set the node that generates abnormal data in the communication network as the abnormal data source node, arrange the risk values γ of each node in W from high to low, and give troubleshooting suggestions to the node management personnel according to the order from high to low risk value. After inspection by the corresponding node management personnel, the abnormal data source node is fed back;
[0096] Step S504: Set the logical distance DL between each node in the communication network, where DL = Lmin + 1, Lmin represents the minimum number of nodes on the path between two nodes, and the logical distance DL between a node and itself is 0;
[0097] Step S505: If the management personnel do not detect the source node with abnormalities in the target troubleshooting node set, set the node with the longest logical distance from the target node as the new target node, enter Step S301 to start the loop, and calculate the target troubleshooting node set again.
[0098] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0099] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for supervising abnormal data transmission based on artificial intelligence, characterized in that, The method includes the following steps: Step S100: Collect the historical operation records of each data transmission node from the historical operation logs of the communication network, record the historical operation records with abnormal data transmission as abnormal transmission records. Each abnormal transmission record includes several operation parameter items. After normalizing the values of each operation parameter item in the abnormal transmission record, record them in the abnormal transmission record set B; Step S200: Classify the abnormal transmission records in the abnormal transmission record set B to obtain several abnormal data transmission types, extract the characteristics of the influence range of each abnormal data transmission type, and obtain the feature vectors of the influence range of each abnormal data transmission type; Step S300: Detect the abnormal data transmission information in the current communication network, set the data transmission node with detected abnormal transmission data as the target node, collect the operation data of the target node, analyze the operation data, obtain the correspondence between the operation data of the target node and the influence range of the abnormal data transmission type, further obtain the first associated nodes of the target node, and set the first weight values of each first associated node; Step S400: Extract the connection content data stream information within the T time period before the target node has abnormal transmission data, calculate the information entropy of the connection content data stream, set the node influence threshold, calculate the separation information entropy, set the node corresponding to the separation information entropy as the second associated node, and set the second weight values of each second associated node according to the information entropy size of communicating with the target node; Step S400 includes: Step S401: Extract the connection content data stream information within the T time period before the target node generates an abnormal transmission data warning. The connection content data stream information includes the transmission path information of the connection content of the target node, and the connection content data stream information includes m parameter items, where m≥1; Step S402: Set the nodes that send data to the target node within the T time period as the sending data nodes, select one of the parameter items in the connection content data stream information parameter items, calculate the joint entropy of the target node connection content data stream information of this item and the information entropy of each sending data node, and use H(x) to represent the joint entropy of the information of this parameter item in the target node connection content data stream; Step S403: Set the node influence threshold ω, and separate n separated information entropies whose information entropies are greater than the threshold ω, which are represented by H(X1), H(X2), H(X3), …… H(X n ) represents the separated information entropy of the first, the second, the third, ……, the nth sending data node for sending the parameter item, and each separated information entropy corresponds to a separated data transmission node; Step S404: Calculate the second weight influence value F of the connection content data stream information parameter item, where where F s represents the second weight influence value of the s-th node among the n sending data nodes, H(X s ) represents the s-th separation information entropy, minH(X) represents the minimum value among the n separation information entropies, and maxH(X) represents the maximum value among the n separation information entropies; Step S405: Calculate the second weight influence value of the connection content data stream information parameter item, collect the second weight influence values of each separated data transmission node, include all nodes with the second weight influence value in the second associated node V2, normalize the second weight influence values of each node in the second associated node, and obtain the second weight value Q of each second associated node; Step S500: Collect all the nodes in the first associated node and the second associated node to form the target troubleshooting node set, set the risk value evaluation function, calculate the node risk value, and give suggestions for the management to troubleshoot. If the management does not detect the node that generates abnormal data in the target troubleshooting node set, set the node with the longest logical distance from the target node as the new target node, return to step S300 for cycling, and give new troubleshooting suggestions again.
2. The method for supervising abnormal data transmission based on artificial intelligence according to claim 1, characterized in that: In step S200, the steps of classifying the abnormal transmission records in the abnormal transmission record set B include: Step S201: Project the abnormal transmission records into points in space. The number of operating parameter items in the abnormal transmission records corresponds to the spatial dimension, and calculate the Euclidean distance between the points corresponding to each abnormal transmission record; Step S202: Select core data from the abnormal transmission records. Set the core data evaluation threshold α and the core data evaluation radius r1. In the space projected by the abnormal transmission records, use x to represent a certain abnormal transmission record. When the number of abnormal transmission records with a Euclidean distance less than r1 from x is greater than α, set x as core data. All abnormal transmission records within the range with x as the center and r1 as the radius form an abnormal data transmission type.
3. The method for abnormal supervision of data transmission based on artificial intelligence according to claim 2, wherein: In step S200, the steps of calculating the feature vectors of the feature ranges of each type of operating data include: Step S203: Set the core data influence radius r2, and set the range with the core data as the center and r2 as the radius as the core data influence range; Step S204: From the operation and maintenance records of the communication network, extract the nodes that generate abnormal data detected after the occurrence of abnormal data transmission. The nodes that generate abnormal data correspond to the historical operation logs. Extract the correspondence between the nodes that generate abnormal data and the abnormal transmission records, and extract a combination of one or more nodes that generate abnormal data corresponding to the core data; Step S205: Set a combination form of nodes of an abnormal data in the historical operation log as an abnormal event, where an abnormal event includes k core data; Step S206: Calculate the abnormal event influence range D of the abnormal event, where D = C1 ∪ C2 ∪ C3 ∪ … ∪ C k , and use C1, C2, C3, … C k to represent the 1st, 2nd, 3rd, …, kth core data influence ranges included in the event influence range D respectively; Step S207: Aggregate the abnormal influence ranges of all abnormal events, denoted as the first abnormal influence range, and calculate the feature vectors of the influence ranges of each abnormal event in the first abnormal influence range.
4. The method for supervising abnormal data transmission based on artificial intelligence according to claim 3, characterized in that: Step S300 includes: Step S301: Set the data transmission node where the currently detected abnormal transmission data is located as the target node, and collect the operating data when the target node has abnormal data transmission. The collected operating data includes at least one operating parameter item; Step S302: After normalizing the collected alarm operating data, set it as the abnormal transmission state vector; Step S303: Project the abnormal transmission state vector and the feature vector of the first abnormal influence range into the same dimensional space, and calculate the included angle θ between the abnormal transmission state vector and the feature vector of the first abnormal influence range; Step S304: Calculate the cosine function value g of the included angle between the abnormal transmission state vector and the eigenvector of the first abnormal influence range cos ; Step S305: Extract g cos with values greater than 0, denoted as g cos There are u values greater than 0 in cos . According to the correspondence between abnormal events and a set of nodes generating abnormal data, extract all u types of nodes generating abnormal data corresponding to the abnormal events, and record them in the first associated node V1; Step S306: Set the first weight value P of each node in the first associated node V1. The first weight value is the cosine value of the included angle between the feature vector of the first abnormal influence range where the node is located and the operating data vector.
5. The method for supervising abnormal data transmission based on artificial intelligence according to claim 4, characterized in that: Step S500 includes: Step S501: Aggregate each node in the first associated node V1 and the second associated node V2 to form a target troubleshooting node set W, where W = V1 ∪ V2. Each target node in the target troubleshooting node set corresponds to a first weight value, a second weight value, or a combination of the first weight value and the second weight value; Step S502: Set the first target investigation coefficient λ and the second target investigation coefficient μ, and calculate the risk value γ of each node in the target investigation node set, where γ = λP w + μQ w , P w is the first weight value corresponding to a target investigation node w in W, and Q w is the second weight value corresponding to w; Step S503: Set the node that generates abnormal data in the communication network as the abnormal data source node. Arrange the risk values γ of each node in W from high to low. According to the order from high to low of the risk values, give troubleshooting suggestions to the node management personnel. After the corresponding node management personnel check, feedback the abnormal data source node; Step S504: Set the logical distance DL between each pair of nodes in the communication network, where DL = Lmin + 1, Lmin represents the minimum number of nodes on the path between two nodes, and the logical distance DL between a node and itself is 0; Step S505: If the source node with an anomaly is not detected by the management staff in the target troubleshooting node set, set the node with the longest logical distance from the target node as the new target node, enter Step S301 to start the loop, and calculate the target troubleshooting node set again.
6. A data transmission anomaly monitoring system for the data transmission anomaly monitoring method based on artificial intelligence according to any one of claims 1-5, characterized in that, The system includes the following modules: a historical operation data preprocessing module, a weight value calculation module, and a data transmission node management module. The historical operation data preprocessing module is used to collect the historical operation data of each data transmission node from the historical operation log and calculate the feature vectors corresponding to each anomaly influence range. The weight value calculation module is used to calculate the first weight value and the second weight value. The data transmission node management module is used to evaluate the risk value of the nodes and troubleshoot and feedback the anomaly data source nodes.
7. The data transmission anomaly supervision system according to claim 6, characterized in that: The historical operation data preprocessing module includes: an abnormal transmission record extraction unit, a node extraction unit for abnormal data, an abnormal transmission record classification unit, a core data extraction unit, a core data influence range calculation unit, an anomaly influence range calculation unit, and an anomaly influence range feature vector calculation unit. The abnormal transmission record extraction unit is used to extract abnormal transmission records. The node extraction unit for abnormal data is used to extract the nodes that generate abnormal data. The abnormal transmission record classification unit is used to classify the abnormal transmission records according to different abnormal events. The core data extraction unit is used to extract core data. The anomaly influence range calculation unit is used to calculate the anomaly influence range of the core data. The anomaly influence range feature vector calculation unit is used to calculate the feature vectors of the anomaly influence range.
8. The data transmission anomaly supervision system according to claim 6, characterized in that: The weight value calculation module includes: an abnormal transmission state vector extraction unit, a vector included angle cosine value calculation unit, a first associated node extraction unit, a first weight value calculation unit, a connection content data stream information extraction unit, an information joint entropy calculation unit, an information entropy separation unit, a source node influence threshold determination unit, and a second weight value calculation unit. The abnormal transmission state vector extraction unit is used to generate an abnormal transmission state vector. The vector included angle cosine value calculation unit is used to calculate the cosine value of the included angle between the abnormal transmission state vector and the feature vector of the first anomaly influence range. The first associated node extraction unit is used to extract the first associated nodes. The first weight value calculation unit is used to calculate the first weight value. The connection content data stream information extraction unit is used to extract the connection content data stream information of the target node within the T time period before the abnormal transmission data alarm is generated. The information joint entropy calculation unit is used to calculate the information joint entropy of the connection content data stream information parameter items. The information entropy separation unit is used to calculate the separation information entropy of the sending parameter items of the sending data node. The source node influence threshold determination unit is used to determine the second associated nodes. The second weight value calculation unit is used to calculate the second weight value.
9. The data transmission anomaly supervision system according to claim 6, characterized in that: The data transmission node management module includes: a target troubleshooting node extraction unit, a node risk value calculation unit, a node risk value sorting unit, a logical distance calculation unit, and an abnormal data source node output unit. The target troubleshooting node extraction unit is used to extract target troubleshooting nodes. The node risk value calculation unit is used to calculate the risk values of the target troubleshooting nodes. The node risk value sorting unit is used to sort the risk values of the target troubleshooting nodes. The logical distance calculation unit is used to calculate the logical distances between data transmission nodes. The abnormal data source node output unit is used to output information about abnormal data source nodes.
Citation Information
Patent Citations
Location method of network node fault
CN104993960A
Intelligent substation communication network anomaly detection method based on multi-dimensional entropy sequence classification
CN105515888A
Information entropy variance analysis-based abnormal traffic detection method
CN105847283A