A Deterministic Network Traffic Identification Method, System, Computer Device, and Medium
By analyzing and processing measurement and event-type packets in the power communication deterministic network, building a multivariate index time series, and combining the graph topology structure for abnormal scores and threshold settings, the problem of difficult to identify normal and abnormal traffic in the existing technology is solved, and more efficient and accurate abnormal traffic recognition is achieved, ensuring the safe and stable operation of the power communication network.
Patent Information
- Application Number
- CN202510489464.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art is difficult to efficiently and accurately identify normal flows and abnormal flows in power communication deterministic networks, resulting in abnormal flow false alarms and missed reports, which cannot meet the security and stable operation needs of power communication networks.
By reliably analyzing and processing measurement messages and event messages in the power communication deterministic network, a multivariate indicator time series is constructed, and a graph topology structure designed based on the correlation relationship between indicators is designed to perform service traffic abnormality scores and abnormal traffic identification threshold settings, so as to fully explore and utilize the semantic information, interactive logic and association relationships of various power communication messages.
It effectively improves the reliability and accuracy of network abnormal traffic identification, and can better ensure the high reliability and high security of the power communication network.
Smart Images

Figure CN120034396B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power communication, and particularly to a method, a system, a computer device and a medium for identifying deterministic network traffic. Background Art
[0002] With the increasing complexity of the power system and the continuous progress of power communication technology, the application of deterministic networks in the power system is becoming increasingly widespread. Especially in fields such as intelligent substations and industrial control, the wide application of standard protocols such as IEC 61850 provides important support for the efficient operation and intelligent management of the power system. However, the complexity and openness of the power communication network also bring new challenges. The destructiveness caused by system failures or malicious network attacks is increasing, seriously threatening the safe and stable operation of the power system.
[0003] In the deterministic network of power communication, abnormal traffic identification is one of the key technologies for monitoring system failures or malicious network attacks and ensuring network security. Currently, abnormal traffic identification technologies mainly rely on the feature analysis and pattern recognition of network traffic. However, due to the complexity and diversity of standard protocols such as IEC 61850, and the characteristics of power communication network traffic such as multi-source heterogeneity, small data length, periodicity, fixed data flow direction, strong time series, and uneven distribution of normal traffic and abnormal traffic, the traditional traffic anomaly detection methods based on simple thresholds, statistical features or machine learning cannot accurately and efficiently identify the normal traffic and abnormal traffic in the power communication network. For example, the rich semantic information and complex interaction logic contained in GOOSE (Generic Object Oriented Substation Event) messages and SV (Sampled Value) messages in the protocol cannot be fully exploited, resulting in insufficient detection accuracy, and it is very easy to have false alarms and missed alarms of abnormal traffic, making it difficult to meet the application requirements of the deterministic network of power communication. Therefore, there is an urgent need to provide a method that can efficiently and accurately identify abnormal traffic in the deterministic network of power communication, so as to provide reliable guarantee for the safe and stable operation of the power communication network. Summary of the Invention
[0004] The object of the present invention is to provide a method for identifying deterministic network traffic. By reliably analyzing and processing measurement messages and event messages in the deterministic network of power communication, a multi-variable index time series related to network traffic anomalies is constructed, and combined with a graph topology structure designed based on the correlation relationship between indexes for service traffic anomaly scoring and abnormal traffic identification threshold setting, it can fully explore and utilize the semantic information, interaction logic and correlation relationship of various power communication messages, effectively improve the reliability and accuracy of network abnormal traffic identification, have good engineering application value, and can effectively guarantee the high reliability and high security of the power communication network.
[0005] To achieve the above object, a method, a system, a computer device and a medium for identifying deterministic network traffic are provided.
[0006] In a first aspect, an embodiment of the present invention provides a method for identifying deterministic network traffic, and the method includes the following steps:
[0007] Obtain network traffic data to be detected according to a preset detection period; the network traffic data to be detected includes measurement message data and event message data;
[0008] Based on the principle that the source and destination addresses are the same, group the network traffic data to be detected to obtain traffic data to be analyzed with multiple groups of address pairs;
[0009] Process and analyze the traffic data to be analyzed for each address pair respectively to obtain corresponding multivariate index time series to be analyzed;
[0010] Based on a preset graph attention prediction network, analyze the multivariate index time series to be analyzed for each address pair respectively to obtain corresponding predicted multivariate index time series, and based on the error between the predicted multivariate index time series and the multivariate index time series to be analyzed, obtain corresponding multi-dimensional anomaly score series;
[0011] Based on the multi-dimensional anomaly score series of each address pair, generate corresponding service association topology graphs, and according to the service association topology graphs, analyze the multi-dimensional anomaly score series respectively to obtain corresponding service traffic anomaly score values;
[0012] Obtain the service traffic anomaly thresholds of each address pair, and based on the comparison result between the service traffic anomaly threshold and the service traffic anomaly score value, obtain the service traffic anomaly identification result.
[0013] Further, the step of processing and analyzing the traffic data to be analyzed for each address pair respectively to obtain corresponding multivariate index time series to be analyzed includes:
[0014] Parse each measurement message data and each event message data in the traffic data to be analyzed respectively to obtain corresponding measurement data sets and event data sets;
[0015] Sort the measurement data corresponding to the same measurement index in all measurement data sets of the traffic data to be analyzed in chronological order to generate corresponding measurement index time series;
[0016] Perform statistical analysis on each measurement index time series respectively to generate corresponding measurement index root mean square value time series and measurement index trapezoidal area sum time series;
[0017] Sort the event data corresponding to the same event metric in all event data sets of the traffic data to be analyzed in chronological order to generate a corresponding event metric time series;
[0018] Summarize all the measurement metric time series, all the root mean square value time series of measurement metrics, all the trapezoidal area sum time series of measurement metrics, and all the event metric time series corresponding to the traffic data to be analyzed to obtain a multi-metric time series;
[0019] Perform normalization processing on each metric series in the multi-metric time series respectively to generate the multi-metric time series to be analyzed.
[0020] Further, the step of summarizing all the measurement metric time series, all the root mean square value time series of measurement metrics, all the trapezoidal area sum time series of measurement metrics, and all the event metric time series corresponding to the traffic data to be analyzed to obtain a multi-metric time series includes:
[0021] Obtain the sequence minimum sampling time interval of each measurement metric time series, and use a preset integer multiple of the minimum value among all the sequence minimum sampling time intervals as the target sequence time scale;
[0022] Based on the target sequence time scale, perform sequence time scale conversion processing on each measurement metric time series, each root mean square value time series of measurement metrics, and each trapezoidal area sum time series of measurement metrics respectively according to the principle of linearly interpolating to fill in missing data to generate corresponding measurement metric time series to be processed, root mean square value time series of measurement metrics to be processed, and trapezoidal area sum time series of measurement metrics to be processed;
[0023] Based on the target sequence time scale, perform sequence time scale conversion processing on each event metric time series respectively according to the principle of filling in missing data with the previous moment's data to generate corresponding event metric time series to be processed;
[0024] Combine all the measurement metric time series to be processed, all the root mean square value time series of measurement metrics to be processed, all the trapezoidal area sum time series of measurement metrics to be processed, and all the event metric time series to be processed to obtain the multi-metric time series.
[0025] Further, the step of performing normalization processing on each metric series in the multi-metric time series respectively to generate the multi-metric time series to be analyzed includes:
[0026] Perform Z-score normalization on each measurement index time series, each root mean square value time series of measurement indexes, and each trapezoidal area sum time series of measurement indexes to generate corresponding normalized measurement index time series, normalized root mean square value time series of measurement indexes, and normalized trapezoidal area sum time series of measurement indexes;
[0027] Perform min-max normalization on each event index time series to generate corresponding normalized event index time series.
[0028] Further, the step of obtaining corresponding multi-dimensional anomaly score sequences based on the errors between the predicted multi-index time series and the multi-index time series to be analyzed includes:
[0029] Perform graph embedding processing on each subsequence in the multi-index time series to be analyzed to generate corresponding graph node representations, and obtain corresponding sequence correlation topological graphs based on all graph node representations corresponding to the multi-index time series to be analyzed;
[0030] Obtain the multi-index error time series corresponding to the predicted multi-index time series and the multi-index time series to be analyzed;
[0031] Respectively obtain the sequence median and interquartile range corresponding to each sub-error time series in the multi-index error time series, and perform normalization processing on the corresponding sub-error time series according to the sequence median and the interquartile range to obtain corresponding sub-error time series to be analyzed;
[0032] Perform scaling processing on the sub-error time series to be analyzed based on the preset sequence weights of each sub-error time series to be analyzed to obtain corresponding sub-anomaly score sequences;
[0033] According to the sequence correlation topological graph, obtain the corresponding sequence correlation degree matrix, and perform fusion processing on each sub-anomaly score sequence according to the sequence correlation degree matrix to generate the multi-dimensional anomaly score sequence.
[0034] Further, the step of analyzing each multi-dimensional anomaly score sequence according to the service correlation topological graph to obtain corresponding service traffic anomaly score values includes:
[0035] Obtain the corresponding service correlation degree matrix according to the service correlation topological graph;
[0036] According to the service correlation degree matrix, perform weighted fusion on the anomaly scores in each multi-dimensional anomaly score sequence to obtain the service traffic anomaly score value.
[0037] Further, the step of obtaining the abnormal threshold of the service traffic for each address pair and obtaining the service traffic anomaly recognition result based on the comparison result between the service traffic abnormal threshold and the service traffic abnormal score value includes:
[0038] Obtain the maximum normal score sequence of each address pair under normal service traffic; each element in the maximum normal score sequence is the maximum historical abnormal score of the subsequence index in the corresponding multi - index time series to be analyzed;
[0039] According to the service association degree matrix corresponding to the service association topology graph, perform weighted fusion on the abnormal scores in the maximum normal score sequence to obtain the corresponding service traffic abnormal threshold;
[0040] Judge whether the service traffic abnormal score value of each address pair is greater than the corresponding service traffic abnormal threshold;
[0041] If so, obtain that the service traffic anomaly recognition result of the corresponding address pair is service traffic anomaly, and obtain the maximum value of the abnormal scores in the multi - dimensional abnormal score sequence of the corresponding address pair. Locate and control the abnormal traffic according to the subsequence index in the multi - index time series to be analyzed corresponding to the maximum value of the abnormal score;
[0042] If not, obtain that the service traffic anomaly recognition result of the corresponding address pair is service traffic normal.
[0043] In a second aspect, an embodiment of the present invention provides a deterministic network traffic recognition system, and the system includes:
[0044] A data collection module, configured to obtain the network traffic data to be detected according to a preset detection period; the network traffic data to be detected includes measurement - type message data and event - type message data;
[0045] A data grouping module, configured to group the network traffic data to be detected based on the principle that the source and destination addresses are the same, and obtain the traffic data to be analyzed for multiple groups of address pairs;
[0046] A data pre - processing module, configured to process and analyze the traffic data to be analyzed for each address pair respectively to obtain the corresponding multi - index time series to be analyzed;
[0047] An index abnormal score module, configured to analyze the multi - index time series to be analyzed for each address pair respectively based on a preset graph attention prediction network to obtain the corresponding predicted multi - index time series, and obtain the corresponding multi - dimensional abnormal score sequence based on the error between the predicted multi - index time series and the multi - index time series to be analyzed;
[0048] A business exception scoring module, which is used to generate a corresponding business association topology graph based on the multi-dimensional exception scoring sequences of each address pair, and analyze each multi-dimensional exception scoring sequence according to the business association topology graph to obtain a corresponding business traffic exception scoring value;
[0049] An identification result generation module, which is used to obtain the business traffic exception threshold of each address pair, and obtain a business traffic exception identification result based on the comparison result between the business traffic exception threshold and the business traffic exception scoring value.
[0050] In a third aspect, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0051] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0052] The present invention provides a method, system, computer device, and medium for identifying deterministic network traffic. By means of the method, it is achieved that according to a preset detection period, network traffic data to be detected including measurement type message data and event type message data is obtained. After grouping the network traffic data to be detected based on the principle that the source and destination addresses are the same to obtain traffic data to be analyzed in multiple groups of address pairs, the traffic data to be analyzed for each address pair is respectively processed and analyzed to obtain corresponding multivariate index time series to be analyzed. Then, based on a preset graph attention prediction network, the multivariate index time series to be analyzed for each address pair is respectively analyzed to obtain corresponding predicted multivariate index time series. Based on the error between the predicted multivariate index time series and the multivariate index time series to be analyzed, corresponding multi-dimensional anomaly score sequences are obtained. And based on the multi-dimensional anomaly score sequences of each address pair, corresponding business association topology graphs are generated. According to the business association topology graphs, the multi-dimensional anomaly score sequences are respectively analyzed to obtain corresponding business traffic anomaly score values, and the business traffic anomaly thresholds for each address pair are obtained. Based on the comparison result between the business traffic anomaly threshold and the business traffic anomaly score value, a technical solution for obtaining a business traffic anomaly identification result is achieved. Compared with the prior art, this method for identifying deterministic network traffic can construct multivariate index time series related to network traffic anomalies through reliable analysis and processing of measurement type messages and event type messages in a power communication deterministic network, and combine a graph topology structure designed based on the association relationship between indicators to perform business traffic anomaly scoring and set anomaly traffic identification thresholds. It can fully exploit and utilize the semantic information, interaction logic, and association relationships of various power communication messages, effectively improve the reliability and accuracy of network anomaly traffic identification, has good engineering application value, and can effectively ensure the high reliability and high security of the power communication network. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic flowchart of the method for identifying deterministic network traffic in an embodiment of the present invention;
[0054] Figure 2 is a flowchart of the method for obtaining the multi-dimensional anomaly score sequence of a single address pair in an embodiment of the present invention;
[0055] Figure 3 is a schematic flowchart of the process for obtaining the business traffic anomaly identification result of a single address pair in an embodiment of the present invention;
[0056] Figure 4 is a schematic diagram of the training performance of the preset graph attention prediction network in an embodiment of the present invention;
[0057] Figure 5 is a schematic diagram of the effect of the design of the business traffic anomaly threshold in an embodiment of the present invention;
[0058] Figure 6It is a schematic diagram for comparing the effects of the traditional abnormal traffic identification method and the deterministic network traffic identification method proposed by the present invention in the business traffic anomaly identification in the substation scenario;
[0059] Figure 7 It is a schematic structural diagram of the deterministic network traffic identification system in an embodiment of the present invention;
[0060] Figure 8 It is an internal structure diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0061] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the following described embodiments are part of the embodiments of the present invention and are only used to illustrate the present invention, but not to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] The deterministic network traffic identification method provided by the present invention can be understood as a technical solution that can efficiently and accurately identify abnormal business traffic in a deterministic network of power communication, which is proposed based on the application status that the existing deterministic network abnormal traffic identification method in the power system cannot fully exploit the rich semantic information and complex interaction logic contained in various communication messages, and cannot effectively solve the impact of the imbalance between normal traffic and abnormal traffic on the reliability of traffic prediction, resulting in extremely easy false alarms and missed alarms of abnormal traffic. The following embodiments will elaborate on the deterministic network traffic identification method of the present invention.
[0063] In one embodiment, as Figure 1 shown, a deterministic network traffic identification method is provided, including the following steps:
[0064] S11. Obtain the network traffic data to be detected according to a preset detection period; the network traffic data to be detected includes measurement message data and event message data; wherein, the preset detection period can be understood as the time interval for performing network traffic anomaly detection determined based on actual application requirements; the network traffic data to be detected obtained regularly can be understood as communication message time series data facilitating the manifestation of the deterministic network traffic characteristics of power communication. Among them, the measurement message data can be understood as the SV (sampled value) message traffic data with continuous values in the IEC 61850 protocol, and the event message data can be understood as the power system event message data with discrete values used for exchange or reporting in the IEC 61850 protocol, such as GOOSE (Generic Object Oriented Substation Event) message traffic data. It should be noted that the sampling time interval for collecting various time series data within the preset detection period can be determined based on the shortest communication period of the corresponding message type in the actual power communication deterministic network. In principle, it can be set in units of milliseconds, seconds, minutes, etc. according to application requirements; the duration interval of the corresponding preset detection period can be set to dozens to hundreds of times the sampling time interval according to actual application requirements to ensure that the obtained network traffic data to be detected has enough data sample points to support subsequent abnormal traffic identification and analysis, which is not specifically limited here.
[0065] S12. Group the network traffic data to be detected based on the principle of the same source and destination addresses to obtain the traffic data to be analyzed for multiple pairs of addresses; wherein, the principle of the same source and destination addresses can be understood as a classification principle for different service traffic data designed considering the characteristics that SV messages and GOOSE messages in the power communication deterministic network are transmitted in a fixed direction and different source and destination addresses actually correspond to different services in the communication network; the traffic data to be analyzed for each pair of obtained addresses can be understood as the set of all measurement message data and all event message data with the same source and destination addresses.
[0066] S13. Process and analyze the traffic data to be analyzed for each pair of addresses respectively to obtain the corresponding multivariate index time series to be analyzed; wherein, the multivariate index time series to be analyzed can be understood as a time series set composed of multiple index time series generated by analyzing, screening, expanding, and combining the dataset content of various messages in the traffic data to be analyzed.
[0067] Specifically, the step of respectively processing and analyzing the traffic data to be analyzed for each pair of addresses to obtain the corresponding multivariate index time series to be analyzed includes:
[0068] Parse each measurement message data and each event message data in the traffic data to be analyzed respectively to obtain corresponding measurement data sets and event data sets. Among them, the measurement data set can be understood as a set of measurement index data carried by a measurement message. Specifically, the type of measurement index data carried by each measurement message can be determined according to actual application requirements and is not specifically limited here. Similarly, the event data set can be understood as a set of event index data reported by an event message. Specifically, the type of event index data carried by each event message can be determined according to actual application requirements and is not specifically limited here. It should be noted that the process of specifically parsing each measurement message data and event message data to obtain the measurement data set and event data set carried by the message can refer to the relevant technical implementation of parsing IEC61850 protocol messages to obtain message data content and will not be elaborated here.
[0069] Sort the measurement data corresponding to the same measurement index in all measurement data sets of the traffic data to be analyzed in chronological order to generate a corresponding measurement index time series. Among them, the measurement index time series can be understood as the time series data obtained by classifying and summarizing various measurement data in all parsed measurement data sets and sorting them according to the actual collected timestamps. If the measurement data set includes voltage values (three-phase) and current values (three-phase) at the same time, corresponding voltage value time series and current value time series can be constructed respectively according to the voltage values (three-phase) and current values (three-phase).
[0070] Perform statistical analysis on each measurement index time series respectively to generate corresponding measurement index root mean square value time series and measurement index trapezoidal area sum time series. Among them, the measurement index root mean square value time series and the measurement index trapezoidal area sum time series can be understood as being obtained by performing root mean square value (RmsValue) index analysis and trapezoidal area sum (TrapAreaSum) index analysis on each measurement index data respectively considering the stability of the power grid and the reliability of power supply quality, and are more representative and characteristic time series data for detecting faults in the power system transmission link. The specific process of obtaining the measurement index root mean square value (RmsValue) time series and the measurement index trapezoidal area sum time series can refer to the calculation technical implementation of the RmsValue index and the TrapAreaSum index in existing statistical analysis and will not be elaborated here.
[0071] Sort the event data corresponding to the same event metric in all event data sets of the traffic data to be analyzed in chronological order to generate a corresponding event metric time series; among them, the event metric time series can be understood as the time series data obtained by classifying and summarizing various event data in all parsed event data sets and sorting them according to the corresponding actual acquisition timestamps. If the event data set includes event data such as temperature alarms, circuit breaker status, and disconnector interlocks at the same time, corresponding temperature alarm time series, circuit breaker status time series, and disconnector interlock time series can be constructed according to events such as temperature alarms, circuit breaker status, and disconnector interlocks. It should be noted that in actual applications, when parsing the event data set for each event class message, data such as the corresponding flow ID and application ID that may change during a malicious attack can also be obtained to further enrich the types of event metric time series and further improve the comprehensiveness of event class message feature extraction and analysis.
[0072] Summarize all the measurement metric time series, all the root mean square value time series of the measurement metrics, all the trapezoidal area sum time series of the measurement metrics, and all the event metric time series corresponding to the traffic data to be analyzed to obtain a multi-metric time series; among them, the multi-metric time series can be understood as a set of multiple metric time series obtained by uniformly processing the time scales of all metric time series.
[0073] Considering that the event class message data is discrete rather than continuous in actual applications, in order to ensure the reliability of all metric sequence analysis, this embodiment preferably determines the sequence time scale of the multi-metric time series based on the sampling time interval of the measurement metric time series; specifically, the step of summarizing all the measurement metric time series, all the root mean square value time series of the measurement metrics, all the trapezoidal area sum time series of the measurement metrics, and all the event metric time series corresponding to the traffic data to be analyzed to obtain a multi-metric time series includes:
[0074] Obtain the minimum sampling duration interval of each measurement metric time series, and use a preset integer multiple of the minimum value among all the minimum sampling duration intervals as the target sequence time scale; among them, the target sequence time scale can be understood as considering the situation that the sampling point moments of each measurement metric time series may be uneven and the sampling frequencies among different measurement metrics are inconsistent in actual applications, and selecting the minimum adjacent data interval among all the measurement metric time series as the target sequence time scale required for generating the multi-metric time series; correspondingly, the preset integer multiple can be set according to actual application requirements and is not specifically limited here.
[0075] Based on the target sequence time scale, sequence time scale conversion processing is respectively performed on each measurement index time series, each root mean square value time series of measurement indexes, and each trapezoidal area sum time series of measurement indexes according to the principle of filling in missing data by linear interpolation, to generate corresponding measurement index time series to be processed, root mean square value time series of measurement indexes to be processed, and trapezoidal area sum time series of measurement indexes to be processed; wherein, the principle of filling in missing data by linear interpolation can be understood as that after screening or merging the data in each measurement index time series with an interval smaller than the target sequence time scale, for the positions of sampling moments where data is missing, the method of linear interpolation is used to complete the data filling.
[0076] Based on the target sequence time scale, sequence time scale conversion processing is respectively performed on each event index time series according to the principle of filling in missing data with the previous moment's data, to generate corresponding event index time series to be processed; wherein, the principle of filling in missing data with the previous moment's data can be understood as a design that takes into account the characteristics that the changes of event indexes are not frequent in practical applications and the event indexes can be considered to maintain the previous state continuously without update, and uses the most recently obtained data before to fill until new data appears at the sampling moment position corresponding to the target sequence time scale.
[0077] Through the above processing of each measurement index sequence and each event index sequence, corresponding time series with uniform sampling in the time dimension can be generated to ensure the reliability of subsequent individual analysis of each index time series and correlation analysis between index time series.
[0078] Combine all measurement index time series to be processed, all root mean square value time series of measurement indexes to be processed, all trapezoidal area sum time series of measurement indexes to be processed, and all event index time series to be processed to obtain the multi-index time series.
[0079] Normalize each index sequence in the multi-index time series respectively to generate the multi-index time series to be analyzed; wherein, the normalization processing can be understood as a time series standardization processing based on the distribution characteristics of different index data to avoid analysis anomalies caused by the dimensional differences of different index data; specifically, the step of normalizing each index sequence in the multi-index time series respectively to generate the multi-index time series to be analyzed includes:
[0080] Perform Z-score normalization on each measurement index time series, each root mean square value time series of measurement indexes, and each trapezoidal area sum time series of measurement indexes to generate corresponding normalized measurement index time series, normalized root mean square value time series of measurement indexes, and normalized trapezoidal area sum time series of measurement indexes. Among them, Z-score normalization can be understood as a processing method that, considering the characteristic that the measurement index values provided by continuous-valued measurement message data (such as SV message data) usually follow a Gaussian or Gaussian-like distribution, preferably eliminates the dimension and makes various measurement data comparable, expressed as:
[0081]
[0082] In the formula,
[0083]
[0084]
[0085] Among them, represents the number of samples of the measurement index time series ; and respectively represent the mean and variance of the measurement index time series ; represents the normalized measurement index time series obtained by performing Z-score normalization on the measurement index time series .
[0086] Perform min-max normalization on each event index time series to generate corresponding normalized event index time series. Among them, min-max normalization is Min-Max normalization. Considering that event message data (such as GOOSE message data) usually has discrete values for the included event states (indexes), for example, only -1, 0, or 1, which is quite different from the Gaussian distribution, it is preferably a processing method that scales the index values corresponding to all events to the range of [0, 1] and converts the data to the same dimension to eliminate the influence between different dimensions, expressed as:
[0087]
[0088] Among them, and respectively represent the minimum and maximum values of the event index time series ; represents the normalized event index time series obtained by performing min-max normalization on the event index time series .
[0089] S14. Analyze the multivariate indicator time series to be analyzed for each address pair based on a preset graph attention prediction network to obtain the corresponding predicted multivariate indicator time series, and obtain the corresponding multi-dimensional anomaly score series based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed. Among them, the preset graph attention prediction network can be understood as a prediction network based on graph attention, and a graph neural network that can meet the function of predicting and analyzing time series data existing in the art can be used. It is trained with a relevant data set constructed by the same method as used to generate the multivariate indicator time series to be analyzed. The network structure can adopt GDN (Graph Deviation Network) or a network structure improved therefrom, and specific limitations are not made here.
[0090] In this embodiment, the process of analyzing the multivariate indicator time series to be analyzed for each address pair based on the preset graph attention prediction network to obtain the corresponding predicted multivariate indicator time series is as Figure 2 shown and includes:
[0091] First, perform graph embedding on each indicator time series in the multivariate indicator time series to be analyzed to obtain the corresponding number of graph vertices. Suppose the multivariate indicator time series to be analyzed for each address pair includes indicator time series , where represents the number of sampling points. Each indicator time series can be graph-embedded (graph node representation) as , where represents the dimension of the graph embedding. In order to obtain rich data features as much as possible, generally take .
[0092] Then, construct the corresponding sequence association topology graph according to the obtained vertices, and use the Pearson correlation coefficient to form the edges of the graph to measure the correlation between the indicator sequences. The Pearson correlation coefficient usually has a better calculation effect when the data normality is insufficient, which can ensure the reliability of the topology graph construction in the power communication network scenario. Among them, the edge between vertices and can be expressed as:
[0093]
[0094] where and respectively represent the graph embeddings corresponding to the th and th vertices in the sequence association topology graph; represents the th and Edges of a vertex.
[0095] On this basis, a corresponding directed graph is constructed, and the adjacency matrix A is used to represent the directed graph, where A ij Indicates whether there is a directed edge from vertex i to vertex j, representing the dependency relationship between sequences; the corresponding A ij The calculation method is as follows:
[0096]
[0097] Among them, Represents the indicator function, Represents taking the indices of the top K values. K is a fixed value and can be determined according to specific application scenarios; that is, all vertices k to vertex i are sorted in descending order of the edge values between them, and it is considered that there is a directed edge from the vertices k corresponding to the top K values to vertex i. The corresponding , otherwise, the corresponding .
[0098] After obtaining the adjacency matrix (i.e., obtaining the graph topological structure), based on this adjacency matrix A and each index time series in the multi-variable index time series to be analyzed, each index time series can be predicted and analyzed through a preset graph attention prediction network to obtain the corresponding predicted index time series, and finally the required predicted multi-variable index time series is obtained. It should be noted that the detailed process of specifically predicting and analyzing the multi-variable index time series to be analyzed based on the preset graph attention prediction network can be implemented with reference to the relevant existing technologies of the specific network structure adopted, which will not be elaborated here.
[0099] Considering that in the actual power communication deterministic network, under normal traffic scenarios, the predicted values of the indicators obtained through prediction and analysis based on historical indicator time series should be within the normal value range in a statistical sense. If the deviation between the predicted value of an indicator at a certain time point and the actual value of the indicator is too large, it can be considered that there is a high possibility of abnormality in the actual value of the indicator at this time point. In this embodiment, preferably, the indicators are abnormally scored based on the prediction errors of each indicator time series, and then the abnormal score values of each address for the corresponding indicator sequences are obtained; specifically, the step of obtaining the corresponding multi-dimensional abnormal score sequence based on the error between the predicted multi-variable index time series and the multi-variable index time series to be analyzed includes:
[0100] Perform graph embedding processing on each subsequence in the multi-variable index time series to be analyzed respectively to generate corresponding graph node representations, and based on all the graph node representations corresponding to the multi-variable index time series to be analyzed, obtain the corresponding sequence association topological graph; among them, the process of obtaining the sequence association topological graph refers to the relevant description of the sequence association topological graph in the previous process of obtaining the predicted multi-variable index time series, which will not be repeated here.
[0101] Obtain the multivariate error time series corresponding to the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; wherein, the multivariate error time series can be understood as a combination of indicator error time series generated by taking the difference of the data at the same moment in each indicator time series in the multivariate indicator time series to be analyzed and the predicted indicator time series in the corresponding predicted multivariate indicator time series, and the element values in each indicator error time series can be expressed as:
[0102]
[0103] wherein, and respectively represent the indicator value and the indicator predicted value corresponding to the t-th moment in the i-th indicator time series in the multivariate indicator time series to be analyzed and the i-th predicted indicator time series in the corresponding predicted multivariate indicator time series; represents the element value (absolute error value) corresponding to the t-th moment in the i-th indicator error time series.
[0104] Respectively obtain the series median and the interquartile range corresponding to each sub-error time series in the multivariate error time series, and perform normalization processing on the corresponding sub-error time series according to the series median and the interquartile range to obtain the corresponding sub-error time series to be analyzed; wherein, the series median can be understood as the element value at the middle position of the time series, and the interquartile range can be understood as the range of the interquartile distance (IQR2) in the series; correspondingly, the element value in the sub-error time series to be analyzed can be expressed as:
[0105]
[0106] In the formula, and respectively represent the series median and the interquartile range of the i-th sub-error time series in the multivariate error time series; represents the normalized error value corresponding to This error normalization method can effectively prevent the situation where the deviation generated by a certain sub-error sequence dominates excessively and reduce the accuracy of subsequent comprehensive analysis of each sub-error sequence.
[0107] Performing scaling processing on the to-be-analyzed sub-error time series based on the preset sequence weights of each to-be-analyzed sub-error time series to obtain corresponding sub-anomaly score series; wherein, the preset sequence weights can be understood as the weights corresponding to the to-be-analyzed sub-error time series determined based on the likelihood of implicit anomaly information in each measurement index and event index in practical applications. For example, the preset sequence weight of the index involved in the measurement type message can be set to 1, and the preset sequence weight of the index involved in the event type message can be set to a value greater than 1 according to the priority of the event itself and application requirements, for amplifying or shrinking the element values in each to-be-analyzed sub-error time series. The element values in each corresponding sub-anomaly score series can be expressed as:
[0108]
[0109] wherein, represents the preset sequence weight of the i-th to-be-analyzed sub-error time series; represents the anomaly score value corresponding to the t-th moment in the i-th sub-anomaly score series.
[0110] In this embodiment, an index data anomaly scoring mechanism that takes the product of the normalized prediction error of each index and the corresponding preset sequence weight as the anomaly score value is designed based on the strong correlation between index anomalies and preset errors, as well as the contribution degree difference of various indexes in anomaly traffic detection. The larger the normalized prediction error of the index, the higher the index anomaly score, and the higher the weight value corresponding to the index sequence, the higher its index anomaly score. This can ensure that the index anomaly score is consistent with the intuitive judgment in practical applications and helps to quickly discover real and important anomaly events.
[0111] According to the sequence association topology graph, obtaining the corresponding sequence association degree matrix, and performing fusion processing on each sub-anomaly score series according to the sequence association degree matrix to generate the multi-dimensional anomaly score series; wherein, the sequence association degree matrix can be obtained based on the adjacency matrix A of the corresponding sequence association topology graph. The non-diagonal elements of this matrix are 0, and the diagonal element d ii is calculated by the following method:
[0112]
[0113] wherein, represents the diagonal element value of the i-th row and i-th column in the sequence association degree matrix D; represents the element value of the i-th row and j-th column in the adjacency matrix A of the sequence association topology graph, and the acquisition process of the adjacency matrix A of the sequence association topology graph can refer to the relevant description in the previous text about obtaining the corresponding predicted index time series by predicting and analyzing the index time series through the preset graph attention prediction network, which will not be elaborated here.
[0114] Considering that various indicators in the actual power communication deterministic network are correlated, and there must also be associations between the corresponding sub - anomaly score sequences. If we directly consider the sub - anomaly score sequences corresponding to each indicator equally without any discrimination as in traditional anomaly traffic scoring, it will inevitably break the relevance and importance differences between the sub - anomaly score sequences, thus deviating from the significance of joint analysis using multiple indicator features simultaneously, and will surely lead to a deviation in anomaly scoring. To improve the reliability of traffic anomaly scoring at each indicator level as much as possible, in this embodiment, preferably, the graph - structure information corresponding to the multivariate time series is used to discover the joint correlation between subsequences, and the importance of the impact of indicators on traffic anomalies is used as the sequence weight to discover the importance differences between subsequences, so that the anomaly scoring is based on more accurate prior features of multivariate time series, thereby achieving more accurate anomaly scoring.
[0115] The multi - dimensional anomaly score sequence in this embodiment can be understood as a multi - indicator anomaly score sequence composed of indicator anomaly score values generated by comprehensively analyzing the anomaly score values at all times in the sub - anomaly score sequences based on a certain address pair. The positions of the indicator anomaly score values in this sequence can be determined according to the pre - set sequence numbers, which are not specifically limited here. Specifically, the element value in the multi - dimensional anomaly score sequence corresponding to a certain address pair (a certain source address and destination address pair) is expressed as:
[0116]
[0117]
[0118] In the formula, represents the i - th sub - anomaly score sequence of the m - th address pair, , represents the total number of address pairs; represents the anomaly score value corresponding to the t - th moment in the i - th sub - anomaly score sequence of the m - th address pair; represents the transpose of the vector; represents the length of the sub - anomaly score sequence; D and respectively represent the sequence correlation degree matrix and the inverse matrix of the sequence correlation degree matrix corresponding to the current sequence association topology graph; represents the trace of the matrix, which is used for normalization; represents the multi - dimensional anomaly score sequence of the m - th address pair The i - th element value (the anomaly score corresponding to the i - th sub - anomaly score sequence) in
[0119] In this embodiment, by comprehensively considering the impact of message types and message correlations on traffic anomalies, and strengthening the role of relatively independent message data in anomaly scoring through the degree matrix, the reliability of anomaly scoring for various data metrics of different address pairs (different services) is effectively guaranteed, providing reliable data support for accurate anomaly scoring of various service traffic in the future.
[0120] S15. Generate a corresponding service association topology graph based on the multi-dimensional anomaly scoring sequences of each address pair, and analyze each multi-dimensional anomaly scoring sequence respectively according to the service association topology graph to obtain the corresponding service traffic anomaly scoring value; among them, after the multi-dimensional anomaly scoring sequences of each address pair are obtained through the above method steps, a service association topology graph can be constructed based on the correlation between the corresponding communication services of the address pair, and the service traffic of each address pair can be scored for anomalies based on the obtained service association topology graph.
[0121] The construction process of the service association topology graph in this embodiment includes: according to the multi-dimensional anomaly scoring sequences of the address pairs , model to obtain a service association topology graph containing M vertices, and define the edge between vertex l and h through the Pearson correlation coefficient, and its expression of the edge is as follows:
[0122]
[0123] On this basis, a directed graph is formed, and the adjacency matrix W is used to represent the directed graph, where represents whether there is a directed edge from vertex l to vertex m, indicating the dependency relationship between multi-dimensional anomaly scoring sequences; The calculation method is as follows:
[0124]
[0125] In the formula, represents the indicator function, represents the index of taking out the first K values, and K is a fixed value, which can be determined according to the specific application scenario; that is, all vertices h to vertex m are sorted in descending order of the edge values between vertex h and vertex m, and it is considered that there is a directed edge from the vertex h corresponding to the first K values to vertex m, and the corresponding , otherwise, the corresponding .
[0126] After obtaining the service - related topological graph of communication services for all source - destination address pairs in the power communication deterministic network, the multi - dimensional anomaly score sequences of each address pair can be comprehensively analyzed at the service level based on the graph topology, and finally the corresponding service flow anomaly score value can be obtained. Specifically, the steps of analyzing each multi - dimensional anomaly score sequence according to the service - related topological graph to obtain the corresponding service flow anomaly score value include:
[0127] According to the service - related topological graph, obtain the corresponding service correlation degree matrix; among them, the service correlation degree matrix G can be obtained based on the adjacency matrix W of the service - related topological graph, and its non - diagonal elements are 0, and the diagonal element g ii is calculated by the following method:
[0128]
[0129] In the formula, represents the value of the diagonal element in the m - th row and m - th column of the service correlation degree matrix G; represents the value of the element in the l - th row and n - th column of the adjacency matrix W of the service - related topological graph.
[0130] According to the service correlation degree matrix, weight - fuse the anomaly scores in each multi - dimensional anomaly score sequence to obtain the service flow anomaly score value; among them, the service flow anomaly score value is expressed as:
[0131]
[0132] Among them, represents the multi - dimensional anomaly score sequence of the m - th address pair; respectively represent the service flow anomaly score values of the m - th address pair, which are actually the comprehensive service anomaly score obtained by the weighted sum of each anomaly score in. Since the influence of service - to - service correlation on the comprehensive score is incorporated, it is more conducive to reflecting the abnormal state of the network system.
[0133] S16. Obtain the service flow anomaly thresholds of each address pair, and based on the comparison result between the service flow anomaly threshold and the service flow anomaly score value, obtain the service flow anomaly recognition result; among them, the service flow anomaly thresholds of each address pair can, in principle, be set as fixed values according to actual experience. However, considering the dynamic variability of service flows in actual applications, in order to ensure that the abnormal flow recognition can adapt to the changes in the flow environment and continuously maintain the accuracy and reliability of the service flow anomaly recognition result, in this embodiment, preferably, the latest service flow anomaly thresholds are dynamically obtained within each detection period and used for the abnormal flow judgment of the corresponding services.
[0134] Specifically, the steps of obtaining the business traffic anomaly thresholds for each address pair and obtaining the business traffic anomaly recognition result based on the comparison result between the business traffic anomaly threshold and the business traffic anomaly score value include:
[0135] Obtain the maximum normal score sequence for each address pair under normal business traffic conditions; each element in the maximum normal score sequence is the maximum historical anomaly score of the subsequence index in the corresponding multi - index time series to be analyzed, and the maximum historical anomaly score can be understood as the maximum index anomaly score value among the index anomaly score values of the same index in the multi - dimensional anomaly score sequences corresponding to multiple historical detection periods before the current detection period; that is, after determining the maximum historical anomaly score corresponding to each index of each address pair under normal traffic conditions, the maximum normal score sequence (vector) can be combined. It should be noted that in practical applications, the acquisition method of the maximum normal score sequence for each address pair can also be adjusted or replaced according to the actual application scenario requirements.
[0136] According to the business association degree matrix corresponding to the business association topology graph, perform weighted fusion on the anomaly scores in the maximum normal score sequence to obtain the corresponding business traffic anomaly threshold; among them, the business traffic anomaly threshold is actually a dynamic threshold with a preset detection period as the change period, that is, the threshold obtained for a certain address pair in a certain detection period can be understood as the optimal threshold applicable to the current detection period, expressed as:
[0137]
[0138] In the formula, represents the maximum normal score sequence of the m - th address pair; represents the inverse matrix of the business association degree matrix G; represents the business traffic anomaly threshold of the m - th address pair.
[0139] Judge whether the business traffic anomaly score value of each address pair is greater than the corresponding business traffic anomaly threshold.
[0140] If so, the recognition result of the service traffic anomaly for the corresponding address pair is that the service traffic is abnormal, and the maximum anomaly score in the multi-dimensional anomaly score sequence corresponding to the address pair is obtained. Based on the subsequence index in the multi-variate index time series to be analyzed corresponding to the maximum anomaly score, anomaly traffic positioning and control are performed. Among them, the maximum anomaly score can be understood as the maximum value among the anomaly scores corresponding to each index in the multi-dimensional anomaly score sequence when the service traffic is abnormal. That is, finding the maximum anomaly score is equivalent to determining the anomaly index and knowing the source and destination address information of the anomaly. Based on the anomaly index and the corresponding source and destination address information, the network anomaly location can be further investigated, and the corresponding anomaly control strategy can be executed according to the obtained network anomaly location.
[0141] If not, the recognition result of the service traffic anomaly for the corresponding address pair is that the service traffic is normal. It should be noted that when determining that the service traffic is normal, the corresponding maximum normal score sequence can also be updated synchronously to provide data support for the service traffic anomaly detection in subsequent detection cycles.
[0142] As provided in the embodiments of the present invention Figure 3 As shown, according to a preset detection period, the network traffic data to be detected including measurement message data and event message data is obtained. After grouping the network traffic data to be detected based on the principle that the source and destination addresses are the same to obtain the traffic data to be analyzed for multiple address pairs, the traffic data to be analyzed for each address pair is processed and analyzed respectively to obtain the corresponding multi-variate index time series to be analyzed. Based on the preset graph attention prediction network, the multi-variate index time series to be analyzed for each address pair is analyzed respectively to obtain the corresponding predicted multi-variate index time series. Based on the error between the predicted multi-variate index time series and the multi-variate index time series to be analyzed, the corresponding multi-dimensional anomaly score sequence is obtained. And based on the multi-dimensional anomaly score sequence of each address pair, the corresponding service association topology graph is generated. According to the service association topology graph, each multi-dimensional anomaly score sequence is analyzed respectively to obtain the corresponding service traffic anomaly score value, and the service traffic anomaly threshold of each address pair is obtained. Based on the comparison result between the service traffic anomaly threshold and the service traffic anomaly score value, the deterministic network traffic anomaly recognition solution for the service traffic anomaly recognition result is obtained. By reliably analyzing and processing the measurement messages and event messages in the power communication deterministic network, the multi-variate index time series related to the network traffic anomaly is constructed, and combined with the graph topology structure designed based on the association relationship between the indexes, the service traffic anomaly scoring and the anomaly traffic recognition threshold setting are carried out, which can fully exploit and utilize the semantic information, interaction logic and association relationship of various power communication messages, effectively improve the reliability and accuracy of network anomaly traffic recognition, have good engineering application value, and can effectively ensure the high reliability and high security of the power communication network.
[0143] To verify the effectiveness of the method of the present invention, this embodiment also tested the abnormal traffic recognition in the substation scenario, and the test results are as Figures 4-6 shown as follows:
[0144] Figure 4 It shows the training performance of the preset graph attention prediction network constructed by training the dataset constructed by the multi-index time series generation method provided by the method of the present invention. It is easy to know that when the method of the present invention is trained with anomaly-free data, the training loss can converge to less than 5% after 40 iterations, that is, it is easy to perform rapid training and deployment for specific power communication network usage scenarios.
[0145] Figure 5 It shows the change of the comprehensive performance F1-Score (F1-Score) of the service traffic anomaly recognition with the set threshold when the source address and the destination address remain unchanged. It can be seen that as the threshold changes, there is an optimal value for the F1-Score, which further reflects the necessity and reliability of the dynamic generation of the service traffic anomaly threshold design provided by the method of the present invention.
[0146] Figure 6 It shows a schematic diagram of the comparison of the service traffic anomaly recognition effects of the traditional abnormal traffic recognition method and the deterministic network traffic recognition method proposed by the present invention applied to the substation scenario. It can be seen that through the service traffic anomaly scoring method of the present invention, the correct detection probability (true positive, true negative) is significantly improved; at the same time, in the power communication network, the harm of missed detection is much higher than that of false alarms, and the missed detection rate of the method of the present invention in the test is only about 0.2%, which has good engineering application value and can effectively ensure the high-reliability and high-security stable operation of the power communication network.
[0147] It should be noted that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders.
[0148] In one embodiment, as Figure 7 shown, a deterministic network traffic recognition system is provided, and the system includes:
[0149] A data acquisition module 1, configured to obtain the network traffic data to be detected according to a preset detection period; the network traffic data to be detected includes measurement type message data and event type message data;
[0150] A data grouping module 2, configured to group the network traffic data to be detected based on the principle that the source and destination addresses are the same, so as to obtain the traffic data to be analyzed for multiple groups of address pairs;
[0151] The data preprocessing module 3 is used to process and analyze the traffic data to be analyzed for each address pair respectively, so as to obtain the corresponding multivariate index time series to be analyzed;
[0152] The index anomaly scoring module 4 is used to analyze the multivariate index time series to be analyzed for each address pair respectively based on a preset graph attention prediction network, so as to obtain the corresponding predicted multivariate index time series, and obtain the corresponding multi-dimensional anomaly scoring sequence based on the error between the predicted multivariate index time series and the multivariate index time series to be analyzed;
[0153] The service anomaly scoring module 5 is used to generate a corresponding service association topology graph based on the multi-dimensional anomaly scoring sequences of each address pair, and analyze each multi-dimensional anomaly scoring sequence respectively according to the service association topology graph, so as to obtain the corresponding service traffic anomaly scoring value;
[0154] The recognition result generation module 6 is used to obtain the service traffic anomaly threshold of each address pair, and obtain the service traffic anomaly recognition result based on the comparison result between the service traffic anomaly threshold and the service traffic anomaly scoring value.
[0155] For the specific limitations of the deterministic network traffic recognition system, reference can be made to the limitations of the deterministic network traffic recognition method in the above text, and the corresponding technical effects can also be equivalently obtained, which will not be elaborated here. Each module in the above deterministic network traffic recognition system can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0156] Figure 8 The internal structure diagram of a computer device in an embodiment is shown. The computer device can specifically be a terminal or a server. As Figure 8As shown in the figure, the computer device includes a processor, a memory, a network interface, a display, a camera, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for identifying deterministic network traffic. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0157] Those of ordinary skill in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computing device may include more or fewer components than those shown in the figure, or combine some components, or have an equivalent component arrangement.
[0158] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above method.
[0159] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the above method.
[0160] In summary, the deterministic network traffic identification method and system provided by the embodiments of the present invention realize obtaining the network traffic data to be detected including measurement message data and event message data according to a preset detection period, grouping the network traffic data to be detected based on the principle of the same source and destination addresses to obtain the traffic data to be analyzed for multiple address pairs, then respectively processing and analyzing the traffic data to be analyzed for each address pair to obtain the corresponding multivariate index time series to be analyzed, and analyzing the multivariate index time series to be analyzed for each address pair based on a preset graph attention prediction network to obtain the corresponding predicted multivariate index time series, obtaining the corresponding multi-dimensional anomaly score series based on the error between the predicted multivariate index time series and the multivariate index time series to be analyzed, and generating the corresponding service association topology graph based on the multi-dimensional anomaly score series for each address pair, analyzing each multi-dimensional anomaly score series based on the service association topology graph to obtain the corresponding service traffic anomaly score value, and obtaining the service traffic anomaly threshold for each address pair, and obtaining the service traffic anomaly identification result based on the comparison result between the service traffic anomaly threshold and the service traffic anomaly score value. This method constructs a multivariate index time series related to network traffic anomalies through reliable analysis and processing of measurement messages and event messages in the power communication deterministic network, and combines the graph topology structure designed based on the association relationship between the indexes to perform service traffic anomaly scoring and set the abnormal traffic identification threshold, which can fully exploit and utilize the semantic information, interaction logic and association relationship of various power communication messages, effectively improve the reliability and accuracy of network abnormal traffic identification, has good engineering application value, and can effectively ensure the high reliability and high security of the power communication network.
[0161] Each embodiment in this specification is described in a progressive manner. For parts that are the same or similar in each embodiment, they can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0162] The above-described embodiments merely represent several preferred embodiments of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the protection scope of the claims described above.
Claims
1. A deterministic network traffic identification method, characterized in that: Applied to an electric power system, the method comprises the following steps: According to a preset detection cycle, the network traffic data to be detected is obtained; the network traffic data to be detected includes measurement message data and event message data; Based on the principle of the same source and destination addresses, the network traffic data to be detected is grouped to obtain traffic data to be analyzed of multiple groups of address pairs; The traffic data to be analyzed of each address pair is processed and analyzed respectively to obtain the corresponding multivariate indicator time series to be analyzed; Based on the preset graph attention prediction network, the time series of the multivariate indicators to be analyzed of each address pair are analyzed respectively to obtain the corresponding predicted time series of the multivariate indicators, and based on the error between the predicted time series of the multivariate indicators and the time series of the multivariate indicators to be analyzed, the corresponding multidimensional anomaly score sequence is obtained; Based on the multidimensional anomaly score sequence of each address pair, a corresponding business association topology map is generated, and according to the business association topology map, each multidimensional anomaly score sequence is analyzed respectively to obtain a corresponding business flow anomaly score value; The business traffic anomaly threshold of each address pair is obtained, and based on the comparison result of the business traffic anomaly threshold and the business traffic anomaly score value, a business traffic anomaly identification result is obtained.
2. The deterministic network traffic identification method according to claim 1, characterized in that: The step of processing and analyzing the traffic data to be analyzed of each address pair to obtain the corresponding multivariate indicator time series to be analyzed comprises: Respectively analyzing each measurement-type message data and each event-type message data in the traffic data to be analyzed to obtain corresponding measurement data sets and event data sets; Sorting the measurement data corresponding to the same measurement indicator in all the measurement data sets of the flow data to be analyzed in chronological order to generate a corresponding measurement indicator time series; Perform statistical analysis on each measurement index time series respectively to generate the corresponding measurement index root mean square value time series and measurement index trapezoidal area and time series; Sorting the event data corresponding to the same event indicator in all event data sets of the traffic data to be analyzed in chronological order to generate a corresponding event indicator time series; Summarize all measurement index time series, all measurement index root mean square value time series, all measurement index trapezoidal area and time series, and all event index time series corresponding to the traffic data to be analyzed to obtain a multivariate index time series; Each indicator series in the multivariate indicator time series is normalized to generate the multivariate indicator time series to be analyzed.
3. The deterministic network traffic identification method according to claim 2, characterized in that: The step of aggregating all measurement index time series, all measurement index root mean square value time series, all measurement index trapezoidal area and time series, and all event index time series corresponding to the traffic data to be analyzed to obtain a multivariate index time series comprises: Obtain the minimum sampling time interval of each measurement indicator time series, and use the preset integer multiple of the minimum value of all the minimum sampling time intervals as the target series time scale; According to the target sequence time scale, based on the principle of linear interpolation to fill in missing data, each measurement indicator time series, each measurement indicator root mean square value time series and each measurement indicator trapezoidal area and time series are respectively converted to a sequence time scale, and the corresponding measurement indicator time series to be processed, the measurement indicator root mean square value time series to be processed and the measurement indicator trapezoidal area and time series to be processed are generated; According to the target sequence time scale, based on the principle of supplementing missing data with the previous moment data, each event indicator time series is converted into a sequence time scale to generate a corresponding event indicator time series to be processed; The multivariate indicator time series is obtained by combining all the measurement indicator time series to be processed, all the root mean square value time series to be processed, all the trapezoidal area and time series to be processed, and all the event indicator time series to be processed.
4. The deterministic network traffic identification method according to claim 2, characterized in that: The step of normalizing each indicator sequence in the multivariate indicator time series to generate the multivariate indicator time series to be analyzed comprises: Perform Z-score normalization processing on each measurement indicator time series, each measurement indicator root mean square value time series, and each measurement indicator trapezoidal area and time series, respectively, to generate corresponding normalized measurement indicator time series, normalized measurement indicator root mean square value time series, and normalized measurement indicator trapezoidal area and time series; Perform minimum and maximum normalization processing on each event indicator time series respectively to generate the corresponding normalized event indicator time series.
5. The deterministic network traffic identification method according to claim 1, characterized in that: The step of obtaining a corresponding multidimensional anomaly score sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed comprises: Performing graph embedding processing on each subsequence in the multivariate indicator time series to be analyzed respectively, generating corresponding graph node representations, and obtaining corresponding sequence association topology graphs based on all graph node representations corresponding to the multivariate indicator time series to be analyzed; Obtaining the multivariate error time series corresponding to the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; Respectively obtain the sequence median and interquartile range corresponding to each sub-error time series in the multivariate error time series, and perform normalization processing on the corresponding sub-error time series according to the sequence median and the interquartile range to obtain the corresponding sub-error time series to be analyzed; Scaling the sub-error time series to be analyzed based on the preset sequence weights of each sub-error time series to be analyzed to obtain a corresponding sub-anomaly score sequence; According to the sequence association topology graph, a corresponding sequence association matrix is obtained, and each sub-anomaly score sequence is fused according to the sequence association matrix to generate the multi-dimensional anomaly score sequence.
6. The deterministic network traffic identification method according to claim 1, characterized in that: The step of analyzing each multi-dimensional anomaly score sequence according to the business association topology diagram to obtain a corresponding business flow anomaly score value comprises: According to the business association topology diagram, obtaining a corresponding business association matrix; According to the business correlation matrix, the anomaly scores in each multi-dimensional anomaly score sequence are weighted and fused to obtain the business flow anomaly score value.
7. The deterministic network traffic identification method according to claim 1, characterized in that: The step of obtaining the abnormal service flow threshold of each address pair and obtaining the abnormal service flow identification result based on the comparison result of the abnormal service flow threshold and the abnormal service flow score value comprises: Obtain the maximum normal score sequence of each address pair when the business traffic is normal; each element in the maximum normal score sequence is the maximum historical abnormal score of the subsequence indicator in the multivariate indicator time series to be analyzed; According to the business association degree matrix corresponding to the business association topology diagram, weighted fusion is performed on the abnormal scores in the maximum normal score sequence to obtain the corresponding business flow abnormal threshold; Determine whether the service flow anomaly score value of each address pair is greater than the corresponding service flow anomaly threshold value; If so, the business traffic anomaly identification result of the corresponding address pair is obtained as business traffic anomaly, and the maximum value of the anomaly score in the multidimensional anomaly score sequence of the corresponding address pair is obtained, and the abnormal traffic is located and controlled according to the subsequence index in the multivariate index time series to be analyzed corresponding to the maximum value of the anomaly score; If not, the business traffic anomaly identification result of the corresponding address pair is that the business traffic is normal.
8. A deterministic network traffic identification system, characterized in that: Applied to a power system, the system comprises: A data acquisition module is used to obtain the network traffic data to be detected according to a preset detection cycle; the network traffic data to be detected includes measurement message data and event message data; A data grouping module, used for grouping the network flow data to be detected based on the same source and destination address principle, to obtain flow data to be analyzed of multiple groups of address pairs; The data preprocessing module is used to process and analyze the traffic data to be analyzed for each address pair to obtain the corresponding multivariate indicator time series to be analyzed; An indicator anomaly scoring module is used to analyze the time series of the multivariate indicators to be analyzed of each address pair based on a preset graph attention prediction network to obtain a corresponding predicted multivariate indicator time series, and obtain a corresponding multidimensional anomaly scoring sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; A business anomaly scoring module is used to generate a corresponding business association topology map based on the multidimensional anomaly scoring sequence of each address pair, and analyze each multidimensional anomaly scoring sequence according to the business association topology map to obtain a corresponding business flow anomaly scoring value; The identification result generation module is used to obtain the business flow anomaly threshold value of each address pair, and obtain the business flow anomaly identification result based on the comparison result of the business flow anomaly threshold value and the business flow anomaly score value.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Network traffic identification method and system
CN117113262A
Malicious traffic identification system and method based on graph neural network and stable learning thought, program, equipment and storage medium
CN118449729A