Deterministic network traffic identification method and system, computer equipment and medium

By analyzing and processing measurement and event-type packets in the power communication deterministic network, building a multivariate index time series, and combining the graph topology structure to perform abnormal scores and threshold settings, the problem of identifying normal and abnormal traffic in the power communication network is solved, the reliability and accuracy of the identification are improved, and the high reliability and security of the network are guaranteed.

CN120034396AActive Publication Date: 2025-05-23STATE GRID ZHEJIANG ELECTRIC POWER CO LTD

Patent Information

Application Number
CN202510489464.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately identify normal flows and abnormal flows in power communication deterministic networks, resulting in abnormal flow false alarms and missed reports, which cannot meet the security and stable operation requirements of power communication deterministic networks.

Method used

By reliably analyzing and processing measurement messages and event messages in the power communication deterministic network, a multivariate indicator time series is constructed, and a graph topology structure designed based on the correlation relationship between indicators is designed to perform service traffic abnormality scores and abnormal traffic identification threshold settings, so as to fully explore and utilize the semantic information, interactive logic and association relationships of various power communication messages.

Benefits of technology

It improves the reliability and accuracy of network abnormal traffic identification, and can effectively ensure the high reliability and high security of the power communication network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034396A_ABST
    Figure CN120034396A_ABST
Patent Text Reader

Abstract

The invention provides a deterministic network traffic identification method and system, computer equipment and a medium, and the method comprises the steps: carrying out the grouping of obtained to-be-detected network traffic data based on a principle that a source address and a destination address are the same, and obtaining to-be-analyzed traffic data of a plurality of groups of address pairs; based on a preset graph attention prediction network, performing prediction analysis on a to-be-analyzed multivariate index time sequence obtained by processing and analyzing each group of to-be-analyzed flow data to obtain a predicted multivariate index time sequence, and obtaining a multi-dimensional abnormal score sequence based on a corresponding prediction error; and based on a service association topological graph generated by all the multi-dimensional abnormal score sequences, analyzing each multi-dimensional abnormal score sequence to obtain a service flow abnormal score value, and based on a comparison result of the service flow abnormal score value and a service flow abnormal threshold value, obtaining a service flow abnormal identification result. According to the method, semantic information, interaction logic and association relationships of various electric power communication messages can be fully mined and utilized, and the reliability and accuracy of network abnormal flow identification are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power communication technology, and in particular to a deterministic network traffic identification method, system, computer equipment and medium. Background Art

[0002] With the increasing complexity of power systems and the continuous advancement of power communication technology, deterministic networks are increasingly used in power systems, especially in the fields of smart substations and industrial control. The widespread application of standard protocols such as IEC 61850 provides important support for the efficient operation and intelligent management of power systems. However, the complexity and openness of power communication networks also bring new challenges. The destructiveness caused by system failures or malicious network attacks is increasing, which seriously threatens the safety and stable operation of power systems.

[0003] In the deterministic network of power communication, abnormal traffic identification is one of the key technologies for monitoring system failures or malicious network attacks and ensuring network security. At present, abnormal traffic identification technology mainly relies on feature analysis and pattern recognition of network traffic; however, due to the complexity and diversity of standard protocols such as IEC 61850, and the characteristics of power communication network traffic such as multi-source heterogeneity, small data length, periodicity, fixed data flow direction, strong timing, and uneven distribution of normal and abnormal traffic, the traditional traffic anomaly detection method based on simple thresholds, statistical features or machine learning cannot accurately and efficiently identify normal and abnormal traffic in power communication networks. For example, the rich semantic information and complex interaction logic contained in the GOOSE (generic object-oriented substation event) message and SV (sampled value) message in the protocol cannot be fully mined, resulting in insufficient detection accuracy, which makes it very easy to have abnormal traffic false alarms and omissions, and it is difficult to meet the application requirements of power communication deterministic networks. Therefore, it is urgent to provide a method that can efficiently and accurately identify abnormal traffic in power communication deterministic networks, thereby providing reliable protection for the safe and stable operation of power communication networks. Summary of the invention

[0004] The purpose of the present invention is to provide a deterministic network traffic identification method, which constructs a multivariate indicator time series related to network traffic anomalies by reliably analyzing and processing measurement messages and event messages in a power communication deterministic network, and performs business traffic anomaly scoring and abnormal traffic identification threshold setting in combination with a graph topology structure designed based on the correlation between indicators. It can fully mine and utilize the semantic information, interaction logic and correlation relationships of various power communication messages, effectively improve the reliability and accuracy of network abnormal traffic identification, have good engineering application value, and can effectively ensure the high reliability and high security of the power communication network.

[0005] In order to achieve the above-mentioned purpose, a deterministic network traffic identification method, system, computer device and medium are provided.

[0006] In a first aspect, an embodiment of the present invention provides a deterministic network traffic identification method, the method comprising the following steps: According to a preset detection cycle, the network traffic data to be detected is obtained; the network traffic data to be detected includes measurement message data and event message data; Based on the principle of the same source and destination addresses, the network traffic data to be detected is grouped to obtain traffic data to be analyzed of multiple groups of address pairs; The traffic data to be analyzed of each address pair is processed and analyzed respectively to obtain the corresponding multivariate indicator time series to be analyzed; Based on the preset graph attention prediction network, the time series of the multivariate indicators to be analyzed of each address pair are analyzed respectively to obtain the corresponding predicted time series of the multivariate indicators, and based on the error between the predicted time series of the multivariate indicators and the time series of the multivariate indicators to be analyzed, the corresponding multidimensional anomaly score sequence is obtained; Based on the multidimensional anomaly score sequence of each address pair, a corresponding business association topology map is generated, and according to the business association topology map, each multidimensional anomaly score sequence is analyzed respectively to obtain a corresponding business flow anomaly score value; The business traffic anomaly threshold of each address pair is obtained, and based on the comparison result of the business traffic anomaly threshold and the business traffic anomaly score value, a business traffic anomaly identification result is obtained.

[0007] Furthermore, the step of processing and analyzing the traffic data to be analyzed of each address pair to obtain the corresponding multivariate indicator time series to be analyzed includes: Respectively analyzing each measurement-type message data and each event-type message data in the traffic data to be analyzed to obtain corresponding measurement data sets and event data sets; Sorting the measurement data corresponding to the same measurement indicator in all the measurement data sets of the flow data to be analyzed in chronological order to generate a corresponding measurement indicator time series; Perform statistical analysis on each measurement index time series respectively to generate the corresponding measurement index root mean square value time series and measurement index trapezoidal area and time series; Sorting the event data corresponding to the same event indicator in all event data sets of the traffic data to be analyzed in chronological order to generate a corresponding event indicator time series; Summarize all measurement index time series, all measurement index root mean square value time series, all measurement index trapezoidal area and time series, and all event index time series corresponding to the traffic data to be analyzed to obtain a multivariate index time series; Each indicator series in the multivariate indicator time series is normalized to generate the multivariate indicator time series to be analyzed.

[0008] Furthermore, the step of aggregating all measurement index time series, all measurement index root mean square value time series, all measurement index trapezoidal area and time series, and all event index time series corresponding to the traffic data to be analyzed to obtain a multivariate index time series includes: Obtain the minimum sampling time interval of each measurement indicator time series, and use the preset integer multiple of the minimum value of all the minimum sampling time intervals as the target series time scale; According to the target sequence time scale, based on the principle of linear interpolation to fill in missing data, each measurement indicator time series, each measurement indicator root mean square value time series and each measurement indicator trapezoidal area and time series are respectively converted to a sequence time scale, and the corresponding measurement indicator time series to be processed, the measurement indicator root mean square value time series to be processed and the measurement indicator trapezoidal area and time series to be processed are generated; According to the target sequence time scale, based on the principle of supplementing missing data with the previous moment data, each event indicator time series is converted into a sequence time scale to generate a corresponding event indicator time series to be processed; The multivariate indicator time series is obtained by combining all the measurement indicator time series to be processed, all the root mean square value time series to be processed, all the trapezoidal area and time series to be processed, and all the event indicator time series to be processed.

[0009] Furthermore, the step of normalizing each indicator series in the multivariate indicator time series to generate the multivariate indicator time series to be analyzed includes: Perform Z-score normalization processing on each measurement indicator time series, each measurement indicator root mean square value time series, and each measurement indicator trapezoidal area and time series, respectively, to generate corresponding normalized measurement indicator time series, normalized measurement indicator root mean square value time series, and normalized measurement indicator trapezoidal area and time series; Perform minimum and maximum normalization processing on each event indicator time series respectively to generate the corresponding normalized event indicator time series.

[0010] Furthermore, the step of obtaining a corresponding multidimensional anomaly score sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed includes: Performing graph embedding processing on each subsequence in the multivariate indicator time series to be analyzed respectively, generating corresponding graph node representations, and obtaining corresponding sequence association topology graphs based on all graph node representations corresponding to the multivariate indicator time series to be analyzed; Obtaining the multivariate error time series corresponding to the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; Respectively obtain the sequence median and interquartile range corresponding to each sub-error time series in the multivariate error time series, and perform normalization processing on the corresponding sub-error time series according to the sequence median and the interquartile range to obtain the corresponding sub-error time series to be analyzed; Scaling the sub-error time series to be analyzed based on the preset sequence weights of each sub-error time series to be analyzed to obtain a corresponding sub-anomaly score sequence; According to the sequence association topology graph, a corresponding sequence association matrix is ​​obtained, and each sub-anomaly score sequence is fused according to the sequence association matrix to generate the multi-dimensional anomaly score sequence.

[0011] Furthermore, the step of analyzing each multi-dimensional anomaly score sequence according to the service association topology diagram to obtain a corresponding service flow anomaly score value includes: According to the business association topology diagram, obtaining a corresponding business association matrix; According to the business correlation matrix, the anomaly scores in each multi-dimensional anomaly score sequence are weighted and fused to obtain the business flow anomaly score value.

[0012] Furthermore, the step of obtaining a service traffic anomaly threshold value for each address pair and obtaining a service traffic anomaly identification result based on a comparison result of the service traffic anomaly threshold value and the service traffic anomaly score value includes: Obtain the maximum normal score sequence of each address pair when the business traffic is normal; each element in the maximum normal score sequence is the maximum historical abnormal score of the subsequence indicator in the multivariate indicator time series to be analyzed; According to the business association degree matrix corresponding to the business association topology diagram, weighted fusion is performed on the abnormal scores in the maximum normal score sequence to obtain the corresponding business flow abnormal threshold; Determine whether the service flow anomaly score value of each address pair is greater than the corresponding service flow anomaly threshold value; If so, the business traffic anomaly identification result of the corresponding address pair is obtained as business traffic anomaly, and the maximum value of the anomaly score in the multidimensional anomaly score sequence of the corresponding address pair is obtained, and the abnormal traffic is located and controlled according to the subsequence index in the multivariate index time series to be analyzed corresponding to the maximum value of the anomaly score; If not, the business traffic anomaly identification result of the corresponding address pair is that the business traffic is normal.

[0013] In a second aspect, an embodiment of the present invention provides a deterministic network traffic identification system, the system comprising: A data acquisition module is used to obtain the network traffic data to be detected according to a preset detection cycle; the network traffic data to be detected includes measurement message data and event message data; A data grouping module, used for grouping the network flow data to be detected based on the same source and destination address principle, to obtain flow data to be analyzed of multiple groups of address pairs; The data preprocessing module is used to process and analyze the traffic data to be analyzed for each address pair to obtain the corresponding multivariate indicator time series to be analyzed; An indicator anomaly scoring module is used to analyze the time series of the multivariate indicators to be analyzed of each address pair based on a preset graph attention prediction network to obtain a corresponding predicted multivariate indicator time series, and obtain a corresponding multidimensional anomaly scoring sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; A business anomaly scoring module is used to generate a corresponding business association topology map based on the multidimensional anomaly scoring sequence of each address pair, and analyze each multidimensional anomaly scoring sequence according to the business association topology map to obtain a corresponding business flow anomaly scoring value; The identification result generation module is used to obtain the business flow anomaly threshold value of each address pair, and obtain the business flow anomaly identification result based on the comparison result of the business flow anomaly threshold value and the business flow anomaly score value.

[0014] In a third aspect, an embodiment of the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.

[0016] The present invention provides a deterministic network traffic identification method, system, computer equipment and medium. The method is used to obtain network traffic data to be detected including measurement message data and event message data according to a preset detection cycle, group the network traffic data to be detected based on the same principle of source and destination addresses to obtain multiple groups of address pairs of traffic data to be analyzed, respectively process and analyze the traffic data to be analyzed of each address pair to obtain a corresponding multivariate indicator time series to be analyzed, and respectively analyze the multivariate indicator time series to be analyzed of each address pair based on a preset graph prediction network to obtain a corresponding predicted multivariate indicator time series, obtain a corresponding multidimensional anomaly score sequence based on an error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed, and generate a corresponding business association topology map based on the multidimensional anomaly score sequence of each address pair, analyze each multidimensional anomaly score sequence according to the business association topology map to obtain a corresponding business traffic anomaly score value, obtain a business traffic anomaly threshold of each address pair, and obtain a business traffic anomaly identification result based on a comparison result of the business traffic anomaly threshold and the business traffic anomaly score value. Compared with the existing technology, this deterministic network traffic identification method constructs a multivariate indicator time series related to network traffic anomalies by reliably analyzing and processing measurement messages and event messages in the power communication deterministic network, and combines the graph topology structure designed based on the correlation between indicators to perform business traffic anomaly scoring and abnormal traffic identification threshold setting. It can fully mine and utilize the semantic information, interaction logic and correlation relationships of various power communication messages, effectively improve the reliability and accuracy of network abnormal traffic identification, have good engineering application value, and can effectively ensure the high reliability and high security of the power communication network. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of a flow chart of a deterministic network traffic identification method in an embodiment of the present invention; Figure 2 is a flow chart of a method for obtaining a multi-dimensional anomaly score sequence for a single address pair in an embodiment of the present invention; Figure 3 It is a schematic diagram of the process of obtaining the abnormal business flow identification result of a single address pair in an embodiment of the present invention; Figure 4 Schematic diagram of the training performance of a preset graph attention prediction network in an embodiment of the present invention; Figure 5 Schematic diagram of the effect of the design of the service flow abnormality threshold in an embodiment of the present invention; Figure 6 1 is a schematic diagram comparing the effects of a conventional abnormal traffic identification method in an embodiment of the present invention and a deterministic network traffic identification method proposed in the present invention applied to abnormal traffic identification in a substation scenario; Figure 7 is a structural diagram of a deterministic network traffic identification system in an embodiment of the present invention; Figure 8 It is a diagram of the internal structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical scheme and beneficial effects of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Obviously, the embodiments described below are part of the embodiments of the present invention and are only used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] The deterministic network traffic identification method provided by the present invention can be understood as a technical solution that can efficiently and accurately identify abnormal business traffic in a power communication deterministic network, based on the current application status that the existing power system deterministic network abnormal traffic identification method cannot fully mine the rich semantic information and complex interactive logic contained in various communication messages, and cannot effectively solve the impact of the uneven distribution of normal traffic and abnormal traffic on the reliability of traffic prediction, resulting in abnormal traffic false alarms and missed alarms. The following embodiments will explain the deterministic network traffic identification method of the present invention in detail.

[0020] In one embodiment, Figure 1 As shown, a deterministic network traffic identification method is provided, comprising the following steps: S11. According to a preset detection cycle, the network traffic data to be detected is obtained; the network traffic data to be detected includes measurement message data and event message data; wherein the preset detection cycle can be understood as a time interval for performing network traffic anomaly detection determined based on actual application requirements; the corresponding regularly obtained network traffic data to be detected can be understood as communication message timing data that is convenient for reflecting the deterministic network traffic characteristics of power communication, wherein the measurement message data can be understood as SV (sampled value) message flow data with continuous values ​​in the IEC 61850 protocol, and the event message data can be understood as power system event message data with discrete values ​​used for exchange or reporting in the IEC 61850 protocol, such as GOOSE (generic object oriented substation event) message flow data. It should be noted that the sampling time interval for collecting various types of time series data within the preset detection period can be determined based on the shortest communication period of the corresponding message type in the actual power communication deterministic network. In principle, it can be set in milliseconds, seconds, minutes, etc. according to application requirements; the corresponding preset detection period duration interval can be set to tens to hundreds of times the sampling time interval according to actual application requirements to ensure that the obtained network traffic data to be detected has enough data sample points to support subsequent abnormal traffic identification and analysis, which is not specifically limited here.

[0021] S12. Based on the principle of identical source and destination addresses, the network traffic data to be detected is grouped to obtain traffic data to be analyzed for multiple groups of address pairs; wherein, the principle of identical source and destination addresses can be understood as a classification principle for different business traffic data designed in consideration of the characteristics that SV messages and GOOSE messages in the power communication deterministic network are all transmitted in a fixed direction and different source and destination addresses actually correspond to different services in the communication network; the traffic data to be analyzed for each group of address pairs obtained can be understood as a collection of all measurement message data and all event message data of the same source and destination addresses.

[0022] S13. Process and analyze the traffic data to be analyzed for each address pair respectively to obtain the corresponding multivariate indicator time series to be analyzed; wherein the multivariate indicator time series to be analyzed can be understood as a time series set of multiple indicator time series generated based on analyzing, screening, expanding and combining the data set content of various types of messages in the traffic data to be analyzed.

[0023] Specifically, the steps of processing and analyzing the traffic data to be analyzed of each address pair to obtain the corresponding multivariate indicator time series to be analyzed include: Parse each measurement message data and each event message data in the traffic data to be analyzed respectively to obtain corresponding measurement data sets and event data sets; among them, the measurement data set can be understood as a set of measurement index data carried by a measurement message, and the specific type of measurement index data carried by each measurement message can be determined according to actual application requirements, and no specific limitation is made here; similarly, the event data set can be understood as a set of event index data reported by an event message, and the specific type of event index data carried by each event message can be determined according to actual application requirements, and no specific limitation is made here; it should be noted that the process of specifically parsing each measurement message data and event message data to obtain the measurement data set and event data set carried by the message can refer to the relevant technical implementation of parsing IEC61850 protocol messages to obtain message data content, and no detailed description is made here.

[0024] Sort the measurement data corresponding to the same measurement index in all measurement data sets of the traffic data to be analyzed in chronological order to generate a corresponding measurement index time series; among them, the measurement index time series can be understood as the time series data obtained by classifying and summarizing various measurement data in all parsed measurement data sets and sorting them according to the actual collected timestamps. If the measurement data set includes voltage values (three-phase) and current values (three-phase) at the same time, corresponding voltage value time series and current value time series can be constructed respectively according to the voltage values (three-phase) and current values (three-phase).

[0025] Perform statistical analysis on each measurement index time series respectively to generate corresponding measurement index root mean square value time series and measurement index trapezoidal area sum time series; among them, the measurement index root mean square value time series and the measurement index trapezoidal area sum time series can be understood as being obtained by performing root mean square value (RmsValue) index analysis and trapezoidal area sum (TrapAreaSum) index analysis on each measurement index data respectively considering the stability of the power grid and the reliability of power supply quality, and are more representative and characteristic time series data for detecting faults in the power system transmission link; the specific process of obtaining the measurement index root mean square value (RmsValue) time series and the measurement index trapezoidal area sum time series can refer to the calculation technical implementation of the RmsValue index and the TrapAreaSum index in existing statistical analysis, and no further description is made here.

[0026] The event data corresponding to the same event indicator in all event data sets of the flow data to be analyzed are sorted in chronological order to generate a corresponding event indicator time series; wherein, the event indicator time series can be understood as the time series data obtained by classifying and summarizing various types of event data in all event data sets obtained by analysis, and sorting them according to the corresponding actual collected timestamps. If the event data set includes event data such as temperature alarm, circuit breaker status, and disconnector interlock, the corresponding temperature alarm time series, circuit breaker status time series, and disconnector interlock time series can be constructed according to events such as temperature alarm, circuit breaker status, and disconnector interlock. It should be noted that in actual applications, when parsing and obtaining event data sets for each event class message, the corresponding flow ID and application ID, which may change when subjected to malicious attacks, can also be obtained to further enrich the type of event indicator time series and further improve the comprehensiveness of event class message feature extraction and analysis.

[0027] All measurement indicator time series, all measurement indicator root mean square value time series, all measurement indicator trapezoidal area and time series and all event indicator time series corresponding to the traffic data to be analyzed are summarized to obtain a multivariate indicator time series; wherein, the multivariate indicator time series can be understood as a collection of multiple indicator time series obtained by unifying the time scales of all indicator time series.

[0028] Taking into account that event message data in practical applications is discrete rather than continuous, in order to ensure the reliability of all indicator sequence analysis, this embodiment preferably determines the sequence time scale of the multivariate indicator time series based on the sampling time interval of the measurement indicator time series; specifically, the step of aggregating all measurement indicator time series, all measurement indicator root mean square value time series, all measurement indicator trapezoidal area and time series and all event indicator time series corresponding to the traffic data to be analyzed to obtain the multivariate indicator time series includes: The minimum sampling time interval of each measurement indicator time series is obtained, and the preset integer multiples of the minimum value of the minimum sampling time intervals of all sequences are used as the target sequence time scale; wherein, the target sequence time scale can be understood as taking into account the situation that the sampling point moments of each measurement indicator time series may be uneven and the sampling frequencies of different measurement indicators may be inconsistent in actual applications, and the minimum adjacent data interval of all measurement indicator time series is selected as the target sequence time scale required to generate the multivariate indicator time series; correspondingly, the preset integer multiples can be set according to the actual application requirements, and no specific limitation is made here.

[0029] According to the target sequence time scale, sequence time scale conversion processing is respectively performed on each measurement index time series, each root mean square value time series of measurement indexes, and each trapezoidal area sum time series of measurement indexes based on the principle of filling in missing data by linear interpolation, to generate corresponding measurement index time series to be processed, root mean square value time series of measurement indexes to be processed, and trapezoidal area sum time series of measurement indexes to be processed; wherein, the principle of filling in missing data by linear interpolation can be understood as that after screening or merging data with intervals smaller than the target sequence time scale in each measurement index time series, for the positions of sampling moments where data is missing, the method of linear interpolation is used to fill in the data.

[0030] According to the target sequence time scale, sequence time scale conversion processing is respectively performed on each event index time series based on the principle of filling in missing data with the previous moment's data, to generate corresponding event index time series to be processed; wherein, the principle of filling in missing data with the previous moment's data can be understood as a design that, considering the characteristics that the changes of event indexes are not frequent in practical applications and the event indexes can be considered to continuously maintain the previous state without update, the previously obtained nearest data is used for filling until new data appears at the sampling moment position corresponding to the target sequence time scale.

[0031] Through the above processing of each measurement index sequence and each event index sequence, corresponding time series with uniform sampling in the time dimension can be generated to ensure the reliability of subsequent individual analysis of each index time series and correlation analysis between index time series.

[0032] Combine all measurement index time series to be processed, all root mean square value time series of measurement indexes to be processed, all trapezoidal area sum time series of measurement indexes to be processed, and all event index time series to be processed to obtain the multi-index time series.

[0033] Normalize each index sequence in the multi-index time series respectively to generate the multi-index time series to be analyzed; wherein, the normalization processing can be understood as a time series standardization processing based on the distribution characteristics of different index data to avoid analysis anomalies caused by the dimensional difference of different index data; specifically, the step of normalizing each index sequence in the multi-index time series respectively to generate the multi-index time series to be analyzed includes: Perform Z-score normalization processing on each measurement indicator time series, each measurement indicator root mean square value time series, and each measurement indicator trapezoidal area and time series, respectively, to generate corresponding normalized measurement indicator time series, normalized measurement indicator root mean square value time series, and normalized measurement indicator trapezoidal area and time series; wherein, Z-score normalization processing can be understood as a processing method that preferably eliminates the dimension and makes various types of measurement data comparable, taking into account the measurement message data (such as SV message data) with continuous values, and the measurement indicator values ​​provided by it usually obey the characteristics of Gaussian or quasi-Gaussian distribution, and is expressed as: In the formula, in, Represents a time series of measurement indicators The number of samples; and Represents the measurement index time series The mean and variance of Represents the time series of measurement indicators Normalized measurement indicator time series obtained by Z-score standardization.

[0034] Each event indicator time series is subjected to minimum and maximum normalization processing to generate the corresponding normalized event indicator time series; wherein, the minimum and maximum normalization processing is Min-Max normalization, which is considered that the event state (indicator) values ​​contained in event message data (such as GOOSE message data) are usually discrete, for example, only -1, 0 or 1, which is quite different from the Gaussian distribution. It is preferred to scale the indicator values ​​corresponding to all events to the range of [0,1], and convert the data to the same dimension to eliminate the influence between different dimensions. The processing method is expressed as: in, and Represents the event indicator time series The minimum and maximum values ​​of Represents the time series of event indicators The normalized event indicator time series obtained by minimum and maximum normalization processing.

[0035] S14. Based on the preset graph attention prediction network, the time series of the multivariate indicators to be analyzed of each address pair is analyzed respectively to obtain the corresponding predicted multivariate indicator time series, and based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed, the corresponding multidimensional anomaly score sequence is obtained; wherein, the preset graph attention prediction network can be understood as a prediction network based on graph attention, and can adopt the existing graph neural network that can meet the function of predicting and analyzing time series data, and can be trained with a related data set constructed using the same method as the above-mentioned method of generating the multivariate indicator time series to be analyzed. The network structure can adopt GDN (Graph Deviation Network) or a network structure obtained by improving it, which is not specifically limited here.

[0036] In this embodiment, based on the preset graph attention prediction network, the time series of the multivariate indicators to be analyzed for each address pair are analyzed respectively, and the process of obtaining the corresponding predicted multivariate indicator time series is as follows: Figure 2 Shown include: First, based on the time series of each indicator in the multivariate indicator time series to be analyzed, the graph is embedded to obtain the corresponding number of graph vertices. If the multivariate indicator time series to be analyzed for each address pair includes Indicator time series , Indicates the number of sampling points, and can convert each indicator time series The graph embedding representation (graph node representation) is , In order to obtain as many data features as possible, it is generally taken , preferably can be set .

[0037] Then, according to the obtained The corresponding sequence association topology graph is constructed by using the Pearson correlation coefficient to form the edge of the graph to measure the correlation between indicator sequences. The Pearson correlation coefficient can usually achieve better calculation results when the data is not standardized enough, which can ensure the reliability of the topology graph construction in the power communication network scenario. and The edges between can be represented as: in, and They represent the and The graph embedding corresponding to each vertex; Indicates and The edge of the vertices.

[0038] On this basis, we construct the corresponding directed graph and use the adjacency matrix A to represent the directed graph, where A ij Indicates whether there is a directed edge from vertex i to vertex j, indicating the dependency between sequences; the corresponding A ij The calculation method is as follows: in, represents the indicative function, It means to take out the index of the first K values ​​in the sorting, K is a fixed value, which can be determined according to the specific application scenario; that is, all the edge values ​​of vertices k and vertex i to vertex i are arranged in descending order, and it is assumed that there is a directed edge from vertex k to vertex i corresponding to the first K values, and the corresponding , otherwise, the corresponding .

[0039] After obtaining the adjacency matrix (i.e., obtaining the graph topology structure), the adjacency matrix A and each indicator time series in the multivariate indicator time series to be analyzed can be used to predict and analyze each indicator time series through the preset graph attention prediction network to obtain the corresponding predicted indicator time series, and finally obtain the desired predicted multivariate indicator time series. It should be noted that the detailed process of predicting and analyzing the multivariate indicator time series to be analyzed based on the preset graph attention prediction network can be referred to the relevant existing technology implementation using the specific network structure, which will not be described in detail here.

[0040] Considering that in an actual power communication deterministic network, the predicted indicator value obtained based on the historical indicator time series prediction analysis under normal traffic scenarios should be within the normal value range in a statistical sense. If the predicted indicator value at a certain time point deviates too much from the actual indicator value, it can be considered that the actual indicator value at that time point has a high probability of abnormality. In this embodiment, the indicators are preferably scored abnormally based on the prediction error of each indicator time series, and then the abnormal score values ​​of each address corresponding to each indicator sequence are obtained; specifically, the step of obtaining the corresponding multidimensional abnormal score sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed includes: Graph embedding processing is performed on each subsequence in the multivariate indicator time series to be analyzed respectively to generate corresponding graph node representations, and based on all graph node representations corresponding to the multivariate indicator time series to be analyzed, the corresponding sequence association topology graph is obtained; wherein, the process of obtaining the sequence association topology graph refers to the relevant description of the sequence association topology graph in the process of obtaining the multivariate indicator time series prediction in the previous text, which will not be repeated here.

[0041] Obtain the multivariate error time series corresponding to the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; wherein the multivariate error time series can be understood as a combination of indicator error time series generated by subtracting the data at the same time in each indicator time series in the multivariate indicator time series to be analyzed and the predicted indicator time series in the corresponding predicted multivariate indicator time series, and the element value in each indicator error time series can be expressed as: in, and They respectively represent the index value and the index forecast value corresponding to the i-th index time series in the multivariate index time series to be analyzed and the i-th forecast index time series in the corresponding forecast multivariate index time series at time t; Represents the element value (absolute error value) corresponding to time t in the i-th indicator error time series.

[0042] The sequence median and interquartile range corresponding to each sub-error time series in the multivariate error time series are obtained respectively, and the corresponding sub-error time series are normalized according to the sequence median and the interquartile range to obtain the corresponding sub-error time series to be analyzed; wherein the sequence median can be understood as the element value at the middle position of the time series, and the interquartile range can be understood as the interval range (IQR2) of the quartiles in the sequence; correspondingly, the element value in the sub-error time series to be analyzed can be expressed as: In the formula, and They represent the median and interquartile range of the ith sub-error time series in the multivariate error time series; Representation and The corresponding normalized error value. This error normalization processing method can effectively prevent the deviation generated by a certain sub-error sequence from excessively dominating and reducing the accuracy of subsequent comprehensive analysis of each sub-error sequence.

[0043] The sub-error time series to be analyzed is scaled based on the preset sequence weights of each sub-error time series to be analyzed to obtain the corresponding sub-anomaly score sequence; wherein the preset sequence weight can be understood as the weight of the corresponding sub-error time series to be analyzed determined based on the probability of each measurement indicator and event indicator implying abnormal information in the actual application. For example, the preset sequence weight of the indicator involved in the measurement message can be set to 1, and the preset sequence weight of the indicator involved in the event message can be set to a value greater than 1 according to the priority of the event itself and the application requirements, which is used to enlarge or reduce the element values ​​in each sub-error time series to be analyzed. The corresponding element values ​​in each sub-anomaly score sequence can be expressed as: in, represents the preset sequence weight of the i-th sub-error time series to be analyzed; Represents the anomaly score value corresponding to time t in the i-th sub-anomaly score sequence.

[0044] In this embodiment, based on the strong correlation between indicator anomalies and preset errors, and the difference in contribution of various indicators in abnormal traffic detection, an indicator data anomaly scoring mechanism is designed that uses the product of the normalized prediction error of each indicator and the corresponding preset sequence weight as the anomaly score value. The larger the normalized prediction error of the indicator, the higher the indicator anomaly score, the higher the weight value corresponding to the indicator sequence, and the higher the indicator anomaly score. This can ensure that the indicator anomaly score is consistent with the intuitive judgment in actual applications, which helps to quickly discover real and important abnormal events.

[0045] According to the sequence association topology graph, the corresponding sequence association matrix is ​​obtained, and each sub-anomaly score sequence is fused according to the sequence association matrix to generate the multidimensional anomaly score sequence; wherein the sequence association matrix can be obtained based on the adjacency matrix A of the corresponding sequence association topology graph, and the non-diagonal elements of the matrix are 0, and the diagonal elements d ii Calculated as follows: in, Represents the diagonal element value of the i-th row and i-th column in the sequence correlation matrix D; The element value of the i-th row and j-th column in the adjacency matrix A of the sequence association topology graph is represented. The acquisition process of the adjacency matrix A of the sequence association topology graph can refer to the relevant description in the previous article that the indicator time series is predicted and analyzed by the preset graph attention prediction network to obtain the corresponding predicted indicator time series, which will not be repeated here.

[0046] Taking into account the correlation between various indicators in the actual power communication deterministic network, and the corresponding sub-anomaly score sequences must also be associated, if the sub-anomaly score sequence corresponding to each indicator is directly considered equally without any distinction as in the traditional abnormal traffic scoring, it is bound to split the correlation and importance difference between the sub-anomaly score sequences, which will deviate from the significance of using multiple indicator features for joint analysis at the same time, and will inevitably cause deviations in anomaly scores. In order to improve the reliability of traffic anomaly scores at each indicator level as much as possible, this embodiment preferably uses the graph structure information corresponding to the multivariate time series to explore the joint correlation between sub-sequences, and uses the importance of the indicator's impact on traffic anomalies as the sequence weight to explore the importance difference between sub-sequences, so that the anomaly score is based on more accurate multivariate time series prior features, thereby achieving more accurate anomaly scores.

[0047] The multi-dimensional anomaly score sequence in this embodiment can be understood as a multi-index anomaly score sequence composed of index anomaly score values ​​generated by comprehensive analysis of the anomaly score values ​​at all times in each sub-anomaly score sequence of a certain address pair. The location of each index anomaly score value in the sequence can be determined according to a pre-set sequence number, which is not specifically limited here. Specifically, the element value in the multi-dimensional anomaly score sequence corresponding to a certain address pair (a certain source address and destination address pair) is expressed as: In the formula, represents the ith sub-anomaly score sequence of the mth address pair, , Indicates the total number of address pairs; represents the anomaly score value corresponding to time t in the i-th sub-anomaly score sequence of the m-th address pair; Represents the transpose of a vector; represents the length of the sub-anomaly score sequence; D and Respectively represent the sequence association matrix and the inverse matrix of the sequence association matrix corresponding to the current sequence association topology graph; Represents the trace of the matrix, used for normalization; Represents the multidimensional anomaly score sequence of the mth address pair The i-th element value in (the anomaly score corresponding to the i-th sub-anomaly score sequence).

[0048] In this embodiment, the impact of message type and message correlation on traffic anomaly is comprehensively considered, and the role of relatively independent message data in anomaly scoring is strengthened through the degree matrix, thereby effectively ensuring the reliability of anomaly scoring of various data indicators of different address pairs (different services), and providing reliable data support for subsequent accurate anomaly scoring of various business traffic.

[0049] S15. Based on the multidimensional anomaly score sequence of each address pair, a corresponding business association topology map is generated, and according to the business association topology map, each multidimensional anomaly score sequence is analyzed respectively to obtain a corresponding business flow anomaly score value; wherein, after the multidimensional anomaly score sequence of each address pair is obtained through the above method steps, a business association topology map can be constructed based on the correlation between the communication services corresponding to the address pairs, and the business flow of each address pair can be scored for anomaly based on the obtained business association topology map.

[0050] The process of constructing the service association topology map in this embodiment includes: Multidimensional anomaly score sequence for address pairs , a business association topology graph containing M vertices is modeled, and the edge between vertices l and h is defined by the Pearson correlation coefficient. The edge expression is as follows: On this basis, a directed graph is constructed and the adjacency matrix W is used to represent the directed graph, where Indicates whether there is a directed edge from vertex l to vertex m, indicating the dependency between multi-dimensional anomaly score sequences; The calculation method is as follows: In the formula, represents the indicative function, It means to take out the index of the first K values ​​in the sorting, K is a fixed value, which can be determined according to the specific application scenario; that is, all the edge values ​​of vertices h and vertex m to vertex m are arranged in descending order, and it is assumed that there is a directed edge from vertex h to vertex m corresponding to the first K values, and the corresponding , otherwise, the corresponding .

[0051] After obtaining the service association topology graph of all source-destination address pairs involved in communication services in the power communication deterministic network, a comprehensive service-level analysis can be performed on the multidimensional anomaly score sequence of each address pair based on the graph topology, and finally the corresponding service flow anomaly score value is obtained. Specifically, the steps of analyzing each multidimensional anomaly score sequence according to the service association topology graph to obtain the corresponding service flow anomaly score value include: According to the business association topology graph, the corresponding business association matrix is ​​obtained; wherein the business association matrix G can be obtained based on the adjacency matrix W of the business association topology graph, and its non-diagonal elements are 0, and the diagonal elements g ii Calculated as follows: In the formula, represents the diagonal element value of the mth row and the mth column in the business relevance matrix G; Represents the element value in the lth row and nth column of the adjacency matrix W of the service association topology graph.

[0052] According to the business correlation matrix, the anomaly scores in each multi-dimensional anomaly score sequence are weighted and fused to obtain the business flow anomaly score value; wherein the business flow anomaly score value is expressed as: in, represents the multidimensional anomaly score sequence of the mth address pair; They represent the abnormal score of the service traffic of the mth address pair, which is actually The comprehensive score of business anomalies obtained by the weighted sum of the anomaly scores in the above formula is more conducive to reflecting the abnormal status of the network system because it incorporates the impact of the correlation between businesses on the comprehensive score.

[0053] S16. Obtain the business traffic anomaly threshold of each address pair, and obtain the business traffic anomaly identification result based on the comparison result of the business traffic anomaly threshold and the business traffic anomaly score value; wherein, the business traffic anomaly threshold of each address pair can be set to a fixed value according to actual experience in principle, but considering the dynamic variability of business traffic in actual applications, in order to ensure that the abnormal traffic identification can adapt to the changes in the traffic environment and continuously maintain the accuracy and reliability of the business traffic anomaly identification result, this embodiment preferably dynamically obtains the latest business traffic anomaly threshold in each detection cycle, and uses it for abnormal traffic judgment of the corresponding business.

[0054] Specifically, the step of obtaining the service traffic anomaly threshold of each address pair and obtaining the service traffic anomaly identification result based on the comparison result of the service traffic anomaly threshold and the service traffic anomaly score value includes: The maximum normal score sequence of each address pair under normal business traffic conditions is obtained; each element in the maximum normal score sequence is the maximum historical abnormal score of the subsequence indicator in the corresponding multivariate indicator time series to be analyzed, and the maximum historical abnormal score can be understood as the maximum indicator abnormal score value among the indicator abnormal score values ​​of the same indicator in the multidimensional abnormal score sequence corresponding to multiple historical detection cycles before the current detection cycle; that is, after determining the maximum historical abnormal score corresponding to each indicator of each address pair under normal traffic conditions, the maximum normal score sequence (vector) can be combined; it should be noted that in actual applications, the method for obtaining the maximum normal score sequence of each address pair can also be adjusted or replaced accordingly according to the requirements of the actual application scenario.

[0055] According to the service association matrix corresponding to the service association topology diagram, the abnormal scores in the maximum normal score sequence are weighted and fused to obtain the corresponding service flow abnormality threshold; wherein the service flow abnormality threshold is actually a dynamic threshold with a preset detection period as a change period, that is, the threshold obtained for a certain address pair in a certain detection period can be understood as the optimal threshold suitable for use in the current detection period, which is expressed as: In the formula, represents the maximum normal score sequence of the mth address pair; represents the inverse matrix of the business relevance matrix G; Indicates the service traffic abnormality threshold of the mth address pair.

[0056] Determine whether the service flow anomaly score value of each address pair is greater than the corresponding service flow anomaly threshold.

[0057] If so, the business traffic anomaly identification result of the corresponding address pair is business traffic anomaly, and the maximum value of the anomaly score in the multidimensional anomaly score sequence of the corresponding address pair is obtained, and the abnormal traffic is located and controlled according to the subsequence indicator in the multivariate indicator time series to be analyzed corresponding to the maximum value of the anomaly score; wherein, the maximum value of the anomaly score can be understood as the maximum value of the anomaly scores corresponding to each indicator in the corresponding multidimensional anomaly score sequence when the business traffic is abnormal, that is, finding the maximum value of the anomaly score is equivalent to determining the abnormal indicator and knowing the source and destination address information of the anomaly, and then based on the abnormal indicator and the corresponding source and destination address information, the network abnormal location can be further checked, and the corresponding abnormal control strategy can be executed according to the obtained network abnormal location.

[0058] If not, the service flow anomaly identification result of the corresponding address pair is that the service flow is normal. It should be noted that when determining that the service flow is normal, the corresponding maximum normal score sequence can also be updated synchronously to provide data support for service flow anomaly detection in subsequent detection cycles.

[0059] The embodiments of the present invention provide Figure 3As shown, according to a preset detection cycle, network traffic data to be detected including measurement message data and event message data is obtained, and after the network traffic data to be detected is grouped based on the same principle of source and destination addresses to obtain multiple groups of address pairs of traffic data to be analyzed, the traffic data to be analyzed of each address pair is processed and analyzed to obtain a corresponding multivariate indicator time series to be analyzed, and based on a preset graph, the multivariate indicator time series to be analyzed of each address pair is analyzed to obtain a corresponding predicted multivariate indicator time series, and a corresponding multidimensional anomaly score sequence is obtained based on an error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed, and a corresponding business association topology map is generated based on the multidimensional anomaly score sequence of each address pair, and each multidimensional anomaly score sequence is analyzed according to the business association topology map. The corresponding business traffic anomaly score is obtained by analysis, and the business traffic anomaly threshold of each address pair is obtained. The deterministic network traffic anomaly identification scheme of the business traffic anomaly identification result is obtained based on the comparison result of the business traffic anomaly threshold and the business traffic anomaly score. The multivariate indicator time series related to network traffic anomaly is constructed by reliably analyzing and processing the measurement type messages and event type messages in the power communication deterministic network, and the business traffic anomaly score and the abnormal traffic identification threshold are set in combination with the graph topology structure designed based on the correlation relationship between indicators. It can fully mine and utilize the semantic information, interaction logic and correlation relationship of various power communication messages, effectively improve the reliability and accuracy of network abnormal traffic identification, have good engineering application value, and can effectively ensure the high reliability and high security of the power communication network.

[0060] In order to verify the effectiveness of the method of the present invention, this embodiment also tests the abnormal flow recognition in the substation scenario. The test results are as follows: Figure 4-6 As shown: Figure 4 A preset graph of the training data set constructed based on the multivariate indicator time series generation method provided by the method of the present invention is given, and attention is paid to the training performance corresponding to the prediction network. It is easy to know that when the method of the present invention is trained using non-abnormal data, the training loss can converge to less than 5% in 40 iterations, that is, it is easy to quickly train and deploy for specific power communication network usage scenarios.

[0061] Figure 5 The comprehensive performance F1 score (F1-Score) of business traffic anomaly identification is given when the source address and the destination address remain unchanged. The change of the set threshold shows that as the threshold changes, the F1-Score has an optimal value, which further reflects the necessity and reliability of the design of dynamically generated business traffic anomaly thresholds provided by the method of the present invention.

[0062] Figure 6A schematic diagram comparing the effects of the traditional abnormal traffic identification method and the deterministic network traffic identification method proposed in the present invention on the identification of abnormal business traffic in the substation scenario is given. It can be seen that the probability of correct detection (true positive and true negative) is significantly improved by the business traffic abnormality scoring method of the present invention; at the same time, in the power communication network, the harm of missed detection is much higher than the harm of false alarm, and the missed detection rate of the method of the present invention in the test is only about 0.2%, which has good engineering application value and can effectively ensure the high reliability and high security of the stable operation of the power communication network.

[0063] It should be noted that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders.

[0064] In one embodiment, Figure 7 As shown, a deterministic network traffic identification system is provided, the system comprising: The data acquisition module 1 is used to obtain the network traffic data to be detected according to a preset detection cycle; the network traffic data to be detected includes measurement message data and event message data; Data grouping module 2, used for grouping the network flow data to be detected based on the same source and destination address principle, to obtain flow data to be analyzed of multiple groups of address pairs; The data preprocessing module 3 is used to process and analyze the traffic data to be analyzed of each address pair respectively to obtain the corresponding multivariate indicator time series to be analyzed; The indicator anomaly scoring module 4 is used to analyze the time series of the multivariate indicators to be analyzed of each address pair based on the preset graph attention prediction network, obtain the corresponding predicted multivariate indicator time series, and obtain the corresponding multidimensional anomaly scoring sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; The service anomaly scoring module 5 is used to generate a corresponding service association topology map based on the multi-dimensional anomaly scoring sequence of each address pair, and analyze each multi-dimensional anomaly scoring sequence according to the service association topology map to obtain a corresponding service flow anomaly scoring value; The identification result generating module 6 is used to obtain the service flow anomaly threshold value of each address pair, and obtain the service flow anomaly identification result based on the comparison result of the service flow anomaly threshold value and the service flow anomaly score value.

[0065] For the specific limitations of the deterministic network traffic identification system, please refer to the limitations of the deterministic network traffic identification method above, and the corresponding technical effects can also be obtained equivalently, which will not be repeated here. Each module in the above-mentioned deterministic network traffic identification system can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0066] Figure 8 FIG. 1 shows an internal structure diagram of a computer device in an embodiment, and the computer device may specifically be a terminal or a server. Figure 8 As shown, the computer device includes a processor, a memory, a network interface, a display, a camera and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a deterministic network traffic identification method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0067] It can be understood by those skilled in the art that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present invention, and does not constitute a limitation on the computer device to which the scheme of the present invention is applied. The specific computing device may include more or less components than those shown in the figure, or combine certain components, or have an equivalent component arrangement.

[0068] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0069] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0070] In summary, the embodiments of the present invention provide a deterministic network traffic identification method and system, wherein the deterministic network traffic identification method realizes obtaining the network traffic data to be detected including the measurement message data and the event message data according to the preset detection period, grouping the network traffic data to be detected based on the same principle of the source and destination addresses to obtain the traffic data to be analyzed of multiple groups of address pairs, respectively processing and analyzing the traffic data to be analyzed of each address pair to obtain the corresponding multivariate indicator time series to be analyzed, and respectively analyzing the multivariate indicator time series to be analyzed of each address pair based on the preset graph prediction network to obtain the corresponding predicted multivariate indicator time series, obtaining the corresponding multidimensional anomaly score sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed, and generating the corresponding business association topology map based on the multidimensional anomaly score sequence of each address pair. According to the business association topology diagram, each multidimensional anomaly score sequence is analyzed to obtain the corresponding business flow anomaly score value, and the business flow anomaly threshold of each address pair is obtained. The technical solution for obtaining the business flow anomaly identification result is based on the comparison result of the business flow anomaly threshold and the business flow anomaly score value. This method constructs a multivariate indicator time series related to network flow anomaly by reliably analyzing and processing measurement messages and event messages in the power communication deterministic network, and combines the graph topology structure designed based on the correlation relationship between indicators to perform business flow anomaly scoring and abnormal flow identification threshold setting. It can fully mine and utilize the semantic information, interaction logic and correlation relationship of various power communication messages, effectively improve the reliability and accuracy of network abnormal flow identification, have good engineering application value, and can effectively ensure the high reliability and high security of the power communication network.

[0071] Each embodiment in this specification is described in a progressive manner, and the same or similar parts of each embodiment can be directly referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. It should be noted that the technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0072] The above-mentioned embodiments only express several preferred implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this technical field, several improvements and substitutions can be made without departing from the technical principle of the present invention, and these improvements and substitutions should also be regarded as the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be based on the protection scope of the claims.

Claims

1. A deterministic network traffic identification method, characterized in that: Applied to an electric power system, the method comprises the following steps: According to a preset detection cycle, the network traffic data to be detected is obtained; the network traffic data to be detected includes measurement message data and event message data; Based on the principle of the same source and destination addresses, the network traffic data to be detected is grouped to obtain traffic data to be analyzed of multiple groups of address pairs; The traffic data to be analyzed of each address pair is processed and analyzed respectively to obtain the corresponding multivariate indicator time series to be analyzed; Based on the preset graph attention prediction network, the time series of the multivariate indicators to be analyzed of each address pair are analyzed respectively to obtain the corresponding predicted time series of the multivariate indicators, and based on the error between the predicted time series of the multivariate indicators and the time series of the multivariate indicators to be analyzed, the corresponding multidimensional anomaly score sequence is obtained; Based on the multidimensional anomaly score sequence of each address pair, a corresponding business association topology map is generated, and according to the business association topology map, each multidimensional anomaly score sequence is analyzed respectively to obtain a corresponding business flow anomaly score value; The business traffic anomaly threshold of each address pair is obtained, and based on the comparison result of the business traffic anomaly threshold and the business traffic anomaly score value, a business traffic anomaly identification result is obtained.

2. The deterministic network traffic identification method according to claim 1, characterized in that: The step of processing and analyzing the traffic data to be analyzed of each address pair to obtain the corresponding multivariate indicator time series to be analyzed comprises: Respectively analyzing each measurement-type message data and each event-type message data in the traffic data to be analyzed to obtain corresponding measurement data sets and event data sets; Sorting the measurement data corresponding to the same measurement indicator in all the measurement data sets of the flow data to be analyzed in chronological order to generate a corresponding measurement indicator time series; Perform statistical analysis on each measurement index time series respectively to generate the corresponding measurement index root mean square value time series and measurement index trapezoidal area and time series; Sorting the event data corresponding to the same event indicator in all event data sets of the traffic data to be analyzed in chronological order to generate a corresponding event indicator time series; Summarize all measurement index time series, all measurement index root mean square value time series, all measurement index trapezoidal area and time series, and all event index time series corresponding to the traffic data to be analyzed to obtain a multivariate index time series; Each indicator series in the multivariate indicator time series is normalized to generate the multivariate indicator time series to be analyzed.

3. The deterministic network traffic identification method according to claim 2, characterized in that: The step of aggregating all measurement index time series, all measurement index root mean square value time series, all measurement index trapezoidal area and time series, and all event index time series corresponding to the traffic data to be analyzed to obtain a multivariate index time series comprises: Obtain the minimum sampling time interval of each measurement indicator time series, and use the preset integer multiple of the minimum value of all the minimum sampling time intervals as the target series time scale; According to the target sequence time scale, based on the principle of linear interpolation to fill in missing data, each measurement indicator time series, each measurement indicator root mean square value time series and each measurement indicator trapezoidal area and time series are respectively converted to a sequence time scale, and the corresponding measurement indicator time series to be processed, the measurement indicator root mean square value time series to be processed and the measurement indicator trapezoidal area and time series to be processed are generated; According to the target sequence time scale, based on the principle of supplementing missing data with the previous moment data, each event indicator time series is converted into a sequence time scale to generate a corresponding event indicator time series to be processed; The multivariate indicator time series is obtained by combining all the measurement indicator time series to be processed, all the root mean square value time series to be processed, all the trapezoidal area and time series to be processed, and all the event indicator time series to be processed.

4. The deterministic network traffic identification method according to claim 2, characterized in that: The step of normalizing each indicator sequence in the multivariate indicator time series to generate the multivariate indicator time series to be analyzed comprises: Perform Z-score normalization processing on each measurement indicator time series, each measurement indicator root mean square value time series, and each measurement indicator trapezoidal area and time series, respectively, to generate corresponding normalized measurement indicator time series, normalized measurement indicator root mean square value time series, and normalized measurement indicator trapezoidal area and time series; Perform minimum and maximum normalization processing on each event indicator time series respectively to generate the corresponding normalized event indicator time series.

5. The deterministic network traffic identification method according to claim 1, characterized in that: The step of obtaining a corresponding multidimensional anomaly score sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed comprises: Performing graph embedding processing on each subsequence in the multivariate indicator time series to be analyzed respectively, generating corresponding graph node representations, and obtaining corresponding sequence association topology graphs based on all graph node representations corresponding to the multivariate indicator time series to be analyzed; Obtaining the multivariate error time series corresponding to the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; Respectively obtain the sequence median and interquartile range corresponding to each sub-error time series in the multivariate error time series, and perform normalization processing on the corresponding sub-error time series according to the sequence median and the interquartile range to obtain the corresponding sub-error time series to be analyzed; Scaling the sub-error time series to be analyzed based on the preset sequence weights of each sub-error time series to be analyzed to obtain a corresponding sub-anomaly score sequence; According to the sequence association topology graph, a corresponding sequence association matrix is ​​obtained, and each sub-anomaly score sequence is fused according to the sequence association matrix to generate the multi-dimensional anomaly score sequence.

6. The deterministic network traffic identification method according to claim 1, characterized in that: The step of analyzing each multi-dimensional anomaly score sequence according to the business association topology diagram to obtain a corresponding business flow anomaly score value comprises: According to the business association topology diagram, obtaining a corresponding business association matrix; According to the business correlation matrix, the anomaly scores in each multi-dimensional anomaly score sequence are weighted and fused to obtain the business flow anomaly score value.

7. The deterministic network traffic identification method according to claim 1, characterized in that: The step of obtaining the abnormal service flow threshold of each address pair and obtaining the abnormal service flow identification result based on the comparison result of the abnormal service flow threshold and the abnormal service flow score value comprises: Obtain the maximum normal score sequence of each address pair when the business traffic is normal; each element in the maximum normal score sequence is the maximum historical abnormal score of the subsequence indicator in the multivariate indicator time series to be analyzed; According to the business association degree matrix corresponding to the business association topology diagram, weighted fusion is performed on the abnormal scores in the maximum normal score sequence to obtain the corresponding business flow abnormal threshold; Determine whether the service flow anomaly score value of each address pair is greater than the corresponding service flow anomaly threshold value; If so, the business traffic anomaly identification result of the corresponding address pair is obtained as business traffic anomaly, and the maximum value of the anomaly score in the multidimensional anomaly score sequence of the corresponding address pair is obtained, and the abnormal traffic is located and controlled according to the subsequence index in the multivariate index time series to be analyzed corresponding to the maximum value of the anomaly score; If not, the business traffic anomaly identification result of the corresponding address pair is that the business traffic is normal.

8. A deterministic network traffic identification system, characterized in that: Applied to a power system, the system comprises: A data acquisition module is used to obtain the network traffic data to be detected according to a preset detection cycle; the network traffic data to be detected includes measurement message data and event message data; A data grouping module, used for grouping the network flow data to be detected based on the same source and destination address principle, to obtain flow data to be analyzed of multiple groups of address pairs; The data preprocessing module is used to process and analyze the traffic data to be analyzed for each address pair to obtain the corresponding multivariate indicator time series to be analyzed; An indicator anomaly scoring module is used to analyze the time series of the multivariate indicators to be analyzed of each address pair based on a preset graph attention prediction network to obtain a corresponding predicted multivariate indicator time series, and obtain a corresponding multidimensional anomaly scoring sequence based on the error between the predicted multivariate indicator time series and the multivariate indicator time series to be analyzed; A business anomaly scoring module is used to generate a corresponding business association topology map based on the multidimensional anomaly scoring sequence of each address pair, and analyze each multidimensional anomaly scoring sequence according to the business association topology map to obtain a corresponding business flow anomaly scoring value; The identification result generation module is used to obtain the business flow anomaly threshold value of each address pair, and obtain the business flow anomaly identification result based on the comparison result of the business flow anomaly threshold value and the business flow anomaly score value.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Network traffic identification method and system

    CN117113262A

  • Malicious traffic identification system and method based on graph neural network and stable learning thought, program, equipment and storage medium

    CN118449729A

  • Computer network anomaly detection method

    CN118784364A

  • Network traffic abnormity monitoring method and device based on BiLSTM-Att network

    CN119232490A

  • Network traffic prediction method and system, computer equipment and storage medium

    CN119728461A

Cited By

  • Charging service real-time monitoring method and device based on service probe

    CN120416093A