A network security protection method, system, device and medium
Through sliding window segmentation and multi-scale alignment technology, combined with bidirectional long and short-term memory networks and cross-scale causal transmission diagrams, the problem of ignoring packet time series anomalies in the existing technology is solved, and efficient detection and suppression of hidden channel attacks is achieved, which significantly improves the recognition accuracy and reliability of network security protection.
Patent Information
- Application Number
- CN202510337205.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The prior art ignores the abnormal monitoring of data packet time series in network security protection, making it difficult to detect malicious behaviors of using time intervals to construct hidden communications in a time-consuming manner, which may lead to the leakage of sensitive information and the existence of illegal communication channels.
The network packet timing is divided by sliding window, the difference sequence of adjacent packet arrival time intervals is extracted, multi-scale alignment is performed to generate standardized timing feature vectors, and input a bidirectional long and short-term memory network for timing mode classification. Next, a cross-scale causal transmission diagram is constructed, continuous co-ordination analysis is performed, the continuous interval of the one-dimensional continuous co-ordination group is extracted, topological stability indicators are calculated, and whether there is a hidden channel attack is determined based on the KL divergence difference of the classification probability distribution is used, and a random delay interference parameter injection data flow is generated based on the Poisson distribution.
It effectively solves the problems of poor adaptability and high error detection rate in complex network environments, significantly improves the recognition accuracy of nonlinear timing abnormalities, reduces the false alarm rate, and achieves the balance between hidden confrontation and communication reliability through dynamic interference mechanism and protocol baseline adaptive update strategy.
Smart Images

Figure CN119854050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network information security technology, and more specifically, to a network security protection method, system, device and medium. Background Art
[0002] The continuous advancement of digital information transmission technology has led to increasingly diverse threats to network security protection. Existing technologies mainly protect against common problems such as data delay, packet loss and content anomalies. However, in the process of digital information transmission, the changes in the tiny time intervals between data packets have not received sufficient attention. This subtle feature is often ignored in traditional intrusion detection systems, resulting in a lack of effective monitoring of malicious behaviors that use time intervals to construct covert communications.
[0003] The existing technology only focuses on monitoring the content of data transmission and conventional traffic indicators, but ignores the monitoring of data packet time series anomalies. This will make it difficult to detect attacks that use time intervals to construct covert channels in a timely manner, which may lead to the leakage of sensitive information and the existence of illegal communication channels. Summary of the invention
[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a network security protection method, system, device and medium to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A network security protection method comprises the following steps:
[0007] The network data packet timing is segmented by sliding windows, the difference sequence of the time intervals between adjacent data packets is extracted, and the difference sequence is aligned at multiple scales to generate a standardized timing feature vector.
[0008] The time series feature vector is input into the bidirectional long short-term memory network to classify the time series pattern, and the hidden state vector set and classification probability distribution are output;
[0009] Perform symbolic discrete mapping on the hidden state vector set and construct a cross-scale causal transfer graph;
[0010] We conduct persistence homology analysis on cross-scale causal transfer graphs, extract persistence intervals corresponding to one-dimensional persistence homology groups, and calculate topological stability indicators based on interval length distributions.
[0011] Based on the KL divergence difference and topological stability index of the classification probability distribution, determine whether there is a covert channel attack;
[0012] When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while the time series statistical characteristics of the preset protocol baseline template are updated.
[0013] In a preferred embodiment, the network data packet timing is segmented by a sliding window, a difference sequence of the time intervals between adjacent data packets is extracted, and a multi-scale alignment is performed on the difference sequence to generate a standardized timing feature vector, including:
[0014] Divide the network data packet arrival time sequence into multiple continuous time windows according to a preset window length, where the window length is dynamically set according to the network protocol type;
[0015] In each time window, the difference sequence of the time intervals between the arrival of adjacent data packets is calculated, and the difference sequence is the absolute difference between the time intervals between the arrival of adjacent data packets;
[0016] The difference sequence is mapped to three time scales: milliseconds, seconds, and minutes. At each time scale, the difference sequence is standardized.
[0017] The standardized difference sequences under three time scales are spliced in chronological order to generate a multi-dimensional time series feature vector.
[0018] In a preferred embodiment, the time series feature vector is input into a bidirectional long short-term memory network for time series pattern classification, and a hidden state vector set and classification probability distribution are output, including:
[0019] Construct a bidirectional long short-term memory network, which includes a forward layer and a backward layer. The forward layer processes the time series feature vector in chronological order, and the backward layer processes the time series feature vector in reverse chronological order.
[0020] The time series feature vector is input into the forward layer to generate a forward hidden state vector sequence, and the time series feature vector is reversely arranged and input into the backward layer to generate a backward hidden state vector sequence;
[0021] Concatenate the vectors of the same time step in the forward hidden state vector sequence and the backward hidden state vector sequence according to the dimension to form a hidden state vector set;
[0022] The hidden state vector set is input into the fully connected layer, and the output of the fully connected layer is processed by the normalized exponential function to obtain the classification probability distribution. The classification probability distribution refers to the probability that the current traffic belongs to the normal protocol type.
[0023] In a preferred embodiment, symbolic discrete mapping is performed on the hidden state vector set to construct a cross-scale causal transfer graph, including:
[0024] Based on the preset symbolization rules, the hidden state vector of each time step is mapped to the corresponding discrete symbol to generate a symbol sequence;
[0025] Dividing the symbol sequence based on different time scales to generate a set of multi-scale symbol subsequences, where each time scale corresponds to a preset time window length;
[0026] Perform Granger causality test on each symbol subsequence in the multi-scale symbol subsequence set, calculate the conditional probability of transition between adjacent symbols, and generate the symbol transition probability matrix corresponding to each time scale;
[0027] The symbol transfer probability matrices at each time scale are vertically spliced according to the scale level, and the time scale identifier is marked inside each matrix to generate a cross-scale causal transfer diagram.
[0028] In a preferred embodiment, a continuous homology analysis is performed on the cross-scale causal transfer graph, the continuous intervals corresponding to the one-dimensional continuous homology group are extracted, and the topological stability index is calculated based on the interval length distribution, including:
[0029] The cross-scale causal transfer graph is subjected to continuous homology analysis, and the symbolic transfer probability matrix in the cross-scale causal transfer graph is used as input to construct a multi-scale filtering complex structure.
[0030] The multi-scale filtering complex structure is constructed by taking the symbols in the symbol transition probability matrix of each time scale as vertices, the transition probability between symbols as the weight of the edge, adding the edges whose edge weights are greater than or equal to the preset filtering parameters to the complex structure according to the preset filtering parameters, and generating the multi-scale filtering sequence according to the hierarchical order of the time scales;
[0031] Based on the multi-scale filtering complex structure, the persistence interval of the one-dimensional persistent homology group is calculated;
[0032] The continuous interval is generated by tracking the generation and disappearance events of the one-dimensional homology group in the multi-scale filtering sequence, recording the starting filter parameters corresponding to each homology generation event and the ending filter parameters corresponding to the disappearance event, and forming the continuous interval interval;
[0033] The topological stability index is calculated based on the length distribution of all persistent intervals.
[0034] In a preferred embodiment, judging whether there is a covert channel attack based on the KL divergence difference of the classification probability distribution and the topological stability index includes:
[0035] Calculate the KL divergence difference between the classification probability distribution of the network traffic to be detected and the classification probability distribution of normal network traffic;
[0036] The KL divergence difference value is calculated as follows: the normal network traffic classification probability distribution is used as the reference distribution, the network traffic classification probability distribution to be detected is used as the comparison distribution, and the probability value of each symbol is calculated by multiplying the logarithm of the ratio of the reference distribution probability to the comparison distribution probability by the reference distribution probability, and the calculation results of all symbols are accumulated and summed;
[0037] Based on the comparison result of the KL divergence difference value with the preset first threshold value and the comparison result of the topological stability index with the preset second threshold value, determining whether there is a covert channel attack;
[0038] When the KL divergence difference value is greater than the first threshold and the topological stability index is less than the second threshold, it is determined that a covert channel attack exists.
[0039] In a preferred embodiment, when there is a covert channel attack, a random delay interference parameter is generated based on Poisson distribution and injected into the target data stream, and the time series statistical characteristics of the preset protocol baseline template are updated at the same time;
[0040] In response to detecting a covert channel attack, generating a time interval sequence that obeys a Poisson probability density function, and using the time interval sequence as a random delay interference parameter;
[0041] The random delay interference parameter is embedded into the data packet transmission timestamp of the target data stream, so that the actual sending interval of adjacent data packets deviates from the original protocol time baseline after the random delay interference parameter is superimposed;
[0042] Collect the transmission time series of the target data stream after the interference is injected, calculate the corresponding mean and variance, and dynamically update the time window sliding average of the preset protocol baseline template based on the mean and variance;
[0043] The expected parameters of the Poisson probability density function in the next period are adjusted according to the updated time window sliding average.
[0044] In another aspect, the present invention provides a network security protection system, comprising:
[0045] Time series vector generation module: Segment the network data packet time series through a sliding window, extract the difference sequence of the time intervals between adjacent data packets, and perform multi-scale alignment on the difference sequence to generate a standardized time series feature vector;
[0046] Temporal pattern classification module: inputs the temporal feature vector into the bidirectional long short-term memory network for temporal pattern classification, and outputs the hidden state vector set and classification probability distribution;
[0047] Symbolic mapping module: symbolic discrete mapping of hidden state vector sets to construct cross-scale causal transfer graphs;
[0048] Continuous homology analysis module: Perform continuous homology analysis on cross-scale causal transfer graphs, extract the continuous intervals corresponding to the one-dimensional continuous homology group, and calculate the topological stability index based on the interval length distribution;
[0049] Covert channel detection module: Based on the KL divergence difference and topological stability index of the classification probability distribution, it determines whether there is a covert channel attack;
[0050] Attack interference injection module: When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while updating the time series statistical characteristics of the preset protocol baseline template.
[0051] On the other hand, the present invention provides a network security protection device, which includes: a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein a network security protection method is implemented when the program or instruction is executed by the processor.
[0052] On the other hand, the present invention provides a network security protection medium, on which a program or instruction is stored, and a network security protection method is implemented when the program or instruction is executed by a processor.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] 1. The present invention effectively solves the problems of poor adaptability and high false positive rate of traditional covert channel detection methods in complex network environments through multi-scale time series feature extraction and cross-protocol dynamic analysis. Aiming at the defect that single time scale analysis in the prior art is difficult to capture microsecond-level timing deviation, sliding window segmentation and multi-granularity normalization processing are adopted, combined with the timing pattern classification of bidirectional long short-term memory network, to achieve high-dimensional feature decoupling of normal protocol traffic and covert channel behavior, significantly improve the recognition accuracy of nonlinear timing anomalies, and further transform the hidden state vector into an interpretable symbol transfer network through symbolic discrete mapping and cross-scale causal transfer graph construction, breaking through the black box limitation of traditional machine learning models and enhancing the generalization ability of unknown covert channel coding rules. Based on the topological stability quantitative index of continuous coherence analysis, the deep anomalies of traffic structure are revealed from the perspective of algebraic topology, and multi-dimensional criteria are formed with KL divergence difference, which greatly reduces the false alarm rate.
[0055] 2. The dynamic interference mechanism and protocol baseline adaptive update strategy of the present invention solve the problem that traditional fixed threshold interference is easily predicted by attackers and destroys protocol communications. Random delay parameters are generated through Poisson distribution and superimposed on the target data stream, which not only destroys the time synchronization that the covert channel relies on, but also ensures that the traffic statistics after interference are consistent with the protocol baseline, achieving a balance between covert confrontation and communication reliability. The dynamic update algorithm of the protocol baseline template quickly responds to network environment fluctuations through sliding mean and variance weight adjustment, avoiding the failure of the detection model due to the solidification of the traffic baseline. In multi-protocol mixed traffic scenarios, this solution can accurately identify timing and storage-type covert channels, and effectively suppress information leakage through dynamic interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A flowchart of a network security protection method of the present invention;
[0057] Figure 2 The present invention is a schematic structural diagram of a network security protection system. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] Embodiment 1: Figure 1 The present invention provides a network security protection method, which comprises the following steps:
[0060] The network data packet timing is segmented by sliding windows, the difference sequence of the time intervals between adjacent data packets is extracted, and the difference sequence is aligned at multiple scales to generate a standardized timing feature vector.
[0061] The time series feature vector is input into the bidirectional long short-term memory network for time series pattern classification, and the hidden state vector set and classification probability distribution are output.
[0062] The set of hidden state vectors is symbolically mapped discretely to construct a cross-scale causal transfer graph.
[0063] The cross-scale causal transfer graph is subjected to persistence homology analysis to extract the persistence intervals corresponding to the one-dimensional persistence homology group, and the topological stability index is calculated based on the interval length distribution.
[0064] Based on the KL divergence difference and topological stability index of the classification probability distribution, it is determined whether there is a covert channel attack.
[0065] When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while the time series statistical characteristics of the preset protocol baseline template are updated.
[0066] The network data packet timing is segmented by sliding windows, the difference sequence of the time intervals between adjacent data packets is extracted, and the difference sequence is aligned at multiple scales to generate a standardized timing feature vector, including:
[0067] The network data packet arrival time series is divided into multiple continuous time windows according to the preset window length, where the window length is dynamically set according to the network protocol type.
[0068] A fixed window length of the number of packets is used for TCP traffic, and a fixed time span is used for UDP traffic. For example, if the network traffic currently being processed is TCP, then whenever the 10th, 20th, 30th, ... packets are captured, the arrival time sequence of the first 10 consecutive packets is divided into a time window.
[0069] Set the window length to a fixed time threshold, which is twice the historical average packet arrival interval of the current network traffic. For example, if the historical average interval is 50 milliseconds, the window length is set to 100 milliseconds, and the arrival time series of all packets in the window are divided into a time window every 100 milliseconds.
[0070] Monitor the arrival of data packets from the network interface card in real time, and determine whether the current traffic belongs to the Transmission Control Protocol or the User Datagram Protocol based on the protocol type field of the data packet; dynamically adjust the window length according to the above rules, and generate an array containing the arrival timestamps of all data packets in the window at the end of each window.
[0071] In each time window, a difference sequence of adjacent data packet arrival time intervals is calculated, and the difference sequence includes the absolute value of the previous data packet arrival time interval minus the next data packet arrival time interval.
[0072] For the generated data packet arrival timestamp array within a single time window, the arrival time intervals between two adjacent data packets are calculated in sequence. For the adjacent interval sequence, the absolute difference between every two consecutive intervals is calculated. Specifically, the interval sequence is traversed from front to back, and the absolute value of the previous interval value minus the next interval value is taken as the difference sequence element.
[0073] The difference sequence is mapped to three time scales: milliseconds, seconds, and minutes. At each time scale, the difference sequence is standardized.
[0074] The generated difference sequence is mapped to the following three time scales: millisecond granularity: directly retain the millisecond unit value of the original difference sequence; second granularity: divide the difference by 1000 and convert it to seconds; minute granularity: divide the difference by 60000 and convert it to minutes.
[0075] In this embodiment, the extracted difference in the arrival time intervals of adjacent data packets may be set to a millisecond-level accuracy.
[0076] The standardization process calculates the mean and standard deviation of the difference sequence in the current time window, and obtains the standardized result by subtracting the mean from each difference and dividing it by the standard deviation. The following operations are performed on the difference sequence at each time scale: the mean and standard deviation of the difference sequence of this granularity in the current time window are calculated; standardization calculation is performed on each difference element: the mean is subtracted from the difference element and divided by the standard deviation to obtain the standardized value; if the standardized value is zero (that is, all differences are the same), the original difference sequence is directly output.
[0077] Assume that the millisecond difference sequence is [5, 8, 6], the mean is 6.33, and the standard deviation is 1.25. The standardized sequence is [(5-6.33) / 1.25, (8-6.33) / 1.25, (6-6.33) / 1.25].
[0078] The standardized difference sequences under three time scales are spliced in chronological order to generate a multi-dimensional time series feature vector.
[0079] The standardized difference sequences at the millisecond, second, and minute levels are concatenated into a one-dimensional array in chronological order; the dimension of the time series feature vector is 3×(the number of data packets in the current time window - 1).
[0080] The time series feature vector is input into the bidirectional long short-term memory network for time series pattern classification, and the hidden state vector set and classification probability distribution are output, including:
[0081] A bidirectional long short-term memory network is constructed. The bidirectional long short-term memory network includes a forward layer and a backward layer. The forward layer processes the time series feature vector in chronological order, and the backward layer processes the time series feature vector in reverse chronological order.
[0082] Forward layer: Process the time series feature vector in chronological order, that is, calculate the hidden state from the first time step to the last time step. The hidden state calculation of each time step depends on the input features of the current time step and the hidden state of the previous time step.
[0083] Backward layer: Process the time series feature vector in reverse time order, that is, calculate the hidden state from the last time step to the first time step. The hidden state calculation of each time step depends on the input features of the current time step and the hidden state of the next time step.
[0084] Bidirectional connection rule: There is no weight sharing between the forward layer and the backward layer, and the hidden state of each time step is calculated independently and only associated at the final concatenation.
[0085] The dimension of the time series feature vector is the feature dimension of each time step multiplied by the total number of time steps. For example, if the feature dimension of each time step is 64 and the time window is divided into 10 time steps, the input dimension is 64×10.
[0086] During input, the forward layer directly receives the original time series feature vector, and the backward layer receives the time series feature vector after reverse arrangement. Reverse arrangement means that the order of time steps in the original time series feature vector is completely reversed. For example, the original time series is a sequence from time step 1 to time step 10, and after reverse arrangement, it is a sequence from time step 10 to time step 1.
[0087] The time series feature vector is input into the forward layer to generate a forward hidden state vector sequence. At the same time, the time series feature vector is reversely arranged and input into the backward layer to generate a backward hidden state vector sequence.
[0088] The time series feature vector is input into the forward layer, and each time step is processed in chronological order: at each time step, the forward layer generates the hidden state of the current time step according to the input features of the current time step and the hidden state of the previous time step through the calculation rules of the long short-term memory unit (including input gate, forget gate, output gate and memory unit update); the hidden state of the initial time step is set to an all-zero vector, and the hidden states of subsequent time steps are calculated recursively in sequence.
[0089] Finally, the forward hidden state vector sequence is output, and the dimension of each hidden state is consistent with the hidden state dimension of the long short-term memory network layer (for example, 64 dimensions).
[0090] The reverse-arranged time series feature vector is input into the backward layer, and each time step is processed in reverse time order: at each time step, the backward layer generates the hidden state of the current time step according to the input features of the current time step and the hidden state of the next time step through the calculation rules of the long short-term memory unit; the hidden state of the initial time step (the last time step after reverse arrangement) is set to an all-zero vector, and the hidden states of subsequent time steps are calculated recursively in sequence.
[0091] The final output is a sequence of backward hidden state vectors, whose hidden state dimension is consistent with the forward layer, and each hidden state corresponds to the corresponding time step in the original time series.
[0092] The vectors of the same time step in the forward hidden state vector sequence and the backward hidden state vector sequence are concatenated by dimension to form a set of hidden state vectors.
[0093] The vectors of the same original time step in the forward hidden state vector sequence and the backward hidden state vector sequence are concatenated according to the feature dimension. For example, the hidden state of the forward layer at time step 1 is concatenated with the hidden state of the backward layer at time step 1 (corresponding to time step 10 after reverse arrangement) to form a double-dimensional hidden state vector.
[0094] Specifically, if the dimension of the forward hidden state is 64 and the dimension of the backward hidden state is 64, the dimension of the concatenated hidden state vector is 128. For example, assuming that the total number of time steps is 10, and the hidden state dimensions of the forward layer and the backward layer are both 64, the hidden state vector set contains 10 vectors of dimension 128, each of which corresponds to the bidirectional feature expression of one time step.
[0095] The hidden state vector set is input into the fully connected layer, and the output of the fully connected layer is processed by the normalized exponential function to obtain the classification probability distribution. The classification probability distribution refers to the probability that the current traffic belongs to the normal protocol type. The output dimension of the classification probability distribution is consistent with the preset number of network traffic categories.
[0096] Construct a fully connected layer, whose input dimension is consistent with the dimension of the concatenated hidden state vector (for example, 128), and the output dimension is set to the number of preset network traffic categories (for example, normal protocol, attack type 1, attack type 2, etc., a total of 5 categories).
[0097] The set of hidden state vectors is input into the fully connected layer, and the classification score is calculated independently for the hidden state vector at each time step. The classification score is generated by linear transformation, that is, each score is obtained by multiplying the hidden state vector by the fully connected layer weight matrix and adding the bias vector.
[0098] A normalized exponential function (Softmax) is applied to the classification scores output by the fully connected layer to convert the classification scores into probability values. Specifically, the probability value of each category is equal to the ratio of the indexed score of the category to the sum of the indexed scores of all categories.
[0099] The final generated classification probability distribution represents the probability that the traffic at the current time step belongs to each preset category, and the sum of all probability values is 1.
[0100] Aggregate the classification probability distributions of all time steps (for example, take the average probability of the time steps) and output the final classification probability distribution of the network traffic in the current sliding window.
[0101] For example, if the total number of time steps is 10 and the average probability of the normal protocol category is 0.95, the current traffic is determined to be of the normal protocol type.
[0102] The hidden state vector set is symbolically mapped discretely to construct a cross-scale causal transfer graph, including:
[0103] Based on the preset symbolization rules, the hidden state vector of each time step is mapped to the corresponding discrete symbol to generate a symbol sequence.
[0104] The symbolization rules include dividing the symbol interval according to the statistical characteristics of the hidden state vector or the dynamic time warping results; dividing the symbol interval according to the statistical characteristics of the hidden state vector (such as mean, variance, quantile). For example, each dimension of the hidden state vector is divided into three intervals of high, medium and low according to the quantile, corresponding to the symbols "H", "M" and "L"; the similarity between the hidden state vector and the preset template sequence is calculated by the dynamic time warping algorithm, and the symbols are divided according to the similarity threshold. For example, if the similarity is greater than 0.8, it is mapped to the symbol "A"; 0.5-0.8 is mapped to "B"; less than 0.5 is mapped to "C".
[0105] For example, in symbol generation, assuming that the dimension of the hidden state vector is 64, and each dimension in the statistical feature partitioning method is divided into three intervals according to the ternary digits, then each hidden state vector is mapped to a 64-dimensional symbol vector (such as [H, M, L, ..., H]).
[0106] After mapping the hidden state vector of each time step into symbols according to the above rules, they are arranged in chronological order to generate a symbol sequence. For example, the symbol sequence from time steps 1 to 5 is "A, B, A, C, B".
[0107] The symbol sequence is divided based on different time scales to generate a set of multi-scale symbol subsequences, where each time scale corresponds to a preset time window length.
[0108] Preset multiple time scales (such as hours, days, weeks), each time scale corresponds to a time window length. For example:
[0109] Hourly scale: The time window length is 6 time steps (assuming each time step represents 10 minutes, the window covers 1 hour).
[0110] Day scale: The time window length is 144 time steps (covering 24 hours).
[0111] The symbol sequence is divided into non-overlapping or partially overlapping subsequences according to the length of each time window. For example, the symbol subsequences at the hour scale are [symbol 1-6], [symbol 7-12], etc.
[0112] The subsequence set of each time scale is generated independently to form a multi-scale symbol subsequence set.
[0113] For example, assuming that the length of the original symbol sequence is 1000 and the time window length is 100 (day scale), it is divided into 10 subsequences, each containing 100 symbols.
[0114] A Granger causality test is performed on each symbol subsequence in the multi-scale symbol subsequence set, the conditional probability of transition between adjacent symbols is calculated, and the symbol transition probability matrix corresponding to each time scale is generated.
[0115] The purpose of the test is to determine whether the previous symbol has statistically significant predictive power for the occurrence of the next symbol, thereby determining the causal relationship between the symbols.
[0116] An autoregressive model is built for the sequence of symbol occurrences in the symbol subsequence, for example: ;in, Represents the time step A discrete symbol (such as "A" or "B"), whose value is a symbol in a preset symbol set; Indicates the lag order, that is, the number of historical symbols included in the regression model, which is set according to the actual scenario (for example, 3 means considering the symbols of the first three time steps); Represents the constant term (intercept term) of the regression model, which is used to adjust the baseline prediction value of the model; Indicates the number of the lag term (e.g. 1 corresponds to the sign of the previous time step), is the corresponding regression coefficient, which is used to quantify the previous The symbol of the time step is the current symbol The weight of influence; represents the error term of the regression model.
[0117] Hypothesis test: The F test is used to determine whether all the lag coefficients are jointly significantly different from zero. If the test result rejects the null hypothesis (i.e., p < 0.05), it is considered that the previous sign has a causal effect on the next sign.
[0118] Significance level: p is the preset significance threshold (for example, p=0.05), which is used to determine whether a causal relationship exists.
[0119] According to the Granger test results, the frequency of transitions between symbols is counted. For example, after the symbol "A" appears, the number of times the symbol "B" appears in the next time step is 50 times, and the total number is 100 times, then the transition probability is 0.5.
[0120] The symbol subsequence of each time scale is calculated separately to obtain the matrix , indicating the scale The probability that the symbol at the current time step will be transferred to the symbol at the next time step.
[0121] The symbol transfer probability matrices at each time scale are vertically spliced according to the scale level, and the time scale identifier is marked inside each matrix to generate a cross-scale causal transfer diagram.
[0122] The symbol transfer probability matrices of each time scale are spliced vertically according to the scale level to form a three-dimensional matrix structure; the scale name (such as "hour", "day", "week") is marked on the third dimension (scale axis) of the matrix.
[0123] The rules for generating the graph structure are as follows: each node represents a discrete symbol (such as "A", "B", "C"); the edge represents the cross-scale transition probability between symbols, and the edge weight is determined by the transition probability matrix of the corresponding scale; the time scale identifier and probability value are recorded in the edge attributes. For example, the probability of the edge "A→B" at the hour scale is 0.3 and at the day scale is 0.5.
[0124] Perform persistence homology analysis on cross-scale causal transfer graphs, extract persistence intervals corresponding to one-dimensional persistence homology groups, and calculate topological stability indicators based on interval length distribution, including:
[0125] The cross-scale causal transfer graph is subjected to continuous homology analysis, and the symbolic transfer probability matrix in the cross-scale causal transfer graph is taken as input to construct a multi-scale filtering complex structure.
[0126] The multi-scale filtering complex structure is constructed as follows: the symbols in the symbol transition probability matrix of each time scale are used as vertices, the transition probability between symbols is used as the weight of the edge, and the edges whose edge weights are greater than or equal to the preset filtering parameters are added to the complex structure according to the preset filtering parameters, and the multi-scale filtering sequence is generated according to the hierarchical order of the time scales.
[0127] When performing continuous homology analysis on a cross-scale causal transfer graph, the symbol transfer probability matrix of different time scales in the cross-scale causal transfer graph is first taken as input data.
[0128] The symbol transfer probability matrix is generated as follows: for each time scale, the frequency of transfers between symbols is counted, and the frequency of transfers between symbols is divided by the total number of times the symbols appear to obtain the conditional probability of transfers between symbols, thus forming the symbol transfer probability matrix for each time scale. For example, at the hourly scale, after symbol "A" appears 100 times, symbol "B" appears 50 times in the next time step, then the transfer probability from symbol "A" to symbol "B" is 50 / 100=0.5, and so on to generate a complete symbol transfer probability matrix.
[0129] When constructing a multi-scale filtering complex structure, the symbols in the symbol transition probability matrix of each time scale are used as vertices, the transition probability between symbols is used as the weight of the edge, and the edges are filtered step by step according to the preset filtering parameter ε. Specifically, the value range of the filtering parameter ε is set to 0 to 1, increasing in steps of 0.1. Under each filtering parameter ε, the edges with edge weights greater than or equal to ε are added to the complex structure, while the vertices and edges in the current complex structure are retained.
[0130] For example, when the filtering parameter ε=0.3, only edges with a transition probability greater than or equal to 0.3 are retained; when ε=0.5, edges with a transition probability less than 0.5 are further filtered out, forming complex structures under different ε values. For multi-scale time levels (such as hours, days, and weeks), the corresponding complex structures are generated in sequence according to the hierarchical order of the time scales to form a multi-scale filtering sequence.
[0131] The hierarchical order is arranged as follows: the complex structure at the hourly scale is the first layer, the complex structure at the daily scale is the second layer, the complex structure at the weekly scale is the third layer, and so on, ensuring that the complex structures at different time scales are arranged from fine to coarse in the filtering sequence according to the time scale.
[0132] Based on the multi-scale filtered complex structure, the persistence interval of the one-dimensional persistent homology group is calculated.
[0133] The continuous interval is generated by tracking the generation and disappearance events of the one-dimensional homology group in the multi-scale filtering sequence, recording the starting filter parameters corresponding to each homology generation event and the ending filter parameters corresponding to the disappearance event, and forming a continuous interval.
[0134] First, for each time scale complex structure, the change process of the complex structure is traversed in the order of the filter parameter ε from small to large. During the traversal process, when a new ring structure (i.e., a one-dimensional homology group) appears for the first time, the corresponding filter parameter value is recorded as the starting parameter ε_start of the homology group; when the ring structure is filled and disappears due to the increase of edges, the corresponding filter parameter value is recorded as the end parameter ε_end, forming the continuous interval [ε_start,ε_end) of the homology group.
[0135] For example, in the hourly-scale complex structure, a ring structure is detected when ε=0.2, and the ring structure disappears when ε=0.6, then the corresponding duration interval is [0.2, 0.6). Perform the above traversal operation on the complex structure of all time scales and count all the duration intervals.
[0136] A generation event refers to the filter parameter value when a new homology group appears, and a disappearance event refers to the parameter value when the homology group is covered by a higher filter parameter.
[0137] The topological stability index is calculated based on the length distribution of all persistent intervals.
[0138] For each continuous interval [ε_start, ε_end), calculate its length L = ε_end - ε_start, for example, the length of the interval [0.2, 0.6) is L = 0.6-0.2 = 0.4. Calculate the mean μ and variance σ of all continuous interval length values. The mean is calculated by dividing the sum of all length values by the total number of intervals, and the variance is calculated by dividing the sum of the squares of the differences between each length value and the mean by the total number of intervals.
[0139] The topological stability index is calculated by multiplying the inverse of the mean by the inverse of the variance to obtain the value of the topological stability index. For example, if the mean μ = 0.3 and the variance σ = 0.02, the topological stability index is (1 / 0.3) × (1 / 0.02) = 166.67.
[0140] Through the above steps, the topological characteristics of symbol transfer in the cross-scale causal transfer graph are quantified as stability indicators, which are used as the basis for subsequent network attack detection.
[0141] Based on the KL divergence difference and topological stability index of the classification probability distribution, determine whether there is a covert channel attack, including:
[0142] Calculate the KL divergence difference between the classification probability distribution of the network traffic to be detected and the classification probability distribution of normal network traffic.
[0143] The KL divergence difference value is calculated as follows: the normal network traffic classification probability distribution is used as the baseline distribution, and the probability distribution of the network traffic classification to be detected is used as the comparison distribution. For the probability value of each symbol, the logarithm of the ratio of the baseline distribution probability to the comparison distribution probability is calculated and multiplied by the baseline distribution probability, and the calculation results of all symbols are accumulated and summed.
[0144] Calculate the difference value of the Kullback-Leibler divergence between the classification probability distribution of the network traffic to be detected and the classification probability distribution of normal network traffic. The classification probability distribution of normal network traffic is generated by counting the frequency of occurrence of each symbol in historical normal traffic. The specific method is: collect all network traffic data in the historical time period, extract the symbol sequence in the traffic, count the total number of times each symbol appears in the historical time period, divide the number of occurrences of each symbol by the total number of occurrences of all symbols, and obtain the probability value of each symbol in normal traffic, forming the classification probability distribution of normal network traffic.
[0145] For example, assuming that symbol "A" appears 200 times and symbol "B" appears 300 times in historical traffic, and the total number of symbols is 1000 times, then the probability of symbol "A" is 200 / 1000=0.2, the probability of symbol "B" is 300 / 1000=0.3, and so on to generate a complete probability distribution.
[0146] The generation method of the classification probability distribution of the network traffic to be detected is consistent with the above process, specifically: collect network traffic data within the current detection time period, extract the symbol sequence, count the number of occurrences of each symbol in the current time period, calculate its proportion of the total number of symbols, and obtain the classification probability distribution of the traffic to be detected.
[0147] For example, in the current traffic, symbol "A" appears 150 times, symbol "B" appears 250 times, and the total number of symbols is 800 times. Then the probability of symbol "A" is 150 / 800=0.1875, and the probability of symbol "B" is 250 / 800=0.3125.
[0148] It is worth noting that KL divergence is the Kullback-Leibler divergence.
[0149] When calculating the Kullback-Leibler divergence difference value, the normal network traffic classification probability distribution is used as the baseline distribution, and the probability distribution of the network traffic classification to be detected is used as the comparison distribution. The following operations are performed for each symbol: the probability value of the symbol in the baseline distribution is taken, recorded as P; the probability value of the same symbol in the comparison distribution is taken, recorded as Q; P is calculated multiplied by the logarithmic function log(P / Q) with the natural logarithm as the base to obtain the contribution value of the symbol; the contribution values of all symbols are accumulated and summed to obtain the final Kullback-Leibler divergence difference value.
[0150] For example, the probability of symbol "A" in the reference distribution is P=0.2, and the probability in the comparison distribution is Q=0.1875, so its contribution value is 0.2×log(0.2 / 0.1875)≈0.2×0.0645≈0.0129; the symbol "B" has P=0.3, Q=0.3125, and the contribution value is 0.3×log(0.3 / 0.3125)≈0.3×(-0.0408)≈-0.0122; add up the positive and negative contribution values of all symbols to get the KL divergence difference value.
[0151] It is worth noting that if a symbol has a probability value in the baseline distribution (for example, P>0) and a probability of 0 in the comparison distribution (Q=0), the contribution value of the symbol is set to a preset maximum value (for example, 10^6) to indicate a significant difference between the distributions.
[0152] Based on the comparison result of the KL divergence difference value with the preset first threshold and the comparison result of the topological stability index with the preset second threshold, it is determined whether there is a covert channel attack.
[0153] When the KL divergence difference value is greater than the first threshold and the topological stability index is less than the second threshold, it is determined that a covert channel attack exists.
[0154] Set specific values for the first threshold and the second threshold. The first threshold is the determination threshold of the KL divergence difference value, which is generated by counting the KL divergence difference values of all time periods in the historical normal traffic, taking the maximum value and multiplying it by the safety factor (for example, 1.2). For example, if the historical maximum difference value is 0.5, the first threshold is set to 0.5×1.2=0.6. The second threshold is the determination threshold of the topology stability index, which is generated by counting the minimum value of the topology stability index of the historical normal traffic and multiplying it by the safety factor (for example, 0.8). For example, if the historical minimum value is 300, the second threshold is set to 300×0.8=240. The safety factor is adjusted according to the actual false alarm rate requirements, and is usually a constant between 0.8 and 1.5.
[0155] When finally determining the covert channel attack, the KL divergence difference value of the traffic to be detected is compared with the first threshold, and the topological stability index is compared with the second threshold.
[0156] If the following two conditions are met at the same time, it is determined that a covert channel attack exists: the KL divergence difference value is greater than the first threshold; the topological stability index is less than the second threshold.
[0157] For example, if the KL divergence difference value is 0.65 (greater than 0.6) and the topology stability index is 200 (less than 240), an attack is triggered. If only one of the conditions is met (such as the difference value exceeds the standard but the topology stability is normal), or both conditions are not met, it is judged as normal traffic. Through dual condition constraints, single feature misjudgment can be avoided. For example, temporary network congestion may cause the difference value to increase, but it is excluded because it does not damage the topology stability (the index is normal).
[0158] When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while the time series statistical characteristics of the preset protocol baseline template are updated.
[0159] In response to detecting a covert channel attack, a time interval sequence obeying a Poisson probability density function is generated, and the time interval sequence is used as a random delay interference parameter.
[0160] After a covert channel attack is detected, a random delay interference parameter generation operation is performed. Specifically, the operation includes: setting the expected parameter λ of the Poisson probability density function according to the statistical characteristics of the historical data packet transmission time interval recorded in the network protocol baseline template, where the initial value of λ is 1 / 3 of the mean value of the adjacent data packet transmission time interval in the historical normal traffic. For example, if the mean value of the adjacent data packet interval in the historical normal traffic is 15 milliseconds, then set λ=5 milliseconds.
[0161] By calling the random number generation interface of the computer system, a set of time interval sequences that obey the Poisson probability density function is generated. In the generation process, a uniformly distributed random number in the interval [0,1) is first generated, and the random number is substituted into the inverse function of the cumulative distribution function of the Poisson probability density function to calculate the time interval value that conforms to the Poisson distribution.
[0162] For example, when λ=5 milliseconds, the generated sequence may contain discrete values such as 3 milliseconds, 6 milliseconds, and 4 milliseconds. The length of the time interval sequence is consistent with the number of packets to be sent in the current target data stream. For example, if the target data stream has 100 packets to be sent, a sequence containing 100 time interval values is generated.
[0163] The random delay interference parameter is embedded into the data packet transmission timestamp of the target data stream, so that the actual sending interval of adjacent data packets deviates from the original protocol time baseline after the random delay interference parameter is superimposed.
[0164] When the generated random delay interference parameters are embedded in the target data stream, the following operations are performed in sequence according to the packet sending order: the planned sending timestamp of the current packet in the original protocol time baseline is taken, and the delay value at the corresponding position in the Poisson distribution time interval sequence is superimposed to obtain the actual sending timestamp.
[0165] For example, the original protocol stipulates that the first data packet is sent at T=0 milliseconds and the second data packet is sent at T=10 milliseconds. If the corresponding random delay is 3 milliseconds, the adjusted sending timestamp is T=0+3=3 milliseconds for the first packet and T=10+3=13 milliseconds for the second packet. For each data packet, its timestamp field is modified to the actual value after superposition before sending, and sent through the network interface. This operation makes the actual sending interval of adjacent data packets equal to the sum of the original protocol interval and the random delay interference parameter. For example, if the original interval is 10 milliseconds, if two consecutive delay values are 3 milliseconds and 5 milliseconds respectively, the actual interval becomes 10+3=13 milliseconds and 10+5=15 milliseconds, thereby destroying the temporal regularity that the covert channel relies on.
[0166] The transmission time series of the target data stream after the interference is injected is collected, the corresponding mean and variance are calculated, and the time window sliding average of the preset protocol baseline template is dynamically updated based on the mean and variance.
[0167] After the interference parameter injection is completed, the actual transmission time sequence of the data packets sent in the target data stream is collected in real time. The collection process lasts for at least one complete data stream transmission cycle, for example, collecting the sending timestamps of the latest 100 data packets.
[0168] Based on the collected timestamp sequence, the arithmetic mean and variance of the intervals between adjacent data packets are calculated. For example, the mean of 100 interval values is 16 milliseconds and the variance is 4 milliseconds squared.
[0169] The mean and variance are input into the update algorithm of the preset protocol baseline template, specifically: take the sliding average of the historical time window stored in the protocol baseline template, and weightedly fuse it with the currently calculated mean according to the weight coefficient to generate an updated sliding average.
[0170] The weight coefficient is dynamically adjusted according to the variance. The larger the variance, the higher the weight of the historical value. For example, if the historical sliding average is 15 milliseconds, the current average is 16 milliseconds, and the variance is 4 milliseconds squared, then the historical weight is set to 0.7, the current weight is 0.3, and the updated sliding average is 15×0.7+16×0.3=15.3 milliseconds.
[0171] The expected parameters of the Poisson probability density function in the next period are adjusted according to the updated time window sliding average, so that the interference parameter generation rule keeps synchronization with the dynamic changes of the protocol baseline template.
[0172] According to the updated protocol baseline template sliding average, adjust the expected parameter λ of the Poisson probability density function in the next transmission cycle. The adjustment rule is: set λ to 1 / 3 of the updated sliding average and make a smooth transition with the λ value of the previous cycle.
[0173] For example, if the updated sliding average is 15.3 milliseconds, the initial value of λ in the next cycle is set to 15.3 / 3≈5.1 milliseconds. At the same time, to avoid parameter mutation, a change rate constraint is imposed on λ so that the change range of adjacent cycles does not exceed 20%.
[0174] For example, if the previous cycle λ=5 milliseconds, the next cycle λ value range is 4 milliseconds to 6 milliseconds. The adjusted λ value will be used to generate a new random delay interference parameter sequence, thus forming a closed-loop feedback control. This process continues until the covert channel attack is resolved.
[0175] Embodiment 2: Figure 2 A schematic diagram of the structure of a network security protection system of the present invention is given, and the network security protection system includes:
[0176] Timing vector generation module: It divides the network data packet timing through a sliding window, extracts the difference sequence of the time intervals between adjacent data packets, and performs multi-scale alignment on the difference sequence to generate a standardized timing feature vector.
[0177] Temporal pattern classification module: inputs the temporal feature vector into the bidirectional long short-term memory network for temporal pattern classification, and outputs the hidden state vector set and classification probability distribution.
[0178] Symbolic mapping module: performs symbolic discrete mapping on the set of hidden state vectors to construct a cross-scale causal transfer graph.
[0179] Persistence homology analysis module: Performs persistence homology analysis on cross-scale causal transfer graphs, extracts persistence intervals corresponding to one-dimensional persistence homology groups, and calculates topological stability indicators based on interval length distribution.
[0180] Covert channel detection module: Based on the KL divergence difference and topological stability index of the classification probability distribution, it determines whether there is a covert channel attack.
[0181] Attack interference injection module: When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while updating the time series statistical characteristics of the preset protocol baseline template.
[0182] Example 3
[0183] A network security protection device comprises: a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, a network security protection method is implemented.
[0184] Example 4
[0185] A network security protection medium stores programs or instructions on the medium, and a network security protection method is implemented when the program or instruction is executed by a processor.
[0186] The above formulas are all dimensionless and numerical calculations. The formula is a formula that is closest to the actual situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.
[0187] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state hard disk.
[0188] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0189] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0190] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0192] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0193] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0194] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A network security protection method, characterized in that: The steps include: The network data packet timing is segmented by sliding windows, the difference sequence of the time intervals between adjacent data packets is extracted, and the difference sequence is aligned at multiple scales to generate a standardized timing feature vector. Input the time series feature vector into the bidirectional long short-term memory network to classify the time series pattern, and output the hidden state vector set and classification probability distribution; Perform symbolic discrete mapping on the hidden state vector set and construct a cross-scale causal transfer graph; We conduct persistence homology analysis on cross-scale causal transfer graphs, extract persistence intervals corresponding to one-dimensional persistence homology groups, and calculate topological stability indicators based on interval length distributions. Based on the KL divergence difference and topological stability index of the classification probability distribution, determine whether there is a covert channel attack, including: Calculate the KL divergence difference between the classification probability distribution of the network traffic to be detected and the classification probability distribution of normal network traffic; The KL divergence difference value is calculated as follows: the normal network traffic classification probability distribution is used as the reference distribution, the network traffic classification probability distribution to be detected is used as the comparison distribution, and the probability value of each symbol is calculated by multiplying the logarithm of the ratio of the reference distribution probability to the comparison distribution probability by the reference distribution probability, and the calculation results of all symbols are accumulated and summed; Based on the comparison result of the KL divergence difference value with the preset first threshold value and the comparison result of the topological stability index with the preset second threshold value, determining whether there is a covert channel attack; When the KL divergence difference value is greater than the first threshold and the topological stability index is less than the second threshold, it is determined that a covert channel attack exists; When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while the time series statistical characteristics of the preset protocol baseline template are updated.
2. A network security protection method according to claim 1, characterized in that: The network data packet timing is segmented by sliding windows, the difference sequence of the time intervals between adjacent data packets is extracted, and the difference sequence is aligned at multiple scales to generate a standardized timing feature vector, including: Divide the network data packet arrival time sequence into multiple continuous time windows according to a preset window length, where the window length is dynamically set according to the network protocol type; In each time window, the difference sequence of the time intervals between the arrival of adjacent data packets is calculated, and the difference sequence is the absolute difference between the time intervals between the arrival of adjacent data packets; The difference sequence is mapped to three time scales: milliseconds, seconds, and minutes. At each time scale, the difference sequence is standardized. The standardized difference sequences under three time scales are spliced in chronological order to generate a multi-dimensional time series feature vector.
3. A network security protection method according to claim 1, characterized in that: The time series feature vector is input into the bidirectional long short-term memory network for time series pattern classification, and the hidden state vector set and classification probability distribution are output, including: Construct a bidirectional long short-term memory network, which includes a forward layer and a backward layer. The forward layer processes the time series feature vector in chronological order, and the backward layer processes the time series feature vector in reverse chronological order. The time series feature vector is input into the forward layer to generate a forward hidden state vector sequence, and the time series feature vector is reversely arranged and input into the backward layer to generate a backward hidden state vector sequence; Concatenate the vectors of the same time step in the forward hidden state vector sequence and the backward hidden state vector sequence according to the dimension to form a hidden state vector set; The hidden state vector set is input into the fully connected layer, and the output of the fully connected layer is processed by the normalized exponential function to obtain the classification probability distribution. The classification probability distribution refers to the probability that the current traffic belongs to the normal protocol type.
4. A network security protection method according to claim 1, characterized in that: The hidden state vector set is symbolically mapped discretely to construct a cross-scale causal transfer graph, including: Based on the preset symbolization rules, the hidden state vector of each time step is mapped to the corresponding discrete symbol to generate a symbol sequence; Dividing the symbol sequence based on different time scales to generate a set of multi-scale symbol subsequences, where each time scale corresponds to a preset time window length; Perform Granger causality test on each symbol subsequence in the multi-scale symbol subsequence set, calculate the conditional probability of transition between adjacent symbols, and generate the symbol transition probability matrix corresponding to each time scale; The symbol transfer probability matrices at each time scale are vertically spliced according to the scale level, and the time scale identifier is marked inside each matrix to generate a cross-scale causal transfer diagram.
5. A network security protection method according to claim 1, characterized in that: Perform persistence homology analysis on cross-scale causal transfer graphs, extract persistence intervals corresponding to one-dimensional persistence homology groups, and calculate topological stability indicators based on interval length distribution, including: The cross-scale causal transfer graph is subjected to continuous homology analysis, and the symbolic transfer probability matrix in the cross-scale causal transfer graph is used as input to construct a multi-scale filtering complex structure. The multi-scale filtering complex structure is constructed by taking the symbols in the symbol transition probability matrix of each time scale as vertices, the transition probability between symbols as the weight of the edge, adding the edges whose edge weights are greater than or equal to the preset filtering parameters to the complex structure according to the preset filtering parameters, and generating the multi-scale filtering sequence according to the hierarchical order of the time scales; Based on the multi-scale filtering complex structure, the persistence interval of the one-dimensional persistent homology group is calculated; The continuous interval is generated by tracking the generation and disappearance events of the one-dimensional homology group in the multi-scale filtering sequence, recording the starting filter parameters corresponding to each homology generation event and the ending filter parameters corresponding to the disappearance event, and forming the continuous interval interval; The topological stability index is calculated based on the length distribution of all persistent intervals.
6. A network security protection method according to claim 1, characterized in that: When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while the time series statistical characteristics of the preset protocol baseline template are updated; In response to detecting a covert channel attack, generating a time interval sequence that obeys a Poisson probability density function, and using the time interval sequence as a random delay interference parameter; The random delay interference parameter is embedded into the data packet transmission timestamp of the target data stream, so that the actual sending interval of adjacent data packets deviates from the original protocol time baseline after the random delay interference parameter is superimposed; Collect the transmission time series of the target data stream after the interference is injected, calculate the corresponding mean and variance, and dynamically update the time window sliding average of the preset protocol baseline template based on the mean and variance; The expected parameters of the Poisson probability density function in the next period are adjusted according to the updated time window sliding average.
7. A network security protection system, used to implement a network security protection method according to any one of claims 1 to 6, characterized in that: include: Time series vector generation module: Segment the network data packet time series through a sliding window, extract the difference sequence of the time intervals between adjacent data packets, and perform multi-scale alignment on the difference sequence to generate a standardized time series feature vector; Temporal pattern classification module: inputs the temporal feature vector into the bidirectional long short-term memory network for temporal pattern classification, and outputs the hidden state vector set and classification probability distribution; Symbolic mapping module: symbolic discrete mapping of hidden state vector sets to construct cross-scale causal transfer graphs; Continuous homology analysis module: Perform continuous homology analysis on cross-scale causal transfer graphs, extract the continuous intervals corresponding to the one-dimensional continuous homology group, and calculate the topological stability index based on the interval length distribution; Covert channel detection module: Based on the KL divergence difference and topological stability index of the classification probability distribution, it determines whether there is a covert channel attack; Attack interference injection module: When there is a covert channel attack, random delay interference parameters are generated based on Poisson distribution and injected into the target data stream, while updating the time series statistical characteristics of the preset protocol baseline template.
8. A network security protection device, characterized in that: The device includes: A processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, a network security protection method as described in any one of claims 1 to 6 is implemented.
9. A network security protection medium, characterized in that: The medium stores programs or instructions, and when the programs or instructions are executed by the processor, a network security protection method as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Deep learning-based side channel attack method and system using SpecAugment technology
CN115037437A
Networked system data hidden attack detection method under differential privacy protection
CN115442160A