Network security evaluation system for network security server based on data monitoring
By adopting traffic spectrum analysis, camouflage traffic detection, access path monitoring and traffic density evaluation modules in the network security data monitoring system, the shortcomings in the evaluation of abnormal traffic and traffic density distribution in the prior art are solved, and accurate identification and protection of potential threats are achieved.
Patent Information
- Application Number
- CN202510288544.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to conduct in-depth frequency characteristic analysis in network security data monitoring, which makes it difficult to identify abnormal traffic dependent on surface features, difficult to deal with complex hidden traffic, and difficult to accurately distinguish normal service traffic from potential threat traffic in the security assessment of traffic density distribution.
The network security assessment system based on data monitoring is adopted, and frequency characteristic analysis is performed through the traffic spectrum analysis module, the camouflage traffic detection module builds the traffic transfer matrix, the access path monitoring module classifies the access path, and the traffic density assessment module performs bimodal structure fitting to identify and evaluate abnormal traffic and potential threats.
It enhances the ability to perceive potential threats, realizes accurate identification of camouflage traffic, reduces the misjudgment rate of camouflage access paths, improves the accurate identification of normal access paths, and is more accurate in distinguishing normal service traffic from potential security risk traffic, improving the timeliness and accuracy of overall protection.
Smart Images

Figure CN120151024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security data monitoring, and particularly to a network security evaluation system for network security servers based on data monitoring. Background Art
[0002] Network security data monitoring is an integral part of network security. By performing real-time monitoring, analysis, and management on data flows in the network environment, it ensures the security, integrity, and availability of the system. The field covers aspects such as traffic monitoring, threat detection, abnormal behavior identification, data leakage prevention, etc. Through various technical means, network data is collected and analyzed to quickly discover and prevent potential security threats.
[0003] Among them, the network security evaluation system for network security servers is a network security protection system designed specifically for servers. By monitoring and managing the network data of the servers, the system can analyze the data traffic of the servers in real time, identify potential security threats and abnormal behaviors, and thus take countermeasures before or at the initial stage of a security incident.
[0004] In the network security protection of existing technologies for data monitoring, it is difficult to perform in-depth frequency feature analysis on the traffic of servers, resulting in over-reliance on surface features for the identification of abnormal traffic, which is not conducive to dealing with complex and concealed traffic. For example, in the face of high-frequency disguised traffic, it is easy to cause missed reports of potential threats. In addition, existing technologies are difficult to accurately perceive minor abnormalities in state transitions in the dynamic monitoring of traffic feature changes. The monitoring method of access paths may ignore some stable access features or mislabel normal traffic as abnormal when dealing with complex path accesses. Finally, for the security evaluation of traffic density distribution, existing technologies are difficult to accurately distinguish normal business traffic and potential threat traffic in complex traffic density structures, which easily affects the traffic security judgment and early warning effect of the system. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and a network security evaluation system for network security servers based on data monitoring is proposed.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: The network security evaluation system for network security servers based on data monitoring includes:
[0007] The traffic spectrum analysis module collects the original traffic data of the network security server to generate a time series, performs uniformly spaced sampling, constructs a spectrogram of network traffic time segment data, sorts by frequency amplitude in descending order, analyzes frequently occurring behavior features, and obtains preliminary abnormal traffic features;
[0008] The disguised traffic detection module extracts feature data from the traffic packets of the network security server according to the preliminary abnormal traffic features, determines the transition probability between states, constructs a traffic transition matrix, and uses the feature state sequence to update the transition probability of the state in real time to obtain the current traffic state matrix. The current traffic state matrix is compared with the state items in the traffic transition matrix to identify the difference points and generate disguised traffic features;
[0009] The access path monitoring module classifies normal traffic and disguised traffic by path according to the difference point information in the disguised traffic features, deletes the disguised access paths, establishes a data set of the normal access paths of the network security server, and generates an access path stability analysis result by analyzing the node information in the normal access paths;
[0010] The traffic density evaluation module performs a bimodal structure fitting on the traffic density data of the normal access paths according to the access path stability analysis result, distinguishes the traffic peaks of normal business activities from the potential security risk traffic peaks, and generates a network traffic security evaluation result.
[0011] As a further solution of the present invention, the specific steps for obtaining the spectrogram of the network traffic time segment data are as follows:
[0012] Based on the original traffic data generated by the server, the traffic data is sampled at a fixed time interval, and by aggregating the data points within each time interval into a sampling point, time segment data for frequency analysis is generated;
[0013] According to the time segment data for frequency analysis, the data in the time domain is converted into the frequency domain. By sorting and analyzing the frequency amplitudes in the spectrogram, the frequency components and frequency nodes are identified, and the spectrogram of the network traffic time segment data is constructed.
[0014] As a further solution of the present invention, the specific steps for obtaining the frequently occurring behavior features are as follows:
[0015] Based on the spectrogram of the network traffic time segment data, the amplitudes of each frequency component are extracted and sorted. An amplitude determination threshold is set, and the frequency nodes exceeding the amplitude determination threshold are screened. The frequency and amplitude information thereof are recorded as the high-frequency behavior feature nodes of this time segment, and a filtered frequency node set is generated;
[0016] Based on the filtered frequency node set, the frequently occurring behavior features are analyzed, and the formula:
[0017]
[0018] Calculate the concentration degree P of high-frequency abnormal behaviors a , and obtain the preliminary abnormal traffic features;
[0019] Among them, A i is the amplitude of the high-amplitude frequency node i selected, m is the total number of high-amplitude frequency nodes exceeding the amplitude determination threshold, n is the total number of frequency nodes in the current time segment, and j represents each target node among all frequency nodes.
[0020] As a further solution of the present invention, the acquisition steps of the constructed traffic transfer matrix are specifically as follows:
[0021] Real-time collect network traffic packets, extract the feature data of the size, time interval, and request frequency of the traffic packets, sort the feature data of the traffic, and combine the information of the preliminary abnormal traffic features to obtain the state sets of normal traffic and abnormal traffic;
[0022] According to the normal traffic data in the state sets of the normal traffic and abnormal traffic, use the formula:
[0023]
[0024] Calculate the probability P'(s a transferring to the next state s a+1 ) of the state s in the characteristic state sequence, and obtain the state transition probability of the feature; a →s a+1 ) to obtain the state transition probability of the feature;
[0025] Among them, s a represents the current feature state, s a+1 represents the next feature state, and respectively represent the feature data of the current state s a and the next state s a+1 ; N is the total number of items representing the feature data, D is the normalization constant, and k is the serial number of the feature data item;
[0026] According to the state transition probability of the feature, fill the transition probability values of the size, time interval, and request frequency feature data into the corresponding positions according to the state indexes of rows and columns in the matrix to construct a traffic transfer matrix.
[0027] As a further solution of the present invention, the acquisition steps of the disguised traffic features are specifically as follows:
[0028] Continuously extract the feature data of traffic packets in the network traffic, including size, time interval, and request frequency, form a new feature state sequence in chronological order, dynamically update the transition probability of each pair of adjacent states, and generate the current traffic state matrix;
[0029] Compare the current traffic status matrix with the status items in the traffic transfer matrix, and identify the status items with significant differences according to a preset fluctuation range to obtain the characteristics of disguised traffic.
[0030] As a further solution of the present invention, the step of obtaining the data set of the normal access path of the network security server is specifically as follows:
[0031] Based on the characteristics of the disguised traffic, extract the information of each access path from the network traffic, including the source address, target address, access frequency of path nodes, and response delay. Compare the path data with the characteristics of the disguised traffic to divide it into normal access paths or disguised access paths, and obtain the classified path data set.
[0032] Based on the classified path data set, delete the data of the disguised access path, and save the characteristic information of the normal access path to obtain the data set of the normal access path of the network security server.
[0033] As a further solution of the present invention, the step of obtaining the analysis result of the access path stability is specifically as follows:
[0034] According to the data set of the normal access path of the network security server, use the formula:
[0035]
[0036] Calculate the path reliability score R to obtain the path reliability information.
[0037] Among them, Q is the occurrence frequency of nodes in the current path, Q min is the minimum value of the node frequency in the normal access path, Q max is the maximum value of the node frequency in the normal access path, T is the request success rate of the current path, T min is the minimum value of the request success rate of the normal access path, T max is the maximum value of the request success rate of the normal access path, H is the response delay of the current path, H min is the minimum response delay of the normal access path, H max is the maximum response delay of the normal access path;
[0038] According to the path reliability information, compare the score with the standard score, judge the stability and reliability status of the current path, record the path score that meets the reliability standard, and store it in the normal access path set to generate the analysis result of the access path stability.
[0039] As a further solution of the present invention, the step of obtaining the network traffic security evaluation result is specifically as follows:
[0040] According to the analysis result of the access path stability, collect the traffic density distribution data on the normal access path, and use the formula:
[0041]
[0042] Calculate the probability density value f(x) at the data point x to obtain the distribution characteristics of the traffic peak;
[0043] where x is the traffic density value to be analyzed, μ is the mean value, representing the central position of the traffic density data set, and σ 2 is the variance, representing the degree of data dispersion;
[0044] According to the distribution characteristics of the traffic peak, extract the mean value and variance of each peak, compare the characteristics of the two peaks, including the central position in the data set, the distribution range, and the deviation degree, and obtain the network traffic security assessment result by judging whether the mean difference between the peaks is significant.
[0045] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0046] In the present invention, the frequency characteristics of the traffic time series are mined through spectrum analysis, and the high-frequency abnormal nodes are screened out by constructing a spectrogram through Fourier transform operations, enhancing the perception ability of potential threats. The extraction of the feature state sequence and its matrix representation of the transition probability enable the system to update the traffic state in real time in a dynamic traffic environment, keenly locate the abnormal change points in the camouflaged traffic, and achieve the accurate identification of the camouflaged traffic. Through the comprehensive analysis of the path nodes, request success rate, and response delay, the system can accurately evaluate the stability and reliability of the access path, reduce the misjudgment rate of the camouflaged access path, and improve the accurate identification of the normal access path. Based on the bimodal structure analysis of the traffic density distribution of the normal access path, the system is more accurate in distinguishing between normal service traffic and potential security risk traffic, improving the timeliness and accuracy of the overall protection. Brief Description of the Drawings
[0047] Figure 1 is the system flow chart of the present invention;
[0048] Figure 2 is the flow chart of constructing the spectrogram of the network traffic time segment data of the present invention;
[0049] Figure 3 is the flow chart of analyzing the frequently occurring behavior characteristics of the present invention;
[0050] Figure 4 is the flow chart of constructing the traffic transfer matrix of the present invention;
[0051] Figure 5 is the flow chart of generating the camouflaged traffic characteristics of the present invention;
[0052] Figure 6 A flow chart of a data set for establishing a normal access path of a network security server in the present invention;
[0053] Figure 7 A flow chart for generating access path stability analysis results for the present invention;
[0054] Figure 8 A flow chart for generating network traffic security assessment results for the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0056] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0057] See also Figure 1 , the network security assessment system for network security servers based on data monitoring includes:
[0058] The traffic spectrum analysis module collects the original traffic data of the network security server to generate a time series, obtains the time distribution characteristics of the traffic data, detects and removes the noise data item by item, and samples the traffic time series data at uniform intervals, calls the data of each time segment for Fourier transform operation, constructs a spectrum diagram of the network traffic time segment data, sorts it by frequency amplitude, scans the high-amplitude frequency nodes in the spectrum diagram according to the frequency amplitude judgment threshold, analyzes the frequently occurring behavior characteristics, and obtains preliminary abnormal traffic characteristics;
[0059] The disguised traffic detection module constructs a state set of normal traffic and abnormal traffic based on the preliminary abnormal traffic characteristics, extracts the characteristic data of size, time interval, and request frequency in the traffic packets of the network security server, sorts the extracted characteristic data to form a characteristic state sequence, calculates the probability of state changes in the characteristic state sequence to determine the transition probability between states, and represents the transition probability between all possible states in a matrix form to construct a traffic transition matrix. During the real-time monitoring of the traffic characteristic data of the network security server, the transition probability of the state is updated in real time using the characteristic state sequence to obtain the current traffic state matrix. The current traffic state matrix is compared with the state items in the traffic transition matrix to identify the difference points and generate the disguised traffic characteristics;
[0060] The access path monitoring module classifies the normal traffic and the disguised traffic by path according to the difference point information in the disguised traffic characteristics to distinguish the normal access path and the disguised access path, deletes the disguised access path, establishes a data set of the normal access path of the network security server, and calculates the reliability score of the normal access path by analyzing the nodes in the normal access path, including the occurrence frequency of servers, routers, or gateways, the path request success rate, and the delay characteristics of the path response, provides a reference basis for the path performance, and generates the access path stability analysis result;
[0061] The traffic density evaluation module statistically analyzes the traffic density of the normal access path according to the normal access path information in the access path stability analysis result at fixed time intervals, collects the traffic density distribution data on the normal access path, fits the traffic density data of the normal access path with a bimodal structure by calculating the mean and variance of the peaks, and conducts a peak feature comparison analysis to further distinguish the traffic peaks of normal business activities from the potential security risk traffic peaks, and generates the network traffic security evaluation result;
[0062] The preliminary abnormal traffic characteristics include characteristic nodes with frequency amplitudes exceeding the threshold, frequently occurring high-frequency behavior patterns, and significant time distribution difference information. The disguised traffic characteristics include abnormal transition probability items different from the normal state, specific size distributions of abnormal traffic packets, and time interval information that does not conform to the normal access frequency. The access path stability analysis result includes the occurrence frequency of high-stability path nodes, the success rate evaluation of path requests, and the statistical characteristics of response time delays. The network traffic security evaluation result includes the mean and variance of the normal traffic peaks in the bimodal traffic density, the frequency distribution characteristics of potential risk traffic peaks, and the traffic density difference analysis between normal business activities and security risks.
[0063] Please refer to Figure 2 , and the specific steps for obtaining the spectrogram of the network traffic time segment data are as follows:
[0064] Based on the original traffic data generated by the server, the traffic data is sampled at fixed time intervals. By aggregating the data points within each time interval into a single sampling point, time segment data for frequency analysis is generated;
[0065] First, the original traffic data generated by the server needs to be sampled and processed. During the specific execution process, first, the traffic data per second is collected through a traffic monitoring tool (such as Wireshark or Snort) to form an original data sequence containing timestamps. To ensure the representativeness of the data and the accuracy of the analysis, the collected original data often contains noise data. Therefore, a filtering method needs to be used for data cleaning. For example, the median filtering or mean filtering method is used to remove the noise. The cleaned data will be sampled at uniform intervals. Assuming that the sampling is set to once every 10 seconds, it means that the traffic data within every 10 seconds generates a sampling point through statistical methods (such as the average value or median), forming a traffic data sequence at 10 - second intervals. This can simplify the data volume, retain the main traffic trends, and reduce unnecessary minor fluctuations.
[0066] According to the time segment data for frequency analysis, the data in the time domain is converted to the frequency domain. By sorting and analyzing the frequency amplitudes in the spectrogram, the frequency components and frequency nodes are identified, and a spectrogram of the network traffic time segment data is constructed;
[0067] First, the time segment data for frequency analysis (such as the traffic data sample every 5 minutes) is imported into the Fast Fourier Transform (FFT) to decompose the data in the time domain into a series of frequency components, revealing different frequency components in the signal. During the specific execution, the sampling data points included in each time segment (such as 600 data points, sampled twice per second for 5 minutes) are input into the FFT. The FFT will decompose these data points into corresponding frequency components, including amplitude and phase information, generating a spectrogram corresponding to this time segment. Assuming that each time segment has 600 data points, after the FFT conversion, 300 frequency components are obtained, with the frequency range from 0Hz to 5Hz (half of the sampling frequency). Each frequency component has a corresponding amplitude, indicating the intensity of that frequency component. According to the output result of the FFT, a spectrogram is plotted, where the horizontal axis represents the frequency (usually from 0Hz to the maximum frequency, depending on half of the sampling rate), and the vertical axis represents the amplitude. The spectrogram can intuitively display the amplitude distribution of different frequency components within the time segment.
[0068] Please refer to Figure 3 , the specific steps for obtaining the frequently occurring behavioral characteristics are as follows:
[0069] Based on the spectrogram of network traffic time segment data, extract the amplitudes of each frequency component and sort them. Set an amplitude determination threshold and filter out the frequency nodes whose amplitudes exceed the amplitude determination threshold. Record their frequency and amplitude information as the high-frequency behavior feature nodes of this time segment, and generate a set of filtered frequency nodes;
[0070] Based on the spectrogram of the generated network traffic time segment data, obtain the amplitudes of each frequency component. For example, assume that the frequency components included in the spectrogram of a time segment are 2Hz, 3Hz, 5Hz, and 8Hz, and the corresponding amplitudes are 12, 20, 8, and 15 respectively. After sorting the amplitudes of these frequency components in descending order, the order is: 3Hz (amplitude 20), 8Hz (amplitude 15), 2Hz (amplitude 12), 5Hz (amplitude 8), ensuring that the most significant frequency components are ranked at the front. Subsequently, set an amplitude determination threshold, for example, set the threshold to 10, and filter out the frequency nodes with higher amplitudes according to this threshold. At this time, the amplitudes of 3Hz, 8Hz, and 2Hz all exceed the threshold. Filter out these frequency nodes and record their frequency and amplitude information as the feature nodes of high-frequency behavior. Through this filtering process, the high-amplitude frequency nodes that frequently appear in this time segment can be locked, and these nodes may reflect the burst traffic characteristics in the network.
[0071] Based on the set of filtered frequency nodes, analyze the frequently occurring behavior characteristics, and use the formula:
[0072]
[0073] Calculate the concentration P of high-frequency abnormal behavior a , and obtain the preliminary abnormal traffic characteristics;
[0074] Among them, P a reflects the concentration of the filtered high-amplitude frequency nodes in the entire time segment. The higher its value, the more concentrated the high-frequency abnormal behavior is, which may represent an abnormal traffic pattern. A i is the amplitude of the high-amplitude frequency node i selected. The amplitude A i reflects the intensity of the frequency component in the traffic data. In the spectrogram of network traffic time segment data, the amplitude of each frequency can be observed. The high-amplitude frequency node means that this frequency component is more significant in the traffic. m is the total number of high-amplitude frequency nodes that exceed the amplitude determination threshold, n is the total number of all frequency nodes in the current time segment, and j is used to represent each target node among all frequency nodes.
[0075] For example, if the amplitudes of the selected high-amplitude frequency nodes are 30, 25, and 20 respectively, and the total amplitude of all frequency nodes in this time segment is 150, the calculation process is as follows:
[0076]
[0077] Calculation result P a = 0.5 indicates that the high-amplitude frequency nodes in this time segment occupy 50% of the total amplitude. By comparing with historical data, for example, the concentration of normal traffic is usually in the range of 20% - 30%. This significantly higher value indicates that there are obvious abnormal traffic behaviors in this time segment.
[0078] Please refer to Figure 4 , and the specific steps for obtaining the traffic transfer matrix are as follows:
[0079] Real-time collect network traffic packets, extract the characteristic data of the size, time interval, and request frequency of the traffic packets, sort the characteristic data of the traffic, and combine the information of the preliminary abnormal traffic characteristics to obtain the state sets of normal traffic and abnormal traffic;
[0080] Use a network monitoring tool (such as Wireshark) to collect network traffic packets in real time, extract the core characteristic data of the traffic packets, including size (bytes), time interval (milliseconds), and request frequency (requests per second), and construct an initial state set of normal traffic. The specific operation is to sort the extracted normal traffic characteristic data (such as size, time interval, request frequency) in ascending order to form a structured characteristic state sequence. This can ensure that the characteristics of normal traffic maintain a stable distribution pattern. After forming the initial state sets of normal traffic and preliminary abnormal traffic, by comparing the characteristic state sequences of the two, the behavior patterns of normal traffic and abnormal traffic are initially identified and distinguished. For example, assume that the time interval characteristics of normal traffic are sorted as 200, 220, 230, 240, 250 milliseconds, and the preliminary abnormal traffic characteristics include high-amplitude frequency nodes, and the corresponding time intervals show frequent low-value fluctuations (such as 150, 180, 170 milliseconds, etc.). Then, the deviation of abnormal traffic in the time interval characteristic can be observed. Through this differential characteristic sequence comparison, the initial state set can be used to identify and distinguish the behavior patterns of normal traffic and preliminary abnormal traffic.
[0081] According to the normal traffic data in the state sets of normal traffic and abnormal traffic, use the formula:
[0082]
[0083] Calculate the probability P'(s a transferring to the next state s a+1 ), and obtain the state transition probability of the characteristic; a →s a+1 ), and obtain the state transition probability of the characteristic;
[0084] Among them, P'(s a→s a+1 ) is used to measure the change significance among various characteristics (such as size, time interval, request frequency) of network traffic packets, s a and s a+1 are adjacent states in the characteristic state sequence, s a represents the current characteristic state, s a+1 represents the next characteristic state. These states can be obtained by analyzing network traffic data, and the state sequence is obtained through the extraction and sorting of characteristic data. and are the characteristic values of each state, representing the characteristic data of the current state s a and the next state s a+1 respectively. For example, in the characteristic data of network traffic packets, specific values such as size, time interval, or request frequency can be extracted as characteristic values, and the characteristic values of each traffic packet can be directly read through a real-time traffic monitoring tool. For example, the size of each traffic packet is recorded within a set time segment. N represents the total number of characteristic data items, which is the number of characteristics of traffic packets extracted within a specified time segment. D is a constant used for result normalization, representing the upper limit of the value range of normal traffic characteristic data, ensuring the comparability of transition probabilities in different time segments, obtained through the analysis of the value range of historical normal traffic data. For example, the value of D is set by statistically analyzing the upper limit of the characteristic data of traffic packet size, time interval, and request frequency under normal conditions. k is the serial number of the characteristic data item, used to traverse all characteristic data items within a specified time segment.
[0085] If 5 traffic packets are monitored within a time segment, the obtained characteristic data is as follows: Size characteristic data: 450 bytes, 500 bytes, 520 bytes, 480 bytes, 510 bytes; Time interval characteristic data: 200 milliseconds, 220 milliseconds, 240 milliseconds, 230 milliseconds, 250 milliseconds; Request frequency characteristic data: 50 times, 52 times, 55 times, 51 times, 53 times.
[0086] For the size characteristic data, set D = 600 (based on the maximum range of traffic packet size under normal conditions). Calculate the transition probabilities of each adjacent state as follows:
[0087]
[0088] For the time interval characteristic data, set D = 300 (based on the maximum value of the time interval under normal conditions). Calculate as follows:
[0089]
[0090] For the request frequency characteristic data, set D = 100 (based on the maximum value of the normal request frequency). Calculate as follows:
[0091]
[0092] The results show that the transition probability of the size feature data is 1.8, which is converted to a percentage of 180%. When compared with the historical benchmark value (usually between 100% - 120%), it can be seen that the change range of the current size feature data exceeds the normal range, showing large fluctuations, which may indicate abnormal fluctuations in the traffic. The transition probability of the time interval feature data is 0.87, which is converted to a percentage of 87%. When compared with the lower limit of the benchmark range (usually between 80% - 100%), it indicates that the change of the time interval is relatively stable, conforming to the normal traffic fluctuation situation. The transition probability of the request frequency feature data is 0.066, which is converted to a percentage of 6.6%. When compared with the benchmark range (usually between 10% - 15%), it indicates that the change of the request frequency in this time segment is extremely small, showing high stability.
[0093] According to the state transition probability of the features, the transition probability values of the size, time interval, and request frequency feature data are filled into the corresponding positions according to the state indexes of rows and columns in the matrix to construct a traffic transition matrix;
[0094] By matrix-representing the transition probabilities between all possible states, a normal traffic transition matrix is constructed. Based on the existing characteristic state sequence, according to the calculation results of the transition probabilities, the probability values of state changes are filled into the matrix item by item. For example, assume that the calculated adjacent transition probability of the size characteristic state sequence is 1.8, the transition probability of the time interval characteristic data is 0.87, and the transition probability of the request frequency characteristic data is 0.066. Using these transition probability values, a complete transition matrix can be constructed for the characteristic state sequence. Construct the matrix structure with each item in the state sequence as the labels of rows and columns. For example, assume there are 5 states, then the transition matrix will be a 5×5 matrix. Each row represents a state, and the probability values of transferring from this state to other states are filled as column values. Substitute the transition probability values of the size, time interval, and request frequency characteristic data into the corresponding row and column positions in the matrix. For the matrix of size characteristic data, the rows and columns represent the indices of specific size states, and each element is filled with the transition probability of adjacent states. Taking the transition probability of 1.8 of the size data in paragraph 2 as an example, in the size transition matrix, the transition probabilities from 450 bytes to 480 bytes and from 480 to 500 bytes in the sequence are filled into the matrix in turn until the transition probabilities of all states are filled. After completing the matrix filling, the transition probability matrices of each characteristic (size, time interval, request frequency) are obtained. Statistical analysis software (such as the Numpy library in Matlab or Python) can be used to store and analyze the transition matrix, and the matrix is saved in a two-dimensional array format for subsequent comparative analysis. The generated traffic transition matrix provides a reference standard for the state changes in the current time segment.
[0095] Please refer to Figure 5 , and the steps for obtaining the characteristics of the disguised traffic are specifically as follows:
[0096] Continuously extract the characteristic data of the traffic packets in the network traffic, including size, time interval, and request frequency, form a new characteristic state sequence in chronological order, dynamically update the transition probability of each pair of adjacent states, and generate the current traffic state matrix;
[0097] First, it is necessary to continuously extract the characteristic data of each traffic packet from the network traffic, including size, time interval, and request frequency, and form a dynamic characteristic state sequence in chronological order. The system analyzes each newly generated characteristic state sequence and records the changes in adjacent states in real time. Then, for each pair of adjacent states in this sequence, the transition probability is updated. The system will dynamically adjust the probability of each state transition, that is, P'(s a →s a+1), so that the current traffic status matrix always reflects the latest traffic characteristics. This process is continuously repeated to ensure that in each time segment when data flows in, the changes in the feature status sequence are updated in real time in the traffic status matrix, so that the matrix always maintains the latest status.
[0098] Compare the current traffic status matrix with the status items in the traffic transfer matrix, and identify the status items with large differences according to the preset fluctuation range to obtain the disguised traffic characteristics;
[0099] Compare each state item in the current traffic status matrix with the corresponding state items in the previously constructed traffic transfer matrix item by item: Assume that the transfer probability matrix of the size feature data in the current time segment and the standard traffic transfer matrix (the matrix representing the normal traffic pattern) are as shown below. Current traffic status matrix (size feature data): The transfer probability from state 450 bytes to 500 bytes is 0.6, the transfer probability from state 500 bytes to 480 bytes is 0.4, and the transfer probability from state 480 bytes to 520 bytes is 0.7. Standard traffic transfer matrix (size feature data): The transfer probability from state 450 bytes to 500 bytes is 0.5, the transfer probability from state 500 bytes to 480 bytes is 0.5, and the transfer probability from state 480 bytes to 520 bytes is 0.5. The comparison process is to compare each state transfer probability in the current traffic status matrix with the corresponding value in the standard matrix item by item: Compare the transfer probability from state 450 bytes to 500 bytes. The current value is 0.6, while the standard value is 0.5, and the difference is 0.1; Compare the transfer probability from state 500 bytes to 480 bytes. The current value is 0.4, while the standard value is 0.5, and the difference is 0.1; Compare the transfer probability from state 480 bytes to 520 bytes. The current value is 0.7, while the standard value is 0.5, and the difference is 0.2. According to the comparison results, identify the state items with relatively large difference values (such as the transfer probability difference from state 480 bytes to 520 bytes being 0.2 is marked as a significantly deviated item). Usually, in traffic detection, for minor changes in normal traffic characteristics (for example, the difference value is below 0.1), it can be considered as normal network fluctuations. A difference of 0.2 has exceeded this normal fluctuation range, indicating that the traffic status in the current time segment significantly deviates from the normal pattern. A difference value of 0.2 means that there is a 40% relative increase in the transfer frequency (from 0.5 to 0.7), which represents an unusual characteristic change in network traffic monitoring. Disguised traffic usually shows abnormal increases or decreases in frequency to interfere with the detection of normal network behavior. Therefore, this deviation in the increase may belong to the disguised characteristics. For example, assume that in the normal traffic pattern, a transfer probability of 0.5 from 480 bytes to 520 bytes indicates that this state change is relatively stable, while in the current traffic, the transfer probability rises to 0.7, meaning that the 480 - byte state suddenly transfers to the 520 - byte state more frequently in certain specific situations. This sudden increase in frequency is inconsistent with the behavior pattern of normal traffic, suggesting that there may be disguised traffic trying to conceal its true intention. Therefore, the difference value of 0.2 is marked as a significant change, exceeding the normal deviation range, showing a sudden increase in the transfer frequency, indicating that there may be disguised or abnormal characteristics in the traffic characteristics in the current time segment.
[0100] Please refer to Figure 6 , the specific steps for obtaining the data set of the normal access path of the network security server are as follows:
[0101] Based on the characteristics of disguised traffic, extract the information of each access path from network traffic, including source address, destination address, access frequency of path nodes, and response latency. Compare the path data with the characteristics of disguised traffic to classify it as a normal access path or a disguised access path, and obtain the classified path dataset;
[0102] First, extract the path information in the traffic through network traffic monitoring to obtain the specific characteristic data of each path, including the traffic source address, destination address, access frequency of nodes in the path, response latency, and success rate of path requests, etc. Then, perform matching according to the predefined disguised feature parameters, compare the path data item by item, mark the paths that conform to the characteristics of disguised traffic as disguised access paths, and mark the paths that do not conform to the disguised characteristics and show stable and abnormal traffic characteristics as normal access paths. In this process, the system judges whether each path belongs to a normal or disguised access path by comparing the node frequencies and response characteristics in the path item by item and contrasting the difference points in the disguised feature library. After classification, separate and record the disguised path data, and mark the normal paths as data to be stored in the database for further analysis.
[0103] Based on the classified path dataset, delete the data of disguised access paths and save the characteristic information of normal access paths to obtain the data set of normal access paths of the network security server;
[0104] After distinguishing the disguised access paths and normal access paths, the system clears the data of disguised access paths and retains the set of normal access paths. Subsequently, format the data set of normal access paths, including detailed characteristics such as the information of each node, occurrence frequency, and response latency of each path. By recording the paths, parse the data set of normal access paths item by item, count the node access frequency, success rate, and average response latency of each path, and record these parameters in turn to establish the data set of normal access paths of the network security server. The system summarizes and stores the path information in the data set to ensure that the subsequent path performance evaluation and traffic control strategies are executed based on normal traffic characteristics.
[0105] Please refer to Figure 7 , and the specific steps for obtaining the analysis results of access path stability are as follows:
[0106] According to the data set of normal access paths of the network security server, use the formula:
[0107]
[0108] Calculate the path reliability score R to obtain the path reliability information;
[0109] Among them, Q is the occurrence frequency of nodes in the current path, representing the relative frequency of access to the server, router, or gateway node in this path within a specified time period. The acquisition method usually involves collecting data through traffic monitoring tools and calculating the proportion of the access count of each node to the total access count. For example, in an actual network monitoring environment, the access frequency of nodes can be recorded through node logs and traffic monitoring systems, and then the proportion is calculated as the occurrence frequency of the node, Q min is the minimum value of the node frequency in the normal access path, representing the lowest occurrence frequency of nodes in the path dataset, Q max is the maximum value of the node frequency in the normal access path, indicating the maximum value of the node frequency in normal traffic. T is the request success rate of the current path, representing the success ratio of all requests in the path within a specific time period, which is obtained by statistical analysis of network request status code data. The success or failure of each request can be analyzed through the log records of the traffic monitoring system, and then the ratio of the number of successful requests to the total number of requests is calculated as the value of the request success rate of the current path, T min is the minimum value of the request success rate of the normal access path, which is the lowest request success rate extracted from the normal path data, T max is the maximum value of the request success rate of the normal access path, representing the maximum standard of the normal path success rate, which is the highest request success rate value extracted from the normal path data. H is the response delay of the current path, indicating the response time of each node (server, router, gateway, etc.) in this path, which is usually obtained through network performance monitoring tools. The response delay of this path can be obtained by measuring the round-trip time of test requests, H min is the minimum response delay of the normal access path, which is the lowest response time extracted from the normal path data, H max is the maximum response delay of the normal access path, representing the maximum value of the response time in the normal path.
[0110] For example, the following data is obtained through statistical analysis of normal traffic historical data:
[0111] The node frequency Q of the current path = 0.85, the minimum value Q of the node frequency in the historical record min = 0.6, and the maximum value Q max = 1.0. The request success rate T of the current path = 0.9, the minimum value T of the success rate in the historical record min = 0.7, and the maximum value T max = 1.0. The response delay H of the current path = 0.15 seconds, the minimum value H of the response delay in the historical record min = 0.05 seconds, and the maximum value H max = 0.5 seconds.
[0112] Normalization of node frequency:
[0113]
[0114] Normalize the request success rate:
[0115]
[0116] Normalize the response delay:
[0117]
[0118] Calculate the reliability score R:
[0119] R = 0.625 + 0.667 - 0.222 = 1.07
[0120] The result shows that the normalized reliability score R = 1.07.
[0121] According to the path reliability information, compare the score with the historical standard score to judge the stability and reliability status of the current path, record the path scores that meet the reliability standards, and store them in the normal access path set to generate the analysis result of the access path stability;
[0122] Based on the obtained path reliability score R = 1.07, the system will compare this score with the historical standard score to determine the stability and reliability status of the current path. Suppose the historical standard score is 1.0 and the current path score is 1.07, which is higher than the standard score. This indicates that the status of the current path shows good reliability during the monitored time period, and the access frequency, request success rate, and response delay of each node in the path are in a relatively stable state. For the specific analysis of this score, first, both the node access frequency and the request success rate are higher than the low threshold range in the system, which means that the access frequency and request success rate of each node in the path conform to the expected normal traffic characteristics; second, the response delay is small, indicating that the overall response speed of the network is fast. Based on the above characteristics, it can be concluded that the current path not only meets the standards of normal access paths but also is relatively higher than the benchmark level. Therefore, this path can be considered reliable and does not require further optimization or processing. Such paths with scores higher than the standard can be recorded as stable paths for reference during subsequent path screening and stored in the normal access path set as a reference data set for normal traffic, providing a discriminant basis for future traffic evaluation and path anomaly detection. At the same time, for paths with similar scores higher than 1.0, they can be directly classified as stable paths to reduce unnecessary consumption of further monitoring resources.
[0123] Please refer to Figure 8 , the specific steps for obtaining the network traffic security assessment result are as follows:
[0124] According to the analysis result of the access path stability, collect the traffic density distribution data on the normal access path and use the formula:
[0125]
[0126] Calculate the probability density value f(x) at the data point x to obtain the distribution characteristics of the traffic peak;
[0127] Among them, x is the traffic density value to be analyzed, which can be used for comparison of different traffic peak characteristics and risk judgment. μ is the mean value, representing the central position of the traffic density data set. Through the formula calculate, x i′ is the i'-th data point in the data set, n' is the total number of data points, and σ 2 is the variance, representing the degree of data dispersion. Through the formula calculate, representing the deviation degree of the data point relative to the mean value. When the variance is large, the traffic density fluctuates greatly. The variance parameter is obtained by collecting traffic data during normal time periods and is used to judge traffic stability.
[0128] Assume that the collected traffic density data is x i′ = [10, 12, 14].
[0129] Mean calculation:
[0130]
[0131] Variance calculation:
[0132]
[0133] Substitute the mean value μ = 12 and the variance σ 2 = 2.67 into the formula:
[0134]
[0135] Calculate the probability density of a specific value:
[0136] The result shows that when x = 10, f(10) ≈ 0.0885; when x = 12, f(12) ≈ 0.2443; when x = 14, f(14) ≈ 0.0885.
[0137] According to the distribution characteristics of the traffic peak, extract the mean value and variance of each peak, compare the characteristics of the two peaks, including the position in the data set, the distribution range, and the deviation degree. By judging whether the mean difference between the peaks is significant, obtain the network traffic security assessment result;
[0138] Based on the probability density value f(x) of the calculated bimodal structure of the traffic density, the mean and variance results of each peak are brought in for peak feature comparison and analysis. First, two main peak features in the bimodal are extracted. For example, the first peak (which may correspond to normal business activities) has a mean μ = 12 and a variance σ 2 = 2.67, and its density distribution function f(x) is 0.0885, 0.2443, and 0.0885 at x = 10, 12, and 14 respectively; while it is assumed that the second peak has a mean μ = 18 and a variance σ 2 = 4.5, and its density function values are also calculated at the corresponding positions. For example, the density values at x = 16, 18, and 20 are 0.065, 0.209, and 0.065 respectively. In the peak feature comparison and analysis, by comparing the means and variances of the two peaks, it can be observed that the data of the first peak (mean 12) is more concentrated, the distribution is compact, and the density function value is higher near the mean, which may correspond to normal business activities; while the second peak (mean 18) has greater fluctuations, a higher variance, and the mean is more deviated from the typical value of normal business activities, which may correspond to potential risk traffic. To further confirm that these differences are statistically significant, a significance test (such as the T-test) is used for verification. The process of the T-test includes calculating the mean and sample standard error of each peak, and then evaluating whether the difference between the two means is large enough to exclude the possibility of random fluctuations based on these values. The T-test generates a p-value, representing whether the difference between the means of the two peaks is statistically significant. Usually, when the p-value is less than the set significance level (such as 0.05), it can be considered that the difference between the means of the two peaks is significant. This indicates that the characteristics of the first peak and the second peak are different. The first peak may represent normal business traffic, while the second peak may correspond to potential abnormal or risk traffic. After the significance test is completed, the comparison results and test data are stored in the traffic security database to generate the final network traffic security assessment result. Through these results, the traffic peaks of normal business activities and the traffic peaks of potential security risks can be marked, providing a reliable basis for security monitoring and risk response.
[0139] The above is only a preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A network security assessment system for a network security server based on data monitoring, characterized in that: The system comprises: The traffic spectrum analysis module collects the original traffic data of the network security server to generate a time series, performs uniform interval sampling, constructs a spectrum diagram of the network traffic time segment data, sorts it by frequency amplitude, analyzes the frequently occurring behavior characteristics, and obtains preliminary abnormal traffic characteristics; The disguised traffic detection module extracts characteristic data from the network security server traffic packet according to the preliminary abnormal traffic characteristics, determines the transition probability between states, constructs a traffic transfer matrix, updates the state transition probability in real time using the characteristic state sequence, obtains the current traffic state matrix, compares the current traffic state matrix with the state items in the traffic transfer matrix, identifies the difference points, and generates disguised traffic characteristics; The access path monitoring module classifies the normal traffic and the disguised traffic by path according to the difference point information in the disguised traffic characteristics, deletes the disguised access path, establishes a data set of the normal access path of the network security server, and generates the access path stability analysis result by analyzing the node information in the normal access path; The traffic density assessment module performs bimodal structure fitting on the traffic density data of the normal access path according to the access path stability analysis result, distinguishes the traffic peak of normal business activities from the traffic peak of potential security risks, and generates a network traffic security assessment result.
2. The network security evaluation system for a network security server based on data monitoring according to claim 1 is characterized in that: The steps for obtaining the spectrum diagram of network traffic time segment data are specifically as follows: Based on the original traffic data generated by the server, the traffic data is sampled at fixed time intervals, and the data points in each time interval are aggregated into a sampling point to generate time segment data for frequency analysis; According to the time segment data used for frequency analysis, the time domain data is converted into the frequency domain, and the frequency components and frequency nodes are identified by sorting and analyzing the frequency amplitudes in the spectrum graph, so as to construct a spectrum graph of the network traffic time segment data.
3. The network security evaluation system for a network security server based on data monitoring according to claim 2 is characterized in that: The steps for obtaining the frequently occurring behavior features are specifically as follows: Based on the spectrum diagram of the network traffic time segment data, the amplitude of each frequency component is extracted and sorted, an amplitude determination threshold is set and frequency nodes exceeding the amplitude determination threshold are screened, and their frequency and amplitude information are recorded as high-frequency behavior feature nodes of the time segment, and a screened frequency node set is generated; Based on the filtered frequency node set, the frequently occurring behavior characteristics are analyzed using the formula: Calculate the concentration P of high-frequency abnormal behavior a , get the preliminary abnormal traffic characteristics; Among them, A i is the amplitude of the screened high-amplitude frequency node i, m is the total number of high-amplitude frequency nodes exceeding the amplitude determination threshold, n is the total number of frequency nodes in the current time segment, and j represents each target node in all frequency nodes.
4. The network security evaluation system for a network security server based on data monitoring according to claim 3 is characterized in that: The acquisition steps for constructing the traffic transfer matrix are specifically as follows: Collect network traffic packets in real time, extract the size, time interval and request frequency feature data of the traffic packets, sort the feature data of the traffic, and combine the information of the preliminary abnormal traffic features to obtain the state set of normal traffic and abnormal traffic; According to the normal flow data in the state set of the normal flow and the abnormal flow, the formula is adopted: Calculate the state s in the characteristic state sequence a Transition to next state s a+1 The probability P'(s a →s a+1 ), and obtain the state transition probability of the feature; Among them, s a Indicates the current feature state, s a+1 represents the next feature state, x sa and x sa+1 Respectively represent the current state s a and the next state s a+1 The characteristic data of, N is the total number of characteristic data items, D is a normalized constant, and k is the sequence number of the characteristic data item; According to the state transition probability of the feature, the transition probability values of the size, time interval and request frequency feature data are filled into corresponding positions in the matrix according to the state index of the row and column to construct a traffic transfer matrix.
5. The network security evaluation system for a network security server based on data monitoring according to claim 4 is characterized in that: The steps for obtaining the camouflaged traffic characteristics are specifically as follows: Continuously extract traffic packet feature data from network traffic, including size, time interval, and request frequency, form a new feature state sequence in chronological order, dynamically update the transition probability of each pair of adjacent states, and generate the current traffic state matrix; The current traffic state matrix is compared with the state items in the traffic transfer matrix, and the state items with large differences are identified according to a preset fluctuation range to obtain the camouflaged traffic characteristics.
6. The network security evaluation system for a network security server based on data monitoring according to claim 5, characterized in that: The steps for acquiring the data set for establishing the normal access path of the network security server are specifically as follows: Based on the camouflaged traffic features, information of each access path is extracted from the network traffic, including the source address, the destination address, the access frequency of the path node, and the response delay, and the path data is compared with the camouflaged traffic features to be classified into a normal access path or a camouflaged access path, thereby obtaining a classified path data set; Based on the classified path data set, the disguised access path data is deleted, the characteristic information of the normal access path is saved, and the data set of the normal access path of the network security server is obtained.
7. The network security evaluation system for a network security server based on data monitoring according to claim 6 is characterized in that: The steps for obtaining the access path stability analysis result are specifically as follows: According to the data set of the normal access path of the network security server, the formula is adopted: Calculate the path reliability score R and obtain the path reliability information; Among them, Q is the frequency of occurrence of the node in the current path, Q min is the minimum value of the node frequency in the normal access path, Q max is the maximum value of the node frequency in the normal access path, T is the request success rate of the current path, and T min is the minimum value of the request success rate of the normal access path, T max is the maximum value of the request success rate of the normal access path, H is the response delay of the current path, and H min is the minimum response delay of the normal access path, H max is the maximum response delay of the normal access path; According to the path reliability information, the score is compared with the standard score to determine the stability and reliability status of the current path, the path score that meets the reliability standard is recorded and stored in the normal access path set to generate an access path stability analysis result.
8. The network security evaluation system for a network security server based on data monitoring according to claim 7 is characterized in that: The steps for obtaining the network traffic security assessment result are specifically as follows: According to the access path stability analysis results, the traffic density distribution data on the normal access path is collected, using the formula: Calculate the probability density value f(x) at the data point x to obtain the distribution characteristics of the traffic peak; Among them, x is the flow density value to be analyzed, μ is the mean, which indicates the center position of the flow density data set, and σ 2 is the variance, which indicates the degree of data dispersion; According to the distribution characteristics of the traffic peaks, the mean and variance of each peak are extracted, and the characteristics of the two peaks are compared, including the location in the data set, the distribution range and the degree of deviation. By judging whether the mean difference between the peaks is significant, the network traffic security assessment result is obtained.
Citation Information
Cited By
Industrial internet data encryption transmission method, system and device and storage medium
CN120614182A