Real-time data security monitoring system based on online data modeling
By constructing a real-time data security monitoring system based on online data modeling, the system can identify and respond to low-intensity probing attacks, solving the problem of interference with the model's discrimination boundary and improving the self-repair and continuous defense capabilities of the real-time data security monitoring system.
Patent Information
- Application Number
- CN202511448114.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-11
AI Technical Summary
In existing real-time data security monitoring systems, malicious actors interfere with the model's boundary judgment by using low-intensity probing samples, which reduces the model's sensitivity to identifying real attack behaviors and affects the system's long-term stability and attack defense capabilities.
A real-time data security monitoring system based on online data modeling is adopted, including modules for data acquisition, feature extraction, feature space construction, offset analysis, attenuation detection, and risk assessment. It identifies feature variation trends caused by low-intensity probing attack samples and triggers a security reconstruction mechanism.
It enhances the sensitivity and self-recovery capability of the real-time data security monitoring system to potential induced attacks, and strengthens the system's robustness and continuous defense capability in dynamic environments.
Smart Images

Figure CN120956523A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and more specifically, to a real-time data security monitoring system based on online data modeling. Background Technology
[0002] In existing real-time data security monitoring, online learning models generally rely on continuously input data for dynamic updates to adapt to changes in the network environment. However, malicious actors can gradually interfere with the model's discrimination boundaries by continuously injecting low-intensity probing samples, causing the model to mistakenly learn abnormal behavior patterns as new "normal distributions" during regular training cycles. This leads to a gradual decrease in the model's sensitivity to identifying real attack behaviors and an irreversible decline in its recognition ability, severely impacting the long-term stability and attack defense capabilities of online security monitoring systems.
[0003] To address the aforementioned problems, a technical solution is provided. Summary of the Invention
[0004] In order to overcome the above-mentioned deficiencies of the prior art, embodiments of the present invention provide a real-time data security monitoring method and system based on online data modeling to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A real-time data security monitoring system based on online data modeling includes;
[0007] Data acquisition module: Acquires real-time network stream data, extracts features from the real-time network stream data, and outputs feature sequences that characterize the behavior patterns of the network stream data;
[0008] Feature Space Module: Constructs a feature space based on feature sequences, generating feature drift mapping sequences and abnormal behavior recognition sensitivity change trend sequences;
[0009] Migration Analysis Module: Employs anomaly distribution migration analysis method to analyze feature drift mapping sequences and identify feature variation trends caused by low-intensity probing attack samples;
[0010] Attenuation Detection Module: This module uses an anomaly sensitivity attenuation detection method to analyze the trend sequence of changes in the sensitivity of anomaly behavior recognition, and identifies the decay characteristics of the real-time security monitoring classifier's ability to recognize real attack events.
[0011] Risk assessment module: Based on feature variation trends and decay characteristics, assess the degree of risk of induced drift in the real-time safety monitoring classifier;
[0012] Reconstruction Trigger Module: Based on the risk level of induced drift in the real-time security monitoring classifier, determine the risk level of the real-time data security monitoring system and trigger the security reconstruction mechanism of the real-time security monitoring classifier.
[0013] In a preferred embodiment, real-time network stream data is collected, features are extracted from the real-time network stream data, and a feature sequence characterizing the behavioral patterns of the network stream data is output, specifically as follows:
[0014] Collect real-time network stream data from network communication links, perform multi-dimensional feature extraction operations on the real-time network stream data, and construct the original feature matrix of behavior patterns;
[0015] The original feature matrix of the behavior pattern is standardized and reconstructed into a time series, forming a continuous feature sequence according to a preset time window sliding strategy.
[0016] In a preferred embodiment, a feature space is constructed based on the feature sequence, and a feature drift mapping sequence and an abnormal behavior recognition sensitivity change trend sequence are generated, specifically:
[0017] By performing embedding mapping processing based on distribution features on continuous feature sequences, the continuous feature sequences are projected into a multi-dimensional feature space to construct a feature space that characterizes the evolution process of real-time network flow data behavior patterns.
[0018] In the feature space, the distribution difference of continuous feature sequences within adjacent time windows is calculated, and a feature drift mapping sequence is generated using a mapping function based on the distribution distance metric.
[0019] In the feature space, the output response of the abnormal behavior recognition model is monitored. Combined with the recognition response records of historical real attack samples, the recognition sensitivity change parameters are calculated, and a sequence of abnormal behavior recognition sensitivity change trends is constructed based on time sequence.
[0020] In a preferred embodiment, anomaly distribution shift analysis is used to analyze the feature drift mapping sequence and identify the feature variation trend caused by low-intensity probing attack samples, specifically:
[0021] The feature drift mapping sequence is divided into multiple consecutive analysis windows, and the statistical distribution characteristic parameters of the feature drift mapping sequence within each analysis window are calculated.
[0022] Based on the statistical distribution characteristic parameters of the feature drift mapping sequence in different analysis windows, a distribution offset index sequence is constructed to describe the variation law of feature distribution.
[0023] Anomaly distribution offset analysis is used to decompose the distribution offset index sequence into trends. By comparing the similarity between the distribution offset index sequence and the feature distribution changes of historical anomaly attack samples, the feature variation trend caused by low-intensity probing attack samples can be identified.
[0024] In a preferred embodiment, an anomaly sensitivity decay detection method is used to analyze the trend sequence of changes in the sensitivity of anomaly behavior recognition, and to identify the decay characteristics of the real-time security monitoring classifier's ability to recognize real attack events, specifically:
[0025] The abnormal behavior recognition sensitivity change trend sequence is divided into multiple continuous analysis periods according to a predetermined detection window, and the statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence are calculated in each analysis period.
[0026] Based on the statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence and the difference between the preset abnormal behavior recognition sensitivity benchmark sequence, a sensitivity decay index sequence is constructed.
[0027] An abnormal sensitivity decay detection method was used to perform time-series trend analysis on the sensitivity decay index sequence to identify the decay characteristics of the real-time security monitoring classifier's ability to identify real attack events.
[0028] In a preferred embodiment, the risk of induced drift in the real-time security monitoring classifier is assessed based on feature variation trends and decay characteristics, specifically as follows:
[0029] The feature variation magnitude index is calculated based on the feature variation trend caused by low-intensity exploratory attack samples.
[0030] Based on the decay characteristics of the real-time security monitoring classifier's ability to identify real attack events, an index of sensitivity decay is calculated.
[0031] A comprehensive index of induced drift risk is calculated by combining the characteristic variation amplitude index and the sensitivity decay index.
[0032] In a preferred embodiment, the risk level of the real-time data security monitoring system is determined based on the degree of risk of induced drift in the real-time security monitoring classifier, and a security reconstruction mechanism for the real-time security monitoring classifier is triggered, specifically as follows:
[0033] Pre-set the risk level classification thresholds for the real-time data security monitoring system;
[0034] Based on the comparison results between the comprehensive risk index of induced drift and the risk level classification threshold, the risk level of the real-time data security monitoring system is classified.
[0035] Based on the risk level of the real-time data security monitoring system, the corresponding real-time security monitoring classifier's security reconstruction mechanism is triggered.
[0036] The technical effects and advantages of the real-time data security monitoring system based on online data modeling of this invention are as follows:
[0037] The data acquisition module ensures the continuity and feature representation capabilities of network flow data; the feature space module, by constructing a feature space, captures the evolution of network behavior patterns; the offset analysis module identifies feature variation trends caused by low-intensity probing attacks, helping to discover potential induced behaviors; the decay detection module identifies the decline in the real-time security monitoring classifier's ability to recognize real attack events, promptly detecting a decrease in its ability to identify real attacks; the risk assessment module integrates feature variation trends and decay characteristics to comprehensively assess the risk level of induced drift in the real-time security monitoring classifier; and the reconstruction trigger module determines the risk level of the real-time data security monitoring system based on the risk level and triggers the security reconstruction mechanism of the real-time security monitoring classifier. This enhances the real-time data security monitoring system's sensitivity, interpretability, and self-recovery capabilities against potential induced attacks, and strengthens its robustness and continuous defense capabilities in dynamic environments. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the structure of the real-time data security monitoring system based on online data modeling of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0040] Example 1
[0041] Figure 1 The present invention provides a real-time data security monitoring system based on online data modeling, comprising:
[0042] Data acquisition module: Acquires real-time network stream data, extracts features from the real-time network stream data, and outputs feature sequences that characterize the behavior patterns of the network stream data;
[0043] Feature Space Module: Constructs a feature space based on feature sequences, generating feature drift mapping sequences and abnormal behavior recognition sensitivity change trend sequences;
[0044] Migration Analysis Module: Employs anomaly distribution migration analysis method to analyze feature drift mapping sequences and identify feature variation trends caused by low-intensity probing attack samples;
[0045] Attenuation Detection Module: This module uses an anomaly sensitivity attenuation detection method to analyze the trend sequence of changes in the sensitivity of anomaly behavior recognition, and identifies the decay characteristics of the real-time security monitoring classifier's ability to recognize real attack events.
[0046] Risk assessment module: Based on feature variation trends and decay characteristics, assess the degree of risk of induced drift in the real-time safety monitoring classifier;
[0047] Reconstruction Trigger Module: Based on the risk level of induced drift in the real-time security monitoring classifier, determine the risk level of the real-time data security monitoring system and trigger the security reconstruction mechanism of the real-time security monitoring classifier.
[0048] Collect real-time network stream data, extract features from the real-time network stream data, and output feature sequences characterizing the behavioral patterns of the network stream data, including:
[0049] Collect real-time network stream data from network communication links, perform multi-dimensional feature extraction operations on the real-time network stream data, and construct the original feature matrix of behavior patterns;
[0050] A network communication link is a channel used for information transmission between network devices, including data transmission paths between servers, routing devices, and switching devices. A data acquisition unit is installed at the entry point of the network communication link. The data acquisition unit is a data acquisition device capable of real-time copying and transmission of network data. The data acquisition unit is configured to capture and copy network stream data packets according to a specific transmission protocol; the captured data packets are the real-time network stream data. Real-time network stream data includes transmission protocol type information, packet length information, packet arrival time information, connection source port information, connection destination port information, number of data packets, and time interval information between data packets.
[0051] The collected real-time network stream data is processed using a multi-dimensional feature extraction method. This method involves extracting the corresponding transmission protocol type, data packet length, data packet arrival time, connection source port, connection destination port, number of data packets, and time interval between data packets for each real-time network stream data packet. Statistical indicators for each feature are calculated, including mean, maximum, minimum, median, and variance, forming a set of statistical parameters for each feature dimension.
[0052] The average value of the data packet length information dimension feature is the sum of the length values of all data packets within the collection period divided by the number of data packets; the variance of the data packet length information dimension feature is the sum of the squares of the differences between each data packet length value and the average data packet length divided by the number of data packets; the average value of the time interval between data packets dimension feature is the sum of the arrival time intervals of adjacent data packets divided by the number of data packet intervals within the collection period, and the statistical indicators of other dimension features are calculated in the same way as above.
[0053] After extracting features from multiple dimensions, a set of statistical parameters for each real-time network stream data packet in each feature dimension is obtained. The set of statistical parameters for each data packet is then used to construct an original feature matrix for behavioral patterns. Each row of the original feature matrix represents network stream data from one acquisition period, and each column represents different statistical parameters for each feature dimension.
[0054] The original feature matrix of the behavior pattern is standardized and reconstructed over time to form a continuous feature sequence according to a preset time window sliding strategy;
[0055] The original feature matrix of the behavior pattern is processed using a standardization method. The standardization process is as follows: for all data in each column of the original feature matrix of the behavior pattern, first calculate the mean and standard deviation of the data in each column, then subtract the mean of the data in that column from each data point, and then divide by the standard deviation of the data in that column to obtain the standardized data matrix of the original feature matrix of the behavior pattern.
[0056] Time series reconstruction is performed based on a standardized data matrix. Specifically, the time series reconstruction involves employing a sliding window strategy on the standardized data matrix, according to a pre-defined time window length and sliding step size. The sliding window strategy involves sliding down the standardized data matrix row by row along the time sequence, starting from the first row, with each row movement being the pre-defined sliding step size. All data covered by each sliding window constitutes the time series data for that window. The time window length is set to a fixed length for multiple rows of data, and the sliding step size is less than or equal to the time window length to ensure overlap or continuity between consecutive windows.
[0057] For example, if the preset time window length is ten data acquisition cycles and the sliding step size is two data acquisition cycles, then the first time window is the data from row 1 to row 10 in the data matrix, the second time window is the data from row 3 to row 12, and so on, with a certain degree of overlap between each window. Through the sliding window strategy, a continuous feature sequence that reflects the dynamic changes in the behavior patterns of network stream data is formed.
[0058] A feature space is constructed based on the feature sequence, generating a feature drift mapping sequence and an abnormal behavior recognition sensitivity change trend sequence, including:
[0059] By performing embedding mapping processing based on distribution features on continuous feature sequences, the continuous feature sequences are projected into a multi-dimensional feature space to construct a feature space that characterizes the evolution process of real-time network flow data behavior patterns.
[0060] Continuous feature sequences use data sequences from multiple time windows as basic units. Each time window contains multiple feature dimensions of the original feature matrix of behavioral patterns after standardization. When performing embedding mapping on continuous feature sequences, an embedding mapping method based on statistical distribution features is used to project the features.
[0061] The embedding mapping method based on statistical distribution features is as follows: First, the statistical distribution features of all feature dimensions within each time window of the continuous feature sequence are calculated. The statistical distribution features include mean, variance, skewness, and kurtosis. These statistical distribution features are combined to form a feature vector, and a multidimensional space mapping algorithm is used to project the feature vector onto the constructed multidimensional feature space to obtain a feature space that represents the time-varying behavior patterns of network flow data.
[0062] In the feature space, the distribution difference of continuous feature sequences within adjacent time windows is calculated, and a feature drift mapping sequence is generated using a mapping function based on the distribution distance metric.
[0063] In a multidimensional feature space, to evaluate the changes in real-time network streaming data behavior patterns between adjacent time windows, the statistical distribution difference of feature vectors between adjacent time windows is calculated. The calculation of this statistical distribution difference uses a mapping function based on a distribution distance metric. The mapping function is a function for calculating the statistical distribution distance between two feature vectors in the feature space, using the Euclidean distance between the feature vectors.
[0064] For example, feature vectors within two adjacent time windows correspond to two specific locations in the feature space. The distribution distance is calculated as follows: calculate the difference between the values of each dimension of the first feature vector and the values of the same dimension of the second feature vector, square the differences respectively, sum them up, and then take the square root of the sum to obtain the distribution distance value. This process is repeated for all adjacent feature vectors within multiple consecutive time windows to obtain the distribution distance value between each pair of adjacent feature vectors, forming a feature drift mapping sequence.
[0065] In the feature space, the output response of the abnormal behavior recognition model is monitored. Combined with the recognition response records of historical real attack samples, the recognition sensitivity change parameters are calculated, and a sequence of abnormal behavior recognition sensitivity change trends is constructed based on the time sequence.
[0066] Based on the feature space, an anomaly behavior recognition model is deployed to identify abnormal behaviors in real-time network flow data. The anomaly behavior recognition model is a classification and recognition model pre-trained using historical real attack samples. Historical real attack samples are clearly labeled attack data, including feature sequences of historical network flow data where the attack type is clearly defined and has actually occurred. The anomaly behavior recognition model is deployed in the feature space. The feature vector of each time window is input into the anomaly behavior recognition model, and the model outputs an anomaly recognition result in real time. The output response includes recognition category information for whether the anomaly behavior detection is true or false, i.e., anomaly behavior category or normal behavior category.
[0067] To calculate the sensitivity variation parameters, the anomaly detection model's response records corresponding to historical real attack samples are first extracted. These records include accuracy, false negative rate, and false positive rate. For the output response of the anomaly detection model in the current time window, the accuracy, false negative rate, and false positive rate are calculated. Accuracy is defined as the ratio of the sum of samples correctly identified as anomalous behavior and samples correctly identified as normal behavior by the anomaly detection model within the current time window to the total number of input samples in the current time window. False negative rate is defined as the ratio of the number of samples that are actually anomalous behavior but were not identified by the anomaly detection model within the current time window to the total number of actual anomalous behavior samples within the current time window. False positive rate is defined as the ratio of the number of samples that are actually normal behavior but were incorrectly identified as anomalous behavior by the anomaly detection model within the current time window to the total number of actual normal behavior samples within the current time window. Next, the recognition accuracy, false negative rate, and false positive rate of the current time window are compared with the recognition accuracy, false negative rate, and false positive rate of historical real attack samples. The changes in recognition accuracy, false negative rate, and false positive rate are calculated respectively. These changes are used as parameters for the change in recognition sensitivity.
[0068] The sensitivity variation parameters of multiple consecutive time windows are arranged sequentially in chronological order to form a sequence of abnormal behavior recognition sensitivity variation trends. This sequence characterizes the changing ability of the abnormal behavior recognition model to recognize abnormal behavior as real-time network stream data changes.
[0069] Anomaly distribution shift analysis is used to analyze feature drift mapping sequences and identify feature variation trends caused by low-intensity probing attack samples, including:
[0070] The feature drift mapping sequence is divided into multiple consecutive analysis windows, and the statistical distribution characteristic parameters of the feature drift mapping sequence within each analysis window are calculated.
[0071] The feature drift mapping sequence is divided into several analysis windows of equal length. The length of the analysis window is usually determined based on the stability requirements of the actual network data feature changes; for example, ten consecutive feature drift mapping values can be used as one analysis window.
[0072] Within each analysis window, statistical distribution parameters are calculated for the feature-shifted mapping sequence. These parameters include the mean, variance, skewness, and kurtosis of the feature-shifted mapping sequence within the current analysis window. These statistical parameters characterize the central tendency, dispersion, deviation from symmetry, and concentration of extreme values of the data within the current analysis window, respectively.
[0073] Based on the statistical distribution characteristic parameters of the feature drift mapping sequence in different analysis windows, a distribution offset index sequence is constructed to describe the variation law of feature distribution.
[0074] The differences in statistical distribution characteristics between adjacent analysis windows are measured, and the distribution offset index of statistical distribution characteristics between each pair of adjacent analysis windows is calculated. The calculation of the distribution offset index includes a distance calculation method between statistical distribution parameters. The distance calculation method is to calculate the difference between each index of statistical distribution characteristics in the two analysis windows, square the difference, and then sum the squared values and take the arithmetic square root of the sum. The distribution offset indices between all adjacent analysis windows are arranged in chronological order to form a distribution offset index sequence.
[0075] Anomaly distribution offset analysis is used to decompose the distribution offset index sequence into trends. By comparing the similarity between the distribution offset index sequence and the feature distribution changes of historical anomaly attack samples, the feature variation trend caused by low-intensity probing attack samples can be identified.
[0076] The abnormal distribution migration analysis method used is trend decomposition analysis. Specifically, trend decomposition analysis involves: extracting the trend of the distribution migration index sequence and decomposing it into a long-term trend component and a volatility component; the long-term trend component represents the long-term trend of the distribution migration index sequence over a time window, while the volatility component represents the short-term fluctuations of the distribution migration index sequence over a time window.
[0077] The implementation method of trend decomposition analysis is as follows: First, the long-term trend component of the distribution offset index sequence is extracted using the moving average method; then, the fluctuation component of the distribution offset index sequence is obtained by subtracting the value of the long-term trend component from the original value of the distribution offset index sequence. The moving average method is to sum the values of multiple data points before and after each data point and then divide by the total number of data points.
[0078] The data on the changes in the feature distribution of historical anomaly attack samples is used for comparison. The data on the changes in the feature distribution of historical anomaly attack samples refers to the feature drift distribution offset features corresponding to historically known attack events, and the trend decomposition process is the same as above.
[0079] The comparison process involves calculating the similarity between the trend components of the current distribution offset index sequence and the trend components of historical anomaly attack samples. The similarity is calculated by summing the squares of the differences between the current trend component's value at each time point and the corresponding value of the historical trend component, and then taking the square root. This similarity calculation determines whether the current trend component exhibits a long-term trend similar to that of historical anomaly attack samples.
[0080] For example, if the calculated trend similarity value is lower than the preset similarity threshold, it means that the current feature variation trend is similar to the feature variation trend of historical low-intensity probing attack samples, thereby identifying the feature variation trend caused by the current low-intensity probing attack samples.
[0081] The feature variation trend is the long-term trend component of the distribution offset index sequence after trend decomposition, representing the long-term change trend of real-time network flow data features under the continuous influence of low-intensity probing attacks.
[0082] An anomaly sensitivity decay detection method was used to analyze the trend sequence of changes in the sensitivity of anomaly behavior recognition, and to identify the decay characteristics of the real-time security monitoring classifier's ability to recognize real attack events, including:
[0083] The abnormal behavior recognition sensitivity change trend sequence is divided into multiple continuous analysis periods according to a predetermined detection window, and the statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence are calculated in each analysis period.
[0084] The abnormal behavior recognition sensitivity change trend sequence consists of the recognition accuracy change value, false negative rate change value, and false positive rate change value of multiple consecutive time windows, reflecting the change of the abnormal recognition model's ability to recognize abnormal behavior in real-time network stream data over time.
[0085] The sequence of changes in sensitivity for abnormal behavior detection is divided into multiple continuous analysis periods by a detection window length. The detection window length is a fixed length pre-set based on the observation requirements of real-time security monitoring for changes in sensitivity for abnormal behavior detection, typically using the sensitivity change parameters of multiple consecutive time windows as a single detection window. For example, the detection window length could be set to the sequence data of ten consecutive sensitivity change parameters.
[0086] Within each analysis period, the statistical characteristics of the changing trend sequence of abnormal behavior recognition sensitivity are calculated. These statistical characteristics include the mean and variance of the changes in recognition accuracy, the mean and variance of the changes in false negative rate, and the mean and variance of the changes in false positive rate.
[0087] Statistical calculations were performed for each of the above analysis periods to obtain the statistical characteristics of the abnormal behavior recognition sensitivity change sequence for all analysis periods.
[0088] Based on the statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence and the difference between the preset abnormal behavior recognition sensitivity benchmark sequence, a sensitivity decay index sequence is constructed.
[0089] The abnormal behavior identification sensitivity benchmark sequence is a standard identification sensitivity statistical index calculated based on the identification response records of historical real attack samples. It includes the mean and variance benchmark values of the change in identification accuracy, the mean and variance benchmark values of the change in false negative rate, and the mean and variance benchmark values of the change in false positive rate.
[0090] The process of constructing an index for the degree of sensitivity decay during a single analysis period is as follows:
[0091] The statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence within each analysis period are compared with the corresponding statistical characteristic benchmark values of the abnormal behavior recognition sensitivity benchmark sequence. Specifically, the differences between the mean and variance of the recognition accuracy change, the mean and variance of the false negative rate change, and the mean and variance of the false positive rate change within the current analysis period and their corresponding benchmark values are calculated.
[0092] The differences are squared, summed, and the square root is taken to obtain the sensitivity decay index for each analysis period. A higher single-analysis-period sensitivity decay index indicates a more significant difference between the current anomaly detection model's ability and the historical baseline, reflecting the degree of degradation in detection capability. The single-analysis-period sensitivity decay indices for each analysis period are combined to obtain a sensitivity decay index sequence.
[0093] An abnormal sensitivity decay detection method was used to perform time-series trend analysis on the sensitivity decay index sequence to identify the decay characteristics of the real-time security monitoring classifier's ability to identify real attack events.
[0094] The abnormal sensitivity decay detection method is a time-series trend analysis method, which combines trend decomposition and trend testing to identify decay characteristics.
[0095] The process of time series trend analysis is as follows:
[0096] A trend decomposition method is used on the sensitivity decay index sequence to decompose it into a long-term trend component and a short-term fluctuation component. The trend decomposition method adopts the moving average method, which is to divide the sum of the values of each data point and its two adjacent data points in the sensitivity decay index sequence by the number of data points taken, thereby obtaining the smooth long-term trend component and then the short-term fluctuation component.
[0097] A trend test is performed on the long-term trend component. The trend test method is to determine whether the value of the long-term trend component gradually increases with the increase of the time analysis period. Specifically, the slope of the trend change of the long-term trend component over multiple consecutive analysis periods is calculated, and the trend slope is calculated through linear trend regression analysis.
[0098] If the calculated trend slope is greater than the preset positive slope threshold, it indicates that the real-time security monitoring classifier's ability to identify real attack events is declining, thus identifying the decline characteristics of the real-time security monitoring classifier.
[0099] For example, if, after trend decomposition, the long-term trend component of the sensitivity decay index sequence shows a monotonically increasing trend over a continuous analysis period, and the calculated trend slope is a clearly positive value exceeding a preset positive slope threshold, then it is determined that the abnormal behavior recognition capability of the real-time security monitoring classifier has significantly decreased over time, identifying the decay characteristic of the real-time security monitoring classifier's ability to recognize real attack events. The decay characteristic of the real-time security monitoring classifier's recognition capability is manifested in a decrease in the recognition accuracy of the abnormal behavior recognition model and an increase in the false negative or false positive rate over a continuous analysis period.
[0100] Based on feature variation trends and decay characteristics, the risk of induced drift in real-time security monitoring classifiers is assessed, including:
[0101] The feature variation magnitude index is calculated based on the feature variation trend caused by low-intensity exploratory attack samples.
[0102] The feature variation magnitude index is used to quantify the degree to which real-time network flow data features change after being affected by low-intensity probing attacks. The feature variation trend is the long-term trend component caused by the identified low-intensity probing attack samples, reflecting the long-term change status of network flow data features within the analysis window.
[0103] The method for calculating the characteristic variation amplitude index is as follows:
[0104] The difference between the trend values of each analysis window in the feature variation trend sequence is calculated. Specifically, the absolute value of the difference in trend values between consecutive analysis windows is calculated, and the absolute values of the difference in trend values between all consecutive analysis windows are summed to obtain the cumulative sum of trend value differences. Then, the sum of trend value differences is divided by the total number of analysis windows to obtain the feature variation amplitude index. The feature variation amplitude index is used to quantitatively represent the cumulative change in the degree of feature shift of real-time network flow data under the continuous influence of low-intensity probing attack samples. Trend values refer to the values of specific data points in the long-term trend component.
[0105] Based on the decay characteristics of the real-time security monitoring classifier's ability to identify real attack events, an index of sensitivity decay is calculated.
[0106] The sensitivity decay index is used to quantify the degree to which a real-time security monitoring classifier's ability to identify real attack events deteriorates over time, and it is derived from the sensitivity decay index sequence.
[0107] The sensitivity attenuation index is calculated as follows:
[0108] The long-term trend component of the sensitivity decay index sequence is accumulated. The values of all analysis periods (i.e., the sensitivity decay index for each single analysis period) of the long-term trend component are summed one by one, and then the accumulated value is divided by the total number of analysis periods to obtain the sensitivity decay index. The sensitivity decay index represents the degree of decline in the overall recognition capability of the real-time security monitoring classifier over a continuous analysis period.
[0109] A comprehensive index of induced drift risk is calculated by combining the characteristic variation amplitude index and the sensitivity decay index.
[0110] The Induced Drift Risk Comprehensive Index is used to comprehensively and quantitatively assess the overall risk level of induced drift in a real-time security monitoring classifier under the combined effects of low-intensity probing attacks and its own sensitivity degradation.
[0111] The calculation method for the comprehensive index of induced drift risk is as follows:
[0112] The characteristic variation amplitude index and the sensitivity decay index were normalized respectively.
[0113] The normalized feature variation amplitude index and sensitivity decay index are multiplied by preset risk weight parameters to obtain the comprehensive index of induced drift risk. The risk weight parameters are set based on the induced drift risk formation mechanism of the real-time safety monitoring classifier, and are usually equal to 0.5.
[0114] The higher the comprehensive index of induced drift risk, the higher the overall risk of induced drift in the real-time security monitoring classifier.
[0115] Based on the risk level of induced drift in the real-time security monitoring classifier, the risk level of the real-time data security monitoring system is determined, and a security reconstruction mechanism for the real-time security monitoring classifier is triggered, including:
[0116] Pre-set the risk level classification thresholds for the real-time data security monitoring system;
[0117] The risk level classification threshold serves as the specific boundary for the comprehensive index of induced drift risk, used to distinguish the state of the real-time data security monitoring system under different risk levels. The specific value of the risk level classification threshold is determined through analysis of the characteristics of abnormal behavior samples in historical network flow data and the performance data of historical security monitoring classifiers.
[0118] First, historical data samples of abnormal network behavior were collected, and the characteristic variation amplitude index and sensitivity decay index were calculated using these historical data samples. Then, a comprehensive index of induced drift risk was calculated based on the historical data samples. After statistically sorting all historical comprehensive index data of induced drift risk, a percentile method was used to define low-risk, medium-risk, and high-risk thresholds.
[0119] The low-risk threshold represents the boundary value at which the risk level of a real-time data security monitoring system changes from a safe state to a slightly risky state; the medium-risk threshold represents the boundary value at which the real-time data security monitoring system changes from a slightly risky state to a moderately risky state; and the high-risk threshold represents the boundary value at which the real-time data security monitoring system changes from a moderately risky state to a severely risky state.
[0120] Based on the comparison results between the comprehensive risk index of induced drift and the risk level classification threshold, the risk level of the real-time data security monitoring system is classified.
[0121] The method for classifying the risk level of a real-time data security monitoring system is as follows: The comprehensive index value of induced drift risk is compared with the risk level classification threshold to determine the risk level of the real-time data security monitoring system.
[0122] If the comprehensive index value of induced drift risk is less than or equal to the low risk threshold, the risk level of the real-time data security monitoring system is low risk.
[0123] If the comprehensive index value of induced drift risk is greater than the low risk threshold and less than or equal to the medium risk threshold, then the risk level of the real-time data security monitoring system is medium risk.
[0124] If the comprehensive index value of induced drift risk is greater than the medium risk threshold and less than or equal to the high risk threshold, then the risk level of the real-time data security monitoring system is high risk.
[0125] If the comprehensive index value of induced drift risk is greater than the high-risk threshold, the risk level of the real-time data security monitoring system is classified as severe risk.
[0126] Based on the risk level of the real-time data security monitoring system, the corresponding security reconstruction mechanism of the real-time security monitoring classifier is triggered.
[0127] The security reconstruction mechanism is a processing mechanism for reconstructing and restoring the recognition capabilities of real-time security monitoring classifiers. The security reconstruction mechanism includes operations such as updating security model parameters, retraining the abnormal behavior recognition model, and adjusting recognition thresholds. Specifically:
[0128] If the risk level of the real-time data security monitoring system is low, the security reconstruction mechanism will not be triggered.
[0129] If the real-time data security monitoring system is classified as medium risk, a security model parameter update mechanism is triggered. This mechanism involves fine-tuning or calibrating relevant model parameters in the real-time security monitoring classifier to improve its accuracy in identifying abnormal behavior in real-time network stream data.
[0130] If the real-time data security monitoring system is classified as high-risk, a retraining mechanism for the abnormal behavior recognition model is triggered. This retraining mechanism involves retraining the model using the latest collected real-time network stream data feature sequences and historical real attack sample data. The retraining method includes using supervised machine learning, training by matching sample labels with feature vectors, to restore the real-time security monitoring classifier's high-efficiency anomaly recognition performance.
[0131] If the risk level of the real-time data security monitoring system is severe, the joint operation of the identification threshold adjustment mechanism and the abnormal behavior identification model retraining mechanism will be triggered. Specifically, the abnormal behavior identification model will be retrained first, and then the identification threshold of the abnormal behavior identification model will be adjusted according to the performance results of the trained model. The identification threshold adjustment method includes optimizing the balance between the false detection rate and the false negative rate between the model output response and the actual attack sample label to achieve the maximum accuracy of detecting abnormal behavior in real-time network stream data.
[0132] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0133] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0134] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0135] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0136] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0137] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0138] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0139] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0141] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A real-time data security monitoring system based on online data modeling, characterized in that, include: Data acquisition module: Acquires real-time network stream data, extracts features from the real-time network stream data, and outputs feature sequences that characterize the behavior patterns of the network stream data; Feature space module: A feature space is constructed based on the feature sequence, generating a feature drift mapping sequence and a sequence showing the changing trend of sensitivity in abnormal behavior recognition. Migration Analysis Module: Employs anomaly distribution migration analysis method to analyze feature drift mapping sequences and identify feature variation trends caused by low-intensity probing attack samples; Attenuation Detection Module: This module uses an anomaly sensitivity attenuation detection method to analyze the trend sequence of changes in the sensitivity of anomaly behavior recognition, and identifies the decay characteristics of the real-time security monitoring classifier's ability to recognize real attack events. Risk assessment module: Based on feature variation trends and decay characteristics, assess the degree of risk of induced drift in the real-time safety monitoring classifier; Reconstruction Trigger Module: Based on the risk level of induced drift in the real-time security monitoring classifier, determine the risk level of the real-time data security monitoring system and trigger the security reconstruction mechanism of the real-time security monitoring classifier.
2. The real-time data security monitoring system based on online data modeling according to claim 1, characterized in that, Collect real-time network stream data, extract features from the real-time network stream data, and output feature sequences that characterize the behavioral patterns of the network stream data, specifically: Collect real-time network stream data from network communication links, perform multi-dimensional feature extraction operations on the real-time network stream data, and construct the original feature matrix of behavior patterns; The original feature matrix of the behavior pattern is standardized and reconstructed into a time series, forming a continuous feature sequence according to a preset time window sliding strategy.
3. The real-time data security monitoring system based on online data modeling according to claim 2, characterized in that, A feature space is constructed based on the feature sequence, generating a feature drift mapping sequence and an abnormal behavior recognition sensitivity change trend sequence, specifically: By performing embedding mapping processing based on distribution features on continuous feature sequences, the continuous feature sequences are projected into a multi-dimensional feature space to construct a feature space that characterizes the evolution process of real-time network flow data behavior patterns. In the feature space, the distribution difference of continuous feature sequences within adjacent time windows is calculated, and a feature drift mapping sequence is generated using a mapping function based on the distribution distance metric. In the feature space, the output response of the abnormal behavior recognition model is monitored. Combined with the recognition response records of historical real attack samples, the recognition sensitivity change parameters are calculated, and a sequence of abnormal behavior recognition sensitivity change trends is constructed based on time sequence.
4. The real-time data security monitoring system based on online data modeling according to claim 3, characterized in that, Anomaly distribution offset analysis is used to analyze feature drift mapping sequences and identify feature variation trends caused by low-intensity probing attack samples. Specifically: The feature drift mapping sequence is divided into multiple consecutive analysis windows, and the statistical distribution characteristic parameters of the feature drift mapping sequence within each analysis window are calculated. Based on the statistical distribution characteristic parameters of the feature drift mapping sequence in different analysis windows, a distribution offset index sequence is constructed to describe the variation law of feature distribution. Anomaly distribution offset analysis is used to decompose the distribution offset index sequence into trends. By comparing the similarity between the distribution offset index sequence and the feature distribution changes of historical anomaly attack samples, the feature variation trend caused by low-intensity probing attack samples can be identified.
5. The real-time data security monitoring system based on online data modeling according to claim 4, characterized in that, An anomaly sensitivity decay detection method was used to analyze the trend sequence of changes in the sensitivity of anomaly behavior recognition, and to identify the decay characteristics of the real-time security monitoring classifier's ability to recognize real attack events, specifically: The abnormal behavior recognition sensitivity change trend sequence is divided into multiple continuous analysis periods according to a predetermined detection window, and the statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence are calculated in each analysis period. Based on the statistical characteristics of the abnormal behavior recognition sensitivity change trend sequence and the difference between the preset abnormal behavior recognition sensitivity benchmark sequence, a sensitivity decay index sequence is constructed. An abnormal sensitivity decay detection method was used to perform time-series trend analysis on the sensitivity decay index sequence to identify the decay characteristics of the real-time security monitoring classifier's ability to identify real attack events.
6. The real-time data security monitoring system based on online data modeling according to claim 5, characterized in that, Based on feature variation trends and decay characteristics, the risk of induced drift in real-time security monitoring classifiers is assessed, specifically as follows: The feature variation magnitude index is calculated based on the feature variation trend caused by low-intensity exploratory attack samples. Based on the decay characteristics of the real-time security monitoring classifier's ability to identify real attack events, an index of sensitivity decay is calculated. A comprehensive index of induced drift risk is calculated by combining the characteristic variation amplitude index and the sensitivity decay index.
7. The real-time data security monitoring system based on online data modeling according to claim 6, characterized in that, Based on the risk level of induced drift in the real-time security monitoring classifier, the risk level of the real-time data security monitoring system is determined, and the security reconstruction mechanism of the real-time security monitoring classifier is triggered, specifically as follows: Pre-set the risk level classification thresholds for the real-time data security monitoring system; Based on the comparison results between the comprehensive risk index of induced drift and the risk level classification threshold, the risk level of the real-time data security monitoring system is classified. Based on the risk level of the real-time data security monitoring system, the corresponding real-time security monitoring classifier's security reconstruction mechanism is triggered.
Citation Information
Patent Citations
Unknown network attack behavior drift detection method based on ensemble learning
CN116886337A
Concept drift detection method and system for network attack identification
CN116938514A
Big data risk assessment system based on AI
CN120298098A
Equipment abnormity early warning system based on AI intelligent analysis
CN120496250A
Regional orderly power utilization dynamic optimization monitoring method based on self-adaptive threshold value
CN120638621A