Big data-oriented online data security compliance supervision method and system, and medium
Through sliding time window monitoring and compliance judgment, combined with the aggregation density analysis of non-compliant data nodes, the efficiency and coverage of data security compliance supervision in the big data environment are solved, and the effect of quickly positioning non-compliant data and improving detection rate is achieved.
Patent Information
- Application Number
- CN202510676906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
In the big data environment, traditional data security compliance supervision methods have become an unfeasible task due to the huge data scale and dynamic transmission paths. Random sampling or focusing on special nodes is difficult to fully cover potential risk points, which weakens the objectivity and effectiveness of compliance supervision.
Through the sliding time window, multiple online data flows are monitored in real time, abnormal data nodes are obtained, and compliance judgments are made based on data compliance standards for data loss, data abnormalities and data tampering, non-compliant data characteristics are determined, non-compliant data nodes are traversed, aggregation density is calculated, and the attacked time period is judged, and multiple data flows are subject to compliance investigation to determine the reasons for the occurrence of non-compliant data.
It has achieved rapid positioning of critical time periods for attacks to generate non-compliant data, improved the detection rate of non-compliant data with inconspicuous data abnormalities, and improved the objectivity and effectiveness of data security and compliance supervision.
Smart Images

Figure CN120200855A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security compliance supervision, and particularly to an online data security compliance supervision method, system and medium for big data. Background Art
[0002] Some business-related data of enterprises involves information privacy, trade secrets, audit trails or intellectual property rights, so data security compliance supervision is required. For example, data such as production information and sales turnover of manufacturing enterprises, processing transaction records and customer identity information of financial enterprises, stored patient electronic medical records and health monitoring data of medical institutions, and user behavior logs and payment information of e-commerce enterprises all need to perform real-time compliance judgments and supervision on data loss, data anomalies and data tampering.
[0003] For the supervision of data security compliance, the traditional mode relies on comparing the current data item by item with the preset compliance standards to identify deviations or violations. However, in the big data environment, the data scale grows exponentially and the transmission path is highly dynamic, making it an infeasible task to compare data segment by segment. To address this challenge, existing technologies usually adopt a sampling strategy, that is, randomly select some data nodes, or focus on special nodes where significant changes occur in the data transmission state (such as peak traffic periods, cross-regional transmission nodes) for compliance analysis.
[0004] However, this random sampling may result in the omission of key data nodes; the selection of special state nodes is easily affected by subjective judgment and it is difficult to comprehensively cover potential risk points. In summary, in the context of a large amount of data, the limitation of only focusing on the compliance of special data nodes with significant changes not only weakens the objectivity and effectiveness of compliance supervision, but also may pose a hidden danger to data security. Summary of the Invention
[0005] To solve the above problems, the present invention provides an online data security compliance supervision method for big data, which can quickly locate the critical time period when non-compliant data is generated due to an attack, and then perform data investigation on multiple data streams to delete non-compliant data or correct the data.
[0006] To achieve the above object, the present invention provides the following technical solutions.
[0007] An online data security compliance supervision system for big data includes the following steps: Real-time monitor multiple online data streams respectively through a sliding time window, obtain the abnormal data nodes where data anomalies occur in each real-time online data stream, and extract the data segments corresponding to the real-time online data streams where the abnormal data nodes are located; Call the data compliance standard configuration file, and based on the data compliance standard, judge the compliance of data loss, data anomalies, and data tampering in the data segment, determine the non-compliant data features that generate loss features, anomaly features, and tampering features in the data segment, and regard the data nodes with non-compliant data features in the real-time online data stream as non-compliant data nodes; Traverse the non-compliant data nodes, and determine the aggregation density of a preset number of non-compliant data nodes on the time axis in each fixed time period; when the aggregation density exceeds the preset threshold, the continuous time period between the time points where the head and tail non-compliant data nodes are located is the attacked time period; Retrieve all the data contents corresponding to the attacked time periods of multiple real-time online data streams for compliance investigation to determine the reasons for the generation of non-compliant data.
[0008] Preferably, the steps of respectively monitoring multiple online data streams in real time through a sliding time window and obtaining the abnormal data nodes that generate data anomalies in each real-time online data stream include the following steps: Receive the online data stream in real time, and set a sliding window with a fixed time size to slide and detect the data in the window; Analyze the data in the sliding window in real time, calculate the statistical information of the data, and judge whether there are data points exceeding the threshold in the current sliding window according to the preset anomaly judgment threshold; Perform outlier detection based on the threshold on the data in the sliding window to obtain abnormal data points; Regard the data points exceeding the threshold and the abnormal data points as abnormal data nodes.
[0009] Preferably, the steps of extracting the data segment corresponding to the real-time online data stream where the abnormal data nodes are located include the following steps: Mark the positions of the abnormal data nodes through the timestamps in the data stream; According to the size of the sliding window and the timestamps of the abnormal data nodes, extract the continuous data of the previous and next n time points containing the abnormal data points from the data stream to form a data segment to be checked for compliance.
[0010] Preferably, the steps of judging the compliance of data loss, data anomalies, and data tampering in the data segment include the following steps: Load the compliance standard related to the data stream from the configuration management system; Based on the loss standard defined in the configuration file, judge whether there are missing data nodes in the data segment to obtain the data loss compliance judgment result; Based on the anomaly threshold in the configuration file, judge whether there are data nodes exceeding the normal range in the data segment to obtain the data anomaly compliance judgment result; Perform matching verification based on the hash value set in the configuration file to determine whether there are tampered data nodes in the data segment, and obtain the judgment result of data tampering compliance; Label the non-compliant data nodes that cause data loss, data anomalies, and data tampering, and store the non-compliant data nodes and their corresponding information in all data streams; the corresponding information of the non-compliant data nodes includes the corresponding timestamp, data value, and the data stream ID to which they belong.
[0011] Preferably, traversing the non-compliant data nodes to determine the aggregation density of a preset number of non-compliant data nodes on the time axis in each fixed time period includes the following steps: Traverse all real-time data streams, mark each non-compliant data node, and store the non-compliant data node and its corresponding information; Construct a list of all non-compliant data nodes based on the non-compliant data nodes and their corresponding information, corresponding to the timestamps of the non-compliant data nodes; For a preset fixed time period, based on the timestamps in the data stream, starting from each non-compliant data node as the starting time point, divide the data into multiple time periods of fixed length; For each time period, count the number of non-compliant data nodes in that time period; According to the number of non-compliant data nodes in each time period and the length of the time period, determine the aggregation density of the non-compliant data nodes: Aggregation density = number of non-compliant data nodes / length of the time period.
[0012] Preferably, when the aggregation density exceeds the preset threshold, the continuous time period between the time points of the head and tail non-compliant data nodes is the attacked time period, including the following steps: For each time period, determine whether its aggregation density exceeds the preset threshold. If it exceeds the threshold, mark that time period as a time period with an attack; Extract the timestamps of the head and tail non-compliant data nodes of the time period with an attack; Based on the timestamps of the head and tail non-compliant data nodes of each time period with an attack, determine the time interval of the time period with an attack; if the interval between the time periods with an attack is less than or equal to the preset maximum duration, regard it as a continuous attacked time period; if the interval between the time periods with an attack exceeds the preset time period, regard it as multiple independent attack events; According to the timestamps of the head and tail non-compliant data nodes of the attack event, obtain the total duration of the attacked time period, and mark and store the attack information; the attack information includes the timestamp, duration, and the data streams involved.
[0013] Preferably, retrieving all data contents corresponding to the attacked time periods of multiple real-time online data streams for compliance investigation and determining the reasons for non-compliant data includes the following steps: Retrieve all real-time data stream data related to the attacked time period, conduct cross-checks, obtain multiple data stream abnormal states, and analyze whether the multiple data stream abnormalities are caused by external factors; Combined with network monitoring, device logs, and external data, trace back to the attacked time period and analyze whether there are problems such as network attacks, hardware failures, or transmission errors; Correct or delete multiple data stream abnormal states during the attacked time period.
[0014] The present invention also provides an online data security compliance supervision system for big data, and the system includes: A processor; A memory, on which a computer program that can run on the processor is stored; Wherein, when the computer program is executed by the processor, the steps of the online data security compliance supervision method for big data are implemented.
[0015] The present invention also provides a computer-readable storage medium, on which a data processing program is stored, and when the data processing program is executed by a processor, the steps of the online data security compliance supervision method for big data are implemented.
[0016] Advantages of the present invention: The present invention proposes an online data security compliance supervision method for big data. This method quickly identifies data points that generate abnormalities through a sliding time window, and determines whether non-compliant data is generated around the data points, including compliance judgments on data loss, data anomalies, and data tampering. Further determine the characteristics of non-compliant data, and use this as a quick way to locate non-compliant data nodes and extract the time periods when non-compliant data is generated. Based on the characteristics of non-compliant data aggregation, this method determines whether it has been attacked by the outside world according to the aggregation density of non-compliant data nodes, locates the time of the attack, and analyzes the data compliance of other data streams during this period, effectively improving the detection rate of non-compliant data with unobvious data anomalies. This method can quickly judge data segments with a relatively high probability of non-compliance, take the data segments where non-compliant characteristic data nodes are concentrated as key data segments, locate the key time periods where non-compliant data streams are likely to be generated, and investigate the reasons for the generation of non-compliant data. Description of the Drawings
[0017] Figure 1 is the method flow chart of an embodiment of the present invention; Figure 2 is the method flow chart of the method for judging concentrated attacks by aggregation density in an embodiment of the present invention. Detailed implementation manners
[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] Embodiment 1 For the supervision of data security compliance, the traditional mode relies on comparing the current data item by item with the preset compliance standards to identify deviations or violations. However, in the big data environment, the data scale grows exponentially and the transmission path is highly dynamic, making it an infeasible task to compare data segment by segment. To address this challenge, regulatory agencies or enterprises usually adopt a sampling strategy, that is, randomly select some data nodes, or focus on special nodes where significant changes occur in the data transmission status (such as peak traffic periods, cross-regional transmission nodes) for compliance analysis. Although this method can improve efficiency, its inherent defects cannot be ignored: random sampling may lead to the omission of key data nodes; the selection of special status nodes is easily affected by subjective judgment and it is difficult to comprehensively cover potential risk points.
[0020] When an enterprise receives business data streams, when it is attacked or tampered with from the outside, non-compliant data transmissions will be concentratedly generated in a short period of time, or due to the end of the attack or the correction of defense measures, the data will return to compliant data within a certain specific range during subsequent transmissions. However, for the compliance supervision of big data, it is not possible to perform data compliance comparison item by item based on time nodes. However, when the data stream is attacked from the outside, a large number of non-compliant data will appear, that is, the characteristic of non-compliant data points gathering appears. Based on the characteristic of non-compliant data points gathering, it is necessary to quickly judge the data segments with a relatively high probability of non-compliance generation, take the data segments where non-compliance characteristic data nodes are concentrated as key data segments, locate the key time periods where non-compliant data streams are likely to be generated, and investigate the reasons for the generation of non-compliant data. For this purpose, this embodiment proposes an online data security compliance supervision method for big data, and the specific steps are as Figure 1 shown, including: S1: Real-time monitor multiple online data streams respectively through a sliding time window, and obtain abnormal data nodes that generate data anomalies in each real-time online data stream.
[0021] S2: Extract the data segments corresponding to the real-time online data streams where the abnormal data nodes are located.
[0022] S3: Call the data compliance standard configuration file, and based on the data compliance standards, perform compliance judgments on the data segments for data loss, data anomalies and data tampering, and determine the non-compliance data characteristics that generate loss characteristics, anomaly characteristics and tampering characteristics in the data segments.
[0023] S4: Take the data nodes with non-compliant data features in the real-time online data stream as non-compliant data nodes.
[0024] S5: Traverse the non-compliant data nodes and determine the aggregation density of a preset number of non-compliant data nodes on the time axis in each fixed time period.
[0025] S6: When the aggregation density exceeds the preset threshold, the continuous time period between the time points of the head and tail non-compliant data nodes is the attacked time period.
[0026] S7: Retrieve all the data contents corresponding to the attacked time periods of multiple real-time online data streams for compliance investigation to determine the cause of non-compliant data.
[0027] In S1, the present invention first performs real-time monitoring of outliers, and obtains outlier data points through threshold judgment and data point outlier analysis, specifically including: S1.1: Receive the online data stream in real time and set a sliding window with a fixed time size to slide and detect the data within the window.
[0028] S1.2: Perform real-time analysis on the data within the sliding window, calculate the statistical information of the data, and judge whether there are data points exceeding the threshold in the current sliding window according to the preset outlier judgment threshold.
[0029] S1.3: Perform outlier detection based on the threshold on the data within the sliding window to obtain outlier data points.
[0030] S1.4: Take the data points exceeding the threshold and the outlier data points as outlier data nodes, and mark the abnormal and threshold-exceeding data points in the data window in real time for subsequent compliance judgment.
[0031] Further, in S2, taking the outlier data nodes as the benchmark, call the continuous data before and after to be used as the minimum unit for compliance check. Specifically, mark the position of the outlier data nodes through the timestamps in the data stream; according to the size of the sliding window and the timestamps of the outlier data nodes, extract the continuous data of the previous and next n time points containing the outlier data points from the data stream to form the data segment to be checked for compliance.
[0032] The present invention gives the abnormal forms of non-compliance generated after the data is attacked, including loss, abnormality and tampering, specifically including the following steps: S3.1: Load the compliance standards related to the data stream from the configuration management system.
[0033] S3.2: Based on the loss criteria defined in the configuration file, judge whether there are missing data nodes in the data segment to obtain the compliance judgment result of data loss.
[0034] S3.3: Based on the anomaly threshold in the configuration file, determine whether there are data nodes in the data segment that exceed the normal range, and obtain the data anomaly compliance judgment result.
[0035] S3.4: Based on the hash value set in the configuration file, perform matching verification to determine whether there are tampered data nodes in the data segment, and obtain the data tampering compliance judgment result.
[0036] S3.5: Label the non-compliant data nodes that cause data loss, data anomalies, and data tampering, and store the non-compliant data nodes and their corresponding information in all data streams; the corresponding information of the non-compliant data nodes includes the corresponding timestamp, data value, and the data stream ID to which they belong.
[0037] Specifically, the non-compliant anomaly forms specifically include: (1) Data Loss Data loss generally refers to the situation where data that should originally exist fails to be successfully recorded, stored, or transmitted due to certain reasons. It may lead to the absence, omission, or incompleteness of data, especially when problems occur during data transmission, storage, backup, etc.
[0038] The manifestation characteristics of data loss are missing fields or records, missing time series data, and discontinuous records. For example, the absence of some fields or records in a dataset may result in null values (NULL, NaN, etc.) during data analysis, or the absence of some items that should exist in the dataset. Or, there are no data records within a certain time period, resulting in the inability to analyze some time points. Or, if the data is stored in chronological order, the missing data usually appears as an interruption or jump in the time series. For example, in an Internet of Things device, if sensor data is lost, it may be found that the temperature record at a certain time point is null or NaN. In a database, if a data record fails to be committed, it may lead to the absence of that record.
[0039] Common methods for detecting data loss in the prior art are null value checking and time series checking. Null value detection: Detect null values (NaN, NULL) in the data, such as using the isna() or isnull() functions in pandas. Time series checking: Compare whether data is missing at adjacent time points, and use timestamp sorting to check the continuity of time series data.
[0040] (2) Data Anomalies Data anomalies generally refer to the situation where, under certain conditions, the data presents values or trends different from most data points. Such anomalies may be caused by system failures, sensor errors, external environmental interference, etc.
[0041] The performance characteristics of data anomalies are abnormal numerical fluctuations, extreme values, sudden trend changes, and excessive fluctuation amplitudes. For example, the data shows obvious fluctuations that deviate from the normal range, such as sensor readings jumping, and unreasonable values appear within a short period. Or, some data points are significantly beyond the normal range, and these values far exceed the range of historical data. Or, the trend of the data suddenly changes, for example, a long-term stable temperature data suddenly jumps up or drops sharply at a certain time point. Or, within certain time periods, the fluctuation amplitude of the data increases abnormally, which may be a sign of equipment failure or external interference.
[0042] Existing technologies usually use the standard deviation method, IQR method, and time series anomaly detection for judgment. Standard deviation method: By calculating the mean and standard deviation of the data, identify the abnormal points that deviate from the mean by more than multiple standard deviations. IQR method: Use the quartile analysis method to identify the data points outside the upper quartile and lower quartile. Time series anomaly detection: Detect trend changes through time series analysis methods (such as ARIMA models, LSTM neural networks, etc.).
[0043] (3)Data Tampering Data tampering refers to the unauthorized modification of data during storage, transmission, or processing. The tampered data may lead to data distortion, affecting the accuracy and credibility of the data.
[0044] The performance characteristics of data tampering are data inconsistency, historical record changes, unexpected checksum or digital changes, and suspicious timestamp modifications. For example, the same data is inconsistent at different time points or in different systems, which may be caused by data tampering or incorrect modification. For example, the transaction amounts in different records of transaction data are inconsistent. The historical data of the record is modified or deleted, which may lead to the inconsistency between the data and the actual events. Another example is that the hash value or checksum of the data changes, which may indicate that the data has been tampered with during transmission. Another example is that the timestamp is modified, and the time of the data record does not match the actual time. Existing technologies usually use hash verification.
[0045] Based on the characteristics of non-compliant data point aggregation, the present invention provides a technical means for judging centralized attacks by aggregation density. The flow chart is as Figure 2 shown, specifically including: S5.1: Traverse all real-time data streams, mark each non-compliant data node, and store the non-compliant data node and its corresponding information.
[0046] S5.2: Construct a list of all non-compliant data nodes according to the non-compliant data nodes and their corresponding information, corresponding to the timestamps of the non-compliant data nodes.
[0047] S5.3: A preset fixed time period. Based on the timestamps in the data stream, starting from each non-compliant data node as the starting time point, the data is divided into multiple time periods of fixed length. For example, if the timestamps in the data stream range from t1 to tn, the time periods can be defined as [t1, t1 + 5], [t1 + 5, t1 + 10], with the unit being minutes, and so on.
[0048] S5.4: For each time period, count the number of non-compliant data nodes within that time period. For example, for the time period [t1, t1 + 5], count the number of non-compliant data nodes that appear within that time period.
[0049] S5.5: Based on the number of non-compliant data nodes within each time period and the length of the time period, determine the aggregation density of non-compliant data nodes: Aggregation density = Number of non-compliant data nodes / Length of the time period. For example, if a time period is 5 minutes and there are 10 non-compliant data nodes within that time period, the aggregation density is 2.
[0050] When the aggregation density exceeds the preset threshold, the continuous time period between the time points of the first and last non-compliant data nodes is the attacked time period. Specifically, it includes the following steps: S6.1: For each time period, determine whether its aggregation density exceeds the preset threshold. If it exceeds the threshold, mark that time period as a time period with an attack.
[0051] S6.2: Extract the timestamps of the first and last non-compliant data nodes of the time period with an attack. For example, for the time period [t_start, t_end], it represents the preliminary time range of the attacked time period.
[0052] S6.3: Based on the timestamps of the first and last non-compliant data nodes of each time period with an attack, determine the time interval of the time period with an attack. If the interval between the time periods with an attack is less than or equal to the preset maximum duration, consider it as a continuous attacked time period; if the interval between the time periods with an attack exceeds the preset time period, consider it as multiple independent attack events.
[0053] S6.4: Based on the timestamps of the first and last non-compliant data nodes of the attack event, obtain the total duration of the attacked time period, and mark and store the attack information; the attack information includes timestamps, duration, and the data streams involved.
[0054] Furthermore, retrieve all the data contents corresponding to the attacked time periods of multiple real-time online data streams for compliance investigation to determine the reasons for the generation of non-compliant data: Retrieve all real-time data stream data related to the attacked time period, conduct cross-checks, obtain multiple data stream abnormal states, and analyze whether multiple data stream abnormalities are caused by external factors. Combine network monitoring, device logs, and external data to trace back the attacked time period and analyze whether there are problems such as network attacks, hardware failures, or transmission errors. Correct or delete multiple data stream abnormal states during the attacked time period.
[0055] In this embodiment, the method can quickly identify data points that generate abnormalities by adopting the technology of a sliding time window, and determine whether non-compliant data is generated by analyzing the surrounding areas of the data points. Specifically, the method can not only judge whether there are problems such as data loss, abnormality, or tampering, but also conduct detailed compliance judgments on each type of non-compliant data, thereby further determining its non-compliance characteristics. This process helps to quickly locate non-compliant data nodes, thereby extracting the time periods during which non-compliant data is generated. These time periods are considered potential abnormal intervals, which can provide accurate positioning for subsequent analysis.
[0056] On this basis, the method pays special attention to the aggregation characteristics of non-compliant data, and judges whether there are signs of external attacks by analyzing the aggregation density of non-compliant data nodes. Specifically, by calculating the aggregation density of non-compliant data within each time period, when the density exceeds the set threshold, the system can identify potential attack behaviors and further determine the time points when the attacks occur. This provides effective support for further analyzing the compliance of other data streams during this period, especially when data abnormalities are not easily detectable, greatly improving the detection rate of non-compliant data. With the help of this technology, the system can not only accurately identify non-compliant data nodes, but also deeply investigate the reasons for the generation of non-compliant data.
[0057] Based on the aggregation characteristics of non-compliant data, the method makes the identification of abnormal data more efficient. Especially when facing data streams with unobvious abnormalities, it can extract problem data segments through density analysis. Specifically, the method judges which time periods of the data stream may have a higher non-compliance risk according to the aggregation density of non-compliant data nodes. These time periods are considered key data segments, and the system will pay special attention to the data compliance of these time periods and deeply explore the root causes that may lead to data non-compliance. This monitoring method based on time period division effectively improves the detection ability of non-compliant data, can timely identify and respond to external attacks, thereby ensuring the security and compliance of the data stream.
[0058] In this way, the system can quickly identify potential problems in the data stream and accurately locate the time point when data anomalies are most likely to occur. This method not only provides an efficient solution for real-time data monitoring but also offers strong protection for data security, especially applicable to scenarios where abnormal behaviors in the data stream are difficult to detect. By continuously monitoring and analyzing the compliance of the data stream, the system can actively identify and alert potential attack events, prevent data security risks in advance, and provide strong support for subsequent processing and response measures.
[0059] The above is an online data security compliance supervision method for big data provided by an embodiment of this embodiment. Based on the same idea, this embodiment also provides a corresponding online data security compliance supervision system for big data. For the specific limitations of the online data security compliance supervision system for big data, reference can be made to the limitations on the online data security compliance supervision method for big data in the above text, which will not be elaborated here. Each module in the above online data security compliance supervision system for big data can be implemented in whole or in part through software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0060] This embodiment also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 provided online data security compliance supervision method for big data.
[0061] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0062] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An online data security compliance supervision method for big data, characterized in that, It includes the following steps: Monitor multiple online data streams in real time through a sliding time window respectively, obtain the abnormal data nodes with data anomalies in each real-time online data stream, and extract the data segments corresponding to the real-time online data streams where the abnormal data nodes are located; Call the data compliance standard configuration file, and based on the data compliance standards, judge the compliance of the data segments for data loss, data anomalies, and data tampering, determine the non-compliant data features that generate loss features, anomaly features, and tampering features in the data segments, and regard the data nodes with non-compliant data features in the real-time online data stream as non-compliant data nodes; Traverse the non-compliant data nodes and determine the aggregation density of a preset number of non-compliant data nodes on the time axis in each fixed time period; When the aggregation density exceeds the preset threshold, the continuous time period between the time points of the head and tail non-compliant data nodes is the attacked time period; Retrieve all the data contents corresponding to the attacked time periods of multiple real-time online data streams for compliance investigation to determine the reasons for the generation of non-compliant data.
2. The online data security compliance supervision method for big data according to claim 1, characterized in that The step of respectively monitoring multiple online data streams in real time through a sliding time window and obtaining the abnormal data nodes with data anomalies in each real-time online data stream includes the following steps: Receive the online data stream in real time and set a sliding window with a fixed time size to slide and detect the data within the window; Perform real-time analysis on the data within the sliding window, calculate the statistical information of the data, and judge whether there are data points exceeding the threshold in the current sliding window based on the preset anomaly judgment threshold; Perform outlier detection based on the threshold on the data within the sliding window to obtain abnormal data points; Regard the data points exceeding the threshold and the abnormal data points as abnormal data nodes.
3. The online data security compliance supervision method for big data according to claim 2, characterized in that, The step of extracting the data segments corresponding to the real-time online data streams where the abnormal data nodes are located includes the following steps: Mark the positions of the abnormal data nodes through the timestamps in the data stream; According to the size of the sliding window and the timestamps of the abnormal data nodes, extract the continuous data of the previous and next n time points containing the abnormal data points from the data stream to form the data segment to be checked for compliance.
4. The online data security compliance supervision method for big data according to claim 3, wherein The step of judging the compliance of the data segments for data loss, data anomalies, and data tampering includes the following steps: Load the compliance standards related to the data stream from the configuration management system; Based on the loss standards defined in the configuration file, judge whether there are missing data nodes in the data segment to obtain the data loss compliance judgment result; Based on the anomaly threshold in the configuration file, judge whether there are data nodes exceeding the normal range in the data segment to obtain the data anomaly compliance judgment result; Based on the hash values set in the configuration file for matching verification, judge whether there are tampered data nodes in the data segment to obtain the data tampering compliance judgment result; Label the types of non-compliant data nodes that generate data loss, data anomalies, and data tampering, and store the non-compliant data nodes and their corresponding information in all data streams; the corresponding information of the non-compliant data nodes includes the corresponding timestamps, data values, and the data stream IDs to which they belong.
5. The online data security compliance supervision method for big data according to claim 4, characterized in that Traversing the non-compliant data nodes and determining the aggregation density of a preset number of non-compliant data nodes on the time axis in each fixed time period includes the following steps: Traverse all real-time data streams, mark each non-compliant data node, and store the non-compliant data node and its corresponding information; Construct a list of all non-compliant data nodes based on the non-compliant data nodes and their corresponding information, corresponding to the timestamps of the non-compliant data nodes; For a preset fixed time period, based on the timestamps in the data stream, with each non-compliant data node as the starting time point, divide the data into multiple time periods of a fixed length; For each time period, count the number of non-compliant data nodes in that time period; Determine the aggregation density of the non-compliant data nodes according to the number of non-compliant data nodes in each time period and the length of the time period: Aggregation density = number of non-compliant data nodes / length of the time period.
6. The online data security compliance supervision method for big data according to claim 1, wherein When the aggregation density exceeds the preset threshold, the continuous time period between the time points of the non-compliant data nodes at the head and tail is the attacked time period, including the following steps: For each time period, judge whether its aggregation density exceeds the preset threshold. If it exceeds the threshold, mark this time period as a time period with an attack; Extract the timestamps of the non-compliant data nodes at the head and tail of the time period with an attack; Based on the timestamps of the non-compliant data nodes at the head and tail of each time period with an attack, judge the time interval of the time period with an attack; if the interval between the time periods with an attack is less than or equal to the preset maximum duration, regard it as a continuous attacked time period; if the interval between the time periods with an attack exceeds the preset time period, regard it as multiple independent attack events; According to the timestamps of the non-compliant data nodes at the head and tail of the attack event, obtain the total duration of the attacked time period, and mark and store the attack information; the attack information includes timestamps, duration, and the data streams involved.
7. The online data security compliance supervision method for big data according to claim 1, characterized in that Retrieving all data contents corresponding to the attacked time periods of multiple real-time online data streams for compliance investigation to determine the reasons for the generation of non-compliant data includes the following steps: Retrieve the data of all real-time data streams related to the attacked time period, conduct cross-checks, obtain the abnormal states of multiple data streams, and analyze whether the abnormal states of multiple data streams are caused by external factors; Combined with network monitoring, device logs, and external data, trace back the attacked time period, and analyze whether there are problems such as network attacks, hardware failures, or transmission errors; Correct or delete the abnormal states of multiple data streams in the attacked time period.
8. An online data security compliance supervision system for big data, characterized in that, The system includes: A processor; A memory, on which a computer program that can run on the processor is stored; Wherein, when the computer program is executed by the processor, it implements the steps of the online data security compliance supervision method for big data as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, A data processing program is stored on the computer-readable storage medium, and when the data processing program is executed by the processor, it implements the steps of the online data security compliance supervision method for big data as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Network abnormal node detection method based on node multi-dimensional features
CN112804255A
Visual management method and system for distributed data center
CN117037022A
Abnormal data monitoring method and device, equipment, medium and program product
CN117573477A
Network security monitoring system based on big data analysis
CN119728279A
Measurement abnormal data detection method and system based on rule engine
CN119807985A
Cited By
Industrial end node data tamper-proofing method and system based on Internet of Things
CN121125353A