Data leakage protection method, device and equipment and storage medium

By combining a parallel detection system and a distributed message processing cluster, data streams are acquired and analyzed in real time, solving the problems of insufficient timeliness and automated response in traditional data leakage protection solutions, and achieving efficient and accurate data leakage risk detection and blocking.

CN118965339BActive Publication Date: 2026-03-27HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional data breach prevention solutions lack timeliness and automation, resulting in inaccurate and inefficient data breach risk detection.

Method used

The system employs a first detection system (such as a network application firewall) and a second detection system (such as a data risk identification system) in parallel to perform real-time data leakage risk detection. By cooperating with a distributed message processing cluster through subscription relationships, it acquires and analyzes data streams in real time, uses a consistent hashing algorithm and preset risk identification rules to detect data leakage risks, and performs data blocking operations when a risk is detected.

Benefits of technology

It enables real-time detection and timely response to data breach risks, improves detection efficiency and accuracy, ensures that data is not leaked or misused, and ensures the real-time nature and integrity of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118965339B_ABST
    Figure CN118965339B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of data security, and provides a data leakage protection method, device, equipment and storage medium, the method comprises the following steps: acquiring a to-be-detected data stream in real time. The to-be-detected data stream is sent to a first detection system and a distributed message processing cluster in parallel; wherein the distributed message processing cluster and a second detection system are configured with a subscription relationship, and the second detection system is used for extracting the to-be-detected data stream received by the distributed message processing cluster based on the subscription relationship. Real-time data leakage risk detection is performed on the to-be-detected data stream by using the first detection system and / or the second detection system. When any one of the first detection system or the second detection system detects that the to-be-detected data stream has a data leakage risk, a data blocking operation is performed. The scheme can respond to the data leakage risk in time while processing a large amount of to-be-detected data, and ensures that the data is not leaked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and in particular to a data leakage prevention method, apparatus, device and storage medium. Background Technology

[0002] In the era of digital transformation, data breaches have become a serious security threat, potentially leading to severe consequences such as the leakage of user privacy and trade secrets.

[0003] Traditional security solutions typically use technologies such as Intrusion Detection Systems (IDS), Intrusion Prevention Systems (IPS), and Application Programming Interface (API) risk monitoring platforms to address data breaches.

[0004] However, traditional solutions for data breach protection mainly suffer from the inability to guarantee the timeliness of data analysis and the fact that once a risk is discovered, intervention can only be carried out by human intervention, which cannot effectively protect data. Summary of the Invention

[0005] This application provides a data leakage prevention method, apparatus, device, and storage medium, which can solve the technical problem of how to efficiently and timely prevent data leakage.

[0006] In a first aspect, embodiments of this application provide a data leakage prevention method, including:

[0007] Real-time acquisition of the data stream to be detected.

[0008] The data stream to be detected is sent in parallel to the first detection system and the distributed message processing cluster; wherein, the distributed message processing cluster and the second detection system are configured with a subscription relationship, and the second detection system is used to extract the data stream to be detected received by the distributed message processing cluster based on the subscription relationship.

[0009] The first detection system and / or the second detection system are used to perform real-time data leakage risk detection in parallel on the data stream to be detected.

[0010] When either the first or second detection system detects a risk of data leakage in the data stream to be detected, a data blocking operation is performed.

[0011] In one embodiment, the first detection system is a network application firewall, and the second detection system is a data risk identification system. Real-time data leakage risk detection of the data stream to be detected is performed in parallel using the first and / or second detection systems. This includes: using the network application firewall to perform real-time data leakage risk detection on the data stream to be detected, obtaining a first detection result; and using the data risk identification system to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules, obtaining a second detection result. The preset risk identification rules include at least one of the following: sensitive data identification rules using regular expressions, access request frequency analysis rules, and access volume and access time analysis rules.

[0012] In one embodiment, the first detection system includes a first risk identification device, and the second detection system includes a second risk identification device. The data risk identification system performs real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules to obtain a second detection result. This includes: determining the virtual node number of the second risk identification device using a consistent hashing algorithm; performing hash calculations on the number name of the first risk identification device and the address name of the user terminal to obtain an association table, which records the association relationships between the first risk identification device, the second risk identification device, and the user terminal, where the user terminal represents the source of the data stream to be detected; and distributing the data stream to be detected to the corresponding virtual node of the second risk identification device according to the association table for real-time data leakage risk detection to obtain the second detection result.

[0013] In one embodiment, a data risk identification system is used to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules to obtain a second detection result. This includes: retrieving historical detection data streams from the data risk identification system; and performing parallel analysis of the historical detection data streams and the data stream to be detected according to preset risk identification rules to obtain a second detection result.

[0014] In one embodiment, real-time acquisition of the data stream to be detected includes: configuring preset data acquisition rules, which include at least one of the following: target name, target port, and target protocol; and acquiring the data stream to be detected in real time using a network application firewall according to the preset data acquisition rules.

[0015] In one embodiment, when a data leakage risk is detected in the data stream to be detected, a data blocking operation is performed, including: when a first detection result indicates a data leakage risk, terminating the network connection of the user address with the data leakage risk using a network application firewall.

[0016] In one embodiment, when the second detection result indicates a risk of data leakage, based on the subscription relationship configured between the second detection system and the distributed message processing cluster, the second detection system sends an early warning notification to the distributed message processing cluster. The distributed message processing cluster is used to instruct the network application firewall to terminate the network connection of the user address that is at risk of data leakage according to the early warning notification.

[0017] Secondly, embodiments of this application provide a data leakage prevention device that has the function of implementing the method in the first aspect or any possible implementation thereof. Specifically, the device includes units for implementing the method in the first aspect or any possible implementation thereof.

[0018] In one embodiment, the device includes:

[0019] The acquisition unit is used to acquire the data stream to be detected in real time.

[0020] The processing unit is used to send the data stream to be detected in parallel to the first detection system and the distributed message processing cluster; wherein, the distributed message processing cluster and the second detection system are configured with a subscription relationship, and the second detection system is used to extract the data stream to be detected received by the distributed message processing cluster based on the subscription relationship.

[0021] The processing unit is also used to perform real-time data leakage risk detection in parallel using the first detection system and / or the second detection system on the data stream to be detected.

[0022] The processing unit is also used to perform a data blocking operation when either the first detection system or the second detection system detects a risk of data leakage in the data stream to be detected.

[0023] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it causes the computer device to implement any of the implementation methods of the first aspect described above.

[0024] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a computer device, causes the computer device to implement the method of any of the implementations of the first aspect described above.

[0025] Fifthly, embodiments of this application provide a computer program product that, when run on a computer device, causes the computer device to execute any of the implementation methods of the first aspect described above.

[0026] The beneficial effects of this application embodiment compared with the prior art are as follows: by acquiring the data stream to be detected in real time, the latest data stream can be obtained in a timely manner, ensuring the real-time nature and accuracy of data leakage risk detection; by sending the data stream to be detected to the first detection system and the distributed message processing cluster in parallel, the processing efficiency of the system can be improved; the second detection system extracts the data stream of the distributed message processing cluster based on the subscription relationship, which can ensure the integrity and consistency of the data; the first detection system and / or the second detection system can improve the detection efficiency and timeliness by performing data leakage risk detection in parallel, and can process a relatively large amount of data to be detected at the same time; when a data leakage risk is detected, a data blocking operation is performed, which can respond to the data leakage risk in a timely manner and ensure that the data is not leaked or misused. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of a data leakage scenario provided in an embodiment of this application.

[0028] Figure 2 This is a flowchart illustrating a data leakage prevention method provided in an embodiment of this application.

[0029] Figure 3 This is a schematic diagram of a hash ring provided in an embodiment of this application.

[0030] Figure 4 This is an interactive flowchart of a data leakage prevention method provided in an embodiment of this application.

[0031] Figure 5 This is a schematic diagram of the structure of a data leakage protection device provided in an embodiment of this application.

[0032] Figure 6 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0033] A data breach refers to the unauthorized or unintentional disclosure, access to, or release of sensitive data to an unauthorized third party. With the digital transformation of enterprises and the increase in various business access channels, such as websites, mobile applications, and APIs, more opportunities for data breaches are created.

[0034] The following is combined Figure 1 Let me explain a specific data breach scenario.

[0035] Figure 1 This is a schematic diagram of a data leakage scenario provided in an embodiment of this application.

[0036] like Figure 1As shown, when a user accesses the network, if there are security vulnerabilities in the application the user is using, attackers can exploit these vulnerabilities to bypass security measures and obtain the user's private data.

[0037] To address the aforementioned issues, this application proposes a data breach prevention method that can efficiently and in real-time detect data breach risks and promptly block data leakage.

[0038] To further illustrate the technical solution of this application, specific embodiments are described below.

[0039] Figure 2 This is a flowchart illustrating a data leakage prevention method provided in an embodiment of this application.

[0040] like Figure 2 As shown, the above method includes the following steps S201 to S204.

[0041] S201. Real-time acquisition of the data stream to be detected.

[0042] A data stream refers to a continuous sequence of data packets or information transmitted between networks or systems in the form of a stream. In other words, a data stream to be monitored refers to a data stream that needs to be monitored and detected to ensure that it does not contain sensitive information or pose a risk of data leakage.

[0043] S202, Send the data stream to be detected to the first detection system and the distributed message processing cluster in parallel.

[0044] The distributed message processing cluster and the second detection system are configured with a subscription relationship. The second detection system is used to extract the data stream to be detected received by the distributed message processing cluster based on the subscription relationship.

[0045] The first detection system and the second detection system refer to two independent data detection systems or software used to perform security checks on the data stream to be detected.

[0046] As an example and not a limitation, the first detection system can refer to a Web Application Firewall (WAF); the second detection system can be a processing system that performs risk analysis on network data streams in a time sequence. This system can identify unauthorized access, sensitive data identification, excessive return of sensitive data at one time, sensitive data pagination requests, etc. It can identify the risk in a specific time sequence and, based on the situation, determine whether there is a risk of data leakage in the entire protection system.

[0047] Web Application Firewalls (WAFs) identify and protect against malicious activity in website, app, and API traffic. After cleaning and filtering the traffic, they return normal and secure traffic to the server, preventing malicious intrusions that could lead to performance issues and ensuring website business and data security. WAFs primarily focus on protecting against network request attacks and lack the ability to identify data types, data flows, and data leakage risks. For example, they cannot detect unauthorized data access, excessive data acquisition at once, abnormal flow of sensitive data, logins from different locations to acquire large amounts of data, or abnormal data outflows from overseas.

[0048] It's understandable that WAFs are used for data collection and blocking. Because WAFs process network packets sequentially, performing security checks in a time-bound manner, they ensure timely data extraction; that is, requests are only allowed after being deemed risk-free. WAFs access the network through transparent or reverse proxies, giving them absolute control over network blocking. When data risks are detected, they can be blocked in real time, and subsequent access requests can be restricted. WAFs are typically deployed in perimeter security applications to protect application-layer network security and prevent hackers from breaching the network through the application layer.

[0049] As an example, and not a limitation, a distributed message processing cluster can refer to the Kafka message processing system. Kafka is a distributed streaming message processing system that can extract and process streaming data in real time. Streaming data refers to data continuously generated from thousands of data sources, typically sending data records simultaneously. Streaming platforms need to process this continuously flowing data, processing it sequentially. Kafka's time-series-based, high-performance processing capabilities are used for data relay and storage. This ensures that the operation of the original WAF system is not affected; furthermore, it allows the risk identification cluster to perform efficient subscription processing, guaranteeing the real-time and efficient nature of data processing.

[0050] As can be understood, a subscription relationship refers to the second detection system registering a subscription with the distributed message processing cluster in order to receive the data streams received by the cluster. Once the subscription relationship is established, the second detection system can obtain the data streams from the distributed message processing cluster in real time and perform relevant processing.

[0051] Through the subscription relationship, the second detection system can receive data streams from the distributed message processing cluster in real time, ensuring data synchronization and consistency.

[0052] S203. Utilize the first detection system and / or the second detection system to perform real-time data leakage risk detection in parallel on the data stream to be detected.

[0053] Data breach risk detection can be used to detect whether sensitive data such as personally identifiable information, financial information, and intellectual property information have been leaked in the data stream. As an example, and not a limitation, the data stream to be detected can be analyzed and tested using methods such as data classification, keyword matching, and regular expressions to determine whether a data breach has occurred.

[0054] It is understandable that using the first and / or second detection systems to perform real-time data leakage risk detection on the data stream to be detected in parallel means that the two detection systems can independently detect the data stream. If both systems indicate a potential leakage risk, the data leakage situation can be more confirmed. At the same time, parallel processing can save time and shorten the time for detecting the data stream to be processed, which helps to discover potential data leakage risks more quickly.

[0055] S204. When either the first or second detection system detects a risk of data leakage in the data stream to be detected, a data blocking operation is performed.

[0056] Data blocking refers to taking immediate measures to interrupt data transmission or access upon detecting a risk of data breach, in order to prevent the leakage of sensitive information. For example, when an attacker is detected attempting to download internal database information to an external source, the system can automatically block this behavior, protecting the security of user data.

[0057] The aforementioned data leakage prevention method ensures the real-time and accurate detection of data leakage risks by acquiring the latest data stream in real time. Parallel transmission of the data stream to be detected to the first detection system and the distributed message processing cluster improves system processing efficiency, while data backup to the distributed message processing cluster ensures data reliability and durability. The second detection system extracts the data stream from the distributed message processing cluster based on subscription relationships, guaranteeing data integrity and consistency while reducing intrusive operations on the original data. Parallel data leakage risk detection by the first and / or second detection systems enhances detection efficiency and timeliness, ensuring early detection of potential data leakage risks. Upon detection of a data leakage risk, data blocking operations are executed, enabling timely response to the risk and effective measures to protect data security, preventing data leakage or misuse.

[0058] In one embodiment, the first detection system is a network application firewall, and the second detection system is a data risk identification system. Real-time data leakage risk detection of the data stream to be detected is performed in parallel using the first and / or second detection systems, including: using the network application firewall to perform real-time data leakage risk detection on the data stream to be detected to obtain a first detection result; and using the data risk identification system to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules to obtain a second detection result. The preset risk identification rules include at least one of the following: sensitive data identification rules using regular expressions, access request frequency analysis rules, and access volume and access time analysis rules.

[0059] In one implementation, the Web application firewall identifies the data stream to be inspected through the ModSecurity module or other risk identification technologies, such as machine learning-based anomaly detection, used to identify and respond to various cybersecurity threats and risks.

[0060] ModSecurity is an open-source web application firewall that can detect and block malicious attacks in web applications, such as SQL injection and cross-site scripting attacks. ModSecurity uses a rule engine to analyze HTTP requests and responses, identify potential malicious behaviors, and take corresponding defensive measures based on a predefined set of rules, such as blocking, logging, and alerting.

[0061] For example, if a website is under attack by malicious SQL injection, ModSecurity can be configured with rules to detect and block such attacks. When a request containing malicious SQL statements is sent to the website, ModSecurity can analyze the request content, identify the characteristics of the SQL injection attack, and intercept or log the request according to pre-defined rules, thereby protecting the website from the attack.

[0062] The following section details several preset risk identification rules used by the second detection system.

[0063] (1) Set sensitive data identification rules using regular expressions

[0064] In the sensitive data identification rules using regular expressions, some common sensitive data patterns can be set, such as ID card numbers, credit card numbers, and mobile phone numbers. Regular expressions are used to match the content in the data stream. Once a matching sensitive data pattern is found, a data leakage risk alert is triggered. For example, a regular expression can be set to match all 15-digit and 18-digit ID card numbers. If the data stream contains content that matches this regular expression, the secondary detection system will detect a potential data leakage risk.

[0065] (2) Access Request Frequency Analysis Rules

[0066] Access request frequency analysis rules can be used to detect whether a user or system accesses data at an abnormal frequency. For example, a rule can be set up so that if a user frequently accesses a protected site within a short period of time, exceeding the normal access frequency range, it will be considered a data breach risk. Such rules can help detect potential malicious access behavior or data breach events.

[0067] (3) Rules for analyzing website traffic and access time

[0068] Access volume and access time analysis rules can be used to detect patterns and anomalies in data access. For example, a rule can be set to monitor the total amount of data accessed each day. If the access volume on a particular day significantly exceeds the average level, or if a large number of accesses occur outside of working hours, there may be a risk of data leakage. By analyzing patterns in access volume and access time, anomalies can be detected promptly, and appropriate measures can be taken.

[0069] In summary, the combination of the first detection system's use of a Web application firewall risk identification module and the second detection system's use of pre-defined risk identification rules such as sensitive data identification rules using regular expressions, access request frequency analysis rules, and access volume and access time analysis rules can improve the accuracy and efficiency of data leakage detection, promptly identify and respond to potential data leakage risks, and protect the security and integrity of data.

[0070] In one embodiment, the first detection system includes a first risk identification device, and the second detection system includes a second risk identification device. The data risk identification system performs real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules to obtain a second detection result. This includes: determining the virtual node number of the second risk identification device using a consistent hashing algorithm; performing hash calculations on the number name of the first risk identification device and the address name of the user terminal to obtain an association table, which records the association relationships between the first risk identification device, the second risk identification device, and the user terminal, where the user terminal represents the source of the data stream to be detected; and distributing the data stream to be detected to the corresponding virtual node of the second risk identification device according to the association table for real-time data leakage risk detection to obtain the second detection result.

[0071] The following is combined Figure 3 Let me explain in detail.

[0072] Figure 3 This is a schematic diagram of a hash ring provided in an embodiment of this application.

[0073] like Figure 3As shown, assume there are 3 WAF devices (first risk identification devices) with ID codes IDF1, IDF2, and IDF3 respectively; and 3 data processing servers (second risk identification devices) with codes A, B, and C respectively. The performance of the data processing servers can be identical, or their performance can be weighted according to specific criteria; this is not a limitation here.

[0074] In the initial distribution, a consistent hashing algorithm is used to distribute the nodes across 2... 32 On the annulus of data processing servers, assuming the number of virtual nodes is increased to 3, we can obtain:

[0075] Add numbers to node A to make it a virtual node, resulting in virtual node numbers: A-01, A-02, A-03;

[0076] Add numbers to node B to make it a virtual node, resulting in virtual node numbers: B-01, B-02, B-03;

[0077] Add numbers to node C to make it a virtual node, resulting in virtual node numbers: C-01, C-02, C-03.

[0078] Next, the value of the virtual node of each data processing server is recorded. Then, a hash calculation is performed based on the WAF ID and the last segment of the user's IP address, and the association table is recorded as shown in Table 1 below.

[0079] Table 1

[0080] Data risk identification server WAF+IP encoded packets A 1、4、7、11...... B 2、5、8、12...... C 3、6、9、13......

[0081] As shown in Table 1, the data streams to be detected are evenly distributed based on the number of data processing servers. Each data processing server only needs to subscribe to the distributed values. In other words, each data processing server subscribes to its corresponding data stream. For example, server A subscribes to the data stream allocated to node A after hash calculation, server B subscribes to the data stream allocated to node B after hash calculation, and so on.

[0082] The following describes the mechanism when a WAF device or data processing server malfunctions or changes.

[0083] If a new WAF device is added, it is only necessary to perform a hash calculation based on the WAF device's code and the last digit of its IP address, and then associate it with the data risk identification node. The pressure will still be balanced.

[0084] If a WAF device stops working, the previously allocated data does not need to be modified. Subscriptions simply mean there is no data to process, and it does not affect the processing of the entire service cluster.

[0085] If a new data processing server is added, the association table is recalculated. When the original client IP is disconnected or not connected, the new server subscribes, thus smoothly transitioning the data from the original processing server.

[0086] If the data processing server crashes, the association table will be recalculated, and the newly added service processing entity will be subscribed to immediately.

[0087] It is understandable that parallel data processing across multiple devices ensures high concurrency, fully utilizing system resources and maximizing data processing efficiency. Secondly, by establishing a correlation table using the data processing server's ID, WAFID, and user IP address, and distributing data based on this table, it is ensured that all data from the same user IP address is handled by the same data risk identification system, maintaining consistency and accuracy of processing results. Finally, the design incorporates a fault-to-load (or change) mechanism, which automatically redistributes tasks when individual data risk identification systems encounter anomalies, ensuring the continuity of data processing and high system availability, and achieving flexible rebalancing of the processing load.

[0088] In this way, the balanced processing of the data stream to be detected is ensured. Each data processing server only needs to focus on the data stream relevant to itself and will not be interfered with by the data processed by other servers. This can effectively utilize the performance of each data processing server and achieve distributed and efficient data risk identification and processing.

[0089] In one embodiment, a data risk identification system is used to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules to obtain a second detection result. This includes: retrieving historical detection data streams from the data risk identification system; and performing parallel analysis of the historical detection data streams and the data stream to be detected according to preset risk identification rules to obtain a second detection result.

[0090] As an example, and not a limitation, historical detection data streams can be extracted from databases, log files, or other data storage media. These streams contain information such as known malicious patterns and attack characteristics.

[0091] Then, based on the preset risk identification rules and related table mentioned above, the data stream to be detected and the historical detection data stream are simultaneously distributed to the corresponding second risk identification device for real-time data leakage risk detection, and the second detection result is obtained.

[0092] It is understandable that newly received data is analyzed in parallel with historically processed data, integrating information from different time periods to comprehensively assess potential data breach risks. This time-span data fusion analysis ensures the accuracy and comprehensiveness of the assessment.

[0093] In one embodiment, real-time acquisition of the data stream to be detected includes: configuring preset data acquisition rules, which include at least one of the following: target name, target port, and target protocol; and acquiring the data stream to be detected in real time using a network application firewall according to the preset data acquisition rules.

[0094] As an example, and not a limitation, you can configure data retrieval rules in a WAF to ensure that it can capture and process the required data streams. For example, you can define the target name (client or server), select the target protocol (such as HTTP, HTTPS), and set the target port. Understandably, the specific data retrieval rules can be customized according to your actual needs.

[0095] Then, the WAF captures the transmitted data stream through a network interface or proxy method according to the configured data acquisition rules.

[0096] Through the above steps, WAF can capture and organize the data stream to be detected, thereby providing support for subsequent data analysis, security detection, and protection.

[0097] In one embodiment, when a data leakage risk is detected in the data stream to be detected, a data blocking operation is performed, including: when a first detection result indicates a data leakage risk, terminating the network connection of the user address with the data leakage risk using a network application firewall.

[0098] Understandably, once a WAF determines that there is a risk of data leakage, it will accurately locate the associated network connection based on the corresponding user's IP address and immediately terminate the network connection.

[0099] In one example, based on preset policies, temporary access restrictions can be implemented or the IP address can be permanently added to the blacklist, effectively preventing potential future threats and strengthening the network security defense.

[0100] In one embodiment, when the second detection result indicates a risk of data leakage, based on the subscription relationship configured between the second detection system and the distributed message processing cluster, the second detection system sends an early warning notification to the distributed message processing cluster. The distributed message processing cluster is used to instruct the network application firewall to terminate the network connection of the user address that is at risk of data leakage according to the early warning notification.

[0101] When a data breach risk is identified, the second detection system is responsible for constructing a blocking command and pushing it to the Kafka message queue to initiate a network disconnection operation. Once Kafka receives the blocking command, it immediately notifies the relevant subscribed WAFs to pull the data. Subsequently, the WAF quickly obtains the data and executes emergency network connection blocking measures to ensure the timeliness and effectiveness of the response.

[0102] The following is combined Figure 4 Let's take a comprehensive look at the data breach prevention methods mentioned above.

[0103] Figure 4 This is an interactive flowchart of a data leakage prevention method provided in an embodiment of this application.

[0104] like Figure 4 As shown, the method is divided into two stages: the preparation stage and the processing stage.

[0105] The preparation phase mainly involves making advance preparations for high-performance operation, and includes the following:

[0106] S401 and WAF use connection pooling technology to establish and maintain long-term connections with the Kafka cluster.

[0107] Connection pools are used to manage the reuse and management of database connections, network connections, or other resource connections. In this context, connection pools are used to manage long-lived connections between the WAF (Web Application Firewall) and the Kafka cluster to improve connection reusability and efficiency.

[0108] Specifically, upon system startup, a certain number of connections are created and stored in a connection pool. These connections can be established with the Kafka cluster for data transfer and communication. When the system needs to communicate with the Kafka cluster, it obtains an available connection from the connection pool. If no connection is available in the pool, it can choose to wait or create a new connection (depending on the specific configuration). The system uses the connection to communicate with the Kafka cluster, transferring data and receiving responses. After communication is complete, the connection is returned to the connection pool for reuse.

[0109] By using connection pooling technology, the overhead of establishing and closing connections for each communication can be reduced, improving system performance and resource utilization. Connection pooling can also limit the number of connections, avoiding the problem of insufficient system resources caused by too many connections, and dynamically expand connections when needed, ensuring an optimal balance of resource usage while maintaining high performance.

[0110] S402, WAF subscription blocking messages.

[0111] WAF interacts with Kafka using a subscription model, ensuring it receives personalized blocking commands instantly. This allows it to quickly disconnect potentially vulnerable clients, effectively mitigating data breach risks. Specifically, the data risk processing system pushes messages to Kafka based on Topic (specific topic: killLink + WAFID) design, achieving precise target communication and command execution, enhancing the timeliness and accuracy of security responses.

[0112] S403, Push Risk Rules.

[0113] The rule setting system predefines the rules for risk identification, enabling the data risk identification system to accurately judge various risks, such as setting sensitive data identification using regular expressions, analyzing access requests by frequency, and analyzing data traffic by access volume and time.

[0114] It is understandable that designing an independent rule management module that supports dynamic rule editing allows for flexible adjustment of verification rules based on actual needs without interfering with the normal operation of the online system. This enables precise local deployment and verification, greatly improving the system's flexibility and stability.

[0115] The risk identification system preloads customized rule sets and pre-builds an optimized processing flow module matrix to ensure that efficient verification operations can be carried out immediately in each parallel module when data flows in, quickly identifying potential threats of data leakage and greatly improving the timeliness and accuracy of risk detection.

[0116] Understandably, the distributed architecture design of the data risk identification system not only ensures high scalability and excellent performance but also provides a robust foundation for system elasticity. This enables seamless integration of container orchestration technology when deployed in a cloud environment, allowing for dynamic scaling of resources to meet ever-changing business needs and maximize the flexibility and efficiency of cloud resources.

[0117] S404, Subscription Data.

[0118] The risk identification system establishes a stable connection with Kafka through a connection pool mechanism and pre-subscribes to the data streams that each module needs to process (with subsequent dynamic adjustments).

[0119] The subscription logic employs a ring distribution algorithm that integrates the WAF's ID, client identifier (IP), and risk identification system performance metrics weights (CPU, memory, network resources) (combined with the above). Figure 3 (As shown in the image), enabling smart subscription.

[0120] It is understandable that distributing data using the WAF ID and client identifier (IP) ensures that data originating from the same source is consistently processed by the same detection system, enhancing the accuracy of risk identification. Secondly, by utilizing WAF, Kafka, and the data risk identification system in parallel to identify data leakage risks, resource scheduling is optimized, the cluster's processing potential is fully explored, and data is guaranteed to be processed in real time with ultra-low latency. Finally, using the association table method mentioned above for data distribution ensures that even if a node in the cluster experiences a sudden failure, service continuity remains unaffected, improving the overall stability and resilience of the system.

[0121] The processing phase mainly involves efficient verification of all application-layer network data (client requests, server responses), and immediate blocking upon detection of problems. This includes the following:

[0122] S405, Receive network data packets.

[0123] This step involves the WAF collecting data from the network (either from the user or the server) and organizing the data stream to ensure that it can be identified and returned to the corresponding terminal in subsequent processing.

[0124] S406A, Data Risk Management.

[0125] WAF uses ModSecurity or other risk identification technologies to detect data breach risks.

[0126] S406B, synchronous data push.

[0127] Data is pushed to Kafka. The topic of the push message includes key elements such as data feature identifier (DATA field), WAF unique identifier (WAFID), and user source IP address to ensure that the information delivery is targeted and accurately matches the processing flow.

[0128] In other words, after the WAF receives and identifies the data, the data is processed in parallel by simultaneously executing steps S406A and S406B.

[0129] Since the data is not split here, smart pointer technology is used to distribute the data in parallel without copying it, thus avoiding the consumption of memory, CPU, and time, and further improving the real-time performance of the processing.

[0130] S407, Get subscription data.

[0131] Once Kafka receives the data, it immediately triggers a notification mechanism based on the previously configured subscription, sending a targeted notification to the associated data risk identification system. This system then responds quickly and retrieves the data, ensuring a seamless and efficient start-up of the risk assessment process.

[0132] S408, Data Risk Handling.

[0133] In this step, the system performs parallel analysis of newly received data and historically processed data based on a pre-defined risk rule framework, comprehensively considering both aspects to fully assess potential data breach risks. This time-span data fusion analysis ensures the accuracy and comprehensiveness of the judgment. Once risk indicators are identified, the system immediately triggers a pre-defined response mechanism, which may include issuing an alarm or activating step S409 to disconnect the network connection and synchronize alarm operations, quickly and effectively containing security threats.

[0134] S409, Risky, network blocked.

[0135] When step S408 determines that network blocking needs to be implemented, the system will activate this process, which is responsible for constructing the blocking command and efficiently pushing it to the Kafka message queue, thereby initiating the network disconnection operation.

[0136] S410, risky, network blocked.

[0137] Once Kafka receives the blocking command, it immediately notifies the relevant subscribed WAFs to pull data. Subsequently, the WAF quickly obtains the data and implements emergency network connection blocking measures to ensure the timeliness and effectiveness of the response.

[0138] S411, Block or Return.

[0139] Once the WAF receives a risk alert, it will accurately locate the associated network link based on the user's IP address contained in the message and immediately terminate the session. At the same time, according to preset policies, it can implement temporary access restrictions or permanently add the IP to the blacklist, effectively preventing potential future threats and strengthening the network security defense. It also ensures an immediate response to any identified security threats (including attacks or data breaches).

[0140] The foregoing mainly describes a data leakage prevention method according to an embodiment of this application with reference to the accompanying drawings. It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially, these steps are not necessarily executed in the order shown in the figures. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the steps or stages of other steps. The following describes an apparatus according to an embodiment of this application with reference to the accompanying drawings. For brevity, appropriate omissions will be made when describing the apparatus below; relevant content can be referred to in the relevant descriptions of the methods above, and will not be repeated.

[0141] Figure 5 This is a schematic diagram of the structure of a data leakage protection device provided in an embodiment of this application.

[0142] like Figure 5 As shown, the device 1000 includes the following units.

[0143] The acquisition unit 1001 is used to acquire the data stream to be detected in real time.

[0144] The processing unit 1002 is used to send the data stream to be detected to the first detection system and the distributed message processing cluster in parallel; wherein, the distributed message processing cluster and the second detection system are configured with a subscription relationship, and the second detection system is used to extract the data stream to be detected received by the distributed message processing cluster based on the subscription relationship.

[0145] The processing unit 1002 is also used to perform real-time data leakage risk detection in parallel using the first detection system and / or the second detection system on the data stream to be detected.

[0146] The processing unit 1002 is also used to perform a data blocking operation when either the first detection system or the second detection system detects a risk of data leakage in the data stream to be detected.

[0147] In one implementation, the device 1000 further includes a storage unit 1003, which can be used to store instructions and / or data, thereby implementing the method in the above embodiments.

[0148] In one embodiment, the storage unit 1003 is further configured to perform real-time data leakage risk detection on the data stream to be detected in parallel using a first detection system and / or a second detection system, including: the first detection system uses a Web application firewall risk identification module to perform real-time data leakage risk detection on the data stream to be detected and obtains a first detection result; the second detection system performs real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules and obtains a second detection result, wherein the preset risk identification rules include at least one of the following: sensitive data identification rules using regular expressions, access request frequency analysis rules, and access volume and access time analysis rules.

[0149] It is understandable that the first and second detection systems can perform data breach risk detection asynchronously, improving system efficiency and performance, and effectively preventing risks arising from vulnerabilities and incomplete rules. By using different detection rules and methods, the coverage of various data breach risks can be increased, avoiding potential security vulnerabilities.

[0150] In one embodiment, the storage unit 1003 is further configured to: determine the virtual node number of the second risk identification device using a consistent hashing algorithm; perform hash calculation on the number name of the first risk identification device and the address name of the user terminal to obtain an association table, the association table being used to record the association relationship between the first risk identification device, the second risk identification device and the user terminal, the user terminal being used to indicate the source of the data stream to be detected; distribute the data stream to be detected to the corresponding second risk identification device according to the association table, perform real-time data leakage risk detection, and obtain a second detection result.

[0151] It is understandable that consistent hashing algorithms can dynamically distribute data streams to different risk identification devices based on data characteristics, achieving load balancing, preventing any single device from being overloaded, and improving the overall performance and efficiency of the system. Distributing data streams to the appropriate risk identification devices based on an association table ensures that data streams are delivered to the most suitable device for detection, improving detection efficiency and accuracy. Recording the relationships between devices through the association table allows for adjustments to the mapping relationships between devices at any time, enabling dynamic configuration and management of devices, improving system flexibility and scalability. Recording the relationships between devices through the association table also ensures that even if a device fails or malfunctions, the system can quickly distribute the data stream to other backup devices for processing, ensuring the smooth progress of data detection tasks. By establishing an association table, it is guaranteed that the data processing on different devices is consistent, avoiding inconsistencies in results caused by differences in data processing on different devices.

[0152] In one embodiment, the storage unit 1003 is further configured to use the second detection system to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules, and obtain a second detection result, including: retrieving historical detection data streams from the second detection system; and performing parallel analysis on the historical detection data streams and the data stream to be detected according to preset risk identification rules to obtain a second detection result.

[0153] It is understandable that by performing parallel analysis of historical detection data streams and data streams to be detected, risks in the data streams can be detected in greater detail, improving the accuracy and efficiency of detection. At the same time, historical data comparisons can be used to better understand the current data stream situation, providing a more comprehensive risk identification service.

[0154] In one embodiment, the acquisition unit 1001 is further configured to acquire the data stream to be detected in real time, including: configuring a preset data acquisition rule, the preset data acquisition rule including at least one of the following: target name, target port, target protocol; and acquiring the data stream to be detected in real time using a network application firewall according to the preset data acquisition rule.

[0155] It's understandable that configuring data acquisition rules allows for the precise acquisition of the data streams that need to be monitored. This ensures that the system obtains critical data for risk identification, improving system efficiency and accuracy.

[0156] In one embodiment, the processing unit 1002 is further configured to perform a data blocking operation when a data stream to be detected is found to have a data leakage risk, including: when a first detection result indicates a data leakage risk, using a network application firewall to terminate the network connection of the user terminal address with the data leakage risk.

[0157] Understandably, when the initial detection results indicate a risk of data leakage, promptly terminating the network connection of the at-risk user address can effectively prevent further data leakage and protect data security.

[0158] In one embodiment, the processing unit 1002 is further configured to, when the second detection result indicates a risk of data leakage, send an early warning notification to the distributed message processing cluster; the distributed message processing cluster instructs the network application firewall to terminate the network connection of the user address that is at risk of data leakage according to the early warning notification.

[0159] Understandably, when the second detection result indicates a risk of data breach, the system can promptly send an early warning notification to the distributed message processing cluster and instruct the first detection system to terminate the network connection of the risky client address. This enables rapid response and handling of data breach events, preventing further data leakage and greater losses. Timely early warnings and automated processing improve the efficiency and accuracy of responding to data breach events.

[0160] It should be noted that the information interaction and execution process between the above-mentioned units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0161] Figure 6 This is a schematic diagram of the structure of the computer device provided in an embodiment of this application. Figure 6 As shown, the computer device 6000 of this embodiment includes: at least one processor 6100 ( Figure 6 (Only one is shown) a processor, a memory 6200, and a computer program 6210 stored in the memory 6200 and executable on at least one processor 6100, wherein when the processor 6100 executes the computer program 6210, the computer device performs the steps described in the above embodiments.

[0162] The processor 6100 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0163] In some embodiments, memory 6200 may be an internal storage unit of computer device 6000, such as a hard disk or memory of computer device 6000. In other embodiments, memory 6200 may be an external storage device of computer device 6000, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on computer device 6000. Furthermore, memory 6200 may include both internal and external storage units of computer device 6000. Memory 6200 is used to store operating system, application programs, boot loader data, and other programs, such as program code for computer programs. Memory 6200 may also be used to temporarily store data that has been output or will be output.

[0164] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units is merely an example. In practical applications, the above functions can be assigned to different functional units or modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0165] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a computer device, it enables the computer device to perform the steps described in the above-described method embodiments.

[0166] This application provides a computer program product that, when run on a computer device, enables the computer device to implement the methods described above.

[0167] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it enables a computer device to implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0168] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In the description, specific details such as particular system structures and technologies are set forth for illustrative purposes rather than for limiting purposes, so as to provide a thorough understanding of the embodiments of this application. However, those skilled in the art should understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary details.

[0169] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0170] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0171] Furthermore, in the description of this application and the appended claims, the terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0172] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0173] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0174] In the embodiments provided in this application, it should be understood that the disclosed apparatus, computer equipment, and methods can be implemented in other ways. For example, the apparatus and computer equipment embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0175] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A data leakage prevention method, characterized in that, include: Configure preset data acquisition rules, which include at least one of the following: target name, target port, and target protocol; Based on preset data acquisition rules, the network application firewall is used to acquire the data stream to be detected in real time. The data stream to be detected is sent in parallel to the first detection system and the distributed message processing cluster; wherein, the distributed message processing cluster and the second detection system are configured with a subscription relationship, and the second detection system is used to extract the data stream to be detected received by the distributed message processing cluster based on the subscription relationship; The first detection system and / or the second detection system are used to perform real-time data leakage risk detection on the data stream to be detected in parallel. When either the first detection system or the second detection system detects a risk of data leakage in the data stream to be detected, a data blocking operation is performed.

2. The method according to claim 1, characterized in that, The first detection system is a network application firewall, and the second detection system is a data risk identification system. The step of using the first detection system and / or the second detection system to perform parallel real-time data leakage risk detection on the data stream to be detected includes: The network application firewall is used to perform real-time data leakage risk detection on the data stream to be detected, and a first detection result is obtained; The data risk identification system is used to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules, and a second detection result is obtained. The preset risk identification rules include at least one of the following: sensitive data identification rules with regular expression, access request frequency analysis rules, and access volume and access time analysis rules.

3. The method according to claim 2, characterized in that, The first detection system includes a first risk identification device, and the second detection system includes a second risk identification device. The step of using the data risk identification system to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules, and obtaining a second detection result, includes: The virtual node number of the second risk identification device is determined using a consistent hashing algorithm; A hash calculation is performed on the number name of the first risk identification device and the address name of the user terminal to obtain an association table. The association table is used to record the association relationship between the first risk identification device, the second risk identification device and the user terminal. The user terminal is used to indicate the source of the data stream to be detected. The data stream to be detected is distributed to the virtual node of the corresponding second risk identification device according to the association table, and real-time data leakage risk detection is performed to obtain the second detection result.

4. The method according to claim 2, characterized in that, The process of using the data risk identification system to perform real-time data leakage risk detection on the data stream to be detected according to preset risk identification rules, and obtaining a second detection result, includes: Retrieve historical detection data streams from the data risk identification system; The historical detection data stream and the data stream to be detected are analyzed in parallel according to the preset risk identification rules to obtain the second detection result.

5. The method according to any one of claims 1-4, characterized in that, When either the first detection system or the second detection system detects a data leakage risk in the data stream to be detected, a data blocking operation is performed, including: If the first detection result indicates a risk of data leakage, the network connection of the user address with the risk of data leakage is terminated using a network application firewall.

6. The method according to claim 5, characterized in that, The method further includes: When the second detection result indicates a risk of data leakage, based on the subscription relationship configured between the second detection system and the distributed message processing cluster, the second detection system sends an early warning notification to the distributed message processing cluster. The distributed message processing cluster then instructs the network application firewall to terminate the network connection of the user address that poses a risk of data leakage, according to the early warning notification.

7. A data leakage protection device, characterized in that, include: The acquisition unit is used to acquire the data stream to be detected in real time; A processing unit is configured to send the data stream to be detected in parallel to a first detection system and a distributed message processing cluster; wherein the distributed message processing cluster and the second detection system are configured to have a subscription relationship, and the second detection system is configured to extract the data stream to be detected received by the distributed message processing cluster based on the subscription relationship; The processing unit is also used to perform real-time data leakage risk detection on the data stream to be detected in parallel using the first detection system and / or the second detection system. The processing unit is also configured to perform a data blocking operation when either the first detection system or the second detection system detects a data leakage risk in the data stream to be detected. The acquisition unit is used for: Configure preset data acquisition rules, which include at least one of the following: target name, target port, and target protocol; Based on preset data acquisition rules, the network application firewall is used to acquire the data stream to be detected in real time.

8. A computer device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the computer device performs the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a computer device, implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Pervasive, domain and situational-aware, adaptive, automated, and coordinated analysis and control of enterprise-wide computers, networks, and applications for mitigation of business and operational risks and enhancement of cyber security

    US20130104236A1

  • Distributed security testing system, method and device, and storage medium

    WO2021097713A1