A method, product, device, and medium for detecting data leakage from a disk array
By acquiring audit logs within the disk array system and analyzing the characteristics of outgoing data packets, the problem of low accuracy and high resource consumption in disk array data leak detection is solved, achieving efficient and accurate data leak detection and timely alerts.
Patent Information
- Application Number
- CN202511188337.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-25
AI Technical Summary
In existing technologies, disk array data leakage detection has low accuracy and excessive resource consumption. DLP data leakage protection tools cannot understand the unique operations of disk array systems, resulting in a high false alarm rate, inability to detect encrypted traffic, and a decrease in storage network throughput when deployed in data centers.
By obtaining audit log information from within the disk array system to detect operational behavior, and combining this with data feature analysis of outgoing data packets, normal and risky operational behaviors can be identified, improving detection accuracy. Furthermore, detection can be performed internally within the system to reduce resource overhead.
It improves the accuracy of disk array data leak detection, reduces resource overhead, avoids impacting storage network throughput, and can promptly detect data leaks and issue alerts.
Smart Images

Figure CN120723561B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, product, device and medium for detecting data leakage in disk arrays. Background Technology
[0002] Business data in the operational systems is primarily stored on disk arrays. To ensure data security, it's necessary to monitor data during disk array system operation to prevent data leaks. One common technology involves deploying Data Leakage Prevention (DLP) tools in the data center. These tools detect potential data leaks by analyzing network traffic to identify sensitive data transmission and monitoring abnormal storage device operations. However, DLP tools cannot understand the unique operations of disk array systems, leading to an inability to accurately distinguish between normal data operations and malicious data leaks. Furthermore, they cannot detect data entering the data transmission channel, resulting in low detection accuracy. Additionally, deploying DLP tools within the data center reduces storage network throughput, leading to excessive resource consumption.
[0003] Therefore, improving the accuracy of disk array data leakage detection and reducing resource overhead have become problems that need to be solved by those skilled in the art. Summary of the Invention
[0004] This application provides a method, computer program product, device, and computer-readable storage medium for detecting data leakage in disk arrays, in order to at least solve the problems of low accuracy and excessive resource consumption in the related art for detecting data leakage in disk arrays.
[0005] This application discloses a method for detecting data leakage in a disk array, applied to a disk array system. The method includes:
[0006] Obtain audit log content information, perform data leakage detection on the operational behavior of the audit log content information, and obtain the first detection result; the audit log content information includes at least one audit log, which is generated by various business models in the disk array system when performing operations;
[0007] If the initial detection result indicates no risk of data leakage, obtain the outgoing data packets generated by the disk array system;
[0008] Data leakage detection is performed on outgoing data packets based on their data characteristics to obtain a second detection result;
[0009] If the second test result indicates that there is no risk of data leakage, it is determined that no data leakage has occurred in the disk array system;
[0010] If the first detection result indicates a risk of data leakage, and / or the second detection result indicates a risk of data leakage, then a data leakage is determined to have occurred in the disk array system.
[0011] This application also provides a disk array data leakage detection device, applied to a disk array system, the device comprising:
[0012] The log analysis module is used to obtain audit log content information, perform data leakage detection on the operational behavior of the audit log content information, and obtain the first detection result. The audit log content information includes at least one audit log, which is generated by the various business models in the disk array system when they perform operations.
[0013] The first acquisition module is used to acquire outgoing data packets generated by the disk array system when the first detection result indicates that there is no risk of data leakage.
[0014] The data packet analysis module is used to perform data leakage detection on outgoing data packets by analyzing their data characteristics, and to obtain a second detection result.
[0015] The first determining module is used to determine that no data leakage has occurred in the disk array system if the second detection result indicates that there is no risk of data leakage.
[0016] The second determining module is used to determine that a data breach has occurred in the disk array system when the first detection result indicates a risk of data leakage, and / or the second detection result indicates a risk of data leakage.
[0017] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for detecting data leakage from a disk array.
[0018] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the disk array data leakage detection methods described above.
[0019] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the disk array data leakage detection methods described above.
[0020] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows:
[0021] The disk array data leakage detection method provided in this application is applied within the disk array system. Since various operations of each business module in the disk array system are recorded in the audit log, this application can obtain the audit log information and perform data leakage detection on the operational behavior of the audit log information to accurately identify normal and risky operations, thereby obtaining a first detection result. If the first detection result obtained by detecting the audit log information indicates that there is no data leakage risk, this application can obtain the outgoing data packets generated by the disk array system and perform data leakage detection on the data characteristics of the outgoing data packets before they enter the transmission channel, which can further obtain a second detection result indicating whether there is a data leakage risk. If the second detection result indicates that there is no data leakage risk, it can be determined that the disk array system has not experienced a data leakage. If the first detection result indicates a data leakage risk, and / or the second detection result indicates a data leakage risk, it can be determined that the disk array system has experienced a data leakage. Therefore, this application can improve the accuracy of data leakage detection by capturing audit log content information and outgoing data packets within the disk array system. Furthermore, since this method is applied within the disk array system, it does not affect the throughput of the storage network and can reduce the resource overhead of the disk array system.
[0022] Furthermore, this invention also provides corresponding computer program products, electronic devices, and computer-readable storage media for the detection method of disk array data leakage, further making the method more practical, and the device, electronic device, and computer-readable storage media have corresponding advantages. Attached Figure Description
[0023] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating a method for detecting data leakage in a disk array, provided in an embodiment of this application;
[0025] Figure 2 This application provides a schematic diagram of a data leakage detection process for manipulating audit log content information, as part of an embodiment of the present application.
[0026] Figure 3 This is a schematic diagram illustrating a data leakage detection process that uses data characteristics to analyze outgoing data packets, as provided in an embodiment of this application.
[0027] Figure 4 This is a schematic diagram illustrating a data breach detection process using a data breach detection model, provided as an embodiment of this application.
[0028] Figure 5 This is a schematic diagram of the structure of a disk array data leakage detection device provided in an embodiment of this application;
[0029] Figure 6 This is a schematic diagram of another disk array data leakage detection device provided in an embodiment of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0031] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0032] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] The specific application environment architecture or specific hardware architecture on which the detection method for data leakage of combined disk arrays depends is described here.
[0034] It's important to note that with the development and expanding application of technologies such as cloud computing, the Internet of Things, and mobile computing, the requirements and emphasis on information security are increasing. Business system data is a core asset, primarily stored on storage devices, mainly disk array systems, which have become the core infrastructure of current data centers. As the amount of data stored on disk arrays increases, the risk of data leakage also grows. Data stored on disk arrays can be divided into two categories: structured data and unstructured data. Structured data refers to structured information such as databases and tables, while unstructured data refers to data types such as text, images, audio, and video that lack a fixed format and organizational structure. Data leakage can have a significant impact on enterprises; therefore, dynamically detecting potential data leaks during the operation of disk array systems is becoming increasingly important.
[0035] In related technologies, data leakage detection in disk array systems primarily involves deploying DLP (Data Leakage Prevention) tools in data centers. These tools detect potential data leaks by analyzing network traffic to identify sensitive data transmission and monitoring abnormal storage device operations. However, this approach often suffers from several problems: First, DLP tools cannot understand the specific operations of disk array systems (e.g., the binary logs of the RAID (Redundant Arrays of Independent Disks) controller, RAID stripe offsets, and cache states), leading to an inaccurate distinction between normal data operations and malicious data leakage, resulting in a high false positive rate. Second, DLP tools are based on traffic flow detection; they cannot detect data leaks transmitted via HTTPS (Hypertext Transfer Protocol Secure) or encrypted tunnels, further reducing overall detection accuracy. Third, using DLP tools for full-traffic deep testing of disk array systems can cause a more than 30% drop in storage network throughput, resulting in excessive resource consumption.
[0036] Therefore, this application provides a method for detecting data leakage in disk arrays, which can improve detection accuracy, reduce resource consumption, and enhance the overall security of the disk defragmentation system. The method is described in detail below with reference to its execution flow. Please refer to... Figure 1 The flowchart shown illustrates a method applied to a disk array system, comprising the following steps S110 to S150.
[0037] S110: Obtain audit log content information, perform data leakage detection on the operation behavior of the audit log content information, and obtain the first detection result; the audit log content information includes at least one audit log, which is generated by each business model in the disk array system when performing operations.
[0038] It should be noted that in practical applications, during the operation of a disk array system, all operations of various business modules (whether configuration management operations performed by administrators or internal operations automatically performed by the software (such as cache refresh, stripe reconstruction, etc.)) are recorded in the audit log file. The recorded audit logs can cover all critical behaviors, including but not limited to login, logout, locking, unlocking, user management, password management, authorization management, core security configuration, access control policies, automatic update policies, security monitoring policies, access to important resources (such as the creation, deletion, and modification of logical volumes), software updates, establishment of remote replication relationships, cache refresh, stripe reconstruction, etc. In addition to the critical behaviors mentioned above, the information recorded in the audit logs may also include, but is not limited to, the time of the event, user ID (including associated terminal, port, IP address, data, device, etc.), event type, name of the accessed resource, and the result of the event (success or failure of the operation).
[0039] In addition, different types of audit logs generated during the operation of the disk array system may include, but are not limited to, the binary log of the RAID controller, the stripe offset anomaly log of the RAID controller, the cache status anomaly log of the disk array system, and the abnormal stripe reconstruction log.
[0040] Understandably, since audit logs record detailed information corresponding to operations, data leakage detection of operational behaviors can be performed by obtaining the audit log content (including at least one audit log entry) from the disk array system and using this audit log content. This detection can determine whether the corresponding operation is a normal operation or a risky operation, thus obtaining the corresponding initial detection result. Because the audit log content information is related to the operation of the disk array system, data leakage detection of operational behaviors using audit log content can accurately determine whether there is a data leakage risk in the disk array system.
[0041] S120: If the first detection result indicates that there is no risk of data leakage, obtain the outgoing data packets generated by the disk array system.
[0042] It should be noted that all traffic transmission operations of the disk array system must undergo data organization by the protocol packet assembly module to obtain outgoing data packets to be sent. In this application, after performing data leakage detection on operational behavior by analyzing the audit log content, if the first detection result indicates that there is no data leakage risk in the disk array system, it only means that the operational behavior itself is free from data leakage risk. Further determination is still needed to determine whether the disk array's traffic transmission operations pose a data leakage risk. Therefore, in this application, outgoing data packets generated by the disk array system can be obtained even if the first detection result indicates no data leakage risk (i.e., the operational behavior itself is free from data leakage risk).
[0043] S130: Perform data leakage detection on outgoing data packets based on data characteristics to obtain a second detection result.
[0044] It is understandable that, after obtaining the outgoing data packet, this application can further perform data leakage detection on the data packet's data characteristics to obtain a second detection result. Since this application obtains the outgoing data packet and performs data leakage detection on its data characteristics before the traffic enters the transmission channel (e.g., HTTPS or encrypted tunnel) for transmission, this application can solve the problem that DLP data leakage prevention tools in related technologies cannot detect encrypted traffic.
[0045] S140: If the second test result indicates that there is no risk of data leakage, it is determined that no data leakage has occurred in the disk array system.
[0046] It should be noted that after obtaining the second detection result, it can be used to determine whether there is a risk of data leakage in the data traffic. If the second detection result indicates that there is no risk of data leakage, it means that the transmission of data traffic is secure and there is no risk of data leakage. Therefore, in this application, by detecting the audit log content and outgoing data packets, and confirming that there is no risk of data leakage in the operational behavior and the data traffic, it can be determined that the disk array system has not experienced data leakage, thereby ensuring the accuracy of the detection.
[0047] S150: If the first detection result indicates a risk of data leakage, and / or the second detection result indicates a risk of data leakage, it is determined that a data leakage has occurred in the disk array system.
[0048] Of course, if the first detection result after performing data leakage detection on the audit log content indicates a data leakage risk, meaning the operation itself poses a data leakage risk, then it can be determined that the disk array system has experienced a data leakage. In this case, it is unnecessary to obtain outgoing data packets generated by the disk array system for data feature-based data leakage detection. If the first detection result indicates no data leakage risk, and a second detection result is obtained after obtaining outgoing data packets and performing data feature-based data leakage detection, and this second detection result indicates a data leakage risk, then it can be determined that the disk array system has experienced a data leakage.
[0049] It should also be noted that, in order to enable staff to quickly detect data breaches and take timely countermeasures, this application specifies that in the event of a data breach in the disk array system, an alarm will be issued. For example, data breach alarms can be issued via email, SMS, alarm device notifications, etc., so that staff can promptly carry out maintenance and protection of the disk array system.
[0050] Therefore, this application can obtain audit log content information and perform data leakage detection on the operational behavior of the audit log content information to accurately identify normal and risky operational behaviors, thereby obtaining a first detection result. If the first detection result obtained by detecting the audit log content information indicates that there is no data leakage risk, this application can obtain outgoing data packets generated by the disk array system and perform data leakage detection on the data characteristics of the outgoing data packets before they enter the transmission channel, which can further obtain a second detection result indicating whether there is a data leakage risk. If the second detection result indicates that there is no data leakage risk, it can be determined that the disk array system has not experienced a data leakage. If the first detection result indicates a data leakage risk, and / or the second detection result indicates a data leakage risk, it can be determined that the disk array system has experienced a data leakage. Thus, this application improves the accuracy of data leakage detection by capturing audit log content information and outgoing data packets within the disk array system. Furthermore, since this method is applied within the disk array system, it does not affect the throughput of the storage network and can reduce the resource overhead of the disk array system.
[0051] Based on the above embodiments, the embodiments of this application further illustrate and optimize the technical solution.
[0052] In one implementation, please refer to Figure 2 The process of obtaining audit log content information in S110 of this application, performing data leakage detection on the operation behavior of the audit log content information, and obtaining the first detection result may include the following steps S210 to S290.
[0053] S210: If at least one business module in the disk array system generates audit logs, obtain the newly added audit logs.
[0054] It should be noted that, in this embodiment of the application, when it is detected that a business module in the disk array system has generated audit logs, the newly added audit log can be obtained from the audit log file. That is, the audit log content information obtained at this time includes the newly added audit log. Specifically, by identifying whether the audit date file has changed, if the audit log file has changed, it can be determined that a business module has generated audit logs, and at this time, the newly added audit log can be obtained.
[0055] S220: Determine whether the corresponding operation is an attack based on the newly added audit log. If the operation is an attack, the first detection result is determined to be that there is a risk of data leakage.
[0056] Understandably, the operation corresponding to the newly added audit log can be determined based on the log information in the new audit log, and then it can be further determined whether the operation is an attack. If the operation is an attack, it can be determined that the disk array system has a data leakage risk. That is, if the operation is an attack, the first detection result is obtained based on the determination that the operation is an attack, and the first detection result is that there is a data leakage risk.
[0057] In one implementation, the process of determining whether the corresponding operation behavior is an attack behavior based on the newly added audit log in S220 may include: determining the corresponding operation behavior based on the operation behavior information in the newly added audit log; matching the operation behavior with each attack behavior in a pre-established attack behavior database; if there is an attack behavior in the attack behavior database that is consistent with the operation behavior, then the operation behavior is determined to be an attack behavior; if there is no attack behavior in the attack behavior database that is consistent with the operation behavior, then the operation behavior is determined not to be an attack behavior.
[0058] It should be noted that in practical applications, an attack behavior database can be pre-established. For example, this can be achieved by collecting attack behaviors corresponding to various attack methods suffered by different disk array systems. This database can include the operations and attack methods corresponding to each attack behavior. Therefore, the corresponding operation behavior can be determined based on the operation behavior information in the newly obtained audit logs. Then, the operation behavior corresponding to this behavior can be matched with the operations corresponding to each attack behavior in the database. If a match is found, the operation behavior can be identified as an attack behavior, and the corresponding attack method can be further identified, triggering an alert. If the matching reveals that there is no corresponding attack behavior in the database, the operation behavior is not an attack behavior and is considered normal. For example, if the operation behavior information includes frequent use of SQL keywords such as SELECT * or BULK INSERT, or frequent use of special characters (such as single quotes or comment characters), then the operation behavior is an attack behavior, indicating a data leak. Frequent use of SELECT * or special characters (such as single quotes or comment characters) indicates malicious data extraction or injection attacks.
[0059] The method described in this application can easily and efficiently identify whether an operation is an attack, thereby improving detection efficiency and accuracy.
[0060] S230: If the operation is not an attack, then obtain each first audit log within the first preset time period before the new audit log is added.
[0061] In this embodiment of the application, in order to further improve the accuracy of data leakage detection, if it is determined above that the operation behavior corresponding to the newly added audit log is not an attack behavior, then each first audit log in the audit log file located within a first preset time period (e.g., 5 minutes) before the newly added audit log can be further obtained.
[0062] S240: Based on the newly added audit logs and each first audit log, determine whether the operation behavior deviates from the corresponding normal behavior baseline.
[0063] It should be noted that this application pre-establishes corresponding normal behavior baselines for different normal operating behaviors. By combining the operating behaviors in the newly added audit logs and the corresponding operating behaviors in each of the first audit logs, the overall operating behavior within this period can be determined. Based on this overall operating behavior and the normal behavior baseline corresponding to the operating behavior, it can be determined whether the operating behavior deviates from the normal behavior baseline. Specifically, based on the operation type of the operating behavior, it can be determined whether there is a preceding operation. If there is a preceding operation, the actual preceding operation can be determined based on the overall operating behavior. Then, the actual preceding operation is further compared with the preceding operation in the corresponding normal behavior baseline. If the two are consistent, it means that the operating behavior has not deviated from the normal behavior baseline. If the actual preceding operation is different from the preceding operation in the corresponding normal behavior baseline, it means that the operating behavior has deviated from the normal behavior baseline. If the operation type determines that the operation has no prerequisite behavior, but it is necessary to limit the operation frequency of the operation to not exceed the corresponding normal behavior baseline, the operation frequency of the operation can be determined based on the overall operation behavior situation. For example, for access operations, the actual access frequency can be determined, and then the actual access frequency can be compared with the corresponding normal behavior baseline (normal access frequency). If the actual access frequency is greater than the corresponding normal behavior baseline, it means that the operation has deviated from the corresponding normal behavior continuation baseline.
[0064] S250: If the operation deviates from the corresponding normal behavior baseline, the first detection result is determined to be a risk of data leakage.
[0065] Understandably, if an operation deviates from the normal behavior baseline—for example, if the operation frequency exceeds the corresponding normal frequency value, or if the preceding behavior of the operation is inconsistent with the preceding behavior in the continuation of normal behavior—it indicates that the disk array system has a data leakage risk. Therefore, in this embodiment, when it is determined that the operation deviates from the normal behavior baseline, a first detection result is generated based on the conclusion that the operation deviates from the normal behavior baseline. This first detection result indicates that there is a data leakage risk.
[0066] For example, in practical applications, audit logs from the daytime can be correlated with audit logs from the nighttime, and audit logs from login events can be correlated with audit logs from accessing data, so as to detect complex data leakage behaviors such as hackers probing the system during the day and stealing data at night.
[0067] S260: If the operation does not deviate from the corresponding normal behavior baseline, then obtain each second audit log within the second preset time period before the newly added audit log.
[0068] Furthermore, if the operational behavior does not deviate from the corresponding normal behavior baseline, in order to further improve the detection accuracy in this embodiment, various second audit logs within a second preset time period prior to the newly added audit log can be obtained. This second preset time period can be longer than the first preset time period, and can be one day, etc. This allows for the acquisition of more audit logs, enabling subsequent correlation analysis of multi-time period and multi-type audit logs to better uncover deeper data leakage behaviors.
[0069] S270: Correlation analysis will be performed on newly added audit logs and each secondary audit log to determine the corresponding behavioral links.
[0070] It is understood that, in this embodiment of the application, the behavioral links between these operational behaviors can be determined according to the time sequence of each audit log (i.e., the time sequence of the operational behaviors) based on the operational behaviors corresponding to the newly added log and the operational behaviors corresponding to each of the second audit logs.
[0071] In practical applications, an attack behavior link library can be pre-established. This library can include multiple attack behavior links, which are formed by various operational behaviors arranged in chronological order. In this application, after obtaining a behavior link, it can be matched against each attack behavior link in the pre-established attack behavior link library. If an attack behavior link matching the behavior link exists in the library, the behavior link is determined to be an attack link; otherwise, it is determined to be a normal link. For example, if various second audit logs indicate multiple failed login attempts followed by subsequent operational behaviors indicating successful access to sensitive data, such behavior links can be further matched against each attack behavior link in the attack behavior link library to discover internal leaks or potential attack chains.
[0072] S280: If the behavioral link is an attack link, the first detection result is determined to be a risk of data leakage.
[0073] It should be noted that if the link is identified as an attack link, the first detection result can be obtained based on the conclusion that the link is an attack link, which indicates that there is a risk of data leakage.
[0074] S290: If the behavioral link is a normal link, the first detection result is determined to be that there is no risk of data leakage.
[0075] If the link is determined to be a legitimate connection and not an attack link, then a first detection result can be obtained based on the conclusion that the link is legitimate, indicating that there is no risk of data leakage. In this embodiment, not only are the operational behaviors of newly added audit logs considered, but also the operational behaviors corresponding to each first audit log within a first preset time period, or the operational behaviors corresponding to each second audit log within a second preset time period, can be combined for overall analysis to detect deeper levels of data leakage and improve detection accuracy.
[0076] In one implementation, the process of performing data leakage detection on outgoing data packets using data characteristics in S130 to obtain a second detection result may include:
[0077] Based on the outgoing data packets, determine whether the data is subject to at least one of the following: sudden increase in traffic, batch data reading behavior at unusual time periods, unusual external connections, data slicing of different protocols, abnormal disk array command streams, or data theft disguised as disk failure. If at least one of these conditions exists, the second detection result indicates a risk of data leakage; if none of these conditions exist, the second detection result indicates no risk of data leakage.
[0078] It should be noted that in practical applications, after obtaining the outgoing data packet, the disk array system's data can be further analyzed based on the outgoing data packet to determine if at least one of the above conditions exists, and a second detection result can be generated.
[0079] In one implementation, the process of determining whether there is at least one of the following situations based on outgoing data packets: a sudden increase in data traffic, batch data reading behavior during unusual time periods, unusual external connections, data slicing of different protocols, abnormal disk array command streams, or data theft by faking disk failures, may include the following steps.
[0080] Based on the size of the outgoing data packets and the baseline of normal traffic for a single session, determine whether there is a sudden increase in data traffic.
[0081] It should be noted that in this embodiment, a normal traffic baseline for a single session can be preset. After obtaining the outgoing data packet, the size of the outgoing data packet can be further determined. Based on the size of the outgoing data packet and the normal traffic baseline for a single session, it can be determined whether there is a sudden increase in data traffic in the disk array system.
[0082] In one implementation, the process of determining whether there is a traffic surge based on the size of the outgoing data packet and the normal traffic baseline of a single session may include:
[0083] Check if the size of outgoing data packets exceeds the normal traffic baseline for a single session;
[0084] If the size of the outgoing data packet is greater than the normal traffic baseline for a single session, it indicates that there is a sudden increase in data traffic.
[0085] The size of the outgoing data packet is compared to the normal traffic baseline for a single session. If the size of the outgoing data packet is larger than the normal traffic baseline for a single session, it indicates a sudden increase in data traffic on the disk array system. Understandably, in the presence of a traffic surge, a second detection result can be further confirmed as a risk of data leakage.
[0086] In addition, to further improve detection accuracy, the method may also include:
[0087] If the size of the outgoing data packet is less than or equal to the normal traffic baseline of a single session, retrieve all other outgoing data packets located within the fifth preset time period before the outgoing data packet;
[0088] Based on the size of the outgoing data packets and the size of each other outgoing data packet within the fifth preset time period preceding the outgoing data packets, determine whether the total outgoing data packet size is greater than the total normal traffic baseline.
[0089] If the total size of outgoing data packets exceeds the baseline of normal traffic, a sudden increase in data traffic is identified.
[0090] It should be noted that in this embodiment, a total normal traffic baseline can be preset. If the size of the outgoing data packet is less than or equal to the size of the single-session normal traffic baseline (where a single session transmits GB-level data), that is, the outgoing data packet does not exceed the single-session normal traffic baseline, in order to further detect whether there is a sudden increase in the data of the disk array system, each other outgoing data packet within the fifth preset time period before the outgoing data packet can be obtained, and then the size of each other outgoing data packet can be determined. The sizes of each other outgoing data packet are summed to obtain the total outgoing data packet size, and the total outgoing data packet size is compared with the size of the total normal traffic baseline. If the total outgoing data packet size is greater than the size of the total normal traffic baseline, it indicates that there is a sudden increase in the data traffic of the disk array system.
[0091] Of course, in practical applications, the size of the total outgoing data packets can be subtracted from the size of the total normal traffic baseline to obtain the difference. Then, the ratio of the difference to the total normal traffic baseline can be calculated. If the ratio is greater than the preset ratio (such as 80%), it indicates that there is a sudden increase in data traffic in the disk array system.
[0092] Understandably, in the case of a sudden surge in traffic, the second detection result can be further confirmed as a risk of data leakage.
[0093] And / or, based on the current time corresponding to the outgoing data packet and the time period corresponding to the current time, obtain the data standard corresponding to the time period, and determine whether the outgoing data packet conforms to the data standard. If the outgoing data packet does not conform to the data standard, determine that there is a batch data reading behavior in an irregular time period.
[0094] It is understandable that, in order to detect whether there is batch data reading behavior in the disk array system during unusual time periods, multiple time periods can be pre-divided, such as dividing a day into multiple time periods. Based on the historical operation of the disk array system, the data standard corresponding to each time period under normal operation can be determined, such as the curve of data volume changing over time. Then, based on the current time corresponding to the acquired outgoing data packet, the time period and corresponding data standard of the current time can be determined. Then, the outgoing data packet is compared with the data standard. If the outgoing data packet does not conform to the data standard, for example, if the data volume of the outgoing data packet exceeds the maximum data volume of the data standard, it can be determined that the outgoing data packet does not conform to the data standard, thereby determining that there is batch data reading behavior in the disk array system during very high time periods. For example, if the time period corresponding to the current time is the early morning period, and the number of times the outgoing data packet accesses the database LUN (Logical Unit Number) reaches M times, but the number of times the database LUN is accessed in the data standard corresponding to the early morning period does not exceed N times, where N is less than M, it indicates that there is batch data reading behavior during very high time periods.
[0095] And / or, determine whether there is an unconventional outbound connection based on the destination Internet Protocol and / or destination port of the outbound data packet.
[0096] It is understandable that in the process of determining whether there is an irregular external connection based on the outgoing data packet, the destination IP and destination port to which the outgoing data packet needs to be sent can be determined based on the outgoing data packet, and then the existence of an irregular external connection can be further identified based on the destination IP and destination port.
[0097] In one implementation, the process of determining whether data has an unconventional outbound connection based on the destination Internet Protocol and / or destination port of the outbound data packet may include:
[0098] The system detects whether the destination Internet Protocol of the outgoing data packet is an abnormal Internet Protocol, and / or whether the destination port of the outgoing data packet is an abnormal port. If the destination Internet Protocol is an abnormal Internet Protocol, and / or the destination port is an abnormal port, then it is determined that there is an abnormal outgoing connection. If the destination Internet Protocol is not an abnormal Internet Protocol and the destination port is not an abnormal port, then it is determined that there is no abnormal outgoing connection.
[0099] It's important to note that the destination IP address of outgoing data packets can be used to determine if it's abnormal. For example, a pre-defined range of normal IP addresses can be used. If the destination IP address is outside this range, it's abnormal; if it is, it's normal. Similarly, a normal port number can be set first, and then compared to the normal port number. If they don't match, it's an abnormal port; if they match, it's normal. In practical applications, if both the destination IP and / or port are determined to be abnormal, it indicates that the disk array system has unauthorized external connections.
[0100] And / or, based on the protocol type and size of the outgoing data packet, combined with the protocol type and size of each other outgoing data packet within a third preset time period preceding the outgoing data packet, determine whether there are data slices with different protocols.
[0101] It should be noted that in the process of determining whether data exists in different protocol data slices based on outgoing data packets, the protocol type and size of the outgoing data packet can be determined based on the content information of the outgoing data packet. Additionally, other outgoing data packets within a third preset time period preceding the outgoing data packet can be obtained, and their protocol types and sizes can be determined. Therefore, based on the protocol type and size of the outgoing data packet, as well as the protocol types and sizes of the other outgoing data packets within the third time period, it can be determined whether data in the disk array system exists in different protocol data slices.
[0102] In one implementation, the process of determining whether there are data slices with different protocols based on the protocol type and size of the outgoing data packet, combined with the protocol type and size of other outgoing data packets within a third preset time period preceding the outgoing data packet, may include the following content.
[0103] If the size of the outgoing data packet is less than the preset data packet size, determine whether the number of other target outgoing data packets with a size less than the preset data packet size and the same protocol type as the outgoing data packet in the third preset time period before the outgoing data packet reaches the preset number. If it does, determine that there are data slices with different protocols.
[0104] Understandably, to more accurately identify whether there are data slices with different protocols in the disk array system, the size of the outgoing data packet can be compared with the preset data packet size. If the size of the outgoing data packet is smaller than the preset data packet size, it means that the outgoing data packet is a small block of data. In order to determine whether there are multiple small outgoing data packets with the same protocol, the protocol type and size of each other outgoing data packet in the third preset time period before the outgoing data packet can be obtained. From these other outgoing data packets, the target other outgoing data packets with the same protocol type as the outgoing data packet and smaller than the preset data packet size can be identified. The number of these target other outgoing data packets can be counted to see if it reaches the preset number. If it reaches the preset number, it means that the same protocol is dividing the data into multiple small data blocks. That is, the same protocol divides a large data packet into multiple small data packets for transmission, forming multiple data slices. Therefore, it can be determined that there are data slices with different protocols in the disk array system.
[0105] And / or, based on the various data transmission and reception instructions of the system within the fourth preset time period, determine the instruction flow relationship, and combine it with the instruction list corresponding to the protocol type of the outgoing data packet to determine whether there is an abnormal disk array instruction flow.
[0106] Understandably, when determining whether an abnormal disk array command flow exists in the disk array system based on outgoing data packets, various data transmission and reception commands generated by the disk array system within a fourth preset time period can be obtained. Based on these data transmission and reception commands and their chronological order, the corresponding command flow relationship can be determined. The protocol type corresponding to the obtained outgoing data packets can be determined, and then a command list corresponding to that protocol type can be determined. A correspondence between protocol types and command lists can be pre-established. After determining the protocol type of the outgoing data packets, the corresponding command list is determined from the correspondence between protocol types and command lists. This command list includes each command in sequential order. Therefore, after obtaining the command flow relationship, it can be further compared with each command in the command list. If the execution relationship between the commands in the command flow relationship and the commands in the command list are different, or if the commands themselves are different, it indicates that an abnormal disk array command flow exists.
[0107] And / or, based on the outgoing data packets, determine the corresponding target system, determine whether there is synchronization traffic between the target system and each disk or storage system of the disk array system, and if not, determine that the data is being stolen by faking a disk failure.
[0108] It should be noted that, in this embodiment, when determining whether data theft by faking a disk failure exists in the disk array system based on the outgoing data packet, the target system to which the outgoing data packet will be sent can be determined based on the content information in the outgoing data packet. Then, it is further identified whether there is synchronization traffic between the target system and each disk or storage system of the disk array system. Since a hot spare disk is activated in the event of a disk failure, and the hot spare disk of the disk array system itself usually synchronizes with the disks or storage systems of the disk array system, there will be synchronization traffic. Therefore, this application can determine whether the target system is faking a disk failure to steal data by identifying whether there is synchronization traffic between the target system and each disk or storage system of the disk array system. If there is no synchronization traffic between the target system and each disk or storage system of the disk array system, it indicates that the target system is faking a disk failure to steal data, thus confirming that data theft by faking a hard drive failure exists in the disk array system. If there is synchronization traffic between the target system and each disk or storage system of the disk array system, it indicates that the target system is performing normal data backup, thus confirming that data theft by faking a hard drive failure has not occurred in the disk array system.
[0109] Please refer to Figure 3In one implementation, after receiving the outgoing data packet, the system can first determine whether there is a sudden increase in data traffic. If a sudden increase in traffic is found, it can be determined that the disk array system is at risk of data leakage, and a data leakage alarm can be issued, and the operation can be terminated. If there is no sudden increase in traffic, the system can further determine whether there are any unconventional external connections. If unconventional external connections are found, it can be determined that the disk array system is at risk of data leakage, and a data leakage alarm can be issued, and the operation can be terminated. If there are no unconventional external connections, the system can further determine whether there are data slices using different protocols. If different protocol data slices are found, it can be determined that the disk array system is at risk of data leakage, and a data leakage alarm can be issued, and the operation can be terminated. If there are no different protocol data slices, the system can further determine whether there are any abnormal disk array command streams. If abnormal disk array command streams are found, it can be determined that the disk array system is at risk of data leakage, and a data leakage alarm can be issued, and the operation can be terminated. If no abnormal disk array command stream is found, the outgoing data packets can be used to further determine if data theft is being carried out by posing as a disk failure. If data theft by posing as a disk failure is found, it indicates a data leakage risk in the system, and a data leakage alarm can be issued, at which point the operation can be terminated. If no data theft by posing as a disk failure is found, the operation can be terminated. The specific implementation process of each detection step can be found in the corresponding content described in the above embodiments, and will not be repeated here.
[0110] Furthermore, this embodiment of the application can further detect whether the disk array system is maliciously damaging system data, thereby improving system security. In practical applications, data packets generated within the disk array system can be acquired, and then it can be further identified whether the data packets are written directly to the data disk without being verified. If the data packets are written directly to the data disk without verification, it is determined that abnormal striping has occurred, indicating that the disk array system is maliciously damaging system data. At this time, an alarm record can be generated and an alarm notification can be issued, so that staff can maintain the data of the disk array system according to the alarm information, thereby improving system security.
[0111] In one implementation, please refer to Figure 4 The method may further include the following steps S410 to S430.
[0112] S410: Periodically retrieve the current audit logs, outgoing data packets, and system context environment status information.
[0113] S420: A pre-trained data leakage detection model is used to perform data leakage analysis on the audit logs, outgoing data packets, and system context environment status information at the current moment, and the analysis results are obtained. The data leakage detection model is pre-trained based on multiple sets of historical audit logs, historical outgoing data packets, historical system context environment status information, and corresponding data leakage situations of multiple disk array systems.
[0114] S430: If the analysis results indicate that the disk array system has a data leakage risk, issue a data leakage alarm.
[0115] It should be noted that, in order to further ensure the accuracy of data leakage detection of the disk array system in this embodiment of the application, a data leakage detection model can be pre-established and started periodically. That is, the audit logs, outgoing data packets, and system context environment status information of the disk array system at the current moment can be obtained periodically, and then the audit logs, outgoing data packets, and system context environment status information at the current moment can be input into the data leakage detection model for analysis to obtain the analysis results. If the analysis results indicate that there is a data leakage risk in the disk array system, a data leakage alarm record can be generated and a data leakage alarm prompt can be issued.
[0116] Understandably, during the training of a data leakage detection model, multiple sets of historical audit logs, historical outgoing data packets, historical system context environment state information, and corresponding data leakage situations from multiple disk array systems can be pre-acquired. The model is then trained based on these multiple sets of data to obtain a well-trained data leakage detection model. The historical system context environment state information includes metadata and physical block access records, and the current system context environment state information also includes metadata and physical block access records. Since metadata and physical block access records are unique to the disk array system—for example, every change in business data will cause changes in metadata; if business data is accessed through abnormal channels, it will cause abnormal changes in metadata, and the same applies to physical block access records—the data leakage detection model trained in this application based on multiple sets of historical audit logs, historical outgoing data packets, historical system context environment state information, and corresponding data leakage situations can accurately detect data leakage behavior, thus improving system security.
[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0118] Please refer to Figure 5Embodiments of this application also provide a disk array data leakage detection device, applied to a disk array system, the device comprising:
[0119] Log analysis module 11 is used to obtain audit log content information, perform data leakage detection on the operation behavior of the audit log content information, and obtain the first detection result; the audit log content information includes at least one audit log, which is generated by each business model in the disk array system when performing operations;
[0120] The first acquisition module 12 is used to acquire outgoing data packets generated by the disk array system when the first detection result indicates that there is no risk of data leakage.
[0121] The packet analysis module 13 is used to perform data leakage detection on outgoing data packets based on data characteristics, and obtain a second detection result;
[0122] The first determining module 14 is used to determine that no data leakage has occurred in the disk array system if the second detection result indicates that there is no risk of data leakage.
[0123] The second determining module 15 is used to determine that a data leak has occurred in the disk array system when the first detection result indicates a risk of data leakage, and / or the second detection result indicates a risk of data leakage.
[0124] In one implementation, the log analysis module 11 includes:
[0125] The first acquisition unit is used to acquire newly added audit logs when at least one business module in the disk array system generates audit logs.
[0126] The first determination unit is used to determine whether the corresponding operation behavior is an attack behavior based on the newly added audit log. If the operation behavior is an attack behavior, the first determination unit is triggered.
[0127] The first determining unit is used to determine that the obtained first detection result indicates a risk of data leakage.
[0128] In one embodiment, the device may further include:
[0129] The second acquisition unit is used to acquire each first audit log within a first preset time period before the newly added audit log if the operation is not an attack.
[0130] The second determining unit is used to determine whether the operation behavior deviates from the corresponding normal behavior baseline based on the newly added audit log and each first audit log;
[0131] The third determining unit is used to determine that if the operation deviates from the corresponding normal behavior baseline, the first detection result obtained is that there is a risk of data leakage.
[0132] The third acquisition unit is used to acquire each second audit log within the second preset time period before the newly added audit log if the operation behavior does not deviate from the corresponding normal behavior baseline.
[0133] The fourth determination unit is used to perform correlation analysis on the newly added audit logs and each of the second audit logs to determine the corresponding behavioral links;
[0134] The fifth determining unit is used to determine that the first detection result obtained is a risk of data leakage when the behavioral link is an attack link;
[0135] The sixth determining unit is used to determine, when the behavior link is a normal link, that the first detection result obtained is that there is no risk of data leakage.
[0136] In one embodiment, the fourth determining unit includes:
[0137] The first determining subunit is used to obtain the behavior links in chronological order based on the operation behavior corresponding to the newly added audit log and the operation behavior corresponding to each of the second audit logs.
[0138] The first matching subunit is used to match the behavior link with each attack behavior link in the pre-established attack behavior link library. If there is an attack behavior link that matches the behavior link, the second determination subunit is triggered.
[0139] The second determining subunit is used to determine whether a behavior link is an attack link; if there is no attack behavior link that matches the behavior link, then the behavior link is determined to be a normal link.
[0140] In one embodiment, the first determination unit includes:
[0141] The third determination subunit is used to determine the corresponding operation behavior based on the operation behavior information in the newly added audit log;
[0142] The second matching subunit is used to match the operation behavior with each attack behavior in the pre-established attack behavior library. If there is an attack behavior in the attack behavior library that matches the operation behavior, the fourth determination subunit is triggered; if there is no attack behavior in the attack behavior library that matches the operation behavior, the fifth determination subunit is triggered.
[0143] The fourth determining subunit is used to determine whether the operation is an attack.
[0144] The fifth determining subunit is used to determine whether the operation is an attack.
[0145] In one implementation, the packet analysis module 13 is used for:
[0146] Based on the outgoing data packets, determine whether the data is subject to at least one of the following: sudden increase in traffic, batch data reading behavior at unusual time periods, unusual external connections, data slicing of different protocols, abnormal disk array command streams, or data theft disguised as disk failure. If at least one of these conditions exists, the second detection result indicates a risk of data leakage; if none of these conditions exist, the second detection result indicates no risk of data leakage.
[0147] In one implementation, the packet analysis module 13 includes:
[0148] The second determination unit is used to determine whether there is a sudden increase in data traffic based on the size of the outgoing data packet and the normal traffic baseline of a single session.
[0149] And / or, the seventh determining unit is used to obtain the data standard corresponding to the time period based on the current time corresponding to the outgoing data packet and the time period corresponding to the current time, and determine whether the outgoing data packet conforms to the data standard. If the outgoing data packet does not conform to the data standard, it is determined that there is a batch data reading behavior of an irregular time period.
[0150] And / or, the third determination unit is used to determine whether there is an unconventional external connection for the data based on the destination Internet Protocol and / or destination port of the outgoing data packet;
[0151] And / or, the fourth determination unit is used to determine whether there are different protocol data slices based on the protocol type and size of the outgoing data packet, combined with the protocol type and size of each other outgoing data packet located in the third preset time period before the outgoing data packet;
[0152] And / or, the fifth determination unit is used to determine the instruction flow relationship based on the various data transmission and reception instructions of the system within the fourth preset time period, and to determine whether there is an abnormal disk array instruction flow in the data by combining the instruction list corresponding to the protocol type of the outgoing data packet;
[0153] And / or, the sixth determination unit is used to determine the corresponding target system based on the outgoing data packet, and to determine whether there is synchronization traffic between the target system and each disk or storage system of the disk array system. If not, the eighth determination unit is triggered.
[0154] The eighth determining unit is used to determine whether data theft is a disguised disk failure.
[0155] In one embodiment, the second determination unit includes:
[0156] The first determination subunit is used to detect whether the size of the outgoing data packet is greater than the normal traffic baseline of a single session;
[0157] The sixth determination subunit is used to determine if there is a sudden increase in data traffic when the size of the outgoing data packet is greater than the normal traffic baseline for a single session.
[0158] In one embodiment, the device may further include:
[0159] The first acquisition subunit is used to acquire each other outgoing data packet within the fifth preset time period before the outgoing data packet when the size of the outgoing data packet is less than or equal to the normal traffic baseline of a single session;
[0160] The second determination subunit is used to determine whether the total outgoing data packet size is greater than the total normal traffic baseline based on the size of the outgoing data packet and the size of each other outgoing data packet within the fifth preset time period before the outgoing data packet.
[0161] The seventh determination subunit is used to determine if there is a sudden increase in data traffic when the total size of outgoing data packets is greater than the total normal traffic baseline.
[0162] In one implementation, the third determination unit is used for:
[0163] The system detects whether the destination Internet Protocol of the outgoing data packet is an abnormal Internet Protocol, and / or whether the destination port of the outgoing data packet is an abnormal port. If the destination Internet Protocol is an abnormal Internet Protocol, and / or the destination port is an abnormal port, then it is determined that there is an abnormal outgoing connection. If the destination Internet Protocol is not an abnormal Internet Protocol and the destination port is not an abnormal port, then it is determined that there is no abnormal outgoing connection.
[0164] In one implementation, the fourth determination unit is used for:
[0165] If the size of the outgoing data packet is less than the preset data packet size, determine whether the number of other target outgoing data packets with a size less than the preset data packet size and the same protocol type as the outgoing data packet in the third preset time period before the outgoing data packet reaches the preset number. If it does, determine that there are data slices with different protocols.
[0166] In one embodiment, the device may further include:
[0167] The second acquisition module is used to periodically acquire the audit logs, outgoing data packets, and system context environment status information at the current moment;
[0168] The analysis module is used to perform data leakage analysis on the audit logs, outgoing data packets, and system context environment status information at the current moment using a pre-trained data leakage detection model, and obtain the analysis results. The data leakage detection model is pre-trained based on multiple sets of historical audit logs, historical outgoing data packets, historical system context environment status information, and corresponding data leakage situations from multiple disk array systems.
[0169] In one embodiment, in practical applications, the disk array data leakage detection device may further include an alarm module, which provides a data leakage alarm when it is determined that a data leakage has occurred in the disk array system.
[0170] It should be noted that, as Figure 6 As shown, in this application, each business processing module generates audit logs. The log analysis module analyzes the audit logs, and the log segmentation module obtains a first detection result after analyzing the audit logs. If the first detection result indicates a data leakage risk, an alarm record message is generated, and the alarm module is triggered to issue a data leakage alarm. If the first detection result indicates no data leakage risk, the data packet analysis module performs data feature detection and analysis on the acquired outgoing data packets to obtain a second detection result. If the second detection result indicates a data leakage risk, an alarm record message is generated, and the alarm module is triggered to issue a data leakage alarm. If the second detection result indicates no data leakage risk, it can be determined that no data leakage has occurred. In addition, to further improve detection accuracy, this application can also set up an intelligent analysis module. The data leakage detection model in the intelligent analysis module analyzes the audit logs, outgoing data packets, and system context environment status information obtained at the current moment to obtain the analysis results. If the analysis results indicate that there is a data leakage risk, an alarm record message can still be generated and the alarm module can be triggered to issue a data leakage alarm. The alarm can be sent to the system administrator through the software interface and various alarm interfaces (such as mobile phone text messages, email, etc.) so that the administrator can perform system maintenance as soon as possible.
[0171] It should be noted that the description of the features of the disk array data leakage detection device provided in this application embodiment can be found in the relevant description of the disk array data leakage detection method embodiment, and will not be repeated here.
[0172] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the disk array data leakage detection method.
[0173] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the disk array data leakage detection method when it runs.
[0174] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0175] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the disk array data leakage detection method.
[0176] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the disk array data leakage detection method embodiments described above.
[0177] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0178] The foregoing has provided a detailed description of a method, product, device, and medium for detecting data leakage in a disk array, as provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for detecting data leakage in a disk array, characterized in that, Applied to a disk array system, the method includes: The audit log content information is obtained, and data leakage detection of the operation behavior is performed on the audit log content information to obtain a first detection result; the audit log content information includes at least one audit log, which is generated by each business model in the disk array system when performing operations. If the first detection result indicates that there is no risk of data leakage, obtain the outgoing data packets generated by the disk array system; Data leakage detection based on data characteristics of the outgoing data packets is performed to obtain a second detection result; If the second detection result indicates that there is no risk of data leakage, it is determined that the disk array system has not experienced data leakage. If the first detection result indicates a risk of data leakage, and / or the second detection result indicates a risk of data leakage, then it is determined that a data leakage has occurred in the disk array system; wherein: The process of obtaining audit log content information, performing data leakage detection on the operation behavior of the audit log content information, and obtaining a first detection result includes: If at least one service module in the disk array system generates an audit log, the newly added audit log is obtained; Based on the newly added audit logs, determine whether the corresponding operation is an attack. If the operation is an attack, then determine that the first detection result is that there is a risk of data leakage. Data leakage detection is performed on the outgoing data packets based on their data characteristics to obtain a second detection result, including: Based on the outgoing data packets, determine whether the data exhibits at least one of the following conditions: sudden increase in traffic, batch data reading behavior during unusual time periods, unusual external connections, data slicing with different protocols, abnormal disk array command streams, or data theft disguised as disk failure. If at least one condition exists, the second detection result indicates a risk of data leakage; if none of these conditions exist, the second detection result indicates no risk of data leakage.
2. The method for detecting data leakage in a disk array according to claim 1, characterized in that, Also includes: If the operation is not an attack, then obtain each first audit log within a first preset time period prior to the newly added audit log; Based on the newly added audit logs and each of the first audit logs, determine whether the operation deviates from the corresponding normal behavior baseline; If the operation deviates from the corresponding normal behavior baseline, the first detection result is determined to indicate a risk of data leakage. If the operation does not deviate from the corresponding normal behavior baseline, then obtain each second audit log within the second preset time period before the newly added audit log; A correlation analysis will be performed on the newly added audit logs and each of the second audit logs to determine the corresponding behavioral links; If the behavioral link is an attack link, the first detection result is determined to be that there is a risk of data leakage. If the behavioral link is a normal link, the first detection result is determined to be that there is no risk of data leakage.
3. The method for detecting data leakage in a disk array according to claim 2, characterized in that, The newly added audit logs and each of the second audit logs will undergo correlation analysis to determine the corresponding behavioral links, including: Based on the operation behavior corresponding to the newly added audit log and the operation behavior corresponding to each of the second audit logs, the behavior links are obtained in chronological order. After determining the corresponding behavior link, the following is also included: The behavioral link is matched with each attack behavior link in a pre-established attack behavior link library. If an attack behavior link that matches the behavioral link exists, the behavioral link is determined to be an attack link; otherwise, the behavioral link is determined to be a normal link.
4. The method for detecting data leakage in a disk array according to claim 1, characterized in that, Based on the newly added audit logs, determine whether the corresponding operation is an attack, including: The corresponding operation behavior is determined based on the operation behavior information in the newly added audit log; The operation is matched with various attack behaviors in a pre-established attack behavior database. If an attack behavior that matches the operation behavior exists in the database, the operation behavior is determined to be an attack behavior; if no attack behavior that matches the operation behavior exists in the database, the operation behavior is determined not to be an attack behavior.
5. The method for detecting data leakage in a disk array according to claim 1, characterized in that, Based on the outgoing data packets, determine whether the data exhibits at least one of the following conditions: sudden increase in traffic, batch data reading behavior at unusual time periods, unusual external connections, data slicing using different protocols, abnormal disk array command streams, or data theft disguised as disk failure. Based on the size of the outgoing data packets and the baseline of normal traffic in a single session, determine whether there is a sudden increase in data traffic; And / or, based on the current time corresponding to the outgoing data packet and the time period corresponding to the current time, obtain the data standard corresponding to the time period, and determine whether the outgoing data packet conforms to the data standard. If the outgoing data packet does not conform to the data standard, determine that there is a batch data reading behavior in an irregular time period. And / or, based on the destination Internet Protocol and / or destination port of the outgoing data packet, determine whether the data has an unconventional outgoing connection; And / or, based on the protocol type and size of the outgoing data packet, and in conjunction with the protocol type and size of each other outgoing data packet located within a third preset time period preceding the outgoing data packet, determine whether there are data slices with different protocols; And / or, based on the various data transmission and reception instructions of the system within the fourth preset time period, determine the instruction flow relationship, and combine it with the instruction list corresponding to the protocol type of the outgoing data packet to determine whether there is an abnormal disk array instruction flow in the data; And / or, based on the outgoing data packet, determine the corresponding target system, determine whether there is synchronization traffic between the target system and each disk or storage system of the disk array system, and if not, determine that the data is being stolen by faking a disk failure.
6. The method for detecting data leakage in a disk array according to claim 5, characterized in that, Based on the size of the outgoing data packets and the baseline of normal traffic in a single session, determine whether there is a sudden increase in data traffic, including: Detect whether the size of the outgoing data packet is greater than the normal traffic baseline for a single session; If the size of the outgoing data packet is greater than the normal traffic baseline for a single session, a sudden increase in data traffic is determined.
7. The method for detecting data leakage in a disk array according to claim 6, characterized in that, Also includes: If the size of the outgoing data packet is less than or equal to the normal traffic baseline of a single session, acquire all other outgoing data packets located within a fifth preset time period before the outgoing data packet; Based on the size of the outgoing data packet and the size of each other outgoing data packet within the fifth preset time period preceding the outgoing data packet, determine whether the total outgoing data packet size is greater than the total normal traffic baseline. If the total size of the outgoing data packets is greater than the total normal traffic baseline, a sudden increase in data traffic is determined.
8. The method for detecting data leakage in a disk array according to claim 5, characterized in that, Based on the destination Internet Protocol and / or destination port of the outgoing data packet, determine whether the data has any unconventional outgoing connections, including: The system detects whether the destination Internet Protocol of the outgoing data packet is an abnormal Internet Protocol, and / or whether the destination port of the outgoing data packet is an abnormal port. If the destination Internet Protocol is an abnormal Internet Protocol, and / or the destination port is an abnormal port, then it is determined that the data has an abnormal outgoing connection. If the destination Internet Protocol is not an abnormal Internet Protocol and the destination port is not an abnormal port, then it is determined that the data does not have a very high outgoing connection.
9. The method for detecting data leakage in a disk array according to claim 5, characterized in that, Based on the protocol type and size of the outgoing data packet, and in conjunction with the protocol type and size of other outgoing data packets within a third preset time period preceding the outgoing data packet, it is determined whether data slices with different protocols exist, including: If the size of the outgoing data packet is less than the preset data packet size, determine whether the number of other target outgoing data packets with a size smaller than the preset data packet size and the same protocol type as the outgoing data packet in each other outgoing data packet within a third preset time period before the outgoing data packet reaches a preset number. If it does, then determine that there are data slices with different protocols.
10. The method for detecting data leakage in a disk array according to any one of claims 1 to 9, characterized in that, Also includes: Periodically retrieve the current audit logs, outgoing data packets, and system context environment status information; A pre-trained data leakage detection model is used to perform data leakage analysis on the audit logs, outgoing data packets, and system context environment status information at the current moment, and the analysis results are obtained. The data leakage detection model is pre-trained based on multiple sets of historical audit logs, historical outgoing data packets, historical system context environment status information, and corresponding data leakage situations of multiple disk array systems.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the disk array data leakage detection method as described in any one of claims 1 to 10.
12. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the disk array data leakage detection method as described in any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the disk array data leakage detection method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Data leakage detection method and device, electronic equipment and readable storage medium
CN114640530A
Data leakage risk assessment system
CN120217434A