Distributed directory controller and method for performing data leakage protection by using distributed directory

By combining a distributed directory controller and an AI-powered directory, the efficiency and accuracy issues of data leak detection in data center systems are resolved, achieving highly efficient protection against data leaks and enhancing data security and the protection of mobile office environments.

CN120937301APending Publication Date: 2025-11-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380096015.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-05-15
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently, accurately, and quickly detecting data breaches in data center systems, leading to data contamination and loss, especially in mobile office and cloud systems where there are detection delays and data breach risks.

Method used

Employing a distributed directory controller, it leverages artificial intelligence directories and machine learning to detect data breaches. By monitoring input and output operations, it identifies unexpected patterns, generates alerts, and takes protective measures, including generating lists of suspicious objects and risk level reports, to prevent data breaches and malware attacks.

Benefits of technology

It improves data security in data center systems, reduces fake threats, prevents data leaks and malware contamination in a timely manner, and protects the Bring Your Own Device environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937301A_ABST
    Figure CN120937301A_ABST
Patent Text Reader

Abstract

A controller for operating in a data center system comprising one or more data nodes. The controller is further configured to: receive a malware indication that malware is running on one of the one or more data nodes; determining data blocks with risks; determining an attacker of the data block with the risk; determining zero or more other data blocks with risks by determining other data blocks accessed by the attacker; an alert is generated indicating the data block at risk, the other data blocks at risk, and the attacker. In addition, the controller is a distributed directory controller. Thus, the controller is used to provide data protection efficiently and reliably, e.g., by using an artificial intelligence (AI) directory to detect any potential sensitive data leakage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates generally to the field of data security, and more specifically to controllers and methods operating in a data center system comprising one or more data nodes. Background Technology

[0002] Typically, different organizations need to maintain data related to multiple topics, customers, and prospects. However, continuous protection of this data is necessary to prevent any loss, for example, through data leakage prevention (DLP) techniques. DLP refers to the processes used to detect and prevent data breaches, leaks, and any accidental damage to data. Therefore, different organizations use DLP to comply with data privacy regulations and protect relevant data (e.g., personally identifiable information, intellectual property information, etc.) to prevent any data loss or leakage. Furthermore, different organizations use DLP to achieve data visibility and provide a secure environment, for example, by protecting mobile workers, protecting bring-your-own-device (BYOD) environments, and protecting cloud systems.

[0003] Data breaches are generally categorized into three types: insider threats, external breaches by attackers, and unintentional (or negligent) data disclosure. These are all causes of data breaches. Insider threat data breaches correspond to attacks that allow access to and transmission of personally identifiable information (PII) outside organizational boundaries by disclosing any privileged user account. External breaches by attackers correspond to attacks targeting sensitive data (e.g., PII), such as gaining access through phishing, malware, or code injection. Unintentional (or negligent) data disclosures correspond to data breaches that occur due to a lack of restrictions within different organizations. Many attempts have been made to prevent such data breaches, such as protecting data in motion, protecting endpoints, protecting data at rest, protecting data in use, data identification, and data breach detection. However, these attempts focus on preventing data breaches, such as backing up data through firewalls and antivirus software, which can lead to data contamination, which is undesirable. In some scenarios, traditional methods for detecting data breaches may result in detection delays, increasing the potential damage caused by the breach. Therefore, there is a technical challenge of how to efficiently, accurately, and quickly detect data breaches to prevent data contamination.

[0004] Therefore, based on the above discussion, there is a need to overcome the disadvantages associated with conventional controllers. Summary of the Invention

[0005] This invention provides a controller and method for operation in a data center system comprising one or more data nodes. The invention provides a solution to the existing problem of how to efficiently, accurately, and quickly detect data breaches to prevent data contamination. The object of this invention is to provide a solution that at least partially overcomes the problems encountered in the prior art, and to provide an improved controller and improved method for operation in a data center system comprising one or more data nodes to provide data breach protection through the use of an artificial intelligence (AI) catalog.

[0006] One or more objects of the present invention are achieved by means of the technical solutions provided in the appended independent claims. Advantageous embodiments of the invention are further defined in the dependent claims.

[0007] In one aspect, the present invention provides a controller for operation in a data center system comprising one or more data nodes. Furthermore, the controller is configured to: receive an indication that malware is running on one of the one or more data nodes; identify a risky data block; identify an attacker associated with the risky data block; identify zero or more other risky data blocks by determining other data blocks accessed by the attacker; and generate an alert indicating the risky data block, the other risky data blocks, and the attacker. Furthermore, the controller is a distributed directory controller.

[0008] The controller is used to efficiently and reliably provide data protection by detecting and protecting against any sensitive data (e.g., personally identifiable information, PII) leaks using an artificial intelligence (AI) catalog. Furthermore, the controller is used to track, monitor, and collect all input-output (I / O) operations (e.g., write and read I / O operations), for example, by using kernel drivers or by using user-space drivers to detect data leaks. Subsequently, the controller is used to detect unexpected patterns and further generate a list of sensitive objects accessed by an attacker to indicate any potential threats, which are then used to handle any potential data leaks. Additionally, the controller is used to reduce any false threats by comparing actual patterns with expected patterns. Furthermore, the controller is used to generate a report including a list of suspicious objects tagged with sensitive data, such as personally identifiable information (PII). Information (PII) and any critical data related to intellectual property can be further used to send alerts to users. Additionally, the controller uses an AI catalog to determine the presence of all suspicious files in the storage system (e.g., using advanced views) and then generates a report. The generated report includes a suspicion score and corresponding risk level for each file (or object). The risk level depends on the data type, such as the sensitivity of the data stored in the file and the probability of a data breach. Therefore, the controller is used to prevent data center systems from being contaminated by any potential data breaches or malware, and to protect Bring Your Own Device (BYOD) environments by efficiently and reliably detecting data breaches by reducing false detections.

[0009] In one implementation, the controller is further configured to receive the malware indication from the device. Furthermore, the malware indication includes an indication of the type of malware.

[0010] In another implementation, the controller is also configured to receive a malware indication from the device, wherein the type of malware is ransomware.

[0011] In this implementation, the controller's instructions to the ransomware are used to take necessary steps to prevent such malware attacks and improve the overall data security of the data center system.

[0012] In another implementation, the device is a monitoring device for monitoring the one or more data nodes.

[0013] The device monitors the one or more data nodes to detect malware attacks on the one or more data nodes, for example by identifying any deviations from the expected I / O pattern, and further instructs the controller on information about the corresponding malware (e.g., ransomware).

[0014] In another implementation, the controller is further configured to receive the expected I / O pattern by determining the expected I / O pattern based on previously received monitoring information. Furthermore, the deviation includes at least one indication of a deviation I / O operation, which differs from the I / O operation in the expected pattern, including an indication of a data block at risk and an indication of an entity accessing the data block at risk.

[0015] In this implementation, deviations in I / O patterns are identified to determine potential data breaches and data security threats, enabling data center systems to detect any potential data breaches and data security threats that should be mitigated before any data loss occurs.

[0016] In another implementation, the controller is also used to utilize machine learning to determine the expected I / O pattern based on previously received monitoring information.

[0017] By leveraging machine learning, the controller predicts anticipated I / O patterns, enabling it to prevent corresponding I / O operations that might include malware. This improves the overall data security of the data center system.

[0018] In another implementation, the controller is also configured to identify the attacker as an entity that accesses risky data blocks in the monitoring information.

[0019] In this implementation, the controller is also configured to identify an attacker as an entity in the monitoring information that accesses the risky data block, in order to prevent the attacker from causing any damage to the data.

[0020] In another implementation, the controller is also used to assign the priority label based on the type of data in the risky data block.

[0021] Advantageously, the assignment of priority labels allows the controller to prioritize responses and protective measures that need to be taken to protect the data accordingly. This improves the overall data security of the data center system.

[0022] In another aspect, the present invention provides a method for a data center system comprising one or more data nodes. Furthermore, the method includes: receiving an indication that malware is running on one of the one or more data nodes; identifying a risky data block; identifying an attacker of the risky data block; identifying zero or more other risky data blocks by determining other data blocks accessed by the attacker; and generating an alert indicating the risky data block, the other risky data blocks, and the attacker, wherein the method is performed by a controller.

[0023] The method described herein achieves all the advantages and technical effects of the controller of the present invention.

[0024] It should be understood that all of the above implementation methods can be combined together.

[0025] It should be noted that all devices, elements, circuits, units, and apparatuses described in this application can be implemented in software elements or hardware elements, or any combination thereof. All steps performed by the various entities described in this application, and the functions to be performed by the various entities described, are intended to indicate that the respective entities are suitable for or used to perform the respective steps and functions. Even in the description of the following specific embodiments, if a particular function or step to be performed by an external entity is not reflected in the description of the specific detailed elements of the entity performing that particular step or function, it will be clear to those skilled in the art that these methods and functions can be implemented in the corresponding software or hardware elements, or in any combination of such elements. It is understood that the features of the present invention are readily combined in various combinations without departing from the scope of the invention as defined by the appended claims.

[0026] Other aspects, advantages, features, and objects of the invention will become apparent from the accompanying drawings and the detailed description of illustrative implementations as explained in conjunction with the following appended claims. Attached Figure Description

[0027] The foregoing description of the invention and the following detailed description of illustrative embodiments can be better understood when read in conjunction with the accompanying drawings. Exemplary structures of the invention are shown in the drawings to illustrate the invention. However, the invention is not limited to the specific methods and means disclosed herein. Furthermore, those skilled in the art will understand that these drawings are not drawn to scale. Where possible, the same elements are represented by the same reference numerals.

[0028] The embodiments of the present invention will now be described by way of example only, in conjunction with the following accompanying drawings, in which:

[0029] Figure 1 A block diagram describing a controller operating in a data center system, provided for embodiments of the present invention;

[0030] Figure 2 A flowchart of a method for a data center system provided in an embodiment of the present invention; and

[0031] Figure 3 An exemplary diagram of a system architecture for malware indication provided in an embodiment of the present invention.

[0032] In the accompanying drawings, underlined reference numerals indicate the item to which the underlined reference numeral is located or the item adjacent to the underlined reference numeral. Ununderlined reference numerals are associated with the item identified by the line linking the ununderlined reference numeral to that item. When a reference numeral is ununderlined but has an associated arrow, the ununderlined reference numeral identifies the general item pointed to by the arrow. Detailed Implementation

[0033] The following detailed description illustrates embodiments of the present invention and ways in which these embodiments can be implemented. While some embodiments of the invention have been disclosed, those skilled in the art will recognize that other embodiments for carrying out or practicing the invention can also be implemented.

[0034] The following detailed description illustrates embodiments of the present invention and ways in which these embodiments can be implemented. While some embodiments of the invention have been disclosed, those skilled in the art will recognize that other embodiments for carrying out or practicing the invention can also be implemented.

[0035] Figure 1 A block diagram illustrating a controller operating in a data center system, provided for embodiments of the present invention. (See reference...) Figure 1 The diagram shows a block diagram 100 describing a controller 104 for operation in a data center system 102, the data center system 102 including one or more data nodes 106, such as a first data node 106A, a second data node 106B to an nth data node 106N.

[0036] Controller 104 is used to operate in data center system 102. Examples of controller 104 may include, but are not limited to: microcontroller, microprocessor, central processing unit (CPU), complex instruction set computing (CISC) processor, application-specific integrated circuit (ASIC) processor, reduced instruction set (RISC) processor, very long instruction word (VLIW) processor, data processing unit, and other processors or control circuits.

[0037] In operation, controller 104 is used to receive indications that malware is running on one of one or more data nodes 106. In one example, controller 104 is used to receive indications that the malware is running on the first data node 106A. In another example, controller 104 is used to receive indications that the malware is running on a second data node 106B. Similarly, controller 104 is used to receive indications that the malware is running on the nth data node 106N. In one implementation, controller 104 is used to receive malware indications from device 108. In another implementation, controller 104 is used to receive malware indications from any external system. Furthermore, the malware indication includes an indication of the malware type. The malware indication corresponds to an alert received by controller 104 that enables controller 104 to take further appropriate actions to mitigate the risk of malware attacks and prevent any data breaches. In this implementation, the type of malware is ransomware. Ransomware corresponds to, for example, malware that prevents users from accessing data by encrypting it. Therefore, the controller 104's indication of ransomware is used to take necessary steps to prevent such attacks and improve the overall data security of the data center system 102. In one implementation, the device 108 is a monitoring device for monitoring the one or more data nodes 106. The monitoring of the one or more data nodes 106 by the device 108 is used to detect malware attacks in the one or more data nodes 106 and further indicate the corresponding malware, such as ransomware, to the controller 104. In this implementation, the device 108 is one of the one or more data nodes 106. For example, a first data node 106A among the one or more data nodes is used to transmit malware indications to the controller 104 of the data center system 102. Similarly, a second data node 106B among the one or more data nodes 106 is used to transmit malware indications to the controller 104 of the data center system 102. In another implementation, the monitoring device may correspond to a data node from another data center system without affecting, for example, Figure 3The scope of the invention is shown and described in detail below. Therefore, the execution of malicious software can be prevented, thereby improving the overall data security of the data center system 102. According to another embodiment, the controller 104 is also configured to receive the monitoring information from the monitored node. For example, the controller 104 is configured to receive monitoring information from a first data node 106A. Similarly, the controller 104 is configured to receive monitoring information from a second data node 106B without affecting the scope of the invention. Advantageously, receiving monitoring information from the monitored node is used to detect any potential threats, for example, by identifying any deviations in the monitoring information within the monitored node. Therefore, the overall data security of the data center system 102 is improved.

[0038] According to one embodiment, controller 104 is configured to receive malware indications by receiving monitoring information from monitored nodes in one or more data nodes 106. First, device 108 is configured to monitor one or more data nodes 106 of data center system 102. Subsequently, device 108 is configured to send monitoring information of the monitored nodes to controller 104. For example, device 108 (or first data node 106A) is configured to transmit malware indications to controller 104 by monitoring a second data node 106B among one or more data nodes 106. Furthermore, the monitoring information includes indications of input / output (I / O) operations of the monitored nodes, the expected I / O pattern of the monitored nodes, and deviations in the monitoring information compared to the expected I / O pattern. As a result, by using indications of I / O operations, expected I / O patterns, and deviations in the monitoring information to mitigate potential threats, these potential threats can be prevented to improve the overall data security of data center system 102. According to one embodiment, the monitoring information includes metadata of the I / O operations but does not include content data of the I / O operations. In one example, the metadata for I / O operations is updated periodically based on indications of the I / O operations of the monitored nodes. Furthermore, the metadata for I / O operations is used to detect malware, for example, by identifying any deviations in I / O patterns. Therefore, potential malware is detected efficiently and accurately before execution, improving the overall data security of data center system 102.

[0039] According to one embodiment, controller 104 is further configured to have a global scope of the data center system by receiving monitoring information from one or more data nodes 106. Advantageously, the global scope of data center system 102 enables controller 104 to efficiently and accurately compare actual I / O patterns with expected I / O patterns. Therefore, false malware indications can be detected efficiently and accurately. In one implementation, controller 104 is further configured to receive the expected I / O pattern by determining the expected I / O pattern based on previously received monitoring information. The deviation includes at least one indication of a deviation I / O operation that differs from the I / O operation in the expected pattern, the deviation I / O operation including an indication of a data block at risk and an indication of an entity accessing the data block at risk. The I / O operation in the expected pattern refers to the expected behavior of the I / O operation expected based on the monitoring information of controller 104 of data center system 102. Therefore, controller 104 is used to identify potential data breaches and data security threats, thereby enabling data center system 102 to take appropriate actions necessary to prevent or mitigate such data breaches and data security threats. In another implementation, controller 104 is further configured to utilize machine learning to determine the expected I / O pattern based on previously received monitoring information. In other words, controller 104 receives monitoring information from device 108, monitoring devices, or monitored nodes. Subsequently, controller 104 determines the expected I / O pattern based on the corresponding monitoring information, for example, by utilizing machine learning. For instance, a machine learning model can be used to train controller 104 to determine the expected I / O pattern. Therefore, by utilizing machine learning, controller 104 predicts the expected I / O pattern, enabling controller 104 to prevent the execution of corresponding I / O operations that may include malware. This improves the overall data security of data center system 102. In yet another implementation, controller 104 is further configured to determine deviations based on monitoring information according to one or more of the following: resource monitoring of data nodes in data center system 102, changes in resource consumption over time, changes in data blocks on compressed or uncompressed volumes, detected randomization patterns, volume changes, size changes, data block change rates, read counts within a window time frame, changes in block segmentation dispersion, and / or monitoring information history. In one example, controller 104 is used to determine deviations based on resource monitoring of data nodes in data center system 102. In another example, controller 104 is used to determine the deviations based on the monitoring information based on changes in resource consumption over time. In yet another example, controller 104 is used to determine the deviations based on the monitoring information based on changes in data blocks on compressed or uncompressed volumes.In another example, controller 104 is used to determine the deviation based on detected randomization patterns, volume changes, size changes, or data block change rates, without affecting the scope of the invention. Similarly, in another example, controller 104 is used to determine the deviation based on the number of reads in a window time frame. In yet another example, controller 104 is used to determine the deviation based on changes in block segment dispersion rate and / or historical monitoring information.

[0040] Furthermore, controller 104 is used to determine deviations based on resource monitoring and changes in resource consumption over time of data nodes in data center system 102. Similarly, in another example, controller 104 is used to determine deviations based on changes in data blocks on compressed or uncompressed volumes, as well as detected randomization patterns, volume changes, size changes, and data block change rates, without affecting the scope of the invention. In yet another example, controller 104 is used to determine deviations based on the number of reads in a window time frame, changes in block segment dispersion, and / or the history of monitoring information. Therefore, the determination of monitoring information is used to protect data, such as personally identifiable information (PII) of data center system 102.

[0041] In one implementation, an artificial intelligence (AI) catalog service is used to determine biases based on monitoring information, which are then used to further identify at-risk data blocks and attackers. In another implementation, the AI ​​catalog service refers to an unstructured data management service that provides a single, centralized viewpoint across the entire enterprise storage system. The AI ​​catalog service runs different types of queries, performs analyses, and provides insights into the customer's storage enterprise. Furthermore, the AI ​​catalog service is used to periodically (e.g., through collectors) collect information on unstructured data from various sources across the entire enterprise storage system, including NAS, S3, and VMs. Additionally, the AI ​​catalog service collects data for resource monitoring of data center system 102 and one or more data nodes 106 of data center system 102. Furthermore, the AI ​​catalog service detects changes in resource consumption over time, changes in data block segmentation on compressed and uncompressed volumes, detects randomization patterns, dispersion in volume changes, size changes, incremental file system scans, and the status of another file such as size, last access, calculates the rate of change, and calculates cross-file I / O temperature by calculating the number of reads within a calculation window time frame. This improves the overall data security of data center system 102.

[0042] Controller 104 is also configured to identify data blocks at risk. In other words, controller 104 is configured to identify data blocks at risk to protect the corresponding data blocks from malware attacks. In one implementation, controller 104 is configured to identify data blocks at risk by retrieving them from malware indications. Malware indications include indications of data blocks at risk. In one example, controller 104 is configured to identify data blocks at risk by malware indications obtained through receiving monitoring information about the corresponding data blocks at risk. In another implementation, controller 104 is also configured to identify data blocks at risk by retrieving them from the deviation operation. Identifying data blocks at risk by malware indications is used to identify data blocks at risk, for example, by displaying data blocks that deviate from the expected pattern. Therefore, the impact of malware attacks can be mitigated. According to one embodiment, controller 104 is also configured to identify the file to which the data block at risk belongs as a file at risk. For example, controller 104 is configured to identify files at risk based on malware indications or by deviation operations. Therefore, the overall data security of the data center system can be improved, for example, by reducing the risk of malware attacks.

[0043] Furthermore, controller 104 is used to identify attackers who may be accessing risky data blocks. The attackers who may be accessing risky data blocks are identified by controller 104 to identify and prevent any potential data breaches. In one implementation, controller 104 is also used to identify the attacker as an entity accessing the risky data blocks in monitoring information. In other words, controller 104 is used to receive monitoring information from one or more data nodes 106, for example, via device 108 or a monitored node, to identify the attacker as an entity accessing the risky data blocks. This prevents the attacker from causing any further damage to the data. In another implementation, controller 104 is also used to identify the attacker by retrieving the attacker from a malware indication, wherein the malware indication includes instructions for the attacker. In one example, device 108 or a monitored node is used to transmit monitoring information indicating the attacker. Therefore, controller 104 is used to prevent the attacker from causing any damage to the data center system 102.

[0044] Controller 104 is also configured to determine zero or more other data blocks at risk by identifying other data blocks accessed by the attacker. The determination of data blocks accessed by the attacker is used to identify other data blocks at risk to protect them from any potential attacks. According to one embodiment, controller 104 is further configured to determine zero or more other data blocks at risk by generating a list of data blocks and assigning priority labels to the data blocks in the list. First, controller 104 is configured to identify the attacker accessing the data blocks at risk. Subsequently, controller 104 is configured to identify data blocks accessed by the attacker, in addition to the identified data blocks at risk. For example, the first data block and the second data block are identified by controller 104 as data blocks accessed by the attacker. Furthermore, controller 104 is configured to generate a list of data blocks accessed by the attacker. Finally, controller 104 is configured to assign priority labels to the data blocks in the generated list. In one implementation, controller 104 is also configured to assign priority labels based on the type of malware. For example, controller 104 is configured to assign priority labels based on data blocks affected by ransomware. Similarly, controller 104 is configured to assign priority labels based on the type of malware. In another implementation, controller 104 is also configured to assign the priority label based on the data type of the data block at risk. For example, if a data block contains sensitive information, then controller 104 assigns a higher priority label to the corresponding data block compared to data blocks with less sensitive information. Therefore, the assignment of priority labels allows controller 104 to prioritize responses and protective measures required to protect the data accordingly. In yet another implementation, controller 104 is also configured to assign priority labels based on the attacker. Assignment of attacker-based priority labels serves to notify controller 104 of any potential threats, enabling controller 104 to take necessary actions to mitigate risk and protect data center system 102 from any harm.

[0045] Controller 104 is also used to generate alerts indicating at-risk data blocks, other at-risk data blocks, and attackers. In one example, controller 104 generates alerts indicating at-risk data blocks, other at-risk data blocks, and attackers. The alerts generated by controller 104 are used to protect data center system 102 from potential threats, malware, and data security breaches. Furthermore, controller 104 is used to take certain necessary actions to protect itself from such potential threats, malware attacks, and data breaches. Additionally, controller 104 is a distributed directory controller. The distributed directory controller includes the original master distributed directory (i.e., the AI ​​directory service) and maintains one or more local directories (e.g., the AI ​​directory), which are created when a user accesses the master distributed directory. Furthermore, the distributed directory controller maintains a certain level of control over the AI ​​directory, enabling it to create local copies of the AI ​​directory. In one implementation, controller 104 may be deployed inside the AI ​​directory. In another implementation, controller 104 may be deployed outside the AI ​​directory as an additional service. In this implementation, controller 104 communicates with the AI ​​directory using an application programming interface (API). Therefore, controller 104, acting as a distributed directory controller, provides data leakage protection through the AI ​​directory.

[0046] Controller 104 is used to provide efficient and reliable data protection, for example, by detecting any potential sensitive data breaches using an artificial intelligence (AI) catalog. Furthermore, controller 104 is used to track, monitor, and collect all input-output (IO) operations (e.g., write and read IO operations), for example, by using kernel drivers or by using user-space drivers to detect data breaches. Subsequently, controller 104 is used to detect unexpected patterns and further generate a list of sensitive objects accessed by an attacker to indicate any potential threats, which are then used to address any potential data breaches. Additionally, controller 104 is used to reduce any false threats by comparing actual patterns with expected patterns. Furthermore, controller 104 is used to generate a report including a list of suspicious objects, which includes tags for sensitive data such as personally identifiable information (PII) and any critical data related to intellectual property, which can be further used to send alerts to users. Controller 104 uses the AI ​​catalog to determine the presence of all suspicious files in the storage system (e.g., using an advanced view) and then generates a report. Furthermore, the generated report includes a suspicion score and corresponding risk level for each file (or object). The risk level depends on the data type, such as the sensitivity of the data stored in files and the probability of data leakage. Therefore, controller 104 is used to prevent data center system 102 from being contaminated by any potential data leakage or malware, and to further protect the Bring Your Own Device (BYOD) environment by reducing false detections and thus detecting data leakage more efficiently and reliably.

[0047] Figure 2 A flowchart illustrating a method for use in a data center system according to an embodiment of the present invention. (See also...) Figure 2 A flowchart of a method 200 for use in a data center system 102 is shown, the data center system 102 including ( Figure 1 (One or more data nodes 106, for example, a first data node 106A, a second data node 106B, and an Nth data node 106N.) Method 200 includes steps 202 to 210.

[0048] In operation, for example, in step 202, method 200 includes receiving an indication that malware is running on one of one or more data nodes 106. In one example, method 200 includes receiving the indication that the malware is running on the first data node 106A. In another example, method 200 includes receiving an indication that the malware is running on a second data node 106B. Similarly, method 200 includes receiving an indication that the malware is running on the nth data node 106N. In one implementation, method 200 includes receiving the malware indication via device 108. In another implementation, method 200 includes receiving the malware indication via a monitored node. Therefore, malware can be blocked before it is executed. Furthermore, in step 204, method 200 includes identifying data blocks at risk. In other words, method 200 includes identifying data blocks at risk to protect the corresponding data blocks from malware attacks. Furthermore, in step 206, method 200 includes identifying the attacker of the data blocks at risk. The attacker of the at-risk data blocks is identified by controller 104 to prevent the attacker from committing any potential data breach. In step 208, method 200 includes identifying zero or more other at-risk data blocks by determining other data blocks accessed by the attacker. The identification of data blocks accessed by the attacker is used to identify other at-risk data blocks to protect these data blocks from any potential attacks. In step 210, method 200 includes generating alerts indicating the at-risk data blocks, other at-risk data blocks, and the attacker. Furthermore, method 200 is performed by controller 104. The alerts generated by controller 104 are used to protect data center system 102 from potential threats, malware, and data security breaches. In addition, controller 104 is used to take certain necessary actions to protect data center system 102 from such potential threats, malware attacks, and data breaches. Furthermore, method 200 is performed by a distributed directory controller. The distributed directory controller includes the original main distributed directory (i.e., AI directory service) for maintaining one or more local directories (e.g., AI directories) that are created when a user accesses the main distributed directory. Furthermore, the distributed directory controller maintains a certain level of control over the AI ​​directory, enabling it to create local copies of the AI ​​directory. Therefore, controller 104, acting as the distributed directory controller, provides data leakage protection through the use of the AI ​​directory.

[0049] Method 200 is used to provide efficient and reliable data protection, for example, by detecting any potential sensitive data breaches using an artificial intelligence (AI) catalog. Furthermore, Method 200 is used to track, monitor, and collect all input-output (IO) operations (e.g., write and read IO operations), for example, by using kernel drivers or by using user-space drivers to detect data breaches. Subsequently, Method 200 is used to detect unexpected patterns and further generate a list of sensitive objects accessed by an attacker to indicate any potential threats, which are further used to address any potential data breaches. Additionally, Method 200 is used to reduce any false threats by comparing actual patterns with expected patterns. Furthermore, Method 200 is used to generate a report including a list of suspicious objects, which includes tags on sensitive data, such as personally identifiable information (PII) and any critical data related to intellectual property, which can be further used to send alerts to users. Therefore, the method 200 is used to prevent the data center system 102 from being contaminated by any potential data leaks or malware, and to further detect data leaks efficiently and reliably by reducing false detections, thereby protecting the bring your own device (BYOD) environment.

[0050] Steps 202 to 210 are merely illustrative, and other alternatives may be provided, in which one or more steps are added, one or more steps are deleted, or one or more steps are provided in a different order, without departing from the scope of the claims herein.

[0051] A computer-readable medium including computer program instructions is provided, which, when executed by one or more processors, cause the distributed directory controller system to perform method 200. In one example, the instructions may be implemented on a computer-readable medium, including but not limited to electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital (SD) cards, solid-state drives (SSDs), computer-readable storage media, and / or CPU cache. In one example, the instructions are generated by a computer program that implements method 200 and is used to implement method 200 in the distributed directory controller system.

[0052] Figure 3 This is an exemplary diagram of a system architecture for malware indication provided in an embodiment of the present invention. (In conjunction with...) Figure 1 Component description Figure 3 . refer to Figure 3 The illustration describes the existence of the first data center system 302A to the nth data center system 302N (e.g., ( Figure 1 Figure 300 illustrates the process of a data center system 102.

[0053] In one implementation scenario, the system architecture includes a first data center system 302A through an nth data center system 302N. Each data center system, for example, the first data center system 302A through the nth data center system 302N, includes multiple collector components. For example, the first data center system 302A includes one or more data nodes, such as a first data node (i.e., NAS) 304, a second data node 306 (i.e., S3), etc. Similarly, the nth data center system 302N includes one or more data nodes, such as a first data node 308, a production ESX server 310, a production MSSQL server 312, a production Oracle server 314, etc. Furthermore, one or more data nodes are used to track the I / O patterns and report any deviations to an Artificial Intelligence (AI) catalog service 316. In addition, the AI ​​catalog service 316 is used to track and record I / O operations (i.e., write and read operations), access the I / O patterns and trends of each data node in one or more data nodes 106, for example, the data nodes include the first data node (i.e., NAS) 304, the second data node 306 (i.e., S3) in the first data center system 302A, and the first data node 308, production ESX server 310, production MSSQL 312, and production Oracle 314 in the nth data center system 302N.

[0054] For example, the AI ​​catalog service 316 is used to collect corresponding data through the internal collector 320 to record I / O behavior, such as the dispersion rate and I / O pattern of each data block. However, if the AI ​​catalog service 316 detects any deviation from the expected I / O pattern, then in this case, the AI ​​catalog service 316 is used to add the corresponding I / O pattern along with all attackers to a suspicious list, so as to trigger an alert, for example, through the AI ​​service engine 318. For example, the controller 104 is used to receive indications that ransomware is running on one of one or more data nodes 106. In addition, the controller 104 is used to identify risky data blocks, attackers, and other data blocks to further generate alerts indicating said data blocks, other risky data blocks, and attackers. As a result, each data center system in the first data center system 302A to the nth data center system 302N is used to provide data protection efficiently and reliably, for example, by using an artificial intelligence (AI) catalog to detect any potential sensitive data leaks. This improves the overall data security of the corresponding data center system.

[0055] Modifications to the embodiments of the invention described above may be made without departing from the scope of the invention as defined by the appended claims. Expressions such as “comprising,” “including,” “integrating,” “having,” and “are” used to describe and claim the invention should be interpreted in a non-exclusive manner, meaning that items, parts, or elements not explicitly described are also permitted. Singular references should also be interpreted in relation to the plural. The word “exemplary” as used herein means “as an example, instance, or illustration.” Any embodiment described as “exemplary” is not necessarily construed as preferential or superior to other embodiments and / or excluding features in conjunction with other embodiments. The word “optionally” as used herein means “provided in some embodiments and not in others.” It should be understood that certain features of the invention described in the context of a single embodiment for clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for clarity may also be provided individually or in any suitable combination or, where appropriate, as embodiments of any other described aspect of the invention.

Claims

1. A controller (104), characterized in that, For operation in a data center system (102) including one or more data nodes (106), the controller (104) is further configured to: Receive an indication that malware is running on one of the one or more data nodes (106). Identify data blocks that are at risk. The attacker who identified the at-risk data block By identifying other data blocks accessed by the attacker, zero or more other data blocks that are at risk are identified, and The controller (104) generates alerts indicating the data block at risk, other data blocks at risk, and the attacker, wherein the controller (104) is a distributed directory controller.

2. The controller (104) according to claim 1, characterized in that, The controller (104) is also configured to receive the malware indication from the device (108), wherein the malware indication includes an indication of the type of malware.

3. The controller (104) according to claim 2, characterized in that, The controller (104) is also configured to receive the malware instruction from the device (108), wherein the type of malware is ransomware.

4. The controller (104) according to claim 2, characterized in that, The device (108) is a monitoring device (108) used to monitor the one or more data nodes (106).

5. The controller (104) according to claim 2, characterized in that, The device (108) is one of the data nodes (106).

6. The controller (104) according to any one of claims 2 to 5, characterized in that, The controller (104) is also configured to determine the attacker by retrieving the attacker from the malware indication, wherein the malware indication includes an indication to the attacker.

7. The controller (104) according to any one of claims 2 to 5, characterized in that, The controller (104) is further configured to determine the risky data block by retrieving the risky data block from the malware indication, wherein the malware indication includes an indication of the risky data block.

8. The controller (104) according to any one of the preceding claims, characterized in that, The controller (104) is also configured to receive the malware indication by receiving monitoring information of the monitored node in one or more data nodes (106), receiving the expected I / O mode of the monitored node, and determining the deviation of the monitoring information from the expected I / O mode, wherein the monitoring information includes indications of input / output I / O operations of the monitored node.

9. The controller (104) according to claim 8, characterized in that, The controller (104) is further configured to receive the expected I / O pattern by determining the expected I / O pattern based on previously received monitoring information, wherein the deviation includes at least one indication of a deviation I / O operation that is different from the I / O operation in the expected pattern, the deviation I / O operation including an indication of the data block at risk and an indication of the entity accessing the data block at risk.

10. The controller (104) according to claim 9, characterized in that, The controller (104) is also used to determine the expected I / O pattern based on previously received monitoring information using machine learning.

11. The controller (104) according to any one of claims 8 to 10, characterized in that, The controller (104) is also configured to determine the deviation based on the monitoring information according to one or more of the following: Resource monitoring of the data node (106) in the data center. Resource consumption changes over time Changes to data blocks on compressed or uncompressed volumes Detected randomization patterns, Volume changes, Size change, Rate of change of data blocks Number of reads within the window time frame Changes in block segment dispersion rate, and / or Historical monitoring information.

12. The controller (104) according to any one of claims 8 to 11, characterized in that, The monitoring information includes metadata of the I / O operation, but does not include content data of the I / O operation.

13. The controller (104) according to any one of claims 8 to 12, characterized in that, The controller (104) is also used to receive the monitoring information from the monitored node.

14. The controller (104) according to any one of claims 8 to 13, characterized in that, The controller (104) is also configured to identify the attacker as an entity in the monitoring information that accesses the risky data block.

15. The controller (104) according to any one of claims 8 to 14, characterized in that, The controller (104) is also configured to determine the risky data block by retrieving the risky data block from the deviation operation.

16. The controller (104) according to any one of the preceding claims, characterized in that, The controller (104) is also configured to identify the file to which the risky data block belongs as a risky file.

17. The controller (104) according to any one of the preceding claims, characterized in that, The controller (104) is also configured to have global scope of the data center system (102) by receiving monitoring information from the one or more data nodes (106).

18. The controller (104) according to any one of the preceding claims, characterized in that, The controller (104) is also configured to identify the zero or more other data blocks at risk by generating a list of data blocks, and to assign priority tags to the data blocks in the list.

19. The controller (104) according to claim 18, characterized in that, The controller (104) is also used to assign the priority label based on the type of the malware.

20. The controller (104) according to claim 18 or 19, characterized in that, The controller (104) is also used to assign the priority label based on the type of data in the data block that is at risk.

21. The controller (104) according to claim 18 or 20, characterized in that, The controller (104) is also used to assign the priority label based on the attacker.

22. A method (200), characterized in that, For a data center system (102) including one or more data nodes (106), the method (200) includes: Receive an indication that malware is running on one of the one or more data nodes (106). Identify data blocks that are at risk. The attacker who identified the at-risk data block By identifying other data blocks accessed by the attacker, zero or more other data blocks that are at risk are identified, and The method (200) generates alerts indicating the data block at risk, other data blocks at risk, and the attacker, wherein the method is executed by a distributed directory controller.

23. A computer program product, characterized in that, Includes program instructions that, when executed by one or more processors in a distributed directory controller system, perform the method (200) according to claim 22.