Statistical ranking of data repositories with policy compliance risk in unstructured data management

The system addresses the challenge of managing unstructured data compliance by using AI and ML for statistical ranking and automated classification, ensuring efficient and compliant data management across hybrid cloud environments.

WO2026064766A1PCT designated stage Publication Date: 2026-03-26DATA DYNAMICS INC +5
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

The challenge of managing unstructured data across various repositories while ensuring compliance with organizational, industry, and regulatory policies is complex, with existing solutions lacking comprehensive risk scoring and statistical analysis, leading to inefficiencies and potential violations of data sovereignty and security.

Method used

A system that provides statistical ranking of data repositories based on policy compliance risk, utilizing AI and ML techniques for accurate and actionable compliance risk assessment, including metadata analysis, content scanning, and automated data classification to prioritize remediation efforts.

Benefits of technology

Enables efficient and compliant data management by providing granular insights and automated workflows, ensuring data security, sovereignty, and regulatory compliance across hybrid cloud environments, reducing human error and optimizing storage utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025047493_26032026_PF_FP_ABST
    Figure US2025047493_26032026_PF_FP_ABST
Patent Text Reader

Abstract

In some embodiments, a method comprises reading a metadata policy. The method can comprise identifying files stored in a data store. The files can have metadata matching a criterion of a metadata policy. The method can comprise associating a risk type for each file of the plurality of files. The method can comprise reading a statistical sampling policy, the statistical sampling policy providing a threshold for files of the risk type. The method can comprise selecting a subset of files of the plurality of files. The subset of files selected by sampling the files based on the risk type of each file according to the statistical sampling policy. The method can comprise associating a first mitigation action with each file of the subsets of file based on a mitigation policy. The method can comprise performing the first mitigation action for each file of the subset of files.
Need to check novelty before this filing date? Find Prior Art

Description

DDT-00225STATISTICAL RANKING OF DATA REPOSITORIES WITH POLICY COMPLIANCE RISK IN UNSTRUCTURED DATA MANAGEMENTBACKGROUND

[0001] Today, digital trust has a larger impact than physical trust. According to recent surveys, 76% of consumers desire greater control over their data, 86% of individuals believe humans should be ultimately accountable for artificial intelligence (Al) decisions, employees with data control report increased trust in their organization's data practices, and 63% of business leaders believe data democratization is crucial for fostering a data-driven culture. For an enterprise to become a trusted custodian of data requires digital trust. Digital trust can be described as one or more of an enterprise’s commitment to data privacy, ethical Al, data sovereignty and security, and compliance with social and environmental regulations.BRIEF SUMMARY

[0002] According to embodiments of the present disclosure, methods of and computer program products for statistical ranking of data repositories with policy compliance risk in unstructured data management are provided.

[0003] In some embodiments, a method comprises reading a metadata policy. The method can comprise identifying files stored in a data store. The files can have metadata matching a criterion of a metadata policy. The method can comprise associating a risk type for each file of the plurality of files. The method can comprise reading a statistical sampling policy, the statistical sampling policy providing a threshold for files of the risk type. The method can comprise selecting a subset of files of the plurality of files. The subset of files selected by sampling the files based on the risk type of each file according to the statistical sampling policy. The method can comprise associating a first mitigation action with each file of the subsets of fileDDT-00225 based on a mitigation policy. The method can comprise performing the first mitigation action for each file of the subset of files.

[0004] In some embodiments, the method can comprise reading a content policy. The content policy can provide a rule for acceptable content for the data store. The method can comprise assigning a second mitigation action for each of the subset of files based on the content policy. The method can comprise performing the second mitigation action for each file of the subset of files.

[0005] In some embodiments, the metadata can include native file metadata, injected policy metadata, metadata from a catalog, or custom data from a tag repository.

[0006] In some embodiments, the metadata policy associates a risk type with the criterion.

[0007] In some embodiments, the metadata policy includes a plurality of criteria, the metadata policy further association each of the criteria with a risk type.

[0008] In some embodiments, determining the first subset by sampling from the plurality of files according to the statistical sampling police further generates a sample of files from the plurality of files based on behavior analysis and user input.

[0009] In some embodiments, the metadata can include file type, access control, file size, user permission, tags, location, access history, unusual access, or unusual access anomaly. In some embodiments, assigning the risk value is based on determining a correlation or a pattern based on the metadata.

[0010] In some embodiments, the first mitigation action is part of a workflow. Selecting the first mitigation can be determined based on the first subset determined by sampling the plurality of files.DDT-00225

[0011] In some embodiments, the statistical sampling policy includes one or more of a random sampling policy, a stratified sampling policy, and a weighted sampling policy.

[0012] In some embodiments, the random sampling policy includes a metadata parameter and a target percentage value that metadata parameter.

[0013] In some embodiments, the weighted sampling policy a metadata parameter, a priority of that parameter, and a weight of that parameter.

[0014] In some embodiments, the stratified sampling policy includes a metadata parameter.

[0015] In some embodiments, a system comprising a computing node comprising a computer readable storage medium having program instructions embodied therewith. The program instructions can be executable by a processor of the computing node to cause the processor to perform any of the above methods.

[0016] A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing node to cause the computing node to perform the method of any of the above methods.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0017] Fig. 1 is a flowchart illustrating a method for enterprise data management according to embodiments of the present disclosure.

[0018] Figs. 2A-C are user interfaces of an enterprise data management system according to embodiments of the present disclosure.

[0019] Fig. 3 is a flowchart illustrating a method according to embodiments of the present disclosure.DDT-00225

[0020] Fig. 4 is a flowchart illustrating a risk type identification service reading stored metadata and applying a risk analysis according to embodiments of the present disclosure.

[0021] Fig. 5 is a flowchart illustrating generate sampling result based on a sampling policy according to embodiments of the present disclosure.

[0022] Fig. 6 is a diagram of a computing node according to an embodiment of the present disclosure.

[0023] Fig. 7 is a flowchart illustrating a process of statistical sampling according to embodiments of the present disclosure.DETAILED DESCRIPTION

[0024] Decentralized identities are on the rise, with over 100 million users projected by 2025. • By 2025, 75% of enterprises will need to implement Explainable Al (XAI) solutions for user trust. By 2025, 15% of large enterprises will begin transitioning to quantum-resistant cryptography to safeguard their data. By 2030, 73% of consumers will prioritize ethical data practices.

[0025] The growing need for Al, automation, and personalization clashes with a rising tide of citizens who demand control over their personal information. This clash creates a strategic dilemma for businesses. The lack of visibility caused by siloed data systems and fragmented point solutions creates a distrust between central IT and business units (BUs), which can also spill over to consumers. Consumers can feel a disconnect between what companies promise, what they do with data, and what users actually see.

[0026] Fig. 1 is a flowchart 100 illustrating a method for enterprise data management according to embodiments of the present disclosure. First, a data workflow made of data stores 104, data processing 106 and data action 108 is created (102).DDT-00225

[0027] In some embodiments, data can be scanned for metadata by a policy or storage administrator (120). The metadata can be provided to a dashboard updated with data usage analysis (e.g., for a Chief Infrastructure Officer, Chief Data Officer, or Data Owner) (126). The method can then optimize storage (e.g., for the data custodian and data owner) (128). The method can then generate an action plan based on corporate policies (e.g., for the data custodian and data owner) (130). The method can then execute the action (e.g., archive, transform, delete, migrate) via a workflow (e.g., for the data custodian and data owner) (132). Then, the method can update the dashboard for the various personas with a reduced risk and optimized storage insights (136).

[0028] In some embodiments, the metadata is provided to a content scan performed by a policy administrator, data custodian, or data owner (122). The method can then generate a dashboard updated with risk exposure insights for a chief information security officer and a data owner (110). The method can then mitigate risk based on those insights (e.g., by the data custodian, data owner) (112). The method can further generate an action plan based on corporate policies to mitigate those risks (e.g., by the data custodian, data owner) (114). The method can then execute the action (e.g., manage permissions, archive, quarantine, delete) via a workflow (124). Then, the method can update the dashboard for the various personas with a reduced risk and optimized storage insights (136).

[0029] Figs. 2A-C are user interfaces of an enterprise data management system according to embodiments of the present disclosure. In Fig. 2A, the user interface illustrates a dashboard having an overall risk of the entire enterprise across all locations according to embodiments of the present disclosure. In addition, four gauges in the dashboard can display metrics of access control, data redundancy, data sensitivity, and retention compliance. An additional gauge canDDT-00225 display a list of top sensitive data, organized by department. Yet another gauge can display statuses (e.g., pending, overdue, in progress, failed, completed) of risk mitigation tasks.

[0030] In Fig. 2B, the user interface illustrates a dashboard having an overall storage efficiency metric according to embodiments of the present disclosure. In addition, five gauges can illustrate data redundancy, cold data, orphaned data, expired data, and junk data. Another gauge can display redundant, obsolete, or trivial (ROT) data by department of the enterprise. Other gauges can display on-premise storage capacity by data store, cloud storage usage by service provider, file distribution by type, file distribution by age, file count trend, and a sustainability metric (e.g., carbon emissions).

[0031] In Fig. 2C, the user interface illustrates a dashboard having an overall risk score of the entire enterprise across all locations according to embodiments of the present disclosure. In addition, five gauges in the dashboard can display metrics of access control, data redundancy, data sensitivity, retention compliance, and encryption strength. Yet another gauge can display statuses (e.g., pending, overdue, in progress, failed, completed) of risk mitigation tasks. Other gauges can display sustainability metrics (e.g., by carbon emissions), top sensitive data by department, and file distribution by file type or age and by department.

[0032] It can be recognized that the gauges of Figs. 2A-C can be combined in any manner, and that the data displayed is merely exemplary and non-limiting. For example, drop down menus for each gauge can change the grouping or ranking of the data. Gauges can be moved, removed, or added to a given dashboard according to a user’s preference.

[0033] In some embodiments, systems for enterprise data management are provided. Such systems can guide businesses in orchestrating data democratization and ethical data usage within an Al-driven, climate- focused world.DDT-00225

[0034] In some embodiments, the systems can provide unparalleled understanding, correlation, consistency and actionability for stakeholders across an enterprise or organization. The systems can apply a nimble and scalable architecture and employ generative Al for modelling and insights and visibility across the organization via role-based access. In addition, the system can address risk, privacy, sovereignty and compliance. At the same time, the system can ensure optimized and sustainable use of storage infrastructure.

[0035] In some embodiments, the system enables people having roles ranging from C-suite executives, business, technology managers, and data owners to all utilize the same data repository. Those people can then make and execute upon informed and intelligent decisions. As a result, data owners can become trusted data champions, and the enterprise or organization meets its responsibility as a trusted custodian.

[0036] In some embodiments, the system can provide data orchestration and trust in four categories. Those categories can include data security and ethical usages, data transformation and sustainability, data governance and compliance, and data democratization and Al empowerment. Each category can be implemented with various aspects, including the aspects detailed below.

[0037] Data Security and Ethical Usages• Privacy Protection for Ethical Al• Risk Analysis & Remediation for Sensitive Information• Data Access Management• Data Security Orchestration, Automation & Actionability

[0038] Data Transformation & Sustainability• Hybrid-Cloud Data MobilityDDT-00225• Data Aggregation and Harmonization• Cloud-Optimized Data Management• Data Lifecycle Management and Footprint Reduction

[0039] Data Governance & Compliance• AI / ML-based Data Observability and Root Cause Analysis(RCA)• RBAC Driven Process & Controls• Data Usage & Traceability• Policy-based Data Governance Framework & Insight

[0040] Data Democratization & Al Empowerment• Data Owner Observability, Control and Actionability• Self-Service Analytics and Insights• Data Privacy by Design• Data Wrangling and Curation for AlStorage Optimization and Lifecycle Management (LCM) across Hybrid Cloud Infrastructure

[0041] In some embodiments, a system can provide storage optimization and lifecycle management (LCM) across hybrid cloud infrastructure. The ever-expanding spread of data across hybrid cloud environments can present a challenge: storage optimization and LCM. Unmanaged data sprawl can lead to wasted resources, skyrocketing storage costs, and compliance nightmares. In some embodiments, the system can empowering organizations to streamline storage, automate data governance, and gain valuable insights.

[0042] In some embodiments, the system can include an Al-powered self-service data management software. The software can empowers enterprises by enabling users across all levels, from C-suite to data owners, to discover, define, act, transform, and audit data through aDDT-00225 user-friendly interface. The system provides correlation, consistency, and standardization across enterprises by delivering granular insights, deriving recommended workflows, and automating actions using personalized policies and role-based access control (RBAC)-driven processes. This transformation fosters a culture of data ownership, where everyone becomes a data champion, and the organization fulfills its responsibility as a data custodian. The software can overcome these challenges through six key capabilities, strategically aligned with industry trends, to help enterprises navigate the complexities of hybrid cloud storage management. It empowers the enterprise to harness the power of the enterprise’s data while keeping storage costs and security risks under control.Data Discovery and Classification

[0043] In some embodiments, the system can provide automated data classification and policy- driven storage tiering. In the method illustrated by Fig. 1, data can be classified by the metadata scan 120 and identified by the content scan 122. Shadow IT, the unauthorized use of cloud services or applications, can lead to sensitive data being stored and accessed outside of an enterprise’s control, jeopardizing data sovereignty. Additionally, unstructured data that can reside in disparate locations across the global footprint of the enterprise can remain unidentified and unclassified, making it vulnerable to unauthorized access. Unclassified data can make identifying and store critical information inefficient. Manual data classification can be timeconsuming and error-prone. Additionally, enterprises can lack a unified storage strategy across on-premises and cloud environments, leading to underutilized resources and wasted storage costs.

[0044] In some embodiments, the system can provide metadata and content analytics that work together to automatically classify data based on type, sensitivity, and how often the data isDDT-00225 accessed. The provided metadata and content analytics can further allow intelligent storage management using data tiering and placement features. For example, frequently accessed data can be automatically identified and placed on high-performance storage for optimal retrieval speed. Meanwhile, less used data can be intelligently archived to cost-effective cloud storage tiers like S3 or OneDrive. This automated approach can optimize storage utilization, reduce costs associated with underutilized resources, and strengthen data governance by ensuring sensitive data resides on secure storage tiers.

[0045] In some embodiments, the system can combine Metadata Analytics and Natural Language Processor (NLP)-powered Content Analytics for data discovery and classification. This combination can automatically identify / discover and classify sensitive data (e.g., PII, PHI, etc.) across a data landscape. As a result, the system can provide data exposure risks before those risks become a data leak, etc. Data owners can take control with a self-service Data Classification, allowing them to further categorize sensitive data based on their specific regulations. Discovery and risk management can be seamlessly integrated to Cloud Integrated Risk Controls (e.g., Microsoft Information Protection). Additionally, the system’s Data Usage & Traceability can provide a centralized record of all data movement activities, making audit trails and compliance reporting a breeze. In addition, the system can ensure data privacy and regulatory adherence by automating data lifecycle management tasks. Based on retention policies, the system can automatically delete or anonymize data.

[0046] For example, pharmaceutical companies can face challenges complying with GDPR regulations regarding patient data privacy in clinical trials. In some embodiments, the system can automatically identify sensitive patient information within research databases. Researchers can then leverage self-service tools to further categorize this data based on specific trialDDT-00225 protocols. This granular control can maintain patient privacy while allowing for essential research to continue. Additionally, the system can provide a centralized record of all data access and movement, simplifying audit trails and compliance reporting. This can empower the company to mitigate data privacy risks and ensure regulatory compliance

[0047] Metadata Analytics automatically scans an enterprise’s network to identify data across repositories across all locations, including file shares, cloud storage buckets, and data lakes. Content Analytics can then analyze the content of the discovered data to pinpoint sensitive information like Personally Identifying Information (PII), Protected Health Information (PHI), and intellectual property (IP). After these steps, an inventory of data assets can be generated. Data Classification enables categorization of data based on sensitivity level (e.g., high-risk PII / PHI and business sensitive data like passwords or network addresses like IP or MAC addresses), regulatory requirements (e.g., HIPAA, GDPR, etc.), and specific business needs. The system enables user definition of custom classification rules, enabling tagging of all data according to the enterprise’s standards. Therefore, the system facilitates automatic risk prioritization, and allows enterprises to focus remediation efforts first on critical areas.

[0048] For example, financial institutions like banks can struggle with managing a mix of critical and less-used data. In some embodiments, the system can automatically classify regulatory documents, loan applications, and emails, placing sensitive data on secure high-performance storage while archiving less frequently accessed info to cost-effective cloud tiers. This optimizes storage utilization and reduces costs.

[0049] As an example, suppose a global bank suspects shadow IT is storing customer data (e.g., PII such as social security numbers, account numbers, account details, etc.) in unauthorized cloud locations (e.g., scattered across emails, loan applications, and internal documents). InDDT-00225 some embodiments, the system can scan the network to identify all data repositories, and use Content Analytics to find PII such as names, social security numbers, and account information. This helps the bank remediate shadow IT practices, ensuring customer data remains controlled and compliant with data sovereignty regulations.

[0050] As yet another example, financial institutions can struggle with the volume customer data (e.g., loan applications, account statements, transaction records). Manually sorting through this data is a time-consuming and expensive task. However, the system’s Al-powered analytics automatically classify the data, saving the bank time and money. Less-frequently accessed data can be archived to cost-effective storage tiers, freeing up valuable space for real-time transaction data. This allows the bank to focus on what matters most: delivering exceptional customer service.

[0051] Preparing data for Al and machine learning projects can be time-consuming and complex. Data scientists can struggle to find, access, and clean data from diverse sources across the organization. In some embodiments, the system’s Data Aggregation and Harmonization can consolidate data from file storage, object storage, and data lakes, eliminating time-consuming manual collection. Data owners and data scientists can work together seamlessly using self- service tools to define granular access permissions, keeping data secure and compliant to regulations and other rules or policies. The system can initiate data migration and even transform it (e.g., file to object) for an even smoother workflow. Automated data cleansing workflows ensure pristine data quality by removing inconsistencies and errors that could reduce the accuracy of AI / ML models. This translates to streamlined data preparation, faster AI / ML project development, and unlocking the true potential of data for advanced analytics.DDT-00225Data Redundancy Elimination and Archival

[0052] In some embodiments, the system can provide Data Redundancy Elimination and Archival for Improved Efficiency. Data sprawl can lead to redundant, obsolete, and trivial (ROT) data consuming valuable storage space. Manual identification and deletion of redundant data is cumbersome and error-prone. Traditional archiving processes can be complex and hinder data accessibility. In the method illustrated by Fig. 1, data redundancy elimination and archival can be performed by the actions executed by the workflow in steps 124 and 132, including archival, deletion, migration, quarantining, etc.

[0053] In some embodiments, the system can find hidden data and help manage that data efficiently. With Data Redundancy, an enterprise can identify and eliminate duplicate data through intelligent classification and tagging. This can free valuable storage space for active data. In addition, cleaning up such data does not remove access to it. The system can intelligently converts redundant files into space-saving object formats e.g., for archiving). These archived objects can still be easily retrieved using software, a RESTful API (e.g., for audits and legal discovery). Plus, the system provides S3 / NFS / SMB-compliant access, ensuring compatibility with existing infrastructure and making data retrieval a breeze.

[0054] For example, a pharmaceutical enterprise can have redundant clinical trial data, which wastes storage space for active research. The above system can identify and eliminate duplicate data sets based on intelligent tags, freeing up valuable storage space for ongoing research projects and improving data accessibility for scientists.DDT-00225Automated Data Lifecycle Management

[0055] In some embodiments, the system can provide Automated Data Lifecycle Management with Retention Compliance. Manually managing data retention policies across hybrid cloud environments can be complex and error-prone. Inconsistent data handling practices and manual enforcement of data policies can lead to security vulnerabilities and compliance gaps. Non- compliance with data retention regulations can sometimes lead to hefty fines. Therefore, risk mitigation and regulatory compliance can be desired.

[0056] In some embodiments, the system provides Retention Compliance that seamlessly integrates the metadata analysis with data policy creation and workflow automation. In relation to the method illustrated by Fig. 1, workflows can be automated by the creation of a data workflow by step 102 and the executing action via workflow 124 or 132. This powerful combination can allow enterprises to discover, classify, and tag data based on specific retention requirements. The software can automate data lifecycle management tasks based on pre-defined policies. The system can handle deletion, archiving, or migration, thereby ensuring effortless compliance with data retention regulations. Therefore, the system minimizes risks associated with non-compliance and streamlines storage management by eliminating the need for manual tasks.

[0057] For example, hospitals can face challenges ensuring compliance with complex healthcare data privacy regulations. In some embodiments, the system’s Retention Compliance utilizes metadata analysis to classify patient medical records and automate data lifecycle management based on retention requirements. This can ensure HIPAA compliance, minimizes regulatory risks, and streamlines storage management.DDT-00225

[0058] Without a clear data governance framework, data transfers across borders can be haphazard, potentially violating data sovereignty regulations. Inconsistent data handling practices can also lead to confusion and difficulty in demonstrating compliance.

[0059] In some embodiments, the system seamlessly integrates with an existing data governance framework to establish data policies that uphold data sovereignty best practices. The system’s Data Policy Creation and Workflow functionality simplifies defining clear rules for data movement and usage. These policies can be tailored based on data sensitivity, location, and retention requirements for different data types. Furthermore, the system gives data owners control over their data through Self-Service Data Classifications and Migrations. This feature streamlines data management by allowing data owners to classify and migrate data according to established policies. The system can then automates the data migration process, ensuring data reaches the appropriate storage locations while adhering to data sovereignty regulations. This not only simplifies governance but also reduces the risk of errors.

[0060] For example, a multinational corporation with worldwide offices collects customer data from various regions. In some embodiments, the system can define clear policies for data movement and usage that comply with data sovereignty regulations in each country they operate. The system’s Self-Service Data Classifications and Migrations can empower business units to classify and migrate their data according to these established policies. Therefore, consistent data handling practices compliance with data sovereignty regulations across the global organization are ensured.

[0061] As another example, hospitals are under constant pressure to comply with strict regulations regarding patient data retention. Manual processes for managing patient records can be error-prone and lead to hefty fines. In some embodiments, the system can automate dataDDT-00225 retention tasks, ensuring hospitals stay compliant and avoid costly penalties. Additionally, the system can provide complete transparency into how patient data is used, fostering trust with patients and regulators. This allows hospitals to focus on their core mission: delivering quality patient care.

[0062] In some embodiments, the system can ensure consistent data handling across an enterprise with the system’s Data Policy Creation and Workflow. The Data police can clearly define comprehensive data policies that translate into automated workflows. These workflows can enforce access controls, data encryption, and other security measures, minimize human error, and guarantee consistent policy application. These workflows work with the results of Data Discovery and Classification. By classifying data based on sensitivity and regulations, the system can use data policies and workflows tailored to the specific needs of the data as it is classified. This can streamline data governance, simplify compliance, and strengthen data security.

[0063] For example, consider an insurance company that implements new data governance policies to comply with industry regulations regarding data privacy and security. The system’s Data Policy Creation and Workflow functionality can empower them to personalize and translate these policies into automated workflows. These workflows can include automatic data discovery and classification for sensitive research data, access management, migration, retention compliance, archival, deletion, and more. Automating these processes can minimize human error and ensures consistent enforcement of data governance policies across the organization.Self-service data classification

[0064] In some embodiments, the system provides Self-Service Data Classification and Migration for Democratized Data Management. Complex data management processes andDDT-00225 limited visibility into data storage locations can hinder business users from effectively managing their data. This can lead to data silos and hinder collaboration efforts. Alternative data governance and migration point solutions, often complex and inaccessible, can create data silos, hinder collaboration, and limit the value businesses can extract from their information.Enterprises can move forward from a “lift and shift” approach and embrace a “data-driven” strategy for infrastructure management. In some embodiments, in the method illustrated by Fig. 1, data can be classified by an enterprise or a user providing rules to the metadata scan 120 or content scan 122.

[0065] In some embodiments, the system’s Data Dynamics' Al-powered self-service data management software can bring a fresh approach to privacy, security, compliance, governance and optimization in the world of Al-led workloads.

[0066] In some embodiments, the system can provide business users the ability to take control of their data with Self-Service Data Classifications and Migrations. This intuitive functionality can allow these business users to classify data based on their department's specific needs and business context. User-friendly tools can enable business users to initiate data migrations across on-premises and cloud storage effortlessly, fostering better collaboration and data governance. Improved data visibility through self-service tools can empower users to take ownership of their data and manage it effectively. Use of these self-service tools can dismantle data silos and can facilitate seamless collaboration across departments, ensuring everyone has access to the information they need.

[0067] For example, insurance companies can have trouble collaborating with data silos. In some embodiments, the system’s Self-Service Data Classifications can empower underwriters and claims adjusters to classify data by insurance product or policyholder. They can initiate dataDDT-00225 migrations between on-premises and cloud storage, fostering collaboration across departments and improving data governance.

[0068] This solution brief explores sixkeyuse casesthathighlighthow Zubin addresses critical industry challenges, such as data silos and inefficient data wrangling. We will delve into the software’s role in empowering organizations to automate data classification, streamline migrations, and break down datasilos. By enabling self-servicedataownershipandfostering collaboration, Zubin paves the way for a future where data is readily available for analysis, leading to better decision-making and improved businessoutcomes.Risk Insights

[0069] In some embodiments, the system can provide data usage and traceability with risk insights for enhanced security. Static data security measures may not be sufficient to address evolving threats and vulnerabilities. Lack of visibility into data usage across hybrid cloud environments can make it difficult to identify and mitigate security risks. Unclear data ownership can lead to confusion about access control responsibilities. Varying data privacy regulations across regions can sometimes require data localization policies. Without proper risk assessment and data localization strategies, an enterprise may be inadvertently violating data sovereignty laws. In some embodiments, in the method illustrated by Fig. 1, the dashboard 110 illustrates risk insights based on the analysis described herein.

[0070] Tracking and maintaining a clear audit trail for data movement across borders is essential for demonstrating compliance with data sovereignty regulations. Without this visibility, regulators may question data handling practices and ability to ensure data remains within desired jurisdictions.DDT-00225

[0071] In some embodiments, the system can help locate data and help secure it. The system’s Data Usage and Traceability feature can provide a central hub or centralized data index for all data movement activities, giving complete visibility of an enterprise’s data. This central hub can be implemented with a dashboard to show data owners how their data is being used. It meticulously tracks every transfer, capturing the source, destination, purpose, and timestamp, along with the user or system responsible. This comprehensive audit trail enables clear and responsible data movement practices. It simplifies compliance with data residency requirements by providing a clear picture of where data resides and how the data is being used. This not only fosters trust but also ensures the enterprise have the information needed to manage the enterprise’s data effectively.

[0072] With Risk Exposure Insights, a powerful analytics tool, an enterprise can identify potential security threats by analyzing data usage patterns. Risk Exposure Insights can identify exposed data and prioritize threats to address first. Risk Exposure Insights can leverage data classification and advanced content analytics powered by Al and Machine Learning (AI / ML) to assess the true risk level of exposed data. Beyond basic risk scoring, Risk Exposure Insights can consider factors such as the type of data exposed and the likelihood of exploitation based on realtime threat intelligence. This allows enterprises to focus remediation efforts on the most critical threats, ensuring the enterprise addresses sensitive data at a highest risk of misuse first.

[0073] For example, consider an energy company that suspects a cyberattack on its industrial control systems (ICS) managing power grids, potentially compromising sensor data, configuration files, and employee credentials. In some embodiments, the system can analyze the data access patterns and can classify those patterns based on sensitivity assigning the highest risk to configuration files due to the potential for widespread power outages. This allows theDDT-00225 enterprise to prioritize securing the ICS and critical infrastructure, then address employee credentials and sensor data, minimizing power grid disruption.

[0074] Further strengthening an enterprise’s defenses, RBAC Down to the Data Owner Layer provides many levels of access control. This system allows for granular permission management based on who owns the data, ensuring only authorized users can access it. RBAC can access data at the data owner layer. Therefore, the system seamlessly integrates with existing access control features to provide unparalleled granular control. As a result, an enterprise can define which users can access specific data and at what level they can access it (e.g., read, write, edit, read and write, etc.). The system can enforce the principle of least privilege, ensuring users only have the access permissions necessary for their tasks. This significantly reduces the points of attack for a breach, minimizing the potential damage caused by compromised credentials or malicious insiders.

[0075] As an example, consider a pharmaceutical company that develops a new drug. Sensitive research data (e.g., formulas, clinical trial results, etc.) can be shared with collaborators while maintaining strict access controls. Pharmaceutical companies also generate vast amounts of data during drug development, but much of that data is hidden and unused. The system’s RBAC Down to the Data Owner Layer can empower researchers to define granular access permissions for collaborators, ensuring only authorized personnel can access specific data elements. This reduces the attack surface and safeguards valuable intellectual property. The system can further shed light on this “dark data” (e.g. , the hidden or unused data) by providing a centralized view of all research data. Data owners, like research scientists, can see exactly who has accessed their data and for what purpose. This transparency fosters trust and accountability within the researchDDT-00225 team. Additionally, the system can empower researchers and scientists to identify potentially valuable data that can be used for further analysis, accelerating drug discovery and development.

[0076] As another example, data governance for insurance companies has been a complex task accessible only to IT specialists. This often led to data silos and hindered collaboration between departments. In some embodiments, the system empowers non-IT users, such as underwriters and claims adjusters, to take charge of their data. User-friendly tools allow users to classify data based on risk factors and claim types, fostering a culture of data ownership. Additionally, granular access controls ensure data security and compliance. This breakdown of data silos encourages collaboration and empowers business users to make faster, more informed decisions.

[0077] This comprehensive approach creates a multi-layered shield for the enterprise’s data. The enterprise can have clear visibility into movement, the ability to identify potential risks through usage analysis, and ironclad access controls thanks to granular ownership permissions. This can significantly minimize points of attack and prevent data breaches, thereby keeping the enterprise’s information safe.

[0078] In some embodiments, the system employs content analytics to classify data and assess the risk associated with where the data is stored and processed. This analysis considers factors including data sensitivity, applicable regulations, and the security strength of storage locations. In some embodiments, the system allows prioritization of risks, developing targeted mitigation strategies, and defining policies to automatically route specific data types (e.g., PII) to designated storage locations within specific regions. This ensures compliance with data residency requirements. In addition, in some embodiments, the system can integrate with leading cloud storage providers like Microsoft Azure to effortlessly enforce data residency policies.DDT-00225

[0079] As another example, consider an insurance company that operates multiple European countries. The European Union’s (EU) General Data Protection Regulation (GDPR) mandates that citizen data be stored within the EU. Manual processes for isolating sensitive data can be slow and inefficient, which can potentially allowing breaches. In some embodiments, the system can automatically discover and classify customer data (e.g., health information) spread across various locations and assess risks associated with it. Data Containment and Isolation can then allow the enterprise to define policies to automatically route sensitive customer data (e.g., health records) to designated storage locations within the EU, ensuring compliance with GDPR and data sovereignty regulations. This feature can automate workflows to quarantine high-risk sensitive data the moment it is detected. By restricting access and preventing dissemination, Data Containment minimizes potential damage from exposed data. Furthermore, the Data Encryption adds another layer of defense by safeguarding sensitive data at rest and in transit. This ensures data remains protected throughout the data’s lifecycle.

[0080] As yet another example, consider a hospital chain with facilities in the US and Canada that needs to comply with HIPAA regulations, which requires strict controls on patient data movement, or other regulations such as the California Consumer Privacy Act (CCPA). The system’s Data Usage & Traceability provides a comprehensive audit trail for all patient data transfers, including the source, destination, and time of each transfer. This allows the hospital chain to show to regulators that patient data is only transferred across borders for legitimate medical purposes and in accordance with HIPAA guidelines.

[0081] For example, unauthorized access to energy data pipelines can be disastrous. In some embodiments, the system’s Data Usage & Traceability provides a centralized record of data movement activities. Combined with Risk Exposure Insights, it can analyze data usage patternsDDT-00225 to identify suspicious access attempts. Additionally, RBAC Down to the Data Owner Layer can strengthen access controls, preventing unauthorized access and minimizing the attack surface for data breaches.

[0082] As yet another example, consider a disgruntled employee with access to a hospital’s radiology department. That disgruntled employee, in this example, attempts to steal patient data. In some embodiments, the system’s Data Classification can be configured to identify and tag patient medical images. Data Containment and Isolation can then automatically quarantine these files, preventing any such rogue access (e.g., by the disgruntled employee) and minimizing potential patient privacy violations. The hospital can then investigate the attempted incident, revoke the employee’s access, and restore the quarantined data from secure backups.Data Observability and Root Cause Analysis

[0083] In some embodiments, the system can provide Data Observability & Root Cause Analysis for Performance Optimization. Performance bottlenecks and data quality issues can hinder critical applications and analytics across hybrid cloud environments. Identifying the root cause of these issues can be time-consuming and complex. In addition, monitoring data usage and detecting anomalies across a global data footprint can be complex. Without this data observability, an enterprise may not be alerted to unauthorized access or potential data breaches that could compromise data sovereignty. In some embodiments, in the method illustrated by Fig. 1, the dashboard 126 provides data usage analysis that can be used for performance optimization 128.

[0084] The system’s Data Observability and Root Cause Analysis functionality can utilize metadata & content analytics powered by AI / ML to continuously monitor data usage within pipelines across the enterprise’s global footprint. This includes analyzing ROT, user accessDDT-00225 patterns, data transfer activities, and changes to data permissions. By identifying the anomalies, organizations can take proactive steps to optimize data storage performance and ensure data quality for downstream analytics. Additionally, the system can provide executive dashboards and reports providing risk and data usage analysis across the enterprise, business units, LOBs, and geographies, complemented by action plans and real-time status updates. The system can provide aggregate data risk reporting and data usage analysis by teams and individual data owners via an intuitive dashboard, thereby enabling visibility, analysis and actionability, with real-time status updates and reminders.

[0085] For example, manufacturing companies rely on real-time data for efficient operations. In some embodiments, the system’s Data Observability utilizes machine learning to continuously monitor data pipelines, detecting anomalies or inconsistencies in sensor data readings. By identifying the root cause, manufacturers can optimize data storage performance and ensure data quality for downstream analytics tasks like predictive maintenance. This improves overall data efficiency and avoids production slowdowns.

[0086] The system’s Data Observability and Root Cause Analysis can employ AI / ML, metadata, and content analytics to continuously monitor data pipelines globally. This includes analyzing user access patterns, data transfer activities, and changes to data permissions. The system can further provide dashboards and reports including insightful risk and data usage analysis across the entire enterprise, from business units to individual teams and locations. The reports can include action plans and real-time status updates to keep everyone in the enterprise informed and moving forward. The system’s dashboard can aggregate data risk reporting and usage analysis for teams and data owners. This translates to clear visibility, actionable insights, and real-time updates and reminders, ensuring the enterprise’s data stays secure and compliant.DDT-00225

[0087] For example, a common challenge faced by banks can be building accurate credit risk assessment models. If the training data used for these models contains errors, such as incorrect income information or missing loan repayment history, that model may misclassify borrowers. In some embodiments, the system’s Data Observability & Root Cause Analysis utilizes machine learning to continuously monitor data quality within pipelines and pinpoint anomalies like ROT, unprotected sensitive business data or unauthorized access control across geographies, business units, and LOBs. This allows them to identify and rectify errors in the training data, leading to more accurate and reliable credit risk models.Statistical Ranking Of Data Repositories With Policy Compliance Risk In Unstructured Data Management

[0088] Managing unstructured data across various repositories while ensuring compliance with organizational, industry, and regulatory policies is complex. There is a need for a system to evaluate and rank data repositories based on their compliance risk to prioritize remediation efforts effectively.

[0089] Alternative implementations include manual audits, basic compliance monitoring tools, and policy enforcement software. While these solutions provide some level of compliance management, they lack comprehensive risk scoring and statistical analysis capabilities for an entire enterprise.

[0090] These alternative implementations have several drawbacks. For example, manual audits are time-consuming and prone to human error. Basic monitoring tools often miss nuanced compliance issues and fail to provide actionable insights. Alternative software solutions do not offer detailed statistical analysis or risk ranking, making it challenging to prioritize remediation.DDT-00225

[0091] In some embodiments, a system that monitors compliance and provides a statistical ranking of data repositories based on policy compliance risk is provided. The system can leverage automation and statistical analysis to offer a more accurate and actionable compliance risk assessment.

[0092] In some embodiments, a system determines a statistical ranking of data repositories based on policy compliance risk. The system includes identifying and employing compliance and governance models, assigning weights to variables based on various parameters, assessing compliance, and performing statistical analysis to quantify and rank compliance risks. In some embodiments, Al and ML techniques are integrated to enhance accuracy and efficiency in data classification, anomaly detection, and risk prediction.

[0093] Fig. 3 is a flowchart 300 illustrating a method according to embodiments of the present disclosure. A user can create policies (302) including one or more of a policy for a metadata scan 304, a policy for statistical sampling 306, a policy for sampling based content scan 308 and a policy for sampling based mitigation actions 310.

[0094] In some embodiments, examples of metadata can include native file metadata as well as injected policy metadata, metadata from a catalog, or custom data from tag repository.

[0095] In some embodiments, analysis of metadata can include 1) file type correlation, 2) access control user correlation, 3) file size and metadata correlation, 4) user permission and metadata tag correlation, 5) file type pattern, 6) location based pattern, 7) access (ACL) history pattern, 8) metadata tag patterns, 9) unusual access anomaly pattern.

[0096] In some embodiments, the policy associates a criterion or criteria with a risk type or risk types.DDT-00225

[0097] In some embodiments, the method can begin a metadata scan (312) after receiving the policy for the metadata scan 304. The method can store and update results in an elastic search database (314). The method can also read the stored metadata (e.g., from step 314) and apply a risk analysis (316). In some embodiments, Kafka or another distributed event streaming platform can provoke the risk type identification service to apply or update the risk analysis. Fig. 4, described in further detail below, illustrates step 316 in further detail according to embodiments of the present disclosure. Based on the results of the metadata scan 314, the method can fetch the results from the elastic search database based on a policy identifier (318). In some embodiments, this can be performed using a call query service using a fast API. The method can generate a sampling result based on the sampling policy or policies 306 (320). Fig. 5, described in further detail below, illustrates step 320 in further detail according to embodiments of the present disclosure. The method can generate a manifest based on the sampling result (328), and then store that manifest result in a postgres database (330).

[0098] The method can further update the sampling result against respective files for the respective policy ID using the Kafka service (322). The method can update sampling results in the elastic search database using the policy for sampling mitigation actions 310 (324). The method can also perform action sampling, using the generated sampling result from the content scan service and mitigate actions services, based on policies such as the policy for sampling based content scans 308 (326).

[0099] Fig. 4 is a flowchart 400 illustrating a risk type identification service reading stored metadata and applying a risk analysis according to embodiments of the present disclosure. In some embodiments, the flowchart 400 is an expansion of step 316 according to embodiments of the present disclosure. Metadata information can be read from a database (402). From that readDDT-00225 metadata, the analysis can be performed at least one of the following steps. In some embodiments, a risk analysis can be performed based on file metadata (404). In some embodiments, a risk analysis can be performed based on an abnormal file metadata pattern and metadata trends using an Al engine (406). In some embodiments, a risk analysis can be performed based on a sensitive information file pattern and a content scan using an Al engine (408). The method can then generate a risk type for files based on one or more of the risks analyses.

[0100] Fig. 5 is a flowchart 500 illustrating generate sampling result based on a sampling policy according to embodiments of the present disclosure. In some embodiments, the flowchart 500 is an expansion of step 320 according to embodiments of the present disclosure. A policy for performing statistical sampling can be loaded (e.g., received from a user, loaded from a database) 504. The method fetches metadata of files encompassed by the policy using an API of query service from the elastic search database (506). The method then extracts relative information from metadata (508).

[0101] In some embodiments, sampling can then be performed by one or more of the following steps. Sampling can be performed based on risk category and risk types provided in the policy definition (510). Sampling can be performed based on a file extension provided in policy definition (512). Sampling can be performed based on NTF tags provided in the policy definition (514). Sampling can be performed randomly without any policy criteria (516).

[0102] Once sampling is performed by one or more of steps 510, 512, 514, and / or 516, the method determines a sampling percentage based on the combination of inputs for the sampling type (518). The method accepts weighting percentage of files to include in the sample which doDDT-00225 not follow the sampling criteria (520). The method then creates a data frame of relative information performing random sampling and produces a desired sample file (522).

[0103] In some embodiments, various features of the system can be used to identify compliance and form governance models. Data access policies can provide roles, permissions, authentication, authorizations, and mechanisms. These data access controls can be implemented on one or more data stores for an enterprise, including privilege access principles. In some embodiments, the data access policies can be implemented on data stores or other system.

[0104] In some embodiments, data access can be monitored with access logs. In some embodiments, an Al or ML model can be trained to detect unusual access patterns and potential data or access breaches. In some embodiments, the Al model can be implemented to detect unusual access patterns and potential data or access breaches.

[0105] In some embodiments, compliance can be based on data retention policies. For example, a data retention police can set retention periods, identify data for archiving or deletion, and / r implement automated retention policies. In some embodiments, AI / ML models can be employed to predict optimal retention periods based on data usage patterns.

[0106] In some embodiments, compliance can be based on data disclosure policies. For example, a data disclosure policy can define data sharing rules, monitor transfer logs, and use data masking or anonymization techniques. In some embodiments, AI / ML models can ensure data sharing compliance and detect unauthorized disclosures. In some embodiments, the data disclosure policy can provide for the use of encryption for certain classes of data (e.g., PII, PHI, IP).

[0107] In some embodiments, compliance can be based on data classification policies. A data classification policy can include classification framework. When implemented, the dataDDT-00225 classification can be automated tools for data scanning and provide tagging and labelling based on the data classification policy. In some embodiments, ML models can accurately classify data based on sensitivity and regulatory requirements.

[0108] In some embodiments, thresholds and weights can be configured for data classification and identification. In some embodiments, thresholds and weights can be configured based on geography, such as local regulations and compliance requirements for a region. These thresholds and weights can be defined for data sovereignty and geopolitical risks (e.g., regional instability, change of governments, etc.). In some embodiments, ML / AI models can dynamically adjust weights based on regulatory changes.

[0109] In some embodiments, thresholds and weights can be configured based on data types.For example, higher weights can be assigned to sensitive data. In some embodiments, structured data and unstructured data can be weighted differently. In some embodiments, a ML / AI model can identify and categorize data types accurately.

[0110] In some embodiments, an associated business unit can allocate higher weights to critical business units (e.g., data stores associated with those business units) handling sensitive data and evaluate the impact of non-compliance for those business units. In some embodiments, a ML / AI model can analyze the business impact of compliance risks.

[0111] In some embodiments, compliance can be based on other parameters. For example, other Parameters can include data volume, growth rate, historical compliance performance, technological infrastructure, and incident history. In some embodiments, a ML / AI model can predict future compliance risks based on historical data.

[0112] In some embodiments, compliance can be assessed for an enterprise having data stored in data stores. In some embodiments, compliance can be assessed by determining policyDDT-00225 adherence. For example, based on one or more policy (e.g., the one or more above described policies), audits can be conducted, automated tools can continuously monitor data storage and use, access can be validated, data can be encrypted, and retention can be monitored and enforced. In some embodiments, a ML / Al model can automate compliance checks and identify policy deviations in real time.

[0113] In some embodiments, compliance can be assessed by identifying violations of policies. For example, a system can track policy non-compliance, categorize violations, investigate root causes, and implement corrective actions. In some embodiments, a ML / AI model can prioritize violations based on severity and potential impact.

[0114] In some embodiments, users can be notified of areas of non-compliance. A notification can highlight non-compliant repositories, document reasons, and prioritize areas needing attention (e.g., via a user interface, a push notification, a dashboard, etc.'). Use ML to recommend remediation actions based on past successful interventions.

[0115] In some embodiments, risk scoring can be employed. In some embodiments, scoring systems can be developed to quantify compliance risk, such as using a weighted model to reflect the importance of different factors. In some embodiments, a ML / AI model can fine-tune the scoring model based on new data. Compliance risk can be quantified with one or more factors including a number of violations, severity, data sensitivity, historical compliance trends, incident history, or any combination of the foregoing. Use ML to identify emerging risk factors and adjust scores accordingly.

[0116] In some embodiments, statistical analysis can be employed. Statistical analysis can include calculations of one or more of a mean, a median, and / or a standard deviation of riskDDT-00225 scores. These statistical analyses can aid understanding distribution. In some embodiments, a ML / AI model can be employed to detect outliers and anomalies.

[0117] In some embodiments, repositories of data (e.g, data stores, databases, servers, etc.) can be ordered by risk scores. A rank of repositories can highlight high-risk repositories for immediate attention or mitigation action. In some embodiments, a ML / AI model can be used to continuously or to periodically update rankings as new data is received.

[0118] In some embodiments, a cluster analysis can group repositories with similar risk profiles using k-means and hierarchical clustering. In some embodiments, ML / AI models can identify patterns and correlations within clusters.

[0119] In some embodiments, trend analysis can track changes in compliance risk over time, analyze the impact of policy updates, and detect emerging risks. In some embodiments, ML / AI models can predict future trends and proactively address potential issues.

[0120] Fig. 7 is a flowchart 700 illustrating a process of statistical sampling according to embodiments of the present disclosure. When the process starts, the method can define a sample percentage. The sample percentage can optionally be defined by e.g., by direct user input, reading from a file, reading from a configuration, etc.), but can be defined by other ways (e.g., an automated definition). The method can then determine whether criteria for the sampling policy is selected, such as file extensions, tags, or risk types. If there are no criteria, then the sampling policy generates a sample of files from a random sample. The sampling stops, and a summary of the sample is generated.

[0121] If criteria is selected, those files can be selected by applying the selected criteria. The method can then determine whether to unbiased the sample with files from outside of the selection. If files outside of the selection are to be used, a weighted sampling is used. TheDDT-00225 weighted sampling can define weights applied to files for meeting or not meeting each criterion. If there are sufficient files that match the criteria, the weighted sampling can generate a sample with a defined weightage of matching files and random files. If there are not enough files available, the weighted sampling can generate a sample available matching files, while the rest are randomly selected. In both cases, the sampling stops, and a summary of the sample is generated.

[0122] If no files outside of the selection are to be used (e.g., only files matching the criteria are to be used), then stratified sampling is used. Stratified sampling can generate a biased sample with only matching files. The method can determine whether enough matching files are in the selection to meet the sample size. If so, the stratified sampling generates a sample with matching files. If not, the stratified sampling method can generate a partial sample with available matching files.

[0123] Referring now to Fig. 6, a schematic of an example of a computing node is shown. Computing node 10 is only one example of a suitable computing node and is not intended to suggest any limitation as to the scope of use or functionality of embodiments described herein. Regardless, computing node 10 is capable of being implemented and / or performing any of the functionality set forth hereinabove.

[0124] In computing node 10 there is a computer system / server 12, which is operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set topDDT-00225 boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.

[0125] Computer system / server 12 may be described in the general context of computer systemexecutable instructions, such as program modules, being executed by a computer system.Generally, program modules may include routines, programs, objects, components, logic, data structures, and so on that perform particular tasks or implement particular abstract data types. Computer system / server 12 may be practiced in distributed cloud computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.

[0126] As shown in Fig. 6, computer system / server 12 in computing node 10 is shown in the form of a general-purpose computing device. The components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to processor 16.

[0127] Bus 18 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards AssociationDDT-00225(VESA) local bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Express (PCIe), and Advanced Microcontroller Bus Architecture (AMBA).

[0128] Computer system / server 12 typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system / server 12, and it includes both volatile and non-volatile media, removable and non-removable media.

[0129] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer system / server 12 may further include other removable / non-removable, volatile / non- volatile computer system storage media. By way of example only, storage system 34 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media can be provided. In such instances, each can be connected to bus 18 by one or more data media interfaces. As will be further depicted and described below, memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the disclosure.

[0130] Program / utility 40, having a set (at least one) of program modules 42, may be stored in memory 28 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data or some combination thereof, may include an implementation of a networking environment. ProgramDDT-00225 modules 42 generally carry out the functions and / or methodologies of embodiments as described herein.

[0131] Computer system / server 12 may also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer system / server 12; and / or any devices (e.g., network card, modem, etc.) that enable computer system / server 12 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interfaces 22. Still yet, computer system / server 12 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via network adapter 20. As depicted, network adapter 20 communicates with the other components of computer system / server 12 via bus 18. It should be understood that although not shown, other hardware and / or software components could be used in conjunction with computer system / server 12. Examples, include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0132] The present disclosure may be embodied as a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0133] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storageDDT-00225 device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0134] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0135] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-settingDDT-00225 data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0136] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0137] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of theDDT-00225 computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0138] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0139] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the blockDDT-00225 diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0140] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

DDT-00225CLAIMSWhat is claimed is:

1. A method comprising: reading a metadata policy comprising a criterion; identifying a plurality of files stored in a datastore, the plurality of files having metadata matching the criterion of the metadata policy; associating a risk type with each file of the plurality of files based on its respective metadata; assigning a risk value to each risk type for each file of the plurality of files; reading a statistical sampling policy, the statistical sampling policy comprising a risk value threshold for each risk type; determining a first subset by sampling from the plurality of files based on the risk type of each file according to the statistical sampling policy; selecting a first mitigation action for each file of the first subset based on its risk type and risk value; and performing the first mitigation action for each file of the first subset.

2. The method of Claim 1, further comprising: reading a content policy, the content policy providing a rule for acceptable content for the data store; and assigning a second mitigation action for each of the subset of files based on the content policy; and performing the second mitigation action for each file of the subset of files.DDT-002253. The method of Claim 1, wherein the metadata can include native file metadata, injected policy metadata, metadata from a catalog, or custom data from a tag repository.

4. The method of Claim 1, wherein the metadata policy associates a risk type with the criterion.

5. The method of Claim 1, wherein the metadata policy includes a plurality of criteria, the metadata policy further association each of the criteria with a risk type.

6. The method of Claim 1, wherein determining the first subset by sampling from the plurality of files according to the statistical sampling police generates a sample of files from the plurality of files based on behavior analysis and user input.

7. The method of Claim 1, wherein the metadata can include file type, access control, file size, user permission, tags, location, access history, unusual access, or unusual access anomaly.

8. The method of Claim 7, wherein assigning the risk value is based on determining a correlation or a pattern based on the metadata.

9. The method of Claim 1, wherein the first mitigation action is part of a workflow.

10. The method of Claim 9, wherein selecting the first mitigation is determined based on the first subset determined by sampling the plurality of files.

11. The method of Claim 1, wherein the statistical sampling policy includes one or more of a random sampling policy, a stratified sampling policy, and a weighted sampling policy.DDT-0022512. The method of Claim 11, wherein the random sampling policy includes a metadata parameter and a target percentage value that metadata parameter.

13. The method of Claim 11, wherein the weighted sampling policy a metadata parameter, a priority of that parameter, and a weight of that parameter.

14. The method of Claim 11, wherein the stratified sampling policy includes a metadata parameter.

15. A system comprising: a computing node comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor of the computing node to cause the processor to perform any of the method of any one of Claims 1-14.

16. A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing node to cause the computing node to perform the method of any one of Claims 1-14:

Citation Information

Patent Citations

  • Ranking content items provided as search results by a search application

    US20170357725A1

  • Data processing systems for migrating data between data centers

    US20200267125A1

  • Detection of sensitive database information

    US20210182607A1

  • Method and System for Segmenting Unstructured Data Sources for Analysis

    US20230214403A1

  • Statistical database query using random sampling of records

    US5878426A