Decentralized workflow for unstructured data management via file tagging, sensitivity labelling, and compliance policies
The decentralized workflow system addresses inefficiencies in unstructured data management by using AI/ML for automated tagging and compliance, ensuring secure and compliant data handling across hybrid cloud environments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-03-26
AI Technical Summary
Existing data management systems face challenges in managing unstructured data across hybrid cloud environments, leading to inefficiencies, increased costs, and non-compliance with data sovereignty and regulatory requirements due to siloed data systems and manual, error-prone processes.
A decentralized workflow system that employs AI/ML-driven file tagging, sensitivity labeling, and compliance policies to automate data classification, remediation, and mobility, ensuring data remains localized while adhering to regional regulations, optimizing storage, and enhancing security.
The system provides efficient, scalable, and compliant data management by automating classification and remediation, reducing operational costs, minimizing errors, and ensuring compliance with diverse regulatory requirements.
Smart Images

Figure US2025047494_26032026_PF_FP_ABST
Abstract
Description
DDT-00325DECENTRALIZED WORKFLOW FOR UNSTRUCTURED DATA MANAGEMENT VIA FILE TAGGING, SENSITIVITY LABELING, AND COMPLIANCE POLICIESBACKGROUND
[0001] Today, digital trust has a larger impact than physical trust. According to recent surveys, 76% of consumers desire greater control over their data, 86% of individuals believe humans should be ultimately accountable for artificial intelligence (Al) decisions, employees with data control report increased trust in their organization's data practices, and 63% of business leaders believe data democratization is crucial for fostering a data-driven culture. For an enterprise to become a trusted custodian of data requires digital trust. Digital trust can be described as one or more of an enterprise’s commitment to data privacy, ethical Al, data sovereignty and security, and compliance with social and environmental regulations.BRIEF SUMMARY
[0002] According to embodiments of the present disclosure, methods of and computer program products for decentralized workflow for unstructured data management via file tagging, sensitivity labeling, and compliance policies are provided.
[0003] In some embodiments, a computer-implemented method can include loading a storage management policy from a server. The method can include scanning a first datastore based on the storage management policy, thereby providing scan results. The method can include providing the scan results to the server. The method can include responsive to providing the scan results to the server, receiving a plurality of tasks for bringing the datastore into compliance with the storage management policy. The method can include executing the plurality of tasks at the datastore.Page 1 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0004] In some embodiments, scanning is based on metadata or contents of files stored in the first datastore.
[0005] In some embodiments, receiving the plurality of tasks can further include receiving at least one query based on one of the plurality of tasks and providing the at least one query to a query service at the datastore.
[0006] In some embodiments, the computer-implemented method can further include displaying, in a user interface, the plurality of tasks and receiving, in the user interface, an authorization to execute the plurality of tasks.
[0007] In some embodiments, the task comprises one or more workflows, each workflow comprising an action or processing to be performed on at least one file of the first datastore.
[0008] In some embodiments, the computer-implemented method of Claim 1 can further include indexing the scan results to an elastic search database, and querying the elastic search database.
[0009] In some embodiments, a storage data policy includes a remediation policy, a quarantine policy, a migration policy, a data redundancy policy.
[0010] In some embodiments, receiving the tasks further includes receiving the tasks from the server.
[0011] In some embodiments, the computer-implemented method can include executing a first task of the at least one tasks on at least one file stored by the first datastore. In some embodiments, the computer-implemented method further includes executing, at a second data store, a second task of the at least one tasks on the at least one file.
[0012] In some embodiments, the storage management policy includes a rule and at least one tasks.Page 2 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0013] In some embodiments, the computer-implemented method can further include displaying, in a user interface, a representation of an alarm triggered based on the scan results.
[0014] In some embodiments, a computing node comprises a computer readable storage medium having program instructions embodied therewith. The program instructions can be executable by a processor of the computing node to cause the processor to perform any of the above methods.
[0015] In some embodiments, a computer program product comprises a computer readable storage medium having program instructions embodied therewith. The program instructions can be executable by a computing node to cause the computing node to perform any of the above methods.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0016] Fig. 1 is a flowchart illustrating a method for enterprise data management according to embodiments of the present disclosure.
[0017] Figs. 2A-C are user interfaces of an enterprise data management system according to embodiments of the present disclosure.
[0018] Fig. 3 is a block diagram illustrating a decentralized workflow according to embodiments of the present disclosure.
[0019] Fig. 4 is a block diagram illustrating a decentralized workflow according to embodiments of the present disclosure.
[0020] Fig. 5 is a block diagram illustrating a decentralized workflow according to embodiments of the present disclosure.
[0021] Fig. 6 is a diagram of a computing node according to an embodiment of the present disclosure.Page 3 of 59FOLEYHOAGUS 12449860.3DDT-00325DETAILED DESCRIPTION
[0022] Decentralized identities are on the rise, with over 100 million users projected by 2025. • By 2025, 75% of enterprises will need to implement Explainable Al (XAI) solutions for user trust. By 2025, 15% of large enterprises will begin transitioning to quantum-resistant cryptography to safeguard their data. By 2030, 73% of consumers will prioritize ethical data practices.
[0023] The growing need for Al, automation, and personalization clashes with a rising tide of citizens who demand control over their personal information. This clash creates a strategic dilemma for businesses. The lack of visibility caused by siloed data systems and fragmented point solutions creates a distrust between central IT and business units, which can also spill over to consumers. Consumers can feel a disconnect between what companies promise, what they do with data, and what users actually see.
[0024] Fig. 1 is a flowchart 100 illustrating a method for enterprise data management according to embodiments of the present disclosure. First, a data workflow made of data stores 104, data processing 106 and data action 108 is created (102).
[0025] In some embodiments, data can be scanned for metadata by a policy or storage administrator (120). The metadata can be provided to a dashboard updated with data usage analysis (e.g., for a Chief Infrastructure Officer, Chief Data Officer, or Data Owner) (126). The method can then optimize storage e.g., for the data custodian and data owner) (128). The method can then generate an action plan based on corporate policies (e.g., for the data custodian and data owner) (130). The method can then execute the action (e.g., archive, transform, delete, migrate) via a workflow (e.g., for the data custodian and data owner) (132). Then, the methodPage 4 of 59FOLEYHOAGUS 12449860.3DDT-00325 can update the dashboard for the various personas with a reduced risk and optimized storage insights (136).
[0026] In some embodiments, the metadata is provide to a content scan performed by a policy administrator, data custodian, or data owner (122). The method can then generate a dashboard updated with risk exposure insights for a chief information security officer and a data owner (110). The method can then mitigate risk based on those insights (e.g., by the data custodian, data owner) (112). The method can further generate an action plan based on corporate policies to mitigate those risks (e.g., by the data custodian, data owner) (114). The method can then execute the action (e.g., manage permissions, archive, quarantine, delete) via a workflow (124). Then, the method can update the dashboard for the various personas with a reduced risk and optimized storage insights (136).
[0027] Figs. 2A-C are user interfaces of an enterprise data management system according to embodiments of the present disclosure. In Fig. 2A, the user interface illustrates a dashboard having an overall risk of the entire enterprise across all locations according to embodiments of the present disclosure. In addition, four gauges in the dashboard can display metrics of access control, data redundancy, data sensitivity, and retention compliance. An additional gauge can display a list of top sensitive data, organized by department. Yet another gauge can display statuses (e.g., pending, overdue, in progress, failed, completed) of risk mitigation tasks.
[0028] In Fig. 2B, the user interface illustrates a dashboard having an overall storage efficiency metric according to embodiments of the present disclosure. In addition, five gauges can illustrate data redundancy, cold data, orphaned data, expired data, and junk data. Another gauge can display redundant, obsolete, or trivial (ROT) data by department of the enterprise. Other gauges can display on-premise storage capacity by data store, cloud storage usage by service provider,Page 5 of 59FOLEYHOAGUS 12449860.3DDT-00325 file distribution by type, file distribution by age, file count trend, and a sustainability metric (e.g., carbon emissions).
[0029] In Fig. 2C, the user interface illustrates a dashboard having an overall risk score of the entire enterprise across all locations according to embodiments of the present disclosure. In addition, five gauges in the dashboard can display metrics of access control, data redundancy, data sensitivity, retention compliance, and encryption strength. Yet another gauge can display statuses (e.g., pending, overdue, in progress, failed, completed) of risk mitigation tasks. Other gauges can display sustainability metrics (e.g., by carbon emissions), top sensitive data by department, and file distribution by file type or age and by department.
[0030] It can be recognized that the gauges of Figs. 2A-C can be combined in any manner, and that the data displayed is merely exemplary and non-limiting. For example, drop down menus for each gauge can change the grouping or ranking of the data. Gauges can be moved, removed, or added to a given dashboard according to a user’s preference.
[0031] In some embodiments, systems for enterprise data management are provided. Such systems can guide businesses in orchestrating data democratization and ethical data usage within an Al-driven, climate-focused world.
[0032] In some embodiments, the systems can provide unparalleled understanding, correlation, consistency and actionability for stakeholders across an enterprise or organization. The systems can apply a nimble and scalable architecture and employ generative Al for modelling and insights and visibility across the organization via role-based access. In addition, the system can address risk, privacy, sovereignty and compliance. At the same time, the system can ensure optimized and sustainable use of storage infrastructure.Page 6 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0033] In some embodiments, the system enables people having roles ranging from C-suite executives, business, technology managers, and data owners to all utilize the same data repository. Those people can then make and execute upon informed and intelligent decisions. As a result, data owners can become trusted data champions, and the enterprise or organization meets its responsibility as a trusted custodian.
[0034] In some embodiments, the system can provide data orchestration and trust in four categories. Those categories can include data security and ethical usages, data transformation and sustainability, data governance and compliance, and data democratization and Al empowerment. Each category can be implemented with various aspects, including the aspects detailed below.
[0035] Data Security and Ethical Usages• Privacy Protection for Ethical Al• Risk Analysis & Remediation for Sensitive Information• Data Access Management• Data Security Orchestration, Automation & Actionability
[0036] Data Transformation & Sustainability• Hybrid-Cloud Data Mobility• Data Aggregation and Harmonization• Cloud-Optimized Data Management• Data Lifecycle Management and Footprint Reduction
[0037] Data Governance & Compliance• AI / ML-based Data Observability and Root Cause Analysis(RCA)• RBAC Driven Process & ControlsPage 7 of 59FOLEYHOAGUS 12449860.3DDT-00325• Data Usage & Traceability• Policy-based Data Governance Framework & Insight
[0038] Data Democratization & Al Empowerment• Data Owner Observability, Control and Actionability• Self-Service Analytics and Insights• Data Privacy by Design• Data Wrangling and Curation for AlStorage Optimization and Lifecycle Management (LCM) across Hybrid Cloud Infrastructure
[0039] In some embodiments, a system can provide storage optimization and lifecycle management (LCM) across hybrid cloud infrastructure. The ever-expanding spread of data across hybrid cloud environments can present a challenge: storage optimization and LCM. Unmanaged data sprawl can lead to wasted resources, skyrocketing storage costs, and compliance nightmares. In some embodiments, the system can empowering organizations to streamline storage, automate data governance, and gain valuable insights.
[0040] In some embodiments, the system can include an Al-powered self-service data management software. The software can empowers enterprises by enabling users across all levels, from C-suite to data owners, to discover, define, act, transform, and audit data through a user-friendly interface. The system provides correlation, consistency, and standardization across enterprises by delivering granular insights, deriving recommended workflows, and automating actions using personalized policies and role-based access control (RBAC)-driven processes. This transformation fosters a culture of data ownership, where everyone becomes a data champion, and the organization fulfills its responsibility as a data custodian. The software can overcome these challenges through six key capabilities, strategically aligned with industryPage 8 of 59FOLEYHOAGUS 12449860.3DDT-00325 trends, to help enterprises navigate the complexities of hybrid cloud storage management. It empowers the enterprise to harness the power of the enterprise’s data while keeping storage costs and security risks under control.Data Discovery and Classification
[0041] In some embodiments, the system can provide automated data classification and policy- driven storage tiering. In the method illustrated by Fig. 1, data can be classified by the metadata scan 120 and identified by the content scan 122. Shadow IT, the unauthorized use of cloud services or applications, can lead to sensitive data being stored and accessed outside of an enterprise’s control, jeopardizing data sovereignty. Additionally, unstructured data that can reside in disparate locations across the global footprint of the enterprise can remain unidentified and unclassified, making it vulnerable to unauthorized access. Unclassified data can make identifying and store critical information inefficient. Manual data classification can be timeconsuming and error-prone. Additionally, enterprises can lack a unified storage strategy across on-premises and cloud environments, leading to underutilized resources and wasted storage costs.
[0042] In some embodiments, the system can provide metadata and content analytics that work together to automatically classify data based on type, sensitivity, and how often the data is accessed. The provided metadata and content analytics can further allow intelligent storage management using data tiering and placement features. For example, frequently accessed data can be automatically identified and placed on high-performance storage for optimal retrieval speed. Meanwhile, less used data can be intelligently archived to cost-effective cloud storage tiers like S3 or OneDrive. This automated approach can optimize storage utilization, reducePage 9 of 59FOLEYHOAGUS 12449860.3DDT-00325 costs associated with underutilized resources, and strengthen data governance by ensuring sensitive data resides on secure storage tiers.
[0043] In some embodiments, the system can combine Metadata Analytics and Natural Language Processor (NLP)-powered Content Analytics for data discovery and classification. This combination can automatically identify / discover and classify sensitive data (e.g, PII, PHI, etc.) across a data landscape. As a result, the system can provide data exposure risks before those risks become a data leak, etc. Data owners can take control with a self-service Data Classification, allowing them to further categorize sensitive data based on their specific regulations. Discovery and risk management can be seamlessly integrated to Cloud Integrated Risk Controls (e.g., Microsoft Information Protection). Additionally, the system’s Data Usage & Traceability can provide a centralized record of all data movement activities, making audit trails and compliance reporting a breeze. In addition, the system can ensure data privacy and regulatory adherence by automating data lifecycle management tasks. Based on retention policies, the system can automatically delete or anonymize data.
[0044] For example, pharmaceutical companies can face challenges complying with GDPR regulations regarding patient data privacy in clinical trials. In some embodiments, the system can automatically identify sensitive patient information within research databases. Researchers can then leverage self-service tools to further categorize this data based on specific trial protocols. This granular control can maintain patient privacy while allowing for essential research to continue. Additionally, the system can provide a centralized record of all data access and movement, simplifying audit trails and compliance reporting. This can empower the company to mitigate data privacy risks and ensure regulatory compliancePage 10 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0045] Metadata Analytics automatically scans an enterprise’s network to identify data across repositories across all locations, including file shares, cloud storage buckets, and data lakes. Content Analytics can then analyze the content of the discovered data to pinpoint sensitive information like Personally Identifying Information (PII), Protected Health Information (PHI), and intellectual property (IP). After these steps, an inventory of data assets can be generated. Data Classification enables categorization of data based on sensitivity level (e.g., high-risk PII / PHI and business sensitive data like passwords or network addresses like IP or MAC addresses), regulatory requirements (e.g., HIPAA, GDPR, etc.), and specific business needs. The system enables user definition of custom classification rules, enabling tagging of all data according to the enterprise’s standards. Therefore, the system facilitates automatic risk prioritization, and allows enterprises to focus remediation efforts first on critical areas.
[0046] For example, financial institutions like banks can struggle with managing a mix of critical and less-used data. In some embodiments, the system can automatically classify regulatory documents, loan applications, and emails, placing sensitive data on secure high-performance storage while archiving less frequently accessed info to cost-effective cloud tiers. This optimizes storage utilization and reduces costs.
[0047] As an example, suppose a global bank suspects shadow IT is storing customer data (e.g., PII such as social security numbers, account numbers, account details, etc.) in unauthorized cloud locations (e.g., scattered across emails, loan applications, and internal documents). In some embodiments, the system can scan the network to identify all data repositories, and use Content Analytics to find PII such as names, social security numbers, and account information. This helps the bank remediate shadow IT practices, ensuring customer data remains controlled and compliant with data sovereignty regulations.Page 11 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0048] As yet another example, financial institutions can struggle with the volume customer data (e.g., loan applications, account statements, transaction records). Manually sorting through this data is a time-consuming and expensive task. However, the system’s Al-powered analytics automatically classify the data, saving the bank time and money. Less-frequently accessed data can be archived to cost-effective storage tiers, freeing up valuable space for real-time transaction data. This allows the bank to focus on what matters most: delivering exceptional customer service.
[0049] Preparing data for Al and machine learning projects can be time-consuming and complex. Data scientists can struggle to find, access, and clean data from diverse sources across the organization. In some embodiments, the system’s Data Aggregation and Harmonization can consolidate data from file storage, object storage, and data lakes, eliminating time-consuming manual collection. Data owners and data scientists can work together seamlessly using self- service tools to define granular access permissions, keeping data secure and compliant to regulations and other rules or policies. The system can initiate data migration and even transform it (e.g., file to object) for an even smoother workflow. Automated data cleansing workflows ensure pristine data quality by removing inconsistencies and errors that could reduce the accuracy of AI / ML models. This translates to streamlined data preparation, faster AI / ML project development, and unlocking the true potential of data for advanced analytics.Data Redundancy Elimination and Archival
[0050] In some embodiments, the system can provide Data Redundancy Elimination and Archival for Improved Efficiency. Data sprawl can lead to redundant, obsolete, and trivial (ROT) data consuming valuable storage space. Manual identification and deletion of redundant data is cumbersome and error-prone. Traditional archiving processes can be complex and hinderPage 12 of 59FOLEYHOAGUS 12449860.3DDT-00325 data accessibility. In the method illustrated by Fig. 1, data redundancy elimination and archival can be performed by the actions executed by the workflow in steps 124 and 132, including archival, deletion, migration, quarantining, etc.
[0051] In some embodiments, the system can find hidden data and help manage that data efficiently. With Data Redundancy, an enterprise can identify and eliminate duplicate data through intelligent classification and tagging. This can free valuable storage space for active data. In addition, cleaning up such data does not remove access to it. The system can intelligently converts redundant files into space-saving object formats (e.g., for archiving). These archived objects can still be easily retrieved using software, a RESTful API (e.g., for audits and legal discovery). Plus, the system provides S3 / NFS / SMB-compliant access, ensuring compatibility with existing infrastructure and making data retrieval a breeze.
[0052] For example, a pharmaceutical enterprise can have redundant clinical trial data, which wastes storage space for active research. The above system can identify and eliminate duplicate data sets based on intelligent tags, freeing up valuable storage space for ongoing research projects and improving data accessibility for scientists.Automated Data Lifecycle Management
[0053] In some embodiments, the system can provide Automated Data Lifecycle Management with Retention Compliance. Manually managing data retention policies across hybrid cloud environments can be complex and error-prone. Inconsistent data handling practices and manual enforcement of data policies can lead to security vulnerabilities and compliance gaps. Non- compliance with data retention regulations can sometimes lead to hefty fines. Therefore, risk mitigation and regulatory compliance can be desired.Page 13 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0054] In some embodiments, the system provides Retention Compliance that seamlessly integrates the metadata analysis with data policy creation and workflow automation. In relation to the method illustrated by Fig. 1, workflows can be automated by the creation of a data workflow by step 102 and the executing action via workflow 124 or 132. This powerful combination can allow enterprises to discover, classify, and tag data based on specific retention requirements. The software can automate data lifecycle management tasks based on pre-defined policies. The system can handle deletion, archiving, or migration, thereby ensuring effortless compliance with data retention regulations. Therefore, the system minimizes risks associated with non-compliance and streamlines storage management by eliminating the need for manual tasks.
[0055] For example, hospitals can face challenges ensuring compliance with complex healthcare data privacy regulations. In some embodiments, the system’s Retention Compliance utilizes metadata analysis to classify patient medical records and automate data lifecycle management based on retention requirements. This can ensure HIPAA compliance, minimizes regulatory risks, and streamlines storage management.
[0056] Without a clear data governance framework, data transfers across borders can be haphazard, potentially violating data sovereignty regulations. Inconsistent data handling practices can also lead to confusion and difficulty in demonstrating compliance.
[0057] In some embodiments, the system seamlessly integrates with an existing data governance framework to establish data policies that uphold data sovereignty best practices. The system’s Data Policy Creation and Workflow functionality simplifies defining clear rules for data movement and usage. These policies can be tailored based on data sensitivity, location, and retention requirements for different data types. Furthermore, the system gives data ownersPage 14 of 59FOLEYHOAGUS 12449860.3DDT-00325 control over their data through Self-Service Data Classifications and Migrations. This feature streamlines data management by allowing data owners to classify and migrate data according to established policies. The system can then automates the data migration process, ensuring data reaches the appropriate storage locations while adhering to data sovereignty regulations. This not only simplifies governance but also reduces the risk of errors.
[0058] For example, a multinational corporation with worldwide offices collects customer data from various regions. In some embodiments, the system can define clear policies for data movement and usage that comply with data sovereignty regulations in each country they operate. The system’s Self-Service Data Classifications and Migrations can empower business units to classify and migrate their data according to these established policies. Therefore, consistent data handling practices compliance with data sovereignty regulations across the global organization are ensured.
[0059] As another example, hospitals are under constant pressure to comply with strict regulations regarding patient data retention. Manual processes for managing patient records can be error-prone and lead to hefty fines. In some embodiments, the system can automate data retention tasks, ensuring hospitals stay compliant and avoid costly penalties. Additionally, the system can provide complete transparency into how patient data is used, fostering trust with patients and regulators. This allows hospitals to focus on their core mission: delivering quality patient care.
[0060] In some embodiments, the system can ensure consistent data handling across an enterprise with the system’s Data Policy Creation and Workflow. The Data police can clearly define comprehensive data policies that translate into automated workflows. These workflows can enforce access controls, data encryption, and other security measures, minimize human error,Page 15 of 59FOLEYHOAGUS 12449860.3DDT-00325 and guarantee consistent policy application. These workflows work with the results of Data Discovery and Classification. By classifying data based on sensitivity and regulations, the system can use data policies and workflows tailored to the specific needs of the data as it is classified. This can streamline data governance, simplify compliance, and strengthen data security.
[0061] For example, consider an insurance company that implements new data governance policies to comply with industry regulations regarding data privacy and security. The system’s Data Policy Creation and Workflow functionality can empower them to personalize and translate these policies into automated workflows. These workflows can include automatic data discovery and classification for sensitive research data, access management, migration, retention compliance, archival, deletion, and more. Automating these processes can minimize human error and ensures consistent enforcement of data governance policies across the organization.Self-service data classification
[0062] In some embodiments, the system provides Self-Service Data Classification and Migration for Democratized Data Management. Complex data management processes and limited visibility into data storage locations can hinder business users from effectively managing their data. This can lead to data silos and hinder collaboration efforts. Alternative data governance and migration point solutions, often complex and inaccessible, can create data silos, hinder collaboration, and limit the value businesses can extract from their information. Enterprises can move forward from a “lift and shift” approach and embrace a “data-driven” strategy for infrastructure management. In some embodiments, in the method illustrated by Fig. 1, data can be classified by an enterprise or a user providing rules to the metadata scan 120 or content scan 122.Page 16 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0063] In some embodiments, the system’s Data Dynamics' Al-powered self-service data management software can bring a fresh approach to privacy, security, compliance, governance and optimization in the world of Al-led workloads.
[0064] In some embodiments, the system can provide business users the ability to take control of their data with Self-Service Data Classifications and Migrations. This intuitive functionality can allow these business users to classify data based on their department's specific needs and business context. User-friendly tools can enable business users to initiate data migrations across on-premises and cloud storage effortlessly, fostering better collaboration and data governance. Improved data visibility through self-service tools can empower users to take ownership of their data and manage it effectively. Use of these self-service tools can dismantle data silos and can facilitate seamless collaboration across departments, ensuring everyone has access to the information they need.
[0065] For example, insurance companies can have trouble collaborating with data silos. In some embodiments, the system’s Self-Service Data Classifications can empower underwriters and claims adjusters to classify data by insurance product or policyholder. They can initiate data migrations between on-premises and cloud storage, fostering collaboration across departments and improving data governance.
[0066] This solution brief explores sixkeyuse casesthathighlighthow Zubin addresses critical industry challenges, such as data silos and inefficient data wrangling. We will delve into the software’s role in empowering organizations to automate data classification, streamlinemigrations, andbreak down datasilos. By enabling self-servicedataownershipandfostering collaboration, Zubin paves the way for a future where data is readily available for analysis, leading to better decision-making and improved businessoutcomes.Page 17 of 59FOLEYHOAGUS 12449860.3DDT-00325Risk Insights
[0067] In some embodiments, the system can provide data usage and traceability with risk insights for enhanced security. Static data security measures may not be sufficient to address evolving threats and vulnerabilities. Lack of visibility into data usage across hybrid cloud environments can make it difficult to identify and mitigate security risks. Unclear data ownership can lead to confusion about access control responsibilities. Varying data privacy regulations across regions can sometimes require data localization policies. Without proper risk assessment and data localization strategies, an enterprise may be inadvertently violating data sovereignty laws. In some embodiments, in the method illustrated by Fig. 1, the dashboard 110 illustrates risk insights based on the analysis described herein.
[0068] Tracking and maintaining a clear audit trail for data movement across borders is essential for demonstrating compliance with data sovereignty regulations. Without this visibility, regulators may question data handling practices and ability to ensure data remains within desired jurisdictions.
[0069] In some embodiments, the system can help locate data and help secure it. The system’s Data Usage and Traceability feature can provide a central hub or centralized data index for all data movement activities, giving complete visibility of an enterprise’s data. This central hub can be implemented with a dashboard to show data owners how their data is being used. It meticulously tracks every transfer, capturing the source, destination, purpose, and timestamp, along with the user or system responsible. This comprehensive audit trail enables clear and responsible data movement practices. It simplifies compliance with data residency requirements by providing a clear picture of where data resides and how the data is being used. This not onlyPage 18 of 59FOLEYHOAGUS 12449860.3DDT-00325 fosters trust but also ensures the enterprise have the information needed to manage the enterprise’s data effectively.
[0070] With Risk Exposure Insights, a powerful analytics tool, an enterprise can identify potential security threats by analyzing data usage patterns. Risk Exposure Insights can identify exposed data and prioritize threats to address first. Risk Exposure Insights can leverage data classification and advanced content analytics powered by Al and Machine Learning (AI / ML) to assess the true risk level of exposed data. Beyond basic risk scoring, Risk Exposure Insights can consider factors such as the type of data exposed and the likelihood of exploitation based on realtime threat intelligence. This allows enterprises to focus remediation efforts on the most critical threats, ensuring the enterprise addresses sensitive data at a highest risk of misuse first.
[0071] For example, consider an energy company that suspects a cyberattack on its industrial control systems (ICS) managing power grids, potentially compromising sensor data, configuration files, and employee credentials. In some embodiments, the system can analyze the data access patterns and can classify those patterns based on sensitivity assigning the highest risk to configuration files due to the potential for widespread power outages. This allows the enterprise to prioritize securing the ICS and critical infrastructure, then address employee credentials and sensor data, minimizing power grid disruption.
[0072] Further strengthening an enterprise’s defenses, RBAC Down to the Data Owner Layer provides many levels of access control. This system allows for granular permission management based on who owns the data, ensuring only authorized users can access it. RBAC can access data at the data owner layer. Therefore, the system seamlessly integrates with existing access control features to provide unparalleled granular control. As a result, an enterprise can define which users can access specific data and at what level they can access it (e.g., read, write, edit,Page 19 of 59FOLEYHOAGUS 12449860.3DDT-00325 read and write, etc.). The system can enforce the principle of least privilege, ensuring users only have the access permissions necessary for their tasks. This significantly reduces the points of attack for a breach, minimizing the potential damage caused by compromised credentials or malicious insiders.
[0073] As an example, consider a pharmaceutical company that develops a new drug. Sensitive research data (e.g., formulas, clinical trial results, etc.) can be shared with collaborators while maintaining strict access controls. Pharmaceutical companies also generate vast amounts of data during drug development, but much of that data is hidden and unused. The system’s RBAC Down to the Data Owner Layer can empower researchers to define granular access permissions for collaborators, ensuring only authorized personnel can access specific data elements. This reduces the attack surface and safeguards valuable intellectual property. The system can further shed light on this “dark data” e.g. , the hidden or unused data) by providing a centralized view of all research data. Data owners, like research scientists, can see exactly who has accessed their data and for what purpose. This transparency fosters trust and accountability within the research team. Additionally, the system can empower researchers and scientists to identify potentially valuable data that can be used for further analysis, accelerating drug discovery and development.
[0074] As another example, data governance for insurance companies has been a complex task accessible only to IT specialists. This often led to data silos and hindered collaboration between departments. In some embodiments, the system empowers non-IT users, such as underwriters and claims adjusters, to take charge of their data. User-friendly tools allow users to classify data based on risk factors and claim types, fostering a culture of data ownership. Additionally, granular access controls ensure data security and compliance. This breakdown of data silos encourages collaboration and empowers business users to make faster, more informed decisions.Page 20 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0075] This comprehensive approach creates a multi-layered shield for the enterprise’s data. The enterprise can have clear visibility into movement, the ability to identify potential risks through usage analysis, and ironclad access controls thanks to granular ownership permissions. This can significantly minimize points of attack and prevent data breaches, thereby keeping the enterprise’s information safe.
[0076] In some embodiments, the system employs content analytics to classify data and assess the risk associated with where the data is stored and processed. This analysis considers factors including data sensitivity, applicable regulations, and the security strength of storage locations. In some embodiments, the system allows prioritization of risks, developing targeted mitigation strategies, and defining policies to automatically route specific data types (e.g., PII) to designated storage locations within specific regions. This ensures compliance with data residency requirements. In addition, in some embodiments, the system can integrate with leading cloud storage providers like Microsoft Azure to effortlessly enforce data residency policies.
[0077] As another example, consider an insurance company that operates multiple European countries. The European Union’s (EU) General Data Protection Regulation (GDPR) mandates that citizen data be stored within the EU. Manual processes for isolating sensitive data can be slow and inefficient, which can potentially allowing breaches. In some embodiments, the system can automatically discover and classify customer data (e.g., health information) spread across various locations and assess risks associated with it. Data Containment and Isolation can then allow the enterprise to define policies to automatically route sensitive customer data (e.g., health records) to designated storage locations within the EU, ensuring compliance with GDPR and data sovereignty regulations. This feature can automate workflows to quarantine high-risk sensitive data the moment it is detected. By restricting access and preventing dissemination,Page 21 of 59FOLEYHOAGUS 12449860.3DDT-00325Data Containment minimizes potential damage from exposed data. Furthermore, the Data Encryption adds another layer of defense by safeguarding sensitive data at rest and in transit. This ensures data remains protected throughout the data’s lifecycle.
[0078] As yet another example, consider a hospital chain with facilities in the US and Canada that needs to comply with HIPAA regulations, which requires strict controls on patient data movement, or other regulations such as the California Consumer Privacy Act (CCPA). The system’s Data Usage & Traceability provides a comprehensive audit trail for all patient data transfers, including the source, destination, and time of each transfer. This allows the hospital chain to show to regulators that patient data is only transferred across borders for legitimate medical purposes and in accordance with HIPAA guidelines.
[0079] For example, unauthorized access to energy data pipelines can be disastrous. In some embodiments, the system’s Data Usage & Traceability provides a centralized record of data movement activities. Combined with Risk Exposure Insights, it can analyze data usage patterns to identify suspicious access attempts. Additionally, RBAC Down to the Data Owner Layer can strengthen access controls, preventing unauthorized access and minimizing the attack surface for data breaches.
[0080] As yet another example, consider a disgruntled employee with access to a hospital’s radiology department. That disgruntled employee, in this example, attempts to steal patient data. In some embodiments, the system’s Data Classification can be configured to identify and tag patient medical images. Data Containment and Isolation can then automatically quarantine these files, preventing any such rogue access (e.g., by the disgruntled employee) and minimizing potential patient privacy violations. The hospital can then investigate the attempted incident, revoke the employee’s access, and restore the quarantined data from secure backups.Page 22 of 59FOLEYHOAGUS 12449860.3DDT-00325Data Observability and Root Cause Analysis
[0081] In some embodiments, the system can provide Data Observability & Root Cause Analysis for Performance Optimization. Performance bottlenecks and data quality issues can hinder critical applications and analytics across hybrid cloud environments. Identifying the root cause of these issues can be time-consuming and complex. In addition, monitoring data usage and detecting anomalies across a global data footprint can be complex. Without this data observability, an enterprise may not be alerted to unauthorized access or potential data breaches that could compromise data sovereignty. In some embodiments, in the method illustrated by Fig. 1, the dashboard 126 provides data usage analysis that can be used for performance optimization 128.
[0082] The system’s Data Observability and Root Cause Analysis functionality can utilize metadata & content analytics powered by AI / ML to continuously monitor data usage within pipelines across the enterprise’s global footprint. This includes analyzing ROT, user access patterns, data transfer activities, and changes to data permissions. By identifying the anomalies, organizations can take proactive steps to optimize data storage performance and ensure data quality for downstream analytics. Additionally, the system can provide executive dashboards and reports providing risk and data usage analysis across the enterprise, business units, LOBs, and geographies, complemented by action plans and real-time status updates. The system can provide aggregate data risk reporting and data usage analysis by teams and individual data owners via an intuitive dashboard, thereby enabling visibility, analysis and actionability, with real-time status updates and reminders.
[0083] For example, manufacturing companies rely on real-time data for efficient operations. In some embodiments, the system’s Data Observability utilizes machine learning to continuouslyPage 23 of 59FOLEYHOAGUS 12449860.3DDT-00325 monitor data pipelines, detecting anomalies or inconsistencies in sensor data readings. By identifying the root cause, manufacturers can optimize data storage performance and ensure data quality for downstream analytics tasks like predictive maintenance. This improves overall data efficiency and avoids production slowdowns.
[0084] The system’s Data Observability and Root Cause Analysis can employ AI / ML, metadata, and content analytics to continuously monitor data pipelines globally. This includes analyzing user access patterns, data transfer activities, and changes to data permissions. The system can further provide dashboards and reports including insightful risk and data usage analysis across the entire enterprise, from business units to individual teams and locations. The reports can include action plans and real-time status updates to keep everyone in the enterprise informed and moving forward. The system’s dashboard can aggregate data risk reporting and usage analysis for teams and data owners. This translates to clear visibility, actionable insights, and real-time updates and reminders, ensuring the enterprise’s data stays secure and compliant.
[0085] For example, a common challenge faced by banks can be building accurate credit risk assessment models. If the training data used for these models contains errors, such as incorrect income information or missing loan repayment history, that model may misclassify borrowers. In some embodiments, the system’s Data Observability & Root Cause Analysis utilizes machine learning to continuously monitor data quality within pipelines and pinpoint anomalies like ROT, unprotected sensitive business data or unauthorized access control across geographies, business units, and LOBs. This allows them to identify and rectify errors in the training data, leading to more accurate and reliable credit risk models.Page 24 of 59FOLEYHOAGUS 12449860.3DDT-00325Decentralized Workflow for Unstructured Data Management via File Tagging, Sensitivity Labeling, and Compliance Policies
[0086] Managing large volumes of unstructured data across various locations and systems presents challenges in ensuring data integrity, security, and regulatory compliance. Organizations often struggle to balance operational efficiency with stringent data sovereignty laws and diverse regional regulations.
[0087] Alternative solutions for unstructured data management include centralized data management systems, manual tagging and labeling, and ad-hoc compliance management. In these alternative solutions, a centralized system handles data in a single location. Such a single location for handling data help providing uniform compliance and management, but is not scalable as the volume of unstructured data grows. Further, a single location may not comply with data sovereignty laws, which is unsuitable for global enterprises.
[0088] Manual tagging and labeling relies heavily on human intervention to categorize and manage data. Such manual processes are time-consuming, labor-intensive, and increase operational costs. Manual tagging is also prone to operator / human error, which leads to inconsistent data categorization. In addition, manual tagging is not scalable for large datasets, which limits its effectiveness in large enterprises.
[0089] Ad-hoc compliance management disparate tools and processes to meet regulatory requirements. Ad-hoc compliance can result in inconsistent enforcement of regulatory requirements and risks non-compliance. Ad-hoc compliance further is an inefficient use of resources and leads to higher operational costs. Ad-hoc compliance further is not standardized and results in fragmented and unreliable data management processes.Page 25 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0090] In some embodiments, a system that combines automated tagging, sensitivity labeling, and compliance policies within a decentralized architecture is desired. Such a system can ensure data compliance, improve operational efficiency, and provide better data management capabilities across multiple regions and jurisdictions.
[0091] In some embodiments, the system provides for a decentralized workflow system for managing unstructured data that combines file metadata, tagging, sensitivity labelling, and dynamic compliance policies. By decoupling data storage from its analysis, the system can enable localized processing of data while maintaining centralized control over rules and workflows. This hybrid approach can ensure that data is managed securely and in compliance with regional, security, and domain-specific regulations. Leveraging artificial intelligence (Al) and ML, the system automates the classification, remediation, and mobility of files based on real-time insights from file tags, labels, and content scans. Additionally, the system can evolve by learning from the decisions and exceptions raised by data custodians (e.g., by training Al and ML models on those decisions and exceptions), thereby improving future recommendations and reducing false positives. This approach can optimize both data security and operational efficiency, while providing flexibility to adapt to the diverse and ever-evolving landscape of global compliance requirements.
[0092] In some embodiments, a system using a decentralized workflow for unstructured data management leverages file tagging, sensitivity labelling, AI / ML models, and compliance policies to provide a dynamic, scalable, and compliant data management solution. A system can employ this decentralized workflow to automate the classification, processing, and remediation of data, enhancing both security and efficiency.
[0093] Decoupled Storage and AnalysisPage 26 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0094] In some embodiments, the system of the present disclosure separates data storage from its processing, allowing data to remain in its original location while processing and analysis occur locally. This approach ensures: a. Data Locality: Sensitive data can remain within its geographic or organizational boundaries, minimizing risk and adhering to local compliance regulations. b. Efficient Processing: Data processing and analysis can be performed by distributed engines deployed near the data source, reducing latency and network overhead. c. Enhanced Security: By avoiding central storage of sensitive information, the system can mitigate risks associated with data breaches and non-compliance with data localization laws.
[0095] This separation of managing rules and policies centrally while executing them locally allows for more flexible and compliant data management, aligning with the needs of modem data governance.
[0096] Automated Classification through File Tagging and Sensitivity Labelling
[0097] The system can automate the classification of unstructured data by using AI / ML models to analyze both existing and newly added tags. These tags are based on metadata, content scans, and compliance requirements. The AI / ML models can continuously leam from the tags and sensitivity labels applied to data files and objects, thereby determining the most appropriate compliance-centric model to apply based on geography, security, or domain.
[0098] The labeling process can relies on a combination of metadata information and a deep content scan of the files. This can enable the system to automatically identify sensitive information (such as PI / PII) and recommend remediation actions based on the data’s ownership,Page 27 of 59FOLEYHOAGUS 12449860.3DDT-00325 location, and compliance context. The model can dynamically adjust based on new information, providing more accurate recommendations as it processes more data.
[0099] AI / ML-Driven Compliance Models
[0100] The system can implement the compliance policies with AIZML-driven models that adapt based on the data’s location, content, and applied tags. These models are tailored to specific regulations such as geography-based compliance (e.g., CCPA, GDPR, PDPA), security frameworks (e.g., SOC, PCI DSS), and domain-specific policies (e.g., HIPAA).
[0101] The AI / ML models not only classify data but also recommend actions such as data mobility (e.g., quarantine, archive, migrate) or remediation (e.g., delete, audit log, change access permissions) based on the compliance requirements tied to the tags and labels. The model can continue to be trained over time by learning from data custodian decisions and common patterns, refining its understanding of the actions most suited to particular data sets.
[0102] Risk Type Identifier, Task Assignment, and Action Recommendation
[0103] The Risk Type Identifier, Task Assignment, and Action Recommendation process integrates Al and ML to provide a dynamic and efficient approach to data remediation and management. The process includes the following components:
[0104] Score-Based Identification:
[0105] Policy Administrators can establish and provide scores for different data sets based on risk factors such as sensitivity, compliance needs, and security threats. These scores are used to categorizing data according to its risk profile. Al and ML models can analyze historical data and scoring patterns to refine the scoring criteria. These models can that the risk identification is accurate and reflective of current conditions.
[0106] Workflow-Based Action Recommendation:Page 28 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0107] Data sets are routed through predefined or defined workflows (e.g., data pipelines) that align with the scoring and risk profiles. These workflows can define remediation or mobility actions needed based on the type of risk identified. The system employs Al or ML models to dynamically select a most appropriate workflow based on real-time data characteristics, past actions, and emerging patterns. Therefore, recommendations can be contextually relevant and optimized for efficiency.
[0108] Task Assignment:
[0109] After determining the appropriate workflows, tasks can be generated for data custodians. Each task includes one or more actions based on the data’s risk profile and the associated workflow. Al models can provide task prioritization and assignment by predicting an impact of different actions and optimizing task distribution based on custodian expertise and workload.
[0110] Action Recommendation:
[0111] The system can provide recommendations for remediation actions as part of the tasks assigned to custodians. These recommendations can be based on the combined input of risk scores, workflows, and historical data. ML models can continuously learn (e.g., be trained, retrained, or fine-tuned) from custodian actions, feedback, and outcomes to enhance the recommendation engine. This learning process can refine the accuracy of future recommendations and can adapt to new data patterns.
[0112] Task Tracking and Feedback:
[0113] Tasks can be tracked within the system to monitor progress and ensure timely completion. Al model-driven analytics can provide insights into task performance and effectiveness. Custodians can provide feedback on the recommendations and actions taken, which is used by the system to improve the Al and ML models, reducing false positives andPage 29 of 59FOLEYHOAGUS 12449860.3DDT-00325 improving overall system accuracy. This process can employ Al and ML models to enhance the decision-making capabilities of the system. These models can ensuring that remediation and data management actions are both effective and adaptive to evolving data landscapes.
[0114] Data Custodian Consent and Governance
[0115] Data Custodians play a crucial role in governing the workflows and actions recommended by the system. In some examples, a decision board can be presented to a data custodian, and the data custodian can review aggregated groups of files tagged for remediation or mobility. The system, through the decision board or user interface, provides multiple paths for action, such as data movement, deletion, or retention. Custodians have the authority to: a. Approve workflows for specific files or data sets, b. Raise exceptions if files are incorrectly flagged as security risks or sensitive, and c. Acknowledge that certain files should remain in their current state, even when recommended for remediation.
[0116] The AI / ML model can learn from these custodian interactions, including approved actions and exceptions, improving its accuracy for future scans. Over time, this can reduce false positives and increases the relevance of recommendations, ensuring that only truly sensitive or non-compliant files are flagged for further action.
[0117] Dynamic Policy Application via Distributed Data Engines
[0118] To ensure compliance and performance, distributed data engines can be deployed near the data source. These engines apply localized rules for data discovery, classification, and remediation based on the AIZML-driven models. Each distributed engine can process data within the constraints of its geographic location or security environment, adhering to relevant governance, risk and compliance (GRC) policies. This can ensure:Page 30 of 59FOLEYHOAGUS 12449860.3DDT-00325 a. Compliance with regional or domain-specific data regulations. b. Data security by providing that sensitive files remain within secure environments. c. Efficiency with reduced latency and fewer network hops due to localized processing.
[0119] Centralized Control with Localized Execution
[0120] The system can provide a central application for governance (e.g., managing policies and rules separate from data), while execution of those governance rules occurs at distributed data engines at or near the data’s location. This offers the following advantages:
[0121] Central Oversight: The central application (e.g., deployed via Kubemetes) controls the system’s rules, policies, and workflows. This can ensure uniformity in compliance enforcement, and allow application of global updates or policy changes across multiple regions. The centralized control cam also simplify management, allow administrators to oversee vast and diverse data sets without needing to access each location individually.
[0122] Local Processing: Although policies are set centrally, all data can be processed locally at distributed engines close to the data’s physical storage. This can reduce the need to transfer sensitive data, ensuring compliance with data localization laws (e.g., GDPR, CCPA). Localized execution can enhance security by keeping sensitive information within secure environments, while also reducing latency, optimizing processing speeds, and minimizing network load.
[0123] Efficient Scaling: As the system scales, new data centers or geographies can be easily added. Therefore, there is no need to migrate data to a central location, as the distributed engines handle localized execution. This modular design can allow for effortless expansion across different regions, ensuring the system can grow in line with organizational needs while maintaining full compliance with regional data regulations.Page 31 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0124] Learning and Evolution of Models
[0125] The AI / ML models can continuously learn from: a. Custodian decisions: The paths chosen by Custodians for data mobility or remediation, b. Exceptions raised: When Custodians flag files as incorrectly identified, and c. Data patterns: Recurring actions across different data sets, locations, and contexts.
[0126] As the system processes more data and interacts with Data Custodians, it can improve the accuracy of recommendations, prioritizing the right actions for each set of files and reducing the occurrence of false positives.
[0127] In some embodiments, an advantageous feature of the present disclosure is automated compliance. AI / ML models can dynamically apply compliance rules based on file tagging, sensitivity labelling, and data location.
[0128] In some embodiments, an advantageous feature of the present disclosure is data security. Sensitive data can remain within its original location, with secure, localized processing reducing the risk of breaches.
[0129] In some embodiments, an advantageous feature of the present disclosure is improved accuracy. AI / ML learning can be retrained, fine-tuned, etc. based on custodian decisions and frequent behaviors to reduce errors.
[0130] In some embodiments, an advantageous feature of the present disclosure is efficient processing. Distributed engines can minimize network latency and overhead by processing data near its source. It also reduces the data duplication that would be required for other analytic techniques.Page 32 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0131] In some embodiments, an advantageous feature of the present disclosure is scalability. The decentralized architecture allows the system to scale across different geographies and domains.
[0132] In summary, embodiments of the present disclosure can transform unstructured data management by combining decentralized workflows, AI / ML-driven compliance models, and real-time governance from Data Custodians. It can provide a scalable, flexible, and compliant framework for managing sensitive data across distributed environments.
[0133] In some embodiments, automated tagging and labeling can employ ML and Al models to automatically tag and label unstructured data based on content, metadata, context, user behavior, and organizational behavior. In some embodiments, a tagging system can perform automated tagging using ML / AI models configured to automatically tag unstructured data based on one or more of content, metadata, context, user behavior, business unit behavior, organization behavior, or a combination of any of the foregoing. In some cases, automated tagging can identify edge cases where the tagging system has a low level of confidence of the tag (e.g., a level of confidence below a certain threshold). In such cases, the user interface can prompt a user to manually tag or confirm a tag. In some examples, the user interface can also allow the user to manually tag an incorrectly tagged item, even if the level of confidence is not below the certain threshold.
[0134] Further, an automated sensitivity labeling system can employ AI / ML models configured to automatically label unstructured data based on similar criteria. In some embodiments, the tags of the system can be structured in a tag hierarchy. Such a tag hierarchy, or hierarchical tagging structure, can enable granular data categorization and easier management.Page 33 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0135] In some embodiments, decentralized policy enforcement can ensure compliance with data sovereignty laws by processing at local data stores, while maintaining centralized control over those data stores. In some embodiments, centralized policies can be defined that can be enforced locally, ensuring compliance with regional regulations. For example, a centralized policy can include policies for different geographic regions, types of data, etc. that are employed across an entire enterprise.
[0136] In some embodiments, there can be multiple policy types. A data retention policy can include rules to retain data for a specified period based on legal and business requirements. A data minimization policy can include rules to minimize data actively being stored by archiving, quarantining, transforming, deleting, or managing permissions. A data discovery policy can include rules for discovering sensitive or critical data, with capabilities to remediate or acknowledge errors.
[0137] In some embodiments, policies can be enforced in a decentralized manner to ensure compliance respecting data sovereignty and local regulations. The system can implement the centralized policies to support large-scale data management across multiple regions, improving operational efficiency and data governance. The system provides centralized control for defining policies and managing data compliance across all regions of an enterprise. The system can implement the centralized policies through a federated and distributed architecture. The system performs local processing to comply with data sovereignty laws and reduce latency. That is, the system allows each local data store to analyze the centralized policies according to its unique location, settings, content, etc. In addition, data stewards can be provided a dashboard illustrating a view of data compliance across different regions / jurisdictions while maintaining localized enforcement.Page 34 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0138] In some embodiments, the system manages workflow with a data ingestion phase, a metadata extraction phase, and an actionable insights phase. The data ingestion stage can include reading unstructured data from various sources such as SMB, NFS, S3 Object stores, OneDrive, SharePoint, Data Bricks, Snowflake, and Data Lakes, or any combination of the foregoing sources. The metadata extraction phase can include extract and analyze metadata from existing data to support tagging and compliance policies. The actional insights phase includes providing (e.g., in a user interface such as a dashboard) data usage, compliance status, and potential risks.
[0139] In some embodiments, the system can provide users one or more remediation action or acknowledge errors to ensure data compliance and integrity.
[0140] In some embodiments, the system can perform audits including tracking file-level changes for accountability and transparency.
[0141] Fig. 3 is a block diagram 300 illustrating a decentralized workflow according to embodiments of the present disclosure. A metadata scanner 302 scans a data store for metadata, such as a policy identifier, a policy execution identifier, a query target, or a workflow identifier. A content scanner (not shown) can scan the datastore for content. The metadata can be forwarded using an event streaming service 312a such as a Kafka service. In some embodiments, the results of the metadata scan are forwarded to an indexing system 318 that indexes those results to elastic search. In some embodiments, results of the content scan are forwarded to an indexing system 318 that indexes those results to elastic search.
[0142] A request for intelligence (RTI) 304 module can perform risk identification and associate that risk with each file or document. The RTI module 304 can publish messages to KafkaPage 35 of 59FOLEYHOAGUS 12449860.3DDT-00325 regarding the risk identification to the indexing system 318 that indexes the risk identifications associated with each file or document to the elastic search.
[0143] In some embodiments, a task creation service 308 can generate a query based on the type of request, where that query is sent to the query service 306. The query service 306 can issue the query to the indexed elastic search (e.g., in response to executing a task or workflow).
[0144] In some embodiments, workflows can be provided to the dashboard 314 to be displayed to user(s), such as a recommended workflow 316. The workflows can be provided in response to a file matching a storage management policy, where the workflows include one or more tasks to bring the data store into compliance with the storage management policy. The workflows and / or recommended workflow can also be provided to the indexing system 318.
[0145] Fig. 4 is a block diagram 400 illustrating a decentralized workflow according to embodiments of the present disclosure. A scanner, such as a metadata scanner 402 or a content scanner 404, determine properties of the data, such as a policy number, a total count of files, whether the data includes personally identifying information (PII), whether there is a risky or sensitive file, an age of the file, and a type of scan the data originated from to a dashboard 410 via a streaming module 408. The scanner (e.g., metadata scanner 402 or content scanner 404) also provides its results to an index that is recorded to an elastic search. In some embodiments, the query be issued to the elastic search periodically (e.g., every 15 seconds to fetch records). The dashboard 410 can receive workflows input by the user, generate them using real time intelligence, or load them from a database. For example, a workflow can include (1) calculating an alarm, (2) call a recommended workflow API to retrieve details, and (3) insert a record in postgres. As another example, the workflow can provide for a thread / scheduler running in the dashboard that checks a data store if all records are fetched and placed in the database. As yetPage 36 of 59FOLEYHOAGUS 12449860.3DDT-00325 another example, a workflow can include identifying files having a last modified date beyond a threshold and moving those identified files to an archive database. As yet another example, a workflow can include identifying files having PII, and encrypting those files.
[0146] In some embodiments, the dashboard can also provide a request to a task creation service 412. For example, the request can be to group by records based on a recommended workflow, match alarm type, or by data owner. Once completed, a message relaying a result of the requests (e.g., successful, failed, incomplete, etc.) can be returned to the dashboard 410.
[0147] In some embodiments, the dashboard 410 can also receive an alarm configuration update 414. The alarm configuration can trigger the dashboard to recalculate all alarms. For example, an alarm can be configured to In addition, it can trigger the dashboard to inform the task creation service 412 to delete pending, overdue, and new tasks.
[0148] In some embodiments, the alarm can be configured by assigning different types of security and efficiencies. For example, default alarms can be configured for security. As an example of a security alarm, if a file is accessed that is meant for a department by a single member outside of the department, an alarm can be triggered, and a follow up action can be executed. As another example, if a number of files exceeds a recommended list, an alarm can be displayed on the dashboard.
[0149] The dashboard can be configured based on the data in the elastic search. The dashboard can have alarms configured based on whether data matches a criteria of the alarm. For example, a file or a class of file may have a threshold of a number of users accessing it before it becomes risk. If more users access that file, then the alarm can be triggered. In another example, an alarm can be configured to be triggered if a percentage of files have overexposed data permissions. In some embodiments, alarms can be configured based on data redundancy criteria.Page 37 of 59FOLEYHOAGUS 12449860.3DDT-00325For example, if the data includes duplicate files above a threshold, an alarm can be triggered. In some embodiments, alarms can be configured based on other criteria, such as shared storage criteria, data sensitivity, etc.. The dashboard can identify and display alarms that are triggered. If there is an alarm, the dashboard prompts the user to take an action.
[0150] Fig. 5 is a block diagram 500 illustrating a decentralized workflow according to embodiments of the present disclosure. A scanner (e.g., a metadata scanner 502 or content scanner (not shown)) provides metadata to an indexing module that indexes the results in an elastic search. The elastic search is accessible by a real time intelligence 504 module. For example, the metadata of a policy exec identifier, a query target, and a workflow ID can be provided. A query with rules, such as if the policy exec ID’s “isFileltem” variable is false, the provided response includes an owner list, a folder list, and a workflow execution identifier, can be issued to the datastore to identify files complying with or not complying with a policy.
[0151] In some embodiments, the RTI module 504 determines whether the workflow identifier is null, and if so, it directly calculates a risk type for a file. If the identifier is not null, the RTI module 504 can check whether the received request includes a migration component. If it does, the RTI module 504 calls a workflow API and retrieves details to pass a policy to the task creation service 504. If there is no migration policy, the task creation service 504 provides a recommended workflow, such as calling a query API. The RTI module 504 can also check if the request’s “assign this workflow to a data owner” variable is true or false.
[0152] In some embodiments, the task creation service 504 can create a task for migrating data. After receiving an owners list from a query response, the task creation service 504 can iterate through the list and creates a task for each data owner. The task can include a request having a policy, an “isFileltem” variable, an owner of the file, access path, and a data action, etc. . APage 38 of 59FOLEYHOAGUS 12449860.3DDT-00325 workflow execution identifier and a workflow identifier can then be associated with the path. The task can also associate all folders to each task and its count.
[0153] For example, for a migration request, the task creation service 504 creates a task for each data owner and associates the list of folders to all tasks. For Data Share 1 having three folders for a user Sanjeev and 2 folders for a user Prativesh, then Sanjeev’s and Prativesh’s tasks will be associated with the 5 folders. The user that creates the task or is assigned the task can migrate the folder. In some embodiments, multiple users can be assigned a task, and any user who is assigned a task can initiate the task’s execution.
[0154] In some embodiments, an action card of the task creation service can be as follows:
[0155] Once created, a user interface can cause the task API to be called by the task ID and migration value. The task creation service can send an execution request to a workflow via a producer and track the status of the workflow via a consumer.
[0156] Architecture
[0157] In some embodiments, an architecture of this decentralized unstructured data management system is designed to balance centralized governance with localized execution. The architecture comprises three layers: (1) Central Application, (2) Distributed Data Engines, and (3) Dashboard and Visualization. These layers working together to manage and process data efficiently while ensuring compliance with applicable regulations.
[0158] Central Application (Kubemetes Deployment)Page 39 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0159] The Central Application can be responsible for defining and controlling rules, workflows, and compliance policies. The central application can include the ability to perform: a. Policy Management can provide centralized control of file tagging, sensitivity labelling, and compliance rules based on geography, security frameworks, and industry regulations. b. Workflow Configuration can provide definition of workflows for data mobility, remediation, and discovery, which are executed by distributed engines. c. Rules Engine can determine how data should be handled based on metadata, content scans, and dynamic compliance models. d. Thresholds and scoring can assign scores to discovered data to trigger specific workflows based on compliance risks, security, and sensitivity levels.
[0160] The central application layer operates as a single point of governance, allowing global rules to be set while delegating execution to the localized engines.
[0161] Distributed Data Engines
[0162] Distributed data engines are local processing units deployed near data sources to execute the policies and workflows defined by the central application. These distributed data engines can have close proximity to data. The distributed data engines can be distributed geographically to remain close to where the data is stored, thereby reducing latency and ensuring compliance with data sovereignty laws. The distributed data engines can provide for localized execution.Localized execution can executes data discovery, remediation, and mobility actions without needing to move the data out of its native environment. The distributed data engines can provide for compliance models. They can dynamically apply the appropriate compliance model based on the file’s attributes (tags, labels, or content). The distributed data engines can provide for Al andPage 40 of 59FOLEYHOAGUS 12449860.3DDT-00325ML model integration. Such integration can provide for automated file tagging, sensitivity labelling, and classification are performed using AI / ML models that continuously learn from custodian decisions and data patterns.
[0163] The distributed data engine layer can enable sensitive data remaining secure within its physical boundaries, while applying local rules to ensure both compliance and efficiency.
[0164] Dashboard and Visualization
[0165] The Dashboard can provide an aggregated, high-level overview of the system’s operations and offers decision-making tools for data custodians. The dashboard can provide for a. Data Visualization, which can displays metadata and insights from across the distributed data engine and highlight files that require attention for remediation or mobility. b. Recommendation Engine, which can provide statistical, ML-model, or Al model- driven suggestions for workflow actions based on file tags, sensitivity labels, and compliance scores. c. Decision Board, which can allow custodians to interact with the dashboard to select workflow paths, raise exceptions, or approve actions for specific file groups.
[0166] The dashboard layer can provide transparency, control, and insights for managing large- scale data across multiple locations while adhering to compliance standards.
[0167] Data Flow Overview
[0168] The data flow can begin with a data discovery phase. In data discovery, the distributed engines can scan files and objects for metadata and content. Based on these scans, the engines assign tags and labels, determining the sensitivity and compliance requirements.Page 41 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0169] The data flow can continue with policy application. In policy application, the central application can send the appropriate compliance model to the distributed engines. These models are dynamically selected based on file attributes, location, and security requirements.
[0170] The data flow can continue with execution. The distributed engines can execute the necessary workflow actions (e.g., data mobility, remediation) without transferring the data outside its original location.
[0171] The data flow can continue with custodian review. Data custodians can review recommendations and make final decisions through the dashboard, with the option to override suggestions and provide feedback, which feeds back into the learning model (e.g., with training, retraining, or fine-tuning).
[0172] The data flow can continue with ongoing learning. The system can continuously refine its AI / ML models based on custodian feedback and the evolving data landscape (e.g. , with training, retraining, or fine-tuning), improving future recommendations.
[0173] File Tagging and Sensitivity Labelling Mechanism
[0174] The File Tagging and Sensitivity Labelling Mechanism component can classify and manage unstructured data effectively. This mechanism can combine predefined tagging with dynamic sensitivity labelling to ensure accurate classification and compliance with various data protection standards.
[0175] Tagging Mechanism
[0176] The tagging process can involve assigning tags to files based on predefined criteria and a dynamic tagging repository. Therefore, files can be categorized according to their content, sensitivity, and compliance requirements.Page 42 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0177] Tags can be sourced from a predefined set provided out-of-the-box, covering common data attributes and compliance requirements (e.g., predefined tags). These tags are designed to standardize the classification process across various data sets.
[0178] In some embodiments, custom tags can be employed, however. In addition to predefined tags, a tagging repository can be employed to manage and store the custom tags. This repository allows for the creation of new tags and the updating of existing ones based on evolving data management needs and compliance requirements.
[0179] In some embodiments, tags are assigned to files either manually by data custodians or automatically through Al and ML models. The assignment can be based on file content, metadata, and compliance policies.
[0180] In some embodiments, assigned tags are validated to ensure they accurately represent the file's content and compliance status. Validation can includes cross-referencing with existing tags and performing consistency checks.
[0181] Sensitivity Labelling
[0182] Sensitivity labelling can be a part of the tagging mechanism. Sensitivity labelling is used to classify files based on their sensitivity and compliance requirements. For example, with object storage, object looking tags would be used. As another example, with Microsoft 365 documents, Azure Information Protection (AIP) with a Data Loss Prevention (DLP) product tags would be used. The example below describes AIP providing advanced classification and labelling capabilities.
[0183] Azure Information Protection (AIP) can be utilized to apply sensitivity labels to files. These labels can be based on a combination of content analysis and predefined sensitivity levels defined within the AIP framework. These sensitivity labels can be applied to files during thePage 43 of 59FOLEYHOAGUS 12449860.3DDT-00325 tagging process. In some embodiments, the sensitivity labels indicate the file’s sensitivity level and guide how the file should be handled, protected, and managed according to compliance policies. Sensitivity labels can be validated to ensure they accurately reflect the file's sensitivity level and compliance requirements. This includes checking the alignment of labels with the file's content and metadata.
[0184] Integration with Al and ML
[0185] Al and ML models can enhance the tagging and labeling process by automating classification and improving accuracy. Al models can analyze file content to determine appropriate tags and sensitivity labels. This analysis can include scanning for sensitive information, compliance-related data, and other relevant attributes.
[0186] ML algorithms utilize metadata to refine tagging and labeling decisions. This includes integrating file properties and contextual information to ensure accurate classification.
[0187] Al and ML models can continuously learn from tagging patterns and feedback to improve accuracy. This adaptive learning can reduce false positives and enhancing the effectiveness of the tagging and labeling mechanism.
[0188] Tagging and Labeling Workflow
[0189] The tagging and labeling process can be integrated into the broader data management workflow, influencing how files are processed and managed.
[0190] Tags and sensitivity labels can be used to determine the workflow or data pipeline through which a file is routed. This can affects the file's remediation, mobility actions, and overall management. Based on the assigned tags and sensitivity labels, tasks are generated for data custodians. These tasks include actions that can be required or recommended for compliance, remediation, or further processing.Page 44 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0191] The system can generate action recommendation(s) based on tags and labels. These recommendations can guide custodians on how to handle files in accordance with their sensitivity and compliance requirements.
[0192] Tag Management and Governance
[0193] Effective tag management and governance can ensure the accuracy and relevance of tags and labels. Policy administrators can manage predefined tags and update the tagging repository as needed. This ensures that tagging remains relevant and aligned with evolving data management requirements. Sensitivity labels can be managed and updated within the AIP framework. This can include adjusting labels based on changes in compliance requirements and sensitivity criteria. Regular audits can be conducted to verify the accuracy of tags and labels. This includes reviewing tag assignments, sensitivity labels, and compliance with governance policies.
[0194] The File Tagging and Sensitivity Labeling Mechanism can combine predefined tagging with dynamic sensitivity labeling using Azure Information Protection (AIP) to ensure accurate classification and compliance of unstructured data. By integrating Al and ML technologies, the system can enhance the efficiency and accuracy of tagging and labeling, while robust management and governance practices ensure ongoing relevance and compliance.
[0195] Decentralized Execution
[0196] The Decentralized Execution model is a core pillar of the unstructured data management system, enabling localized processing and compliance enforcement while maintaining centralized control. This architecture can leverage distributed data engines, deployed near the data’s physical location, ensuring that sensitive information is processed within compliant environments.Page 45 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0197] Distributed Data Engines
[0198] In a decentralized execution model, file discovery, tagging, labeling, mobility, and remediation can be performed by distributed data engines that are physically closer to the data source. These engines process files locally, which has several benefits. First, the actual scanning and content analysis can occur locally (e.g., within the data's compliant geography). This reduces the risk of non-compliance with data sovereignty laws (e.g., GDPR, CCPA) by ensuring that files containing sensitive personal information do not move out of their geographic boundaries. Second, distributed data engines can apply specific compliance models that are tied to the geographic location of the data. These models are chosen dynamically, based on tags and sensitivity labels associated with the files, ensuring that only compliant actions (e.g., mobility, retention, or remediation) are applied. Third, by processing data locally, the system reduces network latency and resource consumption. This optimizes the time and effort required for operations such as metadata extraction, content scanning, and remediation.
[0199] Localized Policy Enforcement
[0200] Each data engine can enforce policies and workflows that are specific to the region, domain, or security framework. These localized policies can be dynamically selected based on tags, labels, and content analysis. Policies related to data mobility, processing, and compliance can be enforced based on the context of where the data resides. For example, data in the EU may trigger a GDPR-compliant workflow, while data in the US may follow a CCPA workflow.
[0201] Each data engine can execute predefined rules and policies that are filtered by the data discovery and processing outcomes. These rules can define what actions should be taken (e.g., remediation, archiving, deletion), ensuring that actions are compliant with local governance frameworks.Page 46 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0202] Files tagged as sensitive (e.g., PII / PI data) can be processed and remediated within the localized environment. For example, sensitive health records can be handled according to HIPAA if located in the US, or PDPA if in Singapore, with mobility restricted to within compliant boundaries.
[0203] Autonomous Al Model Driven Learning
[0204] The distributed data engines can include Al and ML models that learn from the behavior of data custodians and adapt accordingly. These models can operate independently at each location, but their continuous learning feeds into the broader system:
[0205] Each data engine’s Al model observes local data patterns, tagging behavior, and custodian decisions. As the system processes more files, it learns which actions custodians prefer for different data types, improving future recommendations at a local level. In other words, in some embodiments, the local data engines can include a local Al model that is trained, updated, re-trained, or fine-tuned based on local data.
[0206] Through machine learning, the system can refine the accuracy of its tagging and labeling processes, reducing the likelihood of incorrectly flagged files. This improvement is particularly useful in compliance scenarios, where false positives can cause unnecessary disruption to business processes.
[0207] The Al models running on the data engines can adjust their behavior based on the unique compliance environment of each location, meaning that the actions recommended or automatically performed are context-sensitive and improve with each iteration.
[0208] Aggregation and Synchronization
[0209] While data engines can be decentralized and operate independently, the system ensures that aggregated results are centrally available. The results of data scans, including tags, labels,Page 47 of 59FOLEYHOAGUS 12449860.3DDT-00325 and actions taken, can be indexed centrally (e.g., in Elasticsearch). However, only metadata and analysis outcomes are centrally stored, ensuring compliance with geographic restrictions on data movement. Task generation, action recommendations, and workflow paths can be continuously synchronized between the distributed engines and the central system. This synchronization can provide for centralized monitoring and reporting, while maintaining localized execution of the workflows.
[0210] Compliance-Centric Models
[0211] As part of decentralized execution, compliance-centric models are deployed on each distributed data engine. Geography-based models can be designed to handle country or regionspecific laws (e.g. , GDPR, CCPA, PDPA) and are dynamically selected based on file tags or content analysis. For example, files tagged with “EU” can trigger GDPR-compliant workflows. Security and domain-based models can be designed around specific security standards (e.g., SOC 2, PCI DSS) or domain-specific regulations (e.g., HIPAA). The data engine can dynamically apply the relevant model to ensure that all workflows and actions meet the required standards.
[0212] The Decentralized Execution model can ensure that data is processed locally, adhering to specific compliance requirements while optimizing resources and improving overall system efficiency. With the help of distributed data engines, localized policy enforcement, and continuous Al-driven learning, the system can flexibly adapt to the unique challenges of managing unstructured data across multiple geographies, domains, and security frameworks. This approach can ensure that data custodians can make decisions within their local context, while maintaining centralized control and oversight through aggregated results and synchronized recommendations.Page 48 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0213] Referring now to Fig. 6, a schematic of an example of a computing node is shown. Computing node 10 is only one example of a suitable computing node and is not intended to suggest any limitation as to the scope of use or functionality of embodiments described herein. Regardless, computing node 10 is capable of being implemented and / or performing any of the functionality set forth hereinabove.
[0214] In computing node 10 there is a computer system / server 12, which is operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.
[0215] Computer system / server 12 may be described in the general context of computer systemexecutable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, and so on that perform particular tasks or implement particular abstract data types. Computer system / server 12 may be practiced in distributed cloud computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.Page 49 of 59FOLEYHOAGUS 12449860.3DDT-00325
[0216] As shown in Fig. 6, computer system / server 12 in computing node 10 is shown in the form of a general-purpose computing device. The components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to processor 16.
[0217] Bus 18 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Express (PCIe), and Advanced Microcontroller Bus Architecture (AMBA).
[0218] Computer system / server 12 typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system / server 12, and it includes both volatile and non-volatile media, removable and non-removable media.
[0219] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer system / server 12 may further include other removable / non-removable, volatile / non- volatile computer system storage media. By way of example only, storage system 34 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM,Page 50 of 59FOLEYHOAGUS 12449860.3DDT-00325DVD-ROM or other optical media can be provided. In such instances, each can be connected to bus 18 by one or more data media interfaces. As will be further depicted and described below, memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the disclosure.
[0220] Program / utility 40, having a set (at least one) of program modules 42, may be stored in memory 28 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data or some combination thereof, may include an implementation of a networking environment. Program modules 42 generally carry out the functions and / or methodologies of embodiments as described herein.
[0221] Computer system / server 12 may also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer system / server 12; and / or any devices (e.g., network card, modem, etc.) that enable computer system / server 12 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interfaces 22. Still yet, computer system / server 12 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via network adapter 20. As depicted, network adapter 20 communicates with the other components of computer system / server 12 via bus 18. It should be understood that although not shown, other hardware and / or software components could be used in conjunction with computer system / server 12. Examples, include, but are not limited to: microcode, device drivers,Page 51 of 59FOLEYHOAGUS 12449860.3DDT-00325 redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0222] The present disclosure may be embodied as a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0223] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0224] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to anPage 52 of 59FOLEYHOAGUS 12449860.3DDT-00325 external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0225] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readablePage 53 of 59FOLEYHOAGUS 12449860.3DDT-00325 program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0226] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0227] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0228] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, otherPage 54 of 59FOLEYHOAGUS 12449860.3DDT-00325 programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0229] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0230] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.Page 55 of 59FOLEYHOAGUS 12449860.3
Claims
DDT-00325CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: loading a storage management policy from a server; scanning a first datastore based on the storage management policy, thereby providing scan results; providing the scan results to the server; responsive to providing the scan results to the server, receiving a plurality of tasks for bringing the datastore into compliance with the storage management policy; and executing the plurality of tasks at the datastore .
2. The computer-implemented method of Claim 1, wherein scanning is based on metadata or contents of files stored in the first datastore.
3. The computer-implemented method of Claim 1, wherein receiving the plurality of tasks further comprises: receiving at least one query based on one of the plurality of tasks; providing the at least one query to a query service at the datastore.
4. The computer-implemented method of Claim 1, further comprising: displaying, in a user interface, the plurality of tasks; and receiving, in the user interface, an authorization to execute the plurality of tasks.Page 56 of 59FOLEYHOAGUS 12449860.3DDT-003255. The computer-implemented method of Claim 1, wherein the task comprises one or more workflows, each workflow comprising an action or processing to be performed on at least one file of the first datastore.
6. The computer-implemented method of Claim 1, further comprising: indexing the scan results to an elastic search database; and querying the elastic search database.
7. The computer-implemented method of Claim 1, wherein storage data policy includes a remediation policy, a quarantine policy, a migration policy, a data redundancy policy.
8. The computer-implemented method of Claim 1, wherein receiving the tasks further includes receiving the tasks from the server.
9. The computer-implemented method of Claim 1, further comprising: executing a first task of the at least one tasks on at least one file stored by the first datastore.
10. The computer-implemented method of Claim 9, further comprising: executing, at a second data store, a second task of the at least one tasks on the at least one file.
11. The computer-implemented method of Claim 1, wherein the storage management policy includes a rule and at least one tasks.
12. The computer-implemented method of Claim 1, further comprising:Page 57 of 59FOLEYHOAGUS 12449860.3DDT-00325 displaying, in a user interface, a representation of an alarm triggered based on the scan results.
13. A system comprising: a computing node comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor of the computing node to cause the processor to perform the method of any one of Claims 1-12:
14. A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing node to cause the computing node to perform the method of any one of Claims 1-12:Page 58 of 59FOLEYHOAGUS 12449860.3
Citation Information
Patent Citations
Trusted Communication Network
US20070107059A1
Intelligent management and compliance verification in distributed work flow environments
US20140222521A1
System and method for dynamic document matching and merging
US20190074072A1
Management of erasure or retention of user data stored in data stores
US20220382713A1