Federated and distributed architecture for data classification to meet data sovereignty and local data regulations
A federated and distributed data classification architecture addresses the challenge of siloed data systems by automating data classification and management, ensuring compliance and security, and empowering users to manage data effectively, thus fostering a culture of data ownership and compliance with local regulations.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-03-26
AI Technical Summary
The growing need for AI, automation, and personalization clashes with consumers' demand for control over their personal information, leading to a strategic dilemma for businesses due to siloed data systems and fragmented point solutions, causing distrust between central IT and business units and consumers.
A federated and distributed architecture for data classification that includes metadata and content analytics, automated data classification, and self-service data management tools to optimize storage, ensure compliance, and empower users to manage data effectively, while adhering to local regulations and ensuring data sovereignty.
Enables efficient data management, compliance with local regulations, and enhanced data security by automating data classification, reducing storage costs, and fostering a culture of data ownership, thereby ensuring enterprises meet their responsibility as trusted custodians of data.
Smart Images

Figure US2025047504_26032026_PF_FP_ABST
Abstract
Description
DDT-00425FEDERATED AND DISTRIBUTED ARCHITECTURE FOR DATA CLASSIFICATION TO MEET DATA SOVEREIGNTY AND LOCAL DATA REGULATIONSBACKGROUND
[0001] Today, digital trust has a larger impact than physical trust. According to recent surveys, 76% of consumers desire greater control over their data, 86% of individuals believe humans should be ultimately accountable for artificial intelligence (Al) decisions, employees with data control report increased trust in their organization's data practices, and 63% of business leaders believe data democratization is crucial for fostering a data-driven culture. For an enterprise to become a trusted custodian of data requires digital trust. Digital trust can be described as one or more of an enterprise’s commitment to data privacy, ethical Al, data sovereignty and security, and compliance with social and environmental regulations.BRIEF SUMMARY
[0002] According to embodiments of the present disclosure, methods of and computer program products for a federated and distributed architecture for data classification to meet data sovereignty and local data regulations are provided.
[0003] In some embodiments, a method comprises receiving, at a first data store having a first location, a request for data analysis. The request can be sent from a central server at a second location. In some embodiments, the method can comprise extracting text from one or more of content and metadata of files stored at the first data store. In some embodiments, the method can comprise analyzing the extracted text according to at least one parameter in the request, thereby producing an analysis. In some embodiments, the method can comprise reading a policy of the first location. The policy can include at least one rule of storing the files at the first data store.Page 1 of 43FOLEYHOAGUS13107213.1DDT-00425In some embodiments, the method can include applying the at least one rule of the policy to the analysis, thereby removing content from the analysis that is noncompliant with the at least one rule. In some embodiments, the method can include providing the analysis to the central server.
[0004] In some embodiments, the data analysis request includes data store information, an identifier of the data store, a criterion of the data store, and a user input.
[0005] In some embodiments, the data store information includes one or more of a host name and credentials to access the data store.
[0006] In some embodiments, the criterion is a filter criterion of attributes of the data store.
[0007] In some embodiments, user input is one or more of keywords, patterns, an entity present, and a sensitivity label.
[0008] In some embodiments, the entity present represents an entity present in a named entity recognition (NER) model.
[0009] In some embodiments, the request is issued using POST and Kafka or an API gateway.
[0010] In some embodiments, the method further includes adding the analysis to an elastic search database.
[0011] In some embodiments, the data analysis request includes a list of paths of files on the data store.
[0012] In some embodiments, the policy includes one or more of keywords, file patterns, and Boolean logic.
[0013] In some embodiments, removing content that is noncompliant with the policy further comprises removing values of the content from the analysis and retaining a type of information of that content.Page 2 of 43FOLEYHOAGUS13107213.1DDT-00425
[0014] In some embodiments, the content that is noncompliant with the policy is personal information, personally identifying information, or a sensitive entry.
[0015] In some embodiments, determining noncompliance is based on an alarm policy configuration.
[0016] In some embodiments, the data store is a managed document services (MDS) system.
[0017] In some embodiments, the method further includes extracting text from one or more of content and metadata of files stored at the data store.
[0018] In some embodiments, the analysis is a report and / or an aggregation of the at least one file stored at the data store.
[0019] In some embodiments, a system includes an analytics server and a first data store. The first data store can be configured to perform any of the above methods.
[0020] In some embodiments, the system can include a second data store. The second data store configured to perform any of the above methods.
[0021] In some embodiments, a system can include a computing node comprising a computer readable storage medium having program instructions embodied therewith. The program instructions executable by a processor of the computing node to cause the processor to perform any of the above methods.
[0022] In some embodiments, a computer program product comprises a computer readable storage medium having program instructions embodied therewith The program instructions are executable by a computing node to cause the computing node to perform the method of any of the above methods.Page 3 of 43FOLEYHOAGUS13107213.1DDT-00425BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0023] Fig. 1 is a flowchart illustrating a method for enterprise data management according to embodiments of the present disclosure.
[0024] Figs. 2A-C are user interfaces of an enterprise data management system according to embodiments of the present disclosure.
[0025] Fig. 3 is a diagram illustrating a federated and distributed architecture for data classification to meet data sovereignty and local data regulations according to embodiments of the present disclosure.
[0026] Fig. 4 is a network diagram illustrating a federated and distributed architecture for data classification to meet data sovereignty and local data regulations according to embodiments of the present disclosure.
[0027] Fig. 5 is a block diagram illustrating a federated and distributed architecture for data classification to meet data sovereignty and local data regulations according to embodiments of the present disclosure.
[0028] Fig. 6 is a diagram of a computing node according to an embodiment of the present disclosure.DETAILED DESCRIPTION
[0029] Decentralized identities are on the rise, with over 100 million users projected by 2025.By 2025, 75% of enterprises will need to implement Explainable Al (XAI) solutions for user trust. By 2025, 15% of large enterprises will begin transitioning to quantum-resistant cryptography to safeguard their data. By 2030, 73% of consumers will prioritize ethical data practices.Page 4 of 43FOLEYHOAGUS13107213.1DDT-00425
[0030] The growing need for Al, automation, and personalization clashes with a rising tide of citizens who demand control over their personal information. This clash creates a strategic dilemma for businesses. The lack of visibility caused by siloed data systems and fragmented point solutions creates a distrust between central IT and business units, which can also spill over to consumers. Consumers can feel a disconnect between what companies promise, what they do with data, and what users actually see.
[0031] Fig. 1 is a flowchart 100 illustrating a method for enterprise data management according to embodiments of the present disclosure. First, a data workflow made of data stores 104, data processing 106 and data action 108 is created (102). In some embodiments, a data store can be referred to as a data silo. A data store can be a database, a server having data storage capabilities, a facility of multiple databases and / or servers having data storage capabilities, etc.
[0032] In some embodiments, data, including files and objects can be scanned for metadata by a policy or storage administrator (120). The metadata can be provided to a dashboard updated with data usage analysis (e.g, for a Chief Infrastructure Officer, Chief Data Officer, or Data Owner) (126). The method can then optimize storage (e.g, for the data custodian and data owner) (128). The method can then generate an action plan based on corporate policies (e.g., for the data custodian and data owner) (130). The method can then recommend the action plan generated for approval to a user (not shown). The method can receive approval / consent of the recommended plan before executing the sequence of actions described in the plan. The method can then execute the action (e.g, archive, transform, delete, migrate) via a workflow (e.g., for the data custodian and data owner) (132). Then, the method can update the dashboard for the various personas with a reduced risk and optimized storage insights (136).Page 5 of 43FOLEYHOAGUS13107213.1DDT-00425
[0033] In some embodiments, the metadata is provide to a content scan performed by a policy administrator, data custodian, or data owner (122). The method can then generate a dashboard updated with risk exposure insights for a chief information security officer and a data owner (110). The method can then mitigate risk based on those insights (e.g., by the data custodian, data owner) (112). The method can further generate an action plan based on corporate policies to mitigate those risks (e.g., by the data custodian, data owner) (114). The method can then execute the action after receiving approval from the data owner or data custodian (e.g., manage permissions, archive, quarantine, delete) via a workflow (124). Then, the method can update the dashboard for the various personas with a reduced risk and optimized storage insights (136).
[0034] Figs. 2A-C are user interfaces of an enterprise data management system according to embodiments of the present disclosure. In Fig. 2A, the user interface illustrates a dashboard having an overall risk of the entire enterprise across all locations according to embodiments of the present disclosure. In addition, four gauges in the dashboard can display metrics of access control, data redundancy, data sensitivity, and retention compliance. An additional gauge can display a list of top sensitive data, organized by department. Yet another gauge can display statuses (e.g., pending, overdue, in progress, failed, completed) of risk mitigation tasks.
[0035] In Fig. 2B, the user interface illustrates a dashboard having an overall storage efficiency metric according to embodiments of the present disclosure. In addition, five gauges can illustrate data redundancy, cold data, orphaned data, expired data, and junk data. Another gauge can display redundant, obsolete, or trivial (ROT) data by department of the enterprise. Other gauges can display on-premise storage capacity by data store, cloud storage usage by service provider, file distribution by type, file distribution by age, file count trend, and a sustainability metric (e.g., carbon emissions).Page 6 of 43FOLEYHOAGUS13107213.1DDT-00425
[0036] In Fig. 2C, the user interface illustrates a dashboard having an overall risk score of the entire enterprise across all locations according to embodiments of the present disclosure. In addition, five gauges in the dashboard can display metrics of access control, data redundancy, data sensitivity, retention compliance, and encryption strength. Yet another gauge can display statuses (e.g., pending, overdue, in progress, failed, completed) of risk mitigation tasks. Other gauges can display sustainability metrics (e.g., by carbon emissions), top sensitive data by department, and file distribution by file type or age and by department.
[0037] It can be recognized that the gauges of Figs. 2A-C can be combined in any manner, and that the data displayed is merely exemplary and non-limiting. For example, drop down menus for each gauge can change the grouping or ranking of the data. Gauges can be moved, removed, or added to a given dashboard according to a user’s preference.
[0038] In some embodiments, systems for enterprise data management are provided. Such systems can guide businesses in orchestrating data democratization and ethical data usage within an Al-driven, climate- focused world.
[0039] In some embodiments, the systems can provide unparalleled understanding, correlation, consistency and actionability for stakeholders across an enterprise or organization. The systems can apply a nimble and scalable architecture and employ generative Al for modelling and insights and visibility across the organization via role-based access. In addition, the system can address risk, privacy, sovereignty and compliance. At the same time, the system can ensure optimized and sustainable use of storage infrastructure.
[0040] In some embodiments, the system enables people having roles ranging from C-suite executives, business, technology managers, and data owners to all utilize the same data repository. Those people can then make and execute upon informed and intelligent decisions.Page 7 of 43FOLEYHOAGUS13107213.1DDT-00425As a result, data owners can become trusted data champions, and the enterprise or organization meets its responsibility as a trusted custodian.
[0041] In some embodiments, the system can provide data orchestration and trust in four categories. Those categories can include data security and ethical usages, data transformation and sustainability, data governance and compliance, and data democratization and Al empowerment. Each category can be implemented with various aspects, including the aspects detailed below.
[0042] Data Security and Ethical Usages• Privacy Protection for Ethical Al• Risk Analysis & Remediation for Sensitive Information• Data Access Management• Data Security Orchestration, Automation & Actionability
[0043] Data Transformation & Sustainability• Hybrid-Cloud Data Mobility• Data Aggregation and Harmonization• Cloud-Optimized Data Management• Data Lifecycle Management and Footprint Reduction
[0044] Data Governance & Compliance• AI / ML-based Data Observability and Root Cause Analysis(RCA)• RBAC Driven Process & Controls• Data Usage & Traceability• Policy-based Data Governance Framework & Insight
[0045] Data Democratization & Al EmpowermentPage 8 of 43FOLEYHOAGUS13107213.1DDT-00425• Data Owner Observability, Control and Actionability• Self-Service Analytics and Insights• Data Privacy by Design• Data Wrangling and Curation for AlStorage Optimization and Lifecycle Management (LCM) across Hybrid Cloud Infrastructure
[0046] In some embodiments, a system can provide storage optimization and lifecycle management (LCM) across hybrid cloud infrastructure. The ever-expanding spread of data across hybrid cloud environments can present a challenge: storage optimization and LCM. Unmanaged data sprawl can lead to wasted resources, skyrocketing storage costs, and compliance nightmares. In some embodiments, the system can empowering organizations to streamline storage, automate data governance, and gain valuable insights.
[0047] In some embodiments, the system can include an Al-powered self-service data management software. The software can empowers enterprises by enabling users across all levels, from C-suite to data owners, to discover, define, act, transform, and audit data through a user-friendly interface. The system provides correlation, consistency, and standardization across enterprises by delivering granular insights, deriving recommended workflows, and automating actions using personalized policies and role-based access control (RBAC)-driven processes. This transformation fosters a culture of data ownership, where everyone becomes a data champion, and the organization fulfills its responsibility as a data custodian. The software can overcome these challenges through six key capabilities, strategically aligned with industry trends, to help enterprises navigate the complexities of hybrid cloud storage management. It empowers the enterprise to harness the power of the enterprise’s data while keeping storage costs and security risks under control.Page 9 of 43FOLEYHOAGUS13107213.1DDT-00425Data Discovery and Classification
[0048] In some embodiments, the system can provide automated data classification and policy- driven storage tiering. In the method illustrated by Fig. 1, data can be classified by the metadata scan 120 and identified by the content scan 122. Shadow IT, the unauthorized use of cloud services or applications, can lead to sensitive data being stored and accessed outside of an enterprise’s control, jeopardizing data sovereignty. Additionally, unstructured data that can reside in disparate locations across the global footprint of the enterprise can remain unidentified and unclassified, making it vulnerable to unauthorized access. Unclassified data can make identifying and store critical information inefficient. Manual data classification can be timeconsuming and error-prone. Additionally, enterprises can lack a unified storage strategy across on-premises and cloud environments, leading to underutilized resources and wasted storage costs.
[0049] In some embodiments, the system can provide metadata and content analytics that work together to automatically classify data based on type, sensitivity, and how often the data is accessed. The provided metadata and content analytics can further allow intelligent storage management using data tiering and placement features. For example, frequently accessed data can be automatically identified and placed on high-performance storage for optimal retrieval speed. Meanwhile, less used data can be intelligently archived to cost-effective cloud storage tiers like S3 or OneDrive. This automated approach can optimize storage utilization, reduce costs associated with underutilized resources, and strengthen data governance by ensuring sensitive data resides on secure storage tiers.
[0050] In some embodiments, the system can combine Metadata Analytics and Natural Language Processor (NLP)-powered Content Analytics for data discovery and classification.Page 10 of 43FOLEYHOAGUS13107213.1DDT-00425This combination can automatically identify / discover and classify sensitive data (e.g., PII, PHI, etc.) across a data landscape. As a result, the system can provide data exposure risks before those risks become a data leak, etc. Data owners can take control with a self-service Data Classification, allowing them to further categorize sensitive data based on their specific regulations. Discovery and risk management can be seamlessly integrated to Cloud Integrated Risk Controls (e.g., Microsoft Information Protection). Additionally, the system’s Data Usage & Traceability can provide a centralized record of all data movement activities, making audit trails and compliance reporting a breeze. In addition, the system can ensure data privacy and regulatory adherence by automating data lifecycle management tasks. Based on retention policies, the system can automatically delete or anonymize data.
[0051] For example, pharmaceutical companies can face challenges complying with GDPR regulations regarding patient data privacy in clinical trials. In some embodiments, the system can automatically identify sensitive patient information within research databases. Researchers can then leverage self-service tools to further categorize this data based on specific trial protocols. This granular control can maintain patient privacy while allowing for essential research to continue. Additionally, the system can provide a centralized record of all data access and movement, simplifying audit trails and compliance reporting. This can empower the company to mitigate data privacy risks and ensure regulatory compliance
[0052] Metadata Analytics automatically scans an enterprise’s network to identify data across repositories across all locations, including file shares, cloud storage buckets, and data lakes. Content Analytics can then analyze the content of the discovered data to pinpoint sensitive information like Personally Identifying Information (PII), Protected Health Information (PHI), and intellectual property (IP). After these steps, an inventory of data assets can be generated.Page 11 of 43FOLEYHOAGUS13107213.1DDT-00425Data Classification enables categorization of data based on sensitivity level (e.g., high-risk PII / PHI and business sensitive data like passwords or network addresses like IP or MAC addresses), regulatory requirements (e.g., HIPAA, GDPR, etc.), and specific business needs. The system enables user definition of custom classification rules, enabling tagging of all data according to the enterprise’s standards. Therefore, the system facilitates automatic risk prioritization, and allows enterprises to focus remediation efforts first on critical areas.
[0053] For example, financial institutions like banks can struggle with managing a mix of critical and less-used data. In some embodiments, the system can automatically classify regulatory documents, loan applications, and emails, placing sensitive data on secure high-performance storage while archiving less frequently accessed info to cost-effective cloud tiers. This optimizes storage utilization and reduces costs.
[0054] As an example, suppose a global bank suspects shadow IT is storing customer data (e.g., PII such as social security numbers, account numbers, account details, etc.) in unauthorized cloud locations (e.g., scattered across emails, loan applications, and internal documents). In some embodiments, the system can scan the network to identify all data repositories, and use Content Analytics to find PII such as names, social security numbers, and account information. This helps the bank remediate shadow IT practices, ensuring customer data remains controlled and compliant with data sovereignty regulations.
[0055] As yet another example, financial institutions can struggle with the volume customer data (e.g., loan applications, account statements, transaction records). Manually sorting through this data is a time-consuming and expensive task. However, the system’s Al-powered analytics automatically classify the data, saving the bank time and money. Less-frequently accessed data can be archived to cost-effective storage tiers, freeing up valuable space for real-time transactionPage 12 of 43FOLEYHOAGUS13107213.1DDT-00425 data. This allows the bank to focus on what matters most: delivering exceptional customer service.
[0056] Preparing data for Al and machine learning projects can be time-consuming and complex. Data scientists can struggle to find, access, and clean data from diverse sources across the organization. In some embodiments, the system’s Data Aggregation and Harmonization can consolidate data from file storage, object storage, and data lakes, eliminating time-consuming manual collection. Data owners and data scientists can work together seamlessly using self- service tools to define granular access permissions, keeping data secure and compliant to regulations and other rules or policies. The system can initiate data migration and even transform it (e.g., file to object) for an even smoother workflow. Automated data cleansing workflows ensure pristine data quality by removing inconsistencies and errors that could reduce the accuracy of AI / ML models. This translates to streamlined data preparation, faster AI / ML project development, and unlocking the true potential of data for advanced analytics.Data Redundancy Elimination and Archival
[0057] In some embodiments, the system can provide Data Redundancy Elimination and Archival for Improved Efficiency. Data sprawl can lead to redundant, obsolete, and trivial (ROT) data consuming valuable storage space. Manual identification and deletion of redundant data is cumbersome and error-prone. Traditional archiving processes can be complex and hinder data accessibility. In the method illustrated by Fig. 1, data redundancy elimination and archival can be performed by the actions executed by the workflow in steps 124 and 132, including archival, deletion, migration, quarantining, etc.
[0058] In some embodiments, the system can find hidden data and help manage that data efficiently. With Data Redundancy, an enterprise can identify and eliminate duplicate dataPage 13 of 43FOLEYHOAGUS13107213.1DDT-00425 through intelligent classification and tagging. This can free valuable storage space for active data. In addition, cleaning up such data does not remove access to it. The system can intelligently converts redundant files into space-saving object formats (e.g., for archiving). These archived objects can still be easily retrieved using software, a RESTful API (e.g., for audits and legal discovery). Plus, the system provides S3 / NFS / SMB-compliant access, ensuring compatibility with existing infrastructure and making data retrieval a breeze.
[0059] For example, a pharmaceutical enterprise can have redundant clinical trial data, which wastes storage space for active research. The above system can identify and eliminate duplicate data sets based on intelligent tags, freeing up valuable storage space for ongoing research projects and improving data accessibility for scientists.Automated Data Lifecycle Management
[0060] In some embodiments, the system can provide Automated Data Lifecycle Management with Retention Compliance. Manually managing data retention policies across hybrid cloud environments can be complex and error-prone. Inconsistent data handling practices and manual enforcement of data policies can lead to security vulnerabilities and compliance gaps. Non- compliance with data retention regulations can sometimes lead to hefty fines. Therefore, risk mitigation and regulatory compliance can be desired.
[0061] In some embodiments, the system provides Retention Compliance that seamlessly integrates the metadata analysis with data policy creation and workflow automation. In relation to the method illustrated by Fig. 1, workflows can be automated by the creation of a data workflow by step 102 and the executing action via workflow 124 or 132. This powerful combination can allow enterprises to discover, classify, and tag data based on specific retention requirements. The software can automate data lifecycle management tasks based on pre-definedPage 14 of 43FOLEYHOAGUS13107213.1DDT-00425 policies. The system can handle deletion, archiving, or migration, thereby ensuring effortless compliance with data retention regulations. Therefore, the system minimizes risks associated with non-compliance and streamlines storage management by eliminating the need for manual tasks.
[0062] For example, hospitals can face challenges ensuring compliance with complex healthcare data privacy regulations. In some embodiments, the system’s Retention Compliance utilizes metadata analysis to classify patient medical records and automate data lifecycle management based on retention requirements. This can ensure HIPAA compliance, minimizes regulatory risks, and streamlines storage management.
[0063] Without a clear data governance framework, data transfers across borders can be haphazard, potentially violating data sovereignty regulations. Inconsistent data handling practices can also lead to confusion and difficulty in demonstrating compliance.
[0064] In some embodiments, the system seamlessly integrates with an existing data governance framework to establish data policies that uphold data sovereignty best practices. The system’s Data Policy Creation and Workflow functionality simplifies defining clear rules for data movement and usage. These policies can be tailored based on data sensitivity, location, and retention requirements for different data types. Furthermore, the system gives data owners control over their data through Self-Service Data Classifications and Migrations. This feature streamlines data management by allowing data owners to classify and migrate data according to established policies. The system can then automates the data migration process, ensuring data reaches the appropriate storage locations while adhering to data sovereignty regulations. This not only simplifies governance but also reduces the risk of errors.Page 15 of 43FOLEYHOAGUS13107213.1DDT-00425
[0065] For example, a multinational corporation with worldwide offices collects customer data from various regions. In some embodiments, the system can define clear policies for data movement and usage that comply with data sovereignty regulations in each country they operate. The system’s Self-Service Data Classifications and Migrations can empower business units to classify and migrate their data according to these established policies. Therefore, consistent data handling practices compliance with data sovereignty regulations across the global organization are ensured.
[0066] As another example, hospitals are under constant pressure to comply with strict regulations regarding patient data retention. Manual processes for managing patient records can be error-prone and lead to hefty fines. In some embodiments, the system can automate data retention tasks, ensuring hospitals stay compliant and avoid costly penalties. Additionally, the system can provide complete transparency into how patient data is used, fostering trust with patients and regulators. This allows hospitals to focus on their core mission: delivering quality patient care.
[0067] In some embodiments, the system can ensure consistent data handling across an enterprise with the system’s Data Policy Creation and Workflow. The Data police can clearly define comprehensive data policies that translate into automated workflows. These workflows can enforce access controls, data encryption, and other security measures, minimize human error, and guarantee consistent policy application. These workflows work with the results of Data Discovery and Classification. By classifying data based on sensitivity and regulations, the system can use data policies and workflows tailored to the specific needs of the data as it is classified. This can streamline data governance, simplify compliance, and strengthen data security.Page 16 of 43FOLEYHOAGUS13107213.1DDT-00425
[0068] For example, consider an insurance company that implements new data governance policies to comply with industry regulations regarding data privacy and security. The system’s Data Policy Creation and Workflow functionality can empower them to personalize and translate these policies into automated workflows. These workflows can include automatic data discovery and classification for sensitive research data, access management, migration, retention compliance, archival, deletion, and more. Automating these processes can minimize human error and ensures consistent enforcement of data governance policies across the organization.Self-service data classification
[0069] In some embodiments, the system provides Self-Service Data Classification and Migration for Democratized Data Management. Complex data management processes and limited visibility into data storage locations can hinder business users from effectively managing their data. This can lead to data stores and hinder collaboration efforts. Alternative data governance and migration point solutions, often complex and inaccessible, can create data stores, hinder collaboration, and limit the value businesses can extract from their information.Enterprises can move forward from a “lift and shift” approach and embrace a “data-driven” strategy for infrastructure management. In some embodiments, in the method illustrated by Fig. 1, data can be classified by an enterprise or a user providing rules to the metadata scan 120 or content scan 122.
[0070] In some embodiments, the system’s Data Dynamics' Al-powered self-service data management software can bring a fresh approach to privacy, security, compliance, governance and optimization in the world of Al-led workloads.
[0071] In some embodiments, the system can provide business users the ability to take control of their data with Self-Service Data Classifications and Migrations. This intuitive functionality canPage 17 of 43FOLEYHOAGUS13107213.1DDT-00425 allow these business users to classify data based on their department's specific needs and business context. User-friendly tools can enable business users to initiate data migrations across on-premises and cloud storage effortlessly, fostering better collaboration and data governance. Improved data visibility through self-service tools can empower users to take ownership of their data and manage it effectively. Use of these self-service tools can dismantle data stores and can facilitate seamless collaboration across departments, ensuring everyone has access to the information they need.
[0072] For example, insurance companies can have trouble collaborating with data stores. In some embodiments, the system’s Self-Service Data Classifications can empower underwriters and claims adjusters to classify data by insurance product or policyholder. They can initiate data migrations between on-premises and cloud storage, fostering collaboration across departments and improving data governance.
[0073] This solution brief explores sixkeyuse casesthathighlighthow Zubin addresses critical industry challenges, such as data stores and inefficient data wrangling. Wewill delve into the software’s role in empowering organizations to automate data classification, streamline migrations, andbreakdowndata stores. By enabling self-servicedataownershipandfosteringcollaboration, Zubin paves the way for a future where data is readily available for analysis, leading to better decision-making and improved businessoutcomes.Risk Insights
[0074] In some embodiments, the system can provide data usage and traceability with risk insights for enhanced security. Static data security measures may not be sufficient to address evolving threats and vulnerabilities. Lack of visibility into data usage across hybrid cloud environments can make it difficult to identify and mitigate security risks. Unclear dataPage 18 of 43FOLEYHOAGUS13107213.1DDT-00425 ownership can lead to confusion about access control responsibilities. Varying data privacy regulations across regions can sometimes require data localization policies. Without proper risk assessment and data localization strategies, an enterprise may be inadvertently violating data sovereignty laws. In some embodiments, in the method illustrated by Fig. 1, the dashboard 110 illustrates risk insights based on the analysis described herein.
[0075] Tracking and maintaining a clear audit trail for data movement across borders is essential for demonstrating compliance with data sovereignty regulations. Without this visibility, regulators may question data handling practices and ability to ensure data remains within desired jurisdictions.
[0076] In some embodiments, the system can help locate data and help secure it. The system’s Data Usage and Traceability feature can provide a central hub or centralized data index for all data movement activities, giving complete visibility of an enterprise’s data. This central hub can be implemented with a dashboard to show data owners how their data is being used. It meticulously tracks every transfer, capturing the source, destination, purpose, and timestamp, along with the user or system responsible. This comprehensive audit trail enables clear and responsible data movement practices. It simplifies compliance with data residency requirements by providing a clear picture of where data resides and how the data is being used. This not only fosters trust but also ensures the enterprise have the information needed to manage the enterprise’s data effectively.
[0077] With Risk Exposure Insights, a powerful analytics tool, an enterprise can identify potential security threats by analyzing data usage patterns. Risk Exposure Insights can identify exposed data and prioritize threats to address first. Risk Exposure Insights can leverage data classification and advanced content analytics powered by Al and Machine Learning (AI / ML) toPage 19 of 43FOLEYHOAGUS13107213.1DDT-00425 assess the true risk level of exposed data. Beyond basic risk scoring, Risk Exposure Insights can consider factors such as the type of data exposed and the likelihood of exploitation based on realtime threat intelligence. This allows enterprises to focus remediation efforts on the most critical threats, ensuring the enterprise addresses sensitive data at a highest risk of misuse first.
[0078] For example, consider an energy company that suspects a cyberattack on its industrial control systems (ICS) managing power grids, potentially compromising sensor data, configuration files, and employee credentials. In some embodiments, the system can analyze the data access patterns and can classify those patterns based on sensitivity assigning the highest risk to configuration files due to the potential for widespread power outages. This allows the enterprise to prioritize securing the ICS and critical infrastructure, then address employee credentials and sensor data, minimizing power grid disruption.
[0079] Further strengthening an enterprise’s defenses, RBAC Down to the Data Owner Layer provides many levels of access control. This system allows for granular permission management based on who owns the data, ensuring only authorized users can access it. RBAC can access data at the data owner layer. Therefore, the system seamlessly integrates with existing access control features to provide unparalleled granular control. As a result, an enterprise can define which users can access specific data and at what level they can access it (e.g., read, write, edit, read and write, etc.). The system can enforce the principle of least privilege, ensuring users only have the access permissions necessary for their tasks. This significantly reduces the points of attack for a breach, minimizing the potential damage caused by compromised credentials or malicious insiders.
[0080] As an example, consider a pharmaceutical company that develops a new drug. Sensitive research data (e.g., formulas, clinical trial results, etc.) can be shared with collaborators whilePage 20 of 43FOLEYHOAGUS13107213.1DDT-00425 maintaining strict access controls. Pharmaceutical companies also generate vast amounts of data during drug development, but much of that data is hidden and unused. The system’s RBAC Down to the Data Owner Layer can empower researchers to define granular access permissions for collaborators, ensuring only authorized personnel can access specific data elements. This reduces the attack surface and safeguards valuable intellectual property. The system can further shed light on this “dark data” (e.g. , the hidden or unused data) by providing a centralized view of all research data. Data owners, like research scientists, can see exactly who has accessed their data and for what purpose. This transparency fosters trust and accountability within the research team. Additionally, the system can empower researchers and scientists to identify potentially valuable data that can be used for further analysis, accelerating drug discovery and development.
[0081] As another example, data governance for insurance companies has been a complex task accessible only to IT specialists. This often led to data stores and hindered collaboration between departments. In some embodiments, the system empowers non-IT users, such as underwriters and claims adjusters, to take charge of their data. User-friendly tools allow users to classify data based on risk factors and claim types, fostering a culture of data ownership. Additionally, granular access controls ensure data security and compliance. This breakdown of data stores encourages collaboration and empowers business users to make faster, more informed decisions.
[0082] This comprehensive approach creates a multi-layered shield for the enterprise’s data. The enterprise can have clear visibility into movement, the ability to identify potential risks through usage analysis, and ironclad access controls thanks to granular ownership permissions. This can significantly minimize points of attack and prevent data breaches, thereby keeping the enterprise’s information safe.Page 21 of 43FOLEYHOAGUS13107213.1DDT-00425
[0083] In some embodiments, the system employs content analytics to classify data and assess the risk associated with where the data is stored and processed. This analysis considers factors including data sensitivity, applicable regulations, and the security strength of storage locations. In some embodiments, the system allows prioritization of risks, developing targeted mitigation strategies, and defining policies to automatically route specific data types (e.g., PII) to designated storage locations within specific regions. This ensures compliance with data residency requirements. In addition, in some embodiments, the system can integrate with leading cloud storage providers like Microsoft Azure to effortlessly enforce data residency policies.
[0084] As another example, consider an insurance company that operates multiple European countries. The European Union’s (EU) General Data Protection Regulation (GDPR) mandates that citizen data be stored within the EU. Manual processes for isolating sensitive data can be slow and inefficient, which can potentially allowing breaches. In some embodiments, the system can automatically discover and classify customer data (e.g., health information) spread across various locations and assess risks associated with it. Data Containment and Isolation can then allow the enterprise to define policies to automatically route sensitive customer data (e.g., health records) to designated storage locations within the EU, ensuring compliance with GDPR and data sovereignty regulations. This feature can automate workflows to quarantine high-risk sensitive data the moment it is detected. By restricting access and preventing dissemination, Data Containment minimizes potential damage from exposed data. Furthermore, the Data Encryption adds another layer of defense by safeguarding sensitive data at rest and in transit. This ensures data remains protected throughout the data’s lifecycle.
[0085] As yet another example, consider a hospital chain with facilities in the US and Canada that needs to comply with HIPAA regulations, which requires strict controls on patient dataPage 22 of 43FOLEYHOAGUS13107213.1DDT-00425 movement, or other regulations such as the California Consumer Privacy Act (CCPA). The system’s Data Usage & Traceability provides a comprehensive audit trail for all patient data transfers, including the source, destination, and time of each transfer. This allows the hospital chain to show to regulators that patient data is only transferred across borders for legitimate medical purposes and in accordance with HIPAA guidelines.
[0086] For example, unauthorized access to energy data pipelines can be disastrous. In some embodiments, the system’s Data Usage & Traceability provides a centralized record of data movement activities. Combined with Risk Exposure Insights, it can analyze data usage patterns to identify suspicious access attempts. Additionally, RBAC Down to the Data Owner Layer can strengthen access controls, preventing unauthorized access and minimizing the attack surface for data breaches.
[0087] As yet another example, consider a disgruntled employee with access to a hospital’s radiology department. That disgruntled employee, in this example, attempts to steal patient data. In some embodiments, the system’s Data Classification can be configured to identify and tag patient medical images. Data Containment and Isolation can then automatically quarantine these files, preventing any such rogue access (e.g., by the disgruntled employee) and minimizing potential patient privacy violations. The hospital can then investigate the attempted incident, revoke the employee’s access, and restore the quarantined data from secure backups.Data Observability and Root Cause Analysis
[0088] In some embodiments, the system can provide Data Observability & Root Cause Analysis for Performance Optimization. Performance bottlenecks and data quality issues can hinder critical applications and analytics across hybrid cloud environments. Identifying the root cause of these issues can be time-consuming and complex. In addition, monitoring data usagePage 23 of 43FOLEYHOAGUS13107213.1DDT-00425 and detecting anomalies across a global data footprint can be complex. Without this data observability, an enterprise may not be alerted to unauthorized access or potential data breaches that could compromise data sovereignty. In some embodiments, in the method illustrated by Fig. 1, the dashboard 126 provides data usage analysis that can be used for performance optimization 128.
[0089] The system’s Data Observability and Root Cause Analysis functionality can utilize metadata & content analytics powered by AI / ML to continuously monitor data usage within pipelines across the enterprise’s global footprint. This includes analyzing ROT, user access patterns, data transfer activities, and changes to data permissions. By identifying the anomalies, organizations can take proactive steps to optimize data storage performance and ensure data quality for downstream analytics. Additionally, the system can provide executive dashboards and reports providing risk and data usage analysis across the enterprise, business units, LOBs, and geographies, complemented by action plans and real-time status updates. The system can provide aggregate data risk reporting and data usage analysis by teams and individual data owners via an intuitive dashboard, thereby enabling visibility, analysis and actionability, with real-time status updates and reminders.
[0090] For example, manufacturing companies rely on real-time data for efficient operations. In some embodiments, the system’s Data Observability utilizes machine learning to continuously monitor data pipelines, detecting anomalies or inconsistencies in sensor data readings. By identifying the root cause, manufacturers can optimize data storage performance and ensure data quality for downstream analytics tasks like predictive maintenance. This improves overall data efficiency and avoids production slowdowns.Page 24 of 43FOLEYHOAGUS13107213.1DDT-00425
[0091] The system’s Data Observability and Root Cause Analysis can employ AI / ML, metadata, and content analytics to continuously monitor data pipelines globally. This includes analyzing user access patterns, data transfer activities, and changes to data permissions. The system can further provide dashboards and reports including insightful risk and data usage analysis across the entire enterprise, from business units to individual teams and locations. The reports can include action plans and real-time status updates to keep everyone in the enterprise informed and moving forward. The system’s dashboard can aggregate data risk reporting and usage analysis for teams and data owners. This translates to clear visibility, actionable insights, and real-time updates and reminders, ensuring the enterprise’s data stays secure and compliant.
[0092] For example, a common challenge faced by banks can be building accurate credit risk assessment models. If the training data used for these models contains errors, such as incorrect income information or missing loan repayment history, that model may misclassify borrowers. In some embodiments, the system’s Data Observability & Root Cause Analysis utilizes machine learning to continuously monitor data quality within pipelines and pinpoint anomalies like ROT, unprotected sensitive business data or unauthorized access control across geographies, business units, and LOBs. This allows them to identify and rectify errors in the training data, leading to more accurate and reliable credit risk models.
[0093] Organizations operating under high compliance and governance models, such as BFSI, healthcare, and legal sectors, often have a global presence. This results in siloed repositories of unstructured data, complicating data management and governance. These organizations can benefit from a comprehensive view of their data ecosystem for effective data discovery, mobility, and mitigation, all while adhering to stringent data sovereignty and local regulations.Page 25 of 43FOLEYHOAGUS13107213.1DDT-00425
[0094] Alternative solutions can include centralized data processing, which can conflict with regulations that mandate data remain within specific geographical boundaries or comply with specific domain and infrastructure guidelines. In these alternative solutions, centralized data processing can lead to non-compliance with data sovereignty laws. In addition, in these alternative solutions, the lack of localized processing increases the risk of data breaches and unauthorized access. Further, in these alternative solutions, inefficiencies and delays due to the necessity of moving large amounts of data across regions. Last, in these alternative solutions, inadequate support for diverse and dynamic compliance requirements across different regions and domains.
[0095] In some embodiments, a system that provides for localized data processing within data stores while providing a complete view of the enterprise’s data across multiple locations is desired. This approach ensures compliance with local regulations and provides the necessary global oversight for auditors from international and local governance bodies.
[0096] In some embodiments, the advantages over known solutions include compliance being insured by processing data locally and adhering to diverse data sovereignty and local regulations. The advantages also can include enhanced security because localized processing reduces the risk of data breaches and unauthorized data transfers. The advantages also can include improved efficiency because proximity to data stores improves processing efficiency, reduces latency, and can additionally save energy. The advantages further include preventing deduplication for processing. The advantages can further include comprehensive reporting including central aggregation of de-identified data allows for detailed, organization-wide reporting on key metrics. The advantages can further include scalability and flexibility, including having the federated architecture be adaptable to various regulatory environments, suitable for global organizations.Page 26 of 43FOLEYHOAGUS13107213.1DDT-00425In some embodiments, the system bridges the gap between localized data processing requirements and the need for a unified organizational view, offering a robust solution for modem data governance challenges.
[0097] In some embodiments, the system provides federated deployment. Data engines can be strategically deployed close to data stores based on geographical, domain, and infrastructure compliance requirements. However, these engines can process data locally, ensuring that data does not leave the compliant boundaries of the region, domain, or infrastructure.
[0098] In some embodiments, the system provides a federated and distributed architecture for data classification. This data classification can enable enterprises to comply with data sovereignty and local regulations and maintain a unified view of their global data ecosystem. The system can perform all data processing and actions locally within data stores and can aggregate de-identified (e.g., anonymized) data for centralized reporting.
[0099] In some embodiments, the system provides federated deployment of data engines. A data engine is a server, computer, or other processing entity that processes data (e.g., from a data store). Data engines can be deployed close to data stores, which ensures compliance with local regulations by processing data locally.
[0100] In some embodiments, deployment can strategically consider compliance. Deployment strategies can consider various compliance requirements, such as geographical regulations (e.g., GDPR, CCPA, PDPA, DPDP), domain-specific regulations (e.g., HIPAA), and infrastructure constraints (e.g., data should not be transferred out of firewall-protected locations).
[0101] In some embodiments, an Al inference engine can also be employed. The Al inference engine can employ localized data processing. For example, the Al inference engine can performPage 27 of 43FOLEYHOAGUS13107213.1DDT-00425 data discovery. The system can perform deep content scans each data store (e.g., data store) to identify and classify data within each data store based on content, context, and metadata.
[0102] In some embodiments, the Al inference engine can ensure secure data mobility within compliant boundaries, thereby facilitating efficient data movement without violating local regulations. In some embodiments, the system facilitates the movement of data within compliant boundaries, ensuring secure and efficient data transfer.
[0103] In some embodiments, the Al inference engine can implement risk mitigation measures in accordance with local regulatory requirements and organizational policies. In some embodiments, the system can apply localized risk mitigation strategies to protect data according to local regulations and organizational policies.
[0104] In some embodiments, the system can perform centralized aggregation and reporting.For example, the system can perform de-identified data aggregation. Data can be de-identified using local processing that is in accordance with local rules and regulations. Once it is de- identified according with those rules and cleared for central storage as de-identified or anonymized data, it can be sent to a central analytical database. Such de-identification can be performed in all data stores according to each data store’s local rules, and then aggregated to a central storage after such de-identification. This can provide a central analytics database with a de-identified set of data representing the entire or most of the enterprise that also complies with local rules and regulations. It can present a comprehensive view of the organization's data landscape while also being compliant.
[0105] In some embodiments, the system can provide real-time or near real-time reporting risk, security, privacy, compliance, and infrastructure status, accessible to data stewards (CISO, CIO, CDO offices). The central database can provide near real-time reporting on risk, security,Page 28 of 43FOLEYHOAGUS13107213.1DDT-00425 privacy, compliance, and infrastructure metrics. Data stewards, including members of CISO, CIO, and CDO offices, can access detailed reports to ensure compliance and inform decisionmaking.
[0106] Fig. 3 is a diagram 300 illustrating a federated and distributed architecture for data classification to meet data sovereignty and local data regulations according to embodiments of the present disclosure. A central data classification analytics server 302 is operatively connected to data stores 304a-d. As can be seen by Fig. 3, data store 304a exists in California and is holds California State Insurance Personal Data and California State Health Care Personal Data. Data store 304b exists in Texas and holds Texas State Insurance Personal Data, Texas State Health Care Personal Data, and Texas State Banking Personal Data. Data store 304c exists in Germany, and holds the Germany Citizen Insurance Personal Data and the Germany Citizen Health Care Personal Data. Data store 304d exists in the United Kingdom and includes United Kingdom Insurance Personal Data, United Kingdom Health Care Personal Data, and United Kingdom Banking Personal Data. Each set of data can be de-identified or anonymized according to the respective laws of the data store 304a-d. For example, data store 304a complies with USA and California law, data store 304b complies with USA and Texas law, data store 304c complies with EU and German law, and data store 304d complies with EU and UK law. Once deidentification is complete, that de-identified data can be sent to the central data classification analytics server 302.
[0107] It can be recognized that in de-identifying the data, the type of the Pl / PII / Sensitive entity found is maintained. In other words, metadata regarding the type of entry removed can be maintained in the file(s). However, the exact value of the Pl / PII / sensitive entry, such as the name of the person, social security number (SSN), internet protocol (IP) address, etc. arePage 29 of 43FOLEYHOAGUS13107213.1DDT-00425 removed / not retained to avoid creating another point of breach. This also includes the assigned sensitivity value configured by the user during policy definition. For example, a SSN of 123-45- 6789 in a file can be denoted with a tag of “SSN,” and metadata noting a count of occurrences of SSNs in that file can be appended, while removing the actual value of the SSN of “123-45- 6789”.
[0108] Fig. 4 is a network diagram 400 illustrating a federated and distributed architecture for data classification to meet data sovereignty and local data regulations according to embodiments of the present disclosure. A user layer 402 (e.g., a user interface receiving instructions from a user remotely or directly) can instruct a control panel to install a universal data engine 406 for one or more region(s). The control panel can then allocate a group name and install that the universal data engine 406 region-wise. The user layer 402 can then instruct the control panel 404 to install a file content parsing engine 408 for one or more region(s). The control panel can then allocate a group name and install the file content parsing engine 408 region-wise. The user layer 402 can then instruct the control panel 404 to install a file content analysis engine 410 for one or more region(s). The control panel then can install the file content analysis engine 410 region-wise.
[0109] In some embodiments, during the installation of data engines, a unique group name is allocated to all engines being installed. Universal Data Engines are installed on a region-wise basis, including components for parsing and text analysis involving data science. The data engine group name can be unique across regions. Region-specific components are configured to listen to the pattern of the group name. This ensures that region-specific consumers receive the desired requests pertinent to their region.Page 30 of 43FOLEYHOAGUS13107213.1DDT-00425
[0110] In some embodiments, the user layer 402 can instruct the control panel 404 to perform a metadata scan, and the user layer 402 can instruct the control panel 404 to perform a content scan. The user layer 402 can provide parameters in its instructions that define what to scan for in the file content. In response, the control panel can initiate the metadata scan and the content scan at the file content parsing engine 408. File content parsing and file content analysis happen within the region. The file content parsing engine can retrieve a raw file stream from a network attached storage (NAS) 414. In response, the NAS 414 then can return the raw file stream to the file content parsing stream 410. The file content parsing engine 410 can direct a text extraction layer 416 to extract text from the raw file stream, and the text extraction layer 416 returns the extracted text to the file content parsing engine 410. The file content parsing engine 410 can direct a text analysis layer 418 to analyze the extracted text, and the text analysis layer 418 returns results of the analysis to the file content parsing engine 410. Those analysis results are then stored at the central database 420.
[0111] In some embodiments, the control panel 404 can instruct the universal data engine 406 to index data category names (e.g., social security numbers, etc.). In response, the universal data engine 406 can store the index category names at a central database 420. By indexing the category names, the data itself stays in the region, but the category names and the analysis are stored in the central database 420.
[0112] In some embodiments, the user layer 402 can request to review analysis results via the control panel 404. The control panel 404 can retrieve the analysis results from the central database 420, and the central database 420 can return them to the control panel 404. The control panel 404 can then display those results, for example, to the user layer 402.Page 31 of 43FOLEYHOAGUS13107213.1DDT-00425
[0113] Fig. 5 is a block diagram 500 illustrating a federated and distributed architecture for data classification to meet data sovereignty and local data regulations according to embodiments of the present disclosure. A data store 502 of Region 1 and data store 504 of Region 2 each store respective data in a local data store. Each data store 502 and 504 comprises universal data engine, a parser, a keyword service, and a pattern service for local analysis. In some embodiments, a control panel 506 can install the universal data engines for each data store 502 and 504. In some embodiments, a control panel 506 can install content parsing engines (e.g., the parser) for each data store 502 and 504.
[0114] In some embodiments, the steps of the network diagram illustrated by Fig. 4 can be implemented in the system illustrated by Fig. 5. With reference to Fig. 5, the content parsing engine can perform metadata scans and content scans on a raw file stream provided by the local PG and provide analysis of that raw file stream using a respective keyword service and pattern service. The results of those analyses can be stored in the central database 510. The universal data engine 508a-b can provide indexed category names to that can also be stored in the central database.
[0115] Referring now to Fig. 6, a schematic of an example of a computing node is shown. Computing node 10 is only one example of a suitable computing node and is not intended to suggest any limitation as to the scope of use or functionality of embodiments described herein. Regardless, computing node 10 is capable of being implemented and / or performing any of the functionality set forth hereinabove.
[0116] In computing node 10 there is a computer system / server 12, which is operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / orPage 32 of 43FOLEYHOAGUS13107213.1DDT-00425 configurations that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.
[0117] Computer system / server 12 may be described in the general context of computer systemexecutable instructions, such as program modules, being executed by a computer system.Generally, program modules may include routines, programs, objects, components, logic, data structures, and so on that perform particular tasks or implement particular abstract data types. Computer system / server 12 may be practiced in distributed cloud computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.
[0118] As shown in Fig. 6, computer system / server 12 in computing node 10 is shown in the form of a general-purpose computing device. The components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to processor 16.
[0119] Bus 18 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and notPage 33 of 43FOLEYHOAGUS13107213.1DDT-00425 limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Express (PCIe), and Advanced Microcontroller Bus Architecture (AMBA).
[0120] Computer system / server 12 typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system / server 12, and it includes both volatile and non-volatile media, removable and non-removable media.
[0121] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer system / server 12 may further include other removable / non-removable, volatile / non- volatile computer system storage media. By way of example only, storage system 34 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media can be provided. In such instances, each can be connected to bus 18 by one or more data media interfaces. As will be further depicted and described below, memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the disclosure.
[0122] Program / utility 40, having a set (at least one) of program modules 42, may be stored in memory 28 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data or somePage 34 of 43FOLEYHOAGUS13107213.1DDT-00425 combination thereof, may include an implementation of a networking environment. Program modules 42 generally carry out the functions and / or methodologies of embodiments as described herein.
[0123] Computer system / server 12 may also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer system / server 12; and / or any devices (e.g., network card, modem, etc.) that enable computer system / server 12 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interfaces 22. Still yet, computer system / server 12 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via network adapter 20. As depicted, network adapter 20 communicates with the other components of computer system / server 12 via bus 18. It should be understood that although not shown, other hardware and / or software components could be used in conjunction with computer system / server 12. Examples, include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0124] The present disclosure may be embodied as a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0125] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storagePage 35 of 43FOLEYHOAGUS13107213.1DDT-00425 device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0126] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0127] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machinePage 36 of 43FOLEYHOAGUS13107213.1DDT-00425 instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0128] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0129] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processingPage 37 of 43FOLEYHOAGUS13107213.1DDT-00425 apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0130] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0131] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order,Page 38 of 43FOLEYHOAGUS13107213.1DDT-00425 depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0132] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.Page 39 of 43FOLEYHOAGUS13107213.1
Claims
DDT-00425CLAIMSWhat is claimed is:
1. A method comprising: receiving a data analysis request from an analytics server, at a data store, the data store having a location; reading a policy associated with the location; based on the data analysis request, analyzing at least one file stored at the data store, thereby producing an analysis; removing content from the analysis that is noncompliant with the policy; and providing the analysis to the analytics server.
2. The method of Claim 1, wherein the data analysis request includes data store information, an identifier of the data store, a criterion of the data store, and a user input.
3. The method of Claim 2, wherein the data store information includes one or more of a host name and credentials to access the data store.
4. The method of Claim 2, wherein the criterion is a filter criterion of attributes of the data store.
5. The method of Claim 2, wherein user input is one or more of keywords, patterns, an entity present, and a sensitivity label.
6. The method of Claim 5, wherein the entity present represents an entity present in a named entity recognition (NER) model.Page 40 of 43FOLEYHOAGUS13107213.1DDT-004257. The method of Claim 1, wherein the request is issued using POST and Kafka or an API gateway.
8. The method of Claim 1, further comprising adding the analysis to an elastic search database.
9. The method of Claim 1, wherein the data analysis request includes a list of paths of files on the data store.
10. The method of Claim 1, wherein the policy includes one or more of keywords, file patterns, and Boolean logic.
11. The method of Claim 1 , wherein removing content that is noncompliant with the policy further comprises removing values of the content from the analysis and retaining a type of information of that content.
12. The method of Claim 11 , wherein the content that is noncompliant with the policy is personal information, personally identifying information, or a sensitive entry.
13. The method of Claim 1, wherein determining noncompliance is based on an alarm policy configuration.
14. The method of Claim 1, wherein the data store is a managed document services (MDS) system.
15. The method of Claim 1, further comprising extracting text from one or more of content and metadata of files stored at the data store.Page 41 of 43FOLEYHOAGUS13107213.1DDT-0042516. The method of Claim 1, wherein the analysis is a report and / or an aggregation of the at least one file stored at the data store.
17. A system comprising : an analytics server; and a first data store, the first data store configured to perform the method of any one of claims 1-16.
18. The system of claim 17, further comprising: a second data store, the second data store configured to perform the method of any one of Claims 1-16.
19. A system comprising: a computing node comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor of the computing node to cause the processor to perform the method of any one of Claims 1-16:
20. A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing node to cause the computing node to perform the method of any one of Claims 1-16:Page 42 of 43FOLEYHOAGUS13107213.1
Citation Information
Patent Citations
Systems and methods for identifying and mapping sensitive data on an enterprise
US20180063182A1
Method and system for enabling log record consumers to comply with regulations and requirements regarding privacy and the handling of personal data
US20190340388A1
System and method for sensitive data retirement
US20210209251A1
Systems and methods for highly scalable system log analysis, deduplication and management
US9122694B1