Compliance Violation Detection via Schema-Based Data Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage technologies fail to automatically identify and enforce compliance with legal and business regulations, particularly in large distributed networks, making it difficult to audit and ensure policy compliance across diverse data types and geographies.
Innovation Solution
The system recursively discovers network data, groups it by type, identifies data schemas, and applies policy rules to determine compliance, allowing for efficient scanning of only relevant data portions to generate compliance reports and remediate violations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If comprehensive scanning of all network data is performed to ensure policy compliance, then compliance detection accuracy is improved, but computational resource consumption increases
Solution Approach 1:
The patent segments the scanning process by dividing network data into different types (structured, unstructured, semi-structured) and applying type-specific scanning strategies. It further segments files into portions based on data schema locations, scanning only relevant portions rather than entire files, thereby reducing computational resource consumption while maintaining compliance detection accuracy.
Solution Approach 2:
The patent applies local quality by adapting the scanning approach to the specific characteristics of each data type and file portion. Different scanning methods are used for different data types (e.g., regex for structured data, heuristic analysis for unstructured data), and only relevant file portions containing data schemas are scanned intensively, optimizing resource usage while ensuring accurate compliance detection.
2Measurement precision
If recursive discovery of all network data is performed to ensure complete compliance auditing, then compliance coverage is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary actions by first recursively discovering and categorizing all network data, then pre-identifying data schemas and their locations within files. This preliminary structuring allows subsequent compliance scanning to focus only on relevant data portions, significantly reducing processing time while maintaining comprehensive compliance coverage.
Solution Approach 2:
The patent segments the data discovery and scanning process into distinct phases: recursive discovery, data typing, schema identification, and targeted scanning. This segmentation allows the system to process large volumes of network data efficiently by applying appropriate methods to each segment, reducing overall processing time while ensuring complete compliance coverage.
3Measurement precision
If detailed data schema identification is performed to accurately apply policy rules, then policy rule application accuracy is improved, but data processing complexity increases
Solution Approach 1:
The patent applies local quality by identifying and analyzing only the specific portions of files that contain data schemas, rather than processing entire files uniformly. Different identification techniques are applied based on the local characteristics of each data type and file format, reducing processing complexity while maintaining high accuracy in policy rule application.
Solution Approach 2:
The patent performs partial action by scanning only the portions of files that contain data schemas rather than entire files. This selective approach reduces data processing complexity while maintaining sufficient accuracy for policy rule application, as the critical information for compliance determination is located only in specific schema-defined portions.
4Measurement precision
If scanning of multiple files in a grouping is performed to ensure complete compliance verification, then compliance verification thoroughness is improved, but scanning efficiency decreases
Solution Approach 1:
The patent segments the file scanning process by identifying data schemas in one or more files and then scanning only the relevant portions of additional files based on those schema definitions. This segmentation maintains thorough compliance verification while improving scanning efficiency by avoiding redundant scanning of entire files when only specific data portions are relevant.
Solution Approach 2:
The patent applies partial action by scanning only the portions of multiple files that contain data relevant to the identified schemas and applicable policy rules, rather than scanning entire files. This approach maintains verification thoroughness for compliance-critical data while significantly improving scanning efficiency across file groupings.
Data Source
AI summary
Non-limiting examples of the present disclosure describe systems and methods for scanning of data for policy compliance. In one example, network data is evaluated to generate one or more groupings. A grouping may be based on file type of the network data. Data identification rules are applied to identify one or more data schemas from file data of a grouping. One or more policy rules that apply to content of the data schema may be determined. At least one file of the file data may be scanned to determine compliance with the one or more policy rules. A report of compliance with the one or more policy rules may be generated based on a result of a file scan. Other examples are also described.


