Compliance Violation Detection via Schema-Based Data Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage technologies fail to automatically identify and enforce compliance with legal and business regulations, particularly in large distributed networks, making it difficult to audit and ensure policy compliance across diverse data types and geographies.

Innovation Solution

The system recursively discovers network data, groups it by type, identifies data schemas, and applies policy rules to determine compliance, allowing for efficient scanning of only relevant data portions to generate compliance reports and remediate violations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If comprehensive scanning of all network data is performed to ensure policy compliance, then compliance detection accuracy is improved, but computational resource consumption increases

Engineering Contradiction:
Improvecompliance detection accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the scanning process by dividing network data into different types (structured, unstructured, semi-structured) and applying type-specific scanning strategies. It further segments files into portions based on data schema locations, scanning only relevant portions rather than entire files, thereby reducing computational resource consumption while maintaining compliance detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by adapting the scanning approach to the specific characteristics of each data type and file portion. Different scanning methods are used for different data types (e.g., regex for structured data, heuristic analysis for unstructured data), and only relevant file portions containing data schemas are scanned intensively, optimizing resource usage while ensuring accurate compliance detection.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If recursive discovery of all network data is performed to ensure complete compliance auditing, then compliance coverage is improved, but processing time increases

Engineering Contradiction:
Improvecompliance coverageVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by first recursively discovering and categorizing all network data, then pre-identifying data schemas and their locations within files. This preliminary structuring allows subsequent compliance scanning to focus only on relevant data portions, significantly reducing processing time while maintaining comprehensive compliance coverage.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the data discovery and scanning process into distinct phases: recursive discovery, data typing, schema identification, and targeted scanning. This segmentation allows the system to process large volumes of network data efficiently by applying appropriate methods to each segment, reducing overall processing time while ensuring complete compliance coverage.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If detailed data schema identification is performed to accurately apply policy rules, then policy rule application accuracy is improved, but data processing complexity increases

Engineering Contradiction:
Improvepolicy rule application accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by identifying and analyzing only the specific portions of files that contain data schemas, rather than processing entire files uniformly. Different identification techniques are applied based on the local characteristics of each data type and file format, reducing processing complexity while maintaining high accuracy in policy rule application.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs partial action by scanning only the portions of files that contain data schemas rather than entire files. This selective approach reduces data processing complexity while maintaining sufficient accuracy for policy rule application, as the critical information for compliance determination is located only in specific schema-defined portions.

Inventive Principle:
Principle #16Partial or excessive action

4Measurement precision

If scanning of multiple files in a grouping is performed to ensure complete compliance verification, then compliance verification thoroughness is improved, but scanning efficiency decreases

Engineering Contradiction:
Improvecompliance verification thoroughnessVSAvoidscanning efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the file scanning process by identifying data schemas in one or more files and then scanning only the relevant portions of additional files based on those schema definitions. This segmentation maintains thorough compliance verification while improving scanning efficiency by avoiding redundant scanning of entire files when only specific data portions are relevant.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by scanning only the portions of multiple files that contain data relevant to the identified schemas and applicable policy rules, rather than scanning entire files. This approach maintains verification thoroughness for compliance-critical data while significantly improving scanning efficiency across file groupings.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11755529B2Compliance violation detection
Publication Date: 2023.09.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11755529B2 patent drawing
  • US11755529B2 patent drawing
  • US11755529B2 patent drawing

AI summary

Non-limiting examples of the present disclosure describe systems and methods for scanning of data for policy compliance. In one example, network data is evaluated to generate one or more groupings. A grouping may be based on file type of the network data. Data identification rules are applied to identify one or more data schemas from file data of a grouping. One or more policy rules that apply to content of the data schema may be determined. At least one file of the file data may be scanned to determine compliance with the one or more policy rules. A report of compliance with the one or more policy rules may be generated based on a result of a file scan. Other examples are also described.