Structured Content Data Loss Protection via Flattened Stream Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing user productivity applications face challenges in identifying and protecting sensitive data within structured documents like spreadsheets and presentations, where sensitive information may be split across multiple data entities, making it difficult to prevent data loss and ensure compliance with privacy policies.

Innovation Solution

A data loss protection framework that receives structured user content, transforms it into flattened representations, and uses mapping information to identify sensitive data, marking it within the user interface and providing options for obfuscation, allowing for efficient detection and protection of sensitive content across various applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If sensitive data is distributed across multiple structured data entities (cells, objects, tables), then data flexibility and document complexity are improved, but data loss protection effectiveness deteriorates due to difficulty in identification and tracking

Engineering Contradiction:
Improvedata distribution flexibilityVSAvoiddata loss protection effectiveness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments sensitive data identification into two phases: (1) extraction of content from structured data entities into a linearized stream, and (2) classification of sensitive portions within that stream. This segmentation allows the system to handle distributed data across multiple cells and objects by processing it as a unified sequence while tracking original locations through offset information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a content stream as an intermediary representation between the structured data entities and the sensitive data classification process. This content stream serves as a mediator that linearizes the distributed data while preserving location information through offsets, enabling effective DLP analysis without requiring direct manipulation of the complex structured formats.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If traditional DLP methods are applied to structured documents, then implementation simplicity is maintained, but detection accuracy deteriorates because sensitive data may be split across multiple data entities

Engineering Contradiction:
ImproveDLP implementation complexityVSAvoidsensitive data detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transforms the detection problem from a two-dimensional structured format (rows and columns in spreadsheets, slides in presentations) to a one-dimensional linearized content stream. This dimensional transformation simplifies the classification process while maintaining the ability to accurately detect sensitive data through offset mapping back to the original structured locations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If comprehensive sensitive data scanning is performed on all content, then data protection coverage is improved, but processing time increases due to the need to analyze all structured data entities

Engineering Contradiction:
Improvedata protection coverageVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the relevant content portions from the structured data entities into a linearized stream for classification analysis. By extracting content rather than processing entire structured documents, the system achieves comprehensive scanning of all potential sensitive data while significantly reducing processing time through more efficient data representation and analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10671753B2Sensitive data loss protection for structured user content viewed in user applications
Publication Date: 2020.06.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10671753B2 patent drawing
  • US10671753B2 patent drawing
  • US10671753B2 patent drawing

AI summary

Systems, methods, and software for sensitive data handling frameworks for user applications are provided herein. An exemplary method includes receiving subsets of structured user content consolidated into associated flattened representations, the associated flattened representations having a mapping to the structured user content and accompanied by at least lengths and offset information relating to the mapping. The method includes individually parsing the subsets of structured user content to classify portions as comprising sensitive content corresponding to one or more predetermined data schemes and, for each of the portions, identifying an associated offset and length for the portion relating to the subsets of structured user content, and indicating at least the associated offset and length to the user application for marking of the sensitive content in a user interface to the user application.