Privacy Preserving Document Analysis via Stamp Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital analytics systems are unable to collect and extract insights from documents where users expect privacy, as they require access to the content of the digital documents, violating user privacy.

Innovation Solution

The system employs privacy-preserving document analysis techniques that capture visual or contextual features of digital documents to create a stamp representation, which is then projected into a stamp embedding space using a machine learning model. This allows for the derivation of insights about the document without accessing its content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional digital analytics systems access the content of digital documents to generate insights, then measurement precision and productivity are improved, but user privacy is violated

Engineering Contradiction:
Improveinsight accuracyVSAvoidprivacy violation
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary visual and contextual features from digital documents to create stamp representations, separating the essential analytical information from the actual document content. This extraction process enables analytics systems to obtain insights without accessing or storing the original document content, thereby maintaining measurement precision while protecting user privacy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces stamp representations as an intermediary between the original document and the analytics system. These stamp representations serve as a mediator that captures document characteristics and contextual information without containing the actual document content, allowing analytics to be performed on the intermediary rather than the original sensitive material.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If digital analytics systems process original document content, then insight quality is improved, but device complexity and data security requirements increase

Engineering Contradiction:
Improveinsight qualityVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the document analysis process into distinct components: feature extraction, stamp representation generation, and analytics processing. This segmentation allows each component to handle only the specific task it is designed for, reducing the complexity of individual components while maintaining overall insight quality. The stamp representation acts as an intermediate segmented form that preserves necessary information without requiring the full original document.

Inventive Principle:
Principle #1Segmentation

3Productivity

If conventional systems monitor all user interactions with digital content, then productivity and data collection are improved, but user satisfaction and privacy expectations are compromised

Engineering Contradiction:
Improvedata collection efficiencyVSAvoiduser satisfaction
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent applies partial action by collecting only the specific visual and contextual features necessary for analytics rather than monitoring all user interactions and document content. This partial data collection approach maintains productivity by gathering sufficient information for insights while improving user satisfaction by respecting privacy expectations and avoiding excessive monitoring.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12267305B2Privacy preserving document analysis
Publication Date: 2025.04.01 ADOBE INC
  • US12267305B2 patent drawing
  • US12267305B2 patent drawing
  • US12267305B2 patent drawing

AI summary

Systems and techniques for privacy preserving document analysis are described that derive insights pertaining to a digital document without communication of the content of the digital document. To do so, the privacy preserving document analysis techniques described herein capture visual or contextual features of the digital document and creates a stamp representation that represents these features without included the content of the digital document. The stamp representation is projected into a stamp embedding space based on a stamp encoding model generated through machine learning techniques capturing feature patterns and interaction in the stamp representations. The stamp encoding model exploits these feature interactions to define similarity of source documents based on location within the stamp embedding space. Accordingly, the techniques described herein can determine a similarity of documents without having access to the documents themselves.