Document Template Learning for Sensitive Data Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to efficiently manage and protect sensitive information in unstructured and semi-structured documents across data sources, failing to identify and protect personal information effectively due to their complexity and size.

Innovation Solution

A system and method for exemplar learning that includes a processing subsystem with modules like a scanner, pre-processor, extraction module, ranking module, and determining module to automatically identify and classify documents with similar templates, using AI algorithms like Deep Neural Networks to refine and rank sequences and sub-sequences, generating feature vectors, and determining threshold values for classifiers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual keyword-based classifiers are used to identify personal information in documents, then the solution can be implemented with existing platforms, but the efficiency and quality of information analysis deteriorate due to the extreme size and complexity of modern data sources

Engineering Contradiction:
Improvequality of information analysisVSAvoidefficiency of information processing
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces manual keyword-based classification with an automated machine learning system that uses trained models to identify and classify personal information in documents. The system automatically processes documents through multiple modules including scanning, preprocessing, template identification, and classification, eliminating the need for manual keyword configuration and significantly improving both efficiency and analysis quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If existing platforms are used to manage unstructured and semi-structured documents, then the implementation is straightforward, but the ability to efficiently identify and protect personal information deteriorates due to the platforms' failure to handle document complexity

Engineering Contradiction:
Improveease of implementationVSAvoidpersonal information protection capability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent divides the document processing system into multiple specialized modules: a scanning module to retrieve documents, a preprocessing module to clean and standardize data, a template identification module to detect document structures, and a classification module to identify personal information. This segmented architecture maintains ease of implementation while significantly improving the reliability of personal information protection through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If a comprehensive solution is built to handle all document types and data sources, then the coverage and protection capability are improved, but the system complexity and development difficulty worsen

Engineering Contradiction:
Improvedocument coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal document processing framework that can handle multiple document types (structured, semi-structured, and unstructured) and various data sources through a single integrated system. The framework uses adaptable template matching and machine learning classifiers that can be trained on different document formats, providing broad coverage without requiring separate specialized systems for each document type, thus managing complexity while maintaining versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12602538B2Method and system for exemplar learning for templatizing documents across data sources
Publication Date: 2026.04.14 SECURITI LLC
  • US12602538B2 patent drawing
  • US12602538B2 patent drawing
  • US12602538B2 patent drawing

AI summary

A system and method for templatizing documents across data sources is disclosed. The system includes a scanner to scan a plurality of files to retrieve a plurality of textual content. The system includes a pre-processor module to refine an efficacy corresponding to the scanned plurality of files by using a classifier to identify textual content with a similar template. The system includes an extraction module to extract a plurality of common sequences and sub-sequences from the plurality of files. The system includes a ranking module to rank the plurality of common sequences and sub-sequences based on a score. The system includes a feature vector generating module to generate a feature vector from the plurality of common sequences and sub-sequences. The system includes a determining module to determine a threshold value for the classifier thereby developing the classifier automatically to search for positive files with similar templates in the organization.