Document Clustering via Structural Paths and Classification Terms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification methods relying on supervised learning face challenges when a large labeled dataset is impractical or impossible to obtain, particularly due to privacy concerns, making it difficult to effectively classify documents like emails without human intervention.

Innovation Solution

The method involves clustering documents based on classification terms and structural paths without human knowledge, identifying classification terms from labeled communications, and determining clusters of communications that include these terms, even if unlabeled, to generate feature sets for automatic classification and data extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised learning techniques are used for document classification, then classification accuracy can be improved, but the requirement for large quantities of labeled documents increases, which may be impractical or impossible to obtain due to privacy considerations

Engineering Contradiction:
Improveclassification accuracyVSAvoidquantity of labeled documents
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary clustering of documents based on structural paths and classification terms before final classification. This preliminary organization groups similar documents together, allowing the classification model to learn from cluster characteristics rather than requiring every document to be individually labeled, thus reducing the need for large quantities of labeled documents while maintaining classification accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces clustering as an intermediary step between raw documents and final classification. Documents are first grouped into clusters based on structural similarity and classification terms, then classification is performed on clusters rather than individual documents. This intermediary approach reduces the labeling burden while preserving classification accuracy through cluster-level representations

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If human review and labeling of communications is performed, then labeled training data can be obtained, but privacy considerations prevent many emails and private documents from being provided for human review

Engineering Contradiction:
Improvelabeled training dataVSAvoidprivacy concerns
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The system enables documents to classify themselves through self-organization into clusters based on their inherent structural paths and classification terms. Documents automatically group with similar documents without requiring human intervention or exposure of private content, thus obtaining labeled training data while preserving privacy. The clustering process is self-service in that it uses intrinsic document features rather than external human labeling

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of human review and labeling with an automated computational clustering system. Instead of physically examining and labeling documents by human reviewers, the system uses algorithms to automatically group documents based on structural paths and classification terms, eliminating the need for human exposure to private documents while still producing labeled training data

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Extent of automation

If clustering based on classification terms is performed, then automatic classification without human intervention is enabled, but the complexity of identifying and processing structural paths increases

Engineering Contradiction:
Improveautomatic classificationVSAvoidcomplexity of processing structural paths
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent segments the document analysis process into distinct components: structural path extraction, classification term identification, and clustering. By dividing the complex task of automatic classification into these manageable segments, the system can process structural paths more efficiently. Each segment handles a specific aspect, reducing the overall complexity compared to attempting to perform all functions in a single undivided process

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10007717B2Clustering communications based on classification
Publication Date: 2018.06.26 GOOGLE LLC
  • US10007717B2 patent drawing
  • US10007717B2 patent drawing
  • US10007717B2 patent drawing

AI summary

Methods and apparatus related to clustering documents based on one or more classification terms and optionally based on similarity of structural paths of the documents. In some implementations, the documents are communications such as structured emails or other structured communications. In some of those implementations, clustering the communications includes identifying a plurality of classification terms indicative of a classification, identifying a corpus of communications that includes communications that are not labeled with an association to the classification, and determining a cluster of the communications based on occurrence of one or more of the classification terms in the communications of the cluster.