Document Classification via Discourse Trees for Cloud Security

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification systems for data loss prevention, relying on keyword analysis and machine learning, fail to accurately classify public and private documents due to limitations in analyzing writing style and grammar, leading to incorrect classifications and false positives.

Innovation Solution

The use of communicative discourse trees, which combine rhetoric information with communicative actions, is employed in conjunction with machine learning models to improve document classification by representing documents in a richer feature set, enabling better recognition of rhetorical relationships and discourse styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If keyword analysis and machine learning are used for document classification, then the system can automatically classify documents, but the classification accuracy deteriorates due to keywords appearing in both public and private documents

Engineering Contradiction:
Improveautomatic document classificationVSAvoidclassification accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent transitions from keyword-level analysis to discourse-tree-level analysis, adding a hierarchical dimension to the classification approach. By constructing discourse trees that represent the rhetorical structure and communicative intent of documents, the system moves beyond flat keyword matching to a multi-dimensional analysis that captures the nuanced differences between public and private documents.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the fundamental parameters used for classification from simple keyword presence/absence to complex discourse features including rhetorical relationships, communicative actions, and tree structure patterns. This parameter transformation enables the system to distinguish between public and private documents based on their underlying discourse characteristics rather than surface-level keywords.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If keyword-based classification is used, then the system is simple to implement, but it fails to recognize sensitive information and produces false positives

Engineering Contradiction:
Improvesystem implementation simplicityVSAvoiddetection accuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces discourse trees as an intermediary representation layer between the raw document text and the classification decision. These discourse trees serve as a mediator that captures the rhetorical structure and communicative intent, allowing the system to make more reliable classification decisions without requiring direct complex analysis of the original text.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary discourse analysis and tree construction before the actual classification step. By pre-processing documents into structured discourse representations, the system prepares enriched feature sets that facilitate more accurate classification while maintaining a clear separation between the analysis phase and the decision phase.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If discourse analysis is used to improve classification accuracy, then the system can better distinguish public and private documents, but the system complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem structural complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the document analysis process into distinct modular components: discourse segmentation into elementary discourse units, construction of discourse trees with rhetorical relationships, extraction of communicative actions, and final classification. This segmentation allows each component to be developed and optimized independently, managing overall system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal discourse tree representation that can handle various types of documents and classification tasks. The discourse tree framework serves multiple functions including structural analysis, rhetorical relationship capture, and feature extraction, reducing the need for separate specialized systems for different document types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12141177B2Data loss prevention system for cloud security based on document discourse analysis
Publication Date: 2024.11.12 ORACLE INT CORP
  • US12141177B2 patent drawing
  • US12141177B2 patent drawing
  • US12141177B2 patent drawing

AI summary

Systems, devices, and methods of the present invention are related to determining a document classification. For example, a document classification application generates a set of discourse trees, each discourse tree corresponding to a sentence of a document and including a rhetorical relationship that relates two elementary discourse units. The document classification application creates one or more communicative discourse trees from the discourse trees by matching each elementary discourse unit in a discourse tree that has a verb to a verb signature. The document classification application combines the first communicative discourse tree and the second communicative discourse tree into a parse thicket and applies a classification model to the parse thicket in order to determine whether the document is public or private.