Document Classification via Discourse Trees for Cloud Security
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document classification systems for data loss prevention, relying on keyword analysis and machine learning, fail to accurately classify public and private documents due to limitations in analyzing writing style and grammar, leading to incorrect classifications and false positives.
Innovation Solution
The use of communicative discourse trees, which combine rhetoric information with communicative actions, is employed in conjunction with machine learning models to improve document classification by representing documents in a richer feature set, enabling better recognition of rhetorical relationships and discourse styles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If keyword analysis and machine learning are used for document classification, then the system can automatically classify documents, but the classification accuracy deteriorates due to keywords appearing in both public and private documents
Solution Approach 1:
The patent transitions from keyword-level analysis to discourse-tree-level analysis, adding a hierarchical dimension to the classification approach. By constructing discourse trees that represent the rhetorical structure and communicative intent of documents, the system moves beyond flat keyword matching to a multi-dimensional analysis that captures the nuanced differences between public and private documents.
Solution Approach 2:
The patent changes the fundamental parameters used for classification from simple keyword presence/absence to complex discourse features including rhetorical relationships, communicative actions, and tree structure patterns. This parameter transformation enables the system to distinguish between public and private documents based on their underlying discourse characteristics rather than surface-level keywords.
2Ease of manufacture
If keyword-based classification is used, then the system is simple to implement, but it fails to recognize sensitive information and produces false positives
Solution Approach 1:
The patent introduces discourse trees as an intermediary representation layer between the raw document text and the classification decision. These discourse trees serve as a mediator that captures the rhetorical structure and communicative intent, allowing the system to make more reliable classification decisions without requiring direct complex analysis of the original text.
Solution Approach 2:
The patent performs preliminary discourse analysis and tree construction before the actual classification step. By pre-processing documents into structured discourse representations, the system prepares enriched feature sets that facilitate more accurate classification while maintaining a clear separation between the analysis phase and the decision phase.
3Measurement precision
If discourse analysis is used to improve classification accuracy, then the system can better distinguish public and private documents, but the system complexity increases
Solution Approach 1:
The patent segments the document analysis process into distinct modular components: discourse segmentation into elementary discourse units, construction of discourse trees with rhetorical relationships, extraction of communicative actions, and final classification. This segmentation allows each component to be developed and optimized independently, managing overall system complexity through modular architecture.
Solution Approach 2:
The patent creates a universal discourse tree representation that can handle various types of documents and classification tasks. The discourse tree framework serves multiple functions including structural analysis, rhetorical relationship capture, and feature extraction, reducing the need for separate specialized systems for different document types.
Data Source
AI summary
Systems, devices, and methods of the present invention are related to determining a document classification. For example, a document classification application generates a set of discourse trees, each discourse tree corresponding to a sentence of a document and including a rhetorical relationship that relates two elementary discourse units. The document classification application creates one or more communicative discourse trees from the discourse trees by matching each elementary discourse unit in a discourse tree that has a verb to a verb signature. The document classification application combines the first communicative discourse tree and the second communicative discourse tree into a parse thicket and applies a classification model to the parse thicket in order to determine whether the document is public or private.


