Rule Intent Relation Graph for Unstructured Document Mining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data mining techniques fail to effectively extract and represent rule intents from unstructured natural language documents in a structured format, leading to difficulties in machine analysis and comprehension due to noise elimination, rule sentence identification, and relationship extraction among rule intents.
Innovation Solution
A method and system that identify relationships among rule intents by extracting sentences, mining rule intents using heuristic rules, and creating pair-wise relation graphs optimized using trained classifiers, displayed in Semantics of Business Vocabulary and Rules (SBVR) format, which enables machine analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual extraction of rules from unstructured documents is performed, then rule intents can be extracted, but the process becomes tedious and time-consuming due to document size and structure
Solution Approach 1:
The patent replaces manual mechanical extraction processes with automated text mining techniques using machine learning classifiers. The system automatically identifies rule sentences, extracts rule intents, and determines relationships among intents, eliminating the need for tedious manual extraction while maintaining high accuracy through trained classification models.
Solution Approach 2:
The system enables self-service extraction by training classifiers on annotated documents that automatically identify and extract rule intents and their relationships. Once trained, the system can independently process new documents without human intervention, performing extraction, relationship identification, and formal representation generation autonomously.
2Extent of automation
If existing text mining techniques use predefined templates for structured documents, then extraction can be automated, but they fail to eliminate noise and identify rule sentences in unstructured documents
Solution Approach 1:
The patent employs dynamic classification approaches where multiple trained classifiers work in sequence to identify rule sentences and extract intents. The system adapts to different document types and structures by using flexible classification criteria rather than rigid predefined templates, maintaining reliability across varied unstructured document formats.
Solution Approach 2:
The extraction process is segmented into distinct stages: rule sentence identification, rule intent extraction, relationship identification, and formal representation. Each stage uses specialized classifiers trained for specific tasks, allowing the system to handle unstructured documents effectively by breaking down the complex extraction problem into manageable components.
3Ease of manufacture
If rule intents are extracted without identifying relationships among them, then extraction is simpler, but machine analysis and comprehension cannot be performed
Solution Approach 1:
The patent introduces relationship identification as an intermediary step between intent extraction and machine analysis. Trained classifiers identify semantic relationships (such as implication, equivalence, or contradiction) among extracted intents, creating a structured representation that enables subsequent machine analysis while building upon the simpler extraction foundation.
Solution Approach 2:
The system maintains continuous processing through automated pipeline operations: extraction feeds into relationship identification, which feeds into formal representation in SBVR format. This continuous automated workflow preserves extraction simplicity while progressively building machine-analyzable structures without manual intervention at each stage.
4Ease of operation
If extracted rules are not represented in structured format, then extraction is easier, but machines cannot perform analysis and comprehension tasks
Solution Approach 1:
The patent transforms extracted rule intents from unstructured text into structured formal representations using SBVR (Semantics of Business Vocabulary and Rules) notation. This parameter change from natural language to formal logic structure enables machine analysis while preserving all essential information through systematic translation rules that maintain semantic equivalence.
Data Source
AI summary
A method and a system for mining rule intents from documents is provided, wherein the rule intents are basic atomic facts present in a sentence. The proposed method and system for identification of relation among rule intents from a document is performed in multiple stages that include identification and optimization of a pair-wise relation graph from rule intents based on a plurality of relation optimizing heuristic rules. The relations identified among the rule intents are displayed in Semantics of Business Vocabulary and Rules (SBVR) format, which can be easily analyzed by machines as SBVR is a comprehensive standard for business rule representation by Object Management Group (OMG) in accordance with set of a standard pre-defined vocabularies.


