Rule Intent Relation Graph for Unstructured Document Mining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data mining techniques fail to effectively extract and represent rule intents from unstructured natural language documents in a structured format, leading to difficulties in machine analysis and comprehension due to noise elimination, rule sentence identification, and relationship extraction among rule intents.

Innovation Solution

A method and system that identify relationships among rule intents by extracting sentences, mining rule intents using heuristic rules, and creating pair-wise relation graphs optimized using trained classifiers, displayed in Semantics of Business Vocabulary and Rules (SBVR) format, which enables machine analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual extraction of rules from unstructured documents is performed, then rule intents can be extracted, but the process becomes tedious and time-consuming due to document size and structure

Engineering Contradiction:
Improveextraction accuracyVSAvoidextraction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical extraction processes with automated text mining techniques using machine learning classifiers. The system automatically identifies rule sentences, extracts rule intents, and determines relationships among intents, eliminating the need for tedious manual extraction while maintaining high accuracy through trained classification models.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service extraction by training classifiers on annotated documents that automatically identify and extract rule intents and their relationships. Once trained, the system can independently process new documents without human intervention, performing extraction, relationship identification, and formal representation generation autonomously.

Inventive Principle:
Principle #25Self-service

2Extent of automation

If existing text mining techniques use predefined templates for structured documents, then extraction can be automated, but they fail to eliminate noise and identify rule sentences in unstructured documents

Engineering Contradiction:
Improveextraction automationVSAvoidrule identification accuracy
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent employs dynamic classification approaches where multiple trained classifiers work in sequence to identify rule sentences and extract intents. The system adapts to different document types and structures by using flexible classification criteria rather than rigid predefined templates, maintaining reliability across varied unstructured document formats.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The extraction process is segmented into distinct stages: rule sentence identification, rule intent extraction, relationship identification, and formal representation. Each stage uses specialized classifiers trained for specific tasks, allowing the system to handle unstructured documents effectively by breaking down the complex extraction problem into manageable components.

Inventive Principle:
Principle #1Segmentation

3Ease of manufacture

If rule intents are extracted without identifying relationships among them, then extraction is simpler, but machine analysis and comprehension cannot be performed

Engineering Contradiction:
Improveextraction simplicityVSAvoidmachine analysis capability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces relationship identification as an intermediary step between intent extraction and machine analysis. Trained classifiers identify semantic relationships (such as implication, equivalence, or contradiction) among extracted intents, creating a structured representation that enables subsequent machine analysis while building upon the simpler extraction foundation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system maintains continuous processing through automated pipeline operations: extraction feeds into relationship identification, which feeds into formal representation in SBVR format. This continuous automated workflow preserves extraction simplicity while progressively building machine-analyzable structures without manual intervention at each stage.

Inventive Principle:
Principle #20Continuity of useful action

4Ease of operation

If extracted rules are not represented in structured format, then extraction is easier, but machines cannot perform analysis and comprehension tasks

Engineering Contradiction:
Improveextraction easeVSAvoidstructure information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent transforms extracted rule intents from unstructured text into structured formal representations using SBVR (Semantics of Business Vocabulary and Rules) notation. This parameter change from natural language to formal logic structure enables machine analysis while preserving all essential information through systematic translation rules that maintain semantic equivalence.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11048881B2Method and system for identification of relation among rule intents from a document
Publication Date: 2021.06.29 TATA CONSULTANCY SERVICES LTD
  • US11048881B2 patent drawing
  • US11048881B2 patent drawing
  • US11048881B2 patent drawing

AI summary

A method and a system for mining rule intents from documents is provided, wherein the rule intents are basic atomic facts present in a sentence. The proposed method and system for identification of relation among rule intents from a document is performed in multiple stages that include identification and optimization of a pair-wise relation graph from rule intents based on a plurality of relation optimizing heuristic rules. The relations identified among the rule intents are displayed in Semantics of Business Vocabulary and Rules (SBVR) format, which can be easily analyzed by machines as SBVR is a comprehensive standard for business rule representation by Object Management Group (OMG) in accordance with set of a standard pre-defined vocabularies.