Document Impact Analysis via Semantic Clause Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current natural language processing methods fail to accurately interpret the meaning of unstructured text in a contextually relevant manner, particularly in understanding the impact of a document on a specific concept, due to limitations in statistical models being black boxes, requiring large homogeneous training sets, and inability to interpret fine-grained context and interrelationships.

Innovation Solution

A method and system that generate clusters of semantically similar clauses, identify representative concepts, and compute impact using semantic parameters such as impact phrases, intensity, and location, with a configurable impact analysis engine driven by externalized rules, allowing for adaptable and extendable context-based interpretation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If statistical models are used to classify documents, then automated processing is achieved, but the models become black boxes that cannot be explained to users

Engineering Contradiction:
Improveautomated document processingVSAvoidmodel interpretability
Core Design Contradiction:
Extent of automationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the document processing into distinct modules: clause extraction, concept identification, impact analysis, and result generation. Each module operates independently and can be traced, transforming the black box into a transparent series of steps where users can see exactly how conclusions are reached.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer of semantic concepts and relationships between clauses and concepts. This intermediary structure serves as a bridge between the raw document text and the final classification results, making the reasoning process visible and explainable to users while maintaining automated processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If statistical models are trained on large homogeneous training sets, then classification accuracy is improved, but the system becomes less adaptable to new contexts and concepts

Engineering Contradiction:
Improveclassification accuracyVSAvoidcontext adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic concept identification where concepts are automatically discovered and defined during processing rather than being fixed in a static training set. The system adapts concepts based on the specific document context, allowing accurate processing of new domains without requiring retraining on homogeneous data.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters dynamically by adjusting concept thresholds, relationship weights, and analysis depth based on the specific document being processed. This allows the model to maintain high accuracy across diverse contexts without being constrained by a single homogeneous training distribution.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If word-based statistical methods are used, then processing speed is maintained, but the system cannot interpret fine-grained context and ambiguous words accurately

Engineering Contradiction:
Improveprocessing speedVSAvoidcontextual interpretation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments text processing at the clause level rather than just the word level, creating meaningful semantic units that preserve context. This segmentation maintains processing efficiency while enabling accurate interpretation of ambiguous words through their contextual relationships within clauses and documents.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from one-dimensional word frequency analysis to multi-dimensional analysis that includes clause-level semantics, concept relationships, and contextual patterns. This dimensional expansion enables accurate contextual interpretation without sacrificing processing speed, as the system processes these additional dimensions in parallel.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Measurement precision

If the system analyzes detailed contextual relationships, then interpretation accuracy is improved, but the processing time and computational complexity increase

Engineering Contradiction:
Improveimpact analysis accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the analysis process into independent clause-level units that can be processed in parallel. Each clause is analyzed for its concepts and relationships separately, then combined to produce the overall document impact. This segmentation enables detailed contextual analysis while maintaining processing time efficiency through parallel computation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial analysis to the most important clauses and concepts first, with the option to perform more extensive analysis on specific segments based on user needs. This allows high-accuracy impact analysis of critical content while avoiding unnecessary computational overhead in less important areas.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9792277B2System and method for determining the meaning of a document with respect to a concept
Publication Date: 2017.10.17 GENPACT USA INC
  • US9792277B2 patent drawing
  • US9792277B2 patent drawing
  • US9792277B2 patent drawing

AI summary

A computerized method for determining an impact of a document on the specific concept of interest. The method can be configured to identify a cluster of clauses or sentences from a plurality of semantically similar clauses of the document and determine one or more representative concepts for the cluster of the document. An impact of each clause of the cluster is determined using one or more semantic parameters and impact analysis rules. The impact of the each sentence of the cluster is then determined using the impact of the respective clauses and subsequently, the impact of the cluster is determined using the impact of the respective sentences. Based on the impact of the cluster, an impact of the document on the one or more representative concepts is determined.