NLP System for Automated Contract Entity Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing contracts and other structured documents are inefficient, as they require manual extraction of data, making it difficult to identify clauses and terms for actionable analytics.

Innovation Solution

A system and method for natural language processing of structured documents using a server with a computer processor that parses documents, generates ontologies, extracts entities, identifies relationships, and creates vector representations, enabling automated data extraction and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual extraction of data from contracts is used, then data extraction accuracy can be maintained, but processing time and labor costs increase significantly

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data extraction with an automated NLP system comprising entity recognition engines, relation extraction engines, and statistical parsers that process contract documents automatically, eliminating human labor while maintaining extraction accuracy through machine learning models

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data extraction where the contract document itself provides the structured data through automated processing pipelines that extract entities, relationships, and clauses without requiring human analysts to manually review each document

Inventive Principle:
Principle #25Self-service

2Productivity

If automated NLP processing is implemented, then processing speed and productivity improve, but system complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The NLP system is divided into distinct modular components including entity recognition engines, relation extraction engines, statistical parsers, and ontology generators, allowing each module to be developed, maintained, and optimized independently while working together to achieve high processing speed

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary components such as ontologies and structured schemas that mediate between raw contract text and extracted data, simplifying the overall processing pipeline by providing standardized intermediate representations that reduce system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If comprehensive data extraction is performed on all contract components, then data completeness improves, but processing time and computational resources increase

Engineering Contradiction:
Improvedata completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system extracts only the most relevant entities, relationships, and clauses from contract documents using targeted extraction rules and machine learning models that identify and extract critical information while ignoring redundant content, achieving data completeness for key elements without processing every detail

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Different extraction strategies and levels of detail are applied to different sections of the contract based on their importance, with critical clauses receiving more comprehensive analysis while less important sections receive streamlined processing, optimizing the balance between completeness and processing time

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11170179B2Systems and methods for natural language processing of structured documents
Publication Date: 2021.11.09 JPMORGAN CHASE BANK NA
  • US11170179B2 patent drawing
  • US11170179B2 patent drawing
  • US11170179B2 patent drawing

AI summary

Systems and methods for natural language processing of structured documents. In another embodiment, in an information processing apparatus comprising at least one computer processor, a method for processing a structured document may include: (1) receiving a document; (2) parsing the document into a plurality of components using a statistical parser; (3) extracting a plurality of entities from each component; (4) identifying a potential relationship between two of the plurality of entities; (5) generating a numeric representation for the potential relationship; (6) confirming the potential relationship with a logical regression model; and (7) generating and storing a unified structured file for the document.