Contract Metadata Extraction Using ML Segmentation and Rule Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual tracking of contract obligations and service level agreements in contract documents is laborious and time-consuming, especially for large organizations with numerous documents, making it difficult to extract relevant information effectively.

Innovation Solution

A data processing system and method that extracts performance segments and metadata values from contract documents using machine learning models and a rule-based engine, identifying segment types and extracting relevant metadata based on predefined models, reducing the need to process the entire document and increasing accuracy by validating outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tracking of contract obligations and service level agreements is performed, then information extraction can be done with human judgment, but the process becomes laborious and time-consuming

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidtracking time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical tracking with an automated computer-based system that uses natural language processing and machine learning models to extract performance segments and metadata from contract documents, eliminating human labor while maintaining extraction accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service extraction by automatically identifying and extracting relevant contract obligations and service level agreements without human intervention, using trained models to process documents independently

Inventive Principle:
Principle #25Self-service

2Loss of information

If the entire contract document is processed to extract relevant information, then comprehensive information can be obtained, but the processing time increases significantly

Engineering Contradiction:
Improveinformation completenessVSAvoiddocument processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent divides the contract document into multiple segments and uses natural language processing to identify and extract only the relevant performance segments containing obligations and service level agreements, rather than processing the entire document, thus reducing processing time while maintaining information completeness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts only the necessary performance segments and metadata from the contract document using targeted extraction models, obtaining comprehensive relevant information without processing unnecessary portions of the document

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If automated extraction systems are implemented, then processing time is reduced, but extraction accuracy may decrease without proper validation

Engineering Contradiction:
Improveextraction speedVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a rule-based validation engine that provides feedback on the extracted performance segments and metadata, validating the automated extraction results against predefined rules and business logic to ensure accuracy while maintaining high processing speed

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system combines multiple approaches - machine learning models for extraction and rule-based validation for verification - creating a composite extraction system that leverages the strengths of both methods to achieve high speed and high accuracy simultaneously

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS11482027B2Automated extraction of performance segments and metadata values associated with the performance segments from contract documents
Publication Date: 2022.10.25 SIRIONLABS PTE LTD
  • US11482027B2 patent drawing
  • US11482027B2 patent drawing
  • US11482027B2 patent drawing

AI summary

A data processing system for extracting metadata values is described. The data processing system includes an input unit and a processor communicably coupled to the input unit. The input unit is configured to receive a contract document. The processor is configured to extract at least one segment from the contract document and identify a type of the at least one segment. The processor is further configured to extract at least one metadata value from the at least one segment based on a model, wherein the model is determined based on the identified type of the at least one segment.