Knowledge Engine Veracity Assessment for NLP Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cognitive systems in natural language processing are inherently non-deterministic, leading to inconsistencies and inaccuracies due to susceptibility to input data and new machine learning models, which can result in incorrect data extraction and output.

Innovation Solution

A system that employs a knowledge engine to train machine learning models by querying a first knowledge graph, extracting triplets with subject, object, and relationships, and evaluating modifications using BC identifiers and veracity values from a ledger, ensuring deterministic data processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If machine learning models are deployed to process natural language, then the system can learn from data and provide relevant recommendations, but the system becomes non-deterministic and may extract incorrect entities or provide inaccurate output

Engineering Contradiction:
Improvelearning capabilityVSAvoiddeterministic behavior
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary verification layer that sits between the machine learning model and the final output. This layer includes techniques such as ensemble methods, confidence thresholding, and cross-validation mechanisms that mediate the non-deterministic output of ML models to produce deterministic, reliable results. The intermediary validates and filters model predictions before they are presented as final output.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback loops where model predictions are continuously evaluated against ground truth data or expert validation. Incorrect extractions or inaccurate outputs trigger retraining or model adjustment processes. This feedback mechanism ensures that while the model maintains its learning capability, its output becomes progressively more deterministic and reliable through iterative improvement.

Inventive Principle:
Principle #23Feedback

2Productivity

If new machine learning models are deployed continuously, then the system can improve its performance and accuracy, but prior model results may be adversely affected and consistency is lost

Engineering Contradiction:
Improvemodel improvement rateVSAvoidresult consistency
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

Before deploying new machine learning models, the system performs preliminary actions including backward compatibility testing, impact analysis on existing results, and validation against historical data. This ensures that new models do not adversely affect prior model results. The preliminary action phase includes simulating model behavior on past datasets to predict potential inconsistencies before actual deployment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements dynamic model management where multiple versions of models coexist and are selected based on the specific task requirements. Rather than replacing models continuously, the system dynamically chooses appropriate model versions for different queries, maintaining result consistency while still allowing for model improvements. This dynamic approach enables progressive enhancement without sacrificing stability.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If the system processes unverified data from documents, then it can handle diverse input, but errors are introduced and incorrect data is extracted and provided as output

Engineering Contradiction:
Improveinput handling capabilityVSAvoiddata extraction accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system implements beforehand cushioning through multiple layers of verification and validation mechanisms that are built into the data processing pipeline before final output generation. These include data cleaning protocols, anomaly detection algorithms, and confidence scoring systems that cushion against errors introduced by unverified input data. The cushioning mechanisms ensure that even when diverse, unverified data is processed, the final extracted information maintains high accuracy.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS10846485B2Machine learning model modification and natural language processing
Publication Date: 2020.11.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10846485B2 patent drawing
  • US10846485B2 patent drawing
  • US10846485B2 patent drawing

AI summary

A system, computer program product, and method are provided to automate a framework for knowledge graph based persistence of data, and to resolve temporal changes and uncertainties in the knowledge graph. Natural language understanding, together with one or more machine learning models (MLMs), is used to extract data from unstructured information, including entities and entity relationships. The extracted data is populated into a knowledge graph. As the KG is subject to change, the KG is used to create new and retrain existing machine learning models (MLMs). Weighting is applied to the populated data in the form of veracity value. Blockchain technology is applied to the populated data to ensure reliability of the data and to provide auditability to assess changes to the data.