Knowledge Graph and Blockchain for NLP Data Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cognitive systems in natural language processing are inherently non-deterministic, leading to inconsistent data extraction and output due to susceptibility to input information and new machine learning models, which can result in incorrect data extraction and provision of inaccurate results.
Innovation Solution
A system that includes a processing unit with an artificial intelligence platform, a knowledge engine, and a machine learning model manager, which queries natural language input against a knowledge graph and blockchain ledger to extract and verify triplets, sorting them based on veracity values and augmenting machine learning models for deterministic data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If machine learning models are deployed to process natural language, then the system can learn from data and provide relevant recommendations, but the system becomes non-deterministic and may extract inconsistent entities
Solution Approach 1:
The patent introduces an intermediary verification layer that sits between the machine learning model and the final output. This layer includes confidence score thresholds and consistency checks that mediate the non-deterministic ML output, filtering results to ensure only high-confidence, consistent extractions are returned, thus resolving the contradiction between adaptability and reliability
Solution Approach 2:
The system implements feedback mechanisms where extraction results are verified against multiple criteria including confidence scores and consistency with previous extractions. This feedback loop allows the system to learn from its own outputs and adjust, maintaining deterministic behavior while preserving learning capabilities
2Adaptability or versatility
If new machine learning models are deployed to improve processing capabilities, then the system can handle new patterns, but prior model results may be adversely affected and consistency is lost
Solution Approach 1:
The patent applies preliminary action by establishing a baseline of extracted entities and relationships before deploying new models. When new models are introduced, their outputs are compared against this baseline to ensure consistency, preventing adverse effects on prior results while allowing capability improvements
Solution Approach 2:
The system changes parameters such as confidence thresholds and verification criteria when integrating new models. This allows the system to adapt to new model outputs while maintaining stable extraction results through adjusted parameter settings that ensure consistency
3Productivity
If the system processes all extracted data without verification, then processing speed is maintained, but incorrect data may be extracted and provided as output
Solution Approach 1:
The patent applies partial action by implementing selective verification - not all extracted data undergoes the same level of scrutiny. High-confidence extractions above a threshold are processed quickly with minimal verification, while lower-confidence extractions receive more extensive checking, maintaining both speed and accuracy
Data Source
AI summary
A system, computer program product, and method are provided to automate a framework for knowledge graph based persistence of data, and to resolve temporal changes and uncertainties in the knowledge graph. Natural language understanding, together with one or more machine learning models (MLMs), is used to extract data from unstructured information, including entities and entity relationships. The extracted data is populated into a knowledge graph. As the KG is subject to change, the KG is used to create new and retrain existing machine learning models (MLMs). Weighting is applied to the populated data in the form of veracity value. Blockchain technology is applied to the populated data to ensure reliability of the data and to provide auditability to assess changes to the data.


