Discrepancy Curator for Cognitive Computing Corpus
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional search engines are ineffective in interpreting unstructured data and detecting discrepancies between documents, leading to user confusion when finding relevant information, as they rely on keyword-based searches without understanding context or grammar.
Innovation Solution
A cognitive computing system that employs Natural Language Processing (NLP) and machine learning to analyze unstructured data, detect discrepancies, and present ranked answers by confidence, using a document ingestion pre-processor to flag and resolve contradictions within the corpus, enabling the detection and reconciliation of conflicting information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional keyword-based search engines are used to find information in unstructured data, then the search process is simple and fast, but the system cannot interpret context, detect disagreements between documents, or determine correctness of findings
Solution Approach 1:
The patent introduces an intermediary component (discrepancy detection module) that sits between the search engine and the user. This module automatically compares multiple search results, detects contradictions using natural language processing, and presents resolved information. The intermediary handles the complexity of interpretation and conflict detection, allowing the search engine itself to remain relatively simple while achieving reliable, accurate results.
Solution Approach 2:
The patent replaces manual information verification (mechanical human process) with automated natural language processing and discrepancy detection algorithms. Instead of requiring users to manually compare multiple search results and determine correctness, the system uses computational methods to automatically interpret unstructured data, detect disagreements, and resolve conflicts, thereby improving reliability without requiring complex user intervention.
2Loss of information
If multiple documents are retrieved by keyword search, then more information is available to the user, but the user must manually determine which documents are correct and resolve contradictions
Solution Approach 1:
The patent implements a self-service mechanism where the system automatically performs discrepancy detection and resolution without requiring user intervention. The discrepancy detection module autonomously compares retrieved documents, identifies contradictions, and presents resolved information. This allows the system to maintain completeness of information from multiple sources while eliminating the manual effort users would otherwise need to spend verifying and reconciling conflicting information.
3Measurement precision
If cognitive computing systems with NLP and discrepancy detection are implemented, then accuracy and reliability of information retrieval improve, but the system complexity and processing requirements increase
Solution Approach 1:
The patent segments the information retrieval system into distinct functional modules: a search engine for retrieving documents, a natural language processing module for interpreting content, and a discrepancy detection module for comparing and resolving contradictions. Each module performs a specific function with optimized complexity, rather than requiring the entire system to handle all tasks simultaneously. This segmentation allows high precision in discrepancy detection while managing overall system complexity through modular architecture.
Data Source
AI summary
Curation of a corpus of a cognitive computing system is performed by reporting to a user a cluster model of a parse tree structure of discrepancies and corresponding assigned confidence factors detected between at least a portion of a first electronic document and a second or more electronic documents in the information corpus. Responsive to a selection by the user of a discrepancy cluster model, drill-down details regarding the discrepancy are returned to the user, for subsequent user selection of an administrative action option for handling the detected discrepancy to curate the information corpus.


