Discrepancy Curator for Cognitive Computing Corpus

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional search engines are ineffective in interpreting unstructured data and detecting discrepancies between documents, leading to user confusion when finding relevant information, as they rely on keyword-based searches without understanding context or grammar.

Innovation Solution

A cognitive computing system that employs Natural Language Processing (NLP) and machine learning to analyze unstructured data, detect discrepancies, and present ranked answers by confidence, using a document ingestion pre-processor to flag and resolve contradictions within the corpus, enabling the detection and reconciliation of conflicting information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional keyword-based search engines are used to find information in unstructured data, then the search process is simple and fast, but the system cannot interpret context, detect disagreements between documents, or determine correctness of findings

Engineering Contradiction:
Improveaccuracy of information retrievalVSAvoidcomplexity of search system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary component (discrepancy detection module) that sits between the search engine and the user. This module automatically compares multiple search results, detects contradictions using natural language processing, and presents resolved information. The intermediary handles the complexity of interpretation and conflict detection, allowing the search engine itself to remain relatively simple while achieving reliable, accurate results.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual information verification (mechanical human process) with automated natural language processing and discrepancy detection algorithms. Instead of requiring users to manually compare multiple search results and determine correctness, the system uses computational methods to automatically interpret unstructured data, detect disagreements, and resolve conflicts, thereby improving reliability without requiring complex user intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If multiple documents are retrieved by keyword search, then more information is available to the user, but the user must manually determine which documents are correct and resolve contradictions

Engineering Contradiction:
Improvecompleteness of informationVSAvoidease of use
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent implements a self-service mechanism where the system automatically performs discrepancy detection and resolution without requiring user intervention. The discrepancy detection module autonomously compares retrieved documents, identifies contradictions, and presents resolved information. This allows the system to maintain completeness of information from multiple sources while eliminating the manual effort users would otherwise need to spend verifying and reconciling conflicting information.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If cognitive computing systems with NLP and discrepancy detection are implemented, then accuracy and reliability of information retrieval improve, but the system complexity and processing requirements increase

Engineering Contradiction:
Improveprecision of information accuracyVSAvoidcomplexity of processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the information retrieval system into distinct functional modules: a search engine for retrieving documents, a natural language processing module for interpreting content, and a discrepancy detection module for comparing and resolving contradictions. Each module performs a specific function with optimized complexity, rather than requiring the entire system to handle all tasks simultaneously. This segmentation allows high precision in discrepancy detection while managing overall system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11308143B2Discrepancy curator for documents in a corpus of a cognitive computing system
Publication Date: 2022.04.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11308143B2 patent drawing
  • US11308143B2 patent drawing
  • US11308143B2 patent drawing

AI summary

Curation of a corpus of a cognitive computing system is performed by reporting to a user a cluster model of a parse tree structure of discrepancies and corresponding assigned confidence factors detected between at least a portion of a first electronic document and a second or more electronic documents in the information corpus. Responsive to a selection by the user of a discrepancy cluster model, drill-down details regarding the discrepancy are returned to the user, for subsequent user selection of an administrative action option for handling the detected discrepancy to curate the information corpus.