Deterministic Query Processing via Concept Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cognitive systems in natural language processing are inherently non-deterministic, leading to inconsistent results due to susceptibility to input information and new machine learning models, which can introduce errors in data extraction and output.

Innovation Solution

An AI platform with a request manager and cluster manager is used to identify lexical answer types and concepts, forming clusters of documents based on these elements and sorting them to provide deterministic query results, including representative passages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If cognitive systems process natural language based on knowledge acquired from training data, then the system can provide intelligent responses, but the results become incorrect or inaccurate due to language peculiarities and errors in training data

Engineering Contradiction:
Improvenatural language processing capabilityVSAvoidaccuracy of processing results
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary verification process between the cognitive system's natural language processing and the final output. This includes using multiple independent processing paths, cross-validating results against established knowledge bases, and implementing a review layer that checks for consistency and accuracy before presenting results, thereby mediating between the flexible NLP capability and reliable output

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback mechanisms where processing results are continuously evaluated against ground truth data and user corrections. Error patterns are fed back into the training process to improve future accuracy, and real-time feedback allows the system to adjust its processing approach based on the reliability of incoming data, creating a closed-loop system that improves both adaptability and reliability

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If new machine learning models are deployed to improve system capabilities, then the system can handle new tasks, but there is no guarantee that the system will extract the same entities as done previously, leading to inconsistent results

Engineering Contradiction:
Improvemodel update capabilityVSAvoidconsistency of data extraction
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent implements a dynamic model management system where multiple versions of machine learning models are maintained and can be selectively activated. When new models are deployed, the system dynamically adjusts by comparing outputs across model versions, using ensemble methods to combine results, and implementing gradual rollouts that allow for consistency checking against established extraction patterns, thereby maintaining stability while enabling adaptation

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Before deploying new machine learning models, the system performs preliminary validation by testing them against a held-out validation set of known correct extractions. This preliminary action ensures that new models meet consistency thresholds before being activated, and establishes baseline expectations for entity extraction that the new models must satisfy, preventing inconsistency from propagating into production

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If error is introduced through a document, then the document may provide new information, but the system extracts incorrect data and provides incorrect output

Engineering Contradiction:
Improvedata processing flexibilityVSAvoiderror propagation
Core Design Contradiction:
Adaptability or versatilityVSObject-generated harmful factors

Solution Approach 1:

The patent implements preliminary anti-action by introducing error detection and prevention mechanisms before errors can propagate through the system. This includes pre-processing validation that checks documents for known error patterns, cross-referencing extracted data against multiple sources to detect inconsistencies, and implementing confidence thresholding that prevents low-confidence extractions from being processed further, thereby countering potential errors before they spread

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The system converts the potential harm of errors into benefit by using error detection opportunities to improve overall system robustness. When errors are detected in processed documents, the system learns from these instances by adding them to training datasets with correction labels, uses error patterns to refine validation rules, and leverages failed extractions to identify and fix underlying issues in the processing pipeline, turning harmful errors into opportunities for improvement

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS11562029B2Dynamic query processing and document retrieval
Publication Date: 2023.01.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11562029B2 patent drawing
  • US11562029B2 patent drawing
  • US11562029B2 patent drawing

AI summary

Embodiments relate to an intelligent computer platform to identify a lexical answer type (LAT), a first concept relevant to the received request and a second concept related to the identified first concept. The LAT, together with the first and second concepts are utilized to create a first and second cluster. Documents are selectively populated into the clusters. The clusters are subject to sorting based on a relevancy protocol.