Corpus Gap Probability Modeling for Knowledge Base Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing artificial intelligence systems face challenges in accurately processing natural language due to static structures that fail to account for undefined or undiscovered relationships, leading to potential gaps in knowledge bases and incorrect outputs.

Innovation Solution

A system utilizing a knowledge engine with tools like an organization manager, query manager, and machine learning manager to assess probability gaps in a knowledge base, generating maps and adjusting probabilities for accurate query responses through probabilistic modeling and natural language processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If static structures are used to process natural language, then the system can provide determined outputs for given inputs, but the system cannot identify undefined or undiscovered relationships leading to knowledge gaps

Engineering Contradiction:
Improveaccuracy of query responsesVSAvoidability to handle undefined relationships
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static knowledge base structure into a dynamic probabilistic model. Instead of fixed relationships, the system uses probability values that can be updated and adjusted based on query analysis, allowing the system to adapt to undefined relationships while maintaining deterministic output capabilities.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system introduces probability as a new parameter to represent the state of knowledge gaps. By changing from binary (known/unknown) to continuous probability values, the system can express degrees of certainty and identify areas needing improvement in the knowledge base.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If probabilistic modeling is used to identify knowledge gaps, then the system can detect undefined relationships, but the system complexity increases

Engineering Contradiction:
Improveidentification of knowledge gapsVSAvoidsystem architecture
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent divides the complex task of knowledge base analysis into distinct functional modules: the organization manager for corpus analysis, the query manager for probability calculation, and the director for gap identification. This segmentation reduces overall system complexity by making each component's responsibility clear and manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a probabilistic model as an intermediary layer between the static knowledge base and the query processing system. This intermediary translates static data into dynamic probability information without requiring fundamental changes to the underlying knowledge base structure.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If corpus analysis is conducted to form probabilistic models, then inter-corpora associations can be identified, but the processing time increases

Engineering Contradiction:
Improveprobability assessment accuracyVSAvoidquery processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs corpus analysis and probability model formation in advance, before actual queries are processed. The organization manager pre-analyzes corpora to establish baseline probability relationships, so that during query processing, the system only needs to apply pre-computed models rather than analyzing entire corpora from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system focuses probability analysis on local relevant areas rather than uniformly processing entire corpora. The query manager identifies specific taxonomic areas and corpora relevant to each query context, applying probabilistic modeling only where needed to maintain precision while reducing overall processing time.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11443216B2Corpus gap probability modeling
Publication Date: 2022.09.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11443216B2 patent drawing
  • US11443216B2 patent drawing
  • US11443216B2 patent drawing

AI summary

A system, computer program product, and method are provided to conduct gap probability mapping to predict presence and location of one or more gaps in a corpus. A probabilistic model is formed to represent inter-corpora associations of objects, which the model leverages to process query submissions. As queries are received and processed, the model creates an adjustment. Subject to evaluation, confidence of the adjustment is evaluated, and a response correlated to the confidence is returned.