Conceptual Sets Logic for Disambiguating Polysemous Terms in Data Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search methods for large data corpora rely on text-based keyword searches, which often yield imprecise results due to the inability to determine the meaning of words, especially for polysemous terms, leading to inefficient information retrieval and inaccurate translations.

Innovation Solution

A system that uses a Conceptual Index Dictionary to assign Conceptual Numerical Indexes (CNIs) to words based on their meanings and grammatical roles, grouping them using Conceptual Sets Logic (CET Logic) to form Conceptual Sets, which are then stored and queried for precise information extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If text-based keyword searches are used for large data corpora, then search coverage is improved, but measurement precision deteriorates due to inability to determine word meanings

Engineering Contradiction:
Improvesearch coverageVSAvoidsearch precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent introduces Conceptual Sets as an intermediary layer between keywords and documents. Instead of directly matching keywords to documents, the system first converts keywords into Conceptual Sets using a Conceptual Index Dictionary, then matches these Conceptual Sets against Conceptual Sets extracted from documents. This intermediary representation enables semantic understanding while maintaining comprehensive search coverage.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the search parameter from simple text keywords to structured Conceptual Sets with numerical identifiers. By changing the representation parameter from raw text to standardized Conceptual Sets with CNIs, the system achieves both comprehensive coverage (through systematic indexing) and precise matching (through structured comparison).

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If conventional text-based search methods are used, then ease of operation is improved, but loss of information increases due to imprecise results

Engineering Contradiction:
Improvesearch simplicityVSAvoidinformation retrieval accuracy
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent performs preliminary action by pre-processing the entire corpus into Conceptual Sets and storing them in a structured format before actual search queries are executed. The Conceptual Index Dictionary is built in advance, and document representations are converted to Conceptual Sets beforehand. This preliminary structuring enables fast, accurate retrieval without complicating the user's search operation.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If polysemous terms are searched using keyword methods, then search coverage is improved, but manufacturing precision deteriorates due to word sense disambiguation failures

Engineering Contradiction:
Improvesearch coverageVSAvoidconceptual matching accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments the search process into distinct phases: keyword-to-Conceptual-Set conversion, Conceptual-Set matching, and result extraction. By segmenting the handling of polysemous terms into separate Conceptual Sets with different CNIs, the system can process multiple meanings systematically while maintaining precise control over which meaning is applied in each context.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9588963B2System and method of grouping and extracting information from data corpora
Publication Date: 2017.03.07 ENTIGENLOGIC LLC
  • US9588963B2 patent drawing
  • US9588963B2 patent drawing
  • US9588963B2 patent drawing

AI summary

A system for annotating words of a data corpus based upon their particular concept and their corresponding grammatical sense with Conceptual Numerical Identifiers (CNIs) from a Conceptual Dictionary, pairing the words based on conceptual inter-relating network (CIRN) rules, and determining if a selected plurality of paired words are grammatically, syntactically, and linguistically correct by matching CNIs from each pair of words.