Lexical Enrichment for Structured Data Search Precision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing keyword search technologies face challenges in effectively searching structured and semi-structured data due to poor precision and recall, lack of support for relevancy scoring, inability to represent compound concepts, and incoherent linguistic representation, especially when dealing with coded data and mixed structured/unstructured content.

Innovation Solution

Lexical enrichment techniques are employed to generate linguistically complete and well-formed content by utilizing schema, metadata, and code definitions to create predicate phrases that align with user queries, leveraging full text search technology and ontological categorizations to enhance search precision and relevance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If ad hoc heuristics are used to extract semantic content from schemas, then keyword search can be enabled on structured data, but precision and recall deteriorate due to naive assumptions about schema properties

Engineering Contradiction:
Improvekeyword search capabilityVSAvoidsearch precision and recall
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary layer of lexical enrichment that transforms structured data into linguistically well-formed content. This intermediary process uses schema information and ontology mappings as mediators to bridge the gap between structured data formats and natural language search queries, enabling accurate keyword matching without relying on naive heuristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the representational parameters of structured data by transforming it into linguistically enriched content that mirrors natural language semantics. This parameter transformation involves converting schema-based representations into lexically complete forms that can be properly indexed and searched by full-text search engines, thereby improving search precision.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If keyword search is applied to structured data, then search functionality is provided, but relevancy scoring is not supported requiring new mechanisms instead of leveraging full text search engines

Engineering Contradiction:
Improvesearch functionalityVSAvoidrelevancy calculation mechanisms
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent creates a copy of the structured data in the form of lexically enriched content that replicates the semantic information in a format suitable for full-text search engines. This copying process preserves all relevant information while transforming it into a format that leverages existing, highly evolved full-text search algorithms for relevancy scoring, avoiding the need to create new scoring mechanisms.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The lexical enrichment acts as an intermediary that translates structured data into a format compatible with full-text search engines. This intermediary layer enables the use of mature relevancy scoring algorithms from full-text search engines without requiring custom mechanisms, thereby reducing system complexity while maintaining search functionality.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If structured queries are used for searching, then precise data retrieval is achieved, but compound concepts cannot be represented and keyword order/proximity cannot be specified

Engineering Contradiction:
Improvedata retrieval precisionVSAvoidrepresentation of compound concepts
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the search query into individual lexical terms that correspond to schema elements. By segmenting the query this way, the system can independently match each term against the lexically enriched data while maintaining the ability to express compound concepts through combinations of segmented terms, thereby preserving both precision and versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a lexical dimension to structured data representation by enriching it with linguistic content. This dimensional transformation enables the data to be searched using natural language keywords while maintaining the structured retrieval precision, allowing representation of compound concepts through lexical combinations without sacrificing data retrieval accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Quantity of substance

If coded data is used in databases, then data efficiency is improved, but lexical incoherence prevents effective keyword search

Engineering Contradiction:
Improvedata efficiencyVSAvoidkeyword search capability
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent introduces lexical enrichment as an intermediary layer that translates coded data into linguistically coherent representations. This intermediary process maintains the efficiency of coded data storage while adding lexical content that enables effective keyword search, thereby resolving the contradiction between data efficiency and search capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary lexical enrichment of coded data during the data processing stage. By preparing the lexical representations in advance, the system enables effective keyword search without requiring real-time translation during query processing, thus maintaining data efficiency while improving search operationability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9418151B2Lexical enrichment of structured and semi-structured data
Publication Date: 2016.08.16 RAYTHEON CO
  • US9418151B2 patent drawing
  • US9418151B2 patent drawing
  • US9418151B2 patent drawing

AI summary

Generally discussed herein are systems and methods for lexically enriching structured and semi-structured data. In one or more embodiments, a method can include receiving a code, lexicalizing the code, lexically combining the lexicalized code with a lexical descriptor, and sending the lexical combination to a keyword database.