Conversational Lexicon Analyzer for Entity Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language processing techniques are limited in identifying the entity associated with conversational data, as they focus on determining the subject matter or content without providing insight into the entity providing or interacting with the data.

Innovation Solution

A conversational lexicon analysis system that retrieves and analyzes conversational data from various sources to generate training language maps, which are used to identify the entity by comparing received data to a corpus of training maps, determining the entity with the highest confidence value based on lexical features and colloquial terms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional language processing techniques are used to determine subject matter of language data, then the subject matter identification is achieved, but the ability to identify the entity associated with the language data is limited

Engineering Contradiction:
Improveentity identification accuracyVSAvoidlanguage processing capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments language data into atomic units (individual words or phrases) and creates separate language maps for different entities. Each language map captures the unique lexical patterns, colloquial terms, and phrasing characteristics specific to that entity. This segmentation allows the system to analyze and identify entities based on their distinct linguistic fingerprints rather than treating all language data uniformly.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary action by building comprehensive training language maps for multiple entities before actual entity identification occurs. These training language maps are constructed in advance using collected language data from various sources, capturing entity-specific lexical features, colloquialisms, and communication patterns. When new language data arrives, the system can immediately compare it against the pre-built training maps to identify the source entity without needing to learn patterns in real-time.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If language data is analyzed only for subject matter determination, then content classification is achieved, but insight about the entity providing or interacting with the data is lost

Engineering Contradiction:
Improveentity information retentionVSAvoidlanguage processing system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent creates a universal language map structure that can represent multiple entities simultaneously. Each language map contains entity identifiers and captures lexical features that are specific to particular entities while using a consistent framework applicable to all entities in the system. This multi-functional approach allows the same processing mechanism to handle identification of different entities (people, organizations, groups) without requiring separate specialized systems for each entity type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The language map serves as an intermediary structure between raw language data and entity identification results. Instead of directly analyzing language data to determine subject matter only, the system uses language maps as intermediate representations that capture entity-specific linguistic patterns. These maps act as mediators that preserve entity information while enabling comparison and identification, thus preventing loss of entity information without requiring direct complex analysis of every language input.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8527269B1Conversational lexicon analyzer
Publication Date: 2013.09.03 YAHOO ASSETS LLC
  • US8527269B1 patent drawing
  • US8527269B1 patent drawing
  • US8527269B1 patent drawing

AI summary

A system and a method for analyzing conversational data comprising colloquial or informal terms and having an informal structure. A corpus of training language maps, each associated with an entity, is generated from conversational data retrieved from sources previously associated with entities. Subsequently received conversational data is processed to generate a conversational language map which is compared to a plurality of the stored training language maps. A confidence value is generated describing the similarity of the conversational language map to each of the plurality of the stored training language maps. The entity associated with the training language map having the highest confidence value is then associated with the conversational language map.