Conversational Lexicon Analyzer for Entity Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional language processing techniques are limited in identifying the entity associated with conversational data, as they focus on determining the subject matter or content without providing insight into the entity providing or interacting with the data.
Innovation Solution
A conversational lexicon analysis system that retrieves and analyzes conversational data from various sources to generate training language maps, which are used to identify the entity by comparing received data to a corpus of training maps, determining the entity with the highest confidence value based on lexical features and colloquial terms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional language processing techniques are used to determine subject matter of language data, then the subject matter identification is achieved, but the ability to identify the entity associated with the language data is limited
Solution Approach 1:
The patent segments language data into atomic units (individual words or phrases) and creates separate language maps for different entities. Each language map captures the unique lexical patterns, colloquial terms, and phrasing characteristics specific to that entity. This segmentation allows the system to analyze and identify entities based on their distinct linguistic fingerprints rather than treating all language data uniformly.
Solution Approach 2:
The system performs preliminary action by building comprehensive training language maps for multiple entities before actual entity identification occurs. These training language maps are constructed in advance using collected language data from various sources, capturing entity-specific lexical features, colloquialisms, and communication patterns. When new language data arrives, the system can immediately compare it against the pre-built training maps to identify the source entity without needing to learn patterns in real-time.
2Loss of information
If language data is analyzed only for subject matter determination, then content classification is achieved, but insight about the entity providing or interacting with the data is lost
Solution Approach 1:
The patent creates a universal language map structure that can represent multiple entities simultaneously. Each language map contains entity identifiers and captures lexical features that are specific to particular entities while using a consistent framework applicable to all entities in the system. This multi-functional approach allows the same processing mechanism to handle identification of different entities (people, organizations, groups) without requiring separate specialized systems for each entity type.
Solution Approach 2:
The language map serves as an intermediary structure between raw language data and entity identification results. Instead of directly analyzing language data to determine subject matter only, the system uses language maps as intermediate representations that capture entity-specific linguistic patterns. These maps act as mediators that preserve entity information while enabling comparison and identification, thus preventing loss of entity information without requiring direct complex analysis of every language input.
Data Source
AI summary
A system and a method for analyzing conversational data comprising colloquial or informal terms and having an informal structure. A corpus of training language maps, each associated with an entity, is generated from conversational data retrieved from sources previously associated with entities. Subsequently received conversational data is processed to generate a conversational language map which is compared to a plurality of the stored training language maps. A confidence value is generated describing the similarity of the conversational language map to each of the plurality of the stored training language maps. The entity associated with the training language map having the highest confidence value is then associated with the conversational language map.


