Text Analytics System for Unstructured Data Group Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text analytics techniques fail to analyze the context of unstructured data across multiple groups of records, missing significant slang or jargon terms and requiring user input, which can lead to overlooking important terms or expressions.
Innovation Solution
A method and system for text analytics that filters unstructured data into groups, calculates the proportion of term occurrences, and compares these proportions to identify statistically significant differences, allowing for the analysis of alphanumeric characters and symbols, including slang or jargon, without requiring user-input terms of interest.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional text analytics techniques are used to analyze unstructured data within a single record, then the analysis is simple and requires minimal user input, but the technique fails to analyze the context of data across multiple groups of records
Solution Approach 1:
The patent segments the unstructured data into multiple groups of records and analyzes each group separately while comparing them across groups. This segmentation allows the system to maintain simplicity within each group analysis while capturing contextual information across groups through comparison, resolving the contradiction between simple analysis and contextual understanding.
Solution Approach 2:
The patent adds a new dimension of analysis by introducing group-level comparison alongside record-level analysis. Instead of analyzing data in isolation or only within single records, the system now operates in a multi-dimensional space that includes both individual record context and cross-group contextual relationships, enabling comprehensive analysis without excessive complexity.
2Measurement precision
If Boolean search techniques are used requiring user input of terms of interest, then the search is precise for known terms, but the user may overlook new slang expressions or un-encountered terms
Solution Approach 1:
The patent implements self-service by automatically identifying terms of interest through statistical analysis of term frequencies across groups without requiring user input. The system serves itself by detecting patterns, slang, and significant terms autonomously through proportion comparison, eliminating the need for users to pre-define search terms while maintaining precision through statistical significance testing.
Solution Approach 2:
The patent employs feedback mechanisms where the system continuously analyzes term frequencies, compares proportions across groups, and identifies significant terms that emerge from the data itself. This feedback loop enables the system to adapt to new slang and expressions dynamically, adjusting its term identification based on actual data patterns rather than relying on pre-programmed search criteria.
3Ease of operation
If user-defined Boolean searches are used, then the search can be tailored to specific known terms, but the system cannot automatically identify statistically significant terms across groups
Solution Approach 1:
The patent enables the system to serve itself by automatically performing statistical analysis to identify significant terms without user input. The system computes term frequencies, calculates proportions across groups, determines statistical significance through proportion comparison, and presents results autonomously, eliminating the need for users to define search criteria while maintaining high measurement precision through rigorous statistical testing.
Solution Approach 2:
The patent changes the parameters of analysis from user-defined Boolean conditions to statistically computed proportions and significance levels. Instead of relying on fixed user inputs, the system dynamically adjusts its analysis parameters based on actual data distributions, computing term frequencies, proportions, and statistical significance metrics to automatically identify meaningful terms across groups.
4Productivity
If analysis is performed within a single record or file, then the processing is fast and simple, but the relative context provided by comparison amongst different groups of records is lost
Solution Approach 1:
The patent segments the data into multiple groups while maintaining efficient processing by analyzing each group independently and then comparing results. This segmentation strategy preserves processing speed through modular analysis while capturing relative context across groups through systematic comparison of term proportions, resolving the contradiction between fast processing and comprehensive contextual analysis.
Solution Approach 2:
The patent performs preliminary actions by pre-processing and organizing data into groups before conducting the analysis. By preparing the data structure in advance and establishing group boundaries beforehand, the system enables efficient processing while maintaining the ability to compare across groups, thus preserving relative context without sacrificing processing speed during the actual analysis phase.
Data Source
AI summary
A method of text analytics includes filtering a plurality of unfiltered records having unstructured data into at least a first group and a second group. The first group and said second group each include at least two records and the first group is different than the second group. The method includes determining a first proportion of occurrence for a term by comparing a first number of records having at least one occurrence of the term in the first group to a first total number of records in the first group, determining a second proportion of occurrence for the term by comparing a second number of records having at least one occurrence of the term in said second group to a second total number of records in the second group, and comparing the first proportion of occurrence to the second proportion of occurrence to yield a resultant comparison occurrence.


