Unstructured Response Extraction via Frequency Deviation Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for analyzing unstructured text responses are time-consuming and difficult to interpret, especially when comparing large numbers of responses from different populations, and there is a need for efficient assessment and comparison of discrete units of text.
Innovation Solution
A system and method that generates reference data from a first set of unstructured comments, identifies significant words in a second set of unstructured comments, determines their frequency of occurrence, and visualizes these words on a graphical user interface, highlighting deviations from the first set's frequency, allowing users to select and view additional data for each significant word.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If text mining techniques are used to analyze unstructured responses, then information extraction capability is improved, but time consumption and interpretation difficulty increase
Solution Approach 1:
The patent extracts only the most significant words from unstructured text responses using statistical analysis comparing frequency against reference data. This extraction approach retrieves key information without requiring complete manual reading or complex text mining, thus improving information extraction efficiency while reducing time consumption.
Solution Approach 2:
The system creates a simplified visual representation (word cloud) that copies only the essential characteristics of the unstructured data - the significant words and their relative frequencies. This visual copy allows rapid interpretation without processing the full text volume, resolving the contradiction between thorough information extraction and time efficiency.
2Measurement precision
If all unstructured responses are analyzed in detail, then interpretation accuracy is improved, but productivity decreases
Solution Approach 1:
Instead of uniformly analyzing all text with the same depth, the system applies local quality by identifying and highlighting only the significant words that contribute most to interpretation accuracy. The visual display concentrates analytical resources on key terms while ignoring redundant content, maintaining interpretation precision while dramatically improving productivity.
Solution Approach 2:
The patent performs partial action by analyzing only the portions of unstructured text that contain significant words, rather than processing every word equally. This selective approach achieves sufficient interpretation accuracy for decision-making while avoiding the excessive time investment required for complete detailed analysis of all responses.
3Loss of information
If hybrid approaches combining structured and unstructured questions are used, then information depth is improved, but response analysis complexity increases
Solution Approach 1:
The patent segments the analysis process into distinct components: collecting structured responses, collecting unstructured responses, generating reference data from structured responses, identifying significant words in unstructured responses, and creating visual displays. This segmentation reduces analysis complexity by breaking down the hybrid approach into manageable, automated steps while preserving the information depth benefits of combining both question types.
Solution Approach 2:
The system introduces an intermediary processing layer that automatically generates reference data from structured responses and uses this reference to identify significant words in unstructured responses. This intermediary automation mediates between the two data types, reducing manual analysis complexity while maintaining the complementary information depth that hybrid approaches provide.
Data Source
AI summary
In one embodiment, the invention can be a method for assessing unstructured comments, the method including providing reference data generated from a first set of unstructured comments from a first group; receiving a second set of unstructured comments from a second group; identifying a significant word within each unstructured comment of the second set of unstructured comments; for each significant word identified within the second set of unstructured comments, determining a frequency of occurrence of the significant word; and generating a visualization including a portion of the identified significant words, wherein for each visualized significant word, a first aspect of an appearance of the significant word is based on an extent to which the frequency of occurrence deviates from a frequency of occurrence of the significant word in the first set of unstructured comments.


