Text Processing System for Automated Theme Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Companies face challenges in efficiently extracting themes from large volumes of textual data related to customer service interactions, which is time-consuming and often lacks comprehensive analytics.

Innovation Solution

A text processing system that cleanses textual data, extracts phrases, clusters them into hierarchical themes, and generates a graphical representation for visualization, utilizing techniques such as language detection, spell checking, and embedding representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual review of customer service interactions is performed, then data analysis can be conducted, but it is extremely time-consuming and inefficient

Engineering Contradiction:
Improvedata analysis efficiencyVSAvoidmanual review time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical review processes with automated computational systems including text cleansing modules, phrase extraction algorithms, clustering mechanisms, and visualization tools. This substitution transforms the time-consuming manual analysis into efficient automated processing while maintaining comprehensive data examination capabilities

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service automated theme extraction and visualization generation without requiring manual intervention. The text processing system automatically cleanses data, extracts meaningful phrases, clusters them into themes, and generates visual representations, allowing the data to analyze itself rather than requiring human reviewers

Inventive Principle:
Principle #25Self-service

2Loss of information

If comprehensive analytics are generated from textual data, then valuable insights can be obtained, but the complexity of processing and analyzing the data increases

Engineering Contradiction:
Improveanalytics comprehensivenessVSAvoidprocessing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the complex analysis process into distinct modular components: text cleansing module, phrase extraction module, clustering module, and visualization module. Each module handles a specific aspect of the analysis, reducing overall system complexity while enabling comprehensive analytics through coordinated operation of these specialized components

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary processing layers including text cleansing that transforms raw text into standardized format, and phrase extraction that serves as a bridge between raw text and thematic clusters. These intermediaries simplify the complexity by creating structured intermediate representations that are easier to process and analyze

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If text data is cleansed through multiple processing steps, then extraction accuracy is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvetheme extraction accuracyVSAvoidtext processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary text cleansing operations including removing non-ASCII characters, expanding contractions, removing numbers and punctuation, and lemmatizing text before phrase extraction. This preliminary processing prepares the data in advance, improving subsequent extraction accuracy while establishing a efficient processing foundation that reduces overall computational burden

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250181834A1Extracting themes from textual data
Publication Date: 2025.06.05 AMERICAN EXPRESS (INDIA) PTE LTD
  • US20250181834A1 patent drawing
  • US20250181834A1 patent drawing
  • US20250181834A1 patent drawing

AI summary

Disclosed herein are apparatus, system, method, and computer-readable medium aspects for extracting themes from textual data. Textual data is initially cleansed from its submitted form into a simplified form for improved accuracy of topic extraction. From the cleansed text, phrases are extracted. Embeddings of the phrases are then determined so that similarities can be identified between different phrases within the text. Using these embodiments, clustering is performed on the embeddings to reveal the topics included within the text submission, as well as their frequency and relationship to one another. This clustering processing can be repeated at multiple levels of granularity for improved accuracy. Based on an analysis of the resulting clusters, a graphical representation of the clusters at the various levels is generated to provide an easy-to-understand indication of the body of text and the topics and themes included therein.