Text Document Organization via Probabilistic Topic Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for organizing electronic text documents are expensive, time-consuming, and inflexible, often failing to accurately classify documents due to limitations in handling polysemy and synonymy, and require manual human review or extensive training for classification algorithms.

Innovation Solution

A content management system that automatically categorizes electronic text documents by user-specified topics without human intervention, identifies emergent topics, and uses probabilistic language models to handle polysemy and synonymy, reducing manual effort and training requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual human review is used to classify electronic text documents, then classification accuracy is improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing the text documents to extract features and build topic models before actual classification. This preparation work enables the classification algorithm to quickly and accurately categorize documents without requiring manual review of each document, thus reducing time consumption while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary classification algorithm that acts as a mediator between the text documents and the final classification results. This algorithm processes documents automatically using trained topic models, serving as an intermediate step that replaces manual human review while preserving classification accuracy through sophisticated natural language processing techniques.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If classification algorithms are used to organize electronic text documents, then time and cost are reduced, but accuracy decreases due to inability to handle polysemy and synonymy

Engineering Contradiction:
Improvetime consumptionVSAvoidclassification accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system applies parameter changes by dynamically adjusting the topic model parameters and classification thresholds based on the specific characteristics of the text documents being processed. This allows the algorithm to adapt to different contexts and accurately handle polysemy and synonymy by changing the parameters that control topic assignment and classification decisions.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent employs composite materials concept by combining multiple processing techniques - including topic modeling, feature extraction, and classification algorithms - into a unified hybrid system. This composite approach integrates the strengths of different methods to accurately handle linguistic complexities like polysemy and synonymy while maintaining efficient automated processing.

Inventive Principle:
Principle #40Composite materials

3Device complexity

If predetermined topics are used for classification, then classification process is simplified, but flexibility and adaptability to emergent topics are reduced

Engineering Contradiction:
Improveclassification process complexityVSAvoidflexibility
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system implements dynamics by making the topic structure adaptable and changeable rather than fixed. The topic models can be dynamically updated and refined based on new data and emerging patterns in the text documents, allowing the classification system to evolve and adapt to new topics while maintaining the simplified structure of predetermined topics for stable classification.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies universality by designing a classification system that serves multiple functions: it handles both predetermined topics and emergent topics, processes various types of text documents, and adapts to different classification requirements. This multi-functional approach maintains simplicity while providing flexibility through the universal topic modeling framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11714835B2Organizing survey text responses
Publication Date: 2023.08.01 QUALTRICS LLC
  • US11714835B2 patent drawing
  • US11714835B2 patent drawing
  • US11714835B2 patent drawing

AI summary

Embodiments of the present disclosure relate generally to organizing electronic text documents. In particular, one or more embodiments comprise a content management system that improves the organization of electronic text documents by intelligently and accurately categorizing electronic text documents by topic. The content management system organizes electronic text documents based on one or more topics, without the need for a human reviewer to manually classify each electronic text document, and without the need for training a classification algorithm based on a set of manually classified electronic text documents. Further, the content management system identifies and suggests topics for electronic text documents that relate to new or emerging topics.