Text Classification via Term Mapping and Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies struggle to accurately identify and classify the subject matter of text, especially when text is fragmented across multiple documents with limited text in each document, and machine-based manipulations are inadequate for summarizing content.
Innovation Solution
The use of identified groups and a machine-learning model to classify text documents, where the groups can be curated by individuals or another machine-learning classifier, and the model can be retrained based on adjustments to these groups, allowing for better summarization and training of machine-learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine-based manipulations are used to analyze text, then analysis speed is improved, but subject matter identification accuracy deteriorates
Solution Approach 1:
The patent introduces term frequency weights and similarity scores as intermediary metrics between raw text and classification results. These intermediaries bridge the gap between automated processing and accurate subject matter identification by quantifying term importance and document similarity in a structured manner
Solution Approach 2:
The patent replaces traditional keyword-matching mechanical systems with a machine learning classification model that uses trained parameters and algorithms to automatically identify subject matter, achieving both speed and accuracy through computational intelligence
2Ease of operation
If text is divided into multiple separate documents, then information organization is improved, but summarization capability deteriorates
Solution Approach 1:
The patent creates a universal classification model that can process both individual documents and collections of documents using the same trained parameters. The model calculates similarity scores across the entire collection, enabling it to identify common themes and generate summaries that work effectively at both individual and collection levels
Solution Approach 2:
The patent merges information from multiple documents by calculating aggregate term frequencies and similarity scores across the entire document collection. This combining approach allows the system to identify subject matter that spans multiple documents while maintaining the organizational benefits of separate document structures
Data Source
AI summary
A method, apparatus, and computer-readable medium are described that identify subject matter of text using identified groups and a machine-learning model. Using the combination of the identified subject matter and the machine-learning model, classifications may be adjusted over time. Based on the adjusted classifications, the machine-learning model may be retrained to better classify previously unclassified text. One or more benefits may include better summarization of text documents and/or better training of machine-learning models that are then used to assist in the summarization of the text documents. The resulting classifications may be used to improve resource allocations for future tasks.


