Automated Text Classification via Precision-Weighted Query Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for classification and tagging of textual data are inefficient due to the need for manual steps and lack of automation in identifying services based on machine-readable text.
Innovation Solution
A system and method for automatically generating queries based on extracted features from a corpus of documents, calculating precision scores, and selecting queries to accurately identify and tag services in machine-readable text, using a processor to generate queries, calculate precision scores, and apply them to machine-readable text to assign labels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual steps are used for classification and tagging of textual data, then precision in service identification can be maintained through human judgment, but productivity is reduced due to time-consuming manual processes
Solution Approach 1:
The system performs automatic query generation, precision scoring, and label assignment without requiring manual intervention. The computational system serves itself by generating queries from extracted features, evaluating their precision scores, and automatically applying labels to textual data, eliminating the need for human operators to perform these repetitive classification tasks manually.
Solution Approach 2:
The patent replaces manual human operations with an automated computational system that uses machine learning algorithms, precision score calculations, and automated query generation. The mechanical process of manual text analysis and labeling is substituted with electronic data processing, algorithmic query evaluation, and automated decision-making systems that can process large volumes of text rapidly.
2Productivity
If automated query generation is implemented, then productivity in text classification is improved, but device complexity increases due to multiple processing steps including feature extraction, query generation, and precision scoring
Solution Approach 1:
The system divides the text classification task into distinct modular components: feature extraction module, query generation module, precision scoring module, and label assignment module. Each module performs a specific function and can be independently optimized, maintained, and scaled. This segmentation allows the complex system to be managed through separate, well-defined processing stages rather than a monolithic approach.
Solution Approach 2:
The automated query generation system serves multiple functions simultaneously: it extracts features from text, generates relevant queries, evaluates precision scores, and assigns labels. This multi-functional approach consolidates what could be separate systems into a unified platform that handles the entire classification pipeline, reducing the need for multiple independent tools and processes.
3Measurement precision
If precision scoring is applied to generated queries, then measurement precision in service identification is improved, but loss of time occurs during the scoring and selection process
Solution Approach 1:
The system calculates precision scores for all generated queries but selectively applies only those queries that meet a predetermined precision threshold to the textual data. This partial action approach avoids the time cost of evaluating and applying every possible query, while still ensuring that only high-precision queries are used for labeling, thus maintaining measurement precision without unnecessary time expenditure on low-quality queries.
Data Source
AI summary
Provided herein are systems, methods and computer readable media for classification and tagging of textual data. An example method may include accessing a corpus comprising a plurality of documents, each document having one or more labels indicative of services offered by a merchant, generating a query based on extracted features and the documents, generating a precision score for at least a portion of the generated query and selecting a subset of the generated queries based on an assigned precision score satisfying a precision score threshold, the selected subset of the generated queries configured to provide an indication of one or more labels to be applied to machine readable text. A second example method, utilized for tagging machine readable text with unknown labels, may include assigning a label to textual portions of the machine readable text based on results of the application of the queries.


