Text Mining System for Unifying Part Name Variants
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the airline industry, data records containing free-form text often include misspellings, acronyms, and variations in part names, making it challenging to consistently identify and analyze part names, which can lead to inaccurate or incomplete data and operational issues.
Innovation Solution
A system that unifies terms of interest, such as part names, using a combination of domain knowledge, linguistic analysis, and machine learning algorithms to identify and normalize variants, enabling more accurate data analytics and improving the quality of information extracted from electronic documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If free-form text is used to populate data records, then ease of data entry and flexibility are improved, but consistency in identifying part names deteriorates due to misspellings, acronyms, and variations
Solution Approach 1:
The patent introduces an intermediary text mining system that sits between the free-form text input and the data analysis processes. This system includes a database of part names, a text mining module that scans data records, and a unifying module that resolves variations. The intermediary translates diverse free-form expressions into standardized part names, maintaining ease of data entry while achieving consistency in identification.
Solution Approach 2:
The patent replaces manual mechanical processes of data standardization with automated electronic text mining algorithms. Instead of manually reviewing and standardizing each data record, the system uses computer-based text mining modules with machine learning algorithms to automatically identify, extract, and unify part names across thousands of data records, significantly improving efficiency and consistency.
2Adaptability or versatility
If multiple variants of part names are allowed in data records, then adaptability and versatility are improved, but data analytics accuracy deteriorates due to inability to consistently identify the same part
Solution Approach 1:
The patent merges multiple variants of part names into unified representations through the unifying module. The system collects all variants identified by the text mining module, analyzes their relationships using similarity algorithms, and consolidates them into single standardized part names. This merging process maintains adaptability to different expressions while ensuring that all variants are treated as the same part in data analytics.
Solution Approach 2:
The patent changes the parameter of part name representation from multiple possible variants to a single standardized form. The unifying module transforms diverse expressions (acronyms, misspellings, different formats) into consistent standardized part names from the database, changing the state of the data from variable to uniform, thereby improving analytics accuracy while preserving the ability to accept various input forms.
3Measurement precision
If manual review and standardization of part names is performed, then measurement precision and reliability are improved, but productivity and time consumption deteriorate
Solution Approach 1:
The patent implements a self-service system where the text mining module automatically scans data records, identifies part names, and the unifying module automatically resolves variations without human intervention. The system uses machine learning algorithms that continuously learn from the data, improving their performance over time. This automation eliminates the need for manual review while maintaining high accuracy, thereby preserving both precision and productivity.
Solution Approach 2:
The patent substitutes manual mechanical review processes with automated electronic text mining and machine learning systems. The computer-based algorithms process thousands of data records rapidly, identifying and standardizing part names automatically. This electronic substitution of manual processes maintains measurement precision through sophisticated algorithms while dramatically improving productivity by processing vast amounts of data in fractions of the time required for manual review.
4Measurement precision
If sophisticated text mining algorithms are implemented, then measurement precision and reliability are improved, but device complexity increases
Solution Approach 1:
The patent segments the text mining system into distinct functional modules: a text mining module that scans and extracts potential part names, a unifying module that resolves variations, and a database of standardized part names. Each module performs a specific function with well-defined inputs and outputs. This segmentation reduces system complexity by creating manageable, independent components that can be developed, tested, and maintained separately while achieving high measurement precision through their coordinated operation.
Data Source
AI summary
A method is provided for analyzing and interpreting a dataset composed of electronic documents including free-form text. The method includes unifying terms of interest in the collection of terms of interest to identify variants of the terms of interest. This includes identifying candidate variants of a term of interest based on semantic similarity between the term of interest and other terms in the database, determined using an unsupervised machine learning algorithm. Linguistic features and contextual features of the term of interest and its candidate variants are extracted, at least the contextual features being extracted using the unsupervised machine learning algorithm. And a supervised machine learning algorithm is used with the linguistic features and contextual features to identify variants of the term of interest from the candidate variants, such as for application to generate features of the documents for data analytics performed thereon.


