Data Pattern Classification for Mixed Database Columns
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In databases with mixed columns containing ununified information types, such as addresses and telephone numbers in remarks columns, existing methods struggle to correctly associate data between databases due to human error or unspecified input styles, leading to interference during database combination.
Innovation Solution
A data-pattern classification method that extracts a branch column using heuristic and statistical techniques to group patterns based on overlapping character strings, allowing for the identification of shared information types between databases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If information is inputted without unified expression in remarks columns, then flexibility in data entry is improved, but data association accuracy deteriorates
Solution Approach 1:
The patent changes the parameter of data representation by extracting multiple patterns from mixed columns and grouping them into equivalence classes. Instead of requiring unified input expressions, the system transforms diverse input formats into categorized pattern groups, enabling accurate data association while maintaining input flexibility.
Solution Approach 2:
The patent introduces an intermediary classification system that acts as a mediator between diverse input formats and data association requirements. By extracting patterns and creating equivalence classes as intermediate representations, the system enables accurate matching without requiring unified input expressions.
2Ease of manufacture
If conventional rule-based methods are used for data classification, then simple input patterns are handled, but complex mixed information types cannot be properly classified
Solution Approach 1:
The patent segments mixed columns into multiple distinct patterns by extracting characteristic substrings and grouping similar entries. This segmentation divides complex mixed information into manageable pattern equivalence classes, enabling the classification system to handle diverse information types systematically while maintaining methodological simplicity.
Solution Approach 2:
The patent extracts multiple patterns from each mixed column entry, potentially more than strictly necessary. By identifying and grouping multiple characteristic patterns, the system ensures comprehensive coverage of mixed information types while maintaining a straightforward classification approach.
Data Source
AI summary
To easily find columns containing information to be shared between databases found when the databases (DBs) are combined.A data-pattern classification method for a database containing character strings regularly overlapping between the input values of mixed columns and information inputted in another specific column, the mixed column including types of information, the method including: extracting means for extracting a branch column for changing the type of information to be inputted to the mixed column, the branch column being extracted by one of a heuristic first technique and a second technique, the first technique using timing to change the pattern of the mixed column in the database including the mixed column, the pattern referring to a column where a character string overlaps the input value of the mixed column, the second technique using a statistical technique of a likelihood test; and classification to obtain the number of types of information stored in the mixed columns, by grouping, according to information indicated by the patterns, the patterns obtained from the mixed columns based on the extracted branch column.


