Data Pattern Classification for Mixed Database Columns

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In databases with mixed columns containing ununified information types, such as addresses and telephone numbers in remarks columns, existing methods struggle to correctly associate data between databases due to human error or unspecified input styles, leading to interference during database combination.

Innovation Solution

A data-pattern classification method that extracts a branch column using heuristic and statistical techniques to group patterns based on overlapping character strings, allowing for the identification of shared information types between databases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If information is inputted without unified expression in remarks columns, then flexibility in data entry is improved, but data association accuracy deteriorates

Engineering Contradiction:
Improveflexibility in data entryVSAvoiddata association accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of data representation by extracting multiple patterns from mixed columns and grouping them into equivalence classes. Instead of requiring unified input expressions, the system transforms diverse input formats into categorized pattern groups, enabling accurate data association while maintaining input flexibility.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary classification system that acts as a mediator between diverse input formats and data association requirements. By extracting patterns and creating equivalence classes as intermediate representations, the system enables accurate matching without requiring unified input expressions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If conventional rule-based methods are used for data classification, then simple input patterns are handled, but complex mixed information types cannot be properly classified

Engineering Contradiction:
Improvesimplicity of classification methodVSAvoidcapability to handle mixed information types
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent segments mixed columns into multiple distinct patterns by extracting characteristic substrings and grouping similar entries. This segmentation divides complex mixed information into manageable pattern equivalence classes, enabling the classification system to handle diverse information types systematically while maintaining methodological simplicity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts multiple patterns from each mixed column entry, potentially more than strictly necessary. By identifying and grouping multiple characteristic patterns, the system ensures comprehensive coverage of mixed information types while maintaining a straightforward classification approach.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12141239B2Classification method of data pattern and classification system of data pattern
Publication Date: 2024.11.12 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12141239B2 patent drawing
  • US12141239B2 patent drawing
  • US12141239B2 patent drawing

AI summary

To easily find columns containing information to be shared between databases found when the databases (DBs) are combined.A data-pattern classification method for a database containing character strings regularly overlapping between the input values of mixed columns and information inputted in another specific column, the mixed column including types of information, the method including: extracting means for extracting a branch column for changing the type of information to be inputted to the mixed column, the branch column being extracted by one of a heuristic first technique and a second technique, the first technique using timing to change the pattern of the mixed column in the database including the mixed column, the pattern referring to a column where a character string overlaps the input value of the mixed column, the second technique using a statistical technique of a likelihood test; and classification to obtain the number of types of information stored in the mixed columns, by grouping, according to information indicated by the patterns, the patterns obtained from the mixed columns based on the extracted branch column.