Unicode Data Integration for Incompatible Character Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multinational organizations face difficulties in integrating and sharing mission-critical business information due to data being stored in incompatible character sets, which impedes processes like order tracking and unified data views across systems using different character sets.
Innovation Solution
A method and system that convert data from multiple external systems using incompatible character sets into a common Unicode character set, enabling unified processing and integration for business processes such as inventory management, sales forecasting, and customer views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is stored in different character sets for regional systems, then each system can represent its local language characters, but the systems cannot share or integrate mission-critical business information
Solution Approach 1:
The patent applies Unicode as a universal character set that can represent characters from multiple languages and alphabets simultaneously. This allows the system to maintain compatibility with regional character sets while enabling data integration across all systems, making the system multi-functional in handling diverse linguistic data
Solution Approach 2:
The patent introduces Unicode as an intermediary character set that mediates between incompatible regional character sets. Data from different regional systems is converted to Unicode for integration and processing, then converted back to the appropriate regional character set when needed, enabling seamless information sharing without loss
2Loss of information
If conversion between character sets is implemented, then data integration between systems is enabled, but conversion complexity and processing overhead increase
Solution Approach 1:
The patent implements preliminary conversion of regional character sets to Unicode at the point of data entry or storage. By performing the conversion early in the data flow, the system avoids repeated conversions during subsequent processing operations, reducing overall conversion complexity and processing overhead
Solution Approach 2:
The patent changes the character set parameter from multiple incompatible regional encodings to a single unified Unicode encoding. This parameter change simplifies the conversion process by establishing Unicode as the common denominator, reducing the complexity of data integration operations
3Ease of manufacture
If regional systems maintain their own character sets, then each system operates independently with simple data storage, but business processes requiring unified data views cannot be accomplished
Solution Approach 1:
The patent introduces Unicode as an intermediary layer that enables efficient business processes requiring unified data views. Regional systems continue to use their native character sets for local operations, while Unicode facilitates data integration and cross-system business processes, combining simplicity with efficiency
Solution Approach 2:
The patent segments the data handling process into regional components (using local character sets) and integrated components (using Unicode). This segmentation allows regional systems to maintain their simplicity while enabling complex integrated business processes through the Unicode intermediary layer
Data Source
AI summary
An embodiment of the present invention describes a method and system for using related data from external systems employing incompatible character sets to affect a business process. For one embodiment, a first external system uses a first character set. A first data set is received from the first external system, the first data set using the first character set. A second external system uses a second character set. A second data set is received from the second external system, the second data set using the second character set. The first data set and the second data set are converted to use a third character set, the third character set a superset of the first character set and the second character set. The first data set and the second data set, as converted and integrated, are then used to effect one or more business processes. For one embodiment, the third character set is Unicode.


