Ideogrammatic Data Matching with Granular Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing search and match systems for databases face challenges in efficiently differentiating match quality, especially in non-phonetic and ideogrammatic writing systems, leading to costly manual intervention and limited granularity in feedback, which is inadequate for languages like Japanese and Chinese.
Innovation Solution
A system and method that enhance data matching by using polylogogrammatic semantic disambiguation, hanzee acronym expansion, and kanji acronym expansion to convert input data into optimized keys, generating matchgrade patterns and confidence codes for granular feedback, and incorporating business rules for automated decisioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional search and match systems are used for ideogrammatic data, then basic retrieval functionality is provided, but match quality differentiation is insufficient leading to costly manual intervention
Solution Approach 1:
The patent segments the match feedback into multiple granular levels (A, B, C matches with further subdivisions) rather than providing a single binary match result. This segmentation enables more precise differentiation of match quality, allowing the system to identify high-confidence matches that can be automatically processed versus those requiring manual review, thereby reducing manual intervention costs while improving measurement precision.
Solution Approach 2:
The patent implements a comprehensive feedback mechanism that provides detailed match quality information including match grades, confidence scores, and specific attribute comparisons. This feedback enables automated decisioning by clearly indicating which matches are sufficient for automatic processing and which require manual review, resolving the contradiction between precise match differentiation and reducing manual intervention.
2Ease of operation
If coarse-grained match feedback (A, B, C categories) is provided, then automated processing is simplified, but granularity is insufficient for differentiating among multiple matches in the same category
Solution Approach 1:
The patent further segments each match category (A, B, C) into subcategories with distinct match grades and confidence scores. For example, A matches are divided into A1, A2, A3 with varying levels of confidence. This multi-level segmentation preserves detailed match quality information while maintaining automated decisioning capability by providing clear hierarchical structure for algorithmic processing.
Solution Approach 2:
The patent adds additional dimensions to the feedback structure by introducing multiple attributes beyond simple category classification, including match grades, confidence scores, attribute-level comparison results, and specificity indicators. This multi-dimensional feedback provides comprehensive match quality differentiation while remaining suitable for automated processing through structured data formats.
3Reliability
If manual review is performed for all B category matches, then match accuracy is improved, but processing time and cost increase significantly
Solution Approach 1:
The patent applies partial automated review by selectively processing only those B category matches that meet specific confidence thresholds or exhibit certain characteristics, rather than manually reviewing all B matches. This partial automation maintains acceptable match accuracy for high-confidence cases while reducing processing time and costs for lower-confidence matches that can be handled with automated methods or lower priority review.
Solution Approach 2:
The patent changes the parameters used to determine review requirements by introducing confidence scores and match grades as decision criteria. Instead of treating all B matches uniformly, the system uses these parameter changes to automatically identify which B matches require manual review versus those that can be processed automatically, thereby improving efficiency while maintaining reliability.
4Measurement precision
If comprehensive match analysis is performed on all data elements, then matching precision is improved, but computational complexity and processing time increase
Solution Approach 1:
The patent segments the data elements into different importance categories and processes them with varying levels of analysis depth. Critical elements such as proper nouns and key identifiers receive comprehensive analysis, while less critical elements receive simplified processing. This segmentation maintains high matching precision for important fields while reducing overall computational complexity through selective analysis.
Solution Approach 2:
The patent applies different quality levels of analysis to different data elements based on their importance and characteristics. High-importance fields like names and identifiers receive detailed, comprehensive analysis with multiple comparison methods, while lower-importance fields receive more efficient, simplified processing. This local quality approach optimizes the balance between matching precision and computational complexity by allocating resources according to need.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method of searching and matching non-phonetic or ideogrammatic input data to stored data, including the steps of receiving input data comprising a search string having a plurality of elements, converting a subset of the elements into a set of terms, generating an optimized plurality of keys from the set of terms, retrieving stored data based on the optimized keys corresponding to most likely candidates for match, and selecting a best match from the plurality of candidates. At least some of the ideogrammatic elements form part of an ideogrammatic writing system. The method may also include dividing the search string into a plurality of overlapping sub-segments and identifying sub-segments having inferred semantic meaning as well as sub-segments having no semantic meaning in the ideogrammatic writing system, and using the various sub- segments to generate the optimized keys.