Data Aggregation System Using Nomenclature Translation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital networking and communication systems face challenges in aggregating and comparing data from different sources due to inconsistent labeling structures, leading to inaccurate comparisons and conclusions.
Innovation Solution
A system that uses a comparison module to identify undefined data labels and translates them into a common nomenclature using a synonyms database, allowing for accurate aggregation and comparison of data sets with different labeling structures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data from different sources is aggregated using different labeling structures, then the quantity and quality of available data increases, but the accuracy of data comparison and aggregation deteriorates
Solution Approach 1:
The patent introduces a nomenclature database as an intermediary layer between diverse data sources and the aggregation system. This mediator contains standardized labels that map to multiple possible source labels, enabling accurate comparison while accepting data in various formats. The translation module uses this intermediary to convert undefined labels from different data sources into standardized nomenclature labels, resolving the contradiction between accepting diverse data and maintaining comparison accuracy.
Solution Approach 2:
The system changes the parameter of data labels by translating them from their original form into a standardized nomenclature. The translation module modifies the label parameter by looking up undefined labels in the nomenclature database and replacing them with standardized equivalents, thereby maintaining data accuracy across different sources while preserving the ability to accept diverse input formats.
2Measurement precision
If a common labelling scheme is used for data aggregation, then data comparison accuracy is improved, but the ability to process unfamiliar data sets deteriorates
Solution Approach 1:
The system performs preliminary action by pre-populating the nomenclature database with standardized labels and their mappings to various source labels before data aggregation occurs. This advance preparation enables the system to quickly translate and adapt unfamiliar data sets without requiring real-time complex processing, thus maintaining both accuracy and adaptability.
Solution Approach 2:
The nomenclature database serves as an intermediary that bridges common labelling schemes and unfamiliar data sets. It contains pre-defined mappings that allow the system to translate novel or unfamiliar labels into standardized forms, maintaining the ability to process diverse data sources while ensuring comparison accuracy through the standardized nomenclature layer.
3Reliability
If data labels are strictly validated against a fixed nomenclature, then data quality is improved, but the flexibility to handle varied data sources deteriorates
Solution Approach 1:
The translation module acts as an intermediary that handles the conflict between strict validation and flexibility. It first attempts to translate unfamiliar labels using the nomenclature database, and only when translation fails does it prompt for user input. This layered approach maintains data quality through validation while preserving flexibility to handle varied data sources.
Solution Approach 2:
The system implements feedback by monitoring translation success and prompting users when translation fails. This feedback loop allows the system to maintain strict validation standards while adapting to new data sources through user-provided label mappings, which can then be added to the nomenclature database for future automatic translation.
Data Source
AI summary
A system and method for outputting modified input data for storage comprises a communication interface, a comparison module, a translation module and an output module. The communication interface is arranged to receive an input data set comprising a plurality of data labels. The comparison module is arranged to compare the data labels to a plurality of nomenclature-labels in a nomenclature database and identify an undefined data label by determining that at least one of the data labels is not present in the nomenclature database, based on the comparison. The translation module is arranged to translate the undefined data label into a nomenclature-label using a synonyms database. The output module is arranged to output a modified data set based on the input data set and the translated undefined label for storage.


