Neural Network Data Standardization Module
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data standardization techniques are inadequate for large quantities of data attributes from diverse sources, requiring manual efforts that are resource-intensive and ineffective for existing data collections with different standardization schemas.
Innovation Solution
A data standardization module that uses neural networks for unsupervised and supervised learning to transform data elements into vectors, determining distances to classify and standardize data types, enabling automatic classification and modification of input data elements based on semantic similarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual classification methods are used to standardize data attributes, then data standardization can be achieved, but the resource consumption and time required become prohibitive for large datasets
Solution Approach 1:
The patent replaces manual mechanical classification processes with an automated neural network system. The neural network learns from training data and automatically classifies new data attributes into standardized types, eliminating the need for human annotators to manually review and categorize each attribute individually.
Solution Approach 2:
The system performs self-service by automatically classifying data attributes without requiring continuous human intervention. The neural network model, once trained on a representative sample, autonomously processes large quantities of data attributes, assigning standardized types based on learned patterns and semantic relationships.
2Adaptability or versatility
If existing data collections with different standardization schemas are processed, then data from diverse sources can be standardized, but the complexity of handling multiple schemas increases
Solution Approach 1:
The neural network classifier is designed with universal applicability to handle multiple data schemas and domains simultaneously. It learns robust feature representations that enable it to classify attributes from different sources (e.g., customer data, product data, transaction data) under various schemas into a unified standardized taxonomy without requiring schema-specific processing logic.
Solution Approach 2:
The system adapts to different data schemas by dynamically adjusting its internal representation parameters. The neural network modifies its embedding space and classification thresholds based on the input data distribution, allowing it to map diverse attribute naming conventions and structures to consistent standardized types while maintaining operational simplicity.
3Productivity
If large quantities of data attributes are classified automatically, then productivity increases, but the accuracy and precision of classification may decrease without manual verification
Solution Approach 1:
The system incorporates feedback mechanisms where the accuracy of automated classifications can be verified and corrected. Users can review and correct misclassifications, and these corrections feed back into the training process to improve the neural network's performance. This feedback loop enables the system to maintain high accuracy while processing large volumes of data.
Solution Approach 2:
The system applies partial automation by first attempting automatic classification and then providing selective manual review for uncertain or high-value cases. This hybrid approach processes the majority of data attributes automatically for high productivity while reserving manual verification for edge cases where precision is critical, achieving both speed and accuracy.
Data Source
AI summary
A standardized data model (“SDM”) includes standardized data types that indicate classifications of data elements. In a data service platform, such as a marketing data platform, a data standardization module classifies received data elements. One or more components included in the data standardization module are trained using supervised or unsupervised learning techniques to classify received data elements into a standardized data type included in the SDM. In some cases, an output of an unsupervised learning phase is provided as an input to a supervised learning phase. In some cases, a classified data element is modified by the data standardization module to indicate the standardized data type into which the data element is classified.


