ANN Data Normalization and Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data cleansing techniques for Artificial Neural Networks (ANN) are inadequate for stochastic data generated by ANNs, as they rely on deterministic data validation methods that are time-consuming and prone to errors, and fail to effectively normalize and validate data from diverse sources.
Innovation Solution
A system and method that normalizes and validates data from different formats by mapping sources to a common format, performs format and restriction validation, and restarts parsing operations until the required format is achieved, using a deterministic data model like OneTest Data to train the ANN model and generate stochastic data, with a data validation component for validating the stochastic data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If deterministic data validation methods are used for ANN stochastic data, then data validation can be performed, but the process is time-consuming and prone to errors
Solution Approach 1:
The patent changes the validation approach from deterministic fixed rules to probabilistic parameter ranges. Instead of requiring exact matches, the system validates data by checking if values fall within acceptable probability distributions and ranges, enabling efficient validation of stochastic ANN data while maintaining reliability
Solution Approach 2:
The patent replaces the mechanical deterministic validation system with a statistical probabilistic validation system. This substitution allows the system to handle the randomness inherent in ANN-generated data by using probability distributions and statistical thresholds rather than rigid deterministic rules
2Reliability
If manual human validation of data is performed, then data quality can be ensured, but the process is time-consuming and prone to errors
Solution Approach 1:
The patent implements self-service validation where the system automatically validates its own generated data using probabilistic models. The validation mechanism serves itself by using the training data distributions to automatically assess data quality, eliminating the need for manual human intervention and significantly increasing throughput while maintaining quality
3Stability of the object's composition
If data from different formats is normalized to common format, then data consistency is improved, but the normalization process becomes complex
Solution Approach 1:
The patent creates a universal normalization framework that handles multiple data formats through a single probabilistic validation layer. Instead of requiring separate normalization rules for each format, the system uses probability distributions that can accommodate various formats, simplifying the overall process while maintaining consistency
Data Source
AI summary
Disclosed is a method and system for modifying data cleansing techniques for training and validating an Artificial Neural Network (ANN) model. The method comprises normalizing and validating data of different formats, obtained from different sources. The ANN model is trained using the normalized and validated data. Alternatively, the ANN model could be trained using data of a common format obtained from a deterministic data model. The trained ANN model is used to generate ANN stochastic data. Data validation component from the deterministic data model is reused for the normalizing and the validating of the data, for validating the ANN stochastic data.


