ANN Data Normalization and Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data cleansing techniques for Artificial Neural Networks (ANN) are inadequate for stochastic data generated by ANNs, as they rely on deterministic data validation methods that are time-consuming and prone to errors, and fail to effectively normalize and validate data from diverse sources.

Innovation Solution

A system and method that normalizes and validates data from different formats by mapping sources to a common format, performs format and restriction validation, and restarts parsing operations until the required format is achieved, using a deterministic data model like OneTest Data to train the ANN model and generate stochastic data, with a data validation component for validating the stochastic data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deterministic data validation methods are used for ANN stochastic data, then data validation can be performed, but the process is time-consuming and prone to errors

Engineering Contradiction:
Improvedata validation accuracyVSAvoidvalidation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent changes the validation approach from deterministic fixed rules to probabilistic parameter ranges. Instead of requiring exact matches, the system validates data by checking if values fall within acceptable probability distributions and ranges, enabling efficient validation of stochastic ANN data while maintaining reliability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical deterministic validation system with a statistical probabilistic validation system. This substitution allows the system to handle the randomness inherent in ANN-generated data by using probability distributions and statistical thresholds rather than rigid deterministic rules

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual human validation of data is performed, then data quality can be ensured, but the process is time-consuming and prone to errors

Engineering Contradiction:
Improvedata qualityVSAvoidvalidation throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements self-service validation where the system automatically validates its own generated data using probabilistic models. The validation mechanism serves itself by using the training data distributions to automatically assess data quality, eliminating the need for manual human intervention and significantly increasing throughput while maintaining quality

Inventive Principle:
Principle #25Self-service

3Stability of the object's composition

If data from different formats is normalized to common format, then data consistency is improved, but the normalization process becomes complex

Engineering Contradiction:
Improvedata format consistencyVSAvoidnormalization process complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent creates a universal normalization framework that handles multiple data formats through a single probabilistic validation layer. Instead of requiring separate normalization rules for each format, the system uses probability distributions that can accommodate various formats, simplifying the overall process while maintaining consistency

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11847559B2Modifying data cleansing techniques for training and validating an artificial neural network model
Publication Date: 2023.12.19 HCL AMERICA INC
  • US11847559B2 patent drawing
  • US11847559B2 patent drawing
  • US11847559B2 patent drawing

AI summary

Disclosed is a method and system for modifying data cleansing techniques for training and validating an Artificial Neural Network (ANN) model. The method comprises normalizing and validating data of different formats, obtained from different sources. The ANN model is trained using the normalized and validated data. Alternatively, the ANN model could be trained using data of a common format obtained from a deterministic data model. The trained ANN model is used to generate ANN stochastic data. Data validation component from the deterministic data model is reused for the normalizing and the validating of the data, for validating the ANN stochastic data.