Machine Learning Network Training with Human Guesstimates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning algorithms face limitations in accuracy and utility when training sets are weak, sparse, or of poor quality, leading to poor performance in predicting outcomes, especially when the relationship between the training data and the situation to be predicted is not well-represented.
Innovation Solution
A system and method that utilize human votes or guesstimates to initialize and adjust the strength of nodes and connections in a machine learning network, allowing it to learn from synthetic data sets created from user opinions, thereby improving training efficiency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional supervised learning is used to train machine learning networks, then the network can learn from data and make predictions, but the training process requires hundreds or thousands of iterations and may fail when training sets are weak, sparse, or of poor quality
Solution Approach 1:
The patent applies preliminary action by using human experts to pre-label a subset of training data before the machine learning network begins its training process. This pre-labeling of approximately 10% of the data provides initial guidance that significantly reduces the number of training iterations needed, cutting training time from hundreds or thousands of cycles to a fraction of that time while maintaining high prediction accuracy.
Solution Approach 2:
The patent introduces human experts as an intermediary between the raw data and the machine learning network. These experts provide annotated labels for a portion of the training data, serving as a mediator that bridges the gap between unprocessed data and the network's learning requirements, thereby enabling more efficient training with fewer iterations.
2Adaptability or versatility
If machine learning networks are trained with limited or poor quality training data, then the network can still operate, but prediction accuracy deteriorates and performance approaches random behavior
Solution Approach 1:
The patent applies preliminary action by having human experts pre-label approximately 10% of the training data before the machine learning network begins processing. This pre-labeling provides crucial initial guidance that enables the network to achieve high prediction accuracy even when the overall training set is limited or of poor quality, preventing the performance degradation that would otherwise occur.
Solution Approach 2:
The patent introduces human experts as an intermediary that bridges the gap between limited/poor quality data and the machine learning network. These experts provide annotated labels that serve as a mediator, enabling the network to extract meaningful patterns from otherwise insufficient or noisy data, thereby maintaining high prediction accuracy despite data limitations.
3Ease of operation
If explicit rules are used to describe data relationships, then the system can make predictions based on stated rules, but the rules may fail to recognize subtle relationships and cannot handle complex non-linear problems
Solution Approach 1:
The patent replaces the mechanical system of explicit human-generated rules with a machine learning network that automatically learns patterns from data. Instead of relying on pre-stated rules that may miss subtle relationships, the network processes data through multiple layers, automatically discovering complex non-linear relationships and patterns that would be difficult or impossible to encode manually.
Solution Approach 2:
The patent applies self-service by enabling the machine learning network to automatically learn and adapt to data patterns without requiring explicit programming of rules. The network self-adjusts its internal parameters during training, automatically discovering optimal patterns and relationships in the data, thereby eliminating the limitations of manual rule creation while maintaining ease of operation.
Data Source
AI summary
Described is a system and method for training a machine learning network. The method comprises initializing at least one of nodes in a machine learning network and connections between the nodes to a predetermined strength value, wherein the nodes represent factors determining an output of the network, providing a first set of questions to a plurality of users, the first set of questions relating to at least one of the factors, receiving at least one of choices and guesstimates from the users in response to the first set of questions and adjusting the predetermined strength value as a function of the choices/guesstimates. The real and simulated examples presented demonstrate that synthetic training sets derived from expert or non-expert human guesstimates can replace or augment training data sets comprised of actual training exemplars that are too limited in size, scope, or quality to otherwise generate accurate predictions.


