Neural Network Intermediate Result Extraction for Training Data Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning systems, such as neural networks, face challenges in efficiently collecting and processing large amounts of data for training and evaluation, leading to high transfer and storage costs, as well as data protection concerns, due to the need for manual or machine-based selection of relevant data features and contents.
Innovation Solution
A small neural network detector is used to recognize and automatically capture relevant data during operation, reducing memory and transfer resources by identifying suitable data points within the distribution of training data, allowing for more reliable data acquisition without requiring additional hardware or intervals between data points.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If permanent data are collected and transferred to a server for manual or machine selection, then relevant training data can be identified, but enormous data transfer and storage costs are incurred and data protection is compromised
Solution Approach 1:
The patent extracts only the essential intermediate results from the machine learning system's processing pipeline, rather than transferring all raw data. This selective extraction of critical processing stages allows for effective data selection while dramatically reducing the volume of data that needs to be transferred and stored.
Solution Approach 2:
The patent introduces an intermediary evaluation mechanism that assesses the quality of training data based on intermediate processing results. This intermediary layer enables accurate data selection without requiring the transfer of complete datasets, as the intermediary metrics provide sufficient information for evaluation.
2Measurement precision
If manual or machine selection of relevant data features is performed, then training data quality improves, but the process becomes complex and time-consuming
Solution Approach 1:
The patent applies partial action by evaluating only specific intermediate results from the machine learning processing pipeline rather than analyzing all possible data features. This selective approach maintains training data quality while significantly reducing the complexity of the selection process.
Solution Approach 2:
The patent changes the parameters being evaluated from raw data features to intermediate processing metrics. This parameter transformation simplifies the selection process by focusing on meaningful processing stages that directly indicate data quality, rather than dealing with the full complexity of raw data features.
3Reliability
If all collected data are transferred and stored on a server, then complete data availability is achieved, but transfer costs and memory costs increase significantly
Solution Approach 1:
The patent extracts only the necessary intermediate results that are sufficient for data quality evaluation, rather than transferring and storing all collected data. This extraction approach maintains the reliability needed for effective training data selection while dramatically reducing transfer and storage costs.
4Productivity
If data are collected over intervals between distributions, then data processing load is reduced, but relevant data points within distributions may be missed
Solution Approach 1:
The patent performs preliminary evaluation of data quality using intermediate results as data points are generated during processing. This preliminary action allows for immediate identification of relevant data points within distributions without requiring interval-based sampling, thereby maintaining both processing efficiency and detection accuracy.
Data Source
AI summary
A method for detecting whether an input variable for a machine learning system is suitable as an additional training datum or test datum for the machine learning system for retraining and testing. The method includes: processing a detected input variable by way of the machine learning system, intermediate results which are ascertained during the processing of the input variable by the machine learning system being stored; processing the stored intermediate results by way of an anomaly detector, the anomaly detector outputting an output variable which characterizes whether the detected input variable associated with the intermediate results yields an anomalous behavior of the machine learning system; based on the output variable of the network, the input variable of the network and the additional input variables defined as relevant are stored/selected. A computer system, computer program, and a machine-readable memory element on which the computer program is stored are also described.


