AI Learning Data Classification for Performance Impact
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack the ability to distinguish and manage learning data's influence on artificial intelligence performance, leading to potential misuse or inefficiency in data learning processes.
Innovation Solution
An information processing apparatus that records learning data in a manner distinguishing between influencing and non-influencing data, using attribute information to control AI learning and classify data based on performance impact, allowing for selective data usage and management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If learning data is recorded without distinction, then storage capacity is maximized, but the ability to identify influential data is lost
Solution Approach 1:
The patent segments learning data into two distinct categories: influencing learning data and non-influencing learning data. This segmentation is achieved by recording attribute information that identifies whether each data item has influenced AI model performance. The segmentation allows the system to treat different types of data differently, enabling precise identification of influential data without overwhelming complexity in data management.
2Productivity
If all learning data is used for training, then data utilization is maximized, but performance degradation may occur from non-influential data
Solution Approach 1:
The patent applies preliminary action by evaluating and categorizing learning data before it is used for training the AI model. Attribute information is recorded in advance to identify which data items have influenced model performance. This preliminary classification enables the system to selectively use only influencing learning data for subsequent training, preventing performance degradation from non-influential data while maximizing the utility of valuable training data.
3Device complexity
If learning data is not classified, then data storage is simplified, but data security and misuse prevention are compromised
Solution Approach 1:
The patent applies local quality by assigning different attributes and management rules to different portions of the learning data based on their influence on AI performance. Influencing learning data is marked with specific attribute information that distinguishes it from non-influencing data. This localized differentiation enables targeted security measures and controlled access policies to be applied specifically to influential data, preventing misuse while maintaining simple storage structures for the overall dataset.
Data Source
AI summary
An information processing apparatus includes a controller that causes learning data learned by an artificial intelligence to be recorded in a recording unit in such manner that influencing learning data that has influenced performance of an artificial intelligence and non-influencing learning data that has not influenced performance of an artificial intelligence are distinguishable.


