Relevance Metric Data Sample Selection for ML Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for data sample processing in communication networks face challenges in efficiently reducing signaling and data processing overhead, particularly in wireless communication networks, which are crucial for machine learning applications.
Innovation Solution
A computer-implemented method that involves obtaining relevance metrics and criteria in a data processing network node to identify and signal relevant data samples to a data collecting network node, thereby reducing the need for unnecessary data transmission and processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all data samples are transmitted from data collecting network node to data processing network node, then the predictive performance of machine learning models is improved, but the signaling overhead and communication resources increase significantly
Solution Approach 1:
The patent extracts only the most relevant data samples from the complete dataset by computing relevance metrics that quantify the predictive value of each sample. The data processing network node identifies and requests only those samples that exceed a relevance threshold, thereby extracting the essential information needed for model training while discarding redundant data. This resolves the contradiction by maintaining predictive performance through selective extraction rather than transmitting all samples.
Solution Approach 2:
The patent applies local quality by assigning different relevance scores to different data samples based on their individual characteristics and predictive value. Instead of treating all samples uniformly, the system evaluates each sample's contribution to predictive performance and prioritizes transmission of high-value samples. This differential treatment maintains model accuracy while reducing overall communication overhead.
2Measurement precision
If all data samples are stored locally in the data processing network node, then the predictive performance is improved, but the hardware cost and energy consumption increase
Solution Approach 1:
The system extracts only the essential data samples needed for maintaining predictive performance by computing relevance metrics. Instead of storing complete datasets locally, the data processing network node stores only those samples that contribute significantly to model accuracy, thereby reducing storage requirements and associated energy consumption while maintaining model effectiveness.
Solution Approach 2:
The patent applies partial action by storing and processing only a subset of data samples rather than the complete dataset. The relevance metric computation identifies the minimum necessary sample set required to achieve acceptable predictive performance, avoiding the excessive storage and processing of redundant data, thus reducing energy consumption.
3Productivity
If data samples are transmitted wirelessly from wireless device to data processing network node, then the data collection capability is improved, but the signaling overhead becomes prohibitive
Solution Approach 1:
The patent extracts only the most informative data samples for wireless transmission by computing relevance metrics at the data collecting network node. The system identifies samples with high predictive value and requests only those for transmission to the data processing network node, thereby maintaining data collection capability while dramatically reducing the volume of wireless signaling required.
Solution Approach 2:
The relevance metric computation and sample selection are performed in advance at the data collecting network node before wireless transmission. This preliminary filtering of data samples reduces the amount of information that needs to be transmitted over the wireless channel, thereby reducing signaling overhead while preserving the essential data collection capability.
4Measurement precision
If the number of data samples is increased, then the machine learning model performance is improved, but the memory footprint and processing overhead increase
Solution Approach 1:
The patent extracts the most valuable data samples by computing relevance metrics that quantify each sample's contribution to model performance. Instead of storing large numbers of redundant samples, the system identifies and retains only those samples with high predictive value, thereby achieving good model performance with a compact dataset that requires less memory storage and processing resources.
Data Source
AI summary
A computer-implemented method for acquiring new data samples and for maintaining a set of data samples in a database, wherein the set of data samples are configured to form input to a function associated with a predictive performance, the method including obtaining at least one relevance metric (M), where the relevance metric is indicative of an increase in the predictive performance of the function when using a data sample as input together with the set of data samples compared to when not using the data sample, obtaining a relevance criterion (C), where the relevance criterion identifies relevant data samples in a set of data samples based on the at least one relevance metric, signaling the at least one relevance metric (M) and the relevance criterion (C) to a data collecting network node.


