First Node Data Distribution Detection for ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for training Machine Learning (ML) models in 3GPP networks often rely on data quality checks performed after the models have been trained, leading to potential performance degradation due to low-quality data. Additionally, these checks are typically implemented outside the network and only after data has been transferred, causing communication overhead and strain on the network.
Innovation Solution
A computer-implemented method where a first node in a communications system obtains data sets from a second node, annotates them with an indication of their distribution representation, and determines if there has been a change in the data distribution before using it to train a predictive ML model. This node then decides whether to send the data to a third node, ensuring only high-quality data is used for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data quality checks are performed after ML model training, then the training process can proceed without interruption, but the model accuracy and reliability deteriorate due to low-quality data
Solution Approach 1:
The patent implements data quality assessment and distribution change detection before the ML model training process begins. The first node evaluates data quality metrics and determines whether data distribution has changed relative to the training phase, preventing low-quality data from being used in training, thereby maintaining both training continuity and model accuracy
2Device complexity
If data quality checks are implemented outside the network after data transfer, then the network infrastructure remains simple, but communication overhead increases and network resources are strained
Solution Approach 1:
The patent performs data quality assessment and distribution change detection within the network at the first node before data is used for training. This preliminary check prevents unnecessary transmission of low-quality data, reducing communication overhead and network resource consumption while maintaining a relatively simple network infrastructure
Solution Approach 2:
The patent introduces a first node as an intermediary between the data source (second node) and the ML model training process (third node). This intermediary performs data quality assessment and distribution change detection, filtering data before it enters the training pipeline, thereby reducing the burden on network resources and external checking mechanisms
Data Source
AI summary
A method by a first node (111) for handling data. The first node (111) obtains (204) one or more first sets of data corresponding to one or more first features used in a first predictive machine learning, ML, model. The data is annotated with an indication. The indication indicates a respective representation of a distribution of the data. The obtaining (204) is performed before the data are used to train the first ML model. The first node (111) determines (205) whether there has been a change in a respective representation of the distribution of the data, before the data are used to train the first ML model. The first node (111) determines (206) whether to send the data and sends (207) a second indication of the data to a third node (113) in response to the determining (206).


