First Node Data Distribution Detection for ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for training Machine Learning (ML) models in 3GPP networks often rely on data quality checks performed after the models have been trained, leading to potential performance degradation due to low-quality data. Additionally, these checks are typically implemented outside the network and only after data has been transferred, causing communication overhead and strain on the network.

Innovation Solution

A computer-implemented method where a first node in a communications system obtains data sets from a second node, annotates them with an indication of their distribution representation, and determines if there has been a change in the data distribution before using it to train a predictive ML model. This node then decides whether to send the data to a third node, ensuring only high-quality data is used for training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data quality checks are performed after ML model training, then the training process can proceed without interruption, but the model accuracy and reliability deteriorate due to low-quality data

Engineering Contradiction:
Improvetraining process continuityVSAvoidmodel accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements data quality assessment and distribution change detection before the ML model training process begins. The first node evaluates data quality metrics and determines whether data distribution has changed relative to the training phase, preventing low-quality data from being used in training, thereby maintaining both training continuity and model accuracy

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If data quality checks are implemented outside the network after data transfer, then the network infrastructure remains simple, but communication overhead increases and network resources are strained

Engineering Contradiction:
Improvenetwork infrastructure complexityVSAvoidcommunication overhead
Core Design Contradiction:
Device complexityVSLoss of energy

Solution Approach 1:

The patent performs data quality assessment and distribution change detection within the network at the first node before data is used for training. This preliminary check prevents unnecessary transmission of low-quality data, reducing communication overhead and network resource consumption while maintaining a relatively simple network infrastructure

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a first node as an intermediary between the data source (second node) and the ML model training process (third node). This intermediary performs data quality assessment and distribution change detection, filtering data before it enters the training pipeline, thereby reducing the burden on network resources and external checking mechanisms

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250047570A1First node, second node, third node, fourth node and methods performed thereby for handling data
Publication Date: 2025.02.06 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20250047570A1 patent drawing
  • US20250047570A1 patent drawing
  • US20250047570A1 patent drawing

AI summary

A method by a first node (111) for handling data. The first node (111) obtains (204) one or more first sets of data corresponding to one or more first features used in a first predictive machine learning, ML, model. The data is annotated with an indication. The indication indicates a respective representation of a distribution of the data. The obtaining (204) is performed before the data are used to train the first ML model. The first node (111) determines (205) whether there has been a change in a respective representation of the distribution of the data, before the data are used to train the first ML model. The first node (111) determines (206) whether to send the data and sends (207) a second indication of the data to a third node (113) in response to the determining (206).