Synchronized Training Dataset Generation for ML Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to effectively synchronize time-varying data from multiple sources for generating a curated training dataset, leading to inefficient data correlation and increased computing load on User Equipment (UE) due to intrusive monitoring applications and lack of passive data collection mechanisms.
Innovation Solution
A method and system for generating synchronized labelled training datasets by determining timing parameters for User Equipment (UE) and network nodes, calculating a timing advance factor for synchronization, and correlating user experience data with network KPI data to create a synchronized dataset, thereby reducing computational load and maintaining user privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Difficulty of detecting and measuring
If monitoring applications are introduced in UE to collect data, then data collection capability is improved, but computing load on UE increases and user privacy is compromised
Solution Approach 1:
The patent introduces a network node as an intermediary that collects data from multiple sources including UE without requiring monitoring applications on the UE. The network node acts as a mediator that gathers user experience data and network KPI data, eliminating the need for resource-intensive local monitoring applications while preserving user privacy.
Solution Approach 2:
The system enables self-service data collection where the network infrastructure automatically collects and correlates data from multiple sources without requiring active participation or monitoring applications on the UE. The UE simply provides data through normal operations while the network node performs the collection and correlation tasks.
2Quantity of substance
If data is collected from multiple data sources with independent timing clocks, then data comprehensiveness is improved, but data correlation complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-synchronizing the timing clocks of multiple data sources before data collection begins. The network node configures and synchronizes timing parameters of UE, content servers, and other network elements in advance, so that all data sources operate on a common time reference. This eliminates the need for complex post-hoc correlation of time-stamped data from unsynchronized sources.
Solution Approach 2:
The system implements feedback mechanisms where the network node continuously monitors timing parameters and adjusts synchronization settings based on observed timing deviations. This closed-loop approach maintains synchronization accuracy across multiple data sources, simplifying the correlation process while preserving comprehensive data collection.
3Difficulty of detecting and measuring
If external probes are introduced into network infrastructure for data collection, then data measurement capability is improved, but system intrusiveness and cost increase
Solution Approach 1:
The patent makes the network node universal by enabling it to perform multiple functions: data collection, timing synchronization, data correlation, and analysis. Instead of introducing specialized external probes, the existing network nodes are enhanced to handle all these tasks, reducing system intrusiveness and cost while maintaining comprehensive measurement capability.
Data Source
AI summary
Disclosed subject matter relates to supervised machine learning including a method and system for generating synchronized labelled training dataset for building a learning model. The training data generation system determines a timing advance factor to achieve time synchronization between User Equipment (UE) and network nodes, by signalling the UE to initiate playback of the multimedia content based on the timing advance factor. The training data generation system receives network Key Performance Indicator (KPI) data from the network nodes and a user experience data from the UE, concurrently, for the streamed multimedia content, and performs timestamp based correlation to generate a synchronized labelled training dataset for building a learning model. The learning model is further deployed in external analytics system to act as a non-intrusive passive probe to predict real-time user experience, without intruding into the UE, thereby sustaining privacy of the user and eliminating additional computing load on the UE.


