KPI Anonymization for Wireless ML Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wireless communication networks face challenges in sharing data for training machine learning models due to reluctance to expose network performance, fault management responses, and network configurations, leading to siloed usable data and difficulty in training effective machine learning systems.
Innovation Solution
A method is introduced to anonymize network Key Performance Indicators (KPIs) by retrieving, filtering, identifying correlations, grouping, sorting, and tokenizing KPIs to create anonymized data sets accessible for training machine learning models, ensuring network data security while enabling effective model training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If network data is shared for machine learning training, then model training effectiveness is improved, but network security and data privacy are compromised
Solution Approach 1:
The patent creates synthetic KPI data that copies the statistical properties and patterns of real network data without containing actual network information. The synthetic data is generated by transforming real KPIs through randomization and aggregation processes, resulting in data that maintains the necessary training value while eliminating sensitive network details, thus resolving the contradiction between sharing data for training and protecting network security
Solution Approach 2:
The patent introduces an intermediary processing layer between real network data and machine learning models. This intermediary system includes KPI aggregation functions, randomization mechanisms, and synthetic data generation components that mediate the transformation from sensitive real data to safe synthetic data, enabling model training without direct exposure to network-sensitive information
2Quantity of substance
If raw network KPIs are provided for training, then data availability is improved, but data sensitivity and exposure of network performance are increased
Solution Approach 1:
The patent segments the data processing into distinct stages: retrieving real KPIs, aggregating them into higher-level metrics, randomizing the aggregated values, and generating synthetic KPIs. Each segment transforms the data at a different level of abstraction, progressively removing sensitive information while preserving statistical properties needed for training, thus resolving the contradiction between data availability and information protection
Solution Approach 2:
The patent changes the parameters of the KPI data through systematic transformation. Real KPI values are aggregated, randomized, and rescaled to create synthetic KPIs that have different numerical properties but maintain the underlying statistical distributions and patterns. This parameter transformation enables data availability for training while preventing exposure of actual network performance values
3Object-affected harmful factors
If network data is anonymized through tokenization, then data security is improved, but data complexity and processing requirements are increased
Solution Approach 1:
The patent performs preliminary actions by pre-defining KPI aggregation rules, randomization strategies, and synthetic data generation parameters before the actual anonymization process. These pre-configured transformation rules simplify the complex task of anonymization by providing clear guidance on how to process different types of KPIs, thus reducing the operational complexity despite maintaining high security standards
Data Source
AI summary
Various embodiments comprise a wireless communication network configured to anonymize network Key Performance Indicators (KPIs) to train a machine learning model. In some examples, the wireless communication network comprises a network analytics system and a KPI store. The network analytics system retrieves the network KPIs generated by network KPI sources. The analytics system filters the network KPIs based on an intended function of the machine learning model. The analytics system identifies correlated ones of the filtered KPIs and groups the correlated ones of the filtered KPIs into KPI groups based on a network condition. For each KPI group, the analytics system sorts the filtered KPIs into KPI ranges. For each KPI range, the analytics system tokenizes the KPI range by converting the filtered KPIs that compose the KPI range into strings. The KPI store stores the tokenized KPIs in a KPI database accessible by the machine learning model for training.


