Communication Network Data Acquisition Orchestration with Synthetic Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for acquiring communication network data for training AI/ML models face challenges such as high resource costs, bandwidth usage, privacy concerns, and regulatory restrictions, making it difficult to collect and generate data efficiently and effectively.
Innovation Solution
An orchestration node uses an orchestration ML model to determine the optimal split between collecting data from sources and generating data using a generative model, considering factors like resource availability and privacy requirements, to provide a balanced dataset for training AI/ML models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If large volumes of communication network data are collected from data sources, then the quality and performance of AI/ML models is improved, but network bandwidth consumption increases and privacy compliance becomes more difficult
Solution Approach 1:
The patent uses generative models to create synthetic copies of communication network data that replicate the statistical properties and patterns of real data without containing actual sensitive information. These synthetic data copies can be used to train AI/ML models while avoiding privacy compliance issues associated with transferring and storing real user data across borders.
Solution Approach 2:
The patent introduces an intermediary processing layer where real data is first used to train generative models, which then produce synthetic data for model training. This intermediary approach allows the system to access data characteristics without directly handling the actual data, resolving the contradiction between needing data for training and avoiding privacy violations.
2Measurement precision
If real data is collected from communication network sources, then training accuracy is improved, but resource costs and transfer requirements increase
Solution Approach 1:
Instead of transferring and storing large volumes of real communication network data across different locations and devices, the system creates synthetic copies through generative models. These synthetic data copies maintain the necessary training accuracy while dramatically reducing the resource consumption associated with data collection, transfer, and storage operations.
Solution Approach 2:
The system performs preliminary training of generative models using a subset of real data, then uses these pre-trained models to generate synthetic data for subsequent training phases. This preliminary action allows the system to achieve training accuracy without the ongoing resource costs of continuous real data collection and transfer.
3Adaptability or versatility
If data is transferred outside the country where it was collected, then access to data for training is improved, but regulatory compliance becomes more difficult
Solution Approach 1:
The patent creates synthetic data copies that replicate the statistical properties of real data without containing actual sensitive information. These synthetic copies can be freely shared and used for training AI/ML models across different countries and jurisdictions without triggering privacy restrictions or data sovereignty laws that apply to real data transfers.
Solution Approach 2:
The generative model acts as an intermediary that transforms real data into synthetic data representations. This intermediary transformation allows the system to maintain data accessibility and training capabilities while eliminating the regulatory compliance burdens associated with cross-border data transfers of actual personal or sensitive information.
4Ease of manufacture
If generative models are used to generate training data, then data acquisition cost is reduced, but computational resource requirements increase
Solution Approach 1:
The system performs preliminary training of generative models using a relatively small amount of real data, creating a foundation that can then generate large volumes of synthetic data. This preliminary action concentrates the high computational resource requirements into an initial phase, after which the synthetic data generation can proceed with lower ongoing computational costs compared to continuous real data collection and transfer.
Data Source
Figure 1
Figure 2
Figure 3a
AI summary
A method (200) is disclosed for orchestrating acquisition of a quantity of communication network data for training a target Machine Learning (ML) model for use by a communication network node. The method comprises obtaining a representation of a data acquisition state for the communication network data (210) and using an orchestration ML model to map the representation of the data acquisition state to a first amount of the communication network data to be collected from sources of the communication network data, and a remaining amount of the communication network data to be generated using a generative model (220). The method further comprises, when sufficient data has been collected (240), causing a generative model for the communication network data to be trained using the collected communication network data (250), and causing the remaining amount of the communication network data to be generated using the trained generative model (260).