Data Reduction Engine for Network ML Exploration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The machine learning development process in data center networks is hindered by the large volume and complexity of data, leading to time-consuming exploration phases and high costs, with existing methods struggling to efficiently reduce data while maintaining acceptable accuracy for machine learning models.
Innovation Solution
A data reduction engine that utilizes network-structural insights and a grammar based on network topology to perform a structured search of reduction functions, identifying a subset of transformations that reduce input data while ensuring sufficient accuracy for machine learning tasks, thereby speeding up the development process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If data volume is reduced for machine learning development, then processing time and cost are reduced, but data accuracy may deteriorate
Solution Approach 1:
The patent extracts and removes redundant or less important features from the network data while retaining the essential information needed for machine learning tasks. This is achieved through automated feature selection algorithms that identify and eliminate unnecessary data elements, thereby reducing data volume while preserving accuracy.
Solution Approach 2:
The patent applies different processing strategies to different portions of the data based on their importance and characteristics. Critical network metrics are preserved with high fidelity, while less important data is aggregated or simplified. This localized quality approach ensures that data accuracy is maintained where it matters most while reducing overall data volume.
2Productivity
If complex reduction functions are applied to reduce data volume, then data processing efficiency improves, but system complexity increases
Solution Approach 1:
The patent implements self-service through automated feature selection and data reduction algorithms that autonomously identify and apply appropriate reduction techniques without requiring manual intervention. The system automatically evaluates data characteristics, selects optimal reduction functions, and adjusts parameters based on performance feedback, thereby improving efficiency while managing complexity through automation.
Solution Approach 2:
The patent dynamically adjusts reduction parameters and thresholds based on the specific characteristics of the network data and the requirements of the machine learning task. By changing parameters such as aggregation levels, sampling rates, and feature selection criteria, the system optimizes the balance between data reduction efficiency and system complexity for different scenarios.
Data Source
AI summary
The systems and methods may use a data reduction engine to reduce a volume of input data for machine learning exploration for computer networking related problems. The systems and methods may receive input data related to a network and obtain a network topology. The systems and methods may perform a structured search of a plurality of reduction functions based on a grammar to identify a subset of reduction functions. The systems and methods may generate transformed data by applying the subset of reduction functions to the input data and may determine whether the transformed data meets or exceeds a threshold. The systems and methods may output the transformed data in response to the transformed data meeting or exceeding the threshold.


