Parallelizable Distributed Data Preservation for ML Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current platforms face challenges in processing 'Big Data' due to their inability to leverage distributed processing, leading to inefficiencies in machine learning algorithms that require multiple iterations to converge, and existing methods for generating pseudo-random data lack the ability to parallelize and distribute processing effectively, limiting the flexibility of analytical algorithms.
Innovation Solution
The Parallelizable Distributed Data Preservation (PDDP) system transforms original data into smaller datasets with similar statistical characteristics, allowing for parallel processing and generating new datasets that can be used for machine learning, while also enabling incremental learning and real-time data mining through Data Learning Analytics (DLA) and Real-Time Bidding Data Monitoring (RTBD) technologies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If distributed processing is not leveraged, then data processing can be performed on current platforms, but processing efficiency and scalability are limited
Solution Approach 1:
The patent segments the original large dataset into multiple smaller partitions that can be processed independently and in parallel across distributed computing nodes. This segmentation enables the system to leverage distributed processing while maintaining manageable complexity at each node level.
Solution Approach 2:
The patent introduces a new dimension of parallelism by transforming sequential data processing into multi-dimensional parallel processing across multiple nodes and iterations. This allows the system to scale processing capacity by adding more computational nodes rather than increasing the complexity of single-node processing.
2Measurement precision
If multiple iterations are used for machine learning convergence, then model accuracy improves, but processing time and resource consumption increase
Solution Approach 1:
The patent performs preliminary actions by pre-processing and partitioning the dataset before the main machine learning iterations begin. This preparation work enables faster convergence during the actual training iterations by ensuring data is already in the optimal format and distributed across nodes, reducing the time penalty of multiple iterations.
Solution Approach 2:
The patent maintains continuity of useful action by implementing parallel data processing and model training across multiple nodes simultaneously. Instead of sequential processing where each iteration waits for the previous one to complete, the system continuously performs useful computational work across the distributed network throughout the iterative process.
3Quantity of substance
If pseudo-random data generation is used, then data size is reduced, but the ability to parallelize and distribute processing is lost
Solution Approach 1:
The patent segments the data preservation process into distinct parallelizable stages: original data partitioning, pseudo-random generation per partition, and result aggregation. This segmentation allows each stage to be independently optimized and distributed across computing nodes, maintaining processing flexibility while achieving data size reduction.
Solution Approach 2:
The patent creates multiple copies of processed data partitions across distributed nodes, where each node generates pseudo-random data for its local partition. This copying approach enables parallel execution of the data transformation process while maintaining the ability to distribute work across the network, preserving adaptability.
Data Source
AI summary
The Parallelizable Distributed Data Preservation Apparatuses, Methods and Systems (“PDDP”) transforms an ad impression event, a bidding invite, original data set, original data distribution estimation, symetry ML BET table inputs via PDDP components into real-time mobile bid, mobile ad placement, pseudo random datastet, build classifier structure, build regression structure outputs. In one example embodiment, the PDDP includes an apparatus. The PDDP's apparatus' instructions include obtaining original data set and determine appropriate symmetry ML basic element table, generating original data distribution estimation structure and generate new dataset random generation structure, generating new random dataset transformation structure and transforming original data with the symmetry ML basic element table into pseudo random dataset. The PDDP also provides pseudo random dataset to machine learning component and to generate build classifier and build regression structures from the machine learning component.


