Parallelizable Distributed Data Preservation for ML Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current platforms face challenges in processing 'Big Data' due to their inability to leverage distributed processing, leading to inefficiencies in machine learning algorithms that require multiple iterations to converge, and existing methods for generating pseudo-random data lack the ability to parallelize and distribute processing effectively, limiting the flexibility of analytical algorithms.

Innovation Solution

The Parallelizable Distributed Data Preservation (PDDP) system transforms original data into smaller datasets with similar statistical characteristics, allowing for parallel processing and generating new datasets that can be used for machine learning, while also enabling incremental learning and real-time data mining through Data Learning Analytics (DLA) and Real-Time Bidding Data Monitoring (RTBD) technologies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If distributed processing is not leveraged, then data processing can be performed on current platforms, but processing efficiency and scalability are limited

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidprocessing architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the original large dataset into multiple smaller partitions that can be processed independently and in parallel across distributed computing nodes. This segmentation enables the system to leverage distributed processing while maintaining manageable complexity at each node level.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of parallelism by transforming sequential data processing into multi-dimensional parallel processing across multiple nodes and iterations. This allows the system to scale processing capacity by adding more computational nodes rather than increasing the complexity of single-node processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If multiple iterations are used for machine learning convergence, then model accuracy improves, but processing time and resource consumption increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and partitioning the dataset before the main machine learning iterations begin. This preparation work enables faster convergence during the actual training iterations by ensuring data is already in the optimal format and distributed across nodes, reducing the time penalty of multiple iterations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent maintains continuity of useful action by implementing parallel data processing and model training across multiple nodes simultaneously. Instead of sequential processing where each iteration waits for the previous one to complete, the system continuously performs useful computational work across the distributed network throughout the iterative process.

Inventive Principle:
Principle #20Continuity of useful action

3Quantity of substance

If pseudo-random data generation is used, then data size is reduced, but the ability to parallelize and distribute processing is lost

Engineering Contradiction:
Improvedata sizeVSAvoidprocessing flexibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments the data preservation process into distinct parallelizable stages: original data partitioning, pseudo-random generation per partition, and result aggregation. This segmentation allows each stage to be independently optimized and distributed across computing nodes, maintaining processing flexibility while achieving data size reduction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates multiple copies of processed data partitions across distributed nodes, where each node generates pseudo-random data for its local partition. This copying approach enables parallel execution of the data transformation process while maintaining the ability to distribute work across the network, preserving adaptability.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12141662B1Parallelizable distributed data preservation apparatuses, methods and systems
Publication Date: 2024.11.12 CADENT LLC
  • US12141662B1 patent drawing
  • US12141662B1 patent drawing
  • US12141662B1 patent drawing

AI summary

The Parallelizable Distributed Data Preservation Apparatuses, Methods and Systems (“PDDP”) transforms an ad impression event, a bidding invite, original data set, original data distribution estimation, symetry ML BET table inputs via PDDP components into real-time mobile bid, mobile ad placement, pseudo random datastet, build classifier structure, build regression structure outputs. In one example embodiment, the PDDP includes an apparatus. The PDDP's apparatus' instructions include obtaining original data set and determine appropriate symmetry ML basic element table, generating original data distribution estimation structure and generate new dataset random generation structure, generating new random dataset transformation structure and transforming original data with the symmetry ML basic element table into pseudo random dataset. The PDDP also provides pseudo random dataset to machine learning component and to generate build classifier and build regression structures from the machine learning component.