Data Science Package Containers for Efficient Operation Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data science techniques require significant computing power, are inefficient, and inflexible, necessitating trained data scientists and complex processes for handling large data sets, which limits accessibility and efficiency.

Innovation Solution

A data science system that generates component descriptors to create lightweight containers for data science packages, enabling easier transfer, processing, and execution of data science operations, allowing users to create and execute customized data science operations without extensive programming knowledge.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If trained data scientists use sophisticated computing processes and frameworks to perform data science operations on large data sets, then the quality and accuracy of data science operations are improved, but the device complexity and computing resources required increase significantly

Engineering Contradiction:
Improvequality of data science operationsVSAvoidcomputing process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates template data science packages that capture and store proven data science operations, algorithms, and processing steps. These templates can be copied and reused multiple times without requiring the original complex computing processes to be recreated, thereby maintaining operation quality while reducing the complexity of each execution instance

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs data science operations in advance and packages the results into reusable templates. By pre-processing and packaging operations before they are needed, the system eliminates the need to repeatedly execute complex computing processes, reducing device complexity while preserving operational quality through the use of pre-validated templates

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If trained data scientists manually clean and program algorithms for large data sets, then the precision and reliability of data science operations are improved, but the time and labor required increase significantly

Engineering Contradiction:
Improvedata cleaning precisionVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs data cleaning, validation, and algorithm programming in advance as part of creating template data science packages. By completing these time-consuming tasks beforehand and storing them in reusable templates, the system eliminates the need for manual data preparation each time data science operations are executed, thereby maintaining precision while dramatically reducing time requirements

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates reusable templates that capture precisely cleaned data and programmed algorithms. These templates can be copied and applied to new data sets without requiring repeated manual cleaning and programming, thus preserving measurement precision while eliminating the time loss associated with repetitive manual work

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If general-purpose frameworks are used for large-scale data science computations, then the adaptability and versatility of data science operations are improved, but the ease of operation decreases due to complexity

Engineering Contradiction:
Improvedata science operation flexibilityVSAvoidframework usability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent divides complex data science operations into discrete, manageable template packages. Each template encapsulates specific algorithms, data cleaning steps, and processing logic as separate reusable units. This segmentation allows users to select and combine only the templates needed for their specific tasks, maintaining versatility while improving ease of operation by eliminating the need to navigate complex general-purpose frameworks

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces template data science packages as intermediaries between users and the underlying complex computing frameworks. Users interact with simplified templates rather than directly with complex frameworks, thereby maintaining adaptability and versatility while significantly improving ease of operation through the intermediary layer that abstracts away framework complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If dedicated storage space is provisioned for large data sets, then the reliability of data science operations is improved, but the device complexity and infrastructure requirements increase

Engineering Contradiction:
Improvedata storage reliabilityVSAvoidstorage infrastructure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines data, algorithms, and processing instructions into integrated template packages. By merging these elements into unified reusable units, the system eliminates the need for separate dedicated storage infrastructure for raw data, processed data, and algorithm definitions, thereby maintaining data storage reliability while reducing device complexity through consolidation

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10353882B2Packaging data science operations
Publication Date: 2019.07.16 ADOBE INC
  • US10353882B2 patent drawing
  • US10353882B2 patent drawing
  • US10353882B2 patent drawing

AI summary

The present disclosure relates to a data science system that packages data science operations. The data science system packages a data science operation with a component descriptor or service descriptor to allow the data science system to easily apply and execute the data science operations using data from a variety sources. As described herein, the data science system also enables a user to provide data science packages to a marketplace as well as retrieve data science packages created by other users from the marketplace. Further, the data science system can customize a data science package obtained from the marketplace to perform data science operations using data belonging to the user or using user-specified parameters.