Automated Predictive Model Fitting for Spreadsheet Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analysts without specialized skills face difficulties in analyzing large datasets due to the complexity of working with big data, requiring advanced skills in SQL, programming, and cloud computing, leading to missed opportunities and organizational value decline.
Innovation Solution
A system comprising a server, execution server, and model server that automatically proposes and applies predictive models to spreadsheet files, utilizing parallel processing and machine learning to generate predictive outputs from datasets, making it accessible to users with spreadsheet skills.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If analysts use traditional spreadsheet applications to analyze large datasets, then ease of operation is maintained, but productivity deteriorates due to application freezes and crashes
Solution Approach 1:
The patent introduces a server as an intermediary between the spreadsheet application and the large dataset. The server handles data processing tasks that would otherwise overwhelm the local spreadsheet application, allowing analysts to maintain the familiar spreadsheet interface while achieving the computational power needed for large datasets.
Solution Approach 2:
The patent moves the computational workload from the local machine dimension to the server dimension. By distributing data processing across multiple servers and utilizing cloud infrastructure, the system handles large datasets without freezing or crashing the local spreadsheet application.
2Productivity
If analysts acquire advanced skills in SQL, programming, and cloud computing to work with large datasets, then productivity improves, but device complexity increases
Solution Approach 1:
The system performs automatic model fitting and algorithm selection without requiring user expertise in SQL, programming, or cloud computing. The server automatically processes the dataset, fits appropriate predictive models, and generates results, allowing analysts to maintain their spreadsheet skills while achieving big data analytics capabilities.
Solution Approach 2:
The patent divides the complex data analysis task into separate automated components: data preprocessing, model selection, model fitting, and result generation. Each component is handled automatically by the server system, eliminating the need for analysts to master multiple complex technologies while maintaining high productivity.
3Ease of operation
If spreadsheet applications attempt to process large datasets directly, then ease of operation is maintained, but reliability deteriorates due to errors and crashes
Solution Approach 1:
The server acts as a reliable intermediary that handles the computational burden of processing large datasets. This prevents spreadsheet application crashes and errors while maintaining the familiar and easy-to-use spreadsheet interface for analysts.
Solution Approach 2:
By relocating data processing to the server dimension with adequate computational resources, the system achieves reliable processing of large datasets without compromising the stability of the local spreadsheet application that analysts interact with daily.
Data Source
AI summary
Disclosed is a system, a method, and/or a device of automatic fitting and/or proposal of prediction models to data entries of a spreadsheet file representative of a larger dataset. In one embodiment, a system for automatic determination of a predictive model for scaled data analysis includes two or more servers that process a spreadsheet file including data from a dataset, each data entry of the spreadsheet file including one or more independent variables in one or more cells and a dependent variable. The system automatically determines the predictive model fits a data entry of the spreadsheet. The system proposes an algorithm in response to a fitting of the predictive model, the algorithm accepting as inputs the one or more independent variables and outputting the dependent variable. The system applies the algorithm against the dataset utilizing parallel processing to generate the dependent variable for each data entry of the dataset.


