Data Analysis Support Apparatus for Series Data Preprocessing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analysis techniques face inefficiencies in preprocessing series data, particularly in extracting series characteristic amounts from large datasets with varying information structures, leading to a heavy workload for analysts.
Innovation Solution
A data analysis support apparatus that identifies analytical records and generates analytical series data by adding records associated with response and explanatory variables, facilitating data transformation and analysis through aggregation and correlation coefficient calculations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If analysts perform manual preprocessing and data transformation on large amounts of series data, then they can extract series characteristic amounts, but the workload becomes heavy and time-consuming
Solution Approach 1:
The system performs preliminary actions by automatically identifying analytical records and generating analytical series data before the actual analysis work. The data transformation and structuring are done in advance, so when analysts need to extract series characteristic amounts, the data is already prepared in the required format, eliminating the need for manual preprocessing and significantly reducing both time and workload.
Solution Approach 2:
The system enables self-service by automatically performing data transformation and structuring operations that would otherwise require manual analyst intervention. The automated identification of analytical records and generation of analytical series data allows the system to serve itself in preparing the data, freeing analysts from heavy preprocessing work while maintaining high extraction accuracy.
2Ease of manufacture
If series data is stored in conventional table formats, then data storage is simple, but the data structure is not suited for efficient extraction of series characteristic amounts
Solution Approach 1:
The system segments the series data by identifying analytical records and generating analytical series data with specific structures. Instead of treating all data uniformly, the system divides and structures data based on analytical requirements, creating optimized data segments that are tailored for efficient extraction of series characteristic amounts while maintaining storage simplicity.
Solution Approach 2:
The system changes the data structure parameters by transforming conventional table formats into analytical series data formats. This parameter change involves reorganizing data attributes and relationships to optimize for extraction efficiency, allowing the same data to be stored simply while being structured for high-productivity analysis operations.
3Reliability
If analysts perform data transformation on large datasets, then they can prepare data for analysis, but the preprocessing workload increases significantly
Solution Approach 1:
The system performs self-service by automatically executing data transformation and preparation operations. Instead of requiring analysts to manually perform complex preprocessing steps, the system autonomously identifies analytical records, transforms data formats, and generates analytical series data, thereby maintaining high data preparation quality while eliminating the complexity burden from analysts.
Solution Approach 2:
The system introduces an intermediary layer that automatically handles the complex transformation between raw series data and analysis-ready data structures. This intermediary process manages the preprocessing complexity internally, providing high-quality prepared data to analysts without exposing them to the underlying complexity of transformation operations.
Data Source
AI summary
Series data in a table form and which includes a plurality of records each having a value of a response variable, response variable series identification information, a value of an explanatory variable, and explanatory variable series identification information identifying series of the explanatory variable associated with one another is stored. An analytical record that is any of the records and that contains the value of the response variable or the value of the explanatory variable possibly influencing the response variable at a time of analyzing the response variable is identified; and an additional record having the value of the response variable or the value of the explanatory variable in the identified analytical record associated with the value of the response variable in a predetermined record that is any of the records is generated, and analytical series data obtained by adding the generated additional record to the series data is generated.


