Library Search Apparatus for Data Analysis Profile Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data scientists face complexity in selecting appropriate preprocessing and post-processing libraries for data analysis due to varying input data contents across companies and information systems, leading to difficulties in achieving accurate results, especially for novice users.
Innovation Solution
A library search apparatus that utilizes a storage device to store execution result information and past analysis data statistical profiles, allowing a processor unit to specify an analysis target data statistical profile and calculate an effectiveness degree based on similarity, prioritizing library outputs with high effectiveness for data analysis tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data scientists manually search and select appropriate preprocessing and post-processing libraries from thousands of available libraries, then they can achieve accurate data analysis results, but the operation becomes complicated and time-consuming especially for novice users
Solution Approach 1:
The system automatically performs library selection and configuration without requiring manual intervention from data scientists. The automated library selection unit searches for appropriate preprocessing and post-processing libraries based on the analysis purpose and input data characteristics, and the automated library configuration unit automatically configures these libraries, eliminating the need for manual library selection and configuration operations
Solution Approach 2:
The patent introduces an automated library selection unit and automated library configuration unit as intermediary components between the data scientist and the library repository. These intermediary units handle the complex task of library selection and configuration by analyzing input data characteristics and automatically selecting appropriate libraries from thousands of available options, thereby simplifying the user interface and operation for data scientists
2Measurement precision
If preprocessing and post-processing libraries are changed according to specific input data contents, then accurate analysis results can be achieved, but the operation becomes more complex and time-consuming
Solution Approach 1:
The system performs preliminary analysis of input data characteristics before library selection. The automated library selection unit examines the input data characteristics in advance and pre-identifies suitable preprocessing and post-processing libraries, so that when data analysis is needed, the appropriate libraries are already selected and configured, reducing the time required for library selection
Solution Approach 2:
The system uses feedback from input data characteristics to automatically adjust library selection. The automated library selection unit continuously monitors input data characteristics and uses this feedback to select the most appropriate preprocessing and post-processing libraries, enabling rapid adaptation to different data types without manual intervention
3Reliability
If novice data scientists attempt to understand and select from multiple types of analysis libraries, then they can achieve better analysis quality, but the operation becomes difficult and time-consuming
Solution Approach 1:
The automated library selection and configuration system performs self-service by automatically analyzing input data characteristics and selecting appropriate libraries without requiring novice data scientists to understand the complexities of different library types. The system handles the entire library selection and configuration process autonomously, eliminating the need for users to navigate complex library selection procedures
Solution Approach 2:
The patent introduces automated library selection and configuration units as intermediary components that shield novice users from the complexity of library selection. These intermediary units handle the complex task of matching input data characteristics with appropriate library types, preventing users from being overwhelmed by the complexity of choosing from multiple types of analysis libraries
Data Source
AI summary
A computer 101 for searching a library to be used for an application to be used for a data analysis includes a memory 103 for storing information and a CPU 102 connected to the memory 103. The memory 103 stores execution result information including information for specifying a library utilized for a data analysis in the past and a past analysis data statistical profile which is a statistical profile with respect to past analysis data on which the data analysis is performed by using the library. The CPU 102 specifies an analysis target data statistical profile which is a statistical profile with respect to analysis target data that is a data analysis object, specifies an effectiveness degree indicating effectiveness of utilization of the analysis target data for the data analysis with respect to the library based on a similarity between the past analysis data statistical profile of the past analysis data on which the data analysis is performed by using the library and the analysis target data statistical profile, and preferentially outputs information relating to a library where effectiveness is high.


