Library Search Apparatus for Data Analysis Profile Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data scientists face complexity in selecting appropriate preprocessing and post-processing libraries for data analysis due to varying input data contents across companies and information systems, leading to difficulties in achieving accurate results, especially for novice users.

Innovation Solution

A library search apparatus that utilizes a storage device to store execution result information and past analysis data statistical profiles, allowing a processor unit to specify an analysis target data statistical profile and calculate an effectiveness degree based on similarity, prioritizing library outputs with high effectiveness for data analysis tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data scientists manually search and select appropriate preprocessing and post-processing libraries from thousands of available libraries, then they can achieve accurate data analysis results, but the operation becomes complicated and time-consuming especially for novice users

Engineering Contradiction:
Improveaccuracy of data analysis resultsVSAvoidcomplexity of library selection operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically performs library selection and configuration without requiring manual intervention from data scientists. The automated library selection unit searches for appropriate preprocessing and post-processing libraries based on the analysis purpose and input data characteristics, and the automated library configuration unit automatically configures these libraries, eliminating the need for manual library selection and configuration operations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an automated library selection unit and automated library configuration unit as intermediary components between the data scientist and the library repository. These intermediary units handle the complex task of library selection and configuration by analyzing input data characteristics and automatically selecting appropriate libraries from thousands of available options, thereby simplifying the user interface and operation for data scientists

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If preprocessing and post-processing libraries are changed according to specific input data contents, then accurate analysis results can be achieved, but the operation becomes more complex and time-consuming

Engineering Contradiction:
Improveaccuracy of analysis resultsVSAvoidtime to find appropriate libraries
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of input data characteristics before library selection. The automated library selection unit examines the input data characteristics in advance and pre-identifies suitable preprocessing and post-processing libraries, so that when data analysis is needed, the appropriate libraries are already selected and configured, reducing the time required for library selection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from input data characteristics to automatically adjust library selection. The automated library selection unit continuously monitors input data characteristics and uses this feedback to select the most appropriate preprocessing and post-processing libraries, enabling rapid adaptation to different data types without manual intervention

Inventive Principle:
Principle #23Feedback

3Reliability

If novice data scientists attempt to understand and select from multiple types of analysis libraries, then they can achieve better analysis quality, but the operation becomes difficult and time-consuming

Engineering Contradiction:
Improvequality of data analysis applicationVSAvoidcomplexity of library selection process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The automated library selection and configuration system performs self-service by automatically analyzing input data characteristics and selecting appropriate libraries without requiring novice data scientists to understand the complexities of different library types. The system handles the entire library selection and configuration process autonomously, eliminating the need for users to navigate complex library selection procedures

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces automated library selection and configuration units as intermediary components that shield novice users from the complexity of library selection. These intermediary units handle the complex task of matching input data characteristics with appropriate library types, preventing users from being overwhelmed by the complexity of choosing from multiple types of analysis libraries

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11010393B2Library search apparatus, library search system, and library search method
Publication Date: 2021.05.18 HITACHI LTD
  • US11010393B2 patent drawing
  • US11010393B2 patent drawing
  • US11010393B2 patent drawing

AI summary

A computer 101 for searching a library to be used for an application to be used for a data analysis includes a memory 103 for storing information and a CPU 102 connected to the memory 103. The memory 103 stores execution result information including information for specifying a library utilized for a data analysis in the past and a past analysis data statistical profile which is a statistical profile with respect to past analysis data on which the data analysis is performed by using the library. The CPU 102 specifies an analysis target data statistical profile which is a statistical profile with respect to analysis target data that is a data analysis object, specifies an effectiveness degree indicating effectiveness of utilization of the analysis target data for the data analysis with respect to the library based on a similarity between the past analysis data statistical profile of the past analysis data on which the data analysis is performed by using the library and the analysis target data statistical profile, and preferentially outputs information relating to a library where effectiveness is high.