Machine Learning Task Recommendation From Common-Format Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Non-experts in machine learning face difficulties in selecting appropriate machine learning tasks suitable for their data and purposes due to the complexity and variety of available recipes, leading to ineffective model training.
Innovation Solution
A machine learning utilization support apparatus that converts user datasets into a common data format, retrieves relevant machine learning tasks from a database, and suggests these tasks with associated descriptions, using techniques like caption generation and vector similarity retrieval to facilitate task selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple machine learning recipes are provided to support task selection, then the coverage of supported tasks increases, but the difficulty of selecting an appropriate task increases for non-experts
Solution Approach 1:
The patent introduces an automated intermediary system that mediates between the user and the multiple machine learning recipes. The system automatically analyzes uploaded data, identifies appropriate tasks, and retrieves relevant recipes, eliminating the need for users to manually navigate through numerous recipe options. This intermediary automation resolves the contradiction by maintaining comprehensive task coverage while removing selection complexity from the user side.
Solution Approach 2:
The system enables self-service by automatically performing data analysis and task identification without requiring user expertise. The automated machine learning task selection process autonomously evaluates the uploaded data characteristics and selects appropriate machine learning tasks and recipes, allowing non-experts to access the full range of supported tasks without facing selection difficulties.
2Ease of operation
If automated machine learning simplifies the modeling process, then the accessibility to non-experts improves, but the ability to handle complex task selection deteriorates
Solution Approach 1:
The system performs preliminary actions by automatically analyzing the uploaded data and identifying suitable machine learning tasks before the user needs to make selections. The automated task selection process is executed in advance, preparing task recommendations and retrieving appropriate recipes so that when users interact with the system, the complex task selection has already been handled, maintaining both accessibility and handling capability.
Solution Approach 2:
The patent replaces the mechanical system of manual task selection with an automated computational system. Instead of requiring users to manually evaluate and select from multiple machine learning tasks, the system uses automated data analysis and algorithmic task identification to handle complex task selection, thereby maintaining accessibility for non-experts while preserving the ability to handle complex scenarios through computational intelligence.
3Adaptability or versatility
If machine learning task selection is left to user discretion, then the flexibility of task choice increases, but the time required for task selection increases
Solution Approach 1:
The system implements feedback mechanisms where the automated task selection process continuously analyzes user uploads, data characteristics, and task requirements to provide accurate task recommendations. This feedback loop enables the system to quickly identify appropriate tasks and retrieve relevant recipes, maintaining flexible task choice while minimizing selection time through intelligent, data-driven recommendations rather than requiring users to explore all available options.
Data Source
AI summary
According to one embodiment, a machine learning utilization support apparatus includes a database and a first processing circuit. The database stores a candidate of a machine learning task and second data having a common data format and related to a description of the candidate in association with each other in advance. The first processing circuit receives a dataset, converts each sample of the dataset to first data in the common data format, retrieves a candidate of a machine learning task associated with the second data corresponding to the first data in the database based on the first data, and generates suggestion data including a candidate of the machine learning task and a description of the candidate based on a retrieval result of the candidate.


