Descriptor Creation Unit for Multi-Granularity Feature Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data mining methods face challenges in generating feature candidates efficiently using multiple sets of table data with varying granularities, requiring significant labor and expertise from experienced technicians to define processing methods.
Innovation Solution
An information processing system that includes a table storage unit and a descriptor creation unit, which generates feature descriptors by combining mapping conditions and reduction methods for rows across multiple tables, enabling the creation of feature candidates that can influence an objective variable.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple sets of table data with different granularities are used to generate feature candidates, then the quantity and diversity of feature candidates increase, but the complexity of processing and the labor required from technicians increase significantly
Solution Approach 1:
The patent segments the complex feature generation process into distinct modular components: a descriptor creation unit that automatically generates feature descriptors from multiple data sets, and a feature generation unit that applies these descriptors. This segmentation allows each component to handle specific tasks independently, reducing overall processing complexity while maintaining the ability to process multiple data sets with different granularities.
Solution Approach 2:
The patent performs preliminary action by pre-defining descriptor templates that specify how features should be generated from multiple data sets. These templates include pre-configured mapping conditions and combination rules that are established before the actual feature generation process, eliminating the need for technicians to manually define processing methods for each feature candidate.
2Reliability
If multiple sets of table data with different granularities are used to generate feature candidates, then the quality and comprehensiveness of predictive features improve, but the manual labor and expertise required increase tremendously
Solution Approach 1:
The patent implements self-service through the automatic descriptor creation unit that autonomously generates feature descriptors by processing multiple data sets according to predefined templates. The system performs self-adjustment by automatically handling the complexity of integrating data with different granularities, eliminating the need for experienced technicians to manually define processing methods while maintaining high predictive accuracy.
Solution Approach 2:
The patent replaces the mechanical system of manual feature engineering with an automated information processing system. The descriptor creation unit uses algorithmic processing to automatically generate feature descriptors from multiple data sets, substituting the manual expertise and labor of technicians with automated computational processes that maintain or improve predictive accuracy.
3Productivity
If feature descriptors are created by generating combinations of mapping conditions and reduction methods, then the number of feature candidates increases efficiently, but the computational complexity increases
Solution Approach 1:
The patent applies universality by creating descriptor templates that can be universally applied to generate multiple feature candidates from the same set of data. Each descriptor template defines a reusable pattern for combining mapping conditions and reduction methods, allowing the system to efficiently generate numerous feature candidates without proportionally increasing computational complexity, as the same template logic is applied repeatedly.
Data Source
AI summary
A table storage unit 81 stores a first table including an objective variable and a second table different in granularity from the first table. A descriptor creation unit 82 creates a feature descriptor for generating a feature which is a variable that can influence the objective variable, from the first table and the second table. The descriptor creation unit 82 creates a plurality of feature descriptors, each by generating a combination of a mapping condition element indicating a mapping condition for rows in the first table and the second table and a reduction method element indicating a reduction method for reducing, for each objective variable, data of each column included in the second table.


