Descriptor Creation Unit for Multi-Granularity Feature Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data mining methods face challenges in generating feature candidates efficiently using multiple sets of table data with varying granularities, requiring significant labor and expertise from experienced technicians to define processing methods.

Innovation Solution

An information processing system that includes a table storage unit and a descriptor creation unit, which generates feature descriptors by combining mapping conditions and reduction methods for rows across multiple tables, enabling the creation of feature candidates that can influence an objective variable.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple sets of table data with different granularities are used to generate feature candidates, then the quantity and diversity of feature candidates increase, but the complexity of processing and the labor required from technicians increase significantly

Engineering Contradiction:
Improvequantity of feature candidatesVSAvoidprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the complex feature generation process into distinct modular components: a descriptor creation unit that automatically generates feature descriptors from multiple data sets, and a feature generation unit that applies these descriptors. This segmentation allows each component to handle specific tasks independently, reducing overall processing complexity while maintaining the ability to process multiple data sets with different granularities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-defining descriptor templates that specify how features should be generated from multiple data sets. These templates include pre-configured mapping conditions and combination rules that are established before the actual feature generation process, eliminating the need for technicians to manually define processing methods for each feature candidate.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple sets of table data with different granularities are used to generate feature candidates, then the quality and comprehensiveness of predictive features improve, but the manual labor and expertise required increase tremendously

Engineering Contradiction:
Improvepredictive accuracyVSAvoidease of feature generation
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements self-service through the automatic descriptor creation unit that autonomously generates feature descriptors by processing multiple data sets according to predefined templates. The system performs self-adjustment by automatically handling the complexity of integrating data with different granularities, eliminating the need for experienced technicians to manually define processing methods while maintaining high predictive accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual feature engineering with an automated information processing system. The descriptor creation unit uses algorithmic processing to automatically generate feature descriptors from multiple data sets, substituting the manual expertise and labor of technicians with automated computational processes that maintain or improve predictive accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If feature descriptors are created by generating combinations of mapping conditions and reduction methods, then the number of feature candidates increases efficiently, but the computational complexity increases

Engineering Contradiction:
Improvefeature generation efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies universality by creating descriptor templates that can be universally applied to generate multiple feature candidates from the same set of data. Each descriptor template defines a reusable pattern for combining mapping conditions and reduction methods, allowing the system to efficiently generate numerous feature candidates without proportionally increasing computational complexity, as the same template logic is applied repeatedly.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10885011B2Information processing system, descriptor creation method, and descriptor creation program
Publication Date: 2021.01.05 DOTDATA INC
  • US10885011B2 patent drawing
  • US10885011B2 patent drawing
  • US10885011B2 patent drawing

AI summary

A table storage unit 81 stores a first table including an objective variable and a second table different in granularity from the first table. A descriptor creation unit 82 creates a feature descriptor for generating a feature which is a variable that can influence the objective variable, from the first table and the second table. The descriptor creation unit 82 creates a plurality of feature descriptors, each by generating a combination of a mapping condition element indicating a mapping condition for rows in the first table and the second table and a reduction method element indicating a reduction method for reducing, for each objective variable, data of each column included in the second table.