Interpretable Feature Engineering with Knowledge-Guided Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in automating feature engineering for machine learning algorithms, particularly in selecting optimal and interpretable features, which is time-consuming and costly, and requires manual intervention by data scientists.

Innovation Solution

A scalable solution that automates feature engineering by using a knowledge graph and a deep reinforcement learning network to derive interpretable and statistically-important features, balancing model performance and interpretability through a bi-objective optimization process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated feature selection systems are implemented, then productivity is improved, but manufacturing precision deteriorates

Engineering Contradiction:
Improvefeature selection efficiencyVSAvoidfeature selection accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces a domain knowledge base as an intermediary between automated feature selection algorithms and domain experts. This knowledge base contains pre-encoded domain-specific rules, relationships, and constraints that guide the automated system to generate features that are both computationally efficient and domain-appropriate, thereby maintaining precision while improving productivity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary encoding of domain knowledge into structured rules and relationships before the feature selection process begins. This pre-prepared knowledge framework enables the automated system to make informed decisions about feature generation without requiring extensive trial-and-error validation, thus improving efficiency while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If manual feature selection is performed, then manufacturing precision is improved, but productivity deteriorates

Engineering Contradiction:
Improvefeature selection accuracyVSAvoidfeature selection efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system enables domain experts to define their own domain knowledge rules and relationships without requiring programming expertise or deep understanding of machine learning algorithms. The automated system then uses these self-defined rules to generate features, combining the domain expertise of experts with the efficiency of automation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The domain knowledge base serves as an intermediary that captures and structures the expertise of manual feature selectors. This knowledge repository allows automated systems to leverage human expertise at scale, maintaining the precision of manual selection while achieving the productivity of automation

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If feature interpretability is enhanced, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvemodel interpretabilityVSAvoidfeature engineering complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the feature engineering process into distinct components: domain knowledge definition, feature generation based on rules, and automated validation. This segmentation allows the system to maintain interpretability through rule-based generation while managing complexity through modular processing and automated workflows

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250284981A1Feature engineering based on feature interpretability
Publication Date: 2025.09.11 SAP SE
  • US20250284981A1 patent drawing
  • US20250284981A1 patent drawing
  • US20250284981A1 patent drawing

AI summary

Systems and methods include generation of a first set of features based on a second set of features and a learning network, determination of an interpretability value for each of the first set of features, determination of a performance of a model trained using the first set of features, determination of a reward based on the performance and the interpretability values, and generation of a third set of features based on the first set of features, the learning network and the reward.