Feature Taxonomization for Predictive Model Reliability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As online systems with numerous users and features grow in complexity, it becomes challenging to determine which inputs significantly impact predictive capabilities, making it difficult to improve performance metrics for content presentation.

Innovation Solution

An online system computes importance scores for features based on their influence on performance metrics and categorizes them into ranked categories and sub-categories, allowing for a semantically meaningful grouping and reporting of features that impact prediction models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the number of features and inputs in prediction models is increased, then predictive capabilities improve, but the complexity and difficulty of measuring impact increase

Engineering Contradiction:
Improvepredictive capabilitiesVSAvoidcomplexity of inputs and features
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments features into hierarchical categories (first category, second category, third category) with multiple levels. This segmentation organizes the increasing number of features into manageable groups, making it easier to analyze and measure the impact of individual features while maintaining the overall predictive capability of the model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to feature organization, adding category levels beyond simple feature listing. This dimensional change allows the system to handle increased feature complexity by structuring features in multiple layers, thereby maintaining predictive accuracy while improving measurability of feature impacts.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If the number of features is increased, then predictive capabilities improve, but it becomes challenging to determine which inputs modify or change to further improve performance

Engineering Contradiction:
Improvepredictive capabilitiesVSAvoidimpact of different inputs
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

By segmenting features into hierarchical categories, the patent enables systematic measurement of input impacts at different levels. This segmentation allows the system to detect which specific features or category-level groupings most significantly affect performance metrics, making the measurement process more manageable despite the large number of features.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The hierarchical category structure acts as an intermediary between individual features and performance metrics. This intermediary layer facilitates the measurement and analysis of feature impacts by providing structured aggregation points that simplify the detection and measurement process while preserving the detailed information needed for optimization.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11868429B1Taxonomization of features used in prediction models according to sub-categories in a list of ranked categories
Publication Date: 2024.01.09 META PLATFORMS INC
  • US11868429B1 patent drawing
  • US11868429B1 patent drawing
  • US11868429B1 patent drawing

AI summary

An online system accesses a list of features used as input into a predictor to predict a performance metric for content presented to users. The online system computes importance scores for one or more of the features. A ranked list of categories is created, with each category having one or more sub-categories. For each feature having a computed importance score, the online system assigns, for each attribute in the ranked list of attributes for that feature, the feature to a sub-category in one of the categories in the ranked list of categories that has the same rank as the attribute in the ranked list of attributes for the feature, where the sub-category is associated with a label that corresponds with the attribute. For each sub-category in each category, a cumulative score is computed for the sub-category based on the importance scores of the features of that sub-category.