Neural Key Phrase Extraction Using Hybrid Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing key phrase extraction methods struggle to generalize well across diverse domains due to their narrow domain training, and they face challenges in handling non-cohesive web page structures and varied content formats, such as lists and media captions, which are common on the World Wide Web.

Innovation Solution

A neural key phrase extraction system that uses a transformer-based architecture to model language properties, incorporates visual features, and employs hybrid word embeddings by concatenating ELMo embeddings, position embeddings, and visual features, and then converts these embeddings into n-gram embeddings using convolutional transformers, ultimately calculating key phrase scores using a feedforward layer trained with cross-entropy loss.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a neural model is trained on narrow domain data with author-assigned key phrases, then the model achieves good performance on that specific domain, but the model's ability to generalize to other domains deteriorates

Engineering Contradiction:
Improvekey phrase extraction accuracyVSAvoiddomain generalization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a neural model architecture that can handle multiple document types and domains through a unified framework. The model processes diverse web page structures (lists, media captions, text fragments) using the same neural network components, enabling it to generalize across domains while maintaining extraction accuracy through domain-specific training data

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If web pages are modeled as simple word sequences, then the model architecture remains simple, but the model fails to capture the complex structures and visual features present in diverse web page formats

Engineering Contradiction:
Improvemodel architecture complexityVSAvoidstructural and visual feature information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies dimensionality change by extending the traditional word sequence model to incorporate multiple dimensions of information. Specifically, it adds visual feature dimensions (through CNN-visual features) and positional dimensions (through position embeddings) to the linguistic sequence, transforming a one-dimensional word sequence into a multi-dimensional representation that captures spatial, visual, and structural properties of web pages

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Use of energy by moving object

If the model uses only linguistic features, then the processing remains computationally efficient, but the model cannot effectively capture visual prominence and layout information that indicate key phrases

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidkey phrase identification accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent applies merging by combining multiple feature types (linguistic embeddings from ELMo, visual features from CNN, and position embeddings) into a unified hybrid embedding representation. This integrated approach allows the model to leverage both computational efficiency of text processing and the discriminative power of visual features, achieving high key phrase identification accuracy through fused multi-modal representations

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11657223B2Keyphase extraction beyond language modeling
Publication Date: 2023.05.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11657223B2 patent drawing
  • US11657223B2 patent drawing
  • US11657223B2 patent drawing

AI summary

A system for extracting a key phrase from a document includes a neural key phrase extraction model (“BLING-KPE”) having a first layer to extract a word sequence from the document, a second layer to represent each word in the word sequence by ELMo embedding, position embedding, and visual features, and a third layer to concatenate the ELMo embedding, the position embedding, and the visual features to produce hybrid word embeddings. A convolutional transformer models the hybrid word embeddings to n-gram embeddings, and a feedforward layer converts the n-gram embeddings into a probability distribution over a set of n-grams and calculates a key phrase score of each n-gram. The neural key phrase extraction model is trained on annotated data based on a labeled loss function to compute cross entropy loss of the key phrase score of each n-gram as compared with a label from the annotated dataset.