Corpus Pair Embedding for Automated Term Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual data entry for SEO and keyword mapping is inefficient, leading to long lead times and scaling issues, as different companies use varying terminology for similar concepts, making it difficult to optimize content for search engines effectively.

Innovation Solution

The system facilitates automated and unsupervised outside-in term mapping for taxonomies and content using corpus pairs, leveraging trained models to generate embedded representations and affinity scores for equivalent terms across multiple corpora, enabling efficient content personalization and search optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data entry is used for SEO and keyword mapping, then accuracy of term mapping can be maintained, but productivity is reduced and lead times increase

Engineering Contradiction:
Improveterm mapping accuracyVSAvoidmapping process efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs automated term mapping by having the computational model independently analyze corpora, extract terms, generate embeddings, and identify equivalent terms without human intervention. The model serves itself by learning from the data and making mapping decisions autonomously, eliminating the need for manual data entry while maintaining consistency through algorithmic processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual data entry with an automated computational system. Instead of human operators manually entering and mapping keywords, the system uses trained models to generate embedded representations, compute affinity scores, and automatically identify equivalent terms across corpora, substituting human mechanical work with automated information processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual keyword mapping is performed across multiple corpora, then term equivalence can be accurately identified, but the process cannot scale effectively

Engineering Contradiction:
Improveterm equivalence identificationVSAvoidscaling capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system creates a universal term mapping framework that can process multiple corpora simultaneously using the same computational approach. The trained model generates embedded representations and affinity scores that work across different domain corpora, allowing the system to handle diverse datasets with varying terminology while maintaining consistent mapping quality and enabling scalable deployment across numerous corpora.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transforms the mapping problem into a computational parameter optimization task by using embedded representations and affinity scores. By changing from manual categorical mapping to continuous vector space operations with learnable parameters, the system achieves both precision in term equivalence identification and scalability through automated parameter-based comparisons that can be rapidly computed across large numbers of corpora.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If different terminology variations are used across companies for similar concepts, then company-specific expertise is preserved, but content searchability and optimization become difficult

Engineering Contradiction:
Improveterminology diversityVSAvoidcontent searchability
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system introduces embedded representations as an intermediary layer between different company terminologies. Instead of directly comparing diverse terminology strings, the system transforms terms into vector embeddings that capture semantic meaning, then uses affinity scores as a mediator to measure equivalence. This intermediary approach preserves terminology diversity while enabling searchability through the shared embedding space where similar concepts converge regardless of surface-level wording differences.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11436487B2Joint embedding of corpus pairs for domain mapping
Publication Date: 2022.09.06 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11436487B2 patent drawing
  • US11436487B2 patent drawing
  • US11436487B2 patent drawing

AI summary

Techniques for outside-in mapping for corpus pairs are provided. In one example, a computer-implemented method comprises: inputting first keywords associated with a first domain corpus; extracting a first keyword of the first keywords; inputting second keywords associated with a second domain corpus; generating an embedded representation of the first keyword via a trained model and generating an embedded representation of the second keywords via the trained model; and scoring a joint embedding affinity associated with a joint embedding. The scoring the joint embedding affinity comprises: transforming the embedded representation of the first keyword and the embedded representation of the second keywords via the trained model; determining an affinity value based on comparing the first keyword to the second keywords; and based on the affinity value, aggregating the joint embedding of the embedded representation of the first keyword and the embedded representation of the second keywords within the second domain corpus.