Corpus Pair Embedding for Automated Term Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual data entry for SEO and keyword mapping is inefficient, leading to long lead times and scaling issues, as different companies use varying terminology for similar concepts, making it difficult to optimize content for search engines effectively.
Innovation Solution
The system facilitates automated and unsupervised outside-in term mapping for taxonomies and content using corpus pairs, leveraging trained models to generate embedded representations and affinity scores for equivalent terms across multiple corpora, enabling efficient content personalization and search optimization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data entry is used for SEO and keyword mapping, then accuracy of term mapping can be maintained, but productivity is reduced and lead times increase
Solution Approach 1:
The system performs automated term mapping by having the computational model independently analyze corpora, extract terms, generate embeddings, and identify equivalent terms without human intervention. The model serves itself by learning from the data and making mapping decisions autonomously, eliminating the need for manual data entry while maintaining consistency through algorithmic processes.
Solution Approach 2:
The patent replaces the mechanical process of manual data entry with an automated computational system. Instead of human operators manually entering and mapping keywords, the system uses trained models to generate embedded representations, compute affinity scores, and automatically identify equivalent terms across corpora, substituting human mechanical work with automated information processing.
2Measurement precision
If manual keyword mapping is performed across multiple corpora, then term equivalence can be accurately identified, but the process cannot scale effectively
Solution Approach 1:
The system creates a universal term mapping framework that can process multiple corpora simultaneously using the same computational approach. The trained model generates embedded representations and affinity scores that work across different domain corpora, allowing the system to handle diverse datasets with varying terminology while maintaining consistent mapping quality and enabling scalable deployment across numerous corpora.
Solution Approach 2:
The system transforms the mapping problem into a computational parameter optimization task by using embedded representations and affinity scores. By changing from manual categorical mapping to continuous vector space operations with learnable parameters, the system achieves both precision in term equivalence identification and scalability through automated parameter-based comparisons that can be rapidly computed across large numbers of corpora.
3Adaptability or versatility
If different terminology variations are used across companies for similar concepts, then company-specific expertise is preserved, but content searchability and optimization become difficult
Solution Approach 1:
The system introduces embedded representations as an intermediary layer between different company terminologies. Instead of directly comparing diverse terminology strings, the system transforms terms into vector embeddings that capture semantic meaning, then uses affinity scores as a mediator to measure equivalence. This intermediary approach preserves terminology diversity while enabling searchability through the shared embedding space where similar concepts converge regardless of surface-level wording differences.
Data Source
AI summary
Techniques for outside-in mapping for corpus pairs are provided. In one example, a computer-implemented method comprises: inputting first keywords associated with a first domain corpus; extracting a first keyword of the first keywords; inputting second keywords associated with a second domain corpus; generating an embedded representation of the first keyword via a trained model and generating an embedded representation of the second keywords via the trained model; and scoring a joint embedding affinity associated with a joint embedding. The scoring the joint embedding affinity comprises: transforming the embedded representation of the first keyword and the embedded representation of the second keywords via the trained model; determining an affinity value based on comparing the first keyword to the second keywords; and based on the affinity value, aggregating the joint embedding of the embedded representation of the first keyword and the embedded representation of the second keywords within the second domain corpus.


