Multilingual Embedding Alignment for Domain-Independent NLP Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cognitive systems face challenges in scaling language understanding models across multiple languages and domains due to their language-dependent and domain-specific nature, requiring extensive resources and expertise for customization and translation.

Innovation Solution

A multi-lingual/domain embedding system aligns embeddings from different languages and domains using parallel vocabularies to generate a transformation matrix, creating cross-domain, multilingual embeddings that can be used to build language and domain-independent artificial intelligence models, enabling applications in new languages and domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional language-dependent preprocessing and feature engineering techniques are used, then model accuracy for specific language and domain is improved, but scalability to multiple languages and domains deteriorates

Engineering Contradiction:
Improvemodel accuracyVSAvoidscalability to multiple languages and domains
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal embedding space that can represent multiple languages and domains simultaneously. By training embeddings on multilingual and multi-domain data, the system achieves a single model architecture that adapts to different languages and domains without requiring separate preprocessing pipelines or feature engineering for each combination, thus resolving the contradiction between specificity and scalability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the parameter representation by using continuous embedding vectors instead of discrete language-specific or domain-specific features. This transformation allows the model to capture semantic relationships across languages and domains through vector space operations, enabling the same model to handle diverse linguistic and domain-specific data without sacrificing accuracy

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If human translation is used to translate data from existing language to another language, then translation accuracy is improved, but time consumption and labor intensity increase

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtime consumption and labor intensity
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses machine translation to create copied versions of training data in multiple languages. Instead of manually translating each dataset, the system automatically generates translated copies using machine translation models, significantly reducing time and labor while maintaining sufficient accuracy for training purposes through subsequent embedding alignment

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of human translation with automated machine translation systems. This substitution eliminates the need for human translators while maintaining the ability to generate high-quality translated training data through computational methods, thereby resolving the contradiction between accuracy and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If machine translation is used to translate data, then time consumption is reduced, but translation reliability and quality deteriorate

Engineering Contradiction:
Improvetranslation speedVSAvoidtranslation reliability and quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces embedding spaces as an intermediary layer between source and target languages. Instead of relying solely on machine translation output, the system maps translated text into embedding vectors that capture semantic meaning, then aligns these embeddings across languages using parallel vocabulary. This intermediary representation preserves translation speed while improving reliability by focusing on semantic equivalence rather than literal translation accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If cognitive systems are customized for specific tasks and domains, then task-specific performance is improved, but resource requirements and complexity increase

Engineering Contradiction:
Improvetask-specific performanceVSAvoidresource requirements and customization complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal model architecture that can perform multiple tasks across different domains using the same embedding space. By training on diverse multilingual and multi-domain data, the system achieves a single customizable platform that can be adapted to specific tasks through prompt engineering or fine-tuning rather than requiring separate customized systems, thus reducing overall complexity while maintaining task-specific performance

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11222176B2Method and system for language and domain acceleration with embedding evaluation
Publication Date: 2022.01.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11222176B2 patent drawing
  • US11222176B2 patent drawing
  • US11222176B2 patent drawing

AI summary

A method, system and a computer program product are provided for generating a natural language model that is substantially independent of languages and domains by transforming monolingual embeddings into a multilingual embeddings in a first shared embedding space using a cross-lingual learning process, and then transforming the multilingual embeddings into cross-domain, multilingual embeddings in a second shared embedding space using a cross-domain learning process, where the multilingual embeddings and/or cross-domain, multilingual embeddings are evaluated to measure a degree to which the embeddings associate a set of target concepts with a set of attribute words.