Universal Embedding Space for Multilingual Intent Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing intent classification systems face challenges when dealing with machine-translated content, particularly when the machine translator is trained on out-of-domain data, and it is impractical to train separate classifiers for every language encountered.
Innovation Solution
The development of a universal word embedding space that maps words from different languages to the same semantic meaning, allowing for intent classification without the need for language-specific training, using a code-switching corpus and fine-tuning techniques to maintain accuracy across languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If machine translation is used to classify content in languages without dedicated classifiers, then classification capability is extended to multiple languages, but translation accuracy deteriorates for out-of-domain content
Solution Approach 1:
The patent creates a universal classifier that can handle multiple languages directly without requiring separate language-specific classifiers or machine translation. The classifier is trained on multilingual data and can process input in any language while maintaining consistent performance, eliminating the need for translation intermediaries and their associated accuracy problems in out-of-domain contexts.
Solution Approach 2:
The patent eliminates machine translation as an intermediary step by training the classifier to work directly with source language data. Instead of translating content to a target language for classification, the system learns to classify content in its original language, removing the translation intermediary and its associated limitations.
2Measurement precision
If separate classifiers are trained for every language, then classification accuracy is maintained for each language, but device complexity increases significantly
Solution Approach 1:
The patent trains a single universal classifier on multilingual data that can accurately classify content across multiple languages. This eliminates the need for separate classifiers for each language, reducing system complexity from multiple specialized models to one multi-lingual model that maintains high accuracy across different languages.
Solution Approach 2:
The patent merges multiple language-specific classification capabilities into a single unified classifier. By combining training data from multiple languages and merging the learning processes, the system achieves multilingual classification capability without requiring separate classifier systems for each language.
3Adaptability or versatility
If machine translation is used for low-resource languages, then classification capability is provided, but performance deteriorates due to limited training data
Solution Approach 1:
The universal classifier is trained on diverse multilingual data that includes low-resource languages, enabling it to handle these languages directly without machine translation. The classifier learns patterns across languages simultaneously, which improves performance on low-resource languages by leveraging data from resource-rich languages while maintaining accuracy in low-resource contexts.
Solution Approach 2:
The patent changes the training approach by using multilingual training data with varying language resources. The classifier adapts to different language data densities and quality levels during training, which improves its ability to handle low-resource languages with limited training data while maintaining overall classification performance.
Data Source
AI summary
Exemplary embodiments relate to techniques to classify or detect the intent of content written in a language for which a classifier does not exist. These techniques involve building a code-switching corpus via machine translation, generating a universal embedding for words in the code-switching corpus, training a classifier on the universal embeddings to generate an embedding mapping/table; accessing new content written in a language for which a specific classifier may not exist, and mapping entries in the embedding mapping/table to the universal embeddings. Using these techniques, a classifier can be applied to the universal embedding without needing to be trained on a particular language. Exemplary embodiments may be applied to recognize similarities in two content items, make recommendations, find similar documents, perform deduplication, and perform topic tagging for stories in foreign languages.


