Parallel Corpora Generation for Ecommerce Translation Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Ecommerce transactions face challenges in ensuring accurate information translation across different languages, leading to potential lost sales or unhappy customers due to inaccuracies in responses to buyer inquiries.
Innovation Solution
The use of context information from past purchases and communications to improve machine translation in ecommerce transactions, leveraging metadata such as item IDs, titles, and barcodes to create parallel corpora and align item descriptions across languages, facilitating accurate translation through a machine translation system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine translation is used for ecommerce transactions, then translation speed and automation are improved, but translation accuracy deteriorates leading to lost sales or unhappy customers
Solution Approach 1:
The system performs preliminary actions by collecting and storing translation memories from past ecommerce transactions, product descriptions, and communications in a structured database. This pre-prepared contextual data is then automatically retrieved during translation tasks to improve accuracy without sacrificing speed
Solution Approach 2:
The patent introduces translation memory as an intermediary between the machine translation system and the final translation output. The translation memory stores and retrieves contextual information from past transactions, acting as a mediator that enhances translation accuracy by providing domain-specific terminology and phrasing patterns
2Device complexity
If generic machine translation is used, then implementation simplicity is improved, but translation quality for domain-specific terms deteriorates
Solution Approach 1:
The system applies local quality by creating domain-specific translation memories for different ecommerce categories (e.g., electronics, fashion, groceries). Each category has its own contextual data and terminology patterns stored in the translation memory, allowing the system to maintain simplicity while achieving high-quality translations for specific domains
Solution Approach 2:
The patent changes parameters by dynamically adjusting translation strategies based on the detected domain and context. The system modifies translation parameters such as terminology selection, formality level, and cultural adaptation based on the specific ecommerce context, enabling generic system architecture to deliver specialized translation quality
Data Source
AI summary
A method of forming parallel corpora comprises receiving sets of items in first language and second languages, each of the sets having one or more associated descriptions and metadata. The metadata is collected from the two sets of items and are aligned using the metadata. The aligned metadata are mapped from the first language to the second language for each of the sets. The descriptions of two items are fetched and the structural similarity of the descriptions is measured to assess whether two items are likely to be translations of each other. For mapped items with structurally similar descriptions, the mapped item descriptions are formed into respective sentences in first language and in the second language. The sentences are parallel corpora which may be used to translate an item from the first language to the second language, and also to train a machine translation system.


