Multilingual Ranking Model Training for Balanced Product Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cross-border e-commerce websites face challenges in providing accurate search rankings due to uneven data distribution across languages, where languages with large data volumes dominate, leading to less frequent feedback for minority languages, affecting querying and sales.
Innovation Solution
A method to construct a ranking model by deleting search-associated data of target language sites, filtering feature data according to a preset rule, and training the model on balanced training data to account for different language sites, using techniques like neural networks and activation functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If search-associated data from all language sites is used for training the ranking model, then the model can learn from diverse linguistic patterns, but the uneven data distribution causes majority languages to dominate the training process, leading to poor ranking performance for minority languages
Solution Approach 1:
The patent segments the training data by language group, separating majority languages from minority languages. This segmentation allows the system to apply different sampling strategies to different language groups, preventing majority languages from overwhelming the training process while still maintaining multi-language support capability.
Solution Approach 2:
The patent applies local quality by implementing language-specific sampling ratios. Minority languages are assigned higher sampling ratios to ensure adequate representation in the training set, while majority languages use lower sampling ratios. This localized adjustment of data quality ensures each language group receives appropriate attention during model training.
2Quantity of substance
If data from majority languages is included in the training set, then the model benefits from larger data volume, but minority language product objects receive insufficient training attention and rank lower in search results
Solution Approach 1:
The patent changes the parameter of data sampling ratio based on language category. By dynamically adjusting the sampling ratio parameter for different language groups, the system maintains an appropriate balance between majority and minority language data in the training set, ensuring both sufficient overall data volume and adequate minority language representation.
Solution Approach 2:
The patent applies partial action by selectively emphasizing minority language data through higher sampling ratios. This partial emphasis ensures that minority languages receive sufficient training attention without completely excluding majority language data, achieving a balanced training approach that addresses the data volume imbalance.
3Ease of manufacture
If the ranking model is trained on unfiltered search-associated data, then the training process is simpler, but the model fails to account for language-specific characteristics and performs poorly across different language sites
Solution Approach 1:
The patent applies preliminary action by pre-categorizing languages into majority and minority groups before training. This preliminary classification enables the system to apply different sampling strategies tailored to each language group's characteristics, improving cross-language ranking performance while maintaining relatively simple training procedures through automated categorization.
Data Source
AI summary
Embodiments of the present disclosure provide ranking model modeling, a product object search method, a device, and a medium; The method includes: deleting search-associated data of a target language site from search-associated data to obtain target search-associated data; filtering feature data from the target search-associated data according to a preset rule to construct training data; constructing a ranking model; and training the ranking model based on the training data to obtain a ranking model that satisfies specified conditions. This can remove the influence of the target language on the data, make the target search-associated data for each language more balanced, eliminate the impact of uneven data distribution across various sites on low-traffic languages, and result in a ranking model that supports various languages, enabling accurate feedback for searches in various languages;


