Multilingual Ranking Model Training for Balanced Product Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cross-border e-commerce websites face challenges in providing accurate search rankings due to uneven data distribution across languages, where languages with large data volumes dominate, leading to less frequent feedback for minority languages, affecting querying and sales.

Innovation Solution

A method to construct a ranking model by deleting search-associated data of target language sites, filtering feature data according to a preset rule, and training the model on balanced training data to account for different language sites, using techniques like neural networks and activation functions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If search-associated data from all language sites is used for training the ranking model, then the model can learn from diverse linguistic patterns, but the uneven data distribution causes majority languages to dominate the training process, leading to poor ranking performance for minority languages

Engineering Contradiction:
Improvemulti-language support capabilityVSAvoidranking accuracy for minority languages
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the training data by language group, separating majority languages from minority languages. This segmentation allows the system to apply different sampling strategies to different language groups, preventing majority languages from overwhelming the training process while still maintaining multi-language support capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by implementing language-specific sampling ratios. Minority languages are assigned higher sampling ratios to ensure adequate representation in the training set, while majority languages use lower sampling ratios. This localized adjustment of data quality ensures each language group receives appropriate attention during model training.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If data from majority languages is included in the training set, then the model benefits from larger data volume, but minority language product objects receive insufficient training attention and rank lower in search results

Engineering Contradiction:
Improvetraining data volumeVSAvoidranking accuracy for minority languages
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent changes the parameter of data sampling ratio based on language category. By dynamically adjusting the sampling ratio parameter for different language groups, the system maintains an appropriate balance between majority and minority language data in the training set, ensuring both sufficient overall data volume and adequate minority language representation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies partial action by selectively emphasizing minority language data through higher sampling ratios. This partial emphasis ensures that minority languages receive sufficient training attention without completely excluding majority language data, achieving a balanced training approach that addresses the data volume imbalance.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of manufacture

If the ranking model is trained on unfiltered search-associated data, then the training process is simpler, but the model fails to account for language-specific characteristics and performs poorly across different language sites

Engineering Contradiction:
Improvemodel training simplicityVSAvoidcross-language ranking performance
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-categorizing languages into majority and minority groups before training. This preliminary classification enables the system to apply different sampling strategies tailored to each language group's characteristics, improving cross-language ranking performance while maintaining relatively simple training procedures through automated categorization.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260105379A1Ranking model modeling, product object search method, device, and medium
Publication Date: 2026.04.16 HANGZHOU ALIBABA INT INTERNET IND CO LTD
  • US20260105379A1 patent drawing
  • US20260105379A1 patent drawing
  • US20260105379A1 patent drawing

AI summary

Embodiments of the present disclosure provide ranking model modeling, a product object search method, a device, and a medium; The method includes: deleting search-associated data of a target language site from search-associated data to obtain target search-associated data; filtering feature data from the target search-associated data according to a preset rule to construct training data; constructing a ranking model; and training the ranking model based on the training data to obtain a ranking model that satisfies specified conditions. This can remove the influence of the target language on the data, make the target search-associated data for each language more balanced, eliminate the impact of uneven data distribution across various sites on low-traffic languages, and result in a ranking model that supports various languages, enabling accurate feedback for searches in various languages;