A domain multi-modal neural machine translation debiasing method based on time-series coupling reinforcement learning

CN122287657APending Publication Date: 2026-06-26KUNMING UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2026-03-19
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing domain-specific multimodal neural machine translation models present a contradiction between achieving accuracy in domain terminology translation and global text coherence. Multimodal large language models lack fine-grained visual and text adaptation, leading to contextual bias, while domain-specific models lack general cross-modal knowledge, making it difficult to balance translation accuracy and coherence.

Method used

We employ a temporally coupled reinforcement learning approach, extracting general and domain-specific cross-modal features through a dual-encoder structure. We utilize a lightweight visual filter and a cross-modal fusion indicator for visual region selection and dynamic fusion, and optimize the translation results by transferring cross-modal knowledge from a multimodal large language model through knowledge distillation.

Benefits of technology

It significantly improves the fidelity and practicality of domain-specific multimodal translation, enhances translation accuracy and text coherence, while greatly improving reasoning efficiency and solving the context bias problem in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287657A_ABST
    Figure CN122287657A_ABST
Patent Text Reader

Abstract

This invention relates to a temporally coupled reinforcement learning-based bias correction method for domain-specific multimodal neural machine translation. It aims to balance the accuracy of domain terminology translation with textual semantics by integrating domain-related visual features and textual semantics, thereby addressing the problem of insufficient fidelity in translation results caused by contextual bias in existing multimodal large language models and domain-specific multimodal neural machine translation methods. This technique proposes a temporally coupled reinforcement learning framework, combining lightweight domain visual indicators and cross-modal domain-guided fusion indicators to achieve accurate selection and dynamic cross-modal fusion of domain-related visual regions, and ensures inference efficiency through a knowledge distillation strategy. This method can generate high-quality translation results in both image-assisted domain-specific and general scenarios, demonstrating excellent accuracy in domain terminology translation and inference efficiency on various benchmark datasets.
Need to check novelty before this filing date? Find Prior Art