Context-aware translation system, method and computer program product
Patent Information
- Application Number
- TW114140931
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-10-21
Smart Images

Figure TWG2TB001909081_001 
Figure TWG2TB001909081_002 
Figure TWG2TB001909081_003
Abstract
Claims
1. A context-aware translation system, comprising: The system comprises: an acquisition module for acquiring multimodal input signals including raw speech, environmental data, speech paralinguistics, nonverbal cues, and dialogue metadata; a 3D environment perception module for generating scene context, tone parameters, and cultural adaptation parameters based on the environmental data, dialogue metadata, speech paralinguistics, and nonverbal cues acquired by the acquisition module; a dynamic strategy combination module for generating an optimal translation version based on the raw speech, dialogue metadata, speech paralinguistics, nonverbal cues, and the scene context, tone parameters, and cultural adaptation parameters generated by the 3D environment perception module; and a self-optimizing feedback closed-loop module for generating feedback information based on the dialogue metadata, speech paralinguistics, nonverbal cues, the scene context generated by the 3D environment perception module, and the optimal translation version generated by the dynamic strategy combination module, to feed this feedback back to the 3D environment perception module and the dynamic strategy combination module, thereby optimizing the scene context, tone parameters, and / or the optimal translation version. The 3D environment perception module includes: The social relationship parsing unit is used to generate the tone parameter based on the scene context, the dialogue metadata, the nonverbal cues, and the speech paralanguage, wherein the tone parameter includes tone type identifier, social relationship description, and tone adjustment parameters; and the cultural adaptation unit is used to generate the cultural adaptation parameter based on the tone parameter and the scene context, wherein the cultural adaptation parameter includes sensitive expression markers, alternative expression rules, and cultural norm metadata.
2. The context-aware translation system as described in claim 1, wherein, The speech sub-language system includes speech rate, tone, and pauses extracted from the original speech, while the non-verbal cues include facial expressions and body language extracted from the original image. The dialogue metadata system includes participant identity, dialogue history, and sentence records. The multimodal input signal also includes the original text.
3. The context-aware translation system as described in claim 1, wherein, The 3D environment perception module system includes: a physical environment sensing unit, used to generate the scene context based on the environment data and the dialogue metadata, wherein the scene context includes scene classification identifiers, domain terminology library references, and scene metadata.
4. The context-aware translation system as described in claim 1, wherein, The dynamic strategy combination module system includes: a multi-strategy parallel generation unit, used to generate multiple translation versions in parallel based on the original speech and the scene context, tone parameters, and cultural adaptation parameters generated by the 3D environment perception module, wherein the multiple translation versions include translation content, version type, target language, and translation metadata; a dialogue impact prediction unit, used to generate a selected translation version based on the dialogue metadata, nonverbal cues, speech paralanguage, and the multiple translation versions generated by the multi-strategy parallel generation unit, wherein the selected translation version includes selected translation content, output format guidelines, and version metadata; and a cross-modal data integration unit, used to generate the optimal translation version based on the nonverbal cues, speech paralanguage, and the selected translation version generated by the dialogue impact prediction unit, wherein the optimal translation version is speech or text.
5. A context-aware translation system as described in claim 1, wherein, The self-optimizing feedback closed-loop module system includes: a micro-correction unit, used to update the personal speech model and output a personal speech model update format based on editing behavior, the dialogue metadata, the optimal translation version generated by the dynamic strategy combination module, and the scene context generated by the 3D environment perception module, wherein the personal speech model update format includes language preference adjustment parameters, user feedback guidance, and model update metadata; a macro-evolution unit, used to generate knowledge base expansion content based on the dialogue metadata, the speech paralinguistics and nonverbal cues, the personal speech model update format output by the micro-correction unit, the optimal translation version generated by the dynamic strategy combination module, and the scene context generated by the 3D environment perception module, wherein the knowledge base expansion content includes terminology and translation mapping, cultural norm updates, scene-specific rules, and metadata; and a cross-user knowledge transfer unit, used to generate anonymous terminology update data based on the knowledge base expansion content generated by the macro-evolution unit and the scene context generated by the 3D environment perception module, wherein the knowledge base expansion content and the anonymous terminology update data constitute the feedback information.
6. A context-aware translation method, comprising: Acquire multimodal input signals including raw speech, environmental data, speech paralinguistics, nonverbal cues, and dialogue metadata; Based on the environmental data, the dialogue metadata, the speech paralinguistics, and the nonverbal cues, a scene context, tone parameters, and cultural adaptation parameters are generated; based on the original speech, the scene context, the tone parameters, the cultural adaptation parameters, the dialogue metadata, the speech paralinguistics, and the nonverbal cues, an optimal translation version is generated; and based on the optimal translation version, the scene context, the dialogue metadata, the speech paralinguistics, and the nonverbal cues, feedback information is generated to provide feedback, thereby optimizing the scene context, the tone parameters, and / or the optimal translation version. The steps of generating the scene context, tone parameters, and cultural adaptation parameters include: generating the tone parameters based on the scene context, the dialogue metadata, the nonverbal cues, and the speech paralinguistics, wherein the tone parameters include tone type identifiers, social relationship descriptions, and tone adjustment parameters; and generating the cultural adaptation parameters based on the tone parameters and the scene context, wherein the cultural adaptation parameters include sensitive expression markers, alternative expression rules, and cultural norm metadata.
7. The context-aware translation method as described in request item 6, wherein, The speech sub-language system includes speech rate, tone, and pauses extracted from the original speech, while the non-verbal cues include facial expressions and body language extracted from the original image. The dialogue metadata system includes participant identity, dialogue history, and sentence records. The multimodal input signal also includes the original text.
8. The context-aware translation method as described in request item 6, wherein, The steps of generating scene context, tone parameters, and cultural adaptation parameters include: generating scene context based on the environmental data and the dialogue metadata, wherein the scene context includes scene classification identifier, domain terminology library reference, and scene metadata.
9. The context-aware translation method as described in request item 6, wherein, The step of generating an optimal translation version includes: generating multiple translation versions in parallel based on the scene context, the tone parameter, the cultural adaptation parameter, and the original speech, wherein the multiple translation versions include translation content, version type, target language, and translation metadata; generating a selected translation version based on the multiple translation versions, the dialogue metadata, the nonverbal cues, and the speech paralanguage, wherein the selected translation version includes selected translation content, output format guidelines, and version metadata; and generating the optimal translation version based on the selected translation version, the nonverbal cues, and the speech paralanguage, wherein the optimal translation version is speech or text.
10. The context-aware translation method as described in request item 6, wherein, The steps for generating feedback information include: updating the personal speech model to output a personal speech model update format based on the editing behavior, the optimal translation version, the scene context, and the dialogue metadata, wherein the personal speech model update format includes language preference adjustment parameters, user feedback guidance, and model update metadata; generating knowledge base expansion content based on the personal speech model update format, the optimal translation version, the scene context, the dialogue metadata, the speech paralinguistics, and the nonverbal cues, wherein the knowledge base expansion content includes terminology and translation mapping, cultural norm updates, scene-specific rules, and metadata; and generating anonymous terminology update data based on the knowledge base expansion content and the scene context, wherein the knowledge base expansion content and the anonymous terminology update data constitute the feedback information.
11. A computer program product, which is loaded onto a computer to execute the context-aware translation method described in any one of requests 6 to 10.
Citation Information
Patent Citations
Methods and systems for multilingual intelligent voice dialogue
CN111128126B
Real-time translating machine capable of selecting words according to context and use method of real-time translating machine
CN116629278A
A natural speech translation system
CN118430513B
An intelligent message reply and delivery system based on generative ai
TW202505418A
System for real-time fact-checking of multimedia content
TW202533071A