Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7 results about "Parallel corpora" patented technology

Parallel Corpora. The term parallel corpora is typically used in linguistic circles to refer to texts that are translations of each other. And the term comparable corpora refers to texts in two languages that are similar in content, but are not translations.

Novel translation model reasoning method and system based on rwkv

PendingCN122287660AImplement reasoning methodsscale upComputation complexityTheoretical computer science
The RWKV-based novel translation model inference method and system includes the following steps: 1) Collecting novel translations and extracting parallel corpora, and using dynamic MicroBatch concatenation technology for sequence compression; 2) Introducing a lightweight group query attention mechanism on the basis of the RWKV architecture, directly obtaining KV information from the Embedding layer to build the model; 3) Employing a sublinear complexity hybrid parallel training mode, combined with a global scalar scaling FP16 mixed precision strategy for training; 4) Applying a hierarchical distributed heterogeneous architecture, offloading the optimizer to low-performance devices and performing gradient compression transmission; 5) Outputting the translation using a joint decoder and dynamic batch inference technology. This invention, through the above method and system, effectively reduces the computational complexity and memory usage in the long text translation process, improves model training efficiency and inference throughput, and significantly improves the translation efficiency and contextual coherence of ultra-long texts.
Owner:LIAONING UNIVERSITY

Automatic construction method of myanmar-english parallel corpus in large model based on multi-dimensional evaluation

PendingCN122287656AData setModel translation
This invention relates to a method for automatically constructing a large-scale Chinese-Myanmar parallel corpus based on multidimensional evaluation, belonging to the field of natural language processing technology. To address the scarcity of Chinese-Myanmar parallel corpus resources, this invention proposes a method for automatically constructing a large-scale Chinese-Myanmar parallel corpus based on multidimensional evaluation. First, Burmese corpora are manually collected to construct a basic dataset. Second, prompt words are designed to generate a candidate set of pseudo-parallel Chinese-Burmese corpora using a large-scale model. Then, the candidate set is filtered using semantic similarity constraints. Finally, the large-scale model is used to score the filtered candidate set from multiple dimensions, and the corpora are ranked and hierarchically filtered based on the scores to obtain high-quality pseudo-parallel corpora. Experimental results show that the corpus constructed by this invention can effectively improve the model's translation performance and significantly enhance its translation quality in low-resource Chinese-Myanmar language environments.
Owner:KUNMING UNIV OF SCI & TECH

Method for constructing myanmar-chinese parallel corpus based on multi-step thinking large model

This invention relates to a method for constructing a large-scale Burmese-Chinese parallel corpus based on multi-step thinking. The invention includes: translating existing Chinese-English parallel corpora using currently available translation models to obtain original English-Burmese and Chinese-Burmese parallel sentence pairs; calculating double-confidence intervals for semantic similarity and perplexity between aligned Burmese-Chinese sentence pairs based on publicly available high-quality Burmese-Chinese parallel corpora; using the selected double-confidence intervals to perform preliminary screening of Burmese-Chinese parallel sentence pairs, forming pre-processed pseudo-parallel sentence pairs; designing a multi-step thinking chain to guide the large-scale model to progressively optimize the pre-processed pseudo-parallel sentence pairs, thereby generating high-quality Burmese-Chinese parallel corpora for training the translation model, thus effectively improving the performance of Burmese-Chinese machine translation. This invention significantly enhances the ability of large language models to construct corpora in Burmese, a low-resource language, and provides an interpretable and transferable technical paradigm for corpus construction in other low-resource languages.
Owner:KUNMING UNIV OF SCI & TECH +4

A method and system for AI-powered intelligent management of sensitive dialogue information on overseas labor dispatch platforms

This invention discloses an AI-powered intelligent management method and system for sensitive information in dialogues on overseas labor dispatch platforms. The method includes: acquiring dialogue flow information from the overseas labor dispatch platform; establishing an AI recognition model based on the dialogue flow information; performing reverse reasoning based on the AI ​​recognition model to generate sensitive risk tracing paths for each statement in the dialogue flow; performing dynamic pruning and cross-cultural context relabeling on the dialogue flow information based on the sensitive risk tracing paths to obtain a relabeled sensitive risk score and a sensitive management reliability index; and generating and outputting a sensitive information management report using the relabeled sensitive risk score and the sensitive management reliability index. This invention utilizes a cross-cultural context transfer mechanism based on semantic preservation constraints and cultural attribute transfer constraints. It trains cultural adaptation mapping parameters using cross-linguistic parallel corpora in the labor dispatch field, transferring source cultural semantic representations to the target cultural semantic space, and combining this with a set of sensitive semantic anchor points corresponding to the destination country for relabeling and scoring.
Owner:HENGZHI SMART DIGITAL TECHNOLOGY (SHANGHAI) CO LTD

Machine translation polysemous word translation evaluation method based on semantic item trigger word replacement

PendingCN122088522AReveal errors effectivelyImprove targetingNatural language translationSemantic analysisSentence pairWord sense
Aiming at the defects of an existing test method in the field of word sense disambiguation test of machine translation in the aspects of triggering effective word sense conversion, maintaining original sense and maintaining semantic consistency, the invention provides a machine translation polysemous word translation evaluation method based on semantic item trigger word replacement. The method comprises the following steps: firstly, constructing a polysemy semantic item library, and screening source language sentences containing polysemy from a parallel corpus according to the semantic item library; performing natural language processing on the sentence, and identifying a trigger word of a polysemy special definition item in the sentence; replacing the trigger word with a replacement word to generate a variant sentence; performing machine translation and alignment on the original sentences and the variant sentences to obtain expressions of polysemy words in translations; and calculating the similarity of the polysemy translation expression, and judging whether a polysemy disambiguation error exists or not. Through a trigger word replacement strategy, a semantic controlled contrast sentence pair can be constructed, potential errors of a system in polysemous word translation are effectively revealed, and the pertinence and effectiveness of evaluation are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Text translation method and device, electronic equipment and storage medium

The present application relates to the technical field of data processing, and provides a text translation method and device, electronic equipment and storage medium, comprising: obtaining a source language text to be translated, performing semantic retrieval on the source language text to be translated based on a pre-constructed whitelist knowledge base and a general bilingual parallel corpus, and obtaining a whitelist reference set; extracting a keyword set from the source language text to be translated, comparing the keyword set with the whitelist reference set, and determining a to-be-supplemented word set that is not covered by the whitelist reference set; for the words in the to-be-supplemented word set, performing retrieval in the general bilingual parallel corpus based on a vocabulary retrieval mode, and obtaining a supplemented reference set; constructing enhanced prompt information based on the whitelist reference set and the supplemented reference set, inputting the enhanced prompt information and the source language text to be translated into a large language model, and obtaining a target language translation result.
Owner:MIDEA NETWORK INFORMATION SERVICE (SHENZHEN) CO LTD

A Chinese-English Machine Translation Method for the Vertical Field of Traditional Chinese Medicine

ActiveCN115660000BEngineeringData mining
This invention discloses a Chinese-English machine translation method in the vertical domain of Traditional Chinese Medicine (TCM), comprising the following steps: 1. Construction of a parallel TCM corpus; 2. Building a neural machine translation model using transfer learning; 3. Processing a TCM terminology database; 4. Construction of a remotely supervised knowledge base; 5. Comprehensive utilization. The advantages of this invention compared to existing technologies are: better utilization of transfer learning strategies, optimization of model parameters, and improvement of model structure. This allows for significant improvements in model training accuracy and efficiency while fully inheriting the advantages of the original pre-trained model and its massive parameters, resulting in a Chinese-English translation model with TCM linguistic characteristics. It utilizes remote supervision to integrate high-quality TCM Chinese-English parallel corpus resources, professional Chinese-English terminology resources, and synonym / synonym resources into a knowledge base. The target language can be translated using only the knowledge base with extremely high accuracy, and it also has excellent merging capabilities for synonyms and synonyms.
Owner:INST OF INFORMATION ON TRADITIONAL CHINESE MEDICINE CACMS