Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Parallel corpora" patented technology

Parallel Corpora. The term parallel corpora is typically used in linguistic circles to refer to texts that are translations of each other. And the term comparable corpora refers to texts in two languages that are similar in content, but are not translations.

Machine translation professional data enhancement method, system and equipment and storage medium

The invention relates to the technical field of computers, and discloses a machine translation professional data enhancement method, system and device and a storage medium, and the method comprises the steps: extracting source language terms and corresponding target language terms from a professional dictionary, forming a structured term pair set and constructing a generation prompt, generating a source language text based on a generation model, and storing the source language text into a database; inputting a translation prompt containing a term pair mapping relation in the structured term pair set into a translation model, generating a corresponding target language text, taking the generated source language text and the target language text as sentence pairs, performing evaluation based on a judgment model, obtaining a comprehensive score, comparing the comprehensive score with a preset threshold value, and if the comprehensive score is not lower than the preset threshold value, outputting the comprehensive score. If yes, the sentence pairs are stored in the parallel corpus, and if not, control parameters in the generation prompt or the translation prompt are adjusted according to the evaluation result, and the step of generating the source language text and the target language text is executed again. According to the method, high-quality bilingual corpora can be generated, and the method has self-checking and optimizing capabilities.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Java-to-Cangjie code translation method based on large model and compiling feedback

The invention discloses a Java-to-Cangjie code translation method based on a large model and compilation feedback, and particularly belongs to the technical field of software engineering program language processing, and the method comprises the following steps: step 1, performing structured semantic pre-training by constructing a grammar knowledge base of a target language, and injecting grammar prior knowledge of the target language; step 2, performing semantic enhanced supervision fine tuning training by constructing a high-quality data set containing semantic information, and enhancing semantic alignment and cross-language migration ability of the model; 3, introducing an AST structure perception embedded prompt mechanism in a parallel corpus supervision fine tuning training stage, and guiding the model to perform structure perception translation; and 4, establishing a compiler feedback repair loop, and iteratively correcting output based on error information to form a self-optimized closed-loop system. According to the method, an efficient training path is constructed, dependence on large-scale parallel corpora is effectively reduced, and an extensible and high-reliability technical path is provided for cross-language code translation of low-resource programming languages.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Novel translation model reasoning method and system based on rwkv

PendingCN122287660AImplement reasoning methodsscale upComputation complexityTheoretical computer science
The RWKV-based novel translation model inference method and system includes the following steps: 1) Collecting novel translations and extracting parallel corpora, and using dynamic MicroBatch concatenation technology for sequence compression; 2) Introducing a lightweight group query attention mechanism on the basis of the RWKV architecture, directly obtaining KV information from the Embedding layer to build the model; 3) Employing a sublinear complexity hybrid parallel training mode, combined with a global scalar scaling FP16 mixed precision strategy for training; 4) Applying a hierarchical distributed heterogeneous architecture, offloading the optimizer to low-performance devices and performing gradient compression transmission; 5) Outputting the translation using a joint decoder and dynamic batch inference technology. This invention, through the above method and system, effectively reduces the computational complexity and memory usage in the long text translation process, improves model training efficiency and inference throughput, and significantly improves the translation efficiency and contextual coherence of ultra-long texts.
Owner:LIAONING UNIVERSITY

Non-parallel corpus-oriented cross-language document theme alignment method

The invention belongs to the technical field of natural language processing, and discloses a non-parallel corpus-oriented cross-language document theme alignment method, which comprises the following steps of: 1, performing data preprocessing and data enhancement on a non-parallel corpus cross-language data set to obtain document pairs with similar semantics; the method comprises the following steps of 1, establishing a topic inference network, 2, encoding a document to obtain a document vector representation, 3, establishing the topic inference network, and inputting the document vector representation into the network to obtain document-topic distribution; 4, training the network by taking Dirichlet prior loss, intra-language comparison loss and cross-language comparison loss as a joint optimization target; according to the method, by constraining the theme consistency of the same-language documents and the theme distribution closeness of cross-language semantic similar documents at the same time, different-language documents are aligned in a shared theme space, so that cross-language document theme alignment under the non-parallel corpus condition is realized. According to the method and the device, cross-language topic alignment and consistent topic representation can be realized under the condition that parallel corpora or dictionary resources are not needed.
Owner:NANJING UNIV OF POSTS & TELECOMM

Course guiding type multi-task learning large language model fine tuning method

The invention provides a course-guided multi-task learning large language model fine tuning method, which comprises the following steps of: inputting monolingual corpus into a large language model, and performing restoration clean text training by minimizing a prediction error to obtain a first model; inputting the parallel corpora into the first model, and performing specified term generation training through cross entropy loss to obtain a second model; and inputting the bilingual parallel corpora into the second model, and performing cross-language output training through a regularization item loss function to obtain a large language fine tuning model. According to the method, semantic modeling is enhanced through a de-noising auto-encoder task, term alignment knowledge is integrated through a translation task limited by a vocabulary, cross-language format alignment is enhanced through a machine translation task, and a multi-stage shift arrangement training strategy is introduced, so that task interference and disastrous forgetting in multi-task learning are effectively reduced; and the security BLEU and the term coverage rate are obviously improved.
Owner:XINJIANG UNIVERSITY

Automatic construction method of myanmar-english parallel corpus in large model based on multi-dimensional evaluation

PendingCN122287656AData setModel translation
This invention relates to a method for automatically constructing a large-scale Chinese-Myanmar parallel corpus based on multidimensional evaluation, belonging to the field of natural language processing technology. To address the scarcity of Chinese-Myanmar parallel corpus resources, this invention proposes a method for automatically constructing a large-scale Chinese-Myanmar parallel corpus based on multidimensional evaluation. First, Burmese corpora are manually collected to construct a basic dataset. Second, prompt words are designed to generate a candidate set of pseudo-parallel Chinese-Burmese corpora using a large-scale model. Then, the candidate set is filtered using semantic similarity constraints. Finally, the large-scale model is used to score the filtered candidate set from multiple dimensions, and the corpora are ranked and hierarchically filtered based on the scores to obtain high-quality pseudo-parallel corpora. Experimental results show that the corpus constructed by this invention can effectively improve the model's translation performance and significantly enhance its translation quality in low-resource Chinese-Myanmar language environments.
Owner:KUNMING UNIV OF SCI & TECH

Method for constructing myanmar-chinese parallel corpus based on multi-step thinking large model

This invention relates to a method for constructing a large-scale Burmese-Chinese parallel corpus based on multi-step thinking. The invention includes: translating existing Chinese-English parallel corpora using currently available translation models to obtain original English-Burmese and Chinese-Burmese parallel sentence pairs; calculating double-confidence intervals for semantic similarity and perplexity between aligned Burmese-Chinese sentence pairs based on publicly available high-quality Burmese-Chinese parallel corpora; using the selected double-confidence intervals to perform preliminary screening of Burmese-Chinese parallel sentence pairs, forming pre-processed pseudo-parallel sentence pairs; designing a multi-step thinking chain to guide the large-scale model to progressively optimize the pre-processed pseudo-parallel sentence pairs, thereby generating high-quality Burmese-Chinese parallel corpora for training the translation model, thus effectively improving the performance of Burmese-Chinese machine translation. This invention significantly enhances the ability of large language models to construct corpora in Burmese, a low-resource language, and provides an interpretable and transferable technical paradigm for corpus construction in other low-resource languages.
Owner:KUNMING UNIV OF SCI & TECH +4

A method and system for AI-powered intelligent management of sensitive dialogue information on overseas labor dispatch platforms

This invention discloses an AI-powered intelligent management method and system for sensitive information in dialogues on overseas labor dispatch platforms. The method includes: acquiring dialogue flow information from the overseas labor dispatch platform; establishing an AI recognition model based on the dialogue flow information; performing reverse reasoning based on the AI ​​recognition model to generate sensitive risk tracing paths for each statement in the dialogue flow; performing dynamic pruning and cross-cultural context relabeling on the dialogue flow information based on the sensitive risk tracing paths to obtain a relabeled sensitive risk score and a sensitive management reliability index; and generating and outputting a sensitive information management report using the relabeled sensitive risk score and the sensitive management reliability index. This invention utilizes a cross-cultural context transfer mechanism based on semantic preservation constraints and cultural attribute transfer constraints. It trains cultural adaptation mapping parameters using cross-linguistic parallel corpora in the labor dispatch field, transferring source cultural semantic representations to the target cultural semantic space, and combining this with a set of sensitive semantic anchor points corresponding to the destination country for relabeling and scoring.
Owner:HENGZHI SMART DIGITAL TECHNOLOGY (SHANGHAI) CO LTD

Large model adaptive teaching method and system based on causal guided decoupling and multi-modal perceptual learning

The invention discloses a large-model adaptive teaching method and system based on causal guided decoupling and multi-modal perceptual learning. According to the method, two control vectors are decoupled through comparative learning, causal difference extraction and a robust aggregation algorithm according to a non-parallel plain text corpus, a cognitive coefficient is determined according to the knowledge level of a learner, a style coefficient is determined through the emotion of the input voice of the learner, and the learning effect of the learner is improved. The two coefficients perform linear superposition on the two vectors in a shared semantic space of the multi-modal large language model and are injected into the model through an activation vector guiding technology, and the guided activation vectors are synchronously fed back to a plurality of downstream decoders to generate multi-modal output. According to the method, two orthogonal control vectors capable of being continuously adjusted are injected, so that'content-style 'two-dimensional refined control of multi-modal teaching output of AI teachers is realized, the pain points of cognitive mismatch and emotion segmentation in AI teaching are solved, and the dependence on parallel corpora is thoroughly gotten rid of.
Owner:ZHEJIANG UNIV

Machine translation polysemous word translation evaluation method based on semantic item trigger word replacement

PendingCN122088522AReveal errors effectivelyImprove targetingNatural language translationSemantic analysisSentence pairWord sense
Aiming at the defects of an existing test method in the field of word sense disambiguation test of machine translation in the aspects of triggering effective word sense conversion, maintaining original sense and maintaining semantic consistency, the invention provides a machine translation polysemous word translation evaluation method based on semantic item trigger word replacement. The method comprises the following steps: firstly, constructing a polysemy semantic item library, and screening source language sentences containing polysemy from a parallel corpus according to the semantic item library; performing natural language processing on the sentence, and identifying a trigger word of a polysemy special definition item in the sentence; replacing the trigger word with a replacement word to generate a variant sentence; performing machine translation and alignment on the original sentences and the variant sentences to obtain expressions of polysemy words in translations; and calculating the similarity of the polysemy translation expression, and judging whether a polysemy disambiguation error exists or not. Through a trigger word replacement strategy, a semantic controlled contrast sentence pair can be constructed, potential errors of a system in polysemous word translation are effectively revealed, and the pertinence and effectiveness of evaluation are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Speech synthesis adaptation method and device, equipment and storage medium

The embodiment of the invention provides a speech synthesis adaptation method and device, equipment and a storage medium, which can be applied to customer service systems in the insurance fields of finance, medical treatment and the like. According to the method, basic training data and fine tuning data can be obtained firstly. And converting the training text into international phonetic symbols, and learning a cross-language general mapping relation on a phoneme level to obtain general mapping data. The base model is then trained based on the generic mapping data, and fine-tuned using the fine-tuning data. And performing online preference optimization on the basic model according to the prompt set and the multi-target reward function to obtain an adaptive model. According to the method, model fine tuning can be performed through multi-language IPA basic model training in combination with a small amount of paired data in a target language environment, and GRPO online preference optimization based on multi-index rewards is performed, so that the intelligibility, speaker consistency and tone quality of low-resource language synthesis are improved on the premise that large-scale parallel corpora are not needed.
Owner:PING AN TECH (SHENZHEN) CO LTD

Text translation method and device, electronic equipment and storage medium

The present application relates to the technical field of data processing, and provides a text translation method and device, electronic equipment and storage medium, comprising: obtaining a source language text to be translated, performing semantic retrieval on the source language text to be translated based on a pre-constructed whitelist knowledge base and a general bilingual parallel corpus, and obtaining a whitelist reference set; extracting a keyword set from the source language text to be translated, comparing the keyword set with the whitelist reference set, and determining a to-be-supplemented word set that is not covered by the whitelist reference set; for the words in the to-be-supplemented word set, performing retrieval in the general bilingual parallel corpus based on a vocabulary retrieval mode, and obtaining a supplemented reference set; constructing enhanced prompt information based on the whitelist reference set and the supplemented reference set, inputting the enhanced prompt information and the source language text to be translated into a large language model, and obtaining a target language translation result.
Owner:MIDEA NETWORK INFORMATION SERVICE (SHENZHEN) CO LTD

A Chinese-English Machine Translation Method for the Vertical Field of Traditional Chinese Medicine

ActiveCN115660000BEngineeringData mining
This invention discloses a Chinese-English machine translation method in the vertical domain of Traditional Chinese Medicine (TCM), comprising the following steps: 1. Construction of a parallel TCM corpus; 2. Building a neural machine translation model using transfer learning; 3. Processing a TCM terminology database; 4. Construction of a remotely supervised knowledge base; 5. Comprehensive utilization. The advantages of this invention compared to existing technologies are: better utilization of transfer learning strategies, optimization of model parameters, and improvement of model structure. This allows for significant improvements in model training accuracy and efficiency while fully inheriting the advantages of the original pre-trained model and its massive parameters, resulting in a Chinese-English translation model with TCM linguistic characteristics. It utilizes remote supervision to integrate high-quality TCM Chinese-English parallel corpus resources, professional Chinese-English terminology resources, and synonym / synonym resources into a knowledge base. The target language can be translated using only the knowledge base with extremely high accuracy, and it also has excellent merging capabilities for synonyms and synonyms.
Owner:INST OF INFORMATION ON TRADITIONAL CHINESE MEDICINE CACMS

System and method for automatically generating and screening low-resource language parallel corpora

The invention relates to the technical field of natural language processing, and discloses a system and method for automatically generating and screening low-resource language parallel corpora, and the method comprises the steps: constructing a monolingual corpus which comprises a source language sentence set; based on a large language model LLM, multi-candidate translation is conducted on the source language sentence, multi-candidate target language translated sentences are generated, and prompt information of the LLM comprises part-of-speech tagging POS information and inter-line annotation text IGT information obtained after linguistic enhancement is conducted on the source language sentence and an artificially-constructed example parallel sentence pair. According to the system and the method for automatically generating and screening the parallel corpora of the low-resource language, the cross-language semantic similarity can be measured more robustly and accurately by fusing the complementary characteristics of the heterogeneous double encoders. On a noise evaluation set of Uzburk language-Chinese, the method realizes an F1 score of 0.813, and is significantly superior to a baseline method of a single LaBSE or LASER encoder.
Owner:XINJIANG NORMAL UNIVERSITY