Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8 results about "Bilingual corpus" patented technology

Machine translation professional data enhancement method, system and equipment and storage medium

The invention relates to the technical field of computers, and discloses a machine translation professional data enhancement method, system and device and a storage medium, and the method comprises the steps: extracting source language terms and corresponding target language terms from a professional dictionary, forming a structured term pair set and constructing a generation prompt, generating a source language text based on a generation model, and storing the source language text into a database; inputting a translation prompt containing a term pair mapping relation in the structured term pair set into a translation model, generating a corresponding target language text, taking the generated source language text and the target language text as sentence pairs, performing evaluation based on a judgment model, obtaining a comprehensive score, comparing the comprehensive score with a preset threshold value, and if the comprehensive score is not lower than the preset threshold value, outputting the comprehensive score. If yes, the sentence pairs are stored in the parallel corpus, and if not, control parameters in the generation prompt or the translation prompt are adjusted according to the evaluation result, and the step of generating the source language text and the target language text is executed again. According to the method, high-quality bilingual corpora can be generated, and the method has self-checking and optimizing capabilities.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

A large language model machine translation optimization method and system

This invention relates to a method and system for optimizing machine translation using a large language model, belonging to the field of machine translation technology. It includes: generating a corresponding first translation from a source sentence in a bilingual corpus; segmenting the source sentence into words, counting the frequency of easily misspelled words in each segment, and calculating the easily misspelled word score, obtaining an easily misspelled word set based on the score; calculating the semantic similarity between the sentence to be translated and multiple candidate examples, and calculating the quality scores of the multiple candidate examples; selecting the optimal k candidate examples from the multiple candidate examples based on semantic similarity and quality scores, and constructing them as prompt templates; obtaining training sentence pairs from the bilingual corpus based on the easily misspelled word scores to construct a training set; and using the training set and prompt templates to perform low-rank adaptive training on a large language model to obtain an optimized large language model. This invention not only improves translation quality but also enhances the interpretability of the translation process.
Owner:SUZHOU UNIV

Bilingual local term extraction and alignment combined learning method based on mixed graph neural network

The invention discloses a bilingual native term extraction and alignment joint learning method based on a mixed graph neural network. The method comprises the steps of obtaining a bilingual corpus data set; constructing a bilingual local term extraction and alignment joint learning model based on the mixed graph neural network, wherein the bilingual local term extraction and alignment joint learning model comprises an input layer, a semantic coding layer, a multi-channel graph generation layer, a multi-head graph attention network layer, a constraint decoding layer, an alignment joint layer and an output layer; dividing the bilingual corpus data set into a training set, a verification set and a test set according to a preset proportion; obtaining an optimal model through the training set and the verification set; and bilingual local terms of the test set are extracted and aligned through the optimal model. The problem that bilingual local term extraction and alignment efficiency is low and precision is insufficient due to the fact that an existing method lacks a dynamic optimization mechanism, semantic sparseness and ambiguity exist, semantic expression is limited by static expression, semantic hierarchy modeling is insufficient and feedback-driven continuous learning is lacked is solved.
Owner:DALIAN UNIVERSITY OF FOREIGN LANGUAGES

A method for constructing a bilingual knowledge-fused green behavior knowledge graph

The application discloses a kind of fusion bilingual knowledge's green-washing behavior knowledge graph construction method, belong to natural language processing and knowledge graph field.The method includes: S1, from enterprise English environmental protection report, Chinese social responsibility report and so on multi-source heterogeneous data collection bilingual corpus constructs bilingual green-washing corpus;S2, based on bilingual pre-training language model carries out named entity recognition and cross-language entity alignment;S3, based on bilingual pre-training language model and green-washing behavior mode constraint rule carries out relationship extraction and cross-language consistency verification;S4, bilingual knowledge is fused, conflict resolution strategy is handled contradictory information, and knowledge reasoning is completed;S5, bilingual green-washing knowledge is stored to graph database in the form of knowledge graph, and based on knowledge graph carries out green-washing behavior evaluation.The application breaks through the language barrier limit, realizes the integrated identification and structured knowledge representation of bilingual green-washing behavior of multinational enterprise.
Owner:LIANYUNGANG OPEN UNIV

A data management system, method, and device for terminology and corpus in translation projects.

This invention provides a data management system for terminology and corpus in translation projects. It includes a terminology identification and organization module for extracting terms from the source text to be translated and obtaining a first high-frequency terminology; a corpus extraction and organization module for crawling monolingual corpus data of the same language type from relevant websites based on the first high-frequency terminology using web crawling technology, and inputting the monolingual corpus into the organization module to obtain a second high-frequency terminology; a terminology and corpus translation module for merging all high-frequency terms and monolingual corpus into a single monolingual file and translating it to obtain a bilingual file; and a database creation and maintenance module for creating a bilingual database of bilingual terms and corpus, storing bilingual terminology pairs and bilingual corpus pairs in the relevant database using a bilingual parallel structure. This solution solves the problem in existing technologies where it is impossible to continuously expand the relevant corpus based on terms from a specific domain identified in the source text, resulting in inaccurate translation results and low translation efficiency.
Owner:SOUTH CHINA NORMAL UNIV

A bilingual corpus alignment method, system, terminal and medium

The application discloses a bilingual corpus alignment method and system, a terminal and a medium, and relates to the technical field of translation, and the technical scheme points are as follows: a first hit rate of an original text sentence and each translation text sentence is calculated, when the first hit rate meets a first threshold group, the original text sentence and the translation text sentence with the highest first hit rate are output according to the original text sentence sequence number. When the first hit rate meets a second threshold group, the second hit rate of the translation text sentence with the highest first hit rate and each original text sentence is calculated, when the second hit rate meets a third threshold group, the original text sentence and the translation text sentence with the highest second hit rate are output according to the original text sentence sequence number. When the second hit rate meets a fourth threshold group, the translation text sentence and the original text sentence with the highest hit rate of translation words and original text sentence words are output according to the original text sentence sequence number. According to the hit rate between the original text sentence and the translation text sentence, the original text and the translation text are aligned in combination with the original text sequence, so that the alignment accuracy is improved, and the alignment efficiency is improved.
Owner:CHENGDU YOUYI INFORMATION TECH CO LTD

Translation processing method, apparatus, device and medium

ActiveUS12572760B2Natural language translationSemantic analysisLanguage representationData science
Embodiments of the present disclosure relate to a translation processing method and apparatus, a device and a medium. The method comprises: generating a multilingual representation model by training according to a monolingual corpus of each language among a plurality of languages, and generating a multilingual generation model according to the monolingual corpus of each language; concatenating the multilingual representation model and the multilingual generation model with a first translation model respectively to generate a target model to be trained; and generating a second translation model by training the target model according to a bilingual corpus among the plurality of languages, and performing translation processing on target information to be processed, according to the second translation model.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

A system and method for Tibetan-Chinese bilingual corpus collaborative annotation and versioned release

The application belongs to the technical field of natural language processing, and relates to a Tibetan-Chinese bilingual corpus collaborative labeling and versioned publishing system and method. Through the system architecture formed by the front-end interaction module, the business processing module, the data storage module and the basic management and control module, relying on the Tibetan-Chinese bilingual labeling, intelligent collaborative management and control, corpus version management and standardized publishing core units integrated by the business processing module, cooperating with the distributed correlation index storage mechanism of the data storage module, the fine-grained permission and operation log management and control function of the basic management and control module, the core technical problems of the Tibetan language characteristics adaptation deficiency, the low efficiency and frequent conflicts of multi-person collaborative labeling, the lack of version tracing and quality grading control of the corpus, the non-standard publishing process and the disconnection of the corpus and the downstream model in the prior art are solved. The accuracy of Tibetan-Chinese bilingual corpus labeling and storage is ensured through exclusive Tibetan adaptation processing, and the orderly promotion of multi-person labeling is realized through modular collaborative management and control.
Owner:SICHUAN TIANFU GAOCHI INFORMATION TECHNOLOGY CO LTD