Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

12 results about "Bilingual corpus" patented technology

Machine translation professional data enhancement method, system and equipment and storage medium

The invention relates to the technical field of computers, and discloses a machine translation professional data enhancement method, system and device and a storage medium, and the method comprises the steps: extracting source language terms and corresponding target language terms from a professional dictionary, forming a structured term pair set and constructing a generation prompt, generating a source language text based on a generation model, and storing the source language text into a database; inputting a translation prompt containing a term pair mapping relation in the structured term pair set into a translation model, generating a corresponding target language text, taking the generated source language text and the target language text as sentence pairs, performing evaluation based on a judgment model, obtaining a comprehensive score, comparing the comprehensive score with a preset threshold value, and if the comprehensive score is not lower than the preset threshold value, outputting the comprehensive score. If yes, the sentence pairs are stored in the parallel corpus, and if not, control parameters in the generation prompt or the translation prompt are adjusted according to the evaluation result, and the step of generating the source language text and the target language text is executed again. According to the method, high-quality bilingual corpora can be generated, and the method has self-checking and optimizing capabilities.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

A large language model machine translation optimization method and system

This invention relates to a method and system for optimizing machine translation using a large language model, belonging to the field of machine translation technology. It includes: generating a corresponding first translation from a source sentence in a bilingual corpus; segmenting the source sentence into words, counting the frequency of easily misspelled words in each segment, and calculating the easily misspelled word score, obtaining an easily misspelled word set based on the score; calculating the semantic similarity between the sentence to be translated and multiple candidate examples, and calculating the quality scores of the multiple candidate examples; selecting the optimal k candidate examples from the multiple candidate examples based on semantic similarity and quality scores, and constructing them as prompt templates; obtaining training sentence pairs from the bilingual corpus based on the easily misspelled word scores to construct a training set; and using the training set and prompt templates to perform low-rank adaptive training on a large language model to obtain an optimized large language model. This invention not only improves translation quality but also enhances the interpretability of the translation process.
Owner:SUZHOU UNIV

Domain bilingual sentence pair selection method and system based on theme information

The invention discloses a subject information-based field bilingual sentence pair selection method, which is used for selecting a sentence pair subset related to a to-be-translated text from a large-scale bilingual corpus mixed with fields by virtue of subject correlation between a bilingual sentence pair and a target field so as to train a specific field translation system and improve the translation quality of the field text. The method comprises the following steps: firstly, learning topic vectors of phrase pairs by using context words of the phrase pairs in a bilingual corpus; secondly, for a target domain development set and a candidate bilingual sentence pair, obtaining topic vectors of the target domain development set and the candidate bilingual sentence pair by utilizing the extracted phrase pair set; and finally, the topic relevancy between the candidate bilingual sentence pairs and the field development set text is calculated, and the sentence pairs with high relevancy are preferentially selected as target field training data. The invention further discloses a domain bilingual sentence pair selection system based on the theme information. According to the method, the field-related bilingual sentence pairs are selected by means of the topic relevancy of the text, and the problem of insufficient training data in a specific field is solved.
Owner:ANHUI RADIO & TV UNIV

Bilingual local term extraction and alignment combined learning method based on mixed graph neural network

The invention discloses a bilingual native term extraction and alignment joint learning method based on a mixed graph neural network. The method comprises the steps of obtaining a bilingual corpus data set; constructing a bilingual local term extraction and alignment joint learning model based on the mixed graph neural network, wherein the bilingual local term extraction and alignment joint learning model comprises an input layer, a semantic coding layer, a multi-channel graph generation layer, a multi-head graph attention network layer, a constraint decoding layer, an alignment joint layer and an output layer; dividing the bilingual corpus data set into a training set, a verification set and a test set according to a preset proportion; obtaining an optimal model through the training set and the verification set; and bilingual local terms of the test set are extracted and aligned through the optimal model. The problem that bilingual local term extraction and alignment efficiency is low and precision is insufficient due to the fact that an existing method lacks a dynamic optimization mechanism, semantic sparseness and ambiguity exist, semantic expression is limited by static expression, semantic hierarchy modeling is insufficient and feedback-driven continuous learning is lacked is solved.
Owner:DALIAN UNIVERSITY OF FOREIGN LANGUAGES

A multimodal heterogeneous model retrieval enhancement method and system

The application provides a multimodal heterogeneous model retrieval enhancement method and system, comprising: based on user multimodal query, constructing knowledge and application example bilingual corpus, designing joint retrieval mechanism to get result set; through text, image, audio special processing channel and segmented trapezoidal topology structure Spiking neural network, mapping scheduling to get feature representation; constructing three-level cascade architecture of basic model, advanced model and human expert, combining with deferred and waiver decision mechanism, getting decision path and answer candidate set; using Hamilton graph network to represent multimodal relationship, using gradient-free descent method to quickly train and optimize model parameters; through cross-modal semantic alignment and dynamic retrieval window adjustment, getting enhanced retrieval result; through context perception sorting and retrieval enhancement reasoning, getting high-quality response. The application improves the multimodal information retrieval processing efficiency and heterogeneous model reasoning response quality.
Owner:贵州中汇科技发展有限公司

A method for constructing a bilingual knowledge-fused green behavior knowledge graph

The application discloses a kind of fusion bilingual knowledge's green-washing behavior knowledge graph construction method, belong to natural language processing and knowledge graph field.The method includes: S1, from enterprise English environmental protection report, Chinese social responsibility report and so on multi-source heterogeneous data collection bilingual corpus constructs bilingual green-washing corpus;S2, based on bilingual pre-training language model carries out named entity recognition and cross-language entity alignment;S3, based on bilingual pre-training language model and green-washing behavior mode constraint rule carries out relationship extraction and cross-language consistency verification;S4, bilingual knowledge is fused, conflict resolution strategy is handled contradictory information, and knowledge reasoning is completed;S5, bilingual green-washing knowledge is stored to graph database in the form of knowledge graph, and based on knowledge graph carries out green-washing behavior evaluation.The application breaks through the language barrier limit, realizes the integrated identification and structured knowledge representation of bilingual green-washing behavior of multinational enterprise.
Owner:LIANYUNGANG OPEN UNIV

Large language model machine translation optimization method and system

The invention relates to a large language model machine translation optimization method and system, and belongs to the technical field of machine translation. Comprising the steps of generating a corresponding first translation for a source end sentence in a bilingual corpus; performing word segmentation processing on the source end sentence, counting the occurrence frequency of error-prone vocabularies in each segmented word, calculating error-prone word scores of the error-prone vocabularies, and obtaining an error-prone word set according to the scores; calculating semantic similarity between the sentence to be translated and the plurality of candidate examples, and calculating quality scores of the plurality of candidate examples; according to the semantic similarity and the quality score, screening out the optimal k candidate examples from the plurality of candidate examples, and constructing the optimal k candidate examples into a prompt template; according to the error-prone word score, obtaining training sentence pairs from the bilingual corpus to construct a training set; and performing low-rank adaptation training on the large language model by using the training set and the prompt template to obtain an optimized large language model. According to the method, the translation quality is improved, and the interpretability of the translation process is enhanced.
Owner:SUZHOU UNIV

A data management system, method, and device for terminology and corpus in translation projects.

This invention provides a data management system for terminology and corpus in translation projects. It includes a terminology identification and organization module for extracting terms from the source text to be translated and obtaining a first high-frequency terminology; a corpus extraction and organization module for crawling monolingual corpus data of the same language type from relevant websites based on the first high-frequency terminology using web crawling technology, and inputting the monolingual corpus into the organization module to obtain a second high-frequency terminology; a terminology and corpus translation module for merging all high-frequency terms and monolingual corpus into a single monolingual file and translating it to obtain a bilingual file; and a database creation and maintenance module for creating a bilingual database of bilingual terms and corpus, storing bilingual terminology pairs and bilingual corpus pairs in the relevant database using a bilingual parallel structure. This solution solves the problem in existing technologies where it is impossible to continuously expand the relevant corpus based on terms from a specific domain identified in the source text, resulting in inaccurate translation results and low translation efficiency.
Owner:SOUTH CHINA NORMAL UNIV

A bilingual corpus alignment method, system, terminal and medium

The application discloses a bilingual corpus alignment method and system, a terminal and a medium, and relates to the technical field of translation, and the technical scheme points are as follows: a first hit rate of an original text sentence and each translation text sentence is calculated, when the first hit rate meets a first threshold group, the original text sentence and the translation text sentence with the highest first hit rate are output according to the original text sentence sequence number. When the first hit rate meets a second threshold group, the second hit rate of the translation text sentence with the highest first hit rate and each original text sentence is calculated, when the second hit rate meets a third threshold group, the original text sentence and the translation text sentence with the highest second hit rate are output according to the original text sentence sequence number. When the second hit rate meets a fourth threshold group, the translation text sentence and the original text sentence with the highest hit rate of translation words and original text sentence words are output according to the original text sentence sequence number. According to the hit rate between the original text sentence and the translation text sentence, the original text and the translation text are aligned in combination with the original text sequence, so that the alignment accuracy is improved, and the alignment efficiency is improved.
Owner:CHENGDU YOUYI INFORMATION TECH CO LTD

Translation processing method, apparatus, device and medium

Embodiments of the present disclosure relate to a translation processing method and apparatus, a device and a medium. The method comprises: generating a multilingual representation model by training according to a monolingual corpus of each language among a plurality of languages, and generating a multilingual generation model according to the monolingual corpus of each language; concatenating the multilingual representation model and the multilingual generation model with a first translation model respectively to generate a target model to be trained; and generating a second translation model by training the target model according to a bilingual corpus among the plurality of languages, and performing translation processing on target information to be processed, according to the second translation model.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

A machine translation trace detection method

The application discloses a machine translation trace detection method, comprising the following steps: obtaining the similarity of each sentence pair in a bilingual corpus array with respect to a forward translation array, the similarity of each sentence pair in the bilingual corpus array with respect to a round-trip translation array, and the similarity of the round-trip translation array and the forward translation array; and averaging the three similarity values to obtain a score of the degree of machine translation. The machine translation trace is detected based on the similarity of the given to-be-detected translation, the forward translation array and the round-trip translation array, the similarity is calculated by using a classical algorithm, the threshold value is easy to determine through practice, the machine round-trip translation trace can be detected, and the detection result is more accurate, that is, the possibility that the to-be-detected text comes from a machine translation engine can be accurately represented.
Owner:BESTEASY (BEIJING) TRANSLATION CORP

A system and method for Tibetan-Chinese bilingual corpus collaborative annotation and versioned release

The application belongs to the technical field of natural language processing, and relates to a Tibetan-Chinese bilingual corpus collaborative labeling and versioned publishing system and method. Through the system architecture formed by the front-end interaction module, the business processing module, the data storage module and the basic management and control module, relying on the Tibetan-Chinese bilingual labeling, intelligent collaborative management and control, corpus version management and standardized publishing core units integrated by the business processing module, cooperating with the distributed correlation index storage mechanism of the data storage module, the fine-grained permission and operation log management and control function of the basic management and control module, the core technical problems of the Tibetan language characteristics adaptation deficiency, the low efficiency and frequent conflicts of multi-person collaborative labeling, the lack of version tracing and quality grading control of the corpus, the non-standard publishing process and the disconnection of the corpus and the downstream model in the prior art are solved. The accuracy of Tibetan-Chinese bilingual corpus labeling and storage is ensured through exclusive Tibetan adaptation processing, and the orderly promotion of multi-person labeling is realized through modular collaborative management and control.
Owner:SICHUAN TIANFU GAOCHI INFORMATION TECHNOLOGY CO LTD