Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "Parallel corpora" patented technology

Parallel Corpora. The term parallel corpora is typically used in linguistic circles to refer to texts that are translations of each other. And the term comparable corpora refers to texts in two languages that are similar in content, but are not translations.

Data model establishment method based on large language model and local knowledge base

The invention discloses a data model establishment method based on a large language model and a local knowledge base. The method comprises the following steps: S1, collecting parallel corpora through a distributed crawler technology; s2, converting data in a non-text format into data in a text format; s3, cleaning and preprocessing all the collected data; s4, denoising, segmenting and aligning the content to construct a parallel corpus; s5, storing the processed parallel corpora and terminology codes in a vector database; s6, constructing a retriever, and retrieving related information from the vector database; and S7, combining the retrieved related parallel corpora with the to-be-translated content in combination with the footstone model to realize high-quality translation generation. According to the method, the flexibility and practicability of the system are greatly improved, the method is suitable for various application scenes such as received text processing and manuscript translation, cross-language communication is more convenient and efficient, and the method has extremely high practical value and wide application prospects.
Owner:张一帆

Simultaneous interpretation method and system based on large model and electronic equipment

The invention discloses a simultaneous interpretation method and system based on a large model and electronic equipment, and the method comprises the steps: extracting bilingual parallel corpora related to terms from professional resources based on a standardized professional dictionary, obtaining qualified corpora through data enhancement processing and manual screening, and constructing a multi-level corpus according to the levels of words, sentences and paragraphs; the method comprises the following steps: receiving an input audio stream in real time, extracting acoustic features through preprocessing, inputting a pre-established large-scale speech recognition model, and carrying out incremental decoding on the acoustic features in a sliding window mode; and calling a sentence boundary prediction network to judge a pause point, and outputting a text stream with a timestamp. Performing fine tuning on the large-scale speech recognition model by using a multi-level corpus, translating a text stream based on the fine-tuned large-scale speech recognition model, and constraining term translation according to a standardized professional dictionary; and synchronously displaying the audio output in the translation result and the subtitles. According to the scheme, the terminology recognition and translation accuracy is improved, and simultaneous interpretation delay is reduced.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Machine translation professional data enhancement method, system and equipment and storage medium

The invention relates to the technical field of computers, and discloses a machine translation professional data enhancement method, system and device and a storage medium, and the method comprises the steps: extracting source language terms and corresponding target language terms from a professional dictionary, forming a structured term pair set and constructing a generation prompt, generating a source language text based on a generation model, and storing the source language text into a database; inputting a translation prompt containing a term pair mapping relation in the structured term pair set into a translation model, generating a corresponding target language text, taking the generated source language text and the target language text as sentence pairs, performing evaluation based on a judgment model, obtaining a comprehensive score, comparing the comprehensive score with a preset threshold value, and if the comprehensive score is not lower than the preset threshold value, outputting the comprehensive score. If yes, the sentence pairs are stored in the parallel corpus, and if not, control parameters in the generation prompt or the translation prompt are adjusted according to the evaluation result, and the step of generating the source language text and the target language text is executed again. According to the method, high-quality bilingual corpora can be generated, and the method has self-checking and optimizing capabilities.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Java-to-Cangjie code translation method based on large model and compiling feedback

The invention discloses a Java-to-Cangjie code translation method based on a large model and compilation feedback, and particularly belongs to the technical field of software engineering program language processing, and the method comprises the following steps: step 1, performing structured semantic pre-training by constructing a grammar knowledge base of a target language, and injecting grammar prior knowledge of the target language; step 2, performing semantic enhanced supervision fine tuning training by constructing a high-quality data set containing semantic information, and enhancing semantic alignment and cross-language migration ability of the model; 3, introducing an AST structure perception embedded prompt mechanism in a parallel corpus supervision fine tuning training stage, and guiding the model to perform structure perception translation; and 4, establishing a compiler feedback repair loop, and iteratively correcting output based on error information to form a self-optimized closed-loop system. According to the method, an efficient training path is constructed, dependence on large-scale parallel corpora is effectively reduced, and an extensible and high-reliability technical path is provided for cross-language code translation of low-resource programming languages.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method and system for mixed language text understanding for generative artificial intelligence (GENAI) models

This disclosure relates to method and system for mixed language text understanding for Generative Artificial Intelligence (GenAI) models. The method may include receiving a raw parallel corpus of two languages. The method may further include generating a cross-domain codemix parallel corpus and a first set of linguistic features from the raw parallel corpus using statistical and linguistic techniques. The method may further include determining a complexity of each of the plurality of samples of the cross-domain codemix parallel corpus based on a set of complexity parameters. The method may further include sequentially fine-tuning a pre-trained multilingual translation model using each of the plurality of samples in the curriculum learning dataset to obtain a generic pre-trained codemix understanding model.
Owner:WIPRO LTD

Large language model translation system based on encoder and decoder architecture

The invention discloses a large language model translation system based on an encoder and decoder architecture, which comprises the following steps: a data processing stage: collecting massive multilingual bilingual corpora for preprocessing, and constructing high-quality fine-tuning parallel corpora; constructing an encoder-decoder structure by using the pre-trained large language model, and determining the number of layers reserved at a decoder end and the connection mode of an encoder and a decoder by adopting a deep encoding-shallow decoding mode; performing model training by using massive multilingual bilingual corpora and high-quality fine-tuning parallel corpora obtained in the data processing stage to obtain a machine translation model; in the decoding stage, the encoder of the machine translation model encodes the source statement, and then the decoder decodes the source statement to generate a target language sentence. According to the method, the strong context understanding and generating capability of the large language model is utilized, the defect of low reasoning speed is overcome, the translation quality and effect of the model are improved, the convergence speed of the model is increased, the robustness of the model is improved, and the benefits brought by the pre-training method are improved.
Owner:XIAONIU FANYI

Large model Mian machine translation method based on language knowledge retrieval

The invention relates to a large model Mian machine translation method based on language knowledge retrieval, and belongs to the field of natural language processing. Comprising the following steps: constructing tree root nodes of a tree-shaped retrieval structure for parallel corpora from Chinese to Muranju language; dividing the parallel corpora from Chinese to Myanforn into long sentence pairs and short sentence pairs according to sentence lengths; constructing a preliminary tree-shaped retrieval structure through the short sentence pair; inserting the long sentence pair into a tree-shaped retrieval structure; in the retrieval stage, firstly, a to-be-translated text is coded through an LASER model and then matched with related words and sentence pairs in a tree retrieval structure, candidate sentence pairs are selected through a BM25 algorithm, and the cosine similarity between the candidate sentence pairs and the to-be-translated text is further calculated; before the large language model is used for final translation, the text reordering model is used for reordering the context prompt sentence pairs obtained through retrieval, and the context prompt sentence pairs and the to-be-translated text are input into the large language model to obtain a final translation result. The accuracy and smoothness of the translation result can be improved.
Owner:KUNMING UNIV OF SCI & TECH

Weak semantic low-resource character machine translation method based on semantic enhancement

Taking a translation task from Naxi Dongba to Chinese as an example, the invention provides a weak semantic low-resource character machine translation method based on semantic enhancement, which comprises the following steps: S1, designing a Naxi Dongba encoding system, and establishing a Naxi Dongba electronic dictionary; s2, a sufficient number of Naxi Dongba text-Chinese parallel sentence pairs are collected and marked, and a Naxi Dongba text-Chinese parallel corpus is constructed; s3, dividing the data set into a fine adjustment data set and a test data set, and further dividing the fine adjustment data set into a training set and a verification set; s4, constructing a semantic enhancement model based on fine tuning and custom word list embedding; s5, providing an iterative reverse translation method combined with word replacement, and constructing an extended data set; s6, constructing a weak semantic low-resource text machine translation model based on semantic enhancement, adopting an increment updating mechanism, taking the high-quality pseudo-parallel corpus generated in the step S5 as increment, inputting the increment into the semantic enhancement model in the step S3, and adjusting and optimizing the weight of the model through parameters; and S7, inputting the Naxi Dongba coded sentences to be translated into the updated model for translation, and outputting a result. According to the method, translation research from the Naxi Dongba text to Chinese is carried out based on traditional expert experience, automatic translation of the Naxi Dongba text can be achieved, meanwhile, the method has the capacity of continuous learning and adapting to new data, the machine translation effect of weak-semantic low-resource characters is improved, and technical support is provided for research in related fields.
Owner:SOUTHWEST UNIV

Style migration method for enhancing text anonymity

The invention provides a style migration method for enhancing text anonymity, which can be applied to the technical field of natural language processing and privacy protection. The method comprises the following steps: generating a prompt instruction set by utilizing a preset large language model; iteratively matching and translating a plurality of pseudo-parallel corpora with the same meaning but different attributes to obtain a query text library and a reference text library; combining the query text library and the prompt instruction set, inputting the combined query text library and prompt instruction set into a pre-trained text style rewriting model for processing to obtain a generated text, and screening the prompt instruction set through the semantic similarity between the generated text and the reference text library to obtain a screened prompt instruction set; the screened prompt instruction set and the to-be-processed text are combined and then input into a pre-trained text style rewriting model, and a plurality of output texts are generated; and carrying out anonymity effectiveness evaluation and screening on the plurality of output texts by utilizing a preset style migration evaluation index to obtain a text after anonymity enhancement and style migration of the to-be-processed text.
Owner:UNIV OF SCI & TECH OF CHINA

Novel translation model reasoning method and system based on rwkv

PendingCN122287660AImplement reasoning methodsscale upComputation complexityTheoretical computer science
The RWKV-based novel translation model inference method and system includes the following steps: 1) Collecting novel translations and extracting parallel corpora, and using dynamic MicroBatch concatenation technology for sequence compression; 2) Introducing a lightweight group query attention mechanism on the basis of the RWKV architecture, directly obtaining KV information from the Embedding layer to build the model; 3) Employing a sublinear complexity hybrid parallel training mode, combined with a global scalar scaling FP16 mixed precision strategy for training; 4) Applying a hierarchical distributed heterogeneous architecture, offloading the optimizer to low-performance devices and performing gradient compression transmission; 5) Outputting the translation using a joint decoder and dynamic batch inference technology. This invention, through the above method and system, effectively reduces the computational complexity and memory usage in the long text translation process, improves model training efficiency and inference throughput, and significantly improves the translation efficiency and contextual coherence of ultra-long texts.
Owner:LIAONING UNIVERSITY

Non-parallel corpus-oriented cross-language document theme alignment method

The invention belongs to the technical field of natural language processing, and discloses a non-parallel corpus-oriented cross-language document theme alignment method, which comprises the following steps of: 1, performing data preprocessing and data enhancement on a non-parallel corpus cross-language data set to obtain document pairs with similar semantics; the method comprises the following steps of 1, establishing a topic inference network, 2, encoding a document to obtain a document vector representation, 3, establishing the topic inference network, and inputting the document vector representation into the network to obtain document-topic distribution; 4, training the network by taking Dirichlet prior loss, intra-language comparison loss and cross-language comparison loss as a joint optimization target; according to the method, by constraining the theme consistency of the same-language documents and the theme distribution closeness of cross-language semantic similar documents at the same time, different-language documents are aligned in a shared theme space, so that cross-language document theme alignment under the non-parallel corpus condition is realized. According to the method and the device, cross-language topic alignment and consistent topic representation can be realized under the condition that parallel corpora or dictionary resources are not needed.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method, device and equipment for cross-language migration of language model and storage medium

The invention provides a cross-language language model migration method and device, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: acquiring a pseudo parallel corpus; adding a first feedforward neural network in each layer of the first model to obtain a second model; and training the second model based on the pseudo parallel corpus. According to the scheme, the pseudo parallel corpus is constructed by using the vocabularies of the first language and the vocabularies of the second language, so that the data acquisition and labeling cost is remarkably reduced, and the problem of high labeling cost of the parallel corpus is solved. Moreover, the pseudo parallel corpus comprises vocabularies of the first language and vocabularies of the second language, the first feedforward neural network added in the first model is used for processing the second language, and the original second feedforward neural network in the first model is used for processing the first language. Therefore, the second model obtained by training has better capability under the second language, namely, a better migration effect is realized.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

A method for transferring user offensive comment style based on unsupervised learning

This paper discloses a method for transferring the style of user offensive comments based on unsupervised learning. The method first encodes the input comment text using a bidirectional encoding-attention mechanism. The bidirectional encoding is used to capture the contextual information of the text sequence, and the attention mechanism is used to preserve the core information of the text. Secondly, the method uses a generator-discriminator adversarial training to address the lack of specific keywords in non-parallel corpora. Finally, the method constructs a reconstruction loss algorithm through recurrent reinforcement learning to ensure the accuracy of the style labels of the converted offensive comments, as well as the integrity and readability of the text content. This method can effectively address the problems of semantic loss, lack of parallel corpora, and low content retention in the style transfer of offensive comments.
Owner:SOUTHEAST UNIV

A method and system for translating graph data query language for knowledge base question answering

The present invention discloses a method and system for translating graph data query languages for knowledge base question answering, which realizes the translation between graph data query languages (SPARQL, Cypher, and nGQL). Taking the translation from SPARQL to Cypher as an example, the corresponding method includes: 1) preprocessing SPARQL and Cypher statements; 2) training a structure-to-sequence translation model from the source SPARQL query skeleton graph structure to the target Cypher query skeleton sequence on parallel corpora; 3) inputting the query skeleton graph structure obtained by preprocessing the input source SPARQL statement into the translation model to generate the target Cypher query skeleton sequence, and restoring the skeleton sequence to the target Cypher statement according to the mapping dictionary. The corresponding system includes a preprocessing module, a source query language encoding module, and a target query language generation module, which helps with the migration of the knowledge base question answering system. The present invention uses a graph neural network to encode the structural semantic information of the source query statement, realizes end-to-end translation between graph data query languages, helps save labor costs, and contributes to the construction and migration of the knowledge base question answering system.
Owner:EAST CHINA NORMAL UNIV

Machine translation strengthening method based on bilingual dictionary injection

The invention discloses a machine translation strengthening method based on bilingual dictionary injection, and belongs to the technical field of machine translation strengthening. The problem that in the prior art, a traditional machine translation strengthening method is poor in model performance for translation in the special field is solved. The method comprises the following steps: performing bilingual alignment on large-scale unsupervised monolingual corpora to generate a bilingual dictionary; parallel corpora are introduced into the bilingual dictionary, the hit rate of each word pair in the bilingual dictionary in the parallel corpora is counted, a Memory Bank is established and the hit rate is recorded, the importance of the word pairs is sorted according to the hit rate, and the sorted bilingual dictionary is obtained; and performing data enhancement on the sorted source end data in the bilingual dictionary through Memory Bank, and inputting the data into a deep adversarial network model for model training to obtain a trained deep adversarial network model. According to the method, the parallel corpora are effectively subjected to data enhancement, the generation quality of a machine translation system is improved, and the method can be applied to machine translation modeling.
Owner:HARBIN INST OF TECH

Course guiding type multi-task learning large language model fine tuning method

The invention provides a course-guided multi-task learning large language model fine tuning method, which comprises the following steps of: inputting monolingual corpus into a large language model, and performing restoration clean text training by minimizing a prediction error to obtain a first model; inputting the parallel corpora into the first model, and performing specified term generation training through cross entropy loss to obtain a second model; and inputting the bilingual parallel corpora into the second model, and performing cross-language output training through a regularization item loss function to obtain a large language fine tuning model. According to the method, semantic modeling is enhanced through a de-noising auto-encoder task, term alignment knowledge is integrated through a translation task limited by a vocabulary, cross-language format alignment is enhanced through a machine translation task, and a multi-stage shift arrangement training strategy is introduced, so that task interference and disastrous forgetting in multi-task learning are effectively reduced; and the security BLEU and the term coverage rate are obviously improved.
Owner:XINJIANG UNIVERSITY

A model training, machine translation method, device, equipment and storage medium

The embodiment of the present invention discloses a model training, machine translation method, device, equipment and storage medium. The model training method may include: obtaining original parallel corpus including original source end data and original target end data; using the original source end data as exchange target end data, and using the original target end data as exchange source end data to obtain exchange parallel corpus, and training the original translation model based on multiple groups of original parallel corpora and multiple groups of exchange parallel corpora to obtain an intermediate translation model; training the intermediate translation model based on multiple groups of original parallel corpora to obtain a machine translation model. The technical solution of the embodiment of the present invention can improve the training effect of the machine translation model.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Automatic construction method of myanmar-english parallel corpus in large model based on multi-dimensional evaluation

PendingCN122287656AData setModel translation
This invention relates to a method for automatically constructing a large-scale Chinese-Myanmar parallel corpus based on multidimensional evaluation, belonging to the field of natural language processing technology. To address the scarcity of Chinese-Myanmar parallel corpus resources, this invention proposes a method for automatically constructing a large-scale Chinese-Myanmar parallel corpus based on multidimensional evaluation. First, Burmese corpora are manually collected to construct a basic dataset. Second, prompt words are designed to generate a candidate set of pseudo-parallel Chinese-Burmese corpora using a large-scale model. Then, the candidate set is filtered using semantic similarity constraints. Finally, the large-scale model is used to score the filtered candidate set from multiple dimensions, and the corpora are ranked and hierarchically filtered based on the scores to obtain high-quality pseudo-parallel corpora. Experimental results show that the corpus constructed by this invention can effectively improve the model's translation performance and significantly enhance its translation quality in low-resource Chinese-Myanmar language environments.
Owner:KUNMING UNIV OF SCI & TECH

Method, device and equipment for training translation model, medium and product

The embodiment of the invention relates to a method, device and equipment for training a translation model, a medium and a product. The method includes generating, by a model, a translated text sequence based on the initial text sequence. The method further includes translating the text sequence through the model prediction. The method further includes determining a first loss function based on a prediction result of the translated text sequence. The method further includes predicting, by the model, an initial text sequence based on a prediction result of the translated text sequence. The method further includes determining a second loss function based on the prediction of the initial text sequence. The method further comprises the step of carrying out weighted summation on the first loss function and the second loss function to obtain a total loss function. And finally, updating parameters of the model based on the total loss function. Through the method, the multi-language ability of the model can be improved by effectively utilizing a sample only containing a single language, and the dependence on large-scale and high-quality parallel corpora is reduced.
Owner:BEIJING FEISHU TECH CO LTD

Mongolian-Chinese neural machine translation method based on multi-dimensional prompt optimization of large language model

The invention discloses a Mongolian-Chinese neural machine translation method based on multi-dimensional prompt optimization of a large language model. A model outputs a better Mongolian-Chinese translation result through rich semantic and syntactic information; screening and extracting a plurality of keywords for each Mongolian sentence from the existing Mongolian-Chinese parallel corpus; respectively translating all the Mongolian keywords into various corresponding high-resource language word meanings through a multi-language semantic network, and constructing a word meaning chain prompt; mapping the Mongolian source sentence and the whole Chinese target corpus to the same semantic space by using a semantic embedding model; k Chinese sentence examples with the highest similarity score are selected, semantic features of the Chinese sentence examples are analyzed, a semantic association structure is constructed, and syntactic mode prompts are formed; designing clear instructions to describe task translation requirements, namely defining instruction type prompts, and directly guiding the model to perform translation operation; splicing the three prompts to construct a semantic association structure for guiding a large language model to obtain a final translation result; by utilizing the method, the overall effect of Mongolian-Chinese translation can be effectively improved.
Owner:INNER MONGOLIA UNIV OF TECH

A Mongolian-Chinese non-autoregressive machine translation method based on multi-task learning

A Mongolian-Chinese non-autoregressive machine translation method based on multi-task learning includes preprocessing Mongolian-Chinese parallel corpora; dividing the preprocessed Mongolian-Chinese parallel corpus dataset into a training set, a validation set, and a test set; building an autoregressive translation model with a shared encoder, and forming a multi-task learning framework consisting of the shared encoder, the autoregressive translation model decoder, and the non-autoregressive translation model decoder; within the multi-task learning framework, training the non-autoregressive translation model based on the training set, thereby transferring knowledge from the autoregressive translation model to the non-autoregressive translation model. The obtained non-autoregressive translation model can then be used to perform Mongolian-Chinese translation. This method improves the quality of Mongolian-Chinese translation while ensuring an increased translation rate.
Owner:INNER MONGOLIA UNIV OF TECH

A machine translation method capable of learning future information

The present invention discloses a machine translation method capable of learning future information. The steps are as follows: adding a future information network module at the decoder end to construct a machine translation model capable of learning future information; processing training data and using a word embedding model to obtain word embedding representations; initializing parameters and optimizing model training; in the encoder, calculating the word embeddings and obtaining more information in the word embedding vectors. After operating n times, the model has learned the feature information of the sentence; the machine translation model capable of learning future information learns the correlation information between the source language and the target language, and the future information network sends the learned information back to the decoder to assist the decoder in decoding; using the trained machine translation model capable of learning future information to perform machine translation to implement the machine translation method capable of learning future information. The method of the present invention improves the deficiencies of the existing neural machine translation paradigm and enhances the information capture ability of the neural machine translation model for parallel corpora and the translation performance of the model.
Owner:XIAONIU FANYI

Model generation method, word sense disambiguation method, device, medium, and equipment

ActiveCN115017986BSemantic analysisWord-sense disambiguationParaphrase
The present disclosure relates to a model generation method, a word sense disambiguation method, an apparatus, a medium, and a device. The model generation method includes: obtaining multiple groups of parallel corpora, where each group of parallel corpora includes a first text and a second text that are translations of each other, the first text belongs to a first language, and the second text belongs to a second language; determining multiple first samples according to the multiple groups of parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase in the preset paraphrase set that matches a translation word of the first original word, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language; and generating a first classification model according to the multiple first samples. The present disclosure can reduce the data dependence for generating the first classification model.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Method for constructing myanmar-chinese parallel corpus based on multi-step thinking large model

This invention relates to a method for constructing a large-scale Burmese-Chinese parallel corpus based on multi-step thinking. The invention includes: translating existing Chinese-English parallel corpora using currently available translation models to obtain original English-Burmese and Chinese-Burmese parallel sentence pairs; calculating double-confidence intervals for semantic similarity and perplexity between aligned Burmese-Chinese sentence pairs based on publicly available high-quality Burmese-Chinese parallel corpora; using the selected double-confidence intervals to perform preliminary screening of Burmese-Chinese parallel sentence pairs, forming pre-processed pseudo-parallel sentence pairs; designing a multi-step thinking chain to guide the large-scale model to progressively optimize the pre-processed pseudo-parallel sentence pairs, thereby generating high-quality Burmese-Chinese parallel corpora for training the translation model, thus effectively improving the performance of Burmese-Chinese machine translation. This invention significantly enhances the ability of large language models to construct corpora in Burmese, a low-resource language, and provides an interpretable and transferable technical paradigm for corpus construction in other low-resource languages.
Owner:KUNMING UNIV OF SCI & TECH +4

Machine translation model training method and apparatus, device, medium, and product

ActiveCN115563992BData setEngineering
The application relates to a machine translation model training method and device, equipment, medium and product, the method comprising: obtaining a data set comprising a plurality of parallel corpora; constructing a plurality of variant models of a machine translation model, determining a single variant model as a student model and the rest as teacher models; training each teacher model to a convergent state using parallel corpora in the data set; constructing a knowledge distillation network, inputting parallel corpora in the data set into the knowledge distillation network to implement training, and jointly supervising the student model training to a convergent state through a plurality of teacher models. Based on the machine translation model, a plurality of variant models with different advantage reasoning capabilities are prepared in advance as teacher models, and then the advantage reasoning capabilities of the plurality of teacher models are transmitted to the same student model through the knowledge distillation mode, so that the student model has more comprehensive translation reasoning capability and is suitable for translating commodity information of various commodity categories.
Owner:GUANGZHOU HUADUO NETWORK TECH

A method and system for AI-powered intelligent management of sensitive dialogue information on overseas labor dispatch platforms

This invention discloses an AI-powered intelligent management method and system for sensitive information in dialogues on overseas labor dispatch platforms. The method includes: acquiring dialogue flow information from the overseas labor dispatch platform; establishing an AI recognition model based on the dialogue flow information; performing reverse reasoning based on the AI ​​recognition model to generate sensitive risk tracing paths for each statement in the dialogue flow; performing dynamic pruning and cross-cultural context relabeling on the dialogue flow information based on the sensitive risk tracing paths to obtain a relabeled sensitive risk score and a sensitive management reliability index; and generating and outputting a sensitive information management report using the relabeled sensitive risk score and the sensitive management reliability index. This invention utilizes a cross-cultural context transfer mechanism based on semantic preservation constraints and cultural attribute transfer constraints. It trains cultural adaptation mapping parameters using cross-linguistic parallel corpora in the labor dispatch field, transferring source cultural semantic representations to the target cultural semantic space, and combining this with a set of sensitive semantic anchor points corresponding to the destination country for relabeling and scoring.
Owner:HENGZHI SMART DIGITAL TECHNOLOGY (SHANGHAI) CO LTD

Speech recognition algorithm for small languages

InactiveCN120526756ASpeech recognitionRecognition algorithmParallel corpora
The invention relates to the technical field of speech recognition, and discloses a speech recognition algorithm for a small language, which comprises the following steps: S1, constructing a multi-language keyword parallel corpus; s2, analyzing the multi-language keyword parallel corpus to obtain keyword pronunciation similarity indexes of other languages and a target small language, and screening out a plurality of first reference languages; s3, constructing a comprehensive parallel corpus; s4, corpus information of the comprehensive parallel corpus is divided into a plurality of corpus analysis units; s5, screening out a target corpus analysis unit; s6, screening out a target migration language of each target corpus analysis unit according to corpus information data of the comprehensive parallel corpus; s7, performing cross-language migration modeling according to the target migration language of each target corpus analysis unit; the target of accurately positioning the optimal migration source from massive languages is achieved; the problem of model training difficulty caused by data scarcity of the target minority language is effectively solved.
Owner:SHENZHEN WEIOU TECH CO LTD

Large model adaptive teaching method and system based on causal guided decoupling and multi-modal perceptual learning

The invention discloses a large-model adaptive teaching method and system based on causal guided decoupling and multi-modal perceptual learning. According to the method, two control vectors are decoupled through comparative learning, causal difference extraction and a robust aggregation algorithm according to a non-parallel plain text corpus, a cognitive coefficient is determined according to the knowledge level of a learner, a style coefficient is determined through the emotion of the input voice of the learner, and the learning effect of the learner is improved. The two coefficients perform linear superposition on the two vectors in a shared semantic space of the multi-modal large language model and are injected into the model through an activation vector guiding technology, and the guided activation vectors are synchronously fed back to a plurality of downstream decoders to generate multi-modal output. According to the method, two orthogonal control vectors capable of being continuously adjusted are injected, so that'content-style 'two-dimensional refined control of multi-modal teaching output of AI teachers is realized, the pain points of cognitive mismatch and emotion segmentation in AI teaching are solved, and the dependence on parallel corpora is thoroughly gotten rid of.
Owner:ZHEJIANG UNIV

Machine translation polysemous word translation evaluation method based on semantic item trigger word replacement

PendingCN122088522AReveal errors effectivelyImprove targetingNatural language translationSemantic analysisSentence pairWord sense
Aiming at the defects of an existing test method in the field of word sense disambiguation test of machine translation in the aspects of triggering effective word sense conversion, maintaining original sense and maintaining semantic consistency, the invention provides a machine translation polysemous word translation evaluation method based on semantic item trigger word replacement. The method comprises the following steps: firstly, constructing a polysemy semantic item library, and screening source language sentences containing polysemy from a parallel corpus according to the semantic item library; performing natural language processing on the sentence, and identifying a trigger word of a polysemy special definition item in the sentence; replacing the trigger word with a replacement word to generate a variant sentence; performing machine translation and alignment on the original sentences and the variant sentences to obtain expressions of polysemy words in translations; and calculating the similarity of the polysemy translation expression, and judging whether a polysemy disambiguation error exists or not. Through a trigger word replacement strategy, a semantic controlled contrast sentence pair can be constructed, potential errors of a system in polysemous word translation are effectively revealed, and the pertinence and effectiveness of evaluation are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Bilingual Dictionary Inference Method, Apparatus and Storage Medium

The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, and a storage medium for bilingual dictionary inference. The method includes: extracting a target dictionary from parallel corpora; training a target bilingual dictionary inference model based on the extracted target dictionary and a pre-configured initial dictionary, where the target bilingual dictionary inference model is a neural network model capable of translating source-side words into target-side words; wherein both the target dictionary and the initial dictionary include a plurality of aligned word pairs, and the aligned word pairs include source-side words and target-side words. By introducing parallel corpora on the basis of the initial dictionary in the embodiments of the present disclosure and using the target dictionary extracted from the parallel corpora to enrich the training information of the target bilingual dictionary inference model, the subsequent bilingual dictionary inference effect is improved.
Owner:NANJING UNIV