A language translation method, electronic device and system based on big data

Through the method of semantic analysis and feedback optimization of the initial translation, the problem of word order and context in the translation results of the NMT algorithm is solved, which significantly improves the accuracy and fluency of the translation.

CN119783697BActive Publication Date: 2025-05-27YIBAIFEN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510288308.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-05-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing neural machine translation (NMT) algorithms still have word order and context problems in the translation results, resulting in insufficient translation accuracy and fluency.

Method used

By obtaining corpus data, initially train the NMT model, use crawler technology to build a network corpus, semantic analysis of each word unit in the initial translation, obtain sentence structure parameters and semantic eigenvalues, and combine these features to feedback and iterative optimization of the NMT model to improve the translation quality.

Benefits of technology

Effectively identify and optimize translation fuzzy points, improve translation accuracy and contextual adaptability, and improve the semantic expression and fluency of translation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783697B_ABST
    Figure CN119783697B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of natural language processing, and particularly relates to a language translation method, an electronic device and a system based on big data, including: constructing an initial NMT model and translating a sentence segment to be translated to obtain an initial translation; constructing a network corpus, combining the network corpus to perform semantic analysis on the relationship between word units in the sentence segment to be translated and the initial translation in terms of sentence structure, obtaining sentence structure parameters of the word units, and performing replacement on the word units in the sentence by using synonyms and near-synonyms of the word units in the network corpus, analyzing the semantic change characteristics before and after replacement, obtaining semantic feature values of the word units, combining the sentence structure parameters of the word units with the semantic feature values to obtain the translation degree of the word units, thereby providing feedback to the NMT algorithm, and iteratively translating the sentence segment to be translated to obtain a final translation. The present invention effectively optimizes the semantic expression of the translation and improves the accuracy and context adaptability of the translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a language translation method, an electronic device and a system based on big data. Background Art

[0002] The increasing globalization and cross-border exchanges have made language translation an indispensable part of people's daily lives and business activities. Moreover, the amount of data globally has shown an explosive growth, and the emergence of big data technology has provided new opportunities for the improvement of language translation systems. By leveraging big data, translation systems can more accurately capture the relationships between languages, analyze contexts, identify lexical collocations and common expressions, thereby improving the accuracy and naturalness of translations.

[0003] Currently, language translation technologies based on the NMT (Neural Machine Translation) algorithm have gradually become mainstream. Compared with traditional statistical translation methods, neural network translation can not only better capture the grammatical information at the sentence level, but also perform more natural and fluent translations through context information. However, although significant progress has been made in the accuracy and fluency of neural network-based translation methods, there are still problems with word order and context in the translation results, making the NMT algorithm still have certain limitations. Summary of the Invention

[0004] The present invention provides a language translation method, an electronic device and a system based on big data to solve the existing problems.

[0005] The following technical solutions are adopted for a language translation method, an electronic device and a system based on big data of the present invention:

[0006] An embodiment of the present invention provides a language translation method based on big data, the method comprising the following steps:

[0007] Obtain corpus data and preliminarily train an NMT model to obtain an initial NMT model;

[0008] Obtain a sentence segment to be translated and translate the sentence segment to be translated through the initial NMT model to obtain an initial translation;

[0009] Use web crawler technology to crawl web corpus data, thereby constructing a web corpus;

[0010] Semantically analyze the relationship of each word unit in the initial translation in terms of sentence structure in combination with the web corpus, obtain the sentence structure parameters of each word unit in the initial translation, and replace each word unit in the sentence with its synonyms and near-synonyms in the web corpus, analyze the semantic change characteristics before and after replacement, obtain the semantic feature values of each word unit in the initial translation, and combine the sentence structure parameters of any word unit in the initial translation with the semantic feature values to obtain the translation degree of the word unit;

[0011] Feedback to the NMT model in combination with the translation degree of the word unit, construct a feedback function and iteratively translate the sentence segment to be translated to obtain the final translation.

[0012] Furthermore, the specific method for semantically analyzing the relationship of each word unit in the initial translation in terms of sentence structure in combination with the web corpus and obtaining the sentence structure parameters of each word unit in the initial translation is as follows:

[0013] Use Jieba to segment the initial translation, and each phrase in the segmentation result is used as a word unit; for any word unit, obtain several sentence segments in the web corpus that contain the any word unit, and record them as the corpus sentence segments of the any word unit;

[0014] Conduct text structure analysis on the initial translation of any word unit and the corpus sentence segments, and obtain the cosine structure sequences of the initial translation and each corpus sentence segment;

[0015] According to the similarity of the cosine structure sequences corresponding between the initial translation and all the corpus sentence segments corresponding to each word unit in the initial translation, and in combination with the structural characteristics of the corpus sentence segments corresponding to each word unit, obtain the sentence structure parameters of any word unit.

[0016] Furthermore, the specific method for conducting text structure analysis on the initial translation of any word unit and the corpus sentence segments and obtaining the cosine structure sequences of the initial translation and each corpus sentence segment is as follows:

[0017] First, segment the corpus sentence segments to obtain several phrases in any corpus sentence segment, which are called corpus word units, and use the word2vec algorithm to perform word vector conversion on the word units in the initial translation and the corpus word units in all corpus sentence segments respectively to obtain the word vectors of the word units and the corpus word units;

[0018] Then, obtain the cosine value of the word vectors corresponding to any adjacent word units in the initial translation, and obtain the sequence formed by the cosine values between the word vectors of all adjacent word units in the initial translation, which is recorded as the cosine structure sequence of the initial translation; and so on, obtain the cosine structure sequences of any corpus sentence segments.

[0019] Furthermore, the specific method for obtaining the sentence structure parameters of any word unit is:

[0020] Obtain the structural factor between the initial translation and the corpus sentence segments;

[0021] The specific calculation method of the sentence structure parameters of the word unit is as follows:

[0022]

[0023] where is the sentence structure parameter of the th word unit in the initial translation; is the DTW distance between the cosine structure sequences corresponding to the th word unit in the initial translation and the th corpus sentence segment respectively; is the number of corpus sentence segments of the th word unit in the initial translation; is the structural factor between the initial translation and the th word unit in the initial translation and the th corpus sentence segment; is the number of all word segments in the th word unit in the initial translation and the th corpus sentence segment that are the same as the word units in the initial translation; is the number of all word segments in the th word unit in the initial translation and the th corpus sentence segment; is the standard deviation function.

[0024] Furthermore, the specific method for obtaining the structural factor between the initial translation and the corpus sentence segments is as follows:

[0025] Obtain the same word segments in the initial translation and the corpus sentence segments; take a word segment as a node, connect them in the order of the word segments in the corresponding text, and take the backward order as the corresponding connection direction;

[0026] Take the interval between adjacent word segments in the corresponding text as the edge weight of the corresponding connection edge, and respectively form the chain-like word segment graph structures corresponding to the initial translation and the corpus sentence segments;

[0027] Obtain the graph similarity between the word segment graph structures corresponding to the initial translation and the corpus sentence segments through the graph edit distance, and use it as the structural factor between the initial translation and the corpus sentence segments.

[0028] Furthermore, the specific method for obtaining the semantic feature value of the word unit is as follows:

[0029] Obtain the synonyms and near-synonyms of any word unit, collectively referred to as the approximate words of the corresponding word unit, and obtain the corpus segments containing the approximate words of the word unit in the web corpus, which are called the approximate corpus segments of the word unit;

[0030] Combine the initial translation and the approximate corpus segments of each word unit to perform semantic feature analysis on the word unit, and obtain the semantic feature value of the word unit.

[0031] Furthermore, the specific method for obtaining the semantic feature value of the word unit is as follows:

[0032] For any word unit, obtain all the approximate corpus segments corresponding to the word unit, obtain the approximate words corresponding to the word unit in the approximate corpus segments and the word segments adjacent to the left and right of the approximate word as the adjacent word segments of the approximate word, and obtain the phrase formed by the cosine value of the approximate word and each corresponding adjacent word segment with respect to the word vector, which is denoted as the semantic feature group of the approximate word, where is a preset quantity parameter;

[0033] For any word unit, use all the approximate words of the word unit to replace the word unit to obtain a new initial translation, and obtain the semantic feature groups of all the approximate words of the word unit in the new initial translation; and so on, obtain the semantic feature group corresponding to the word unit;

[0034] Based on the cosine value between the semantic feature groups obtained from the word unit and the corresponding approximate words in different texts, calculate the semantic feature value of the word unit.

[0035] Furthermore, the specific calculation method for the semantic feature value of the word unit is as follows:

[0036]

[0037] where, represents the semantic feature value of the th word unit in the initial translation; is the semantic feature group of the th word unit in the initial translation; is the semantic feature group of the approximate word after the th word unit in the initial translation is replaced by the th corresponding approximate word; is the semantic feature group of the th approximate word corresponding to the th word unit in the th approximate corpus segment; is the quantity of the approximate corpus segments containing the th approximate word; is the number of approximate words for the first word unit in the initial translation; is the standard deviation function; is the cosine function; is the exponential function with the natural constant as the base.

[0038] An embodiment of the present invention provides a language translation electronic device based on big data, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the described language translation method based on big data.

[0039] An embodiment of the present invention provides a language translation system based on big data, and the system includes the following modules:

[0040] An initial model module, configured to obtain corpus data and preliminarily train an NMT model to obtain an initial NMT model;

[0041] A preliminary translation module, configured to obtain a sentence segment to be translated and translate the sentence segment to be translated through the initial NMT model to obtain an initial translation;

[0042] A corpus collection module, configured to crawl network corpus data by using web crawler technology, thereby constructing a network corpus;

[0043] A translation analysis module, configured to perform semantic analysis on the relationship of each word unit in the initial translation in terms of sentence structure in combination with the network corpus, obtain the sentence structure parameters of each word unit in the initial translation, and perform replacement on the synonyms and near-synonyms of the word unit in the network corpus in the corresponding sentence, analyze the semantic change characteristics before and after replacement, obtain the semantic feature values of each word unit in the initial translation, and combine the sentence structure parameters and semantic feature values of any word unit in the initial translation to obtain the translation degree of the word unit;

[0044] A feedback translation module, configured to feedback to the NMT model in combination with the translation degree of the word unit, construct a feedback function and iteratively translate the sentence segment to be translated to obtain a final translation.

[0045] The beneficial effect of the technical solution of the present invention is that after constructing the initial NMT model and the network corpus, through semantic contrast analysis of the word units in the initial translation with the synonyms and near-synonyms in the network corpus in terms of syntactic structure and synonym replacement, a translation degree index is further quantified to effectively identify the translation fuzzy points in the initial translation. By integrating the translation degree parameter into the NMT model weight adjustment through an iterative optimization mechanism, the progressive improvement of translation quality is realized. This method can significantly improve the translation quality by combining big data technology, thereby optimizing the semantic expression of the translation and enhancing the accuracy and context adaptability of the translation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0047] Figure 1 It is a flowchart of the steps of a language translation method based on big data according to the present invention;

[0048] Figure 2 It is a schematic flowchart of obtaining a structure factor provided by an embodiment of the present invention;

[0049] Figure 3 It is a schematic diagram of the positional relationship between word units, approximate words and corresponding adjacent word segments provided by an embodiment of the present invention;

[0050] Figure 4 It is a schematic diagram of a semantic feature space provided by an embodiment of the present invention;

[0051] Figure 5 It is a structural block diagram of a language translation system based on big data according to the present invention. Detailed implementation manners

[0052] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and effects of a language translation method, an electronic device and a system based on big data according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0054] The following specifically describes the specific solutions of a language translation method, an electronic device and a system based on big data provided by the present invention with reference to the accompanying drawings.

[0055] Please refer to Figure 1 , which shows a flowchart of the steps of a language translation method based on big data provided by an embodiment of the present invention. The method includes the following steps:

[0056] Step S001: Obtain corpus data and preliminarily train an NMT model to obtain an initial NMT model.

[0057] It should be noted that in language translation, generally, the NMT model needs to be trained with a large amount of bilingual parallel corpus. These corpora include source language and target language paired sentences, such as English and Chinese parallel corpora.

[0058] Specifically, to implement a language translation method based on big data proposed in this embodiment, first, it is necessary to collect bilingual parallel corpus data and train the NMT model. The specific process is as follows:

[0059] Step S101, obtain corpus data and preprocess it.

[0060] In an embodiment of the present invention, the specific steps for obtaining and preprocessing corpus data include:

[0061] First, collect bilingual parallel corpus to obtain corpus data.

[0062] Then, preprocess the bilingual parallel corpus, including: data cleaning, word segmentation, normalization, and data splitting; among them, special characters, HTML tags, blank lines, etc. in the bilingual parallel corpus are removed through data cleaning; word segmentation processing is performed on the bilingual parallel corpus, Jieba is used to process the Chinese corpus, and Nltk is used to process the English corpus; the normalization process is to uniformly encode the vocabulary formed by the words after word segmentation of the bilingual parallel corpus based on the Subword Unit (subword unit) technology, and divide the data set formed by the corpus data into a training set, a validation set, and a test set.

[0063] It should be noted that in the embodiment of the present invention, the proportions of the data set divided into the training set, the validation set, and the test set are 80%, 10%, and 10% respectively. The specific division ratio can be adjusted according to the actual situation, and the embodiment of the present invention does not make specific limitations.

[0064] Step S102, construct an NMT model and train the NMT model in combination with the corpus data to obtain an initial NMT model.

[0065] First, construct an NMT neural network model based on the encoder-decoder architecture (Encoder-Decoder) and combined with the self-attention mechanism and the Transformer structure as the NMT model, and further implement the NMT model through a deep learning framework.

[0066] As an embodiment of the present invention, the deep learning framework is PyTorch to implement the NMT model.

[0067] Then, input the source language sentences in the training set into the NMT model, with the target language sentences as labels, and adjust the model parameters by optimizing the cross-entropy loss function during the training process.

[0068] So far, an initial NMT model that has been initially trained is obtained through the above method.

[0069] Step S002: Obtain the sentence segment to be translated, and translate the sentence segment to be translated through the initial NMT model to obtain an initial translation.

[0070] In an embodiment of the present invention, the specific steps for obtaining the initial translation are as follows:

[0071] First, obtain the sentence segment to be translated, input the sentence segment to be translated into the initial NMT model, and generate a context vector by the encoder in the initial NMT model.

[0072] Then, the decoder generates a sentence in the target language according to the context vector. During the generation process, the greedy algorithm (Greedy Search) is used to generate the optimal translation result.

[0073] It should be noted that since the greedy algorithm is an existing optimal solution seeking algorithm, the embodiments of the present invention will not be specifically described herein.

[0074] Finally, output the initial translation of the sentence segment to be translated.

[0075] So far, the initial translation of the sentence segment to be translated is obtained through the above method.

[0076] Step S003: Use web crawler technology to crawl web corpus data, thereby constructing a web corpus.

[0077] It should be noted that when the NMT model trained through an existing fixed corpus is performing language translation, it will be affected by the number of words in the training data, resulting in underfitting or overfitting problems when translating some words, causing the semantics of some words in the translation result to be inappropriate and the translation result to be unsmooth. Therefore, in the embodiments of the present invention, semantic analysis is performed on other corpus data and combined with the words in the initial translation, thereby improving the translation effect of the translation.

[0078] To achieve the purpose of subsequent semantic analysis, the embodiments of the present invention need to obtain web corpus data and construct a web corpus, including:

[0079] First, based on the Scrapy crawler library, use web crawler technology to obtain sentences, paragraphs, phrases, and words on the Internet platform.

[0080] The Internet platform includes news websites, forums, and blogs.

[0081] Then, clean the crawled web corpus data, remove irrelevant content such as advertisements and navigation bars on the web pages; remove duplicate content.

[0082] Finally, store all the processed phrases, sentences, paragraphs, etc. as text files and store them in the database as a corpus, denoted as the web corpus.

[0083] Thus, the web corpus is obtained through the above method.

[0084] Step S004: Perform semantic analysis on the relationship of each word unit in the initial translation in terms of sentence structure in combination with the web corpus, obtain the sentence structure parameters of each word unit in the initial translation, and replace the word units in the sentence with their synonyms and near-synonyms in the web corpus, analyze the semantic change characteristics before and after replacement, obtain the semantic characteristic values of each word unit in the initial translation, and combine the sentence structure parameters of any word unit in the initial translation with the semantic characteristic values to obtain the translation degree of the word unit.

[0085] It should be noted that during the process of translating the sentence segment to be translated by the NMT model, the NMT may have an incorrect understanding of the sentence segment to be translated, or the attributive in the translation result is too long, which does not conform to the Chinese expression habit, resulting in an unsatisfactory translation result. Therefore, in order to improve the translation effect of the NMT, it is selected to analyze the translation effect of the initial translation of the NMT by using the web corpus and obtain the translation degree of the word units in the initial translation.

[0086] As an embodiment, the specific method for obtaining the translation degree includes:

[0087] Step S401, obtain a number of word units of the initial translation, and calculate the sentence structure parameters of each word unit in the initial translation in combination with the sentence structure of the word unit in the web corpus.

[0088] It should be noted that when translating text, it is not only necessary to translate the vocabulary in one language into the corresponding vocabulary in another language, but also necessary to capture the structural and semantic properties of the language. That is, language is essentially composed of hierarchical structures and semantics containing implicit meanings. Then, when translating vocabulary and generating sentences, the core task is to capture these structural and semantic information. Therefore, in the embodiments of the present invention, it is selected to perform semantic analysis on the relationship of each word unit in the initial translation in terms of sentence structure in combination with the web corpus to obtain the sentence structure parameters of each word unit in the initial translation.

[0089] As an embodiment, the method for obtaining the sentence structure parameters of the word unit includes:

[0090] Step S401.1: Use Jieba word segmentation to segment the initial translation, and each phrase in the segmentation result is used as a word unit; for any word unit, obtain several sentence segments in the web corpus that contain the said any word unit, and record them as the corpus sentence segments of the said any word unit.

[0091] Step S401.2: Perform text structure analysis on the initial translation of any word unit and the corpus sentence segments, and obtain the cosine structure sequences of the initial translation and each corpus sentence segment.

[0092] As an embodiment, the specific process of obtaining the cosine structure sequence includes:

[0093] First, segment the corpus sentence segments to obtain several phrases in any corpus sentence segment, which are called corpus word units, and use the word2vec algorithm to perform word vector conversion on the word units in the initial translation and the corpus word units in all corpus sentence segments respectively, to obtain the word vectors of the word units and the corpus word units.

[0094] Then, obtain the cosine values of the corresponding word vectors of any adjacent word units in the initial translation, and obtain the sequence formed by the cosine values between the word vectors of all adjacent word units in the initial translation, which is recorded as the cosine structure sequence of the initial translation.

[0095] And so on, obtain the cosine structure sequences of any corpus sentence segments.

[0096] Finally, obtain the DTW distance between the cosine structure sequences corresponding to the initial translation and any corpus sentence segment respectively.

[0097] It should be noted that the word units in the initial translation and the corpus word vectors in the corpus sentence segments are essentially through word vector conversion to extract the semantic information in the initial translation or corpus sentence segments, that is, through word vector conversion, each word can be transformed into a dense vector representation, and these vectors can retain the semantic similarity between the words. For example, the word vectors of "apple" and "orange" are relatively close in the semantic space, while the word vectors of "apple" and "car" will be farther apart; in addition, by calculating the cosine similarity of all adjacent word units, it reflects the similarity distribution between words within the whole sentence, and can further help to identify the structural features of the translation, so as to facilitate the subsequent analysis of the compliance degree between the text structure of the initial translation and the common expression methods in terms of semantic structure.

[0098] Step S401.3: According to the similarity between the initial translation and the cosine structure sequences corresponding to all corpus sentence segments corresponding to each word unit in the initial translation, and in combination with the sentence structure features between the initial translation and the corpus sentence segments corresponding to each word unit in the initial translation, obtain the sentence structure parameters of any word unit.

[0099] As an embodiment, the specific calculation method of the sentence structure parameters of the word unit is as follows:

[0100]

[0101] Wherein, is the sentence structure parameter of the th word unit in the initial translation; is the DTW distance between the cosine structure sequences respectively corresponding to the initial translation and the th word unit and the th corpus sentence segment in the initial translation; is the number of corpus sentence segments of the th word unit in the initial translation; is the structure factor between the initial translation and the th word unit and the th corpus sentence segment in the initial translation; is the number of all word segmentations in the th word unit and the th corpus sentence segment in the initial translation that are the same as the word unit in the initial translation; is the number of all word segmentations in the th word unit and the th corpus sentence segment in the initial translation; is the standard deviation function.

[0102] As an embodiment, the specific method for obtaining the structure factor includes:

[0103] First, obtain the same word segmentations in the initial translation and the corpus sentence segment; take a word segmentation as a node, connect them in the order of the word segmentations in the corresponding text, and take the backward order as the corresponding connection direction.

[0104] Then, take the interval between adjacent word segmentations in the corresponding text as the edge weight of the corresponding connection edge, and respectively form the chain-like word segmentation graph structures corresponding to the initial translation and the corpus sentence segment.

[0105] Finally, obtain the graph similarity between the word segmentation graph structures corresponding to the initial translation and the corpus sentence segment through the graph edit distance, as the structure factor between the initial translation and the corpus sentence segment.

[0106] As shown in Figure 2 is the schematic flow chart for obtaining the structure factor.

[0107] It should be noted that the structure factor reflects the similarity between the initial translation and the sentence segments in the corpus in terms of the sentence structure formed by the word segmentation order. The sentence structure parameters of word units reflect the characteristics of these word units in terms of sentence structure in the initial translation. The more similar text structures formed by word segmentation in the initial translation exist in the network corpus, and the more similar the structures are, the more in line with the language habits the text structure of the corresponding word unit in the initial translation is, and the larger the corresponding sentence structure parameter is.

[0108] Step S402: Obtain the sentence segments in the network corpus that contain synonyms and near-synonyms of the word unit, and thus, for the synonyms, near-synonyms and the corresponding word unit contained in the sentence segments, conduct semantic feature analysis by means of replacement to obtain the semantic feature value of the word unit.

[0109] It should be noted that usually, although the semantic information of the corresponding synonyms and near-synonyms of a vocabulary is similar to that of the vocabulary itself in a sentence segment, in some cases, the semantic information may change. Therefore, in order to more accurately describe the semantic features reflected by the vocabulary, the present invention selects to combine the synonyms and near-synonyms of the word unit to further obtain the semantic features of the word unit.

[0110] As an embodiment, the specific method for obtaining the semantic feature value of the word unit includes:

[0111] Step S402.1: Obtain the synonyms and near-synonyms of any word unit, collectively referred to as the approximate words of the corresponding word unit, and obtain the sentence segments in the network corpus that contain the approximate words of the word unit, which are called the approximate sentence segments of the word unit.

[0112] Step S402.2: Conduct semantic feature analysis on the word unit by combining the initial translation and the approximate sentence segments of each word unit to obtain the semantic feature value of the word unit.

[0113] As an embodiment, the specific steps for obtaining the semantic features of the word unit include:

[0114] First, for any word unit, obtain all the approximate sentence segments corresponding to the word unit, obtain the approximate words corresponding to the word unit in the approximate sentence segments and the word segments adjacent to the left and right of the approximate word as the adjacent word segments of the approximate word, and obtain the phrase formed by the cosine values of the word vectors of the approximate word and each corresponding adjacent word segment, which is called the semantic feature group of the approximate word, where is a preset quantity parameter.

[0115] It should be noted that according to experience, the preset quantity parameter is 3, which can be adjusted according to the actual situation, and the embodiments of the present invention do not make specific limitations.

[0116] Then, for any word unit, use all the approximate words of this word unit to replace this word unit to obtain a new initial translation, and obtain the semantic feature group of all the approximate words of this word unit in the new initial translation. And so on, obtain the semantic feature group corresponding to the word unit.

[0117] As Figure 3 shown in the schematic diagram of the positional relationship between the word unit, the approximate words and the corresponding adjacent segmented words, which includes the initial translation of the word unit , and the approximate words of the word unit and the approximate corpus sentence segments to which they belong. The two cells adjacent to the left and right of the word unit and the approximate words and are the adjacent segmented words of the word unit or the approximate words in the corresponding sentence segment.

[0118] It should be noted that what meaning a word expresses in the text has a high correlation with the combination relationship formed with the adjacent words, but the number of adjacent words cannot be too large, otherwise it will be affected by the excessive amount of information to be conveyed in the text and cannot effectively represent this correlation feature. Therefore, the present invention selects to obtain the semantic feature group constructed by the segmented words within a preset range to represent the feature space in semantics. Taking a three-dimensional space as an example for display, as Figure 4 shown in the schematic diagram of the semantic feature space, where is the word vector of the approximate word, , and are the word vectors corresponding to the adjacent segmented words of this approximate word. Then the corresponding , and are the included angles between the word vectors. Then the cosine values obtained from the corresponding included angles form the semantic feature group of this approximate word. The smaller the included angle between the word vectors, the closer the semantic information between the corresponding words, that is, the closer the expressed meanings. Then the semantic feature group formed by the cosine values of the word vectors between multiple adjacent words in a sentence segment represents the semantic feature of the word in the meaning conveyed by the corresponding sentence segment. Therefore, the embodiments of the present invention further use the semantic feature group to quantify the semantic feature of the word vectors in the initial translation.

[0119] Finally, for any word unit in the initial translation, calculate the semantic feature value of this word unit based on the cosine value between the semantic feature groups obtained from this word unit and the corresponding approximate words in different texts.

[0120] As an embodiment, the specific calculation method of the semantic feature value of the word unit is: ​

[0121]

[0122] Among them, represents the semantic feature value of the th word unit in the initial translation; is the semantic feature group of the th word unit in the initial translation; is the semantic feature group of the approximate word after the th word unit in the initial translation is replaced by the corresponding th approximate word; is the semantic feature group of the th approximate word corresponding to the th word unit in the th approximate corpus sentence segment; is the number of approximate corpus sentence segments containing the th approximate word; is the number of approximate words of the th word unit in the initial translation; is the standard deviation function; is the cosine function; is the exponential function with the natural constant as the base.

[0123] It should be noted that represents the cosine value between the semantic feature group of the word unit and the semantic feature group of the initial translation after the word unit is replaced by the approximate word, represents the cosine value between the semantic feature group of the word unit and the corresponding semantic feature group of the approximate word in the approximate corpus sentence segment to which the approximate word belongs. The cosine value between the semantic feature groups reflects the degree of semantic change of the word unit and its approximate word in different contexts. Then reflects the discrete situation of the semantic changes of all approximate words corresponding to the word units in the initial translation. That is to say, considering whether there will be a problem in the translation process that due to the change of context, and the semantics of the word unit itself and the approximate word are easily affected by the change of context, resulting in an excessive degree of discrete semantic change and making the translation inappropriate. Furthermore, by obtaining the semantic feature value to reflect the stable feature of the word unit in semantics, when the initial translation before and after the word unit is replaced by the corresponding approximate word (i.e., synonyms and near-synonyms), and the semantic information expressed by the corresponding approximate word in the sentence segment to which it belongs do not change significantly, that is, the semantics is relatively stable, then the translation of the word unit in the translation result is accurate enough, and the corresponding semantic feature value is larger.

[0124] Step S403: For any word unit in the initial translation, comprehensively analyze the sentence structure parameters and semantic feature values of the word unit, and calculate the translation degree of the word unit.

[0125] As an embodiment, the specific calculation method for the translation degree of any word unit in the initial translation is as follows:

[0126]

[0127] Wherein, is the translation degree of the word unit; is the sentence structure parameter of the word unit; is the semantic feature value of the word unit; is the sigmoid normalization function.

[0128] It should be noted that the translation degree is the degree to which the word unit affects the semantic information in the initial translation during the translation process. The larger the translation degree, the more the word unit conforms to the corresponding language habits in terms of semantic structure and semantic information expression in the initial translation, and the better the translation effect.

[0129] Thus far, the translation degree of each word unit in the initial translation is obtained through the above method.

[0130] Step S005: Feedback to the NMT model in combination with the translation degree of the word unit, construct a feedback function, and iteratively translate the sentence segment to be translated to obtain the final translation.

[0131] It should be noted that through the translation degree, the degree of conformity of each word unit in the initial translation to the corresponding semantic information in the sentence segment to be translated in terms of the expression habits of the corresponding language is described. By quantifying this degree of conformity, the translation effect of the initial translation can be effectively clarified, and thus fed back to the initial NMT model to continuously adjust the initial translation further, improve the translation effect of the initial NMT model, and make the words in the translation more conform to the expression habits of the target language.

[0132] As an embodiment, the specific method for obtaining the final translation includes:

[0133] First, construct a feedback function.

[0134] It should be noted that according to the translation degree of the word unit in the initial translation, a feedback function is constructed. The purpose of this feedback function is to adjust the current translation result according to the translation degree, so that important word units are better processed in the translation.

[0135] As an embodiment, the basic structure of the feedback function can be expressed as:

[0136]

[0137] Wherein, is the translation degree of the th word unit; is the translation result of this word unit; Represents the total feedback after weighting the translation degrees of all word units.

[0138] It should be noted that the feedback function adjusts the translation result by weighting the translation degrees of each word unit, thereby optimizing the target translation.

[0139] Then, after obtaining the feedback function, the translation model is adjusted through an optimization algorithm. The specific steps are as follows:

[0140] Calculate the loss function: According to the feedback function, calculate the gap between the translation result and the expected translation (for example, using the cross-entropy loss function).

[0141] Gradient update: Calculate the gradient through the backpropagation algorithm and update the model parameters to reduce the loss.

[0142] Retranslation process: Use the optimized model to re-translate the source sentence to generate a new translation.

[0143] Finally, iterate the optimization to obtain the final translation.

[0144] It should be noted that each iteration optimizes the model based on the previous round's translation result and the feedback function, making the translation result more accurate and fluent. After multiple rounds of iterative adjustment, the finally generated translation should fully consider the translation degrees of each word unit, that is, the translation quality of important word units is relatively high, while less important words can be reasonably processed in the translation, so as to generate a high-quality target language translation.

[0145] Through the above steps, the translation of the sentence segment to be translated is completed to obtain the final translation.

[0146] In other embodiments of the present invention, a language translation electronic device based on big data is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the content of steps S001 to S005 in the above-mentioned language translation method based on big data.

[0147] Further, in an alternative embodiment, the memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory may be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, Static RAM (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0148] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0149] Please refer to Figure 5 , which shows a block diagram of a big data-based language translation system provided by an embodiment of the present invention. The system includes the following modules:

[0150] An initial model module for obtaining corpus data and preliminarily training an NMT model to obtain an initial NMT model;

[0151] A preliminary translation module for obtaining a sentence segment to be translated and translating the sentence segment to be translated through the initial NMT model to obtain an initial translation;

[0152] A corpus collection module for using web crawler technology to crawl web corpus data to construct a web corpus;

[0153] A translation analysis module for semantically analyzing the relationship of each word unit in the initial translation in terms of sentence structure in combination with the web corpus, obtaining the sentence structure parameters of each word unit in the initial translation, and replacing the word unit with its synonyms and near-synonyms in the web corpus in the corresponding sentence, analyzing the semantic change characteristics before and after replacement, obtaining the semantic feature values of each word unit in the initial translation, and combining the sentence structure parameters of any word unit in the initial translation with the semantic feature values to obtain the translation degree of the word unit;

[0154] A feedback translation module for feeding back to the NMT model in combination with the translation degree of the word unit, constructing a feedback function and iteratively translating the sentence segment to be translated to obtain a final translation.

[0155] After constructing the initial NMT model and the web corpus in this embodiment, through semantic contrast analysis of the word units in the initial translation with their synonyms and near-synonyms in the web corpus in terms of syntactic structure and synonym replacement, a translation degree index is further quantitatively generated to effectively identify the translation ambiguity points in the initial translation. The translation degree parameter is incorporated into the weight adjustment of the NMT model through an iterative optimization mechanism to achieve a progressive improvement in translation quality. This method can significantly improve translation quality by combining big data technology, thereby optimizing the semantic expression of the translation and enhancing the accuracy and context adaptability of the translation.

[0156] It should be noted that the model used in this embodiment is only used to represent a negative correlation relationship and to constrain the result of the model output to be within the interval. In specific implementation, it can be replaced with other models with the same purpose. This embodiment only takes the model as an example for description and does not make specific limitations on it. Among them, refers to the input of the model.

[0157] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A language translation method based on big data, characterized in that: The method comprises the following steps: Obtain corpus data and preliminarily train the NMT model to obtain the initial NMT model; Obtain the sentence segment to be translated, and translate it through the initial NMT model to obtain the initial translation; Use crawler technology to crawl network corpus data to build a network corpus; Combined with the network corpus, semantic analysis is performed on the relationship between each word unit in the initial translation and the sentence structure parameters of each word unit in the initial translation are obtained. The synonyms and near synonyms of the word unit in the network corpus are used to replace the word unit in the corresponding sentence, and the semantic change characteristics before and after the replacement are analyzed to obtain the semantic feature value of each word unit in the initial translation. The sentence structure parameters of any word unit in the initial translation are combined with the semantic feature value to obtain the translation degree of the word unit. Feedback is given to the NMT model based on the translation degree of word units, a feedback function is constructed, and the sentences to be translated are translated iteratively to obtain the final translation. The method of combining the network corpus to semantically analyze the relationship between each word unit in the initial translation and the sentence structure to obtain the sentence structure parameters of each word unit in the initial translation includes the following specific methods: The initial translation is segmented using Jieba word segmentation, and each phrase in the segmentation result is taken as a word unit; for any word unit, several sentence segments containing the arbitrary word unit in the network corpus are obtained and recorded as the corpus sentence segments of the arbitrary word unit; Perform text structure analysis on the initial translation of any word unit and the corpus sentence segment to obtain the cosine structure sequence of the initial translation and each corpus sentence segment; According to the similarity of the cosine structure sequences corresponding to the initial translation and all the corpus sentence segments corresponding to each word unit in the initial translation, and in combination with the structural features of the corpus sentence segments corresponding to each word unit, the sentence structure parameters of any word unit are obtained; The specific method for obtaining the sentence structure parameters of the arbitrary word unit is: Obtain the structural factors between the initial translation and the corpus segment; The specific calculation method of the sentence structure parameters of the word unit is: in, The first Sentence structure parameters of word units; The initial translation and the first word unit The DTW distance between the cosine structure sequences corresponding to the corpus sentences; The first The number of corpus segments with word units; The initial translation and the first word unit Structural factors between corpus segments; The first word unit The number of all the participles in the corpus segment that are identical to the word units in the initial translation; The first word unit The number of all participles in a corpus segment; is the standard deviation function; The specific method for obtaining the structural factor between the initial translation and the corpus sentence segment is: Get the same participles in the initial translation and the corpus segment; take a participle as a node, connect them according to the order of the participles in the corresponding text, and take the backward order as the corresponding connection direction; The intervals between adjacent segmented words in the corresponding text are used as the edge weights of the corresponding connecting edges, forming chain segmented graph structures corresponding to the initial translation and the corpus sentence segments; The graph similarity between the word segmentation graph structures corresponding to the initial translation and the corpus segment is obtained through the graph edit distance as a structural factor between the initial translation and the corpus segment.

2. According to the language translation method based on big data as claimed in claim 1, it is characterized in that: The text structure analysis of the initial translation and the corpus sentence segment of any word unit is performed to obtain the cosine structure sequence of the initial translation and each corpus sentence segment, including the specific method of: First, the corpus segments are segmented to obtain several phrases in any corpus segment, called corpus word units. The word2vec algorithm is used to convert the word units in the initial translation and the corpus word units in all corpus segments into word vectors to obtain word units and word vectors of corpus word units. Then, the cosine value of the word vector corresponding to any adjacent word unit in the initial translation is obtained, and a sequence formed by the cosine values ​​between the word vectors of all adjacent word units in the initial translation is obtained, which is recorded as the cosine structure sequence of the initial translation; and so on, the cosine structure sequence of any corpus sentence segment is obtained.

3. According to the language translation method based on big data as claimed in claim 1, it is characterized in that: The specific method for obtaining the semantic feature value of the word unit is: Obtain synonyms and near-synonyms of any word unit, collectively referred to as approximate words of the corresponding word unit, and obtain corpus segments containing the approximate words of the word unit in the network corpus, referred to as approximate corpus segments of the word unit; The semantic feature analysis of each word unit is performed based on the initial translation and the approximate corpus sentence of each word unit to obtain the semantic feature value of the word unit.

4. According to the language translation method based on big data as claimed in claim 3, it is characterized in that: The specific method for obtaining the semantic feature value of the word unit is: For any word unit, obtain all approximate corpus segments corresponding to the word unit, obtain the approximate words corresponding to the word unit in the approximate corpus segments, and the words adjacent to the approximate words on the left and right. The word segmentation is used as the neighboring word segmentation of the approximate word, and the word group formed by the cosine value of the word vector of the approximate word and each corresponding neighboring word segmentation is obtained, which is recorded as the semantic feature group of the approximate word, where is the preset quantity parameter; For any word unit, replace the word unit with all the similar words of the word unit to obtain a new initial translation, and obtain the semantic feature group of all the similar words of the word unit in the new initial translation; and so on, obtain the semantic feature group corresponding to the word unit; The semantic feature value of the word unit is calculated based on the cosine value between the semantic feature groups of the word unit and the corresponding similar words obtained in different texts.

5. According to the language translation method based on big data as claimed in claim 4, it is characterized in that: The specific calculation method of the semantic feature value of the word unit is: in, Indicates the first The semantic feature value of each word unit; The first The semantic feature group of each word unit; The first The word unit is corresponding to After the similar words are replaced, the semantic feature group of the similar words; The first The word unit corresponds to the Similar words in the The semantic feature groups in the similar corpus segments; To contain The number of similar corpus segments with similar words; The first The number of similar words to the word unit; is the standard deviation function; is the cosine function; is an exponential function with a natural constant as its base.

6. A language translation electronic device based on big data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the language translation method based on big data as described in any one of claims 1 to 5 are implemented.

7. A language translation system based on big data, using a language translation method based on big data as claimed in any one of claims 1 to 5, characterized in that: The system includes the following modules: The initial model module is used to obtain corpus data and preliminarily train the NMT model to obtain the initial NMT model; The preliminary translation module is used to obtain the sentence segments to be translated, and translate the sentence segments to be translated through the initial NMT model to obtain the initial translation; Corpus collection module, used to crawl network corpus data using crawler technology, so as to build a network corpus; The translation analysis module is used to perform semantic analysis on the relationship between each word unit in the initial translation and the sentence structure in combination with the network corpus, obtain the sentence structure parameters of each word unit in the initial translation, and replace the word unit in the corresponding sentence with the synonyms and near synonyms in the network corpus, analyze the semantic change characteristics before and after the replacement, obtain the semantic feature value of each word unit in the initial translation, and combine the sentence structure parameters and semantic feature values ​​of any word unit in the initial translation to obtain the translation degree of the word unit; The feedback translation module is used to provide feedback to the NMT model based on the translation degree of word units, build a feedback function and iteratively translate the sentences to be translated to obtain the final translation.

Citation Information

Patent Citations

  • Data-enhanced machine translation method based on similar word and synonym replacement

    CN108920473A

  • Training corpus set construction method, translation model training method and translation method

    CN115587590A