English translation system and method based on cloud technology
By setting up a taboo detection module in the cloud server, the problem of inaccurate translation of taboo content caused by cultural differences is solved, and the interpretation information is automatically replaced or added during the translation process, ensuring the accuracy of translation and smooth communication.
Patent Information
- Application Number
- CN202510255266.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing English translation system fails to effectively deal with taboo content caused by differences in culture, customs, etc., resulting in misunderstandings in communication.
Set up a taboo detection module in the cloud server. Through the taboo detection module, determine whether there are taboo words in the word segmentation, and find replacement words in the replacement corpus or obtain additional information for translation to ensure the accuracy of the translation.
By automatically replacing or adding interpretation information, misunderstandings caused by taboo terms are avoided, and the accuracy of the translated content and smooth communication are ensured.
Smart Images

Figure CN120373321A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of online translation technology, and particularly to an English translation system and method based on cloud technology. Background Art
[0002] English is one of the world's common languages. In many international occasions, such as international academic conferences and exchanges, real-time and accurate English translation is required to enable both parties to communicate to accurately and quickly understand each other's intentions, thereby ensuring the smooth progress of the communication. The traditional translation method is for one of the two parties to communicate to learn the other's language, or for a professional translator to perform real-time translation. This translation method has the problem of high cost.
[0003] With the development of modern network technology, real-time online translation services have become possible. After using online translation technology, users only need to input the content to be translated in the form of voice or text, and with the help of cloud technology, it can be transmitted to the cloud server, where it is translated into the target language and finally transmitted to the user's electronic device. This translation technology has almost no threshold and cost, so it has been widely used. When using computer technology for English translation, literal translation and free translation are usually adopted. Literal translation is to abide by the words and structures of the original text, convert each word in the original text into a word in the target language, and then connect them to form a sentence. This translation method is relatively simple and direct, but it is not satisfactory in terms of the smoothness of the sentence and the expression of meaning. Therefore, in many cases, free translation is adopted. Free translation means that there is no need to abide by the words and structures of the original text, as long as the meaning expressed after translation is the same as the original text, which provides a large room for play for translation, and it can be accepted in terms of the smoothness of the sentence and the accuracy of the meaning expression.
[0004] However, whether literal translation or free translation is adopted, the first criterion for current English translation work is to ensure the accuracy of the translation result. However, in addition to accuracy, there are often huge differences in culture, customs, etc. between the two parties to the communication. If only accuracy is required, it will inevitably lead to misunderstandings when there are taboo contents in the communication content, seriously affecting the smooth progress of the communication. Summary of the Invention
[0005] The embodiments of this application provide an English translation system and method based on cloud technology to solve the problem in the prior art that the translation of taboo contents caused by differences in culture, customs, etc. is inaccurate.
[0006] On the one hand, the embodiments of this application provide an English translation system based on cloud technology, including:
[0007] A translation terminal for obtaining the original text;
[0008] A cloud server is used to split the original text into multiple word segments, determine the part-of-speech of each word segment and the structure of the original text, input the word segments, part-of-speech, and structure into a pre-trained translation model to obtain the corresponding English translation. During the translation process, a taboo detection module is built into the cloud server. The taboo detection module is used to determine whether there are taboo terms in the word segments. If there are, it determines whether there are replacement words corresponding to the taboo terms in the replacement corpus. If there are, it replaces the word segments belonging to the taboo terms with the replacement words, and then the translation model continues to perform the translation. If there are no replacement words corresponding to the taboo terms in the replacement corpus, the cloud server sends a notification message to the translation terminal to remind the sender to input additional information for assisting in translation. After the cloud server obtains the additional information, it combines the additional information with the original text to form replacement content, and uses the translation model to translate the replacement content.
[0009] After obtaining the English translation, the cloud server sends the English translation to the translation terminal, and the translation terminal displays the English translation.
[0010] On the other hand, an embodiment of the present application also provides an English translation method based on cloud technology, including:
[0011] Obtain the original text;
[0012] Split the original text into multiple word segments, determine the part-of-speech of each word segment and the structure of the original text, input the word segments, part-of-speech, and structure into a pre-trained translation model to obtain the corresponding English translation;
[0013] During the translation process, determine whether there are taboo terms in the word segments. If there are, determine whether there are replacement words corresponding to the taboo terms in the replacement corpus. If there are, replace the word segments belonging to the taboo terms with the replacement words, and then the translation model continues to perform the translation; if there are no replacement words corresponding to the taboo terms in the replacement corpus, obtain the additional information for assisting in translation input by the sender, combine the additional information with the original text to form replacement content, and use the translation model to translate the replacement content;
[0014] Display the English translation.
[0015] An English translation system and method based on cloud technology in the present application have the following advantages:
[0016] By setting up a taboo detection module, when there are taboo terms in the Chinese text to be translated, it can automatically perform replacement or eliminate the misunderstanding of the recipient by adding explanatory information, avoiding possible misunderstandings while ensuring the accuracy of the translation content. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of the functional modules of an English translation system based on cloud technology provided by an embodiment of the present application. Specific embodiments
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.
[0020] Figure 1 It is a schematic diagram of the composition of an English translation system based on cloud technology provided by an embodiment of the present application. An embodiment of the present application provides an English translation system based on cloud technology, including:
[0021] A translation terminal for obtaining the original text;
[0022] A cloud server for splitting the original text into multiple word segments, determining the part of speech of each word segment and the structure of the original text, inputting the word segments, part of speech, and structure into a pre-trained translation model to obtain the corresponding English translation; during the translation process, a taboo detection module is built into the cloud server. The taboo detection module is used to determine whether there are taboo terms in the word segments. If there are, it is determined whether there are replacement words corresponding to the taboo terms in the replacement corpus. If there are, the word segments belonging to the taboo terms are replaced with the replacement words, and then the translation model continues to translate. If there are no replacement words corresponding to the taboo terms in the replacement corpus, the cloud server sends a notification message to the translation terminal to remind the sender to input additional information for assisting translation. After the cloud server obtains the additional information, it combines the additional information with the original text to form replacement content, and uses the translation model to translate the replacement content;
[0023] After obtaining the English translation, the cloud server sends the English translation to the translation terminal, and the translation terminal displays the English translation.
[0024] Exemplarily, the translation terminal can adopt any electronic device capable of connecting to the network, such as a smart phone, a tablet computer, a desktop computer, etc. The cloud server is set in the cloud, on which a source text segmentation module, a translation model, a taboo detection module and a replacement corpus are deployed. The source text segmentation module therein is used to segment the source text to determine each word segment in the source text, and then analyze the part of speech of each word segment and the structure of the source text sentence. A word segment is the smallest unit that can be obtained after segmenting a sentence, which can be a single character or a word composed of only a few characters, and the part of speech includes nouns, verbs, adjectives, auxiliary words, conjunctions, etc. When analyzing the part of speech, it can be classified according to a pre-set part-of-speech comparison table. The structure of a sentence includes subject-predicate-object, subject-predicate-object-complement, etc. The structure of a sentence needs to be determined based on the part of speech, where the subject and object are usually nouns, and the predicate is usually a verb.
[0025] In an embodiment of the present application, the translation model is established based on the Transformer model. In the Transformer model, it includes a preprocessing layer, an encoder and a decoder. The preprocessing layer first converts each word segment into a vector representation, and then performs position encoding on each word segment based on the structure of the sentence. The encoder uses the self-attention mechanism to calculate the correlation of each vector representation, further updates the vector representation based on this correlation, uses a feed-forward neural network to further process and transform the vector representation of each word segment, and finally enhances the stability of the translation model through added residual connections and layer normalization. The decoder then uses the self-attention masking and cross-attention mechanisms to consider relevant word segments, and finally uses a feed-forward neural network to further process the information output by the decoding module in the decoder, converts it into the distribution probability of English words through the Softmax layer, and uses this distribution probability to sequentially determine each word in the English translation.
[0026] The method for the taboo detection module to determine whether there is a taboo term in the word segment includes: comparing each word segment with each candidate word in the replacement corpus. If the word segment is the same as any one of the candidate words, it is considered that the word segment is a taboo term, and the candidate word is used as the replacement word.
[0027] The replacement corpus therein is a word library pre-established by experts, which stores taboo terms and corresponding candidate words in various different regions. Due to differences in living habits, customs, cultures, etc. in different regions, the same taboo term may need to be replaced with different candidate words. Therefore, there is a many-to-many correspondence relationship between taboo terms and candidate words.
[0028] Specifically, when the taboo detection module compares the word segment and the candidate word, it can first convert them into word vectors in the vector space respectively, and then calculate the distance between the two word vectors in the vector space. If the distance is less than the set distance threshold, it can be considered that the word segment and the candidate word are the same.
[0029] However, there is a many-to-many relationship between the taboo terms and candidate words in this application. If the specific situation of the recipient is not considered, there may be many candidate words available. Therefore, before comparing the segmented words with the candidate words, this application also obtains the above text information, analyzes the above text information to determine the background information of the recipient, filters the candidate words according to the background information, and then compares the segmented words with the candidate words among the filtered candidate words.
[0030] The above text information can be extracted from the content of the recipient's reply. When extracting, some key segmented words can be analyzed. These key segmented words can be words representing country names, region names, and some unique words in certain countries and regions. If the above text information contains these key segmented words, the region where the recipient is located can be directly determined through these key segmented words, and then the background information of the recipient can be obtained.
[0031] After determining the background information of the recipient, the candidate words can be filtered. Specifically, when filtering the candidate words, corresponding possibility scores are set for each candidate word according to the background information, and the candidate words with possibility scores exceeding the score threshold are retained. In the embodiments of this application, each candidate word is set with multiple tags, and each tag is set with a pre-determined possibility score, where the tag corresponds to a region. When the region in the background information matches any one of the tags of the candidate word, the candidate word can be set with the possibility score corresponding to the tag.
[0032] For example, when it is found through analyzing the above text information of the recipient that the region of the recipient is Mumbai, the background information of the recipient can be set to India. The candidate word "meat" in the replacement corpus has a tag for India. If the segmented word "beef" appears in the original text, since there is a corresponding relationship between "beef" and India, the taboo detection module will consider "beef" to be a taboo term. At this time, the candidate word "meat" will be set with the possibility score corresponding to the tag, such as 1.0, while other candidate words, such as "pork", "chicken", etc., will not be set with possibility scores because these candidate words do not have a tag for India.
[0033] If the obtained possibility score exceeds the pre-set score threshold, such as 0.6, it can be considered that the possibility of the candidate word being used as a replacement word is very high, and then these candidate words will all be retained. Then, the vector distance between each candidate word and the segmented word is calculated, and the most accurate candidate word is selected from multiple candidate words as the replacement word.
[0034] Further, when analyzing the above information, extract the keywords in the above information and determine whether each keyword corresponds to any region. If it does, summarize all the regions corresponding to the keywords to form background information. In the above example, "Mumbai" is the keyword, and this keyword corresponds to the region "India", so the background information includes India.
[0035] Further, there is a corresponding relationship between the candidate word and the keyword belonging to the region. When there are taboo requirements in the region to which the keyword belongs, set a high possibility score for the candidate word corresponding to the keyword. When there are no taboo requirements in the region to which the keyword belongs, set a low possibility score for the candidate word corresponding to the keyword, and the high possibility score is higher than the low possibility score.
[0036] Still taking the above example for illustration. The keyword "Mumbai" is a region name. Geographically, Mumbai belongs to India, so the keyword "Mumbai" has a corresponding relationship with the background information India. And the candidate word "meat" has the label "India", so the keyword "Mumbai" has a corresponding relationship with the candidate word "meat". Since in Indian culture, cows cannot be slaughtered, beef cannot be eaten either, which is a taboo requirement. Therefore, a high possibility score, such as 1.0 as mentioned above, can be set for the candidate word "meat". If the participle "chicken" appears in the original text, since there is no taboo requirement for this participle in Indian culture, a low possibility score, such as 0.1, can be set for the candidate word "meat".
[0037] After obtaining the replacement word, the participle in the original text in Chinese form can be replaced with the replacement word to obtain the replacement content. This replacement content only replaces the participle involving taboo terms in the original text with the replacement word, but the overall sentence structure remains unchanged and the meaning expression is also almost unchanged. Therefore, when it is input into the translation model and the corresponding English translation is obtained, the accuracy of the translation result can be effectively ensured.
[0038] However, in not all cases can a replacement word be found for the participle. For example, after analyzing the above information, it is found that the recipient's region is Europe, and the original text contains "How old are you". After taking "Europe" as the background information, there are taboo requirements regarding age sensitivity. However, no corresponding replacement word for the participle "age" can be found in the replacement corpus, and directly omitting this original sentence will lead to inaccurate translation results. Therefore, this application adopts the method of adding additional information. In this case where it must be translated and there is no replacement word, use additional information to explain such taboo terms, so that the recipient can obtain the explanation information while getting the English translation of these taboo terms, avoiding misunderstandings.
[0039] The notification information sent to the translation terminal can be: "The original text mentions age, but the recipient is located in Europe, which may lead to misunderstandings. You are required to provide additional information for explanation." After the translation terminal displays this notification information, the sender can input additional information, such as: "Recommended work". After the cloud server combines the additional information with the original text, it also adds a preset explanatory statement to the combined content to form a replacement content. For example, the explanatory statement can be: "Sorry, but this is for xxx". Combining this explanatory statement with the additional information, we can get "Sorry, but this is for recommended work". Inputting the replacement content containing this information into the translation model can output the English translation.
[0040] Furthermore, the translation terminal obtains the input content in the form of voice or text through a microphone or keyboard. When the input content is in the form of voice, the translation terminal converts the voice into text and sends the text-form original text to the cloud server. Correspondingly, after obtaining the English translation, the translation terminal can also convert it into voice, or of course, directly display it in text form.
[0041] The embodiment of the present application also provides an English translation method based on cloud technology. The method includes the following steps:
[0042] Obtain the original text;
[0043] Segment the original text into multiple word segments, determine the part of speech of each word segment and the structure of the original text, and input the word segments, part of speech, and structure into a pre-trained translation model to obtain the corresponding English translation;
[0044] During the translation process, determine whether there are taboo terms in the word segments. If there are, determine whether there are replacement words corresponding to the taboo terms in the replacement corpus. If there are, replace the word segments belonging to the taboo terms with the replacement words, and then the translation model continues to translate; if there are no replacement words corresponding to the taboo terms in the replacement corpus, obtain the additional information input by the sender for assisting translation, combine the additional information with the original text to form a replacement content, and use the translation model to translate the replacement content;
[0045] Display the English translation.
[0046] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0047] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. An English translation system based on cloud technology, characterized in that Including: A translation terminal for obtaining the original text; A cloud server for splitting the original text into multiple word segments, determining the part of speech of each word segment and the structure of the original text, inputting the word segments, the part of speech and the structure into a pre-trained translation model to obtain the corresponding English translation; During the translation process, a taboo detection module is built in the cloud server. The taboo detection module is used to determine whether there are taboo terms in the word segments. If so, it is determined whether there are replacement words corresponding to the taboo terms in the replacement corpus. If there are, the word segments belonging to the taboo terms are replaced with the replacement words, and then the translation model continues to perform the translation. If there are no replacement words corresponding to the taboo terms in the replacement corpus, the cloud server sends a notification message to the translation terminal to remind the sender to input additional information for assisting in the translation. After the cloud server obtains the additional information, the additional information is combined with the original text to form replacement content, and the translation model is used to translate the replacement content; After obtaining the English translation, the cloud server sends the English translation to the translation terminal, and the translation terminal displays the English translation.
2. The English translation system based on cloud technology according to claim 1, characterized in that, The method for determining whether there are taboo terms in the word segments includes: Comparing each word segment with each candidate word in the replacement corpus. If the word segment is the same as any one of the candidate words, the word segment is considered a taboo term, and the candidate word is used as the replacement word.
3. The English translation system based on cloud technology according to claim 2, wherein, Before comparing the word segment with the candidate word, context information is also obtained, the context information is analyzed to determine the background information of the recipient, the candidate words are screened according to the background information, and then the word segment is compared with the candidate words in the screened candidate words.
4. An English translation system based on cloud technology according to claim 3, characterized in that, When screening the candidate words, corresponding possibility scores are set for each candidate word according to the background information, and the candidate words with possibility scores exceeding the score threshold are retained.
5. An English translation system based on cloud technology according to claim 3, characterized in that, The background information includes the region where the recipient is located. When analyzing the context information, keywords in the context information are extracted, and it is determined whether each keyword corresponds to any region. If so, the regions corresponding to all the keywords are summarized to form the background information.
6. An English translation system based on cloud technology according to claim 5, characterized in that, The candidate word has a corresponding relationship with the keyword belonging to the region. When there are taboo requirements in the region to which the keyword belongs, a high possibility score is set for the candidate word corresponding to the keyword. When there are no taboo requirements in the region to which the keyword belongs, a low possibility score is set for the candidate word corresponding to the keyword. The high possibility score is higher than the low possibility score.
7. An English translation system based on cloud technology according to claim 1, characterized in that, After combining the additional information with the original text, a preset explanatory statement is also added to the combined content to form the replacement content.
8. An English translation system based on cloud technology according to claim 1, characterized in that The translation terminal obtains input content in the form of voice or text through a microphone or keyboard. When the input content is in the form of voice, the translation terminal converts the voice into text, and the translation terminal sends the original text in text form to the cloud server.
9. A cloud technology-based English translation system according to claim 1, characterized in that, The translation model is established based on the Transformer model.
10. A method applied to an English translation system based on cloud technology according to any one of claims 1-9, characterized in that, It includes: Obtain the original text; Segment the original text into multiple word tokens, determine the part of speech of each word token and the structure of the original text, and input the word tokens, the part of speech, and the structure into a pre-trained translation model to obtain the corresponding English translation; During the translation process, determine whether there are taboo words among the word tokens. If there are, determine whether there are replacement words corresponding to the taboo words in the replacement corpus. If there are, replace the word tokens belonging to the taboo words with the replacement words, and then continue the translation by the translation model; if there are no replacement words corresponding to the taboo words in the replacement corpus, obtain the additional information input by the sender for assisting translation, combine the additional information with the original text to form replacement content, and use the translation model to translate the replacement content; Display the English translation.
Citation Information
Patent Citations
Machine translation method and device based on term replacement
CN112541365A
Chinese humor classification model based on reverse translation
CN112818118A
Machine translation system and machine translation method
CN114528859A
Term recognition method for multi-language translation
CN116822517A
Translation method, translation model generation method and related equipment
CN117010416A