An English translation system and method based on cloud technology

By using a cloud-based taboo detection module and replacement corpus, the problem of inaccurate translation of taboo content caused by cultural differences has been solved, achieving more accurate and culturally appropriate English translation.

CN120373321BActive Publication Date: 2026-04-24SHAANXI TECHN INST OF DEFENSE IND
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHAANXI TECHN INST OF DEFENSE IND
Filing Date
2025-03-05
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Current English translation technologies fail to effectively handle taboo content caused by cultural and customary differences, leading to misunderstandings in communication.

Method used

The cloud-based translation system uses a taboo detection module to identify taboo words and searches for replacement words or obtains additional information from the replacement corpus to ensure translation accuracy.

Benefits of technology

Effectively eliminate misunderstandings caused by taboo terms, ensuring that the translated content is more reliable in terms of accuracy and cultural adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373321B_ABST
    Figure CN120373321B_ABST
Patent Text Reader

Abstract

The application discloses an English translation system and method based on cloud technology, and relates to the online translation technical field.The system comprises a translation terminal used for obtaining original text, and a cloud server used for translating the original text by using a translation model to obtain English translation text.The cloud server is internally provided with a taboo detection module, which is used for judging whether there is taboo language in the original text.If there is, the corresponding replacement word is inquired from a replacement corpus, the word segmentation is replaced by the replacement word, and then the translation model continues to translate.The application can automatically replace the taboo language in the Chinese text to be translated or eliminate the misunderstanding of the receiver by adding explanation information, so that the possible misunderstanding is avoided on the basis of ensuring the accuracy of the translation content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of online translation technology, and in particular to an English translation system and method based on cloud technology. Background Technology

[0002] English is one of the world's most widely used languages. In many international settings, such as international academic conferences and exchanges, real-time and accurate English translation is essential to ensure that both parties understand each other's intentions quickly and accurately, thus guaranteeing smooth communication. Traditional translation methods involve one party learning the other's language or having a professional translator provide real-time translation. However, these methods are costly.

[0003] With the development of modern network technology, real-time online translation services have become possible. Using online translation technology, users only need to input the content to be translated in the form of voice or text. With the help of cloud technology, this information is transmitted to a cloud server, translated into the target language, and finally transmitted to the user's electronic device. This translation technology is virtually barrier-free and cost-free, thus it has been widely adopted. When using computer technology for English translation, literal translation and free translation are commonly used. Literal translation adheres to the words and structure of the original text, converting each word in the original text into words in the target language and then connecting them to form a sentence. This translation method is relatively simple and direct, but it is unsatisfactory in terms of sentence fluency and meaning expression. Therefore, free translation is often used. Free translation does not need to adhere to the words and structure of the original text; it only needs to ensure that the translated meaning is consistent with the original. This provides a great deal of room for interpretation, and it is generally acceptable in terms of sentence fluency and accuracy of meaning expression.

[0004] However, regardless of whether literal or free translation is used, the primary principle of current English translation work is to ensure the accuracy of the translation result. But besides accuracy, significant cultural and customary differences often exist between the communicating parties. If absolute accuracy is demanded, misunderstandings are inevitable when taboo content is included in the communication, seriously hindering the smooth progress of the exchange. Summary of the Invention

[0005] This application provides an English translation system and method based on cloud technology to solve the problem of inaccurate translation of taboo content caused by cultural and customary differences in the prior art.

[0006] On the one hand, embodiments of this application provide an English translation system based on cloud technology, including:

[0007] Translation terminal, used to obtain the source text;

[0008] The cloud server is used to segment the original text into multiple words, determine the part-of-speech tag of each word and the structure of the original text, and input the word segments, part-of-speech tags, and structure into a pre-trained translation model to obtain the corresponding English translation. During the translation process, the cloud server has a built-in taboo detection module, which is used to determine whether there are taboo words in the word segments. If there are, it determines whether there is a replacement word corresponding to the taboo word in the replacement corpus. If there is, the word segment belonging to the taboo word is replaced with the replacement word, and then the translation model continues to translate. If there is no replacement word corresponding to the taboo word in the replacement corpus, the cloud server sends a notification message to the translation terminal, reminding the sender to input additional information to assist in the translation. After obtaining the additional information, the cloud server combines the additional information with the original text to form the replacement content, and uses the translation model to translate the replacement content.

[0009] After receiving the English translation, the cloud server sends the English translation to the translation terminal, which then displays the English translation.

[0010] On the other hand, embodiments of this application also provide an English translation method based on cloud technology, including:

[0011] Get the original text;

[0012] The original text is segmented into multiple words, the part of speech of each word is determined, and the structure of the original text is determined. The word segmentation, part of speech, and structure are then input into a pre-trained translation model to obtain the corresponding English translation.

[0013] During the translation process, it is determined whether there are any taboo words in the word segmentation. If so, it is checked whether there are any corresponding replacement words in the replacement corpus. If so, the word segment containing the taboo word is replaced with the replacement word, and then the translation model continues to translate. If there are no corresponding replacement words in the replacement corpus, additional information input by the sender to assist in the translation is obtained. The additional information is combined with the original text to form the replacement content, and the translation model is used to translate the replacement content.

[0014] Display the English translation.

[0015] The cloud-based English translation system and method described in this application have the following advantages:

[0016] By setting up a taboo detection module, the system can automatically replace or add explanatory information when taboo terms are present in the Chinese text to be translated, thus eliminating potential misunderstandings for the recipient while ensuring the accuracy of the translation. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the functional modules of an English translation system based on cloud technology, provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] Figure 1 This is a schematic diagram illustrating the composition of a cloud-based English translation system provided in this application embodiment. This application embodiment provides a cloud-based English translation system, including:

[0021] Translation terminal, used to obtain the source text;

[0022] The cloud server is used to segment the original text into multiple words, determine the part-of-speech tag of each word and the structure of the original text, and input the word segments, part-of-speech tags, and structure into a pre-trained translation model to obtain the corresponding English translation. During the translation process, the cloud server has a built-in taboo detection module, which is used to determine whether there are taboo words in the word segments. If there are, it determines whether there is a replacement word corresponding to the taboo word in the replacement corpus. If there is, the word segment belonging to the taboo word is replaced with the replacement word, and then the translation model continues to translate. If there is no replacement word corresponding to the taboo word in the replacement corpus, the cloud server sends a notification message to the translation terminal, reminding the sender to input additional information to assist in the translation. After obtaining the additional information, the cloud server combines the additional information with the original text to form the replacement content, and uses the translation model to translate the replacement content.

[0023] After receiving the English translation, the cloud server sends the English translation to the translation terminal, which then displays the English translation.

[0024] For example, the translation terminal can be any network-connected electronic device, such as a smartphone, tablet, or desktop computer. The cloud server is located in the cloud and deploys a text segmentation module, a translation model, a tabu detection module, and a substitution corpus. The text segmentation module segments the original text to identify each word, then analyzes the part of speech and sentence structure of each word. A word segment is the smallest unit that can be obtained after segmenting a sentence; it can be a single character or a word composed of only a few characters. Parts of speech include nouns, verbs, adjectives, particles, conjunctions, etc., and can be classified according to a pre-defined part-of-speech lookup table. Sentence structure includes subject-verb-object, subject-verb-object-complement, etc., and the sentence structure needs to be determined based on parts of speech. The subject and object are usually nouns, while the predicate is usually a verb.

[0025] In the embodiments of this application, the translation model is based on the Transformer model. The Transformer model includes a preprocessing layer, an encoder, and a decoder. The preprocessing layer first converts each word segment into a vector representation, and then performs positional encoding on each word based on the sentence structure. The encoder uses a self-attention mechanism to calculate the relevance of each vector representation, and further updates the vector representation based on this relevance. A feedforward neural network is used to further process and transform the vector representation of each word. Finally, residual connections and layer normalization are added to enhance the stability of the translation model. The decoder uses self-attention masking and cross-attention mechanisms to consider relevant word segments. Finally, a feedforward neural network is used to further process the information output by the decoding module in the decoder, converting it into a probability distribution of English words through a Softmax layer. This probability distribution is then used to determine each word in the English translation in sequence.

[0026] The method for the taboo detection module to determine whether there are taboo words in the word segmentation includes: comparing each word segment with each candidate word in the replacement corpus; if the word segment is the same as any candidate word, the word segment is considered to be a taboo word and the candidate word is used as the replacement word.

[0027] The replacement corpus is a pre-built lexicon by experts, which stores taboo terms and corresponding candidate words from various regions. Due to differences in living habits, customs, and culture across regions, the same taboo term may need to be replaced by different candidate words, thus there is a many-to-many correspondence between taboo terms and candidate words.

[0028] Specifically, when comparing word segments and candidate words, the tabu detection module can first convert them into word vectors in the vector space, and then calculate the distance between the two word vectors in the middle of the vector. If the distance is less than the set distance threshold, the word segments and candidate words can be considered to be the same.

[0029] However, there is a many-to-many relationship between the prohibited terms and candidate words in this application. If the specific circumstances of the recipient are not considered, there may be many candidate words to choose from. Therefore, before comparing the word segment with the candidate words, this application also obtains the context information, analyzes the context information to determine the background information of the recipient, filters the candidate words based on the background information, and then compares the word segment with the candidate words from the filtered candidate words.

[0030] The preceding information can be extracted from the recipient's reply. During extraction, key word segments can be analyzed. These key word segments may represent country names, region names, or words unique to certain countries or regions. If the preceding information contains these key word segments, the recipient's location can be directly determined, thus obtaining the recipient's background information.

[0031] After determining the recipient's background information, candidate words can be filtered. Specifically, during the filtering process, a corresponding probability score is assigned to each candidate word based on the background information, and candidate words with probability scores exceeding a threshold are retained. In the embodiments of this application, each candidate word is assigned multiple tags, and each tag is assigned a pre-determined probability score. The tags correspond to regions, and when a region in the background information matches any tag of a candidate word, that candidate word can be assigned the probability score corresponding to that tag.

[0032] For example, if the analysis of the recipient's preceding information reveals that the recipient's region is Mumbai, the recipient's background information can be set to India, replacing the label for the candidate word "meat" in the corpus that indicates India. If the word "beef" appears in the original text, since "beef" corresponds to India, the taboo detection module will consider "beef" a taboo term. In this case, the candidate word "meat" will be set with a probability score corresponding to the label, such as 1.0. Other candidate words, such as "pork" and "chicken," will not have a probability score set because they do not have the label for India.

[0033] If the probability score obtained from the screening exceeds a pre-set scoring threshold, such as 0.6, it can be considered that the candidate word has a high probability of being used as a replacement word, and these candidate words will be retained. Then, the vector distance between each candidate word and the word segment is calculated, and the most accurate candidate word is selected from multiple candidate words as the replacement word.

[0034] Furthermore, when analyzing the preceding information, keywords are extracted, and it is determined whether each keyword corresponds to any given region. If a correspondence is found, all regions corresponding to the keywords are aggregated to form background information. In the example above, "Mumbai" is a keyword, and this keyword corresponds to the region "India," therefore the background information includes India.

[0035] Furthermore, there is a correspondence between candidate words and keywords belonging to a region. When there are restrictions in the region to which a keyword belongs, the candidate words corresponding to the keyword will be given a high probability score. When there are no restrictions in the region to which a keyword belongs, the candidate words corresponding to the keyword will be given a low probability score. The high probability score is higher than the low probability score.

[0036] Let's illustrate with the example above. The keyword "Mumbai" is a regional name; geographically, Mumbai is in India, therefore the keyword "Mumbai" corresponds to the background information of India. The candidate word "meat" has the tag "India," so the keyword "Mumbai" corresponds to the candidate word "meat." Since cows cannot be slaughtered in Indian culture, beef is also forbidden, which is a taboo. Therefore, the candidate word "meat" can be assigned a high probability score, such as 1.0 as mentioned above. However, if the original text contains the segment "chicken," since there is no taboo regarding this segment in Indian culture, the candidate word "meat" can be assigned a low probability score, such as 0.1.

[0037] After obtaining the replacement words, the word segments in the original Chinese text can be replaced with the replacement words to obtain the replacement content. This replacement content only replaces the word segments in the original text that contain taboo words with the replacement words, but the overall sentence structure remains unchanged, and the meaning is almost unchanged as well. Therefore, after inputting it into the translation model and obtaining the corresponding English translation, the accuracy of the translation result can be effectively ensured.

[0038] However, word segmentation cannot always find a replacement word. For example, after analyzing the preceding information, it was found that the recipient's region is Europe. The original text contains the phrase "How old are you?". With "Europe" as background information, there is a taboo requirement sensitive to age. However, the word segment for "age" cannot find a corresponding replacement word in the replacement corpus. Simply omitting this sentence from the original text would lead to inaccurate translation results. Therefore, this application adopts the method of adding additional information. In cases where translation is necessary and no replacement word exists, additional information is used to explain the taboo term, so that the recipient receives the explanation information along with the English translation of these taboo terms, avoiding misunderstandings.

[0039] The notification message sent to the translation terminal could be: "The original text mentions age, but the recipient is located in Europe, which may lead to misunderstanding. Further information is needed for clarification." After the translation terminal displays this notification, the sender can enter additional information, such as "job recommendation." The cloud server combines this additional information with the original text and adds a pre-defined explanatory statement to the combined content, forming a replacement message. For example, the explanatory statement could be: "Sorry, but this is for xxx." Combining this explanatory statement with the additional information yields "Sorry, but this is for job recommendation." Inputting this replacement message, containing these elements, into the translation model will output the English translation.

[0040] Furthermore, the translation terminal acquires input in the form of speech or text via a microphone or keyboard. When the input is in speech form, the translation terminal converts the speech into text and sends the original text to the cloud server. Similarly, after receiving the English translation, the translation terminal can also convert it back into speech, or it can display it directly as text.

[0041] This application also provides a cloud-based English translation method, which includes the following steps:

[0042] Get the original text;

[0043] The original text is segmented into multiple words, the part of speech of each word is determined, and the structure of the original text is determined. The word segmentation, part of speech, and structure are then input into a pre-trained translation model to obtain the corresponding English translation.

[0044] During the translation process, it is determined whether there are any taboo words in the word segmentation. If so, it is checked whether there are any corresponding replacement words in the replacement corpus. If so, the word segment containing the taboo word is replaced with the replacement word, and then the translation model continues to translate. If there are no corresponding replacement words in the replacement corpus, additional information input by the sender to assist in the translation is obtained. The additional information is combined with the original text to form the replacement content, and the translation model is used to translate the replacement content.

[0045] Display the English translation.

[0046] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0047] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A cloud-based English translation system, characterized in that, include: Translation terminal, used to obtain the source text; A cloud server is used to segment the original text into multiple words, determine the part-of-speech tag of each word and the structure of the original text, and input the word segments, the part-of-speech tag, and the structure into a pre-trained translation model to obtain the corresponding English translation. During the translation process, the cloud server has a built-in taboo detection module. This module is used to determine whether there are taboo words in the word segmentation. If there are, it determines whether there is a replacement word corresponding to the taboo word in the replacement corpus. If there is, the word segment containing the taboo word is replaced with the replacement word, and then the translation model continues to translate. If there is no replacement word corresponding to the taboo word in the replacement corpus, the cloud server sends a notification message to the translation terminal, reminding the sender to input additional information to assist in the translation. After obtaining the additional information, the cloud server combines the additional information with the original text to form replacement content, and then uses the translation model to translate the replacement content. After receiving the English translation, the cloud server sends the English translation to the translation terminal, and the translation terminal displays the English translation. The methods for determining whether the prohibited term exists in the word segmentation include: Each word segment is compared with each candidate word in the replacement corpus. If a word segment is the same as any candidate word, the word segment is considered a taboo word, and the candidate word is used as the replacement word. When comparing word segments and candidate words, the taboo detection module first converts them into word vectors in the vector space, and then calculates the distance between the two word vectors in the vector space. If the distance is less than a set distance threshold, the word segment and candidate word are considered to be the same. Before comparing the word segment with the candidate word, the preceding context information is also obtained, the preceding context information is analyzed to determine the background information of the receiver, the candidate word is filtered according to the background information, and then the word segment is compared with the candidate word in the filtered candidate word. When filtering the candidate words, a corresponding probability score is set for each candidate word according to the background information, and the candidate words whose probability scores exceed the score threshold are retained; each candidate word is set with multiple tags, and each tag is set with a predetermined probability score, wherein the tags correspond to regions, and when the region in the background information matches any tag of the candidate word, the candidate word is set to the probability score corresponding to that tag; The background information includes the region where the recipient is located. When analyzing the above information, keywords are extracted from the above information, and it is determined whether each keyword corresponds to any region. If they correspond, the regions corresponding to all the keywords are summarized to form the background information. After combining the additional information with the original text, a preset explanatory statement is added to the combined content to form the replacement content.

2. The English translation system based on cloud technology according to claim 1, characterized in that, The candidate words correspond to the keywords of the regions to which they belong. When there are taboo requirements in the region to which the keyword belongs, the candidate words corresponding to the keyword are given a high probability score. When there are no taboo requirements in the region to which the keyword belongs, the candidate words corresponding to the keyword are given a low probability score. The high probability score is higher than the low probability score.

3. The English translation system based on cloud technology according to claim 1, characterized in that, The translation terminal acquires input content in the form of voice or text through a microphone or keyboard. When the input content is in the form of voice, the translation terminal converts the voice into text and sends the original text in text form to the cloud server.

4. The English translation system based on cloud technology according to claim 1, characterized in that, The translation model is based on the Transformer model.

5. Applicable to claim 1 4. A method for an English translation system based on cloud technology as described in any one of the claims, characterized in that, include: Get the original text; The original text is segmented into multiple words, the part of speech of each word is determined, and the structure of the original text is determined. The word segments, the part of speech, and the structure are input into a pre-trained translation model to obtain the corresponding English translation. During the translation process, it is determined whether there are any taboo words in the word segmentation. If so, it is determined whether there is a corresponding replacement word in the replacement corpus. If so, the word segment containing the taboo word is replaced with the replacement word, and then the translation model continues to translate. If there is no corresponding replacement word in the replacement corpus, additional information input by the sender to assist in the translation is obtained. The additional information is combined with the original text to form replacement content, and the translation model is used to translate the replacement content. The English translation is displayed.

Citation Information

Patent Citations

  • Term recognition method for multi-language translation

    CN116822517A

  • Translation method, translation model generation method and related equipment

    CN117010416A

  • Information processing apparatus and non-transitory computer readable medium

    US20190087416A1

  • Server for providing global language translation service, method for providing global language translation service, and program for providing global language translation service

    WO2023075274A1