Tourism translation explanation standardized translation method based on corpus

Through the analysis of overall emotion vector and semantic correlation based on the corpus, the final English translation of tourism translation explanation was determined, which solved the problem that existing translation engines could not express emotions and context, and achieved a more accurate and emotionally rich translation effect.

CN120337950AActive Publication Date: 2025-07-18GUIZHOU BUSINESS SCHOOL
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510824891.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

In the existing travel translation explanation, the existing translation engine cannot effectively express the emotions and context of the original text, resulting in poor translation results.

Method used

Based on the corpus, we determine the overall emotional vector and semantic correlation of the original text of the tourism translation interpretation, integrate emotional correlation and semantic correlation, obtain the translation normativeness, and select the final English translation.

Benefits of technology

It improves the accuracy of translation and emotional expression ability, making the translation results more in line with the emotions and context of the original text, and improves the translation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337950A_ABST
    Figure CN120337950A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine-aided translation, in particular to a standardized translation method for tourism translation explanation based on a corpus, which comprises the following steps: determining to-be-selected English translation of a tourism translation explanation original text based on the corpus; determining an overall sentiment vector of the tourism translation original text, and obtaining sentiment correlation between the tourism translation original text and to-be-selected English translation; fusing the emotion relevance and the semantic relevance between the tourism translation explanation original text and the English translation to be selected to obtain a translation specification degree of the English translation to be selected; and determining a final English translation from the English translations to be selected according to the translation specification degree. According to the method, the emotion factors are considered at the same time, so that the final English translation is accurately obtained from the English translations to be selected, the obtained final English translation is more related to and accurate to the tourism translation original text, the emotion of the tourism translation original text can be well expressed, and the translation effect of tourism translation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine-assisted translation, and particularly relates to a standardized translation method for tourism translation interpretation based on a corpus. Background Art

[0002] Chinese-English translation is becoming more and more common in tourism translation interpretation. Among them, it is very important to accurately express the meaning of Chinese interpretation in Chinese-English translation. Currently, artificial intelligence assistance is often used to assist in manual translation. When translating the original text of tourism interpretation, the existing technology often gives translation results according to the translation engine. However, most translation engines can only do literal translation rigidly, and only translate and recommend vocabulary references according to the lexical semantics. In the process of translating the original text of some interpretations with strong emotional colors, the effect of literal translation often cannot express the original emotion well, and cannot effectively express the context and emotion of the original text, resulting in poor translation effect. Summary of the Invention

[0003] In order to solve the technical problem of poor translation effect of existing tourism translation interpretation, the purpose of the present invention is to provide a standardized translation method for tourism translation interpretation based on a corpus, and the specific technical solutions adopted are as follows: The present invention provides a standardized translation method for tourism translation interpretation based on a corpus, including: Based on the corpus, determine the candidate English translations of the original text of tourism translation interpretation, and the candidate English translations include the initial English translation and at least one candidate English translation; Determine the overall emotional vector of the original text of tourism translation interpretation, and obtain the emotional correlation between the original text of tourism translation interpretation and the candidate English translations; Fuse the emotional correlation and the semantic correlation between the original text of tourism translation interpretation and the candidate English translations to obtain the translation standard degree of the candidate English translations; According to the translation standard degree, determine the final English translation from the candidate English translations.

[0004] In an exemplary embodiment, the process of obtaining the overall emotional vector includes: Based on the number of emotional intensities existing in the same dimension of the emotional vectors of each original word in the original text of tourism translation interpretation, determine the emotional intensity reference index of each dimension; According to the numerical values of the emotional intensities of each dimension in the emotional vectors of each original word, obtain the emotional intensity characteristics of the same dimension for each original word; Fuse the emotional intensity characteristics of the same dimension of all original words to obtain the comprehensive emotional intensity characteristics of each dimension; Fuse the emotional intensity reference index of each dimension and the comprehensive emotional intensity characteristics to obtain the emotional sub-vector intensity of each dimension; Sort the emotional score vector intensities of each dimension to obtain the overall emotional vector.

[0005] In an exemplary embodiment, the process of obtaining the emotional intensity reference index includes: normalizing the number of emotional intensities existing in the same dimension, and the obtained result is the emotional intensity reference index of this dimension.

[0006] In an exemplary embodiment, the process of obtaining the emotional intensity feature includes: determining the emotional intensity weights of each dimension in the emotional vectors of each original text word, and weighting the numerical values of the emotional intensities according to the emotional intensity weights to obtain the emotional intensity feature.

[0007] In an exemplary embodiment, the process of obtaining the emotional intensity weight includes: Obtain the emotional vector similarity of every two adjacent original text words in the original text of the tourism translation commentary to obtain a similarity sequence; Based on the similarity sequence, segment the original text of the tourism translation commentary to obtain multiple original text word segments; According to the number of original text words in the original text word segment and the average value of the emotional vector similarity corresponding to the original text word segment, obtain the emotional intensity weights of each dimension in the emotional vectors of each original text word in the original text word segment; the emotional intensity weight is directly proportional to both the number of original text words and the average value of the emotional vector similarity.

[0008] In an exemplary embodiment, the segmenting the original text of the tourism translation commentary based on the similarity sequence to obtain multiple original text word segments includes: Starting from the first emotional vector similarity in the similarity sequence, when an emotional vector similarity less than a preset similarity threshold appears, use the position between the two original text words corresponding to this emotional vector similarity as a segmentation point to segment the original text of the tourism translation commentary until traversing to the last emotional vector similarity in the similarity sequence, thereby obtaining multiple original text word segments.

[0009] In an exemplary embodiment, the process of obtaining the emotional correlation includes: Obtain the emotional similarity between the overall emotional vector of the original text of the tourism translation commentary and the emotional vectors of each candidate translation word in the target candidate English translation to obtain the emotional correlation corresponding to each candidate translation word in the target candidate English translation; the target candidate English translation is any one of the candidate English translations; The process of obtaining the semantic correlation includes: Obtain the semantic similarity between the semantic vectors of each original text word and the semantic vectors of the corresponding candidate translation words in the target candidate English translation, and obtain the semantic relevance corresponding to each candidate translation word in the target candidate English translation.

[0010] In an exemplary embodiment, the process of obtaining the translation standardization degree includes: Fuse the emotional relevance and semantic relevance of the same candidate translation word in the target candidate English translation to obtain the lexical relevance of the candidate translation word; Based on the lexical relevance of all candidate translation words in the target candidate English translation, obtain the translation standardization degree of the target candidate English translation.

[0011] In an exemplary embodiment, the obtaining the translation standardization degree of the target candidate English translation based on the lexical relevance of all candidate translation words in the target candidate English translation includes: Fuse the lexical relevance of all candidate translation words in the target candidate English translation to obtain the comprehensive relevance of the target candidate English translation; Compare the part-of-speech of each candidate translation word in the target candidate English translation with the corresponding original text word, and determine the number of target candidate translation words in the target candidate English translation; the number of target candidate translation words is the number of candidate translation words whose part-of-speech has not changed; According to the comprehensive relevance and the number of target candidate translation words, obtain the translation standardization degree of the target candidate English translation; the translation standardization degree is directly proportional to the comprehensive relevance and inversely proportional to the number of target candidate translation words.

[0012] In an exemplary embodiment, the process of obtaining the candidate English translation includes: Based on the corpus, determine at least one candidate translation word for each original text word in the original text of the tourism translation commentary, and respectively combine one of the candidate translation words of each original text word to obtain a candidate English translation.

[0013] The present invention has the following beneficial effects: In addition to the initial English translation corresponding to the original text of tourism translation and interpretation, at least one candidate English translation is obtained, which together constitute the candidate English translations to be selected. By obtaining the overall sentiment vector of the original text of tourism translation and interpretation, the sentiment correlation between the original text of tourism translation and interpretation and each candidate English translation is obtained. At the same time, the semantic correlation between the original text of tourism translation and interpretation and each candidate English translation is obtained. Analyzing from both aspects of the two sentiment correlations and semantic correlations, the translation standardization degree of each candidate English translation is obtained. Finally, according to the translation standardization degree, the final English translation is determined from the candidate English translations. Compared with the existing method of only determining the translation result based on semantics, the present invention takes into account both semantic factors and sentiment factors, accurately obtains the final English translation from multiple candidate English translations, and the obtained final English translation is more relevant and accurate to the original text of tourism translation and interpretation, and can well express the sentiment of the original text of tourism translation and interpretation, improving the translation effect of tourism translation and interpretation. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 FIG. is a flowchart of a method for standardizing translation of tourism translation and interpretation based on a corpus provided by an embodiment of the present invention; Figure 2 FIG. is a flowchart of obtaining an overall sentiment vector provided by an embodiment of the present invention; Figure 3 FIG. is a flowchart of obtaining a sentiment intensity weight provided by an embodiment of the present invention; Figure 4 FIG. is a flowchart of obtaining a translation standardization degree provided by an embodiment of the present invention; Figure 5 FIG. is a specific implementation flowchart of step S3-2 provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and their effects of the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The data information collected in this application is obtained with full consent and authorization.

[0017] This embodiment provides a corpus-based standardized translation method for tourism translation interpretations, which can be applied to Chinese-English translation devices for tourists whose native language is English.

[0018] The corpus used in this embodiment needs to include a large number of Chinese words and a large number of English words, and for each Chinese word, at least one corresponding English word, that is, at least one English word with the same meaning as the Chinese word, can be obtained from it to achieve the translation effect. For any word in the large number of Chinese words and a large number of English words in the corpus, each word contains two word vectors, which are respectively denoted as semantic vectors and sentiment vectors . The semantic vector of each word can be obtained by existing model algorithms, such as through the Word2Vec model. The sentiment vector of each word has the same number of dimensions (denote the number of dimensions as ), and it is a multi-dimensional vector. The values of different dimensions represent the sentiment intensity in different sentiment directions. In this embodiment, the value range of each dimension is limited to for convenient calculation. The larger the value of the dimension, the stronger the sentiment intensity in the corresponding sentiment direction. For example, the sentiment vector of a certain word is . It should be understood that for the dimension with a value of 0, it means that there is no sentiment intensity in this dimension, and for the dimension with a value greater than 0, it means that there is sentiment intensity in this dimension. Based on the above requirements, a corpus is constructed.

[0019] As Figure 1 shown, a corpus-based standardized translation method for tourism translation interpretations provided by this embodiment includes: Step S1: Based on the corpus, determine the candidate English translations of the original text of the tourism translation interpretation. The candidate English translations include the initial English translation and at least one candidate English translation; Step S2: Determine the overall sentiment vector of the original text of the tourism translation interpretation, and thereby obtain the sentiment correlation between the original text of the tourism translation interpretation and the candidate English translations; Step S3: Integrate the sentiment correlation and the semantic correlation between the original text of the tourism translation interpretation and the candidate English translations to obtain the translation norm degree of the candidate English translations; Step S4: Determine the final English translation from the candidate English translations according to the translation norm degree.

[0020] The specific implementation processes of each step are described below in combination with the accompanying drawings.

[0021] Step S1: Based on the corpus, determine the candidate English translations of the original text of the tourism translation interpretation. The candidate English translations include the initial English translation and at least one candidate English translation.

[0022] The original text of the tourist translation and explanation is the Chinese original text that needs to be translated, such as the Chinese explanation of tourist attractions expressed by the tour guide. The purpose of the present invention is to determine the final English translation of the original text of the tourist translation and explanation.

[0023] This embodiment uses an existing translation engine, which is connected to a corpus. The original travel translation explanation is input into the translation engine, and the translation engine outputs an English translation. It should be understood that the English translation output by the translation engine is used as the initial English translation.

[0024] Due to the different expression habits between Chinese and English, some parts of speech may change during the translation process. For example, when translating "We need to solve this problem" into English, we get "We need a solution to this problem." The Chinese verb "解决问题" is translated into the noun "solution" during the translation process. Although the part of speech changes, the semantics of the expression remains unchanged.

[0025] For the sake of convenience, each word in the original text of the tourism translation and commentary is defined as each original word. In order to standardize the translation, it is first necessary to match the semantics of each translated word (i.e., English word) with the original word one by one to obtain a semantic matching pair for subsequent change standardization. Due to the phenomenon of part of speech changes in the above-mentioned translation process, matching cannot be performed only based on part of speech. The degree of semantic consistency between words must also be considered to match the original word with the translated word.

[0026] Take a sentence and its translation as an example to match the synonyms of the two. Specifically: First, the original words need to be separated. Since there is no clear space symbol to separate words in Chinese, they need to be analyzed word by word. The original sentence is stored as a string and recorded as a string , from the first character of the string At the beginning, set a temporary and initialized empty string as , the characters join in Inside, with Search the corpus to see if there is a matching string in the corpus (if it exists, it is recorded as ) or a string with it as a substring (if it exists, it is recorded as ).

[0027] There is a containment relationship between some words, such as "肝胆" and "肝胆相照". After identifying the matching words, it is still necessary to further determine whether they are contained in another word. Search the resulting string in the corpus and , and use this to perform lexical judgment and division.

[0028] If only exists, it indicates that the string of the original text vocabulary is not a complete vocabulary, and it is necessary to continue to add the next character in the original text sentence and perform the retrieval again.

[0029] If and both exist, it indicates that the string of the original text vocabulary is a complete vocabulary, but this vocabulary may still be part of another vocabulary, and it is necessary to continue to add characters for retrieval and judgment.

[0030] If only exists, it indicates that is a complete vocabulary and is not included in other vocabularies, and the retrieval is stopped. Divide this vocabulary, set to be empty, and add the next character after this vocabulary in the original text sentence for the division of the next character.

[0031] If and both do not exist, it indicates that the previous retrieved vocabulary is a complete vocabulary, and no new other vocabulary can be obtained by combining it with its subsequent characters in the original text. Divide the previous obtained vocabulary, set to be empty and perform the division of the next character.

[0032] Divide each vocabulary in the original text sentence using the above character division method, and record the number of the original text vocabularies obtained after division as .

[0033] In an exemplary embodiment, taking the obtained th original text vocabulary as an example, record the semantic vector of this original text vocabulary as , obtain the semantic vectors of each translation vocabulary in the corpus, obtain the cosine similarity between the th original text vocabulary and the semantic vectors of each translation vocabulary in the corpus, take the maximum cosine similarity, and record it as , and use the translation vocabulary corresponding to the maximum cosine similarity as the initial translation vocabulary of the th original text vocabulary, so as to obtain the initial translation vocabularies of each original text vocabulary. Combine the initial translation vocabularies of each original text vocabulary into a sentence (the problem of grammar errors needs to be avoided), and the obtained sentence is the initial English translation of the tourism translation commentary original text.

[0034] Then, for the th original text vocabulary, set a cosine similarity threshold, which is ( is the semantic selection threshold, which can be customized by the user. The setting of needs to ensure that each original word has at least one candidate translation word. In this embodiment, ) to obtain the th original word and the cosine similarity of the semantic vectors of each translation word in the corpus, and obtain at least one translation word whose cosine similarity is greater than the cosine similarity threshold . Define each translation word other than the initial translation word among these translation words as a candidate translation word, so as to obtain at least one candidate translation word for the th original word. Thus, at least one candidate translation word for each original word is obtained.

[0035] In principle, by selecting one candidate translation word from the multiple candidate translation words of each original word for sentence combination, a candidate English translation will be obtained, and thus multiple candidate English translations will be obtained. The number of candidate English translations obtained through permutation and combination is the product of the number of candidate translation words of each original word. For example: if there are 4 original words, and the number of candidate translation words of these four original words are 2, 3, 3, and 4 respectively, then, through permutation and combination, the total number of sentence combinations that can be obtained is: 2×3×3×4 = 72. However, some of the sentences obtained through permutation and combination may not meet the grammatical requirements, that is, some sentences have grammar errors. Then, delete these combination methods with grammar errors and retain the combination methods with correct grammar, so as to obtain at least one candidate English translation of the original text of the tourism translation commentary. Collectively refer to the initial English translation of the original text of the tourism translation commentary and each candidate English translation as the candidate English translations of the original text of the tourism translation commentary. The present invention selects one English translation from the candidate English translations as the final English translation of the original text of the tourism translation commentary.

[0036] Step S2: Determine the overall emotional vector of the original text of the tourism translation commentary, and thereby obtain the emotional correlation between the original text of the tourism translation commentary and the candidate English translations.

[0037] For the original text of the tourism translation commentary, it may contain a large number of original words with emotional colors. The initial English translation obtained by literal translation may not fully conform to the original emotional expression, resulting in a significant reduction in the translation effect. Since different words have their own emotional intensities, combined with the context in the original text of the tourism translation commentary, the overall emotional color emphasized by each original word can be inferred, so as to standardize the words in the English translation.

[0038] According to the emotional vectors of each original word in the original text of the tourism translation commentary, obtain the overall emotional vector of the original text of the tourism translation commentary. In an exemplary embodiment, as Figure 2 shown, the following gives a specific acquisition process of the overall emotional vector: Step S2-1: Determine the emotional intensity reference index for each dimension based on the number of emotional intensities in the same dimension among the emotional vectors of each original word in the tourism translation commentary original text.

[0039] For any dimension, obtain the number of emotional intensities in this dimension among the emotional vectors of each original word in the tourism translation commentary original text. As can be seen from the above, if the emotional intensity of this dimension is 0, it means that there is no emotional intensity in this dimension. If the emotional intensity of this dimension is greater than 0, it means that there is emotional intensity in this dimension. Based on this, determine whether there is emotional intensity in this dimension in the emotional vector of each original word. Then, based on the emotional vectors of all original words, obtain the number of original words with emotional intensity in this dimension, and get the number of existing emotional intensities corresponding to this dimension. Then normalize the number of existing emotional intensities in this dimension, and the result obtained is the emotional intensity reference index for this dimension.

[0040] In an exemplary embodiment, a normalization method is given as follows: ; where, represents the emotional intensity reference index of the i-th dimension, represents the number of existing emotional intensities of the i-th dimension, represents the number of existing emotional intensities of the b-th dimension, represents the number of dimensions of the emotional vector.

[0041] Step S2-2: Obtain the emotional intensity characteristics of the same dimension for each original word according to the numerical values of the emotional intensities of each dimension in the emotional vector of each original word.

[0042] The numerical values of each dimension in the emotional vector of the original word are the emotional intensities of each dimension. The larger the numerical value, the greater the emotional intensity. According to the numerical values of the emotional intensities of each dimension in the emotional vector of each original word, obtain the emotional intensity characteristics of the same dimension for each original word. In an exemplary embodiment, determine the emotional intensity weights of each dimension in the emotional vector of each original word, and weight the numerical values of the emotional intensities according to the emotional intensity weights to obtain the emotional intensity characteristics.

[0043] In an exemplary embodiment, as Figure 3 shown, a specific acquisition process of the emotional intensity weight is given as follows: Step S2-2-1: Obtain the similarity of the emotional vectors of every two adjacent original words in the tourism translation commentary original text to obtain a similarity sequence.

[0044] In the original text of tourism translation interpretation, obtain the emotional vector similarity of every two adjacent original text words. Taking the th original text word as an example, obtain the emotional vector similarity between the emotional vector of the th original text word and the emotional vector of the th original text word, where . is the number of original text words. The higher the emotional vector similarity, the higher the emotional consistency of two adjacent original text words.

[0045] In this embodiment, the emotional vector similarity is specifically the cosine similarity, that is, obtain the cosine similarity between the emotional vector of the th original text word and the emotional vector of the th original text word. Since the value range of the cosine similarity is from -1 to 1, in order to ensure that the cosine similarity is within the range of 0 - 1 for convenient data processing, this embodiment normalizes the cosine similarity, and all cosine similarities involved in this embodiment are normalized cosine similarities. The normalization method of the cosine similarity can be: (cosine similarity + 1) / 2.

[0046] Arrange each emotional vector similarity in the order of the original text words to obtain a similarity sequence.

[0047] Step S2-2-2: Based on the similarity sequence, segment the original text of tourism translation interpretation to obtain multiple original text segments.

[0048] The similarity sequence represents the emotional vector similarity of every two adjacent original text words. Then, according to the similarity sequence, segment the original text of tourism translation interpretation to obtain multiple original text segments, so that the emotional vector similarities of the original text words in the same original text segment are relatively high, and the emotional vector similarities of the original text words between different original text segments are relatively low.

[0049] In an exemplary embodiment, a similarity threshold is preset. This preset similarity threshold is used to confirm whether the emotional vector similarity is relatively high. The value range of this preset similarity threshold is 0 - 1, and its specific value is set according to actual needs, such as 0.7.

[0050] Starting from the first emotional vector similarity in the similarity sequence, compare it with the preset similarity threshold respectively. Specifically: if the first emotional vector similarity is greater than or equal to the preset similarity threshold, then continue to compare the second emotional vector similarity with the preset similarity threshold. If the second emotional vector similarity is greater than or equal to the preset similarity threshold, then continue to compare the third emotional vector similarity with the preset similarity threshold until an emotional vector similarity less than the preset similarity threshold appears. Then, take the position between the two original words corresponding to the emotional vector similarity less than the preset similarity threshold as the segmentation point to segment the original text of the tourism translation commentary. The first original word of the original text of the tourism translation commentary to the previous original word of these two original words is used as an original word segment. Then, based on the remaining original words, continue to compare the subsequent emotional vector similarities with the preset similarity threshold, and then follow the above process until an emotional vector similarity less than the preset similarity threshold appears. Take the position between the two original words corresponding to the emotional vector similarity less than the preset similarity threshold as the segmentation point to segment the remaining original words. Then, the first original word of the remaining original words to the previous original word of these two original words is used as an original word segment. And so on until the last emotional vector similarity in the similarity sequence is traversed. Thus, the segmentation of the original text of the tourism translation commentary is completed, and multiple original word segments are obtained. It should be understood that there is a possibility that a single original word forms an original word segment in this embodiment.

[0051] A specific example is given as follows: Suppose there are 10 original words, and the number of emotional vector similarities in the similarity sequence is 9. Starting from the first emotional vector similarity, compare it with the preset similarity threshold. If the third emotional vector similarity is less than the preset similarity threshold, since the two original words corresponding to the third emotional vector similarity are the third original word and the fourth original word, then the first original word to the third original word is used as an original word segment. Then, starting from the fourth original word and from the fourth emotional vector similarity, compare it with the preset similarity threshold. If the fifth emotional vector similarity is less than the preset similarity threshold, since the two original words corresponding to the fifth emotional vector similarity are the fifth original word and the sixth original word, then the fourth original word and the fifth original word are used as an original word segment. Then, starting from the sixth original word and from the sixth emotional vector similarity, compare it with the preset similarity threshold until the last emotional vector similarity is greater than or equal to the preset similarity threshold. Then, all the remaining original words are used as an original word segment.

[0052] Step S2-2-3: Obtain the emotional intensity weights of each dimension in the emotional vector of each original word in the original word segment according to the number of original words in the original word segment and the average value of the emotional vector similarity corresponding to the original word segment.

[0053] For any original word segment, the emotions of the original words in this original word segment are consistent and continuous. Moreover, the greater the number of original words in this original word segment, the higher the emotional intensity weights of each dimension in the emotional vector of each original word in this original word segment.

[0054] Calculate the average value of the emotional vector similarity corresponding to this original word segment. The higher the average value of the emotional vector similarity, the higher the emotional intensity weights of each dimension in the emotional vector of each original word in this original word segment. Then, the emotional intensity weights are directly proportional to both the number of original words and the average value of the emotional vector similarity. The higher the emotional intensity weights, the higher the continuous concentration degree of the emotional words in the original word segment.

[0055] It should be understood that if two original words both belong to this original word segment, the emotional vector similarity corresponding to these two original words belongs to the emotional vector similarity of this original word segment. Based on the above example, for the original word segment composed of the 1st to 3rd original words, the emotional vector similarity of this original word segment includes: the emotional vector similarity between the 1st and 2nd original words, and the emotional vector similarity between the 2nd and 3rd original words; for the original word segment composed of the 4th and 5th original words, the emotional vector similarity of this original word segment is the emotional vector similarity between the 4th and 5th original words. Then, for an original word segment containing only one original word, the average value of the emotional vector similarity of this type of original word segment is set as the average value of the emotional vector similarities of each original word segment including at least two original words.

[0056] In an exemplary embodiment, a specific quantification method of the emotional intensity weights is given as follows: ; where, represents the emotional intensity weight of each dimension in the emotional vector of each original word in the original word segment where the -th original word is located, represents the number of original words in the original word segment where the -th original word is located, represents the average value of the emotional vector similarity corresponding to the original word segment where the -th original word is located.

[0057] represents the normalization function. Indicates the normalization of. In this embodiment, the normalization method here can be: obtain the maximum and minimum values in the corresponding to each original text word, and then adopt the maximum-minimum normalization method to normalize the corresponding to the of the nth original text word.

[0058] It can be seen from the above calculation method that for the original text word segment where the nth original text word is located, the emotional intensity weights of each dimension in the emotional vectors of all original text words belonging to the same original text word segment are the same, and are all equal to . Then, can also represent the emotional intensity weights of each dimension in the emotional vector of the nth original text word.

[0059] Therefore, the emotional intensity weights of all dimensions in the emotional vector of the same original text word are the same, and the emotional intensity weights of each dimension in the emotional vectors of each original text word belonging to the same original text word segment are the same.

[0060] Then, the emotional intensity feature of the ith dimension in the emotional vector of the nth original text word is equal to: . Among them, is the value of the emotional intensity of the ith dimension in the emotional vector of the nth original text word.

[0061] Step S2-3: Integrate the emotional intensity features of the same dimension of all original text words to obtain the comprehensive emotional intensity features of each dimension.

[0062] Calculate the average value of the emotional intensity features of the same dimension of all original text words, and the result obtained is the comprehensive emotional intensity feature of this dimension. The calculation formula is as follows: ; Among them, is the comprehensive emotional intensity feature of the ith dimension.

[0063] Step S2-4: Integrate the emotional intensity reference indicators and comprehensive emotional intensity features of each dimension to obtain the emotional sub-vector intensity of each dimension.

[0064] For any dimension, calculate the product of the emotional intensity reference indicator of this dimension and the comprehensive emotional intensity feature of this dimension as the emotional sub-vector intensity of this dimension. The calculation formula is as follows: ; Among them, It represents the intensity of the sentiment score vector for the i-th dimension. In this way, the intensity of the sentiment score vector for each dimension is obtained. It represents the reference index of the sentiment intensity for the i-th dimension.

[0065] Step S2-5: Sort the intensity of the sentiment score vector for each dimension to obtain the overall sentiment vector.

[0066] Sort the intensity of the sentiment score vector for each dimension in the order of the dimensions to obtain the overall sentiment vector. , which is . Among them, represents the intensity of the sentiment score vector for the 1st dimension, represents the intensity of the sentiment score vector for the 2nd dimension, represents the intensity of the sentiment score vector for the last dimension.

[0067] Each original word has its own unique emotional color, and moreover, based on each dimension of the sentiment vector of the original word, the same original word can have multiple emotional expressions. In a sentence composed of multiple original words, as a whole, the multiple original words together constitute the lexical distribution trend of the overall sentence, that is, the overall sentiment vector.

[0068] The more the words selected in the translation conform to the overall sentiment vector of the original text, the higher the emotional consistency and the more suitable for translation interpretation. Then, after obtaining the overall sentiment vector of the original text of the tourism translation interpretation, according to the overall sentiment vector of the original text of the tourism translation interpretation, obtain the emotional correlation between the original text of the tourism translation interpretation and each candidate English translation.

[0069] In an exemplary embodiment, for the sake of convenience of explanation, set the target candidate English translation as any one of the candidate English translations, so as to obtain the sentiment vectors of each candidate translation word in the target candidate English translation. Each candidate translation word in the target candidate English translation corresponds one by one to each original word in the original text of the tourism translation interpretation.

[0070] Then, for any candidate translation word in the target candidate English translation, obtain the emotional similarity between the overall sentiment vector of the original text of the tourism translation interpretation and the sentiment vector of this candidate translation word. This emotional similarity is specifically the cosine similarity, that is, obtain the cosine similarity between the overall sentiment vector of the original text of the tourism translation interpretation and the sentiment vector of this candidate translation word, and the obtained cosine similarity is used as the emotional correlation between the original text of the tourism translation interpretation and this candidate translation word. In this way, obtain the emotional correlation between the original text of the tourism translation interpretation and each candidate translation word in the target candidate English translation.

[0071] For translation work, semantic consistency remains the most fundamental guarantee. Then, this embodiment also obtains the semantic relevance of each candidate translation word in the original text of the tourist translation commentary and the target candidate English translation. In an exemplary embodiment, for any candidate translation word in the target candidate English translation, the original word corresponding to the candidate translation word is determined. Then, the semantic similarity between the semantic vector of the original word and the semantic vector of the candidate translation word is obtained. The semantic similarity is specifically the cosine similarity, and the obtained cosine similarity is used as the semantic relevance between the original word and the candidate translation word. In this way, the semantic relevance between each original word in the original text of the tourist translation commentary and the corresponding candidate translation words in the target candidate English translation is obtained.

[0072] Step S3: Integrate the emotional relevance and the semantic relevance between the original text of the tourist translation commentary and the candidate English translation to obtain the translation standard degree of the candidate English translation.

[0073] According to Step S2, the emotional relevance corresponding to each candidate translation word in the original text of the tourist translation commentary and the target candidate English translation, and the semantic relevance between each original word in the original text of the tourist translation commentary and the corresponding candidate translation words in the target candidate English translation are obtained. Then, these two aspects of relevance are integrated to obtain the translation standard degree of the target candidate English translation. In an exemplary embodiment, as Figure 4 shown, a specific process for obtaining the translation standard degree is given: Step S3-1: Integrate the emotional relevance and the semantic relevance of the same candidate translation word in the target candidate English translation to obtain the word relevance of the candidate translation word.

[0074] For any candidate translation word in the target candidate English translation, the emotional relevance and the semantic relevance corresponding to the candidate translation word are obtained, and the average value of the emotional relevance and the semantic relevance is calculated. The obtained result is the word relevance of the candidate translation word. In this way, the word relevance of each candidate translation word in the target candidate English translation is obtained.

[0075] Step S3-2: Based on the word relevance of all the candidate translation words in the target candidate English translation, obtain the translation standard degree of the target candidate English translation.

[0076] Integrate the word relevance of all the candidate translation words in the target candidate English translation to obtain the translation standard degree of the target candidate English translation. In an exemplary embodiment, the average value of the word relevance of all the candidate translation words in the target candidate English translation can be calculated, and this average value is directly used as the translation standard degree of the target candidate English translation.

[0077] As other implementation manners, in order to simultaneously consider the influence of the number of candidate translation words with unchanged parts of speech in the candidate translation words on the translation standardization of the target candidate English translation, such as Figure 5 shown, another implementation process of the translation standardization of the target candidate English translation is provided as follows: Step S3-2-1: Integrate the lexical relevance of all candidate translation words in the target candidate English translation to obtain the comprehensive relevance of the target candidate English translation.

[0078] Integrating the lexical relevance of all candidate translation words in the target candidate English translation specifically means calculating the average value of the lexical relevance of all candidate translation words in the target candidate English translation as the comprehensive relevance of the target candidate English translation.

[0079] Step S3-2-2: Compare the parts of speech of each candidate translation word in the target candidate English translation with the corresponding original word to determine the number of target candidate translation words of the target candidate English translation.

[0080] Since the original words and the candidate translation words in the target candidate English translation are in one-to-one correspondence. Then, in the correspondence relationship between each original word and its corresponding candidate translation word, after translating from Chinese to English, obtain the number of candidate translation words with unchanged parts of speech, and define it as the number of target candidate translation words, that is, the number of target candidate translation words is the number of candidate translation words with unchanged parts of speech. The more the number of candidate translation words with unchanged parts of speech, the more literal the translation is, the less convenient it is for English tourists to understand, and the lower the translation standardization of the target candidate English translation, that is, the translation standardization is inversely proportional to the number of target candidate translation words.

[0081] Moreover, the stronger the comprehensive relevance of the target candidate English translation, the more suitable the target candidate English translation is for the translation of the original sentence, and the higher the translation standardization of the target candidate English translation, that is, the translation standardization is directly proportional to the comprehensive relevance.

[0082] Step S3-2-3: Obtain the translation standardization of the target candidate English translation according to the comprehensive relevance and the number of target candidate translation words.

[0083] Obtain the translation standardization of the target candidate English translation according to the comprehensive relevance of the target candidate English translation and the number of target candidate translation words of the target candidate English translation. In an exemplary embodiment, a quantification method is given as follows: ; wherein, represents the translation standardization of the x-th candidate English translation, represents the comprehensive relevance of the x-th candidate English translation, represents the number of target candidate translation words for the x-th candidate English translation. represents the exponential function with the natural constant e as the base. represents the negative correlation normalization of

[0084] Step S4: Determine the final English translation from the candidate English translations according to the translation standard degree.

[0085] Adopt Step S3 to obtain the translation standard degrees of each candidate English translation. The higher the translation standard degree, the more suitable it is to be used as the translation of the original text for tourism translation interpretation. Then, determine the maximum translation standard degree from the translation standard degrees of each candidate English translation, and determine the candidate English translation corresponding to the maximum translation standard degree. Output the candidate English translation corresponding to the maximum translation standard degree as the final English translation of the original text for tourism translation interpretation.

[0086] It should be noted that the above order of the embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0087] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments.

Claims

1. A corpus-based standardized translation method for tourism translation and interpretation, characterized in that, Including: Based on a corpus, determine candidate English translations for the original text of tourist translation commentary. The candidate English translations include an initial English translation and at least one alternative English translation; Determine the overall sentiment vector of the original text of tourist translation commentary, and thereby obtain the sentiment correlation between the original text of tourist translation commentary and the candidate English translations; Integrate the sentiment correlation and the semantic correlation between the original text of tourist translation commentary and the candidate English translations to obtain the translation standard degree of the candidate English translations; Determine the final English translation from the candidate English translations according to the translation standard degree; 2. The corpus-based standardized translation method for tourism translation interpretation according to claim 1, characterized in that, The process of obtaining the overall sentiment vector includes: Based on the number of sentiment intensities existing in the same dimension among the sentiment vectors of each original word in the original text of tourist translation commentary, determine the sentiment intensity reference index for each dimension; According to the numerical values of the sentiment intensities of each dimension in the sentiment vectors of each original word, obtain the sentiment intensity characteristics of the same dimension for each original word; Integrate the sentiment intensity characteristics of the same dimension of all original words to obtain the comprehensive sentiment intensity characteristics of each dimension; Integrate the sentiment intensity reference index and the comprehensive sentiment intensity characteristics of each dimension to obtain the sentiment sub-vector intensity of each dimension; Sort the sentiment sub-vector intensities of each dimension to obtain the overall sentiment vector; 3. The corpus-based standardized translation method for tourism translation and interpretation according to claim 2, characterized in that, The process of obtaining the sentiment intensity reference index includes: normalizing the number of sentiment intensities existing in the same dimension, and the result is the sentiment intensity reference index for this dimension; 4. The corpus-based standardized translation method for tourism translation interpretation according to claim 2, characterized in that, The process of obtaining the sentiment intensity characteristics includes: determining the sentiment intensity weight of each dimension in the sentiment vector of each original word, and weighting the numerical value of the sentiment intensity according to the sentiment intensity weight to obtain the sentiment intensity characteristics; 5. The corpus-based standardized translation method for tourism translation interpretation according to claim 4, characterized in that, The process of obtaining the sentiment intensity weight includes: Obtain the similarity of the sentiment vectors of every two adjacent original words in the original text of tourist translation commentary to obtain a similarity sequence; Based on the similarity sequence, segment the original text of tourist translation commentary to obtain multiple original word segments; According to the number of original words in the original word segment and the average value of the sentiment vector similarity corresponding to the original word segment, obtain the sentiment intensity weight of each dimension in the sentiment vector of each original word in the original word segment; the sentiment intensity weight is directly proportional to both the number of original words and the average value of the sentiment vector similarity; 6. The corpus-based standardized translation method for tourism translation interpretation according to claim 5, characterized in that, The segmenting the original text of tourist translation commentary based on the similarity sequence to obtain multiple original word segments includes: Starting from the first sentiment vector similarity in the similarity sequence, when the sentiment vector similarity is less than a preset similarity threshold, use the position between the two original words corresponding to this sentiment vector similarity as a segmentation point to segment the original text of tourist translation commentary until traversing to the last sentiment vector similarity in the similarity sequence, thereby obtaining multiple original word segments; 7. The corpus-based standardized translation method for tourism translation interpretation according to claim 1, characterized in that, The process of obtaining the sentiment correlation includes: Obtain the emotional similarity between the overall emotional vector of the original text of the tourism translation commentary and the emotional vectors of each candidate translation word in the target candidate English translation, and obtain the emotional relevance corresponding to each candidate translation word in the target candidate English translation; the target candidate English translation is any one of the candidate English translations. The process of obtaining the semantic relevance includes: Obtain the semantic similarity between the semantic vectors of each original word and the semantic vectors of the corresponding candidate translation words in the target candidate English translation, and obtain the semantic relevance corresponding to each candidate translation word in the target candidate English translation.

8. The corpus-based standardized translation method for tourism translation interpretation according to claim 7, characterized in that, The process of obtaining the translation standardization includes: Fuse the emotional relevance and semantic relevance of the same candidate translation word in the target candidate English translation to obtain the lexical relevance of the candidate translation word. Based on the lexical relevance of all candidate translation words in the target candidate English translation, obtain the translation standardization of the target candidate English translation.

9. The corpus-based standardized translation method for tourism translation interpretation according to claim 8, characterized in that, The obtaining the translation standardization of the target candidate English translation based on the lexical relevance of all candidate translation words in the target candidate English translation includes: Fuse the lexical relevance of all candidate translation words in the target candidate English translation to obtain the comprehensive relevance of the target candidate English translation. Compare the part-of-speech of each candidate translation word in the target candidate English translation with the corresponding original word, and determine the number of target candidate translation words in the target candidate English translation; the number of target candidate translation words is the number of candidate translation words with unchanged part-of-speech. According to the comprehensive relevance and the number of target candidate translation words, obtain the translation standardization of the target candidate English translation; the translation standardization is directly proportional to the comprehensive relevance and inversely proportional to the number of target candidate translation words.

10. The corpus-based standardized translation method for tourism translation interpretation according to claim 1, characterized in that, The process of obtaining the candidate English translation includes: Based on the corpus, determine at least one candidate translation word for each original word in the original text of the tourism translation commentary, and respectively combine one of the candidate translation words of each original word to obtain a candidate English translation.

Citation Information

Patent Citations

  • Automatic name translation system and method

    CN107861953A

  • Automatic student answer scoring method for English examination translation questions

    CN112085985A

  • Machine translation method fused with sentiment analysis

    CN117077692A

  • Emotion understanding model training method and device, emotion understanding method and equipment

    CN119474389A

  • Translation precision optimization method and system based on artificial intelligence

    CN119849514A