A corpus-based approach to the standardization of touristic translation commentary

By analyzing the overall sentiment vector and semantic relevance based on a corpus, the final English translation of tourism interpretation is determined, which solves the problem of poor expression of emotion and context in existing technologies and achieves a more accurate translation effect.

CN120337950BActive Publication Date: 2025-12-30GUIZHOU BUSINESS SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510824891.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-12-30
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively convey the emotions and context of the original text in tourism translation and interpretation, resulting in poor translation quality.

Method used

Based on the corpus, the overall sentiment vector and semantic relevance of the original text of the tourism translation are determined. By integrating the sentiment relevance and semantic relevance, the standardization of the candidate English translations is obtained, and the final English translation is selected.

Benefits of technology

It improves the translation quality of tourism interpretation, making the English translation more relevant and accurate to the emotions and context of the original text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337950B_ABST
    Figure CN120337950B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of machine-aided translation, and particularly relates to a tourism translation and explanation standardization translation method based on a corpus, which comprises the following steps: determining a tourism translation and explanation original text's candidate English translation based on the corpus; determining the tourism translation and explanation original text's overall sentiment vector, and obtaining the sentiment correlation between the tourism translation and explanation original text and the candidate English translation based on the overall sentiment vector; fusing the sentiment correlation and the semantic correlation between the tourism translation and explanation original text and the candidate English translation to obtain the translation text standardization degree of the candidate English translation; and determining a final English translation from the candidate English translation according to the translation text standardization degree. The present application considers the sentiment factor, so that the final English translation can be accurately obtained from multiple candidate English translations, the final English translation is more relevant and accurate to the tourism translation and explanation original text, the sentiment of the tourism translation and explanation original text can be well expressed, and the translation effect of the tourism translation and explanation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine-aided translation technology, specifically to a corpus-based standardized translation method for tourism interpretation. Background Technology

[0002] English-Chinese translation is becoming increasingly common in tourism interpretation. Among the key aspects, the ability to clearly and accurately convey the meaning of the Chinese interpretation is crucial. Currently, artificial intelligence is often used to assist human translation. When translating original tourism interpretation texts, existing technologies often rely on translation engines to provide results. However, most translation engines can only achieve a rigid literal translation, merely translating based on semantic meaning and recommending vocabulary. In translating emotionally charged interpretation texts, literal translations often fail to effectively express the original emotions and context, resulting in poor translation quality. Summary of the Invention

[0003] To address the technical problem of unsatisfactory translation quality in existing tourism interpretation methods, the present invention aims to provide a corpus-based standardized translation method for tourism interpretation. The specific technical solution adopted is as follows:

[0004] This invention provides a corpus-based standardized translation method for tourism interpretation, comprising:

[0005] Based on a corpus, candidate English translations for the original tourist narration text are determined, including an initial English translation and at least one candidate English translation.

[0006] Determine the overall sentiment vector of the original tourism translation text, and use it to obtain the sentiment correlation between the original tourism translation text and the candidate English translations;

[0007] By integrating the aforementioned emotional relevance and the semantic relevance between the original tourism translation and the candidate English translation, the standardization of the candidate English translation is obtained.

[0008] Based on the standardization of the translation, the final English translation is determined from the candidate English translations.

[0009] In an exemplary embodiment, the process of obtaining the overall sentiment vector includes:

[0010] Based on the number of emotional intensities in the same dimension of the emotional vector of each word in the original text of the tourism translation and explanation, a reference index for the emotional intensity of each dimension is determined.

[0011] Based on the numerical values ​​of the sentiment intensity of each dimension in the sentiment vector of each original word, the sentiment intensity feature of the same dimension for each original word is obtained.

[0012] By integrating the sentiment intensity features of all original words in the same dimension, a comprehensive sentiment intensity feature across all dimensions is obtained.

[0013] By integrating the emotional intensity reference indicators from various dimensions and the comprehensive emotional intensity features, the emotional score vector intensity of each dimension is obtained.

[0014] The overall sentiment vector is obtained by sorting the intensity of the sentiment vectors in each dimension.

[0015] In an exemplary embodiment, the process of obtaining the emotional intensity reference index includes: normalizing the number of emotional intensities in the same dimension, and obtaining the result as the emotional intensity reference index for that dimension.

[0016] In an exemplary embodiment, the process of obtaining the emotional intensity feature includes: determining the emotional intensity weight of each dimension in the emotional vector of each original word, and weighting the emotional intensity value according to the emotional intensity weight to obtain the emotional intensity feature.

[0017] In an exemplary embodiment, the process of obtaining the emotional intensity weight includes:

[0018] Obtain the sentiment vector similarity of every two adjacent words in the original text of the tourism translation and explanation to obtain a similarity sequence;

[0019] Based on the similarity sequence, the original text of the tourism translation and explanation is segmented to obtain multiple original text vocabulary segments;

[0020] Based on the number of words in the original text segment and the mean similarity of the sentiment vectors corresponding to the original text segment, the sentiment intensity weights of each dimension in the sentiment vector of each word in the original text segment are obtained; the sentiment intensity weights are both proportional to the number of words in the original text and the mean similarity of the sentiment vectors.

[0021] In an exemplary embodiment, the original text of the tourism translation is segmented based on the similarity sequence to obtain multiple vocabulary segments, including:

[0022] Starting from the first sentiment vector similarity in the similarity sequence, when a sentiment vector similarity is less than a preset similarity threshold, the position between the two original words corresponding to that sentiment vector similarity is used as a segmentation point to segment the original text of the tourism translation and explanation, until the last sentiment vector similarity in the similarity sequence is reached, thereby obtaining multiple original word segments.

[0023] In one exemplary embodiment, the process of obtaining the emotional relevance includes:

[0024] The sentiment similarity between the overall sentiment vector of the original tourism translation and the sentiment vector of each candidate translation word in the target candidate English translation is obtained, thus obtaining the sentiment relevance corresponding to each candidate translation word in the target candidate English translation; the target candidate English translation is any one of the candidate English translations.

[0025] The process of obtaining the semantic relevance includes:

[0026] Obtain the semantic similarity between the semantic vectors of each word in the original text and the semantic vectors of the corresponding words in the target English translation, and obtain the semantic relevance between them and the corresponding words in the target English translation.

[0027] In an exemplary embodiment, the process of obtaining the translation standardization includes:

[0028] By integrating the sentiment relevance and semantic relevance of the same candidate translation words in the target English translation, the lexical relevance of the candidate translation words is obtained;

[0029] Based on the lexical relevance of all candidate translation terms in the target English translation, the standardization of the target English translation is obtained.

[0030] In an exemplary embodiment, obtaining the standardization of the target English translation based on the lexical relevance of all candidate translation terms in the target candidate English translation includes:

[0031] By integrating the lexical relevance of all candidate translation terms in the target English translation, the comprehensive relevance of the target English translation is obtained;

[0032] By comparing the part of speech of each candidate translation word in the target candidate English translation with the corresponding words in the original text, the number of target candidate translation words in the target candidate English translation is determined; the number of target candidate translation words is the number of candidate translation words whose part of speech has not changed.

[0033] Based on the comprehensive relevance and the number of target candidate translation words, the standardization degree of the target candidate English translation is obtained; the standardization degree of the translation is directly proportional to the comprehensive relevance and inversely proportional to the number of target candidate translation words.

[0034] In an exemplary embodiment, the process of obtaining the candidate English translation includes:

[0035] Based on the corpus, at least one candidate translation word is determined for each original word in the original tourism translation and explanation text. Then, one candidate translation word for each original word is combined into a sentence to obtain a candidate English translation.

[0036] This invention has the following beneficial effects: In addition to the initial English translation corresponding to the original tourism interpretation text, this invention also obtains at least one candidate English translation, which together constitute the candidate English translations. By obtaining the overall sentiment vector of the original tourism interpretation text, the sentiment relevance between the original tourism interpretation text and each candidate English translation is obtained. At the same time, the semantic relevance between the original tourism interpretation text and each candidate English translation is also obtained. By analyzing both sentiment and semantic relevance, the translation standardization of each candidate English translation is obtained. Finally, based on the translation standardization, the final English translation is determined from the candidate English translations. Compared with the existing method of determining the translation result based only on semantics, this invention considers both semantic and sentiment factors, accurately obtains the final English translation from multiple candidate English translations, and the obtained final English translation is more relevant and accurate to the original tourism interpretation text, and can better express the sentiment of the original tourism interpretation text, thus improving the translation effect of tourism interpretation. Attached Figure Description

[0037] Figure 1 This is a flowchart of a corpus-based standardized translation method for tourism interpretation provided in one embodiment of the present invention;

[0038] Figure 2 This is a flowchart of the process for obtaining the overall sentiment vector provided in one embodiment of the present invention;

[0039] Figure 3 This is a flowchart illustrating the process of obtaining emotional intensity weights according to an embodiment of the present invention;

[0040] Figure 4 This is a flowchart illustrating the process of obtaining translation standardization according to an embodiment of the present invention;

[0041] Figure 5 This is a flowchart illustrating a specific implementation of step S3-2 provided in one embodiment of the present invention. Detailed Implementation

[0042] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. All data and information collected in this application have been obtained with full consent.

[0044] This embodiment provides a corpus-based standardized translation method for tourism interpretation, which can be applied to Chinese-English translation devices for use by tourists whose native language is English.

[0045] The corpus used in this embodiment needs to include a large number of Chinese and English words, and for each Chinese word, at least one corresponding English word (i.e., at least one English word with the same meaning as the Chinese word) can be obtained to achieve the translation effect. For any single word from the large number of Chinese and English words in the corpus, each word contains two word vectors, denoted as semantic vectors. and emotional vectors The semantic vector of each word This can be obtained using existing model algorithms, such as the Word2Vec model. The sentiment vector for each word. All have the same number of dimensions (denoted as ). (), is a multidimensional vector, where the values ​​of different dimensions represent the intensity of emotion in different emotional directions. In this embodiment, the value range of each dimension is limited to... For ease of calculation, the larger the value of the dimension, the stronger the emotional intensity in the corresponding sentiment direction. For example, the sentiment vector of a certain word is... It should be understood that a dimension with a value of 0 indicates that there is no emotional intensity in that dimension, while a dimension with a value greater than 0 indicates that there is emotional intensity in that dimension. Based on the above requirements, a corpus is constructed.

[0046] like Figure 1 As shown in this embodiment, a corpus-based standardized translation method for tourism interpretation includes:

[0047] Step S1: Based on the corpus, determine the candidate English translations for the original tourism translation text. The candidate English translations include the initial English translation and at least one alternative English translation.

[0048] Step S2: Determine the overall sentiment vector of the original tourism translation and explanation text, and use it to obtain the sentiment relevance between the original tourism translation and explanation text and the candidate English translation;

[0049] Step S3: Integrate the emotional relevance and the semantic relevance between the original tourism translation and the candidate English translation to obtain the standardization of the candidate English translation;

[0050] Step S4: Determine the final English translation from the candidate English translations based on the standardization of the translation.

[0051] The specific implementation process of each step is described below with reference to the accompanying drawings.

[0052] Step S1: Based on the corpus, determine the candidate English translations of the original text for tourism translation and interpretation. The candidate English translations include the initial English translation and at least one alternative English translation.

[0053] The original text for tourism translation and interpretation is the Chinese original text that needs to be translated, such as the Chinese interpretation of tourist attractions expressed by a tour guide. The purpose of this invention is to determine the final English translation of the original text for tourism translation and interpretation.

[0054] In this embodiment, an existing translation engine is adopted, and this translation engine is connected to the corpus. The original text for tourism translation and interpretation is input into the translation engine, and the translation engine outputs an English translation. It should be understood that this English translation directly output by the translation engine is used as the initial English translation.

[0055] Due to the different expression habits of Chinese and English, during the translation process, some changes in word forms may occur. For example, translating "我们需要解决这个问题" into English gives "We need a solution to this problem." Here, the Chinese verb "解决" is translated as the noun "solution" during the translation process. Although the word form changes, the semantic meaning expressed remains the same.

[0056] For the sake of convenience in explanation, each word in the original text for tourism translation and interpretation is defined as each original word. In order to standardize the translation, first, it is necessary to match the semantics of each translated word (i.e., English word) with that of the original word one by one to obtain semantic matching pairs for subsequent change standardization. Due to the phenomenon of word form changes in the above translation process, matching cannot be based solely on word forms, and the degree of semantic consistency between words also needs to be considered to match the original word and the translated word.

[0057] Take a certain original sentence and its translation as an example to match their synonymous words. Specifically: First, each original word needs to be segmented. Since there is no clear space symbol in Chinese to separate words, each character needs to be analyzed one by one. Store the original sentence as a string denoted as string starting from the first character of the string and denote the temporarily set and initially empty string as Add the character to and search in the corpus to see if there is a string that matches it (if there is, denote it as ) or a string of which it is a substring (if there is, denote it as ). ).

[0058] There is an inclusion relationship between some words, such as "肝胆" and "肝胆相照". After identifying the matching words, it is still necessary to further determine whether it is included in another word. For the above according to Strings retrieved from the corpus and This is used to determine and classify vocabulary.

[0059] If only If it exists, it indicates the string of words in the original text. It's not a complete word; we need to add the next character from the original sentence. And perform the search again.

[0060] like and All exist, indicating strings representing words from the original text. It is a complete word, but it may still be part of another word, so we need to add more characters to make a search.

[0061] If only If it exists, it means If the term is a complete word and not contained within any other word, the search will stop. The word will be segmented and... To set an empty string, append the character following the word in the original text. Proceed to the next character segment.

[0062] like and If none of them exist, it indicates that the previously retrieved term was a complete term, and that combining it with subsequent characters in the original text could not yield any new terms. Therefore, the previously retrieved term is divided and left blank. Then proceed with the next character segmentation.

[0063] The original text is divided into words using the character segmentation method described above. The number of words in the original text after segmentation is denoted as . .

[0064] In one exemplary embodiment, the obtained first Taking a single word from the original text as an example, let the semantic vector of that word be denoted as . Obtain the semantic vector of each translated word in the corpus, and obtain the first... The cosine similarity between the semantic vectors of each original word and each translated word in the corpus is taken, and the highest cosine similarity is denoted as . The translation word corresponding to the highest cosine similarity is taken as the first... The initial translation of each original word is obtained by combining the initial translations of each original word into sentences (while avoiding grammatical errors). The resulting sentences are the initial English translations of the original tourist narration.

[0065] Then, for the first Given a set of original words, a cosine similarity threshold is set for each word. ( The semantic selection threshold can be customized by the user. The settings need to ensure that each source word has at least one candidate translation word. This embodiment sets... ), get the The cosine similarity between the semantic vectors of each original word and each translated word in the corpus is calculated, and the cosine similarity is greater than a certain threshold. At least one translation word is obtained, and all other translation words besides the initial translation word are defined as candidate translation words, thus obtaining the first... For each original word, at least one candidate translation word is obtained.

[0066] In principle, selecting one candidate translation word from multiple candidate translation words for each original word and combining them into sentences yields one candidate English translation, thus resulting in multiple candidate English translations. The number of candidate English translations obtained through permutation and combination is the product of the number of candidate translation words for each original word. For example, if there are 4 original words, and the number of candidate translation words for these four words are 2, 3, 3, and 4 respectively, then the total number of sentence combinations that can be obtained through permutation and combination is 2 × 3 × 3 × 4 = 72. However, some sentences obtained through permutation and combination may not meet grammatical requirements, i.e., some sentences contain grammatical errors. These grammatically incorrect combinations are then deleted, and the grammatically correct combinations are retained, thus obtaining at least one candidate English translation for the original tourism translation text. The initial English translation of the original tourism translation text and all candidate English translations are collectively referred to as the candidate English translations for the original tourism translation text. This invention selects one English translation from the candidate English translations as the final English translation of the original tourism translation text.

[0067] Step S2: Determine the overall sentiment vector of the original tourism translation and explanation text, and use it to obtain the sentiment relevance between the original tourism translation and explanation text and the candidate English translation.

[0068] The original text for tourism translation may contain a large number of emotionally charged words. A literal translation may not fully capture the emotional expression of the original text, resulting in a significantly diminished translation quality. Since different words have their own emotional intensity, by combining the context of the original tourism translation text, we can infer the overall emotional tone emphasized by each word, thereby standardizing the vocabulary in the English translation.

[0069] Based on the sentiment vectors of individual words in the original tourism translation text, the overall sentiment vector of the original tourism translation text is obtained. In an exemplary embodiment, such as... Figure 2 As shown below, a specific process for obtaining the overall sentiment vector is given:

[0070] Step S2-1: Based on the number of emotional intensities in the same dimension of the emotional vector of each word in the original text of the tourism translation, determine the reference index of emotional intensity for each dimension.

[0071] For any given dimension, we obtain the number of words with emotional intensity in the sentiment vector of each original word in the travel translation text under that dimension. As mentioned above, if the sentiment intensity of a dimension is 0, it means that there is no emotional intensity in that dimension; if the sentiment intensity of a dimension is greater than 0, it means that there is emotional intensity in that dimension. Based on this, we determine whether there is emotional intensity in the sentiment vector of each original word in that dimension. Then, based on the sentiment vectors of all original words, we obtain the number of original words with emotional intensity in that dimension, thus obtaining the number of words with emotional intensity corresponding to that dimension. Finally, we normalize the number of words with emotional intensity in that dimension, and the result is the reference index for the emotional intensity of that dimension.

[0072] In one exemplary embodiment, a normalization method is given as follows:

[0073] ;

[0074] in, This represents the reference index for the intensity of emotion in the i-th dimension. This represents the number of emotional intensities present in the i-th dimension. This represents the number of emotional intensities present in the b-th dimension. This represents the number of dimensions in the sentiment vector.

[0075] Step S2-2: Based on the emotional intensity values ​​of each dimension in the emotional vector of each original word, obtain the emotional intensity feature of the same dimension for each original word.

[0076] The numerical values ​​of each dimension in the sentiment vector of each original word represent the sentiment intensity of that dimension; the larger the value, the greater the sentiment intensity. Based on the sentiment intensity values ​​of each dimension in the sentiment vector of each original word, a sentiment intensity feature for the same dimension of each original word is obtained. In an exemplary embodiment, the sentiment intensity weights of each dimension in the sentiment vector of each original word are determined, and the sentiment intensity values ​​are weighted according to these weights to obtain the sentiment intensity feature.

[0077] In one exemplary embodiment, such as Figure 3As shown below, a specific process for obtaining the emotional intensity weight is given:

[0078] Step S2-2-1: Obtain the sentiment vector similarity of every two adjacent words in the original text of the tourism translation and explanation, and obtain the similarity sequence.

[0079] In the original text of tourism translation and explanation, the sentiment vector similarity of every two adjacent words in the original text is obtained, with the first... Taking the first original word as an example, obtain the first... The sentiment vector of the first original word and the second The sentiment vector similarity of each original word's sentiment vector, where... . This represents the number of words in the original text. The higher the sentiment vector similarity, the higher the sentiment consistency between two adjacent words in the original text.

[0080] In this embodiment, the sentiment vector similarity is specifically cosine similarity, that is, obtaining the first... The sentiment vector of the first original word and the second The cosine similarity of the sentiment vectors of each original word. Since the numerical range of cosine similarity is -1 to 1, to ensure that the cosine similarity is within the range of 0-1 for easy data processing, this embodiment normalizes the cosine similarity, and all cosine similarities involved in this embodiment are normalized cosine similarities. The normalization method for cosine similarity can be: (cosine similarity + 1) / 2.

[0081] The similarity of each sentiment vector is arranged according to the order of the words in the original text to obtain a similarity sequence.

[0082] Step S2-2-2: Based on the similarity sequence, the original text of the tourism translation is segmented to obtain multiple original text vocabulary segments.

[0083] The similarity sequence represents the similarity of the sentiment vectors of every two adjacent words in the original text. Based on the similarity sequence, the original text of the tourism translation is segmented to obtain multiple word segments. This results in a higher similarity of the sentiment vectors of words within the same word segment and a lower similarity of the sentiment vectors of words between different word segments.

[0084] In an exemplary embodiment, a preset similarity threshold is used to determine whether the similarity of the emotion vectors is high. The preset similarity threshold has a value range of 0-1, and its specific value is set according to actual needs, such as 0.7.

[0085] Starting with the first sentiment vector similarity in the similarity sequence, each is compared with a preset similarity threshold. Specifically: if the first sentiment vector similarity is greater than or equal to the preset similarity threshold, the second sentiment vector similarity is compared with the preset similarity threshold. If the second sentiment vector similarity is greater than or equal to the preset similarity threshold, the third sentiment vector similarity is compared with the preset similarity threshold, and so on, until a sentiment vector similarity is less than the preset similarity threshold. Then, the position between the two original words corresponding to the sentiment vector similarity less than the preset similarity threshold is used as the segment point to segment the original text of the tourism translation and explanation. The original text of the tourism translation and explanation is divided into a segment from the first original word to the first original word of the two original words. Then, based on the remaining original words, the sentiment vector similarity is compared with a preset similarity threshold. This process is repeated until a sentiment vector similarity is less than the preset similarity threshold. The position between the two original words corresponding to the sentiment vector similarity less than the preset similarity threshold is used as a segmentation point to segment the remaining original words. The segment from the first original word in the remaining original words to the preceding original word between these two words is considered as one original word segment. This process is repeated until the last sentiment vector similarity in the similarity sequence is reached. This completes the segmentation of the original text for tourism translation and explanation, resulting in multiple original word segments. It should be understood that in this embodiment, there is a possibility that a single original word may be considered as one original word segment.

[0086] Here's a specific example: Assuming there are 10 words in the original text, the number of sentiment vector similarities in the similarity sequence is 9. Starting with the first sentiment vector similarity, it is compared with a preset similarity threshold. If the third sentiment vector similarity is less than the preset similarity threshold, since the two original words corresponding to the third sentiment vector similarity are the third and fourth original words, then the first to the third original words are considered as one original word segment. Then, starting with the fourth original word and its sentiment vector similarity, it is compared with the preset similarity threshold. If the fifth sentiment vector similarity is less than the preset similarity threshold, since the two original words corresponding to the fifth sentiment vector similarity are the fifth and sixth original words, then the fourth and fifth original words are considered as one original word segment. Then, starting from the 6th original word and the 6th sentiment vector similarity, it is compared with the preset similarity threshold until the last sentiment vector similarity is greater than or equal to the preset similarity threshold. Then, all the remaining original words are considered as one original word segment.

[0087] Step S2-2-3: Based on the number of words in the original text segment and the mean similarity of the sentiment vectors corresponding to the original text segment, obtain the sentiment intensity weights of each dimension in the sentiment vector of each word in the original text segment.

[0088] For any given segment of source text, the sentiment of each word in that segment is consistent and continuous. Furthermore, the more words in that segment, the higher the sentiment intensity weight of each dimension in the sentiment vector of each word in that segment.

[0089] The mean sentiment vector similarity is calculated for each sentiment vector corresponding to the original text segment. A higher mean sentiment vector similarity indicates a higher sentiment intensity weight for each dimension of the sentiment vector for each word in the original text segment. Therefore, the sentiment intensity weight is directly proportional to both the number of original text words and the mean sentiment vector similarity. A higher sentiment intensity weight indicates a higher degree of continuity and concentration of sentiment words in the original text segment.

[0090] It should be understood that the sentiment vector similarity between two original words only belongs to the sentiment vector similarity of the entire original word segment if both words belong to it. Based on the example above, for an original word segment consisting of the first to third original words, the sentiment vector similarity of that segment includes: the sentiment vector similarity between the first and second original words, and the sentiment vector similarity between the second and third original words; for an original word segment consisting of the fourth and fifth original words, the sentiment vector similarity is the sentiment vector similarity between the fourth and fifth original words. Therefore, for an original word segment containing only one original word, the mean sentiment vector similarity of this type of segment is set as the average of the mean sentiment vector similarity of all original word segments containing at least two original words.

[0091] In one exemplary embodiment, a specific method for quantifying the emotional intensity weight is given below:

[0092] ;

[0093] in, Indicates the first The sentiment intensity weight of each dimension in the sentiment vector of each word in the original word segment. Indicates the first The number of original words in the original word segment where each original word is located. Indicates the first The mean similarity of the sentiment vectors of the original words in the original text.

[0094] Represents the normalization function. Indicates to Normalization. In this embodiment, the normalization method here can be: obtaining the normalization corresponding to each original word. The maximum and minimum values ​​in the range are then normalized using the maximum and minimum values ​​method to calculate the value of the first value. The original words correspond to Normalize.

[0095] As can be seen from the above calculation method, for the th For each word in the original text, the sentiment intensity weights of all dimensions in the sentiment vector of all words belonging to that original text segment are equal to 0. .So, It can also represent the first The sentiment intensity weights of each dimension in the sentiment vector of each original word.

[0096] Therefore, the sentiment intensity weights of all dimensions in the sentiment vector of the same original word are the same, and the sentiment intensity weights of all dimensions in the sentiment vectors of each original word belonging to the same original word segment are also the same.

[0097] So, the first The sentiment intensity feature of the i-th dimension in the sentiment vector of each original word is equal to: .in, For the first The numerical value of the sentiment intensity of the i-th dimension in the sentiment vector of each original word.

[0098] Step S2-3: Integrate the sentiment intensity features of all original words in the same dimension to obtain the comprehensive sentiment intensity features of each dimension.

[0099] The average of the sentiment intensity features of all words in the original text in the same dimension is calculated to obtain the comprehensive sentiment intensity feature for that dimension. The calculation formula is as follows:

[0100] ;

[0101] in, is the comprehensive emotional intensity feature of the i-th dimension.

[0102] Step S2-4: Integrate the emotional intensity reference indicators and comprehensive emotional intensity features from various dimensions to obtain the emotional score vector intensity of each dimension.

[0103] For any dimension, the product of the emotional intensity reference index for that dimension and the comprehensive emotional intensity feature for that dimension is calculated as the emotional subvector intensity for that dimension. The calculation formula is as follows:

[0104] ;

[0105] in, Let represent the intensity of the sentiment vector in the i-th dimension. In this way, the intensity of the sentiment vector in each dimension can be obtained. This represents the reference index for the emotional intensity of the i-th dimension.

[0106] Step S2-5: Sort the intensity of the sentiment vectors in each dimension to obtain the overall sentiment vector.

[0107] The intensity of the sentiment vectors for each dimension is sorted according to their order of importance to obtain the overall sentiment vector. ,for .in, This represents the intensity of the sentiment vector in the first dimension. This represents the intensity of the sentiment vector in the second dimension. This represents the intensity of the sentiment vector in the last dimension.

[0108] Each original word has its own unique emotional connotation. Moreover, based on the various dimensions of the emotional vector of the original words, the same original word can have multiple emotional expressions. In a sentence composed of multiple original words, the sentence as a whole, the multiple original words together constitute the overall word distribution trend of the sentence, that is, the overall emotional vector.

[0109] The closer the vocabulary selected in the translation is to the overall sentiment vector of the original text, the higher its sentiment consistency, and the more suitable it is for translation and interpretation. Therefore, after obtaining the overall sentiment vector of the original text for tourism translation and interpretation, we can obtain the sentiment relevance between the original text for tourism translation and interpretation and each candidate English translation based on the overall sentiment vector of the original text for tourism translation and interpretation.

[0110] In one exemplary embodiment, for ease of explanation, the target candidate English translation is set to any one of the candidate English translations, thereby obtaining the sentiment vector of each candidate translation word in the target candidate English translation. Each candidate translation word in the target candidate English translation corresponds one-to-one with each word in the original text of the tourism translation and explanation.

[0111] Then, for any candidate word in the target English translation pool, the sentiment similarity between the overall sentiment vector of the original tourism translation text and the sentiment vector of the candidate word is obtained. Specifically, this sentiment similarity is the cosine similarity, i.e., the cosine similarity between the overall sentiment vector of the original tourism translation text and the sentiment vector of the candidate word. The obtained cosine similarity is used as the sentiment relevance between the original tourism translation text and the candidate word. In this way, the sentiment relevance between the original tourism translation text and each candidate word in the target English translation pool is obtained.

[0112] For translation work, semantic consistency remains the most fundamental requirement. Therefore, this embodiment also obtains the semantic relevance between the original text of the tourism translation and the various candidate translation terms in the target English translation. In an exemplary embodiment, for any candidate translation term in the target English translation, the corresponding original text term is determined. Then, the semantic similarity between the semantic vector of the original text term and the semantic vector of the candidate translation term is obtained. This semantic similarity is specifically cosine similarity, and the obtained cosine similarity is used as the semantic relevance between the original text term and the candidate translation term. In this way, the semantic relevance between each original text term in the tourism translation and the corresponding candidate translation terms in the target English translation is obtained.

[0113] Step S3: Integrate the emotional relevance and the semantic relevance between the original tourism translation and the candidate English translation to obtain the standardization of the candidate English translation.

[0114] According to step S2, the sentiment relevance between the original tourism translation and the candidate translation words in the target English translation is obtained, as well as the semantic relevance between the original tourism translation and the corresponding candidate translation words in the target English translation. Then, these two relevances are fused to obtain the standardization of the target English translation. In an exemplary embodiment, such as... Figure 4 As shown, a specific process for obtaining the standardization of a translation is presented:

[0115] Step S3-1: Integrate the sentiment relevance and semantic relevance of the same candidate translation words in the target candidate English translation to obtain the lexical relevance of the candidate translation words.

[0116] For any candidate word in the target English translation pool, obtain its sentiment relevance and semantic relevance, calculate the average of the sentiment relevance and semantic relevance, and obtain the lexical relevance of that candidate word. Repeat this process to obtain the lexical relevance of each candidate word in the target English translation pool.

[0117] Step S3-2: Based on the lexical relevance of all candidate translation words in the target candidate English translation, obtain the standardization of the target candidate English translation.

[0118] By integrating the lexical relevance of all candidate translation terms in the target English translation, the standardization degree of the target English translation can be obtained. In an exemplary embodiment, the average value of the lexical relevance of all candidate translation terms in the target English translation can be calculated, and this average value can be directly used as the standardization degree of the target English translation.

[0119] As another implementation method, in order to simultaneously consider the impact of the number of candidate words whose parts of speech have not changed on the standardization of the translation, such as... Figure 5 As shown, the following is another implementation process for achieving the standardization of the target candidate English translation:

[0120] Step S3-2-1: Integrate the lexical relevance of all candidate translation words in the target candidate English translation to obtain the comprehensive relevance of the target candidate English translation.

[0121] The lexical relevance of all candidate translation terms in the target English translation is integrated. Specifically, the average lexical relevance of all candidate translation terms in the target English translation is calculated as the comprehensive relevance of the target English translation.

[0122] Step S3-2-2: Compare the parts of speech of each candidate translation word in the target candidate English translation with the corresponding words in the original text to determine the number of target candidate translation words in the target candidate English translation.

[0123] Since there is a one-to-one correspondence between the source text vocabulary and the candidate translation vocabulary in the target English translation, the number of candidate translation vocabulary whose parts of speech remain unchanged after translating from Chinese to English is defined as the target candidate translation vocabulary number. In other words, the target candidate translation vocabulary number is the number of candidate translation vocabulary whose parts of speech remain unchanged. A higher number of candidate translation vocabulary whose parts of speech remain unchanged indicates a more literal translation, making it less convenient for English-speaking readers to understand, and resulting in a lower standardization of the target English translation. In other words, the standardization of the translation is inversely proportional to the number of target candidate translation vocabulary.

[0124] Moreover, the stronger the overall relevance of the target English translation, the more suitable the target English translation is for translating the original text, and the higher the standardization of the target English translation. That is, the standardization of the translation is directly proportional to the overall relevance.

[0125] Step S3-2-3: Based on the comprehensive relevance and the number of target candidate translation words, obtain the standardization of the target candidate English translation.

[0126] The standardization degree of the target English translation is obtained based on the comprehensive relevance of the target English translation candidate and the number of target English translation candidate words. In an exemplary embodiment, the following quantification method is given:

[0127] ;

[0128] in, This indicates the standardization of the x-th candidate English translation. This indicates the overall relevance of the x-th candidate English translation. This represents the number of target vocabulary words for the x-th candidate English translation. This represents an exponential function with base e. Indicates to The negative correlation normalization.

[0129] Step S4: Determine the final English translation from the candidate English translations based on the standardization of the translation.

[0130] Step S3 is used to obtain the standardization score of each candidate English translation. The higher the standardization score, the more suitable it is for use as the original text of the tourism interpretation. Therefore, the highest standardization score is determined from the candidate English translations, and the corresponding candidate English translation is selected. This candidate English translation with the highest standardization score is then used as the final English translation of the tourism interpretation and output.

[0131] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A corpus-based tourism translation and interpretation standardization translation method, characterized in that, The method comprises the following steps: Based on the corpus, determine the candidate English translation of the original text of the tourism translation commentary, which includes the initial English translation and at least one candidate English translation; Determine the overall sentiment vector of the original text of the tourism translation commentary, and obtain the sentiment relevance of the original text of the tourism translation commentary and the candidate English translation; The overall sentiment vector is obtained by sorting the sentiment sub-vector strength of each dimension, the sentiment sub-vector strength is obtained by fusing the sentiment strength reference index of each dimension and the comprehensive sentiment strength feature, the comprehensive sentiment strength feature is obtained by fusing the sentiment strength features of all original text words in the same dimension, and the sentiment strength reference index is determined based on the number of sentiment strength in the same dimension in the sentiment vector of each original text word in the original text of the tourism translation commentary; The sentiment strength feature is obtained by weighting the numerical value of the sentiment strength according to the sentiment strength weight of each dimension in the sentiment vector of each original text word; The process of obtaining the sentiment strength weight comprises: obtaining the similarity of the sentiment vectors of every two adjacent original text words in the original text of the tourism translation commentary to obtain a similarity sequence; Based on the similarity sequence, the original text of the tourism translation commentary is segmented to obtain a plurality of original text word segments; According to the number of original text words in the original text word segment and the average value of the sentiment vector similarity corresponding to the original text word segment, the sentiment strength weight is obtained; Fusing the sentiment relevance and the semantic relevance between the original text of the tourism translation commentary and the candidate English translation, the translation standard of the candidate English translation is obtained; According to the translation standard, the final English translation is determined from the candidate English translation.

2. The corpus-based tour translation and explanation normalization translation method according to claim 1, wherein, The process of obtaining the sentiment strength reference index comprises: normalizing the number of sentiment strength in the same dimension to obtain the sentiment strength reference index of the dimension.

3. The corpus-based tour translation and explanation standardization translation method according to claim 1, wherein, The sentiment strength weight is directly proportional to the number of original text words and the average value of the sentiment vector similarity.

4. The corpus-based tour translation and explanation normalization translation method according to claim 1, wherein, Based on the similarity sequence, the original text of the tourism translation commentary is segmented to obtain a plurality of original text word segments, which comprises: Starting from the first sentiment vector similarity in the similarity sequence, when the sentiment vector similarity is less than the preset similarity threshold, the position between the two original text words corresponding to the sentiment vector similarity is taken as the segmentation point, and the original text of the tourism translation commentary is segmented, until the last sentiment vector similarity in the similarity sequence is traversed, thereby obtaining a plurality of original text word segments.

5. The corpus-based tour translation and explanation normalization translation method according to claim 1, wherein, The process of obtaining the sentiment relevance comprises: Obtain the sentiment similarity between the overall sentiment vector of the original text of the tourism translation commentary and the sentiment vector of each candidate translation word in the target candidate English translation to obtain the sentiment relevance corresponding to each candidate translation word in the target candidate English translation; The target candidate English translation is any one of the candidate English translations; The process of obtaining the semantic relevance comprises: Obtain the semantic similarity between the semantic vector of each original text word and the semantic vector of the corresponding candidate translation word in the target candidate English translation to obtain the semantic relevance corresponding to each candidate translation word in the target candidate English translation.

6. The corpus-based tour translation and explanation normalization translation method according to claim 5, wherein, The process of obtaining the translation standard comprises: fuse the sentiment relevance and the semantic relevance of the same candidate translation vocabulary in the target candidate English translation to obtain vocabulary relevance of the candidate translation vocabulary; obtain translation standard degree of the target candidate English translation based on the vocabulary relevance of all candidate translation vocabularies in the target candidate English translation.

7. The corpus-based standardized translation method for tourism interpretation as described in claim 6, characterized in that, The obtaining of the translation standard degree of the target candidate English translation based on the vocabulary relevance of all candidate translation vocabularies in the target candidate English translation comprises: fuse the vocabulary relevance of all candidate translation vocabularies in the target candidate English translation to obtain comprehensive relevance of the target candidate English translation; compare the parts of speech of each candidate translation vocabulary in the target candidate English translation with the corresponding original text vocabulary to determine the number of target candidate translation vocabularies in the target candidate English translation; the number of target candidate translation vocabularies is the number of candidate translation vocabularies whose parts of speech have not changed; obtain the translation standard degree of the target candidate English translation according to the comprehensive relevance and the number of target candidate translation vocabularies; the translation standard degree is directly proportional to the comprehensive relevance and inversely proportional to the number of target candidate translation vocabularies.

8. The corpus-based tour translation and explanation normalization translation method according to claim 1, wherein, The obtaining process of the candidate English translation comprises: determine at least one candidate translation vocabulary for each original text vocabulary in the tourism translation commentary original text based on a corpus, and combine each candidate translation vocabulary of each original text vocabulary to obtain a candidate English translation.

Citation Information

Patent Citations

  • Automatic name translation system and method

    CN107861953A

  • Machine translation method fused with sentiment analysis

    CN117077692A