A Multilingual Translation Method and System

By identifying and processing culturally sensitive content, combining semantic unit splitting and semantic combination technology, the problem that traditional multilingual translation technology is difficult to convey cultural connotation is solved, and higher cultural adaptability and translation quality are achieved.

CN119647486BActive Publication Date: 2025-05-27MINGTAI (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510147221.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Traditional multilingual translation technology is difficult to accurately convey the cultural connotation and emotional colors in the source language, resulting in a lack of cultural adaptability in the translation results.

Method used

By identifying and processing culturally sensitive content, using semantic unit splitting and semantic coherence-based evaluation mechanisms, the target language expression is extracted, combined with the target language semantic knowledge base for semantics, generate translation results, and optimize the final results through automatic evaluation and manual review.

Benefits of technology

It improves the cultural adaptability and expression accuracy of the translation results, enhances the reliability and fluency of the translation results, and ensures that the translation results are not only accurate in grammar and word use, but also accurately conveys the cultural connotation and deep meaning of the source language.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119647486B_ABST
    Figure CN119647486B_ABST
Patent Text Reader

Abstract

A multilingual translation method and system, which relates to the field of electrical digital data processing. In this method, cultural feature information and semantic structure information are obtained; the source language text is split into several semantic units based on the semantic structure information; cultural connotation information with a correlation degree greater than a preset threshold with the semantic units is retrieved; the semantic units are determined as culturally sensitive semantic units; the target language expression is extracted based on the cultural connotation information; the target language expression is semantically combined with the standard translation results of other semantic units in the source language text except the culturally sensitive semantic units to obtain a translation result; it is judged whether the semantic coherence of the translation result meets a preset condition; if the preset condition is met, the translation result is output; if the preset condition is not met, the step of semantic combination is executed again. This application improves the accuracy of conveying the cultural connotations and emotional colors contained in the source language, thereby improving the cultural adaptability of the translation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic digital data processing, and particularly relates to a multi-language translation method and system. Background Art

[0002] With the continuous development of economic globalization, cross-language communication has become increasingly frequent, and multi-language translation technology plays an increasingly important role in people's production and life. The traditional manual translation method not only takes a lot of time and effort and is inefficient in processing large-scale texts, but also is prone to problems such as understanding deviations and inaccurate expressions, and has been difficult to meet the rapidly developing international needs today.

[0003] In related technologies, an encoder-decoder model can be established to encode the source language text into a vector representation and then decode it to generate the target language text, achieving significant improvements in both translation speed and accuracy. This technology uses an attention mechanism to dynamically focus on important information in the source text and uses deep learning algorithms to continuously optimize the parameters of the translation model, making the translation results more accurate and fluent.

[0004] However, when translating texts containing proverbs, sayings and other texts with profound cultural connotations, there is often a problem of literal translation and loss of the deep meaning of the original text, making it difficult to accurately convey the cultural connotations and emotional colors contained in the source language, resulting in a lack of cultural adaptability in the translation results. Summary of the Invention

[0005] This application provides a multi-language translation method and system for improving the accuracy of conveying the cultural connotations and emotional colors contained in the source language, thereby improving the cultural adaptability of the translation results.

[0006] In a first aspect, this application provides a multi-language translation method, which obtains the source language text to be translated and the target language type;

[0007] Obtain the cultural feature information corresponding to the source language text from a preset cultural knowledge base; obtain the semantic structure information of the source language text from a preset source language semantic knowledge base;

[0008] Based on the semantic structure information, split the source language text into several semantic units;

[0009] For each semantic unit, retrieve the cultural connotation information with a correlation degree greater than a preset threshold associated with the semantic unit in a preset cultural knowledge base;

[0010] When it is determined that the semantic unit contains a cultural feature identifier, determine that the semantic unit is a culture-sensitive semantic unit;

[0011] Extract from a preset cultural knowledge base a target language expression whose cultural connotation matching degree with the cultural connotation of the culture-sensitive semantic unit is greater than a preset matching degree based on cultural connotation information;

[0012] Invoke the target language semantic knowledge base, and semantically combine the target language expression with the standard translation results of other semantic units in the source language text except for the culture-sensitive semantic unit to obtain a translation result;

[0013] Judge whether the semantic coherence of the translation result meets the preset conditions according to the preset translation evaluation rules;

[0014] If the preset conditions are met, output the translation result;

[0015] If the preset conditions are not met, re-execute the step of semantically combining the target language expression with the standard translation results of other semantic units in the source language text except for the culture-sensitive semantic unit to obtain a translation result.

[0016] By adopting the above technical solution, through the identification and special processing of culture-sensitive content, the translation result can accurately convey the cultural connotation in the source language. By adopting the splitting method of semantic units and the evaluation mechanism based on semantic coherence, the overall coherence of the translation result is improved, and the translation quality and expression accuracy in cross-cultural communication are improved. When encountering a situation where the preset conditions of semantic coherence are not met, the system will search for a better solution by recombining, enhancing the reliability and expression fluency of the translation result. The accuracy of conveying the cultural connotation and emotional color contained in the source language is improved, and thus the cultural adaptability of the translation result is improved.

[0017] Combined with some embodiments of the first aspect, in some embodiments, when it is determined that the semantic unit contains a cultural feature identifier, it specifically includes:

[0018] Obtain a preset cultural feature identifier dictionary, where the preset cultural feature identifier dictionary includes multiple cultural scenarios and the feature words corresponding to each cultural scenario;

[0019] Perform word segmentation processing on the semantic unit to obtain a word sequence;

[0020] Match each word in the word sequence with the feature words in the cultural feature identifier dictionary;

[0021] When the number of matched feature words is greater than the preset number, determine the cultural scenario corresponding to the semantic unit according to the matched feature words;

[0022] Extract cultural feature rules from the cultural knowledge base based on the cultural scenario;

[0023] When it is determined that the semantic structure of the semantic unit meets the composition conditions of the cultural feature identifier according to the cultural feature rules, it is confirmed that the semantic unit contains the cultural feature identifier.

[0024] By adopting the above technical solutions, a corresponding relationship is established between the cultural scenario and the feature words, the semantic unit is segmented and matched with the feature words, the cultural scenario is determined based on the matching result and the corresponding cultural feature rules are extracted, and then it is judged whether the semantic unit contains the cultural feature identifier. Through the dual verification of the quantity threshold control of the feature words and the verification of the cultural feature rules, the accuracy and reliability of cultural feature recognition are improved. The dictionary-based management method of feature words facilitates the accumulation and update of knowledge, enables the system to handle different types of cultural scenarios, and improves the coverage and adaptability of cultural feature recognition. Through the verification of the semantic structure by the cultural feature rules, the misjudgment rate is reduced, and the pertinence and accuracy of subsequent translation processing are improved.

[0025] Combined with some embodiments of the first aspect, in some embodiments, the target language expression is semantically combined with the standard translation results of other semantic units in the source language text except for the culture-sensitive semantic units to obtain the translation result, which specifically includes:

[0026] Construct a target language syntactic tree template, which includes multiple semantic slots and the syntactic relationships between the semantic slots;

[0027] Obtain the syntactic structure information of the target language expression from the target language semantic knowledge base;

[0028] Fill the target language expression into the corresponding semantic slots according to the syntactic structure information;

[0029] Obtain a semantic association degree calculation model, which is used to calculate the semantic association strength between different semantic units;

[0030] Calculate the semantic association degree between the standard translation results of other semantic units and the content in the filled semantic slots based on the semantic association degree calculation model;

[0031] Sort the standard translation results of other semantic units according to the semantic association degree to obtain a sorting result;

[0032] Fill the standard translation results of other semantic units into the remaining semantic slots in descending order according to the sorting result;

[0033] Determine the conjunctions between the semantic slots according to the syntactic tree template;

[0034] Fill the conjunctions between the semantic slots and convert them into a linear text sequence to obtain the translation result.

[0035] By adopting the above technical solutions, a structured translation organization method based on syntactic trees is used, which clearly defines the syntactic relationships between semantic components, reducing the problem of word order chaos that may be caused by simple splicing. Through the quantitative calculation of semantic association degrees and the sorting mechanism based on association strengths, the rationality and coherence of each semantic unit in the combination process are improved. The application of syntactic tree templates makes the translation results have good syntactic structures, improving the standardization and readability of expressions. While maintaining the semantic integrity of the original text, the expression quality of the translation results is enhanced through a structured combination method.

[0036] In combination with some embodiments of the first aspect, in some embodiments, after re - executing the step of semantically combining the target - language expression with the standard translation results of other semantic units in the source - language text except for culture - sensitive semantic units to obtain the translation results, the method further includes:

[0037] Obtain a historical translation sample library, where the historical translation sample library contains multiple sample pairs of source - language texts and corresponding target - language translation texts;

[0038] Calculate the similarity between the source - language text and the source - language texts in the historical translation sample library based on a preset similarity calculation rule;

[0039] Screen out the sample pairs with similarity greater than a preset similarity threshold from the historical translation sample library;

[0040] Extract the semantic combination patterns of the target - language translation texts in the sample pairs;

[0041] Apply the semantic combination patterns to the translation results to generate alternative translation results;

[0042] Compare the semantic coherence scores of the alternative translation results and the translation results;

[0043] Select the one with a higher semantic coherence score as the optimized translation result.

[0044] By adopting the above technical solutions, the translation results are optimized using existing high - quality translation experience, improving translation efficiency. The sample screening mechanism based on the similarity threshold ensures the relevance and applicability of reference cases, reducing the negative impacts that may be brought by inappropriate references. Through the extraction and application of semantic combination patterns, successful translation experience is transformed into reusable knowledge, enhancing the learning ability and optimization effect of the translation system, and also enhancing the practicality and reliability of the system while improving translation quality.

[0045] In combination with some embodiments of the first aspect, in some embodiments, calculating the similarity between the source - language text and the source - language texts in the historical translation sample library based on a preset similarity calculation rule specifically includes:

[0046] Extract features from the source language text to generate a text feature vector;

[0047] Extract features from the source language texts in the historical translation sample library to generate sample feature vectors;

[0048] Calculate the Euclidean distances between the text feature vector and the sample feature vectors in each dimension;

[0049] Perform a weighted sum of the Euclidean distances in each dimension to obtain the similarity.

[0050] By adopting the above technical solution, extracting features from the source language text and the source language texts in the historical translation sample library to generate feature vectors, calculating the Euclidean distances between the two feature vectors in each dimension and performing a weighted sum to obtain the similarity, it is possible to accurately quantify the similarity degree between texts from a mathematical perspective, capture the feature differences of texts in multiple dimensions such as word frequency, syntactic structure, and semantic association, and has higher accuracy compared to simple text matching. By setting weight coefficients for different dimension features, this solution can highlight the influence of important features and weaken the interference of secondary features, making the similarity calculation result more in line with the actual language usage scenario.

[0051] Combined with some embodiments of the first aspect, in some embodiments, after selecting the one with a high semantic coherence score as the optimized translation result, the method further includes:

[0052] Evaluate the optimized translation result according to a preset automatic evaluation rule to obtain an automatic evaluation result;

[0053] Send the automatic evaluation result to the detection terminal;

[0054] Receive the review and evaluation feedback from the detection terminal, and modify the optimized translation result according to the review and evaluation feedback to obtain the final translation result.

[0055] By adopting the above technical solution, the technical solution of automatically evaluating the optimized translation result and sending it to the detection terminal for manual review realizes the dual check of machine evaluation and manual evaluation. The automatic evaluation system can quickly detect common errors and unreasonable expressions in the translation result, improving the evaluation efficiency. The manual review link can discover deep semantic problems and cultural differences that cannot be recognized by automatic evaluation, ensuring the translation quality, enabling the translation system to modify and improve the translation result according to the evaluation feedback, and ensuring that the finally output translation result is not only accurate and standard in grammar and word usage, but also can accurately convey the cultural connotation and deep meaning of the source language, improving the translation quality.

[0056] Combined with some embodiments of the first aspect, in some embodiments, modifying the optimized translation result according to the review and evaluation feedback to obtain the final translation result specifically includes:

[0057] Obtain the modification suggestions in the review and evaluation feedback to get a set of modification suggestions;

[0058] Sort the set of modification suggestions according to the importance level to obtain a prioritized list;

[0059] Based on the prioritized list, use the target - language expression template to replace the optimized translation result to obtain the final translation result.

[0060] By adopting the above - mentioned technical solution, a technical solution for sorting the importance level of the modification suggestions in the review and evaluation feedback and replacing them based on the priority using the target - language expression template enables the system to process various modification suggestions in the optimal order. The application of the target - language expression template ensures that the modified expression conforms to the language habits and expression norms of the target language, improves the efficiency of translation modification, and makes the final translation result more natural and fluent while maintaining the original meaning.

[0061] In a second aspect, an embodiment of the present application provides a multilingual translation system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0062] In a third aspect, an embodiment of the present application provides a computer - readable storage medium, including instructions, when the above - mentioned instructions run on the system, enabling the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0063] In a fourth aspect, an embodiment of the present application provides a computer program product, when the computer program product runs on the system, enabling the system to execute the method described in any possible implementation manner in the first aspect.

[0064] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0065] 1. The present application provides a multi - language translation method. By identifying and specially processing culture - sensitive content, the translation result can accurately convey the cultural connotations in the source language. Adopting the splitting method of semantic units and the evaluation mechanism based on semantic coherence, the overall coherence of the translation result is improved, and the translation quality and expression accuracy in cross - cultural communication are enhanced. When encountering a situation that does not meet the preset conditions of semantic coherence, the system will search for a better solution through recombination, enhancing the reliability and expression fluency of the translation result. The accuracy of conveying the cultural connotations and emotional colors contained in the source language is improved, and thus the cultural adaptability of the translation result is enhanced.

[0066] 2. The present application provides a multi - language translation method. It optimizes the translation result by using existing high - quality translation experience, improving the translation efficiency. The sample screening mechanism based on the similarity threshold ensures the relevance and applicability of the reference cases, reducing the negative impacts that may be brought by inappropriate references. By extracting and applying semantic combination patterns, the successful translation experience is transformed into reusable knowledge, enhancing the learning ability and optimization effect of the translation system, and also enhancing the practicality and reliability of the system while improving the translation quality.

[0067] 3. The present application provides a multi - language translation method, which is a technical solution of automatically evaluating the optimized translation result and sending it to the detection terminal for manual review, realizing double - check of machine evaluation and manual evaluation. The automatic evaluation system can quickly detect common errors and unreasonable expressions in the translation result, improving the evaluation efficiency. The manual review link can discover deep - level semantic problems and cultural differences that cannot be recognized by automatic evaluation, ensuring the translation quality. The translation system can modify and improve the translation result targeted according to the evaluation feedback, ensuring that the finally output translation result is not only accurate and standard in grammar and word - using, but also can accurately convey the cultural connotations and deep meanings of the source language, improving the translation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a flowchart of a multi - language translation method in an embodiment of the present application.

[0069] Figure 2 It is a flowchart of an optimization method based on historical translation samples in an embodiment of the present application.

[0070] Figure 3 It is a schematic structural diagram of an entity device of a multi - language translation system provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations including one or more of the listed items.

[0072] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0073] Next, an embodiment is used in combination with Figure 1 , to describe a multi-language translation method in the embodiments of the present application:

[0074] Please refer to Figure 1 , which is a schematic flowchart of a multi-language translation method in the embodiments of the present application.

[0075] S101. Obtain the source language text to be translated and the target language type;

[0076] The system first needs to obtain the source language text to be translated input or uploaded by the user. The source language text can be provided to the system by the user in various ways such as keyboard input, speech recognition conversion, image recognition conversion, and document upload. At the same time, the system also needs to obtain the target language type selected or set by the user. The target language type can be selected by the user from the alternative language list provided by the system through a drop-down menu, radio button, etc., or can be the language name or code directly input by the user.

[0077] The system can implement this step by constructing an input interface. The input interface includes a text input box, a voice input button, an image upload button, a document upload button, etc. The user can provide the source language text to the system through these interfaces. At the same time, the input interface also includes a language selection drop-down menu or radio button, and the user can select the target language type through these controls. After the system obtains the source language text and the target language type, it passes them to the subsequent processing module.

[0078] S102. Obtain the cultural feature information corresponding to the source language text from the preset cultural knowledge base; obtain the semantic structure information of the source language text from the preset source language semantic knowledge base;

[0079] The system needs to obtain information related to the source language text from two preset knowledge bases. First, the system needs to obtain the cultural feature information corresponding to the source language text from the preset cultural knowledge base. The cultural feature information can include the cultural background, customs, traditions, values, etc. involved in the source language text. Second, the system needs to obtain the semantic structure information of the source language text from the preset source language semantic knowledge base. The semantic structure information can include the part of speech, syntactic structure, semantic roles, etc. of the source language text.

[0080] The system can implement this step by constructing a knowledge base query interface. The knowledge base query interface receives the source language text as input, then searches in the preset cultural knowledge base and source language semantic knowledge base, and returns the retrieved cultural feature information and semantic structure information. The knowledge base can store and organize data in various forms such as relational databases and graph databases, and the query interface can be implemented using various query languages such as SQL and SPARQL.

[0081] S103. Split the source language text into several semantic units based on the semantic structure information;

[0082] The system needs to split the source language text into several independent semantic units according to the semantic structure information obtained from the source language semantic knowledge base. The semantic units can be language fragments of different lengths and complexities such as words, phrases, clauses, etc., and each semantic unit can express a relatively complete semantics. The purpose of splitting the semantic units is to perform more fine-grained processing and translation in subsequent steps.

[0083] The system can implement this step by constructing a semantic analysis model. The semantic analysis model receives the source language text and its corresponding semantic structure information as input, and then cuts the text into several semantic units according to the information such as part of speech and syntactic role in the semantic structure. Common semantic analysis models include rule-based methods and statistic-based methods. The former divides semantic units by manually defining rules, and the latter automatically learns the rules for dividing semantic units from a large corpus through machine learning algorithms.

[0084] S104. Retrieve the cultural connotation information with a correlation degree greater than a preset threshold with respect to each semantic unit in the preset cultural knowledge base;

[0085] In this step, the system needs to retrieve for each semantic unit in the preset cultural knowledge base and find the cultural connotation information with a relatively high correlation degree with this semantic unit. The cultural connotation information can include the cultural background, customs, traditions, values, etc. related to this semantic unit. The system filters out the cultural connotation information with a correlation degree greater than the preset threshold as the cultural feature of this semantic unit by comparing the correlation degree between the semantic unit and the cultural connotation information.

[0086] The system can implement this step by constructing a semantic matching model. The semantic matching model receives semantic units and candidate cultural connotation information as inputs, and obtains the correlation degree between the two by calculating their similarity or distance in the semantic space. Common semantic matching methods include methods based on vector space models (such as cosine similarity) and methods based on deep learning (such as siamese networks). The system can set a correlation degree threshold and retain the cultural connotation information with a correlation degree greater than this threshold.

[0087] In this step, if the preset cultural knowledge base is not comprehensive enough or the accuracy of the semantic matching model is insufficient, it may lead to the inconsistency between the retrieved cultural connotation information and the actual cultural characteristics of the semantic unit. To solve this problem, the system can adopt the method of transfer learning, use knowledge in other fields (such as common sense, encyclopedia) to expand the cultural knowledge base, and fine-tune the semantic matching model with a small amount of labeled data to improve the accuracy and comprehensiveness of cultural connotation information retrieval.

[0088] S105. When it is determined that the semantic unit contains a cultural feature identifier, determine that the semantic unit is a culture-sensitive semantic unit;

[0089] Regarding how the system determines that the semantic unit contains a cultural feature identifier, the steps are as follows: Obtain a preset cultural feature identifier dictionary, and the preset cultural feature identifier dictionary includes multiple cultural scenarios and the corresponding feature words for each cultural scenario;

[0090] Perform word segmentation on the semantic unit to obtain a word sequence;

[0091] Match each word in the word sequence with the feature words in the cultural feature identifier dictionary;

[0092] When the number of matched feature words is greater than the preset number, determine the cultural scenario corresponding to the semantic unit according to the matched feature words;

[0093] Extract cultural feature rules from the cultural knowledge base based on the cultural scenario;

[0094] When it is determined according to the cultural feature rules that the semantic structure of the semantic unit meets the composition conditions of the cultural feature identifier, it is confirmed that the semantic unit contains the cultural feature identifier. Subsequently, the system determines that the semantic unit is a culture-sensitive semantic unit.

[0095] In this step, the system needs to determine whether a semantic unit is a culture-sensitive semantic unit based on whether it contains a cultural feature identifier. Cultural feature identifiers can be keywords, idioms, proverbs, etc. that reflect specific cultural connotations. If a cultural feature identifier corresponding to the cultural connotation information matching it appears in a semantic unit, it can be considered that the semantic unit is sensitive to this cultural connotation and needs special treatment during translation.

[0096] For how to determine that a semantic unit contains a cultural feature identifier, the system can adopt the following steps: First, the system needs to obtain a preset cultural feature identifier dictionary, which contains various cultural scenarios and their corresponding feature words. Then, the system performs word segmentation on the semantic unit to obtain a word sequence. Next, the system matches each word in the word sequence with the feature words in the cultural feature identifier dictionary. If the number of matching feature words exceeds a certain threshold, the cultural scenario corresponding to the semantic unit can be determined based on these feature words. Finally, the system extracts the feature rules of this cultural scenario from the cultural knowledge base and determines whether the semantic structure of the semantic unit meets the composition conditions of the cultural feature identifier according to the feature rules. If it meets, the semantic unit is confirmed as a culture-sensitive semantic unit.

[0097] S106. Extract from the preset cultural knowledge base the target language expressions whose cultural connotation matching degree with the cultural connotation of the culture-sensitive semantic unit is greater than the preset matching degree;

[0098] In this step, the system needs to find, for each culture-sensitive semantic unit, the target language expressions that match its cultural connotation from the preset cultural knowledge base. The cultural connotation matching degree indicates the similarity between the target language expression and the source language culture-sensitive semantic unit in terms of semantics, emotion, rhetoric, etc. The higher the matching degree, the more accurately the expression can convey the cultural connotation contained in the source language semantic unit. The system screens all the target language expressions in the knowledge base by setting a matching degree threshold and extracts the expressions with a matching degree greater than the threshold.

[0099] The system can achieve this step by constructing a cross-language cultural connotation matching model. This model takes the source language culture-sensitive semantic unit and the target language candidate expression as inputs, performs cross-language semantic encoding on them through a deep learning network, and calculates their similarity in the semantic space to obtain the cultural connotation matching degree. The training data of the model can come from an artificial translation corpus or can be automatically mined from a large-scale bilingual corpus in an unsupervised manner.

[0100] S107. Invoke the target language semantic knowledge base and combine the target language expressions semantically with the standard translation results of other semantic units in the source language text except the culture-sensitive semantic units to obtain the translation result;

[0101] The system calls the target language semantic knowledge base, and semantically combines the target language expression with the standard translation results of other semantic units except the culture-sensitive semantic units in the source language text to obtain the translation result, specifically including: constructing a target language syntactic tree template, which includes multiple semantic slots and the syntactic relationships between the semantic slots; obtaining the syntactic structure information of the target language expression from the target language semantic knowledge base; filling the target language expression into the corresponding semantic slots according to the syntactic structure information; obtaining a semantic relevance calculation model, which is used to calculate the semantic relevance strength between different semantic units; calculating the semantic relevance between the standard translation results of other semantic units and the content in the filled semantic slots based on the semantic relevance calculation model; sorting the standard translation results of other semantic units according to the semantic relevance to obtain a sorting result; filling the standard translation results of other semantic units into the remaining semantic slots in descending order according to the sorting result; determining the conjunctions between the semantic slots according to the syntactic tree template; filling the conjunctions between the semantic slots and converting them into a linear text sequence to obtain the translation result.

[0102] In this step, the system needs to semantically combine the target language expressions of the culture-sensitive semantic units with the standard translation results of other semantic units in the source language text to obtain the complete target language translation result. The standard translation results can be obtained through a conventional machine translation system. During the semantic combination process, the system needs to call the target language semantic knowledge base and use the grammar and semantic rules therein to perform operations such as reordering, inserting, and replacing the translation results of different semantic units, and add appropriate conjunctions to finally form a target language text with coherent semantics and reasonable structure.

[0103] The system can implement this step by constructing a semantic combination model based on the syntactic tree. This model first constructs a target language syntactic tree template, which contains several semantic slots and the syntactic dependency relationships between the slots. Then, the model fills the target language expressions of the culture-sensitive semantic units into the corresponding semantic slots. Next, according to the semantic relevance calculation results, the model fills the standard translation results of other semantic units into the remaining semantic slots in descending order. Finally, the model adds conjunctions between the slots according to the structure of the syntactic tree and linearizes the slot content into a target language text. The semantic relevance can be calculated through a pre-trained language model or directly perform end-to-end relevance learning on the syntactic tree through methods such as graph neural networks.

[0104] In this step, if the rules of the target language semantic knowledge base are not comprehensive enough, or the semantic combination model cannot well coordinate the translation results of different semantic units, translation results with incoherent semantics and chaotic structures may be generated. To solve these problems, the system can introduce a large-scale target language corpus to expand the rules of the semantic knowledge base in an unsupervised or semi-supervised manner; in addition, the system can add some flexibility and fault tolerance to the syntactic tree template, allowing local adjustments to the syntactic structure to meet the needs of the translation results of different semantic units; at the same time, the system can combine technologies such as reinforcement learning, and by setting appropriate reward functions, guide the semantic combination model to generate more fluent and natural translation results.

[0105] S108. Judge whether the semantic coherence of the translation result meets the preset conditions according to the preset translation evaluation rules;

[0106] This step needs to evaluate the quality of the translation result output in step S107 to determine whether it meets the preset conditions of the system to ensure the readability and comprehensibility of the translation. The translation evaluation rules can include indicators in multiple aspects such as grammar correctness, semantic integrity, logical consistency, and stylistic appropriateness. Each indicator has corresponding judgment criteria and thresholds. Only when the translation result meets the standards in each indicator can it be determined to meet the preset conditions. The preset conditions refer to a series of measurement standards and thresholds used by the system when judging whether the translation result meets the requirements of semantic coherence. These preset conditions are set by the system according to specific task requirements and domain knowledge before applying this translation method and are used as the basis for evaluating the translation quality. Specifically, the preset conditions can include the following aspects:

[0107] Grammar correctness condition: The translation result needs to conform to the grammar rules of the target language, such as word order, sentence pattern, tense, number agreement, etc., to avoid grammar errors or unreasonable expressions. The system can set specific types of grammar errors and quantity thresholds, and translation results exceeding the thresholds will be determined not to meet the grammar correctness condition.

[0108] Semantic integrity condition: The translation result needs to completely convey the semantic information of the source language text without problems such as missing translation, mistranslation, or adding irrelevant information. The system can set quantitative indicators and thresholds for semantic similarity or coverage to judge whether the translation result meets the semantic integrity condition.

[0109] Logical consistency condition: The translation result needs to be consistent in logic and fact, avoiding self-contradictory or expressions that do not conform to objective facts. The system can set some common sense rules and reasoning rules to check the logical consistency of the translation result and set corresponding threshold conditions.

[0110] Appropriateness condition of language style: The style and tone of the translation result need to match the source language text and conform to the expression habits of a specific register and communication scenario. The system can pre-define the language style features in different registers and scenarios and set a similarity threshold to judge whether the translation result meets the appropriateness condition of language style.

[0111] Fluency condition of the translation: The translation result needs to conform to the expression habits of the target language, read smoothly and naturally, without problems such as stiffness, verbosity, and repetition. The system can establish a language model, calculate the naturalness score of the translation result, and set a fluency score threshold to judge whether it meets the fluency condition of the translation.

[0112] The system can achieve this step by constructing a multi-index translation quality evaluation model. This model takes the translation result as input and judges and scores it from different perspectives. In terms of grammatical correctness, the model can use rule-based or statistical grammar checkers to identify grammar errors in the translation result; in terms of semantic integrity, the model can rely on pre-trained semantic representation models to judge whether the translation result completely conveys the semantic information of the source language text; in terms of logical consistency, the model can use techniques such as common sense reasoning and anaphora resolution to check whether the translation result is logically and factually self-contradictory; in terms of appropriateness of language style, the model can combine register classification and sentiment analysis to judge whether the style and tone of the translation result match the source language text. Finally, the model sums up the scores of each index after weighting and compares them with the preset threshold to obtain the final evaluation result.

[0113] S109. Output the translation result.

[0114] If the preset conditions are met, the translation result is output. If the preset conditions are not met, step S107 is executed again.

[0115] This step mainly realizes outputting the translation result that has passed the quality evaluation in step S108 to the user. If the evaluation result meets the preset semantic coherence condition, the system directly presents the translation result to the user; if the evaluation result does not meet the standard, the system needs to return to step S107, adjust and optimize the result of semantic combination, and then conduct quality evaluation again until a translation result that meets the requirements is generated.

[0116] The output form of the translation result can be customized according to the user's needs and usage scenarios. The system can support multiple output methods such as text, speech, and images, facilitating the user to obtain the translation result on different devices and in different environments. When outputting the translation result, the system can also provide some auxiliary information, such as the original text content, keyword explanations, cultural background introductions, etc., to help the user better understand and use the translation result. In addition, the system can provide interactive functions for the user, allowing the user to evaluate and give feedback on the translation result, and continuously optimize the system's translation strategy.

[0117] In the above embodiments, by identifying and specially processing culture-sensitive content, the translation result can accurately convey the cultural connotations in the source language. By adopting the splitting method of semantic units and the evaluation mechanism based on semantic coherence, the overall coherence of the translation result is improved, and the translation quality and expression accuracy in cross-cultural communication are enhanced. When encountering situations that do not meet the preset conditions of semantic coherence, the system will search for a better solution through recombination, enhancing the reliability and expression fluency of the translation result. The accuracy of conveying the cultural connotations and emotional colors contained in the source language is improved, thereby enhancing the cultural adaptability of the translation result.

[0118] After completing the basic translation process, in order to further improve the translation quality, this application also provides an optimization method based on historical translation samples. By extracting similarity matching and semantic combination patterns, more reference and optimization space are provided for the current translation task. The following combines Figure 2 to describe an optimization method based on historical translation samples in the embodiments of this application:

[0119] Please refer to Figure 2 which is a schematic flowchart of an optimization method based on historical translation samples in the embodiments of this application.

[0120] S201. Obtain a historical translation sample library;

[0121] The system obtains a historical translation sample library, which contains multiple pairs of samples of source language texts and corresponding target language translation texts.

[0122] In this step, the system needs to obtain a historical translation sample library, which contains a large number of source language texts and corresponding target language translation texts, and these text pairs constitute a set of translation samples. The historical translation sample library can come from multiple channels, such as a manually translated corpus, a parallel corpus generated by a machine translation system, a translation memory library, etc. The scale and quality of the sample library have an important impact on subsequent optimization of the translation result.

[0123] The system can obtain and construct a historical translation sample library in various ways. A common way is to extract samples from existing bilingual or multilingual parallel corpora, which can be public data sets or corpora accumulated within the system. Another way is to use web crawler technology to crawl a large number of bilingual web pages from the Internet and form a sample library after alignment and filtering. In addition, the system can also access a manual translation interface, allowing users to upload and manage their own translation samples, continuously enriching and improving the sample library.

[0124] S202. Calculate the similarity between the source language text and the source language texts in the historical translation sample library based on a preset similarity calculation rule;

[0125] The system calculates the similarity between the source language text and the source language texts in the historical translation sample library based on a preset similarity calculation rule, which specifically includes: extracting features from the source language text to generate a text feature vector; extracting features from the source language texts in the historical translation sample library to generate sample feature vectors; calculating the Euclidean distance between the text feature vector and the sample feature vectors in each dimension; and performing a weighted sum of the Euclidean distances in each dimension to obtain the similarity.

[0126] In this step, the system needs to calculate the similarity between the current source language text to be translated and all the source language texts in the historical translation sample library as the basis for subsequent screening of similar samples. The similarity calculation rule can be based on various text representation methods and distance metrics, such as the vector space model, edit distance, etc. Different calculation rules are applicable to different text characteristics and task requirements.

[0127] Specifically, the system can adopt the following steps to calculate the similarity between the source language text and the source language texts in the historical translation sample library: First, extract features from the current source language text and convert it into a text feature vector with a fixed dimension; then, perform the same feature extraction on each source language text in the historical translation sample library to obtain a series of sample feature vectors; next, the system calculates the Euclidean distance between the text feature vector and each sample feature vector in each dimension to obtain a distance vector; finally, the system performs a weighted sum of the components in the distance vector to obtain the similarity between the source language text and the sample text.

[0128] S203. Screen out the sample pairs with a similarity greater than the preset similarity threshold from the historical translation sample library;

[0129] In this step, the system needs to screen out the translation sample pairs with relatively high similarity according to the similarity between the source language text and the source language texts in the historical translation sample library, providing a reference for subsequent extraction of semantic combination patterns. The setting of the similarity threshold needs to balance the quantity and quality of the samples, ensuring both that enough similar samples are screened out and that too many irrelevant noise samples are not introduced.

[0130] The system can use sorting and truncation methods to achieve the screening of similar samples. Specifically, the system first sorts all the samples in descending order according to the similarity scores, and then selects the top N samples with scores greater than or equal to the threshold as the similar sample subset according to the preset similarity threshold. The size of the threshold can be set according to the scale and distribution characteristics of the sample library, or can be automatically optimized through methods such as cross-validation. In addition, the system can further filter the screened similar samples, such as removing duplicate samples, low-quality samples, etc., to improve the purity of the similar samples.

[0131] S204. Extract the semantic combination patterns of the target language translation text in the sample pairs;

[0132] In this step, the system needs to extract the semantic combination patterns contained in the target language translation text from the selected similar translation sample pairs as a reference for optimizing the current translation result. The semantic combination patterns reflect the conventional structures and expression ways of the target language when expressing specific semantics, which can help the system generate more target language - accustomed translation results.

[0133] The system can extract semantic combination patterns through the following methods: First, perform syntactic analysis on the target language translation text to obtain a syntactic dependency tree; then, identify important components in the syntactic tree, such as the subject, predicate, object, etc., and extract the semantic relationships between these components; next, the system abstracts the semantic relationships into semantic combination templates, which contain information such as the types of semantic components, semantic roles, and combination orders; finally, the system generalizes and summarizes the semantic combination templates extracted from multiple sample texts to obtain some common semantic combination patterns.

[0134] S205. Apply the semantic combination patterns to the translation result to generate alternative translation results;

[0135] In this step, the system needs to apply the semantic combination patterns extracted from the similar samples to the current translation result to generate one or more alternative translation results. The alternative translation results are semantically equivalent to the original translation result, but may be more in line with the idiomatic usage of the target language in terms of expression. By generating diverse alternative translation results, the system can provide more choices for users and also provide a greater optimization space for subsequent translation optimization.

[0136] Specifically, the system can adopt the following steps to apply the semantic combination patterns to the translation result: First, perform semantic analysis on the original translation result to extract key semantic components, such as terms, named entities, semantic roles, etc.; then, retrieve the combination templates that match these semantic components in the semantic combination pattern library, and fill the semantic components into the corresponding positions of the template according to the semantic relationships and combination orders in the template to form a new semantic combination structure; finally, the system generates one or more alternative translation results based on the new semantic combination structure, which adopt different expression ways while maintaining the original semantics.

[0137] S206. Compare the semantic coherence scores of the alternative translation results and the translation result;

[0138] In this step, the system needs to evaluate the semantic coherence of the alternative translation results and the original translation results, calculate their respective coherence scores, and use them as the basis for optimizing the translation results. Semantic coherence reflects the consistency and continuity of the translation results in terms of logic, semantics, pragmatics, etc., and is one of the important indicators for measuring translation quality. The higher the coherence score, the better the readability and comprehensibility of the translation results.

[0139] The system can adopt various methods to calculate the semantic coherence score. Common methods include language model-based methods and discourse analysis-based methods. The language model-based method calculates the probability or perplexity of the translation results under the n-gram language model of the target language by establishing the n-gram language model of the target language. The larger the probability or the smaller the perplexity, the more in line with the usage habits of the target language the translation results are, and the better the coherence. The discourse analysis-based method calculates the integrity and rationality of the discourse relations by identifying the discourse relations in the translation results, such as reference, connection, substitution, etc. The closer and more reasonable the discourse relations are, the better the coherence of the translation results.

[0140] In the process of semantic coherence evaluation, it may be that the training corpus of the language model is insufficient or not matched with the target domain, resulting in inaccurate coherence scores; the discourse analysis method misidentifies the discourse structure and semantic roles of the translation results, affecting the reliability of coherence calculation; the coherence scores are inconsistent with the manual evaluation results and are difficult to be used as a direct basis for translation optimization. To solve these problems, the system can use domain adaptation technology to dynamically adjust the parameters of the language model according to the characteristics of the target domain and improve the domain adaptability of the language model; at the same time, the system can introduce a discourse analysis model based on deep learning to improve the accuracy of discourse relation recognition and semantic role annotation through end-to-end joint learning; in addition, the system can also combine the coherence scores with other translation quality indicators through methods such as multi-objective learning and reinforcement learning to learn a comprehensive translation optimization model, so that the optimized translation results are improved in multiple dimensions.

[0141] S207. Select the one with a high semantic coherence score as the optimized translation result.

[0142] In this step, the system needs to select the one with the highest score as the final optimized result according to the semantic coherence scores of the alternative translation results and the original translation results. While ensuring the semantic correctness of the optimized translation results, the language expression is more fluent, natural, and more in line with the expression habits of the target language.

[0143] Specifically, the system can sort the original translation result and all alternative translation results according to the semantic coherence score to obtain an ordered candidate set; then, the system selects the result with the highest score from the candidate set as the preliminary selection result; next, the system further checks and verifies the preliminary selection result, such as grammar checking, term consistency checking, etc., to ensure that there are no obvious errors or unreasonable points; finally, the system outputs the result that passes the check as the optimized translation result and updates the translation memory to provide reference for subsequent translation tasks.

[0144] In the above embodiments, the translation result is optimized by using existing high-quality translation experience, improving the translation efficiency. The sample screening mechanism based on the similarity threshold ensures the relevance and applicability of the reference cases, reducing the negative impact that may be brought by improper reference. Through the extraction and application of semantic combination patterns, the successful translation experience is transformed into reusable knowledge, enhancing the learning ability and optimization effect of the translation system, and also enhancing the practicality and reliability of the system while improving the translation quality.

[0145] Furthermore, after selecting the translation result with a high semantic coherence score as the optimized translation result, the system can also evaluate the optimized translation result according to the preset automatic evaluation rules to obtain an automatic evaluation result; send the automatic evaluation result to the detection terminal; receive the review and evaluation feedback from the detection terminal, and modify the optimized translation result according to the review and evaluation feedback to obtain the final translation result, specifically including: obtaining the modification suggestions in the review and evaluation feedback to get a set of modification suggestions; sorting the set of modification suggestions according to the importance level to obtain a priority sorted list; based on the priority sorted list, using the target language expression template to replace the optimized translation result to obtain the final translation result.

[0146] The system can automatically evaluate the optimized translation results according to preset automatic evaluation rules, which may include grammar checking, term consistency checking, semantic integrity checking, etc. By applying these rules, the system generates comprehensive automatic evaluation results. Then, the system sends the automatic evaluation results to the detection terminal, and human evaluators review the automatic evaluation results. Based on their professional knowledge and experience, the human evaluators confirm, correct, and supplement the automatic evaluation results to form a review evaluation feedback and send the feedback back to the system. Finally, the system receives the review evaluation feedback returned by the detection terminal and modifies the optimized translation results according to the modification suggestions in the feedback to obtain the final translation results. During the modification process, the system first extracts all the modification suggestions from the review evaluation feedback to form a set of modification suggestions. Then, the system evaluates the importance of each suggestion in the set of modification suggestions according to the preset importance evaluation rules to obtain the importance score of each suggestion. Next, the system sorts the set of modification suggestions according to the importance scores to obtain a priority sorted list arranged from high to low in importance. Finally, the system traverses the priority sorted list and applies the modification suggestions in the list to the optimized translation results in turn using the expression templates of the target language to finally obtain the modified final translation results.

[0147] In the above embodiments, the technical solution of automatically evaluating the optimized translation results and sending them to the detection terminal for manual review realizes the dual review of machine evaluation and manual evaluation. The automatic evaluation system can quickly detect common errors and unreasonable expressions in the translation results, improving the evaluation efficiency. The manual review link can discover deep semantic problems and cultural differences that cannot be recognized by automatic evaluation, ensuring the translation quality, enabling the translation system to make targeted modifications and improvements to the translation results according to the evaluation feedback, and ensuring that the finally output translation results are not only accurate and standard in grammar and word usage but also can accurately convey the cultural connotations and deep meanings of the source language, thus improving the translation quality.

[0148] The following describes the system in the embodiments of the present invention application from the perspective of hardware processing. Please refer to Figure 3 which is a schematic structural diagram of an entity device of a multi-language translation system provided by an embodiment of the present application.

[0149] It should be noted that Figure 3 the structure of the system shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present invention.

[0150] As Figure 3As shown, the system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the methods in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0151] The following components are connected to the I / O interface 305: an input section 306 including a camera, an infrared sensor, etc.; an output section 307 including a Liquid Crystal Display (LCD) and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.

[0152] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the Central Processing Unit (CPU) 301, various functions defined in the present invention are executed.

[0153] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.

[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0155] As another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or may exist alone without being assembled into the system. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of a system, the system implements the method provided in the above embodiments.

[0156] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

[0157] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted as "if determining...", "in response to determining...", "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0158] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.

[0159] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The aforementioned storage media include various media that can store program codes, such as ROM, random access memory (RAM), magnetic disks, or optical discs.

Claims

1. A multilingual translation method, characterized in that: include: Get the source language text and target language type to be translated; Acquiring cultural feature information corresponding to the source language text from a preset cultural knowledge base; Acquiring semantic structure information of the source language text from a preset source language semantic knowledge base; Splitting the source language text into a plurality of semantic units based on the semantic structure information; For each of the semantic units, cultural connotation information having a correlation degree with the semantic unit greater than a preset threshold is retrieved from the preset cultural knowledge base; In the case where it is determined that the semantic unit contains a cultural feature identifier, determining that the semantic unit is a culturally sensitive semantic unit; Extracting a target language expression having a cultural connotation matching degree greater than a preset matching degree of the cultural connotation of the culturally sensitive semantic unit from the preset cultural knowledge base based on the cultural connotation information; Calling a target language semantic knowledge base, and semantically combining the target language expression with standard translation results of other semantic units in the source language text except the culturally sensitive semantic unit, to obtain a translation result; Determining whether the semantic coherence of the translation result meets a preset condition according to a preset translation evaluation rule; If the preset condition is met, outputting the translation result; If the preset condition is not met, the step of semantically combining the target language expression with the standard translation results of other semantic units in the source language text except the culturally sensitive semantic unit to obtain the translation result is re-executed, specifically including: Constructing a target language syntax tree template, wherein the syntax tree template includes a plurality of semantic slots and syntax relations between the semantic slots; Acquire syntactic structure information of the target language expression from the target language semantic knowledge base; fill the target language expression into the corresponding semantic slot according to the syntactic structure information; Acquire a semantic relevance calculation model, wherein the semantic relevance calculation model is used to calculate the semantic relevance strength between different semantic units; Calculating the semantic relevance between the standard translation results of the other semantic units and the content in the filled semantic slots based on the semantic relevance calculation model; Sorting the standard translation results of the other semantic units according to the semantic association to obtain a sorting result; According to the order of the sorting results from high to low, the standard translation results of the other semantic units are sequentially filled into the remaining semantic slots; Determining the connectives between the semantic slots according to the syntax tree template; The connecting words are filled between the semantic slots and converted into a linear text sequence to obtain a translation result.

2. The method according to claim 1, characterized in that In the case where it is determined that the semantic unit contains a cultural feature identifier, the method specifically includes: Obtaining a preset cultural feature identification dictionary, wherein the preset cultural feature identification dictionary includes a plurality of cultural scenes and a feature word corresponding to each of the cultural scenes; Performing word segmentation processing on the semantic unit to obtain a word sequence; Matching each word in the word sequence with a characteristic word in the cultural characteristic identification dictionary; When the number of matched feature words is greater than a preset number, determining the cultural scene corresponding to the semantic unit according to the matched feature words; Extracting cultural feature rules from the cultural knowledge base based on the cultural scenario; When it is determined according to the cultural feature rule that the semantic structure of the semantic unit satisfies the constituent conditions of the cultural feature identifier, it is confirmed that the semantic unit includes the cultural feature identifier.

3. The method according to claim 1, characterized in that: After the step of re-performing the step of semantically combining the target language expression with the standard translation results of other semantic units in the source language text except the culturally sensitive semantic unit to obtain the translation result, the method further includes: Acquire a historical translation sample library, wherein the historical translation sample library includes a plurality of sample pairs of the source language text and the corresponding target language translation text; Calculating the similarity between the source language text and the source language text in the historical translation sample library based on a preset similarity calculation rule; Screening out sample pairs whose similarity is greater than a preset similarity threshold from the historical translation sample library; extracting semantic combination patterns of target language translation texts in the sample pairs; Applying the semantic combination pattern to the translation result to generate an alternative translation result; comparing the semantic coherence score of the alternative translation result with that of the translation result; The translation result with a high semantic coherence score is selected as the optimized translation result.

4. The method according to claim 3, characterized in that The calculating the similarity between the source language text and the source language text in the historical translation sample library based on the preset similarity calculation rule specifically includes: Extracting features from the source language text to generate a text feature vector; Extracting features from the source language text in the historical translation sample library to generate a sample feature vector; The Euclidean distance between the text feature vector and the sample feature vector in each dimension is calculated; and the Euclidean distance in each dimension is weighted and summed to obtain the similarity.

5. The method according to claim 3, characterized in that: After selecting the translation result with a high semantic coherence score as the optimized translation result, the method further includes: Evaluate the optimized translation result according to preset automatic evaluation rules to obtain an automatic evaluation result; and send the automatic evaluation result to a detection terminal; The review and evaluation feedback from the detection terminal is received, and the optimized translation result is modified according to the review and evaluation feedback to obtain a final translation result.

6. The method according to claim 5, characterized in that The step of modifying the optimized translation result according to the review and evaluation feedback to obtain the final translation result specifically includes: Obtain modification suggestions in the review and evaluation feedback to obtain a set of modification suggestions; Sorting the set of modification suggestions according to importance to obtain a priority sorting list; Based on the priority sorting list, the optimized translation result is replaced with a target language expression template to obtain a final translation result.

7. A multilingual translation system, characterized in that: The system comprises: One or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method as described in any one of claims 1-6.

8. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a system, the system is caused to execute the method according to any one of claims 1 to 6.

9. A computer program product, characterized in that When the computer program product is executed on a system, the system is caused to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Term recognition method for multi-language translation

    CN116822517A