Text translation method and device, equipment and storage medium

By establishing the full text index of historical translation data under the target language and performing similarity calculations, the problem of inefficient translation efficiency in commodity internationalization is solved, an efficient translation process is achieved, and the overseas supply and marketing level and efficiency of commodities are improved.

CN120181104APending Publication Date: 2025-06-20创优数字科技(广东)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510418621.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the process of internationalization of goods, translation for each product separately leads to inefficient translation efficiency, which in turn affects the degree of overseas supply and marketing of goods.

Method used

By obtaining the historical translation data in the target language and the initial text to be translated for the target product, establishing a full text index of the historical translation data, searching and determining each first word and sentence, selecting the corresponding translated word and sentence, and determining the target translated word and sentence through similarity calculation, replacing the initial text to be translated to generate the target translated text.

Benefits of technology

It improves translation efficiency, reduces the workload of manual translation, reduces the translation cost, and ensures the quality of translation, and improves the overseas supply and marketing level and efficiency of goods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181104A_ABST
    Figure CN120181104A_ABST
Patent Text Reader

Abstract

The invention discloses a text translation method and device, equipment and a storage medium, and the method comprises the steps: obtaining historical translation data in a target language and an initial to-be-translated text of a target commodity, and building a full-text index of the historical translation data; retrieving the initial to-be-translated text based on the full-text index to determine each first word and sentence; selecting translated words and sentences corresponding to the first words and sentences from the historical translation data; screening other second words and sentences except the first words and sentences in the initial to-be-translated text; performing similarity calculation on each second word and sentence and historical translation data to determine each target translation word and sentence; and in the initial to-be-translated text, replacing each first word and phrase with respective corresponding translated word and phrase, and meanwhile, replacing each second word and phrase with respective corresponding target translated word and phrase to obtain a target translated text. The translation efficiency is greatly improved, the overseas supply and marketing degree and efficiency of commodities are further improved, and manpower and material resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text translation, and in particular to a text translation method, apparatus, device and storage medium. Background Art

[0002] With the expansion of global business, the localization translation of product packaging has become a core part of the internationalization strategy of enterprises. Generally, in the system or platform of product life cycle management, the translation needs of product packaging usually involve complex scenarios of multiple languages ​​and multiple countries. For example, when a product is sold to Russia, the United States, France and other countries, the instruction text on its outer packaging must be translated into Russian, English, French, etc. respectively to meet the needs of local users.

[0003] However, for different types of goods, or different products of the same type, there will be more or less repeated sentences in the same translation language. If translation is performed separately for each product, the translation efficiency will be greatly reduced, and the overseas supply and sales of the products will also be greatly reduced. Summary of the invention

[0004] In view of this, the present application provides a text translation method, apparatus, device and storage medium for solving the problem that if translation is performed for each commodity separately, the translation efficiency will be greatly reduced and the overseas supply and sales of the commodity will also be greatly reduced.

[0005] To achieve the above objectives, the proposed solution is as follows:

[0006] In a first aspect, a text translation method comprises:

[0007] Acquire historical translation data in the target language and the initial text to be translated of the target product, and establish a full-text index of the historical translation data;

[0008] Retrieving the initial text to be translated based on the full-text index to determine each first word or sentence;

[0009] Selecting each translated word or phrase corresponding to each first word or phrase from the historical translation data;

[0010] Filtering each second word or sentence other than the first word or sentence in the initial text to be translated;

[0011] Calculating the similarity between each of the second words and sentences and the historical translation data to determine each target translation word and sentence;

[0012] In the initial text to be translated, each of the first words and sentences is replaced by a corresponding translated word and sentence, and each of the second words and sentences is replaced by a corresponding target translation word and sentence, so as to obtain a target translation text.

[0013] Preferably, retrieving the initial text to be translated based on the full-text index to determine each first phrase includes:

[0014] Performing word segmentation on the initial text to be translated, and deleting stop words after word segmentation to obtain a first text to be translated;

[0015] Converting the first text to be translated into a lowercase format uniformly to obtain a second text to be translated;

[0016] Removing special characters from the second text to be translated to obtain a third text to be translated;

[0017] Retrieving each corresponding first phrase in the third text to be translated based on the full-text index.

[0018] Preferably, selecting each translated phrase corresponding to each first phrase from the historical translation data includes:

[0019] Determining the manufacturer of the target commodity;

[0020] Screening the commodity with the highest historical sales volume among all commodities of the same category as the target commodity by the manufacturer;

[0021] Extracting the word usage characteristics, sentence structures, and logical relationship expression characteristics of the commodity with the highest historical sales volume in the target language;

[0022] Selecting each translated phrase corresponding to each first phrase from the historical translation text according to the word usage characteristics, sentence structures, and logical relationship expression characteristics.

[0023] Preferably, selecting each translated phrase corresponding to each first phrase from the historical translation data includes:

[0024] Extracting the original text and the translated text from the historical translation data;

[0025] Splitting the original text to obtain each piece of original data, and at the same time splitting the translated text to obtain each piece of translated data;

[0026] Performing standardization processing on each piece of translated data to obtain each piece of standard translated data;

[0027] Combining each piece of standard translated data with its corresponding original data to form each first data pair;

[0028] Removing duplicate data pairs from each first data pair to obtain each second data pair;

[0029] Screening each second data pair to obtain each target data pair;

[0030] Take each translation text in each of the target data pairs as each translated sentence.

[0031] Preferably, screening each of the second data pairs to obtain each target data pair includes:

[0032] Screen out each second data pair with the same original text as each third data pair;

[0033] Obtain the translation time corresponding to the translation text in each of the third data pairs, and at the same time obtain the first product type of the target product and the second product type corresponding to each of the third data pairs;

[0034] Delete each third data pair whose translation time is earlier than the preset time to obtain each fourth data pair;

[0035] Determine whether there is a second product type that is the same as the first product type;

[0036] If so, take each fourth data pair corresponding to the second product type and each other second data pair as each target data pair.

[0037] Preferably, calculating the similarity between each of the second sentences and the historical translation data to determine each target translation sentence includes:

[0038] Encode each of the second sentences to obtain each untranslated vector;

[0039] Encode the historical translation data to obtain each historical translation vector;

[0040] For each untranslated vector, calculate the similarity between the untranslated vector and each of the historical translation vectors to obtain a similarity value;

[0041] Sort in descending order of the similarity value, and select the top N historical translation vectors as target translation vectors;

[0042] Decode each of the target translation vectors to obtain each target translation sentence.

[0043] Preferably, use a pre-trained encoding model to encode each of the second sentences or historical translation data;

[0044] The encoding model includes a basic encoding layer, a hierarchical semantic encoding layer, and a hierarchical fusion layer;

[0045] The input end of the basic encoding layer is used as the input end of the encoding model, and the hierarchical fusion layer is used as the output layer of the encoding model;

[0046] The output end of the basic coding layer is connected to the input end of the hierarchical semantic coding layer, and the output end of the hierarchical semantic coding layer is connected to the input end of the hierarchical fusion layer.

[0047] In a second aspect, a text translation device includes:

[0048] A data acquisition and index establishment module, configured to acquire historical translation data in a target language and an initial text to be translated of a target commodity, and establish a full-text index of the historical translation data;

[0049] A retrieval module, configured to retrieve the initial text to be translated based on the full-text index to determine each first phrase;

[0050] A selection module, configured to select each translated phrase corresponding to each of the first phrases from the historical translation data;

[0051] A screening module, configured to screen each other second phrase in the initial text to be translated except the first phrases;

[0052] A similarity calculation module, configured to calculate the similarity between each of the second phrases and the historical translation data respectively to determine each target translation phrase;

[0053] A replacement module, configured to replace each of the first phrases with its corresponding translated phrase in the initial text to be translated, and at the same time replace each of the second phrases with its corresponding target translation phrase to obtain a target translation text.

[0054] In a third aspect, a text translation device includes a memory and a processor;

[0055] The memory is configured to store a program;

[0056] The processor is configured to execute the program to implement each step of the text translation method described in any item of the first aspect.

[0057] In a fourth aspect, a storage medium stores a computer program, and when the computer program is executed by a processor, each step of the text translation method described in any item of the first aspect is implemented.

[0058] As can be seen from the above technical solution, in this application, historical translation data in the target language and an initial text to be translated of the target commodity are obtained, and a full-text index of the historical translation data is established; the initial text to be translated is retrieved based on the full-text index to determine each first phrase; each translated phrase corresponding to each first phrase is selected from the historical translation data; each other second phrase in the initial text to be translated except the first phrase is screened; the similarity between each second phrase and the historical translation data is calculated respectively to determine each target translation phrase; in the initial text to be translated, each first phrase is replaced with its corresponding translated phrase, and at the same time each second phrase is replaced with its corresponding target translation phrase to obtain a target translation text. In this application, the historical translation data in the target language is used as a translation resource to establish a full-text index, so that the initial text to be translated can be retrieved based on the full-text index, improving the retrieval efficiency to determine each first phrase and the translated sentences already recorded in the historical translation data. For sentences that have not been translated, the historical translation data can also be used for translation, and the target translation sentences are determined by calculating the similarity, so as to obtain the target translation text after translation. In this application, all the content of the initial text to be translated is translated in this separated way, greatly improving the translation efficiency. It is not necessary to translate each commodity one by one as in the existing method, further improving the degree and efficiency of overseas supply and sales of commodities, saving manpower and material resources. This way of reusing and referring to historical translation data can make full use of existing translation resources, reduce the workload of manual translation, lower the translation cost, and ensure the translation quality at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0060] Figure 1 It is an optional flowchart of a text translation method provided by an embodiment of this application;

[0061] Figure 2 It is a schematic structural diagram of a text translation device provided by an embodiment of this application;

[0062] Figure 3 It is a schematic structural diagram of a text translation device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0064] The present invention can be used in many general or special computing device environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or equipment, and so on.

[0065] An embodiment of the present invention provides a text translation method, which can be applied to various computer terminals or intelligent terminals, and the execution subject can be a processor or a server of a computer terminal or an intelligent terminal. The method flow chart of the method is as Figure 1 shown and specifically includes:

[0066] S1: Obtain historical translation data in the target language and the initial text to be translated of the target commodity, and establish a full-text index of the historical translation data.

[0067] The historical translation data in the target language can be obtained by summarizing each historical translation document in the PLM (Product Lifecycle Management) system. For the usage scenario of the present application, if you want to sell the commodity to English-speaking countries such as the United States or the United Kingdom, the target language is English. If you want to translate imported overseas products in China, the target language is Chinese.

[0068] The initial text to be translated of the target commodity can be regarded as the original text description of the commodity, or the original product description, or the additional description text of the commodity, advertising language, etc. This embodiment does not limit this.

[0069] The historical translation data includes multiple fields: source text source_text, translated text translated_text, target language target_lang, translation time timestamp, translator ID translator_id, commodity classification product_category, etc. During the index establishment process, these fields can be configured for indexing, such as the selection of the tokenizer, the index storage location, etc., and then the index creation operation is executed, so as to perform word segmentation processing on the content of the specified field and establish an index data structure.

[0070] S2: Retrieve the initial text to be translated based on the full-text index to determine each first sentence or phrase.

[0071] Retrieving the initial text to be translated by full-text indexing is equivalent to creating an accurate map for the massive amount of initial text to be translated. There is no need to search word by word in sequence as in the traditional way. The full-text index can be used to quickly locate the corresponding text to determine the first word or sentence. It can also improve the accuracy of the search and will not miss any possible first word or sentence.

[0072] For larger amounts of initial text to be translated, full-text indexing can be used to handle a variety of target languages, making it more adaptable.

[0073] S3: Selecting translated words and sentences corresponding to the first words and sentences from the historical translation data.

[0074] Since the historical translation data is already a translated historical record, and each first word or sentence in the initial text to be translated is determined based on the historical translation data, it is easy to determine each corresponding translated word or sentence from the historical translation data.

[0075] It is understandable that although the translation is carried out for the target language, no matter it is a character, a phrase or a sentence, there will be different translated phrases. For example, the Chinese word "妈妈" has the corresponding English words "mam", "mother", etc. Therefore, a first phrase may correspond to one or more translated phrases.

[0076] It should be noted that for each first sentence, there is an original sentence corresponding to the translated sentence in the historical translation record, and the first sentence is exactly the same as the original sentence, so the translated sentence can be directly used as the translation version of the first sentence in the target language.

[0077] S4: Filtering the second words and sentences other than the first words and sentences in the initial text to be translated.

[0078] In addition to the completely corresponding first words and sentences, the initial text to be translated may also contain other second words and sentences that cannot completely correspond to the historical translation data. Therefore, each second word and sentence needs to be translated. In this case, the historical translation data is also used for translation, such as using similarity translation. This can improve the translation speed while ensuring that there are no translation problems.

[0079] Of course, if no corresponding words or phrases that can be translated with similarity can be found in the historical translation data, a third-party translation software can be used for translation.

[0080] S5: Calculate the similarity between each of the second words and sentences and the historical translation data to determine each target translation word and sentence.

[0081] Similarity calculations can be performed according to indicators such as the semantic similarity and correlation between phrases and sentences, so as to screen out each target translated phrase in the historical translation data that meets the similarity requirements.

[0082] Similarly, a second phrase may correspond to one or more target translated phrases. For example, there are several translated phrases in the historical translation data with relatively high similarity to the second phrase, or two or more translated phrases have the same similarity degree to the second phrase. Therefore, these translated phrases can be used as target translated phrases.

[0083] S6: In the initial text to be translated, replace each of the first phrases with its corresponding translated phrase, and at the same time replace each of the second phrases with its corresponding target translated phrase to obtain the target translated text.

[0084] In the above steps, the translated phrases corresponding to each phrase in the initial text to be translated have been found, so they can be replaced simultaneously to obtain the target translated text.

[0085] As can be seen from the above technical solution, in this application, historical translation data in the target language and the initial text to be translated of the target commodity are obtained, and a full-text index of the historical translation data is established; the initial text to be translated is retrieved based on the full-text index to determine each first phrase; each translated phrase corresponding to each first phrase is selected from the historical translation data; other second phrases in the initial text to be translated except the first phrases are screened; the similarity between each second phrase and the historical translation data is calculated respectively to determine each target translated phrase; in the initial text to be translated, each first phrase is replaced with its corresponding translated phrase, and at the same time each second phrase is replaced with its corresponding target translated phrase to obtain the target translated text. In this application, the historical translation data in the target language is used as a translation resource to establish a full-text index, so that the initial text to be translated can be retrieved based on the full-text index, improving the retrieval efficiency to determine each first phrase and the translated sentences already recorded in the historical translation data. For sentences that have not been translated, the historical translation data can also be used for translation, and the target translated sentences are determined by calculating the similarity, so as to obtain the target translated text after translation. In this application, all the content of the initial text to be translated is translated in this separate way, greatly improving the translation efficiency. It is not necessary to translate each commodity one by one as in the existing method, further improving the degree and efficiency of overseas supply and sales of commodities, saving manpower and material resources. This way of reusing and referring to historical translation data can make full use of existing translation resources, reduce the workload of manual translation, reduce the translation cost, and ensure the translation quality at the same time.

[0086] In the method provided by the embodiment of the present invention, the process of retrieving the initial text to be translated based on the full-text index to determine each first phrase and sentence is specifically described as follows:

[0087] Perform word segmentation on the initial text to be translated, and delete stop words after word segmentation to obtain the first text to be translated;

[0088] Convert the first text to be translated into a lowercase format uniformly to obtain the second text to be translated;

[0089] Remove special characters in the second text to be translated to obtain the third text to be translated;

[0090] Retrieve each corresponding first phrase and sentence in the third text to be translated based on the full-text index.

[0091] Specifically, since it is necessary to determine each first phrase and sentence, it is necessary to first perform word segmentation on the initial text to be translated, that is, distinguish it according to the format of words or phrases and sentences, and then delete the words belonging to the preset stop words, such as "de", "shi", etc. Stop words usually do not carry important information. Deleting them can reduce data noise, reduce computational complexity, and improve the accuracy of retrieval and translation; converting the text to a lowercase format is to unify the case, which can avoid retrieval or matching problems caused by inconsistent case. For example, "Apple" and "apple" will be regarded as different words in the case-sensitive situation. After converting to lowercase, it can be ensured that they are regarded as the same word, thereby improving the consistency and accuracy of retrieval; special characters (such as punctuation marks, numbers, special symbols, etc.) usually do not participate in semantic analysis. Removing them can further simplify the second text to be translated and reduce unnecessary interference. This step helps to improve the purity of the text and makes the subsequent retrieval and translation more focused on meaningful vocabulary.

[0092] The above steps work together to ensure the efficiency of text preprocessing and the accuracy of subsequent retrieval and translation.

[0093] The following will elaborate on the process of selecting each translated phrase and sentence corresponding to each first phrase and sentence from the historical translation data in the present application in the following two ways:

[0094] (1) Select from the perspective of the manufacturer's translation style.

[0095] Determine the manufacturer of the target commodity;

[0096] Screen the commodity with the highest historical sales volume among all commodities of the same type as the target commodity by the manufacturer;

[0097] Extract the word-using characteristics, sentence structure, and logical relationship expression characteristics of the commodity with the highest historical sales volume in the target language;

[0098] Select the translated sentences corresponding to each of the first sentences from the historical translation text according to the word - using characteristics, sentence structures, and logical relationship expressions described above.

[0099] Specifically, the manufacturer of the target product can be determined by analyzing relevant information about the target product (such as brand, model, manufacturer, etc.). Since different manufacturers have different product description styles, translation can be carried out based on the translation style of the manufacturer of the target product. Considering the sales volume and promotion volume of the product, among products of the same type as the target product, there must be a product with the highest historical sales volume. The product with the highest sales volume usually reflects the market preference for this type of product. Therefore, to a certain extent, the highest historical sales volume is inseparable from the advantages of the product itself and also from the accurate and attractive translated product description. Then it can be considered that its corresponding translation text may be more in line with the habits and needs of target - language users. Therefore, by analyzing the translation characteristics of high - sales - volume products, it can provide a reference for the translation of the target product and improve the market adaptability of the translation.

[0100] Therefore, through natural language processing techniques (such as word - frequency statistics, syntactic analysis, semantic analysis, etc.), extract the word - using characteristics, sentence structures, and logical relationship expressions of the product with the highest historical sales volume in the target language. Select the translated sentences corresponding to each of the first sentences from the historical translation text according to these characteristics, structures, and expressions, ensuring that the translation text is consistent with the translation style of the product with the highest historical sales volume in terms of word - using, sentence pattern, and logic, and improving the coherence and professionalism of the overall translation.

[0101] (2) Select from the original text and the translation text in the historical translation data.

[0102] Extract the original text and the translation text from the historical translation data.

[0103] Split the original text to obtain each piece of original data, and at the same time split the translation text to obtain each piece of translation data.

[0104] Perform standardization processing on each piece of the translation data to obtain each piece of standardized translation data.

[0105] Combine each piece of the standardized translation data with its corresponding original data to form each first data pair.

[0106] Remove duplicate data pairs from each of the first data pairs to obtain each second data pair.

[0107] Screen each of the second data pairs to obtain each target data pair.

[0108] Take each translation text in each of the said target data pairs as each translated sentence or phrase.

[0109] Specifically, after extracting the original text and the translation text from the historical translation data, it is necessary to first ensure the integrity of the corresponding relationship between the original text and the translation text to avoid data loss or misalignment. Then, split the original text and the translation text into independent sentences or paragraphs (i.e., "each piece of data"). That is to say, split the large text into smaller units for easy precise matching and processing, improving the efficiency and accuracy of subsequent steps. At the same time, for subsequent matching and analysis, to reduce errors caused by inconsistent formats, it is also necessary to perform standardized and normalized processing on the translation data, including unifying case, removing special characters, standardizing terms, etc., which can reduce data noise and improve data consistency.

[0110] To provide structured data for subsequent matching and retrieval, a mapping relationship can be established between the standard translation data and their respective corresponding original data. Since there will be exactly the same first data pairs, in order to lighten the data, reduce data redundancy, and lower storage and computing costs, the redundant and repeated first data pairs can be deleted, and only one pair needs to be retained.

[0111] Among them, for the process of screening each of the said second data pairs to obtain each target data pair, it can specifically include the following steps:

[0112] Screen out each second data pair with the same original text as each third data pair;

[0113] Obtain the translation time corresponding to the translation text in each of the said third data pairs, and at the same time obtain the first product type of the target product and the second product types corresponding to each of the said third data pairs;

[0114] Delete each third data pair whose translation time is earlier than the preset time to obtain each fourth data pair;

[0115] Judge whether there is a second product type that is the same as the first product type;

[0116] If so, take each fourth data pair corresponding to the second product type and each other second data pair as each target data pair.

[0117] Specifically, in order to make the target data pair more accurate, the second data pair can be further screened. For the second data pair with the same original text, consider the translation time and the product type. First, for the translation time, as languages evolve, streamline, and become more internationalized, many languages will optimize translations. Therefore, for the case where the original text is the same but the translation texts are different, eliminate the translations with too early time and retain the newer translation forms to ensure that the translation results conform to the current language habits and industry standards.

[0118] In addition, for the commodity type, selecting the same translation style and method for the same commodity type can ensure that the translation result is consistent with the type of the target commodity and improve the professionalism of the translation. If there is no second commodity type that is the same as the first commodity type, then select a similar commodity type. For example, a dress and a skirt are similar types of clothes.

[0119] The following embodiments will explain in detail the steps of calculating the similarity between each of the second phrases and the historical translation data to determine each target translation phrase in the present application.

[0120] Encode each of the second phrases to obtain each untranslated vector;

[0121] Encode the historical translation data to obtain each historical translation vector;

[0122] For each of the untranslated vectors, calculate the similarity between the untranslated vector and each of the historical translation vectors to obtain a similarity value;

[0123] Sort in descending order of the similarity values, and select the top N historical translation vectors as the target translation vectors; N can be a number such as 5 or 10.

[0124] Decode each of the target translation vectors to obtain each target translation phrase.

[0125] Specifically, converting the high-dimensional historical translation data and the second phrases into low-dimensional vector representations can reduce the computational complexity, standardize the data format at the same time, and the encoding methods of the two are the same, thereby ensuring the effectiveness of the similarity calculation.

[0126] A vector library can be constructed for historical translation vectors to facilitate quick retrieval and matching. Then, the semantic similarity between the untranslated vector and the historical translation vectors is quantified by calculating the similarity between them. The higher the similarity value, the closer the untranslated text is semantically to the historical translation data, and the better the matching effect. Among them, an appropriate text similarity algorithm can be selected, such as cosine similarity: the similarity is measured by calculating the cosine value between vectors; Jaccard similarity (Jaccard similarity coefficient): the similarity is calculated based on the overlap degree of the vocabulary sets; edit distance: the edit distance between two sentences is calculated, and the smaller the distance, the higher the similarity; TF-IDF weighted similarity (term frequency–inverse document frequency): the sentence similarity is calculated by weighting the vocabulary with TF-IDF; pre-trained model: sentence vectors are generated using pre-trained models such as BERT (Bidirectional Encoder Representations from Transformers), and then the similarity is calculated.

[0127] Among them, cosine similarity is an index to measure the similarity between vectors. It is calculated based on the angle between vectors, and its value range is from -1 to 1. The closer the value is to 1, the more similar the two vectors are.

[0128] Screening according to the similarity value from high to low can ensure that the selected translation data is highly relevant semantically to the untranslated text. The value of N can also be adjusted to control the number of returned translation results, balance precision and diversity, and improve the flexibility of screening.

[0129] Furthermore, a pre-trained encoding model can be used to encode each of the second sentences or historical translation data. For the encoding model, its structure is as follows:

[0130] The encoding model includes a basic encoding layer, a hierarchical semantic encoding layer, and a hierarchical fusion layer;

[0131] The input end of the basic encoding layer serves as the input end of the encoding model, and the hierarchical fusion layer serves as the output layer of the encoding model;

[0132] The output end of the basic encoding layer is connected to the input end of the hierarchical semantic encoding layer, and the output end of the hierarchical semantic encoding layer is connected to the input end of the hierarchical fusion layer.

[0133] Specifically, the basic encoding layer of the encoding model contains a basic encoder (such as mBERT or XLM-R), which can generate word-level context representations and send them to the hierarchical semantic encoding layer. The hierarchical semantic encoding layer can add hierarchical semantic encoding on the basis of the data output by the basic encoding layer, encode the semantics at the word level, sentence level, and paragraph level respectively, and generate richer text representations by capturing semantic information at the word, sentence, and paragraph levels. For the encoding at the sentence level, sentence pooling or attention mechanism can be used to aggregate the word-level representations into sentence-level representations. For the encoding at the paragraph level, for long texts (such as paragraphs or documents), hierarchical attention mechanism is used to further aggregate the representations of multiple sentences into paragraph-level representations. For example, each sentence is encoded first, and then attention mechanism is used to weight and fuse the sentence representations. The hierarchical fusion module fuses the representations at the word level, sentence level, and paragraph level to generate the final text representation, including processes such as concatenation, weighted sum attention fusion. Attention fusion refers to dynamically adjusting the weights of different-level representations using the attention mechanism.

[0134] Furthermore, after obtaining the target translation text, the target translation text can be sent to the display interface of the user client. Web Worker can be used to perform matching calculations asynchronously to avoid interface lag, and the user can use debounce to request data to reduce repeated calculations.

[0135] corresponding to Figure 1 the method described above, an embodiment of the present invention also provides a text translation device for Figure 1 the specific implementation of the method in. The text translation device provided by the embodiment of the present invention can be in a computer terminal or various mobile devices. Combining Figure 2 , the text translation device is introduced. As Figure 2 shown, the device may include:

[0136] A data acquisition and index establishment module 10, configured to acquire historical translation data in the target language and the initial text to be translated of the target commodity, and establish a full-text index of the historical translation data;

[0137] A retrieval module 20, configured to retrieve the initial text to be translated based on the full-text index to determine each first sentence or phrase;

[0138] A selection module 30, configured to select each translated sentence or phrase corresponding to each of the first sentences or phrases from the historical translation data;

[0139] A screening module 40, configured to screen each of the other second phrases in the initial text to be translated except the first phrase;

[0140] A similarity calculation module 50, configured to calculate the similarity between each of the second phrases and the historical translation data respectively to determine each target translated phrase;

[0141] A replacement module 60, configured to replace each of the first phrases with its corresponding translated phrase in the initial text to be translated, and at the same time replace each of the second phrases with its corresponding target translated phrase to obtain a target translated text.

[0142] It can be seen from the above technical solutions that in this application, by obtaining historical translation data in a target language and an initial text to be translated of a target commodity, and establishing a full-text index of the historical translation data; retrieving the initial text to be translated based on the full-text index to determine each first phrase; selecting each corresponding translated phrase from the historical translation data for each of the first phrases; screening each of the other second phrases in the initial text to be translated except the first phrase; calculating the similarity between each of the second phrases and the historical translation data respectively to determine each target translated phrase; in the initial text to be translated, replacing each of the first phrases with its corresponding translated phrase, and at the same time replacing each of the second phrases with its corresponding target translated phrase to obtain a target translated text. This application uses the historical translation data in the target language as a translation resource and establishes a full-text index, so that the initial text to be translated can be retrieved based on the full-text index, improving the retrieval efficiency to determine each first phrase and the translated sentences already recorded in the historical translation data. For sentences that have not been translated, the historical translation data can also be used for translation, and the target translated sentences are determined by calculating the similarity, so as to obtain the target translated text after translation. This application translates all the content of the initial text to be translated in this separated way, greatly improving the translation efficiency, without the need to translate each commodity one by one as in the existing method, further improving the overseas supply and sales degree and efficiency of the commodity, saving manpower and material resources. This way of reusing and referring to historical translation data can make full use of existing translation resources, reduce the workload of manual translation, reduce the translation cost, and ensure the translation quality at the same time.

[0143] Furthermore, an embodiment of this application provides a text translation device. Optionally, Figure 3 The hardware structure block diagram of the text translation device is shown. Referring to Figure 3 , the hardware structure of the text translation device may include: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.

[0144] In the embodiments of the present application, the number of the processor 01, the communication interface 02, the memory 03, and the communication bus 04 is at least one, and the processor 01, the communication interface 02, and the memory 03 complete the communication with each other through the communication bus 04.

[0145] The processor 01 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.

[0146] The memory 03 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0147] Wherein, the memory stores a program, and the processor can call the program stored in the memory. The program is used to execute the following text translation method, including:

[0148] Obtain the historical translation data in the target language and the initial text to be translated of the target commodity, and establish a full-text index of the historical translation data;

[0149] Retrieve the initial text to be translated based on the full-text index to determine each first sentence or phrase;

[0150] Select each translated sentence or phrase corresponding to each first sentence or phrase from the historical translation data;

[0151] Screen each other second sentence or phrase in the initial text to be translated except the first sentence or phrase;

[0152] Calculate the similarity between each second sentence or phrase and the historical translation data respectively to determine each target translation sentence or phrase;

[0153] In the initial text to be translated, replace each first sentence or phrase with its corresponding translated sentence or phrase, and at the same time replace each second sentence or phrase with its corresponding target translation sentence or phrase to obtain the target translation text.

[0154] Optionally, the refined functions and extended functions of the program can refer to the description of the text translation method in the method embodiments.

[0155] The embodiments of the present application also provide a storage medium. The storage medium can store a program suitable for the processor to execute. When the program runs, it controls the device where the storage medium is located to execute the following text translation method, including:

[0156] Obtain historical translation data in the target language and the initial text to be translated of the target commodity, and establish a full-text index of the historical translation data;

[0157] Retrieve the initial text to be translated based on the full-text index to determine each first sentence or phrase;

[0158] Select each translated sentence or phrase corresponding to each first sentence or phrase from the historical translation data;

[0159] Screen out each other second sentence or phrase in the initial text to be translated except the first sentence or phrase;

[0160] Calculate the similarity between each second sentence or phrase and the historical translation data respectively to determine each target translation sentence or phrase;

[0161] In the initial text to be translated, replace each first sentence or phrase with its corresponding translated sentence or phrase, and at the same time replace each second sentence or phrase with its corresponding target translation sentence or phrase to obtain the target translation text.

[0162] Specifically, the storage medium may be a computer-readable storage medium, and the computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM.

[0163] Optionally, the refined functions and extended functions of the program may refer to the description of the text translation method in the method embodiments.

[0164] In addition, in each embodiment of the present disclosure, each functional module may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a live broadcast device, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present disclosure.

[0165] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0166] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0167] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A text translation method, characterized in that: include: Acquire historical translation data in the target language and the initial text to be translated of the target product, and establish a full-text index of the historical translation data; Retrieving the initial text to be translated based on the full-text index to determine each first word or sentence; Selecting each translated word or phrase corresponding to each first word or phrase from the historical translation data; Filtering each second word or sentence other than the first word or sentence in the initial text to be translated; Calculating the similarity between each of the second words and sentences and the historical translation data to determine each target translation word and sentence; In the initial text to be translated, each of the first words and sentences is replaced by a corresponding translated word and sentence, and each of the second words and sentences is replaced by a corresponding target translation word and sentence, so as to obtain a target translation text.

2. The method according to claim 1, characterized in that The searching the initial text to be translated based on the full-text index to determine each first word or sentence includes: Segmenting the initial text to be translated, and deleting stop words after segmentation to obtain a first text to be translated; Converting the first text to be translated into lowercase format to obtain a second text to be translated; removing special characters from the second text to be translated to obtain a third text to be translated; Based on the full-text index, the first words and sentences corresponding to the third text to be translated are retrieved.

3. The method according to claim 1, characterized in that The selecting each translated phrase corresponding to each first phrase from the historical translation data includes: Determine the manufacturer of the target product; Filter the products of the manufacturer with the highest historical sales among the products of the same category as the target product; Extracting the wording characteristics, sentence structure and logical relationship expression features of the product with the highest historical sales in the target language; According to the wording characteristics, sentence structure and logical relationship expression features, each translated word or sentence corresponding to each first word or sentence is selected from the historical translation text.

4. The method according to claim 1, characterized in that: The selecting each translated phrase corresponding to each first phrase from the historical translation data includes: Extracting the original text and the translated text from the historical translation data; Splitting the original text to obtain individual pieces of original text data, and splitting the translated text to obtain individual pieces of translated data; Performing standardization processing on each of the translation data to obtain each of the standard translation data; Combining each of the standard translation data and the corresponding original data into first data pairs; Removing duplicate data pairs from each of the first data pairs to obtain each of the second data pairs; Screening each of the second data pairs to obtain each target data pair; The translated texts in the target data pairs are taken as translated words and sentences.

5. The method according to claim 4, characterized in that The step of screening each of the second data pairs to obtain each target data pair includes: Filter out the second data pairs with the same original text as the third data pairs; Obtaining the translation time corresponding to the translated text in each of the third data pairs, and simultaneously obtaining the first commodity type of the target commodity and the second commodity type corresponding to each of the third data pairs; Deleting each third data pair whose translation time is earlier than a preset time to obtain each fourth data pair; Determining whether there is a second commodity type that is the same as the first commodity type; If so, each fourth data pair corresponding to the second commodity type and each other second data pair are used as target data pairs.

6. The method according to any one of claims 1 to 5, characterized in that The calculating the similarity between each of the second words and sentences and the historical translation data to determine each target translation word and sentence includes: Encode each of the second words and sentences respectively to obtain each untranslated vector; Encoding the historical translation data to obtain each historical translation vector; For each of the untranslated vectors, similarity calculation is performed between the untranslated vector and each of the historical translation vectors to obtain a similarity value; Sort the similarity values ​​from high to low, and select the first N historical translation vectors as target translation vectors; Each of the target translation vectors is decoded respectively to obtain each target translation phrase.

7. The method according to claim 6, characterized in that Encoding each of the second words or historical translation data using a pre-trained encoding model; The coding model includes a basic coding layer, a hierarchical semantic coding layer, and a hierarchical fusion layer; The input end of the base coding layer serves as the input end of the coding model, and the hierarchical fusion layer serves as the output layer of the coding model; The output end of the basic coding layer is connected to the input end of the hierarchical semantic coding layer, and the output end of the hierarchical semantic coding layer is connected to the input end of the hierarchical fusion layer.

8. A text translation device, characterized in that: include: A data acquisition and index building module, used to acquire historical translation data in the target language and the initial text to be translated of the target product, and to build a full-text index of the historical translation data; A retrieval module, configured to retrieve the initial text to be translated based on the full-text index to determine each first word or sentence; A selection module, configured to select, from the historical translation data, each translated phrase corresponding to each first phrase; A screening module, used for screening all second words and sentences other than the first words and sentences in the initial text to be translated; A similarity calculation module, used to calculate the similarity between each of the second words and sentences and the historical translation data to determine each target translation word and sentence; The replacement module is used to replace each of the first words and sentences with the corresponding translated words and sentences in the initial text to be translated, and replace each of the second words and sentences with the corresponding target translation words and sentences to obtain a target translation text.

9. A text translation device, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the text translation method according to any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the text translation method according to any one of claims 1 to 7 is implemented.