Text translation method and its device, equipment, medium, and product

Through iterative pre-training and fine-tuning training of the text translation model, a special dictionary was constructed, and the problem of poor translation effect of low-resource corpus in Traditional Chinese was solved, precise mutual translation between Traditional Chinese and English was achieved, and the translation quality and user experience of cross-border e-commerce platforms were improved.

CN114757211BActive Publication Date: 2025-08-12GUANGZHOU HUADUO NETWORK TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210259679.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2025-08-12
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

In the prior art, the low-resource corpus translation technology of Traditional Chinese has problems with low translation loyalty and fluency, especially in cross-border e-commerce platforms, the parallel corpus resources in specific regions using Traditional Chinese are poor, resulting in poor translation results.

Method used

Using a text translation model that is pre-trained to a convergent state, through iterative pre-training and fine-tuning training, the first training data set and the second training data set are used to build a special dictionary, replace synonyms, and achieve accurate mutual translation between traditional Chinese and English.

Benefits of technology

It improves the loyalty and fluency of translation, meets the localized expression needs of users of cross-border e-commerce platforms, and improves user experience and user stickiness of the platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114757211B_ABST
    Figure CN114757211B_ABST
Patent Text Reader

Abstract

The present application discloses a text translation method and its apparatus, device, medium, and product. The method comprises: obtaining a text to be translated; translating the text to be translated using a text translation model pre-trained to a convergent state to obtain a result text, wherein the training process of the text translation model comprises the following steps: iteratively pre-training the text translation model using a first training sample from a first training data set to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus; and iteratively fine-tuning the text translation model using a second training sample from a second training data set to achieve convergence, wherein the second training sample is a second parallel corpus consisting of the first language corpus and a second dialect of the second language. The present application utilizes a text translation model to achieve accurate translation between different languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of e-commerce translation technology, and in particular to a text translation method and its corresponding apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] For major cross-border e-commerce platforms, translation is the key bridge to maintain communication between buyers and sellers on the cross-border e-commerce platforms, and it is also the most basic rigid demand of cross-border e-commerce platforms. Accurate e-commerce translation can play a vital role in scenarios such as product display description, search, and recommendation. On the one hand, it greatly increases the degree of user familiarity with the product. On the other hand, it can assist cross-border e-commerce platform users to establish product labels with their localized expression texts, promote the matching between the product labels and the expression texts that accurately translate user needs, and make the searched and recommended products meet user needs.

[0003] When translating text from a specific dialect of Traditional Chinese into English, existing technologies for translating low-resource corpora such as Traditional Chinese face the following challenges:

[0004] 1. Due to the small population and small scale of e-commerce businesses in specific regions where Traditional Chinese is used, the current accumulation of parallel corpora in industry and academia for specific regions using Traditional Chinese is relatively small, and corpus resources are relatively scarce. In addition, the models trained with this parallel corpus are only barely able to understand the original meaning, and the translation fidelity and fluency are low, and have not yet reached the stage where they can be put into use.

[0005] 2. An existing technical solution first uses a parallel corpus of simplified Chinese, which has relatively rich corpus resources, to train a model for translation between simplified Chinese and English. Secondly, it uses a simplified-traditional mapping between simplified and traditional Chinese for translation, ultimately achieving translation results between traditional Chinese and English. However, this technical solution still has significant errors. For some brand words and attribute terms, for example, the English expression "instant noodles" is translated into "convenient noodles" in simplified Chinese. After the simplified-traditional mapping, it is still converted into "convenient noodles" in Traditional Chinese, while the specific term used in specific regions using Traditional Chinese is "convenient noodles" or "doll noodles".

[0006] Similar situations also arise in translation scenarios where different dialects of the same language have similar but partially different synonymous corpora. To achieve accurate translation between a dialect with low-resource corpora and another language, the applicant has made corresponding explorations. Summary of the Invention

[0007] The primary purpose of this application is to solve at least one of the above problems and provide a text translation method and its corresponding device, computer equipment, computer-readable storage medium, and computer program product.

[0008] In order to meet the various objectives of this application, this application adopts the following technical solutions:

[0009] A text translation method provided to meet one of the purposes of this application includes the following steps:

[0010] Get the text to be translated;

[0011] The text to be translated is translated using a text translation model that has been pre-trained to a convergent state to obtain a result text, wherein the text to be translated is a text expressed in a first language or a second dialect of a second language, and correspondingly, the result text is a text expressed in the second dialect of the second language or a text expressed in the first language. The training process of the text translation model includes the following steps:

[0012] Iterative pre-training of the text translation model is performed using a first training sample in a first training dataset to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus, and the second language corpus is a translation of the first language corpus into a first dialect of a second language with individual words replaced by synonyms in the second dialect of the second language;

[0013] The text translation model is iteratively fine-tuned and trained by calling a second training sample in a second training data set to converge the model, wherein the second training sample is a second parallel corpus consisting of a corpus in the first language and a second dialect corpus in the second language.

[0014] In a further embodiment, the text translation model training process includes the following steps:

[0015] Obtaining a first translated text in a first dialect of a second language and a second translated text in a second dialect corresponding to each original text expressed in the first language in the preset data set, as a first corpus and a second corpus respectively;

[0016] Comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between the different translated versions of the synonyms into a special dictionary;

[0017] Calling a corpus resource library, the corpus resource library includes a first parallel corpus, the first parallel corpus consisting of two corresponding texts: a first language corpus and a second language first dialect corpus;

[0018] According to the mapping relationship data of the special dictionary, the translated versions of the synonymous words belonging to the first dialect of the second language in the first parallel corpus are replaced with the translated versions of the synonymous words in the second dialect in the special dictionary to form the second language corpus, so that the corpus resource library constitutes the first training data set.

[0019] In a further embodiment, the text translation model training process further includes the following steps:

[0020] Each original text expressed in the first language in the preset data set and its corresponding translated text in the second language and second dialect are constructed into a parallel corpus, and the parallel corpus is constructed into a second training data set.

[0021] In a further embodiment, the method of comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus and constructing mapping relationship data between the different translated versions of the synonyms into a special dictionary includes the following steps:

[0022] comparing the difference texts between different translations of the same original text in the first corpus and the second corpus;

[0023] Using an entity noun classification model that has been pre-trained to a convergent state, the difference text in the first corpus is judged to determine whether it is a specific type of noun;

[0024] The difference texts in the first corpus and the second corpus corresponding to nouns belonging to a specific type are determined as different translation versions of synonyms, constructed into mapping relationship data, and stored in a pre-constructed special dictionary.

[0025] In a further embodiment, based on the mapping relationship data of the special dictionary, replacing the translated versions of the synonymous words belonging to the first dialect of the second language in the first parallel corpus with the translated versions of the synonymous words in the second dialect in the special dictionary to form the second language corpus includes the following steps:

[0026] searching the corresponding target text in the full parallel corpus in the corpus resource library according to the translated versions of the synonyms corresponding to the first dialect of the second language in the special dictionary;

[0027] A translation version of the second dialect of the second language corresponding to the synonym in a special dictionary is obtained, and a randomly selected portion of the target text is replaced with the translation version of the second dialect.

[0028] In an extended embodiment, a text translation model that has been pre-trained to a convergent state is used to translate the text to be translated. After obtaining the result text, the following steps are included:

[0029] In response to the text replacement instruction acting on the text to be translated, the text to be translated is replaced with the result text for display on a current interface for obtaining the text to be translated;

[0030] In response to a pointing event acting on the display area of the result text, a text prompt box is displayed on the current interface, and the text to be translated is displayed in the text prompt box.

[0031] A text translation device provided to meet one of the purposes of the present application includes: a text acquisition module, a model translation module, a model pre-training module, and a model fine-tuning module, wherein the text acquisition module is used to acquire a text to be translated; the model translation module is used to translate the text to be translated using a text translation model pre-trained to a convergent state to obtain a result text, wherein the text to be translated is a text expressed in a first language or a second dialect of a second language, and correspondingly, the result text is a text expressed in a second dialect of the second language or a text expressed in the first language, and the training process of the text translation model includes: the model pre-training module The block is used to call a first training sample in a first training data set to perform iterative pre-training on the text translation model to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus, and the second language corpus is the result of replacing individual words in a translation text of the first language corpus corresponding to a first dialect of the second language with synonyms in the second dialect of the second language; the model fine-tuning module is used to call a second training sample in a second training data set to perform iterative fine-tuning training on the text translation model to achieve convergence, wherein the second training sample is a second parallel corpus consisting of the first language corpus and a second dialect of the second language.

[0032] In a further embodiment, before the training process of the text translation model, the method includes: a bidirectional translation module for obtaining a first translation text in a first dialect of a second language and a second translation text in a second dialect corresponding to each original text expressed in the first language in a preset data set, as a first corpus and a second corpus, respectively; a dictionary construction module for comparing synonyms between different translation texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between different translation versions of the synonyms into a special dictionary; a resource library calling module for calling a corpus resource library, which includes a first parallel corpus, which is composed of two corresponding texts of a corpus in the first language and a corpus in the first dialect of the second language; and a synonym replacement module for replacing the translated versions of synonyms belonging to the first dialect of the second language in the first parallel corpus with translated versions of synonyms in the second dialect in the special dictionary according to the mapping relationship data of the special dictionary to form a second language corpus, so that the corpus resource library constitutes the first training data set.

[0033] In a further embodiment, before the training process of the text translation model, it also includes: a corpus construction module, which is used to construct each original text expressed in the first language in the preset data set and its corresponding translation text in the second language and second dialect into a parallel corpus, and construct the parallel corpus into a second training data set.

[0034] In a further embodiment, the dictionary construction module includes: a difference comparison submodule, which is used to compare the difference texts between different translated texts of the same original text in the first corpus and the second corpus; a type judgment submodule, which is used to use an entity noun classification model pre-trained to a convergent state to judge the difference texts in the first corpus to determine whether they are nouns of a specific type; a relationship construction submodule, which is used to determine the difference texts in the first corpus and the second corpus corresponding to nouns of a specific type as different translation versions of synonyms, construct them into mapping relationship data, and store them in a pre-constructed special dictionary.

[0035] In a further embodiment, the synonym replacement module includes: a target retrieval submodule, which is used to retrieve the corresponding target text in the full parallel corpus in the corpus resource library based on the translation version of the synonym corresponding to the first dialect of the second language in the special dictionary; a random replacement submodule, which is used to obtain the translation version of the second dialect of the second language corresponding to the synonym in the special dictionary, and randomly select part of the target text to replace it with the translation version of the second dialect.

[0036] In an extended embodiment, the model translation module includes: a text replacement submodule for responding to a text replacement instruction acting on the text to be translated, replacing the text to be translated with the result text for display on the current interface for obtaining the text to be translated; and a text prompt submodule for responding to a pointing event acting on the display area of the result text, displaying a text prompt box on the current interface, and displaying the text to be translated in the text prompt box.

[0037] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, wherein the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the text translation method described in the present application.

[0038] A computer-readable storage medium is provided to meet another purpose of the present application, which stores a computer program implemented according to the text translation method in the form of computer-readable instructions. When the computer program is called and executed by a computer, the steps included in the method are executed.

[0039] A computer program product provided to meet another purpose of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the method described in any embodiment of the present application.

[0040] According to the typical embodiments and alternative embodiments of the present application, it can be seen that the technical solution of the present application has many advantages, including but not limited to the following aspects:

[0041] First, the present application uses a large amount of first parallel corpus in the first training data set for text translation model training, so that the model learns the basic language expression ability such as word order and grammar of the second language and first dialect with relatively rich corpus resources; on this basis, since individual words in the translated text of the second language and first dialect in the first parallel corpus in the first training data set are replaced with synonyms of the second language and second dialect, the text translation model is pre-trained with the first training data set, so that the text translation model can accurately extract synonym features and translate the text to be translated into the expression words of the second language and second dialect with relatively low corpus resources, so that the translation effect is more grounded, which is convenient for users in the location of the second dialect to understand, efficiently assists users to use, improves user experience, and increases user stickiness of the e-commerce platform.

[0042] Secondly, the second parallel corpus in the second training dataset is composed of the first language corpus and the second language second dialect corpus. Although its corpus resources may be relatively small, the text translation model has already acquired a certain translation capability during the pre-training phase. Therefore, after fine-tuning the text translation model using the second training dataset, the basic Chinese language expression of the model is further adjusted to the language expression of the second dialect, thereby improving the translation fidelity and fluency of the model and achieving accurate mutual translation between the first language and the second language second dialect.

[0043] In addition, the technical solution of the present application, on the one hand, uses a simple and easy-to-implement text translation model architecture with low computational cost and low load, making it suitable for deployment on the client side and applicable to a variety of e-commerce environment application scenarios. On the other hand, the execution efficiency is higher and the implementation cost is lower, and it is also suitable for deployment on the background server to respond to massive concurrent demands, thereby obtaining economies of scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0045] Figure 1 This is a flowchart of a typical embodiment of the text translation method of the present application;

[0046] Figure 2This is a schematic diagram of the process of constructing a first training data set in an embodiment of the present application;

[0047] Figure 3 A schematic diagram of the process of constructing a special dictionary in an embodiment of the present application;

[0048] Figure 4 Schematic diagram of the process of synonym replacement in the embodiment of the present application;

[0049] Figures 5(a) and 5(b) are schematic diagrams of a graphical user interface of a terminal device according to an embodiment of the present application, respectively illustrating a product interface before translation and a product interface after translation;

[0050] Figure 6 This is a flowchart of an extended embodiment of the text translation method of the present application;

[0051] Figure 7 This is a functional block diagram of the text translation device of this application;

[0052] Figure 8 This is a schematic diagram of the structure of a computer device used in this application. DETAILED DESCRIPTION

[0053] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0054] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0055] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0056] Those skilled in the art will appreciate that the terms "client," "terminal," and "terminal device" as used herein include both devices that are wireless signal receivers, i.e., devices that only have wireless signal receivers without transmission capabilities, and devices that have receiving and transmitting hardware capable of two-way communication over a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers and tablet computers, which have single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service), which may combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, a pager, Internet / Intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; and conventional laptop and / or palmtop computers or other devices, which have and / or include a radio frequency receiver. As used herein, the terms "client," "terminal," or "terminal device" may be portable, transportable, or installed in a vehicle (air, sea, and / or land), or may be adapted and / or configured to operate locally, and / or in a distributed manner, at any other location on Earth and / or in space. As used herein, the terms "client," "terminal," or "terminal device" may also refer to a communication terminal, an Internet terminal, or a music / video playback terminal, such as a PDA, an MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or may refer to a smart TV, a set-top box, or other device.

[0057] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with capabilities equivalent to those of a personal computer. It is a hardware device that has the necessary components revealed by the von Neumann principle, such as a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. Computer programs are stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input and output devices to complete specific functions.

[0058] It should be noted that the concept of "server" referred to in this application can also be extended to server clusters. Based on the network deployment principles understood by those skilled in the art, the servers described should be logically divided. In physical space, these servers can be independent of each other but callable through interfaces, or integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method of this application.

[0059] Unless expressly specified, one or more technical features of the present application can be deployed on a server for implementation and accessed by a client through a remote call to obtain an online service interface provided by the server, or can be directly deployed and run on a client for implementation.

[0060] Unless expressly specified otherwise, the neural network models referenced or may be referenced in this application may be deployed on a remote server and remotely called on the client, or may be deployed and directly called on a client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence may be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.

[0061] Unless explicitly specified, the various data involved in this application can be stored remotely on a server or on a local terminal device, as long as they are suitable for being called by the technical solution of this application.

[0062] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus exhibit commonality, unless otherwise specified, these methods can be independently executed. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept. Therefore, concepts with the same expression, as well as concepts that are appropriately transformed for convenience despite different expression, should be understood as equivalent.

[0063] Unless expressly stated to be mutually exclusive, the various embodiments disclosed in this application may be cross-combined with the relevant technical features of the various embodiments to flexibly construct new embodiments, as long as such combination does not deviate from the creative spirit of this application and can meet the needs of the prior art or resolve certain deficiencies in the prior art. Those skilled in the art should be aware of such flexibility.

[0064] A text translation method of the present application can be programmed as a computer program product and deployed in a client or server for execution. For example, in the e-commerce platform application scenario of the present application, it is generally deployed in a server for implementation. The method can be executed by accessing an interface opened after the computer program product is run and performing human-computer interaction with the process of the computer program product through a graphical user interface.

[0065] See also Figure 1 In a typical embodiment, the text translation method of the present application includes the following steps:

[0066] Step S1100: obtaining a text to be translated;

[0067] For e-commerce platforms, the need to translate the text to be translated may occur in various specific business scenarios. For example, when an e-commerce platform user needs to translate the text on the product display interface, such as the product title text, product details text, etc.; or when an e-commerce platform user enters text in the product search input box and submits it to the server, etc., all of these can trigger the translation of the text to be translated. It is not difficult to understand that according to different specific business scenarios, the source of the text to be translated is also different, and those skilled in the art should be aware of this.

[0068] Those skilled in the art will understand that when an e-commerce platform user triggers the translation of a text to be translated, the corresponding server and client of the e-commerce platform can obtain the text to be translated and perform translation processing.

[0069] Step S1200: Using a text translation model pre-trained to a convergent state, the text to be translated is translated to obtain a result text, wherein the text to be translated is a text expressed in a first language or a second dialect of a second language, and correspondingly, the result text is a text expressed in the second dialect of the second language or a text expressed in the first language. The training process of the text translation model includes the following steps:

[0070] The text to be translated is a text expressed in a specific regional language of Traditional Chinese or a text expressed in English. It can be understood that the expression type depends on the expression type corresponding to the language used locally by users of the e-commerce platform.

[0071] When the server is the translation execution end, it obtains the text to be translated submitted by the client, and calls the text translation model to perform translation processing on the text to be translated; when the client is the translation execution end, it downloads the text translation model from the server in advance and deploys it locally, so that when obtaining the text to be translated, it calls the text translation model to perform translation processing. The text translation model is a model pre-trained to a convergence state. Its specific training process will be disclosed in the subsequent steps of this typical embodiment, and this step will not be discussed for the time being.

[0072] Step S1210: Calling a first training sample in a first training data set to perform iterative pre-training on the text translation model to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus, and the second language corpus is a result of replacing individual words in a translation text of the first language corpus corresponding to a first dialect of a second language with synonyms in the second dialect of the second language;

[0073] The first language corpus is English, and the translation text of the first dialect of the second language corresponding to the first language corpus is a translation text expressed in a specific regional language of simplified Chinese. The second dialect of the second language is a specific regional language of traditional Chinese. The second language corpus is the result of replacing individual words in the translation text of the English corpus in the specific regional language of simplified Chinese with synonyms in the specific regional language of simplified Chinese. The individual words include but are not limited to words corresponding to product attribute words and product brand words. For example, the individual words are product attribute words. Accordingly, the English language corpus is "spicy instant noodles", and its corresponding translation text in the specific regional language of simplified Chinese is "spicy instant noodles". The individual word "instant noodles" in the translation text is replaced with the synonym "instant noodles" in the specific regional language of traditional Chinese. The result of the replacement is "spicy instant noodles" as the second language corpus. It is easy to understand that the first language corpus and the second language corpus have the same semantics, and the two constitute the first parallel corpus. The first training sample is the first parallel corpus, and the first training set includes multiple first training samples.

[0074] Those skilled in the art will appreciate that the first language corpus and its corresponding second language first dialect translation constitute a parallel corpus. This parallel corpus resource is a high-resource corpus, meaning that the existing parallel corpus resources are relatively abundant and publicly available through relevant agencies such as the United Nations' public news and language research laboratories, making it easy to collect and apply. Therefore, it can be understood that the first parallel corpus is also a high-resource corpus.

[0075] The iterative pre-training is to iteratively train the text translation model on the data of the first training sample according to a set number of iterations. Specifically, first, the text translation model performs forward propagation, that is, along the order from the input layer to the output layer of the translation model, the weights and bias terms of the text translation model are calculated and stored in sequence. Secondly, the text translation model performs backward propagation, that is, along the order from the output layer to the input layer of the translation model, the gradients corresponding to the weights and bias terms of the text translation model are calculated and stored in sequence. Therefore, in each iterative training, according to the mutual dependence of the forward propagation and the forward propagation and the loss values corresponding to each layer of the text translation model, the weights and bias terms of the text translation model are continuously adjusted so that their corresponding gradient descent trends are obtained.

[0076] After the text translation model calls the first training sample in the first training data set to implement the iterative training, the overall loss value of the text translation model no longer changes or changes extremely slowly, which indicates that the text translation model has converged, thereby obtaining a pre-trained to converged text translation model.

[0077] The first parallel corpus structure includes but is not limited to a word pair structure, a parallel phrase pair structure, and a parallel syntactic structure.

[0078] For the specific construction of the first training sample in the first data set, the identification and replacement of the individual words, reference may be made to other embodiments to be disclosed later in this application, and this step will not be discussed here for the time being.

[0079] In summary, the pre-trained to convergent text translation model has the ability to extract semantic features corresponding to the language expression characteristics such as word order and grammar of the text to be translated and its corresponding translated text with the help of high-resource parallel corpus with relatively rich corpus resources. It also has the ability to extract semantic features of individual words in the text to be translated, so that the text to be translated can be translated using the language expression characteristics of the first dialect of the second language, and in particular, individual words in the text to be translated can be translated into the second dialect of the second language.

[0080] Step S1220: Call the second training sample in the second training data set to perform iterative fine-tuning training on the text translation model to achieve convergence, wherein the second training sample is a second parallel corpus consisting of the first language corpus and the second dialect corpus of the second language.

[0081] The second training set includes a plurality of second training samples. The second parallel corpus structure includes, but is not limited to, a word pair structure, a parallel phrase pair structure, and a parallel syntactic structure. The second training samples are a second parallel corpus consisting of an English language corpus and a corpus in a specific regional language using Traditional Chinese, where the English language corpus and the corpus in the specific regional language using Traditional Chinese have the same semantics.

[0082] Those skilled in the art should be aware that such parallel corpus resources disclosed by relevant agencies such as the United Nations' public news and language research laboratories can be collected and put into use as the second parallel corpus; texts corresponding to product titles and product details from major cross-border e-commerce platforms can also be collected and then handed over to professionals who use traditional Chinese for translation of specific regional languages for manual translation, thereby constructing an English language corpus and its corresponding specific regional language corpus using traditional Chinese, and the two together constitute the second parallel corpus for use.

[0083] It can be seen from the typical embodiments of the present application that the technical solution of the present application has many advantages, including but not limited to the following aspects:

[0084] First, the present application uses a large amount of first parallel corpus in the first training data set for text translation model training, so that the model learns the basic language expression ability such as word order and grammar of the second language and first dialect with relatively rich corpus resources; on this basis, since individual words in the translated text of the second language and first dialect in the first parallel corpus in the first training data set are replaced with synonyms of the second language and second dialect, the text translation model is pre-trained with the first training data set, so that the text translation model can accurately extract synonym features and translate the text to be translated into the expression words of the second language and second dialect with relatively low corpus resources, so that the translation effect is more grounded, which is convenient for users in the location of the second dialect to understand, efficiently assists users to use, improves user experience, and increases user stickiness of the e-commerce platform.

[0085] Secondly, the second parallel corpus in the second training dataset is composed of the first language corpus and the second language second dialect corpus. Although its corpus resources may be relatively small, the text translation model has already acquired a certain translation capability during the pre-training phase. Therefore, after fine-tuning the text translation model using the second training dataset, the basic Chinese language expression of the model is further adjusted to the language expression of the second dialect, thereby improving the translation fidelity and fluency of the model and achieving accurate mutual translation between the first language and the second language second dialect.

[0086] In addition, the technical solution of the present application, on the one hand, uses a simple and easy-to-implement text translation model architecture with low computational cost and low load, making it suitable for deployment on the client side and applicable to a variety of e-commerce environment application scenarios. On the other hand, the execution efficiency is higher and the implementation cost is lower, and it is also suitable for deployment on the background server to respond to massive concurrent demands, thereby obtaining economies of scale.

[0087] In a further embodiment, before step S1200, the training process of the text translation model, the following steps are included:

[0088] Step S1000: Acquire a first translated text in a first dialect of a second language and a second translated text in a second dialect corresponding to each original text in a first language in a preset data set, as a first corpus and a second corpus respectively;

[0089] The preset data set includes multiple original texts expressed in the first language. The original texts can be collected by technical personnel in this field, including titles, details, and comments of products expressed in the first language, i.e., English, from major cross-border e-commerce platforms. For example, 250,000 products can be collected, each of which has one title, one detail, and five comments, totaling 1.75 million original text data expressed in the first language.

[0090] Each original text expressed in the first language in the preset data set is translated into a first dialect of the second language and a second dialect of the second language to obtain corresponding first translated texts and second translated texts, and a first corpus and a second corpus are constructed respectively based on them.

[0091] Step S1010: comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between different translated versions of the synonyms into a special dictionary;

[0092] In order to accurately compare the differences between different translated texts of the same original text in the first corpus and the second corpus and eliminate the influence of simplified and traditional fonts, first, the translated texts in the second corpus are converted from simplified to traditional to obtain translated texts with the same font as the translated texts in the first corpus, that is, traditional Chinese is converted into simplified Chinese. Secondly, the translated texts in the second corpus are compared with the corresponding translated texts in the first corpus, and the differences between the two are compared to obtain synonyms that are different in the translated texts of the two. It is not difficult to understand that the synonyms are semantically consistent but there are differences in the expressions corresponding to the localized expressions in different languages. Furthermore, a one-to-one mapping relationship is established between the corresponding synonyms in the first corpus and the second corpus, and then this is constructed into a special dictionary.

[0093] Step S1020: calling a corpus resource library, the corpus resource library including a first parallel corpus, the first parallel corpus consisting of two corresponding texts: a corpus in a first language and a corpus in a first dialect of a second language;

[0094] The corpus resource library includes a first parallel corpus, which is composed of two corresponding texts: a first language corpus and a second language first dialect corpus. The two texts are semantically consistent and are translated texts of each other.

[0095] The first parallel corpus in the corpus resource library can be directly used here by collecting such parallel corpus resources disclosed by relevant agencies such as public news from the United Nations, language research laboratories, etc., or by collecting the first corpus or the corpus of the first dialect of the second language and manually translating it accordingly to construct the parallel corpus. Those skilled in the art can flexibly select it so that it can be called in this step.

[0096] Step S1030: Based on the mapping relationship data of the special dictionary, the translated versions of the synonymous words belonging to the first dialect of the second language in the first parallel corpus are replaced with the translated versions of the synonymous words in the second dialect in the special dictionary to form a second language corpus, so that the corpus resource library constitutes a first training data set.

[0097] According to the one-to-one mapping relationship between the translated versions of the synonyms in the second language and the first dialect and the second language in the special dictionary, all the synonyms in the second language and the first dialect in the first parallel corpus of the corpus resource library are replaced with the synonyms in the second language and the second dialect, so as to complete the replacement of the translated versions of the synonyms. At this point, the corpus resource library constitutes a first training data set.

[0098] In this embodiment, a special dictionary is constructed to complete the replacement of translation versions of synonyms in the first parallel corpus of the corpus resource library to form a first training data set for the text translation model to perform pre-training, so that the text translation model can accurately extract the features of synonyms and translate them into expressions in the second language and the second dialect, thereby achieving localized and accurate translation.

[0099] In a further embodiment, before step S1200, the training process of the text translation model, the following steps are further included:

[0100] Step S1001: construct each original text expressed in a first language in the preset data set and its corresponding translated text in a second language and a second dialect into a parallel corpus, and construct the parallel corpus into a second training data set.

[0101] For each original text expressed in the first language in the preset data set, a person skilled in the art can collect the titles, details, and comment texts of products expressed in the first language on major cross-border e-commerce platforms. For example, 250,000 products can be collected, of which each product has one title, one detail, and five comments, totaling 1.75 million original text data expressed in the first language. Furthermore, in one embodiment, a pre-packaged translation interface disclosed in the prior art can be called to translate the original text data expressed in the first language into a translated text expressed in a second dialect of a second language, and then the translated text is submitted to a professional for manual review and correction to ensure the accuracy of the translation. Afterwards, a one-to-one mapping relationship is established between the translated text and the original text, and the two are constructed into the parallel corpus based on this, and the respective parallel corpora are constructed as the second training data set.

[0102] In this embodiment, a training data set is constructed through manual annotation, so that the text translation model can extract more accurate semantic features at the second language and second dialect expression level during fine-tuning training, thereby improving the fidelity and fluency of the model translation.

[0103] In a further embodiment, step S1010, comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between the different translated versions of the synonyms into a special dictionary, includes the following steps:

[0104] Step S1011, comparing the difference texts between different translated texts of the same original text in the first corpus and the second corpus;

[0105] The first corpus and the second corpus respectively store different translated texts of the same original text, i.e., the English expression text. Correspondingly, the translated text of the first corpus is a translated text of the first dialect of the second language, i.e., expressed in a specific region using simplified Chinese, and the translated text of the first corpus is a translated text of the second dialect of the second language, i.e., expressed in a specific region using traditional Chinese. Since the difference between the first corpus and the second corpus is compared to obtain the corresponding difference text, the difference is compared in terms of language expression rather than font. Therefore, the translated text of the first corpus or the second corpus is translated into the text expressed by the other so as to maintain the consistency of the glyph types of the two. Further, on this basis, the consistency between the translated texts of the first corpus or the second corpus is compared to obtain the difference text. For the above exemplary example, the original text is "I li ke instant nood l es”, the corresponding translation text of the first corpus is “I love instant noodles”, and the translation text of the second corpus is “I love instant noodles”. Therefore, the difference text obtained by comparison is “instant noodles” corresponding to the first corpus and “instant noodles” in the second corpus.

[0106] Step S1012: using an entity noun classification model that has been pre-trained to a convergent state to judge the difference text in the first corpus to determine whether it is a specific type of noun;

[0107] The difference text in the first corpus is input into an entity noun classification model that has been pre-trained to a convergent state to extract the corresponding semantic features of the type, and then the similarity between the semantic features of the type and the semantic features corresponding to the specific type pre-extracted and stored by the model is calculated to obtain a normalized similarity result between the two. Based on the similarity result, it is determined whether the type corresponding to the noun is a specific type such as a product attribute word, a product brand word, etc. The similarity result judgment standard can be set to whether it reaches 0.7 or above. The specific numerical setting can be flexibly set by those skilled in the art.

[0108] Step S1013: Determine the difference texts in the first corpus and the second corpus corresponding to nouns of a specific type as different translation versions of synonyms, construct them into mapping relationship data, and store them in a pre-constructed special dictionary.

[0109] Furthermore, the difference texts in the first corpus corresponding to nouns of a specific type and the difference samples in the second corpus corresponding thereto are determined to be different translation versions of synonyms, and then the two are correspondingly constructed into one-to-one mapping relationship data, and the mapping relationship data is stored in a pre-constructed special dictionary.

[0110] In this embodiment, the difference texts corresponding to specific types of nouns in the first feature library and the second feature library are screened out and correspondingly constructed as mapping relationship data and stored in a special dictionary, so that the model trained according to the first training set constructed by the special dictionary in the subsequent steps can extract special types of nouns such as product attribute words and product brand words in the text to be translated, and translate them into nouns expressed in the second language and the second dialect. In this way, the products expressed in the first language and the second language and the second dialect on the cross-border e-commerce platform can be marked with labels such as product attributes and brands expressed in the other language or dialect, which is convenient for assisting the implementation of various e-commerce application scenarios such as subsequent accurate product recommendations and accurate product searches, and further deepens the localization level of translation and improves the accuracy of translation.

[0111] In a further embodiment, step S1030, based on the mapping relationship data of the special dictionary, replacing the translated versions of the synonymous words belonging to the first dialect of the second language in the first parallel corpus with the translated versions of the synonymous words in the second dialect in the special dictionary to form the second language corpus, includes the following steps:

[0112] Step S1031: searching for corresponding target texts in the full parallel corpus in the corpus resource library according to the translated versions of the synonyms corresponding to the first dialect of the second language in the special dictionary;

[0113] According to the translation version of the synonym corresponding to the first dialect of the second language in the special dictionary, the parallel corpus in which the synonym appears in the full parallel corpus in the corpus resource library is retrieved, thereby obtaining the target text corresponding to the synonym from the parallel corpus, and the parallel corpus structure includes but is not limited to word pair structure, parallel phrase pair structure, and parallel syntactic structure.

[0114] Step S1032: Obtain a translation version of the second dialect of the second language corresponding to the synonym in the special dictionary, and randomly select a portion of the target text to replace it with the translation version of the second dialect.

[0115] Then, using the target text corresponding to the synonymous word in step S1031 as an index and a one-to-one mapping relationship between the target text and the synonymous word expressed in the second dialect of the second language, the text of the translated version of the synonymous word in the second dialect of the second language in the special dictionary is obtained. Further, a synonym replacement operation is performed to replace the translated version of the synonymous word, and a parallel corpus in which the translated version of the synonymous word in the first dialect of the second language appears in the full parallel corpus in the corpus resource library is randomly selected, and the synonymous word therein is replaced with the synonymous word in the translated version of the second dialect in the aforementioned second language.

[0116] In this embodiment, by randomly selecting parallel corpora from the corpus resource library and replacing the translation versions of synonyms therein, when applying them to model training, the model's recognition of synonyms is improved, which helps improve the model's generalization ability.

[0117] In an extended embodiment, step S1200, translating the text to be translated using a text translation model pre-trained to a convergent state, and obtaining a result text, includes the following steps:

[0118] Step S1300: In response to the text replacement instruction acting on the text to be translated, the text to be translated is replaced with the result text for display on the current interface for obtaining the text to be translated;

[0119] In the present application, this step can be added to realize that the result text corresponding to the translation of the text to be translated is displayed on the corresponding interface. Specifically, a translation control can be provided in the graphical user interface of the computer program product of the present application, and when the user touches the translation control, the corresponding text replacement instruction can be triggered. The graphical user interface can be a product display graphical user interface displayed on the user terminal device. Specifically, for example, it can be set in the function selection area of the product display graphical user interface, or it can be in other unmentioned interfaces, as long as it is accessible to the client user.

[0120] When browsing the product display graphical user interface, an e-commerce platform user can touch a "translation" control provided in the product display interface as shown in Figure 5(a)100 or Figure 5(b)200 as needed, thereby triggering the text replacement instruction acting on the text to be translated. In response to the text replacement instruction, in this embodiment, the text to be translated is stored in the cache. At the same time, after obtaining the text to be translated and inputting it into the text translation model, the result text translated by the text translation model is then replaced by the result text and displayed at the corresponding position in the product display interface as shown in Figure 5(a)101.

[0121] Step S1400: In response to a pointing event acting on the display area of the result text, a text prompt box is displayed on the current interface, and the text to be translated is displayed in the text prompt box.

[0122] In the present application, this step can be added to realize that the text to be translated corresponding to the result text before it is translated is displayed on the corresponding interface. Specifically, a text prompt box can be provided in the graphical user interface of the computer program product of the present application, and the user can trigger the corresponding pointing event after pointing to the result text display area in the graphical user interface. The graphical user interface can be a product display graphical user interface displayed on the user terminal device. For example, it can be set in the product title text area of the product display graphical user interface, or it can be in other unmentioned interfaces, as long as it is accessible to the client user.

[0123] When browsing the product display graphical user interface that displays the result text, the user of the e-commerce platform can point to the display area of the result text in the product display interface as needed, as shown in Figure 5(a)101, thereby triggering the pointing event acting on the display area of the result text. In response to the event, the corresponding translation text of the result text before it is translated, i.e., the text to be translated, is obtained from the cache and displayed in a text prompt box preset in the display area of the result text, as shown in Figure 5(a)102. The text prompt box can be provided with a corresponding close control in the prompt box for the user to manually close the prompt box, or the prompt box can be set to close immediately when the display area loses the pointing. Those skilled in the art can flexibly select and implement it.

[0124] In this embodiment, the text to be translated is replaced by the corresponding result text after the translation of the text to be translated at the corresponding location on the product display interface to achieve a text translation effect on the interface. At the same time, the text to be translated is cached, so that when the display area of the result text on the product display interface is pointed to by the user, the corresponding text to be translated is quickly and directly placed in a text prompt box near the display area without the user having to wait for the loading time of the secondary translation, making it convenient for the user to know the text before and after the translation, and also convenient for the user to compare the text before and after the translation.

[0125] See also Figure 7, a text translation device provided to meet one of the purposes of the present application is a functional embodiment of the text translation method of the present application, the device includes: a text acquisition module 1100, a model translation module 1200, a model pre-training module 1300, and a model fine-tuning module 1400, wherein the text acquisition module 1100 is used to acquire the text to be translated; the model translation module 1200 is used to translate the text to be translated using a text translation model pre-trained to a convergent state to obtain a result text, wherein the text to be translated is a text expressed in a first language or a second dialect of a second language, and correspondingly, the result text is a text expressed in a second dialect of the second language or a text expressed in the first language, and the text translation module 1200 is used to translate the text to be translated using a text translation model pre-trained to a convergent state to obtain a result text. The training process of the translation model includes: the model pre-training module 1300 is used to call the first training sample in the first training data set to perform iterative pre-training on the text translation model to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus, and the second language corpus is the result of replacing individual words in the translation text of the first language corpus corresponding to the first dialect of the second language with synonyms in the second dialect of the second language; the model fine-tuning module 1400 is used to call the second training sample in the second training data set to perform iterative fine-tuning training on the text translation model to achieve convergence, wherein the second training sample is a second parallel corpus consisting of the first language corpus and the second dialect of the second language.

[0126] In a further embodiment, before the training process of the text translation model, the method includes: a bidirectional translation module for obtaining a first translation text in a first dialect of a second language and a second translation text in a second dialect corresponding to each original text expressed in the first language in a preset data set, as a first corpus and a second corpus, respectively; a dictionary construction module for comparing synonyms between different translation texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between different translation versions of the synonyms into a special dictionary; a resource library calling module for calling a corpus resource library, which includes a first parallel corpus, which is composed of two corresponding texts of a corpus in the first language and a corpus in the first dialect of the second language; and a synonym replacement module for replacing the translated versions of synonyms belonging to the first dialect of the second language in the first parallel corpus with translated versions of synonyms in the second dialect in the special dictionary according to the mapping relationship data of the special dictionary to form a second language corpus, so that the corpus resource library constitutes the first training data set.

[0127] In a further embodiment, before the training process of the text translation model, it also includes: a corpus construction module, which is used to construct each original text expressed in the first language in the preset data set and its corresponding translation text in the second language and second dialect into a parallel corpus, and construct the parallel corpus into a second training data set.

[0128] In a further embodiment, the dictionary construction module includes: a difference comparison submodule, which is used to compare the difference texts between different translated texts of the same original text in the first corpus and the second corpus; a type judgment submodule, which is used to use an entity noun classification model pre-trained to a convergent state to judge the difference texts in the first corpus to determine whether they are nouns of a specific type; a relationship construction submodule, which is used to determine the difference texts in the first corpus and the second corpus corresponding to nouns of a specific type as different translation versions of synonyms, construct them into mapping relationship data, and store them in a pre-constructed special dictionary.

[0129] In a further embodiment, the synonym replacement module includes: a target retrieval submodule, which is used to retrieve the corresponding target text in the entire corpus resource library based on the translation version of the synonym corresponding to the first dialect of the second language in the special dictionary; a random replacement submodule, which is used to obtain the translation version of the second dialect of the second language corresponding to the synonym in the special dictionary, and randomly select part of the target text to replace it with the translation version of the second dialect.

[0130] In an extended embodiment, the model translation module 1200 includes: a text replacement submodule for responding to a text replacement instruction acting on the text to be translated, replacing the text to be translated with the result text for display on the current interface for obtaining the text to be translated; and a text prompt submodule for responding to a pointing event acting on the display area of the result text, displaying a text prompt box on the current interface, and displaying the text to be translated in the text prompt box.

[0131] In order to solve the above technical problems, the embodiment of the present application also provides a computer device. Figure 8As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions, and the database may store a control information sequence, and when the computer-readable instructions are executed by the processor, the processor may implement a text translation method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor may execute the text translation method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0132] In this embodiment, the processor is used to execute Figure 7 The memory stores the program code and various data required to execute the modules and submodules. The network interface is used to transmit data between user terminals and servers. The memory in this embodiment stores the program code and data required to execute all modules and submodules in the text translation device of this application. The server can call the server's program code and data to execute the functions of all submodules.

[0133] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the text translation method of any embodiment of the present application.

[0134] The present application also provides a computer program product, comprising a computer program / instruction, which implements the steps of the method described in any embodiment of the present application when executed by one or more processors.

[0135] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the method. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0136] To sum up, first, the present application uses a large amount of first parallel corpus in the first training data set for text translation model training, so that the model learns the basic language expression ability such as word order and grammar of the second language and first dialect with relatively rich corpus resources; on this basis, since individual words in the translated text of the second language and first dialect in the first parallel corpus in the first training data set are replaced with synonyms of the second language and second dialect, the text translation model is pre-trained with the first training data set, so that the text translation model can accurately extract synonym features and translate the text to be translated into the expression words of the second language and second dialect with relatively low corpus resources, so that the translation effect is more grounded, which is convenient for users in the second dialect to understand, efficiently assists users in use, improves user experience, and increases user stickiness of the e-commerce platform.

[0137] Secondly, the second parallel corpus in the second training dataset is composed of the first language corpus and the second language second dialect corpus. Although its corpus resources may be relatively small, the text translation model has already acquired a certain translation capability during the pre-training phase. Therefore, after fine-tuning the text translation model using the second training dataset, the basic Chinese language expression of the model is further adjusted to the language expression of the second dialect, thereby improving the translation fidelity and fluency of the model and achieving accurate mutual translation between the first language and the second language second dialect.

[0138] Those skilled in the art will appreciate that the steps, measures, and schemes in the various operations, methods, and processes discussed in this application may be interchanged, modified, combined, or deleted. Furthermore, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be interchanged, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and schemes in the prior art that are similar to those disclosed in this application may also be interchanged, modified, rearranged, decomposed, combined, or deleted.

[0139] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A text translation method, characterized in that: The steps include: Get the text to be translated; The text to be translated is translated using a text translation model that has been pre-trained to a convergent state to obtain a result text, wherein the text to be translated is a text expressed in a first language or a second dialect of a second language, and correspondingly, the result text is a text expressed in the second dialect of the second language or a text expressed in the first language. The training process of the text translation model includes the following steps: Iterative pre-training of the text translation model is performed using a first training sample in a first training dataset to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus, and the second language corpus is a translation of the first language corpus into a first dialect of a second language with individual words replaced by synonyms in the second dialect of the second language; Calling a second training sample in a second training data set to perform iterative fine-tuning training on the text translation model to achieve convergence, wherein the second training sample is a second parallel corpus consisting of a corpus in the first language and a corpus in a second dialect of the second language; The text translation model training process includes: Obtaining a first translated text in a first dialect of a second language and a second translated text in a second dialect corresponding to each original text expressed in the first language in the preset data set, as a first corpus and a second corpus respectively; Comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between the different translated versions of the synonyms into a special dictionary; Calling a corpus resource library, the corpus resource library includes a first parallel corpus, the first parallel corpus consisting of two corresponding texts: a first language corpus and a second language first dialect corpus; According to the mapping relationship data of the special dictionary, the translated versions of the synonymous words belonging to the first dialect of the second language in the first parallel corpus are replaced with the translated versions of the synonymous words in the second dialect in the special dictionary to form the second language corpus, so that the corpus resource library constitutes the first training data set, including: searching the corresponding target text in the full parallel corpus in the corpus resource library according to the translated versions of the synonyms corresponding to the first dialect of the second language in the special dictionary; A translation version of the second dialect of the second language corresponding to the synonym in a special dictionary is obtained, and a randomly selected portion of the target text is replaced with the translation version of the second dialect.

2. The text translation method according to claim 1, characterized in that Before the training process of the text translation model, the following steps are also included: Each original text expressed in the first language in the preset data set and its corresponding translated text in the second language and second dialect are constructed into a parallel corpus, and the parallel corpus is constructed into a second training data set.

3. The text translation method according to claim 1, characterized in that Comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between the different translated versions of the synonyms into a special dictionary, includes the following steps: comparing the difference texts between different translations of the same original text in the first corpus and the second corpus; Using an entity noun classification model that has been pre-trained to a convergent state, the difference text in the first corpus is judged to determine whether it is a specific type of noun; The difference texts in the first corpus and the second corpus corresponding to nouns belonging to a specific type are determined as different translation versions of synonyms, constructed into mapping relationship data, and stored in a pre-constructed special dictionary.

4. The text translation method according to claim 1, characterized in that The text to be translated is translated using a text translation model that has been pre-trained to a convergent state. After obtaining the result text, the following steps are included: In response to the text replacement instruction acting on the text to be translated, the text to be translated is replaced with the result text for display on a current interface for obtaining the text to be translated; In response to a pointing event acting on the display area of the result text, a text prompt box is displayed on the current interface, and the text to be translated is displayed in the text prompt box.

5. A text translation device, characterized in that: include: A text acquisition module is used to obtain the text to be translated; The model translation module is used to translate the text to be translated using a text translation model that has been pre-trained to a convergent state to obtain a result text. The text to be translated is a text expressed in a first language or a second dialect of a second language. Accordingly, the result text is a text expressed in the second dialect of the second language or a text expressed in the first language. The training process of the text translation model includes the following steps: a model pre-training module configured to perform iterative pre-training on the text translation model by calling a first training sample in a first training dataset to achieve convergence, wherein the first training sample is a first parallel corpus consisting of a first language corpus and a second language corpus, and the second language corpus is a translation of the first language corpus into a first dialect of a second language with individual words replaced by synonyms in the second dialect of the second language; a model fine-tuning module, configured to call a second training sample in a second training data set to perform iterative fine-tuning training on the text translation model to achieve convergence, wherein the second training sample is a second parallel corpus consisting of a corpus in the first language and a corpus in a second dialect of the second language; The text translation model training process includes: Obtaining a first translated text in a first dialect of a second language and a second translated text in a second dialect corresponding to each original text expressed in the first language in the preset data set, as a first corpus and a second corpus respectively; Comparing synonyms between different translated texts of the same original text in the first corpus and the second corpus, and constructing mapping relationship data between the different translated versions of the synonyms into a special dictionary; Calling a corpus resource library, the corpus resource library includes a first parallel corpus, the first parallel corpus consisting of two corresponding texts: a first language corpus and a second language first dialect corpus; According to the mapping relationship data of the special dictionary, the translated versions of the synonymous words belonging to the first dialect of the second language in the first parallel corpus are replaced with the translated versions of the synonymous words in the second dialect in the special dictionary to form the second language corpus, so that the corpus resource library constitutes the first training data set, including: searching the corresponding target text in the full parallel corpus in the corpus resource library according to the translated versions of the synonyms corresponding to the first dialect of the second language in the special dictionary; A translation version of the second dialect of the second language corresponding to the synonym in a special dictionary is obtained, and a randomly selected portion of the target text is replaced with the translation version of the second dialect.

6. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that It stores a computer program implemented according to the method described in any one of claims 1 to 4 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • E-commerce title translation method and corresponding device, equipment and medium

    CN113435214A