System and method for training a translation model using examples of source expansion training

By training a translation model with source augmentation using labels like domains and URLs, the model can emulate high-quality translation styles and reduce computational complexity, addressing the variability in training data quality and domain-specific translation needs.

JP2025520752AActive Publication Date: 2025-07-03GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024575758
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-07-03
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

Neural machine translation models are affected by the quality and quantity of training data, which is often of varying quality and requires human supervision, making it difficult to ensure consistent translation quality across different domains.

Method used

A translation model is trained using source augmentation training examples, where labels such as Internet domains, subdomains, URLs, or IP addresses are included to associate translation styles with their sources, allowing the model to emulate high-quality sources and reduce the need for manual data filtering.

Benefits of technology

This approach enables the generation of a translation model that can efficiently emulate different translation qualities and styles by adjusting source labels during inference, reducing computational complexity and the need for multiple domain-specific models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520752000001_ABST
    Figure 2025520752000001_ABST
Patent Text Reader

Abstract

A system and method for training a translation model based on a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on the second text sequence. In some examples, the label may include an internet domain, an internet subdomain, a uniform resource locator, a website name, or an IP address. In some examples, the label may further indicate the source of the first text sequence. In some examples, each given training example may be automatically generated by sampling a first text sequence from a first page of a given internet domain, sampling a second text sequence from a second page of the given internet domain, and generating a label based on all or part of the source data of the second page.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The quality of translations generated by neural machine translation models can be affected by both the quantity and quality of the data used to train the models. Unfortunately, while large amounts of training data can be collected using various automated methods, it can be difficult to guarantee the quality of such data, and often requires human supervision. For example, a system may be configured to crawl the Internet to identify sets of pages published in multiple languages (e.g., pages from domains en.websight.com and es.website.com may have the same content published in English and Spanish respectively), and separate corresponding sequences of text from which training examples can be generated. However, training examples from some websites or web pages may be of relatively high or low quality depending on various factors, such as whether the translation was created or supervised by a human translator, whether the translation is more concise or more verbose, etc. Similarly, training examples from some websites or web pages may use specific jargon, making them more or less desirable for training a given translation model (e.g., web pages targeted at a specific region may use region-specific dialects, and web pages targeted at scientific or legal content may use terms that have different meanings in non-scientific or non-legal contexts, etc.).

Summary of the Invention

[0002] The present technology relates to a system and method for training a translation model using source augmentation training examples so that a model can learn to associate a particular translation style with the source of each example. For example, in some aspects of the present technology, a translation model can be trained based on a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on the source of the second text sequence. In some aspects, the label can include an Internet domain, an Internet subdomain, a Uniform Resource Locator (“URL”), a website name, or an IP address related to the source of the second text sequence. Similarly, in some aspects, the label can further indicate the source of the first text sequence. Further, in some aspects of the present technology, each given training example of a plurality of training examples can be automatically generated by sampling a first text sequence from a first page of a given Internet domain, sampling a second text sequence from a second page of the given Internet domain, and generating a label based on the source of the second text sequence and / or the first text sequence (e.g., all or part of the URL, Internet domain, Internet subdomain, website name, or IP address of the first and / or second page).

[0003] Accordingly, the present technology can generate a translation model that can encourage emulating, during inference, a translation of a particular high-quality source or otherwise desirable source by simply including a label of that source along with the input text sequence. These high-quality or desirable sources can be identified after training by repeatedly providing a validation set of examples to the trained translation model using different labels and comparing the quality of the generated translations (e.g., using an automatic quality metric, human scorers, or a combination thereof). In this way, the present technology can reduce or eliminate the amount of filtering required for a given set of training data, thereby enabling training of the translation model using a large dataset of automatically collected, generated, and / or filtered synthetic training examples. Similarly, the present technology can be used to generate a translation model that can be flexibly and efficiently "tuned" to emulate different translation qualities and / or styles by simply changing which source labels are used during inference. Accordingly, the present technology can solve the technical problem of how to control the output of a translation model trained on multiple sources or domains to generate a translation based on characteristics of a particular source or domain of interest. Further, in various exemplary embodiments, this can be achieved by training only a single model (rather than one or more models for each domain of interest), thereby reducing technical complexity and computational cost.

[0004] In one aspect, the present disclosure is about training a translation model, and the training involves: (1) for each given training example among a plurality of training examples, where the given training example includes a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on the source of the second text sequence, using the translation model to generate a predicted text sequence based at least in part on the first text sequence and the label of the given training example, and using one or more processors of a processing system to compare the predicted text sequence with the second text sequence to generate a loss value for the given training example; and (2) using one or more processors to modify one or more parameters of the translation model based at least in part on the loss values generated for each of the plurality of training examples. Some aspects, the label includes an Internet domain. Some aspects, the label includes an Internet subdomain. Some aspects, the label includes a uniform resource locator. Some aspects, the label includes a website name. Some aspects, the label includes an IP address. Some aspects, the label further indicates the source of the first text sequence. Some aspects, the source of the first text sequence is within a first subdomain of a given Internet domain, and the source of the second text sequence is within a second subdomain of the given Internet domain. Some aspects, the method further includes generating each given training example of the plurality of training examples by using one or more processors to sample a first text sequence from a first page of a given Internet domain, sampling a second text sequence from a second page of the given Internet domain, and generating a label based on all or part of the uniform resource locator of the second page.In some embodiments, the method further includes generating each given training example of a plurality of training examples by using one or more processors to sample a first text sequence from a first page of a given Internet domain, sampling a second text sequence from a second page of the given Internet domain, and generating a label based on all or part of the IP address of the second page.

[0005] In another embodiment, the present disclosure describes a computer program product that includes computer-readable instructions that, when executed by a processing system, cause the processing system to perform any of the methods described in the preceding paragraphs.

[0006] In another aspect, the present disclosure describes a processing system including: (1) a memory storing a translation model; and (2) one or more processors coupled to the memory and configured to train the translation model according to a training method including: for each given training example of a plurality of training examples, the given training example including a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on a source of the second text sequence, using the translation model to generate a predicted text sequence based at least in part on the first text sequence and the label of the given training example; comparing the predicted text sequence with the second text sequence to generate a loss value for the given training example; and modifying one or more parameters of the translation model based at least in part on the loss value generated for each of the plurality of training examples. In some aspects, the one or more processors are configured to train the translation model with each given training example including a label including an Internet domain according to the training method. In some aspects, the one or more processors are configured to train the translation model with each given training example including a label including an Internet subdomain according to the training method. In some aspects, the one or more processors are configured to train the translation model with each given training example including a label including a uniform resource locator according to the training method. In some aspects, the one or more processors are configured to train the translation model with each given training example including a label including a website name according to the training method. In some aspects, the one or more processors are configured to train the translation model with each given training example including a label including an IP address according to the training method.In some aspects, one or more processors are configured to train a translation model on each given training example that includes labels indicating the source of a first text sequence and the source of a second text sequence according to a training method. In some aspects, one or more processors are further configured to generate each given training example of a plurality of training examples by sampling a first text sequence from a first page of a given Internet domain, sampling a second text sequence from a second page of the given Internet domain, and generating a label based on all or part of the uniform resource locator of the second page. In some aspects, one or more processors are further configured to generate each given training example of a plurality of training examples by sampling a first text sequence from a first page of a given Internet domain, sampling a second text sequence from a second page of the given Internet domain, and generating a label based on all or part of the IP address of the second page.

[0007] In another aspect, the present disclosure includes (1) a memory storing a translation model, and (2) one or more processors coupled to the memory and configured to use the translation model to generate a predicted translation of an input text sequence based on the input text sequence and a label, wherein for each given training example of a plurality of training examples, the translation model includes: (a) a given training example including a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on a source of the second text sequence, and using the translation model to generate a predicted text sequence based at least in part on the first text sequence and the label of the given training example; comparing the predicted text sequence with the second text sequence to generate a loss value for the given training example; and (b) modifying one or more parameters of the translation model based at least in part on the loss value generated for each of the plurality of training examples. The one or more processors are trained to generate a predicted translation according to a training method including the above steps, and the present disclosure describes a processing system including the one or more processors.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0009] Next, the present technology will be described with respect to the following exemplary systems and methods. Common reference numerals between the figures illustrated and described below are intended to identify the same features.

[0010] Exemplary system FIG. 1 shows a high-level system diagram 100 of an exemplary processing system 102 for executing the method described herein. The processing system 102 may include one or more processors 104 and a memory 106 that stores instructions 108 and data 110. The instructions 108 and data 110 may include a translation model, as further described below. Additionally, the data 110 may store training examples (e.g., those used in pre-training, training, or fine-tuning), training signals and / or loss values generated during training, and / or predicted text sequences generated by the translation model that are used when training the translation model.

[0011] The processing system 102 may be present on a single computing device. For example, the processing system 102 may be a server, a personal computer, or a mobile device, and thus, the translation model may be local to that single computing device. Similarly, the processing system 102 may be present on a cloud computing system or other distributed system. In such a case, the translation model may be distributed across two or more different physical computing devices. For example, the processing system may include a first computing device that stores layers 1 to n of a translation model having m layers, and a second computing device that stores layers n to m of the translation model. In such a case, the first computing device may be a computing device (e.g., a personal computer, a mobile phone, a tablet, etc.) having less memory and / or processing power compared to the memory and / or processing power of the second computing device, or vice versa. Similarly, in some aspects of the present technology, (as further described below with respect to, for example, the exemplary method 500 of FIG. 5), the processing system may include one or more computing devices that store the translation model, and one or more separate computing devices configured to collect and / or generate training examples. Further, in some aspects of the present technology, data used by the translation model (e.g., training data, labels used during inference, etc.) may be stored on a computing device different from the translation model.

[0012] Furthermore, in this regard, FIG. 2 shows a high-level system diagram 200 in which the exemplary processing system 102 just described is distributed across two computing devices 102a and 102b, each of the two computing devices 102a and 102b including one or more processors (104a, 104b) and memories (106a, 106b) that store instructions (108a, 108b) and data (110a, 110b). A processing system 102 including computing devices 102a and 102b that communicate with one or more websites and / or remote storage systems via one or more networks 202 including a website 204 and a remote storage system 212 is shown. In this example, the website 204 includes one or more servers 206a-206n. Each of the servers 206a-206n may have one or more processors (e.g., 208) and an associated memory (e.g., 210) that stores instructions and data including the content of one or more web pages. Similarly, although not shown, the remote storage system 212 may also include one or more processors and a memory that stores instructions and data. In some aspects of the present technology, the processing system 102 including computing devices 102a and 102b may be configured to retrieve data from one or more of the website 204 and / or the remote storage system 212 for use in training a translation model. For example, in some aspects, the first computing device 102a may be configured to retrieve training examples from the remote storage system 212 for use in pre-training, training, or fine-tuning a translation model stored in the first computing device 102a and / or the second computing device 102b.Similarly, in some aspects, as further described below with respect to the exemplary method 500 of FIG. 5, the first computing device 102a may be configured to store a translation model, and the second computing device 102b may be configured to collect data from the website 204 and generate training examples based on the data retrieved for use in training the translation model. Further, in such a case, the second computing device 102b may be configured to store one or more of the generated training examples on the remote storage system 212 for retrieval by the first computing device 102a.

[0013] The processing systems described herein may be implemented on any type of computing device(s), such as any type of general-purpose computing device, server, or set thereof, and may further include other components typically present in general-purpose computing devices or servers. Similarly, the memory of such a processing system may be of any non-transitory type capable of storing information accessible by the processor(s) of the processing system. For example, the memory may include non-transitory media such as hard drives, memory cards, optical disks, solid state, tape memory, and the like. Computing devices suitable for the roles described herein may include the different combinations described above, whereby different portions of instructions and data are stored on different types of media.

[0014] In any case, the computing device described in this specification may further include any other components commonly used in connection with a computing device, such as a user interface subsystem. The user interface subsystem may include one or more user inputs (e.g., a mouse, keyboard, stylus, touch screen, and / or microphone), as well as one or more electronic displays (e.g., a monitor with a screen, or any other electrical device operable to display information). Output devices other than electronic displays, such as speakers, lights, and vibration elements, pulse elements, or tactile elements, may also be included in the computing device described in this specification.

[0015] One or more processors included in each computing device may be any conventional processor, such as a commercially available central processing unit (“CPU”), graphics processing unit (“GPU”), tensor processing unit (“TPU”), etc. Alternatively, one or more processors may be a dedicated device such as an ASIC or other hardware-based processor. Each processor may have multiple cores operable in parallel. The processor(s), memory, and other elements of a single computing device may be housed within a single physical housing or may be distributed among two or more housings. Similarly, the memory of a computing device may include a hard drive or other storage medium located within a housing different from that of the processor(s), such as within an external database or a networked storage device. Thus, a reference to a processor or a computing device is understood to include a collection of processors, computing devices, or memories that may or may not operate in parallel, and a reference to one or more servers of a load balancing server farm or a cloud-based system.

[0016] The computing devices described herein may store instructions (such as machine code) directly executable by a processor(s) or instructions (such as scripts) indirectly executable by a processor(s). The computing device may also store data that may be retrieved, stored, or modified by one or more processors according to the instructions. The instructions may be stored as computing device code on a computing device-readable medium. In that regard, the terms "instructions" and "program" may be used interchangeably herein. The instructions may also be in object code form for direct processing by a processor(s), or may include any other computing device language such as a script or a collection of independent source code modules that are interpreted on demand or pre-compiled. By way of example, the programming language may be C#, C++, JAVA®, or another computer programming language. Similarly, any component of the instructions or program may be implemented in a computer scripting language such as JavaScript®, PHP, ASP, or any other computer scripting language. Further, any one of these components may be implemented using a combination of a computer programming language and a computer scripting language.

[0017] Exemplary method FIG. 3 is a flowchart 300 showing how an exemplary training example can be generated based on pages of a website, according to an aspect of the present disclosure. In the example of FIG. 3, the website in question is assumed to be the exemplary website 204 of FIG. 2 described above. Further, the website 204 is assumed to include two web pages 302a and 302b. In this example, the web page 302a is from the URL "http: / / en.website.com / " and includes English text, and the web page 302b is from the URL "http: / / es.website.com / " and includes corresponding Spanish text. Thus, the web pages 302a and 302b are different subdomains of the same root domain (website.com).

[0018] FIG. 3 further shows a training example 304 that can be generated from the content of the web pages 302a and 302b. In this case, the training example 304 includes a first text sequence including a sentence from the web page 302a that states in English "This page is also available in other languages", a second text sequence including the corresponding sentence from the web page 302b that states in Spanish "Esta pagina esta disponible", and a label including a part of the URL of the web page 302b. As will be appreciated, it would also be possible to generate a second training example where the sentence from the web page 302b is the "first text sequence", the sentence from the web page 302a is the "second text sequence", and the label includes a part of the URL of the web page 302a.

[0019] The label for the example of FIG. 3 uses the full domain name of web page 302b, although the label may be based on any suitable information regarding the source of web page 302b. For example, in some embodiments, the label for training example 304 may include the full URL of web page 302b (e.g., http: / / es.website.com / ), the domain and / or subdomain of web page 302b (e.g., "es.website.com", "website.com", "es", "website", or "com"), the name of the website (e.g., "Website"), the IP address of web page 302b, and / or any other suitable information related to the source of web page 302b.

[0020] Similarly, although not reflected in the example of FIG. 3, in some embodiments of the present technology, the label may include information regarding the source of the first text sequence instead of, or in addition to, information based on the source of the second text sequence. For example, in some embodiments, the label for training example 304 may include information regarding the source of web page 302a, such as the full URL of web page 302a (e.g., http: / / en.website.com / ), the domain and / or subdomain of web page 302a (e.g., "en.website.com", "website.com", "en", "website", or "com"), the name of the website (e.g., "Website"), the IP address of web page 302a, and / or any other suitable information related to the source of web page 302a.

[0021] Furthermore, in some embodiments, the label for training example 304 may include information not directly related to web pages 302a and 302b. For example, if web pages 302a and 302b are obtained from a curated set of websites or web pages related to a particular topic (e.g., artificial intelligence, law, sports, etc.), the label for training example 304 may include information related to that topic (either alone or in addition to other source information).

[0022] Labels can be included in the training examples 304 in any suitable way and format. For example, in some aspects of the present technology, the labels can be prepended or appended to the input sequence as vector embeddings, tokenized text, or raw text (thus requiring no extra preprocessing or special vocabulary). In that regard, if the training examples are collected from sources with similar domain names, including the raw text of the domain name in each label can increase the likelihood that the translation model infers the similarity of the training examples of those domains.

[0023] The example of FIG. 3 assumes that the training examples 304 are generated based on text collected from web pages 302a and 302b, but it should be understood that the training examples can be generated from any suitable source available in multiple languages (such as books, user manuals, advertisements, song lyrics, etc.). Thus, as an example, the training examples can be generated from a first text sequence collected from a book, a second corresponding text sequence collected from a translated copy of the book, and a label indicating information based on sources such as the title of the book, the title of the translated copy, the name of the author, the name of the translator, etc.

[0024] FIG. 4 illustrates an exemplary method 400 for training a translation model according to an aspect of the present disclosure.

[0025] In step 402, the processing system (e.g., the processing system 102 of FIG. 1 or FIG. 2) selects a given training example from a plurality of training examples, where the given training example includes a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on the source of the second text sequence. The plurality of training examples may be from any suitable source or collection of sources. For example, the plurality of training examples may include training examples from an existing database of training data, examples generated or supervised by humans, and / or synthetically generated examples (e.g., generated according to the exemplary method 500 of FIG. 5). The label may also include any suitable information regarding the source of the second text sequence, including any of the options described above with respect to the training example 304 of FIG. 3.

[0026] Furthermore, although not reflected in the example of FIG. 4, in some aspects of the present technology, the label may also include other information instead of, or in addition to, information based on the source of the second text sequence, as described above with respect to the training example 304 of FIG. 3. In that regard, in some aspects, the label may include information regarding the source of the first text sequence instead of, or in addition to, information based on the source of the second text sequence. Similarly, in some aspects, the label may include information that is not directly related to the source of the first text sequence or the second text sequence (e.g., the field or group of topics to which the training example belongs) instead of, or in addition to, information based on the source of the second text sequence.

[0027] In step 404, the processing system uses a translation model to generate a predicted text sequence based at least in part on a first text sequence and a label of a given training example (e.g., the first text sequence and label of training example 304 in FIG. 3). The processing system can do this using any suitable type of translation model, architecture, and number of parameters, including a transformer architecture, a long short-term memory (“LSTM”) architecture, a recurrent neural network architecture (“RNN”), a convolutional neural network (“CNN”) architecture, and / or any suitable hybrid thereof. For example, in some aspects of the present technology, the translation model can be a deep LSTM network including a plurality of encoder layers and decoder layers (e.g., a 6-layer LSTM encoder and an 8-layer LSTM decoder, an 8-layer LSTM encoder and an 8-layer LSTM decoder, etc.). Similarly, in some aspects of the present technology, the translation model can be based on a hybrid architecture, such as an architecture that uses a transformer as an encoder and an RNN as a decoder (e.g., a 12-layer transformer encoder and a 2-layer RNN decoder).

[0028] Furthermore, the translation model can generate the predicted text sequence directly or indirectly based on the first text sequence and label of a given training example. Thus, for example, the processing system or the translation model can be configured to first process the first text sequence and / or label to generate a modified version thereof (e.g., a tokenized version of the first text sequence and / or label, a label based on the first text sequence and / or label, a vector, etc.). In such a case, the translation model can generate the predicted text sequence based on the modified version of the first text sequence and / or label (e.g., the tokenized version, the vector, etc.).

[0029] In step 406, the processing system compares the predicted text sequence with a second text sequence of a given training example (e.g., the second text sequence of training example 304 in FIG. 3) to generate a loss value. The processing system may perform this comparison and generate the loss value in any suitable way using any suitable loss function(s). For example, in some aspects of the present technology, the processing system may be configured to compare the predicted text sequence with the second text sequence using a "hard distillation" method that evaluates how similar each string of text is to other strings. Similarly, in some aspects, the processing system may be configured to compare the predicted text sequence with the second text sequence using a loss in a connectionist temporal classification method ("CTC loss") or a cross-entropy loss.

[0030] In step 408, the processing system determines whether there are any further training examples in the batch. In that regard, the plurality of training examples may be divided into multiple batches or may be held as a whole, in which case there would be a single "batch" that includes all the training examples of the plurality of first training examples. In either case, as indicated by the "yes" arrow, if the processing system determines that there are further training examples in the batch, the processing system proceeds to step 410. In step 410, the processing system selects the next given training example from the batch and then repeats steps 404-408 for that newly selected training example. This process then repeats for each next given training example in the batch until the processing system determines in step 408 that there are no further training examples in the batch and thus proceeds to step 412 (as indicated by the "no" arrow).

[0031] As shown in step 412, after the loss value is generated for each given training example within a batch (in step 406), the processing system modifies one or more parameters of the translation model based at least in part on the generated loss value. The processing system may be configured to modify one or more parameters based on these generated loss values in any suitable way and at any suitable interval. For example, an optimization routine such as stochastic gradient descent may be applied to the generated loss values to determine parameter modification. In some aspects of the present technology, each “batch” may include a single training example such that the processing system performs a backpropagation step of modifying one or more parameters of the translation model each time a loss value is generated. Similarly, if each “batch” includes two or more training examples, the processing system may be configured to combine the generated loss values (e.g., by summing or averaging multiple loss values) into a total loss value and modify one or more parameters of the translation model based on the total loss value.

[0032] In step 414, the processing system determines whether there are additional batches within the plurality of training examples. If the plurality of training examples are not split and thus there is a single “batch” that includes every training example within the plurality of training examples, the determination in step 414 automatically becomes “no” and method 400 then ends as shown in step 418. However, if the plurality of training examples are split into two or more batches, the processing system proceeds to step 416 following the “yes” arrow and selects the next given training example from the plurality of training examples. This then initiates a set of other passes through steps 404 - 408 for each training example within the next batch and for other modifications of one or more parameters of the translation model in step 412. This process continues until there are no more additional batches, at which point the processing system proceeds to step 418 following the “no” arrow.

[0033] Method 400 is shown to end at step 418 if all of a plurality of training examples are used to adjust the parameters of the translation model. However, it should be understood that method 400 may be repeated any suitable number of times using the same plurality of training examples until each of the predicted text sequences is close enough to its respective second text sequence of each training example. In that regard, in some aspects of the present technology, the processing system may be configured to repeat method 400 for a plurality of training examples a certain predetermined number of times. Further, in some aspects, the processing system may aggregate all of the loss values generated during a given pass through method 400 and be configured to determine whether to repeat method 400 for the plurality of training examples based on the total loss value. For example, in some aspects of the present technology, the processing system may be configured to repeat method 400 for a plurality of training examples if the total loss value of the most recent pass through method 400 exceeds some predetermined threshold. Similarly, in some aspects, the processing system may be configured to use gradient descent and thus repeat method 400 for a plurality of training examples until the total loss value for a given pass through method 400 is greater than or equal to the total loss value from the previous pass.

[0034] As described above, when a translation model is trained according to method 400, the translation model can be tested using different labels to determine which label will produce the highest quality results for a given validation set with the trained translation model. For example, if a trained translation model is intended to be used for translation between English and French, the validation set can be obtained for that language pair (e.g., from a benchmark translation dataset, from one or more representative websites or books, etc.). Similarly, if a trained translation model is intended to perform translations in a specific topic area, the validation set can be obtained from sources in that topic area (e.g., websites related to that topic, books related to that topic, etc.). Examples of the validation set can then be repeatedly provided to the translation model to generate translations using each different label within a set of candidate labels. The set of translations for each candidate label can then be evaluated and compared for quality to identify which label produced the most desirable results for the translation model. These quality evaluations can be performed in any suitable way, using, for example, any known automatic quality metric (e.g., BLEU, BLEURT, ROUGE, BERT score), a comparison to the target translation (e.g., when using examples from a benchmark training set that includes the target translation for each input text sequence), an evaluation by human raters, or a combination thereof.

[0035] FIG. 5 illustrates an exemplary method 500 for generating a plurality of training examples according to an aspect of the present disclosure. In that regard, in some aspects of the present technology, the exemplary method of FIG. 5 can be used to generate the plurality of training examples referenced in method 400 of FIG. 4.

[0036] In step 502, a processing system (e.g., the processing system 102 of FIG. 1 or FIG. 2, the processing system of method 400 of FIG. 4, etc.) samples a first text sequence (e.g., the first text sequence sampled from web page 302a to generate the training example 304 of FIG. 3) from a first page of a given Internet domain. The processing system may perform this sampling in any suitable manner. For example, in some aspects of the present technology, the processing system may directly sample the first text sequence from the first page. Similarly, in some aspects, the processing system may download the first page (or a portion thereof) and then sample the first text sequence from the downloaded copy or a portion of the first page.

[0037] In step 504, the processing system samples a second text sequence (e.g., the second text sequence sampled from web page 302b to generate the training example 304 of FIG. 3) from a second page of a given Internet domain. Again, the processing system may perform this sampling in any suitable manner. For example, in some aspects of the present technology, the processing system may directly sample the second text sequence from the second page. Similarly, in some aspects, the processing system may download the second page (or a portion thereof) and then sample the second text sequence from the downloaded copy or a portion of the second page.

[0038] In step 506, the processing system generates a label (e.g., a label generated based on the URL of web page 302b to generate training example 304 in FIG. 3) based on the source of the second text sequence. As described above with respect to training example 304 in FIG. 3, the processing system may generate a label based on any suitable information regarding the source of the second text sequence, including any of the options described above with respect to training example 304 in FIG. 3. Thus, in some aspects of the present technology, the processing system may generate a label based on all or part of the URL of the second page (e.g., "http: / / es.website.com / ", "es.website.com", "website.com", "es", "website", or "com"), the name of the website (e.g., "Website"), the IP address of the second page, and / or any other suitable information related to the source of the second page.

[0039] Furthermore, although not reflected in the example of FIG. 5, in some aspects of the present technology, the label may also include other information instead of or in addition to information based on the source of the second text sequence, as described above with respect to training example 304 in FIG. 3. In that regard, in some aspects, the label may include information regarding the source of the first text sequence instead of or in addition to information based on the source of the second text sequence. For example, the label may include all or part of the URL of the first page (e.g., "http: / / en.website.com / ", "en.website.com", "website.com", "es", "website", or "com"), the name of the website (e.g., "Website"), the IP address of the second page, and / or any other suitable information related to the source of the first page. Similarly, in some aspects, the label may include information that is not directly related to the source of the first text sequence or the second text sequence (e.g., the field or group of topics to which the training example belongs) instead of or in addition to information based on the source of the second text sequence.

[0040] Unless otherwise specified, the foregoing alternative examples are not mutually exclusive and can be implemented in various combinations to achieve their respective advantages. These and other variations and combinations of the features described above can be utilized without departing from the subject matter defined by the claims. Therefore, the foregoing description of the exemplary systems and methods should be regarded as illustrative rather than limiting the subject matter defined by the claims. In addition, the provision of the examples described herein and phrases expressed as "such as", "including", "comprising", etc. should not be construed as limiting the subject matter of the claims to the specific examples. Rather, the examples are intended to illustrate only a portion of the many possible embodiments. Further, the same reference numerals in different drawings can identify the same or similar elements.

Claims

1. A method performed by a computer, the method comprising: training a translation model, the training comprising: for each given training example of a plurality of training examples including a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on a source of the second text sequence, using the translation model to generate a predicted text sequence based at least in part on the first text sequence and the label of the given training example; using one or more processors of a processing system to compare the predicted text sequence with the second text sequence to generate a loss value for the given training example; using the one or more processors to modify one or more parameters of the translation model based at least in part on the loss value generated for each of the plurality of training examples.

2. The method of claim 1, wherein the label includes an internet domain.

3. The method of claim 1 or claim 2, wherein the label includes an internet subdomain.

4. The method according to any one of claims 1 to 3, wherein the label includes a uniform resource locator.

5. The method according to any one of claims 1 to 4, wherein the label includes a website name.

6. The method according to any one of claims 1 to 5, wherein the label includes an IP address.

7. The method according to any one of claims 1 to 6, wherein the label further indicates a source of the first text sequence.

8. The method according to any one of claims 1 to 7, wherein the source of the first text sequence is within a first subdomain of a given internet domain, and the source of the second text sequence is within a second subdomain of the given internet domain.

9. using the one or more processors to sample the first text sequence from a first page of a given internet domain; sample the second text sequence from a second page of the given internet domain. generating the label based on all or part of the uniform resource locator of the second page; further comprising, by, generating a respective given training example of the plurality of training examples, the method according to any one of claims 1 to 8. **Claim 10** using the one or more processors, sampling the first text sequence from a first page of a given Internet domain; sampling the second text sequence from a second page of the given Internet domain; generating the label based on all or part of the IP address of the second page; by, generating a respective given training example of the plurality of training examples; further comprising, the method according to any one of claims 1 to 9. **Claim 11** A processing system, a memory storing a translation model; one or more processors coupled to the memory and configured to train the translation model according to a training method, wherein the training method for each given training example of a plurality of training examples including a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on a source of the second text sequence, using the translation model to generate a predicted text sequence based at least in part on the first text sequence and the label of the given training example; comparing the predicted text sequence with the second text sequence to generate a loss value for the given training example; modifying one or more parameters of the translation model based at least in part on the loss value generated for each of the plurality of training examples, the processing system. **Claim 12** The processing system according to claim 11, wherein the one or more processors are configured to train the translation model with each given training example including a label including an Internet domain according to the training method. **Claim 13** The one or more processors are configured to train the translation model with each given training example including a label including an Internet subdomain according to the training method, the processing system according to claim 11 or claim 12.

14. The one or more processors are configured to train the translation model with each given training example including a label including a uniform resource locator according to the training method, the processing system according to any one of claims 11 to 13.

15. The one or more processors are configured to train the translation model with each given training example including a label including a website name according to the training method, the processing system according to any one of claims 11 to 14.

16. The one or more processors are configured to train the translation model with each given training example including a label including an IP address according to the training method, the processing system according to any one of claims 11 to 15.

17. The one or more processors are configured to train the translation model with each given training example including a label indicating the source of the first text sequence and the source of the second text sequence according to the training method, the processing system according to any one of claims 11 to 16.

18. The one or more processors are sampling the first text sequence from a first page of a given Internet domain, sampling the second text sequence from a second page of the given Internet domain, generating the label based on all or part of the uniform resource locator of the second page, Thereby, the processing system according to any one of claims 11 to 17 is further configured to generate each given training example of the plurality of training examples.

19. The one or more processors are sampling the first text sequence from a first page of a given Internet domain, sampling the second text sequence from a second page of the given Internet domain; generating the label based on all or part of the IP address of the second page; The processing system according to any one of claims 11 to 18, further configured to generate a respective given training example of the plurality of training examples.

20. A processing system, a memory storing a translation model; one or more processors coupled to the memory and configured to generate a predicted translation of the input text sequence based on the input text sequence and a label using the translation model; the translation model is trained to generate the predicted translation according to a training method, and the training method includes: for each given training example of a plurality of training examples including a first text sequence in a first language, a second text sequence in a second language different from the first language, and a label based on a source of the second text sequence, using the translation model to generate a predicted text sequence based at least in part on the first text sequence and the label of the given training example; comparing the predicted text sequence with the second text sequence to generate a loss value for the given training example; modifying one or more parameters of the translation model based at least in part on the loss value generated for each of the plurality of training examples.

21. A computer program product including computer-readable instructions that, when executed by a processing system, cause the processing system to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and server for training a machine learning algorithm for translation

    US20200279022A1