Sample construction method and apparatus, electronic device, and readable storage medium

By replacing and labeling keywords in parallel corpus training samples, extended text is generated, which solves the problem of low accuracy in translating non-standard words in existing translation models and achieves more flexible and accurate translation results.

CN116089569BActive Publication Date: 2026-05-29VIVO MOBILE COMM CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2023-02-08
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing translation models have low accuracy when dealing with texts containing non-standard words, and cannot effectively utilize non-standard vocabulary in parallel corpus training samples.

Method used

By acquiring keywords from parallel corpus training samples, replacing them with non-standard words and generating extended text, and replacing the standard type labels with the standard type labels of the non-standard words, target training samples are constructed to enrich the content of the parallel corpus training samples.

Benefits of technology

The vocabulary range of the parallel corpus training samples was expanded, which improved the translation model's ability to identify and translate non-standard words, thereby enhancing translation accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116089569B_ABST
    Figure CN116089569B_ABST
Patent Text Reader

Abstract

The application discloses a sample construction method and device, electronic equipment and readable storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: obtaining parallel corpus training samples, the parallel corpus training samples contain original texts and carry standard type labels corresponding to each keyword in the original texts; replacing a first keyword in the original text with at least one first non-standard word corresponding to the first keyword to generate at least one expanded text; replacing a first standard type label corresponding to the first keyword with a second standard type label corresponding to the first non-standard word to obtain the parallel corpus training sample after label replacement; and constructing a target training sample based on the parallel corpus training sample after label replacement and the at least one expanded text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a sample construction method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] With the development of computer performance and Internet technology, existing translation methods usually employ large-scale bilingual parallel corpora to train translation models and generate translations based on the distribution of real corpora in the text to be translated.

[0003] However, since parallel corpus training samples are often composed of high-quality standard texts, the translation model trained on such parallel corpus training samples can only translate standard texts. When translating texts containing non-standard words, the overall translation accuracy is low.

[0004] Therefore, how to construct richer parallel corpus training samples is an urgent problem to be solved in this application. Summary of the Invention

[0005] The purpose of this application is to provide a sample construction method, apparatus, electronic device, and readable storage medium that can solve the problem of how to construct richer parallel corpus training samples.

[0006] In a first aspect, embodiments of this application provide a sample construction method, which includes: obtaining parallel corpus training samples, wherein the parallel corpus training samples contain original text and carry a normative type label corresponding to each keyword in the original text; replacing a first keyword in the original text with at least one first non-normative word corresponding to the first keyword to generate at least one extended text; replacing the first normative type label corresponding to the first keyword with a second normative type label corresponding to the first non-normative word to obtain a parallel corpus training sample with replaced labels; and constructing a target training sample based on the parallel corpus training sample with replaced labels and at least one extended text.

[0007] Secondly, embodiments of this application provide a sample construction apparatus, comprising: an acquisition module, a processing module, and a construction module; the acquisition module is configured to acquire parallel corpus training samples, the parallel corpus training samples containing original text and carrying a normative type label corresponding to each keyword in the original text; the processing module is configured to replace a first keyword in the original text of the parallel corpus training samples acquired by the acquisition module with at least one first non-normative word corresponding to the first keyword, thereby generating at least one extended text; the processing module is further configured to replace the first normative type label corresponding to the first keyword in the parallel corpus training samples acquired by the acquisition module with a second normative type label corresponding to the first non-normative word, thereby obtaining a parallel corpus training sample with replaced labels; the construction module is configured to construct a target training sample based on the parallel corpus training sample with replaced labels processed by the processing module and at least one extended text.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, parallel corpus training samples are obtained. These parallel corpus training samples contain original text and carry a canonical type label corresponding to each keyword in the original text. The first keyword in the original text is replaced with at least one first non-canonical word corresponding to the first keyword to generate at least one extended text. The first canonical type label corresponding to the first keyword is replaced with a second canonical type label corresponding to the first non-canonical word to obtain a parallel corpus training sample with replaced labels. Based on the parallel corpus training sample with replaced labels and at least one extended text, a target training sample is constructed. Through this scheme, the sample construction device can replace keywords in the original text of the parallel corpus training sample to generate at least one extended text, thereby expanding the vocabulary range covered by the parallel corpus training sample. Simultaneously, the canonical type label corresponding to the keyword is replaced with the canonical type label corresponding to the non-canonical word to obtain a parallel corpus training sample with replaced labels, thereby enriching the content contained in the parallel corpus training sample. Finally, the sample construction device can construct the target training sample based on the parallel corpus training sample with replaced labels and at least one extended text. Therefore, the target training samples can contain non-standard words and their corresponding standard type labels, which can enrich the content of the parallel corpus training samples and make the parallel corpus training samples have more and more flexible training content. Attached Figure Description

[0013] Figure 1 This is an example diagram illustrating a non-standard term provided in an embodiment of this application;

[0014] Figure 2 This is a flowchart of a sample construction method provided in an embodiment of this application;

[0015] Figure 3 This is one of the schematic diagrams illustrating a sample construction method provided in the embodiments of this application;

[0016] Figure 4 This is a second schematic diagram illustrating an example of a sample construction method provided in this application embodiment;

[0017] Figure 5 This is a third example of a sample construction method provided in the embodiments of this application;

[0018] Figure 6 This is a flowchart illustrating a translation model provided in an embodiment of this application.

[0019] Figure 7 This is a schematic diagram of the structure of a sample construction device provided in an embodiment of this application;

[0020] Figure 8It is one of the schematic diagrams of the hardware structure of an electronic device provided by an embodiment of the present application;

[0021] Figure 9 It is the second of the schematic diagrams of the hardware structure of an electronic device provided by an embodiment of the present application. Specific embodiments

[0022] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and do not limit the number of objects. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0024] Some terms / nouns involved in the embodiments of the present application are explained below.

[0025] 1. Cognate words: Among languages or scripts with relatively close language family branches, there are often many words with the same linguistic origin. These words have similar pronunciations, spellings, or meanings, and may be easily confused in terms of glyph composition. For example, Chinese and Japanese that are both written in Chinese characters (e.g., "honor" and "栄誉"), English and German that belong to the West Germanic branch (e.g., "popular" and "populär"), simplified and traditional Chinese, etc. Due to reasons such as input errors, the words in the text to be translated may be replaced by cognates, resulting in a decline in the quality of the translated text.

[0026] 2. Kana: A syllabary of the Japanese language, with two writing forms: hiragana and katakana, which can be converted into each other, and each kana represents a syllable. Chinese characters in Japanese can be transcribed into kana according to their pronunciations, similar to Chinese pinyin. At the same time, kana is also a writing form of the Japanese language, used to represent native Japanese words and grammatical particles, etc.

[0027] 3. Kanji (Japanese Characters): Kanji, used in Japanese, together with kana, form the written script of Japanese. They are often used to represent the names of objects or actions. There are about 2,000-3,000 commonly used kanji in modern Japanese. Their forms are derived from Chinese characters, and there are certain overlaps and differences between them and simplified and traditional Chinese characters.

[0028] 4. Original text: The original text to be translated, with no restrictions on the specific language of the original text.

[0029] 5. Translation: The result of translating the original text using a translation model; there are no restrictions on the specific language of the translation.

[0030] 6. Language Model: A model used to calculate the probability of a sentence (i.e., the probability that a sequence of words can form a normal sentence). Its core is to calculate the probability of the current word appearing by analyzing the first n words in the sentence. Perplexity is usually used as the evaluation metric.

[0031] 7. Perplexity: An indicator for evaluating the quality of a sentence. The higher the perplexity, the more difficult it is to understand, meaning it is less likely to be a fluent and semantically correct sentence.

[0032] 8. Morphology: The study of words in a sentence, including word structure, morphology, and parts of speech, such as nouns, adjectives, adverbs, and singular and plural forms in English.

[0033] 9. Syntactic structure: The relationships between sentence components and the rules or processes by which they form sentences, such as the common "subject-verb-object" structure.

[0034] 10. Sequence labeling: Given a sentence, label each word in the sentence, or predict the category label of the words.

[0035] 11. Word segmentation: A type of sequence labeling task. For languages ​​such as Chinese and Japanese where there are no spaces between words in writing, word segmentation models can segment sentences at the word level and predict the lexical and syntactic structure categories of words. The word segmentation model trained in this scheme also involves predicting extended forms of non-standard words (e.g., pronunciation spelling, cognates, easily confused words, etc.).

[0036] The sample construction method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0037] Existing machine translation methods typically use large-scale bilingual parallel corpora to train translation models and generate translations based on the distribution of real corpora.

[0038] However, since the original texts in the parallel corpus training samples are usually high-quality, standard texts, issues such as words being transliterated into phonetic spellings or cognates are rare. Translation models often don't encounter these non-standard expressions and lack the ability to accurately translate them. However, in certain specific scenarios, the text input to the translation model may contain non-standard words whose expressions do not conform to conventional grammar. For example, in language education scenarios, such as... Figure 1 As shown, words in text may be transcribed into their phonetic spelling in the target language (such as Pinyin or Japanese kana) for teaching or exams. User typing errors can also lead to spelling mistakes, typos, and substitutions of cognates in the text to be translated. In tasks such as image translation and speech translation, the recognition results of pre-processing modules like image text recognition and speech recognition may contain errors related to character shape similarity, pronunciation similarity, and encoding errors, which may also cause downstream translation models to receive non-standard text. Therefore, because these text sequences containing non-standard or erroneous words are often not very common sequences—meaning their expression does not conform to conventional grammar, lexical, or syntactic structures—translation models often struggle to translate such non-standard or erroneous words correctly.

[0039] Taking Japanese as an example, on the one hand, Japanese has two writing systems: kana and kanji. Japanese kanji are highly similar to Chinese kanji, and there are certain overlaps and differences between them and both simplified and traditional Chinese characters, as shown in Table 1. When Chinese users input Japanese, they may replace kanji words with cognates that do not exist in Japanese, or misspelled characters due to convenience, laziness, or confusion of character forms. This may lead to translation errors in the model.

[0040] Table 1

[0041]

[0042] On the other hand, Japanese kana can have meaning on their own and be used in written expression, or they can be used to spell the pronunciation of kanji. In online texts such as those on social media platforms, many users, for convenience, do not spell the standard kanji but directly replace them with the pronunciation of kana, such as... Figure 1 As shown in Table 2, however, kana with the same pronunciation can have many instances of "one word with multiple meanings," resulting in numerous non-standard Japanese kanji expressions. Furthermore, since there are no spaces between words in Japanese writing, and the character set for kana transcribing completely overlaps with the normal text, existing methods struggle to correctly identify and segment non-standard kana words in sentences if a large number of kanji are transcribed into kana. In addition, Japanese also has a large number of homophones; the same kana pronunciation may correspond to multiple different kanji words, as shown in Table 2.

[0043] Table 2

[0044]

[0045] Since most existing text translation methods are trained on standardized corpora, when inputting text with non-standard expressions, the translation model often outputs transliterations of these words, or even random translations, resulting in an inability to obtain accurate translations.

[0046] The sample construction method provided in this application allows the sample construction device to replace keywords in the original text of the parallel corpus training samples, generating at least one extended text to expand the vocabulary covered by the parallel corpus training samples. Simultaneously, it replaces the normative type labels corresponding to the keywords with normative type labels corresponding to non-normative words, obtaining parallel corpus training samples with replaced labels, thus enriching the content contained in the parallel corpus training samples. Finally, the sample construction device can construct a target training sample based on the parallel corpus training samples with replaced labels and at least one extended text. Therefore, the target training sample can contain non-normative words and their corresponding normative type labels, thereby enriching the content of the parallel corpus training samples and enabling them to have more flexible training content.

[0047] The entity executing the sample construction method provided in this application embodiment can be a sample construction device. Exemplarily, the sample construction device can be an electronic device, or a component within that electronic device, such as an integrated circuit or a chip. The sample construction method provided in this application embodiment will be described executively below using a sample construction device as an example.

[0048] This application provides a sample construction method. Figure 2 A flowchart of a sample construction method provided in an embodiment of this application is shown. The execution subject of this method can be a sample construction device. Figure 2 As shown, the sample construction method provided in this application embodiment may include the following steps 201 to 204.

[0049] Step 201: Obtain parallel corpus training samples.

[0050] The parallel corpus training samples mentioned above can contain the original text and carry the canonical type label corresponding to each keyword in the original text.

[0051] In this embodiment of the application, the parallel corpus training samples can be bilingual or multilingual corpora consisting of the original text and its parallel corresponding translation text.

[0052] Optionally, the original text can be text that does not contain non-standard words.

[0053] Optionally, the above-mentioned keyword can be any word in the original text.

[0054] Optionally, the above-mentioned specification type tag can indicate the specification type of the keyword.

[0055] It can be understood that, on the one hand, due to the existence of a large number of homophones in the same language, the extended forms of these words may be the same as other normative words in the normative word list. For example, "さくら" can be either the kana transcription of the surname "佐倉 (Sakura)" or the noun "cherry blossom". Therefore, it is difficult to identify all non-conforming words through rule-based methods. On the other hand, due to the different rules of text sequences between different languages, for example, there are no spaces between words in Japanese, and when a large number of Chinese characters in the text to be translated are transcribed into kana, it is also difficult for rule-based methods to accurately identify the boundaries between words. Therefore, it is difficult to accurately translate all words in the text to be translated through rule-based methods. Therefore, the sample construction device in the sample construction method provided by the embodiments of the present application can adopt text data (i.e., the original text) marked with information such as lexical and syntactic structures, and add the corresponding specification type tags of the keywords on this basis.

[0056] Step 202: Replace the first keyword in the original text with at least one first non-conforming word corresponding to the first keyword to generate at least one extended text.

[0057] Optionally, the sample construction device can replace any keyword in the original text with at least one first non-conforming word corresponding to it, obtaining multiple extended texts with the same semantics but different degrees of specification.

[0058] Optionally, other annotation information such as the词性 (lexical category) and syntactic structure of the extended text can be kept consistent with the annotation information of the original text.

[0059] Optionally, the above-mentioned non-conforming word can be a word whose expression does not conform to conventional grammar, lexicology or syntactic structure.

[0060] Optionally, the above-mentioned non-conforming word can include at least one of the following situations: including pronunciation spelling, including typos, including replacement of cognate characters, including glyph errors.

[0061] Optionally, "replacing the first keyword in the original text with at least one first non-conforming word corresponding to the first keyword" can be understood as: replacing the keyword that conforms to the specification with a non-conforming word whose expression is homologous, has the same or similar pronunciation, or is similar in glyph and does not conform to conventional grammar, lexicology or syntactic structure.

[0062] For example, if the original text contains the keyword "境界", the sample construction device can replace it with the homophonic "教会" or the non-compliant word "きょうかい (bianjie)".

[0063] Optionally, the parallel corpus training sample can be a parallel corpus training sample in the parallel corpus training sample set. The above step 202 may include the following step 202a.

[0064] Step 202a: Based on the word frequency of each keyword in the original text in the parallel corpus training sample set, determine at least one first keyword from the original text, and replace each first keyword in the at least one first keyword in the original text with its corresponding first non-compliant word to generate a first extended text.

[0065] Among them, the first extended text is any one of the above at least one extended text.

[0066] Optionally, the sample construction device can replace the keywords in the original text based on the word frequency of each keyword in the original text in the parallel corpus training sample set.

[0067] It can be understood that the higher the word frequency of a word, the more likely it is to be replaced.

[0068] Specifically, the first keyword in the original text can be replaced with its corresponding first non-compliant word according to its word frequency in the parallel corpus training sample set.

[0069] Exemplarily, as Figure 3 shown, for the keywords "とても (非常地)", "頼もしく (可信赖地)", "優しい (温柔地)", according to their word frequencies in the parallel corpus training sample set, replace "頼もしく (可信赖地)" with the form containing pinyin reading and writing (i.e., its specification type label is pinyin reading and writing - hiragana) "たのもしく (可信lai地)", replace "優しい (温柔地)" with the form containing pinyin reading and writing (i.e., its specification type label is pinyin reading and writing - hiragana) "やさしい (温rou地)" to obtain extended text 1; replace "とても (非常地)" with the form containing pinyin reading and writing (i.e., its specification type label is pinyin reading and writing - katakana) "トテモ (feichangde)", replace "頼もしく" with the form containing cognate characters (i.e., its specification type label is cognate characters - Traditional Chinese) "賴もしく (可信賴地)", replace "優しい" with the form containing cognate characters (i.e., its specification type label is cognate characters - Simplified Chinese) "优しい (温柔地)" to obtain extended text 2.

[0070] Thus, since the sample construction device can replace keywords based on their word frequency in the parallel corpus training sample set, keywords with high word frequency can be replaced more often with at least one non-standard word corresponding to them. This allows the generated extended text to contain as many possible non-standard forms as possible corresponding to the original text, thereby enabling more comprehensive training of the translation model.

[0071] Step 203: Replace the first normative type label corresponding to the first keyword with the second normative type label corresponding to the first non-normative word to obtain parallel corpus training samples after label replacement.

[0072] Optionally, the canonical type label can indicate the canonical type of the word.

[0073] For example, when a word is a compliant word (i.e., the first keyword), its corresponding compliant type tag (i.e., the first compliant type tag) can indicate that it is a compliant word; when a word is a non-compliant word (i.e., the first non-compliant word), its corresponding compliant type tag (i.e., the second compliant type tag) can indicate that it is a non-compliant form.

[0074] For example, as shown in Table 3, the second standard type label can include various forms such as pronunciation spelling - hiragana, pronunciation spelling - katakana, cognates - simplified Chinese, cognates - traditional Chinese, easily confused words - simplified Chinese, easily confused words - traditional Chinese, easily confused words - recombination, etc.

[0075] Table 3

[0076]

[0077] Step 204: Construct target training samples based on parallel corpus training samples with replaced labels and at least one extended text.

[0078] Optionally, the sample construction device can associate non-standard words in the expanded text with their corresponding standard type labels in the parallel corpus training samples after label replacement to obtain target training samples.

[0079] This application provides a sample construction method. The sample construction device can replace keywords in the original text of parallel corpus training samples to generate at least one extended text, thereby expanding the vocabulary covered by the parallel corpus training samples. Simultaneously, it replaces the normative type labels corresponding to the keywords with normative type labels corresponding to non-normative words, obtaining parallel corpus training samples with replaced labels, thus enriching the content contained in the parallel corpus training samples. Finally, the sample construction device can construct a target training sample based on the parallel corpus training samples with replaced labels and at least one extended text. Therefore, the target training sample can contain non-normative words and their corresponding normative type labels, thereby enriching the content of the parallel corpus training samples and enabling them to have more flexible training content.

[0080] Optionally, the number of extended texts is N, where N is a positive integer. After step 202 above, the sample construction method provided in this application embodiment may further include step 205 below.

[0081] Step 205: If the second extended text in the N extended texts contains out-of-vocabulary words that are not included in the parallel corpus training sample set, initialize the feature information of the out-of-vocabulary words.

[0082] The initialization process includes at least one of the following: weighting the feature information of the out-of-vocabulary word according to the first keyword corresponding to the out-of-vocabulary word and the word frequency of each non-standard word corresponding to the first keyword corresponding to the out-of-vocabulary word in each of the N extended texts in the parallel corpus training sample set; weighting the feature information of the out-of-vocabulary word using the feature information of the cognate words corresponding to the out-of-vocabulary word; setting the feature information of the out-of-vocabulary word to 0; and randomly initializing the feature information of the out-of-vocabulary word.

[0083] Optionally, the first extended text and the second extended text may be the same or different.

[0084] Optionally, the sample construction device can convert the obtained extended text into a sequence of word vectors corresponding to model training based on the feature information of the words.

[0085] For example, the sample construction device can obtain word vector sequences through algorithms such as word-to-vector (Word2Vec) algorithm and regression algorithm based on global word frequency statistics (Glove algorithm), or it can obtain word vector sequences through training and iteration in translation models such as Transformer.

[0086] In practice, the sample construction device can obtain the word vector sequence corresponding to the extended text in any possible way, and this application does not impose any specific limitations.

[0087] In this embodiment, for out-of-vocabulary words (i.e., words that have not appeared in the parallel corpus training sample set), the feature information of out-of-vocabulary words can be initialized using any combination of the following methods to obtain the corresponding word vectors: ① The feature information of out-of-vocabulary words is weighted and averaged according to the word frequency of each non-standard word corresponding to the first keyword of the out-of-vocabulary word in each of the N extended texts in the parallel corpus training sample set; ② The feature information of out-of-vocabulary words is weighted and averaged using the feature information of the cognate words corresponding to the out-of-vocabulary words; ③ The feature information of out-of-vocabulary words is set to 0; ④ The feature information of out-of-vocabulary words is randomly initialized.

[0088] In the actual training process of the model, the sample construction device can also randomly initialize the feature information of the canonical type labels corresponding to the out-of-vocabulary words, or combine the canonical type labels and obtain the canonical type labels and their feature information corresponding to the out-of-vocabulary words by weighted averaging of the feature information of the corresponding words.

[0089] Thus, on the one hand, at the data level, since the sample construction device can initialize the feature information of out-of-vocabulary words, it can enhance the training of the translation model; on the other hand, at the model level, since the sample construction device can initialize the feature information of out-of-vocabulary words, the translation model can learn the phonetic relevance between non-standard words and their corresponding standard words during training, thereby improving the translation robustness of the translation model. Therefore, the sample construction method provided in this application embodiment can improve the translation quality and accuracy of the translation model.

[0090] Optionally, after step 204 above, the sample construction method provided in this application embodiment may further include steps 206 and 207 as described below.

[0091] Step 206: Restore at least one non-standard word in the first translated text to a standard word to generate M second translated texts.

[0092] One non-standard word is restored to at least one standard word.

[0093] Optionally, the first translated text mentioned above can be a sentence or a paragraph.

[0094] Optionally, the first translated text can be text entered by the user or text obtained from other devices.

[0095] Optionally, the sample construction device can identify non-standard words in the first translated text using the following three methods: Method 1: an extended vocabulary construction method based on cognates, pronunciations, and character sets; Method 2: a word segmentation model method based on an extended vocabulary enhancement; Method 3: an irregular translation detection method based on language model probabilities.

[0096] Methods 1 to 3 will be described in detail below with reference to specific embodiments.

[0097] Method 1: An extended vocabulary construction method based on cognates, pronunciations, and character sets.

[0098] In this embodiment of the application, the extended form of the non-standard word in the first translated text may include all words matched by the non-standard word in the extended vocabulary.

[0099] It is understandable that when the first translated text includes non-standard words, the characters used in those words may exceed the normal character set of the current language. For example, Chinese characters may contain Pinyin characters outside the standard character set, or the non-standard word may use a spelling that does not exist in the current language. For example, the English word "October" may be spelled using the German word "Oktober," which is from the same language family. Therefore, the sample construction device in the sample construction method provided in this application embodiment can construct an extended vocabulary by mining the similarity of words between different languages. The extended vocabulary is shown in Table 3, for example.

[0100] In this embodiment of the application, taking Japanese as an example, the extended vocabulary may include: common pronunciation spellings of words and their variants; cognate or synonymous characters / words of words in other languages ​​with similar language system branches; easily confused words obtained by recombining words and their cognates; easily confused words obtained by replacing words with words that have similar characters, etc.

[0101] Optionally, if a word has a highly similar dictionary definition to its cognate or synonymous characters / words, cognate construction can be performed by mining dictionary information from various languages.

[0102] Alternatively, easily confused words can be words that do not exist in their original language or cognate languages.

[0103] Optionally, the extended vocabulary may include multiple word sets, each word set may include one or more non-standard words and a set of standard words corresponding to the non-standard word.

[0104] Optionally, the sample construction device can identify non-standard words in the first translated text through methods such as character set detection and extended vocabulary matching, and use the set of words matched in the extended vocabulary as the first word set.

[0105] Method 2: Word segmentation model method based on extended vocabulary enhancement.

[0106] Optionally, such as Figure 3 As shown, words in the first translated text can be replaced with any extended form in the extended vocabulary based on the word frequency setting of the word in the parallel corpus training sample set, and the corresponding standard type label can be replaced. The word segmentation model is trained using the extended form of the corpus and its corresponding standard type label.

[0107] Optionally, prior to step 206 above, the sample construction method provided in this application embodiment may further include step A below.

[0108] Step A: After inputting the first translated text into the word segmentation model, the first translated text is segmented into M words, where M is an integer greater than 1. Each of the M words is then identified as a non-standard word, and the identification result corresponding to each word is obtained. The identification result corresponding to a word is used to characterize whether a word belongs to a non-standard word.

[0109] For example, the word segmentation model can be a word segmentation model that has been augmented and trained.

[0110] For example, the word segmentation model trained with enhancement can predict the canonical type label for each word obtained. If the predicted canonical type label of the word indicates that the word is a non-canonical word, then the word is identified as a non-canonical word.

[0111] Thus, since the sample construction device enables the word segmentation model trained with enhancement to acquire the ability to recognize words, learn the similarity between non-standard words and standard words in terms of lexical, syntactic structure, and contextual information, and predict the standard type label of the output word segmentation, the word segmentation model can accurately segment the first translated text and identify non-standard words in the first translated text.

[0112] Method 3: Irregular translation detection method based on language model probability.

[0113] It's understandable that non-standard words appear less frequently in the parallel corpus training samples, and homophones often differ significantly in meaning and context, making the text less fluent than normal text. Therefore, a language model can be used to calculate the perplexity of the first translated text to determine whether it contains non-standard expressions.

[0114] Optionally, the sample construction device can input the first translated text into the n-gram language model and calculate the current word w using the following formula 1. i The probability associated with the first n words of the first translated text.

[0115] (Formula 1)

[0116] Among them, w i is the current word, and N is the number of words in the first translated text.

[0117] From formula (1), we can see the conditional probability of the current word. The lower the value, the lower the fluency of the first translated text and the higher the perplexity of that first translated text.

[0118] Optionally, prior to step 206 above, the sample construction method provided in this application embodiment may further include steps B1 to B4 as described below.

[0119] Step B1: Segment the first translated text into M words.

[0120] Where M is an integer greater than 1.

[0121] For example, the sample building device can segment the first translated text input into words using an enhanced word segmentation model.

[0122] Step B2: For each of the M word segments, if the conditional probability of a word segment is less than the first preset threshold, obtain the P first standardized words corresponding to that word segment.

[0123] Where P is a positive integer.

[0124] It is understandable that if the conditional probability of a word segment is less than the first preset threshold, it means that the word segment may be a non-standard word.

[0125] Optionally, the P first conforming words can be X conforming words from the set of conforming words matched by the word segment in the expanded vocabulary.

[0126] Step B3: Replace one word in the first translated text with each of the P first standard words, resulting in P replaced first translated texts.

[0127] Step B4: If the first perplexity of any replaced first translated text is less than the second perplexity of the first translated text, and the difference between the first perplexity and the second perplexity is greater than the second preset threshold, then the sample construction device determines a word segment as a non-standard word.

[0128] It can be understood that if the first perplexity of any replaced first translated text is less than the second perplexity of the original first translated text, and the difference between the first and second perplexities is greater than a second preset threshold, then the replaced first translated text is considered more fluent and reasonable. In other words, the original first translated text contained non-standard words.

[0129] Thus, since the sample construction device can replace the potentially non-compliant words in the first translation text with their corresponding first compliant words, and calculate the perplexity of the first translation text before and after the replacement respectively, when the difference in the perplexity of the first translation text after replacement decreases and is greater than the second preset threshold, the word is determined as a non-compliant word. Therefore, the recognition of non-compliant words can be made more accurate, and the first translation text after replacement is made more fluent and reasonable, so that the subsequent translation is more accurate and has a higher correct rate.

[0130] Optionally, the above step 206 can be specifically implemented by the following steps 206a and 206b.

[0131] Step 206a: Obtain a first word set corresponding to at least one non-compliant word.

[0132] Among them, the first word set may include: multiple word subsets. A word subset may include one or more non-compliant words among at least one non-compliant word, and each non-compliant word corresponds to a set of compliant words.

[0133] It can be understood that if there are multiple non-compliant words in at least one non-compliant word, the sets of compliant words corresponding to each non-compliant word among the multiple non-compliant words may be the same or different.

[0134] For example, among the above at least one non-compliant word, there are non-compliant words "already" and "already", and the set of compliant words corresponding to the non-compliant word "already" may be a set including the compliant word "already", and the set of compliant words corresponding to the non-compliant word "already" may also be a set including the compliant word "already".

[0135] Step 206b: For each word subset in the multiple word subsets, perform a reduction mapping of a word subset and the set of compliant words corresponding to each non-compliant word in the word subset in the first translation text to generate at least one second translation text.

[0136] In the embodiment of the present application, "performing a reduction mapping of a word subset and the set of compliant words corresponding to each non-compliant word in the word subset" can be understood as: sequentially restoring each non-compliant word in the above word subset to each compliant word in the set of compliant words corresponding thereto, and traversing all combinations of compliant word restorations.

[0137] For example, the first translated text is: Thinking of saying goodbye to xiaoyuan tomorrow, a deep sense of attachment wells up in my heart. Among them, there are non-conforming words "xiaoyuan" and non-conforming word "shenshen". The set of conforming words corresponding to the non-conforming word "xiaoyuan" includes: campus, small courtyard; the set of conforming words corresponding to the non-conforming word "shenshen" includes: deep, careful. Then, the sample construction device can perform a reduction mapping on the set of conforming words corresponding to each non-conforming word to obtain 6 second translated texts, which are respectively: Thinking of saying goodbye to the campus tomorrow, a deep sense of attachment wells up in my heart; Thinking of saying goodbye to the campus tomorrow, a deep sense of attachment wells up in my heart; Thinking of saying goodbye to the campus tomorrow, a careful sense of attachment wells up in my heart; Thinking of saying goodbye to the small courtyard tomorrow, a deep sense of attachment wells up in my heart; Thinking of saying goodbye to the small courtyard tomorrow, a deep sense of attachment wells up in my heart; Thinking of saying goodbye to the campus tomorrow, a careful sense of attachment wells up in my heart.

[0138] In this way, since the sample construction device can restore the non-conforming words in the first translated text to all possible conforming words to generate at least one second translated text, it is possible to correct the non-conforming words in the first translated text as much as possible, making the subsequent translated text more accurate and smooth.

[0139] Step 207: Input the first feature information corresponding to the first translated text and the second feature information corresponding to X second translated texts among the M second translated texts into the first translation model for text translation to obtain the target translation.

[0140] Among them, the first feature information includes the text feature information of the first translated text and the feature information of the specification type label corresponding to the non-conforming words in the first translated text, and the second feature information includes the text feature information of the second translated text and the feature information of the specification type label corresponding to the non-conforming words in the second translated text.

[0141] In the embodiment of the present application, the first translation model is trained based on a target training sample set, the target training sample set includes multiple target training samples, one target training sample corresponds to one parallel corpus training sample in the parallel corpus training sample set, M and X are positive integers, and X is less than or equal to M.

[0142] Optionally, the above step 207 can be specifically implemented by the following step 207a and step 207b.

[0143] Step 207a: Input X second translated texts among the M second translated texts and the first translated text into the first translation model for text translation, and output L candidate translations.

[0144] Among them, the L candidate translations include the candidate translations corresponding to the X second translation texts and the candidate translation corresponding to the first translation text. One candidate translation corresponds to at least one second translation text. L is a positive integer, and L is less than or equal to X.

[0145] It can be understood that since the enhanced translation model can make the same translation for non-compliant words in different extended forms, the number of candidate translations output by the translation model is less than the number of the input second translation texts.

[0146] Exemplarily, as Figure 4 shown, when the original text (i.e., the first translation text) "両親は学校に勤める (Parents work at school)" is input into the enhanced translation model, the target translation "Parents work at school" can be obtained.

[0147] When the extended text 1 "両親は學校につとめる (Parents work at school)" is input into the enhanced translation model, that is, the "両親 (both parents)" in the original text is replaced with the form recombined with the confusing word (i.e., its specification type label is confusing word - recombination) "両親 (both parents)", the "学校" is replaced with the form containing the cognate word and traditional Chinese characters (i.e., its specification type label is cognate word - traditional Chinese characters) "學校 (school)", and the "勤める (work)" is replaced with the form containing the phonetic reading and writing (i.e., its specification type label is phonetic reading and writing - hiragana) "つとめる (gongzuo)", the target translation "Parents work at school" can also be obtained.

[0148] When the extended text 2 "兩親はがっこうにツトメル (Parents work at school)" is input into the enhanced translation model, that is, the "両親 (both parents)" in the original text is replaced with the form of the confusing word in traditional Chinese characters (i.e., its specification type label is confusing word - traditional Chinese characters) "兩親 (both parents)", the "学校" is replaced with the form containing the phonetic reading and writing (i.e., its specification type label is phonetic reading and writing - hiragana) "がっこう (xuexiao)", and the "勤める (work)" is replaced with the form containing the phonetic reading and writing (i.e., its specification type label is phonetic reading and writing - katakana) "ツトメル (gongzuo)", the target translation "Parents work at school" can also be obtained.

[0149] Step 207b: Determine the candidate translation that meets the first condition among the L candidate translations as the target translation.

[0150] Optionally, the candidate translations that meet the first condition may include at least one of the following:

[0151] Case 1: The candidate translation whose fluency meets the first predetermined condition;

[0152] Case 2: The candidate translation whose translation quality meets the second predetermined condition;

[0153] Case 3: Candidate translations whose relevance meets the third predetermined condition.

[0154] The aforementioned relevance includes at least one of the following: prior probability, similarity, and perplexity.

[0155] For example, the first predetermined condition could be that the perplexity of the candidate translation is less than or equal to a third preset threshold. It can be understood that the lower the perplexity of the candidate translation, the higher its fluency and the more reasonable it is.

[0156] For example, in case 1, the sample construction device can calculate the perplexity of L candidate translations using a language model, and determine the candidate translations with perplexity less than or equal to a third preset threshold as the target translations.

[0157] For example, the second predetermined condition could be that the translation quality of the candidate translation is greater than or equal to a fourth preset threshold. It is understood that the sample construction device can determine candidate translations with translation quality greater than or equal to the fourth preset threshold as the target translation.

[0158] For example, the third predetermined condition could be that the relevance of the candidate translation is greater than or equal to a fifth preset threshold. It is understood that the sample construction device can determine candidate translations with a relevance greater than or equal to the fifth preset threshold as target translations.

[0159] It should be noted that if there are candidate translations that satisfy multiple predetermined conditions in the first condition, the sample construction device can determine the candidate translation that satisfies the most predetermined conditions as the target translation.

[0160] Thus, since the sample construction device can determine the candidate translation with the best evaluation results as the target translation based on the fluency, translation quality and relevance of the candidate translations, the output target translation can be optimized.

[0161] Alternatively, the sample construction device can evaluate the translation quality of candidate translations using representation and feature learning methods.

[0162] For example, after step 207a above, the sample construction method provided in this application embodiment may further include steps 207c and 207d below.

[0163] Step 207c: For each of the L candidate translations, extract the first text feature information of the candidate translation, as well as the first translated text and the second text feature information of the first translated text.

[0164] For example, the first text feature information may include lexical and syntactic features of the candidate translation.

[0165] For example, the second text feature information may include lexical and syntactic features of the second translated text and the first translated text.

[0166] For example, the sample construction device can extract the first text feature information of the candidate translation by training a word segmentation model of the target language, and extract the second translation text corresponding to the candidate translation and the second text feature information of the first translation text by using a word segmentation model of the source language.

[0167] Step 207d: Based on the first text feature information and the second text feature information, calculate the translation quality parameters corresponding to a candidate translation.

[0168] For example, the sample construction apparatus can use a regression algorithm to calculate the quality of the translation results.

[0169] For example, the translation quality parameter corresponding to a candidate translation can be the result value of a regression algorithm.

[0170] It is understandable that regression algorithms can output the probability of the quality of candidate translations: the closer the result of the regression algorithm is to 1, the better the quality of the candidate translation; the closer the result of the regression algorithm is to 0, the worse the quality of the candidate translation.

[0171] Thus, since the sample construction device can calculate the translation quality parameters corresponding to a candidate translation based on the first text feature information of a candidate translation, the first translated text corresponding to the candidate translation, and the second text feature information of the first translated text, it can select candidate translations with better translation quality.

[0172] Optionally, the relevance of the above candidate translations can be obtained by weighting the following six evaluation indicators:

[0173] ① Based on the extended words in the second translated text corresponding to the candidate translations, their prior probabilities are calculated according to their extension type, similarity to the corresponding non-standard words in the first translated text, and word frequency. Candidate translations with higher probabilities are then selected. (Example: If only extension type is considered, let the prior probabilities for phonetic spelling, cognates, and easily confused words be [0.7, 0.2, 0.1], and the prior probabilities for hiragana and katakana in phonetic spelling be [0.8, 0.2], then the prior probability for phonetic spelling-hiragana is 0.7*0.8=0.56). ② Input the second translated text corresponding to the candidate translations into the word segmentation model, calculate the similarity between word segmentation and lexical / syntactic structure annotation information and the first translated text, and select candidate translations corresponding to the second translated text with higher similarity. ③ Calculate the perplexity of the second and first translated texts using a language model, and select candidate translations corresponding to the first translated text whose perplexity is lower than that of the first translated text and whose perplexity difference exceeds a second preset threshold. ④ Input the candidate translations into a word segmentation model, calculate the similarity between the word segmentation and lexical / syntactic structure annotation information and the candidate translations corresponding to the first translated text, and select candidate translations with higher similarity. ⑤ Calculate the string similarity among all candidate translations, and select candidate translations with higher similarity. ⑥ Calculate the similarity between the translations corresponding to extended words in all candidate translations.

[0174] It should be noted that evaluation index ④ can be determined by other evaluation indexes of the candidate translations corresponding to the first translated text. If the fluency and translation quality of the candidate translations corresponding to the text to be translated are poor, the weight of index ④ will also be reduced accordingly.

[0175] Furthermore, since different second translation texts can yield the same candidate translations through the enhanced translation model, the relevance of the candidate translation can be obtained by weighting the evaluation metrics of the candidate translations corresponding to the different second translation texts.

[0176] Optionally, prior to step 207 above, the sample construction method provided in this application embodiment can also use evaluation indicators ①~③ for calculating the relevance of candidate translations to screen at least one second translation text, and screen out X second translation texts from M second translation texts, so as to improve the efficiency during actual translation and reduce the power consumption of the sample construction device.

[0177] In the sample construction method provided in this application embodiment, on the one hand, since this application can restore non-standard words in the first translated text to at least one standard word, generating at least one second translated text, it can restore the first translated text containing non-standard words to a standard first translated text, avoiding translation errors caused by the presence of non-standard words; on the other hand, since this application can simultaneously input part or all of the standard second translated text and the original first translated text into the translation model when translating the first translated text, it can output a more accurate translation result. Thus, the sample construction method provided in this application embodiment can improve the accuracy of the translation model.

[0178] Optionally, prior to step 206 above, the sample construction method provided in this application embodiment may further include step 208 below.

[0179] Step 208: After inputting the first translated text into the first word segmentation model, the first translated text is segmented into K words, and each of the K words is identified as a non-standard word to obtain the recognition result corresponding to each word.

[0180] Among them, the recognition result corresponding to a word segment is used to characterize whether a word segment belongs to a non-standard word. If a word segment belongs to a non-standard word, the recognition result corresponding to a word segment includes the standard type corresponding to the word segment.

[0181] In this embodiment of the application, the first word segmentation model is trained based on the target training sample set, where K is an integer greater than 1.

[0182] For example, the sample construction device can also use the word segmentation model trained in method 2 above to predict the labels of non-standard words in the enhanced text, and introduce the corresponding standard type label vectors in the training of the translation model, so that the model learns the semantic relationship between the extended words and the first keyword corresponding to the out-of-vocabulary words, thereby enhancing the prediction and translation capabilities of the extended words.

[0183] It is understandable that the canonical type label vector and the word vector have the same dimension. The vector of the extended word in the enhanced sentence is added to the vector corresponding to the canonical type label predicted by the word segmentation model to obtain the final representation vector of the extended word.

[0184] Specifically, taking Japanese as an example, as shown in Table 3, extended words include types such as pronunciation spelling, cognates, and easily confused words. Each type has further subcategories such as hiragana, katakana, simplified Chinese, traditional Chinese, and recombinant words. Therefore, there are multiple combinations of word standardization type label types. During training, the label vectors for each standardization type can be randomly initialized, or the weighted average of the word vectors corresponding to each component word (e.g., pronunciation spelling, cognates, hiragana, katakana, etc.) can be used as the initial vector, and the standardization type label vectors can be iteratively optimized through model training.

[0185] It should be noted that the label prediction of the word segmentation model for the extended words may be different from the label of the actual extended form of the word replacement. However, the enhanced sentences with these incorrect label predictions of the standard type are not corrected, but are retained in a certain proportion, thereby enhancing the robustness of the translation model and enabling the model to learn to output the correct translation even when the extended word labels are incorrect.

[0186] Optionally, the sample construction device can use parallel corpus training samples and target training samples to enhance the training of the basic translation model.

[0187] It is understandable that the output translations of each original text and all its corresponding extended texts are identical, thereby enhancing the translation model's robustness to non-standard expressions.

[0188] For example, the enhanced translation model can generate a word vector table containing expanded words and a canonical type label vector table.

[0189] In this embodiment, after the sample construction device inputs at least N second translated texts and the first translated text into the translation model, the input text can first be segmented using an enhanced word segmentation model to identify non-standard words and predict their extended forms. Then, by querying a vector table, the standard word vectors, extended word vectors, and standard type label vectors corresponding to the input text are input into the model, such as... Figure 4 As shown, the generated translation is obtained.

[0190] Thus, on the one hand, at the data level, the sample construction device can construct adversarial training data based on an expanded vocabulary to enhance the training of the translation model; on the other hand, at the model level, the sample construction device incorporates canonical type label vectors into the input encoding layer, allowing the model to learn the expanded forms of non-canonical words and the semantic relevance between non-canonical words and their corresponding canonical words in the text during training. This improves the robustness and quality of the translation model for the first translation text containing non-canonical words, ensuring that the translation model can output the correct translation regardless of whether the first translation text contains non-canonical words.

[0191] Optionally, after step 201 above, the sample construction method provided in this application embodiment may further include steps 301 and 302 as described below. Step 202 above can be specifically implemented through step a below.

[0192] Step 301: Display each compliant word corresponding to the first non-compliant word.

[0193] The first non-standard word can be one or more non-standard words from at least one set of non-standard words. In other words, the first non-standard word mentioned above can be one or more non-standard words.

[0194] For example, the sample construction device can display each conforming word corresponding to the first non-conforming word in descending order of relevance.

[0195] For example, the sample construction device can calculate the relevance of each conforming word corresponding to the first non-conforming word using the following formula (2).

[0196] (Formula 2)

[0197] in, To expand the prior probabilities of words that do not conform to the norm in the vocabulary; The lexical similarity between non-standard words and their restored standard words; Restoring non-standard words; , , These are adjustable weighting coefficients.

[0198] For example, for If only parts of speech are considered, if a non-standard word has the same lexical structure as its restored standard word, then... It can be 1, if the lexical structure of the non-standard word differs from that of its restored standard word, then... It can be 0.

[0199] Step 302: Receive the first input for the target compliant word in the displayed compliant words.

[0200] For example, the target conforming words mentioned above are one or more conforming words among the displayed conforming words.

[0201] In one example, the target conforming word mentioned above can be the conforming word corresponding to the same non-conforming word.

[0202] In one example, the target conforming words mentioned above can include conforming words corresponding to multiple different non-conforming words.

[0203] In one example, when there are multiple compliant words corresponding to non-compliant words in the above-mentioned target compliant words, the sample construction device restores each selected compliant word of the user.

[0204] Exemplarily, the above-mentioned target compliant word can be a compliant word selected by the user to replace the non-compliant word.

[0205] Exemplarily, the above-mentioned first input is used to select the compliant word to be restored from the displayed compliant words.

[0206] Exemplarily, the above-mentioned first input can be a touch input, a specific voice input or a specific gesture input of the user to the target compliant word, and the embodiments of the present application do not limit this.

[0207] For example, the first input can be a click input of the user to the target compliant word.

[0208] Step a: In response to the first input, restore the first non-compliant word in the first translation text to the target compliant word to generate at least one second translation text.

[0209] Exemplarily, if there are non-compliant words in the first translation text that have not been manually restored by the user, the electronic device can restore them according to the above relevant steps to generate at least one second translation text.

[0210] Exemplarily, as Figure 5 shown in (a) of, the sample construction device can display the compliant words "両親(两亲)" and "良心" corresponding to the non-compliant word "りょうしん(liangqin)". Then, the sample construction device receives a click input (i.e., the first input) of the user to the target compliant word "両親(两亲)", as Figure 5 shown in (b) of, restore the non-compliant word "りょうしん(liangqin)" to the target compliant word "両親(两亲)", and generate the second translation text "りょうしんは学校に勤める(两亲 / 父母在学校工作)".

[0211] In this way, since the sample construction device can display the compliant words corresponding to the non-compliant words, and the user selects the target compliant word to be restored through the first input, the generated second translation text can be correspondingly reduced, thereby reducing the power consumption required for translation.

[0212] Optionally, the sample construction method provided by the embodiments of the present application can construct a corresponding extended vocabulary according to the linguistic features of different languages for application to different translation languages and language directions.

[0213] This application provides a sample construction method. Figure 6 The diagram illustrates a flowchart of a translation model provided in an embodiment of this application, whereby the translation model is trained using target training samples. Figure 6 As shown, the sample construction method provided in this application embodiment may include the following steps 601 to 607.

[0214] Step 601: Obtain the text to be translated.

[0215] Step 602: Automatically identify whether there are any non-standard words in the text to be translated.

[0216] Step 603: If there are non-standard words in the text to be translated, sort the restoration results of the non-standard words according to their credibility and present them to the user.

[0217] Step 604: In response to the user's selection of the first input of the restoration result of non-standard words, restore the non-standard words and generate at least one second translation text.

[0218] Step 605: Restore the non-standard words that the user did not select to restore, and generate at least one second translation text.

[0219] Step 606: Input at least one second translation text into the first translation model to perform text translation and obtain at least one candidate translation.

[0220] Step 607: Determine the target translation from at least one candidate translation and output the target translation.

[0221] The sample construction method provided in this application can be executed by a sample construction device. This application uses an example of a sample construction device executing the sample construction method to illustrate the sample construction device provided in this application.

[0222] Figure 7 A schematic diagram of a possible structure of the sample construction apparatus involved in an embodiment of this application is shown. For example... Figure 7 As shown, the sample construction device 70 may include: an acquisition module 71, a processing module 72, and a construction module 73.

[0223] The acquisition module 71 is used to acquire parallel corpus training samples, which contain original text and carry the normative type label corresponding to each keyword in the original text; the processing module 72 is used to replace the first keyword in the original text of the parallel corpus training samples acquired by the acquisition module 71 with at least one first non-normative word corresponding to the first keyword to generate at least one extended text; the processing module 72 is also used to replace the first normative type label corresponding to the first keyword in the parallel corpus training samples acquired by the acquisition module 71 with the second normative type label corresponding to the first non-normative word to obtain the parallel corpus training samples after label replacement; the construction module 73 is used to construct a target training sample based on the parallel corpus training samples after label replacement processed by the processing module 72 and at least one extended text.

[0224] One possible implementation is that the parallel corpus training sample is a single parallel corpus training sample from a set of parallel corpus training samples; the aforementioned processing module 72 is specifically used for:

[0225] Based on the word frequency of each keyword in the original text in the parallel corpus training sample set, at least one first keyword is determined from the original text. Each first keyword in the original text is replaced with its corresponding first non-standard word to generate the first extended text.

[0226] Wherein, the first extended text is any one of the at least one extended texts.

[0227] One possible implementation is that the parallel corpus training sample is a single parallel corpus training sample from the parallel corpus training sample set, and the number of expanded texts is N, where N is a positive integer;

[0228] The aforementioned processing module 72 is further configured to initialize the word feature information of the out-of-vocabulary words when, after replacing the first keyword in the original text with at least one first non-standard word corresponding to the first keyword to generate at least one extended text, the second extended text among the N extended texts contains out-of-vocabulary words not included in the parallel corpus training sample set.

[0229] The initialization process includes one of the following:

[0230] The word feature information of the out-of-vocabulary word is weighted and averaged according to the first keyword corresponding to the out-of-vocabulary word and the word frequency of each non-standard word corresponding to the first keyword corresponding to the out-of-vocabulary word in each of the N extended texts in the parallel corpus training sample set.

[0231] We use the feature information of cognate words corresponding to out-of-vocabulary words to perform a weighted average of the feature information of out-of-vocabulary words;

[0232] Set the feature information of out-of-vocabulary words to 0;

[0233] The feature information of out-of-vocabulary words is randomly initialized.

[0234] One possible implementation is that the parallel corpus sample is a parallel corpus training sample from a set of parallel corpus samples; the above device also includes: a translation module;

[0235] The aforementioned processing module 72 is further configured to, after the construction module 73 constructs the target training sample based on the parallel corpus training sample with the replaced labels and at least one extended text, restore at least one non-standard word in the first translation text to a standard word, so as to generate M second translation texts, and restore one non-standard word to at least one standard word.

[0236] The aforementioned translation module is used to input the first feature information corresponding to the first translated text and the second feature information corresponding to X of the M second translated texts obtained by the processing module 72 into the first translation model for text translation to obtain the target translation. The first feature information includes the text feature information of the first translated text and the feature information of the normative type label corresponding to the non-standard words in the first translated text. The second feature information includes the text feature information of the second translated text and the feature information of the normative type label corresponding to the non-standard words in the second translated text.

[0237] The first translation model is trained based on the target training sample set, which includes multiple target training samples. Each target training sample corresponds to a parallel corpus training sample in the parallel corpus training sample set. M and X are positive integers, and X is less than or equal to M.

[0238] In one possible implementation, the above-mentioned device further includes: a word segmentation module;

[0239] The aforementioned word segmentation module is used to input the first translation text into the first word segmentation model before the processing module 72 restores at least one non-standard word in the first translation text to a standard word to generate M second translation texts. After the first translation text is input into the first word segmentation model, the module performs word segmentation on the first translation text to obtain K words. Then, it identifies non-standard words in each of the K words to obtain the identification result corresponding to each word. The identification result corresponding to a word is used to characterize whether a word belongs to a non-standard word. In the case that a word belongs to a non-standard word, the identification result corresponding to a word includes the standard type corresponding to the word.

[0240] The first word segmentation model is trained based on the target training sample set, where K is an integer greater than 1.

[0241] One possible implementation method is that non-compliant words include at least one of the following: containing pinyin reading and writing, containing misspellings, containing substitutions of cognate words, and containing character shape errors.

[0242] This application provides a sample construction apparatus. This apparatus can replace keywords in the original text of parallel corpus training samples to generate at least one extended text, thereby expanding the vocabulary covered by the parallel corpus training samples. Simultaneously, it replaces the normative type labels corresponding to the keywords with normative type labels corresponding to non-normative words, obtaining parallel corpus training samples with replaced labels, thus enriching the content contained in the parallel corpus training samples. Finally, the sample construction apparatus can construct a target training sample based on the parallel corpus training samples with replaced labels and at least one extended text. Therefore, the target training sample can contain non-normative words and their corresponding normative type labels, thereby enriching the content of the parallel corpus training samples and enabling them to have more flexible training content.

[0243] The sample construction device in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. The embodiments of this application do not specifically limit the device.

[0244] The sample construction apparatus in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0245] The sample construction apparatus provided in this application embodiment can achieve... Figures 2 to 6The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0246] Optionally, such as Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 801 and a memory 802. The memory 802 stores a program or instructions that can run on the processor 801. When the program or instructions are executed by the processor 801, they implement the various steps of the above-described sample construction method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0247] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0248] Figure 9 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0249] The electronic device 900 includes, but is not limited to, components such as: radio frequency unit 901, network module 902, audio output unit 903, input unit 904, sensor 905, display unit 906, user input unit 907, interface unit 908, memory 909, and processor 910.

[0250] Those skilled in the art will understand that the electronic device 900 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 910 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0251] The processor 910 is configured to: acquire parallel corpus training samples, wherein the parallel corpus training samples contain original text and carry a canonical type label corresponding to each keyword in the original text; replace the first keyword in the original text of the acquired parallel corpus training samples with at least one first non-canonical word corresponding to the first keyword to generate at least one extended text; replace the first canonical type label corresponding to the first keyword in the acquired parallel corpus training samples with a second canonical type label corresponding to the first non-canonical word to obtain a parallel corpus training sample with replaced labels; and construct a target training sample based on the parallel corpus training sample with replaced labels and at least one extended text.

[0252] Optionally, the parallel corpus training sample is a single parallel corpus training sample from a set of parallel corpus training samples; the processor 910 described above is specifically used for:

[0253] Based on the word frequency of each keyword in the original text in the parallel corpus training sample set, at least one first keyword is determined from the original text. Each first keyword in the original text is replaced with its corresponding first non-standard word to generate the first extended text.

[0254] Wherein, the first extended text is any one of the at least one extended texts.

[0255] Optionally, the parallel corpus training sample is a parallel corpus training sample in the parallel corpus training sample set, and the number of expanded texts is N, where N is a positive integer;

[0256] The processor 910 described above is further configured to initialize the word feature information of the out-of-vocabulary words when, after replacing the first keyword in the original text with at least one first non-standard word corresponding to the first keyword to generate at least one extended text, the second extended text among the N extended texts contains out-of-vocabulary words not included in the parallel corpus training sample set.

[0257] The initialization process includes one of the following:

[0258] The word feature information of the out-of-vocabulary word is weighted and averaged according to the first keyword corresponding to the out-of-vocabulary word and the word frequency of each non-standard word corresponding to the first keyword corresponding to the out-of-vocabulary word in each of the N extended texts in the parallel corpus training sample set.

[0259] We use the feature information of cognate words corresponding to out-of-vocabulary words to perform a weighted average of the feature information of out-of-vocabulary words;

[0260] Set the feature information of out-of-vocabulary words to 0;

[0261] The feature information of out-of-vocabulary words is randomly initialized.

[0262] Optionally, a parallel corpus sample is a parallel corpus training sample from a set of parallel corpus samples;

[0263] The processor 910 described above is also used to construct a target training sample based on parallel corpus training samples with replacement labels and at least one extended text, and then restore at least one non-standard word in the first translated text to a standard word to generate M second translated texts, where one non-standard word is restored to at least one standard word.

[0264] The processor 910 is further configured to input the first feature information corresponding to the first translated text and the second feature information corresponding to X of the M second translated texts obtained by the processor 910 into the first translation model for text translation to obtain the target translation. The first feature information includes the text feature information of the first translated text and the feature information of the standard type label corresponding to the non-standard words in the first translated text. The second feature information includes the text feature information of the second translated text and the feature information of the standard type label corresponding to the non-standard words in the second translated text.

[0265] The first translation model is trained based on the target training sample set, which includes multiple target training samples. Each target training sample corresponds to a parallel corpus training sample in the parallel corpus training sample set. M and X are positive integers, and X is less than or equal to M.

[0266] Optionally, the processor 910 is used to restore at least one non-standard word in the first translated text to a standard word before generating M second translated texts. After inputting the first translated text into the first word segmentation model, the processor segments the first translated text into K words and identifies non-standard words for each of the K words to obtain an identification result corresponding to each word. The identification result corresponding to a word is used to characterize whether a word belongs to a non-standard word. In the case that a word belongs to a non-standard word, the identification result corresponding to a word includes the standard type corresponding to the word.

[0267] The first word segmentation model is trained based on the target training sample set, where K is an integer greater than 1.

[0268] Optionally, non-standard words include at least one of the following: containing pinyin reading and writing, containing misspellings, containing substitutions of cognate characters, and containing character shape errors.

[0269] This application provides an electronic device that can replace keywords in the original text of a parallel corpus training sample with at least one non-standard word corresponding to the keyword, generate at least one extended text, and replace the standard type label corresponding to the keyword with the standard type label corresponding to the non-standard word, thus obtaining a parallel corpus training sample with replaced labels. Then, the electronic device can construct a target training sample based on the parallel corpus training sample with replaced labels and at least one extended text. Therefore, the target training sample can contain non-standard words and their corresponding standard type labels, enabling the translation model trained on the target training sample to translate non-standard words, thereby improving the accuracy of the translation model.

[0270] It should be understood that, in this embodiment, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The GPU 9041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0271] The memory 909 can be used to store software programs and various data. The memory 909 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 909 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 909 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0272] Processor 910 may include one or more processing units; optionally, processor 910 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 910.

[0273] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described sample construction method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0274] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0275] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described sample construction method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0276] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0277] This application provides a computer program product stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-described sample construction method embodiments and achieve the same technical effects. To avoid repetition, further details are omitted here.

[0278] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0279] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0280] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A sample construction method, characterized in that, The method includes: Obtain parallel corpus training samples, wherein the parallel corpus training samples contain the original text and carry the canonical type label corresponding to each keyword in the original text; The first keyword in the original text is replaced with at least one first non-standard word corresponding to the first keyword to generate at least one extended text; Replace the first standard type label corresponding to the first keyword with the second standard type label corresponding to the first non-standard word to obtain the parallel corpus training sample after label replacement; the parallel corpus training sample is a parallel corpus training sample in the parallel corpus training sample set, and the number of extended texts is N, where N is a positive integer; If the second extended text in the N extended texts contains out-of-vocabulary words not included in the parallel corpus training sample set, the feature information of the out-of-vocabulary words is initialized; wherein the initialization process includes at least one of the following: The feature information of the out-of-vocabulary word is weighted and averaged according to the first keyword corresponding to the out-of-vocabulary word and the word frequency of each non-standard word corresponding to the first keyword corresponding to the out-of-vocabulary word in each of the N extended texts in the parallel corpus training sample set; Using the feature information of the cognate words corresponding to the out-of-vocabulary words, a weighted average of the feature information of the out-of-vocabulary words is calculated. Set the feature information of the unregistered words to 0; Randomly initialize the feature information of the unregistered words; Based on the parallel corpus training samples with replaced labels and the at least one extended text, a target training sample is constructed.

2. The method according to claim 1, characterized in that, The parallel corpus training sample is a single parallel corpus training sample from the set of parallel corpus training samples. The step of replacing the first keyword in the original text with at least one first non-standard word corresponding to the first keyword to generate at least one extended text includes: Based on the word frequency of each keyword in the original text in the parallel corpus training sample set, at least one first keyword is determined from the original text, and each first keyword in the original text is replaced with its corresponding first non-standard word to generate the first extended text. Wherein, the first extended text is any one of the at least one extended texts.

3. The method according to claim 1, characterized in that, The parallel corpus training sample is a single parallel corpus training sample from the set of parallel corpus training samples. After constructing the target training sample based on the parallel corpus training samples with replaced labels and the at least one expanded text, the method further includes: Each extended text in the first translated text is restored to a standard word to generate M second translated texts. A non-standard word is restored to at least one standard word. The first feature information corresponding to the first translated text and the second feature information corresponding to X of the M second translated texts are input into the first translation model to perform text translation in order to obtain the target translation. The first feature information includes the text feature information of the first translated text and the feature information of the standard type label corresponding to the non-standard words in the first translated text. The second feature information includes the text feature information of the second translated text and the feature information of the standard type label corresponding to the non-standard words in the second translated text. The first translation model is trained based on a target training sample set, which includes multiple target training samples. Each target training sample corresponds to a parallel corpus training sample in the parallel corpus training sample set. M and X are positive integers, and X is less than or equal to M.

4. The method according to claim 3, characterized in that, Before restoring at least one non-standard word in the first translated text to a standard word to generate M second translated texts, the method further includes: After inputting the first translated text into the first word segmentation model, the first translated text is segmented into K words. Each of the K words is then identified as a non-standard word, and an identification result is obtained for each word. The identification result for a word is used to characterize whether the word is a non-standard word. If the word is a non-standard word, the identification result for the word includes the standard type corresponding to the word. The first word segmentation model is trained based on the target training sample set, where K is an integer greater than 1.

5. The method according to any one of claims 1 to 3, characterized in that, The non-standard words include at least one of the following: containing phonetic spelling errors, containing misspellings, containing substitutions of cognate words, and containing character shape errors.

6. A sample construction apparatus, characterized in that, The device includes: an acquisition module, a processing module, and a construction module; The acquisition module is used to acquire parallel corpus training samples, which contain original text and carry the standard type label corresponding to each keyword in the original text; The processing module is used to replace the first keyword in the original text of the parallel corpus training sample obtained by the acquisition module with at least one first non-standard word corresponding to the first keyword, so as to generate at least one extended text; the parallel corpus training sample is a parallel corpus training sample in the parallel corpus training sample set, and the number of extended texts is N, where N is a positive integer; The processing module is further configured to initialize the word feature information of the out-of-vocabulary words when, after replacing the first keyword in the original text with at least one first non-standard word corresponding to the first keyword to generate at least one extended text, the second extended text among the N extended texts contains out-of-vocabulary words not included in the parallel corpus training sample set. The processing module is further configured to replace the first normative type label corresponding to the first keyword in the parallel corpus training sample obtained by the acquisition module with the second normative type label corresponding to the first non-normative word, so as to obtain the parallel corpus training sample after label replacement. The construction module is used to construct a target training sample based on the parallel corpus training sample with replaced labels after being processed by the processing module and the at least one extended text. The initialization process includes one of the following: The word feature information of the unregistered word is weighted and averaged according to the first keyword corresponding to the unregistered word and the word frequency of each non-standard word corresponding to the first keyword corresponding to the unregistered word in each of the N extended texts in the parallel corpus training sample set. Using the feature information of the cognate words corresponding to the out-of-vocabulary words, a weighted average of the feature information of the out-of-vocabulary words is calculated. Set the feature information of the unregistered words to 0; The feature information of the unregistered words is randomly initialized.

7. The apparatus according to claim 6, characterized in that, The parallel corpus training sample is a single parallel corpus training sample from the set of parallel corpus training samples. The processing module is specifically used for: Based on the word frequency of each keyword in the original text in the parallel corpus training sample set, at least one first keyword is determined from the original text, and each first keyword in the original text is replaced with its corresponding first non-standard word to generate the first extended text. Wherein, the first extended text is any one of the at least one extended texts.

8. The apparatus according to claim 6, characterized in that, The parallel corpus training sample is a single parallel corpus training sample from the set of parallel corpus training samples. The device further includes: a translation module; The processing module is further configured to, after the construction module constructs the target training sample based on the parallel corpus training sample with the replaced labels and the at least one extended text, restore at least one non-standard word in the first translated text to a standard word, so as to generate M second translated texts, and restore one non-standard word to at least one standard word; The translation module is used to input the first feature information corresponding to the first translated text and the second feature information corresponding to X of the M second translated texts obtained by the processing module into the first translation model for text translation to obtain the target translation. The first feature information includes the text feature information of the first translated text and the feature information of the standard type label corresponding to the non-standard words in the first translated text. The second feature information includes the text feature information of the second translated text and the feature information of the standard type label corresponding to the non-standard words in the second translated text. The first translation model is trained based on a target training sample set, which includes multiple target training samples. Each target training sample corresponds to a parallel corpus training sample in the parallel corpus training sample set. M and X are positive integers, and X is less than or equal to M.

9. The apparatus according to claim 8, characterized in that, The device further includes: a word segmentation module; The word segmentation module is used to input the first translated text into a first word segmentation model before the processing module restores at least one non-standard word in the first translated text to a standard word to generate M second translated texts. The module then segments the first translated text into K words and identifies each of the K words as a non-standard word, obtaining an identification result for each word. The identification result for each word is used to characterize whether the word is a non-standard word. If the word is a non-standard word, the identification result includes the standard type corresponding to the word. The first word segmentation model is trained based on the target training sample set, where K is an integer greater than 1.

10. The apparatus according to any one of claims 6 to 8, characterized in that, The non-standard words include at least one of the following: containing pinyin reading and writing, containing misspellings, containing substitutions of cognate characters, and containing character shape errors.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the sample construction method as described in any one of claims 1 to 5.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the sample construction method as described in any one of claims 1 to 5.