Corpus generation method, medium, electronic equipment and program product

By extracting and utilizing the characteristics of each language, the mixed corpus is generated, and the problem of difficulty in collecting mixed corpus is solved, and efficient and accurate mixed corpus generation is achieved.

CN119940373APending Publication Date: 2025-05-06MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411823196.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

It is difficult to collect mixed corpus, and it is difficult for the prior art to generate mixed corpus efficiently and accurately.

Method used

By extracting the first pronunciation features, lexical features and grammatical features of each language, the corresponding vocabulary is determined, and when generating mixed corpus, the unique vocabulary is selected to retain the pronunciation features, and the number and accuracy of the vocabulary is controlled by the number of characters and grammatical features.

Benefits of technology

It realizes fine control of mixed corpus, improves generation efficiency and accuracy, and ensures the accuracy of the generated corpus in pronunciation and vocabulary expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940373A_ABST
    Figure CN119940373A_ABST
Patent Text Reader

Abstract

The invention relates to a corpus generation method, a medium, electronic equipment and a program product. The method comprises the following steps: extracting at least one of a first pronunciation feature, a vocabulary feature and a grammar feature of a first corpus corresponding to each language; for any first language, determining a first vocabulary corresponding to the first language according to a second pronunciation feature in the first pronunciation features; determining a second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature; determining a third vocabulary corresponding to the first language according to part-of-speech features and semantic features in the grammatical features; and if the second corpus comprises the first vocabulary, a fourth vocabulary is selected from the third corpus to be translated to generate a fourth corpus, and the fourth corpus comprises the second vocabulary and the third vocabulary. The fourth corpus is accurately and efficiently generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular, to a method for generating corpus, a medium, an electronic device and a program product. Background Art

[0002] With the development of Internet technology, content mixed with different languages ​​is constantly being produced. These contents are called lingua franca. Common lingua franca include English lingua franca, Chinese lingua franca and lingua franca of different dialects. A lingua franca can be a mixture of two or more languages. However, due to the limitations of the text and voice data of the lingua franca itself, it is difficult to collect lingua franca corpus. Summary of the invention

[0003] The present disclosure provides a corpus generation method, medium, electronic device and program product to improve the generation efficiency of mixed corpus.

[0004] In a first aspect, the present disclosure provides a method for generating a corpus, the method comprising: Extracting at least one of a first pronunciation feature, a lexical feature, and a grammatical feature of a first corpus corresponding to each language; For any first language, determining a first vocabulary corresponding to the first language according to a second pronunciation feature in the first pronunciation feature, wherein the second pronunciation feature belongs to the first language and does not belong to the second language; determining a second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature, the number of characters being less than or equal to a threshold; Determining a third vocabulary corresponding to the first language according to the part-of-speech feature and the semantic feature in the grammatical feature; If the second corpus includes the first vocabulary, a fourth vocabulary is selected from the third corpus for translation to generate a fourth corpus, wherein the third corpus is a part of the second corpus and does not include the first vocabulary, and the fourth corpus includes the second vocabulary and the third vocabulary.

[0005] In a second aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when the computer program is executed by a processing device.

[0006] In a third aspect, the present disclosure provides an electronic device, including: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the method in the first aspect.

[0007] In a fourth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0008] Through the above technical solution, when the second corpus is obtained, if it is determined that the second corpus includes the first vocabulary, a fourth vocabulary is selected from part of the corpus in the second corpus for translation, and the fourth corpus includes the second vocabulary and the third vocabulary. Here, the first vocabulary is a vocabulary corresponding to the first language determined according to the second pronunciation feature in the first pronunciation feature for any first language. When the second corpus includes the first vocabulary, the first vocabulary is not translated, so that the pronunciation feature of the first vocabulary in the first language is retained in the generated fourth corpus, thereby realizing the adjustment of the pronunciation feature of the generated fourth corpus. The second vocabulary is a vocabulary corresponding to the first language determined according to the number of characters in the vocabulary feature, wherein the number of characters is less than or equal to the threshold value, so that by selecting the second vocabulary with the number of characters less than the threshold value to translate the vocabulary in the second corpus, the number of characters contained in the generated fourth corpus can be limited. In addition, the third vocabulary is a vocabulary corresponding to the first language determined according to the part-of-speech feature and the semantic feature in the grammatical feature, which ensures that after translating part of the vocabulary in the second corpus, a fourth corpus with accurate grammar is obtained. In summary, when converting the second corpus into the fourth corpus, the technical solution provided by the present disclosure can more finely control the generation of the corpus.

[0009] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 FIG. 4 is a diagram showing an example of an application environment of a method for generating corpus according to an exemplary embodiment of the present invention.

[0011] Figure 2 The present invention is a flowchart of a method for generating corpus according to an exemplary embodiment.

[0012] Figure 3 This is an example diagram of obtaining typical features of English in a method for generating corpus according to an exemplary embodiment.

[0013] Figure 4 This is an example diagram of obtaining typical features of Chinese in a method for generating corpus according to an exemplary embodiment.

[0014] Figure 5 This is an example diagram of the relationship between a target vocabulary and a first associated word in a method for generating a corpus according to an exemplary embodiment.

[0015] Figure 6 This is an example diagram of the relationship between a target vocabulary and a second associated word in a method for generating a corpus according to an exemplary embodiment.

[0016] Figure 7 This is an example diagram showing a method for generating a corpus in which the second corpus is in multiple languages ​​according to an exemplary embodiment.

[0017] Figure 8 This is an example diagram showing a method for generating a corpus in which the second corpus is in multiple languages ​​according to an exemplary embodiment.

[0018] Fig. 9 This is an example diagram showing a method for generating a corpus in which the second corpus is in multiple languages ​​according to an exemplary embodiment.

[0019] Fig.10 The figure is an application flow chart of a method for generating corpus according to an exemplary embodiment of the present disclosure.

[0020] Fig.11 It is a detailed process diagram of obtaining typical features in a method for generating a corpus according to an exemplary embodiment of the present disclosure.

[0021] Fig.12 It is a structural block diagram of a corpus generation device according to an exemplary embodiment of the present disclosure.

[0022] Fig.13 It is a structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The specific implementation of the present disclosure is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the present disclosure, and is not used to limit the present disclosure.

[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0025] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0026] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0027] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0028] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0030] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0031] First, some terms in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.

[0032] 1. Mixed languages Mixed languages ​​refer to mixtures of two or more languages ​​formed under certain social conditions. Mixed language types can include mixtures of different large languages, mixtures of dialects, and mixtures of large languages ​​and dialects. For example, mixtures of different large languages ​​can be mixtures of different languages, such as the mixture of English and Chinese, the mixture of English and French, the mixture of Korean and Japanese, etc.; mixtures of dialects can be mixtures of dialects from different regions of the same language, such as the mixture of Cantonese and Min, the mixture of Xiang and Gan; mixtures of large languages ​​and dialects can be mixtures of a certain language and a certain dialect of another language, such as the mixture of English and Cantonese, the mixture of Japanese and Min.

[0033] The corpus composed of mixed languages ​​can be a mixed corpus, which can be a corpus containing texts in multiple languages, which can be used to train multilingual models, study cross-language semantic relations, and process and analyze multilingual data.

[0034] 2. IPA (International Phonetic Alphabet) IPA is a system for phonetic notation based on the Latin alphabet, designed by the International Phonetic Association as a standardized way to represent spoken sounds. IPA is divided into strict phonetic notation and broad phonetic notation.

[0035] Among them, strict phonetic symbols use phonemes to mark pronunciation; broad phonetic symbols are a phoneme system of speech compiled on the basis of strict phonetic symbols, and speech is marked according to phonemes, that is, only phonemes are recorded, and phoneme variants and other non-essential accompanying phenomena are not recorded.

[0036] 3. Naturalness In natural language processing (NLP) and language generation tasks, naturalness usually refers to the naturalness and fluency of the generated text or speech. Criteria for evaluating naturalness include grammatical correctness, semantic consistency, coherence, and comprehensibility.

[0037] 4. Syllables A syllable is the basic unit of pronunciation in a language, usually composed of one or more phonemes. A syllable usually contains a vowel phoneme and possibly a consonant phoneme.

[0038] 5. Phonemes A phoneme is the smallest unit of speech in a language and is used to distinguish the meaning of a word. A phoneme is abstract and does not correspond to a specific pronunciation, but has different phonetic representations in a specific language.

[0039] In the related technologies, language (language) research mainly focuses on single languages, while research on mixed languages ​​is very rare, mainly because the text and voice data of mixed languages ​​are extremely limited. From the above introduction, we know that mixed languages ​​can include not only the mixture of different languages, but also the mixture of different languages ​​and dialects, or the mixture of dialects. Among the few dialect data, there is a problem of few dialect types, and even less dialect data for different scenarios.

[0040] At present, the cost of collecting mixed corpora is much higher than that of monolingual corpora. It takes a long time for third-party data companies to collect mixed corpora, and the quality cannot be guaranteed. In addition, a large amount of mixed corpora is the basis of deep learning, so it is particularly important to generate mixed corpora efficiently, accurately and in batches.

[0041] In view of this, the embodiments of the present disclosure provide a method for generating corpus, which can extract language particularity based on the theory of linguistic typology while respecting the particularity of the language itself, and propose a mixed framework between languages ​​with distant kinship, based on which a large amount of mixed corpus can be generated efficiently and accurately, so as to perform dialect synthesis and dialect variety recognition based on the generated mixed corpus.

[0042] The following introduces the application scenarios of the corpus generation method provided by the embodiments of the present disclosure.

[0043] The method for generating corpus provided in the embodiments of the present disclosure can be applied to any scenario where mixed corpus needs to be generated, and the method should be applied to products in these scenarios. For example, resource recovery systems, medical and health systems, and content recommendation systems. Below, some application scenarios of the generation of corpus in the embodiments of the present disclosure are exemplified: (1) Resource recovery system The resource recovery system is a technical platform for managing and optimizing the debt resource recovery process, helping enterprises to efficiently carry out resource recovery activities and reduce bad debt rates. For multilingual customers, a large amount of mixed corpus can be generated during the resource recovery process, and the mixed corpus can be used to match agents according to their different preferences. On this basis, the language model is trained using the matched sample data to obtain a semantically concise and accurate resource recovery speech model.

[0044] (2) Medical and health system The medical health system is used to provide a series of services and functions to improve and maintain people's health. That is, the medical health system can assist doctors in communicating with different patients. By using automatically generated mixed corpus to train the medical health system, it can not only recognize the voice of patients using mixed languages, but also assist doctors in giving targeted communication suggestions. The communication suggestions can be recommended according to different patients and can also be in mixed languages.

[0045] (3) Marketing system Marketing systems are used to drive sales of products and services, increase brand awareness, and build and maintain relationships with target audiences. To achieve better marketing, more attractive and personalized advertising copy and marketing content can be created through mixed corpus to suit different markets and audiences. For example, social media creation can be done using mixed corpus.

[0046] In order to better understand the method, device, readable medium, electronic device and program product for generating corpus provided by the embodiments of the present disclosure, the application environment applicable to the embodiments of the present disclosure is described below.

[0047] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram showing an application environment of the method for generating corpus provided by an embodiment of the present disclosure. As an implementation mode, the method for generating corpus provided by an embodiment of the present disclosure can be applied to an electronic device. The electronic device can be, for example, Figure 1 The server 110 shown in FIG. 1 may be connected to the terminal device 120 via a network.

[0048] The network is used to provide a medium for a communication link between the server 110 and the terminal device 120. The network may include various connection types, such as a wired communication link, a wireless communication link, etc., which is not limited in the embodiment of the present disclosure.

[0049] It should be understood that Figure 1 The server 110, network, and terminal device 120 are merely illustrative. Depending on the implementation requirements, there may be any number of servers 110, networks, and terminal devices 120. For example, the server 110 may be a physical server 110, or a server 110 cluster composed of multiple servers 110. The terminal device 120 may be a smart phone, a landline phone, a personal computer, a mobile terminal, a tablet computer, and the like. It is understood that the embodiments of the present application may also allow multiple terminal devices 120 to access the server 110 at the same time.

[0050] The embodiments of the present disclosure are further explained below with reference to the accompanying drawings.

[0051] Figure 2 is a flowchart of a method for generating corpus according to an exemplary embodiment. The method can be applied to electronic devices. Figure 2 , the method for generating the corpus may include the following steps: Step S210: extracting at least one of a first pronunciation feature, a vocabulary feature, and a grammatical feature of a first corpus corresponding to each language.

[0052] In the embodiment of the present disclosure, each language may correspond to a plurality of typical features, wherein the plurality of typical features may include at least one of a first pronunciation feature, a vocabulary feature, and a grammatical feature of the first corpus.

[0053] The typical features of a language can be obtained by extracting features of different dimensions using a large number of sample corpora. For example, the typical features of Chinese can be extracted by using a large number of sample corpora related to Chinese. For another example, the typical features of English can be extracted by using a large number of sample corpora related to English.

[0054] In the process of obtaining typical features corresponding to different languages, the embodiments of the present disclosure may input monolingual corpora of all languages ​​into the language specificity extraction model to obtain typical features corresponding to each language. Exemplarily, the language specificity extraction model may be a mixed feature random field (Mixed TRF) language model, which can flexibly support the extraction of multiple language features by combining discrete language features and neural network language features.

[0055] On this basis, the embodiment of the present disclosure can store each language and its corresponding typical features according to the corresponding relationship. Specifically, the embodiment of the present disclosure can obtain a large amount of sample corpora of all languages ​​to form a sample data set. Among them, the sample corpora can be monolingual corpora and / or multilingual corpora. In addition, each sample corpus can include voice data and text data. Exemplarily, the voice data can be in mp3 format, and the text data can be a text version.

[0056] Accordingly, the disclosed embodiment can generate the pinyin and phonetic symbols of each text according to the voice data and text data, that is, convert the text into pinyin, and convert the pinyin into IPA (phonetic symbols). Then, the pinyin and phonetic symbols obtained by the conversion are used as inputs of the language specificity extraction model to obtain the first pronunciation feature of each language.

[0057] The first pronunciation feature of the first corpus corresponding to each language may be a unique property of pronunciation of each language, and the first pronunciation feature uses a unique phonetic system of each language. For example, the first pronunciation feature corresponding to Chinese includes the features of tones, the features of retroflex consonants, and the features of zero initial consonants. For another example, the first pronunciation feature corresponding to English includes the features of vowels, consonants, stresses, and weak pronunciations.

[0058] From the above introduction, we know that the first pronunciation feature corresponding to each language can be extracted through a language specificity extraction model, wherein the language specificity extraction model can include a speech extraction model. Exemplarily, the speech extraction model can be an acoustic model based on LSTM (Long Short-Term Memory) or a model obtained by combining a convolutional neural network and a recurrent neural network, based on which phonetic features can be effectively extracted from the corpus.

[0059] The input data of the speech extraction model can be the phonetic symbols of all speech in each language, and the output data can be a list of all phonemes, a list of all missing phonemes, and a list of unique phonemes in each language, etc. These lists of sound speed can be collectively referred to as the first pronunciation features. For example, the first pronunciation features of Chinese can be obtained based on the speech extraction model, and the first pronunciation features can include a list of all phonemes in Chinese, a list of all missing phonemes in Chinese, and a list of factors unique to Chinese, etc.

[0060] In other words, the embodiments of the present disclosure can extract the speech phonetic symbols corresponding to each language through the speech extraction model, extract features of the speech phonetic symbols, and compare the extracted features with features of other languages ​​to obtain the first pronunciation features corresponding to each language.

[0061] It should be noted that the input data of the speech extraction model (all speech phonetic symbols of the target language) can be pre-stored, that is, the input data can be searched and obtained according to the language in the sample corpus. Specifically, the embodiment of the present disclosure can search for the phonetic symbol information of the sample corpus from a pre-built phonetic symbol database, wherein each language in the phonetic symbol database can correspond to multiple speech phonetic symbols. For example, if the language of the sample corpus is English, the embodiment of the present disclosure can obtain the phonetic symbols of all English speech by searching the phonetic symbol database, and use it as the input of the speech extraction model.

[0062] In the disclosed embodiment, the first pronunciation feature may include a list of all phonemes, a list of all missing phonemes, and a list of unique phonemes for each language, wherein the list of unique phonemes may include multiple unique speeds and unique tones, and the list of missing phonemes may include multiple missing phonemes and missing tones. Among them, the unique phonemes may be phonemes that are unique / unique to the target language compared to other languages; the unique tones may be tones that are unique to the target language compared to other languages; the missing phonemes may be phonemes that the target language does not have compared to other languages; and the missing tones may be tones that the target language does not have compared to other languages.

[0063] Here, the unique phonemes, missing phonemes, unique tones and missing tones of the target language can be obtained by comparing the target language with all other languages, or by comparing the target language with another candidate language. Here, the unique phonemes can include unique initials, unique finals and the like.

[0064] As an example, when the target language is Chinese and the candidate language is English, the unique phonemes of Chinese can be determined by comparing the two. For example, the unique initial consonants "zh", "ch", "sh", etc. of Chinese can all be used as the unique phonemes of Chinese, and the unique finals "an", "en", "ong", etc. can all be used as the unique phonemes of Chinese.

[0065] Continuing with the above examples, the missing phonemes of Chinese compared to English may include consonant clusters, voiceless and unvoiced oppositions, and fricatives, such as the missing "str", "thin", and "then" phonemes of Chinese compared to English. The tones unique to Chinese compared to English include yinping, yangping, shangsheng, and qusheng. In addition, when the candidate language is Vietnamese, the missing tones of Chinese include sharp, interrogative, and stressed tones.

[0066] In the disclosed embodiment, unique phonemes, missing phonemes, unique tones and missing tones may be typical features of the speech dimension, that is, by extracting features of the speech dimension for each language, the first pronunciation features corresponding to each language may be obtained.

[0067] The lexical features of the first corpus corresponding to each language may be features related to vocabulary, which may be classification attributes of vocabulary in each language, and the classification attributes determine the function and usage of vocabulary in a sentence. The lexical features may include the number of characters. For example, the number of characters corresponding to "Apple" is 2, and the number of characters corresponding to "Apple" is 5. It can be seen that the number of characters of vocabulary with the same semantics corresponding to different languages ​​may be different. In addition, the lexical features may include vocabulary and / or missing vocabulary specific to each of the languages.

[0068] Optionally, the language-specific extraction model may also include a vocabulary extraction model, the input data of which may be all words and word vectors in each language, and the output data may be a list of all words in each language, a list of all missing words, and the number of characters in the words. The list of all missing words may include multiple missing words, and the missing words refer to words that are missing in the target language to express a certain concept. For example, based on the vocabulary extraction model, a list of all Chinese words, a list of all missing Chinese words, and the number of words with the same semantics as Chinese words may be obtained.

[0069] In other words, the embodiments of the present disclosure can extract the vocabulary corresponding to each language through the vocabulary extraction model, convert the vocabulary into vectors, and compare the converted vectors with other languages ​​to obtain the vocabulary features corresponding to each language.

[0070] It should be noted that the vocabulary specificity model can be pre-stored, that is, the input data can be obtained by searching for the language in the sample corpus. For example, if the language of the sample corpus is English, the embodiment of the present disclosure can obtain all the English words and word vectors by searching, and use them as input to the vocabulary extraction model. Optionally, the vocabulary list, missing words, and the number of characters of the vocabulary corresponding to each language can also be obtained by extracting the sample corpus.

[0071] In the disclosed embodiment, unique words, missing words and the number of characters in words may be typical features of the vocabulary dimension, that is, by extracting features of the vocabulary dimension for each language, the vocabulary features corresponding to each language may be obtained.

[0072] It can be seen from the above technical solution that the features of the vocabulary dimension can include unique vocabulary, missing vocabulary and the number of characters of vocabulary in each language. Among them, unique vocabulary can be vocabulary that is unique / unique to the target language compared with other languages, and missing vocabulary can be vocabulary that the target language does not have compared with other languages. Here, other languages ​​can be languages ​​other than the target language, or can also be candidate languages; the number of characters of vocabulary can be the number of characters of the translation vocabulary corresponding to each vocabulary in the target language.

[0073] As an example, the target language is Chinese and the candidate language is Eskimo. Through comparison, we know that the target language has words such as yubai, qipao and erhu, and the target language lacks words such as aput (snow on the ground), qana (falling snow) and piqsirpoq (pile of snow).

[0074] The grammatical features of the first corpus corresponding to each language may be a series of attributes representing grammatical relations in each language, which may affect the morphological changes of vocabulary and sentence structure. Specifically, the grammatical features may include part-of-speech features and semantic characteristics. For example, Chinese has quantifiers, but English does not.

[0075] Optionally, the language specificity extraction model may also include a grammar extraction model, the input data of which may be all words and word vectors in each language, and the output data may be the grammatical features of each language. Different languages ​​have different corresponding grammatical features, that is, the input data for obtaining the grammatical features are different. Exemplarily, the grammar extraction model may be a top-down analysis model, or a bottom-up analysis model, etc., wherein the top-down analysis model and the bottom-up analysis model may be a recurrent neural network or a short-term memory network.

[0076] like Figure 3 As shown, the embodiment of the present disclosure can input the first sample corpus into the grammar extraction model to obtain typical features of English. Here, the first sample corpus can be English corpus. Figure 4 As shown, in the embodiment of the present disclosure, the second sample corpus can be input into the grammar extraction model to obtain typical features of Chinese. Here, the second sample corpus can be Chinese corpus.

[0077] In addition, the embodiment of the present disclosure may also simultaneously input sample corpora of different languages ​​into the grammar extraction model to obtain grammatical features of each language.

[0078] The grammatical features in the disclosed embodiment are different from the first pronunciation features and part-of-speech features mentioned above. They mainly focus on the combination of words and words, and belong to the category of syntax. In other words, the grammatical features involve the part of speech of the word, that is, the grammatical expression, which is used to indicate which words the target word can be combined with.

[0079] Specifically, the disclosed embodiment can obtain the first associated word and / or the second associated word associated with each target word, the first associated word can be the word before the target word, and the second associated word can be the word after the target word. On this basis, the part of speech of the first associated word of each target word is compared with the part of speech of the first associated word of another target word, and / or the part of speech of the second associated word of each target word is compared with the part of speech of the second associated word of another target word, and the parts with the same part of speech are used as grammatical features.

[0080] In this process, the disclosed embodiment can compare the same word in a certain language with the synonyms in other languages ​​one by one, and use the same parts in the comparison results as the grammatical characteristics of the language. In other words, multiple target words in the target vocabulary group are compared one by one, and the grammatical characteristics can be obtained based on the comparison results.

[0081] In addition, the embodiment of the present disclosure can obtain the associated words of each target vocabulary, that is, obtain the first associated words and / or the second associated words of the target vocabulary, wherein the first associated words may be words placed before the target vocabulary, and the second associated words may be words placed after the target vocabulary. On this basis, the grammatical features are obtained by combining the first associated words and / or the second associated words.

[0082] It should be noted that the target vocabulary may have only the first associated word, or only the second associated word, or may be connected with both the first associated word and the second associated word.

[0083] In addition, in the process of acquiring grammatical features, the embodiments of the present disclosure may place different target words in the same context, and then determine whether the sentence is valid. If so, the same parts in the comparison results may be used as grammatical features.

[0084] As an example, Figure 5 The target words shown are "happy" and "happy", which are placed in the first context, which is "someone + copula + ___". In the first context, the target word has a first associated word. The target words "happy" and "happy" are placed in the first context respectively, and the sentences obtained are "He is happy" and "He is happy". Figure 5As shown in the figure, the sentence "He is happy" is valid, while the sentence "He is happy" is not valid. The main reason is that happy, as a sentence-ending component, can be preceded by a copula, while happy, as a sentence-ending component, cannot be preceded by a copula. Therefore, the first associated words of the two target words are different.

[0085] As another example, Figure 6 The target words shown are "happy" and "happy", and the two are placed in the second context, which is "___+noun". The target words in the second context have a second associated word. The target words "happy" and "happy" are placed in the second context respectively, and the sentences obtained are "happy boy" and "happy boy". Figure 6 As shown in the figure, the sentence "happy boy" is valid, and the sentence "happy boy" is also valid. Since both "happy" and "happy" can be followed by nouns, the second associated words of these two target words have the same part of speech. It can be seen that in different contexts, the grammatical features of "happy" and "happy" may be different.

[0086] Optionally, the disclosed embodiment may also perform commonality analysis on typical features of the same part of speech in the target language, and use the analysis results as typical features of the part of speech in the language. Here, commonality analysis can be used to place words of the same part of speech in the target language in the same context for comparative analysis, and check the combination rules of these words to obtain analysis results, which can be used as grammatical features corresponding to the language.

[0087] The analysis result may be which types of words the target word can be combined with, which types of words it cannot be combined with, and which combinations have special conditions. Based on the commonality analysis, the disclosed embodiment can verify the generation process of the mixed corpus and the feasibility of synonym replacement.

[0088] Continuing with the above example, "happy" as the final component of a sentence cannot be preceded by a copula, but when it is used to express contrast or emphasis, it can be expressed as: he is happy, others are suffering.

[0089] It should be noted that the typical features of a language may also include pronunciation and vocabulary features, which may be features related to phonetics and vocabulary. Extracting the pronunciation and vocabulary features of each language may include: extracting the correspondence between the vocabulary and phonemes corresponding to each language, and using the correspondence to obtain the pronunciation and vocabulary features corresponding to each language, wherein the pronunciation and vocabulary features may include the proportional relationship between the number of vocabulary and the number of syllables.

[0090] Optionally, the language specificity extraction model may also include a speech and vocabulary extraction model, the input data of which may be the correspondence between all words and phonemes / pinyin in each language, and the output data may be the correspondence between all words and the number of phonemes / syllables in each language, where the number may be an average, a mode, or a weighted average. Exemplarily, the speech and vocabulary extraction model may be an N-gram model (N-gram Model), a Seq2Seq model (Sequence to Sequence Model), or an LSTM model, etc., through which the correspondence between words and phonemes in each language may be extracted.

[0091] In other words, the embodiments of the present disclosure can extract the correspondence between the number of words and the number of syllables in each language by using a speech and vocabulary extraction model, and the correspondence can be used as the pronunciation vocabulary feature corresponding to each language.

[0092] It should also be noted that the correspondence between words and phonemes can be pre-stored, that is, the input data can be obtained by searching for the corresponding language in the sample corpus. For example, if the language of the sample corpus is English, the embodiment of the present disclosure can obtain the correspondence between all English words and phonemes by searching, and use it as the input of the specificity extraction model combining speech and words. Optionally, the correspondence between the number of words and the number of syllables in each language can also be obtained by extracting the sample corpus.

[0093] In the disclosed embodiment, the correspondence between the number of words and the number of syllables may be a typical feature of the voice and vocabulary combination dimension, that is, by extracting features of the voice and vocabulary combination dimension for each language, the pronunciation vocabulary features corresponding to each language may be obtained.

[0094] It can be seen from the above technical solution that the combined dimensions of voice and vocabulary can be used to express the same concept. For example, the ratio between the number of words and the number of syllables is the same as expressing the concept of "water". Chinese uses one syllable [shui], English uses two syllables water[[ˈwɔːtər]], and German uses two syllables Wasser[[[ˈwɔsər]. For example, the ratio between the number of words and the number of syllables corresponding to the word "mobile phone" is "2:2". For another example, the ratio between the number of words and the number of syllables corresponding to the word "water" is "1:2".

[0095] By extracting typical features related to phonemes (speech) and vocabulary and using the pronunciation and vocabulary features to generate the fourth corpus, the matching of speech and vocabulary of the obtained mixed corpus can be made more accurate, that is, it can ensure that the finally generated mixed corpus is more accurate in speech and vocabulary expression and meets the actual needs of users.

[0096] Step S220: For any first language, determine the first vocabulary corresponding to the first language according to the second pronunciation feature in the first pronunciation feature.

[0097] As an optional way, after extracting at least one of the first pronunciation feature, vocabulary feature, and grammar feature of the first corpus corresponding to each language, the embodiments of the present disclosure can, for any first language, determine the first vocabulary corresponding to the first language according to the second pronunciation feature in the first pronunciation feature, where the second pronunciation feature can belong to the first language and not belong to the second language. Exemplarily, the second pronunciation feature can be a list of unique phonemes.

[0098] Specifically, for the first language, the embodiments of the present disclosure can obtain its corresponding first pronunciation feature, which can include a second pronunciation feature and a third pronunciation feature. Among them, the second pronunciation feature can belong to the first language and not belong to the second language, and the third pronunciation feature can belong to both the first language and the second language. For example, when the first language is Chinese, its corresponding second pronunciation feature is tone, which is a pronunciation feature that English does not have. At the same time, the corresponding third pronunciation feature of Chinese can be plosive, which is a pronunciation feature that both Chinese and English have.

[0099] Step S230: Determine the second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature.

[0100] In some embodiments, the fewer the number of characters in the vocabulary feature, the more conducive it is to expression. To improve the efficiency of corpus generation, the embodiments of the present disclosure can determine the second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature. Among them, the number of characters can be less than or equal to a threshold.

[0101] As an example, the number of characters of the Chinese word "苹果" is 2, and the corresponding English includes the words "apple", "redapple", and "green apple". Among them, the number of characters of "apple" is the least. At this time, "apple" can be used as the second vocabulary corresponding to "苹果".

[0102] Step S240: Determine the third vocabulary corresponding to the first language according to the part-of-speech feature and semantic feature in the grammar feature.

[0103] As introduced above, each language has its own unique grammar feature, where the grammar feature can include a part-of-speech feature and a semantic feature. The embodiments of the present disclosure can determine the third vocabulary corresponding to the first language according to the part-of-speech feature and semantic feature in the grammar feature. That is to say, the same vocabulary can correspond to multiple different language words, and the part-of-speech feature and semantic feature can be used to select the word that best matches the target word in terms of part-of-speech feature and semantic feature from multiple words as the third vocabulary.

[0104] As an example, the English expressions corresponding to the Chinese word "happy" include "happy", "Happily", "Merry", and "Enjoy", etc. Among them, the parts of speech of "Enjoy" and "Happily" are different from that of "happy", so they can be directly excluded. Since the meaning of "happy" is closest to that of "happy", "happy" can be used as the third word at this time.

[0105] Step S250: If the second corpus includes the first word, select a fourth word from the third corpus for translation to generate a fourth corpus.

[0106] As an alternative approach, after obtaining the second corpus, if the present disclosure embodiment detects that the second corpus includes the first word, a fourth word can be selected from the third corpus for translation to generate a fourth corpus. The third corpus can be a part of the second corpus and does not include the first word, and the generated fourth corpus can include the second word and the third word.

[0107] From the above introduction, it is known that the first word can be a word determined according to the second pronunciation feature in the first pronunciation features, and the second pronunciation feature belongs to the first language and does not belong to the second language. That is to say, the pronunciation feature of the first word is unique to the first language. For example, compared with English, "light tone" is a pronunciation feature unique to Chinese. For example, "de" can be used as the first word, and this first word does not need to be translated during the process of generating the fourth corpus.

[0108] In the present disclosure embodiment, the second corpus can be a monolingual corpus. At this time, the second corpus can only include one language. After obtaining this monolingual corpus, the present disclosure embodiment can determine whether the monolingual corpus includes the first word based on using a mixed corpus generation model. When it is determined that it includes the first word, a fourth word is selected from the third corpus for translation to generate a fourth corpus. Here, the fourth corpus can be a mixed corpus.

[0109] That is to say, the present disclosure embodiment can use a mixed corpus generation model to generate a mixed corpus corresponding to the monolingual corpus, as detailed in Figure 7 shown. By Figure 7 knowing that the second corpus can be the monolingual corpus 701, the first mixed corpus 702 can be obtained based on the monolingual corpus 701. At this time, the number of languages in the first mixed corpus 702 is more than that in the monolingual corpus 701, and the first mixed corpus 702 can include the language in the monolingual corpus 701. For example, if the language of the monolingual corpus is English, the languages of the generated first mixed corpus 702 can include English and Chinese. It can be seen that the first mixed corpus 702 can include the language English in the second corpus.

[0110] Optionally, the second corpus may also be a mixed corpus. In this case, the second corpus may include at least two languages, such as two languages. After obtaining the second corpus, the embodiment of the present disclosure may use the mixed corpus generation model to generate a second mixed corpus corresponding to the second corpus. The number of languages ​​in the second mixed corpus may be greater than the number of languages ​​in the second corpus. That is, the embodiment of the present disclosure may increase the number of languages ​​in the mixed corpus by using the mixed corpus generation model. For details, Figure 8 As shown. Figure 8 It is known that the second corpus may be a multilingual corpus 801 , and based on the multilingual corpus 801 , a second mixed corpus 802 (fourth corpus) may be obtained.

[0111] At this time, the number of languages ​​in the second mixed corpus 802 may be greater than the number of languages ​​in the multilingual corpus 801, and the second mixed corpus 802 may include the languages ​​in the multilingual corpus 801, such as Figure 8 The languages ​​of the multilingual corpus include "English" and "Chinese", and the languages ​​of the generated second mixed corpus 802 include English "I", French "bin" and Chinese "Locke", that is, the second mixed corpus 802 can include the languages ​​of the multilingual corpus, English and Chinese.

[0112] Optionally, the second corpus may also be a mixed corpus. In this case, the second corpus may include at least two languages, such as three languages. After acquiring the second corpus, the embodiment of the present disclosure may generate a third mixed corpus corresponding to the second corpus using a mixed corpus generation model. The number of languages ​​in the third mixed corpus (fourth corpus) may be less than the number of languages ​​in the second corpus, that is, the embodiment of the present disclosure may reduce the number of languages ​​in the mixed corpus using a mixed corpus generation model. The embodiment of the present disclosure may generate a mixed corpus containing multiple languages ​​according to user needs, and the number of languages ​​of the generated mixed corpus may be greater than the number of languages ​​of the input corpus, or may be less than the number of languages ​​of the input corpus. For example, the input corpus is a corpus containing four languages, and the output may be a corpus containing three languages, which can reduce the number of languages.

[0113] Exemplarily, the mixed corpus generation model can be an encoder-decoder based TTS framework model, such as Tacotron, DeepVoice, etc. These models can be used for training of monolingual data or multilingual data to synthesize multiple languages. For example, Chinese corpus is used to synthesize Chinese-English mixed corpus. For another example, Chinese-Korean corpus is used to synthesize Chinese-English mixed corpus.

[0114] See also Fig. 9, the second corpus may be a multilingual corpus 901, and a third mixed corpus 902 may be obtained based on the multilingual corpus 901. At this time, the number of languages ​​in the third mixed corpus 902 may be less than the number of languages ​​in the multilingual corpus 901, and the third mixed corpus 902 may include some of the languages ​​in the multilingual corpus 901, such as Fig. 9 The languages ​​of the multilingual corpus include English, Chinese and French, and the languages ​​of the generated third mixed corpus 902 include English and Chinese, that is, the third mixed corpus 902 may include the languages ​​of the multilingual corpus, English and Chinese. The first mixed corpus 702 , the second mixed corpus 802 , and the third mixed corpus 902 may be collectively referred to as a fourth corpus.

[0115] In summary, the second corpus in the embodiments of the present disclosure may be a monolingual corpus, that is, it may include only one language; the second corpus may also be a multilingual corpus, and the number of languages ​​of the multilingual corpus may include at least two languages.

[0116] In the embodiment of the present disclosure, the second corpus is mainly used to generate the fourth corpus, which can be a mixed corpus, wherein the second corpus can be obtained by at least one of the following methods: based on a public language data set; based on a web crawler technology; based on a text generator; crawling the second corpus from a social media platform or other online resources; manually generated according to demand. The specific sampling method for generating the second corpus is not explicitly limited here, and can be selected according to actual conditions.

[0117] As an optional method, before converting the second corpus into the fourth corpus, the embodiment of the present disclosure may first determine multiple languages ​​for generating a mixed corpus. Here, the multiple languages ​​for generating a mixed corpus may be generated according to a preset rule. For example, if the preset rule is to generate a mixed corpus including English and Chinese, the corresponding multiple languages ​​are English and Chinese.

[0118] Optionally, the multiple languages ​​used to generate the fourth corpus can also be determined according to the user's personal needs. If the personal needs are different, the corresponding multiple languages ​​used to generate the mixed corpus are also different. For example, the user in A wants to generate the fourth corpus including English and French, and the user in B wants to generate the fourth corpus including German and French. It can be seen that the multiple languages ​​used to generate the fourth corpus here can be input by the user according to personal needs.

[0119] Accordingly, in response to the received second corpus, the embodiment of the present disclosure can identify the language included in the second corpus, and on this basis, output the language and prompt the user to input multiple languages ​​for generating the fourth corpus. By outputting the language in the second corpus, the user can be assisted in inputting the language.

[0120] Optionally, the multiple languages ​​used to generate the fourth corpus can also be determined according to a preset language generation priority. For example, the language generation priority includes English+Chinese, English+German, and English+Chinese+German, etc. Based on the language generation priority, the embodiment of the present disclosure can sequentially determine multiple languages ​​for generating the fourth corpus.

[0121] As an optional method, after obtaining the text fusion template, the embodiment of the present disclosure can transform the second corpus into the fourth corpus according to the features corresponding to each language and the text fusion template. Here, the features of each language can be obtained by extracting the particularity of each language from multiple dimensions.

[0122] Through the above introduction, it is known that before obtaining the second corpus, the embodiment of the present disclosure can pre-acquire features corresponding to each language, and these features can be typical features unique to each language. Therefore, after determining multiple languages ​​for generating the fourth corpus, the embodiment of the present disclosure can obtain the determined typical features corresponding to each language by searching. On this basis, the first vocabulary and the fourth vocabulary can be obtained according to the typical features of these languages, and the fourth vocabulary is translated to obtain the fourth corpus, wherein the fourth corpus can include the second vocabulary and the third vocabulary, and the second vocabulary and the third vocabulary can be obtained by analyzing the lexical characteristics and grammatical characteristics.

[0123] It should be noted that, if the type of the second corpus in the embodiment of the present disclosure is different, the corresponding method of generating the mixed corpus may also be different. Specifically, when the type of the second corpus is a monolingual corpus, the embodiment of the present disclosure can obtain the fourth corpus by increasing the number of languages, as shown in detail. Figure 7 shown.

[0124] Optionally, when the type of the second corpus is a multilingual corpus, the embodiment of the present disclosure may obtain the fourth corpus by increasing the number of languages, as shown in detail. Figure 8 In addition, the embodiment of the present disclosure can also obtain the fourth corpus by reducing the number of languages, as shown in detail. Fig. 9 As shown, at this time, the number of languages ​​of the second corpus can be greater than or equal to three.

[0125] The language in the embodiments of the present disclosure may be referred to as a language, a monolingual corpus may be referred to as a monolingual corpus, a multilingual corpus may be referred to as a multilingual corpus, and the like.

[0126] In addition, after acquiring the second corpus, the embodiment of the present disclosure may input the second corpus into the mixed corpus generation model to obtain the fourth corpus.

[0127] It should be noted that after generating the fourth corpus, the embodiments of the present disclosure can perform cross-lingual transfer learning and voice tone transfer based on the fourth corpus to obtain a richer mixed corpus. Additionally, text-to-speech operations can also be performed based on the fourth corpus to obtain mixed speech.

[0128] In some embodiments, the first pronunciation feature may include phoneme features and / or tone features. Determining the first vocabulary corresponding to the first language according to the second pronunciation feature in the first pronunciation feature of the first language includes the following steps: Determine the phoneme features and / or tone features that belong to the first language and do not belong to the second language as the second pronunciation feature; Determine the vocabulary in the first corpus with the second pronunciation feature as the first vocabulary.

[0129] Specifically, after extracting at least one of the first pronunciation feature, vocabulary feature, and grammar feature of the first corpus corresponding to each language, the embodiments of the present disclosure can traverse each language based on these features to obtain the first vocabulary corresponding to each language. During the traversal process, the language to be traversed can be used as the first language. On this basis, determine the first vocabulary corresponding to the first language according to the second pronunciation feature in the first pronunciation feature corresponding to the first language.

[0130] Specifically, the embodiments of the present disclosure can obtain the phoneme features and / or tone features that belong to the first language and do not belong to the second language as the second pronunciation feature. On this basis, determine the vocabulary in the first corpus with the second pronunciation feature as the first vocabulary.

[0131] Exemplarily, when the first language is Chinese and the second language is English, the embodiments of the present disclosure can obtain the phoneme features and / or tone features that belong to Chinese and do not belong to English as the second pronunciation feature. Then, determine the vocabulary in the first corpus with the second pronunciation feature as the first vocabulary. For example, the retroflex consonants "zh", "ch", "sh", and "r" in Chinese are unique phonemes in Chinese and there are no exactly corresponding phonemes in English. Also, Chinese has four tones, namely the high level, rising tone, falling-rising tone, and falling tone, and such a tone system does not exist in English. These phoneme features and / or tone features can all be used as the second pronunciation feature, and the vocabulary with this second pronunciation feature can be used as the first vocabulary. For example, the Chinese word "山水" can be used as the first vocabulary.

[0132] In other words, if the corpus corresponding to the second language includes the first vocabulary, the embodiments of the present disclosure can select vocabulary from the corpus for translation to generate the fourth corpus, and the selected vocabulary may not include the first vocabulary.

[0133] The disclosed embodiment determines the first vocabulary corresponding to each language by utilizing phoneme features and / or tone features, thereby ensuring translation accuracy while preventing the unique vocabulary in each language from being mistranslated and ensuring that it is not translated. This not only avoids translation errors but also enables the finally generated fourth corpus to retain the unique style of the first language.

[0134] In some implementations, determining a second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature of the first language includes the following steps: Determining a translation of a fourth word in a second language into a plurality of fifth words in a first language; extracting lexical features of each fifth word; A fifth word in the plurality of word features whose number of characters meets a preset threshold is used as the second word.

[0135] Specifically, in the process of obtaining the second vocabulary corresponding to the first language, the embodiment of the present disclosure can determine multiple translation vocabulary of the corresponding vocabulary in the first language for any vocabulary in the second language. On this basis, the translation vocabulary with the least number of characters in the multiple translation vocabulary is used as the second vocabulary, so that when the fourth vocabulary is selected from the third corpus, the second vocabulary is used to translate the fourth vocabulary.

[0136] Through the above introduction, it is known that the embodiment of the present disclosure can determine the second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature, wherein the number of characters can be less than or equal to a threshold. In the process of obtaining the second vocabulary, the embodiment of the present disclosure can first obtain the fourth vocabulary of the second language and translate it into multiple fifth vocabulary of the first language.

[0137] On this basis, the vocabulary characteristics of each fifth vocabulary are extracted, and the number of characters in the vocabulary characteristics is used to determine whether the fifth vocabulary meets the preset condition. If so, the fifth vocabulary can be used as the second vocabulary. The preset condition can be that the number of characters in the vocabulary characteristics meets a preset threshold, such as the number of characters is less than or equal to the preset threshold.

[0138] As an example, the first language is English and the second language is Chinese. After the fourth word "happy" in Chinese is translated into English, the corresponding words are "happy", "joyful", "Merry" and "Happily", etc. These words can be used as the fifth word. The features of the words "happy", "joyful", "Merry" and "Happily" are extracted, and it is determined that the number of characters of "happy" meets the preset conditions, so "happy" and "Merry" can be used as the second words.

[0139] The disclosed embodiment screens multiple fifth vocabularies corresponding to the first language and uses vocabularies that meet a preset threshold as second vocabularies, thereby ensuring that the number of characters in the generated fourth corpus is as small as possible, which not only ensures the readability of the fourth corpus but also facilitates writing.

[0140] In some implementations, determining a third vocabulary corresponding to the first language according to a part-of-speech feature and a semantic feature in a grammatical feature of the first language includes the following steps: A translation word of a fourth word in the second language is determined from the words in the first language according to the part-of-speech features and semantic features in the grammatical features of the first language and the part-of-speech features and semantic features in the grammatical features of the second language, and the translation word is used as the third word.

[0141] From the above introduction, we know that each language has corresponding first pronunciation features, vocabulary features and grammatical features, wherein the grammatical features may include part-of-speech features and semantic features. In order to ensure the accuracy of the third vocabulary acquisition, the embodiment of the present disclosure may acquire the part-of-speech features and semantic features in the grammatical features of the first language, and acquire the part-of-speech features and semantic features in the grammatical features of the second language.

[0142] On this basis, according to the corresponding lexical features and semantic features of the first language and the part-of-speech features and semantic features of the second language, a translation word of the fourth word in the second language is determined from the words of the first language, and the translation word is used as the third word.

[0143] As an example, the first language is English and the second language is Chinese. Regarding the word "happy", it corresponds to "happy" and "Merry" in English. By comparing the part-of-speech features and semantic features of "happy" and "Merry" with the part-of-speech features and semantic features of "happy", it is determined that "happy" is closer, so the translation word "happy" of "happy" in Chinese can be used as the third word.

[0144] By combining the part-of-speech features and semantic features in the grammatical features, the third vocabulary in the or region can be made more accurate, thereby ensuring the accuracy of the fourth corpus generation.

[0145] In some embodiments, before determining a translation vocabulary of a fourth vocabulary in the second language from the vocabulary in the first language based on the part-of-speech feature and the semantic feature in the grammatical feature of the first language and the part-of-speech feature and the semantic feature in the grammatical feature of the second language, the following steps are included: Extract word vectors of vocabulary in a first corpus of a first language and a second corpus of a second language; The distance between the word vectors of the vocabulary is calculated, and the vocabulary whose distance meets the preset distance threshold is determined as a vocabulary set, and the semantics of each word in the vocabulary set are similar.

[0146] In the embodiment of the present disclosure, the vocabulary set can be a similar semantic vocabulary mapping, that is, the vocabulary set can be words with similar semantics between different languages, such as a mapping of translation words, such as "happy" and "happy" can be words with similar semantics between different languages. In the process of constructing the vocabulary set, the embodiment of the present disclosure can extract the word vectors of the vocabulary in the first corpus of the first language, and extract the word vectors of the vocabulary in the second corpus of the second language. On this basis, the distance between the word vectors of the vocabulary is calculated, and it is determined whether the distance meets the preset distance threshold. If the preset distance threshold is met, the corresponding vocabulary can be determined as the vocabulary set.

[0147] Exemplarily, words whose word vector distance is less than a preset threshold are regarded as similar semantic words, and these similar semantic words can form a vocabulary set. Exemplarily, the preset threshold can be determined based on the distance range of a large number of similar words, such as the average distance of multiple similar words can be used as the preset threshold.

[0148] In summary, the vocabulary set can be determined by calculating the distance between the word vectors of the words. That is, based on the word vector, the vector distance between each word and other words can be calculated, and then the vector distance can be used to obtain similar semantic words, and the vocabulary set can be composed of similar semantic words.

[0149] Exemplarily, the vocabulary set may include a mapping relationship for the word "happy", and similar semantic words corresponding to "happy" include: English "happy", French "heureux", German "glücklich", Spanish "feliz" and Italian "felice", etc.

[0150] Optionally, the embodiment of the present disclosure can use a preset number of dimensions as a sliding window to divide the dimension of the word vector of each word in the vocabulary set into multiple dimension segments. On this basis, the distance of the word vector in each dimension segment is calculated, and the similarity of each word in the vocabulary set is determined according to the distance of the word vector in each dimension segment and the weight of each dimension segment. Here, the similarity can be the semantic difference of each word in the vocabulary set.

[0151] As an optional method, after obtaining the vocabulary set, the embodiment of the present disclosure can further screen the semantically similar words in the vocabulary set. Specifically, the dimension of the word vector of each word in the vocabulary set is divided into multiple dimension segments with a preset number of dimensions as the sliding window. On this basis, the word vector distance within each dimension segment is calculated, and the similarity of each word in the vocabulary set is determined based on the distance of the word vector within each dimension segment and the weight of each dimension segment. The smaller the similarity of each word, the more similar the semantics of the two words are.

[0152] In other words, the embodiments of the present disclosure can use the specified value as a sliding window to obtain the vector distance of similar words on the specified dimension, and perform weighted summation of the vector distances in different dimensional segments to obtain the final vector distance. The vector distance in each dimensional segment can be used as the semantic difference in the corresponding segment.

[0153] Exemplarily, the vocabulary set includes three vocabulary groups, wherein the first vocabulary group includes three first words with the same semantics, and the three first words are vocabulary 1, vocabulary 2 and vocabulary 3. The three first words are respectively converted into 128-dimensional word vectors, and the 128-dimensional word vectors can be divided into 13 segments, the first 12 segments can include 10-dimensional data, and the last segment can include 8-dimensional data. On this basis, the word vector distances between the three words are calculated according to the segments.

[0154] For example, the distance between the first segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 1; the distance between the second segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 2; the distance between the third segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 3; the distance between the fourth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 4; the distance between the fifth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 5; the distance between the sixth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 6; the distance between the seventh segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain distance 7; The eighth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain a distance of 8; the ninth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain a distance of 9; the tenth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain a distance of 10; the eleventh segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain a distance of 11; the twelfth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain a distance of 12; the thirteenth segment vector of vocabulary 1 and vocabulary 2 is calculated to obtain a distance of 13, and then these distances are weighted and summed to obtain the similarity between vocabulary 1 and vocabulary 2. The similarity between vocabulary 1 and vocabulary 3, and the similarity between vocabulary 2 and vocabulary 3 are obtained in a similar manner, so they will not be repeated here.

[0155] The embodiment of the present disclosure can more accurately determine the translation vocabulary by screening a plurality of semantically similar vocabulary in the vocabulary set, so that the fourth corpus finally obtained can be more accurate.

[0156] It should be noted that before obtaining the word vectors of the extracted corpus, the embodiment of the present disclosure can first segment the corpus to obtain the word segmentation results. On this basis, the word vector of each word is extracted based on the word segmentation results, and the distance between the words is obtained based on the word vectors, thereby obtaining a vocabulary set.

[0157] In some embodiments, the semantic feature includes similarity of each word in the vocabulary set, and determining a translation word of a fourth word in the second language from the words in the first language according to the part-of-speech feature and the semantic feature in the grammatical feature of the first language and the part-of-speech feature and the semantic feature in the grammatical feature of the second language includes the following steps: When the part-of-speech feature in the grammatical feature of the first language and the part-of-speech feature in the grammatical feature of the second language are consistent, a translation word for the fourth word in the second language is determined from the words in the first language according to the similarity of each word in the vocabulary set.

[0158] Continuing with the above example, the first language is English, the second language is Chinese, the part of speech of the Chinese word "happy" is adjective, and the part of speech of the English words "happy" and "Merry" is also adjective. Through calculation, it is found that the first distance between "happy" and "happy" is smaller than the second distance between "happy" and "Merry", that is, the similarity between "happy" and "happy" is smaller than that between "Merry" and "happy", so "happy" in English can be used as the translation word of "happy" in Chinese, that is, "happy" can be used as the third word.

[0159] To sum up, in the process of determining the translation vocabulary of each fourth vocabulary, the embodiment of the present disclosure can first ensure that the part-of-speech features of the first language are consistent with the part-of-speech features of the second language, and at the same time obtain the similarity of the corresponding vocabulary of the two languages. When the similarity is less than a preset threshold, the vocabulary can be used as the translation vocabulary of the fourth vocabulary.

[0160] In the embodiment of the present disclosure, the similarity of each word in the vocabulary set can be obtained by weighted averaging the vector distances between the words. In addition, the parts of speech of each word in the vocabulary set can be the same.

[0161] As an example, there are 5 first words with the same part of speech in the vocabulary set. By calculation, the vector distances between the 5 first words can be obtained respectively, and then the vector distances are weighted and averaged to obtain the similarity between the first words. On this basis, it is determined whether these similarities are less than the preset threshold. If they are, the first words are retained as target words, otherwise, the first words are deleted. For example, if the similarity between three of the 5 first words with the same part of speech is less than the preset threshold, these three words are retained, and the other two words with similarities greater than the preset threshold are deleted.

[0162] Exemplarily, the embodiments of the present disclosure may input multiple words for comparison into a POS module (part of speech) to obtain the parts of speech of these words. Here, the POS module may be used to obtain the parts of speech of each input word. For example, the words input into the POS module are word 1, word 2, and word 4, and the output is word 1-noun; word 2-noun; word 4-adjective; word 5-noun. By comparison, it is found that the part of speech of word 4 is different from that of word 1, word 2, and word 5. At this time, the embodiments of the present disclosure may delete word 4 and retain word 1, word 2, and word 5. Here, the languages ​​corresponding to word 1, word 2, and word 5 are different.

[0163] In addition, in the process of obtaining the semantic difference between each first word, the embodiment of the present disclosure can obtain the vector distances in different dimensional segments between the first words. On this basis, the embodiment of the present disclosure can perform cluster analysis on these vector distances to lock the word pairs that need to be compared through the distance analysis method, that is, quantify the distance of the word pairs by using the distance calculation method.

[0164] Through the above introduction, it is known that after obtaining the vocabulary set, the embodiment of the present disclosure can place multiple words in the vocabulary set in the same context respectively to obtain the naturalness corresponding to each word. On this basis, the vocabulary set is screened according to the naturalness to obtain a target vocabulary set. Specifically, the embodiment of the present disclosure can use words whose naturalness exceeds a preset threshold as target words, and the target vocabulary set is composed of the target words. Among them, naturalness is used to evaluate whether the semantics is reasonable and whether the grammar is compliant.

[0165] In this process, the embodiment of the present disclosure can first determine the reference document (context) corresponding to the same semantic vocabulary, and on this basis, input multiple identical semantic vocabulary into the reference document respectively to obtain multiple candidate sentences, and then obtain the naturalness of each candidate sentence, and use the sentence whose naturalness exceeds a preset threshold as the target sentence, and the vocabulary in the target sentence can be used as the target vocabulary.

[0166] Here, context refers to the same context, which has no limit on length and can be the same sentence or the same paragraph. The original monolingual text corpus (sample corpus) can be used as context to obtain source data (reference documents). Specifically, the monolingual text containing a certain word is aggregated to form multiple sentences, and then the target word is mined, and the remaining part can be used as context.

[0167] It should be noted that the naturalness can be generated by a naturalness model, that is, the same semantic vocabulary in different languages is put into the same reference document (context), and the reference document is input into the naturalness model respectively to generate the naturalness of each reference document. Exemplarily, the naturalness model can be an N-gram model, a language model based on a recurrent neural network, or a Transformer model, etc. The input data of these models can be the reference document, and the output data can be the naturalness corresponding to the reference document.

[0168] If it is determined that the naturalness is higher than the preset threshold, the corresponding vocabulary can be used as the target vocabulary. On the contrary, if the naturalness is lower than the preset threshold, the vocabulary to be detected cannot be called semantically similar vocabulary, so it can be deleted.

[0169] As another optional way, if it is determined that the first vocabulary does not exist in the second corpus, the embodiment of the present disclosure can select the seventh vocabulary from the fifth corpus for translation to generate the sixth corpus, where the translated vocabulary of the seventh vocabulary can include the first vocabulary.

[0170] In other words, when it is determined that the second corpus does not include the first vocabulary, the embodiment of the present disclosure can select the seventh vocabulary from the fifth corpus and translate the seventh vocabulary to generate the sixth corpus, and the translated vocabulary of the seventh vocabulary does not include the first vocabulary.

[0171] Exemplarily, the language of the second corpus is English. The seventh vocabulary is selected from the fifth corpus in the second corpus, and the seventh vocabulary is translated into Chinese. The translated Chinese vocabulary cannot be the first vocabulary. For example, the translated vocabulary of the seventh vocabulary cannot include the word "landscape".

[0172] In some embodiments, selecting the fourth vocabulary from the third corpus for translation to generate the fourth corpus includes the following steps: Select the fourth vocabulary from the third corpus and translate it according to the preset language.

[0173] As an example, the preset language is Chinese and the second corpus is English. In the process of selecting the fourth vocabulary from the third corpus for translation to generate the fourth corpus, the embodiment of the present disclosure can translate the fourth vocabulary into Chinese. For example, the second corpus is "I am Rock", and the determined fourth vocabulary is "am" and "Rock". At this time, "am" and "Rock" can be translated, and finally "I is Rock" is obtained.

[0174] It can be seen that the embodiment of the present disclosure can select the fourth vocabulary from the third corpus for translation according to the preset first expected information (preset language) for indicating the language to be included in the generated corpus, so as to generate the fourth corpus that meets the first expected information. In addition, the embodiment of the present disclosure can select the fourth vocabulary from the third corpus for translation according to the preset ratio of different languages.

[0175] As an optional manner, after determining multiple languages ​​for generating the fourth corpus, the embodiment of the present disclosure may set a corresponding preset ratio for each language, and select a fourth vocabulary from the third corpus for translation based on the preset ratio.

[0176] Specifically, the embodiments of the present disclosure can set different preset ratios for each language based on the random module of Python, and the preset ratio can be called a preset ratio coefficient. For example, it is determined that the multiple languages ​​used to generate the fourth corpus include "English" and "Chinese", and different preset ratio coefficients can be set for "English" and "Chinese" using the random module of Python, and the preset ratio obtained is English: Chinese = 1:2.

[0177] It should be noted that the preset ratios corresponding to different languages ​​may be the same, such as English: Chinese = 1:1; the preset ratios corresponding to different languages ​​may also be different, such as English: Chinese = 1:2.

[0178] Here, the preset ratio can be used to control the number of words. Optionally, the embodiment of the present disclosure can limit the preset ratio according to the syllable length. Here, the preset ratio of languages ​​can be the length ratio of language A and language B in a text.

[0179] Exemplarily, the second corpus is "I am Rock", and the multiple languages ​​determined for generating the fourth corpus are English and Chinese, respectively. The phon-based random module generates a preset ratio of English and Chinese, that is, English: Chinese = 1:2.

[0180] The embodiment of the present disclosure sets a corresponding preset ratio for each language, and can flexibly and effectively acquire a plurality of different fourth corpora based on the preset ratio, thereby improving the efficiency of acquiring mixed corpora.

[0181] Optionally, according to the preset proportion of different languages, in the process of selecting the fourth vocabulary from the third corpus for translation, the embodiment of the present disclosure can select the fourth vocabulary from different positions of the third corpus for translation. The positions are different, and the corresponding selected fourth vocabulary may also be different, thereby ensuring the flexible generation of the fourth corpus.

[0182] As an optional method, the preset language may be the language used during the conversation between the first user and the agent. For example, the language used during the conversation between the first user and the agent is English, and the language corresponding to the second corpus is Chinese. At this time, the fourth vocabulary may be selected from the third corpus and translated according to English, that is, the fourth vocabulary is translated into English. In other words, the language of the second vocabulary and the third vocabulary corresponding to the generated fourth corpus is English (preset language).

[0183] For example, the second corpus is "I am Locke", the determined preset language is English, and the fourth word is "I am", and the fourth word is translated, and the fourth corpus obtained is "I am Locke".

[0184] In summary, after acquiring the second corpus, the embodiment of the present disclosure can first determine multiple languages ​​for generating a mixed corpus, and set a corresponding preset ratio for each language, according to which a fourth vocabulary can be selected from the third corpus for translation to obtain a fourth corpus.

[0185] From the above introduction, we know that the fourth corpus can be output by a mixed corpus generation model, which can include LSTM (Long Short-Term Memory), which can be a time recursive neural network suitable for processing and predicting important events with relatively long intervals and delays in time series. The bidirectional LSTM network can simultaneously utilize information in both the past and future directions, making the final prediction more accurate.

[0186] The following, combined Fig.10 , a complete example is used to explain the method of generating corpus in detail.

[0187] Step 11: Get the corpus of all languages.

[0188] The corpus of all languages ​​may include monolingual corpus and multilingual corpus. In addition, the corpus may include two formats: voice and text. The voice may be in mp3 format and the text may be in text version. The disclosed embodiment may generate pinyin and phonetic symbols of each text based on the voice and text.

[0189] Step 12: Input the monolingual corpus of all languages ​​into the language specificity extraction model to obtain the typical features of each language.

[0190] Here, the typical features may include a first pronunciation feature, a vocabulary feature, and a grammatical feature. These typical features may be obtained using a language specificity extraction model. The structure of the language specificity extraction model may be as follows: Fig.11 As shown, based on Fig.11It is known that the language specificity extraction model may include a speech extraction model, a vocabulary extraction model and a speech and vocabulary combined extraction model, through which different characteristics of the language can be extracted respectively.

[0191] Step 13: Set different preset ratios, input the preset ratios into the template generation model, and obtain text fusion templates of various ratios and positions.

[0192] Here, the preset ratio can be used to control the number of words, or it can be limited by syllable length. In addition, the above-mentioned various ratios refer to the length ratio of language A and language B in a generated text, and the various positions refer to the beginning, middle and end of a sentence.

[0193] Step 14: Input the single language text into the ascending text fusion template (the number of languages ​​increases) and output the multi-language fusion text.

[0194] Step 15: Input the multilingual text into the ascending text fusion template (the number of languages ​​increases) and output the multilingual fusion text.

[0195] Step 16: Input the multilingual text into the descending text fusion template (the number of languages ​​decreases), and output the multilingual fusion text language.

[0196] Step 17: The multilingual texts may be merged to form a multilingual text corpus, which serves as a text enhancement library for the original single-language text library.

[0197] The embodiments of the present disclosure can be applied to resource recovery scenarios. The specific corpus to be converted can be resource recovery corpus. The specific steps are as follows: Step 21: Obtain resource recovery corpora in all languages.

[0198] Step 22: Input the resource recovery corpus of all languages ​​into the language specificity extraction model to obtain the typical features of each language.

[0199] Step 23: Set a preset scale and input the preset scale coefficient into the template generation model to obtain text fusion templates of various scales and positions.

[0200] Step 24: Input the single language text into the ascending text fusion template (the number of languages ​​increases) and output the multi-language fusion text.

[0201] For example, the monolingual text itself has only one language, and the number of languages ​​of the multilingual text finally formed can be greater than 1, so the process can be successful.

[0202] Step 25: Input the multi-language text into the ascending text fusion template (the number of languages ​​increases), and output the multi-language fusion text.

[0203] For example, the multilingual text itself does not have only one language, so the number of languages ​​of the final multilingual text can be more than the original number of languages, which is called ascent.

[0204] Step 26: Input the multi-language text into the descending text fusion template (the number of languages ​​decreases), and output the multi-language fusion text language; For example, the multilingual text itself does not have only one language, so the number of languages ​​of the final multilingual text can be less than the original number of languages, which is called reduction.

[0205] Step 27: Merge the multilingual texts to form a multilingual text corpus, which serves as a text enhancement library for the original single-language text library.

[0206] For each original single-language text, there can be a corresponding modified multi-language text as its enhanced text.

[0207] Step 28: Classify the enhanced text according to the enhancement type. The classified data forms a classification database, and different databases are matched according to the different preferences of the agents.

[0208] Here, the enhancement type can be language ascending or language descending; after the enhancement in the above steps, the text can be directly stored in the corpus, that is, the text generated by the language ascending method is directly stored in the language ascending enhancement corpus, and the text generated by the language descending method is directly stored in the language descending enhancement corpus.

[0209] For example, the agent's preference is mostly Chinese, or the agent's commonly used vocabulary is obtained, and words in the agent's commonly used vocabulary are used as much as possible.

[0210] Step 29: Use the data matched with different agent preferences as training data and send it to the generative pre-trained language model to train a new resource recovery speech model with the most accurate semantic expression and the least effort.

[0211] Based on the same concept, the embodiment of the present disclosure also provides a corpus generation device, which can be part or all of an electronic device through software, hardware, or a combination of software and hardware. Fig.12 As shown, the corpus generation device 1200 may include: an extraction module 1210 , a first determination module 1220 , a second determination module 1230 , a third determination module 1240 and a translation module 1250 .

[0212] The extraction module 1210 is configured to extract at least one of a first pronunciation feature, a vocabulary feature, and a grammatical feature of a first corpus corresponding to each language; The first determination module 1220 is configured to determine, for any first language, a first vocabulary corresponding to the first language according to a second pronunciation feature in the first pronunciation feature, wherein the second pronunciation feature belongs to the first language and does not belong to the second language; The second determination module 1230 is configured to determine a second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature, the number of characters being less than or equal to a threshold; The third determination module 1240 is configured to determine a third vocabulary corresponding to the first language according to the part-of-speech feature and the semantic feature in the grammatical feature; If the second corpus includes the first vocabulary, the translation module 1250 selects a fourth vocabulary from the third corpus for translation to generate a fourth corpus, where the third corpus is part of the second corpus and does not include the first vocabulary, and the fourth corpus includes the second vocabulary and the third vocabulary.

[0213] In a possible implementation, the first pronunciation feature includes a phoneme feature and / or a tone feature, and the first determination module 1220 is configured to determine the phoneme feature and / or the tone feature that belongs to the first language and does not belong to the second language as the second pronunciation feature; and determine the vocabulary in the first corpus that has the second pronunciation feature as the first vocabulary.

[0214] In a possible implementation, the second determination module 1230 is configured to determine the translation of the fourth word in the second language into multiple fifth words in the first language; extract the lexical features of each of the fifth words; and use the fifth words whose number of characters in the multiple lexical features meets a preset threshold as the second word.

[0215] In a possible implementation, the third determination module 1240 is configured to determine a translation vocabulary of a fourth vocabulary in the second language from the vocabulary of the first language based on the part-of-speech features and semantic features in the grammatical features of the first language and the part-of-speech features and semantic features in the grammatical features of the second language, and use the translation vocabulary as the third vocabulary.

[0216] In a possible implementation, the corpus generating device 1200 may further include: The vocabulary set determination module is configured to extract word vectors of words in the first corpus of the first language and the second corpus of the second language; calculate the distance between the word vectors of the words, and determine the words whose distances meet a preset distance threshold as a vocabulary set, wherein the semantics of each word in the vocabulary set are similar.

[0217] In a possible implementation, the vocabulary set determination module is further configured to divide the dimensions of the word vector of each word in the vocabulary set into multiple dimension segments with a preset number of dimensions as a sliding window; calculate the distance of the word vector within each dimension segment; and determine the similarity of each word in the vocabulary set based on the distance of the word vector within each dimension segment and the weight of each dimension segment, wherein the similarity is the semantic difference between each word in the vocabulary set.

[0218] In a possible implementation, the semantic feature includes similarities of each word in the vocabulary set, and the translation module 1250 is configured to determine a translation word for a fourth word in the second language from the words in the first language according to the similarities of each word in the vocabulary set when a part-of-speech feature in the grammatical feature of the first language is consistent with a part-of-speech feature in the grammatical feature of the second language.

[0219] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0220] Based on the same concept, the embodiment of the present disclosure further provides an electronic device, including: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to perform the steps of any of the above-mentioned corpus generation methods.

[0221] Fig.13 An electronic device (eg, Figure 1 A schematic diagram of the structure of a terminal device or server in the system 1300.

[0222] The electronic device 1300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage device 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the electronic device 1300 are also stored. The processing device 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0223] Typically, the following devices may be connected to the I / O interface 1305: an input device 1306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1309. The communication device 1309 may allow the electronic device 1300 to communicate with other devices wirelessly or by wire to exchange data. Although Fig.13 The electronic device 1300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0224] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a processor. When the computer program is executed by the processor, the steps of the above-mentioned method for generating corpus are implemented.

[0225] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a processor. When the computer program is executed by the processor, the steps of the above-mentioned method for generating corpus are implemented.

[0226] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings; however, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, a variety of simple modifications can be made to the technical solution of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0227] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0228] In addition, various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A method for generating corpus, characterized in that: The method comprises: Extracting at least one of a first pronunciation feature, a lexical feature, and a grammatical feature of a first corpus corresponding to each language; For any first language, determining a first vocabulary corresponding to the first language according to a second pronunciation feature in the first pronunciation feature, wherein the second pronunciation feature belongs to the first language and does not belong to the second language; determining a second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature, the number of characters being less than or equal to a threshold; Determining a third vocabulary corresponding to the first language according to the part-of-speech feature and the semantic feature in the grammatical feature; If the second corpus includes the first vocabulary, a fourth vocabulary is selected from the third corpus for translation to generate a fourth corpus, wherein the third corpus is a part of the second corpus and does not include the first vocabulary, and the fourth corpus includes the second vocabulary and the third vocabulary.

2. The method according to claim 1, characterized in that The first pronunciation feature includes a phoneme feature and / or a tone feature, and determining the first vocabulary corresponding to the first language according to the second pronunciation feature in the first pronunciation feature of the first language includes: determining a phoneme feature and / or a tone feature belonging to the first language and not belonging to the second language as the second pronunciation feature; The vocabulary in the first corpus having the second pronunciation feature is determined as the first vocabulary.

3. The method according to claim 1, characterized in that The step of determining the second vocabulary corresponding to the first language according to the number of characters in the vocabulary feature of the first language includes: determining to translate the fourth word in the second language into a plurality of fifth words in the first language; extracting vocabulary features of each of the fifth words; The fifth vocabulary whose number of characters in the plurality of vocabulary features meets a preset threshold is used as the second vocabulary.

4. The method according to claim 1, characterized in that: The determining of the third vocabulary corresponding to the first language according to the part-of-speech feature and the semantic feature in the grammatical feature of the first language includes: According to the part-of-speech features and semantic features in the grammatical features of the first language and the part-of-speech features and semantic features in the grammatical features of the second language, a translation word of a fourth word in the second language is determined from the words in the first language, and the translation word is used as the third word.

5. The method according to claim 4, characterized in that Before determining a translation vocabulary of a fourth vocabulary in the second language from the vocabulary of the first language based on the part-of-speech features and semantic features in the grammatical features of the first language and the part-of-speech features and semantic features in the grammatical features of the second language, the method includes: Extracting word vectors of vocabulary in a first corpus of the first language and a second corpus of the second language; The distances between the word vectors of the words are calculated, and the words whose distances satisfy a preset distance threshold are determined as a word set, wherein the semantics of the words in the word set are similar.

6. The method according to claim 4, characterized in that After determining the vocabulary set, the method further comprises: Using a preset number of dimensions as a sliding window, the dimension of the word vector of each word in the vocabulary set is divided into a plurality of dimension segments; Calculate the distance between word vectors in each dimension segment; The similarity of each word in the vocabulary set is determined according to the distance of the word vectors in each dimensional segment and the weight of each dimensional segment, and the similarity is the semantic difference of each word in the vocabulary set.

7. The method according to claim 6, characterized in that The semantic features include similarities of the respective words in the vocabulary set, and determining the translation words of the fourth word in the second language from the words in the first language according to the part-of-speech features and the semantic features in the grammatical features of the first language and the part-of-speech features and the semantic features in the grammatical features of the second language, includes: When the part-of-speech features in the grammatical features of the first language are consistent with the part-of-speech features in the grammatical features of the second language, a translation word for a fourth word in the second language is determined from the words in the first language according to the similarity of each word in the vocabulary set.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

9. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.