Bilingual Dictionary Inference Method, Apparatus and Storage Medium
By extracting the target dictionary from parallel corpus and training the bilingual dictionary inference model, the problem of the lack of effect of bilingual dictionary inference on long-distance language pairs is solved, and more efficient word-level alignment information inference is achieved.
Patent Information
- Application Number
- CN202010679242.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-07-15
AI Technical Summary
The existing bilingual dictionary inference methods have poorer results in languages with far distances, especially in language matching between Chinese and English, Chinese and Japanese, and it is difficult to effectively infer word-level alignment information.
By extracting the target dictionary from parallel corpus and combining the preconfigured initial dictionary, the target bilingual dictionary inference model is trained, which is a neural network model with the translation of source-end words into target-end words.
Parallel corpus is introduced on the basis of the initial dictionary, and the extracted target dictionary enriches the training information of the bilingual dictionary inference model, which significantly improves the effect of bilingual dictionary inference, especially in long-distance language pairs.
Smart Images

Figure CN114021551B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a bilingual dictionary inference method, apparatus, and storage medium. Background Art
[0002] The bilingual dictionary inference task can provide word-level alignment information for machine translation under the condition of only a small amount of alignment signals, and is an important part of low-resource or unsupervised machine translation systems.
[0003] The currently commonly used bilingual dictionary inference methods are mainly divided into two steps. The first step is to train monolingual word vectors of two languages on a large-scale monolingual data, and transform each word into a high-dimensional representation. The second step is to train a linear mapping to map the word vectors of the source language into the word vector space of the target language.
[0004] The above-mentioned bilingual dictionary inference methods have better inference effects on language pairs that are relatively close in distance, such as English-French, English-German, etc. However, for language pairs that are relatively far apart, such as Chinese-English, Chinese-Japanese, etc., the bilingual dictionary inference effect is poor. Summary of the Invention
[0005] In view of this, the present disclosure provides a bilingual dictionary inference method, apparatus, and storage medium. The technical solutions include:
[0006] According to one aspect of the present disclosure, a bilingual dictionary inference method is provided, and the method includes:
[0007] Extract a target dictionary from parallel corpora;
[0008] Train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, where the target bilingual dictionary inference model is a neural network model capable of translating source-side words into target-side words;
[0009] Wherein, both the target dictionary and the initial dictionary include a plurality of aligned word pairs, and the aligned word pairs include source-side words and target-side words.
[0010] In a possible implementation manner, the extracting a target dictionary from parallel corpora includes:
[0011] Train an initial bilingual dictionary inference model according to pre-configured monolingual word vectors and the initial dictionary;
[0012] Extract the target dictionary from the parallel corpora according to the initial bilingual dictionary inference model and the word alignment model.
[0013] In another possible implementation, extracting the target dictionary from the parallel corpus according to the initial bilingual dictionary inference model and the word alignment model includes:
[0014] Obtaining a first initialization probability of the word alignment model according to the initial bilingual dictionary inference model;
[0015] Performing word alignment learning on the parallel corpus according to the first initialization probability of the word alignment model to obtain a first word alignment probability;
[0016] Determining the target dictionary according to the first word alignment probability.
[0017] In another possible implementation, obtaining the first initialization probability of the word alignment model according to the initial bilingual dictionary inference model includes:
[0018] According to the initial bilingual dictionary inference model, obtaining the first initialization probability p ini (y|x) of the word alignment model through the following formula:
[0019]
[0020] where x is the source-side word, y is the target-side word, E src (x) is the word vector of the source-side word in the initial bilingual dictionary inference model, E tgt (y) is the word vector of the target-side word in the initial bilingual dictionary inference model, Y(x) represents the translation target of x in the translation table of the word alignment model, τ is used to indicate the sharpness of the initialization distribution, and y′ is any one of the translation targets of x in the translation table of the word alignment model.
[0021] In another possible implementation, determining the target dictionary according to the first word alignment probability includes:
[0022] Obtaining a second initialization probability of the word alignment model according to the initial bilingual dictionary inference model, where the second initialization probability is different from the first initialization probability;
[0023] Performing word alignment learning on the parallel corpus according to the second initialization probability of the word alignment model to obtain a second word alignment probability;
[0024] Performing bidirectional filtering according to the first word alignment probability and the second word alignment probability to obtain the target dictionary.
[0025] In another possible implementation, the method further includes:
[0026] Extract an updated target dictionary from the parallel corpus according to the target bilingual dictionary inference model and the word alignment model;
[0027] Train an updated target bilingual dictionary inference model according to the updated target dictionary and the initial dictionary.
[0028] In another possible implementation, the method further includes:
[0029] Obtain the input source-side word;
[0030] According to the source-side word, call the trained target bilingual dictionary inference model and output the target-side word.
[0031] According to another aspect of the present disclosure, there is provided a bilingual dictionary inference device, the device includes:
[0032] An extraction module, configured to extract a target dictionary from the parallel corpus;
[0033] A training module, configured to train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, and the target bilingual dictionary inference model is a neural network model that translates a source-side word into a target-side word;
[0034] Wherein, both the target dictionary and the initial dictionary include multiple aligned word pairs, and the aligned word pair includes a source-side word and a target-side word.
[0035] In a possible implementation, the extraction module is configured to:
[0036] Train an initial bilingual dictionary inference model according to the pre-configured monolingual word vector and the initial dictionary;
[0037] Extract the target dictionary from the parallel corpus according to the initial bilingual dictionary inference model and the word alignment model.
[0038] In another possible implementation, the extraction module is further configured to:
[0039] Obtain the first initialization probability of the word alignment model according to the initial bilingual dictionary inference model;
[0040] Perform word alignment learning on the parallel corpus according to the first initialization probability of the word alignment model to obtain a first word alignment probability;
[0041] Determine the target dictionary according to the first word alignment probability.
[0042] In another possible implementation, the extraction module is further configured to:
[0043] According to the initial bilingual dictionary inference model, the first initialization probability p of the word alignment model is obtained through the following formula ini (y|x):
[0044]
[0045] where x is the source-side word, y is the target-side word, and E src (x) is the word vector of the source-side word in the initial bilingual dictionary inference model, and E tgt (y) is the word vector of the target-side word in the initial bilingual dictionary inference model, Y(x) represents the translation target of x in the translation table of the word alignment model, τ is used to indicate the sharpness of the initialization distribution, and y' is any one of the translation targets of x in the translation table of the word alignment model.
[0046] In another possible implementation, the extraction module is further configured to
[0047] obtain a second initialization probability of the word alignment model according to the initial bilingual dictionary inference model, where the second initialization probability is different from the first initialization probability;
[0048] perform word alignment learning on the parallel corpus according to the second initialization probability of the word alignment model to obtain a second word alignment probability;
[0049] perform bidirectional filtering according to the first word alignment probability and the second word alignment probability to obtain the target dictionary.
[0050] In another possible implementation, the device further includes an update module; the update module is configured to
[0051] extract an updated target dictionary from the parallel corpus according to the target bilingual dictionary inference model and the word alignment model;
[0052] train an updated target bilingual dictionary inference model according to the updated target dictionary and the initial dictionary.
[0053] In another possible implementation, the device further includes an acquisition module and a call module;
[0054] The acquisition module is configured to acquire an input source-side word;
[0055] The call module is configured to call the trained target bilingual dictionary inference model according to the source-side word and output the target-side word.
[0056] According to another aspect of the present disclosure, there is provided a computer device, which includes: a processor; and a memory for storing processor-executable instructions;
[0057] Wherein, the processor is configured to:
[0058] Extract a target dictionary from parallel corpora;
[0059] Train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, where the target bilingual dictionary inference model is a neural network model capable of translating source-side words into target-side words;
[0060] Wherein, both the target dictionary and the initial dictionary include a plurality of aligned word pairs, and the aligned word pairs include source-side words and target-side words.
[0061] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0062] In the embodiments of the present disclosure, by extracting a target dictionary from parallel corpora and training a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, parallel corpora are introduced on the basis of the initial dictionary, and the training information of the target bilingual dictionary inference model is enriched by using the target dictionary extracted from the parallel corpora, thereby improving the subsequent bilingual dictionary inference effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The drawings included in and constituting a part of the specification illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.
[0064] Figure 1 is a schematic diagram of the principle of the word vector space mapping method in the related art;
[0065] Figure 2 shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present disclosure;
[0066] Figure 3 shows a flowchart of a bilingual dictionary inference method provided by an exemplary embodiment of the present disclosure;
[0067] Figure 4 shows a flowchart of a bilingual dictionary inference method provided by another exemplary embodiment of the present disclosure;
[0068] Figure 5 shows a schematic diagram of the principle of a bilingual dictionary inference method provided by an exemplary embodiment of the present disclosure;
[0069] Figure 6 shows a schematic diagram of the principle of the bilingual dictionary inference method provided by another exemplary embodiment of the present disclosure;
[0070] Figure 7 shows a flowchart of the bilingual dictionary inference method provided by another exemplary embodiment of the present disclosure;
[0071] Figure 8 shows a schematic structural diagram of the bilingual dictionary inference device provided by an exemplary embodiment of the present disclosure;
[0072] Figure 9 is a block diagram of a device for performing the bilingual dictionary inference method shown according to an exemplary embodiment;
[0073] Figure 10 is a block diagram of a device for performing the bilingual dictionary inference method shown according to another exemplary embodiment. Detailed implementation manners
[0074] The following will describe in detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0075] The special term "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.
[0076] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0077] The bilingual dictionary inference task can provide word-level alignment information for machine translation under the condition of only a small amount of alignment signals, and is an important part of low-resource or unsupervised machine translation systems.
[0078] The bilingual dictionary inference task includes automatically generating a large-scale bilingual dictionary in a data-driven manner from a large-scale monolingual corpus and a small bilingual dictionary, and its feasibility is based on: the word vector spaces of different languages can be transformed through linear transformation. Monolingual word vectors represent the relationships between words in one language in the same space, such as Figure 1As shown in the figure, the word vector space on the left represents the word vector space of Chinese, and the word vector space on the right represents the word vector space of English. If the word vector space structures of different languages are similar, a linear transformation or an orthogonal linear transformation can be used to map the word vector of one language to the word vector space of another language, thereby obtaining a cross-language word representation space. Figure 1 The figure only schematically shows how to map Chinese word vectors to the space of English word vectors through a linear transformation representing rotation. For example, the word vector of the Chinese word "狗" is mapped to the space of the English word vector "狗", the word vector of the Chinese word "猫" is mapped to the space of the English word vector "Cat", the word vector of the Chinese word "马" is mapped to the space of the English word vector "马", and the word vector of the Chinese word "羊" is mapped to the space of the English word vector "羊". This cross-language word representation space describes the word relationship within the language and between the two languages. After that, the inference of the bilingual dictionary can be completed by the nearest neighbor search method, that is, for a source word to be translated, the target word closest to it is searched as the translation.
[0079] The commonly used bilingual dictionary inference method is mainly divided into two steps. The first step is to train the monolingual word vectors of the two languages on large-scale monolingual data to transform each word into a high-dimensional representation. The second step is to train a linear mapping to map the word vector of the source language to the word vector space of the target language.
[0080] The above-mentioned bilingual dictionary inference method has a good inference effect on language pairs that are relatively close, such as English-French, English-German, etc. However, for language pairs that are relatively far away, such as Chinese-English, Chinese-Japanese, etc., the effect is far less good than that of language pairs that are relatively far away, especially in distinguishing synonyms. There will be more problems. One of the main reasons for this phenomenon is that the word vector space gap between distant language pairs is large, and more supervisory signals are needed to complete the conversion between word vector spaces. In addition, the only supervisory data considered in related technologies is the dictionary, while in actual scenarios, alignment signals will also appear in other forms: such as parallel corpora. Parallel corpora also contain alignment information that is helpful for bilingual dictionary inference tasks.
[0081] In order to solve the above technical problems, the embodiments of the present disclosure provide a bilingual dictionary inference method, device and storage medium. The embodiments of the present disclosure extract a target dictionary from a parallel corpus, and obtain a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary training, so that the parallel corpus is introduced on the basis of the initial dictionary, and the target dictionary extracted from the parallel corpus is used to enrich the training information of the target bilingual dictionary inference model, thereby improving the subsequent bilingual dictionary inference effect.
[0082] First, the application scenarios involved in the present disclosure are introduced.
[0083] Please refer to Figure 2 , which shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present disclosure.
[0084] The computer device can be a terminal or a server. The terminal includes a tablet computer, a laptop computer, a desktop computer, and so on. The server can be a single server, or a server cluster composed of several servers, or a cloud computing service center.
[0085] As Figure 1 shown, the computer device includes a processor 10, a memory 20, and a communication interface 30. Those skilled in the art can understand that Figure 1 the structure shown in
[0086] does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:
[0087] The memory 20 can be used to store software programs and modules. The processor 10 executes various functional applications and data processing by running the software programs and modules stored in the memory 20. The memory 20 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system 21, virtual modules, and application programs required for at least one function (such as neural network model training, etc.); the data storage area can store data created according to the use of the computer device. The memory 20 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. Correspondingly, the memory 20 can also include a memory controller to provide the processor 10 with access to the memory 20.
[0088] Among them, the processor 20 is used to perform the following functions: extract a target dictionary from the parallel corpus; train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, and the target bilingual dictionary inference model is a neural network model that translates source-side words into target-side words. Next, several exemplary embodiments are used to introduce the bilingual dictionary inference method provided by the embodiments of the present disclosure.
[0089] Please refer to Figure 3 , which shows a flowchart of a bilingual dictionary inference method provided by an exemplary embodiment of the present disclosure. In this embodiment, the bilingual dictionary inference method is applied to the Figure 2 illustrated computer device as an example. The bilingual dictionary inference method includes:
[0090] Step 301, extract a target dictionary from the parallel corpus.
[0091] The computer device extracts a target dictionary from the parallel corpus. Among them, the parallel corpus includes multiple source-side words and the target-side words that are parallel to each of the multiple source-side words.
[0092] Parallel corpus is a bilingual or multilingual corpus including source - side words and their parallel - corresponding target - side words, and its alignment level is at the word level. The parallel corpus is used to indicate the mapping relationship between source - side words and target - side words.
[0093] The target dictionary includes multiple aligned word pairs, and an aligned word pair includes a source - side word and a target - side word.
[0094] Among them, the source - side word is a word in the first language category, the target - side word is a word in the second language category, and the second language is different from the first language. For example, the first language is Chinese and the second language is English. The embodiments of the present disclosure do not limit the types of the first language and the second language.
[0095] Optionally, the source - side word is also called the original word, and the target - side word corresponding to the source - side word is also called the translation word of the source - side word.
[0096] Step 302: Based on the extracted target dictionary and the pre - configured initial dictionary, train a target bilingual dictionary inference model. The target bilingual dictionary inference model is a neural network model that can translate source - side words into target - side words.
[0097] Among them, the initial dictionary is a pre - configured dictionary, which includes multiple aligned word pairs, and an aligned word pair includes a source - side word and a target - side word.
[0098] The computer device introduces the parallel corpus on the basis of the initial dictionary. After extracting the target dictionary from the parallel corpus, the computer device trains a target bilingual dictionary inference model according to the extracted target dictionary and the pre - configured initial dictionary.
[0099] The target bilingual dictionary inference model is a model obtained by training a neural network using the initial dictionary and the target dictionary. That is, the target bilingual dictionary inference model is a bilingual dictionary inference model determined according to the initial dictionary and the target dictionary.
[0100] The target bilingual dictionary inference model is a neural network model that can translate source - side words into target - side words.
[0101] The target bilingual dictionary inference model is used to convert the input source - side word into a target - side word.
[0102] The target bilingual dictionary inference model is used to represent the mapping relationship between source - side words and target - side words.
[0103] The target bilingual dictionary inference model is a preset mathematical model, and this target bilingual dictionary inference model includes the model coefficients between source - side words and target - side words. The model coefficients can be dynamically modified values.
[0104] In summary, in the embodiments of the present disclosure, a target dictionary is extracted from parallel corpora, and a target bilingual dictionary inference model is trained based on the extracted target dictionary and a pre-configured initial dictionary, so that parallel corpora are introduced on the basis of the initial dictionary, and the target dictionary extracted from the parallel corpora is used to enrich the training information of the target bilingual dictionary inference model, improving the subsequent bilingual dictionary inference effect.
[0105] In the embodiments of the present disclosure, a solution is proposed to extract an additional dictionary, i.e., a target dictionary, from parallel corpora to assist in the learning of mapping. In related technologies, the dictionary obtained by a statistical word alignment model has more noise in the case of less parallel corpora, and the learning quality of word alignment for low-frequency words is not good. For this reason, the embodiments of the present disclosure also propose to combine an existing bilingual dictionary inference model and a statistical word alignment model to better extract the target dictionary from parallel corpora, which can ensure the extraction of high-quality word pairs from parallel corpora and further improve the bilingual dictionary inference effect for distant language pairs. Please refer to Figure 4 , which shows a flowchart of a bilingual dictionary inference method provided by another exemplary embodiment of the present disclosure. In this embodiment, the bilingual dictionary inference method is applied to Figure 2 the computer device shown for illustration. The bilingual dictionary inference method includes:
[0106] Step 401, train an initial bilingual dictionary inference model according to pre-configured monolingual word vectors and an initial dictionary.
[0107] The computer device trains an initial bilingual dictionary inference model according to pre-configured monolingual word vectors and an initial dictionary.
[0108] Among them, both the monolingual word vectors and the initial dictionary are pre-configured. The pre-configured monolingual word vectors are word vectors corresponding to two pre-configured languages respectively. The two languages include a first language and a second language. The initial dictionary includes multiple aligned word pairs, and the aligned word pairs include a source-side word and a target-side word. The source-side word is a word of the first language category, and the target-side word is a word of the second language category, and the second language is different from the first language.
[0109] For example, the first language is Chinese and the second language is English. The embodiments of the present disclosure do not limit the types of the first language and the second language.
[0110] The target bilingual dictionary inference model is a model obtained by training a neural network using pre-configured monolingual word vectors and an initial dictionary. That is, the initial bilingual dictionary inference model is a bilingual dictionary inference model determined according to pre-configured monolingual word vectors and an initial dictionary.
[0111] The computer device normalizes the monolingual word vectors corresponding to two pre-configured languages respectively to obtain the initialized source-side word vectors and the initialized target-side word vectors; for each aligned word pair in the initial dictionary, the mapping matrix is trained by minimizing the CSLS distance (cross-domain similarity local scaling) between the source-side word and the target-side word. After the training is completed, the computer device obtains an initial bilingual dictionary inference model according to the source-side word vectors, the target-side word vectors, and the trained mapping matrix.
[0112] Among them, the initial bilingual dictionary inference model includes multiple source-side words and the target-side words corresponding to the multiple source-side words respectively.
[0113] Schematically, the computer device normalizes the pre-configured monolingual word vectors respectively to obtain the initialized source-side word vectors and the initialized target-side word vectors. In the training phase, for the word pair (x i , y i ) in the initial dictionary L, the mapping matrix W is optimized by minimizing the CSLS distance between x i and y i , and the formula is as follows:
[0114] E src (x i ) = E init_src (x i )W, E tgt (y i ) = E init_tgt (y i )
[0115]
[0116] Among them, E init_src (x i ) is the initialized source-side word vector, E init_tgt (y i ) is the initialized source-side word vector, E src (x i ) is, E tgt (y i ) is the source-side word vector, E tgt (y i ) is the target-side word vector.
[0117] After the training is completed, take E src (x i ) = E init_src (x i )W, E tgt (y i ) = E init_tgt (yi ) Obtain an initial bilingual dictionary inference model.
[0118] Among them, N y (x) represents the k-nearest neighbor parameter of the source-side word x in the target-side words, and N x (y) represents the k-nearest neighbor parameter of the target-side word y in the source-side words, and k takes 10.
[0119] Step 402: Extract a target dictionary from the parallel corpus according to the initial bilingual dictionary inference model and the word alignment model.
[0120] The computer device extracts a target dictionary from the parallel corpus according to the initial bilingual dictionary inference model and the word alignment model.
[0121] Optionally, the initial bilingual dictionary inference model is the RCSLS bilingual dictionary inference model, and the word alignment model is the Fast-Align word alignment model. The embodiments of the present disclosure do not limit this.
[0122] Optionally, the computer device obtains a first initialization probability of the word alignment model according to the initial bilingual dictionary inference model; performs word alignment learning on the parallel corpus according to the first initialization probability of the word alignment model to obtain a first word alignment probability; and determines the target dictionary according to the first word alignment probability.
[0123] Optionally, the computer device obtains a first initialization probability of the word alignment model according to the initial bilingual dictionary inference model, including: obtaining the first initialization probability p of the word alignment model through the following formula according to the initial bilingual dictionary inference model ini (y|x):
[0124]
[0125] Among them, x is the source-side word, y is the target-side word, E src (x) is the word vector of the source-side word in the initial bilingual dictionary inference model, and E tgt (y) is the word vector of the target-side word in the initial bilingual dictionary inference model, Y(x) represents the translation target of x in the translation table of the word alignment model, τ is used to indicate the sharpness of the initialization distribution, and y' is any one of the translation targets of x in the translation table of the word alignment model.
[0126] Optionally, Y(x) represents the possible translation targets of x in the translation table of the word alignment model.
[0127] Through the above method of the first initialization probability, bilingual word vectors and information of a small amount of dictionaries are incorporated into word alignment, that is, cross-lingual word representations and a statistical word alignment model are combined. If word pairs are extracted from parallel corpora using cross-lingual and statistical word alignment models respectively and then combined, the effect is not as good as the solution of the first initialization probability provided in this embodiment of the present disclosure. After calculating the first initialization probability of the word alignment model, the computer device performs word alignment learning on the parallel corpus according to the word alignment model, and obtains the first word alignment probability after the word alignment model converges.
[0128] Optionally, the computer device performs word alignment learning on the parallel corpus according to the first initialization probability of the word alignment model to obtain the first word alignment probability, including: the computer device randomly initializes a position alignment probability table according to the first initialization probability of the word alignment model; according to the first initialization probability table and the initialized position alignment probability table, the EM (Expectation-Maximization algorithm) algorithm is used for iterative learning on the parallel corpus to obtain the first word alignment probability.
[0129] The computer device can determine the target dictionary according to the first word alignment probability. However, if the target dictionary is derived only from a single-direction probability table, the noise is still relatively large. Therefore, this embodiment of the present disclosure also provides a two-way filtering method. The computer device obtains the second initialization probability of the word alignment model according to the initial bilingual dictionary inference model, and the second initialization probability is different from the first initialization probability; according to the second initialization probability of the word alignment model, word alignment learning is performed on the parallel corpus to obtain the second word alignment probability; the target dictionary is obtained through two-way filtering according to the first word alignment probability and the second word alignment probability.
[0130] Among them, the second word alignment probability is different from the first word alignment probability. The first word alignment probability and the second word alignment probability are probability tables in two directions.
[0131] The first word alignment probability is the translation probability obtained after word alignment learning from the source end to the target end. The second word alignment probability is the translation probability obtained after word alignment learning from the target end to the source end.
[0132] Optionally, the computer device obtains the second initialization probability of the word alignment model according to the initial bilingual dictionary inference model, including: according to the initial bilingual dictionary inference model, the second initialization probability p ini (x|y) of the word alignment model is obtained through the following formula:
[0133]
[0134] Among them, x is the source-end word, y is the target-end word, and E src(y) is the word vector of the target - side word in the initial bilingual dictionary inference model, E tgt (x) is the word vector of the source - side word in the initial bilingual dictionary inference model, X(y) represents the possible translation targets of y in the word alignment model translation table, τ is used to indicate the sharpness of the initialization distribution, and x′ is any one of the possible translation targets of y in the word alignment model translation table.
[0135] Optionally, the computer device determines the target dictionary according to the intersection of the first word alignment probability and the second word alignment probability. For example, the first word alignment probability is p s2t , and the second word alignment probability is p t2s . The computer device determines the target dictionary L2 through the following formula:
[0136] L2 = {(x,y)|y = argmax y∈Y(x) p s2t (y|x) ∧ x = argmax x∈X(y) p t2s (x|y)}
[0137] where x is the source - side word, y is the target - side word, Y(x) represents the possible translation targets of x in the word alignment model translation table, and X(y) represents the possible translation targets of y in the word alignment model translation table.
[0138] The embodiments of the present disclosure do not limit the manner in which the computer device determines the target dictionary according to the first word alignment probability.
[0139] Step 403: Train a target bilingual dictionary inference model according to the extracted target dictionary and the pre - configured initial dictionary.
[0140] The computer device trains a target bilingual dictionary inference model according to the extracted target dictionary and the pre - configured initial dictionary.
[0141] Among them, both the target dictionary and the initial dictionary include multiple aligned word pairs, and the aligned word pairs include source - side words and target - side words.
[0142] In an illustrative example, as Figure 5 shown, the computer device trains an initial bilingual dictionary inference model 53 according to the pre - configured monolingual word vector 51 and the initial dictionary 52; according to the initial bilingual dictionary inference model 53 and the word alignment model 54, extracts the target dictionary 56 from the parallel corpus 55; and trains the bilingual dictionary inference model 53 according to the extracted target dictionary 56 and the pre - configured initial dictionary 52.
[0143] The computer device obtains multiple aligned word pairs from the extracted target dictionary and the pre-configured initial dictionary, and preprocesses each aligned word pair among the multiple aligned words to obtain source-side word vectors and target-side word vectors. Among them, the preprocessing includes normalization processing.
[0144] In the training stage, for each aligned word pair in the target dictionary and the initial dictionary, the computer device trains the mapping matrix by minimizing the CSLS distance between the source-side word and the target-side word. After the training is completed, the computer device obtains a target bilingual dictionary inference model according to the source-side word vectors, the target-side word vectors, and the trained mapping matrix.
[0145] It should be noted that the process of training the target bilingual dictionary inference model can be analogously referred to the training process of the above-mentioned initial bilingual dictionary inference model, which will not be elaborated here.
[0146] Optionally, the computer device calculates a loss function according to the extracted target dictionary and the pre-configured initial dictionary, and trains a target bilingual dictionary inference model according to the loss function.
[0147] The loss function is the sum of the loss values of the target bilingual dictionary inference model on the initial dictionary and the target dictionary. For example, the loss function loss of the target bilingual dictionary inference model is:
[0148] loss = RCSLS(L1) + RCSLS(L2)
[0149] Among them, the RCSLS loss value of the target bilingual dictionary inference model on the initial dictionary L1 is RCSLS(L1), and the RCSLS loss value of the target bilingual dictionary inference model on the target dictionary L2 is RCSLS(L2).
[0150] After training the target bilingual dictionary inference model, in order to better extract the target dictionary from the parallel corpus, the step of generating the first initialization probability can be re-executed for iteration until the performance of the model converges, and an updated target bilingual dictionary inference model is obtained. That is, the computer device extracts an updated target dictionary from the parallel corpus according to the target bilingual dictionary inference model and the word alignment model; according to the updated target dictionary and the initial dictionary, an updated target bilingual dictionary inference model is trained.
[0151] It should be noted that the process of the computer device extracting an updated target dictionary from the parallel corpus according to the target bilingual dictionary inference model and the word alignment model can be analogously referred to the relevant details of extracting the target dictionary from the parallel corpus above, and the process of the computer device training an updated target bilingual dictionary inference model according to the updated target dictionary and the initial dictionary can be analogously referred to the relevant details of training the target bilingual dictionary inference model above, which will not be elaborated here.
[0152] In a schematic example, as Figure 6 shown, the computer device trains an initial bilingual dictionary inference model 63 based on pre-configured monolingual word vectors 61 and an initial dictionary 62. An s2t initialization probability is obtained according to the initial bilingual dictionary inference model 63; according to the s2t initialization probability, the s2t word alignment model performs word alignment learning on the parallel corpus 64 to obtain an s2t word alignment probability; a t2s initialization probability is obtained according to the initial bilingual dictionary inference model 63; according to the t2s initialization probability, the t2s word alignment model performs word alignment learning on the parallel corpus 64 to obtain a t2s word alignment probability; a target dictionary 65 is determined according to the intersection of the s2t word alignment probability and the t2s word alignment probability. The bilingual dictionary inference model 63 is trained according to the target dictionary 65 and the pre-configured initial dictionary 62.
[0153] Based on the target bilingual dictionary inference model obtained through the above training, the model usage process includes but is not limited to the following steps, as Figure 7 shown:
[0154] Step 701, the computer device obtains the input source-side word.
[0155] Optionally, after receiving a translation instruction, the computer device obtains the input source-side word to be translated. The source-side word is a word of the first language category. For example, the first language is Chinese.
[0156] Step 702, the computer device calls the trained target bilingual dictionary inference model according to the source-side word and outputs the target-side word.
[0157] Optionally, the target bilingual dictionary inference model is applied in the field of machine translation. The source-side word is a word of the first language category to be translated, and the target-side word is a word of the second language category after translation. The second language is different from the first language. For example, the second language is English. The types of the first language and the second language are not limited in the embodiments of the present disclosure.
[0158] The computer device obtains the trained target bilingual dictionary inference model, inputs the source-side word into the target bilingual dictionary inference model, and outputs the translated word, that is, the target-side word. Among them, the target bilingual dictionary inference model is the target bilingual dictionary inference model obtained in the above method embodiments.
[0159] In summary, the embodiments of the present disclosure provide a bilingual dictionary inference method. On the one hand, the embodiments of the present disclosure introduce parallel corpora on the basis of an initial dictionary, and use the target dictionary extracted from the parallel corpora to enrich the training information of the target bilingual dictionary inference model, improving the subsequent bilingual dictionary inference effect. On the other hand, the computer device extracts the target dictionary from the parallel corpora according to the initial bilingual dictionary inference model and the word alignment model, that is, combines cross-lingual word representations and a statistics-based word alignment model, and learns from the data using their respective advantages. On the other hand, the computer device obtains the first initialization probability of the word alignment model according to the initial bilingual dictionary inference model, improving the performance of the word alignment model. On the other hand, the computer device performs bidirectional filtering according to the first word alignment probability and the second word alignment probability to obtain the target dictionary, ensuring the quality of the target dictionary, and thus helping the learning of the target bilingual dictionary inference model. On the other hand, the computer device extracts the updated target dictionary from the parallel corpora according to the target bilingual dictionary inference model and the word alignment model; trains the updated target bilingual dictionary inference model according to the updated target dictionary and the initial dictionary, and iteratively trains the word alignment model and the target bilingual dictionary inference model to fully exploit the information in the parallel corpora.
[0160] From the application perspective, the bilingual dictionary inference method provided by the embodiments of the present disclosure can be applicable to different distant language pairs; can simultaneously utilize existing parallel corpora and dictionary data, adapting to different information scales; does not require training a deep neural network model, with high efficiency; can be compatible with other word alignment models or bilingual dictionary inference models; and trains the word alignment model while training the target bilingual dictionary inference model.
[0161] The following is the device embodiment of the present disclosure. For parts not elaborated in detail in the device embodiment, reference can be made to the technical details disclosed in the above method embodiment.
[0162] Please refer to Figure 8 , which shows a schematic structural diagram of a bilingual dictionary inference device provided by an exemplary embodiment of the present disclosure. The bilingual dictionary inference device can be implemented as all or part of a computer device through software, hardware, and a combination of both. The device includes: an extraction module 810 and a training module 820.
[0163] The extraction module 810 is configured to extract a target dictionary from parallel corpora;
[0164] The training module 820 is configured to train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary. The target bilingual dictionary inference model is a neural network model that translates source-side words into target-side words;
[0165] Among them, both the target dictionary and the initial dictionary include multiple aligned word pairs, and an aligned word pair includes a source-side word and a target-side word.
[0166] In a possible implementation, the extraction module 810 is configured to:
[0167] Train an initial bilingual dictionary inference model based on pre-configured monolingual word vectors and the initial dictionary;
[0168] Extract the target dictionary from the parallel corpus according to the initial bilingual dictionary inference model and the word alignment model.
[0169] In another possible implementation, the extraction module 810 is further configured to:
[0170] Obtain a first initialization probability of the word alignment model according to the initial bilingual dictionary inference model;
[0171] Perform word alignment learning on the parallel corpus according to the first initialization probability of the word alignment model to obtain a first word alignment probability;
[0172] Determine the target dictionary according to the first word alignment probability.
[0173] In another possible implementation, the extraction module 810 is further configured to:
[0174] According to the initial bilingual dictionary inference model, obtain the first initialization probability p ini (y|x) of the word alignment model through the following formula:
[0175]
[0176] where x is the source-side word, y is the target-side word, E src (x) is the word vector of the source-side word in the initial bilingual dictionary inference model, E tgt (y) is the word vector of the target-side word in the initial bilingual dictionary inference model, Y(x) represents the translation target of x in the translation table of the word alignment model, τ is used to indicate the sharpness of the initialization distribution, and y′ is any one of the translation targets of x in the translation table of the word alignment model.
[0177] In another possible implementation, the extraction module 810 is further configured to:
[0178] Obtain a second initialization probability of the word alignment model according to the initial bilingual dictionary inference model, and the second initialization probability is different from the first initialization probability;
[0179] Perform word alignment learning on the parallel corpus according to the second initialization probability of the word alignment model to obtain a second word alignment probability;
[0180] The target dictionary is obtained through bidirectional filtering based on the first-word alignment probability and the second-word alignment probability.
[0181] In another possible implementation, the device further includes an update module; the update module is configured to:
[0182] Extract the updated target dictionary from the parallel corpus according to the target bilingual dictionary inference model and the word alignment model;
[0183] Train the updated target bilingual dictionary inference model according to the updated target dictionary and the initial dictionary.
[0184] In another possible implementation, the device further includes an acquisition module and a call module;
[0185] The acquisition module is configured to acquire the input source-side word;
[0186] The call module is configured to call the trained target bilingual dictionary inference model according to the source-side word and output the target-side word.
[0187] It should be noted that when the device provided in the above embodiments implements its functions, only the division of the above-mentioned respective functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to actual needs, that is, the content structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0188] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0189] An embodiment of the present disclosure further provides a computer device, which includes a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to: implement the steps executed by the computer device in each of the above method embodiments.
[0190] An embodiment of the present disclosure further provides a non-volatile computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the methods in each of the above method embodiments are implemented.
[0191] Figure 9 It is a block diagram of a device 900 for performing a bilingual dictionary inference method shown according to an exemplary embodiment. For example, the device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0192] Refer to Figure 9, device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.
[0193] The processing component 902 generally controls the overall operation of the device 900, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.
[0194] The memory 904 is configured to store various types of data to support the operation of the device 900. Examples of such data include instructions for any application or method operating on the device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0195] The power component 906 provides power to the various components of the device 900. The power component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 900.
[0196] The multimedia component 908 includes a screen that provides an output interface between the device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0197] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC) that is configured to receive external audio signals when the device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.
[0198] The I / O interface 912 provides an interface between the processing component 902 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0199] The sensor component 914 includes one or more sensors for providing an assessment of various aspects of the status of the device 900. For example, the sensor component 914 can detect the on / off state of the device 900, the relative positioning of components, such as the display and keypad of the device 900, the sensor component 914 can also detect a change in the position of the device 900 or a component of the device 900, the presence or absence of user contact with the device 900, the orientation or acceleration / deceleration of the device 900, and the temperature change of the device 900. The sensor component 914 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 914 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 914 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0200] The communication component 916 is configured to facilitate communication between the device 900 and other devices in a wired or wireless manner. The device 900 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0201] In an exemplary embodiment, the apparatus 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0202] In an exemplary embodiment, a non - volatile computer - readable storage medium is also provided, such as a memory 904 including computer program instructions, and the above computer program instructions can be executed by a processor 920 of the apparatus 900 to complete the above method.
[0203] Figure 10 FIG. is a block diagram of an apparatus 1000 for performing a bilingual dictionary inference method according to another exemplary embodiment. For example, the apparatus 1000 may be provided as a server. Referring to Figure 10 , the apparatus 1000 includes a processing component 1022, which further includes one or more processors, and memory resources represented by a memory 1032 for storing instructions executable by the processing component 1022, such as application programs. The application programs stored in the memory 1032 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1022 is configured to execute instructions to perform the above method.
[0204] The apparatus 1000 may also include a power component 1026 configured to perform power management of the apparatus 1000, a wired or wireless network interface 1050 configured to connect the apparatus 1000 to a network, and an input / output (I / O) interface 1058. The apparatus 1000 may operate based on an operating system stored in the memory 1032, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0205] In an exemplary embodiment, a non - volatile computer - readable storage medium is also provided, such as a memory 1032 including computer program instructions, and the above computer program instructions can be executed by a processing component 1022 of the apparatus 1000 to complete the above method.
[0206] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer - readable storage medium having thereon computer - readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0207] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0208] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0209] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0210] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0211] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0212] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0213] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0214] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A bilingual dictionary inference method, characterized in that, The method includes: extracting a target dictionary from parallel corpora; training a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, where the target bilingual dictionary inference model is a neural network model for translating source-side words into target-side words; wherein both the target dictionary and the initial dictionary include multiple aligned word pairs, and the aligned word pairs include source-side words and target-side words; The extracting the target dictionary from parallel corpora includes: training an initial bilingual dictionary inference model according to pre-configured monolingual word vectors and the initial dictionary; extracting the target dictionary from the parallel corpora according to the initial bilingual dictionary inference model and a word alignment model.
2. The method according to claim 1, characterized in that The extracting the target dictionary from the parallel corpora according to the initial bilingual dictionary inference model and the word alignment model includes: obtaining a first initialization probability of the word alignment model according to the initial bilingual dictionary inference model; performing word alignment learning on the parallel corpora according to the first initialization probability of the word alignment model to obtain a first word alignment probability; determining the target dictionary according to the first word alignment probability.
3. The method according to claim 2, wherein The obtaining the first initialization probability of the word alignment model according to the initial bilingual dictionary inference model includes: According to the initial bilingual dictionary inference model, the first initialization probability p of the word alignment model is obtained through the following formula ini (y|x): where x is the source - side word, y is the target - side word, and E src (x) is the word vector of the source - side word in the initial bilingual dictionary inference model, and E tgt (y) is the word vector of the target - side word in the initial bilingual dictionary inference model, Y(x) represents the translation target of x in the translation table of the word alignment model, τ is used to indicate the sharpness of the initialization distribution, and y′ is any one of the translation targets of x in the translation table of the word alignment model.
4. The method according to claim 2, characterized in that The determining the target dictionary according to the first word alignment probability includes: obtaining a second initialization probability of the word alignment model according to the initial bilingual dictionary inference model, where the second initialization probability is different from the first initialization probability; performing word alignment learning on the parallel corpora according to the second initialization probability of the word alignment model to obtain a second word alignment probability; performing bidirectional filtering according to the first word alignment probability and the second word alignment probability to obtain the target dictionary.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: extracting an updated target dictionary from the parallel corpora according to the target bilingual dictionary inference model and the word alignment model; training an updated target bilingual dictionary inference model according to the updated target dictionary and the initial dictionary.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: obtaining an input source-side word; calling the trained target bilingual dictionary inference model according to the source-side word and outputting the target-side word.
7. A bilingual dictionary inference device, characterized in that, The apparatus includes: an extraction module, configured to extract a target dictionary from parallel corpora; a training module, configured to train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, where the target bilingual dictionary inference model is a neural network model for translating source-side words into target-side words; wherein both the target dictionary and the initial dictionary include multiple aligned word pairs, and the aligned word pairs include source-side words and target-side words; The extraction module is further configured to: train an initial bilingual dictionary inference model according to pre-configured monolingual word vectors and the initial dictionary; extract the target dictionary from the parallel corpora according to the initial bilingual dictionary inference model and a word alignment model.
8. A computer device, characterized in that, The computer device includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: Extract a target dictionary from parallel corpora; Train a target bilingual dictionary inference model according to the extracted target dictionary and a pre-configured initial dictionary, where the target bilingual dictionary inference model is a neural network model that translates source-side words into target-side words; Among them, both the target dictionary and the initial dictionary include multiple aligned word pairs, and the aligned word pairs include source-side words and target-side words; The extracting the target dictionary from parallel corpora includes: Train an initial bilingual dictionary inference model according to pre-configured monolingual word vectors and the initial dictionary; Extract the target dictionary from the parallel corpora according to the initial bilingual dictionary inference model and a word alignment model.
9. A non - volatile computer - readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 6 is implemented.