Parallel sentence pair construction method, device, electronic device and storage medium
Through cross-language language model training, semantic features are determined using word participle word meaning relationships of sample sentences of different languages, which solves the problem of low quality of parallel sentences in scarce resource language, and achieves higher sentence embedding accuracy and parallel corpus quality.
Patent Information
- Application Number
- CN202210688236.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-16
AI Technical Summary
Due to the scarcity of training data, machine translation models for scarce resource languages have low accuracy in sentence embedding, which in turn affects the quality of parallel sentence pairs and the construction of parallel corpus.
Using a cross-language language model, the semantic features are determined based on the word meaning relationship between the word participle contained in sample sentences of different languages, and parallel sentence pairs are constructed based on these features.
It improves the accuracy of sentence embedding, improves the quality of parallel sentence pairs, and overcomes the problem of poor model training results caused by sparse parallel corpus in scarce resource language.
Smart Images

Figure CN115062633B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, electronic device and storage medium for constructing parallel sentence pairs. Background Art
[0002] In recent years, machine translation for scarce resource languages has become a hot research direction. Researchers have conducted in-depth research on pivot languages, transfer learning, data enhancement, semi- / unsupervised training, and other aspects, and have achieved good results. However, in the actual training process, the quantity and quality of training data largely determine the performance of the machine translation model. Therefore, it can be seen that the effect of machine translation for scarce resource languages faces great challenges.
[0003] At present, before carrying out machine translation tasks for scarce resource languages, sentence embedding models are often trained through parallel corpora of multiple languages, and parallel sentence pairs are mined by performing similarity calculations on massive Internet text data, thereby expanding the parallel corpora of different languages. However, this method relies on parallel corpora for model training. When it comes to scarce resource languages, the sentence embedding accuracy is low, resulting in low quality of the mined parallel sentence pairs, which in turn makes the quality of the constructed parallel corpus poor, forming a vicious circle. Summary of the invention
[0004] The present invention provides a parallel sentence pair construction method, device, electronic device and storage medium, which are used to solve the defects in the prior art that the parallel corpus of scarce resource languages is scarce, resulting in poor training effect of the model, so that the accuracy of sentence embedding is too low and the quality of the mined parallel sentence pairs is not high.
[0005] The present invention provides a method for constructing parallel sentence pairs, comprising:
[0006] Obtain a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages;
[0007] determining a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is trained based on a semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages;
[0008] Based on the similarity between the first semantic feature and the second semantic feature, a parallel sentence pair is constructed.
[0009] According to a parallel sentence pair construction method provided by the present invention, the cross-language language model is trained based on the following steps:
[0010] Determining, based on an initial language model, an initial first semantic feature of the first sample sentence and an initial second semantic feature of the second sample sentence;
[0011] Determine the word meaning loss based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word meaning feature of each word in the initial first semantic feature and the word meaning feature of each word in the initial second semantic feature;
[0012] Based on the word meaning loss, parameters of the initial language model are iterated to obtain a cross-language language model.
[0013] According to a parallel sentence pair construction method provided by the present invention, the word meaning loss is determined based on the word meaning relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the similarity between the participle features of each participle in the initial first semantic feature and the participle features of each participle in the initial second semantic feature, including:
[0014] Determine a positive sample word pair based on a word pair in which the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence is a synonym or a near-synonym;
[0015] Determine negative sample word pairs based on word pairs of non-synonymous words and synonymous words in semantic relationships between each participle in the first sample sentence and / or each participle in the second sample sentence;
[0016] The word sense loss is determined based on the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature, and the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature.
[0017] According to a parallel sentence pair construction method provided by the present invention, the initial language model is iterated based on the word meaning loss to obtain a cross-language language model, including:
[0018] Determining a mask loss based on a probability of the masked word segmentation indicated by a word segmentation feature of the masked word in an initial semantic feature, wherein the initial semantic feature includes the initial first semantic feature and / or the initial second semantic feature;
[0019] Based on the word meaning loss and the mask loss, parameters of the initial language model are iterated to obtain a cross-language language model.
[0020] According to a parallel sentence pair construction method provided by the present invention, the initial language model is iterated based on the word meaning loss to obtain a cross-language language model, and then the method further includes:
[0021] Based on the sample parallel sentence pairs, the cross-language language model is fine-tuned.
[0022] According to a method for constructing parallel sentence pairs provided by the present invention, sample sentences in any language are determined based on the following steps:
[0023] Determine the search term in any of the above languages;
[0024] Based on the search results of the search term, construct an initial corpus in any language;
[0025] Classifying the language of each sentence in the initial corpus to obtain the language category of each sentence;
[0026] Screening out sentences in the initial corpus whose language category is not the any language, to obtain a corpus in the any language;
[0027] A sample sentence in any language is obtained from a corpus in any language.
[0028] According to a method for constructing parallel sentence pairs provided by the present invention, constructing the initial corpus of any language based on the search results of the search term includes:
[0029] Determine a target website based on the search results of the search term;
[0030] Building an initial corpus in any language based on the website content of the target website;
[0031] The target websites are websites corresponding to the first preset number of search results when they are arranged in descending order of the search term occurrence frequency.
[0032] According to a parallel sentence pair construction method provided by the present invention, the language classification of each sentence in the initial corpus to obtain the language category of each sentence includes:
[0033] Determining the language category of each sentence in the initial corpus based on a language classification model;
[0034] The language classification model includes a masked language layer and a multi-classification layer. The masked language layer is trained based on masked sentences and masked word segmentation of the masked sentences. The multi-classification layer is trained based on the masked sentences and the language category of the masked sentences in combination with the masked language layer.
[0035] According to a parallel sentence pair construction method provided by the present invention, the language classification is performed on each sentence in the initial corpus to obtain the language category of each sentence, and then the method further includes:
[0036] Based on the semantic encoding layer of the boundary discrimination model, semantic encoding is performed on each sentence in the initial corpus to obtain the semantic features of each sentence;
[0037] Based on the boundary discrimination layer of the boundary discrimination model, the semantic features of each sentence are subjected to boundary discrimination under the language category of each sentence to obtain the boundary discrimination result of each sentence;
[0038] Based on the boundary determination results of the sentences, the sentences are divided into sentences.
[0039] The present invention also provides a parallel sentence pair construction device, comprising:
[0040] A sentence acquisition unit, used to acquire a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages;
[0041] a semantic feature extraction unit, configured to determine a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is trained based on a semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages;
[0042] A parallel sentence pair construction unit is used to construct a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature.
[0043] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for constructing parallel sentence pairs as described above is implemented.
[0044] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for constructing parallel sentence pairs.
[0045] The parallel sentence pair construction method, device, electronic device and storage medium provided by the present invention perform model training based on the semantic relationship between the segmentations contained in the sample sentences of different languages. The sufficiency of training data ensures that the model can fully learn the corresponding relationship between the segmentations contained in the sample sentences of different languages during the training process, so that the performance of the trained cross-language language model is better. When facing scarce resource languages, the accuracy of sentence embedding can be greatly improved, which provides key assistance for the construction of parallel sentence pairs. It completely overcomes the defects of the traditional scheme that the training effect of the model is poor due to the scarcity of parallel corpus of scarce resource languages, thereby making the accuracy of sentence embedding too low and the quality of the mined parallel sentence pairs not high. The construction process of parallel sentence pairs is improved, the quality of parallel sentence pair construction is improved, and the coverage of the field is broadened. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 It is a schematic diagram of the process of constructing a parallel sentence pair provided by the present invention;
[0048] Figure 2 is a schematic diagram of the training process of the cross-language language model provided by the present invention;
[0049] Figure 3 is a flow chart of step 220 in the method for constructing parallel sentence pairs provided by the present invention;
[0050] Figure 4 is a flow chart of step 230 in the method for constructing parallel sentence pairs provided by the present invention;
[0051] Figure 5 is a schematic diagram of the initial language model training process provided by the present invention;
[0052] Figure 6 is a schematic diagram of the determination process of the sample sentence provided by the present invention;
[0053] Figure 7 is a flow chart of step 620 in the method for constructing parallel sentence pairs provided by the present invention;
[0054] Figure 8 It is a schematic diagram of the construction process of the initial corpus provided by the present invention;
[0055] Fig. 9is a schematic diagram of the boundary discrimination process provided by the present invention;
[0056] Fig.10 It is a structural schematic diagram of the boundary discrimination model provided by the present invention;
[0057] Fig.11 It is a structural schematic diagram of a parallel sentence pair construction device provided by the present invention;
[0058] Fig.12 It is a result schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] In recent years, machine translation for scarce resource languages has become a hot research direction. Researchers have conducted in-depth research and exploration in pivot languages, transfer learning, data enhancement, semi-supervised and unsupervised training, and multilingual hybrid modeling, and have achieved good results. However, since the performance of machine translation models depends largely on the quantity and quality of training data, machine translation for scarce resource languages still faces great challenges in practical applications.
[0061] With the advancement of large-scale pre-training model technology, parallel corpus mining has made progress. Researchers use parallel corpora of multiple different languages to train cross-language sentence embedding models, and mine parallel sentence pairs by calculating similarities on massive Internet text data, thereby expanding the parallel corpus of different language pairs. However, for scarce resource languages, the accuracy of sentence embeddings obtained by cross-language sentence embedding models is low, resulting in poor quality of mining parallel sentence pairs for scarce resource languages.
[0062] The traditional solution initially uses the different web page tag information in the multilingual website to perform page-level alignment, then divide the web page content into sentences, and then perform sentence-level alignment, thereby constructing a parallel corpus. However, the above solution is too dependent on the web page content and structured information of the original multilingual website, so it cannot effectively utilize the massive Internet data. In addition, the parallel corpus constructed according to this solution mostly has distinct website characteristics, is too limited in scope, and has insufficient coverage.
[0063] Furthermore, there are traditional solutions that use the similarity of sentence embeddings in different languages to mine parallel corpora. In such solutions, the higher the similarity of sentence embeddings, the greater the probability that two sentences in different languages are parallel sentence pairs. However, the above solutions make insufficient use of web page information and rely on parallel corpora for model training. Therefore, when facing scarce resource languages, the accuracy of sentence embeddings obtained by the model is low, resulting in low quality of constructed parallel sentence pairs.
[0064] In view of the above situation, the present invention provides a parallel sentence pair construction method, which aims to use the semantic relationship between the word segments contained in sample sentences of different languages to perform model training, so that the trained cross-language language model can improve the accuracy of sentence embedding when facing scarce resource languages, and achieve the improvement of parallel sentence pair mining quality. Figure 1 is a flow chart of the method for constructing parallel sentence pairs provided by the present invention, such as Figure 1 As shown, the method includes:
[0065] Step 110, obtaining a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages;
[0066] Specifically, since a parallel sentence pair is two sentences with exactly the same semantics but belonging to different languages, before constructing a parallel sentence pair, it is first necessary to determine two sentences in different languages, namely, a first sentence and a second sentence. The first sentence and the second sentence can be sentences in any language, but they must belong to different languages. Specifically in an embodiment of the present invention, the languages corresponding to the first sentence and the second sentence can be Chinese and Hausa, Chinese and Persian, Chinese and Burmese, Chinese and Malay, etc.
[0067] The sentence pair consisting of the first sentence and the second sentence can be one or more, which is not specifically limited in the embodiment of the present invention. In the case of multiple sentence pairs, it is necessary to construct parallel sentence pairs for each sentence pair, that is, it is necessary to determine whether the two sentences in different languages in each sentence pair can form a parallel sentence pair.
[0068] Step 120, determining a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is trained based on a semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages;
[0069] Specifically, after the first sentence and the second sentence are obtained through step 110, step 120 can be executed to apply the cross-language language model to determine the first semantic feature of the first sentence and the second semantic feature of the second sentence. This process specifically includes the following steps:
[0070] First, the first sentence and the second sentence are input into the cross-language language model. Then, the cross-language language model performs semantic extraction on the input first sentence and the second sentence respectively, extracts features in the first sentence and the second sentence that can characterize the semantic information of the corresponding sentences, thereby obtaining the semantic features of the first sentence, i.e., the first semantic features, and the semantic features of the second sentence, i.e., the second semantic features. The semantic features here can also be referred to as sentence embedding vectors of the corresponding sentences.
[0071] Before inputting the first sentence and the second sentence into the cross-language language model, the cross-language language model can also be pre-trained. Different from the traditional scheme of directly using parallel corpora for model training, in the embodiment of the present invention, taking into account the scarcity of parallel corpora of scarce resource languages, using a very small amount of parallel corpora to train the model will cause the model to be unable to accurately learn the mapping relationship between parallel corpora, thereby making the accuracy of the semantic features obtained by the model low. Therefore, the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence is used to train the initial language model, so that the initial language model can fully learn the corresponding relationship between the participles contained in the sample sentences of different languages, and finally obtain a trained cross-language language model. The first sample sentence and the second sample sentence here are sample sentences of different languages.
[0072] The semantic relationship here can be understood as the correlation between the meanings of different participles. For example, the semantic relationship between the participle "adjacent" and the participle "close" is close, that is, they are synonyms or near-synonyms; the semantic relationship between the participle "close" and the participle "distant" is distant, the two are not synonyms or near-synonyms, but antonyms; the semantic relationship between the participle "close" and the participle "semantic" is distant, the two are not synonyms or near-synonyms, and are unrelated words.
[0073] Compared with the parallel corpus used in the traditional solution, the semantic relationship between the segmented words contained in the sample sentences of two different languages used in the embodiment of the present invention is more delicate and easier to obtain; and, precisely because the coarse-grained parallel corpus is scarce and difficult to obtain, the embodiment of the present invention uses the semantic relationship between the fine-grained segmented words that is easy to obtain to train the initial language model, and takes the similarity between the segmented word features corresponding to the semantic relationship as the training target, so that the initial language model can judge the similarity between the segmented word features according to the semantic relationship between the segmented words, so that when the semantic relationship is close, that is, when the two segmented words are synonyms or near synonyms, the similarity between the segmented word features output by the initial language model is as high as possible through training; conversely, when the semantic relationship is distant, that is, when the two segmented words are not synonyms or near synonyms, the similarity between the segmented word features output by the initial language model is as low as possible through training.
[0074] Compared with the parallel corpus which is rare and difficult to obtain, the semantic relationship between the segmented words used in the embodiments of the present invention is very easy to obtain. For example, it can be directly obtained through reference books such as textbooks and dictionaries, and the quantity is sufficient. Sufficient training data is the key to ensuring the performance of the model, that is, it enables the model to fully learn the semantic relationship between the segmented words contained in the sample sentences of different languages during the training process, which provides a key support for the construction process of parallel sentence pairs.
[0075] In the embodiment of the present invention, the semantic relationship between word segmentations is used to train the initial language model, so that the performance of the trained cross-language language model can be better and the extracted semantic features can be more accurate, thereby making the parallel sentence pair construction process more precise and the constructed parallel sentence pairs more accurate.
[0076] Step 130: construct a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature.
[0077] Specifically, in step 120, based on obtaining the first semantic feature of the first sentence and the second semantic feature of the second sentence, step 130 may be performed to construct a parallel sentence pair according to the similarity between the first semantic feature and the second semantic feature. This process may specifically include the following steps:
[0078] First, determining the similarity between the first semantic feature of the first sentence and the second semantic feature of the second sentence, wherein the similarity is calculated by using cosine similarity, Euclidean distance, Minkowski distance, etc. between the semantic features;
[0079] Then, based on the similarity between the first semantic feature and the second semantic feature, a parallel sentence pair can be constructed by comparing the result with a preset similarity threshold. Specifically, the similarity between the first semantic feature and the second semantic feature is compared with the preset similarity to obtain a comparison result. Furthermore, if the comparison result shows that the similarity between the semantic features reaches or even exceeds the preset similarity, it indicates that the semantic information of the first sentence and the second sentence has become consistent, and the two can form a parallel sentence pair.
[0080] Accordingly, if the comparison result shows that the similarity between the semantic features does not reach the preset similarity, it means that there is a difference in the semantic information of the first sentence and the second sentence, and the two cannot form a parallel sentence pair.
[0081] It should be noted that the preset similarity here is a pre-set basis for determining whether two sentences in different languages can form a parallel sentence pair, which can be obtained through experiments.
[0082] The parallel sentence pair construction method provided by the present invention performs model training based on the semantic relationship between the segmentations contained in the sample sentences of different languages. The sufficiency of the training data ensures that the model can fully learn the corresponding relationship between the segmentations contained in the sample sentences of different languages during the training process, so that the performance of the trained cross-language language model is better. When facing the scarce resource language, the accuracy of sentence embedding can be greatly improved, which provides a key boost for the construction of parallel sentence pairs. It completely overcomes the defects of the traditional scheme that the training effect of the model is poor due to the scarcity of parallel corpus of the scarce resource language, thereby making the accuracy of sentence embedding too low and the quality of the mined parallel sentence pairs not high. The construction process of parallel sentence pairs is improved, the quality of parallel sentence pair construction is improved, and the coverage of the field is broadened.
[0083] Based on the above embodiment, in step 130, the process of constructing a parallel sentence pair according to the similarity between the first semantic feature and the second semantic feature specifically includes the following steps:
[0084] First, determine the similarity between the first semantic feature of the first sentence and the second semantic feature of the second sentence, where the similarity can be specifically expressed as cosine similarity, Euclidean distance, Minkowski distance, etc.;
[0085] At the same time, a preset number of sample sentences that are nearest neighbors to the first sentence and the second sentence are determined, and the similarities between any two of the preset number of sample sentences are determined, and then the average of the similarities between all the combinations of any two of them is calculated, thereby obtaining the average similarity of the nearest neighbor samples;
[0086] Then, based on the similarity between the first semantic feature and the second semantic feature and the average similarity of the nearest neighbor samples, the similarity between the first semantic feature of the first sentence and the second semantic feature of the second sentence is determined. Specifically, the ratio of the similarity between the first semantic feature and the second semantic feature to the average similarity of the nearest neighbor samples is used as the similarity between the first semantic feature and the second semantic feature. This similarity can also be called the similarity of the marginal function.
[0087] Subsequently, since the similarity of the marginal function can characterize the degree of similarity between the semantic information of the first sentence and the semantic information of the second sentence, parallel sentence pairs can be constructed based on the similarity of the marginal functions. That is, the larger the ratio, the higher the similarity of the marginal function, and the greater the probability that the first sentence and the second sentence can form a parallel sentence pair; conversely, the smaller the ratio, the lower the similarity of the marginal function, and the less likely the first sentence and the second sentence can form a parallel sentence pair.
[0088] Based on the above embodiments, Figure 2 is a schematic diagram of the training process of the cross-language language model provided by the present invention, such as Figure 2As shown in the figure, the cross-language language model is trained based on the following steps:
[0089] Step 210, determining an initial first semantic feature of the first sample sentence and an initial second semantic feature of the second sample sentence based on the initial language model;
[0090] Step 220, determining the word meaning loss based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word meaning feature of each word in the initial first semantic feature and the word meaning feature of each word in the initial second semantic feature;
[0091] Step 230 , based on the word meaning loss, iterate the parameters of the initial language model to obtain a cross-language language model.
[0092] Specifically, the training process of the cross-language language model may include the following steps:
[0093] First, execute step 210 to determine an initial language model, where the initial language model may be a language mask model or other types of language models, such as a statistical language model (SLM), a neural network language model (NNLM), etc., then input the first sample sentence and the second sample sentence into the initial language model, and perform semantic extraction on the first sample sentence and the second sample sentence by the initial language model, and finally obtain an initial first semantic feature of the first sample sentence output by the initial language model, and an initial second semantic feature of the second sample sentence;
[0094] Subsequently, step 220 is executed to calculate the word meaning loss of the initial language model training process based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word meaning features of each word in the initial first semantic feature and the word meaning features of each word in the initial second semantic feature. Specifically, the similarity between the word meaning features when the word meaning relationship is close and the similarity between the word meaning features when the word meaning relationship is distant are determined respectively; then, the word meaning loss of the initial language model is determined based on the two similarities; here, the word meaning relationship that is close can be understood as synonyms or near-synonyms, and the word meaning relationship that is distant can be understood as non-synonyms and near-synonyms, such as irrelevant words, antonyms, etc.;
[0095] Since the training goal of the initial language model is to make the similarity between the segmentation features of each segmentation in the initial first semantic feature and the segmentation features of each segmentation in the initial second semantic feature as high as possible when the semantic relationship between each segmentation in the first sample sentence and each segmentation in the second sample sentence is close; correspondingly, when the semantic relationship between each segmentation in the first sample sentence and each segmentation in the second sample sentence is distant, make the similarity between the segmentation features of each segmentation in the initial first semantic feature and the segmentation features of each segmentation in the initial second semantic feature as low as possible; therefore, if the similarity between the segmentation features when the semantic relationship is close is high, and the similarity between the segmentation features when the semantic relationship is distant is low, it can be determined that the semantic loss of the initial language model is small; accordingly, if the similarity between the segmentation features when the semantic relationship is close is low, and / or the similarity between the segmentation features when the semantic relationship is distant is high, it can be determined that the semantic loss of the initial language model is large.
[0096] Thereafter, step 230 can be executed to iterate the parameters of the initial language model according to the word meaning loss. Specifically, the parameters of the initial language model can be adjusted according to the word meaning loss so that the adjusted initial language model can determine that the similarity between word segmentation features is as high as possible when the word meaning relationship is close, and determine that the similarity between word segmentation features is as low as possible when the word meaning relationship is distant. After the training is completed, a cross-language language model is obtained.
[0097] Based on the above embodiments, Figure 3 is a flow chart of step 220 in the parallel sentence pair construction method provided by the present invention, such as Figure 3 As shown, step 220 includes:
[0098] Step 221, determining positive sample word pairs based on the word pairs whose semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence is synonymous or near-synonymous;
[0099] Step 222, determining negative sample word pairs based on the word pairs of non-synonymous words and synonymous words in the semantic relationship between each participle in the first sample sentence and / or each participle in the second sample sentence;
[0100] Step 223, determining the word meaning loss based on the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature, and the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature.
[0101] Specifically, in step 220, the process of determining the word meaning loss of the initial language model according to the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word feature of each word in the initial first semantic feature and the word feature of each word in the initial second semantic feature, specifically includes the following steps:
[0102] Step 221, determining positive sample word pairs according to the word pairs in which the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence is synonymous or near-synonymous, that is, selecting the participles with similar semantic relationship, that is, the word pairs in which the semantic relationship is synonymous or near-synonymous, from each participle in the first sample sentence and each participle in the second sample sentence, that is, the positive sample word pairs, wherein the two participles in the positive sample word pairs come from the first sample sentence and the second sample sentence respectively;
[0103] Step 222, according to the word pairs of non-synonymous and synonymous words in the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, determine the negative sample word pairs, that is, from each participle in the first sample sentence and / or each participle in the second sample sentence, randomly select two participles other than the positive sample words as the negative sample word pairs, and the two participles in the negative sample word pairs are from the same sample sentence or from different sample sentences.
[0104] Step 223, determining the similarity between the segmentation features of the positive sample word pair in the initial first semantic feature and the initial second semantic feature, specifically, firstly determining the segmentation feature corresponding to the segmentation feature in the initial first semantic feature of the segmentation in the positive sample word pair originating from the first sample sentence, and the segmentation feature corresponding to the segmentation feature in the initial second semantic feature of the segmentation in the second sample sentence, and then determining the similarity between the two segmentation features;
[0105] At the same time, if the two segmented words in the negative sample word pair are derived from the same sample sentence, the similarity between the two segmented word features corresponding to the negative sample word pair in the initial first semantic feature or the initial second semantic feature is determined; if the two segmented words in the negative sample word pair are derived from different sample sentences, the similarity between the segmented word features corresponding to the negative sample word pair in the initial first semantic feature and the initial second semantic feature is determined;
[0106] After that, the word meaning loss of the initial language model can be determined according to the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature, and the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature, that is, when the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature is high, and the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature is low, it is determined that the word meaning loss of the initial language model is small; correspondingly, when the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature is low, and / or the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature is high, it is determined that the word meaning loss of the initial language model is large.
[0107] Based on the above embodiments, Figure 4 is a flow chart of step 230 in the method for constructing parallel sentence pairs provided by the present invention, such as Figure 4 As shown, step 230 includes:
[0108] Step 231, determining a mask loss based on a probability of a masked word segmentation indicated by a word segmentation feature of the masked word segmentation in the initial semantic feature, wherein the initial semantic feature includes an initial first semantic feature and / or an initial second semantic feature;
[0109] Step 232, based on the word meaning loss and the mask loss, the parameters of the initial language model are iterated to obtain a cross-language language model.
[0110] In an embodiment of the present invention, when training the initial language model, the similarity between the word segmentation features corresponding to the word semantic relationship is introduced on the basis of the word semantic relationship between the word segmentations contained in the sample sentences of different languages. This is used as a training target and trained together with the masking task to improve the embedding representation capability of the cross-language language model.
[0111] Therefore, in step 230, when performing parameter iteration on the initial language model, in addition to applying the word meaning loss, the mask loss of the mask task needs to be determined, so that the two can jointly perform the parameter iteration process for the initial language model. This process specifically includes the following steps:
[0112] Step 231, taking the masked participle in the first sample sentence and / or the masked participle in the second sample sentence as reference when the initial language model is input, determining the participle feature corresponding to the masked participle from each participle feature of the initial first semantic feature and / or each participle feature of the initial second semantic feature, and indicating the probability of the masked participle according to the participle feature corresponding to the masked participle, and determining the loss of the masking task, that is, the masking loss. In other words, the masking loss is determined according to whether the meanings of the masked participle and its participle feature are consistent with each other;
[0113] Step 232, based on the mask loss and the word meaning loss, the parameters of the initial language model are iterated to obtain a cross-language language model. The joint training of multiple tasks can greatly improve the performance of the model, so that the embedding representation ability of the trained cross-language language model is better. In this process, the parameter adjustment process based on the mask loss is actually to make the word segmentation features corresponding to the masked word segmentation in the initial semantic features obtained by the initial language model as close as possible to the masked word segmentation in the sample sentence, that is, the parameters of the initial language model are adjusted with the purpose of making the word segmentation features corresponding to the masked word segmentation in the sample sentence consistent with those in the initial semantic features, thereby improving the initial language model's ability to understand the overall semantics of the sentence.
[0114] Based on the above embodiments, Figure 5 is a schematic diagram of the initial language model training process provided by the present invention, such as Figure 5 As shown, first, from each segmentation of the first sample sentence and each segmentation of the second sample sentence, two segmentations whose semantic relationship is synonymous or near synonymous are selected, namely, X3 and y1 to form a positive sample word pair (x3, y1), and from each segmentation of the first sample sentence and / or each segmentation of the second sample sentence, two segmentations other than the positive sample word are randomly selected, namely, y2 and y3 to form a negative sample word pair (y2, y3);
[0115] The language of the first sample sentence is "th", that is, Thai; the language of the second sample sentence is "zh", that is, Chinese. Figure 5 The "bos" in the string indicates the start symbol, and the "eos" indicates the end symbol.
[0116] Then, the word meaning loss L of the initial language model is determined based on the similarity between the word segmentation feature x'3 corresponding to the positive sample word pair in the initial first semantic feature, the word segmentation feature y'1 corresponding to the positive sample word pair in the initial second semantic feature, and the similarity between the two word segmentation features y'2 and y'3 corresponding to the negative sample word pair in the initial second semantic feature. NCE ;
[0117] At the same time, with the masked participle x2 in the first sample sentence and / or the masked participle y4 in the second sample sentence as reference, the participle feature x'2 corresponding to the masked participle x2 is determined from the initial first semantic feature, and / or the participle feature y'4 corresponding to the masked participle y4 is determined from the initial second semantic feature; then, according to the probability that the participle feature x'2 corresponding to the masked participle x2 indicates the masked participle x2, and / or the participle feature y'4 corresponding to the masked participle y4 indicates the probability of the masked participle y4, the mask loss L of the initial language model is determined. MLM It should be noted that the masked words in the sample sentences are masked word segmentations;
[0118] After that, we can use the mask loss L MLM and word meaning loss L NcE , iterate the parameters of the initial language model to obtain a cross-language language model.
[0119] The loss function of the initial language model training process is:
[0120]
[0121]
[0122] Where log p θ (w i |(WM)) is the mask loss, W represents the input sample sentence sequence, w i represents the i-th word segmentation in the input sample sentence sequence, M is the set of masked word segmentations in the sample sentence, and p θ is the model parameter, L NCE is the word meaning loss, k is the language category set, N is the number of negative sample word pairs, j and k + is a positive sample word pair, k - is a randomly selected negative sample word pair, α and β are the weights of mask loss and word meaning loss, which are obtained through experimental parameter adjustment.
[0123] Based on the above embodiment, in step 230, based on the word meaning loss, the parameters of the initial language model are iterated to obtain a cross-language language model, and then the following steps are further included:
[0124] Fine-tune the cross-language language model based on sample parallel sentence pairs.
[0125] In order to further improve the model's cross-language representation ability for word embedding, in an embodiment of the present invention, after the initial language model training is completed and the cross-language language model is obtained, fine-tuning training can be performed on the cross-language language model, that is, the aforementioned process of iterating parameters of the initial language model based on word meaning loss is regarded as pre-training of the model, and after the pre-training is completed, a small number of sample parallel sentence pairs are used to fine-tune the cross-language language model, that is, only the input of the pre-training process needs to be replaced with the sample parallel sentence pairs, and the rest remains unchanged.
[0126] During the fine-tuning training process, Kullback-Leibler Divergence Loss (relative entropy loss) can be introduced to ensure the stability of the fine-tuning training process.
[0127] Based on the above embodiments, Figure 6 is a schematic diagram of the determination process of the sample sentence provided by the present invention, such as Figure 6 As shown, sample sentences in any language are determined based on the following steps:
[0128] Step 610, determining the search term of the language;
[0129] Step 620, constructing an initial corpus of the language based on the search results of the search term;
[0130] Step 630, classifying the language of each sentence in the initial corpus to obtain the language category of each sentence;
[0131] Step 640, filtering out sentences in the initial corpus that are not in the language category of the language, and obtaining a corpus in the language;
[0132] Step 650: Obtain sample sentences of the language from a corpus of the language.
[0133] Specifically, the process of determining the sample sentences used in the model training process includes the following steps:
[0134] First, step 610 is executed to determine a search term in any language. Specifically, character strings of different lengths are counted in the original monolingual corpus with a length range of 5 to 20 characters, and the characters are used as search terms.
[0135] Then, step 620 is executed to apply the search term obtained in the previous step to search in the search engine to obtain the search results, that is, each website containing the search term, and the content of the website is crawled down as the initial corpus of the language;
[0136] Then, step 630 is executed. Considering that the corpus of the language obtained in step 620 is relatively rough, and may contain unclear sentences or sentences of other languages, it is necessary to further process it to obtain a more standardized corpus of the language. Specifically, each sentence in the initial corpus is language-classified to determine the language to which each sentence belongs and obtain the language category of each sentence. This process can be completed with the help of a language classification model, that is, each sentence in the initial corpus can be input into the language classification model, and the language classification model performs language discrimination on each input sentence, and finally obtains the language type of each sentence output by the language classification model.
[0137] Before inputting each sentence into the classification model, a language classification model can be pre-trained. The training process of the language classification model combines the masking task and the multi-classification task. That is, both the masking task and the multi-classification task are used as training targets to jointly train the initial language classification model to obtain a trained language classification model.
[0138] It should be noted that in order to improve the accuracy of language classification results, in an embodiment of the present invention, before each sentence is input into the language classification model, a character set encoding-based method can also be used to filter each sentence in the initial corpus to improve the language classification process for multiple languages.
[0139] Thereafter, step 640 is executed to filter out sentences according to the language category of each sentence in the initial corpus, that is, to filter out sentences in the initial corpus that are not in the language category of the language, so as to obtain a corpus of the language. The above process is repeated to construct corpora of different languages.
[0140] Finally, step 650 is executed to directly obtain sample sentences of the language from the corpus of the language.
[0141] Based on the above embodiments, Figure 7 is a flow chart of step 620 in the method for constructing parallel sentence pairs provided by the present invention, such as Figure 7 As shown, step 620 includes:
[0142] Step 621, based on the search results of the search term, determine the target website, the target website being the website corresponding to the first preset number of search results when arranged in descending order of the search term's frequency of occurrence;
[0143] Step 622: construct an initial corpus of the language based on the website content of the target website.
[0144] Specifically, in step 620, the process of constructing the initial corpus of the language according to the search results of the search term includes the following steps:
[0145] Step 621, first, the search term of the language is used as an input term to search in a search engine to obtain multiple websites containing the search term, i.e., search results; then, the search results are sorted in descending order according to the frequency of occurrence of the search term to obtain an arrangement order of the search results, and a preset number of search results are selected from the arrangement order as target search results, i.e., target websites;
[0146] It should be noted that the preset number here can be set according to actual needs. As a preference, in the embodiment of the present invention, the preset number is determined to be 3, that is, the websites corresponding to the first 3 search results in the arrangement order are taken as the target websites.
[0147] Step 622 is to construct an initial corpus of the language based on the website content of the target website. Specifically, the first Q internal links in the target website are obtained through a directional crawling algorithm, that is, the first Q internal links in the target website that contain the most search terms. Q can be set according to actual conditions, and the web page content of the website corresponding to these Q internal links is crawled as the initial corpus of the language.
[0148] Based on the above embodiments, Figure 8 It is a schematic diagram of the construction process of the initial corpus provided by the present invention, such as Figure 8 As shown in FIG. 1 , the construction process of the initial corpus of any language may specifically include the following steps:
[0149] First, you can search for information sources, that is, determine a search term in any language, and enter the search term into a search engine to search, and obtain search results, that is, multiple websites containing the search term;
[0150] Then, according to the search results of the search terms, the target website is determined, that is, the search results are sorted in descending order according to the frequency of the search terms to obtain the arrangement order of the search results, and the first preset number of search results are selected from the arrangement order as the target search results, that is, the target website; the preset number here can be set according to actual needs;
[0151] Subsequently, the target website can be preliminarily identified as to the language of the sentences. If the language category is not that language, the content of the target website can be crawled and stored for backup.
[0152] Accordingly, if the language category is this language, information source screening can be performed, that is, the value of the corpus information contained in the target website can be judged to determine the configuration method of the initial corpus. Specifically, all internal links in the target website can be obtained. If the number of internal links exceeds the preset threshold, it means that the corpus information contained in the target website is relatively sufficient and has a high value. At this time, the corpus information can be obtained from the top Q-level internal links of the target website through a refined configuration method, and used as the initial corpus of the language; accordingly, if the number of internal links does not exceed the preset threshold, it means that the corpus information contained in the target website is not sufficient and has a low value. At this time, the corpus information can be obtained from the top Q-level internal links of the target website through an automated expansion method, and used as the initial corpus of the language.
[0153] The preset threshold here can be set accordingly according to actual needs, for example, it can be 5000, 8000, 10000, etc., and preferably, in the embodiment of the present invention, the preset number is determined to be ten thousand, that is, when the number of all internal links in the target website exceeds ten thousand, a refined configuration method is adopted to construct the initial corpus of the language; when the number of all internal links in the target website does not exceed ten thousand, an automated expansion method is adopted to construct the initial corpus of the language.
[0154] Based on the above embodiment, step 630 includes:
[0155] Based on the language classification model, determine the language category of each sentence in the initial corpus;
[0156] The language classification model includes a masked language layer and a multi-classification layer. The masked language layer is trained based on masked sentences and masked word segmentation of masked sentences. The multi-classification layer is trained based on masked sentences and language categories of masked sentences jointly with the masked language layer.
[0157] Specifically, in step 630, the process of performing language classification on each sentence in the initial corpus to obtain the language category of each sentence may specifically be: inputting each sentence in the initial corpus into the language classification model, each sentence input into the language classification model first passes through the masked language layer, the masked language layer performs high-dimensional feature extraction on each sentence to obtain a high-dimensional feature matrix, and then passes through the multi-classification layer, the multi-classification layer performs maximum pooling on the high-dimensional feature matrix to convert the features of the matrix dimension into features of the vector dimension, and performs language discrimination based on the features of the vector dimension, specifically, calculating the probability of each sentence belonging to each language through the fully connected layer and the Softmax layer in the multi-classification layer to obtain the predicted language category of each sentence.
[0158] The language classification model includes a masked language layer and a multi-classification layer, wherein the masked language layer is used to perform masking tasks, and the multi-classification layer is used to perform multi-language language classification tasks. During the model training process, the masking task and the multi-classification task are used as training targets, the masked cross entropy loss and the multi-classification cross entropy loss are integrated, and the unsupervised training data is fully utilized for model training to improve the representation and prediction capabilities of the language classification model; specifically, during training, masked sentences and masked word segmentation of masked sentences can be applied for unsupervised training to obtain the masked language layer, and on the basis of the masked sentences and the language categories of the masked sentences, the masked language layer is combined for supervised training to obtain the multi-classification layer.
[0159] The method provided by the embodiment of the present invention trains a masked language layer through masked sentences and masked word segmentation of masked sentences, and on the basis of masked sentences and their language categories, jointly trains the masked language layer to obtain a multi-classification layer. It takes a multi-task enhanced cross-language hybrid training method as a benchmark, combines massive multi-language data and multi-classification targets for hybrid training, and can greatly improve the language discrimination effect of the language classification model.
[0160] Based on the above embodiment, the loss function of the language classification model training process is:
[0161]
[0162] Where log p θ (w i |(WM)) is the mask loss of the language classification model for the mask task, W is the input mask sentence sequence, w i represents the i-th masked word in the input masked sentence sequence, M is the set of masked words in the masked sentence, and p θ is the model parameter, y j log(p j ) is the classification loss of the language classification model for multi-classification tasks, K is the language category set, y j represents the language category of the jth masked sentence, p j It represents the language category of the j-th masked sentence predicted by the language classification model. α and β are the weights of the mask loss and classification loss, respectively, which are obtained through experimental parameter adjustment.
[0163] Based on the above embodiments, Fig. 9 is a schematic diagram of the boundary discrimination process provided by the present invention, such as Fig. 9 As shown, in step 630, each sentence in the initial corpus is language classified to obtain the language category of each sentence, and then the following is further included:
[0164] Step 910, based on the semantic coding layer of the boundary discrimination model, semantic coding is performed on each sentence in the initial corpus to obtain the semantic features of each sentence;
[0165] Step 920, based on the boundary discrimination layer of the boundary discrimination model, perform boundary discrimination on the semantic features of each sentence under the language category of each sentence to obtain the boundary discrimination result of each sentence;
[0166] Step 930, based on the boundary determination results of each sentence, each sentence is divided into sentences.
[0167] Specifically, after the language category of each sentence in the initial corpus is obtained in step 630, considering that the division accuracy of each sentence may not be high, that is, there is a situation where the sentence segmentation is unclear, therefore, the boundary of each sentence in the initial corpus can be judged, so as to obtain the boundary judgment result of each sentence. The specific process includes the following steps:
[0168] First, execute step 910, semantically encode each sentence in the initial corpus according to the semantic encoding layer in the boundary discrimination model, so as to obtain the semantic features of each sentence. Specifically, each sentence in the initial corpus and the language category of each sentence are input into the semantic encoding layer in the boundary discrimination model, and the semantic encoding layer semantically encodes each input sentence to obtain the semantic features of each sentence output by the semantic encoding layer; the semantic encoding layer here is actually a multi-layer encoder based on the self-attention mechanism;
[0169] Then, step 920 is executed. Considering that the semantic features of different languages may be quite different, a routing module for semantic enhancement for each language may be added during the boundary determination process. This module can make the semantic features of each language more distinct. The specific process may be that the semantic features of each sentence are input into the boundary determination layer in the boundary determination model. The boundary determination layer uses the language category of each sentence as a reference and performs boundary determination according to the semantic features of each sentence, that is, the semantic features of each sentence are subjected to boundary determination under the corresponding language category. It can also be understood that the semantic features of each sentence are divided into different languages, and boundary determination is performed for the semantic features under the same language, thereby obtaining the boundary determination results of each sentence. This can greatly improve the accuracy of the boundary determination process.
[0170] Then, step 930 is executed to segment each sentence in the initial corpus according to the boundary discrimination result of each sentence output by the boundary discrimination model, that is, to segment each sentence in the initial corpus.
[0171] It should be noted that before applying the boundary discrimination model for boundary discrimination, the boundary discrimination model can also be pre-trained. In terms of training data, in the embodiment of the present invention, the boundary information between the context blocks in the web page content crawled through the network is used as the sentence end boundary, and the training data is constructed by intercepting text strings of different lengths before and after, splicing multiple text strings containing this boundary information, etc.; in addition, some data are collected from websites such as TED (Technology, Entertainment, Design) and Wikipedia, and some training data are formed by sorting and summarizing.
[0172] In terms of model training, in an embodiment of the present invention, the semantic encoding layer is first pre-trained as a mask language model using the text block content containing paragraph-level information, and then a small amount of annotated data carrying sentence-end boundaries is applied to train the pre-trained model. This process can make full use of the contextual information in the paragraph-level data during the pre-training process to make up for the lack of sentence-end punctuation training data.
[0173] In the embodiment of the present invention, chapter-level sequence data is used for model training, which significantly improves the performance of the sentence-end boundary discrimination model in the absence of sentence-end punctuation training data.
[0174] Fig.10 is a schematic diagram of the structure of the boundary discrimination model provided by the present invention, such as Fig.10 As shown in the figure, the structure of the boundary discrimination model is similar to Bidirectional Encoder Representation from Transformers; the process of applying the boundary discrimination model to perform boundary discrimination includes the following steps: first, each sentence in the initial corpus is used as input text, and is input into the semantic encoding layer in the boundary discrimination model together with the language category of each sentence for semantic encoding, so as to obtain the semantic features of each sentence output by the semantic encoding layer; the semantic encoding layer here is actually a multi-layer encoder based on the self-attention mechanism;
[0175] Then, semantic enhancement is performed through the routing module, that is, the semantic features of each sentence are semantically enhanced based on the language category of each language, for example, Fig.10 The semantic features of Thai, Burmese, Tibetan and other languages are strengthened respectively; It should be noted that the routing module here consists of two mapping layers and an activation layer between the two mapping layers;
[0176] Subsequently, the boundary discrimination layer in the boundary discrimination model is applied to perform boundary discrimination according to the semantic features of each sentence. The specific process may be that the enhanced semantic features of each sentence first pass through the Long Short-Term Memory (LSTM) layer which has excellent performance in time series modeling, and then, the Softmax layer is used to perform category judgment on each position in each sentence to determine whether each position is the end of the sentence, and finally the boundary discrimination results of each sentence are obtained.
[0177] The parallel sentence pair construction device provided by the present invention is described below. The parallel sentence pair construction device described below and the parallel sentence pair construction method described above can be referenced to each other.
[0178] Fig.11 is a schematic diagram of the structure of the parallel sentence pair construction device provided by the present invention, such as Fig.11 As shown, the device comprises:
[0179] A sentence acquisition unit 1110 is used to acquire a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages;
[0180] A semantic feature extraction unit 1120 is used to determine a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is trained based on a semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages;
[0181] The parallel sentence pair construction unit 1130 is used to construct a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature.
[0182] The parallel sentence pair construction device provided by the present invention performs model training based on the semantic relationship between the segmentations contained in the sample sentences of different languages. The sufficiency of the training data ensures that the model can fully learn the corresponding relationship between the segmentations contained in the sample sentences of different languages during the training process, so that the performance of the trained cross-language language model is better. When facing the scarce resource language, the accuracy of sentence embedding can be greatly improved, which provides a key boost for the construction of parallel sentence pairs. It completely overcomes the defects of the traditional scheme that the training effect of the model is poor due to the scarcity of parallel corpus of the scarce resource language, thereby making the accuracy of sentence embedding too low and the quality of the mined parallel sentence pairs not high. It improves the construction process of parallel sentence pairs, realizes the improvement of the construction quality of parallel sentence pairs, and broadens the field coverage.
[0183] Based on the above embodiment, the device further includes a model training unit, which is used to:
[0184] Determining, based on an initial language model, an initial first semantic feature of the first sample sentence and an initial second semantic feature of the second sample sentence;
[0185] Determine the word meaning loss based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word meaning feature of each word in the initial first semantic feature and the word meaning feature of each word in the initial second semantic feature;
[0186] Based on the word meaning loss, parameters of the initial language model are iterated to obtain a cross-language language model.
[0187] Based on the above embodiment, the model training unit is used to:
[0188] Determine a positive sample word pair based on a word pair in which the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence is a synonym or a near-synonym;
[0189] Determine negative sample word pairs based on word pairs of non-synonymous words and synonymous words in semantic relationships between each participle in the first sample sentence and / or each participle in the second sample sentence;
[0190] The word sense loss is determined based on the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature, and the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature.
[0191] Based on the above embodiment, the model training unit is used to:
[0192] Determining a mask loss based on a probability of the masked word segmentation indicated by a word segmentation feature of the masked word in an initial semantic feature, wherein the initial semantic feature includes the initial first semantic feature and / or the initial second semantic feature;
[0193] Based on the word meaning loss and the mask loss, parameters of the initial language model are iterated to obtain a cross-language language model.
[0194] Based on the above embodiment, the device further includes a model fine-tuning unit, which is used to:
[0195] Based on the sample parallel sentence pairs, the cross-language language model is fine-tuned.
[0196] Based on the above embodiment, the device further includes a sample sentence determining unit, which is used to:
[0197] Determine the search terms for the language;
[0198] Based on the search results of the search term, construct an initial corpus of the language;
[0199] Classifying the language of each sentence in the initial corpus to obtain the language category of each sentence;
[0200] Screening out sentences in the initial corpus that are not in the language category of the language to obtain a corpus in the language;
[0201] Get sample sentences of the language from the corpus of the language.
[0202] Based on the above embodiment, the sample sentence determination unit is used to:
[0203] Determine a target website based on the search results of the search term;
[0204] Building an initial corpus of the language based on the website content of the target website;
[0205] The target websites are websites corresponding to the first preset number of search results when they are arranged in descending order of the search term occurrence frequency.
[0206] Based on the above embodiment, the sample sentence determination unit is used to:
[0207] Determining the language category of each sentence in the initial corpus based on a language classification model;
[0208] The language classification model includes a masked language layer and a multi-classification layer. The masked language layer is trained based on masked sentences and masked word segmentation of the masked sentences. The multi-classification layer is trained based on the masked sentences and the language category of the masked sentences in combination with the masked language layer.
[0209] Based on the above embodiment, the sample sentence determination unit is used to:
[0210] Based on the semantic encoding layer of the boundary discrimination model, semantic encoding is performed on each sentence in the initial corpus to obtain the semantic features of each sentence;
[0211] Based on the boundary discrimination layer of the boundary discrimination model, the semantic features of each sentence are subjected to boundary discrimination under the language category of each sentence to obtain the boundary discrimination result of each sentence;
[0212] Based on the boundary determination results of the sentences, the sentences are divided into sentences.
[0213] Fig.12 An example of a physical result schematic diagram of an electronic device is shown as follows: Fig.12As shown, the electronic device may include: a processor 1210, a communication interface 1220, a memory 1230 and a communication bus 1240, wherein the processor 1210, the communication interface 1220 and the memory 1230 communicate with each other through the communication bus 1240. The processor 1210 may call the logic instructions in the memory 1230 to execute a parallel sentence pair construction method, the method comprising: obtaining a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages; determining a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is obtained by training based on the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, wherein the first sample sentence and the second sample sentence correspond to different languages; and constructing a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature.
[0214] In addition, the logic instructions in the above-mentioned memory 1230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0215] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the parallel sentence pair construction method provided by the above methods, and the method includes: obtaining a first sentence and a second sentence, the first sentence and the second sentence corresponding to different languages; based on a cross-language language model, determining a first semantic feature of the first sentence and a second semantic feature of the second sentence, the cross-language language model is obtained by training based on the word-meaning relationship between each participle in a first sample sentence and each participle in a second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages; based on the similarity between the first semantic feature and the second semantic feature, constructing a parallel sentence pair.
[0216] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the parallel sentence pair construction method provided by the above-mentioned methods, the method comprising: obtaining a first sentence and a second sentence, the first sentence and the second sentence corresponding to different languages; based on a cross-language language model, determining a first semantic feature of the first sentence and a second semantic feature of the second sentence, the cross-language language model being trained based on the semantic relationship between each participle in a first sample sentence and each participle in a second sample sentence, the first sample sentence and the second sample sentence corresponding to different languages; constructing a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature.
[0217] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0218] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0219] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing parallel sentence pairs, characterized in that: include: Obtain a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages; determining a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is trained based on a semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages; constructing a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature; The cross-language language model is trained based on the following steps: Determining, based on an initial language model, an initial first semantic feature of the first sample sentence and an initial second semantic feature of the second sample sentence; Determine the word meaning loss based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word meaning feature of each word in the initial first semantic feature and the word meaning feature of each word in the initial second semantic feature; Based on the word meaning loss, parameters of the initial language model are iterated to obtain a cross-language language model.
2. The method for constructing parallel sentence pairs according to claim 1, characterized in that: The determining of the word meaning loss based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word feature of each word in the initial first semantic feature and the word feature of each word in the initial second semantic feature, includes: Determine a positive sample word pair based on a word pair in which the semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence is a synonym or a near-synonym; Determine negative sample word pairs based on word pairs of non-synonymous words and synonymous words in semantic relationships between each participle in the first sample sentence and / or each participle in the second sample sentence; The word sense loss is determined based on the similarity between the word segmentation features of the positive sample word pairs in the initial first semantic feature and the initial second semantic feature, and the similarity between the word segmentation features of the negative sample word pairs in the initial first semantic feature and / or the initial second semantic feature.
3. The method for constructing parallel sentence pairs according to claim 1, characterized in that: The step of iterating parameters of the initial language model based on the word meaning loss to obtain a cross-language language model includes: Determining a mask loss based on a probability of the masked word segmentation indicated by a word segmentation feature of the masked word in an initial semantic feature, wherein the initial semantic feature includes the initial first semantic feature and / or the initial second semantic feature; Based on the word meaning loss and the mask loss, parameters of the initial language model are iterated to obtain a cross-language language model.
4. The method for constructing parallel sentence pairs according to claim 1, characterized in that: The step of iterating parameters of the initial language model based on the word meaning loss to obtain a cross-language language model further includes: Based on the sample parallel sentence pairs, the cross-language language model is fine-tuned.
5. The method for constructing parallel sentence pairs according to any one of claims 1 to 4, characterized in that: Sample sentences in any language are determined based on the following steps: Determine the search term in any of the above languages; Based on the search results of the search term, construct an initial corpus in any language; Classifying the language of each sentence in the initial corpus to obtain the language category of each sentence; Screening out sentences in the initial corpus whose language category is not the any language, to obtain a corpus in the any language; A sample sentence in any language is obtained from a corpus in any language.
6. The method for constructing parallel sentence pairs according to claim 5, characterized in that: The step of constructing the initial corpus in any language based on the search results of the search term includes: Determine a target website based on the search results of the search term; Building an initial corpus in any language based on the website content of the target website; The target websites are websites corresponding to the first preset number of search results when they are arranged in descending order of the search term occurrence frequency.
7. The method for constructing parallel sentence pairs according to claim 5, characterized in that: The language classification of each sentence in the initial corpus to obtain the language category of each sentence includes: Determining the language category of each sentence in the initial corpus based on a language classification model; The language classification model includes a masked language layer and a multi-classification layer. The masked language layer is trained based on masked sentences and masked word segmentation of the masked sentences. The multi-classification layer is trained based on the masked sentences and the language category of the masked sentences in combination with the masked language layer.
8. The method for constructing parallel sentence pairs according to claim 5, characterized in that: The language classification of each sentence in the initial corpus is performed respectively to obtain the language category of each sentence, and then the following steps are further included: Based on the semantic encoding layer of the boundary discrimination model, semantic encoding is performed on each sentence in the initial corpus to obtain the semantic features of each sentence; Based on the boundary discrimination layer of the boundary discrimination model, the semantic features of each sentence are subjected to boundary discrimination under the language category of each sentence to obtain the boundary discrimination result of each sentence; Based on the boundary determination results of the sentences, the sentences are divided into sentences.
9. A parallel sentence pair construction device, characterized in that: include: A sentence acquisition unit, used to acquire a first sentence and a second sentence, wherein the first sentence and the second sentence correspond to different languages; a semantic feature extraction unit, configured to determine a first semantic feature of the first sentence and a second semantic feature of the second sentence based on a cross-language language model, wherein the cross-language language model is trained based on a semantic relationship between each participle in the first sample sentence and each participle in the second sample sentence, and the first sample sentence and the second sample sentence correspond to different languages; a parallel sentence pair construction unit, configured to construct a parallel sentence pair based on the similarity between the first semantic feature and the second semantic feature; The cross-language language model is trained based on the following steps: Determining, based on an initial language model, an initial first semantic feature of the first sample sentence and an initial second semantic feature of the second sample sentence; Determine the word meaning loss based on the word meaning relationship between each word in the first sample sentence and each word in the second sample sentence, and the similarity between the word meaning feature of each word in the initial first semantic feature and the word meaning feature of each word in the initial second semantic feature; Based on the word meaning loss, parameters of the initial language model are iterated to obtain a cross-language language model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the parallel sentence pair construction method as described in any one of claims 1 to 8 is implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing parallel sentence pairs as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multilingual news subject similarity comparison method based on parallel corpus
CN108519971A
Pre-training method and device of intelligent translation model and storage medium
CN111460838A