A Chinese Pre-training Method Based on Subword Encoding and Inverse Document Frequency Masking
By using the method based on subword encoding and inverse document frequency occlusion, the problems of low word-level encoding efficiency and sparse word-level encoding in Chinese pre-training are solved, and faster model convergence and better pre-training effects are achieved.
Patent Information
- Application Number
- CN202110480038.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-04-30
AI Technical Summary
The existing Chinese pre-training methods lead to excessive long sentence sequences during word-level encoding, degradation of coding efficiency, and the problem of sparse data and unlogged words during word-level encoding, affecting the pre-training performance.
Using a method based on subword encoding and inverse document frequency occlusion, the dictionary and occurrence probability are learned through a monolingual language model, subword encoding is performed, and the inverse document frequency is calculated. The inverse document frequency occlusion prediction task is used for pre-training, and the high-frequency subword elements are obstructed for prediction.
It alleviates the problem of degradation in word-level coding efficiency, eliminates the problem of unlogged words in word-level coding, and makes the pre-trained language model converge faster, has better results, and fully integrates word-level information.
Smart Images

Figure CN115270764B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer information processing, and particularly relates to a Chinese pre-training method based on sub-word encoding and inverse document frequency masking. Background Art
[0002] With the continuous development of information processing technology and artificial intelligence, pre-training methods have been widely used in different natural language processing tasks and play a crucial role. Many large enterprises and research institutions, such as Google, Facebook, Baidu, Alibaba, etc., have conducted large-scale in-depth research on pre-training methods.
[0003] When training a Chinese pre-training model, the methods of English pre-training cannot be completely copied. Different from English, Chinese has no explicit word boundaries and has two different encoding methods, character-level and word-level. Nowadays, most Chinese pre-training methods adopt character-level encoding, which will lead to too long sentence sequences, decreased encoding efficiency, and the need for additional training methods to integrate word-level information. In addition, directly using word-level encoding will cause problems such as data sparsity, out-of-vocabulary words, and error propagation, which will damage the performance of pre-training. Summary of the Invention
[0004] The present invention is made to solve the above problems, and aims to provide a Chinese pre-training method based on sub-word encoding and inverse document frequency masking.
[0005] The present invention provides a Chinese pre-training method based on sub-word encoding and inverse document frequency masking for pre-training a Chinese language model, having the following features, including the following steps: Step 1, collect a large-scale unsupervised Chinese corpus, and learn a unigram language model through an iterative algorithm according to the large-scale unsupervised Chinese corpus to obtain a dictionary and occurrence probabilities for sub-word encoding in the unigram language model;
[0006] Step 2, perform sub-word encoding on the input text of the Chinese language model based on the unigram language model to obtain a sub-word element sequence;
[0007] Step 3, calculate the inverse document frequency of each sub-word element in the sub-word element sequence;
[0008] Step 4, perform pre-training through an inverse document frequency masking prediction task, and the inverse document frequency masking prediction task is to mask the sub-word element with the highest inverse document frequency, and the Chinese language model performs pre-training by predicting the masked sub-word element;
[0009] Step 5: Input the large-scale unsupervised Chinese corpus into the Chinese language model, perform pre-training through the inverse document frequency masking prediction task after subword encoding and calculation of the inverse document frequency, and obtain the trained Chinese language model after large-scale computational training.
[0010] Among them, the unigram language model assumes that each subword element appears independently, and regards a piece of text as a sequence of subword elements. The probability of occurrence of this text is the product of the probability of occurrence of all subword elements in the text:
[0011]
[0012]
[0013] In formula (1), V is a learnable dictionary, x is the input text, and x i For a subword element, the likelihood L of the Chinese language model on the entire data set formed by a large-scale unsupervised Chinese corpus is calculated to optimize the occurrence probability p(x). The formula is as follows:
[0014]
[0015] The Chinese pre-training method based on subword encoding and inverse document frequency masking provided by the present invention may also have the following features: wherein, in step 1, the dictionary and the occurrence probability are calculated by an iterative algorithm, as follows:
[0016] Step 1-1, based on large-scale unsupervised Chinese corpus, a relatively large dictionary is heuristically obtained as the iterative seed dictionary;
[0017] Step 1-2, fix the seed dictionary and use the EM algorithm to optimize the occurrence probability p(x);
[0018] Steps 1-3, for each subword element x in the seed dictionary i Calculate loss, where loss is the decrease in likelihood L on the entire dataset after the corresponding subword element is deleted from the seed dictionary;
[0019] Steps 1-4: sort the subword elements in the seed dictionary according to the loss, and then remove a certain proportion of the subword elements with the smallest loss;
[0020] Step 1-5, repeat step 1-2-step 1-4 until a dictionary V of a reasonable size is obtained.
[0021] In the Chinese pre-training method based on sub-word encoding and inverse document frequency masking provided by the present invention, it may further have the following features: Among them, after obtaining the vocabulary V and the occurrence probability p(x) through iteration, for any piece of text x, through the unigram language model, the sub-word element x* with the highest probability is obtained, and the formula is as follows:
[0022]
[0023] In formula (3), S(x) is the candidate set of sub-word elements in text x obtained according to the vocabulary V, and then the Viterbi algorithm is used to find the optimal solution x*.
[0024] In the Chinese pre-training method based on sub-word encoding and inverse document frequency masking provided by the present invention, it may further have the following features: Among them, the calculation method of the inverse document frequency IDF is:
[0025]
[0026] In formula (4), w is any word, N is the total number of documents in the corpus, and N w is the number of documents in the corpus that contain the word w.
[0027] In the Chinese pre-training method based on sub-word encoding and inverse document frequency masking provided by the present invention, it may further have the following features: Among them, the Chinese language model is the Chinese BERT model.
[0028] Functions and Effects of the Invention
[0029] According to a Chinese pre-training method based on sub-word encoding and inverse document frequency masking involved in the present invention, because sub-word encoding is performed based on the unigram language model, and encoding is performed through data-driven character-word mixing, it can alleviate the problems of sequence process and efficiency decline caused by character-level encoding, and can also eliminate problems such as out-of-vocabulary words caused by word-level encoding; and by calculating the inverse document frequency of each sub-word element in the sub-word element sequence obtained after sub-word encoding, masking the sub-word element with the highest inverse document frequency, and then predicting the masked sub-word element for pre-training, it can make the pre-training language model converge faster and have better effects, and fully integrate word-level information. Description of the Drawings
[0030] Figure 1 is a flowchart of a Chinese pre-training method based on sub-word encoding and inverse document frequency masking in an embodiment of the present invention;
[0031] Figure 2 is a schematic diagram of a method of a Chinese pre-training method based on sub-word encoding and inverse document frequency masking in an embodiment of the present invention;
[0032] Figure 3It is a performance comparison chart of the model using the inverse document frequency masking prediction task and the random masking training method in the embodiments of the present invention after different pre-training steps;
[0033] Figure 4 It is a comparison of the computational cost of sub-word encoding in the embodiments of the present invention with that of a model encoding based on character encoding under different text lengths. Detailed implementation manners
[0034] In order to make the technical means and effects achieved by the present invention easy to understand, the present invention will be specifically described below in conjunction with embodiments and the accompanying drawings.
[0035] <Embodiment>
[0036] Figure 1 It is a flowchart of a Chinese pre-training method based on sub-word encoding and inverse document frequency masking in the embodiments of the present invention.
[0037] As Figure 1 shown, a Chinese pre-training method based on sub-word encoding and inverse document frequency masking in this embodiment is used for pre-training of a Chinese language model. The Chinese language model can be directly used in natural language processing of Chinese, such as tasks like Chinese word segmentation, sentiment recognition, automatic question answering, reading comprehension, etc. This Chinese language model is a Chinese BERT model, and it includes the following steps:
[0038] Step 1, collect a large-scale unsupervised Chinese corpus, and learn a unigram language model through an iterative algorithm based on the large-scale unsupervised Chinese corpus to obtain the dictionary and occurrence probabilities for sub-word encoding in the unigram language model.
[0039] In this embodiment, the large-scale unsupervised Chinese corpus is a publicly open-source Chinese corpus, which can be directly downloaded online. The large-scale unsupervised Chinese corpus of the present invention includes four parts: Wikipedia, news corpus, encyclopedia Q&A, and community Q&A, forming a large dataset that covers formal and informal texts in different fields, ensuring the data scale and diversity.
[0040] The unigram language model assumes that each sub-word element appears independently. Taking a text segment as a sequence of sub-word elements, the occurrence probability of this text segment is the product of the occurrence probabilities of all sub-word elements in the text:
[0041]
[0042]
[0043] In formula (1), V is a learnable dictionary, x is the input text, x iAs a subword element, the occurrence probability p(x) is optimized by calculating the likelihood L of the Chinese language model on the entire dataset formed by large-scale unsupervised Chinese corpora. The formula is as follows:
[0044]
[0045] In step 1, the dictionary and occurrence probability are calculated by an iterative algorithm, specifically as follows:
[0046] Step 1-1: Based on large-scale unsupervised Chinese corpora, a relatively large dictionary is heuristically obtained as the seed dictionary for iteration;
[0047] Step 1-2: Fix the seed dictionary and use the EM algorithm to optimize the occurrence probability p(x);
[0048] Step 1-3: For each subword element x in the seed dictionary i Calculate the loss, where the loss is the decrease in the likelihood L on the entire dataset when the corresponding subword element is removed from the seed dictionary;
[0049] Step 1-4: Sort the subword elements in the seed dictionary according to the loss, and then remove a certain proportion of the subword elements with the smallest loss;
[0050] Step 1-5: Repeat steps 1-2 to 1-4 until a dictionary V of a reasonable size is obtained.
[0051] Step 2: Based on the unigram language model, perform subword encoding on the input text of the Chinese language model to obtain a sequence of subword elements.
[0052] After obtaining the dictionary V and occurrence probability p(x) through iteration, for any piece of text x, through the unigram language model, the subword element x* with the highest probability is obtained. The formula is as follows:
[0053]
[0054] In formula (3), S(x) is the candidate set of subword elements in text x obtained according to dictionary V, and then the Viterbi algorithm is used to find the optimal solution x*.
[0055] Step 3: Calculate the inverse document frequency of each subword element in the sequence of subword elements.
[0056] The calculation method of the inverse document frequency IDF is:
[0057]
[0058] In formula (4), w is any word, N is the total number of documents in the corpus, and N w is the number of documents in the corpus that contain word w.
[0059] Step 4, perform pre-training through the inverse document frequency masking prediction task, which is to mask the sub-word elements with the highest inverse document frequency, and the Chinese language model performs pre-training by predicting the masked sub-word elements.
[0060] In this embodiment, pre-training is performed through the masked language model, that is, some words in the text are masked and replaced with a specific symbol [MASK], and then the Chinese language model is used to predict the masked part based on the context. When training other pre-training models, a random masking method is generally adopted, that is, the words and characters selected for masking are random. However, the inverse document frequency masking prediction task adopted in the present invention is based on the inverse document frequency of words. The higher the inverse document frequency, the fewer times the word appears, the more difficult it is to learn, and the more it needs to be masked.
[0061] The specific method of the inverse document frequency masking prediction task is as follows: for each sentence of text to be input into the model, after sub-word encoding, a sequence of sub-word elements is obtained, and then the inverse document frequency (IDF) of each sub-word element in the input text is statistically obtained on the entire data set and sorted from high to low. Select the K words with the highest IDF, and then randomly select M words from the K words for masking, and let the model predict these M words. It is stipulated that M < K, and generally K = 2 * M. Through the inverse document frequency masking prediction task of the present invention, the prediction ability of the training model for difficult words and rare words can be highlighted, making the entire pre-training task more difficult and enhancing the training effect of the model.
[0062] Step 5, input a large-scale unsupervised Chinese corpus into the Chinese language model, and perform pre-training through the inverse document frequency masking prediction task after sub-word encoding and calculating the inverse document frequency respectively. After large-scale calculation training, a trained Chinese language model is obtained.
[0063] In this embodiment, the trained Chinese BERT model is also fine-tuned on the training data for specific tasks, and then the model performance is tested on the task-specific test set.
[0064] Figure 2 It is a schematic diagram of a method for Chinese pre-training based on sub-word encoding and inverse document frequency masking in an embodiment of the present invention.
[0065] As Figure 2 shown, the input text of the Chinese BERT model is sub-word encoded through a unigram language model to obtain multiple sub-word elements, and then the inverse document frequency of each sub-word element is calculated, and the sub-word elements with the highest inverse document frequency are masked. The Chinese language model performs pre-training by predicting the masked sub-word elements.
[0066] In this embodiment, the Chinese BERT model is also used to compare the performance on downstream tasks after different numbers of pre-training steps (Pre-TrainingSteps) using the inverse document frequency masking prediction task (IDF-Masking) and the random masking training method (Random-Masking) of the present invention respectively. Figure 3 FIG. is a performance comparison diagram of the inverse document frequency masking prediction task and the random masking training method of the present invention after different numbers of pre-training steps.
[0067] As Figure 3 shown, the performance of the model pre-trained using the inverse document frequency masking prediction task of the present invention is significantly better than that of the model pre-trained using the random masking training method.
[0068] In this embodiment, the sub-word encoding (Ours) of the present invention is also compared with the model encoding based on character encoding (BERT) in terms of computational complexity at different text lengths. Figure 4 FIG. is a comparison of the computational complexity of the sub-word encoding of the present invention with the model encoding based on character encoding at different text lengths.
[0069] As Figure 4 shown, the computational complexity of the present invention at different text lengths is significantly less than that of BERT. Since the less the computational complexity, the faster the model speed, the encoding speed of the sub-word encoding of the present invention is significantly better than that of the character-level encoding, and as the sentence length gradually increases, the speed advantage of the present invention becomes more and more obvious.
[0070] Functions and effects of the embodiment
[0071] According to a Chinese pre-training method based on sub-word encoding and inverse document frequency masking involved in this embodiment, because sub-word encoding is performed based on a unigram language model and encoded through data-driven word-character mixing, it can alleviate the problems of sequence process and efficiency decline caused by character-level encoding, and can also eliminate problems such as out-of-vocabulary words caused by word-level encoding; and by calculating the inverse document frequency of each sub-word element in the sub-word element sequence obtained after sub-word encoding, masking the sub-word element with the highest inverse document frequency, and then predicting the masked sub-word element for pre-training, it can make the pre-trained language model converge faster and have better effects, and fully integrate word-level information.
[0072] The above embodiments are preferred cases of the present invention and are not used to limit the protection scope of the present invention.
Claims
1. A Chinese pre-training method based on sub-word encoding and inverse document frequency masking, for pre-training a Chinese language model, characterized in that Including the following steps: Step 1, collect a large-scale unsupervised Chinese corpus, and learn a unigram language model from the large-scale unsupervised Chinese corpus through an iterative algorithm to obtain the dictionary and occurrence probabilities for subword encoding in the unigram language model; Step 2, perform subword encoding on the input text of the Chinese language model based on the unigram language model to obtain a sequence of subword elements; Step 3, calculate the inverse document frequency of each subword element in the sequence of subword elements; Step 4, perform pre-training through an inverse document frequency masking prediction task, where the inverse document frequency masking prediction task is to mask the subword element with the highest inverse document frequency, and the Chinese language model performs pre-training by predicting the masked subword element; Step 5, input the large-scale unsupervised Chinese corpus into the Chinese language model, perform pre-training through the subword encoding, calculating the inverse document frequency, and then through the inverse document frequency masking prediction task. After large-scale computational training, the trained Chinese language model is obtained. Among them, the unigram language model assumes that each subword element appears independently. A text segment is regarded as a sequence of subword elements, and the occurrence probability of this text segment is the product of the occurrence probabilities of all subword elements in the text: In formula (1), V is a learnable dictionary, x is the input text, and x i is a sub-word element. The occurrence probability p(x) is optimized by calculating the likelihood L of the Chinese language model on the entire data set formed by the large-scale unsupervised Chinese corpus. The formula is as follows:
2. The Chinese pre-training method based on subword encoding and inverse document frequency masking according to claim 1, characterized in that: Among them, In step 1, the dictionary and the occurrence probabilities are calculated through an iterative algorithm, specifically as follows: Step 1-1, heuristically obtain a relatively large dictionary from the large-scale unsupervised Chinese corpus as the seed dictionary for iteration; Step 1-2, fix the seed dictionary and optimize the occurrence probability p(x) using the EM algorithm; Step 1-3, for each sub-word element x in the seed dictionary i Calculate the loss, where the loss is the decrease in the likelihood L on the entire dataset when the corresponding sub-word element is deleted from the seed dictionary; Step 1-4, sort the subword elements in the seed dictionary according to the loss, and then remove a certain proportion of the subword elements with the smallest loss; Step 1-5, repeat steps 1-2 to 1-4 until a dictionary V of a reasonable size is obtained.
3. The Chinese pre-training method based on subword encoding and inverse document frequency masking according to claim 1, characterized in that: Among them, After iteratively obtaining the dictionary V and the occurrence probability p(x), for any text segment x, through the unigram language model, the subword element x* with the highest probability is obtained, and the formula is as follows: x * = argmax P(x), x ∈ S(X) (3) In formula (3), S(x) is the candidate set of subword elements in the text segment x obtained according to the dictionary V, and then the Viterbi algorithm is used to find the optimal solution x*.
4. The Chinese pre-training method based on subword encoding and inverse document frequency masking according to claim 1, characterized in that: Among them, The calculation method of the inverse document frequency IDF is: In formula (4), w is any word, N is the total number of documents in the corpus, and N w is the number of documents in the corpus that contain the word w.
5. The Chinese pre-training method based on subword encoding and inverse document frequency masking according to claim 1, characterized in that: Among them, The Chinese language model is a Chinese BERT model.
Citation Information
Patent Citations
Event extraction method and system fusing dependency information and pre-trained language model
CN111897908A
Multi-label text classification method based on statistics and pre-trained language model
CN112214599A