A Language Model Pretraining Method Based on Coreference Resolution
By switching word_learning and phrase_learning modes through the adaptive training mode, the semantic training of pronouns, phrases, and entities is enhanced, and the problem of lack of semantic information in the word vector representation in the common reference digestion task is solved, and the prediction accuracy and semantic representation ability are improved.
Patent Information
- Application Number
- CN202111237852.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-25
AI Technical Summary
In the prior art, in the co-referential digestion task, word vector representation lacks semantic information, resulting in low prediction accuracy, especially low digestion accuracy for pronouns.
Through the adaptive training mode, the word_learning mode and phrase_learning mode are switched according to the loss function, the semantic training of pronouns, phrases, and entities is increased, and the semantic representation ability of the model is enhanced.
The prediction accuracy in the co-reference digestion task is improved, especially in the aspect of pronoun digestion, and the model's ability to capture semantic information is enhanced.
Smart Images

Figure CN113886591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a language model pre-training method based on coreference elimination. Background Art
[0002] The task of coreference resolution is to classify expressions in a text that refer to the same entity (including pronouns, named entities, noun phrases, etc.). At present, advanced end-to-end neural network coreference resolution models all take word vectors as input, obtain span representations based on attention modules, and then score span pairs for coreference, thereby achieving coreference resolution. The implementation of coreference resolution requires the use of contextual information and world knowledge for reasoning, that is, advanced language models are needed to obtain more semantically rich word vector representations. Bert (Bidirectional Encoder Representations from Transformer) is a language model that is currently used more frequently. It predicts the masked words based on the Transformer algorithm framework by randomly masking words in massive text corpora (this article mainly focuses on English corpora, and words refer to a word in English). However, this pre-training method has some disadvantages. For example, in the sentence "Harry Potter is a wonderful work of magic literature", if only "Harry" is covered, it is easy to predict "Potter". In this way, the word vector of "Harry Potter" learned by the model cannot contain information such as "magic literature", that is, the context information is not rich enough. Especially in the field of coreference resolution, language representation with richer semantic information is needed to capture the relationship between entities. In addition, in the coreference resolution task, due to the weak semantics of pronouns themselves, the error rate of pronoun resolution is high. Bert's pre-training method has a low probability of covering pronouns, and more external knowledge is required when resolving them. Therefore, it is necessary to strengthen the model's learning of pronouns.
[0003] Spanbert proposes a pre-training method for span-level tasks such as knowledge question answering and named entity recognition, which randomly masks arbitrary continuous spans. The span length follows the distribution of L~Geo(0.2).
[0004] Baidu's ERNIE model uses a three-stage masking mechanism for pre-training for Chinese, namely basic-level, phrase-level, and entity-level. This hierarchical progression of single words, phrases, and entity granularity embeds phrase and entity knowledge, greatly improving the representation ability of the language model. However, in actual use, it is found that this pre-training method of ERNIE will cause the forgetting of basic-level knowledge in the entity-level training stage, thereby reducing the model's word representation ability.
[0005] Therefore, technicians in this field are committed to developing a language model pre-training method based on coreference elimination. Summary of the invention
[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to provide a word vector with richer semantic information for the language model pre-training method of English coreference resolution, so as to improve the prediction accuracy of coreference resolution.
[0007] The process of human language learning is generally to learn basic words first, then learn phrases, and finally apply them to sentence and paragraph level tasks. However, since the knowledge of the neural network language model is stored in the form of network weights, if you train words first and then phrases, there may be a situation where word granularity information is forgotten. Therefore, the inventor proposes to adaptively train word blocks of different granularities according to the current function loss, and at the same time, to increase the training of pronouns in the language model to address the problem of low accuracy of pronoun resolution in coreference resolution. That is, in the training stage, the first 20% of the steps first use the word_learning mode to learn word information and train words, and the last 80% of the steps adaptively select the word_learning mode to train words or the phrase_learning mode to train phrases according to the loss, and the two modes use different loss functions.
[0008] In one embodiment of the present invention, a language model pre-training method based on coreference elimination is provided, including:
[0009] S100, data preprocessing, extracting pronouns in the corpus through string matching, and processing tools extracting named entities, noun phrases, etc. in the corpus as a set of masking candidates in the training data generation stage;
[0010] S200, training data generation, masking processing is performed through mask_word mode (i.e., word masking mode) and mask_phrase mode (i.e., phrase masking mode), and mask_word training data and mask_phrase training data are generated respectively:
[0011] S300, pre-training, select factor α according to the training modet Adaptively switch to the word_learning mode or the phrase_learning mode for training.
[0012] Optionally, in the co-reference resolution-based language model pre-training method in the above embodiments, step S100 includes:
[0013] S110. Obtain English Wikipedia data;
[0014] S120. Extract pronouns in the corpus and establish a pronoun set PronounSet;
[0015] S130. Extract all named entities in the corpus and establish an entity set EntitySet;
[0016] S140. Extract noun phrases in the corpus and remove the phrases that overlap with EntitySet to obtain a noun phrase set NounPhraseSet.
[0017] Further, in the co-reference resolution-based language model pre-training method in the above embodiments, the extraction tool in step S100 is the entity recognition module in the python natural language processing toolkit Spacy.
[0018] Further, in the co-reference resolution-based language model pre-training method in the above embodiments, the noun phrase extraction module in Spacy is used in step S140.
[0019] Optionally, in the co-reference resolution-based language model pre-training method in any of the above embodiments, step S200 includes:
[0020] S210. Duplicate the data twice and name them Data One and Data Two respectively;
[0021] S220. Create training instances for the texts in Data One and Data Two according to the training data generation method of BERT, and each instance includes multiple sentences;
[0022] S230. The instances created from Data One and Data Two are respectively subjected to masking processing in the mask_word mode (i.e., the word masking mode) and the mask_phrase mode (i.e., the phrase masking mode);
[0023] S240. Generate mask_word training data. Randomly select 15% of the words from the sentences in the instances created from Data One and put them into CandidateSet1 (the first mask candidate word set). Replace each word in CandidateSet1 with "[MASK]" with an 80% probability, replace it with other random words with a 10% probability, and keep it unchanged with a 10% probability;
[0024] S250. Generate mask_phrase training data. Randomly select named entities and noun phrases in the sentences of the instances created by Data 2 and add them to CandidateSet2 (the second masking candidate set). Replace each token in CandidateSet2 with "[MASK]" with an 80% probability, replace it with other random words with a 10% probability, and keep it unchanged with a 10% probability. All word replacement actions in each token should be consistent, that is, either replace all at the same time or keep them all unchanged.
[0025] Further, in the language model pre-training method based on coreference resolution in the above embodiment, step S220 includes limiting the sentence length to 128 English words. For sentences with less than 128 words, pad them, that is, supplement English words to 128. For sentences with more than 128 words, truncate them.
[0026] Further, in the language model pre-training method based on coreference resolution in the above embodiment, in step S240, pronouns in CandidateSet1 account for approximately one-third, and the remaining words account for two-thirds. When the total number of pronouns is less than one-third, replace them with general words.
[0027] Further, in the language model pre-training method based on coreference resolution in the above embodiment, in step S250, the named entities and nouns selected in CandidateSet2 account for 15% of the sentence length, and among them, named entities and noun phrases each account for 50%.
[0028] Optionally, in the language model pre-training method based on coreference resolution in any of the above embodiments, step S300 includes: In the word_learning mode, input the mask_word training data into the BERT network to predict the masked words and calculate the corresponding loss; in the phrase_learning mode, input the mask_phrase training data into the BERT network to predict the masked phrases and calculate the corresponding loss.
[0029] Further, in the language model pre-training method based on coreference resolution in any of the above embodiments, step S300 includes:
[0030] S310. Warm-up training. First, learn basic words. In the first 20% of the training steps, use the word_learning mode for warm-up training and save the initial word_learning prediction loss and the initial phrase_learning prediction loss
[0031] S302. Adaptive training, in the last 80% of the training steps, according to the selection factor α t Determine whether to use the word_learning or phrase_learning mode for training in the (t + 1)-th step, as follows:
[0032]
[0033] Respectively represent the losses of the two modes at the t-th training step. When α t > 0, use the word_learning mode in the (t + 1)-th step, otherwise continue training using the phrase_learning mode.
[0034] The present invention adds semantic training for pronouns, phrases, and entities, and adaptively switches the learning mode, enhancing the semantic representation ability of the model and better applying to the coreference resolution task.
[0035] The following will further illustrate the concept, specific structure, and technical effects of the present invention with reference to the accompanying drawings to fully understand the purpose, features, and effects of the present invention. Description of the Drawings
[0036] Figure 1 is a schematic flowchart of a pre-training method for a language model based on coreference resolution according to an exemplary embodiment. Detailed Embodiments
[0037] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0038] In the drawings, components with the same structure are denoted by the same numerical labels, and components with similar structures or functions are denoted by similar numerical labels. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. To make the drawings clearer, the thickness of some components is schematically exaggerated appropriately in the drawings.
[0039] The inventor designed a pre-training method for a language model based on coreference resolution, as Figure 1 shown, including the following steps:
[0040] S100. Data preprocessing, extract pronouns in the corpus through string matching, and use a processing tool to extract named entities, noun phrases, etc. in the corpus as the masking candidate set in the training data generation stage. The extraction tool is the entity recognition module in the python natural language processing toolkit Spacy; specifically including:
[0041] S110. Obtain English Wikipedia data;
[0042] S120. Extract pronouns in the corpus and establish a pronoun set PronounSet;
[0043] S130. Extract all named entities in the corpus and establish an entity set EntitySet;
[0044] S140. Use the noun phrase extraction module in Spacy to extract noun phrases in the corpus, and remove the phrases that overlap with EntitySet to obtain a noun phrase set NounPhraseSet.
[0045] S200. Training data generation, perform masking processing through the mask_word mode (i.e., word masking mode) and mask_phrase mode (i.e., phrase masking mode) to generate mask_word training data and mask_phrase training data respectively: specifically including:
[0046] S210. Copy the data twice, named data one and data two respectively;
[0047] S220. According to the training data generation method of BERT, create training instances for the texts in data one and data two. Each instance includes multiple sentences, and the sentence length is limited to 128 English words. For those less than 128, fill them up, that is, supplement English words to 128, and for those more than 128, perform truncation processing;
[0048] S230. The instances created from data one and data two are respectively processed by the mask_word mode (i.e., word masking mode) and mask_phrase mode (i.e., phrase masking mode);
[0049] S240. Generate mask_word training data. Randomly select 15% of the words in the sentences of the instances created from data one and put them into CandidateSet1 (masking candidate word set one), where pronouns account for about one-third, and the remaining words account for two-thirds. When the total number of pronouns is less than one-third, use general words to replace them; Replace each word in CandidateSet1 with "[MASK]" with an 80% probability, replace it with other random words with a 10% probability, and keep it unchanged with a 10% probability;
[0050] S250. Generate mask_phrase training data. Randomly select named entities and noun phrases in the sentences of the instances created by Data 2 and add them to CandidateSet2 (the second masked candidate set), where the selected named entities and nouns account for 15% of the sentence length, with named entities and noun phrases each accounting for 50%; replace each token block in CandidateSet2 with "[MASK]" with an 80% probability, replace it with other random words with a 10% probability, and keep it unchanged with a 10% probability. All token replacement actions within each token block should be consistent, that is, either replace all at the same time or keep them all unchanged.
[0051] S300. Pre-training. According to the training mode selection factor α t Adaptively switch to the word_learning mode or the phrase_learning mode for training, including: in the word_learning mode, input the mask_word training data into the BERT network to predict the masked words and calculate the corresponding loss; in the phrase_learning mode, input the mask_phrase training data into the BERT network to predict the masked phrases and calculate the corresponding loss; specifically including:
[0052] S310. Warm-up training. First, learn basic words. In the first 20% of the training steps, use the word_learning mode for warm-up training and save the initial word_learning prediction loss and the initial phrase_learning prediction loss
[0053] S302. Adaptive training. In the last 80% of the training steps, according to the selection factor α t Determine whether to use the word_learning or phrase_learning mode for training at the (t + 1)-th step, specifically as follows:
[0054]
[0055] respectively represent the losses of the two modes at the t-th training step; when α t > 0, use the word_learning mode at the (t + 1)-th step, otherwise continue training using the phrase_learning mode.
[0056] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in this technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A pre-training method for a language model based on coreference resolution, characterized in that It includes the following steps: S100. Data preprocessing: extract pronouns in the corpus through string matching, and use a processing tool to extract named entities and noun phrases in the corpus as the masking candidate set in the training data generation stage; S200. Training data generation: perform masking processing through the mask_word mode and the mask_phrase mode to generate mask_word training data and mask_phrase training data respectively, specifically including: S210. Duplicate the data twice, and name them Data One and Data Two respectively; S220. According to the training data generation method of BERT, create training instances for the texts in Data One and Data Two, and each instance includes multiple sentences; S230. The instances created from Data One and Data Two are respectively processed by the mask_word mode and the mask_phrase mode for masking; S240. Generate mask_word training data: randomly select 15% of the words in the sentences in the instances created from Data One and put them into CandidateSet1, and replace each word in CandidateSet1 with "[MASK]" with an 80% probability, replace it with other random words with a 10% probability, and keep it unchanged with a 10% probability; S250. Generate mask_phrase training data: randomly select named entities and noun phrases in the sentences in the instances created from Data Two and add them to CandidateSet2, and replace each chunk in CandidateSet2 with "[MASK]" with an 80% probability, replace it with other random words with a 10% probability, and keep it unchanged with a 10% probability. All word replacement behaviors in each chunk should be consistent, that is, either replace them all or keep them all unchanged; S300. Pre-training, select according to the training mode factor Adaptively switch to the word_learning mode or the phrase_learning mode for training, specifically including: S310. Preheat training. First, learn basic words. Use the word_learning mode for preheat training in the first 20% of the training steps, and save the initial word_learning prediction loss and the initial phrase_learning prediction loss ; S302. Adaptive training, for the last 80% of the training steps, according to the selected factor decide whether to use the word_learning or phrase_learning mode for training at the (t + 1)-th step, as follows: When the selection factor > 0, the word_learning mode is adopted in the (t + 1)-th step, otherwise the phrase_learning mode is adopted to continue the training.
2. The method for pre-training a language model based on coreference resolution according to claim 1, wherein The step S100 includes: S110. Obtain English Wikipedia data; S120. Extract pronouns in the corpus and establish a pronoun set PronounSet; S130. Extract all named entities in the corpus and establish an entity set EntitySet; S140. Extract noun phrases in the corpus and remove the phrases that overlap with EntitySet to obtain a noun phrase set NounPhraseSet.
3. The method for pre-training a language model based on coreference resolution according to claim 1, wherein The processing tool is the entity recognition module in the python natural language processing toolkit Spacy.
4. The method for pre-training a language model based on coreference resolution according to claim 2, wherein The step S140 uses the noun phrase extraction module in Spacy.
5. The method for pre-training a language model based on coreference resolution according to claim 1, wherein The processing method in the step S220 is: the sentence length is limited to 128 English words. If it is less than 128, it is padded; if it is more than 128, it is truncated.
6. The method for pre-training a language model based on coreference resolution according to claim 1, wherein, In the step S250, the number of named entities and nouns selected in CandidateSet2 accounts for 15% of the sentence length, and named entities and noun phrases each account for 50%.
7. The method for pre-training a language model based on coreference resolution according to claim 1, wherein The step S300 includes: in the word_learning mode, inputting the mask_word training data into the BERT network to predict the masked words and calculating the corresponding loss; in the phrase_learning mode, inputting the mask_phrase training data into the BERT network to predict the masked phrases and calculating the corresponding loss.
Citation Information
Patent Citations
Method and device for realizing anaphora resolution
CN111160006A
Anaphora resolution weak supervised learning method using language model
CN111428490A