A Commonsense Question Answering Method Based on a Lexicon-Enhanced Pre-trained Model
By using dictionary resources and external hop attention mechanism to enhance the pre-training model, the problem of insufficient utilization of dictionary knowledge in the existing technology is solved, and performance improvement in common sense question and answer tasks is achieved.
Patent Information
- Application Number
- CN202210836783.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The existing pre-trained language models lack effective knowledge acquisition methods in common sense Q&A tasks, especially the insufficient utilization of dictionary knowledge, which leads to poor performance in knowledge-driven tasks.
Using the dictionary resources built by experts as external knowledge, the encoder model of the pre-trained model is enhanced through description-entity prediction and entity discrimination pre-training tasks, and combining the external hop attention mechanism and plug-in fine-tuning, a double tower encoder model is built to improve the effect of the model in common sense question and answer tasks.
The effect of the model in knowledge-driven common sense question-and-answer task is significantly improved. By combining the entity knowledge in dictionary knowledge, the model's language comprehension ability and question-and-answer performance are improved.
Smart Images

Figure CN115293142B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and specifically relates to the application of contrastive learning and dictionary-enhanced pre-training models in knowledge-driven question answering and natural language understanding. Background Art
[0002] Pre-trained language models (PLMs), such as BERT, RoBERTa, ALBERT, are popular in both academia and industry due to their state-of-the-art performance on various natural language processing (NLP) tasks. However, since they only capture general language representations learned from large-scale corpora, they are proven to be lacking in knowledge when dealing with knowledge-driven tasks. To address this challenge, many works, such as ERNIE-THU, KEPLER, KnowBERT, K-Adapter, and ERICA, aim to inject knowledge into PLMs for further improvement.
[0003] Common sense question answering is a typical application scenario of pre-trained language models. However, existing knowledge-enhanced PLMs still have some shortcomings. First, few methods focus on the knowledge itself, including what type of knowledge is needed and the feasibility of obtaining this knowledge. On the one hand, some models take for granted the use of knowledge graphs (KGs), which are difficult to obtain in practice and have been shown to be less effective than dictionary knowledge. On the other hand, many methods use Wikipedia, which is easier to obtain but often noisy and has low knowledge density. Second, current K-PLMs mainly focus on one or two types of knowledge-driven tasks. Although they have been shown to be useful on some specific tasks, their language understanding capabilities have either not been further verified on GLUE.
[0004] Therefore, in the field of common sense question answering, how to improve the effect and performance of PLMs is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The purpose of the present invention is to solve the problems existing in the prior art and provide a common sense question answering method based on a dictionary enhanced pre-training model.
[0006] Inspired by the fact that dictionary knowledge is more effective than structured knowledge, the present invention utilizes dictionary resources as external knowledge to improve the efficiency of PLMs. According to relevant experience, the advantages of doing so are as follows: First, it is consistent with human reading habits and cognitive processes; during the reading process, when encountering unfamiliar words, people usually consult dictionaries or encyclopedias. Second, compared with the long texts of Wikipedia, dictionary knowledge is more concise and has a high knowledge density. Third, dictionary knowledge is easier to obtain, which is of great significance for the practical application of K-PLMs. Even in the absence of a dictionary, it can be obtained by simply constructing a generator to summarize and explain the description of a word.
[0007] The specific technical solution adopted by the present invention is as follows:
[0008] A method for commonsense question answering based on a dictionary-enhanced pre-trained model, the steps of which are as follows:
[0009] S1: Obtain multiple dictionary knowledge as training corpora, and preprocess each corpus sample into the same input format; the content of each corpus sample includes a term and the definition description of the term. At the same time, each term also corresponds to a positive sample and a negative sample. The positive sample contains synonyms of the term and the definition descriptions of the synonyms, and the negative sample contains antonyms of the term and the definition descriptions of the antonyms;
[0010] S2: Use BERT or RoBERTa as the original encoder model, and use the training corpora to train the encoder model, update the parameters of the encoder model, and obtain a dictionary-enhanced encoder model; the specific training steps are as in S21~S22:
[0011] S21: Sample the training corpora, and perform masking processing on some of the sampled terms to cover the entity content of the terms, forming a first sample for predicting the term entity through the description, and the remaining sampled terms are directly used as the second sample;
[0012] S22: At the same time, perform iterative training on the encoder model through the description-entity prediction pre-training task and the entity discrimination pre-training task. The total loss of the training is the weighted sum of the losses of the two pre-training tasks;
[0013] In the description-entity prediction pre-training task, send the first sample obtained by sampling in S21 into the encoder model to obtain the corresponding hidden layer state, and then perform masking prediction through the pooling layer and the fully connected layer, and calculate the masking prediction loss as the loss of the description-entity prediction pre-training task;
[0014] In the entity discrimination pre-training task, the second sample obtained by sampling in S21 is used, combined with the corresponding positive and negative samples, for contrastive learning. The encoder model obtains the representations of the entries and definition descriptions corresponding to each sample, and calculates the contrastive learning loss as the loss of the entity discrimination pre-training task to narrow the representation distance of synonyms and separate the representation distances between antonyms;
[0015] S3: After completing the model training in S2, a dual-tower encoder model is formed by combining the dictionary-enhanced encoder model and the original encoder model, and a question-answering task output layer is connected after the dual-tower encoder model to obtain a question-answering model; wherein, the input of the dual-tower encoder model is the question text. The input question text passes through the original encoder model to obtain the first representation. At the same time, the input question text is matched based on the dictionary to identify all the entries in the question text. The identified entries pass through the dictionary-enhanced encoder model to obtain the second representation. The first representation and the second representation are fused and then input into the question-answering task output layer for answer prediction; the original encoder model and the question-answering task output layer in the question-answering model are fine-tuned based on the question-answering dataset;
[0016] S4. Based on the question-answering model after fine-tuning in S3, the answer to the question is predicted according to the input question.
[0017] Preferably, in the question-answering model, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c and the dictionary-enhanced encoder model encodes each identified entry respectively, and finally outputs the word embedding of each entry. The sum of the word embeddings of all entries is used as the second representation
[0018] Preferably, in the question-answering model, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c and the dictionary-enhanced encoder model encodes each identified entry respectively, and finally outputs the word embedding of each entry. The weighted sum of the word embeddings of all entries is calculated through the attention mechanism as the second representation :
[0019]
[0020] where: ATT represents the attention function, h c is used as the key (Key) and value (Value) of the attention function, e i is used as the query (Query) of the attention function, e iDenote the final output obtained by the encoder model enhanced by the dictionary for the \(i\)-th recognized entry or the entry and its definition description, and \(K\) is the total number of entries recognized from the question text.
[0021] Preferably, in the Q&A model, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation \(h\). c , and the dictionary-enhanced encoder model encodes each recognized entry respectively, extracts the outputs of each layer of the original encoder model and the dictionary-enhanced encoder model, and calculates the weighted sum of the word embeddings of all entries of the output of any \(l\)-th layer through the attention mechanism. Then average the weighted sums of word embeddings of all layers to obtain the second representation :
[0022]
[0023]
[0024] where \(h\) l represents the output of the question text input to the original encoder model at the \(l\)-th layer of the model, represents the output of the \(i\)-th recognized entry or the entry and its definition description input to the dictionary-enhanced encoder model at the \(l\)-th layer of the model; ATT represents the attention function, and \(h\) l serves as the key and value of the attention function, and \(e\) i serves as the query of the attention function; \(L\) represents the total number of layers in the original encoder model and the dictionary-enhanced encoder model, and \(K\) is the total number of entries recognized from the question text.
[0025] Preferably, in the Q&A model, through the obtained first representation \(h\) c and the second representation after concatenation, input them into the Q&A task output layer for answer prediction.
[0026] Preferably, in S1, each entry \(e\) and definition description desc in the corpus sample are preprocessed by adding [CLS] and [SEP] into the same input format \(s=\{[CLS]e[SEP]desc[SEP]\}\).
[0027] Preferably, in S22, the masked prediction loss \(L\) dep adopts the cross-entropy loss.
[0028] Preferably, in S22, the contrastive learning loss \(L\) edd is calculated as follows:
[0029]
[0030] Wherein: e represents the entry in the training corpus, and D represents the set of entries for training; The distribution represents the hidden state obtained by concatenating the entries and the definition descriptions of the entries in the corpus sample, positive sample, and negative sample and then feeding them into the encoder model.
[0031] Preferably, in S2, the calculation formula of the total loss function used for training the encoder model is:
[0032] L = λ1L dep + λ2L edd ;
[0033] Where λ1 and λ2 respectively represent the weight values of the loss functions of the two tasks.
[0034] Preferably, the output layer of the question-and-answer task is composed of a Linner layer and a Softmax layer.
[0035] Preferably, the original encoder model is preferably BERT-large.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] Compared with the prior art, the present invention can utilize the knowledge contained in the dictionary constructed by experts, and by using a task-specific output layer to model the characteristics of the common sense question-and-answer task, it can effectively improve the effect of the model in knowledge-driven common sense question-and-answer. Moreover, the present invention can further utilize the entity knowledge in the dictionary knowledge in the two-tower encoder model by combining the external jump attention mechanism and the plug-in fine-tuning means, effectively improving the effect of the pre-trained model in the common sense question-and-answer task. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the steps of a common sense question-and-answer method based on a dictionary-enhanced pre-trained model;
[0039] Figure 2 It is a pre-training flow chart of the method of the present invention;
[0040] Figure 3 It is three different fine-tuning frameworks of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described and explained below with reference to the accompanying drawings and specific embodiments.
[0042] As Figure 1As shown in the figure, in a preferred embodiment of the present invention, a common sense question answering method based on a dictionary-enhanced pre-trained model is provided, and its steps are as shown in S1 to S4:
[0043] S1: Obtain multiple dictionary knowledge as training corpora, and preprocess each corpus sample into the same input format; the content of each corpus sample includes a term and the definition description of the term. At the same time, each term also corresponds to a positive sample and a negative sample. The positive sample contains the synonyms of the term and the definition description of the synonyms, and the negative sample contains the antonyms of the term and the definition description of the antonyms.
[0044] As a preferred implementation manner of the embodiment of the present invention, the term e and the definition description desc in each corpus sample are preprocessed into the same input format s = {[CLS]e[SEP]desc[SEP]} by adding [CLS] and [SEP] tags.
[0045] Since there are actually three types of term entities in the present invention, namely the term Entry, its synonyms Syn, and antonyms Ant, the input formats for the term Entry-term description Desc and the synonyms Syn and antonyms Ant can be constructed: [CLS]Entry[SEP]Desc[SEP], [CLS]Syn[SEP]Desc[SEP], [CLS]Ant[SEP]Desc[SEP].
[0046] S2: Use BERT or RoBERTa as the original encoder model, and use the training corpora to train the encoder model, update the encoder model parameters, and obtain a dictionary-enhanced encoder model; the specific training steps are as follows: S21 to S22:
[0047] S21: Sample the training corpora, and perform masking processing on some of the sampled terms to cover the term entity content, forming a first sample for predicting the term entity through the description, and the remaining sampled terms are directly used as the second sample;
[0048] S22: At the same time, perform iterative training on the encoder model through the description-entity prediction pre-training task and the entity discrimination pre-training task, and the total loss of the training is the weighted sum of the losses of the two pre-training tasks;
[0049] In the description-entity prediction pre-training task, the first sample obtained by sampling in S21 is fed into the encoder model to obtain the corresponding hidden layer state, and then mask prediction is performed through the pooling layer and the fully connected layer, and the mask prediction loss is calculated as the loss of the description-entity prediction pre-training task;
[0050] In the entity discrimination pre-training task, the second sample obtained by sampling in S21 is used, combined with the corresponding positive and negative samples, for contrastive learning. The encoder model obtains the representations of the lemmas and definition descriptions corresponding to each sample, and calculates the contrastive learning loss as the loss of the entity discrimination pre-training task to narrow the distance between the representations of synonyms and separate the representations of antonyms.
[0051] As a preferred implementation manner of an embodiment of the present invention, the above-mentioned masked prediction loss L dep can adopt the cross-entropy loss. The above-mentioned contrastive learning loss L edd The calculation formula can adopt the following form:
[0052]
[0053] where: e represents the lemma in the training corpus, and D represents the set of trained lemmas; The distribution represents the hidden state obtained by concatenating the lemmas and the definition descriptions of the lemmas in the corpus sample, positive sample, and negative sample and then feeding them into the encoder model.
[0054] Thus, the calculation formula of the total loss function L used when training the encoder model can be expressed as:
[0055] L = λ1L dep + λ2L edd
[0056] where λ1 and λ2 respectively represent the weight values of the loss functions of the two tasks, and the specific weight values can be optimized and adjusted according to the actual situation.
[0057] As a preferred implementation manner of an embodiment of the present invention, the sampling data distribution during sampling in the above-mentioned predefined task preferably adopts a uniform distribution, that is, the corpus is uniformly sampled so that all lemmas are likely to be sampled.
[0058] The process of training the above-mentioned dictionary-enhanced encoder model is as Figure 2 shown.
[0059] S3: After completing the model training in S2, a dual-tower encoder model is formed by combining the dictionary-enhanced encoder model and the original encoder model, and a question-answering task output layer is connected after the dual-tower encoder model to obtain a question-answering model; wherein, the input of the dual-tower encoder model is the question text. The input question text passes through the original encoder model to obtain a first representation. At the same time, based on the dictionary, the input question text is matched to identify all the lemmas in the question text. The identified lemmas pass through the dictionary-enhanced encoder model to obtain a second representation. The first representation and the second representation are fused and then input into the question-answering task output layer for answer prediction; the original encoder model and the question-answering task output layer in the question-answering model are fine-tuned based on the question-answering data set.
[0060] S4. Based on the Q&A model after fine-tuning in S3, predict the answer to the question according to the input question.
[0061] It should be noted that the original encoder model in the present invention can be BERT or RoBERTa, and the preferred mode in the subsequent embodiments is BERT-large.
[0062] As a preferred implementation mode of the embodiment of the present invention, in the above Q&A model, different representation combination methods can be set for the first representation and the second representation output by the dual-tower encoder model, mainly including three types: (1) direct concatenation, (2) cross-attention mechanism, and (3) layer-aware cross-attention mechanism. As Figure 3 shown, the following will respectively describe the specific implementations of these three representation combination methods in detail:
[0063] (1) Direct concatenation:
[0064] In the Q&A model adopting this representation combination method, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c , and the dictionary-enhanced encoder model encodes each identified entry respectively, and finally outputs the word embedding of each entry. The sum of the word embeddings of all entries is used as the second representation
[0065] (2) Cross-attention mechanism:
[0066] In the Q&A model adopting this representation combination method, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c , and the dictionary-enhanced encoder model encodes each identified entry respectively, and finally outputs the word embedding of each entry. The weighted sum of the word embeddings of all entries is calculated through the attention mechanism as the second representation :
[0067]
[0068] Where: ATT represents the attention function, h c serves as the key and value of the attention function, e i serves as the query of the attention function, e i represents the i-th identified entry or the final output obtained by the dictionary-enhanced encoder model for the entry and its definition description, and K is the total number of entries identified from the question text.
[0069] (3) Layer-Aware Cross-Out Attention Mechanism:
[0070] In the question-answering model adopting this representation combination method, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c , and the lexicon-enhanced encoder model encodes each identified entry separately, extracts the outputs of each layer of the original encoder model and the lexicon-enhanced encoder model respectively, and calculates the weighted sum of the word embeddings of all entries of the output of any l-th layer through the attention mechanism Then, the weighted sums of the word embeddings of all layers are averaged to obtain the second representation :
[0071]
[0072]
[0073] where h l represents the output of the question text input to the original encoder model at the l-th layer of the model, represents the output of the i-th identified entry or the entry and its definition description input to the lexicon-enhanced encoder model at the l-th layer of the model; ATT represents the attention function, and h l serves as the key and value of the attention function, and e i serves as the query of the attention function; L represents the total number of layers in the original encoder model and the lexicon-enhanced encoder model, and K is the total number of entries identified from the question text.
[0074] It should be noted that in the above (2) cross-out attention mechanism and (3) layer-aware cross-out attention mechanism, the queries of the attention function output by the lexicon-enhanced encoder model have two forms, and the difference lies in the different inputs of the lexicon-enhanced encoder model. The first query form inputs the i-th identified entry, while the second query form inputs the i-th identified entry and its definition description. Therefore, in the above (2) cross-out attention mechanism, when the first query form is adopted, e i represents the final output of the i-th identified entry obtained through the lexicon-enhanced encoder model; when the second query form is adopted, e i represents the final output of the i-th identified entry and its definition description obtained through the lexicon-enhanced encoder model. In the above (3) layer-aware cross-out attention mechanism, when the first query form is adopted, represents the output of the i-th identified entry input to the lexicon-enhanced encoder model at the l-th layer of the model; when the second query form is adopted, It represents the output of the i-th recognized entry and its definition description after being input into the dictionary-enhanced encoder model at the l-th layer of the model.
[0075] In addition, as a preferred implementation manner of the embodiment of the present invention, in the above-mentioned question-and-answer model, the first representation h c and the second representation can be fused in a concatenated manner and then input into the question-and-answer task output layer for answer prediction. The question-and-answer task output layer can be composed of a Linner layer and a Softmax layer. The concatenated and fused representation first passes through the Linner layer, and the output of the Linner layer then passes through the Softmax layer to output the predicted probability distribution, thereby realizing the prediction of the answer.
[0076] Next, the common sense question-and-answer method based on the dictionary-enhanced pre-training model described in S1-S4 above will be applied to a specific example to demonstrate its specific implementation manner and technical effects.
[0077] Embodiment
[0078] A dictionary is a resource that lists the vocabulary of a language, clarifies its meaning through explanatory notes, and often explains its pronunciation, origin, usage, synonyms, antonyms, etc. In the present invention, the entries in the dictionary are the entries, and the explanations of the entries are the definition descriptions. Table 1 shows an example of the English word "forest". In the present invention, four types of information are used for pre-training: each entry, its definition description, synonyms, and antonyms, and knowledge injection pre-training is carried out using the entries in the dictionary and their meanings (i.e., explanatory descriptions). In addition, in order to improve the representativeness of the entries, the synonyms and antonyms of the entries are used for contrastive learning.
[0079] Table 1 Example of dictionary entries
[0080]
[0081] As Figure 1 shown, in this embodiment, according to the process described in S1-S4 above, two new pre-training tasks are used: (1) the dictionary entry prediction task and (2) the entry description discrimination task, that is, the aforementioned description-entity prediction pre-training task and entity discrimination pre-training task, to capture different aspects of dictionary knowledge by further training the pre-trained language model PLM (in this embodiment, BERT is used as the pre-trained encoder model), and then a question-and-answer model is constructed. The following specifically describes the implementation process of this embodiment:
[0082] For the prediction of lemmas, this embodiment follows the design of masked language modeling (MLM) in BERT, but imposes restrictions on the tokens to be masked. Initially, given an input sequence, the MLM task randomly masks a certain proportion of the input tokens with a special [MASK] symbol and then tries to recover them. Inspired by the work of Defsent, to effectively learn lemma representations, this embodiment takes each lemma e = {t1, t2, …, t i , …, t m} and its description desc = {w1, w2,....w n} as inputs, only masks the tokens of the lemma e in the selected input sample s = {[CLS]e[SEP]desc[SEP]}, and finally predicts the masked lemma tokens according to the corresponding description desc. It should be noted that if a lemma e consists of multiple tokens, all the constituent tokens will be masked. In the case of polysemy, a lemma e has multiple meanings (i.e., descriptions), and this embodiment constructs an input sample for each meaning in a similar way. This embodiment can formulate the lemma token prediction as:
[0083] P(t1, t2, …, t i , …, t m | s\{t1, t2, …, t i , …, t m})
[0084] where t i is the i-th symbol of e, and s\{t1, t2, …, t i , …, t m} represents the input symbols of sample s with the tokens y i…m masked. This embodiment initializes the encoder model with the pre-trained checkpoint of BERT-large and uses MLM as one of the optimization objectives, and it uses the cross-entropy loss as the loss function L dep .
[0085] To better capture the semantics of dictionary lemmas, this embodiment introduces lemma description discrimination and attempts to improve the robustness of lemma representations through contrastive learning. Specifically, this embodiment constructs positive (or negative) samples as follows: Given a lemma e and its description desc, this embodiment obtains its synonyms D s = {e syn} (or antonyms D a = {e ant}) from the dictionary source, and pairs each e syn (or e ant ) with its description desc syn (or desc ant) as a positive (or negative) sample. Taking the entry "Forest" in Table 1 as an example, "woodland" and "desert" are one of its synonyms and antonyms, respectively. The corresponding positive and negative samples are shown in Table 2. In the experiment of this embodiment, the same number (for example, 5) of positive and negative samples are used. Please note that currently in this embodiment, only the antonym of one entry is used to construct a strict negative sample, but in the future, it is also possible to explore the construction of negative samples by random selection.
[0086] Table 2 Examples of positive and negative samples
[0087] Positive [CLS]woodland[SEP]Land covered with wood or trees[SEP] Negative [CLS]desert[SEP]arid land with little or no vegetation[SEP]
[0088] In this embodiment, h ori ,h syn ,h ant To represent the original, positive and negative input samples. In order to get closer to h ori and h syn distance, push h ori and h ant , this embodiment designs a comparative target, where (e ori ,e syn ) is considered a positive pair, (e ori ,e ant ) is considered negative. This embodiment uses h c , represents the hidden state of the special symbol [CLS], to represent the representation of the input sample. Define a contrastive target L edd as follows:
[0089]
[0090] Where f(x,y) represents the exponential of the dot product between hidden states x and y. This embodiment adds the dictionary entry prediction task loss and the entry description discrimination task loss to finally obtain the overall loss function L:
[0091] L=λ1L dep +λ2L edd
[0092] Where L dep and L edd Denotes the loss function of the two tasks. In the experiment of this embodiment, λ1=0.4 and λ2=0.6 can be set.
[0093] Taking BERT-large as the original encoder model, the encoder model is trained using the training corpus to update the encoder model parameters. After training until convergence, a dictionary-enhanced encoder model can be obtained, which is named DictBERT in this embodiment. The specific training steps are as described in the foregoing S21 - S22 and will not be repeated here.
[0094] In this embodiment, DictBERT is used as a plugin and a PLM with fixed parameters is used during fine-tuning. In this way, this embodiment can enjoy the flexibility of training different DictBERTs for different dictionaries and avoid the catastrophic forgetting problem of continuous training. Specifically, this embodiment first identifies the dictionary entries from a given input, then uses DictBERT as a KB to retrieve the corresponding entry information (i.e., entry embeddings), and finally injects the retrieved entry information into the original input to obtain an enhanced representation for downstream tasks. In the case where the input consists of multiple sequences (e.g., NLI), this embodiment processes each input sequence separately and then inputs them into the downstream specific question-answering task layer for subsequent processing.
[0095] Specifically, when performing the question-answering task, a dual-tower encoder model can be formed by combining the dictionary-enhanced encoder model DictBERT and the original encoder model BERT-large, and a question-answering task output layer is connected after the dual-tower encoder model to obtain a question-answering model. The input of the dual-tower encoder model is the question text. The input question text passes through the original encoder model to obtain a first representation. At the same time, based on the dictionary, the input question text is matched to identify all the dictionary entries in the question text. The identified dictionary entries pass through the dictionary-enhanced encoder model to obtain a second representation. The first representation and the second representation are fused and then input into the question-answering task output layer for answer prediction. The question-answering task output layer can be composed of a Linner layer and a Softmax layer. The concatenated and fused representation first passes through the Linner layer, and the output of the Linner layer then passes through the Softmax layer to output the predicted probability distribution, thereby realizing the prediction of the answer. This question-answering model needs to be trained. A question-answering dataset with annotations can be used to fine-tune the original encoder model and the question-answering task output layer in the question-answering model. After fine-tuning, it can be used for commonsense question-answering.
[0096] To better utilize the implicit knowledge retrieved in downstream tasks, this embodiment introduces three different knowledge infusion mechanisms in the question-answering model (see Figure 3 ): (1) direct concatenation, (2) out-of-hop attention mechanism, and (3) layer-aware out-of-hop attention mechanism.
[0097] As Figure 3 shown, this embodiment directly concatenates the set output of BERT (i.e., hc ) and the sum of the retrieved entry embeddings from DictBERT (i.e., ) are concatenated. Then, this concatenation (i.e., [h c ; ) is fed into the task-specific layer of the downstream task.
[0098] The simplest way to incorporate the identified entries into the original text is to sum their embeddings and concatenate the result with the text representation. However, this method cannot determine which entry is more important and which sense is more appropriate in the case of polysemous entries.
[0099] Therefore, this embodiment further proposes a skip-out attention mechanism to address this defect. As Figure 3 shown, following Transformer-XH, the hidden state h c of the [CLS] token in the input query is used as the "attention center" to attend to each identified entry in the same input. With the attention weights, when integrating these entries or senses as external knowledge into the original input query, more important entries or senses will be attended to. The formula for the skip-out attention mechanism is as follows:
[0100]
[0101] where e i represents the DictBERT output of the i-th identified entry. K is the number of identified entries in the input query, represents the weighted sum of the retrieved entry embeddings. After obtaining , [hc; can be used for the final inference.
[0102] To further improve the performance, this embodiment extends the skip-out attention of the last layer to each inner layer to make it hierarchical. As Figure 3 shown, the attention scores of each layer are calculated, and finally their average value is used to judge the implicit input knowledge. Specifically, the inter-layer skip-out attention can be expressed as:
[0103]
[0104] where, represents the weighted sum of the output of the l-th layer of DictBERT.
[0105] Next, the above method is applied to a specific dataset. The specific implementation steps are as described above, and mainly its effects are shown below.
[0106] This embodiment uses knowledge-driven question answering such as CommonsenseQA and OpenBookQA to evaluate the performance of DictBERT on this task.
[0107] In this embodiment, different variants of DictBERT were evaluated in experiments. DictBERT+Concat(K) uses a concatenation mechanism, DictBERT+EHA(K) and DictBERT+EHA(K+V) adopt the extra-hop attention mechanism, while Dict-BERT+LWA(K+V) uses the inter-layer attention mechanism. The symbol K represents retrieving entry embeddings from DictBERT using lemmas, i.e., adopting the first query form described above, and K+V represents performing knowledge retrieval using both lemmas and their corresponding definition descriptions, i.e., adopting the second query form described above.
[0108] Table 3. Experimental results of CommonsenseQA and OpenbookQA
[0109]
[0110] The performance of DictBERT on knowledge-driven QA tasks, namely CommonsenseQA and OpenBookQA, is shown in Table 4. Compared with BERT-large, the basic setting of this embodiment, DictBERT+Concat, achieved significant improvements of 6.0% and 4.0% on these two tasks respectively. In addition, this embodiment observed significant increases (2.4% and 1.9%) brought by the extra-hop attention mechanism, verifying again the importance of identifying the sensitive weights of entries in the input samples. Finally, DictBERT+LWA(K+V) achieved the best results on both tasks, obtaining final gains of 9.0% and 7.1% compared with the BERT-large baseline. For more convincing results, this embodiment also compared DictRoBERTa with the original RoBERTa-large on CommonsenseQA and OpenBookQA. As shown in Table 4, this conclusion also holds for RoBERTa. Similarly, DictRoBERTa+LWA(K+V) achieved the best results, ultimately improving by more than 6.4% and 6.5% respectively.
[0111] Table 4. Ablation experiment results
[0112]
[0113] In addition, this embodiment conducts ablation studies on different components of DictBERT. First, this embodiment evaluates BERT-large+Concat(K) and BERT-large+LWA(K+V), which directly use BERT-large instead of the pre-trained Dict-BERT as a plug-in. As can be seen from the results, the improvement is quite limited, confirming the necessity of injecting external knowledge. Second, this embodiment evaluates the effectiveness of each of the two training tasks. DictBERT(DEP)+Concat and DictBERT(DEP+EDD)+Concat. Contrastive learning is helpful to a certain extent (0.4% on average), while only masking token markers is better than masking both token and description markers (+0.3% for all three). Finally, this embodiment examines the necessity of using DictBERT as a plug-in KB instead of directly fine-tuning it for downstream tasks (only DictBERT), and whether the size of the dictionary matters (DictBERT plus). The three knowledge infusion mechanisms of this embodiment can further improve the performance of pure DictBERT, indicating the benefits of using DictBERT as a plug-in. To evaluate the impact of the dictionary size, this embodiment uses a combination of the Cambridge Dictionary, Oxford Dictionary, and Wikipedia Dictionary, with a total number of entries exceeding 1 million. The results show that DictBERT plus+LWA(K+V) can further improve the performance of the three task sets (+0.23% on average).
[0114] This embodiment proposes DictBERT, which enhances the PLM with dictionary knowledge through two novel pre-training tasks and an attention-based knowledge infusion mechanism during fine-tuning. At the same time, a set of sufficient experiments are conducted to prove its effectiveness in common sense question answering tasks. Importantly, the method of the present invention can be easily applied in practice. Moreover, the present invention can further explore more effective pre-training tasks and knowledge infusion mechanisms and apply this method to more knowledge-driven tasks.
[0115] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A common sense question answering method based on a lexicon-enhanced pre-trained model, characterized in that, Here are the steps: S1: Obtain multiple dictionary knowledge as training corpus, and preprocess each corpus sample into the same input format; the content of each corpus sample includes a term and a definition description of the term, and each term also corresponds to a positive sample and a negative sample. The positive sample contains the synonyms of the term and the definition description of the synonym, and the negative sample contains the antonym of the term and the definition description of the antonym; S2: Use BERT or RoBERTa as the original encoder model, train the encoder model using the training corpus, update the encoder model parameters, and obtain a dictionary-enhanced encoder model; the specific training steps are as follows: S21-S22: S21: sampling the training corpus, and performing masking processing on some of the sampled terms to cover the term entity content, thereby forming a first sample for predicting the term entity through description, and the rest of the sampled terms are directly used as the second sample; S22: iteratively training the encoder model through the description-entity prediction pre-training task and the entity discrimination pre-training task at the same time, wherein the total loss of the training is the weighted sum of the losses of the two pre-training tasks; In the description-entity prediction pre-training task, the first sample obtained by sampling in S21 is sent to the encoder model to obtain the corresponding hidden layer state, and then the mask prediction is performed through the pooling layer and the fully connected layer, and the mask prediction loss is calculated as the loss of the description-entity prediction pre-training task; In the entity discrimination pre-training task, the second sample obtained by sampling in S21 is used in combination with the corresponding positive and negative samples for contrastive learning. The encoder model obtains the representation of the entry and definition description corresponding to each sample, and the contrastive learning loss is calculated as the loss of the entity discrimination pre-training task to shorten the representation distance of synonyms and separate the representation distance between antonyms. S3: After completing the model training in S2, the dictionary-enhanced encoder model and the original encoder model are combined to form a dual-tower encoder model, and the question-answering task output layer is connected after the dual-tower encoder model to obtain a question-answering model; wherein, the input of the dual-tower encoder model is the question text, the input question text is passed through the original encoder model to obtain a first representation, and at the same time, the input question text is matched based on the dictionary to identify all the terms in the question text, and the identified terms are passed through the dictionary-enhanced encoder model to obtain a second representation, and the first representation and the second representation are fused and input into the question-answering task output layer for answer prediction; the original encoder model and the question-answering task output layer in the question-answering model are fine-tuned based on the question-answering dataset; S4. Based on the question-answering model fine-tuned in S3, the answer to the question is predicted according to the input question.
2. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 1, wherein, In the described question-and-answer model, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c , and the dictionary-enhanced encoder model encodes each identified entry separately, and finally outputs the word embedding of each entry. The sum of the word embeddings of all entries is used as the second representation 3. The common sense question answering method based on the dictionary enhanced pre-trained model according to claim 1, wherein In the question-and-answer model, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c , and the dictionary-enhanced encoder model encodes each identified entry respectively, and finally outputs the word embedding of each entry. The weighted sum of the word embeddings of all entries is calculated through the attention mechanism as the second representation where: ATT represents the attention function, h c as the key (Key) and value (Value) of the attention function, e i as the query (Query) of the attention function, e i represents the i-th recognized entry or the final output obtained by the encoder model that enhances the entry and its definition description through the dictionary, and K is the total number of entries recognized from the question text.
4. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 1, wherein In the Q&A model, the original encoder model encodes the input question text and finally outputs the hidden state of the [CLS] token as the first representation h c , while the lexicon-enhanced encoder model encodes each recognized entry respectively, extracts the outputs of each layer of the original encoder model and the lexicon-enhanced encoder model, and calculates the weighted sum of the word embeddings of all entries of the output of any l-th layer through the attention mechanism Then, the weighted sums of the word embeddings of all layers are averaged to obtain the second representation Among them, h l represents the output of the original encoder model after the problem text input at the l-th layer of the model, represents the output of the i-th recognized term or the term and its definition description input into the dictionary-enhanced encoder model at the l-th layer of the model; ATT represents the attention function, and h l serves as the key and value of the attention function, and e i serves as the query of the attention function; L represents the total number of layers in the original encoder model and the dictionary-enhanced encoder model, and K is the total number of terms recognized from the problem text.
5. The common sense question answering method based on a lexicon-enhanced pre-trained model according to claim 1, wherein In the Q&A model, the first representation h obtained c and the second representation are concatenated and then input into the output layer of the Q&A task for answer prediction.
6. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 1, wherein In S1, the term e and definition description desc in each corpus sample are preprocessed into the same input format s={[CLS]e[SEP]desc[SEP]} by adding [CLS] and [SEP].
7. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 1, wherein, In the above S22, the mask prediction loss L dep uses cross-entropy loss.
8. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 7, wherein, In S22, the contrastive learning loss L edd The calculation formula is as follows: Where: e represents the entries in the training corpus, and D represents the set of trained entries; They respectively represent the hidden states obtained by concatenating the entries and the definition descriptions of the entries in the corpus sample, positive sample, and negative sample and then feeding them into the encoder model.
9. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 8, wherein In S2, the calculation formula of the total loss function used when training the encoder model is: L = λ1L dep + λ2L edd ; Where λ1 and λ2 represent the weight values of the loss functions of the two tasks respectively.
10. The common sense question answering method based on the lexicon-enhanced pre-trained model according to claim 1, characterized in that, The Q&A task output layer consists of a Linner layer and a Softmax layer.