A Word Representation Learning Method Based on Random Accessible Pointwise Mutual Information

By calculating attention weights based on point mutual information, the problem of unbalanced contribution of context words in Skip-gram and GloVe models is solved, and the quality of word representation and the performance of natural language processing tasks are improved.

CN115952807BActive Publication Date: 2025-07-11XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211623207.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-07-11
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

There are a large number of invalid training examples and the problem of ignoring the differences in semantic contribution of contextual words in the learning process. Especially in Skip-gram and GloVe models, the attention mechanism cannot be effectively applied.

Method used

The point mutual information method based on random access is adopted, and the point mutual information is calculated as attention weight by statistically stating the word co-occurrence matrix and calculating the point mutual information as attention weight. It is applied to the Skip-gram and GloVe models for word representation learning, and the training process is accelerated using the large-scale matrix random access method of the GloVe model.

Benefits of technology

Improve the quality of word representation and improve the performance of the model in word analogy and natural language processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115952807B_ABST
    Figure CN115952807B_ABST
Patent Text Reader

Abstract

A word representation learning method based on randomly accessible point mutual information, which relates to natural language processing. A. Prepare a large-scale unannotated text corpus; B. Scan the corpus and count word pairs to obtain a word co-occurrence matrix; C. Use a large-scale matrix random access method based on the GloVe model to achieve random access to the word co-occurrence matrix and obtain approximate values of the elements of the matrix; D. Calculate point mutual information using the approximate values of the elements of the word co-occurrence matrix obtained by random access; E. Calculate attention weights based on point mutual information and apply the attention weights to Skip-gram or GloVe model word representation learning to obtain the target word representation. A point mutual information attention weight operator is proposed, an attention mechanism suitable for Skip-gram and GloVe models is proposed, and a random access method is proposed for the problem that the co-occurrence matrix used in calculating point mutual information is too large to be fully loaded into memory. Higher-quality word representations are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to natural language processing, and more particularly to a word representation learning method based on randomly accessible point mutual information. Background Art

[0002] Word representations are extremely important in deep learning-based natural language processing systems because various natural language processing tasks, such as question answering systems, machine translation, text summarization, sentiment classification, named entity recognition, etc., require word representations as inputs, and the quality of word representations will directly affect the results of these tasks. To explore the internal relationships between words, Harris (Harris Z S. Distributional structure [J]. Word, 1954, 10(2-3): 146-162.) first proposed the Distributional Hypothesis, which states that words with similar contexts have similar semantics. Firth (Firth J R. A synopsis of linguistic theory, 1930-1955 [J]. Studies in Linguistic Analysis, 1957.) further elaborated on and explained Harris's Distributional Hypothesis, arguing that the semantic information of a word is determined by its context. After that, Hinton (Hinton GE. Learning distributed representations of concepts [C] / / Proceedings of the Eighth Annual Conference of the Cognitive Science Society. 1986, 1: 12.) proposed the idea of Distributed Representation, mapping all words in the vocabulary to a continuous, low-dimensional vector space, which is the so-called word representation.

[0003] Existing word representation methods usually use a fixed-size sliding window to traverse the corpus, select all words in the window except the center word as the context, and treat each word in the context equally. This strategy has the following drawbacks:

[0004] (1) First, not all words within the window necessarily contribute to the semantics of the central word. Especially in the Skip-gram model in Word2Vec (Mikolov T, Chen K, Corrado G, et al. Efficient estimation of word representations in vector space [J]. arXiv preprint arXiv:1301.3781, 2013. and Mikolov T, Sutskever I, Chen K, et al. Distributed representations of words and phrases and their compositionality [J]. Advances in Neural Information Processing Systems, 2013, 26.) and the GloVe (Pennington J, Socher R, Manning CD. Glove: Global vectors for word representation [C] / / Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014:1532 - 1543.) model, where word pairs composed of the central word and words within the window are used as training examples, there will be a large number of word pairs, and there is no dependency relationship between the two words. The dependency relationship between the parent and child nodes in the dependency tree is the main dependency relationship between words in a sentence. For a sentence containing \(n\) words, there are a total of \(n\) bidirectional word pair relationships (word pairs formed by parent and child nodes) in its dependency tree. For a sentence with \(n\) words, there are \(n - 1\) bidirectional dependency relationships in its dependency tree. Assuming that 10 words before and after the central word are used as context, then there will be 20 word pairs in a window. The ratio of the number of word pairs with dependency relationships in a sentence to the number of all word pairs used in training is approximately \(2*(n - 1) / (20*n)≈1 / 10\), which shows that most training examples are invalid and incorrect examples. (In fact, the number of word pairs in the windows at both ends of the sentence will be less than 20, and the dependency relationships in the dependency tree are not necessarily all included in the window, so this ratio is an estimated value.)

[0005] (2) Secondly, the semantic contributions of individual words in the context to the central word vary. Some words in the context have a stronger correlation with the central word, while other words have a weaker correlation with the central word. Therefore, all words in the context cannot be treated equally in the process of word representation learning.

[0006] Some studies have shown (Levy O, Goldberg Y. Dependency-based word embeddings[C] / / Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics(Volume 2:Short Papers).2014:302-308. and Komninos A, Manandhar S. Dependency based embeddings for sentence classification tasks[C] / / Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.2016:1490-1500.) that introducing dependency relationships into the word representation model and selecting the context according to the dependency relationships between words can effectively improve the model performance. However, this method depends to a certain extent on the quality of dependency parsing. If the accuracy of dependency parsing is low, it will directly affect the performance of the word representation model, and this method still does not quantify the semantic contributions of the selected context and ignores the correlation differences between individual words in the context and the central word.

[0007] Regarding the problem of the relevance difference between each word in the context and the central word, some research works use the attention mechanism in the word representation model to distinguish the importance of different contexts. Such models are all improvements on the CBOW model, assigning different attention weights to each word in the context, and then summing the weighted word representations to obtain the context vector. Ling et al. (Ling W, Tsvetkov Y, Amir S, et al. Not all contexts are created equal: Better word representations with variable attention[C] / / Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015:1367-1372.) assign different weights to the context according to the relative position to the central word in the improvement of the CBOW model. Liu et al. (Liu Q, Ling Z H, Jiang H, et al. Part-of-speech relevance weights for learning word embeddings[J]. arXiv preprint arXiv:1603.07695, 2016.) construct a weight matrix based on the part of speech of the word and assign different weights according to the part of speech collocation between the context and the central word. While Sonkar et al. (Sonkar S, Waters A E, Baraniuk R G. Attention word embedding[J]. arXiv preprint arXiv:2006.00988, 2020.) use dot product attention to calculate the corresponding attention weights of the context.

[0008] The above research work shows that by using the attention mechanism to focus on more important contexts, the performance of the word representation model can be improved to a certain extent. However, the above method only applies the attention mechanism to improving the calculation process of the context vector in the CBOW model. Since only word pairs are used in the Skip-gram model and the GloVe model, without multiple context words, the above method cannot apply the attention mechanism to the Skip-gram model and the GloVe model. To address the above deficiencies, a word representation learning method based on stochastically accessible pointwise mutual information (

[18] Church K, Hanks P. Word association norms, mutual information, and lexicography[J]. Computational Linguistics, 1990, 16(1):22-29. and Church K W, Gale W A, Hanks P, et al. Using Statistics in Lexical Analysis[M]. Lawrence Erlbaum Associates, 1991.) is proposed. For the first time, the attention mechanism is applied to the Skip-gram model and the GloVe model, and a method for stochastically accessing large-scale matrices is proposed to speed up the stochastic access of matrices during model training. Finally, a comparative experiment is conducted between the Skip-gram model and the GloVe model with the pointwise mutual information attention mechanism and the original models. The experimental results on the word analogy test set and the GLUE test set show that using the pointwise mutual information as the operator of the attention weight can effectively improve the quality of word representation. Summary of the Invention

[0009] The present invention discloses a word representation learning method based on stochastically accessible pointwise mutual information. First, a word co-occurrence matrix is statistically obtained from a large-scale unannotated text corpus, and then a large-scale matrix stochastic access method based on the GloVe model is used to achieve stochastic access to the word co-occurrence matrix, obtaining an approximate value of the elements of the matrix, and then the pointwise mutual information is calculated from the approximate value. Then, the attention weight is calculated based on the pointwise mutual information, and the attention weight is applied to the Skip-gram or GloVe model for word representation learning to obtain the target word representations, namely Skip-gram-PMI word representation and GloVe-PMI word representation. And a comparative experiment is conducted. The experimental results show that the Skip-gram and GloVe models with the pointwise mutual information attention mechanism obtain higher-quality word representations.

[0010] The present invention provides a word representation learning method based on stochastically accessible pointwise mutual information, including the following steps:

[0011] Step A. Prepare a large-scale unannotated text corpus;

[0012] Step B. Scan the corpus and count word pairs to obtain a word co-occurrence matrix;

[0013] Step C. Implement random access to the word co-occurrence matrix using a large-scale matrix random access method based on the GloVe model to obtain approximate values of the elements of the matrix;

[0014] Step D. Calculate pointwise mutual information using the approximate values of the elements of the word co-occurrence matrix obtained by random access;

[0015] Step E. Calculate attention weights based on pointwise mutual information and apply the attention weights to the Skip-gram or GloVe model for word representation learning to obtain the target word representation.

[0016] In a possible implementation, in Step C, the implementation of random access to the word co-occurrence matrix using a large-scale matrix random access method based on the GloVe model to obtain approximate values of the elements of the matrix specifically includes:

[0017] C1. Use the word co-occurrence matrix, the GloVe model, and Formula 1 to train to obtain word vectors and word vector biases;

[0018] The loss function for training the GloVe model is shown in Formula 1:

[0019]

[0020] where v i and b i represent the word vector and word vector bias of the i-th word, respectively represent the context word vector and context word vector bias of the j-th word, v i and b i , are all training parameters, Value is the matrix to be randomly accessed, this matrix is a non-negative square matrix, Value ij represents the value of the i-th row and j-th column of the matrix to be randomly accessed; Freq is the frequency matrix, Freq ij is the frequency of the element Value ij ; the matrix Value to be randomly accessed currently is the word co-occurrence matrix, and this word co-occurrence matrix is the frequency matrix Freq, and Freq ij = Value ij ;

[0021] C2. Calculate the approximate value of the co-occurrence frequency of word w i and word w j in the word co-occurrence matrix through Formula 2 or Formula 3.

[0022] The goal of model training is to minimize the loss function J, and the value of the function J is non - negative. During training, the value of the function J is made to tend towards 0, resulting in the following equation:

[0023]

[0024] When statistically calculating the word co - occurrence matrix in step B, if the order of words is not ignored, formula 2 is used to calculate Value ij , if the order of words is chosen to be ignored, the statistically obtained co - occurrence matrix will be symmetric. At this time, formula 3 can be used to calculate Value ij :

[0025]

[0026] When the parameters v i , b i , etc. are trained based on the GloVe model, formula 2 or formula 3 can be used to calculate Value ij , so as to achieve fast random access to the elements in the Value matrix. The Value matrix is too large to be loaded into memory, but these trained parameters can be fully loaded into memory.

[0027] Through the above - mentioned method, fast random access to a huge word co - occurrence matrix is achieved; wherein, the huge word co - occurrence matrix refers to a word co - occurrence matrix for which the memory required to load the complete matrix is greater than the memory of the computing device.

[0028] In a possible implementation, in step C, a large - scale matrix random access method based on the GloVe model is used to achieve random access to the word co - occurrence matrix, obtaining an approximate value of the elements of the matrix. It further includes:

[0029] The large - scale matrix random access method based on the GloVe model can also be applied to access any other large - scale matrix. The access method is as follows:

[0030] (1) First, add the same constant to all elements in the large - scale matrix to make it a non - negative matrix Value';

[0031] (2) Add zero elements to the non - negative matrix Value' to further expand it into a non - negative square matrix Value'';

[0032] (3) Use the non - negative square matrix Value'' as the matrix to be randomly accessed, and process it according to steps C1 and C2. Among them, for formula 1 in step C1, if the frequency matrix Freq cannot be obtained, Freq ij can all be set to 1.

[0033] In a possible implementation, in step D, calculating the point mutual information using the approximate values of the elements of the co-occurrence matrix obtained by random access includes:

[0034] Calculating the point mutual information using the approximate values of the elements of the co-occurrence matrix obtained by random access, and the calculation refers to the following formula:

[0035]

[0036] where PMI is the point mutual information between word w i and word w c , u i , b i represent the word vector and word vector bias of word w i , respectively represent the context word vector and context word vector bias of word w i , u c , b c represent the word vector and word vector bias of word w c , respectively represent the context word vector and context word vector bias of word w c . Freq i represents the number of occurrences of word w i in the corpus, Freq c represents the number of occurrences of word w c in the corpus;

[0037] When the corpus is determined, |C| is a constant, and the specific calculation formula is as follows:

[0038]

[0039] where Freq ic represents the co-occurrence frequency of word w i and word w c in the corpus.

[0040] In a possible implementation, in step E, calculating the attention weights based on the point mutual information and applying the attention weights to the Skip-gram or GloVe model for word representation learning to obtain the target word representation includes:

[0041] E1. Obtaining the point mutual information between each context word and the center word. At this time, the context word is word w i , and the center word is word w c ;

[0042] Among them, context words refer to all words within the left and right windows of the central word. For example, in the sentence "I want toeat an apple tomorrow", assuming the window size is 2, when "I" is the central word, the left window is empty and the right window is "want to"; when "want" is the central word, the left window is "I" and the right window is "to eat"; when "to" is the central word, the left window is "I want" and the right window is "eat an".

[0043] E2. Use the exponential normalization formula to transform the point mutual information into the attention weight att(w i ,w c );

[0044] Among them, exponential normalization is to perform exponential normalization on the point mutual information in step E1; the exponential normalization formula used in the present invention is as follows:

[0045]

[0046] Among them, att(w i ,w c ) is the attention weight of word w c to word w i in the current window, and window is the window size.

[0047] E3. Apply the attention weight att(w i ,w c ) to the Skip-gram model to obtain the target word representation, or apply the attention weight att(w i ,w c ) to the GloVe model to obtain the target word representation.

[0048] In a possible implementation manner, in step E3, the applying the attention weight att(w i ,w c ) to the Skip-gram model to obtain the target word representation includes:

[0049] Use the attention weight att(w i ,w c ) as the attention weight operator of the Skip-gram model, and use the coefficient 2×window×att(w i ,w c ) to dynamically adjust the change amount Δ of the word representation and the parameter vector update when performing error backpropagation. The adjusted change amounts of the word representation and the parameter vector update are Δ' = 2*window*att(w i ,w c) * Δ; where att(w i , w c ) represents the weight of word w i for word w c , and window represents the window size;

[0050] The change amount Δ' of the word representation and parameter vector update is such that, on the premise of keeping the total number of training samples (2 * window word pairs) within the window unchanged, it appropriately increases the frequency of the word pairs composed of the context with higher attention weights and the central word used for training, and at the same time reduces the frequency of other word pairs used for training, thereby improving the effectiveness of the training samples; where the training samples are the sentences in the text corpus.

[0051] After using the text corpus for model training and using Δ' as the change amount for word representation and parameter vector update during training, a Skip-gram-PMI word representation is obtained.

[0052] Through the above method, the attention mechanism based on point mutual information is applied to the Skip-gram model to obtain a higher-quality Skip-gram-PMI word representation.

[0053] In a possible implementation manner, in step E3, applying the attention weight att(w i , w c ) to the GloVe model includes:

[0054] Using the attention weight att(w i , w c ) as the attention weight operator and applying it to the word representation learning of the GloVe model; where, during the process of word representation learning, the following formula is used to update the frequency of the word pairs composed of the central word and each word in the context when statistically calculating the word co-occurrence matrix for word representation learning:

[0055]

[0056] where X ic is the value of the i-th row and c-th column in the current word co-occurrence matrix, window represents the window size for statistically calculating the word co-occurrence frequency, and att(w i , w c ) is the attention weight operator calculated through point mutual information.

[0057] Formula 7 can appropriately increase the co-occurrence frequency of word pairs with higher correlation and reduce the co-occurrence frequency of word pairs with lower correlation according to the correlation differences between each word in the context and the central word, thereby selectively extracting more accurate "word-word" co-occurrences from the text corpus and reducing noise.

[0058] Based on the text corpus, point mutual information is applied to the GloVe model through Formula 7, and model training is carried out, where the initial parameters used during training are randomly initialized. After the training is completed, GloVe-PMI word representations are obtained.

[0059] In the large-scale matrix random access method based on the GloVe model used in step C, the GloVe model is independent of the current GloVe model applying attention weights. When the word co-occurrence matrix used during training is small and can be directly loaded into memory, point mutual information PMI can be directly calculated without using the large-scale matrix random access method based on the GloVe model.

[0060] Through the above method, the attention mechanism based on point mutual information is applied to the GloVe model to obtain higher-quality GloVe-PMI word representations.

[0061] The present invention proposes a point mutual information attention weight operator, proposes an attention mechanism suitable for the Skip-gram and GloVe models, and at the same time proposes a random access method for the problem that the co-occurrence matrix used in calculating point mutual information is too large to be fully loaded into memory. Experiments show that the Skip-gram and GloVe models applying point mutual information attention obtain higher-quality word representations. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a flowchart of an embodiment of the present invention.

[0063] Figure 2 It is an implementation flowchart of the random access process. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following describes the method of the present invention in detail with reference to the drawings and embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the implementation manner and specific operation process are given, but the protection scope of the present invention is not limited to the following embodiments.

[0065] See Figure 1 , a flowchart of a word representation learning method based on randomly accessible point mutual information provided by an embodiment of the present application, specifically including:

[0066] Step A. Prepare a large-scale unlabeled text corpus.

[0067] Specifically, the text corpus contains a large number of sentences. For example, the text corpus can be a text in the txt file format, where each sentence is on a line and ends with a carriage return character "\n" at the end of the sentence.

[0068] For example, the corpus used in this experiment is a collection of the One Billion Word Benchmark, the UMBC Webbase Corpus, the News Crawl Corpus, and the English Wikipedia Corpus, with a size of 53 GB.

[0069] Step B. Scan the corpus and count word pairs to obtain a word co-occurrence matrix.

[0070] Specifically, scan the text corpus, select each word as the center word in turn, and determine the left and right windows of the center word; count the co-occurrence times of the center word and each context word within the window.

[0071] For example, for the sentence "I want to eat an apple tomorrow", assuming the window size is 2, when "I" is the center word, the left window is empty and the right window is "want to"; when "want" is the center word, the left window is "I" and the right window is "to eat"; when "to" is the center word, the left window is "I want" and the right window is "eat an". Among them, taking the window with "to" as the center word as an example, the co-occurrence times of the word pairs "to-I", "to-want", "to-eat", "to-an" will be incremented by 1.

[0072] Step C. Implement random access to the word co-occurrence matrix using a large-scale matrix random access method based on the GloVe model to obtain approximate values of the elements of the matrix.

[0073] Specifically, it includes:

[0074] C1. Use the word co-occurrence matrix, the GloVe model, and Formula 1 to train to obtain word vectors and word vector biases;

[0075] The loss function for training the GloVe model of this method is as follows:

[0076]

[0077] Among them, v i , b i represent the word vector and word vector bias of the i-th word, respectively represent the context word vector and context word vector bias of the j-th word, v i , b i , are all training parameters, Value is the matrix to be randomly accessed, and this matrix is a non-negative square matrix, Valueij Represents the value of the element in the \(i\)-th row and \(j\)-th column of the matrix to be randomly accessed. Freq is the frequency matrix, and Freq ij is the frequency of the element Value ij . Since the matrix Value to be randomly accessed currently is the word co-occurrence matrix, and this word co-occurrence matrix is the frequency matrix Freq, therefore, Freq is used in the training of the GloVe model in this method ij = Value ij .

[0078] C2. Calculate the approximate value of the co-occurrence frequency of the words \(w_i\) and \(w_j\) in the word co-occurrence matrix through Formula 2 or Formula 3.

[0079] Since the goal of model training is to minimize the loss function \(J\) and the value of the function \(J\) is non-negative, during training, the value of the function \(J\) will tend to 0, and the following equation can be obtained:

[0080]

[0081] When statistically calculating the word co-occurrence matrix in step B, if the order of words is not ignored, then Formula 2 is used to calculate Value ij . If the order of words is chosen to be ignored, then the statistically obtained co-occurrence matrix will be symmetric. At this time, Value can be calculated using the following Formula 3 ij :

[0082]

[0083] When the parameters \(v\) i , \(b\) i , etc. are trained based on the GloVe model, then Formula 2 or Formula 3 can be used to calculate Value ij , so as to achieve fast random access to the elements in the Value matrix. The Value matrix is too large to be loaded into memory, but these trained parameters can be fully loaded into memory.

[0084] It also includes:

[0085] The large-scale matrix random access method based on the GloVe model can also be applied to access any other large-scale matrix, and the access method is as follows:

[0086] (4) First, add the same constant to all elements in the large-scale matrix to make it a non-negative matrix Value';

[0087] (5) Let the non-negative matrix Value' add zero elements to further expand it into a non-negative square matrix Value'';

[0088] (6) Use the non - negative square matrix "Value" as the matrix to be randomly accessed and process it according to the aforementioned steps C1 and C2. Among them, for formula 1 in step C1, if the frequency matrix Freq cannot be obtained, Freq ij can be set to 1 for all.

[0089] Step D. Calculate the point - mutual information using the approximate values of the elements of the co - occurrence matrix obtained by random access.

[0090] Specifically, it includes:

[0091] Calculate the point - mutual information using the approximate values of the elements of the co - occurrence matrix obtained by random access. The calculation is as follows:

[0092]

[0093] where PMI is the point - mutual information between word w i and word w c ; u i , b i represent the word vector and word - vector bias of word w i ; represent the context word vector and context - word - vector bias of word w i respectively; u c , b c represent the word vector and word - vector bias of word w c ; represent the context word vector and context - word - vector bias of word w c respectively. Freq i represents the number of occurrences of word w i in the corpus, and Freq c represents the number of occurrences of word w c in the corpus.

[0094] When the corpus is determined, |C| is a constant, and the specific calculation formula is as follows:

[0095]

[0096] where Freq ic represents the co - occurrence frequency of word w i and word w c in the corpus.

[0097] Step E. Calculate the attention weights based on the point - mutual information and apply the attention weights to the Skip - gram or GloVe model for word representation learning to obtain the target word representation.

[0098] Specifically, it includes:

[0099] E1. Obtain the pointwise mutual information between each context word and the central word. At this time, the context word is word w i , and the central word is word w c .

[0100] Among them, the context word refers to all the words within the left and right windows of the central word. For example, in the sentence "I want to eat an apple tomorrow", assuming the window size is 2, when "I" is the central word, the left window is empty and the right window is "want to"; when "want" is the central word, the left window is "I" and the right window is "to eat"; when "to" is the central word, the left window is "I want" and the right window is "eat an".

[0101] E2. Use the exponential normalization formula to transform the pointwise mutual information into the attention weight att(w i , w c ).

[0102] Among them, the exponential normalization is to perform exponential normalization on the pointwise mutual information in step E1. The exponential normalization formula used is as follows:

[0103]

[0104] Among them, att(w i , w c ) is the attention weight of word w c to word w i in the current window, and window is the window size.

[0105] E3. Apply the attention weight att(w i , w c ) to the Skip-gram model to obtain the target word representation, or apply the attention weight att(w i , w c ) to the GloVe model to obtain the target word representation.

[0106] Among them, the original model structures of the Skip-gram model and the GloVe model cannot apply the traditional attention mechanism method. To address this issue, the present invention proposes a new attention mechanism, and proposes an attention weight operator based on pointwise mutual information, and applies the attention mechanism based on pointwise mutual information to the Skip-gram model and the GloVe model. And through experiments, it is proved that after applying the attention mechanism based on pointwise mutual information, the quality of the word representations obtained by the Skip-gram model and the GloVe model has been significantly improved.

[0107] The specific experimental situation is as follows:

[0108] (1) Comparative experiments were conducted on the Skip-gram model and GloVe model after applying the attention weight operator of point mutual information, and the original Skip-gram model and original GloVe model on the word analogy dataset. The experimental results are shown in Table 1.

[0109] Table 1

[0110]

[0111] The corpus used in the experiment is a collection of the One Billion Word Benchmark, UMBC Webbase Corpus, News Crawl Corpus, and English Wikipedia Corpus, with a size of 53 GB. The dimension of the experimental word representation is 300. The parameters used in the Skip-gram related experiments are: window size of 8, negative sampling count of 10, and number of iterations of 3. The parameters used in the GloVe related experiments are: window size of 8, number of iterations of 30, maximum co-occurrence frequency in the weight function of 100, and exponent of 0.75. The explanations of the keywords in the table are as follows: Skip-gram and GloVe are the models before improvement, and Skip-gram-PMI and GloVe-PMI are the models improved using the method of the present invention. Skip-gram-Dot-Product, Skip-gram-Scaled-Dot-Product, and Skip-gram-Cos respectively represent the models that use traditional methods such as dot product, scaled dot product, and cosine similarity to calculate the attention weight instead of point mutual information. Square root of Dimension, Dimension, Double Dimension respectively represent the processing of dividing the value after the dot product of the Skip-gram-Scaled-Dot-Product model by the square root of the dimension, dividing by the dimension, and dividing by twice the dimension. For the GloVe-PMI model, since it is impossible to use methods such as dot product, scaled dot product, and cosine similarity to calculate the attention weight corresponding to each word in the context when counting the co-occurrence matrix of the GloVe model, only the original GloVe model and GloVe-PMI model with the same hyperparameter settings are compared.

[0112] (2) Comparative experiments were conducted on the Skip-gram model after applying the attention weight operator of point mutual information and the original Skip-gram model on the GLUE test set. The experimental results are shown in Table 2.

[0113] Table 2

[0114]

[0115] The words used represent a 300-dimensional vector. The evaluation code of GLUE is from the official Github (https: / / github.com / nyu-mll / GLUE-baselines). The evaluation scores are obtained by submitting the predicted results to the official GLUE website (https: / / gluebenchmark.com / ) for evaluation by GLUE. Among them, the MNLI score is the average of the MNLI-matched and MNLI-mismatched test sets. The MRPC and QQP scores are accuracy and F1 scores. The STS-B score is the Pearson and Spearman correlation coefficients. The CoLA score is the Matthews correlation coefficient. The scores of other tasks are all accuracy, and all scores are scaled to a percentage system.

[0116] In step E3, the attention weight att(w i , w c ) is applied to the Skip-gram model to obtain the target word representation, specifically including:

[0117] Using the attention weight att(w i , w c ) as the attention weight operator of the Skip-gram model, and using the coefficient 2×window×att(w i , w c ) to dynamically adjust the change amount Δ of the word representation and the parameter vector update during backpropagation of errors. The adjusted change amounts of the word representation and the parameter vector update are Δ’ = 2*window*att(w i , w c )*Δ; where att(w i , w c ) represents the weight of word w i to word w c , and window represents the window size.

[0118] The adjusted change amount Δ’ of the word representation and the parameter vector update serves to appropriately increase the frequency of word pairs composed of the context with higher attention weights and the central word used for training while keeping the total number of training samples (2*window word pairs) within the window unchanged, and also reduce the frequency of other word pairs used for training, thereby improving the effectiveness of the training samples; where the training samples are sentences in the text corpus.

[0119] After using the text corpus for model training and using Δ’ as the change amount during training to update the word representation and parameter vectors, the Skip-gram-PMI word representation is obtained.

[0120] In step E3, the attention weight att(w i , w c ) is applied to the GloVe model to obtain the target word representation, specifically including:

[0121] Using the attention weight att(w i , w c ) as the attention weight operator and applying it to the word representation learning of the GloVe model; among them, during the process of the word representation learning, when counting the word co-occurrence matrix for word representation learning, the following formula is used to update the frequency of the word pairs composed of the center word and each word in the context:

[0122]

[0123] where X ic is the value of the i-th row and c-th column in the current word co-occurrence matrix, window represents the window size when counting the word co-occurrence frequency, and att(w i , w c ) is the attention weight operator calculated by the point mutual information.

[0124] The formula 7 can appropriately increase the co-occurrence frequency of word pairs with higher relevance and decrease the co-occurrence frequency of word pairs with lower relevance according to the correlation differences between each word in the context and the center word, so as to selectively extract more accurate "word-word" co-occurrences from the text corpus and reduce noise.

[0125] Based on the text corpus, apply the point mutual information to the GloVe model through formula 7 and perform model training, where the initial parameters used during training are randomly initialized. After training is completed, the GloVe-PMI word representation is obtained.

[0126] Note: The GloVe model used in the large-scale matrix random access method based on the GloVe model in step C is independent of the GloVe model to which the attention weight is currently applied. When the word co-occurrence matrix used during training is small and can be directly loaded into memory, the point mutual information PMI can be directly calculated without using the large-scale matrix random access method based on the GloVe model.

Claims

1. A word representation learning method based on randomly accessible point mutual information, characterized in that It includes the following steps: Step A. Prepare a large-scale unannotated text corpus; Step B. Scan the corpus and count word pairs to obtain a word co-occurrence matrix; Step C. Implement random access to the word co-occurrence matrix using a large-scale matrix random access method based on the GloVe model to obtain approximate values of the elements of the matrix, including: C1. Use the word co-occurrence matrix, the GloVe model, and Formula 1 to train word vectors and word vector biases; The loss function for training the GloVe model is as follows: Among them, v i , b i represent the word vector and word vector bias of the i-th word, respectively represent the context word vector and context word vector bias of the j-th word, are all training parameters, Value is the matrix to be randomly accessed, and this matrix is a non-negative square matrix. Value ij represents the value of the i-th row and j-th column of the matrix to be randomly accessed; Freq is the frequency matrix, and Freq ij is the frequency of the element Value ij ; because the matrix value to be randomly accessed currently is the word co-occurrence matrix, and this word co-occurrence matrix is the frequency matrix Freq, so Freq ij = Value ij ; C2. Calculate the approximate value of the co-occurrence frequency of word w i and word w j in the co-occurrence matrix by Formula 2 or Formula 3; Since the goal of model training is to minimize the loss function J and the value of function J is non-negative, during training, the value of function J will tend to 0, resulting in the following equation: When calculating the word co-occurrence matrix in step B, if the order of words is not ignored, formula 2 is used to calculate Value ij If the order of words is chosen to be ignored, the obtained co-occurrence matrix will be symmetric. In this case, formula 3 below is used to calculate Value ij : After training these parameters based on the GloVe model, use Equation 2 or Equation 3 to calculate Value ij so as to achieve fast random access to the elements in the Value matrix. The Value matrix is too large to be loaded into memory, but these trained parameters can be fully loaded into memory; Step D. Calculate point mutual information using the approximate values of the elements of the word co-occurrence matrix obtained by random access; Step E. Calculate attention weights based on point mutual information, apply the attention weights to the Skip-gram or GloVe model for word representation learning, and obtain the target word representation.

2. The word representation learning method based on randomly accessible point mutual information according to claim 1, characterized in that In Step C, implementing random access to the word co-occurrence matrix using a large-scale matrix random access method based on the GloVe model to obtain approximate values of the elements of the matrix further includes: The large-scale matrix random access method based on the GloVe model is also applied to access any other large-scale matrix, and the access method is as follows: (1) First, add the same constant to all elements in the large-scale matrix to make it a non-negative matrix value'; (2) Let the non-negative matrix value' add zero elements to be further extended into a non-negative square matrix value"; (3) Use the non - negative square matrix “value” as the matrix to be randomly accessed and process it according to the aforementioned steps C1 and C2; among them, for formula 1 in step C1, if the frequency matrix Freq cannot be obtained, then set Freq ij to 1 for all.

3. The word representation learning method based on randomly accessible point mutual information according to claim 1, characterized in that In Step D, calculating point mutual information using the approximate values of the elements of the word co-occurrence matrix obtained by random access, and the calculation refers to the following formula: Among them, PMI is the point mutual information between word w i and word w c . u i and b i represent the word vector and word vector bias of word w i respectively. u i and b c respectively represent the context word vector and context word vector bias of word w c . u c and b represent the word vector and word vector bias of word w c respectively. Freq i represents the number of occurrences of word w i in the corpus, and Freq c represents the number of occurrences of word w c in the corpus. When the corpus is determined, |C| is a constant, and the specific calculation formula is as follows: Among them, Freq ic represents the co-occurrence frequency of word w i and word w c in the corpus.

4. The word representation learning method based on randomly accessible point mutual information according to claim 1, characterized in that In Step E, calculating attention weights based on point mutual information, applying the attention weights to the Skip-gram or GloVe model for word representation learning, and obtaining the target word representation, including: E1. Obtain the pointwise mutual information between each context word and the central word. At this time, the context word is word w i , and the central word is word w c ; Among them, context words refer to all words within the left and right windows of the center word; E2. Use the exponential normalization formula to transform the point mutual information into the attention weight att(w i , w c ); Among them, exponential normalization is to perform exponential normalization on the point mutual information in Step E1; the exponential normalization formula used is as follows: Among them, att(w i , w c ) is the attention weight of word w c to word w i in the current window, where window is the window size; E3. Apply the attention weight att(w i , w c ) to the Skip-gram model to obtain the target word representation, or apply the attention weight att(w i , w c ) to the GloVe model to obtain the target word representation.

5. The word representation learning method based on randomly accessible point mutual information according to claim 4, wherein In step E3, apply the attention weight att(w i , w c ) to the Skip-gram model to obtain the target word representation, including: Use the attention weight att(w i , w c ) as the attention weight operator of the Skip-gram model. When performing error backpropagation, use the coefficient 2×window×att(w i , w c ) to dynamically adjust the variation Δ of the word representation and the parameter vector update. The adjusted variation of the word representation and the parameter vector update is Δ’ = 2 * window * att(w i , w c ) * Δ; where att(w i , w c ) represents the weight of word w i on word w c , and window represents the window size; The change amount Δ' of the word representation and parameter vector update is used to appropriately increase the frequency of word pairs composed of context with higher attention weights and the center word being used for training while reducing the frequency of other word pairs being used for training on the premise of keeping the total number of training samples in the window unchanged, thereby improving the effectiveness of training samples; among them, the training samples are sentences in the text corpus; After training the model by using the text corpus and using Δ' as the change amount for word representation and parameter vector update during training, the Skip-gram-PMI word representation is obtained.

6. The word representation learning method based on randomly accessible point mutual information according to claim 4, characterized in that In step E3, the attention weight att(w i , w c ) is applied to the GloVe model to obtain the target word representation, including: Use the attention weight att(w i , w c ) as the attention weight operator and apply it to the word representation learning of the GloVe model. During the word representation learning process, when statistically calculating the word co-occurrence matrix for word representation learning, the following formula is used to update the frequency of the word pairs composed of the central word and each context word: Among them, X ic is the value at the i-th row and c-th column in the current word co-occurrence matrix, window represents the window size when counting word co-occurrence frequencies, att(w i , w c ) is the attention weight operator calculated by the above point mutual information.

7. The word representation learning method based on randomly accessible point mutual information according to claim 6, characterized in that Formula 7 appropriately increases the co-occurrence frequency of word pairs with higher correlation and reduces the co-occurrence frequency of word pairs with lower correlation according to the correlation differences between each word in the context and the center word, thereby selectively extracting more accurate "word-word" co-occurrences from the text corpus and reducing noise.

8. The word representation learning method based on randomly accessible point mutual information according to claim 6, wherein Based on the text corpus, the point mutual information is applied to the GloVe model through Formula 7 for model training. The initial parameters used during training are randomly initialized. After training is completed, the GloVe-PMI word representation is obtained.

9. The word representation learning method based on randomly accessible point mutual information according to claim 1, characterized in that In step C, the GloVe model used in the large-scale matrix random access method based on the GloVe model is independent of the GloVe model currently applying attention weights; when the word co-occurrence matrix used during training is small and can be directly loaded into memory, the point mutual information PMI is directly calculated without using the large-scale matrix random access method based on the GloVe model.

Citation Information

Patent Citations

  • Word vector model based on point mutual information and text classification method based on CNN

    CN109189925A

  • Method and device for training word vector model

    CN110555209A