Word sense disambiguation with convolutional block attention module embedded in Regenerative network
By embedding CBAM into the Regnety network in the field of natural language processing, the problems of traditional algorithms' feature limitations and poor classifier training effect during vocabulary ambiguity are solved, and high-quality feature extraction and correct classification of semantic categories are achieved.
Patent Information
- Application Number
- CN202210720884.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-23
AI Technical Summary
In the field of natural language processing, traditional algorithms have problems such as feature limitations and poor classifier training results when disambiguating vocabulary ambiguity.
The word meaning disambiguation method based on CBAM embedding Regnety network is adopted, and the adjacent vocabulary units of ambiguity are extracted and vectorized, and the semantic classification is used using the CBAMRegnety network.
High-quality disambiguation characteristics are effectively extracted, the accuracy of semantic categories is improved, the amount of data and parameters is reduced, and overfitting is prevented.
Smart Images

Figure QLYQS_1 
Figure QLYQS_2 
Figure QLYQS_3
Abstract
Description
Technical field:
[0001] The present invention relates to a word sense disambiguation method in which a convolutional block attention module is embedded in a Regenerative network, and the method has good application in the technical field of natural language processing. Background technology:
[0002] In the field of natural language processing, words generally have polysemy. The purpose of word sense disambiguation is to determine the semantics of ambiguous words in a specific context. Word sense disambiguation has important applications in machine translation, automatic summarization, information retrieval, and text classification. The performance of these systems is closely related to word sense disambiguation.
[0003] Some common algorithms are often used to disambiguate vocabulary, such as k-means, naive Bayes, classification methods based on association rules, and artificial neural networks. However, traditional algorithms have some shortcomings and deficiencies. The extracted disambiguation features are limited to local areas, and the training effect of the classifier is not very good. In recent years, deep learning algorithms have been widely used in the field of natural language processing. Embedding CBAM into Regenety to process features to obtain more accurate disambiguation features solves the problem of manually extracting disambiguation features. In Regenety, the weights of neurons are shared. This allows neurons to share resources, reduces the complexity of the network model, and prevents overfitting. Embedding CBAM into Regenety can effectively disambiguate ambiguous vocabulary and achieve correct semantic classification. Summary of the invention:
[0004] In order to solve the lexical ambiguity problem in the field of natural language processing, the present invention discloses a word sense disambiguation method based on CBAM embedded in Regenerative network.
[0005] To this end, the present invention provides the following technical solutions:
[0006] 1. Word sense disambiguation method based on CBAM embedded in Regenerative network, ambiguous word m has n semantic categories s1, s2, …, s n , the method comprises the following steps:
[0007] Step 1: Perform word segmentation, part-of-speech tagging, pinyin first letter tagging, tone tagging and semantic class tagging on the training and test corpora of SemEval-2007:Task#5, and select the word form, part-of-speech, semantic class, pinyin first letter and tone of adjacent lexical units with noun, verb, adjective, numeral, quantifier and pronoun parts of speech around the ambiguous word m as disambiguation features.
[0008] Step 2: Use the Word2Vec tool to vectorize the disambiguation features extracted from the training corpus of SemEval-2007:Task#5 to obtain training data, and use the Word2Vec tool to vectorize the disambiguation features extracted from the test corpus of SemEval-2007:Task#5 to obtain test data.
[0009] Step 3: The training includes two processes: forward propagation and back propagation. The training data is used to optimize CBAMRegnety to obtain the optimized CBAMRegnety.
[0010] Step 4: The test process is a forward propagation process, that is, a semantic classification process. On the optimized CBAMRegnety, input the test data and calculate the weight of the ambiguous word m in each semantic category. The semantic category with the largest weight is the semantic category of the ambiguous word.
[0011] 2. The word sense disambiguation method based on CBAM embedded in Regenty network according to claim 1 is characterized in that in step 1, the specific steps are:
[0012] Step 1-1 uses a Chinese word segmentation tool to perform vocabulary segmentation on Chinese sentences;
[0013] Step 1-2 uses a Chinese part-of-speech tagging tool to tag the parts of speech of all words in the sentence;
[0014] Step 1-3 uses Chinese semantic annotation tools to annotate the semantic categories of all words in the sentence;
[0015] Step 1-4: Use the Chinese character to pinyin tool to mark the pinyin first letter and tone of all words in the sentence;
[0016] Step 1-5 selects adjacent lexical units with noun, verb, adjective, numeral, quantifier and pronoun parts of speech on the left and right of the ambiguous word;
[0017] Steps 1-6 use the word form, part of speech, semantic class, pinyin initial letter, and tone of the selected adjacent lexical units as disambiguation features.
[0018] 3. The word sense disambiguation method based on CBAM embedded in Regenty network according to claim 1 is characterized in that in the step 2, the specific steps are:
[0019] Step 2-1 uses the Word2Vec tool to vectorize the disambiguation features extracted from the training corpus of SemEval-2007:Task#5 to obtain training data;
[0020] Step 2-2 uses the Word2Vec tool to vectorize the disambiguation features extracted from the test corpus of SemEval-2007:Task#5 to obtain test data.
[0021] 4. The word sense disambiguation method based on CBAM embedded in Regenty network according to claim 1, characterized in that in step 3, the specific steps are:
[0022] Step 3-1 inputs the training data into the initialized CBAMRegnety;
[0023] Step 3-2 passes through convolution layer 1 and extracts feature X1;
[0024] Step 3-3 extracts feature X2 through a channel attention convolution layer, wherein the channel attention convolution layer includes a channel attention module SE and a convolution layer 2;
[0025] Step 3-4 passes through the convolutional block attention convolutional layer to extract feature X3. The convolutional block attention convolutional layer includes a convolutional block attention module CBAM and a convolutional layer 3. The output calculation process of CBAM is as follows:
[0026]
[0027]
[0028] Among them, σ is the sigmoid function, MLP is a two-layer neural network, AvgPool(X2) means the average pooling operation on X2, MaxPool(X2) means the maximum pooling operation on X2, Mc*X2 is the multiplication operation of Mc and the input feature X2, and the output of CBAM is Mc*(Mc*X2);
[0029] Step 3-5 passes through the adaptive average pooling layer to calculate the ambiguous word m in the semantic category s i The weight w(s i |m), i=1,2,...,n;
[0030] Step 3-6 uses the cross entropy loss function to calculate the error loss. The calculation process of the error loss is as follows:
[0031]
[0032] Among them, loss represents the average error of training data, n is the number of training data, and y k is the label of the kth training data;
[0033] Steps 3-7 are performed based on the error loss back propagation, and the parameters are updated layer by layer. The parameter update process is as follows:
[0034]
[0035] Among them, θ represents the parameter set, θ' represents the updated parameter set, and a is the learning rate;
[0036] Step 3-8 continuously iterates step 3-1 to step 3-7 until the specified number of cycles is reached to obtain the optimized CBAMRegnety.
[0037] 5. The word sense disambiguation method based on CBAM embedded in Regenty network according to claim 1 is characterized in that in step 4, the specific process is:
[0038] Step 4-1 Input the test data into the optimized CBAMRegnety;
[0039] Step 4-2 passes through convolution layer 1 and extracts feature X1;
[0040] Step 4-3 extracts feature X2 through a channel attention convolution layer, wherein the channel attention convolution layer includes a channel attention module SE and a convolution layer 2;
[0041] Step 4-4 passes through the convolutional block attention convolutional layer to extract feature X3. The convolutional block attention convolutional layer includes a convolutional block attention module CBAM and a convolutional layer 3. The output calculation process of CBAM is as follows:
[0042]
[0043]
[0044] Among them, σ is the sigmoid function, MLP is a two-layer neural network, AvgPool(X2) means the average pooling operation on X2, MaxPool(X2) means the maximum pooling operation on X2, Mc*X2 is the multiplication operation of Mc and the input feature X2, and the output of CBAM is Mc*(Mc*X2);
[0045] Step 4-5 passes through the adaptive average pooling layer to calculate the ambiguous word m in the semantic category s i The weight w(s i |m), i=1,2,...,n;
[0046] Step 4-6 selects the semantic category corresponding to the maximum weight as the semantic category of the ambiguous word m:
[0047]
[0048] Among them, s represents the semantic category of the ambiguous word m.
[0049] Beneficial effects:
[0050] 1. The present invention is a word sense disambiguation method based on CBAM embedded in Regnetic network. The word form, part of speech, semantic class, pinyin initial letter and tone of adjacent lexical units with noun, verb, adjective, numeral, quantifier and pronoun parts of speech on the left and right of the ambiguous word are selected as disambiguation features. The disambiguation features are vectorized by using Word2Vec tool, and the extracted disambiguation features have high quality.
[0051] 2. The disambiguation model used in the present invention is CBAM embedded in the Regenety network. Its biggest features are local perception and parameter sharing, which covers more features to be identified. The accuracy of the final semantic category discrimination is also higher. It can handle high-dimensional data well without manually selecting data features. CBAM allows Regenety to learn to focus on key information. It enhances effective features and suppresses invalid features, and can extract more complete disambiguation features, reduce the amount of data and parameters, and prevent over-simulation.
[0052] 3. When training the disambiguation model, the stochastic gradient descent method is used to update the parameters. The error is calculated and the error is returned along the original route through back propagation, that is, from the output layer through each intermediate hidden layer, and the parameters of each layer are updated layer by layer, and finally return to the input layer. Forward propagation and back propagation are continuously performed to reduce the error and update the model parameters until the disambiguation model is trained. As the error back propagation continuously updates the parameters, the CBAM embedded in the Regenerative Network can accurately disambiguate the input data. Description of the drawings:
[0053] Figure 1 A flowchart of Chinese sentence word sense disambiguation in an embodiment of the present invention;
[0054] Figure 2 This is the training process of the word sense disambiguation model based on CBAM embedded in Regenty network in the implementation mode of the present invention.
[0055] Figure 3 This is the testing process of the word sense disambiguation model based on CBAM embedded in Regenty network in the implementation mode of the present invention. Specific implementation method:
[0056] In order to make the technical solutions in the embodiments of the present invention be described clearly and completely, taking the test sentence "The meeting emphasized the need to take the medical reform as an opportunity to continue to deepen the reform of traditional Chinese medical institutions" containing the ambiguous word "traditional Chinese medicine" as an example, the present invention is further described in detail in combination with the drawings in the embodiments. The ambiguous word "traditional Chinese medicine" has two semantic categories, s1: practitioner_of_Chinese_medicine, s2: traditional_Chinese_medical_science.
[0057] The flowchart of the word sense disambiguation method based on CBAM embedded in Regenty network in the embodiment of the present invention is as follows: Figure 1 As shown, the following steps are included:
[0058] Step 1 The process of extracting disambiguation features is as follows:
[0059] Step 1-1 uses the Chinese word segmentation tool to segment the Chinese sentence into words. The specific results are as follows:
[0060] Word segmentation results: The meeting emphasized that we should take the medical reform as an opportunity to continue to deepen the reform of traditional Chinese medical institutions
[0061] Step 1-2 uses the Chinese part-of-speech tagging tool to tag the parts of speech of all words in the sentence. The specific results are as follows:
[0062] Part-of-speech tagging: The meeting emphasized that we should take the medical reform as an opportunity to continue to deepen the reform of traditional Chinese medicine institutions.
[0063] Step 1-3 uses the Chinese semantic annotation tool to annotate the semantic classes of all words in the sentence. The specific results are as follows:
[0064] Semantic annotation: The meeting / n / Di23 emphasized / v / Gb21 to / v / Ag04 take / p / Di02 medical reform / j / -1 as / v / Ih01 an opportunity / n / Ca04 to continue / v / Ig03 to deepen / v / Ih10 the reform / vn / Ih10 of traditional Chinese medicine / n / Ae15 medical / n / Hg20 institutions / n / Di09
[0065] Step 1-4 uses the Chinese character to pinyin tool to mark the pinyin first letter and tone of all words in the sentence. The specific results are:
[0066] Pinyin initials and tone markings: The meeting / n / Di23 / hy / 44 emphasized / v / Gb21 / qd / 24 to / v / Ag04 / y / 4 take / p / Di02 / y / 3 medical reform / j / -1 / yg / 13 as / v / Ih01 / w / 4 an opportunity / n / Ca04 / qj / 41 continue / v / Ig03 / jx / 44 deepen / v / Ih10 / sh / 14 traditional Chinese medicine / n / Ae15 / zy / 11 medical / n / Hg20 / yl / 12 institutions / n / Di09 / jg / 14 reform / vn / Ha04 / gg / 32
[0067] Step 1-5 selects adjacent lexical units with nouns, verbs, adjectives, numerals, quantifiers and pronouns on the left and right of the ambiguous word. The specific results are:
[0068] Meeting / n / Di23 / hy / 44 Emphasis / v / Gb21 / qd / 24 Want / v / Ag04 / y / 4 / v / Ih01 / w / 4 Opportunity / n / Ca04 / qj / 41 Continue / v / Ig03 / jx / 44 Deepen / v / Ih10 / sh / 14 Medical / n / Hg20 / yl / 12 Agency / n / Di09 / jg / 14 Reform / vn / Ha04 / gg / 32
[0069] Steps 1-6 use the word form, part of speech, semantic class, pinyin initials, and tones of the selected adjacent vocabulary units as disambiguation features. The specific results are:
[0070] Word form Part of Speech Semantic Class Pinyin initials tone Meeting n Di23 hy 44 emphasize v Gb21 qd 24 want v Ag04 y 4 for v Ih01 w 4 Opportunity n Ca04 qj 41 continue v Ig03 jx 44 deepen v Ih10 sh 14 Medical n Hg20 yl 12 mechanism n Di09 jg 14 reform v Ha04 gg 32
[0071] Step 2: Get test data and training data:
[0072] Step 2-1 uses the Word2Vec tool to vectorize the disambiguation features extracted from the training corpus of SemEval-2007:Task#5 to obtain training data;
[0073] Step 2-2 uses the Word2Vec tool to vectorize the disambiguation features extracted from the test corpus of SemEval-2007:Task#5 to obtain test data. The result is:
[0074] tensor([[-2.1511e-04,9.0341e-05,2.2727e-03,...,-2.1809e-03,-1.6873e-03,-3.8032e-03],
[0075] [3.5645e-03,-2.1729e-03,3.7584e-03,...,-3.0925e-03,1.5182e-03,9.5325e-05],
[0076] [2.1334e-03,-3.1879e-03,1.7961e-03,...,-1.9394e-03,-8.0814e-04,2.8975e-03],
[0077] ...,
[0078] [1.9837e-03,-4.4545e-03,3.3295e-03,...,2.1458e-04,-1.3078e-03,-5.0350e-04],
[0079] [-4.2148e-03,-1.8479e-03,-1.7771e-03,...,1.3312e-03,4.8244e-04,-2.3722e-03],
[0080] [-2.5976e-03,3.2163e-03,9.0198e-04,...,3.1928e-03,2.0543e-03,1.2940e-03]])
[0081] Step 3 uses the training data to optimize CBAMRegnety;
[0082] Step 3-1 inputs the training data into the initialized CBAMRegnety;
[0083] Step 3-2 passes through convolution layer 1 and extracts feature X1;
[0084] Step 3-3 extracts feature X2 through a channel attention convolution layer, wherein the channel attention convolution layer includes a channel attention module SE and a convolution layer 2;
[0085] Step 3-4 passes through the convolutional block attention convolutional layer to extract feature X3. The convolutional block attention convolutional layer includes a convolutional block attention module CBAM and a convolutional layer 3. The output calculation process of CBAM is as follows:
[0086]
[0087]
[0088] Among them, σ is the sigmoid function, MLP is a two-layer neural network, AvgPool(X2) means the average pooling operation on X2, MaxPool(X2) means the maximum pooling operation on X2, Mc*X2 is the multiplication operation of Mc and the input feature X2, and the output of CBAM is Mc*(Mc*X2);
[0089] Step 3-5 is to calculate the weight w(s) of the ambiguous word "traditional Chinese medicine" in the semantic category s1 = practitioner_of_Chinese_medicine, s2 = traditional_Chinese_medical_science through the adaptive average pooling layer. i|m), i=1,2;
[0090] Step 3-6 uses the cross entropy loss function to calculate the error loss between the actual output and the expected output 中医 , the calculation process is as follows:
[0091] loss 中医 =0.588
[0092] Step 3-7 is based on the error loss 中医 Back propagation, update parameters layer by layer, the parameter update process is as follows:
[0093]
[0094] Among them, θ 中医 represents the parameter set, θ' 中医 represents the updated parameter set, a is the learning rate;
[0095] Step 3-8 continuously iterates step 3-1 to step 3-7 until the specified number of times is reached to obtain the optimized CBAMRegnety.
[0096] Step 4: semantically classify the ambiguous word "traditional Chinese medicine":
[0097] Step 4-1 Input the test data of "TCM" into the optimized CBAMRegnety;
[0098] Step 4-2 passes through convolution layer 1 and extracts feature X1;
[0099] Step 4-3 extracts feature X2 through a channel attention convolution layer, wherein the channel attention convolution layer includes a channel attention module SE and a convolution layer 2;
[0100] Step 4-4 passes through the convolutional block attention convolutional layer to extract feature X3. The convolutional block attention convolutional layer includes a convolutional block attention module CBAM and a convolutional layer 3. The output calculation process of CBAM is as follows:
[0101]
[0102]
[0103] Among them, σ is the sigmoid function, MLP is a two-layer neural network, AvgPool(X2) means the average pooling operation on X2; MaxPool(X2) means the maximum pooling operation on X2; Mc*X2 is the multiplication operation of Mc and the input feature X2, and the output of CBAM is M_s*(M_c*X2);
[0104] Step 4-5 is to calculate the weight w(s) of the ambiguous word "traditional Chinese medicine" in the semantic category s1 = practitioner_of_Chinese_medicine, s2 = traditional_Chinese_medical_science through the adaptive average pooling layer. i |m), i=1,2;
[0105] Output the weights assigned to the ambiguous word "traditional Chinese medicine" under the semantic categories s1 = practitioner_of_Chinese_medicine and s2 = traditional_Chinese_medical_science:
[0106] w(practitioner_of_Chinese_medicine|TCM)=0.0628
[0107] w(traditional_Chinese_medical_science|Traditional Chinese Medicine)=0.0825
[0108] Steps 4-6 output the semantic category with the maximum weight, as follows:
[0109]
[0110] s=traditional_Chinese_medical_science represents the semantic category corresponding to the ambiguous word “traditional Chinese medicine”.
[0111] Through the word sense disambiguation model of the optimized CBAM embedded in the Regenety network, the Chinese sentence "The government strongly advocates continuing to deepen the development of traditional Chinese medical institutions" containing the ambiguous word "traditional Chinese medicine" is disambiguated. The semantic category corresponding to the ambiguous word "traditional Chinese medicine" is traditional_Chinese_medical_science.
[0112] The word sense disambiguation method of embedding CBAM into Regnetic network in the embodiment of the present invention can select accurate disambiguation features and use CBAM embedded into Regnetic network to determine the semantic category of ambiguous words, and has a high accuracy rate.
[0113] The above is a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings, and the specific implementation methods herein are only used to help understand the method of the present invention. For those of ordinary skill in the art, according to the idea of the present invention, changes and modifications can be made in the specific implementation methods and application scopes, so the present invention should not be understood as limiting the present invention.
Claims
1. Word sense disambiguation method based on CBAM embedded in Regenerative network, ambiguous word m has n semantic categories s1, s2, …, s n , the method comprises the following steps: Step 1: Perform word segmentation, part-of-speech tagging, pinyin initials tagging, tone tagging and semantic class tagging on the training and test corpora of SemEval-2007:Task#5, and select the word form, part-of-speech, semantic class, pinyin initials and tones of adjacent lexical units with noun, verb, adjective, numeral, quantifier and pronoun parts of speech around the ambiguous word m as disambiguation features; Step 2: Use the Word2Vec tool to vectorize the disambiguation features extracted from the training corpus of SemEval-2007:Task#5 to obtain training data, and use the Word2Vec tool to vectorize the disambiguation features extracted from the test corpus of SemEval-2007:Task#5 to obtain test data; Step 3: The training includes two processes: forward propagation and back propagation. The training data is used to optimize CBAMRegnety to obtain the optimized CBAMRegnety. The specific steps are as follows: Step 3-1 inputs the training data into the initialized CBAMRegnety; Step 3-2 passes through convolution layer 1 and extracts feature X1; Step 3-3 extracts feature X2 through a channel attention convolution layer, wherein the channel attention convolution layer includes a channel attention module SE and a convolution layer 2; Step 3-4 passes through the convolutional block attention convolutional layer to extract feature X3. The convolutional block attention convolutional layer includes a convolutional block attention module CBAM and a convolutional layer 3. The output calculation process of CBAM is as follows: Among them, σ is the sigmoid function, MLP is a two-layer neural network, AvgPool(X2) means the average pooling operation on X2, MaxPool(X2) means the maximum pooling operation on X2, Mc*X2 is the multiplication operation of Mc and the input feature X2, and the output of CBAM is Mc*(Mc*X2); Step 3-5 passes through the adaptive average pooling layer to calculate the ambiguous word m in the semantic category s i The weight w(s i |m), i=1,2,...,n; Step 3-6 uses the cross entropy loss function to calculate the error loss. The calculation process of the error loss is as follows: Among them, loss represents the average error of training data, n is the number of training data, and y k is the label of the kth training data; Steps 3-7 are performed based on the error loss back propagation, and the parameters are updated layer by layer. The parameter update process is as follows: Among them, θ represents the parameter set, θ' represents the updated parameter set, and a is the learning rate; Step 3-8 continuously iterates step 3-1 to step 3-7 until the specified number of cycles is reached to obtain the optimized CBAMRegnety; Step 4: The test process is a forward propagation process, that is, a semantic classification process; on the optimized CBAMRegnety, input the test data and calculate the weight of the ambiguous word m in each semantic category, where the semantic category with the largest weight is the semantic category of the ambiguous word.
2. The word sense disambiguation method based on CBAM embedded in Regenetic network according to claim 1 is characterized in that: In step 1, the specific steps are: Step 1-1 uses a Chinese word segmentation tool to perform vocabulary segmentation on Chinese sentences; Step 1-2 uses a Chinese part-of-speech tagging tool to tag the parts of speech of all words in the sentence; Step 1-3 uses Chinese semantic annotation tools to annotate the semantic categories of all words in the sentence; Step 1-4: Use the Chinese character to pinyin tool to mark the pinyin first letter and tone of all words in the sentence; Step 1-5 selects adjacent lexical units with noun, verb, adjective, numeral, quantifier and pronoun parts of speech on the left and right of the ambiguous word; Steps 1-6 use the word form, part of speech, semantic class, pinyin initial letter, and tone of the selected adjacent lexical units as disambiguation features.
3. The word sense disambiguation method based on CBAM embedded in Regenty network according to claim 1 is characterized in that: In step 2, the specific steps are: Step 2-1 uses the Word2Vec tool to vectorize the disambiguation features extracted from the training corpus of SemEval-2007:Task#5 to obtain training data; Step 2-2 uses the Word2Vec tool to vectorize the disambiguation features extracted from the test corpus of SemEval-2007:Task#5 to obtain test data.
4. The word sense disambiguation method based on CBAM embedded in Regenetic network according to claim 1 is characterized in that: In step 4, the specific process is: Step 4-1 Input the test data into the optimized CBAMRegnety; Step 4-2 passes through convolution layer 1 and extracts feature X1; Step 4-3 extracts feature X2 through a channel attention convolution layer, wherein the channel attention convolution layer includes a channel attention module SE and a convolution layer 2; Step 4-4 passes through the convolutional block attention convolutional layer to extract feature X3. The convolutional block attention convolutional layer includes a convolutional block attention module CBAM and a convolutional layer 3. The output calculation process of CBAM is as follows: Among them, σ is the sigmoid function, MLP is a two-layer neural network, AvgPool(X2) means the average pooling operation on X2, MaxPool(X2) means the maximum pooling operation on X2, Mc*X2 is the multiplication operation of Mc and the input feature X2, and the output of CBAM is Mc*(Mc*X2); Step 4-5 passes through the adaptive average pooling layer to calculate the ambiguous word m in the semantic category s i The weight w(s i |m), i=1,2,...,n; Step 4-6 selects the semantic category corresponding to the maximum weight as the semantic category of the ambiguous word m: Among them, s represents the semantic category of the ambiguous word m.
Citation Information
Patent Citations
Dependency constraint and knowledge-based adjective meaning disambiguation method and apparatus
CN106202034A
A method of word sense disambiguation in Chinese sentences based on convolution neural network
CN109214007A