Word Sense Disambiguation Method Based on Multi-Channel Graph Sampling Aggregation Neural Network
Through the multi-path sampling and aggregation neural network model, combining word segmentation, part of speech, semantic classes and radical annotation, a word meaning disambiguation feature map is constructed, and a GraphSAGE neural network is used for semantic classification, which solves the problem of insufficient vocabulary ambiguity processing in the existing technology, and achieves higher word meaning disambiguation accuracy and classification effect.
Patent Information
- Application Number
- CN202211133281.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-09-17
AI Technical Summary
When dealing with vocabulary ambiguity, existing natural language processing algorithms have the problem of insufficient extraction of local disambiguation characteristics and poor classification effect, especially in the fields of machine translation, automatic abstracts and information retrieval.
The multi-path sampling aggregation neural network (GraphSAGE) model is used to extract the unambiguous features through word segmentation, part-of-speech annotation, semantic class annotation and radical annotation, and vectorized using Bi-LSTM and Attention networks to construct four word-sense disambiguation feature maps, and semantic classification is used using the optimized GraphSAGE neural network.
It improves the accuracy of word meaning disambiguation, can better determine the semantic categories of ambiguity vocabulary in a specific context, and improves the application effect of machine translation, automatic abstracts, and information retrieval.
Smart Images

Figure QLYQS_1 
Figure QLYQS_4 
Figure QLYQS_5
Abstract
Description
Technical Field:
[0001] The present invention relates to a word sense disambiguation method based on a multi-way graph sampling aggregation neural network, which has good applications in the field of natural language processing. Background Art:
[0002] In natural language, some words often have multiple meanings, and word sense disambiguation has become a major research issue in natural language processing. Its purpose is to determine the correct semantics of ambiguous words in a specific context. Word sense disambiguation has important applications in machine translation, automatic summarization, information retrieval, and text classification. Determining the correct semantics of ambiguous words will have better effects on the applications in these fields.
[0003] Nowadays, some common algorithms are often used to disambiguate and classify words, such as: Naive Bayes, k-means, classification methods based on statistics, association rules, and artificial neural networks, etc. However, these traditional algorithms have some deficiencies. They can only extract local disambiguation features or the extracted disambiguation features are insufficient, and the classification effect of the classifier is poor. With the gradual wide application of deep learning algorithms in the field of natural language processing, models such as recurrent neural networks, convolutional neural networks, and graph neural networks, these deep learning algorithms can better extract disambiguation features. The graph sampling aggregation neural network is a graph neural network model proposed in recent years. This model directly models on the graph. By constructing a word sense disambiguation feature graph, disambiguation features can be better extracted, and the disambiguation features of nodes and their neighboring nodes are aggregated. For ambiguous words, the GraphSAGE network can be well applied for disambiguation to achieve correct semantic classification. Summary of the Invention:
[0004] In order to solve the problem of lexical ambiguity in the field of natural language processing, the present invention discloses a word sense disambiguation method based on a multi-way graph sampling aggregation neural network.
[0005] To this end, the present invention provides the following technical solutions:
[0006] 1. A word sense disambiguation method based on a multi-way graph sampling aggregation neural network, characterized in that the method mainly includes the following steps:
[0007] Step 1: Perform word segmentation, part-of-speech tagging, semantic class tagging, and radical tagging on all Chinese sentences included in the SemEval-2007: Task#5 corpus, select the sentences where the ambiguous words are located, and use the word form, part of speech, semantic class, and radical of the two adjacent lexical units on the left and right of the ambiguous words as disambiguation features.
[0008] Step 2: Use the Bi-LSTM and Attention networks to vectorize the extracted sentence features, use the Word2Vec tool to vectorize the word form, part-of-speech, semantic class, and radical features, use the processed training corpus in SemEval-2007: Task #5 as the training data, and use the processed test corpus in SemEval-2007: Task #5 as the test data.
[0009] Step 3: Take the extracted sentence, as well as the word form, part-of-speech, semantic class, and radical of the two adjacent lexical units on the left and right of the ambiguous word as the nodes of the graph, and construct a sentence-word form word sense disambiguation feature graph, a sentence-part-of-speech word sense disambiguation feature graph, a sentence-semantic class word sense disambiguation feature graph, and a sentence-radical word sense disambiguation feature graph respectively.
[0010] Step 4: In the training process, input the four word sense disambiguation feature graphs constructed from the training data into the multi-way GraphSAGE neural network and optimize it to obtain the optimized multi-way GraphSAGE neural network.
[0011] Step 5: The test process is a semantic classification process. Input the four word sense disambiguation feature graphs constructed from the test data into the optimized multi-way GraphSAGE neural network, and calculate the probability distribution of the ambiguous word under each semantic category. Among them, the semantic category with the highest probability is the semantic category of the ambiguous word.
[0012] 2. The word sense disambiguation method based on the multi-way graph sampling aggregation neural network according to claim 1, wherein in step 1, the Chinese sentence containing the ambiguous word w is segmented, part-of-speech tagged, semantic class tagged, and radical tagged, and the disambiguation features are extracted. The specific steps are as follows:
[0013] Step 1-1: Use a Chinese word segmentation tool to segment the Chinese sentence into words.
[0014] Step 1-2: Use a Chinese part-of-speech tagging tool to tag the part-of-speech of the segmented words.
[0015] Step 1-3: Use a Chinese semantic class tagging tool to tag the semantic class of the segmented words.
[0016] Step 1-4: Use a Chinese radical tagging tool to tag the radical of the segmented words.
[0017] Step 1-5: Select the sentence where the ambiguous word is located, and take the word form, part-of-speech, semantic class, and radical of the two adjacent lexical units on the left and right of the ambiguous word as the disambiguation features.
[0018] 3. The word sense disambiguation method of the multi-channel graph sampling aggregation neural network according to claim 1, characterized in that in step 2, the sentence features are vectorized, and the word form, part of speech, semantic class, and radical features are vectorized to obtain training data and test data. The specific steps are as follows:
[0019] Step 2-1: Use a Bi-LSTM and Attention network to vectorize the extracted sentence features, and use the Word2Vec tool to vectorize the extracted word form, part of speech, semantic class, and radical features respectively. After vectorization, each disambiguation feature corresponds to a 200-dimensional feature vector.
[0020] Step 2-2: Use the processed training corpus in SemEval-2007:Task#5 as training data, and use the processed test corpus in SemEval-2007:Task#5 as test data.
[0021] 4. The word sense disambiguation method based on the multi-channel graph sampling aggregation neural network according to claim 1, characterized in that in step 3, four word sense disambiguation feature graphs are constructed. The specific steps are as follows:
[0022] Step 3-1: Use the sentence with the ambiguous word w and the word forms of the two adjacent lexical units on the left and right of w as the nodes in the sentence-word form graph; use the sentence with w and the parts of speech of the two adjacent lexical units on the left and right of w as the nodes in the sentence-part of speech graph; use the sentence with w and the semantic classes of the two adjacent lexical units on the left and right of w as the nodes in the sentence-semantic class graph; use the sentence with w and the radicals of the two adjacent lexical units on the left and right of w as the nodes in the sentence-radical graph.
[0023] Step 3-2: Use the feature vectors of the disambiguation features obtained in step 2 to perform weight embedding on the sentence nodes and word form nodes in the sentence-word form graph, the sentence nodes and part of speech nodes in the sentence-part of speech graph, the sentence nodes and semantic class nodes in the sentence-semantic class graph, and the sentence nodes and radical nodes in the sentence-radical graph respectively.
[0024] Step 3-3: Establish the edge relationship between the sentence nodes and word form nodes in the sentence-word form graph according to the number of times the word form appears in the sentence, establish the edge relationship between the sentence nodes and part of speech nodes in the sentence-part of speech graph according to the number of times the part of speech appears in the sentence, establish the edge relationship between the sentence nodes and semantic class nodes in the sentence-semantic class graph according to the number of times the semantic class appears in the sentence, and establish the edge relationship between the sentence nodes and radical nodes in the sentence-radical graph according to the number of times the radical appears in the sentence.
[0025] 5. The word sense disambiguation method based on the multi-way graph sampling aggregation neural network according to claim 1, characterized in that in step 4, the multi-way GraphSAGE neural network is optimized, and the specific steps are as follows:
[0026] Step 4-1: Input the sentence-word form graph constructed from the training data into the initialized GraphSAGE0, input the sentence-part-of-speech graph constructed from the training data into the initialized GraphSAGE1, input the sentence-semantic class graph constructed from the training data into the initialized GraphSAGE2, and input the sentence-radical graph constructed from the training data into the initialized GraphSAGE3, where GraphSAGE0, GraphSAGE1, GraphSAGE2, and GraphSAGE3 are neural networks with the same initialization;
[0027] Step 4-2: Respectively pass through the aggregation layers of their respective GraphSAGE neural networks to filter the disambiguation information, extract more complete disambiguation features, and aggregate the features between nodes and their adjacent nodes. The aggregation formula is as follows:
[0028]
[0029] Among them, represents the feature vector of node v, including k-layer aggregation operations, represents the feature vector of node u i including k - 1 layer aggregation operations, and n is the number of adjacent node samplings of node v;
[0030] Step 4-3: Respectively pass through the output layers of their respective GraphSAGEs, and splice the output results of the four output layers. Use the softmax function to calculate the prediction probability of the ambiguous word w in the semantic category s i (i = 1, 2,..., n):
[0031]
[0032] Among them, a i represents the input data of the softmax function;
[0033] Step 4-4: Use the cross-entropy loss function to calculate the error loss between the predicted label category and the actual label category. The calculation formula is as follows:
[0034]
[0035] Among them, x is the predicted probability value of the ambiguous word w in different semantic categories, and y is the true semantic category label of the ambiguous word w;
[0036] Step 4-5 Backpropagate according to the error loss, and update the parameters layer by layer. The parameter update process is as follows:
[0037]
[0038] Among them, θ represents the parameter set, θ’ represents the updated parameter set, and α is the learning rate;
[0039] Step 4-6 Continuously iterate Steps 4-1 to 4-5 until the number of training times is reached, and obtain the optimized multi-way GraphSAGE neural network;
[0040] 6. The word sense disambiguation method based on a multi-way graph sampling aggregation neural network according to claim 1, characterized in that, in the said step 5, semantic classification is performed on the ambiguous word w, and the specific steps are as follows:
[0041] Step 5-1 Input the sentence-word form graph constructed from the test data into the optimized GraphSAGE0, input the sentence-part-of-speech graph constructed from the test data into the optimized GraphSAGE1, input the sentence-semantic class graph constructed from the test data into the optimized GraphSAGE2, and input the sentence-radical graph constructed from the test data into the optimized GraphSAGE3;
[0042] Step 5-2 Pass through the aggregation layers of their respective GraphSAGE neural networks respectively, filter the disambiguation information, extract more complete disambiguation features, and aggregate the features between nodes and their adjacent nodes. The aggregation formula is as follows:
[0043]
[0044] Among them, represents the feature vector of node v, including k-layer aggregation operations, represents the feature vector of node u i with k-1 layer aggregation operations, and n is the number of adjacent node samplings of node v;
[0045] Step 5-3 Pass through the output layers of their respective GraphSAGEs respectively, and splice the output results of the four output layers. Use the softmax function to calculate the prediction probability of the ambiguous word w in the semantic category s i (i = 1, 2,..., n):
[0046]
[0047] Among them, a i represents the input data of the softmax function;
[0048] Step 5-4 selects the semantic class with the highest probability from P(s1|w), P(s2|w),..., P(s n |w) as the semantic class of w:
[0049]
[0050] Beneficial effects:
[0051] 1. The present invention is a word sense disambiguation method based on a multi-way graph sampling aggregation neural network. It performs lexical segmentation, part-of-speech tagging, semantic class tagging, and radical tagging on Chinese sentences. The Word2Vec tool, along with the Bi-LSTM and Attention networks, is used to vectorize the disambiguation features. The extracted disambiguation features have high quality.
[0052] 2. The model used in the present invention is the GraphSAGE neural network. The biggest feature is that it iteratively updates node features with the help of a graph structure. Each node only samples a part of its adjacent nodes to iteratively update its own features. By constructing four types of word sense disambiguation feature graphs and using a multi-way GraphSAGE neural network, better classification results can be obtained.
[0053] 3. The classifier used in the present invention is the softmax classifier, which can not only solve binary classification problems but also multi-classification problems.
[0054] 4. When training the model, the gradient descent method is used to update the weight matrix parameters in the aggregation layer of the model. The error is calculated through the loss function, and the model parameters are updated using gradient descent to obtain an optimized multi-way GraphSAGE neural network, which improves the disambiguation accuracy. Description of the drawings:
[0055] Figure 1 is the flowchart of word sense disambiguation based on a multi-way graph sampling aggregation neural network in the embodiment of the present invention;
[0056] Figure 2 is the model structure diagram of Bi-LSTM with an Attention mechanism in the embodiment of the present invention;
[0057] Figure 3 is the sentence-word form graph constructed from the test sentence in the embodiment of the present invention;
[0058] Figure 4 is the sentence-part-of-speech graph constructed from the test sentence in the embodiment of the present invention;
[0059] Figure 5 is the sentence-semantic class graph constructed from the test sentence in the embodiment of the present invention;
[0060] Figure 6 This is the sentence-radical graph constructed from test sentences in the embodiments of the present invention;
[0061] Figure 7 This is the training process of the word sense disambiguation model based on the multi-way graph sampling aggregation neural network in the embodiments of the present invention;
[0062] Figure 8 This is the testing process of the word sense disambiguation model based on the multi-way graph sampling aggregation neural network in the embodiments of the present invention. Specific implementation manners:
[0063] In order to clearly and completely describe the technical solutions in the embodiments of the present invention, taking the test sentence "Every meal and every dish we eat is hard-won!" containing the ambiguous word "dish" as an example, and in combination with the accompanying drawings in the embodiments, the present invention will be further described in detail. The ambiguous word "dish" has two semantic categories, s1: dish, s2: vegetable. There are a total of 56 sentences in the training corpus of "dish", and a total of 19 sentences in the test corpus of "dish".
[0064] The flowchart of word sense disambiguation based on the multi-way graph sampling aggregation neural network in the embodiments of the present invention is as Figure 1 shown, and includes the following steps.
[0065] Step 1 Process the Chinese sentence containing the ambiguous word "dish" and extract disambiguation features, as follows:
[0066] Step 1-1 Use a Chinese word segmentation tool to segment the Chinese sentence, specifically as follows:
[0067] Word segmentation result: Every meal and every dish we eat is hard-won
[0068] Step 1-2 Use a Chinese part-of-speech tagging tool to tag the parts of speech of the segmented words, specifically as follows:
[0069] Part-of-speech tagging: we / r eat / v of / u one / m meal / n one / m dish / n all / d hard-won / i ah / y
[0070] Step 1-3 Use a Chinese semantic class tagging tool to tag the semantic classes of the segmented words, specifically as follows:
[0071] Semantic class tagging: we / r / Aa02 eat / v / Je13 of / u / Bo29 one / m / Ka16 meal / n / Br04 one / m / Ka16 dish / n / Bh09 all / d / Cb25 hard-won / i / Ed07 ah / y / Ke03
[0072] Step 1-4 Use a Chinese character radical tagging tool to tag the radicals of the segmented words, specifically as follows:
[0073] Radical annotation: The / r / Aa02 / we / v / Je13 / eat / u / Bo29 / what we eat / n / Br04 / a meal / n / Bh09 / a dish / d / Cb25 / are all / i / Ed07 / not easily come by / y / Ke03 / ah
[0074] Step 1-5 Select the sentence where the ambiguous word is located. The word form, part of speech, semantic class, and radical of the two adjacent lexical units on the left and right of the ambiguous word are used as disambiguation features, specifically as follows:
[0075]
[0076] Step 2 Obtain the training data and test data.
[0077] Step 2-1 Use the Bi-LSTM and Attention networks as shown in Figure 2 to vectorize the extracted sentence features. Use the Word2Vec tool to vectorize the extracted word form, part of speech, semantic class, and radical features respectively. After vectorization, each disambiguation feature corresponds to a 200-dimensional feature vector. The results of vectorizing the features extracted from the test sentence "The meal and dish we eat are all not easily come by!" are as follows:
[0078]
[0079] Step 2-2 Use the processed training corpus in SemEval-2007:Task#5 as the training data, and use the processed test corpus in SemEval-2007:Task#5 as the test data;
[0080] Step 3 Construct four disambiguation feature graphs. The four disambiguation feature graphs of "The meal and dish we eat are all not easily come by!" are as shown in Figure 3 , Figure 4 , Figure 5 and Figure 6 shown. The specific steps are as follows:
[0081] Step 3-1 Take the sentence with the ambiguous word "dish" and the word forms of the two adjacent lexical units on the left and right of the ambiguous word as nodes in the sentence-word form graph, take the sentence with "dish" and the parts of speech of the two adjacent lexical units on the left and right of the ambiguous word as nodes in the sentence-part of speech graph, take the sentence with "dish" and the semantic classes of the two adjacent lexical units on the left and right of the ambiguous word as nodes in the sentence-semantic class graph, and take the sentence with "dish" and the radicals of the two adjacent lexical units on the left and right of the ambiguous word as nodes in the sentence-radical graph;
[0082] In step 3-2, the eigenvectors of each disambiguation feature obtained in step 2 are used to perform weight embedding on the sentence nodes and word form nodes in the sentence-word form graph, the sentence nodes and part-of-speech nodes in the sentence-part-of-speech graph, the sentence nodes and semantic class nodes in the sentence-semantic class graph, and the sentence nodes and radical nodes in the sentence-radical graph respectively;
[0083] In step 3-3, edge relationships between sentence nodes and word form nodes are established for the sentence-word form graph according to the number of times the word form appears in the sentence, edge relationships between sentence nodes and part-of-speech nodes are established for the sentence-part-of-speech graph according to the number of times the part of speech appears in the sentence, edge relationships between sentence nodes and semantic class nodes are established for the sentence-semantic class graph according to the number of times the semantic class appears in the sentence, and edge relationships between sentence nodes and radical nodes are established for the sentence-radical graph according to the number of times the radical appears in the sentence;
[0084] In step 4, the multi-way GraphSAGE neural network is optimized using training data, as Figure 7 shown, and the specific steps are as follows:
[0085] In step 4-1, the sentence-word form graph constructed from the training data of "cai" is input into the initialized GraphSAGE0, the sentence-part-of-speech graph constructed from the training data of "cai" is input into the initialized GraphSAGE1, the sentence-semantic class graph constructed from the training data of "cai" is input into the initialized GraphSAGE2, and the sentence-radical graph constructed from the training data of "cai" is input into the initialized GraphSAGE3, where GraphSAGE0, GraphSAGE1, GraphSAGE2, and GraphSAGE3 are neural networks with the same initialization;
[0086] In step 4-2, through the aggregation layers of their respective GraphSAGE neural networks, the disambiguation information is filtered to extract more complete disambiguation features, and the features between nodes and their adjacent nodes are aggregated. The aggregation formula is as follows:
[0087]
[0088] where represents the eigenvector of node v, including k layers of aggregation operations, represents the eigenvector of node u i including k - 1 layers of aggregation operations, and n is the number of adjacent node samplings of node v;
[0089] Step 4-3 passes through the output layers of each GraphSAGE respectively, and concatenates the output results of the four output layers to calculate the prediction probabilities P(s1|菜) and P(s2|菜) of the ambiguous word "菜" under the semantic categories s1=dish, s2=vegetable;
[0090] Step 4-4 uses the cross entropy loss function to calculate the error loss between the predicted label category and the actual label category 菜 , the calculation formula is as follows:
[0091] loss 菜 =0.6826
[0092] Step 4-5: According to the error loss 菜 Back propagation, update parameters layer by layer, the parameter update process is as follows:
[0093]
[0094] Among them, θ 菜 represents the parameter set, θ' 菜 represents the updated parameter set, α is the learning rate;
[0095] Step 4-6: continuously iterate step 4-1 to step 4-5 until the number of training times is reached, and the optimized multi-way GraphSAGE neural network is obtained;
[0096] Step 5: Test process Figure 8 As shown in the figure, the semantic classification of the ambiguous word "菜" is carried out, and the specific steps are as follows:
[0097] Step 5-1: Input the sentence-word form graph constructed by the test data of "菜" into the optimized GraphSAGE0, input the sentence-part-of-speech graph constructed by the test data of "菜" into the optimized GraphSAGE1, input the sentence-semantic class graph constructed by the test data of "菜" into the optimized GraphSAGE2, and input the sentence-radical graph constructed by the test data of "菜" into the optimized GraphSAGE3;
[0098] Step 5-2 is performed through the aggregation layer of each GraphSAGE neural network to filter the disambiguation information, extract more complete disambiguation features, and aggregate the features between the node and its adjacent nodes. The aggregation formula is as follows:
[0099]
[0100] in, Represents the feature vector of node v, including k-layer aggregation operations, It represents the feature vector of node ui, which includes k-1 layer aggregation operations, and n is the number of adjacent node samplings of node v;
[0101] In step 5-3, it passes through the output layers of their respective GraphSAGEs respectively, and splices the output results of the four output layers, and uses the softmax function to calculate the prediction probabilities of the ambiguous word "cai" under the semantic categories s1 and s2:
[0102] P(dish|cai) = 0.552975
[0103] P(vegetable|cai) = 0.447025
[0104] Among them, P(dish|cai) represents the occurrence probability of the ambiguous word "cai" under the semantic category s1 = dish, and P(vegetable|cai) represents the occurrence probability of the ambiguous word "cai" under the semantic category s2 = vegetable;
[0105] In step 5-4, the semantic category with the maximum probability is selected from P(dish|cai) and P(vegetable|cai) as the semantic category of "cai":
[0106]
[0107] Using the optimized multi-way GraphSAGE neural network, word sense disambiguation is performed on the Chinese sentence "Every grain of food and every dish we eat is hard-won!" containing the ambiguous word "cai", and the semantic category corresponding to the ambiguous word "cai" is dish.
[0108] Using the optimized multi-way GraphSAGE neural network to perform disambiguation on the test corpus containing the ambiguous word "cai", the accuracy acc of word sense disambiguation is:
[0109] acc = 17 / 19 ≈ 0.8947
[0110] The word sense disambiguation based on the multi-way graph sampling aggregation neural network in the embodiment of the present invention can select diverse and accurate disambiguation features, determine the semantic category of the ambiguous word by constructing multiple word sense disambiguation feature graphs and integrating the multi-way GraphSAGE neural network, and has a high accuracy.
[0111] The above is a detailed introduction to the embodiments of the present invention in combination with the accompanying drawings. The specific embodiments herein are only used to help understand the method of the present invention. For those of ordinary skill in the art in this technical field, according to the idea of the present invention, changes and modifications can be made within the specific embodiments and application scope. Therefore, this specification of the present invention should not be construed as a limitation on the present invention.
Claims
1. A method for word sense disambiguation based on a multi-way graph sampling aggregation neural network, characterized in that The method mainly includes the following steps: Step 1: Segment, perform part-of-speech tagging, semantic class tagging, and radical tagging on all Chinese sentences contained in the SemEval-2007: Task #5 corpus. Select the sentences where the ambiguous words are located, as well as the word forms, part-of-speech, semantic classes, and radicals of the two adjacent lexical units on the left and right of the ambiguous words as disambiguation features; Step 2: Use the Bi-LSTM and Attention networks to vectorize the extracted sentence features, use the Word2Vec tool to vectorize the word form, part-of-speech, semantic class, and radical features. Take the processed training corpus in SemEval-2007: Task #5 as the training data, and take the processed test corpus in SemEval-2007: Task #5 as the test data; Step 3: Take the extracted sentences, as well as the word forms, part-of-speech, semantic classes, and radicals of the two adjacent lexical units on the left and right of the ambiguous word as the nodes of the graph, and construct a sentence-word form and semantic disambiguation feature graph, a sentence-part-of-speech and semantic disambiguation feature graph, a sentence-semantic class and semantic disambiguation feature graph, and a sentence-radical and semantic disambiguation feature graph respectively; Step 4: Training process. Input the four semantic disambiguation feature graphs constructed from the training data into the multi-way GraphSAGE neural network and optimize it to obtain the optimized multi-way GraphSAGE neural network; Step 5: The testing process is a semantic classification process. Input the four semantic disambiguation feature graphs constructed from the test data into the optimized multi-way GraphSAGE neural network, and calculate the probability distribution of the ambiguous word under each semantic category. Among them, the semantic category with the highest probability is the semantic category of the ambiguous word.
2. The word sense disambiguation method based on a multi-graph sampling aggregation neural network according to claim 1, characterized in that In the above Step 1, segment, perform part-of-speech tagging, semantic class tagging, and radical tagging on the Chinese sentence containing the ambiguous word w, and extract disambiguation features. The specific steps are as follows: Step 1-1: Use a Chinese word segmentation tool to segment the Chinese sentence into words; Step 1-2: Use a Chinese part-of-speech tagging tool to perform part-of-speech tagging on the segmented words; Step 1-3: Use a Chinese semantic class tagging tool to perform semantic class tagging on the segmented words; Step 1-4: Use a Chinese radical tagging tool to perform radical tagging on the segmented words; Step 1-5: Select the sentence where the ambiguous word is located, and the word forms, part-of-speech, semantic classes, and radicals of the two adjacent lexical units on the left and right of the ambiguous word as disambiguation features.
3. The method for word sense disambiguation of the multi-channel graph sampling aggregation neural network according to claim 1, wherein In the above Step 2, vectorize the sentence features, vectorize the word form, part-of-speech, semantic class, and radical features, and obtain the training data and test data. The specific steps are as follows: Step 2-1: Use the Bi-LSTM and Attention networks to vectorize the extracted sentence features, and use the Word2Vec tool to vectorize the extracted word form, part-of-speech, semantic class, and radical features respectively. After vectorization, each disambiguation feature corresponds to a 200-dimensional feature vector; Step 2-2 uses the processed training corpus in SemEval-2007: Task #5 as the training data and the processed test corpus in SemEval-2007: Task #5 as the test data.
4. The word sense disambiguation method based on a multi-graph sampling aggregation neural network according to claim 1, characterized in that, In Step 3, four word sense disambiguation feature graphs are constructed. The specific steps are as follows: Step 3-1 uses the sentence with the ambiguous word w and the word forms of the two adjacent lexical units on the left and right of w as the nodes in the sentence-word form graph; uses the sentence with w and the part-of-speech tags of the two adjacent lexical units on the left and right of w as the nodes in the sentence-part-of-speech graph; uses the sentence with w and the semantic classes of the two adjacent lexical units on the left and right of w as the nodes in the sentence-semantic class graph; uses the sentence with w and the radicals of the two adjacent lexical units on the left and right of w as the nodes in the sentence-radical graph. Step 3-2 uses the feature vectors of the disambiguation features obtained in Step 2 to perform weight embedding on the sentence nodes and word form nodes in the sentence-word form graph, the sentence nodes and part-of-speech nodes in the sentence-part-of-speech graph, the sentence nodes and semantic class nodes in the sentence-semantic class graph, and the sentence nodes and radical nodes in the sentence-radical graph. Step 3-3 establishes the edge relationship between the sentence nodes and word form nodes in the sentence-word form graph according to the number of times the word form appears in the sentence, establishes the edge relationship between the sentence nodes and part-of-speech nodes in the sentence-part-of-speech graph according to the number of times the part-of-speech appears in the sentence, establishes the edge relationship between the sentence nodes and semantic class nodes in the sentence-semantic class graph according to the number of times the semantic class appears in the sentence, and establishes the edge relationship between the sentence nodes and radical nodes in the sentence-radical graph according to the number of times the radical appears in the sentence.
5. The word sense disambiguation method based on a multi-graph sampling aggregation neural network according to claim 1, characterized in that In Step 4, the multi-way GraphSAGE neural network is optimized. The specific steps are as follows: Step 4-1 inputs the sentence-word form graph constructed from the training data into the initialized GraphSAGE0, inputs the sentence-part-of-speech graph constructed from the training data into the initialized GraphSAGE1, inputs the sentence-semantic class graph constructed from the training data into the initialized GraphSAGE2, and inputs the sentence-radical graph constructed from the training data into the initialized GraphSAGE3. Here, GraphSAGE0, GraphSAGE1, GraphSAGE2, and GraphSAGE3 are neural networks with the same initialization. Step 4-2 respectively passes through the aggregation layers of their respective GraphSAGE neural networks to filter the disambiguation information, extract more complete disambiguation features, and aggregate the features between the nodes and their adjacent nodes. The aggregation formula is as follows: Among them, represents the feature vector of node v, which contains k layers of aggregation operations, represents node u i 's feature vector, which contains k - 1 layers of aggregation operations, and n is the number of sampled adjacent nodes of node v; Step 4-3 passes through the output layers of their respective GraphSAGEs, respectively, and concatenates the output results of the four output layers, and uses the softmax function to calculate the prediction probability of the ambiguous word w in the semantic category s i (i = 1, 2, ..., n): where a i represents the input data of the softmax function; Step 4-4 uses the cross-entropy loss function to calculate the error loss between the predicted label class and the actual label class. The calculation formula is as follows: where x is the predicted probability value of the ambiguous word w under different semantic classes, and y is the true semantic class label of the ambiguous word w. Step 4-5 performs backpropagation according to the error loss and updates the parameters layer by layer. The parameter update process is as follows: Among them, θ represents the parameter set, θ' represents the updated parameter set, and α is the learning rate; Steps 4-6 continuously iterate Steps 4-1 to 4-5 until the number of training times is reached, and an optimized multi-way GraphSAGE neural network is obtained.
6. The method for word sense disambiguation based on a multi-way graph sampling aggregation neural network according to claim 1, wherein In the said Step 5, semantic classification is performed on the ambiguous word w. The specific steps are as follows: Step 5-1 Input the sentence-word form graph constructed from the test data into the optimized GraphSAGE0, input the sentence-part-of-speech graph constructed from the test data into the optimized GraphSAGE1, input the sentence-semantic class graph constructed from the test data into the optimized GraphSAGE2, and input the sentence-radical graph constructed from the test data into the optimized GraphSAGE3; Step 5-2 Pass through the aggregation layers of their respective GraphSAGE neural networks respectively, filter the disambiguation information, extract more complete disambiguation features, and aggregate the features between nodes and their adjacent nodes. The aggregation formula is as follows: Among them, represents the feature vector of node v, including k-layer aggregation operations, represents node u i 's feature vector, including k-1 layer aggregation operations, and n is the number of neighbor node samplings of node v; Step 5-3 passes through the output layers of their respective GraphSAGEs, respectively, and splices the output results of the four output layers, and uses the softmax function to calculate the predicted probability of the ambiguous word w under the semantic category s i (i = 1, 2, ..., n): where a i represents the input data of the softmax function; Step 5-4 selects the semantic class with the highest probability from P(s1|w), P(s2|w),..., P(s n |w) as the semantic class of w: 。
Citation Information
Patent Citations
Biomedical text word sense disambiguation method based on attention neural network
CN113065350A
Chinese word sense disambiguation method based on graph convolutional neural network
CN113095087A