Graph Sampling Aggregation Multi-Voting Word Sense Disambiguation Method Embedding Semantic Similarity

Through the graph sampling aggregation multi-voting word meaning disambiguation method embedded with semantic similarity, the GraphSAGE neural network and perceptron linear classifier are used to solve the problem of insufficient word meaning disambiguation in the existing technology, and achieve higher word meaning disambiguation accuracy and classification effect, which is suitable for the field of natural language processing.

CN115759117BActive Publication Date: 2025-07-18HARBIN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211482872.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-07-18
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

The existing natural language processing algorithms have problems in insufficient extraction of local disambiguation characteristics and poor classification effects in terms of word meaning disambiguation, especially in the fields of machine translation, automatic abstracts and information retrieval, and it is difficult to accurately determine the correct semantics of ambiguity vocabulary.

Method used

The graph sampling aggregation multi-voting word meaning disambiguation method embedded with semantic similarity is used to process Chinese sentences through word segmentation, part-of-speech annotation, semantic class annotation and Wubi encoding annotation, and the GraphSAGE neural network and perceptron linear classifier are used to combine multiple disambiguation feature maps and voting mechanisms to determine the semantic categories of ambiguity vocabulary.

Benefits of technology

It improves the accuracy and classification effect of word meaning disambiguation, can better obtain disambiguation characteristics from around ambiguity vocabulary, enhances the accuracy of word meaning disambiguation, and is suitable for applications in the field of natural language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

The present invention relates to a multi-voting word sense disambiguation method of graph sampling aggregation (Graph SAmple and aggreGatE, GraphSAGE) embedded with semantic similarity. First, the Chinese sentences containing ambiguous words are subjected to word segmentation, part-of-speech tagging, semantic class tagging, translation annotation, and Wubi code annotation. Using the sentences containing ambiguous words, as well as the word forms, parts of speech, semantic classes, translations, and Wubi codes contained in the two lexical units on the left and right of the ambiguous word as disambiguation features and as nodes to construct five disambiguation feature graphs. Calculate the semantic similarity between the target ambiguous word and the adjacent vocabulary using the "Thesaurus of Chinese Synonyms", vectorize the features using the Word2Vec tool and the Doc2Vec tool, and embed the semantic similarity weights into the feature vectors. Optimize the multi-GraphSAGE neural network with the training corpus, and use the optimized multi-GraphSAGE neural network combined with the perceptron linear classifier to obtain the semantic classification results of each disambiguation feature graph, and adopt a voting mechanism to determine the semantic category of the ambiguous word. The present invention has a good word sense disambiguation effect and can more accurately judge the true meaning of the ambiguous word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to a graph sampling aggregation multi-voting word sense disambiguation method embedding semantic similarity, which has good applications in the field of natural language processing. Background Art:

[0002] In natural language, some words often have multiple meanings, and word sense disambiguation has become a major research issue in natural language processing. Its purpose is to determine the correct semantics of ambiguous words in a specific context. Word sense disambiguation has important applications in machine translation, automatic summarization, information retrieval, and text classification. Determining the correct semantics of ambiguous words will produce better effects in the applications of these fields.

[0003] Currently, some common algorithms are often used to disambiguate and classify words, such as: Naive Bayes, k-means, classification methods based on statistics, association rules, and artificial neural networks, etc. However, these traditional algorithms have some deficiencies. They can only extract local disambiguation features or the extracted disambiguation features are insufficient, and the classification effect of the classifier is poor. With the gradual wide application of deep learning algorithms in the field of natural language processing, models such as recurrent neural networks, convolutional neural networks, and graph neural networks, these deep learning algorithms can better extract disambiguation features. The graph sampling aggregation neural network is a graph neural network model proposed in recent years. This model directly models on the graph. By constructing the text in the form of a graph, disambiguation features can be better extracted, and the disambiguation features of nodes and their neighboring nodes are aggregated and updated. For ambiguous words, the GraphSAGE network can be well applied for disambiguation to achieve correct semantic classification. Summary of the Invention:

[0004] In order to solve the problem of lexical ambiguity in the field of natural language processing, the present invention discloses a graph sampling aggregation multi-voting word sense disambiguation method embedding semantic similarity.

[0005] For this purpose, the present invention provides the following technical solutions:

[0006] 1. The graph sampling aggregation multi-voting word sense disambiguation method embedding semantic similarity, characterized in that the method mainly includes the following steps:

[0007] Step 1: Perform word segmentation, part-of-speech tagging, semantic class tagging, translation annotation, and Wubi encoding annotation on all Chinese sentences included in the SemEval-2007:Task#5 corpus, select the sentences where the ambiguous words are located, and use the word form, part-of-speech, semantic class, translation, and Wubi encoding of the two adjacent lexical units on the left and right of the ambiguous words as disambiguation features.

[0008] Step 2: Use the "Thesaurus of Chinese Synonyms" to calculate the semantic similarity between the target ambiguous word and its adjacent words, generate the semantic similarity weights, use the Doc2Vec tool to vectorize the extracted sentence features, use the Word2Vec tool to vectorize the extracted word form, part of speech, semantic class, translation, and Wubi coding features, and embed the semantic similarity weights into the feature vectors. Use the processed training corpus in SemEval-2007: Task #5 as the training data, and use the processed test corpus in SemEval-2007: Task #5 as the test data.

[0009] Step 3: Take the extracted sentence, as well as the word form, part of speech, semantic class, translation, and Wubi coding of the two adjacent word units on the left and right of the ambiguous word as the nodes of the disambiguation feature graph, and construct a sentence-word form disambiguation feature graph, a sentence-part of speech disambiguation feature graph, a sentence-semantic class disambiguation feature graph, a sentence-translation disambiguation feature graph, and a sentence-Wubi coding disambiguation feature graph respectively.

[0010] Step 4: Training process, input the five disambiguation feature graphs constructed from the training data into the multi-GraphSAGE neural network and optimize it to obtain the optimized multi-GraphSAGE neural network.

[0011] Step 5: Testing process, input the five disambiguation feature graphs constructed from the test data into the optimized multi-GraphSAGE neural network, obtain the semantic class of each disambiguation feature graph through the perceptron linear classifier, and adopt a voting mechanism to determine the semantic class of the ambiguous word.

[0012] 2. The graph sampling aggregation multi-voting word sense disambiguation method with embedded semantic similarity according to claim 1, wherein in the step 1, the Chinese sentence containing the ambiguous word w is segmented, part-of-speech tagged, semantic-class tagged, translation tagged, and Wubi coding tagged, and the disambiguation features are extracted. The specific steps are as follows:

[0013] Step 1-1: Use a Chinese word segmentation tool to segment the Chinese sentence into words.

[0014] Step 1-2: Use a Chinese part-of-speech tagging tool to tag the part of speech of the segmented words.

[0015] Step 1-3: Use a Chinese semantic-class tagging tool to tag the semantic class of the segmented words.

[0016] Step 1-4: Use a Chinese machine translation tool to tag the translation of the segmented words.

[0017] Step 1-5: Use a Chinese Wubi coding tagging tool to tag the Wubi coding of the segmented words.

[0018] Step 1-6 selects the sentence where the ambiguous word is located, and uses the word form, part of speech, semantic class, translation, and Wubi code of the two adjacent lexical units on the left and right of the ambiguous word as disambiguation features.

[0019] 3. The graph sampling aggregation multi-vote word sense disambiguation method with embedded semantic similarity according to claim 1, characterized in that in step 2, the semantic similarity between the target ambiguous word and adjacent words is calculated, the sentence features are vectorized, and the features of word form, part of speech, semantic class, translation, and Wubi code are vectorized and embedded with semantic similarity weights to obtain training data and test data. The specific steps are as follows:

[0020] Step 2-1 consults the "Thesaurus" to obtain the semantic class codes of the target ambiguous word w and the two adjacent words on the left and right, calculates the semantic similarity between the target ambiguous word w and the adjacent words, selects the maximum similarity as the semantic similarity sim between the ambiguous word w and the adjacent words, and obtains the semantic similarities Q L2 、Q L1 、Q R1 、Q R2 of the target ambiguous word w and the four adjacent words respectively. Take Q L2 +Q L1 +Q R1 +Q R2 as the semantic similarity Q s of the target ambiguous word w and the sentence. The formula for calculating the semantic similarity of words w1 and w2 is as follows:

[0021] If the first layer of the semantic class codes of w1 and w2 is different, then sim(w1, w2) = 0.1

[0022] If the first layer of the semantic class codes of w1 and w2 is the same and the second layer is different, then

[0023] If the first layer of the semantic class codes of w1 and w2 is the same, the second layer is the same, and the third layer is different, then

[0024] If the first layer of the semantic class codes of w1 and w2 is the same, the second layer is the same, the third layer is the same, and the fourth layer is different, then

[0025] If the first layer of the semantic class codes of w1 and w2 is the same, the second layer is the same, the third layer is the same, the fourth layer is the same, and the fifth layer is different, then

[0026] where l represents the shortest distance between words w1 and w2, l = (5 - the number of common levels) × 2, n is the total number of nodes in the branched layer, α and β are coefficients, and β1 < β2 < β3 < β4;

[0027] In step 2-2, the Doc2Vec tool is used to vectorize the extracted sentence features, and the Word2Vec tool is used to vectorize the extracted word form, part of speech, semantic class, translation, and five-stroke code features respectively. Each disambiguation feature corresponds to a 200-dimensional feature vector;

[0028] In step 2-3, the semantic similarity Q is embedded into the sentence feature vector s , and the semantic similarities Q of adjacent words in the word form, part of speech, semantic class, translation, and five-stroke code of the lexical units on both sides of the ambiguous word are respectively embedded L2 , Q L1 , Q R1 , Q R2 .

[0029] In step 2-4, the processed training corpus in SemEval-2007: Task#5 is used as the training data, and the processed test corpus in SemEval-2007: Task#5 is used as the test data.

[0030] 4. The method for graph sampling aggregation multi-vote word sense disambiguation by embedding semantic similarity according to claim 1, wherein in the step 3, five disambiguation feature graphs are constructed, and the specific steps are as follows:

[0031] In step 3-1, the sentence with the ambiguous word w and the word forms of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-word form disambiguation feature graph; the sentence with w and the parts of speech of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-part of speech disambiguation feature graph; the sentence with w and the semantic classes of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-semantic class disambiguation feature graph; the sentence with w and the translations of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-translation disambiguation feature graph; the sentence with w and the five-stroke codes of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-five-stroke code disambiguation feature graph;

[0032] In step 3-2, the feature vectors of the disambiguation features obtained in step 2 are used to perform vector embedding on the sentence nodes and word form nodes in the sentence-word form disambiguation feature graph, the sentence nodes and part of speech nodes in the sentence-part of speech disambiguation feature graph, the sentence nodes and semantic class nodes in the sentence-semantic class disambiguation feature graph, the sentence nodes and translation nodes in the sentence-translation disambiguation feature graph, and the sentence nodes and five-stroke code nodes in the sentence-five-stroke code disambiguation feature graph;

[0033] Step 3-3 establishes the edge relationship between the sentence node and the word form node for the sentence-word form disambiguation feature graph according to the number of times the word form appears in the sentence, establishes the edge relationship between the sentence node and the part-of-speech node for the sentence-part-of-speech disambiguation feature graph according to the number of times the part of speech appears in the sentence, establishes the edge relationship between the sentence node and the semantic class node for the sentence-semantic class disambiguation feature graph according to the number of times the semantic class appears in the sentence, establishes the edge relationship between the sentence node and the translation node for the sentence-translation disambiguation feature graph according to the number of times the translation appears in the sentence, and establishes the edge relationship between the sentence node and the Wubi code node for the sentence-Wubi code disambiguation feature graph according to the number of times the Wubi code appears in the sentence.

[0034] 5. The graph sampling aggregation multi-vote word sense disambiguation method with embedded semantic similarity according to claim 1, wherein in the step 4, the multi-GraphSAGE neural network is trained and optimized, and the specific steps are as follows:

[0035] Step 4-1 inputs the sentence-word form disambiguation feature graph constructed from the training data into the initialized GraphSAGE_1, inputs the sentence-part-of-speech disambiguation feature graph constructed from the training data into the initialized GraphSAGE_2, inputs the sentence-semantic class disambiguation feature graph constructed from the training data into the initialized GraphSAGE_3, inputs the sentence-translation disambiguation feature graph constructed from the training data into the initialized GraphSAGE_4, and inputs the sentence-Wubi code disambiguation feature graph constructed from the training data into the initialized GraphSAGE_5, wherein GraphSAGE_1, GraphSAGE_2, GraphSAGE_3, GraphSAGE_4, and GraphSAGE_5 are neural networks with the same parameter initialization;

[0036] Step 4-2 performs neighbor multi-order sampling on each disambiguation feature graph. First, starting from the source node, randomly sample the first-order neighbor nodes of the source node according to the first-order sampling number to obtain the list of the first-order neighbor nodes of the source node. Take the obtained first-order neighbor nodes of the source node as the starting point and randomly sample their first-order neighbors according to the node sampling number, that is, randomly sample the second-order neighbor nodes of the source node according to the sampling number to obtain the list of the second-order neighbor nodes of the source node. If the number of the first-order neighbor nodes of a certain node is less than the specified sampling number, sampling with replacement will be performed, and duplicate nodes will appear in the sampling result;

[0037] Step 4-3 performs an aggregation operation on the features of the sampled neighbor nodes. Node v is the first-order sampled neighbor node of the central node u, that is N(u) represents the set of the first-order sampled neighbor nodes of node u. Pooling aggregation is performed on the features between the neighbor nodes, and the formula of the pooling aggregation operator is as follows:

[0038]

[0039] Among them, represents the vector value of the sampled first-order neighbor node v i , max represents obtaining the maximum vector value of the vector value set , i ∈ {1, …, m}, m represents the number of first-order sampled neighbor nodes of the central node u, and the calculation formula for the new vector value of the central node u is as follows:

[0040] h u ' = σ{W1h u +W2Pool({h v})+b}

[0041] Among them, W1 and W2 are the weight parameter values learned by the aggregation operation, b is the bias parameter value learned by the aggregation operation, h u is the original feature vector value of the node, h u ' represents the new vector value of node u obtained after calculation, and σ represents the activation function;

[0042] Step 4-4 Each disambiguation feature map G j (j = 1, 2, 3, 4, 5) passes through the output layer of GraphSAGE, and the output vector x j is input into the perceptron linear classifier, and the perceptron linear classification model is used for semantic classification to obtain the semantic classification result The perceptron is expressed as a mapping function from the input space to the output space, as follows:

[0043]

[0044] Among them, w and b are the perceptron model parameters, w ∈ R n is the weight vector, b ∈ R is the bias, wx j represents the inner product of w and x j , and sign is the sign function, and its form is as follows:

[0045]

[0046] Step 4-5 Count the number of votes obtained by semantic category c1 Count the number of votes obtained by semantic category c2 The calculation formula is as follows:

[0047] If then

[0048] If then

[0049] Select the semantic category with the most votes as the predicted semantic category Y of the ambiguous word w predict :

[0050]

[0051] where count represents the vote count.

[0052] In steps 4-6, the cross-entropy loss function is used to calculate the error loss between the predicted label category and the actual label category, and the calculation formula is as follows:

[0053]

[0054] where y i represents the true label, is the predicted label, and N is the number of training sentences;

[0055] In step 4-7, according to the error loss, backpropagation is performed to update the parameters layer by layer, and the parameter update process is as follows:

[0056]

[0057] where θ represents the parameter set, θ' represents the updated parameter set, and α is the learning rate;

[0058] In step 4-8, continuously iterate steps 4-1 to 4-7 until the set number of training times is reached to obtain an optimized multi-GraphSAGE neural network;

[0059] 6. The graph sampling aggregation multi-vote word sense disambiguation method for embedding semantic similarity according to claim 1, characterized in that in the step 5, semantic classification of the ambiguous word w is performed, and the specific steps are as follows:

[0060] In step 5-1, the sentence-word form disambiguation feature graph constructed from the test data is input into the optimized GraphSAGE_1, the sentence-part-of-speech disambiguation feature graph constructed from the test data is input into the optimized GraphSAGE_2, the sentence-semantic category disambiguation feature graph constructed from the test data is input into the optimized GraphSAGE_3, the sentence-translation disambiguation feature graph constructed from the test data is input into the optimized GraphSAGE_4, and the sentence-Wubi coding disambiguation feature graph constructed from the test data is input into the optimized GraphSAGE_5;

[0061] Step 5-2 performs neighbor multi-order sampling on each disambiguation feature map. First, starting from the source node, the first-order neighbor nodes of the source node are randomly sampled according to the first-order sampling number to obtain a list of the first-order neighbor nodes of the source node. Taking the obtained first-order neighbor nodes of the source node as the starting points, they are randomly sampled for the first-order neighbors according to the node sampling number, that is, the second-order neighbor nodes of the source node are randomly sampled according to the sampling number to obtain a list of the second-order neighbor nodes of the source node. When the number of the first-order neighbor nodes of a certain node is less than the specified sampling number, sampling with replacement will be performed, and duplicate nodes will appear in the sampling results;

[0062] Step 5-3 performs an aggregation operation on the sampled neighbor node features. Node v is the first-order sampled neighbor node of the central node u, that is N(u) represents the set of the first-order sampled neighbor nodes of node u, and pooling aggregation is performed on the features between the neighbor nodes. The formula of the pooling aggregation operator is as follows:

[0063]

[0064] where represents the vector value of the sampled first-order neighbor node v i and max represents obtaining the maximum vector value of the vector value set . i ∈ {1, …, m}, where m represents the number of the first-order sampled neighbor nodes of the central node u. The calculation formula of the new vector value of the central node u is as follows:

[0065] h u ' = σ{W1h u + W2Pool({h v}) + b}

[0066] where W1 and W2 are the weight parameter values learned by the aggregation operation, b is the bias parameter value learned by the aggregation operation, h u is the original feature vector value of the node, h u ' represents the new vector value of node u obtained after calculation, and σ represents the activation function;

[0067] Step 5-4 Each disambiguation feature map G j (j = 1, 2, 3, 4, 5) passes through the output layer of GraphSAGE, and the output vector x j is input into the perceptron linear classifier, and the perceptron linear classification model is used for semantic classification to obtain the semantic classification result The perceptron is expressed as a mapping function from the input space to the output space, as follows:

[0068]

[0069] where w and b are the perceptron model parameters, w ∈ Rn is the weight vector, b ∈ R is the bias, wx j represents the inner product of w and x j and sign is the sign function, which has the following form:

[0070]

[0071] Step 5-5: Count the number of votes obtained by semantic category c1 Count the number of votes obtained by semantic category c2 The calculation formula is as follows:

[0072] If then

[0073] If then

[0074] Select the semantic category with the most votes as the predicted semantic category Y of the ambiguous word w predict :

[0075]

[0076] where count represents the vote count.

[0077] Beneficial effects:

[0078] 1. The present invention is a graph sampling aggregation multi-vote word sense disambiguation method embedding semantic similarity. It performs lexical segmentation, part-of-speech tagging, semantic class tagging, translation tagging, and Wubi coding tagging on Chinese sentences. It uses the Word2Vec tool and the Doc2Vec tool to vectorize the disambiguation features. The extracted disambiguation features have high quality.

[0079] 2. The model used in the present invention is the GraphSAGE neural network. The biggest feature is to iterate node features with the help of the graph structure. Each node only samples a part of its neighbor nodes to iterate and update its own features. By constructing five disambiguation feature graphs and using multiple GraphSAGE neural networks, better classification results can be obtained.

[0080] 3. The classifier used in the present invention is the perceptron linear classifier. The voting mechanism is adopted to obtain the classification result, which can enhance the word sense disambiguation effect.

[0081] 4. The present invention uses the "Thesaurus of Chinese Synonyms" to obtain the semantic similarity between the target ambiguous word and its left and right adjacent words and embeds it into the disambiguation feature vector, which can better obtain disambiguation features from around the ambiguous word.

[0082] 5. When training the model, the gradient descent method is used to update the weight matrix parameters in the aggregation layer of the model. The error is calculated through the cross-entropy loss function, and the gradient descent is used to update the model parameters to obtain an optimized multi-GraphSAGE neural network, which improves the disambiguation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS:

[0083] Figure 1 It is a flowchart of graph sampling aggregation multi-vote word sense disambiguation with embedded semantic similarity in the embodiment of the present invention;

[0084] Figure 2 It is a sentence-morphology disambiguation feature graph constructed from test sentences in the embodiment of the present invention;

[0085] Figure 3 It is a sentence-part-of-speech disambiguation feature graph constructed from test sentences in the embodiment of the present invention;

[0086] Figure 4 It is a sentence-semantic class disambiguation feature graph constructed from test sentences in the embodiment of the present invention;

[0087] Figure 5 It is a sentence-translation disambiguation feature graph constructed from test sentences in the embodiment of the present invention;

[0088] Figure 6 It is a sentence-Wubi coding disambiguation feature graph constructed from test sentences in the embodiment of the present invention;

[0089] Figure 7 It is the training process of the graph sampling aggregation multi-vote word sense disambiguation model with embedded semantic similarity in the embodiment of the present invention;

[0090] Figure 8 It is the testing process of the graph sampling aggregation multi-vote word sense disambiguation model with embedded semantic similarity in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION:

[0091] In order to clearly and completely describe the technical solutions in the embodiments of the present invention, taking the test sentence "Although the cigarette dealers have very powerful connections, the law enforcement officers in this case did not waver in the face of sentiment, reason, and law, and insisted on transferring the parties to the judicial authorities." containing the ambiguous word "waver" as an example, and in combination with the accompanying drawings in the embodiments, the present invention will be further described in detail. The ambiguous word "waver" has two semantic categories, c1: shake, c2: vacillate. There are 47 sentences in the training corpus of "waver", and 16 sentences in the testing corpus of "waver".

[0092] The flowchart of graph sampling aggregation multi-vote word sense disambiguation with embedded semantic similarity in the embodiment of the present invention is as Figure 1 shown and includes the following steps.

[0093] Step 1 processes the Chinese sentence containing the ambiguous word "dongyao" and extracts disambiguation features as follows:

[0094] Step 1-1 uses a Chinese word segmentation tool to segment the Chinese sentence as follows:

[0095] Word segmentation result: Although the law enforcement officers handling this case, despite the fact that the cigarette dealers have very powerful connections, did not waver in the face of the law, reason, and human feelings and insisted on transferring the parties to the judicial authorities

[0096] Step 1-2 uses a Chinese part-of-speech tagging tool to tag the segmented words as follows:

[0097] Part-of-speech tagging: Although / c cigarette / n dealer / n relationship / n very / d hard / a this / r bureau / n of / u handling / vn personnel / n at / p sentiment / n reason / n law / n before / f slightest / m not / d waver / v insist / v take / p party / n transfer / v judicial / n authority / n

[0098] Step 1-3 uses a Chinese semantic class tagging tool to tag the segmented words as follows:

[0099] Semantic class tagging: Although / c / Kc10 cigarette / n / Bf05 dealer / n / Ae09 relationship / n / Db01 very / d / Ka01 hard / a / Ee12 this / r / Ed61 bureau / n / Dn08 of / u / Bo29 handling / vn / Hj12 personnel / n / Aa01 at / p / Jd02 sentiment / n / Df07 reason / n / Hc02 law / n / Di02 before / f / Cb04 slightest / m / Eb01 not / d / Jd08 waver / v / Je02 insist / v / Gb02 take / p / Hc10 party / n / Aa01 transfer / v / Hi27 judicial / n / Di25 authority / n / Da01

[0100] Step 1-4 uses a Chinese machine translation tool to annotate the translated text of the segmented words as follows:

[0101] Translation note: Although the relationship between the cigarette dealers is very strong, the personnel handling the case of this bureau did not waver in the slightest in the face of affection, reason, and law, and persisted in transferring the party to the judicial organs.

[0102] Steps 1-5 use the Chinese character Wubi coding annotation tool to annotate the segmented words as follows:

[0103] Wubi coding annotation: although / c / Kc10 / although / nytpcigarette / n / Bf05 / cigarette / oldydealer / n / Ae09 / dealer / mrbbrelationship / n / Db01 / relation / udtxvery / d / Ka01 / very / tveyhard / a / Ee12 / hard / dgjqthis / r / Ed61 / this / ypwhbureau / n / Dn08 / bureau / nnkdof / u / Bo29 / of / rqyyhandle the case / vn / Hj12 / handle the case / lwpvpersonnel / n / Aa01 / personnel / wwkmin / p / Jd02 / in / dhfdaffection / n / Df07 / affection / ngegreason / n / Hc02 / reason / gjfglaw / n / Di02 / law / ifcyin front / f / Cb04 / in front of / dmue the slightest / m / Eb01 / slightest / xxyp no / d / Jd08 / no / imde waver / v / Je02 / waver / fcre insist / v / Gb02 / persist in / jcrf put / p / Hc10 / bundle / rcn the party / n / Aa01 / party / igww transfer / v / Hi27 / transfer / tquq justice / n / Di25 / justice / ngif organ / n / Da01 / organ / smud

[0104] Step 1-6 selects the sentence where the ambiguous word is located, and the word form, part of speech, semantic class, translation and Wubi code of the two adjacent lexical units on the left and right of the ambiguous word are used as disambiguation features, as follows:

[0105]

[0106] Step 2: Get training data and test data.

[0107] Step 2-1 calculates the semantic similarity between the ambiguous word and the adjacent words, and generates the semantic similarity of the adjacent words and the semantic similarity of the sentence. Consult the "Synonym Dictionary" to obtain the semantic class coding of the ambiguous word "摇" and the two adjacent words on the left and right: "有点", "没", "坚持", "把", and calculate the semantic similarity between the target ambiguous word "摇" and the two adjacent words on the left and right: "有点", "没", "坚持", "把", and the calculation results are as follows:

[0108] Q L2 =sim(shake, slightest)=0.1,Q L1 =sim(shake, no)=0.202148,

[0109] QR1 = sim(shake, adhere) = 0.202635, Q R2 = sim(shake, put) = 0.1,

[0110] Q s = Q L2 + Q L1 + Q R1 + Q R2 = 0.604783

[0111] Step 2-2 Use the Doc2Vec tool to vectorize the extracted sentence features, and use the Word2Vec tool to vectorize the extracted word form, part of speech, semantic class, translation, and Wubi code features respectively. Each disambiguation feature corresponds to a 200-dimensional feature vector. The results of vectorizing the features extracted from the test sentence "Despite the strong connections of the tobacco dealers, the law enforcement officers in this bureau did not waver in the face of sentiment, reason, and law, and adhered to transferring the parties to the judicial organs." are as follows:

[0112]

[0113]

[0114] Step 2-3 Embed the semantic similarity Q into the sentence feature vector s , and embed the semantic similarities Q of adjacent words in the word form, part of speech, semantic class, translation, and Wubi code of the lexical units on both sides of "shake" respectively L2 , Q L1 , Q R1 , Q R2 , Q. The results of embedding the semantic similarity into the feature vector of the test sentence "Despite the strong connections of the tobacco dealers, the law enforcement officers in this bureau did not waver in the face of sentiment, reason, and law, and adhered to transferring the parties to the judicial organs." are as follows:

[0115]

[0116]

[0117] Step 2-4 Use the processed training corpus in SemEval-2007: Task#5 as the training data, and use the processed test corpus in SemEval-2007: Task#5 as the test data;

[0118] Step 3 Construct five disambiguation feature maps of "Despite the strong connections of the tobacco dealers, the law enforcement officers in this bureau did not waver in the face of sentiment, reason, and law, and adhered to transferring the parties to the judicial organs.", such as Figure 2 , Figure 3 , Figure 4 , Figure 5and Figure 6 As shown below, the specific steps are as follows:

[0119] Step 3-1: Take the sentence with the ambiguous word "dongyao" (waver) and the word forms of the two adjacent lexical units on the left and right as nodes in the sentence-word form disambiguation feature graph, take the sentence with "dongyao" and the part-of-speech tags of the two adjacent lexical units on the left and right as nodes in the sentence-part-of-speech disambiguation feature graph, take the sentence with "dongyao" and the semantic classes of the two adjacent lexical units on the left and right as nodes in the sentence-semantic class disambiguation feature graph, take the sentence with "dongyao" and the translations of the two adjacent lexical units on the left and right as nodes in the sentence-translation disambiguation feature graph, and take the sentence with "dongyao" and the five-stroke codes of the two adjacent lexical units on the left and right as nodes in the sentence-five-stroke code disambiguation feature graph;

[0120] Step 3-2: Use the feature vectors of each disambiguation feature obtained in Step 2 to perform vector embedding on the sentence nodes and word form nodes in the sentence-word form disambiguation feature graph, the sentence nodes and part-of-speech nodes in the sentence-part-of-speech disambiguation feature graph, the sentence nodes and semantic class nodes in the sentence-semantic class disambiguation feature graph, the sentence nodes and translation nodes in the sentence-translation disambiguation feature graph, and the sentence nodes and five-stroke code nodes in the sentence-five-stroke code disambiguation feature graph;

[0121] Step 3-3: Establish the edge relationship between the sentence nodes and word form nodes in the sentence-word form disambiguation feature graph according to the number of times the word form appears in the sentence, establish the edge relationship between the sentence nodes and part-of-speech nodes in the sentence-part-of-speech disambiguation feature graph according to the number of times the part-of-speech appears in the sentence, establish the edge relationship between the sentence nodes and semantic class nodes in the sentence-semantic class disambiguation feature graph according to the number of times the semantic class appears in the sentence, establish the edge relationship between the sentence nodes and semantic class nodes in the sentence-translation disambiguation feature graph according to the number of times the translation appears in the sentence, and establish the edge relationship between the sentence nodes and five-stroke code nodes in the sentence-five-stroke code disambiguation feature graph according to the number of times the five-stroke code appears in the sentence;

[0122] Step 4: Use the training data to train and optimize the multi-GraphSAGE neural network, as Figure 7 shown below. The specific steps are as follows:

[0123] Step 4-1 Input the sentence-word form disambiguation feature map constructed from the training data of "waver" into the initialized GraphSAGE_1, input the sentence-part of speech disambiguation feature map constructed from the training data of "waver" into the initialized GraphSAGE_2, input the sentence-semantic class disambiguation feature map constructed from the training data of "waver" into the initialized GraphSAGE_3, input the sentence-translation disambiguation feature map constructed from the training data of "waver" into the initialized GraphSAGE_4, and input the sentence-Wubi coding disambiguation feature map constructed from the training data of "waver" into the initialized GraphSAGE_5, where GraphSAGE_1, GraphSAGE_2, GraphSAGE_3, GraphSAGE_4, and GraphSAGE_5 are neural networks with the same parameter initialization;

[0124] Step 4-2 Perform multi-order neighbor sampling on each disambiguation feature map. First, starting from the source node, randomly sample the first-order neighbor nodes of the source node according to the first-order sampling number to obtain the list of the first-order neighbor nodes of the source node. Use the obtained first-order neighbor nodes of the source node as the starting points and randomly sample their first-order neighbors according to the node sampling number, that is, randomly sample the second-order neighbor nodes of the source node according to the sampling number to obtain the list of the second-order neighbor nodes of the source node. If the number of the first-order neighbor nodes of a certain node is less than the specified sampling number, sampling with replacement will be performed, and duplicate nodes will appear in the sampling results;

[0125] Step 4-3 Perform an aggregation operation on the features of the sampled neighbor nodes. Node v is the first-order sampled neighbor node of the central node u, that is N(u) represents the set of the first-order sampled neighbor nodes of node u. Pooling aggregation is performed on the features between the neighbor nodes. The formula of the pooling aggregation operator is as follows:

[0126]

[0127] where represents the vector value of the sampled first-order neighbor node v i max represents obtaining the maximum vector value of the vector value set i ∈ {1,…,m}, m represents the number of the first-order sampled neighbor nodes of the central node u. The calculation formula of the new vector value of the central node u is as follows:

[0128] h u ' = σ{W1h u +W2Pool({h v})+b}

[0129] where W1 and W2 are the weight parameter values learned by the aggregation operation, b is the bias parameter value learned by the aggregation operation, hu is the original feature vector value of the node, h u ' represents the new vector value of node u obtained after calculation, and σ represents the activation function;

[0130] In step 4-4, each disambiguation feature map G j (j = 1, 2, 3, 4, 5) passes through the output layer of GraphSAGE, and the output vector x j is input into the perceptron linear classifier, and the perceptron linear classification model is used for semantic classification to obtain the semantic classification result The perceptron is represented as a mapping function from the input space to the output space, as follows:

[0131]

[0132] where w and b are the perceptron model parameters, w ∈ R n is the weight vector, b ∈ R is the bias, and wx j represents the inner product of w and x j , and sign is the sign function, and its form is as follows:

[0133]

[0134] In step 4-5, count the number of votes obtained by the semantic category c1 = shake Count the number of votes obtained by the semantic category c2 = vacillate The calculation formula is as follows:

[0135] If Then

[0136] If Then

[0137] Select the semantic category with the most votes as the predicted semantic category Y of the ambiguous word "shake" predict :

[0138]

[0139] where count represents the vote count.

[0140] In step 4-6, use the cross-entropy loss function to calculate the error loss between the predicted label category and the actual label category 动摇 :

[0141] loss 动摇 = 0.1675

[0142] In step 4-7, according to the error loss 动摇Backpropagation, update the parameters layer by layer, and the parameter update process is as follows:

[0143]

[0144] Among them, θ 动摇 represents the parameter set, θ' 动摇 represents the updated parameter set, and α is the learning rate;

[0145] Steps 4-8 continuously iterate Steps 4-1 to 4-7 until the set number of training times is reached, and an optimized multi-GraphSAGE neural network is obtained;

[0146] The testing process in Step 5 is as Figure 8 shown. For semantic classification of the ambiguous word "dongyao", the specific steps are as follows:

[0147] Step 5-1 Input the sentence-word form disambiguation feature graph constructed from the test data of "dongyao" into the optimized GraphSAGE_1, input the sentence-pos disambiguation feature graph constructed from the test data of "dongyao" into the optimized GraphSAGE_2, input the sentence-semantic class disambiguation feature graph constructed from the test data of "dongyao" into the optimized GraphSAGE_3, input the sentence-translation disambiguation feature graph constructed from the test data of "dongyao" into the optimized GraphSAGE_4, and input the sentence-Wubi coding disambiguation feature graph constructed from the test data of "dongyao" into the optimized GraphSAGE_5;

[0148] Step 5-2 Perform neighbor multi-order sampling on each disambiguation feature graph. First, starting from the source node, randomly sample the first-order neighbor nodes of the source node according to the first-order sampling number to obtain the list of the first-order neighbor nodes of the source node. Take the obtained first-order neighbor nodes of the source node as the starting point, and randomly sample their first-order neighbors according to the node sampling number, that is, randomly sample the second-order neighbor nodes of the source node according to the sampling number to obtain the list of the second-order neighbor nodes of the source node. If the number of first-order neighbor nodes of a certain node is less than the specified sampling quantity, sampling with replacement will be performed, and duplicate nodes will appear in the sampling results;

[0149] Step 5-3 Perform an aggregation operation on the features of the sampled neighbor nodes. Node v is the first-order sampled neighbor node of the central node u, that is N(u) represents the set of the first-order sampled neighbor nodes of node u. Perform pooling aggregation on the features between neighbor nodes. The formula of the pooling aggregation operator is as follows:

[0150]

[0151] Among them, represents the sampled first-order neighbor node vi The vector value, and max represents obtaining the set of vector values The maximum vector value, where i ∈ {1, …, m}, and m represents the number of first-order sampled neighbor nodes of the central node u. The calculation formula for the new vector value of the central node u is as follows:

[0152] h u ' = σ{W1h u +W2Pool({h v})+b}

[0153] where W1 and W2 are the weight parameter values learned by the aggregation operation, b is the bias parameter value learned by the aggregation operation, h u is the original feature vector value of the node, h u ' represents the new vector value of node u obtained after calculation, and σ represents the activation function;

[0154] In step 5-4, each disambiguation feature map G j (j = 1, 2, 3, 4, 5) passes through the output layer of GraphSAGE, and the output vector x j is input into the perceptron linear classifier, and the perceptron linear classification model is used for semantic classification to obtain the semantic classification result

[0155]

[0156] In step 5-5, count the number of votes obtained by the semantic category c1 = shake Count the number of votes obtained by the semantic category c2 = vacillate That is: Select the semantic category with the most votes as the predicted semantic category Y of the ambiguous word "shake" predict :

[0157]

[0158] Using the optimized multi-GraphSAGE neural network, perform word sense disambiguation on the Chinese sentence "Despite the strong connections of the cigarette dealers, the case-handling officers in this case did not waver in the face of sentiment, reason, and law, and insisted on transferring the parties to the judicial authorities." containing the ambiguous word "shake". The semantic category corresponding to the ambiguous word "shake" is shake.

[0159] Using the optimized multi-GraphSAGE neural network to disambiguate the test corpus containing the ambiguous word "shake", the accuracy rate accuracy of word sense disambiguation is:

[0160]

[0161] The graph sampling aggregation multi-voting word sense disambiguation with embedded semantic similarity in the embodiments of the present invention can select diverse and accurate disambiguation features. By constructing multiple disambiguation feature graphs and embedding the semantic similarity of disambiguation feature vectors, and integrating multiple GraphSAGE neural networks to determine the semantic category of ambiguous words, it has a high accuracy rate.

[0162] The above is a detailed introduction to the embodiments of the present invention in conjunction with the accompanying drawings. The specific implementation manners herein are only used to help understand the method of the present invention. For those of ordinary skill in the art, according to the idea of the present invention, changes and modifications can be made within the specific implementation manners and application scope. Therefore, the present specification should not be construed as a limitation to the present invention.

Claims

1. A graph sampling aggregation multi-voting word sense disambiguation method embedded with semantic similarity, characterized in that The method mainly includes the following steps: Step 1: Perform word segmentation, part-of-speech tagging, semantic class tagging, translation annotation, and Wubi code annotation on all Chinese sentences contained in the SemEval-2007: Task#5 corpus. Select the sentences where the ambiguous words are located, as well as the word form, part-of-speech, semantic class, translation, and Wubi code of the two adjacent lexical units on the left and right of the ambiguous words as disambiguation features; Step 2: Use the "Thesaurus of Chinese Synonyms" to calculate the semantic similarity between the target ambiguous word and the adjacent words, generate the semantic similarity weights. Use the Doc2Vec tool to vectorize the extracted sentence features, use the Word2Vec tool to vectorize the extracted word form, part-of-speech, semantic class, translation, and Wubi code features, and embed the semantic similarity weights into the feature vectors. Use the processed training corpus in SemEval-2007: Task#5 as the training data, and use the processed test corpus in SemEval-2007: Task#5 as the test data; Step 3: Take the extracted sentences, as well as the word form, part-of-speech, semantic class, translation, and Wubi code of the two adjacent lexical units on the left and right of the ambiguous word as the nodes of the disambiguation feature graph, and construct a sentence-word form disambiguation feature graph, a sentence-part-of-speech disambiguation feature graph, a sentence-semantic class disambiguation feature graph, a sentence-translation disambiguation feature graph, and a sentence-Wubi code disambiguation feature graph respectively; Step 4: Training process, input the five disambiguation feature graphs constructed from the training data into the multi-GraphSAGE neural network and optimize it to obtain the optimized multi-GraphSAGE neural network; Step 5: Testing process, input the five disambiguation feature graphs constructed from the test data into the optimized multi-GraphSAGE neural network, obtain the semantic class of each disambiguation feature graph through the perceptron linear classifier, and adopt a voting mechanism to determine the semantic class of the ambiguous word.

2. The graph sampling aggregation multi-voting word sense disambiguation method with embedded semantic similarity according to claim 1, characterized in that In the above Step 1, when performing word segmentation, part-of-speech tagging, semantic class tagging, translation annotation, and Wubi code annotation on the Chinese sentence containing the ambiguous word w and extracting the disambiguation features, the specific steps are as follows: Step 1-1: Use a Chinese word segmentation tool to segment the Chinese sentence into words; Step 1-2: Use a Chinese part-of-speech tagging tool to tag the part-of-speech of the segmented words; Step 1-3: Use a Chinese semantic class tagging tool to tag the semantic class of the segmented words; Step 1-4: Use a Chinese machine translation tool to annotate the translation of the segmented words; Step 1-5: Use a Chinese Wubi code annotation tool to annotate the Wubi code of the segmented words; Step 1-6: Select the sentence where the ambiguous word is located, as well as the word form, part-of-speech, semantic class, translation, and Wubi code of the two adjacent lexical units on the left and right of the ambiguous word as the disambiguation features.

3. The graph sampling aggregation multi-voting word sense disambiguation method embedding semantic similarity according to claim 1, characterized in that In the above Step 2, when calculating the semantic similarity between the target ambiguous word and the adjacent words, vectorizing the sentence features, vectorizing the word form, part-of-speech, semantic class, translation, and Wubi code features and embedding the semantic similarity weights, and obtaining the training data and test data, the specific steps are as follows: Step 2-1: Consult the "Thesaurus of Chinese Synonyms" to obtain the semantic class codes of the target ambiguous word w and its two adjacent words on the left and right. Calculate the semantic similarity between the target ambiguous word w and the adjacent words, and select the maximum similarity value as the semantic similarity sim between the ambiguous word w and the adjacent words. Obtain the semantic similarities Q of the target ambiguous word w with the four adjacent words respectively L2 , Q L1 , Q R1 , Q R2 . Take Q L2 + Q L1 + Q R1 + Q R2 as the semantic similarity Q between the target ambiguous word w and the sentence s . The formula for calculating the semantic similarity of words w1 and w2 is as follows: If the first layer of the semantic class encodings of w1 and w2 is different, then sim(w1, w2) = 0.1 If the first layer of the semantic class encodings of w1 and w2 is the same and the second layer is different, then If the first layer, the second layer of the semantic class encodings of w1 and w2 are the same, and the third layer is different, then If the first layer, the second layer, and the third layer of the semantic class encodings of w1 and w2 are the same, and the fourth layer is different, then If the first layer, the second layer, the third layer, the fourth layer of the semantic class encodings of w1 and w2 are the same, and the fifth layer is different, then Among them, l represents the shortest distance between words w1 and w2, l = (5 - the number of common levels) × 2, n is the total number of nodes in the branch layer, α and β are coefficients, and β1 < β2 < β3 < β4; In step 2-2, the Doc2Vec tool is used to vectorize the extracted sentence features, and the Word2Vec tool is used to vectorize the extracted word form, part of speech, semantic class, translation, and Wubi code features respectively. Each disambiguation feature corresponds to a 200-dimensional feature vector; Step 2-3 embeds the semantic similarity Q into the sentence feature vector s , respectively embed the semantic similarity Q of adjacent words into the word form, part of speech, semantic class, translation, and Wubi code in the lexical units on both sides of the ambiguous word L2 , Q L1 , Q R1 , Q R2; In step 2-4, the processed training corpus in SemEval-2007:Task#5 is used as the training data, and the processed test corpus in SemEval-2007:Task#5 is used as the test data.

4. The graph sampling aggregation multi-voting word sense disambiguation method with embedded semantic similarity according to claim 1, characterized in that In step 3, five disambiguation feature graphs are constructed. The specific steps are as follows: In step 3-1, the sentence with the ambiguous word w and the word forms of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-word form disambiguation feature graph; the sentence with w and the parts of speech of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-part of speech disambiguation feature graph; the sentence with w and the semantic classes of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-semantic class disambiguation feature graph; the sentence with w and the translations of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-translation disambiguation feature graph; the sentence with w and the Wubi codes of the two adjacent lexical units on the left and right of w are used as the nodes in the sentence-Wubi code disambiguation feature graph; In step 3-2, the feature vectors of the disambiguation features obtained in step 2 are used to perform vector embedding on the sentence nodes and word form nodes in the sentence-word form disambiguation feature graph, the sentence nodes and part of speech nodes in the sentence-part of speech disambiguation feature graph, the sentence nodes and semantic class nodes in the sentence-semantic class disambiguation feature graph, the sentence nodes and translation nodes in the sentence-translation disambiguation feature graph, and the sentence nodes and Wubi code nodes in the sentence-Wubi code disambiguation feature graph; In step 3-3, the edge relationship between the sentence nodes and the word form nodes in the sentence-word form disambiguation feature graph is established according to the number of times the word form appears in the sentence, the edge relationship between the sentence nodes and the part of speech nodes in the sentence-part of speech disambiguation feature graph is established according to the number of times the part of speech appears in the sentence, the edge relationship between the sentence nodes and the semantic class nodes in the sentence-semantic class disambiguation feature graph is established according to the number of times the semantic class appears in the sentence, the edge relationship between the sentence nodes and the translation nodes in the sentence-translation disambiguation feature graph is established according to the number of times the translation appears in the sentence, and the edge relationship between the sentence nodes and the Wubi code nodes in the sentence-Wubi code disambiguation feature graph is established according to the number of times the Wubi code appears in the sentence.

5. The graph sampling aggregation multi-voting word sense disambiguation method with embedded semantic similarity according to claim 1, characterized in that, In step 4, the multi-GraphSAGE neural network is trained and optimized. The specific steps are as follows: Step 4-1 Input the sentence-morphological disambiguation feature map constructed from the training data into the initialized GraphSAGE_1, input the sentence-positional disambiguation feature map constructed from the training data into the initialized GraphSAGE_2, input the sentence-semantic class disambiguation feature map constructed from the training data into the initialized GraphSAGE_3, input the sentence-translation disambiguation feature map constructed from the training data into the initialized GraphSAGE_4, and input the sentence-Wubi code disambiguation feature map constructed from the training data into the initialized GraphSAGE_5. Among them, GraphSAGE_1, GraphSAGE_2, GraphSAGE_3, GraphSAGE_4, and GraphSAGE_5 are neural networks with the same parameter initialization; Step 4-2 Perform multi-order neighbor sampling on each disambiguation feature map. First, starting from the source node, randomly sample the first-order neighbor nodes of the source node according to the first-order sampling number to obtain the list of the first-order neighbor nodes of the source node. Take the obtained first-order neighbor nodes of the source node as the starting points and randomly sample their first-order neighbors according to the node sampling number, that is, randomly sample the second-order neighbor nodes of the source node according to the sampling number to obtain the list of the second-order neighbor nodes of the source node. If the number of the first-order neighbor nodes of a certain node is less than the specified sampling quantity, sampling with replacement will be performed, and duplicate nodes will appear in the sampling results; Step 4-3 performs an aggregation operation on the sampled neighbor node features. Node v is a first-order sampled neighbor node of the central node u, that is N(u) represents the set of first-order sampled neighbor nodes of node u. Pooling aggregation is performed on the features between neighbor nodes. The formula of the pooling aggregation operator is as follows: Among them, represents the vector value of the sampled first-order neighbor node v i , max represents obtaining the maximum vector value of the vector value set , i ∈ {1, …, m}, m represents the number of first-order sampled neighbor nodes of the central node u, and the calculation formula for the new vector value of the central node u is as follows: h u ' = σ{W1h u +W2Pool({h v})+b} Among them, W1 and W2 are the weight parameter values learned by the aggregation operation, b is the bias parameter value learned by the aggregation operation, and h u is the original feature vector value of the node, and h u ' represents the new vector value of node u obtained after calculation, and σ represents the activation function; Step 4-4 Each disambiguation feature map G j (j = 1, 2, 3, 4, 5) Pass through the output layer of GraphSAGE, and output the vector x j Input it into the perceptron linear classifier, and use the perceptron linear classification model for semantic classification to obtain the semantic classification result The perceptron is expressed as a mapping function from the input space to the output space, as follows: where \(w\) and \(b\) are the parameters of the perceptron model, \(w\in\mathbb{R}\) n is the weight vector, \(b\in\mathbb{R}\) is the bias, \(wx\) j represents the inner product of \(w\) and \(x\) j and sign is the sign function, whose form is as follows: Step 4-5: Count the number of votes obtained by semantic category c1 Count the number of votes obtained by semantic category c2 The calculation formula is as follows: If then If then Select the semantic category with the most votes as the predicted semantic category Y of the ambiguous word w predict : Among them, count represents the vote count; Step 4-6 Use the cross-entropy loss function to calculate the error loss between the predicted label category and the actual label category. The calculation formula is as follows: Among them, y i represents the true label, is the predicted label, and N is the number of training sentences; Step 4-7 Backpropagate according to the error loss and update the parameters layer by layer. The parameter update process is as follows: Among them, θ represents the parameter set, θ' represents the updated parameter set, and α is the learning rate; Step 4-8 Continuously iterate Steps 4-1 to 4-7 until the set number of training times is reached to obtain the optimized multi-GraphSAGE neural network.

6. The graph sampling aggregation multi-voting word sense disambiguation method embedding semantic similarity according to claim 1, characterized in that, In Step 5, perform semantic classification on the ambiguous word w. The specific steps are as follows: Step 5-1 Input the sentence-morphological disambiguation feature map constructed from the test data into the optimized GraphSAGE_1, input the sentence-positional disambiguation feature map constructed from the test data into the optimized GraphSAGE_2, input the sentence-semantic class disambiguation feature map constructed from the test data into the optimized GraphSAGE_3, input the sentence-translation disambiguation feature map constructed from the test data into the optimized GraphSAGE_4, and input the sentence-Wubi code disambiguation feature map constructed from the test data into the optimized GraphSAGE_5; Step 5-2 performs neighbor multi-order sampling on each disambiguation feature map. First, starting from the source node, randomly sample the first-order neighbor nodes of the source node according to the first-order sampling number to obtain a list of the first-order neighbor nodes of the source node. Take the obtained first-order neighbor nodes of the source node as the starting points and randomly sample their first-order neighbors according to the node sampling number, that is, randomly sample the second-order neighbor nodes of the source node according to the sampling number to obtain a list of the second-order neighbor nodes of the source node. If the number of first-order neighbor nodes of a certain node is less than the specified sampling quantity, sampling with replacement will be performed, and duplicate nodes will appear in the sampling results; Step 5-3 performs an aggregation operation on the sampled neighbor node features. Node v is a first-order sampled neighbor node of the central node u, that is N(u) represents the set of first-order sampled neighbor nodes of node u. Pooling aggregation is performed on the features between neighbor nodes. The formula for the pooling aggregation operator is as follows: Among them, represents the vector value of the sampled first-order neighbor node v i , max represents obtaining the maximum vector value of the vector value set , i ∈ {1, …, m}, m represents the number of first-order sampled neighbor nodes of the central node u, and the calculation formula for the new vector value of the central node u is as follows: h u ' = σ{W1h u +W2Pool({h v}) + b} Among them, W1 and W2 are the weight parameter values learned by the aggregation operation, b is the bias parameter value learned by the aggregation operation, and h u is the original feature vector value of the node, and h u ' represents the new vector value of node u obtained after calculation, and σ represents the activation function; Step 5-4 Each disambiguation feature map G j (j = 1, 2, 3, 4, 5) Pass through the output layer of GraphSAGE, and the output vector x j is input into the perceptron linear classifier, and the semantic classification is performed using the perceptron linear classification model to obtain the semantic classification result The perceptron is expressed as a mapping function from the input space to the output space, as follows: where w and b are the parameters of the perceptron model, w ∈ R n is the weight vector, b ∈ R is the bias, and wx j represents the inner product of w and x j and sign is the sign function, whose form is as follows: Step 5-5: Count the number of votes obtained by semantic category c1 Count the number of votes obtained by semantic category c2 The calculation formula is as follows: If then If then Select the semantic category with the most votes as the predicted semantic category Y of the ambiguous word w predict : Among them, count represents the vote count.

Citation Information

Patent Citations

  • Chinese word sense disambiguation method based on graph convolutional neural network

    CN113095087A

  • Biomedicine English word sense disambiguation method based on graph attention neural network

    CN114186553A