Method and device for measuring word semantic similarity by integrating SAO and Bayesian model

By integrating SAO structure and Bayesian model, the SAO structure of words is extracted from the corpus and combined with the WordNet knowledge base, the accuracy and reliability of word semantic similarity measurements are solved, and more accurate word similarity calculations are achieved.

CN119692358BActive Publication Date: 2025-07-18XIAMEN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510207265.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-18
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing method of semantic similarity measurement of word has limitations in accuracy and reliability, and it is difficult to meet the multi-field application needs of natural language processing technology.

Method used

Combining SAO structure extraction and Bayesian model, the semantic similarity parameters are obtained by extracting the SAO structure of words from the corpus, and using the WordNet knowledge base to adjust the results, and finally calculating the semantic similarity of words in the Bayesian model.

Benefits of technology

It improves the accuracy and reliability of word semantic similarity measurements, can more comprehensively understand and quantify the similarity between words, and is suitable for the development of natural language processing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119692358B_ABST
    Figure CN119692358B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for measuring the semantic similarity of words by integrating the SAO and Bayesian models, which relates to the technical field of calculating the semantic similarity of words. This method is an innovative method for measuring the semantic similarity of words, which integrates the extraction of the SAO (subject-action-object) structure and the Bayesian model. First, the SAO structure is extracted from the text, and then the occurrence frequencies of similar words are statistically counted based on these structures, which is used as a measure of the similarity of words. At the same time, semantic knowledge bases such as WordNet are also utilized to obtain the semantic similarity parameters of words, and these parameters and statistical results are integrated through the Bayesian model to calculate the posterior probability of the semantic similarity of words. The innovation of this method lies in that it not only utilizes the advantages of the statistical-based method but also combines the knowledge-based method, providing a new perspective to understand and quantify the similarity between words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of word semantic similarity calculation, and particularly to a method and device for measuring word semantic similarity by integrating SAO and Bayesian models. Background Art

[0002] In the digital age, with the development of artificial intelligence and natural language processing technologies, accurate measurement of word semantic similarity in text has become particularly important. Such measurement has a profound impact on a variety of intelligent applications, including but not limited to text classification, knowledge retrieval, service matching and other fields, and is a key link in all of them. The core challenge of word semantic similarity measurement lies in how to accurately capture and quantify the subtle differences and connections between words, which is crucial for improving the efficiency of natural language processing (NLP) systems.

[0003] Currently in this field, although researchers have proposed various methods to solve the problem of word semantic similarity measurement, each of them has its own advantages and limitations. Specifically, knowledge-based methods rely on expert experience and predefined semantic relationships, which may lead to strong subjectivity of results and difficulty in adapting to new fields. Corpus-based methods rely on large-scale text data. Although the data is rich, they face the problems of data sparsity and noise interference, affecting the accuracy of measurement results. Word vector methods simplify the calculation process by converting words into points in a vector space, but this method may not be able to fully capture the deep semantic information of words. In addition, although multi-method fusion can integrate the advantages of various technologies, it has a high computational complexity to implement and requires solving the compatibility problem between different models. Briefly speaking, although these methods can provide useful similarity measurement results in specific situations, they still have limitations in practical applications. For example, knowledge-based methods may not cover all semantic relationships, while corpus-based methods may be limited by the quality and scope of available data. In addition, although word vector-based methods can provide fast similarity measurement in some cases, they may not be able to accurately reflect the complex semantic relationships between words.

[0004] Therefore, although many methods have been proposed to solve the problem of word semantic similarity measurement, this field still needs new methods and technologies to improve the accuracy and reliability of measurement. These new methods need to be able to overcome the limitations of existing technologies and provide more accurate and comprehensive semantic similarity measurement results to meet the growing needs of NLP applications.

[0005] In view of this, the present application is proposed. Summary of the Invention

[0006] The present invention provides a method and apparatus for measuring the semantic similarity of words by integrating the SAO and Bayesian models, which can at least partially improve the above problems.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for measuring the semantic similarity of words by integrating the SAO and Bayesian models, comprising:

[0009] Extracting the SAO structure of the words to be compared from a preset corpus based on semantic part-of-speech annotation, and performing statistical preprocessing on the SAO structure to obtain a statistical result;

[0010] Calculating the SAO similarity of words according to the statistical result to obtain an SAO similarity score;

[0011] Calculating the empirical similarity of the semantic similarity between the comparison word and the word to be compared according to a preset knowledge base to generate an empirical similarity, wherein the knowledge base is a WordNet knowledge base;

[0012] Substituting the SAO similarity score and the empirical similarity into the Bayesian formula to obtain the final calculation result of the semantic similarity of words.

[0013] The present invention also provides an apparatus for measuring the semantic similarity of words by integrating the SAO and Bayesian models, comprising:

[0014] An extraction and statistics unit for extracting the SAO structure of the words to be compared from a preset corpus based on semantic part-of-speech annotation, and performing statistical preprocessing on the SAO structure to obtain a statistical result;

[0015] An SAO similarity unit for calculating the SAO similarity of words according to the statistical result to obtain an SAO similarity score;

[0016] An empirical similarity unit for calculating the empirical similarity of the semantic similarity between the comparison word and the word to be compared according to a preset knowledge base to generate an empirical similarity, wherein the knowledge base is a WordNet knowledge base;

[0017] A Bayesian calculation unit for substituting the SAO similarity score and the empirical similarity into the Bayesian formula to obtain the final calculation result of the semantic similarity of words.

[0018] In summary, the method for measuring word semantic similarity by integrating SAO and Bayesian models proposes a model for measuring word semantic similarity by integrating SAO and Bayesian models. This method uses the SAO similarity of words as the basis for determining word similarity, and combines the knowledge base to obtain semantic consistency parameters, which can combine both the empirical and statistical methods for calculating word semantic similarity; and proposes a new similarity measure based on the Bayesian model. Specifically, taking the SAO co-occurrence statistics as the SAO similarity probability, fusing the semantic consistency parameters of the knowledge base, and calculating the posterior probability of word similarity. Through this research, it provides a new idea for judging word semantic similarity, improves the accuracy of similarity judgment, and provides a reference for solving the problem of similarity calculation. Briefly speaking, the method for measuring word semantic similarity by integrating SAO and Bayesian models can combine the statistical method based on the corpus and the method based on the knowledge base in order to obtain a more comprehensive and accurate similarity measure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 FIG. 6 is a schematic flowchart of the method for measuring word semantic similarity by integrating SAO and Bayesian models provided in the first embodiment of the present invention;

[0020] Figure 2 FIG. 10 is a schematic model flowchart of the method for measuring word semantic similarity by integrating SAO and Bayesian models provided in the embodiment of the present invention;

[0021] Figure 3 FIG. 14 is a schematic diagram of the wordnet structure provided in the embodiment of the present invention;

[0022] Figure 4 FIG. 18 is a schematic diagram of the Bayesian model of word semantic similarity provided in the embodiment of the present invention;

[0023] Figure 5 FIG. 22 is a schematic diagram of the node score distribution provided in the embodiment of the present invention;

[0024] Figure 6 FIG. 26 is a schematic diagram of the curve fitting situation provided in the embodiment of the present invention;

[0025] Figure 7 FIG. 30 is the calculation and comparison result of different methods provided in the embodiment of the present invention;

[0026] Figure 8 FIG. 34 is a schematic diagram of the modules of the device for measuring word semantic similarity by integrating SAO and Bayesian models provided in the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the objectives, technical solutions and advantages of the present invention more comprehensible and clear, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0028] Referring to Figure 1 As shown, the first embodiment of the present invention discloses a method for measuring the semantic similarity of words by integrating the SAO and Bayesian models, which can be executed by a device for measuring the semantic similarity of words by integrating the SAO and Bayesian models (hereinafter referred to as the measurement device), and specifically, by one or more processors in the measurement device to implement the following method:

[0029] The method for measuring the semantic similarity of words by integrating the SAO and Bayesian models is an innovative method for measuring the semantic similarity of words. Among them, semantic similarity measures the degree of substitutability of a word in the context of another word. The semantic relationships between words are mainly divided into: 1. Hyponymy relationship, 2. Synonymy relationship, 3. Antonymy relationship, 4. Inclusion relationship. Generally, the synonymy relationship, hyponymy relationship and inclusion relationship will obtain higher semantic similarity scores, while the antonymy relationship score is lower. However, in special cases, the semantic similarity of word pairs with antonymy relationships will show special values, and this situation will be analyzed according to cases in this embodiment.

[0030] The method for measuring the semantic similarity of words by integrating the SAO and Bayesian models provides a computational model for the semantic similarity of words by integrating the SAO and naive Bayesian models. This model is mainly divided into three parts; the first part is to extract the keyword SAO structure from the corpus by using semantic annotation, and to count the SAO structure similarity of the word pairs to be compared to obtain the SAO similarity score; the second part is to obtain the empirical similarity of the semantic similarity of the word pairs to be compared according to the knowledge base; the third part is to incorporate the SAO similarity score and the empirical similarity into the Bayesian model to obtain the final semantic similarity score. The specific implementation steps and the construction process of the Bayesian model for the semantic similarity of words are as Figure 2 shown.

[0031] Specifically:

[0032] S1. Extract the SAO structure of the words to be compared from the preset corpus based on semantic part-of-speech annotation, and perform statistical preprocessing on the SAO structure to obtain statistical results;

[0033] Specifically, step S1 includes: performing sentence splitting on the text in the preset corpus to obtain multiple sentences, and performing semantic part-of-speech annotation on the words in each sentence;

[0034] Extract the SAO structure of the sentence according to the part-of-speech of the labeled words, extract the SAO structures of the two words to be compared, perform lemmatization on the extracted SAO structures, and convert other word forms of the words to be compared into their prototypes.

[0035] Count the occurrence frequencies of the SAO structures corresponding to the two words to be compared in the text respectively, and compare them to obtain the SAO structures that co-occur in the two words to be compared and their occurrence frequencies, and generate the statistical results.

[0036] In this embodiment, the sentences in the corpus are segmented, and then the words in the sentences are further labeled with semantic part-of-speech, and the SAO of the sentences is extracted according to the part-of-speech of the words. Take the SAO structures of the two words to be compared. Since there are some words with multiple word forms in the extraction process, such as past tense, progressive tense, future tense or plural form, etc., it is necessary to perform lemmatization on the obtained SAO; for the convenience of statistics and understanding, it is also necessary to convert other word forms of the words into their prototypes.

[0037] Count the occurrence frequencies of the SAO corresponding to the words to be compared in the text respectively, compare the SAO structures of the two groups of words to be compared, and obtain the SAO structures and their frequencies that co-occur in the two groups. When counting the co-occurring predicates, in the case where the frequency of the general predicate is too large and the influence of other predicates is too small, the method of removing the general predicate is adopted.

[0038] S2. According to the statistical results, calculate the SAO similarity of the words to obtain the SAO similarity score.

[0039] Specifically, step S2 includes: based on the statistical results, when it is determined that the comparison word and the word to be compared are the subject, calculate the corresponding SAO similarity of the words according to the co-occurring predicates, objects, and their frequencies between the comparison word and the word to be compared.

[0040] When it is determined that the comparison word and the word to be compared are the predicate, calculate the corresponding SAO similarity of the words according to the co-occurring subjects, objects, and their frequencies between the comparison word and the word to be compared.

[0041] When it is determined that the comparison word and the word to be compared are the object, calculate the corresponding SAO similarity of the words according to the co-occurring subjects, predicates, and their frequencies between the comparison word and the word to be compared.

[0042] When it is determined that there is no co-occurring SAO structure between the comparison word and the word to be compared, the corresponding SAO similarity is 0, and the Laplace smoothing algorithm is used to adjust the SAO similarity of the words.

[0043] Preferably, the calculation formula for the SAO similarity of the words is: , , ,in, is the word-subject SAO similarity, is the word-predicate SAO similarity, is the word object SAO similarity, is the probability of co-occurrence of subject of word w1 and word w2, is the probability of word w2 co-occurring with the subject of word w1, is the sum of the probabilities of subject co-occurrence between words, is the probability of word w1 and word w2 co-occurring as predicates, is the probability of word w2 co-occurring with word w1 as a predicate, is the sum of the probabilities of predicate co-occurrence between words, is the probability of co-occurrence of word w1 and word w2 object, is the probability of word w2 co-occurring with the object of word w1, is the sum of the probabilities of object co-occurrence between words.

[0044] It should be noted that due to the different calculation methods, and , and , and , are not always equal.

[0045] In this embodiment, the SAO structure refers to a triple extracted from a text sentence, in which the subject and the object appear as the subject, representing the executor and the executor in the event description; the predicate, as the action between the subject and the object, links the two together. As a structure based on part of speech and semantic acquisition, the SAO structure can explain the logical relationship between words and avoid confusion of word features. Specifically, the advantages of using SAO as a word information feature include: avoiding interference from other irrelevant words in the context, and being able to accurately describe the associated information of the words; most importantly, reducing the workload of statistics and calculations, and improving the accuracy of the semantic similarity algorithm.

[0046] In this embodiment, in the word semantic similarity measurement method integrating SAO and Bayesian model, the word SAO similarity is determined by the ratio of the SAO of words with less information contained in the SAO of words with more information.

[0047] Based on the calculation formula of word SAO similarity calculation, assuming that word 1 contains i groups of SAO, and word 2 contains j groups of SAO, , , , and its calculation formula is: , , , , , , , , , where represents the co-occurring subject of Word 1 and Word 2, similarly; represents the subject of Word 1, similarly.

[0048] S3. Calculate the semantic similarity and empirical similarity between the comparison word and the word to be compared according to a preset knowledge base to generate an empirical similarity, where the knowledge base is a WordNet knowledge base;

[0049] Specifically, step S3 includes: calculating the path distance similarity empirical score wu_simlarity and the semantic similarity empirical score lin_simarity based on the path between the comparison word and the word to be compared according to the preset knowledge base to determine the highest score between the comparison word and the word to be compared;

[0050] Calculate the average values of the path distance similarity empirical score wu_simlarity and the semantic similarity empirical score lin_simarity respectively to obtain the empirical similarity. The calculation formula is: , where is the semantic similarity empirical score calculated based on the path distance similarity empirical score wu_similarity, is the semantic similarity empirical score calculated based on the semantic similarity empirical score lin_similarity;

[0051] When it is determined that the scores of the path distance similarity empirical score wu_simlarity and the semantic similarity empirical score lin_simarity are both 1, the dissimilar empirical similarity is 0, and the Laplace smoothing algorithm is used to adjust the empirical similarity.

[0052] Preferably, the calculation formula of the Laplace smoothing algorithm is: , , where is the smoothing parameter, is the observation result with the value of l under the jth characteristic result of the random variable a, K is the number of characteristic types, is the number of characteristic classifications, N is the number of statistical samples, is an indicator function, indicating that in the ith sample, the value of the jth characteristic is equal to , and when the category Y is equal to , the function value is 1, otherwise it is 0. is the k-th category, is the i-th category, is the j-th feature in the i-th sample.

[0053] In this embodiment, is an indicator function, which means that in the i-th sample, the value of the j-th feature is equal to and the category Y is equal to , the function value is 1, otherwise it is 0. Summing this indicator function over all samples gives the number of samples in which the feature takes the value under the category . is also a summing indicator function, representing the number of samples where the category Y is equal to . Then add , where is the number of possible values of the j-th feature . This is also part of Laplace smoothing, used to adjust the denominator to keep the probability smooth.

[0054] Preferably, the calculation formula for the empirical score of path distance similarity wu_simlarity is: , where is the highest common parent node of the words w1 and w2 in the knowledge base, is the shortest path length from the word w1 to the word w2, is the depth of the word in the knowledge base hierarchy tree.

[0055] Preferably, the calculation formula for the empirical score of semantic similarity lin_simarity is: , where is the probability that the information content of the nearest common parent node of the words w1 and w2 accounts for the total description information content, is the probability that the information content of the word w1 accounts for the total information content, is the probability that the information content of the word w2 accounts for the total information content, is the information content of the highest common parent node of the words w1 and w2 in the knowledge base, is the information content of the word w1, is the information content of the word w2.

[0056] In this embodiment, WordNet, which has more comprehensive information and more accurate calculation results, is used as the knowledge base. By calculating and comparing the path-based wu_similarity (hereinafter referred to as wu) and the concept-content-based lin_similarity (hereinafter referred to as lin) between the comparison word and the word to be compared, the highest score of the two is obtained. Further, the average value of wu and lin is calculated to obtain the empirical similarity. It should be noted that since the empirical similarity needs to be put into the Bayesian model, when the scores of wu and lin are both 1, the empirical similarity of dissimilarity is judged to be 0, which will cause the final empirical score to lose its effect. To solve this problem, this method selects Laplace smoothing to adjust the empirical similarity.

[0057] Specifically, please refer to Figure 3 , WordNet is a widely used English semantic knowledge base in the field of NLP. It forms a semantic network with the semantic relationships between words. The semantic relationships between words are expressed through links, making WordNet have a good concept hierarchy and tree structure. Taking the set of hyponyms of the noun "channel" as an example, the semantic relationship examples between words in WordNet are shown as Figure 3 shown. Among them, Figure 3 the upper-level node in is the parent node of the lower-level node, usually representing the hyponymy relationship, etc.; the nodes with the same parent node are sibling nodes to each other, usually representing the near-synonym or synonym relationship; for each definition, there is its word set definition, expressing the specific meaning of the word set.

[0058] Specifically, in this embodiment, the semantic similarity calculation model derived from WordNet mainly considers three aspects: based on the path distance between word pairs, based on the information content (IC), and based on attribute features. Due to the good network performance of WordNet, this method mainly focuses on the number of connecting edges between two concepts. The number of connecting edges is inversely proportional to the semantic similarity. The more connecting edges, the lower the semantic similarity; the number of connecting edges is directly proportional to the semantic distance. The more connecting edges, the farther the semantic distance. The word semantic similarity measurement method integrating SAO and Bayesian model selects Wu_similarity as the empirical score of path distance similarity. The calculation formula of Wu_similarity is: . Taking the schematic diagram in Figure 3 as an example, here is "channel", is the length reaching via (line, channel), is the depth of the channel at this level. Among them, the path distance similarity empirical score Wu_simlarity believes that the lower the level where the concept is located, the smaller the similarity; when calculating, not only the concept path is considered, but also the depth of the highest common parent node in the classification tree is considered. Therefore, the accuracy is relatively better than other algorithms, and the computational complexity is low.

[0059] In addition, Wordnet not only has good network performance, but also contains rich concept information; the semantic similarity calculation method based on IC mainly focuses on the information content of concepts. The method for measuring the semantic similarity of words that integrates SAO and the Bayesian model selects the Lin_simlarity method as the empirical score of semantic similarity based on IC. The calculation formula of Lin_simlarity is: . Take Figure 3 the schematic diagram as an example, the nearest common parent node of and contact is "line", then the Lin_simlarity of the two is the ratio of twice the information amount of "line" to the sum of the information amounts of the two.

[0060] Finally, based on the results of the above two empirical score algorithms, taking word1 and word2 as examples, the formula for calculating the empirical similarity of words based on the knowledge base is obtained .

[0061] In this embodiment, since the semantic similarity calculated in this article is based on the Bayesian model, zero-probability events will lead to extreme situations (0 or 1) in the final calculation result. To avoid such extreme situations, the Laplace smoothing technique is selected to adjust zero-probability events. Among them, the smoothing parameter , generally set to 1, which ensures that the numerator will not be 0.

[0062] Specifically, in the method for measuring the semantic similarity of words that integrates SAO and the Bayesian model, there are two situations where the Laplace smoothing algorithm needs to be introduced, that is, the situation where the score is 1 when calculating the empirical similarity, and the situation where there is no co-occurring SAO when calculating the SAO similarity.

[0063] When calculating the empirical score, if the words are under the same parent node, the empirical score of word similarity is 1, and the non-similar empirical score is 0. When introducing the empirical score into the Bayesian model, the semantic similarity calculation completely depends on the SAO similarity, which obviously violates the original intention of this method. Therefore, Laplace smoothing is used for the situation where the semantic similarity empirical score is 1; the usage method is: divide the empirical score into two categories, that is =2, let the number of statistical samples N = 10, =1, at this time, the number of similar samples counted is adjusted to: similar is equal to 10 + 1, and non-similar is equal to 1. , , after Laplace adjustment, the situation where the empirical score becomes 0 and loses its effect is avoided, which is beneficial to ensuring the accuracy of the algorithm.

[0064] When counting co-occurring SAOs, especially when the words are in an antonymous relationship or an ontological relationship in different domains, the situation of no co-occurring SAO often occurs. At this time , and Once the value is 0, the result of the entire Bayesian model is also 0, resulting in the failure of the algorithm. To avoid this situation, Laplace smoothing can be used to adjust the SAO co-occurrence statistical results. Based on the original statistical results, the statistical results are divided into two categories, namely co-occurrence and non-co-occurrence, that is = 2, let = 1, at this time, , , where T represents the total sum of statistical samples, and B represents the total sum of statistical samples that do not contribute. After the above adjustment, the situation where the probability of co-occurring SAO in statistics is 0 is avoided, removing obstacles for the Bayesian model to calculate similarity, which is beneficial to improving the accuracy of the algorithm.

[0065] S4. Substitute the SAO similarity score and the empirical similarity into the Bayesian formula to obtain the final calculation result of the semantic similarity of words.

[0066] Specifically, step S1 includes: using the SAO similarity score as the probability of judging similarity between word pairs, and the empirical similarity as a regulatory parameter;

[0067] Use the Bayesian formula to calculate the SAO similarity score and the empirical similarity to obtain the final calculation result of the semantic similarity of words. The calculation formula is: , .

[0068] In this embodiment, according to the Bayesian formula, the SAO similarity score of word pairs is used as the probability of judging similarity between word pairs, then the probability of judging dissimilarity is "1 - SAO similarity score". At the same time, considering that there will inevitably be deviations in the similarity method based on statistics and the similarity method based on the corpus, the empirical similarity obtained by the similarity method of words based on the corpus is used as a regulatory parameter. Finally, substitute the above similarity score and regulatory parameter into the Bayesian formula to obtain the calculation result of the semantic similarity of words calculated by the Bayesian model.

[0069] In this embodiment, the SAO statistical model is used as the SAO similarity probability of the Bayesian model, and the similarity empirical score based on the knowledge base is used as the semantic consistency parameter, and finally the similarity calculation result of words is obtained. Among them, the Bayesian model is as follows Figure 4As shown in the figure. Assume that the marginal distribution is P(Y). Then, P(Ssim|Y), P(Asim|Y), and P(Osim|Y) represent the probability distribution of determining semantic similarity between word pairs. Set the prior probabilities P(Y=T) and P(Y=F) to 0.5. Then, the posterior probabilities of determining word similarity based on SAO similarity are P(Y =T| Ssim), P(Y =T| Asim), and P(Y =T| Osim), and the posterior probabilities of determining word dissimilarity are P(Y =F| Ssim), P(Y =F| Asim), and P(Y =F| Osim). Assume that word 1 and word 2 are the subject. Then, the formula for calculating the semantic similarity sim(w1, w2) of word 1 and word 2 jointly constrained by the attribute Asim( ) and the attribute Osim( ) is as shown in Equation .

[0070] According to the conditional probability formula, we can obtain: , . According to the joint probability formula, we can obtain: . For , since and do not affect each other, so .

[0071] Similarly, we can obtain , . Combining P(Y=T)=P(Y=F)=0.5 with the above formulas, we can obtain: , . In addition, if word 1 and word 2 are the predicate, the semantic similarity is calculated by jointly constraining the attributes Ssim( ) and Osim( ); if word 1 and word 2 are the object, the semantic similarity is calculated by jointly constraining the attributes Ssim( ) and Asim( ).

[0072] Specifically, in this embodiment, bathroom texts are selected as the corpus for experimental comparison. Among them, the text structure comparison is shown in Table 1.

[0073] Table 1 Text Structure Comparison

[0074]

[0075] Based on the existing bathroom overview text, extract the SAO structure of the sentences to form a database. Since the SAO database of bathroom overview text contains more than 1.7 million SAO data, the probability of accidental events is reduced, and the credibility of the database is improved. Part of the SAO structure is shown in Table 2.

[0076] Table 2 SAO Structure of Sanitary Ware Text Part

[0077]

[0078] Select a pair of words to be compared, extract the SAO structure of the words according to whether the words are the subject, predicate or object. Since some words have multiple word forms, such as dispenser, dispensers, dispensering, and dispensation, etc., it is necessary to convert multiple word forms into the general form. In this embodiment, the WordNetLemmatizer method in the python NLTK library is used for word form merging. After word form merging, extract the SAOs of similar words respectively, and label the total number of SAOs and the co-occurring SAOs; finally, count and calculate the co-occurring SAOs to obtain the SAO similarity. At the same time, calculate the empirical score of the word pair based on the knowledge base. In this embodiment, the wu_simlarity and lin_simlarity between word pairs are calculated using the wordnet module of python, and the average value of the two algorithms is calculated. The calculation results of the SAO similarity of word pairs and the empirical similarity calculation results based on the knowledge base are shown in Table 3.

[0079] Table 3 Calculation Results of SAO Similarity of Word Pairs and Empirical Similarity Calculation Results Based on Knowledge Base

[0080]

[0081] Finally, fuse the SAO similarity and the empirical score into the Bayesian model to obtain the word pair similarity value. In this embodiment, forty pairs of words are selected for verification, and the final empirical score and the calculation results of the SAO similarity are shown in the "Similarity" column of Table 4.

[0082] Among them, the correlation coefficient is a measure of the strength of the relationship between two variables. In statistics, the main correlation coefficients used are the pearson correlation coefficient and the spearman correlation coefficient. The Pearson correlation coefficient mainly measures the degree of linear correlation between two variables, and its value ranges from [-1, 1]. The closer the correlation coefficient is to 1, the more linearly correlated the variables are. The closer the correlation coefficient is to -1, the more linearly negatively correlated they are. 0 indicates non-linear correlation. The main function of the Spearman correlation coefficient is to describe the degree of the relationship between two variables using a monotonic function, and its calculation method is mainly based on the ranking of variables rather than the original data. The calculation formulas of the Pearson and spearman correlation coefficients are as follows: , , where, represents the pearson correlation coefficient between variable x and variable y, and n represents the number of observations; represents the Spearman correlation coefficient, represents the rank difference of the corresponding variables.

[0083] In this embodiment, the SimLax-999 dataset is selected as the manual judgment standard and compared with the calculation results of the method for measuring the semantic similarity of words integrating the SAO and Bayesian models. The SimLax-999 dataset scoring mainly measures the similarity between word pairs and excludes word pairs with mixed parts of speech that can form phrases (such as "yellow-hair", "play golf") to avoid interference from such noises, which has high reference value and is suitable for the evaluation of semantic similarity calculation. In addition, in addition to the word pairs in simlax-999, the method for measuring the semantic similarity of words integrating the SAO and Bayesian models also selects some data in wordsim-353 and word pairs with high semantic similarity and adds them to the validation set. The calculation results of the method for measuring the semantic similarity of words integrating the SAO and Bayesian models are shown in Table 4 below.

[0084] Table 4 Calculation Results of the Method for Measuring the Semantic Similarity of Words Integrating the SAO and Bayesian Models and Calculation Results of Other Methods

[0085]

[0086] Taking the method for measuring the semantic similarity of words integrating the SAO and Bayesian models and comparing it with the method for calculating semantic similarity based on the knowledge base, the method for calculating semantic similarity based on word vectors, and the manual validation set, the final Pearson correlation coefficient is shown in Table 5. Among them, Wu&lin is the average value of the Wu_simlarity method and the Lin_simlarity method based on the WordNet knowledge base; CBOW is the calculation result of the cosine similarity of word vectors based on the bag-of-words model; Skip Gram is the calculation result of the cosine similarity of word vectors based on the neural network; Similarity is the calculation result of the method in this paper; Manual rating is the manual score. The Spearman correlation coefficient is shown in Table 6.

[0087] Table 5 Pearson Coefficients of Calculation Results of Each Method

[0088]

[0089] Table 6 Spearman Coefficients of Calculation Results of Each Method

[0090]

[0091] In this embodiment, the distribution status and curve fitting status of the nodes and the manual validation set scores are as follows Figure 5As shown, the curve fitting situation is as Figure 6 shown. From the Pearson correlation coefficient reaching 0.823466 and the p-value being 6.84e-11, it can be seen that the method for measuring the semantic similarity of words by integrating the SAO and Bayesian models has a strong correlation and significant effect with the artificial scoring results. From the spearman correlation coefficient reaching 0.819447, it can be seen that the method for measuring the semantic similarity of words by integrating the SAO and Bayesian models has a strong dependence and consistency with the artificial judgment results. From the curve fitting results and the node distribution situation, it can be obtained that the method for measuring the semantic similarity of words by integrating the SAO and Bayesian models has a better semantic similarity calculation effect compared with the semantic similarity calculation method based on the knowledge base and the calculation method based on word vectors. Based on the data in Table 2 and with the artificial annotation scoring results as the benchmark, the comparison results of various algorithms are as Figure 7 shown. From Figure 7 it can be seen that the method for measuring the semantic similarity of words by integrating the SAO and Bayesian models is closer to the artificial judgment results.

[0092] According to the calculation results in Table 4, the semantic similarity calculation model has good results when the semantic similarity of word pairs is high or low, and has a poor calculation effect when the semantic similarity score is in the interval [0.4, 0.5], and there are obvious outliers; such as "achieve-try", "child-adult", "effort-difficulty", "second-minute", "capability-strength", etc. The main reason for such problems is that simlax-999 mainly scores the semantic similarity of word pairs. However, the semantic similarity of the above word pairs is low, but the semantic correlation is high, and their SAO co-occurrence opportunities are more, which can be verified from the scoring results of the semantic similarity algorithm based on the knowledge base. This outlier phenomenon also leads to a slightly lower correlation coefficient score in this paper.

[0093] In summary, the method for measuring the semantic similarity of words by integrating the SAO and Bayesian models aims to solve the challenges of calculating the semantic similarity of words in the field of natural language processing by integrating the SAO structure and the Bayesian model. This method first extracts the SAO structure of the text from the sample set, and statistically calculates the SAO of similar words to obtain the SAO similarity statistical probability; then, uses semantic knowledge bases such as WordNet to obtain the semantic similarity parameters of words, and adjusts the semantic consistency parameters by using the Laplace smoothing technique. Finally, combines the similarity parameters and the SAO similarity statistical posterior probability, and calculates the semantic similarity probability of words through the Bayesian model.

[0094] The method for measuring the semantic similarity of words integrating SAO and Bayesian model can combine subjective judgment based on experience with objective analysis based on statistics, improving the accuracy and reliability of measuring the semantic similarity of words. And in the case study of the above-mentioned bathroom supplies text dataset, this method shows good feasibility and effectiveness. Through this method, the similarity between words can be understood and quantified more accurately, providing a new perspective and tool for the development of natural language processing technology. The method for measuring the semantic similarity of words integrating SAO and Bayesian model has brought new breakthroughs to the field of semantic similarity measurement, and is expected to promote the progress and application of related technologies.

[0095] Please refer to Figure 8 , the second embodiment of the present invention provides a device for measuring the semantic similarity of words integrating SAO and Bayesian model, which includes:

[0096] An extraction and statistics unit 201, configured to extract the SAO structure of the words to be compared from a preset corpus based on semantic part-of-speech annotation, and perform statistical preprocessing on the SAO structure to obtain a statistical result;

[0097] An SAO similarity unit 202, configured to calculate the SAO similarity of words according to the statistical result to obtain an SAO similarity score;

[0098] An empirical similarity unit 203, configured to calculate the empirical similarity of the semantic similarity between the comparison word and the word to be compared according to a preset knowledge base to generate an empirical similarity, where the knowledge base is a WordNet knowledge base;

[0099] A Bayesian calculation unit 204, configured to substitute the SAO similarity score and the empirical similarity into the Bayesian formula to obtain the final calculation result of the semantic similarity of words.

[0100] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for measuring word semantic similarity that integrates the SAO and Bayesian models, characterized in that Including: Extracting the SAO structure of the to-be-compared words from a preset corpus based on semantic part-of-speech annotation, and performing statistical preprocessing on the SAO structure to obtain a statistical result; Calculating the SAO similarity of words according to the statistical result to obtain an SAO similarity score; Calculating the semantic similarity empirical similarity between the comparison word and the to-be-compared word according to a preset knowledge base to generate an empirical similarity, where the knowledge base is the WordNet knowledge base; Substituting the SAO similarity score and the empirical similarity into the Bayesian formula to obtain the final calculation result of the semantic similarity of words; Calculating the SAO similarity according to the statistical result to obtain an SAO similarity score, specifically: Based on the statistical result, when it is determined that the comparison word and the to-be-compared word are subjects, calculating the corresponding SAO similarity of words according to the predicates, objects, and their frequencies that co-occur between the comparison word and the to-be-compared word; When it is determined that the comparison word and the to-be-compared word are predicates, calculating the corresponding SAO similarity of words according to the subjects, objects, and their frequencies that co-occur between the comparison word and the to-be-compared word; When it is determined that the comparison word and the to-be-compared word are objects, calculating the corresponding SAO similarity of words according to the subjects, predicates, and their frequencies that co-occur between the comparison word and the to-be-compared word; When it is determined that there is no co-occurring SAO structure between the comparison word and the to-be-compared word, the corresponding SAO similarity is 0, and the Laplace smoothing algorithm is used to adjust the SAO similarity of words; The calculation formula for the SAO similarity of words is as follows: , , , where is the SAO similarity of the subject of the word, is the SAO similarity of the predicate of the word, is the SAO similarity of the object of the word, is the probability of the co-occurrence of the subject of word w1 and word w2, is the probability of the co-occurrence of the subject of word w2 and word w1, is the total probability of the co-occurrence of the subjects between words, is the probability of the co-occurrence of the predicate of word w1 and word w2, is the probability of the co-occurrence of the predicate of word w2 and word w1, is the total probability of the co-occurrence of the predicates between words, is the probability of the co-occurrence of the object of word w1 and word w2, is the probability of the co-occurrence of the object of word w2 and word w1, is the total probability of the co-occurrence of the objects between words.

2. The method for measuring word semantic similarity by integrating the SAO and Bayesian models according to claim 1, wherein Extracting the SAO structure of the to-be-compared words from a preset corpus based on semantic part-of-speech annotation, and performing statistical preprocessing on the SAO structure to obtain a statistical result, specifically: Performing sentence splitting on the text in the preset corpus to obtain multiple sentences, and performing semantic part-of-speech annotation on the words in each sentence; Performing SAO extraction processing on the sentences according to the annotated word parts of speech, extracting the SAO structures of the two to-be-compared words, performing word form merging processing on the extracted SAO structures, and converting the other word forms of the to-be-compared words into prototypes; Respectively counting the occurrence frequencies of the SAO structures corresponding to the two to-be-compared words in the text and comparing them to obtain the co-occurring SAO structures and their occurrence frequencies among the two to-be-compared words, and generating a statistical result.

3. The method for measuring the semantic similarity of words integrating the SAO and Bayesian models according to claim 1, wherein The calculation formula of the Laplace smoothing algorithm is as follows: , , where is the smoothing parameter, is the observation result with value l under the j-th characteristic result of the random variable a, K is the number of characteristic types, is the number of characteristic classifications, N is the number of statistical samples, is an indicator function, indicating that in the i-th sample, the value of the j-th characteristic is equal to , and when the category Y is equal to , the function value is 1, otherwise it is 0, is the k-th category, is the i-th category, is the j-th characteristic in the i-th sample.

4. The method for measuring word semantic similarity by integrating the SAO and Bayesian models according to claim 3, wherein Calculating the semantic similarity empirical similarity between the comparison word and the to-be-compared word according to a preset knowledge base to generate an empirical similarity, specifically: According to the preset knowledge base, calculating the path distance similarity empirical score wu_simlarity based on the path and the semantic similarity empirical score lin_simarity based on the concept content between the comparison word and the to-be-compared word to determine the highest score between the comparison word and the to-be-compared word; Calculate the average of the empirical scores of path distance similarity wu_simlarity and semantic similarity lin_simarity respectively to obtain the empirical similarity. The calculation formula is as follows: , where is the empirical score of semantic similarity calculated based on the empirical score of path distance similarity wu_similarity, is the empirical score of semantic similarity calculated based on the empirical score of semantic similarity lin_similarity; When it is determined that the scores of both the path distance similarity empirical score wu_simlarity and the semantic similarity empirical score lin_simarity are 1, the dissimilar empirical similarity is 0, and the Laplace smoothing algorithm is used to adjust the empirical similarity.

5. The method for measuring word semantic similarity by integrating the SAO and Bayesian models according to claim 4, wherein The calculation formula for the empirical score wu_simlarity of path distance similarity is as follows: , where is the highest common parent node of words w1 and w2 in the knowledge base, is the shortest path length from word w1 to word w2, is the depth of the word in the hierarchical tree of the knowledge base.

6. The method for measuring word semantic similarity by integrating the SAO and Bayesian models according to claim 4, wherein The calculation formula for the empirical score of semantic similarity lin_simarity is as follows: , where is the probability that the information content of the nearest common parent node of words w1 and w2 accounts for the total description information content, is the probability that the information content of word w1 accounts for the total information content, is the probability that the information content of word w2 accounts for the total information content, is the information content of the highest common parent node of words w1 and w2 in the knowledge base, is the information content of word w1, is the information content of word w2, is the highest common parent node of words w1 and w2 in the knowledge base.

7. The method for measuring word semantic similarity by integrating the SAO and Bayesian models according to claim 4, wherein Substituting the SAO similarity score and the empirical similarity into the Bayesian formula to obtain the final calculation result of the semantic similarity of words, specifically: Take the SAO similarity score as the probability of judging similarity between word pairs, and take the empirical similarity as a tuning parameter; The Bayesian formula is used to calculate the SAO similarity score and the empirical similarity to obtain the final calculation result of word semantic similarity. The calculation formula is as follows: , .

8. A device for measuring the semantic similarity of words integrating SAO and Bayesian model, characterized in that, It includes: An extraction and statistics unit, configured to extract the SAO structures of the words to be compared from a preset corpus based on semantic part-of-speech tagging, and perform statistical preprocessing on the SAO structures to obtain a statistical result; An SAO similarity unit, configured to calculate the SAO similarity of words according to the statistical result to obtain an SAO similarity score; An empirical similarity unit, configured to calculate the semantic similarity empirical similarity between a comparison word and a word to be compared according to a preset knowledge base to generate an empirical similarity, where the knowledge base is a WordNet knowledge base; A Bayesian calculation unit, configured to substitute the SAO similarity score and the empirical similarity into a Bayesian formula to obtain a final calculation result of word semantic similarity; According to the statistical result, perform SAO similarity calculation to obtain an SAO similarity score, specifically: Based on the statistical result, when it is determined that the comparison word and the word to be compared are subjects, calculate the corresponding SAO similarity of the words according to the predicates, objects, and their frequencies that co-occur between the comparison word and the word to be compared; When it is determined that the comparison word and the word to be compared are predicates, calculate the corresponding SAO similarity of the words according to the subjects, objects, and their frequencies that co-occur between the comparison word and the word to be compared; When it is determined that the comparison word and the word to be compared are objects, calculate the corresponding SAO similarity of the words according to the subjects, predicates, and their frequencies that co-occur between the comparison word and the word to be compared; When it is determined that there is no co-occurring SAO structure between the comparison word and the word to be compared, the corresponding SAO similarity is 0, and the Laplace smoothing algorithm is used to adjust the SAO similarity of the words; The calculation formula for the SAO similarity of words is as follows: , , , where is the SAO similarity of the subject of the word, is the SAO similarity of the predicate of the word, is the SAO similarity of the object of the word, is the probability of the co-occurrence of the subjects of word w1 and word w2, is the probability of the co-occurrence of the subjects of word w2 and word w1, is the total probability of the co-occurrence of subjects between words, is the probability of the co-occurrence of the predicates of word w1 and word w2, is the probability of the co-occurrence of the predicates of word w2 and word w1, is the total probability of the co-occurrence of predicates between words, is the probability of the co-occurrence of the objects of word w1 and word w2, is the probability of the co-occurrence of the objects of word w2 and word w1, is the total probability of the co-occurrence of objects between words.