Text semantic similarity calculation method based on multilayer semantic information fusion
By constructing E2vec and C2vec models combined with Word2vec and Siamese BERT models, the multi-layer semantic information is fused, which solves the problem of insufficient semantic information in text similarity calculation and improves the calculation accuracy.
Patent Information
- Application Number
- CN202311440941.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art lacks deep semantic information when calculating text similarity, resulting in low calculation accuracy, especially when facing the problem of polysense of one word.
The entity vector model E2vec and the category vector model C2vec are constructed using the Probase knowledge graph, combined with Word2vec, the word vector is obtained, the text context semantic features are extracted through the Siamese BERT model, and multi-layer semantic information is fused for calculation.
It improves the accuracy of text similarity calculation, solves the problem of multiple meanings of one word, and enhances the vector representation ability of text semantics.
Smart Images

Figure CN120387434A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for calculating text semantic similarity based on multi-layer semantic information fusion. Background Art
[0002] Short text similarity calculation plays an important role in natural language processing and has been widely applied to many language processing fields such as automatic question answering, text classification, machine translation, and information retrieval. For example, in the field of automatic question answering, for a given input short question-answer text, this technology can sort the short texts with a relatively high match with the query short text, score them considering other fuzzy decision factors, and feedback several answers with a relatively high match degree to the user. This method has been practically applied in scenarios such as Baidu Search and Xiaodu Intelligent Question Answering. With the increasing demand for these applications, text similarity calculation has become a hot direction in natural language processing. Although traditional string-based and corpus-based similarity metrics are intuitive, simple, and easy to implement, they only stay on the surface of the text and do not consider the deep semantic information of the text, resulting in poor accuracy in scenarios with complex semantic structures. Therefore, deep learning-based methods can make up for the deficiencies of traditional methods, that is, by generating embedding vectors of the text through neural networks, so as to learn the deep semantics of the text and be able to calculate text similarity more accurately.
[0003] Currently popular text similarity calculation methods based on deep neural networks mostly make different improvements to the neural network model. For example, the Attention-BiLSTM model and the Attention-Transformer model with multi-head attention mechanisms are added, or the Word2vec model is combined with an improved deep learning model. Although these calculation methods can help mine the semantic information of the text, their effects are not very satisfactory when encountering the problem of polysemy. Summary of the Invention
[0004] The present invention provides a method for calculating text semantic similarity based on multi-layer semantic information fusion, aiming to solve the problem of low calculation accuracy caused by the lack of semantic information in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solution: A method for calculating text semantic similarity based on multi-layer semantic information fusion, including the steps of: Using entity relationship triples and entity category triples in the Probase knowledge graph to construct an entity vector model E2vec and a category vector model C2vec respectively; Using the word vector model Word2vec to obtain the word vectors of each word in the text pair to be calculated for similarity; Use the entity library in the Probase knowledge graph as a dictionary to extract the entity set from all the words in the text, and send the entities in the entity set into the entity vector model E2vec and the category vector model C2vec to obtain the entity vectors of each entity and all the category vectors corresponding to each entity; Calculate the cosine similarity between the top 20 category vectors of the entity to be recognized and the context entity vector of the entity to be recognized, and use the category vector with the highest score as the category vector of the target entity in the current text pair; Perform a concatenation operation on the word vector, entity vector, and category vector of the target entity to obtain the final semantic information enhanced word vector for each word; Input the semantic information enhanced word vector into the Siamese BERT model to extract the context semantic features of the text pair and obtain the text pair vector after fusing the full-text semantic information; Calculate the similarity score of the text pair through the text pair vector Preferably, the entity vector model E2vec and the category vector model C2vec include: The training corpus of the E2vec model is the entity relationship triples in Microsoft's Probase knowledge graph, which contains tens of millions of entities and the number of times they appear together; The training corpus of the C2vec model is the entity category triples in the Probase knowledge graph, which contains millions of categories, the entity sets included in each category, and the probability that each entity belongs to each category.
[0006] Preferably, after separately constructing the entity vector model E2vec and the category vector model C2vec, the following steps are further included: Use the Structured Deep Network Embedding (SDNE) algorithm to pre-train the entity relationship triples and entity category triples, learn the semantic relationships in the knowledge graph, and obtain the pre-trained model and pre-trained model parameters; Use the pre-trained model to obtain entity vectors and category vectors.
[0007] Preferably, using the Structured Deep Network Embedding (SDNE) algorithm to pre-train the entity relationship triples and entity category triples includes the following steps: Collect all the entity relationship triples and entity category triples in the Probase knowledge graph, and perform data preprocessing on the entity relationship triples and entity category triples; Construct a network graph structure according to the preprocessed data. In the entity vector model E2vec, use the entities as the nodes of the graph, the relationships between entities as the edges of the graph, and the frequency of co-occurrence of two entities as the edge weights of the adjacency matrix of the graph; Take the entities and categories in the category vector model C2vec as the nodes of the graph, take the relationships between entities and categories as the edges of the graph, and take the probability that an entity belongs to a certain category as the edge weights of the adjacency matrix of the graph; Based on the SDNE algorithm, use the training set to train the pre-trained embedding model, and use the test set to test the pre-trained embedding model; Save the parameters and vectors of the two pre-trained embedding models that have been trained and tested to obtain entity vectors and category vectors.
[0008] Preferably, the training of the pre-trained embedding model using the training set based on the SDNE algorithm includes the following steps: Adjust the parameters of the model to optimize the model performance; Use the multi-classification task evaluation model to further adjust the model parameters and map entities of the same category to a similar vector space.
[0009] Preferably, calculate the cosine similarity between the top 20 category vectors of the entity to be recognized and the context entity vectors of the entity to be recognized, and take the category vector with the highest score as the category vector of the target entity in the current text pair. The specific calculation expression is: Among them, sim ( v 1, v 2) is the cosine similarity; and respectively represent the two vectors for which the similarity needs to be calculated; Set the moving window size for selecting context-adjacent entities to 2, and obtain the category vector information of each entity by combining context semantic information.
[0010] Preferably, when performing a concatenation operation on the word vectors, entity vectors, and category vectors of the target entity, set the entity vectors and category vectors that are not entities in the word to 0; If an entity consists of multiple words, perform an addition operation on the word vectors of each word in the entity, and take the result as the word vector of the entity.
[0011] Preferably, use the MRPC dataset to fine-tune the Siamese BERT model, including the steps: Divide the MRPC dataset into training data and test data. Each group of data is a pair of text samples, and each sample contains two text strings and a label; Construct a Siamese BERT model that includes two BERT encoders and a similarity calculation layer; The Siamese BERT model is trained using training data, and the model parameters are updated through the backpropagation algorithm and the optimization algorithm to optimize the accuracy of the model in calculating similarity.
[0012] Preferably, inputting the semantic information enhanced word vectors into the Siamese BERT model includes the steps of: Taking the semantic information enhanced word vectors of the short text pair as the initial word vectors of the model, and respectively inputting them into the left and right sides of the Siamese BERT model; Learning the context semantic features of the two texts through the shared encoding layer, and outputting the vectors of the two texts after fusing the full text semantic information.
[0013] Preferably, the similarity score of the text pair is calculated through the text pair vectors, and the specific expression is: where x and y respectively represent the feature vectors of the two texts, and respectively represent the means of the two feature vectors, P represents the Pearson correlation coefficient. Through the above formula, the P value is restricted between 0 and 1, representing the similarity score of the model.
[0014] The present invention has the following beneficial effects compared with the prior art: Based on the single word vector, the present invention uses the Word2vec model to obtain word vectors, considers semantic information in multiple aspects, uses the entity library in the Probase knowledge graph as a dictionary to extract the entity set in the word set, and incorporates some semantic knowledge that cannot be included in the Word2vec model, such as the relationship between entities appearing in the text and which category a specific entity belongs to in the text. To a certain extent, it solves the problem of polysemy that may occur during calculation, helps to identify the specific category to which the entity belongs in the text, and combines the Siamese BERT model fine-tuned for the text similarity task to further extract context semantic features, so that the implicit semantic features of the text are fully mined, thereby enhancing the vector representation ability of the text semantics and solving the problem of low calculation accuracy caused by the lack of semantic information. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic flow chart of the steps of the method of the present invention; Figure 2 is a partial structural schematic diagram of the Probase knowledge graph; Figure 3 is a schematic diagram of the Siamese BERT model of the present method. DETAILED DESCRIPTION OF THE INVENTION
[0016] The following further describes the specific embodiments of the present invention in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0017] For the sake of understanding and illustration, a method for calculating text semantic similarity based on multi-layer semantic information fusion in an embodiment of the present invention will be described in detail below.
[0018] As Figure 1 shown, a method for calculating text semantic similarity based on multi-layer semantic information fusion provided by the present invention is described below, including the steps: Step S1: Use the entity relationship triples and entity category triples in the Probase knowledge graph to construct a pre-trained embedding model (entity vector model) E2vec and (category vector model) C2vec respectively.
[0019] This step specifically includes: First, use the Structured Deep Network Embedding (SDNE) algorithm to pre-train the entity relationship triples and entity category triples in the Probase knowledge graph respectively, learn the semantic relationships in the knowledge graph and save the model parameters, obtaining two pre-trained embedding models E2vec and C2vec and pre-trained model parameters. These two models are respectively used to obtain entity vectors and category vectors.
[0020] Among them, the Probase knowledge graph is a large-scale and complex network structure graph proposed by Microsoft, composed of knowledge triples, arranged in the order of (node, edge, node). It includes entity category triples, entity relationship triples, entity attribute triples, etc., as Figure 2 shown, such as (Bob, type, person), (Apple, produces, iPhone), (iPhone, manufacturer, Apple), etc. Probase contains a large number of specific categories such as "popular alternative banking option" and "exciting outdoor activity", which helps to enhance the discrimination between similar categories.
[0021] Among them, the construction process of the pre-trained embedding models E2vec and C2vec is as follows: (1) First, collect all entity-relationship triples and entity-category triples in the Probase knowledge graph, remove invalid and incorrect data, and normalize the entity-relationship weights. This can not only preserve the relative magnitude relationship between the weights but also prevent the problem of gradient explosion during calculation. Then, construct a network graph structure based on the collected data. In E2vec, entities are used as nodes of the graph, relationships between entities are used as edges of the graph, and the frequency of co-occurrence of two entities is used as the edge weight of the adjacency matrix of the graph. In C2vec, entities and categories are used as nodes of the graph, relationships between entities and categories are used as edges of the graph, and the probability that an entity belongs to a certain category is used as the edge weight of the adjacency matrix of the graph. Finally, divide the data into a training set and a test set.
[0022] (2) Select the SDNE algorithm and use the training set to train the model. Optimize the performance of the model by adjusting the model parameters, and evaluate the model with the multi-classification task. Further adjust the model parameters to map entities of the same category to a similar vector space.
[0023] (3) Save the parameters and vectors of the two pre-trained embedding models that have been trained and tested to obtain entity and category vectors.
[0024] The training corpus of the E2vec model is the entity-relationship triples in the Probase knowledge graph from Microsoft, which contains tens of millions of entities and the number of times they co-occur. These entities interact, connect, and depend on each other to form complex semantic relationships, such as (Bob, was born in, Honolulu), etc. The training corpus of the C2vec model is the entity-category triples in the Probase knowledge graph, which contains millions of categories, the set of entities included in each category, and the probability that each entity belongs to each category, such as (Bob, type, person), (visa, type, popular alternative banking option), etc. The Probase knowledge graph contains a large number of specific categories like "popular alternative banking option" and "exciting outdoor activity", which helps to enhance the discrimination between similar categories. Combining the corpus-based method and the deep neural network-based method gives play to their respective advantages and enhances the vector representation ability of text semantics.
[0025] Among them, use the Structured Deep Network Embedding (SDNE) algorithm to pre-train the entity-relationship triples and entity-category triples. The SDNE algorithm includes the steps: For a given entity-relation triple and entity-category triple, the goal of the SDNE algorithm is to represent each node in the node set V using vectors. Among them, the vectors of nodes connected by edges are similar (referred to as first-order similarity), and there is similarity between two nodes that have common neighbor nodes but are not directly connected (referred to as second-order similarity). SDNE uses an autoencoder structure to optimize both first-order and second-order similarity simultaneously, and the learned vector representation can preserve local and global structures.
[0026] For first-order similarity, the loss function is defined as follows: (1) where and are the vector representations of nodes and respectively, K represents the Kth layer, represents the element in the i-th row and j-th column of the adjacency matrix s of the graph, represents calculating the square of the matrix 2-norm. This loss function can make the representations of two adjacent nodes in the embedding space relatively close.
[0027] For second-order similarity, the loss function is defined as follows: (2) where contains the neighbor structure information of node , is the neighbor structure information of node after reconstruction by the autoencoder, is the element-wise product, represents calculating the square of the matrix F-norm, , if , then , otherwise . By assigning a value greater than 1 to , the influence of non-zero elements is amplified. Through such a reconstruction process, vertices with similar structures have similar vector representations.
[0028] To prevent overfitting, regularization is added, and the jointly optimized loss function is: (3) where is the parameter controlling the first-order loss, is the parameter controlling the regularization term, and the regularization term is calculated as follows: (4) where is the parameter matrix of the K-th layer.
[0029] Step S2: Perform data preprocessing such as sentence segmentation, word tokenization, and stop word removal on the short text pairs for which similarity calculation is required, obtain the word sets of the short text pairs, and use the publicly available pre-trained word vector model Word2vec to obtain the word vectors of each word in the text pairs to be calculated for similarity.
[0030] Step S3: Use the entity library in the Probase knowledge graph as a dictionary to extract the entity set in the word set, and send the entities into the entity vector model E2vec and the category vector model C2vec respectively to obtain the entity vectors of each entity and all its corresponding category vectors.
[0031] Among them, the entity library comes from the Probase knowledge graph and contains 4.28M (million) entities, such as some specific things like "cellphone", "milk", "bottled water", etc.
[0032] This step specifically includes: Use the regular expression tokenizer to extract the entity set of the word set, where the dictionary is set to the entity library in Probase and the maximum character length is set to 50. For example, the following two sentences:
[0033] 1. He said the foodservice pie business doesn't fit the company's long-term growth strategy planning. 2. The foodservice pie business does not fit our long-term growth strategy planning. After data preprocessing, the entity set of 1 is ['foodservice', 'pie', 'business', 'fit', 'company', 'growth','strategy'], and the entity set of 2 is ['foodservice', 'pie', 'business', 'fit', 'growth','strategy']. Send the entity sets into E2vec and C2vec respectively to obtain the entity vectors of each entity and their corresponding category vectors.
[0034] Step S4: Calculate the similarity between the top 20 category vectors of the (target) entity to be recognized and the entity vectors in the context of the entity to be recognized (with a moving window size of 2). The category vector with the highest score is considered the category vector of the target entity in the current text.
[0035] The calculation formula is as follows: (5) where and represent the two vectors for which the similarity needs to be calculated respectively. The moving window size for selecting contextually adjacent entities is set to 2. By combining context semantic information, the category information of each entity is initially obtained to solve the problem of polysemy.
[0036] Using this method, entities can be mapped to different semantic concepts (categories), and the similarity scores with each category can be calculated based on the context content of the entities. For example, the word "Oracle" in "Oracle has a very high market value" will be mapped to categories such as "big company", "Silicon Valley giant", "ancient characters", "Ellison", etc. According to the definition of "market value" in the Probase knowledge graph and the context in which it is used, combined with "Oracle", it can be calculated that Oracle here refers to the category of "Silicon Valley giant" rather than "ancient characters".
[0037] Step S5: Concatenate the three types of vectors (word vectors, entity vectors, and category vectors of the target entity) to obtain the final semantically enhanced word vectors.
[0038] This includes words that do not belong to the entity set, and their entity vectors and category vectors are set to 0. If an entity consists of multiple words, the word vectors of each word in the entity need to be summed first, and the result is used as the word vector of this entity.
[0039] Step S6: Input the semantically enhanced word vectors into the fine-tuned pre-trained Siamese BERT model to further extract the context semantic features of the short text pair, and obtain the vectors (x and y) of the short text pair after integrating the full-text semantic information.
[0040] Specifically, it includes: Select the MRPC dataset to fine-tune the pre-trained Siamese BERT model, specifically including: Select the MRPC dataset to fine-tune the pre-trained Siamese BERT model. First, divide the MRPC dataset into training data and test data. Each group of data is a pair of text samples, and each sample contains two text strings and a label, where the label indicates whether the two texts are similar. Then, construct a Siamese BERT model, which includes two BERT encoders and a similarity calculation layer. The two BERT encoders share parameters and are respectively used to encode the two input texts. Finally, use the training data to train the Siamese BERT model, and update the model parameters through the backpropagation algorithm and the optimization algorithm to optimize the accuracy of the model in calculating the similarity. The Siamese BERT model diagram of this method is as shown in Figure 3 shown.
[0041] When fine-tuning the Siamese BERT model, since the similarity label of the dataset is a 0-1 label, where 0 indicates dissimilar and 1 indicates similar; and the similarity score calculated by the model ranges from 0 to 1, a threshold is set for the similarity calculation result. If the result is less than 0.5, it is regarded as dissimilar, and if it is greater than 0.5, it is regarded as similar.
[0042] Input the semantic information enhanced word vectors into the Siamese BERT model, specifically including: Use the semantic information enhanced word vectors of the short text pair as the initial word vectors of the model, and input them into the left and right sides of the Siamese BERT model respectively. Through the shared encoding layer, learn the context semantic features of the two texts, enhance the vector representation ability of the text semantics, and the output is the vectors of the two texts after integrating the full-text semantic information.
[0043] Step S7: Calculate the similarity score of the short text pair, specifically including: Extract the final output of the Siamese BERT model described in step S6, that is, the feature vectors of the two texts and , and obtain the similarity score of the short text pair by calculating the Pearson correlation coefficient of the vectors. The specific formula is as follows: (6) where x and y respectively represent the feature vectors of the two texts, and respectively represent the means of the two feature vectors. The right half of the numerator of formula (6) represents the Pearson correlation coefficient. Through the above formula, the P value is restricted between 0 and 1, and can represent the similarity score of the model.
[0044] The above-described embodiments are only preferred specific embodiments of the present invention, and the protection scope of the present invention is not limited thereto. Any simple variations or equivalent substitutions of technical solutions that can be obviously obtained by those skilled in the art within the technical scope disclosed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for calculating text semantic similarity based on multi-layer semantic information fusion, characterized in that Including the steps: Using the entity-relationship triples and entity-category triples in the Probase knowledge graph, construct an entity vector model E2vec and a category vector model C2vec respectively; Use the word vector model Word2vec to obtain the word vectors of each word in the text pair to be calculated for similarity; Use the entity library in the Probase knowledge graph as a dictionary to extract the entity set in all words in the text, and send the entities in the entity set into the entity vector model E2vec and the category vector model C2vec to obtain the entity vectors of each entity and all category vectors corresponding to each entity; Calculate the cosine similarity between the top 20 category vectors of the entity to be recognized and the context entity vectors of the entity to be recognized, and use the category vector with the highest score as the category vector of the target entity in the current text pair; Perform a concatenation operation on the word vectors, entity vectors, and category vectors of the target entity to obtain the final semantic information enhanced word vectors of each word; Input the semantic information enhanced word vectors into the Siamese BERT model to extract the context semantic features of the text pair and obtain the text pair vector after integrating the full-text semantic information; Calculate the similarity score of the text pair through the text pair vector.
2. The text semantic similarity calculation method based on multi-layer semantic information fusion according to claim 1, characterized in that, The entity vector model E2vec and the category vector model C2vec include: The training corpus of the E2vec model is the entity-relationship triples in Microsoft's Probase knowledge graph, including tens of millions of entities and the number of times they appear together; The training corpus of the C2vec model is the entity-category triples in the Probase knowledge graph, including millions of categories, the entity sets included in each category, and the probability that each entity belongs to each category.
3. The text semantic similarity calculation method based on multi-layer semantic information fusion according to claim 1, characterized in that After respectively constructing the entity vector model E2vec and the category vector model C2vec, it further includes the steps: Using the Structured Deep Network Embedding (SDNE) algorithm to pre-train the entity-relationship triples and entity-category triples, learn the semantic relationships in the knowledge graph, and obtain the pre-trained model and pre-trained model parameters; Use the pre-trained model to obtain entity vectors and category vectors.
4. The text semantic similarity calculation method based on multi-layer semantic information fusion according to claim 3, wherein The use of the Structured Deep Network Embedding (SDNE) algorithm to pre-train the entity-relationship triples and entity-category triples includes the following steps: Collect all entity-relationship triples and entity-category triples in the Probase knowledge graph, and perform data preprocessing on the entity-relationship triples and entity-category triples; Construct a network graph structure according to the preprocessed data. In the entity vector model E2vec, use entities as the nodes of the graph, use the relationships between entities as the edges of the graph, and use the frequency of co-occurrence of two entities as the edge weights of the adjacency matrix of the graph; In the category vector model C2vec, use entities and categories as the nodes of the graph, use the relationships between entities and categories as the edges of the graph, and use the probability that an entity belongs to a certain category as the edge weights of the adjacency matrix of the graph; Based on the SDNE algorithm, use the training set to train the pre-trained embedding model, and use the test set to test the pre-trained embedding model; Save the parameters and vectors of the two pre-trained embedding models that have been trained and tested to obtain entity vectors and category vectors.
5. The text semantic similarity calculation method based on multi-layer semantic information fusion according to claim 4, characterized in that, Based on the SDNE algorithm, use the training set to train the pre-trained embedding model, including the following steps: Adjust the parameters of the model to optimize the model performance; Use the multi-classification task to evaluate the model and further adjust the model parameters to map entities of the same class into a similar vector space.
6. The text semantic similarity calculation method based on multi-layer semantic information fusion according to claim 1, wherein Calculate the cosine similarity between the top 20 category vectors of the entity to be recognized and the context entity vector of the entity to be recognized, and use the category vector with the highest score as the category vector of the target entity in the current text pair. The specific calculation formula is: Among them, sim(v1, v2) is the cosine similarity; v1 and v2 respectively represent the two vectors for which the similarity needs to be calculated; Select the moving window size of the context neighboring entities to be 2, and obtain the category vector information of each entity by combining the context semantic information.
7. The text semantic similarity calculation method based on multi-layer semantic information fusion according to claim 1, characterized in that, When concatenating the word vectors, entity vectors, and category vectors of the target entity, set the entity vectors and category vectors that are not entities in the word to 0; If an entity consists of multiple words, perform an addition operation on the word vectors of each word in the entity, and use the result as the word vector of the entity.
8. A method for calculating text semantic similarity based on multi-layer semantic information fusion according to claim 1, characterized in that Use the MRPC dataset to fine-tune the Siamese BERT model, including the steps: Divide the MRPC dataset into training data and test data. Each group of data is a pair of text samples, and each sample contains two text strings and a label; Construct a Siamese BERT model that includes two BERT encoders and a similarity calculation layer; Use the training data to train the Siamese BERT model, and update the model parameters through the backpropagation algorithm and the optimization algorithm to optimize the accuracy of the model in calculating the similarity.
9. A method for calculating text semantic similarity based on multi-layer semantic information fusion according to claim 1, characterized in that, The input of the semantic information enhanced word vectors into the Siamese BERT model includes the steps: Use the semantic information enhanced word vectors of the short text pair as the initial word vectors of the model, and input them into the left and right sides of the Siamese BERT model respectively; Learn the context semantic features of the two texts through the shared encoding layer, and output the vectors of the two texts after fusing the full-text semantic information.
10. A method for calculating text semantic similarity based on multi-layer semantic information fusion as described in claim 1, characterized in that, Calculate the similarity score of the text pair through the text pair vectors. The specific formula is: where x and y represent the feature vectors of two texts respectively, and y represent the means of the two feature vectors respectively, P represents the Pearson correlation coefficient. Through the above formula, the P value is restricted between 0 and 1, representing the similarity score of the model.