A method for calculating text similarity of multi-dimensional vectorization based on syntactic enhancement
By combining syntactic analysis and word vector methods in text similarity calculation, the WoBERT model and Transformer's self-attention mechanism are used for vectorization, and the dimensionality reduction process is performed through a deep integration transformer, the problem of poor computing efficiency and accuracy in the existing technology is solved, and more efficient and accurate text similarity calculation is achieved.
Patent Information
- Application Number
- CN202510416311.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The prior art has poor computational efficiency and accuracy in text similarity calculation, making it difficult to effectively combine syntactic analysis and semantic features of word vectors.
The multi-dimensional vectorized text similarity calculation method based on syntax enhancement is used to generate word participle results and component labels through word participle and syntax analysis, and vectorized representation is combined with the WoBERT model and Transformer's self-attention mechanism, and dimensionality reduction is performed through the deep integration transformer, and the similarity is finally calculated by the asimil similarity formula and the weighted average method.
It improves the accuracy and efficiency of text processing, optimizes word vector representation, and accurately calculates semantic similarity between words, thereby improving the accuracy and efficiency of similarity calculation.
Smart Images

Figure CN119918526B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to a method for calculating text similarity based on syntactic enhancement and multi-dimensional vectorization. Background Art
[0002] Currently, in the field of natural language processing (NLP), text similarity calculation is an important task for understanding and processing text data. With the rapid development of the Internet and social media, the volume and complexity of text data have been increasing continuously. Traditional similarity calculation methods based on word frequency are difficult to meet the requirements of accuracy and efficiency. In recent years, the application of word vector models (such as Word2Vec, GloVe, etc.) and deep learning technologies has enabled semantic-based text similarity calculation, which can capture the semantic relationships between words by mapping words into a high-dimensional space. However, simply relying on word vectors may ignore the syntactic structure and context information of the text, thus limiting the accuracy of similarity calculation. Syntactic analysis, as an important direction of text processing, reveals the grammatical relationships between words and provides rich context information. However, traditional syntactic analysis methods often cannot effectively combine the semantic features of word vectors. Therefore, how to combine syntactic analysis with word vectors and make full use of the advantages of both to improve the accuracy of text similarity calculation has become an important research topic.
[0003] It can be seen that there is an urgent need for a method for calculating text similarity based on syntactic enhancement and multi-dimensional vectorization that can improve calculation efficiency and accuracy. Summary of the Invention
[0004] In view of this, the embodiments of the present invention provide a method for calculating text similarity based on syntactic enhancement and multi-dimensional vectorization, which at least partially solves the problem of poor calculation efficiency and accuracy in the prior art.
[0005] The embodiments of the present invention provide a method for calculating text similarity based on syntactic enhancement and multi-dimensional vectorization, including:
[0006] Step 1, perform word segmentation and syntactic analysis on the input text to obtain the word segmentation result and its corresponding constituent labels;
[0007] Step 2, respectively use the WoBERT model and the self-attention mechanism of Transformer to perform vectorization representation on the word segmentation result and fuse them to obtain a fused vector;
[0008] Step 3, perform dimensionality reduction processing on the fused vector through a deep integrated transformer to generate a target word vector with strong expression ability;
[0009] Step 4, for the target word vectors with the same constituent labels, calculate the similarity value through the sememe similarity formula;
[0010] Step 5, based on all the similarity values, adjust the final similarity score through weighted average.
[0011] According to a specific implementation manner of an embodiment of the present invention, the step 1 includes:
[0012] Step 1.1, use the jieba model to segment the input text, split the sentences in the input text into independent words, and generate a word segmentation result.
[0013] Step 1.2, input all the word segmentation results into the conditional random field model to obtain the component labels corresponding to each word segmentation result.
[0014] According to a specific implementation manner of an embodiment of the present invention, before the step 1.2, the method further includes:
[0015] Use the jieba model to segment the sample text, split the sentences in the sample text into independent words, generate a word segmentation result and form a word list accordingly. Let the input sentence be S and the word segmentation result be ;
[0016] Use the pre-trained model to perform part-of-speech tagging on the words in the word list W, and output the corresponding part-of-speech tags and form a part-of-speech list accordingly , where the part-of-speech tags include nouns, verbs, and adjectives;
[0017] Construct a corresponding feature set for each word segmentation result, where the feature set includes the word segmentation result, part of speech, context words, word length, and whether it is the first or last word;
[0018] Use the part-of-speech list to indicate the component labels corresponding to each word in the word list, and form a component label set;
[0019] Use the feature set and the component label set to train the conditional random field model, learn the relationship between words and their components by maximizing the conditional probability, and optimize the model parameters to obtain a trained conditional random field model.
[0020] According to a specific implementation manner of an embodiment of the present invention, the step 2 specifically includes:
[0021] Step 2.1, use the WoBERT model to extract the initial word vectors corresponding to each word segmentation result;
[0022] Step 2.2, use the self-attention mechanism of Transformer to generate the dynamic word vectors corresponding to each word segmentation result;
[0023] Step 2.3, fuse the initial word vectors and the dynamic word vectors to obtain a fused vector.
[0024] According to a specific implementation manner of an embodiment of the present invention, step 3 specifically includes:
[0025] Step 3.1, using a forward-looking layer to perform preliminary processing on the fusion vector through a deep neural network architecture, performing linear transformation and non-linear activation on the fusion vector through multiple neurons, and generating an output feature vector H
[0026]
[0027] where H is the output feature vector of the forward-looking layer, represents the fusion vector, is the weight matrix, is the bias vector, is the activation function;
[0028] Step 3.2, introducing an ensemble learning strategy, learning and predicting the feature vector H through multiple sub-models of an autoencoder, and integrating the outputs of different sub-models through an aggregation function to obtain an aggregation vector, where the ensemble learning strategy includes random forest and gradient boosting, and the expression of the aggregation vector is
[0029]
[0030] where is the output of different models, Aggregate is the aggregation function, represents the number of tokenization results;
[0031] Step 3.3, performing refinement processing on the aggregation vector to obtain a dimensionality-reduced vector
[0032]
[0033] where represents the weight matrix for linear transformation, represents the bias vector for linear transformation;
[0034] Step 3.4, combining context information, dynamically adjusting the dimensionality-reduced vector to obtain a target word vector
[0035]
[0036] where represents a function for non-linear mapping of the dimensionality-reduced vector based on context information, represents the context feature vector extracted from the input text.
[0037] According to a specific implementation manner of an embodiment of the present invention, step 3.2 further includes:
[0038] Introduce a sparse coding constraint to construct a loss function, and control the outputs of multiple sub-models of the autoencoder by minimizing the loss function.
[0039] According to a specific implementation manner of an embodiment of the present invention, step 4 specifically includes:
[0040] Let the target word vectors of two identical component labels be and respectively, and calculate the similarity value between them through a sememe similarity formula, where the sememe similarity formula is
[0041]
[0042]
[0043]
[0044]
[0045] where the numerator represents the sum of the concept similarities of the two target word vectors, that is, their concept inner product, and the denominator represents the product of the norms of the two target word vectors, which is used to normalize the inner product value so that the similarity is between 0 and 1, and k is the dimension of the target word vector.
[0046] According to a specific implementation manner of an embodiment of the present invention, step 5 specifically includes:
[0047] Step 5.1, let all the similarity values be , where n is the number of words;
[0048] Step 5.2, assign corresponding weights to different component labels according to the importance of the component labels or the grammatical structure;
[0049] Step 5.3, for the similarities of all component labels, aggregate these similarities through weighted geometric mean to obtain the overall sentence similarity as:
[0050] .
[0051] The multi-dimensional vectorized text similarity calculation scheme based on syntactic enhancement in the embodiments of the present invention includes: Step 1, perform word segmentation and syntactic analysis on the input text to obtain the word segmentation result and its corresponding constituent labels; Step 2, respectively use the WoBERT model and the self-attention mechanism of Transformer to perform vectorized representation on the word segmentation result and fuse them to obtain a fused vector; Step 3, perform dimensionality reduction processing on the fused vector through a deep integrated transformer to generate a target word vector with strong expression ability; Step 4, for the target word vectors with the same constituent labels, calculate the similarity value through the sememe similarity formula; Step 5, based on all the similarity values, adjust the final similarity score through weighted average.
[0052] The beneficial effects of the embodiments of the present invention are as follows: Through the scheme of the present invention, combining word segmentation, syntactic analysis with the WoBERT model and the Transformer self-attention mechanism, introducing a text similarity matching method based on the combination of syntax and word vectors, the accuracy of text processing is improved. Using a deep integrated transformer for dimensionality reduction processing optimizes the word vector representation. At the same time, through the sememe similarity formula and the weighted average method, the semantic similarity between words is accurately calculated, thereby improving the accuracy and efficiency of similarity calculation. Brief Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a schematic flowchart of a method for calculating multi-dimensional vectorized text similarity based on syntactic enhancement provided by the embodiments of the present invention. Detailed Embodiments
[0055] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0056] The following describes the implementation manners of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0057] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present invention, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement a device and / or practice a method. In addition, this device can be implemented and this method can be practiced using other structures and / or functions in addition to one or more of the aspects described herein.
[0058] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0059] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0060] An embodiment of the present invention provides a method for calculating text similarity with multi - dimensional vectorization based on syntactic enhancement. The method can be applied to the process of calculating text similarity in the natural language processing scenario.
[0061] See Figure 1 , which is a schematic flowchart of a method for calculating text similarity with multi - dimensional vectorization based on syntactic enhancement provided by an embodiment of the present invention. As Figure 1 shown, the method mainly includes the following steps:
[0062] Step 1: Segment the input text and perform syntactic analysis to obtain the segmentation results and their corresponding constituent tags.
[0063] Further, Step 1 includes:
[0064] Step 1.1: Use the Jieba model to segment the input text, split the sentences in the input text into independent words, and generate segmentation results.
[0065] Step 1.2: Input all the segmentation results into the conditional random field model to obtain the constituent tags corresponding to each segmentation result.
[0066] Optionally, before Step 1.2, the method further includes:
[0067] Use the Jieba model to segment the sample text, split the sentences in the sample text into independent words, generate segmentation results and form a word list accordingly. Let the input sentence be S, and the segmentation result be ;
[0068] Use the pre-trained model to perform part-of-speech tagging on the words in the word list W, and output the corresponding part-of-speech tags and form a part-of-speech list accordingly , where the part-of-speech tags include nouns, verbs, and adjectives;
[0069] Construct a corresponding feature set for each segmentation result, where the feature set includes the segmentation result, part of speech, context words, word length, and whether it is the first or last word;
[0070] Use the part-of-speech list to indicate the constituent tags corresponding to each word in the word list, and form a constituent tag set;
[0071] Use the feature set and the constituent tag set to train the conditional random field model, learn the relationship between words and their constituents by maximizing the conditional probability, and optimize the model parameters to obtain the trained conditional random field model.
[0072] In specific implementation, in the field of natural language processing, to enhance syntactic relations, the input text is regarded as data, and specific word segmentation and syntactic analysis are performed on each text to extract sentence constituent information. The following is the specific process of segmenting sentences and extracting constituents:
[0073] Step 1.1: Use the Jieba model to segment the input sentence. The segmentation process splits the sentence into independent words and generates a word list. Let the input sentence be S, and the segmentation result be ;
[0074] Step 1.2: Use the pre-trained model (HanLP) to perform part-of-speech tagging on the words in the word list and output the corresponding part-of-speech tags, such as nouns, verbs, adjectives, etc. Perform part-of-speech tagging on the word list W and output the part-of-speech list ;
[0075] Step 1.3: Use the Conditional Random Field (CRF) model for sentence constituent analysis. In this process, first construct a feature set for each word. The features can include the word itself, its part of speech, context words (such as the previous and next words), the length of the word, and whether it is the first or last word. These features will help the model understand the role of each word in the sentence. Next, prepare a labeled training data set, where the label corresponding to each word indicates its constituent in the sentence (such as subject, predicate, object, etc.). These labels will be used to train the CRF model. In the model training stage, input the prepared feature set and labels into the CRF model for training. The CRF learns the relationship between words and their constituents by maximizing the conditional probability and optimizes the model parameters. Construct a feature vector for each word which can be represented as:
[0076]
[0077] where represents the number of characters of the current word, an indication of whether it is a boundary word;
[0078] Let the label set be, and the model is trained by maximizing the conditional probability as follows:
[0079]
[0080] where is the normalization factor, is the feature function, is the sentence constituent;
[0081] Step 1.4: Use the trained model to predict the constituents of the newly input sentence. The model will analyze the new sentence based on the previously constructed feature set and predict the constituent labels of each word. Perform word segmentation and feature extraction on the new sentence and the model outputs the predicted constituent label Y, enhancing the syntactic analysis through this step.
[0082] Step 2: Respectively use the WoBERT model and the self-attention mechanism of Transformer to vectorize and fuse the word segmentation results to obtain a fused vector;
[0083] Based on the above embodiments, the specific steps of step 2 include:
[0084] Step 2.1: Use the WoBERT model to extract the initial word vectors corresponding to each tokenization result;
[0085] Step 2.2: Use the self-attention mechanism of Transformer to generate the dynamic word vectors corresponding to each tokenization result;
[0086] Step 2.3: Fuse the initial word vectors and the dynamic word vectors to obtain the fused vectors.
[0087] In specific implementation, each word in the word list decomposed in Step 1 is vectorized using the self-attention mechanisms of WoBERT and Transformer to reduce the deviation in semantic capture. The process of vectorizing words in these two ways is as follows:
[0088] Step 2.1: Use the WoBERT model for word-level tokenization to extract word-level vector representations, focusing on processing the semantics of words in Chinese. WoBERT is optimized for the characteristics of Chinese vocabulary levels and can effectively capture the context connections between words to generate word vectors :
[0089]
[0090] Step 2.2: Use Transformer-XL to capture long-range dependencies and dynamically adjust the word vectors so that the word vectors can be dynamically updated according to the changes in the context. Use Transformer-XL to generate dynamic word vectors :
[0091]
[0092] Finally, output these two vector representations:
[0093] 。
[0094] Step 3: Perform dimensionality reduction on the fused vectors through a deep integrated transformer to generate target word vectors with strong expressive ability;
[0095] Based on the above embodiments, Step 3 specifically includes:
[0096] Step 3.1: Use the prospective layer to perform preliminary processing on the fused vectors through a deep neural network architecture, perform linear transformation and non-linear activation on the fused vectors through multiple neurons to generate the output feature vector H
[0097]
[0098] where H is the output feature vector of the prospective layer, represents the fused vector, is the weight matrix, is the bias vector, is the activation function;
[0099] Step 3.2: Introduce an ensemble learning strategy. Learn and predict the feature vector H through multiple sub-models of the autoencoder, and integrate the outputs of different sub-models through an aggregation function to obtain an aggregated vector. Among them, the ensemble learning strategy includes random forest and gradient boosting, and the expression of the aggregated vector is
[0100]
[0101] where, are the outputs of different models, Aggregate is the aggregation function, represents the number of word segmentation results;
[0102] Step 3.3: Refine the aggregated vector to obtain a dimensionality-reduced vector
[0103]
[0104] where, represents the weight matrix for linear transformation, represents the bias vector for linear transformation;
[0105] Step 3.4: Combine the context information to dynamically adjust the dimensionality-reduced vector to obtain the target word vector
[0106]
[0107] where, represents the function for non-linearly mapping the dimensionality-reduced vector based on the context information, represents the context feature vector extracted from the input text.
[0108] Furthermore, Step 3.2 further includes:
[0109] Introduce a sparse coding constraint to construct a loss function, and control the outputs of multiple sub-models of the autoencoder by minimizing the loss function.
[0110] In specific implementation, for the high-dimensional vector E generated in Step 2, a deep ensemble transformer (DETs) is used for dimensionality reduction of high-dimensional features. The specific steps are as follows:
[0111] Step 3.1: Use the Forward-Thinking Layer to preliminarily process the input features through a deep neural network architecture, extract key features, and learn complex relationships in the high-dimensional space. At this time, the input features undergo linear transformation and non-linear activation (using the Leaky ReLU activation function) through multiple neurons to generate the output feature vector H:
[0112]
[0113] where H is the output feature vector of the Forward-Thinking Layer, is the weight matrix, is the bias vector, is the activation function (Leaky ReLU);
[0114] Step 3.2: Introduce an ensemble learning strategy, including random forest and gradient boosting, to learn and predict features through multiple models, thereby enhancing the robustness of feature representation. The outputs of different models are integrated through an aggregation function:
[0115]
[0116] where, are the outputs of different models, and Aggregate is the aggregation function;
[0117] Step 3.3: Further refine the processed feature vector to generate a dimensionality-reduced feature with strong expressive ability:
[0118]
[0119] where, represents the weight matrix for linear transformation. Its role is to perform a linear mapping on the aggregated vector compress the information in the high-dimensional space into the low-dimensional space, thereby achieving dimensionality reduction. Each weight value in the matrix is optimized to retain the meaningful information in the aggregated vector to the greatest extent while reducing data redundancy and noise. represents the bias vector for linear transformation. It introduces an offset mechanism during the dimensionality reduction process, making the mapped vector more expressive and better able to adapt to specific task requirements. The bias value can help the model adjust the output more flexibly;
[0120] Step 3.4: Introduce a sparse coding constraint into the loss function of the autoencoder to promote the model to generate more interpretable feature representations, thereby improving the sparsity and generalization ability of the features. Here, L is the total loss, including the reconstruction loss and the weight W of the sparsity regularization. The feature generation is optimized by minimizing this loss function:
[0121]
[0122] Among them, L is the total loss, is the reconstruction loss, is the weight of sparsity regularization, and W is the model weight;
[0123] Step 3.5: Combine the context information to dynamically adjust the dimension-reduced feature vectors so that the final feature representation can adapt to different application scenarios:
[0124] .
[0125] Step 4: For the target word vectors with the same component label, calculate the similarity value through the sememe similarity formula;
[0126] Based on the above embodiments, the specific steps of Step 4 include:
[0127] Let the target word vectors of two same component labels be and respectively, and calculate the similarity value between them through the sememe similarity formula, where the sememe similarity formula is
[0128]
[0129]
[0130]
[0131]
[0132] Among them, the numerator represents the total concept similarity of the two target word vectors, that is, their concept inner product, and the denominator represents the product of the norms of the two target word vectors, which is used to normalize the inner product value so that the similarity is between 0 and 1, and k is the dimension of the target word vector.
[0133] Specifically in implementation, for the word vectors generated in Step 3, calculate the similarity of the word vectors of the same Y in two sentences. If a certain Y only appears in one sentence and not in the other sentence, it is directly removed. Let the word vectors of two same Y be and respectively;
[0134] Step 4.1: For the vectors and , define their similarity as the similarity value obtained by their sememe similarity calculation, and the similarity formula:
[0135]
[0136] molecule represents the sum of the conceptual similarities of two vectors, i.e., their conceptual inner product. The denominator represents the product of the norms of two vectors, which is used to normalize the inner product value so that the similarity is between 0 and 1;
[0137] Step 4.2: The inner product of two vectors and is defined as:
[0138]
[0139] where k is the dimension of the vector;
[0140] Step 4.3: Define the norm of the vector. Define the norm of the vector as the square root of its own inner product, i.e.:
[0141]
[0142] Similarly, define the norm of the vector as:
[0143]
[0144] The design idea of this formula is based on the cosine similarity of vectors, ensuring that the similarity score is between 0 and 1. The closer the value is to 1, the more similar the two vectors are.
[0145] Step 5. Based on all the similarity values, adjust the final similarity score through weighted average.
[0146] Based on the above embodiments, step 5 specifically includes:
[0147] Step 5.1. Let all the similarity values be respectively , where n is the number of words;
[0148] Step 5.2. According to the importance of the component labels or the grammatical structure, assign corresponding weights to different component labels ;
[0149] Step 5.3. For the similarities of all component labels, aggregate these similarities through weighted geometric mean to obtain the overall sentence similarity as:
[0150] .
[0151] In specific implementation, obtain the similarity calculation values of word vectors with the same components from Step 4, and let these similarity values be , where n is the number of words;
[0152] Step 5.1: Assign weights to different components Y according to the importance of the components or the syntactic structure. The weight of each component is determined according to its role in the sentence.
[0153] Step 5.2: For the similarities of all components , aggregate these similarities through weighted geometric mean, then the overall sentence similarity is:
[0154]
[0155] Through the geometric mean method, the similarities of each component can be organically integrated together, and at the same time, the interference of the extreme similarity of a single component on the overall result can be effectively reduced. This method makes the overall similarity of the sentence more robust and avoids the adverse effects of local outliers on the final calculation result.
[0156] The method for calculating text similarity with multi-dimensional vectorization based on syntactic enhancement provided by this embodiment combines word segmentation, syntactic analysis with the WoBERT model, and the Transformer self-attention mechanism, introduces a text similarity matching method based on the combination of syntax and word vectors, improves the accuracy of text processing, uses a deep integrated transformer for dimensionality reduction processing, optimizes the word vector representation, and at the same time accurately calculates the semantic similarity between words through the sememe similarity formula and the weighted average method, thereby improving the accuracy and efficiency of similarity calculation.
[0157] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof.
[0158] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A multi-dimensional vectorized text similarity calculation method based on syntax enhancement, characterized in that: include: Step 1: Perform word segmentation and syntactic analysis on the input text to obtain the word segmentation results and their corresponding component labels; Step 2: Use the WoBERT model and Transformer’s self-attention mechanism to vectorize and fuse the word segmentation results to obtain a fused vector. Step 3: Reduce the dimension of the fused vector through the deep integration transformer to generate a target word vector with strong expressive ability; The step 3 specifically includes: Step 3.1, use the forward-looking layer to perform preliminary processing on the fusion vector through the deep neural network architecture, perform linear transformation and nonlinear activation on the fusion vector through multiple neurons, and generate the output feature vector H Among them, H is the output feature vector of the forward-looking layer, represents the fusion vector, is the weight matrix, is the bias vector, is the activation function; Step 3.2, introduce an integrated learning strategy, learn and predict the feature vector H through multiple sub-models of the autoencoder, integrate the outputs of different sub-models through an aggregation function, and obtain an aggregate vector, wherein the integrated learning strategy includes random forest and gradient boosting, and the expression of the aggregate vector is in, is the output of different models, Aggregate is the aggregation function, Indicates the number of word segmentation results; Step 3.3: Refine the aggregated vector to obtain a reduced dimension vector in, represents the weight matrix used for linear transformation, represents the bias vector used for linear transformation; Step 3.4: Combine the context information and dynamically adjust the dimensionality reduction vector to obtain the target word vector in, represents a function that nonlinearly maps the reduced-dimensional vector based on context information, represents the context feature vector extracted from the input text; Step 4: For the target word vectors with the same component label, calculate the similarity value using the semantic similarity formula; Step 5: Based on all similarity values, adjust the final similarity score by weighted average.
2. The method according to claim 1, characterized in that The step 1 comprises: Step 1.1: Use the Jieba model to segment the input text, split the sentences in the input text into independent words, and generate segmentation results; Step 1.2: Input all the word segmentation results into the conditional random field model to obtain the component label corresponding to each word segmentation result.
3. The method according to claim 2, characterized in that Before step 1.2, the method further comprises: Use the Jieba model to segment the sample text, split the sentences in the sample text into independent words, generate the segmentation results and form a word list based on them. Suppose the input sentence is S and the segmentation result is ; Use the pre-trained model to perform part-of-speech tagging on the words in the word list W, and output the corresponding part-of-speech tags to form a part-of-speech list based on them , wherein the part-of-speech tags include nouns, verbs, and adjectives; Constructing a corresponding feature set for each word segmentation result, wherein the feature set includes the word segmentation result, part of speech, context words, word length, and whether it is the first or last word; The part-of-speech list indicates the component label corresponding to each word in the word list to form a component label set; The conditional random field model is trained using the feature set and component label set. The relationship between words and their components is learned by maximizing the conditional probability, and the model parameters are optimized to obtain a trained conditional random field model.
4. The method according to claim 3, characterized in that The step 2 specifically includes: Step 2.1, use the WoBERT model to extract the initial word vector corresponding to each word segmentation result; Step 2.2, use the Transformer's self-attention mechanism to generate a dynamic word vector corresponding to each word segmentation result; Step 2.3, fuse the initial word vector and the dynamic word vector to obtain a fused vector.
5. The method according to claim 4, characterized in that The step 3.2 also includes: Sparse coding constraints are introduced to construct the loss function, and the outputs of multiple sub-models of the autoencoder are controlled by minimizing the loss function.
6. The method according to claim 5, characterized in that The step 4 specifically includes: Assume that the target word vectors of two identical component labels are and , the similarity value between them is calculated by the semantic similarity formula, where the semantic similarity formula is Among them, the molecule Represents the sum of the conceptual similarities of the two target word vectors, that is, their conceptual inner product, the denominator It represents the product of the norms of the two target word vectors, which is used to normalize the inner product value so that the similarity is between 0 and 1. k is the dimension of the target word vector.
7. The method according to claim 6, characterized in that The step 5 specifically includes: Step 5.1, let all similarity values be , where n is the number of words; Step 5.2: Assign corresponding weights to different component labels based on their importance or grammatical structure. ; Step 5.3: Similarity of all ingredient labels , these similarities are summarized by weighted geometric mean to obtain the overall sentence similarity for: 。
Citation Information
Patent Citations
Text similarity calculation method and device, computer equipment and computer storage medium
CN110287312A
Automatic abstracting method and device based on XLNet
CN111666764A