Word vector construction method and system based on semantic similarity calculation

CN115221870BActive Publication Date: 2026-09-25HEFEI UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210695532.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2026-09-25
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

[0006]针对现有技术的不足,本发明提供了一种基于语义相似性计算的词向量构建方法及系统,解决了现有的词向量构建方法存在准确度差且效率不高的问题

Benefits of technology

[0036]本发明提供了一种基于语义相似性计算的词向量构建方法及系统。与现有技术相比,具备以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115221870B_ABST
    Figure CN115221870B_ABST
Patent Text Reader

Abstract

The application provides a word vector construction method and system based on semantic similarity calculation, and relates to the technical field of word vector construction. The application firstly acquires an initial word vector including a center word vector and a preceding matrix and a following matrix of the center word vector; then utilizes a convolutional neural network to simultaneously perform multiple convolution operations on the preceding matrix and the following matrix to acquire an upper and lower context aggregation matrix, and simultaneously converts the center word vector into a center word vector matrix with the same dimension as the upper and lower context aggregation matrix; finally, the semantic similarity of the upper and lower context aggregation matrix and the center word vector matrix is calculated, and the upper and lower context aggregation matrix when the value of the semantic similarity is the largest is output as the final word vector. The word vector constructed by the application has higher accuracy, and the efficiency of constructing the word vector is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of word vector construction technology, specifically to a word vector construction method and system based on semantic similarity calculation. Background Technology

[0002] Word vector technology is the process of mapping words in a corpus into computable semantic vectors, that is, transforming words into dense vectors. Word vectors provide a foundation for various downstream tasks in natural language processing, such as text representation, text classification, and named entity recognition, and the performance of these downstream tasks largely depends on the quality of word vector construction.

[0003] Currently, the main methods for constructing word vectors include those that model word vectors by statistically analyzing static word frequency information, such as TF-IDF and bag-of-words models; those that predict surrounding words based on the idea of ​​a given center word, such as skip-gram and GloVe models; and those that consider contextual semantic information, such as ELMo word embedding methods.

[0004] However, in reality, word vectors are represented in different forms in different contexts. Reasonable word vector representation and modeling should be based on the context of the corpus. Word vector modeling methods such as TF-IDF and bag-of-words model fail to consider the context, resulting in inaccurate word vectors. Skip-gram and GloVe optimize the word vector matrix by calculating the similarity between the center word and each context word separately, which cannot simultaneously gather information from all context words to calculate the center word vector. This fails to solve the problem of "one word with multiple meanings" and also results in inaccurate word vectors. Although the ELMo model considers contextual information, the sequential input-output characteristics of its LSTM prevent it from being computed in parallel. This requires ELMo to spend a long time training and iterating to update parameters, resulting in low efficiency in word vector acquisition. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a word vector construction method and system based on semantic similarity calculation, which solves the problems of poor accuracy and low efficiency in existing word vector construction methods.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] Firstly, this invention proposes a word vector construction method based on semantic similarity calculation, the method comprising:

[0010] Initial word vectors are obtained based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and the preceding and following context matrices of the center word vector;

[0011] A convolutional neural network with multiple non-shared parameters is used to perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously to obtain a context aggregation matrix; and the center word vector is converted into a center word vector matrix with the same dimension as the context aggregation matrix.

[0012] Calculate the semantic similarity between the context aggregation matrix and the center word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector.

[0013] Preferably, obtaining word vectors based on a pre-constructed dictionary matrix includes:

[0014] S11. Perform text segmentation on the corpus, and use the Skip-gram model to train original word vectors based on the segmented corpus. Construct a dictionary matrix based on the trained original word vectors.

[0015] S12. Based on the real corpus, sample the central words sequentially, and perform word extraction operations on the adjacent words of the central words with a preset window size. Then, query the dictionary matrix to obtain the initial word vector.

[0016] Preferably, the step of using a convolutional neural network with multiple non-shared parameters to simultaneously perform multiple convolution operations on the context matrix and the background matrix to obtain the context aggregation matrix includes:

[0017] S21. Starting from the central word vector, a convolutional neural network with multiple non-shared parameter convolution kernels is used to perform the first convolution operation on the preceding context matrix and the following context matrix to obtain the first context aggregation matrix.

[0018] S22. Then, the first context aggregation matrix is ​​used as the input for the second convolution operation, and the steps of S21 above are repeated until the final context aggregation matrix is ​​obtained after completing a preset number of convolution operations.

[0019] Preferably, a padding operation is performed before each convolution operation.

[0020] Preferably, converting the center word vector into a center word vector matrix with the same dimensions as the context aggregation matrix includes:

[0021] Multiply the center word vector by the weight vector to convert it into a center word vector matrix with the same dimensions as the context aggregation matrix.

[0022] Secondly, this invention also proposes a word vector construction system based on semantic similarity calculation, the system comprising:

[0023] A word vector acquisition module is used to acquire initial word vectors based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and a context matrix and a background matrix of the center word vector;

[0024] The context aggregation matrix and center word vector matrix acquisition module is used to perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously using a convolutional neural network with multiple non-shared parameters to obtain the context aggregation matrix; and to convert the center word vectors into a center word vector matrix with the same dimension as the context aggregation matrix.

[0025] The word vector output module is used to calculate the semantic similarity between the context aggregation matrix and the center word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector.

[0026] Preferably, the word vector acquisition module obtains word vectors based on a pre-constructed dictionary matrix, including:

[0027] S11. Perform text segmentation on the corpus, and use the Skip-gram model to train original word vectors based on the segmented corpus. Construct a dictionary matrix based on the trained original word vectors.

[0028] S12. Based on the real corpus, sample the central words sequentially, and perform word extraction operations on the adjacent words of the central words with a preset window size. Then, query the dictionary matrix to obtain the initial word vector.

[0029] Preferably, the context aggregation matrix and center word vector matrix acquisition module utilizes a convolutional neural network with multiple non-shared parameter convolution kernels to simultaneously perform multiple convolution operations on the preceding context matrix and the following context matrix to obtain the context aggregation matrix, including:

[0030] S21. Starting from the central word vector, a convolutional neural network with multiple non-shared parameter convolution kernels is used to perform the first convolution operation on the preceding context matrix and the following context matrix to obtain the first context aggregation matrix.

[0031] S22. Then, the first context aggregation matrix is ​​used as the input for the second convolution operation, and the steps of S21 above are repeated until the final context aggregation matrix is ​​obtained after completing a preset number of convolution operations.

[0032] Preferably, a padding operation is performed before each convolution operation.

[0033] Preferably, converting the center word vector into a center word vector matrix with the same dimensions as the context aggregation matrix includes:

[0034] Multiply the center word vector by the weight vector to convert it into a center word vector matrix with the same dimensions as the context aggregation matrix.

[0035] (III) Beneficial Effects

[0036] This invention provides a method and system for constructing word vectors based on semantic similarity calculation. Compared with existing technologies, it has the following advantages:

[0037] 1. This invention first obtains the center word vector and its context and context matrices based on a pre-constructed dictionary matrix. Then, it uses a convolutional neural network with multiple non-shared parameter kernels to perform multiple convolution operations on the context and context matrices simultaneously to obtain a context aggregation matrix. Simultaneously, the center word vector is converted into a center word vector matrix with the same dimensions as the context aggregation matrix. Finally, the semantic similarity between the context aggregation matrix and the center word vector matrix is ​​calculated, and the context aggregation matrix with the maximum semantic similarity value is output as the final word vector. The word vectors constructed by this invention have higher accuracy and are more efficient in word vector construction.

[0038] 2. This invention updates word vectors by calculating the similarity between the context aggregation matrix and the central word vector matrix. This can better obtain reasonable semantic representations of words based on different contexts of real corpora, solving the problem of "one word with multiple meanings" and making the final word vectors more accurate.

[0039] 3. This invention performs multi-level convolution on the context matrix and the following context matrix. Each layer generates word vectors that contain different levels of word features. Higher layers can capture context-based word features, while lower layers can learn grammatical information, resulting in higher accuracy of the final word vectors.

[0040] 4. This invention performs convolution operations on the context matrix generated by the surrounding words simultaneously based on the convolution of the convolutional neural network. This parallel computing method can improve the efficiency of model training and optimization and increase the speed of obtaining the final word vectors. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1This is a flowchart of the word vector construction method based on semantic similarity calculation according to the present invention;

[0043] Figure 2 This is a flowchart of the word vector construction method based on semantic similarity calculation in an embodiment of the present invention;

[0044] Figure 3 This is a schematic diagram of the convolution operation in an embodiment of the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] This application provides a word vector construction method and system based on semantic similarity calculation, which solves the problems of poor accuracy and low efficiency in existing word vector construction methods, and achieves the goal of constructing word vectors efficiently and accurately, laying the foundation for the good execution of downstream NLP tasks.

[0047] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows:

[0048] To efficiently and accurately obtain word vectors for predicted words, this application first obtains the center word vector and its context and background matrices based on a constructed dictionary matrix. Then, it uses a convolutional neural network with multiple non-shared parameter kernels to perform multiple convolution operations on the context and background matrices simultaneously, obtaining a context aggregation matrix. Simultaneously, the center word vector is converted into a center word vector matrix with the same dimensions as the context aggregation matrix. Finally, the semantic similarity between the context aggregation matrix and the center word vector matrix is ​​calculated, and the context aggregation matrix with the highest semantic similarity value is output as the final word vector. The word vectors constructed using this method have higher accuracy and are more efficient in word vector construction.

[0049] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0050] Example 1:

[0051] Firstly, this invention proposes a word vector construction method based on semantic similarity calculation, see [link to relevant documentation]. Figure 1-2 The method includes:

[0052] S1. Obtain initial word vectors based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and the context matrix and the preceding context matrix of the center word vector;

[0053] S2. Using a convolutional neural network with multiple non-shared parameters, perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously to obtain a context aggregation matrix; and convert the center word vector into a center word vector matrix with the same dimension as the context aggregation matrix.

[0054] S3. Calculate the semantic similarity between the context aggregation matrix and the central word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector.

[0055] As can be seen, the word vector construction method based on semantic similarity calculation in this embodiment first obtains the center word vector and its context matrix and background matrix based on a pre-constructed dictionary matrix; then, it uses a convolutional neural network with multiple non-shared parameter kernels to perform multiple convolution operations on the context matrix and background matrix simultaneously to obtain the context aggregation matrix, while converting the center word vector into a center word vector matrix with the same dimension as the context aggregation matrix; finally, it calculates the semantic similarity between the context aggregation matrix and the center word vector matrix, and outputs the context aggregation matrix with the maximum semantic similarity value as the final word vector. The word vectors constructed by this invention have higher accuracy and higher efficiency in word vector construction.

[0056] The following is in conjunction with the appendix Figure 1-3 The following detailed explanations of the specific steps S1-S3 are provided to illustrate the implementation process of an embodiment of the present invention. The word vector construction method based on semantic similarity calculation proposed in this embodiment specifically includes the following steps:

[0057] S1. Obtain initial word vectors based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and the context matrix and the background matrix of the center word vector.

[0058] S11. Perform text segmentation on the corpus, and use the Skip-gram model to train original word vectors based on the segmented corpus. Construct a dictionary matrix based on the trained original word vectors.

[0059] First, the Jieba word segmentation tool is used to segment the text in the corpus. Then, the Skip-gram model is used to train the segmented corpus, representing the words in the corpus as an initial word vector dictionary, thus constructing a dictionary matrix. The dictionary matrix contains V words, and each word vector is represented by c. (i) (c (i) ∈R N×1 (i = 1 ~ V).

[0060] S12. Based on the real corpus, sample the center words sequentially, and perform word extraction operations with a window size of d on the adjacent words of the center words. Then, query the dictionary matrix to obtain the initial word vector.

[0061] The process involves sampling a central word from a real corpus, obtaining d adjacent preceding and d following words based on that central word, and then querying a dictionary matrix to obtain the central word vector and the word vectors of the preceding and following words. This yields the corresponding central word vector, as well as the preceding and following word matrices. Specifically:

[0062] After processing each word in the real corpus using S11, a central word is sampled sequentially from the segmented corpus. Then, a word extraction operation with a window size of d (d = 3~10) is performed on the adjacent words of the central word, resulting in 1 central word, d preceding words, and d following words. The word vectors corresponding to the central word, preceding words, and following words are then queried from the dictionary matrix obtained in S11, yielding word vector representations S (S∈R) of (2d+1) words. N×(2d+1) ), Wherein, the center word vector c (center) The context matrix is ​​represented as:

[0063] c (context) =(c L(Context) c R(context) ), (c (context) ∈R N×2d ).

[0064] Represents the matrix to the left of the central word;

[0065] This represents the matrix to the right of the central word.

[0066] S2. Perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously using a convolutional neural network to obtain a context aggregation matrix; obtain a center word vector matrix based on the center word vector.

[0067] See Figure 3 The process of using a convolutional neural network to perform multiple convolution operations on the context matrix simultaneously is as follows:

[0068] S21. Starting from the central word vector, a convolutional neural network with multiple non-shared parameter convolution kernels is used to perform the first convolution operation on the preceding context matrix and the following context matrix simultaneously to obtain the first context aggregation matrix.

[0069] Specifically, starting with the center word vector, multiple non-shared parameter convolution kernels are used to simultaneously perform convolution operations with a stride of 1 on both sides of the feature dimension. The kernel size is 1×2. To ensure that the output feature matrix maintains the same dimension as the input context matrix, padding is performed before each convolution operation, adding a column to the edges of the context matrix on both sides. There are 2d convolution kernels during the convolution process, represented as follows: The formula for convolution operation is:

[0070]

[0071]

[0072] in, The subscript (1) represents the first convolution, and the subscripts i and j represent the values ​​in the i-th row and j-th column of the matrix after convolution; · is to multiply each pair of corresponding elements in the two vectors and then sum all the products.

[0073] After the first convolution operation using a convolutional neural network, the signal is processed by the ReLU non-linear activation function:

[0074] in, Indicates taking 0 and The larger values, Indicates taking 0 and The larger value in the middle yields the first context aggregation matrix after the first convolution operation. in:

[0075] The first superposition matrix after the first convolution operation is represented as:

[0076]

[0077] The first text matrix after the first convolution operation is represented as:

[0078]

[0079] Then, the first context matrix after the first convolution operation is normalized by normalization, normalizing each element of the matrix to between 0 and 1, resulting in N1:

[0080]

[0081] in, This means taking the minimum value among all elements. This indicates taking the maximum value among all elements.

[0082] S22. Then, the first context aggregation matrix is ​​used as the input for the second convolution operation, and the steps of S21 above are repeated until the final context aggregation matrix is ​​obtained after completing a preset number of convolution operations.

[0083] The first context aggregation matrix output after the first convolution is: Then, the normalized matrix N1 is used as the input for the second convolution operation. After performing a preset number of convolution operations in the same manner, the final context aggregation matrix output is... After standardization, N is obtained. n N n This is the final context aggregation matrix, which will be the final N. n use express, The preset number of convolution operations is 3 to 5.

[0084] In real-world corpora, the semantics of words are determined by their context, and these semantics can change depending on the context. To ensure that word vectors possess contextual information, this embodiment updates word vector representations by calculating the similarity between the context aggregation matrix and the central word vector, which effectively addresses the problem of polysemy. Furthermore, each convolutional layer generates a word vector representation during training, with word vectors at different levels having different meanings. Lower-level word vectors capture syntactic information and can be used for part-of-speech tagging; higher-level word vectors capture semantic information related to the context.

[0085] In addition, the convolution operation based on the convolutional neural network performs convolution operation on the context matrix generated by the surrounding words at the same time. This parallel computing method can improve the efficiency of model training and optimization.

[0086] 2) Obtain the center word vector matrix based on the center word vector.

[0087] Since the input center word vector is a column vector, a single dimension cannot fully represent semantic information. Therefore, the center word vector is transformed into a dimensional form with the same dimensions as the embedded context aggregation matrix. The center word vector c... (center) Multiplying by the weight vector u transforms it into a sum Same matrix dimensions, i.e., center word vector matrix N0 = c (center) ×u T , (u∈R 2d×1 After normalization, the final representation of the center word vector is obtained as follows:

[0088] S3. Calculate the semantic similarity between the context aggregation matrix and the central word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector.

[0089] Since the semantics of words in real-world corpora are determined by their context, and these semantics can change depending on the context, this embodiment updates word vector representations by calculating the similarity between the context aggregation matrix and the central word vector to ensure that word vectors possess contextual information. This effectively addresses the problem of polysemy. Specifically,

[0090] When calculating the semantic similarity between the context aggregation matrix and the center word vector matrix, the context aggregation matrix is... With the center word vector matrix The matrix obtained by the dot product is:

[0091]

[0092] Here, ⊙ represents the dot product, also known as the Hadamard product, which operates on matrices of the same shape and produces a third matrix of the same dimensions. The elements of the new matrix are the products of the corresponding elements of the original matrix.

[0093] The larger the sum of the product of the context aggregation matrix and the center word vector matrix, the more similar they are, meaning the final trained context aggregation matrix can represent the semantics of the center word in the current context. Therefore, in this process, minimizing the loss function and using gradient descent to update the weights of the convolutional neural network and the center word vector transformation weights are the objectives, i.e., finding the minimum loss.

[0094]

[0095] When the loss function is minimized, the context aggregation matrix at that moment can represent the semantics of the central word in the current context, which is the word vector to be obtained.

[0096] This completes the entire process of a word vector construction method based on semantic similarity calculation in this embodiment.

[0097] The word vectors obtained in the final embodiment can be used to complete downstream NLP tasks.

[0098] For downstream NLP tasks, the word vectors from all levels of the convolution operation in this embodiment are weighted and merged into a single word vector as input to the downstream task. This can be expressed by the formula:

[0099]

[0100] Where, N i Corresponding to the context aggregation matrix after the i-th convolution, wi The weights corresponding to the i-th context matrix.

[0101] Example 2:

[0102] Secondly, the present invention also provides a word vector construction system based on semantic similarity calculation, the system comprising:

[0103] A word vector acquisition module is used to acquire initial word vectors based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and a context matrix and a background matrix of the center word vector;

[0104] The context aggregation matrix and center word vector matrix acquisition module is used to perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously using a convolutional neural network with multiple non-shared parameters to obtain the context aggregation matrix; and to convert the center word vectors into a center word vector matrix with the same dimension as the context aggregation matrix.

[0105] The word vector output module is used to calculate the semantic similarity between the context aggregation matrix and the center word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector.

[0106] Optionally, the word vector acquisition module obtains word vectors based on a pre-built dictionary matrix, including:

[0107] S11. Perform text segmentation on the corpus, and use the Skip-gram model to train original word vectors based on the segmented corpus. Construct a dictionary matrix based on the trained original word vectors.

[0108] S12. Based on the real corpus, sample the central words sequentially, and perform word extraction operations on the adjacent words of the central words with a preset window size. Then, query the dictionary matrix to obtain the initial word vector.

[0109] Optionally, the context aggregation matrix and center word vector matrix acquisition module utilizes a convolutional neural network with multiple non-shared parameter convolution kernels to simultaneously perform multiple convolution operations on the preceding context matrix and the following context matrix to obtain the context aggregation matrix, including:

[0110] S21. Starting from the central word vector, a convolutional neural network with multiple non-shared parameter convolution kernels is used to perform the first convolution operation on the preceding context matrix and the following context matrix to obtain the first context aggregation matrix.

[0111] S22. Then, the first context aggregation matrix is ​​used as the input for the second convolution operation, and the steps of S21 above are repeated until the final context aggregation matrix is ​​obtained after completing a preset number of convolution operations.

[0112] Optionally, padding can be performed before each convolution operation.

[0113] Optionally, converting the center word vector into a center word vector matrix with the same dimensions as the context aggregation matrix includes:

[0114] Multiply the center word vector by the weight vector to convert it into a center word vector matrix with the same dimensions as the context aggregation matrix.

[0115] It is understood that the word vector construction system based on semantic similarity calculation provided in this embodiment of the invention corresponds to the word vector construction method based on semantic similarity calculation described above. The explanation, examples, and beneficial effects of the relevant content can be referred to the corresponding content in the word vector construction method based on semantic similarity calculation, and will not be repeated here.

[0116] In summary, compared with existing technologies, it has the following beneficial effects:

[0117] 1. This invention first obtains the center word vector and its context and context matrices based on a pre-constructed dictionary matrix. Then, it uses a convolutional neural network with multiple non-shared parameter kernels to perform multiple convolution operations on the context and context matrices simultaneously to obtain a context aggregation matrix. Simultaneously, the center word vector is converted into a center word vector matrix with the same dimensions as the context aggregation matrix. Finally, the semantic similarity between the context aggregation matrix and the center word vector matrix is ​​calculated, and the context aggregation matrix with the maximum semantic similarity value is output as the final word vector. The word vectors constructed by this invention have higher accuracy and are more efficient in word vector construction.

[0118] 2. This invention updates word vectors by calculating the similarity between the context aggregation matrix and the central word vector matrix. This can better obtain reasonable semantic representations of words based on different contexts of real corpora, solving the problem of "one word with multiple meanings" and making the final word vectors more accurate.

[0119] 3. This invention performs multi-level convolution on the context matrix and the following context matrix. Each layer generates word vectors that contain different levels of word features. Higher layers can capture context-based word features, while lower layers can learn grammatical information, resulting in higher accuracy of the final word vectors.

[0120] 4. This invention performs convolution operations on the context matrix generated by the surrounding words simultaneously based on the convolution of the convolutional neural network. This parallel computing method can improve the efficiency of model training and optimization and increase the speed of obtaining the final word vectors.

[0121] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0122] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A word vector construction method based on semantic similarity calculation, characterized in that, The method includes: Initial word vectors are obtained based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and the preceding and following context matrices of the center word vector; A convolutional neural network with multiple non-shared parameters is used to perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously to obtain a context aggregation matrix; and the center word vector is converted into a center word vector matrix with the same dimension as the context aggregation matrix. Calculate the semantic similarity between the context aggregation matrix and the center word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector; The step of using a convolutional neural network with multiple non-shared parameter kernels to simultaneously perform multiple convolution operations on the context matrix and the background matrix to obtain the context aggregation matrix includes: S21. Starting from the central word vector, a convolutional neural network with multiple non-shared parameter convolution kernels is used to perform the first convolution operation on the preceding context matrix and the following context matrix to obtain the first context aggregation matrix. S22. Then, the first context aggregation matrix is ​​used as the input for the second convolution operation, and the steps of S21 above are repeated until the final context aggregation matrix is ​​obtained after completing a preset number of convolution operations. Padding is performed before each convolution operation.

2. The method as described in claim 1, characterized in that, The method of obtaining word vectors based on a pre-constructed dictionary matrix includes: S11. Perform text segmentation on the corpus, and use the Skip-gram model to train original word vectors based on the segmented corpus. Construct a dictionary matrix based on the trained original word vectors. S12. Based on the real corpus, sample the central words sequentially, and perform word extraction operations on the adjacent words of the central words with a preset window size. Then, query the dictionary matrix to obtain the initial word vector.

3. The method as described in claim 1, characterized in that, The step of converting the center word vector into a center word vector matrix with the same dimension as the context aggregation matrix includes: Multiply the center word vector by the weight vector to convert it into a center word vector matrix with the same dimensions as the context aggregation matrix.

4. A word vector construction system based on semantic similarity calculation, characterized in that, The system includes: A word vector acquisition module is used to acquire initial word vectors based on a pre-constructed dictionary matrix; the initial word vectors include a center word vector, and a context matrix and a background matrix of the center word vector; The context aggregation matrix and center word vector matrix acquisition module is used to perform multiple convolution operations on the preceding context matrix and the following context matrix simultaneously using a convolutional neural network with multiple non-shared parameters to obtain the context aggregation matrix; And convert the center word vectors into a center word vector matrix with the same dimensions as the context aggregation matrix; The word vector output module is used to calculate the semantic similarity between the context aggregation matrix and the center word vector matrix, and output the context aggregation matrix with the maximum semantic similarity value as the final word vector.

5. The system as described in claim 4, characterized in that, The word vector acquisition module obtains word vectors based on a pre-constructed dictionary matrix, including: S11. Perform text segmentation on the corpus, and use the Skip-gram model to train original word vectors based on the segmented corpus. Construct a dictionary matrix based on the trained original word vectors. S12. Based on the real corpus, sample the central words sequentially, and perform word extraction operations on the adjacent words of the central words with a preset window size. Then, query the dictionary matrix to obtain the initial word vector.

6. The system as described in claim 4, characterized in that, The context aggregation matrix and center word vector matrix acquisition module utilizes a convolutional neural network with multiple non-shared parameter convolution kernels to simultaneously perform multiple convolution operations on the preceding context matrix and the following context matrix to obtain the context aggregation matrix, including: S21. Starting from the central word vector, a convolutional neural network with multiple non-shared parameter convolution kernels is used to perform the first convolution operation on the preceding context matrix and the following context matrix to obtain the first context aggregation matrix. S22. Then, the first context aggregation matrix is ​​used as the input for the second convolution operation, and the steps of S21 above are repeated until the final context aggregation matrix is ​​obtained after completing a preset number of convolution operations.

7. The system as described in claim 6, characterized in that, Padding is performed before each convolution operation.

8. The system as described in claim 4, characterized in that, The step of converting the center word vector into a center word vector matrix with the same dimension as the context aggregation matrix includes: Multiply the center word vector by the weight vector to convert it into a center word vector matrix with the same dimensions as the context aggregation matrix.

Citation Information

Patent Citations

  • Method and device for training word vector embedding model

    CN111291165A