A Mongolian word embedding method incorporating prior knowledge

CN116187311BActive Publication Date: 2025-07-25INNER MONGOLIA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211152903.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-07-25
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

蒙古语文本因为语料匮乏,无法使用目前主流的bert模型来做词嵌入,只能使用传统的word2vec方法,但是word2vec方法无法准确的表示每个词的意思

Benefits of technology

[0015] The present invention incorporates prior knowledge for verification, dynamically adjusts parameters such as dimensions and weight coefficients, and finally selects a dimension, the corresponding weight matrix, and window size that can accurately map Mongolian knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116187311B_ABST
    Figure CN116187311B_ABST
Patent Text Reader

Abstract

A Mongolian word embedding method incorporating prior knowledge. Words in the Mongolian text are placed into an empty vocabulary table without repetition to obtain the vocabulary table "vocab". The words in "vocab" are encoded using one-hot. The one-hot vectors of all words in the Mongolian text are input into the hidden layer 1 of the neural network. The one-hot vectors of each word are successively multiplied by the central word weight matrix of the hidden layer 1, and the obtained central word vectors are successively multiplied by the predicted word weight matrix of the hidden layer 2 to obtain the predicted word vectors of each word. Then, a softmax operation is performed to output new predicted word vectors. Prior knowledge is introduced to verify the cosine similarity of related words. If it is not satisfied, the hidden layer parameters are returned for dynamic adjustment. Through the verification by incorporating prior knowledge and supplemented by dynamic adjustment of the corresponding parameters, the present invention aims to improve the accuracy of vector representation for text embedding and find a better embedding result in the high-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and relates to the research on word embedding vectors of minority language texts using artificial intelligence, and particularly relates to a Mongolian word embedding method incorporating prior knowledge. Background Art

[0002] With the increasing achievements in natural language processing research based on deep learning, word embedding, as the first step of the neural network in natural language processing, has a great impact on the accuracy of experiments. Due to the lack of corpus, the current mainstream bert model cannot be used for word embedding of Mongolian texts, and only the traditional word2vec method can be used. However, the word2vec method cannot accurately represent the meaning of each word. Summary of the Invention

[0003] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a Mongolian word embedding method incorporating prior knowledge. By incorporating prior knowledge verification and dynamically adjusting the corresponding parameters, it is expected to improve the accuracy of vector representation of text embedding and find a better embedding result in the high-dimensional space.

[0004] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0005] A Mongolian word embedding method incorporating prior knowledge, characterized by comprising the following steps:

[0006] Step 1, crawl the Mongolian product review messages on the Internet to obtain the Mongolian text s. There are N words in the Mongolian text s, which are w1, w2, …, w N ;

[0007] Step 2, put w1, w2, …, w N without repetition into an empty word list to obtain the word list vocab. The number of words in the word list vocab is vocab_size;

[0008] Step 3, encode the words in the word list vocab using one-hot vector encoding;

[0009] Step 4, input the one-hot vectors of all the words in the Mongolian text into the hidden layer 1 of the neural network. Dynamically set the number of neurons in the hidden layer 1 to embedding_size. Each neuron has vocab_size weight coefficients. Use the skip-gram algorithm to traverse all the words in the Mongolian text s, that is, start from w1 until w N, the vector dimension of each word is vocab_size; perform a dimensionality reduction operation by multiplying the one-hot vectors of each word with the center word weight matrix N1 of the first hidden layer of the neural network in sequence. N1 is a matrix with vocab_size rows and embedding_size columns; concatenate the obtained center word vectors row by row in sequence to obtain the center word vector matrix M1. The center word vector matrix M1 has N rows and embedding_size columns;

[0010] Step 5: Input the center word vector matrix M1 into the second hidden layer. The number of neurons in the second hidden layer is vocab_size, and each neuron has embedding_size weight coefficients. Multiply each center word vector with the predicted word weight matrix N2 of the second hidden layer in sequence. N2 is a matrix with embedding_size rows and vocab_size columns; concatenate the obtained predicted word vectors row by row in sequence to obtain the predicted word vector matrix M2. The predicted word vector matrix M2 has N rows and vocab_size columns;

[0011] Split the predicted word vector matrix M2 row by row into the predicted word vectors of each word, perform softmax operation on each predicted word vector, output a new predicted word vector matrix. The size of the new predicted word vector matrix is 1 row and vocab_size columns. Calculate the cross entropy between each new predicted word vector and the corresponding one-hot vector. If the obtained cross entropy does not meet the requirements, return to Step 4, otherwise execute the next step.

[0012] Step 6: Introduce prior knowledge to verify the cosine similarity of relevant words. If it does not meet the requirements, return to Step 4 to dynamically adjust the hidden layer parameters. If it meets the requirements, end the process.

[0013] In the said Step 6, introduce the addition and subtraction logical equations or multiplication and division logical equations of word vectors to verify. Calculate the cosine similarity between the vector obtained from the prior knowledge operation and the vector in the center word matrix M1. If the finally obtained result is less than 0.5, it indicates that there is a corresponding dimensional space that can understand Mongolian, and output the center word matrix M1 as the word embedding matrix; if the finally obtained result is greater than 0.5, it indicates that the current dimension is not sufficient to learn the vector representing Mongolian knowledge, then return to Step 4 to adjust the hidden layer parameters.

[0014] Compared with the prior art, the beneficial effects of the present invention are:

[0015] The present invention incorporates prior knowledge for verification, dynamically adjusts parameters such as dimensions and weight coefficients, and finally selects a dimension, the corresponding weight matrix, and window size that can accurately map Mongolian knowledge.

[0016] The present invention uses prior knowledge for verification, enabling the trained Mongolian word embedding matrix to be directly used in fields such as Mongolian-Chinese translation, Mongolian sentiment analysis, and semantic recognition. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of the present invention.

[0018] Figure 2 It is a schematic diagram of the neural network structure of the present invention. Detailed Embodiment

[0019] The following will describe in detail the embodiments of the present invention with reference to the drawings and examples.

[0020] Word embedding is the process of converting unstructured natural language text into structured word vectors. Most of the existing language word embedding methods are based on the pre-trained word embedding method of bert for large-scale corpora. Although the pre-trained word embedding method is more accurate than the traditional word embedding method (word2vec), Mongolian cannot be used due to the lack of corpus resources. The traditional word embedding method (word2vec) learns a single vector for each word and cannot accurately represent the semantic relationship of each word. To solve this problem, the present invention introduces prior knowledge R to test the result of word embedding. If the measured result meets the prior knowledge R, the number of neurons embedding_size in the hidden layer of the neural network is dynamically adjusted, and the corresponding hidden layer weight matrix is adjusted. This is done until the measured cosine similarity can meet the prior knowledge R. The present invention can make the semantic relationship of the Mongolian word vectors generated by word embedding closer, consider more context information, and provide useful help for the field of Mongolian semantic analysis.

[0021] As Figure 1 shown, a Mongolian sentiment word embedding method incorporating prior knowledge of the present invention includes the following steps:

[0022] Step 1, Crawl Mongolian product review messages on the Internet to obtain Mongolian text s. There are N words in Mongolian text s, which are w1, w2,..., w N .

[0023] Exemplarily, through the date of the review, the content of the review, and different product pages, use the distributed multi-threaded crawler technology to crawl product review messages to achieve data collection, and supplement and modify the crawled data by means of manual collection. Finally, a product review data set is constructed.

[0024] Step 2, Put w1, w2,..., w N without repetition into an empty vocabulary to obtain the vocabulary vocab. The number of words in the vocabulary vocab is vocab_size.

[0025] Specifically, the vocabulary `vocab` is initially empty. Traverse all the words `w1` to `w` in the Mongolian text `s`. N , if the `t`-th word `w` t appears for the first time, put it into the vocabulary `vocab`; otherwise, skip `w`. t , until the `N`-th word `w` N is traversed, and the vocabulary `vocab` is obtained. The number of words in the vocabulary `vocab` is `vocab_size`.

[0026] The algorithm for this step can be described as follows:

[0027] The vocabulary pointer initially has an empty space. Compare `w1`, `w2`, …, `w` N sequentially with the vocabulary pointer . Put the word element `w` that does not exist in t into . The formula is as follows:

[0028]

[0029]

[0030] Pass all the contents in the finally obtained pointer into the vocabulary `vocab`. After the pointer content is passed out, it is cleared itself. The formula is as follows

[0031]

[0032] Step 3: Encode the words in the vocabulary `vocab` using one-hot vector encoding. This encoding is used to represent all the words in the Mongolian text `s`.

[0033] Specifically, for the `t`-th word `w` t in the Mongolian text `s`, if it is in the `k`-th position of the vocabulary `vocab`, then the `k`-th dimension of its one-hot vector is 1, and the rest of the dimensions are all 0.

[0034] Step 4: As Figure 2 shown, construct a simple neural network consisting of an input layer, a hidden layer, and an output layer. The hidden layer has a two-layer structure. The first layer is hidden layer 1, which is used to reduce the dimension of the one-hot vector of the central word; the second layer is hidden layer 2, which is used to output the predicted word vector. By calculating the cross-entropy between the predicted word vector and the original one-hot vector of the word, the accuracy of the neural network word embedding is judged.

[0035] First, input the one-hot vectors of all words in the Mongolian text s into the hidden layer 1 of the neural network. Dynamically set the number of neurons in the hidden layer 1 to be embedding_size. Each neuron has vocab_size weight coefficients. The weight coefficients correspond to the vocabulary vocab, which can be understood as the probability of each central word appearing. The number of neurons and the weight coefficients in the neurons can be learned by the neural network.

[0036] Use the skip-gram algorithm to traverse all words in the Mongolian text s, that is, starting from w1 until w N , multiply the one-hot vectors of each word with the central word weight matrix N1 of the hidden layer 1 of the neural network in sequence. Since the vector dimension of each word is vocab_size, and N1 is a matrix with vocab_size rows and embedding_size columns, it is equivalent to multiplying a 1*vocab_size vector with a vocab_size*embedding_size vector to obtain a 1*embedding_size central word vector. This operation is equivalent to reducing the dimension of the word vector from vocab_size to embedding_size, thus realizing the dimensionality reduction operation, which is equivalent to converting the sparse word vector representation of length vocab_size into a dense word vector representation of length embedding_size.

[0037] Concatenate the obtained central word vectors row by row to get the central word vector matrix M1. The central word vector matrix M1 has N rows and embedding_size columns. Its first row represents the central word vector of w1, and the Nth row represents the central word vector of w N .

[0038] Step 5, input the central word vector matrix M1 into the hidden layer 2. The number of neurons in the hidden layer 2 is vocab_size. Each neuron has embedding_size weight coefficients (learned by the neural network itself). Multiply each central word vector with the predicted word weight matrix N2 of the hidden layer 2 in sequence. N2 is a matrix with embedding_size rows and vocab_size columns. Here, the multiplication is to multiply the N*embedding_size matrix M1 with the embedding_size*vocab_size matrix N2. Concatenate the obtained predicted word vectors row by row to get the predicted word vector matrix M2. The predicted word vector matrix M2 has N rows and vocab_size columns. Its row order represents the predicted vectors of the words in the corresponding order in the Mongolian text. At this time, the dimension-reduced 1*embedding_size word vector becomes a 1*vocab_size vector.

[0039] The predicted word vector matrix M2 is split row - by - row into predicted word vectors for each word. Softmax operation is performed on each predicted word vector, and a new predicted word vector matrix is output. The order of the dimension with the highest probability in the normalized resulting predicted vector is the corresponding order in the vocabulary. The one - hot vector corresponding to this order in the vocabulary is taken out as the true vector.

[0040] The new predicted word vector matrix has 1 row and vocab_size columns. Let K represent all dimensions and K1 represent the first dimension of the predicted word vector, then K vocab_size K1 represents the vocab_size - th dimension of the predicted word vector. Calculate the cross - entropy between each new predicted word vector and the corresponding one - hot vector (also a 1 * vocab_size vector). If the obtained cross - entropy does not meet the requirements, return to step 4; otherwise, proceed to the next step.

[0041] Step 6: Introduce prior knowledge to verify the cosine similarity of relevant words. If it does not meet the requirements, return to step 4 to dynamically adjust the hidden - layer parameters; if it meets the requirements, end the process.

[0042] Specifically, adjusting the hidden - layer parameters here refers to adjusting the parameters of hidden layer 1, that is, its number of neurons and weight coefficients.

[0043] The steps of integrating prior knowledge in this step are as follows:

[0044] Introduce some logical equations of addition and subtraction of basic vectors. If the corresponding rows in M1 of the weight matrix (i.e., the word - vector encodings after dimensionality reduction) are mutually operated on, and the cosine - similarity results of the obtained vector results and the true values are all added up and then divided by the total number of logical equations. If the final result is less than 0.5, it indicates that there is a corresponding dimensional space that can better understand Mongolian. If the final result is greater than 0.5, it indicates that the current dimension is not sufficient to learn vectors that can represent Mongolian knowledge well, and then return to step 4 to adjust the corresponding parameters.

[0045] In one embodiment, the logical equations for integrating prior knowledge are as shown in Equation 1 The word vector of (king, king) minus the word vector of (man, man) and then plus the word vector of (woman, woman) equals the word vector of (queen, queen). That is Get the word vector of (queen, queen). Here, the word vector is obtained from the corresponding row in the weight matrix M1 and belongs to the predicted result. Query (queen, the queen)'s position in the vocabulary is concatenated into a real word vector. Calculate the cosine similarity between the predicted vector and the real vector. The formula is as follows:

[0046]

[0047] k1 represents the first dimension of the vector, k2 represents the second dimension of the vector, and kemmbeding_size represents the kemmbeding_size-th dimension of the vector. k11 represents the coordinate value of the first dimension of the real word vector, k12 represents the coordinate value of the first dimension of the predicted word vector, k21 represents the coordinate value of the second dimension of the real word vector, k22 represents the coordinate value of the second dimension of the predicted word vector, kemmbeding_size1 represents the coordinate value of the kemmbeding_size-th dimension of the real word vector, and kemmbeding_size2 represents the coordinate value of the kemmbeding_size-th dimension of the predicted word vector.

[0048] There are a total of R_size logical equations in the prior knowledge R. Add up the cosine similarities obtained from all such logical expressions in the prior knowledge R, and then divide by the number of logical expressions in the prior knowledge. The analytical formula is If the absolute value of the final result is less than 0.5, it means that this space can well map Mongolian knowledge. If the absolute value is greater than 0.5, return to step 4 to dynamically adjust the relevant parameters.

[0049] In a specific embodiment of the present invention, the Mongolian text s collected from the Internet, the Mongolian text s is In Chinese, it means "I like this book, but this book is not good-looking, so I want to buy a book cover". The text contains the words (I), (this), (book), (picture), (willingly), (comma), (but), (this), (book), (ugly), (comma), (so), (from), (I), (one), (book), (son), (frame), (sell), (will), (Full stop) Compare w1, w2, …, w 21 sequentially with the vocabulary pointer The initial space of the vocabulary pointer is empty, and it does not contain If that is w1 is sent into then the pointer contains then send w2 into and it contains it contains so send w4 into and it contains send w5 into and it contains w6 does not belong to the pointer contains the pointer contains so skip this element, and the elements in it remain unchanged, skip this element, and the elements in it remain unchanged, the pointer contains skip w 11 , and the words in it remain unchanged, send w 12 into it contains it contains skip this word and do not send it into the vocabulary pointer, and the words in it remain unchanged, it contains do not send this word into the vocabulary pointer, and the elements in it remain unchanged, it contains contains contains contains Send w 21 to the vocabulary pointer into contains At this point, all words have been traversed. Pass the elements in the vocabulary pointer into the vocabulary vocab. The vocabulary

[0050] The number of elements in vocab is vocab_size. In this example, vocab_size is 16, that is, the one-hot vectors of the words have 16 dimensions.

[0051] At the 1st position in vocab, the one-hot vector is:

[0052] [1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].

[0053] At the 2nd position in vocab, the one-hot vector is:

[0054] [0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].

[0055] At the 3rd position in vocab, the one-hot vector is:

[0056] [0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].

[0057] At the 4th position in vocab, the one-hot vector is:

[0058] [0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].

[0059] At the 5th position in vocab, the one-hot vector is:

[0060] [0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].

[0061] w6 = "◆" is at the 6th position in the vocab, and the one-hot vector is:

[0062] [0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0].

[0063] At the 7th position in the vocab, the one-hot vector is:

[0064] [0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0].

[0065] At the 2nd position in the vocab, the one-hot vector is:

[0066] [0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0].

[0067] At the 3rd position in the vocabulary, the one-hot vector is:

[0068] [0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0].

[0069] At the 8th position in the vocabulary, the one-hot vector is:

[0070] [0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0].

[0071] The rest is the same:

[0072] w 11 = "◆" one-hot vector is: [0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0].

[0073] The one-hot vector is: [0,0,0,0,0,0,00,1,0,0,0,0,0,0,0].

[0074] At the 10th position in the vocab, the one-hot vector is:

[0075] [0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0].

[0076] At the 1st position in the vocab, the one-hot vector is:

[0077] [1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]。

[0078] w 15 The one-hot vector of is: [0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0].

[0079] w 16 The one-hot vector of is: [0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0].

[0080] The one-hot vector of is: [0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0].

[0081] w 18 The one-hot vector of is: [0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0].

[0082] w 19 The one-hot vector of is: [0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0].

[0083] w 20 The one-hot vector of is: [0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0].

[0084] w 21 The one-hot vector of = "◆◆" is: [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1].

[0085] The hidden layer 1 includes embedding_size neurons, and embedding_size contains vocab_size weight coefficients (the coefficients are randomly generated by the neural network). It can be regarded as a central word matrix of vocab_size * embedding_size. The purpose is to reduce the dimension of a 1 * 16 matrix to a 1 * embedding_size matrix.

[0086] The one-hot encoding of the words in the Mongolian text s is passed into the hidden layer of the neural network. The hidden layer is divided into hidden layer 1 and hidden layer 2. As Figure 2 shown, there are embedding_size neurons in hidden layer 1. Each neuron has vocab_size, that is, 16 weight coefficients. Hidden layer 1 randomly generates these weight coefficients, and uses w1 (that is, the central word w center) The row vector [1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0] is multiplied by the central word weight matrix W in hidden layer 1. The size of the central word matrix is 16 * embedding_size. The multiplication of a 1*16 matrix and a 16*embedding_size matrix results in a 1*embedding_size vector. The result of the operation of w2 in hidden layer 1 is also a 1*embedding_size vector. These 21 vectors are concatenated, resulting in a vector of matrix N*embedding_size. N is 21. The obtained central word matrix after operation is M1.

[0087] There are a total of vocab_size neurons in hidden layer 2, that is, 16 neurons.

[0088] Multiply matrix M1 by the surrounding word weight matrix, that is, multiply the matrix of N*embedding_size by the matrix of embedding_size*vocab_size. Then the prediction matrix M2 is obtained. The size of M2 is N*vocab_size, the size of vocab_size is 16, and the first row of M2 is the prediction matrix of w1. The number of rows corresponds to the prediction vectors of the words in the corresponding order in the text.

[0089] The predicted vector is verified with the prior knowledge logical equation, using the cosine similarity formula to evaluate. If it is satisfied, the central word matrix M1 is output as the word embedding matrix, and each row is the word embedding vector of the word in the corresponding order in the text. If the prior knowledge is not satisfied, it means that the current dimensional space cannot well express the logical relationship of Mongolian. The neural network hidden layer 1 dynamically adjusts the dimension embedding_size and the corresponding weight coefficients.

Claims

1. A Mongolian word embedding method incorporating prior knowledge, characterized in that, It includes the following steps: Step 1: Crawl Mongolian product review messages on the Internet to obtain Mongolian text s. There are N words in Mongolian text s, which are w1, w2, …, w N ; Step 2, put w1, w2, …, w N into an empty vocabulary without repetition to obtain the vocabulary vocab, and the number of words in the vocabulary vocab is vocab_size; Step 3: Encode the words in the vocabulary `vocab` using one-hot vector encoding; Step 4: Input the one-hot vectors of all words in the Mongolian text into the first hidden layer of the neural network. Dynamically set the number of neurons in the first hidden layer to embedding_size. Each neuron has vocab_size weight coefficients. Use the skip-gram algorithm to traverse all words in the Mongolian text s, that is, starting from w1 until w N , and the vector dimension of each word is vocab_size; sequentially multiply the one-hot vectors of each word with the center word weight matrix N1 of the first hidden layer of the neural network for dimensionality reduction operation. N1 is a matrix with vocab_size rows and embedding_size columns; sequentially splice the obtained center word vectors by rows to obtain the center word vector matrix M1. The center word vector matrix M1 has N rows and embedding_size columns; Step 5: Pass the central word vector matrix M1 into the hidden layer 2. The number of neurons in the hidden layer 2 is `vocab_size`, and each neuron has `embedding_size` weight coefficients. Multiply each central word vector with the predicted word weight matrix N2 of the hidden layer 2 in turn. N2 is a matrix with `embedding_size` rows and `vocab_size` columns. Concatenate the obtained predicted word vectors by rows in turn to get the predicted word vector matrix M2. The predicted word vector matrix M2 has N rows and `vocab_size` columns; Split the predicted word vector matrix M2 by rows into the predicted word vectors of each word, perform softmax operation on each predicted word vector, and output a new predicted word vector matrix. The size of the new predicted word vector matrix is 1 row and `vocab_size` columns. Calculate the cross-entropy between each new predicted word vector and the corresponding one-hot vector. If the obtained cross-entropy does not meet the requirements, return to Step 4; otherwise, execute the next step; Step 6: Introduce prior knowledge to verify the cosine similarity of related words. If it is not satisfied, return to Step 4 to dynamically adjust the hidden layer parameters. If it is satisfied, end the process. The process is as follows: Introduce the addition and subtraction logical equations or multiplication and division logical equations of word vectors to verify. Calculate the cosine similarity between the vector obtained by prior knowledge operation and the vector in the central word matrix M1. If the finally obtained result is less than 0.5, it means that there is a corresponding dimensional space that can understand Mongolian, and output the central word matrix M1 as the word embedding matrix. If the finally obtained result is greater than 0.5, it means that the current dimension is not sufficient to learn the vector representing Mongolian knowledge, and return to Step 4 to adjust the hidden layer parameters.

2. The Mongolian word embedding method incorporating prior knowledge according to claim 1, wherein In the said step 2, traverse all words in the Mongolian text s. If the t-th word w t appears for the first time, put it into the vocabulary vocab; otherwise, skip w t , until the N-th word w N is traversed to obtain the vocabulary vocab.

3. The Mongolian word embedding method incorporating prior knowledge according to claim 1, characterized in that, In step 3, for the t-th word w of the Mongolian text s t , if it is located at the k-th position in the vocabulary vocab, then the k-th dimension of its one-hot vector is 1, and the remaining dimensions are all 0.

4. The Mongolian word embedding method incorporating prior knowledge according to claim 1, wherein In Step 4, the neural network includes an input layer, a hidden layer, and an output layer. The hidden layer has a two-layer structure. The first layer is the hidden layer 1, which is used to reduce the dimension of the one-hot vector of the central word. The second layer is the hidden layer 2, which is used to output the predicted word vector. The accuracy of the word embedding of the neural network is judged by calculating the cross-entropy between the predicted word vector and the original one-hot vector of the word.

5. The Mongolian word embedding method incorporating prior knowledge according to claim 1, characterized in that In Step 4, the original size of the one-hot vector of the central word is 1*`vocab_size`, which is passed into the hidden layer 1 and multiplied by N1 to obtain a row vector matrix of 1*`embedding_size`.

6. The Mongolian word embedding method incorporating prior knowledge according to claim 1, wherein In Step 6, adjusting the hidden layer parameters refers to adjusting the parameters of the hidden layer 1.

Citation Information

Patent Citations

  • Text classification method combining dynamic word embedding with part-of-speech tagging

    CN107291795A

  • Mongolian multi-modal fine-grained sentiment analysis method fusing prior knowledge model

    CN113609849A