A spelling correction method, device, equipment and storage medium of Indonesian language
Patent Information
- Application Number
- CN202310269082.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-03-17
AI Technical Summary
这意味着这类技术将无法给出登录词表以外的单词作为纠错方案,当错误单词的正确形式是一个登录词表以外的单词,这部分拼写纠错技术将无法纠正
[0047]与现有技术相比,本发明实施例公开的印尼语的拼写纠错方法、装置、终端设备和计算机可读存储介质,首先,获取待检测的印尼语语句,根据预设的印尼语词典对所述印尼语语句中的单词进行检测,得到所述印尼语语句中的错误单词;通过预设的语义提取模型,对所述错误单词的上下文进行语义提取,得到所述错误单词的上下文特征向量;并基于所述错误单词和预先构建的二元印尼语统计模型,获取所述错误单词对应的候选单词集合,使得所获取的候选单词集合与所述错误单词的相似度较高,并能够获取除语义提取模型或者encoder-decoder框架训练用的登录词表以外的印尼语单词,以提高印尼语错误单词纠错的准确性;然后,通过预先搭建的encoder-decoder框架、所述错误单词和所述上下文特征向量,从字符级上计算所述候选单词集合中每一所述候选单词的第一选取概率;根据每一所述候选单词与所述错误单词的编辑距离,对每一所述候选单词的第一选取概率进行调整,得到每一所述候选单词的第二选取概率;最后,根据每一所述候选单词的第二选取概率,从所述候选单词集合中选择第二选取概率最大的所述候选单词作为所述错误单词的纠错单词。因此,能够从候选单词集合中选取与错误单词相似度最高/可能性最高的候选单词对所述错误单词进行纠正,同时,由于候选单词均是基于现有的印尼语词典进行选择,因此,能够限制纠错方案的可选范围,以保证最后对所述错误单词进行纠正的单词是正确的,以提高印尼语单词拼写纠正的正确率。
Smart Images

Figure CN116362231B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to a method, apparatus, terminal device, and computer-readable storage medium for correcting spelling errors in Indonesian. Background Technology
[0002] Existing natural language spell correction technologies generally treat spell correction as a classification task. When training the deep learning model for spell correction, these technologies pre-define a finite vocabulary, called "vocabulary words," and then train the deep learning model to learn high-dimensional representations of these words, enabling the model to understand the meaning of the input text. During spell correction, the deep model selects the most probable word from the vocabulary as the correction solution. This means that these technologies cannot provide words outside the vocabulary as correction solutions; if the correct form of an incorrect word is a word outside the vocabulary, this spell correction technology will be unable to correct it.
[0003] As a representative low-resource language, Indonesian has an insufficient and sparse existing corpus, which means that existing deep learning models applied to Indonesian can only learn a limited number of Indonesian words. If the correct form of an incorrect word is not learned by the deep learning model, the incorrect word cannot be corrected, resulting in a low accuracy rate for spelling correction in Indonesian. Summary of the Invention
[0004] This invention provides a method, apparatus, terminal device, and computer-readable storage medium for correcting spelling errors in Indonesian, which can improve the accuracy of spelling correction in Indonesian words.
[0005] This invention provides a method for correcting spelling errors in Indonesian, including:
[0006] The Indonesian sentence to be detected is obtained, and the words in the Indonesian sentence are detected according to a preset Indonesian dictionary to obtain the erroneous words in the Indonesian sentence.
[0007] Using a pre-defined semantic extraction model, the context of the erroneous word is semantically extracted to obtain the context feature vector of the erroneous word;
[0008] Based on the erroneous word and a pre-built binary Indonesian statistical model, a candidate word set corresponding to the erroneous word is obtained; wherein, the candidate word set includes several candidate words;
[0009] Using a pre-built encoder-decoder framework, the erroneous word, and the context feature vector, calculate the first selection probability of each candidate word in the candidate word set;
[0010] Based on the edit distance between each candidate word and the erroneous word, the first selection probability of each candidate word is adjusted to obtain the second selection probability of each candidate word;
[0011] Based on the second selection probability of each candidate word, the candidate word with the highest second selection probability is selected from the candidate word set as the correction word for the erroneous word, so as to correct the erroneous word.
[0012] As an improvement to the above solution, the step of obtaining a candidate word set corresponding to the erroneous word based on the erroneous word and a pre-built binary Indonesian statistical model includes:
[0013] Words whose edit distance from the erroneous word is less than a preset threshold are selected to form a first candidate word subset;
[0014] Obtain the preceding and following words of the erroneous word in the Indonesian sentence;
[0015] The word preceding the erroneous word is taken as the preceding word, and the next word of the preceding word is predicted using a pre-constructed binary Indonesian statistical model to obtain a second candidate word subset; wherein, the second candidate word subset includes all words predicted from the preceding word;
[0016] The word following the erroneous word is taken as the next word, and the word preceding the next word is predicted using a pre-built binary Indonesian statistical model to obtain a third candidate word subset; wherein, the third candidate word subset includes all words predicted from the next word;
[0017] The intersection of the first subset of candidate words, the second subset of candidate words, and the third subset of candidate words is used to obtain the candidate word set corresponding to the erroneous word; wherein, the words in the candidate word set are candidate words.
[0018] As an improvement to the above scheme, the encoder-decoder framework consists of an encoder, a decoder, and an attention layer;
[0019] Then, the step of calculating the first selection probability of each candidate word in the candidate word set using the pre-built encoder-decoder framework, the erroneous word, and the context feature vector includes:
[0020] The erroneous word is input character by character into the encoder of the encoder-decoder framework for encoding, to obtain the hidden feature vector of the i-th character of the erroneous word and the context condition vector output at the last moment; where i is greater than or equal to 0;
[0021] The context condition vector and the context feature vector of the erroneous word are added together to obtain the initial encoding feature vector of the erroneous word;
[0022] The attention feature vector of the i-th character of the erroneous word is calculated by using the attention layer of the encoder-decoder framework, the hidden feature vector of the i-th character of the erroneous word, and the initial encoded feature vector of the erroneous word.
[0023] The attention feature vector of the i-th character of the erroneous word is concatenated with the (i-1)-th character of each candidate word to obtain the feature vector to be decoded of the i-th character of each candidate word.
[0024] The feature vector of the i-th character of each candidate word is input into the decoder of the encoder-decoder framework for decoding to obtain the character probability of the i-th character of each candidate word.
[0025] Calculate the mean of the character probabilities of all characters in each candidate word to obtain the first selection probability of each candidate word.
[0026] As an improvement to the above scheme, the step of adjusting the first selection probability of each candidate word based on the edit distance between each candidate word and the erroneous word to obtain the second selection probability of each candidate word includes:
[0027] An adjustment weight is determined based on the edit distance between each candidate word and the erroneous word; wherein the edit distance between the candidate word and the erroneous word is inversely proportional to the adjustment weight.
[0028] The first selection probability of each candidate word is multiplied by the corresponding adjustment weight to obtain the second selection probability of each candidate word.
[0029] As an improvement to the above solution, the step of extracting the context of the erroneous word using a preset semantic extraction model to obtain the context feature vector of the erroneous word includes:
[0030] The erroneous words in the Indonesian sentence are masked to obtain the sentence to be detected;
[0031] The sentence to be detected is input into a preset semantic extraction model, and the context of the erroneous word is semantically extracted to obtain the context feature vector of the erroneous word.
[0032] As an improvement to the above scheme, the semantic extraction model is the BERT-BiLSTM model;
[0033] Then, the step of inputting the statement to be detected into a preset semantic extraction model to perform semantic extraction on the context of the erroneous word and obtain the context feature vector of the erroneous word includes:
[0034] The statement to be detected is input into a preset semantic extraction model, and the BERT model in the semantic extraction model is used to extract the semantics of the statement to be detected to obtain a first semantic feature vector.
[0035] The first semantic feature vector is semantically extracted by the BiLSTM layer in the semantic extraction model to obtain the second semantic feature vector.
[0036] The features corresponding to the position of the erroneous word in the second semantic feature vector are used as the context feature vector of the erroneous word.
[0037] As an improvement to the above scheme, the number of units in the encoder of the encoder-decoder framework is twice the number of units in the BiLSTM layer of the semantic extraction model.
[0038] Accordingly, another embodiment of the present invention provides an Indonesian spelling correction device, comprising:
[0039] The data preprocessing module is used to acquire the Indonesian sentence to be detected, detect the words in the Indonesian sentence according to the preset Indonesian dictionary, and obtain the erroneous words in the Indonesian sentence.
[0040] The semantic extraction module is used to extract the semantics of the context of the erroneous word using a preset semantic extraction model, and obtain the context feature vector of the erroneous word.
[0041] The candidate word acquisition module is used to acquire a set of candidate words corresponding to the erroneous word based on the erroneous word and a pre-built binary Indonesian statistical model; wherein, the set of candidate words includes several candidate words;
[0042] The word generation module is used to calculate the first selection probability of each candidate word in the candidate word set using a pre-built encoder-decoder framework, the erroneous words, and the context feature vector;
[0043] The probability adjustment module is used to adjust the first selection probability of each candidate word based on the edit distance between each candidate word and the erroneous word, so as to obtain the second selection probability of each candidate word.
[0044] The word correction module is used to select the candidate word with the highest second selection probability from the candidate word set as the correction word for the erroneous word, based on the second selection probability of each candidate word, so as to correct the erroneous word.
[0045] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the Indonesian spelling correction method as described in any of the above.
[0046] Another embodiment of the present invention provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the Indonesian spelling correction method as described in any of the above embodiments.
[0047] Compared with existing technologies, the Indonesian spelling correction method, apparatus, terminal device, and computer-readable storage medium disclosed in this invention first acquire an Indonesian sentence to be detected, and detect the words in the Indonesian sentence according to a preset Indonesian dictionary to obtain erroneous words in the Indonesian sentence; then, through a preset semantic extraction model, semantic extraction is performed on the context of the erroneous words to obtain the context feature vector of the erroneous words; and based on the erroneous words and a pre-constructed binary Indonesian statistical model, a candidate word set corresponding to the erroneous words is obtained, such that the obtained candidate word set has a high similarity to the erroneous words, and can obtain information other than the semantic extraction model or encoder-d. The `ecoder` framework is trained using Indonesian words outside the vocabulary to improve the accuracy of Indonesian error correction. Then, using a pre-built encoder-decoder framework, the erroneous word, and the context feature vector, a first selection probability for each candidate word in the candidate word set is calculated at the character level. Based on the edit distance between each candidate word and the erroneous word, the first selection probability of each candidate word is adjusted to obtain a second selection probability. Finally, based on the second selection probability of each candidate word, the candidate word with the highest second selection probability is selected from the candidate word set as the correction word for the erroneous word. Therefore, it is possible to select the candidate word with the highest similarity / probability to the erroneous word from the candidate word set to correct the erroneous word. Furthermore, since the candidate words are selected based on an existing Indonesian dictionary, the range of possible correction schemes can be limited to ensure that the word ultimately corrected is correct, thereby improving the accuracy of Indonesian word spelling correction. Attached Figure Description
[0048] Figure 1 This is a flowchart illustrating an Indonesian spelling correction method provided in an embodiment of the present invention;
[0049] Figure 2 This is the model framework used in an Indonesian spelling correction method provided in this embodiment of the invention;
[0050] Figure 3 This is a schematic diagram of the structure of an Indonesian spelling correction device provided in an embodiment of the present invention;
[0051] Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] See Figure 1 , Figure 1 This is a flowchart illustrating an Indonesian spelling correction method according to an embodiment of the present invention.
[0054] The Indonesian spelling correction method provided in this embodiment of the invention includes the following steps:
[0055] S11. Obtain the Indonesian sentence to be detected, and detect the words in the Indonesian sentence according to the preset Indonesian dictionary to obtain the erroneous words in the Indonesian sentence.
[0056] S12. Using a preset semantic extraction model, perform semantic extraction on the context of the erroneous word to obtain the context feature vector of the erroneous word;
[0057] S13. Based on the erroneous word and the pre-constructed binary Indonesian statistical model, obtain a candidate word set corresponding to the erroneous word; wherein, the candidate word set includes several candidate words;
[0058] S14. Using the pre-built encoder-decoder framework, the erroneous word, and the context feature vector, calculate the first selection probability of each candidate word in the candidate word set;
[0059] S15. Based on the edit distance between each candidate word and the erroneous word, adjust the first selection probability of each candidate word to obtain the second selection probability of each candidate word;
[0060] S16. Based on the second selection probability of each candidate word, select the candidate word with the highest second selection probability from the candidate word set as the correction word for the erroneous word, so as to correct the erroneous word.
[0061] It should be noted that spelling error correction is a fundamental component of many applications, such as essay scoring, search engines, and speech recognition. It is defined as identifying incorrect words and providing correction suggestions. Spelling errors can be divided into two categories: (1) non-word errors, where a word is misspelled as a word that does not exist in reality, such as apple-aple. These errors can usually be detected using dictionary lookup methods; (2) true word errors, where a word is misspelled as another real word, such as apple-apply. This invention mainly focuses on the correction of non-word errors. Therefore, in step S11, after obtaining the Indonesian sentence to be detected, the words in the Indonesian sentence are detected according to a preset Indonesian dictionary to obtain the incorrect words in the Indonesian sentence, i.e., non-words. In addition, before detecting the words in the Indonesian sentence according to the preset Indonesian dictionary, the Indonesian sentence needs to be analyzed to segment the words in the Indonesian sentence. The word segmentation process can refer to other existing technologies, which will not be elaborated on here.
[0062] As a preferred embodiment, obtaining the candidate word set corresponding to the erroneous word based on the erroneous word and a pre-built binary Indonesian statistical model includes:
[0063] Words whose edit distance from the erroneous word is less than a preset threshold are selected to form a first candidate word subset;
[0064] Obtain the preceding and following words of the erroneous word in the Indonesian sentence;
[0065] The word preceding the erroneous word is taken as the preceding word, and the next word of the preceding word is predicted using a pre-constructed binary Indonesian statistical model to obtain a second candidate word subset; wherein, the second candidate word subset includes all words predicted from the preceding word;
[0066] The word following the erroneous word is taken as the next word, and the word preceding the next word is predicted using a pre-built binary Indonesian statistical model to obtain a third candidate word subset; wherein, the third candidate word subset includes all words predicted from the next word;
[0067] The intersection of the first subset of candidate words, the second subset of candidate words, and the third subset of candidate words is used to obtain the candidate word set corresponding to the erroneous word; wherein, the words in the candidate word set are candidate words.
[0068] Specifically, the binary Indonesian statistical model is constructed based on word binary pairs in a pre-set Indonesian dictionary.
[0069] Specifically, the construction steps of the binary Indonesian statistical model are as follows:
[0070] Obtain N Indonesian texts from a preset Indonesian dictionary, and segment the Indonesian texts into words; where N is greater than 1.
[0071] Count the frequency of word pairs <front_word, back_word> in all the Indonesian texts, where front_word and back_word are two adjacent words in the Indonesian text, and front_word is the word preceding back_word;
[0072] Remove word pairs with a frequency of 1 and save the frequency statistics of the final word pairs to obtain the binary Indonesian statistical model.
[0073] Preferably, the preset threshold is 3.
[0074] For example, taking a preset threshold of 3 as an example, step S13 specifically involves: obtaining words whose edit distance to the erroneous word is less than 3 to form a first candidate word subset; obtaining the preceding and following words of the erroneous word in the Indonesian sentence; querying the word pairs when the preceding word (i.e., the preceding context word) of the erroneous word is used as the front_word using the binary Indonesian statistical model, and combining their corresponding back_words to form a second candidate word subset; querying the word pairs when the following word (i.e., the following context word) of the erroneous word is used as the back_word using the binary Indonesian statistical model, and combining their corresponding front_words to form a third candidate word subset; taking the intersection of the first candidate word subset, the second candidate word subset, and the third candidate word subset to obtain the candidate word set L corresponding to the erroneous word. candidate ={w c_1_1 ,w c_1_2 ,...,w c_d_j}; where w c_d_j For the j-th candidate word with an edit distance of d in the candidate word set, w c_d_j ={c1,c2,...,c m}, c m For w c_d_j The m-th character in the string.
[0075] It is worth noting that by analyzing non-word errors in existing Indonesian datasets, we found that erroneous and correct words are highly similar in letter composition. Therefore, we can use edit distance to measure the similarity between erroneous words and words in the dictionary, filter out words with high similarity from the dictionary as candidate words, and further determine candidate words based on the previous and next words of the erroneous word, as well as the related word pairs when the previous and next words appear in existing Indonesian texts. This constrains the selection range of candidate words and improves the accuracy of Indonesian spelling error correction.
[0076] In some preferred embodiments, the encoder-decoder framework consists of an encoder, a decoder, and an attention layer;
[0077] Then, the step of calculating the first selection probability of each candidate word in the candidate word set using the pre-built encoder-decoder framework, the erroneous word, and the context feature vector includes:
[0078] The erroneous word is input character by character into the encoder of the encoder-decoder framework for encoding, to obtain the hidden feature vector of the i-th character of the erroneous word and the context condition vector output at the last moment; where i is greater than or equal to 0;
[0079] The context condition vector and the context feature vector of the erroneous word are added together to obtain the initial encoding feature vector of the erroneous word;
[0080] The attention feature vector of the i-th character of the erroneous word is calculated by using the attention layer of the encoder-decoder framework, the hidden feature vector of the i-th character of the erroneous word, and the initial encoded feature vector of the erroneous word.
[0081] The attention feature vector of the i-th character of the erroneous word is concatenated with the (i-1)-th character of each candidate word to obtain the feature vector to be decoded of the i-th character of each candidate word.
[0082] The feature vector of the i-th character of each candidate word is input into the decoder of the encoder-decoder framework for decoding to obtain the character probability of the i-th character of each candidate word;
[0083] Calculate the mean of the character probabilities of all characters in each candidate word to obtain the first selection probability of each candidate word.
[0084] In this embodiment, the encoder consists of an LSTM network, and the decoder consists of an LSTM network.
[0085] Specifically, the step of inputting the erroneous word character by character into the encoder of the encoder-decoder framework for encoding, to obtain the hidden feature vector of the i-th character of the erroneous word and the context condition vector output at the last moment, is as follows:
[0086] The erroneous word is input character by character into the encoder of the encoder-decoder framework. The erroneous word is encoded according to the following formula to obtain the hidden feature vector of the i-th character of the erroneous word and the context condition vector output at the last moment:
[0087] h i =LSTM encoder (x i )
[0088] output final =h m
[0089] Where, x i For the i-th character in the erroneous word, LSTM encoder h represents the encoder operation. i h is the hidden feature vector of the i-th character of the incorrect word. m The output is the hidden feature vector of the m-th character of the incorrect word. final This is the context condition vector output by the encoder at the last moment.
[0090] Specifically, the initial encoding feature vector of the erroneous word is V. int_state =V hidden_state +output final Among them, V hidden_state This is the context feature vector of the erroneous word.
[0091] It should be noted that the initial encoded feature vector of the erroneous word is used as the initial state of the decoder.
[0092] Further, the step of calculating the attention feature vector of the i-th character of the erroneous word input by the encoder through the attention layer of the encoder-decoder framework, the hidden feature vector of the i-th character of the erroneous word, and the initial encoded feature vector of the erroneous word is specifically as follows:
[0093] The hidden feature vector of the i-th character of the erroneous word and the initial encoded feature vector of the erroneous word are input into the attention layer of the encoder-decoder framework. The attention feature vector of the i-th character of the erroneous word input to the encoder is calculated according to the following formula:
[0094] W attention =Softmax(W3tanh(W1V) int_state +W2h i +b1)+b2)
[0095] V content =∑W attention ·V int_state
[0096] Among them, W attention The encoder is input with the attention weight of the i-th character of the erroneous word, tanh is the activation function, W1, W2, W3, b1, b2 are network optimization parameters, and V... content The encoder is input with the attention feature vector of the i-th character of the erroneous word.
[0097] Specifically, the feature vector to be decoded for the i-th character of each candidate word is calculated using the following formula: X decoder_t =(V content |c i-1 ); where X decoder_t Let c be the feature vector to be decoded for the i-th character of the candidate word. i-1 It is the (i-1)th character of the candidate word.
[0098] It is worth noting that for any candidate word, its 0th character is c0= <START〉。
[0099] Specifically, the character probability of the i-th character of each candidate word is calculated using the following formula: in, LSTM represents the character probability of the i-th character in a candidate word. decoder This is the operation of the decoder.
[0100] Specifically, the first selection probability of each candidate word is calculated using the following formula: in, For the candidate word w c_d_j The first selection probability, m is the candidate word w c_d_j Total number of characters.
[0101] In one specific implementation, adjusting the first selection probability of each candidate word based on the edit distance between each candidate word and the erroneous word to obtain a second selection probability of each candidate word includes:
[0102] An adjustment weight is determined based on the edit distance between each candidate word and the erroneous word; wherein the edit distance between the candidate word and the erroneous word is inversely proportional to the adjustment weight.
[0103] The first selection probability of each candidate word is multiplied by the corresponding adjustment weight to obtain the second selection probability of each candidate word.
[0104] In one specific implementation, taking a preset threshold of 3 as an example, if the edit distance between the candidate word and the erroneous word is 1, the adjustment weight is 1.73; if the edit distance between the candidate word and the erroneous word is 2, the adjustment weight is 1.19; and if the edit distance between the candidate word and the erroneous word is 3, the adjustment weight is 1.03. Understandably, since candidate words with smaller edit distances are more likely to be the correct correction solution than candidate words with larger edit distances, to reflect this characteristic, we adjust the weight based on the candidate word w... c_d_j The edit distance d is used to fine-tune the first selection probability of each candidate word by multiplying it by the corresponding adjustment weight, so that candidate words with smaller edit distances receive more attention.
[0105] It is worth noting that in actual operation, after calculating the second selection probability of each candidate word, the candidate words are sorted in descending order of their second selection probabilities. The top 5 candidate words with the highest second selection probabilities are selected as the recommended error correction scheme from the candidate word set. In this embodiment, the candidate word with the highest second selection probability is used as the default error correction scheme. Of course, the second, third, or fourth candidate word with the highest second selection probability can also be selected as the default error correction scheme depending on the actual situation; no specific limitation is made here.
[0106] It is worth noting that existing spell correction models based primarily on neural networks treat the error correction task as a classification task at the lexical level. This means that existing spell correction models can only classify incorrect words into a single word from a pre-learned glossary. In contrast, this invention treats Indonesian spell correction as a generation task. The character composition of the incorrect word is used as input to a pre-built encoder-decoder framework. Combined with the contextual feature vector of the incorrect word provided by a semantic extraction model, a character-by-character method is used to determine the probability of generating candidate words from the incorrect word, thus generating a correction scheme character by character. This is different from the existing method of selecting a correction scheme from the glossary. Therefore, a correction scheme beyond the glossary can be provided.
[0107] Specifically, the step of extracting the context of the erroneous word using a preset semantic extraction model to obtain the context feature vector of the erroneous word includes:
[0108] The erroneous words in the Indonesian sentence are masked to obtain the sentence to be detected;
[0109] The sentence to be detected is input into a preset semantic extraction model, and the context of the erroneous word is semantically extracted to obtain the context feature vector of the erroneous word.
[0110] For example, given an Indonesian sentence S = {w1, w2, ..., w...} of length n... i ,...,w n Let w be the i-th word in the Indonesian sentence. i For the sake of clarity, the incorrect word is denoted as X, where X = w i The error words in the Indonesian sentence are masked using [MASK], resulting in the sentence S to be detected. [MASK] ={w1,w2,...,[MASK],...,w n}
[0111] Furthermore, the semantic extraction model is the BERT-BiLSTM model;
[0112] Then, the step of inputting the statement to be detected into a preset semantic extraction model to perform semantic extraction on the context of the erroneous word and obtain the context feature vector of the erroneous word includes:
[0113] The statement to be detected is input into a preset semantic extraction model, and the BERT model in the semantic extraction model is used to extract the semantics of the statement to be detected to obtain a first semantic feature vector.
[0114] The first semantic feature vector is semantically extracted by the BiLSTM layer in the semantic extraction model to obtain the second semantic feature vector.
[0115] The features corresponding to the position of the erroneous word in the second semantic feature vector are used as the context feature vector of the erroneous word.
[0116] Specifically, the first semantic feature vector H = BERT(S [MASK] ); where BERT() represents the model operation of the BERT model, S [MASK] The statement to be tested.
[0117] Specifically, the second semantic feature vector H l =BiLSTM(H); where BiLSTM() represents the model operation of the BiLSTM layer, and H l ={h l1 ,h l2 ,...,h lMASK ,...,h ln}, h ln h is the nth feature in the second semantic feature vector. lMASK The features are the positions corresponding to the erroneous words in the second semantic feature vector.
[0118] Specifically, the context feature vector V of the erroneous word hidden_state =h lMASK .
[0119] Specifically, the number of units in the encoder of the encoder-decoder framework is twice the number of units in the BiLSTM layer of the semantic extraction model.
[0120] It is worth noting that if the encoder is an LSTM network, the number of units in the encoder's LSTM network is twice the number of units in the BiLSTM layer of the semantic extraction model, to ensure that the context condition vector of the erroneous word and the context feature vector can be added together.
[0121] It is worth noting that the semantic extraction model is a BERT model with BiLSTM layers. Compared to ordinary network models, the BERT model can better capture contextual feature vectors. The candidate word set is obtained by selecting high-probability candidate words from the existing Indonesian dictionary, which limits the range of error correction schemes and makes the correction more accurate. The word generation module is an encoder-decoder framework that can combine the character composition of the erroneous word with the contextual feature vector to generate an error correction scheme character by character. Therefore, this invention does not rely heavily on the amount of Indonesian labeled data. Since the semantic extraction model used in this invention can be pre-trained, the semantic extraction module and the encoder-decoder framework only need a small amount of data for fine-tuning when training the error correction task. In addition, since the candidate word acquisition uses edit distance and N-gram statistical language models, it does not require data specifically for Indonesian spelling error correction. Compared with existing technologies, the Indonesian spelling error correction method provided in this embodiment has a better error correction effect on Indonesian text spelling errors.
[0122] See Figure 2 This is the model framework used in an Indonesian spelling correction method provided in this embodiment of the invention. The following is a brief description of the Indonesian spelling correction method provided in this embodiment using a specific example:
[0123] Given an Indonesian sentence S = {w1, w2, ..., wn} of length n. i ,...,w n}; where w i The incorrect word X = w is of length m. i X = {x1, x2, ..., x} m Inputting X and S into the corresponding model framework of this invention, the output target word Y = {y1, y2, ..., y} m Since non-word error detection is typically achieved through dictionary-based detection, the task can be formulated as a conditional generation problem, achieved by modeling and maximizing the conditional probability p(Y|XS). See also Figure 2 The model framework shown describes how non-word X in S is masked and then input into a semantic extraction model to obtain the contextual feature vector V of X at the lexical level. hidden_state V hidden_state As the initial state, X is cyclically input into the constructed encoder-decoder framework for encoding, obtaining the context condition vector output. final Then, the correction scheme Y is generated. predictL represents the final selected candidate words. In practical applications, to ensure the generated results are free of spelling errors, a simulated input scoring generation strategy is adopted. Candidate words with an edit distance of less than or equal to 3 from X and an n-gram relationship are selected from the dictionary, forming the candidate word set L. candidate ={w c_1_1 ,w c_1_2 ,...,w c_d_j Then, the decoder in the encoder-decoder framework is used to calculate the confidence score of the candidate words, i.e., the first selection probability, thus obtaining the optimal error correction scheme. Figure 2 p shown ym To generate the m-th character y given the first m-1 characters. m The character probability, P Y The first selection probability for generating the target word Y.
[0124] See Figure 3 This is a schematic diagram of the structure of an Indonesian spelling correction device provided in an embodiment of the present invention.
[0125] The Indonesian spelling correction device provided in this embodiment of the invention includes:
[0126] The data preprocessing module 21 is used to acquire the Indonesian sentence to be detected, detect the words in the Indonesian sentence according to the preset Indonesian dictionary, and obtain the erroneous words in the Indonesian sentence.
[0127] The semantic extraction module 22 is used to extract the semantics of the context of the erroneous word through a preset semantic extraction model, and obtain the context feature vector of the erroneous word.
[0128] The candidate word acquisition module 23 is used to acquire a set of candidate words corresponding to the erroneous word based on the erroneous word and a pre-built binary Indonesian statistical model; wherein, the set of candidate words includes several candidate words;
[0129] The word generation module 24 is used to calculate the first selection probability of each candidate word in the candidate word set using a pre-built encoder-decoder framework, the erroneous word, and the context feature vector;
[0130] The probability adjustment module 25 is used to adjust the first selection probability of each candidate word based on the edit distance between each candidate word and the erroneous word, so as to obtain the second selection probability of each candidate word.
[0131] The word correction module 26 is used to select the candidate word with the highest second selection probability from the candidate word set as the correction word for the erroneous word, based on the second selection probability of each candidate word, so as to correct the erroneous word. As an improvement to the above solution,
[0132] In one preferred embodiment, the candidate word acquisition module 23 is specifically used for:
[0133] Words whose edit distance from the erroneous word is less than a preset threshold are selected to form a first candidate word subset;
[0134] Obtain the preceding and following words of the erroneous word in the Indonesian sentence;
[0135] The word preceding the erroneous word is taken as the preceding word, and the next word of the preceding word is predicted using a pre-constructed binary Indonesian statistical model to obtain a second candidate word subset; wherein, the second candidate word subset includes all words predicted from the preceding word;
[0136] The word following the erroneous word is taken as the next word, and the word preceding the next word is predicted using a pre-built binary Indonesian statistical model to obtain a third candidate word subset; wherein, the third candidate word subset includes all words predicted from the next word;
[0137] The intersection of the first subset of candidate words, the second subset of candidate words, and the third subset of candidate words is used to obtain the candidate word set corresponding to the erroneous word; wherein, the words in the candidate word set are candidate words.
[0138] As one preferred embodiment, the encoder-decoder framework in the word generation module 24 consists of an encoder, a decoder, and an attention layer;
[0139] Therefore, the word generation module 24 includes:
[0140] The feature vector encoding unit is used to input the erroneous word character by character into the encoder of the encoder-decoder framework for encoding, to obtain the hidden feature vector of the i-th character of the erroneous word and the context condition vector output at the last moment; where i is greater than or equal to 0;
[0141] The feature vector initialization unit is used to add the context condition vector of the erroneous word and the context feature vector to obtain the initial encoding feature vector of the erroneous word;
[0142] The attention feature processing unit is used to calculate the attention feature vector of the i-th character of the erroneous word input by the encoder through the attention layer of the encoder-decoder framework, the hidden feature vector of the i-th character of the erroneous word, and the initial encoded feature vector of the erroneous word;
[0143] The feature vector concatenation unit is used to concatenate the attention feature vector of the i-th character of the erroneous word with the (i-1)-th character of each candidate word to obtain the feature vector to be decoded of the i-th character of each candidate word.
[0144] The character probability calculation unit is used to input the feature vector to be decoded of the i-th character of each candidate word into the decoder of the encoder-decoder framework for decoding, so as to obtain the character probability of the i-th character of each candidate word;
[0145] The candidate word probability calculation unit is used to calculate the average value of the character probabilities of all characters in each candidate word to obtain the first selection probability of each candidate word.
[0146] Specifically, the probability adjustment module 25 is used for:
[0147] An adjustment weight is determined based on the edit distance between each candidate word and the erroneous word; wherein the edit distance between the candidate word and the erroneous word is inversely proportional to the adjustment weight.
[0148] The first selection probability of each candidate word is multiplied by the corresponding adjustment weight to obtain the second selection probability of each candidate word.
[0149] Specifically, the semantic extraction module 22 includes:
[0150] The MASK processing unit is used to mask the erroneous words in the Indonesian sentence to obtain the sentence to be detected.
[0151] The semantic extraction unit is used to input the statement to be detected into a preset semantic extraction model, extract the semantics of the context of the erroneous word, and obtain the context feature vector of the erroneous word.
[0152] Furthermore, the semantic extraction model is the BERT-BiLSTM model;
[0153] Therefore, the semantic extraction unit is specifically used for:
[0154] The statement to be detected is input into a preset semantic extraction model, and the BERT model in the semantic extraction model is used to extract the semantics of the statement to be detected to obtain a first semantic feature vector.
[0155] The first semantic feature vector is semantically extracted by the BiLSTM layer in the semantic extraction model to obtain the second semantic feature vector.
[0156] The features corresponding to the position of the erroneous word in the second semantic feature vector are used as the context feature vector of the erroneous word.
[0157] It should be noted that the specific descriptions and beneficial effects of the various embodiments of the Indonesian spelling correction device in this embodiment can be found in the specific descriptions and beneficial effects of the various embodiments of the Indonesian spelling correction method described above, and will not be repeated here.
[0158] See Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of the present invention.
[0159] An embodiment of the present invention provides a terminal device including a processor 10, a memory 20, and a computer program stored in the memory 20 and configured to be executed by the processor 10. When the processor 10 executes the computer program, it implements the Indonesian spelling correction method as described in any of the above embodiments.
[0160] When the processor 10 executes the computer program, it implements the steps in the above-described embodiment of the Indonesian spelling correction method, for example... Figure 1 All steps of the Indonesian spelling correction method shown. Alternatively, when the processor 10 executes the computer program, it implements the functions of each module / unit in the above-described Indonesian spelling correction device embodiment, for example... Figure 3 The functions of each module of the Indonesian spelling correction device are shown.
[0161] For example, the computer program may be divided into one or more modules, which are stored in the memory 20 and executed by the processor 10 to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the terminal device.
[0162] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor 10 and a memory 20. Those skilled in the art will understand that the schematic diagram is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0163] The processor 10 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. The processor 10 is the control center of the terminal device, connecting various parts of the terminal device via various interfaces and lines.
[0164] The memory 20 can be used to store the computer programs and / or modules. The processor 10 implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory 20 and calling the data stored in the memory 20. The memory 20 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0165] Wherein, if the modules / units integrated in the terminal device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the various method embodiments described above. Wherein, the computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0166] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0167] Another embodiment of the present invention provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the Indonesian spelling correction method as described in any of the above method embodiments.
[0168] In summary, the Indonesian spelling correction method, apparatus, terminal device, and computer-readable storage medium provided by this invention first acquire an Indonesian sentence to be detected, and detect the words in the Indonesian sentence according to a preset Indonesian dictionary to obtain erroneous words in the Indonesian sentence; then, through a preset semantic extraction model, semantic extraction is performed on the context of the erroneous words to obtain the context feature vector of the erroneous words; and based on the erroneous words and a pre-constructed binary Indonesian statistical model, a candidate word set corresponding to the erroneous words is obtained, such that the obtained candidate word set has a high similarity to the erroneous words, and can obtain the semantic extraction model or encoder-decoder framework. The training uses Indonesian words outside the registered vocabulary to improve the accuracy of Indonesian error correction and ensure the ability to correct OOV (Out-of-V) words. Then, using a pre-built encoder-decoder framework, the erroneous word, and the context feature vector, a first selection probability of each candidate word in the candidate word set is calculated at the character level. Based on the edit distance between each candidate word and the erroneous word, the first selection probability of each candidate word is adjusted to obtain a second selection probability. Finally, based on the second selection probability of each candidate word, the candidate word with the highest second selection probability is selected from the candidate word set as the correction word for the erroneous word. Therefore, it is possible to select the candidate word with the highest similarity / probability to the erroneous word from the candidate word set to correct the erroneous word. Furthermore, since the candidate words are selected based on an existing Indonesian dictionary, the range of possible correction schemes can be limited to ensure that the word corrected is ultimately correct, thereby improving the accuracy of Indonesian word spelling correction.
[0169] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for correcting spelling errors in Indonesian, characterized in that, include: The Indonesian sentence to be detected is obtained, and the words in the Indonesian sentence are detected according to a preset Indonesian dictionary to obtain the erroneous words in the Indonesian sentence. Using a pre-defined semantic extraction model, the context of the erroneous word is semantically extracted to obtain the context feature vector of the erroneous word; Based on the erroneous word and a pre-built binary Indonesian statistical model, a candidate word set corresponding to the erroneous word is obtained; wherein, the candidate word set includes several candidate words; Using a pre-built encoder-decoder framework, the erroneous word, and the context feature vector, calculate the first selection probability of each candidate word in the candidate word set; Based on the edit distance between each candidate word and the erroneous word, the first selection probability of each candidate word is adjusted to obtain the second selection probability of each candidate word; Based on the second selection probability of each candidate word, the candidate word with the highest second selection probability is selected from the candidate word set as the correction word for the erroneous word, so as to correct the erroneous word; The encoder-decoder framework consists of an encoder, a decoder, and an attention layer. The step of calculating the first selection probability of each candidate word in the candidate word set using the pre-built encoder-decoder framework, the erroneous words, and the contextual feature vectors includes: The erroneous word is input character by character into the encoder of the encoder-decoder framework for encoding, to obtain the first character of the erroneous word. The hidden feature vectors of each character and the context condition vector output at the last moment; where, Greater than or equal to 0; The context condition vector and the context feature vector of the erroneous word are added together to obtain the initial encoding feature vector of the erroneous word; Through the attention layer of the encoder-decoder framework, the first... The hidden feature vector of the nth character and the initial encoded feature vector of the erroneous word are used to calculate the nth character of the erroneous word input by the encoder. Attention feature vectors for each character; The first of the incorrect words The attention feature vector of the nth character and the nth character of each candidate word The characters are concatenated to obtain the first character of each candidate word. The feature vector of the character to be decoded; The first of each of the candidate words The feature vector of the nth character to be decoded is input into the decoder of the encoder-decoder framework for decoding, to obtain the nth character of each candidate word. The probability of a character; Calculate the mean of the character probabilities of all characters in each candidate word to obtain the first selection probability of each candidate word.
2. The Indonesian spelling correction method as described in claim 1, characterized in that, The process of obtaining a candidate word set corresponding to the erroneous word based on the erroneous word and a pre-built binary Indonesian statistical model includes: Words whose edit distance from the erroneous word is less than a preset threshold are selected to form a first candidate word subset; Obtain the preceding and following words of the erroneous word in the Indonesian sentence; The word preceding the erroneous word is taken as the preceding word, and the next word of the preceding word is predicted using a pre-constructed binary Indonesian statistical model to obtain a second candidate word subset; wherein, the second candidate word subset includes all words predicted from the preceding word; The word following the erroneous word is taken as the next word, and the word preceding the next word is predicted using a pre-built binary Indonesian statistical model to obtain a third candidate word subset; wherein, the third candidate word subset includes all words predicted from the next word; The intersection of the first subset of candidate words, the second subset of candidate words, and the third subset of candidate words is used to obtain the candidate word set corresponding to the erroneous word; wherein, the words in the candidate word set are candidate words.
3. The Indonesian spelling correction method as described in claim 1, characterized in that, The step of adjusting the first selection probability of each candidate word based on the edit distance between each candidate word and the erroneous word to obtain the second selection probability of each candidate word includes: An adjustment weight is determined based on the edit distance between each candidate word and the erroneous word; wherein the edit distance between the candidate word and the erroneous word is inversely proportional to the adjustment weight. The first selection probability of each candidate word is multiplied by the corresponding adjustment weight to obtain the second selection probability of each candidate word.
4. The Indonesian spelling correction method as described in claim 1, characterized in that, The step of extracting semantics from the context of the erroneous word using a preset semantic extraction model to obtain the context feature vector of the erroneous word includes: The erroneous words in the Indonesian sentence are masked to obtain the sentence to be detected; The sentence to be detected is input into a preset semantic extraction model, and the context of the erroneous word is semantically extracted to obtain the context feature vector of the erroneous word.
5. The Indonesian spelling correction method as described in claim 4, characterized in that, The semantic extraction model is the BERT-BiLSTM model; Then, the step of inputting the statement to be detected into a preset semantic extraction model to perform semantic extraction on the context of the erroneous word and obtain the context feature vector of the erroneous word includes: The statement to be detected is input into a preset semantic extraction model, and the BERT model in the semantic extraction model is used to extract the semantics of the statement to be detected to obtain a first semantic feature vector. The first semantic feature vector is semantically extracted by the BiLSTM layer in the semantic extraction model to obtain the second semantic feature vector. The features corresponding to the position of the erroneous word in the second semantic feature vector are used as the context feature vector of the erroneous word.
6. The Indonesian spelling correction method as described in claim 5, characterized in that, The encoder-decoder framework has twice the number of units in its encoder compared to the BiLSTM layer of the semantic extraction model.
7. An Indonesian spelling correction device, characterized in that, include: The data preprocessing module is used to acquire the Indonesian sentence to be detected, detect the words in the Indonesian sentence according to the preset Indonesian dictionary, and obtain the erroneous words in the Indonesian sentence. The semantic extraction module is used to extract the semantics of the context of the erroneous word using a preset semantic extraction model, and obtain the context feature vector of the erroneous word. The candidate word acquisition module is used to acquire a set of candidate words corresponding to the erroneous word based on the erroneous word and a pre-built binary Indonesian statistical model; wherein, the set of candidate words includes several candidate words; The word generation module is used to calculate the first selection probability of each candidate word in the candidate word set using a pre-built encoder-decoder framework, the erroneous words, and the context feature vector; The probability adjustment module is used to adjust the first selection probability of each candidate word based on the edit distance between each candidate word and the erroneous word, so as to obtain the second selection probability of each candidate word. The word correction module is used to select the candidate word with the highest second selection probability from the candidate word set as the correction word for the erroneous word, based on the second selection probability of each candidate word, so as to correct the erroneous word; The encoder-decoder framework consists of an encoder, a decoder, and an attention layer, and the word generation module includes: The feature vector encoding unit is used to input the erroneous word character by character into the encoder of the encoder-decoder framework for encoding, to obtain the first character of the erroneous word. The hidden feature vectors of each character and the context condition vector output at the last moment; where, Greater than or equal to 0; The feature vector initialization unit is used to add the context condition vector of the erroneous word and the context feature vector to obtain the initial encoding feature vector of the erroneous word; Attention feature processing unit, used to process the attention layer of the encoder-decoder framework, the first digit of the erroneous word... The hidden feature vector of the nth character and the initial encoded feature vector of the erroneous word are used to calculate the nth character of the erroneous word input by the encoder. Attention feature vectors for each character; The feature vector concatenation unit is used to concatenate the first feature vector of the erroneous word. The attention feature vector of the nth character and the nth character of each candidate word The characters are concatenated to obtain the first character of each candidate word. The feature vector of the character to be decoded; The character probability calculation unit is used to calculate the first character probability of each candidate word. The feature vector of the nth character to be decoded is input into the decoder of the encoder-decoder framework for decoding, to obtain the nth character of each candidate word. The probability of a character; The candidate word probability calculation unit is used to calculate the average value of the character probabilities of all characters in each candidate word to obtain the first selection probability of each candidate word.
8. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the Indonesian spelling correction method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the Indonesian spelling correction method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Error-detecting and error-correcting method and system for indonesian word
CN109145287A
Grammar error correction method for adding spelling error correction function
CN111460794A
Method and apparatus for genome spelling correction and acronym standardization
US20210326526A1