Devanagari word segmentation and recognition method, device, electronic device and storage medium

By considering the current character and the types of characters after it in the Tianchengwen word segmentation method, formulating word segmentation rules suitable for Hindi are solved, and the problem of low accuracy of Hindi text recognition in the prior art is achieved, and a more efficient text recognition effect is achieved.

CN114254638BActive Publication Date: 2025-06-06IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111580309.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-06-06
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The prior art word segmentation method is not suitable for Hindi, resulting in low accuracy in Hindi text recognition.

Method used

A method of word segmentation in Tianchengwen is provided. By obtaining the character sequence of the Tianchengwen text to be divided, word segmentation is performed based on the current character and the type of characters after it, including consonants and vowels, consonants and consonants, vowels and vowels additional symbols, etc.

Benefits of technology

The constructed Tiancheng Wenzi word list can greatly reduce the size of the dictionary, improve the accuracy of Hindi text recognition, and effectively deal with similar words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254638B_ABST
    Figure CN114254638B_ABST
Patent Text Reader

Abstract

The present invention provides a Devanagari word segmentation and recognition method, device, electronic device and storage medium, wherein the word segmentation method comprises: obtaining a character sequence of a Devanagari text to be segmented; based on the type of the current character in the character sequence and the type of the character after the current character, segmenting the current character and the characters after the current character, and updating the last character in the subword obtained by the segmentation to the next character in the character sequence as the current character for word segmentation until the word segmentation is completed. The Devanagari word segmentation and recognition method, device, electronic device and storage medium provided by the embodiment of the present invention, on the basis of analyzing and organizing the basic unit structure, proposes a word segmentation rule suitable for the language structure characteristics of the Devanagari language, taking into account both the type of the current character and the type of the character after the current character, thereby determining the language structure of a segment of characters in the character sequence, and performing word segmentation accordingly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text recognition, and in particular to a Devanagari word segmentation and recognition method, device, electronic equipment and storage medium. Background Art

[0002] Multilingual text recognition plays an important supporting role in barrier-free information communication between humans and between humans and machines. Text recognition technology generally needs to obtain a vocabulary through word segmentation, and use the vocabulary to represent sentences.

[0003] In the prior art, word segmentation methods are mainly used for English or Chinese, where text sentences are mostly based on words and the character structure is regular without compound writing or connected writing.

[0004] However, Hindi is a type of abugida written in Devanagari script. The degree of adhesion between its characters is relatively common, and there are different compound writing rules between vowels and consonants, resulting in a relatively variable structure of its glyphs. Existing word segmentation methods are not suitable for Hindi scenarios, resulting in a low accuracy rate in Hindi text recognition. Summary of the invention

[0005] The present invention provides a Devanagari word segmentation and recognition method, device, electronic device and storage medium, which are used to solve the defect that the word segmentation method in the prior art is not suitable for Hindi, resulting in a low Hindi text recognition accuracy.

[0006] The present invention provides a Devanagari word segmentation method, comprising:

[0007] Get the character sequence of the Devanagari text to be segmented;

[0008] Based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented, and the last character in the sub-word obtained by the segmentation is updated to the next character in the character sequence as the current character for segmentation until the segmentation is completed.

[0009] According to a Devanagari word segmentation method provided by the present invention, the word segmentation of the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character includes:

[0010] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a non-blank character, the current character and the first character after it are written as a subword, and the non-blank character is a character other than a consonant and a vowel eliminater.

[0011] According to a Devanagari word segmentation method provided by the present invention, the word segmentation of the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character includes:

[0012] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a vowel elimination character, the current character and the first character after it are co-written according to the vowel elimination character writing structure, and the co-written characters are updated to the current character, and the current character and the characters after it are segmented.

[0013] According to a Devanagari word segmentation method provided by the present invention, the word segmentation of the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character includes:

[0014] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character after the current character is a vowel remover, then determining the structure of the current character;

[0015] If the structure of the current character is a vowel-elimination conformation structure, the current character and the first character and the second character after it are conformed and updated to the current character; otherwise, the current character and the first character after it are conformed to a subword.

[0016] According to a Devanagari word segmentation method provided by the present invention, the word segmentation of the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character includes:

[0017] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character is a vowel, the current character and the first character after it are written as a subword.

[0018] According to a Devanagari word segmentation method provided by the present invention, the word segmentation of the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character includes:

[0019] If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is a vowel datum, then the current character and the first character after it are written as a subword;

[0020] If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is not a vowel daemon, the current character is regarded as a subword.

[0021] The present invention also provides a Devanagari text recognition method, comprising:

[0022] Obtain an image to be recognized;

[0023] Based on the subword vocabulary, Devanagari text recognition is performed on the image to be recognized to obtain a text recognition result; the subword vocabulary is determined based on any one of the Devanagari word segmentation methods described above.

[0024] The present invention also provides a Devanagari word segmentation device, comprising:

[0025] A text acquisition unit, used for acquiring Devanagari text to be segmented;

[0026] The word segmentation unit is used to segment the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character, and update the last character in the sub-word obtained by the segmentation to the next character in the character sequence as the current character for word segmentation until the word segmentation is completed.

[0027] The present invention also provides a Devanagari text recognition device, comprising:

[0028] An image acquisition unit, used for acquiring an image to be recognized;

[0029] The text recognition unit is used to perform Devanagari text recognition on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on any of the Devanagari word segmentation methods described above.

[0030] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any one of the above-mentioned Devanagari word segmentation methods or Devanagari text recognition methods are implemented.

[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of any of the above-mentioned Devanagari word segmentation methods or Devanagari text recognition methods are implemented.

[0032] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of any one of the above-mentioned Devanagari word segmentation methods or Devanagari text recognition methods are implemented.

[0033] The Devanagari word segmentation and recognition method, device, electronic device and storage medium provided by the present invention, on the basis of analyzing and organizing the basic unit structure, proposes a word segmentation rule suitable for the language structure characteristics of the Devanagari language, taking into account both the type of the current character and the type of the character after the current character, thereby determining the language structure of a segment of characters in the character sequence, and performing word segmentation accordingly. Compared with the BPE word segmentation method, the word segmentation method provided by the embodiment of the present invention is more suitable for the Devanagari language, and the Devanagari sub-word word list constructed thereby can greatly reduce the size of the dictionary and better handle similar words. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 It is a schematic diagram of the Devanagari script;

[0036] Figure 2 It is one of the flow charts of the Devanagari word segmentation method provided by the present invention;

[0037] Figure 3 A schematic diagram of the co-writing of consonants and vowels in the Devanagari word segmentation method provided by the present invention;

[0038] Figure 4 A schematic diagram of the co-writing of double consonants in the Devanagari word segmentation method provided by the present invention;

[0039] Figure 5 A schematic diagram of the multi-consonant co-writing method of the Devanagari script word segmentation method provided by the present invention;

[0040] Figure 6 This is the second flow chart of the Devanagari word segmentation method provided by the present invention;

[0041] Figure 7 It is a schematic diagram of the flow of the Devanagari text recognition method provided by the present invention;

[0042] Figure 8 It is a schematic diagram of the structure of the Devanagari word segmentation device provided by the present invention;

[0043] Fig. 9 It is a structural schematic diagram of the Devanagari text recognition device provided by the present invention;

[0044] Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] Most of the text recognition based on deep learning is to extract image feature information through deep learning network, encode it into feature sequence form, and then get the prediction result through decoding. However, this solution based on general text is not very ideal for the recognition of language characters with unique glyph structure, such as Hindi.

[0047] Both the training and prediction of deep learning networks require the use of a vocabulary to represent sentences. The usual method of constructing a vocabulary is to first segment each sentence, then count the frequency and select the highest N phrases to form a vocabulary. Since the training set contains a large number of words, such as the total number of English words is about 1 million, for the sake of computational efficiency, selecting N words with higher frequencies does not cover all the words in the training set well. The granularity of subword division is between words and characters. For example, "logogram" can be divided into two subwords, "logo" and "gram", and the divided "logo" and "gram" can be used to construct other words.

[0048] At present, text recognition technology is mainly based on the BPE (Byte Pair Encoding) subword modeling method. First, the word is split into the smallest unit, and these smallest units are used as the initial vocabulary. Then, the frequency of adjacent unit pairs in the word is counted on the training corpus, and then the unit pair with the highest frequency is selected and merged into a new subword unit. Repeat the above steps to obtain a pre-set subword vocabulary size. Finally, according to the constructed vocabulary, the Encoder-Decoder model is used to complete the text recognition.

[0049] However, in Hindi text scenarios, because Hindi is a vowel-affidavit-based language, written in Devanagari script, the characters have up-down, left-right, and other structures. In addition, there are different reduplication and merging rules between vowels and consonants. In writing, the letters will overlap after being written together, resulting in changes in the morphology, such as Figure 1 The diagram of Devanagari coinscription is shown. Therefore, the subwords of Devanagari are varied in form, and some subwords after coinscription are very small, with only a dot or a stroke, etc., which makes them visually similar, leading to more recognition errors.

[0050] Based on this, an embodiment of the present invention provides a Devanagari word segmentation method, and the subword vocabulary obtained by the method can be applied to the Devanagari text recognition scenario.

[0051] It should be noted that the Devanagari word segmentation method provided in the embodiment of the present invention can be applied to languages ​​written in Devanagari, such as Hindi, Sanskrit or Nepali, etc. The embodiment of the present invention will take Hindi as an example to explain the Devanagari word segmentation method in detail.

[0052] Figure 2 This is one of the flow charts of the Devanagari word segmentation method provided by the present invention. Figure 2 As shown, the method includes:

[0053] Step 210: Obtain a character sequence of the Devanagari text to be segmented.

[0054] Specifically, the Devanagari text here refers to a text written in Devanagari characters, for example, it can be a Hindi text, a Sanskrit text, or a Nepali text.

[0055] Taking Hindi text as an example, the Hindi text to be segmented may be a text library generated based on Hindi corpus, or may be Hindi text directly input by a user, or may be obtained by voice transcription of collected audio, which is not specifically limited in the embodiment of the present invention.

[0056] Since Hindi is written in Devanagari script, the characters have structures such as up and down, left and right, etc. There are different reduplication and merging rules between vowels and consonants. In writing, letters may overlap after being combined, resulting in changes in shape.

[0057] In order to improve the accuracy of Hindi word segmentation, the Hindi text here can be stored in the computer in the form of Unicode (Unicode, Unicode, unicode). Unicode is an industry standard in the field of computer science, including character sets and encoding schemes. It is created to solve the limitations of traditional character encoding schemes. It sets a unified and unique binary code for each character in each language to meet the requirements of cross-language and cross-platform text conversion and processing.

[0058] After obtaining the Hindi text to be segmented, the characters and character sequences in the Hindi text can be determined according to the Unicode corresponding to the Hindi text. The character sequence is a sequence formed by encoding the characters in order, which is convenient for subsequent segmentation processing.

[0059] Step 220, based on the type of the current character in the character sequence and the type of the character after the current character, segment the current character and the characters after the current character, and update the last character in the sub-word obtained by the segmentation to the next character in the character sequence as the current character for segmentation until the segmentation is completed.

[0060] Specifically, before Hindi word segmentation, it is necessary to sort out the basic unit structure of Hindi.

[0061] Hindi has 54 phonemes, including 11 vowels and 43 consonants. Hindi also has 4 vowel additional characters: 2 additional consonant characters: These additional characters cannot be directly combined with vowels or consonants to form syllables, and are always written after vowels or consonants.

[0062] In Hindi, if you want to express a consonant without any vowels, the earliest method used was to shorten or simplify the consonant syllable. Later, a special symbol appeared - the vowel elimination symbol.

[0063]

[0064] The writing system of abugida is between syllabic writing and phonetic writing. Since the Hindi language rhythm emphasizes long and short syllables, and syllables are based on vowels, no matter how many consonants there are, they are all regarded as one syllable, so multiple consonants must be written as a combined visual unit.

[0065] First of all, its basic unit is the syllable. A basic consonant character has a short vowel a without adding any vowel symbols. Its corresponding Latin transliteration must be represented by two letters, such as Hindi If the consonant is combined with other vowels, a vowel mark is required, such as

[0066] Based on the above analysis of the basic unit structure of Hindi, Hindi text can be segmented according to the duplication and merging rules between Hindi vowels and consonants.

[0067] Here, the current character may be a basic character, for example, the current character is a basic consonant character or a basic vowel character.

[0068] In some embodiments, the current character may also be in the form of a string, including multiple basic characters. For example, in the case of three or more consonants being written together, the current multiple consonant characters written together may be written together as a whole as the current character.

[0069] The type of the current character may be a consonant or a vowel. In the embodiment of the present invention, the type of the current character may be a consonant or a vowel as the starting character for word segmentation.

[0070] It should be noted that if the type of the current character is an additional character such as a vowel remover, the current character is directly segmented and the current character is treated as a subword alone.

[0071] In some embodiments, if the current character includes multiple basic characters, due to the language characteristics of Hindi, at this time, the type of the current character is usually a consonant with vowels eliminated and in accordance with the writing structure.

[0072] According to the type of the current character and the type of the character after the current character, the current character and the characters after the current character are segmented.

[0073] For example, if the type of the current character is a vowel, since syllables in Hindi are based on vowels, the current vowel character can be directly segmented. Specifically, the vowel character can be taken out and added as a subword to the subword word list.

[0074] For another example, if the type of the current character is a consonant, since consonants in Hindi can be written together with other characters according to certain rules, such as the rules for writing consonants together with the vowels that follow them, or the rules for writing consonants together with the consonants that follow them, the current character and the characters that follow it can be segmented according to the type of the character that follows the current character.

[0075] In a specific example, if the type of the current character is a consonant and the first character after it is a vowel, the current character and the first character after it can be written as a subword, and the subword is added to the subword vocabulary.

[0076] Whenever a co-writing is completed and a sub-word is obtained, the next character of the last character in the sub-word in the character sequence can be updated to the current character for word segmentation until the word segmentation is completed.

[0077] The Devanagari word segmentation method provided in the embodiment of the present invention, based on the analysis and arrangement of the basic unit structure, proposes a word segmentation rule suitable for the language structure characteristics of the Devanagari language, taking into account both the type of the current character and the type of the character after the current character, thereby determining the language structure of a segment of characters in the character sequence, and performing word segmentation accordingly. Compared with the BPE word segmentation method, the word segmentation method provided in the embodiment of the present invention is more suitable for the Devanagari language, and the Devanagari sub-word word list constructed thereby can greatly reduce the size of the dictionary and better handle similar words.

[0078] In addition, the subword vocabulary obtained according to the Devanagari word segmentation method provided by the embodiment of the present invention can be applied to Devanagari text recognition, thereby effectively improving the accuracy of Devanagari text recognition.

[0079] Based on the above embodiment, in step 220, based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented, which specifically includes:

[0080] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a non-blank character, then the current character and the first character after it are written as a subword, and the non-blank character is a character other than a consonant and a vowel remover.

[0081] Specifically, considering that in Devanagari, the co-writing rules of consonants and the characters that follow them are relatively complicated, there are three rules that can be summarized, namely, the co-writing rules of consonants and vowels, the co-writing rules of double consonants, and the co-writing rules of multiple consonants.

[0082] The embodiment of the present invention and the following embodiments all specifically illustrate the method of segmenting Devanagari words according to the above three consonant co-writing rules, assuming that the type of the current character in the character sequence is a consonant.

[0083] In this embodiment, the current character is a basic consonant character. If the type of the first character after the current character is a non-blank character, the current character and the first character after it are written as a subword, wherein the non-blank character is a character other than a consonant and a vowel eliminator, for example, the non-blank character can be a vowel character or a consonant suffix, etc.

[0084] Take the vowel character as an example. Figure 3 A schematic diagram of the co-writing of consonants and vowels in the Devanagari word segmentation method provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, a basic consonant character and a subsequent vowel character can be written as a subword, and the subword is added to the subword vocabulary.

[0085] Of course, a basic consonant character and a subsequent consonant additional character may also be written as a subword, and the subword may be added to the subword vocabulary.

[0086] The Devanagari word segmentation method provided in the embodiment of the present invention realizes word segmentation processing of Devanagari text by combining consonants and non-space characters as a subword, wherein the non-space characters are characters other than consonants and vowel elimination symbols.

[0087] Based on the above embodiment, in step 220, based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented, which specifically includes:

[0088] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a vowel elimination character, the current character and the first character after it are co-written according to the vowel elimination character writing structure, and the co-written character is updated as the current character, and the current character and the characters after it are segmented.

[0089] Specifically, the vowel canceller here is a special symbol in Hindi, which is used to represent a consonant without any vowels and is usually used between two basic consonants.

[0090] Considering that there are three or more consonants in Hindi, there is a vowel eliminator between the two basic consonants.

[0091] Here, the current character and the first character after it are co-written according to the vowel elimination conforming writing structure, that is, the current consonant character and the vowel elimination character after it are co-written, and the co-written character can be used as a consonant conforming to the vowel elimination writing structure, and the co-written character is updated to the current character, and then the current character and the characters after it are segmented according to the type of the character after the current character.

[0092] It should be noted that the current character here can be a basic consonant character, or a consonant that conforms to the writing structure after vowel elimination.

[0093] Based on the above embodiment, in step 220, based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented, which specifically includes:

[0094] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character is a vowel remover, then determine the structure of the current character;

[0095] If the structure of the current character is a vowel-elimination conformation structure, the current character and the first character and the second character after it are conformed and updated to the current character; otherwise, the current character and the first character after it are conformed as a subword.

[0096] Specifically, if the writing form of the current character string is "current character (consonant) + the first character after (consonant) + the second character after (vowel eliminater)", the structure of the current character is first determined.

[0097] The structure of the current character may be a basic consonant character, or a consonant whose vowels are eliminated and conforms to the writing structure.

[0098] If the structure of the current character is a vowel-elimination conformation structure, the current character and the first character and the second character after it are conformed and updated to the current character. That is, the current character and the first character and the second character after it are co-constructed, and then the type of the next character is determined, and the current character and the characters after it are segmented according to the type of the next character.

[0099] If the structure of the current character is not a vowel-elimination structure, that is, the structure of the current character is a basic consonant character, the current character and the first character following it are written as a subword.

[0100] At this time, the second character after the current character is a vowel eliminator, which can also be used as a subword alone.

[0101] Figure 4 A schematic diagram of the double consonant co-writing method of the Devanagari script word segmentation method provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the double consonants that appear consecutively can be written as a subword. It can be understood that the double consonants that appear consecutively here are all basic consonant characters, and there is no vowel elimination to meet the writing structure of the characters.

[0102] From the description of the above embodiment, it can be seen that a basic consonant and a basic consonant following it can be written together as a subword. The two basic consonant characters can be any basic consonant characters, and can be the same characters or different characters.

[0103] The method provided in the embodiment of the present invention realizes word segmentation processing of Devanagari text by writing two consecutive basic consonant characters together as a subword.

[0104] Based on the above embodiment, in step 220, based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented, which specifically includes:

[0105] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character is a vowel, then the current character and the first character after it are written as a subword.

[0106] Specifically, when the type of the current character is a consonant, there are also two cases. One is that the current character is a basic consonant character, and the other is that the current character is a vowel and the consonant that meets the writing structure is eliminated. But no matter which case, if the type of the first character after the current character is a consonant, and the type of the second character is a vowel, the current character and the first character after it will be written as a subword. In other words, when consonants are written together (double consonants or multiple consonants), if a vowel is encountered, stop judging the type of the next character, and write the consonant before the vowel together as a subword.

[0107] Figure 5 A schematic diagram of the multi-consonant co-writing method of the Devanagari script word segmentation method provided by an embodiment of the present invention is shown in FIG. Figure 5 As shown, the consonants and vowel elimination symbols can be co-written until a vowel appears and the co-writing stops, and the consonants before the vowel are co-written as a subword.

[0108] The method provided by the embodiment of the present invention, when consonants are written together, especially when three or more consonants are written together (in this case, there is a vowel elimination symbol between two basic consonants), the consonants before the vowels are written together as a subword, thereby realizing word segmentation processing of the Devanagari text.

[0109] Based on the above embodiment, in step 220, based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented, which specifically includes:

[0110] If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is a vowel datum, then the current character and the first character after it are written as a subword;

[0111] If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is not a vowel diacritical mark, the current character is regarded as a subword.

[0112] Specifically, it can be seen from the description of the above embodiment that if the type of the current character in the Hindi text is a consonant, there are 3 corresponding co-writing rules. If the type of the current character in the Hindi text is a vowel, the embodiment of the present invention will describe the word segmentation method in this case.

[0113] If the type of the current character in the Hindi text is a vowel, and the type of the first character after the current character is a vowel daemon, the current vowel character and the vowel daemon after it are written as a subword and added to the subword word list.

[0114] If the type of the current character in the Hindi text is a vowel, and the type of the first character after the current character is not a vowel embellishment, for example, it can be any character except a vowel embellishment, then the current character vowel character is taken as a subword and added to the subword word list.

[0115] The vowels here refer to the basic vowel characters.

[0116] The method provided by the embodiment of the present invention realizes word segmentation processing of Devanagari text by writing vowel characters or vowel characters and vowel datums together as a subword.

[0117] Based on any of the above embodiments, Figure 6 This is the second flow chart of the Devanagari word segmentation method provided by the embodiment of the present invention. Figure 6 As shown, in this embodiment, the implementation of each step has been described in the previous embodiment. For relevant parts, please refer to the previous description.

[0118] The embodiment of the present invention further explains the overall process of the Devanagari word segmentation method.

[0119] Based on any of the above embodiments, Figure 7 is a flow chart of a Devanagari text recognition method provided by an embodiment of the present invention, such as Figure 7 As shown, the method includes:

[0120] Step 710, obtaining an image to be recognized;

[0121] Step 720 , based on the subword vocabulary, perform Devanagari text recognition on the image to be recognized to obtain a text recognition result; the subword vocabulary is determined based on any of the above-mentioned Devanagari word segmentation methods.

[0122] Specifically, the subword vocabulary obtained by the Devanagari word segmentation method introduced in the above embodiment can be applied to Devanagari text recognition.

[0123] The image to be recognized here refers to an image that needs to perform Devanagari text recognition, and the image contains Devanagari text.

[0124] The image to be identified may be a color image or a grayscale image, and the embodiment of the present invention does not limit the specific form of the image to be identified.

[0125] The image to be identified may be an image taken by the user himself, an image downloaded from the Internet, an image received from a device, or an image in a video. The embodiment of the present invention does not limit the image source of the image to be identified.

[0126] The text recognition may be performed on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on any of the above-mentioned Devanagari word segmentation methods.

[0127] For example, Hindi sentences can be effectively encoded and structured using a subword vocabulary, and then a neural network model can be used to decode the text sentence information it contains to obtain text recognition results.

[0128] The method provided by the embodiment of the present invention first determines a subword vocabulary through a Devanagari word segmentation method, and based on the subword vocabulary, performs Devanagari text recognition on the image to be recognized, thereby improving the recognition accuracy of the Devanagari text.

[0129] The Devanagari word segmentation device provided by the present invention is described below. The Devanagari word segmentation device described below and the Devanagari word segmentation method described above can be referenced to each other.

[0130] Based on any of the above embodiments, Figure 8 is a schematic diagram of the structure of the Devanagari word segmentation device provided by an embodiment of the present invention. Figure 8 As shown, the device comprises:

[0131] A text acquisition unit 810 is used to acquire a character sequence of a Devanagari text to be segmented;

[0132] The word segmentation unit 820 is used to segment the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character, and update the last character in the sub-word obtained by the segmentation to the next character in the character sequence as the current character for word segmentation until the word segmentation is completed.

[0133] The Devanagari word segmentation device provided in the embodiment of the present invention, based on the analysis and arrangement of the basic unit structure of Devanagari, proposes a word segmentation rule suitable for the language structure characteristics of Devanagari, taking into account both the type of the current character and the type of the character after the current character, thereby determining the language structure of a segment of characters in the character sequence, and performing word segmentation accordingly. Compared with the BPE word segmentation method, the word segmentation device provided in the embodiment of the present invention is more suitable for Devanagari, and the Devanagari sub-word vocabulary constructed thereby can greatly reduce the size of the dictionary and better handle similar words.

[0134] Based on any of the above embodiments, the word segmentation unit 820 is further used for:

[0135] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a non-blank character, the current character and the first character after it are written as a subword, and the non-blank character is a character other than a consonant and a vowel eliminater.

[0136] Based on any of the above embodiments, the word segmentation unit 820 is further used for:

[0137] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a vowel elimination character, the current character and the first character after it are co-written according to the vowel elimination character writing structure, and the co-written character is updated to the current character, and the current character and the characters after it are segmented.

[0138] Based on any of the above embodiments, the word segmentation unit 820 is further used for:

[0139] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character after the current character is a vowel remover, then determining the structure of the current character;

[0140] If the structure of the current character is a vowel-elimination conformation structure, the current character and the first character and the second character after it are conformed and updated to the current character, and the current character and the characters after it are segmented; otherwise, the current character and the first character after it are conformed as a subword.

[0141] Based on any of the above embodiments, the word segmentation unit 820 is further used for:

[0142] If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character is a vowel, the current character and the first character after it are written as a subword.

[0143] Based on any of the above embodiments, the word segmentation unit 820 is further used for:

[0144] If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is a vowel datum, then the current character and the first character after it are written as a subword;

[0145] If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is not a vowel daemon, the current character is regarded as a subword.

[0146] The Devanagari text recognition device provided by the present invention is described below. The Devanagari text recognition device described below and the Devanagari text recognition method described above can be referred to each other.

[0147] Based on any of the above embodiments, Fig. 9 is a schematic diagram of the structure of a Devanagari text recognition device provided by an embodiment of the present invention. Fig. 9As shown, the device comprises:

[0148] An image acquisition unit 910 is used to acquire an image to be recognized;

[0149] The text recognition unit 920 is used to perform Devanagari text recognition on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on any of the Devanagari word segmentation methods described above.

[0150] The text recognition device provided by the embodiment of the present invention first determines a subword vocabulary by using a Devanagari word segmentation method, and then performs Devanagari text recognition on the image to be recognized based on the subword vocabulary, thereby improving the recognition accuracy of the Devanagari text.

[0151] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10 As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030 and a communication bus 1040, wherein the processor 1010, the communication interface 1020 and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 may call the logic instructions in the memory 1030 to execute the Devanagari word segmentation method or the Devanagari text recognition method.

[0152] The Devanagari word segmentation method includes: obtaining a character sequence of a Devanagari text to be segmented; segmenting the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character, and updating the last character in the subword obtained by the segmentation to the next character in the character sequence as the current character for segmentation until the segmentation is completed.

[0153] The Devanagari text recognition method comprises: obtaining an image to be recognized; performing Devanagari text recognition on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on any one of the Devanagari word segmentation methods described above.

[0154] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0155] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the Devanagari word segmentation method or Devanagari text recognition method provided by the above methods.

[0156] The Devanagari word segmentation method includes: obtaining a character sequence of a Devanagari text to be segmented; segmenting the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character, and updating the last character in the subword obtained by the segmentation to the next character in the character sequence as the current character for segmentation until the segmentation is completed.

[0157] The Devanagari text recognition method includes: obtaining an image to be recognized; performing text recognition on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on any one of the Devanagari word segmentation methods described above.

[0158] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the Devanagari word segmentation method or Devanagari text recognition method provided by the above methods.

[0159] The Devanagari word segmentation method includes: obtaining a character sequence of a Devanagari text to be segmented; segmenting the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character, and updating the last character in the subword obtained by the segmentation to the next character in the character sequence as the current character for segmentation until the segmentation is completed.

[0160] The Devanagari text recognition method includes: obtaining an image to be recognized; performing text recognition on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on any one of the Devanagari word segmentation methods described above.

[0161] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0162] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A Devanagari word segmentation method, It is characterized in that include: Get the character sequence of the Devanagari text to be segmented; Based on the type of the current character in the character sequence and the type of the character after the current character, the current character and the characters after the current character are segmented to obtain a subword, and the last character in the subword is updated to the next character in the character sequence as the current character for segmentation until the segmentation is completed, so as to construct a Devanagari subword vocabulary; Wherein, the type of the current character includes a consonant or a vowel; The step of segmenting the current character and the characters following it includes: writing the current character and the first character following it as a subword, or writing the current character as a subword.

2. The Devanagari word segmentation method according to claim 1, It is characterized in that The word segmentation of the current character and the characters following the current character based on the type of the current character in the character sequence and the type of the characters following the current character comprises: If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a non-blank character, the current character and the first character after it are written as a subword, and the non-blank character is a character other than a consonant and a vowel eliminater.

3. The Devanagari word segmentation method according to claim 1, It is characterized in that The word segmentation of the current character and the characters following the current character based on the type of the current character in the character sequence and the type of the characters following the current character comprises: If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a vowel elimination character, the current character and the first character after it are co-written according to the vowel elimination character writing structure, and the co-written characters are updated to the current character, and the current character and the characters after it are segmented.

4. The Devanagari word segmentation method according to claim 1, It is characterized in that The word segmentation of the current character and the characters following the current character based on the type of the current character in the character sequence and the type of the characters following the current character comprises: If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character after the current character is a vowel remover, then determining the structure of the current character; If the structure of the current character is a vowel-elimination conformation structure, the current character and the first character and the second character after it are conformed and updated to the current character, and the current character and the characters after it are segmented; otherwise, the current character and the first character after it are conformed as a subword.

5. The Devanagari word segmentation method according to claim 1, It is characterized in that The word segmentation of the current character and the characters following the current character based on the type of the current character in the character sequence and the type of the characters following the current character comprises: If the type of the current character in the character sequence is a consonant, and the type of the first character after the current character is a consonant, and the type of the second character is a vowel, the current character and the first character after it are written as a subword.

6. The Devanagari word segmentation method according to claim 1, It is characterized in that The word segmentation of the current character and the characters following the current character based on the type of the current character in the character sequence and the type of the characters following the current character comprises: If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is a vowel datum, then the current character and the first character after it are written as a subword; If the type of the current character in the character sequence is a vowel, and the type of the first character after the current character is not a vowel daemon, the current character is regarded as a subword.

7. A Devanagari text recognition method, It is characterized in that include: Obtain an image to be recognized; Based on the subword vocabulary, performing Devanagari text recognition on the image to be recognized to obtain a text recognition result; The subword vocabulary is determined based on the Devanagari word segmentation method described in any one of claims 1 to 6.

8. A Devanagari word segmentation device, It is characterized in that include: A text acquisition unit, used for acquiring a character sequence of a Devanagari text to be segmented; A word segmentation unit is used to segment the current character and the characters after the current character based on the type of the current character in the character sequence and the type of the character after the current character to obtain a subword, and update the last character of the subword to the next character in the character sequence as the current character to perform word segmentation until the word segmentation is completed, so as to construct a Devanagari subword vocabulary; Wherein, the type of the current character includes a consonant or a vowel; The step of segmenting the current character and the characters following it includes: writing the current character and the first character following it as a subword, or writing the current character as a subword.

9. A Devanagari text recognition device, It is characterized in that include: An image acquisition unit, used for acquiring an image to be recognized; A text recognition unit is used to perform Devanagari text recognition on the image to be recognized based on a subword vocabulary to obtain a text recognition result; the subword vocabulary is determined based on the Devanagari word segmentation method according to any one of claims 1 to 6.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the Devanagari word segmentation method according to any one of claims 1 to 6 or the Devanagari text recognition method according to claim 7 are implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the Devanagari word segmentation method according to any one of claims 1 to 6 or the Devanagari text recognition method according to claim 7 are implemented.

Citation Information

Patent Citations

  • A Lao character segmentation method

    CN109255120A

  • Multi-task Thai word segmentation method based on syllable segmentation and word segmentation combined learning

    CN112883726A