Dongba scripture page identification method, device and equipment and storage medium
By obtaining the content projection value of the Dongba scripture image page and combining morphological opening operations and Hoff transformation, identifying the homepage, introduction page and text translation page of the Dongba scripture, the problem of low digital recognition efficiency of Dongba scripture is solved and automatic reading with high accuracy is achieved.
Patent Information
- Application Number
- CN202311805817.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-08-05
AI Technical Summary
The digital scanning and document image recognition of the Middle East Pakistan Meridian Volume in the prior art are large in work, and the recognition effect is not good.
By obtaining the content projection value of the current image page of the Dongba scriptures, we can determine whether it meets the preset homepage and introduction page conditions, and use the preset height threshold and text line spacing to identify the Dongba scriptures, literal texts of the scriptures and voluntary texts of the scriptures, and combine morphological opening operations and Hoff transformation to process the image page.
It improves the accuracy and reliability of Dongba scripture page recognition, and can automatically identify the home page, introduction page and text translation page, reduce image interference and improve recognition rate.
Smart Images

Figure CN120431584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document recognition, and in particular to a method, device, equipment and storage medium for identifying pages of a Dongba scripture. Background Art
[0002] The Complete Collection of Naxi Dongba Ancient Texts and Annotations utilizes a four-level translation and annotation system: original Dongba text in pictographic script, Naxi phonetic notation with international phonetic notation, Chinese literal translation with annotations, and Chinese paraphrase translation. This provides a foundation for interpreting the numerous Dongba scriptures lost to the public and overseas. However, since the Complete Collection of Naxi Dongba Ancient Texts and Annotations comprises 100 volumes, each containing approximately 10 volumes, totaling 1,000 volumes, and the volumes are categorized into only five broad categories, the workload of digitizing and scanning the corresponding document images is substantial.
[0003] Therefore, how to improve the recognition effect of Dongba scriptures is a problem to be solved in this field. Summary of the Invention
[0004] In view of this, the present invention aims to provide a Dongba scripture page recognition method, apparatus, device, and storage medium. This method can identify the first page of a Dongba scripture based on pixel value conditions, facilitate automatic reading of the translation and annotation pages, and improve the reliability of the recognition results. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a method for identifying pages of Dongba scriptures, comprising:
[0006] Get the content projection value corresponding to the current image page of the Dongba scripture;
[0007] Determining whether the content projection value meets the preset Dongba scripture homepage condition;
[0008] If the content projection value does not meet the preset Dongba Sutra homepage condition, determining whether the content projection value meets the preset Dongba Sutra introduction page condition;
[0009] If the content projection value does not meet the preset Dongba Sutra introduction page condition, then determining that the current image page is the main text translation and annotation page of the Dongba Sutra;
[0010] According to the preset height threshold, the first preset text line spacing and the second preset text line spacing, the Dongba scriptures, the literal translation text and the free translation text in the main text annotation page are respectively identified to obtain the corresponding annotation page recognition results.
[0011] Optionally, before obtaining the content projection value corresponding to the current image page of the Dongba scripture, the method further includes:
[0012] Performing morphological opening operations on the document images of the Dongba scripture in sequence to obtain pages after the operation;
[0013] Determining the boundary of the calculated page according to a vertical projection segmentation algorithm, and adjusting the angle of the calculated page using a Hough transform to obtain a complete image page;
[0014] The page structure features of the complete image page are eliminated to obtain a target image page corresponding to the Dongba scripture.
[0015] Optionally, obtaining the content projection value corresponding to the current image page of the Dongba scripture includes:
[0016] determining a current image page from the target image pages;
[0017] A horizontal projection algorithm and a vertical projection algorithm are respectively used to perform projection segmentation processing on the current image page to obtain a first content projection value and a second content projection value corresponding to the current image page.
[0018] Optionally, determining whether the content projection value meets a preset Dongba scripture homepage condition includes:
[0019] Determine whether there is content information in which the blank spacing between the upper and lower parts of the first content projection value is not less than the first spacing in the preset Dongba scripture homepage condition;
[0020] If the first content projection value contains content information in which the blank spacing between the upper and lower parts is not less than the first spacing, the current image page is determined to be the first page of the Dongba scripture;
[0021] If the first content projection value does not contain content information with upper and lower blank spaces not less than the first space, then determining whether the second content projection value contains content information with a height not less than a first height of a preset Dongba scripture homepage condition;
[0022] If the second content projection value includes content information whose height is not less than a first height of a preset Dongba Sutra homepage condition, the current image page is determined to be the homepage of the Dongba Sutra;
[0023] If the second content projection value does not contain content information with a height not less than the first height of the preset Dongba Sutra homepage condition, it is determined that the content projection value does not meet the preset Dongba Sutra homepage condition.
[0024] Optionally, the determining whether the content projection value meets a preset Dongba scripture introduction page condition includes:
[0025] Determine a target text line in which the text line spacing in the first content projection value is not greater than the second spacing in the preset Dongba scripture introduction page condition;
[0026] Determining whether the ratio of the target text line to all text lines in the first content projection value is greater than a first threshold in the preset Dongba scripture introduction page;
[0027] If yes, the current image page is determined to be the introduction page of the Dongba scripture;
[0028] If not, it is determined that the current image page does not meet the preset Dongba scripture introduction page condition.
[0029] Optionally, the step of identifying the Dongba scriptures, the literal translation text, and the paraphrase text in the main text annotation page according to a preset height threshold, a first preset text line spacing, and a second preset text line spacing includes:
[0030] Determining, based on the first content projection value, that content information in the main text translation and annotation page whose pixel height is not less than a preset height threshold is Dongba scripture;
[0031] Determining, based on the first content projection value, content information in which the spacing between adjacent text lines in the main text annotation page is not greater than a first preset text line spacing as a literal translation text;
[0032] According to the first content projection value, content information in which the spacing between adjacent text lines in the main text annotation page is greater than the first preset text line spacing and not greater than the second preset text line spacing is determined as the scripture paraphrase text.
[0033] Optionally, obtaining the corresponding translation annotation page recognition result includes:
[0034] The content of the paraphrase text of the scripture is recognized by using a preset character recognition tool to obtain a corresponding content recognition result; the preset character recognition tool is a tool obtained by merging a preset number of character recognition tools using a preset merging formula.
[0035] In a second aspect, the present application provides a Dongba scripture page recognition device, comprising:
[0036] The projection value acquisition module is used to obtain the content projection value corresponding to the current image page of the Dongba scripture;
[0037] The first judgment module is used to judge whether the content projection value meets the preset Dongba scripture homepage condition;
[0038] a second judgment module, configured to judge whether the content projection value meets the preset Dongba Sutra introduction page condition when the content projection value does not meet the preset Dongba Sutra homepage condition;
[0039] a third judgment module, configured to determine that the current image page is a text translation and annotation page of the Dongba Sutra when the content projection value does not meet the preset Dongba Sutra introduction page condition;
[0040] The page recognition module is used to recognize the Dongba scriptures, the literal translation text and the free translation text in the main text annotation page according to a preset height threshold, a first preset text line spacing and a second preset text line spacing, and obtain corresponding annotation page recognition results.
[0041] In a third aspect, the present application provides an electronic device, comprising:
[0042] Memory, used to store computer programs;
[0043] A processor is used to execute the computer program to implement the Dongba scripture page recognition method as described above.
[0044] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the Dongba scripture page recognition method as described above.
[0045] It can be seen that the present application first obtains the content projection value corresponding to the current image page of the Dongba sutra; determines whether the content projection value meets the preset Dongba sutra homepage condition; if the content projection value does not meet the preset Dongba sutra homepage condition, then determines whether the content projection value meets the preset Dongba sutra introduction page condition; if the content projection value does not meet the preset Dongba sutra introduction page condition, then determines that the current image page is the main text translation and annotation page of the Dongba sutra; then, according to the preset height threshold, the first preset text line spacing and the second preset text line spacing, the Dongba sutra, the literal translation text and the paraphrase text in the main text translation and annotation page are identified to obtain the corresponding translation and annotation page recognition result. In this way, the present application can judge the type of the image page of the Dongba sutra to obtain the main text translation and annotation page of the Dongba sutra, and then perform content recognition on the relevant main text translation and annotation content, so that the homepage, introduction page and main text translation and annotation page of the Dongba sutra can be accurately separated, which is convenient for automatic reading of the Dongba sutra and improves the accuracy and reliability of the recognition result. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0047] Figure 1 This is a flow chart of a Dongba scripture page recognition method disclosed in this application;
[0048] Figure 2 This is a schematic diagram of the layout of the front page of a specific Dongba scripture disclosed in this application;
[0049] Figure 3 This is another specific schematic diagram of the layout of the front page of the Dongba Sutra disclosed in this application;
[0050] Figure 4 This is a schematic diagram of a specific Dongba scripture page segmentation result disclosed in this application;
[0051] Figure 5 This is a schematic structural diagram of a Dongba scripture page recognition device disclosed in this application;
[0052] Figure 6 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] See also Figure 1 As shown, the embodiment of the present invention discloses a method for identifying pages of Dongba scriptures, comprising:
[0055] Step S11: Obtain the content projection value corresponding to the current image page of the Dongba scripture.
[0056] In this application, a current image page of a Dongba scripture is first obtained. It should be noted that the Dongba scripture includes multiple page structures such as the homepage and the main text. When identifying the translation and annotation pages of the Dongba scripture, in order to improve the accuracy of content recognition, it is necessary to first obtain the content projection value of the image page of the Dongba scripture. In a specific embodiment, before obtaining the content projection value corresponding to the current image page of the Dongba scripture, the following steps may also be included: performing a morphological opening operation on the document image of the Dongba scripture in sequence to obtain a post-operation page; determining the boundary of the post-operation page according to a vertical projection segmentation algorithm, and adjusting the angle of the post-operation page using a Hough transform to obtain a complete image page; and removing the page structure features of the complete image page to obtain a target image page corresponding to the Dongba scripture. Specifically, for multiple Dongba scripture document images, morphological calculations are first performed sequentially on each document image to obtain the final page. It is understood that the morphological opening operation can reduce shadow interference in the image. Subsequently, the final image is projected and segmented to determine the content boundaries, and a Hough transform is performed to adjust the angle of the content within the page to obtain a complete image page. This reduces image interference and improves the accuracy of subsequent page recognition. In a specific embodiment, for the initial document image of the Dongba scripture, a morphological opening operation is first performed to assess the background brightness of the document image and extract the document background. The background is then removed from the document image and the image contrast is adjusted, ensuring image validity while also removing any interference from document shadows on document segmentation. A vertical projection segmentation algorithm is then used to determine the boundaries of individual document pages within the document image, completing the extraction of the individual document pages. The features of the document header segmentation line are then combined with a Hough transform to detect the header segmentation line. Simultaneously, the tilt angle of the line is calculated to correct for the tilt of the individual document, effectively avoiding interference from the complex layout of the entire page on line detection and improving accuracy. Furthermore, the binary image and grayscale image of the two pages can be stored independently. The binary image is conducive to the subsequent analysis of the layout structure characteristics of the document, while the grayscale image can retain the characteristics of the Dongba scripture to the greatest extent.
[0057] In another specific embodiment, obtaining a content projection value corresponding to a current image page of the Dongba scripture scroll includes: determining a current image page from the target image pages; and performing projection segmentation processing on the current image page using a horizontal projection algorithm and a vertical projection algorithm, respectively, to obtain a first content projection value and a second content projection value corresponding to the current image page. Specifically, when performing page identification, a page is first selected from the target image pages as the current image page. The current image page is then projected, performing horizontal and vertical projection processing on the current image page, respectively, to obtain a first content projection value for the horizontal projection and a second content projection value for the vertical projection.
[0058] Step S12: determining whether the content projection value meets the preset Dongba scripture homepage condition.
[0059] In the present application, whether the current image page is the first page of the Dongba Sutra scroll can be determined based on the content projection value corresponding to the current image page and the preset Dongba Sutra scroll first page condition. In a specific embodiment, determining whether the content projection value meets the preset Dongba Sutra scroll first page condition can include: determining whether there is content information in the first content projection value whose upper and lower blank spaces are not less than a first spacing in the preset Dongba Sutra scroll first page condition; if there is content information in the first content projection value whose upper and lower blank spaces are not less than the first spacing, then determining the current image page is the first page of the Dongba Sutra scroll; if there is no content information in the first content projection value whose upper and lower blank spaces are not less than the first spacing, then determining whether there is content information in the second content projection value whose height is not less than a first height in the preset Dongba Sutra scroll first page condition; if there is content information in the second content projection value whose height is not less than the first height in the preset Dongba Sutra scroll first page condition, then determining the current image page is the first page of the Dongba Sutra scroll; if there is no content information in the second content projection value whose height is not less than the first height in the preset Dongba Sutra scroll first page condition, then determining that the content projection value does not meet the preset Dongba Sutra scroll first page condition. Specifically, for the horizontally projected first content projection value, it can be determined whether there is content information (text block) with a blank spacing between its upper and lower parts not less than the first spacing specified in the preset Dongba Sutra first page condition, that is, whether the blank spacing between the upper and lower parts of the content information is not less than the first spacing. If the first content projection value contains content information that meets the first spacing, the current image page can be determined as the first page of the Dongba Sutra. Correspondingly, if the first content projection value does not contain content information that meets the first spacing, then for the vertically projected second content projection value, it can be determined whether there is content information with a height not less than the first height specified in the preset Dongba Sutra first page condition, that is, whether the height corresponding to the content information is not less than the first height. If the second content projection value contains content information that meets the first height, the current image page can be determined as the first page of the Dongba Sutra. Furthermore, if the second content projection value corresponding to the current image page also does not contain content information that meets the first height, then it is determined that the current page image is not the first page of the Dongba Sutra, and the first page determination process is then repeated for the image page next to the current image page. It should be pointed out that there are two types of layouts for the front page of Dongba scriptures: one is that the cover of the original text is placed horizontally, consistent with the layout characteristics of the main text page; the other is that the cover of the original text is placed vertically, and other components are placed on its left, such as Figure 2As shown. The first type of Dongba Sutra homepage has a high similarity with the layout characteristics of the main text page, but the layout characteristics of the "Compiler Information" section can be used to distinguish the Dongba Sutra homepage from the main text page. According to statistics, the blank spacing between the upper and lower parts of each text line in the "Compiler Information" section is greater than 480 pixels. The second type of Dongba Sutra homepage is because the original Dongba Sutra cover is placed vertically, as shown in the figure. Figure 3 As shown, a text line with a large line height (over 1200 pixels) is formed, which is significantly different from other document pages. Therefore, when the text block height H(i) obtained by horizontal segmentation is greater than 1200 pixels, the page can be determined to be the first page of the Dongba Sutra.
[0060] That is, for the k (k < m) text lines that make up the document image page, if we assume that the height of the i-th text line is H(i), its upper spacing is uw(i), and its lower spacing is dw(i), then the judgment formula for the first page of the scripture volume can be:
[0061]
[0062] Step S13: If the content projection value does not meet the preset Dongba Sutra homepage condition, determine whether the content projection value meets the preset Dongba Sutra introduction page condition.
[0063] Furthermore, if it is determined that the current image page does not meet the requirements for a homepage, the type of the current image page can be determined based on preset Dongba scripture introduction page requirements. In a specific embodiment, the introduction page following the homepage can also be identified. In this specific embodiment, when identifying the introduction page, a horizontal projection segmentation algorithm can be combined to extract the text lines in the introduction page. Since the introduction page only contains Chinese and English text descriptions, the text line type can be directly classified as "scripture paraphrase", thereby preventing the independent characteristics of the introduction page from affecting the identification of the main text translation and annotation page of the Dongba scripture.
[0064] In a specific embodiment, determining whether the content projection value meets the preset Dongba Sutra introduction page conditions includes: determining a target text line whose text line spacing in the first content projection value is not greater than the second spacing in the preset Dongba Sutra introduction page conditions; determining whether the ratio of the target text line to all text lines in the first content projection value is greater than a first threshold in the preset Dongba Sutra introduction page conditions; if so, determining that the current image page is the introduction page of the Dongba Sutra; if not, determining that the current image page does not meet the preset Dongba Sutra introduction page conditions. Specifically, first determining whether the spacing between text lines in the first content projection value is not greater than the second spacing in the preset Dongba Sutra introduction page conditions; if such a spacing exists, then determining whether the ratio of the corresponding target text line to all text lines in the first content projection value of the entire current image page is greater than a first threshold in the preset Dongba Sutra introduction page conditions; further, if the ratio of the target text line is greater than the first threshold, the current image page can be determined as the introduction page of the Dongba Sutra. Accordingly, if the ratio of the target text lines to the total text lines is not greater than the first threshold, the current image page does not qualify as an introduction page for a Dongba scripture. It is understood that an introduction page consists of a title and an introduction, the spacing between text lines is relatively uniform, and the number of text lines in the title is much smaller than that in the introduction. Therefore, the ratio of the introduction to the total text content can be used to determine whether the current image page is an introduction page.
[0065] Step S14: If the content projection value does not meet the preset Dongba scripture introduction page condition, then determine that the current image page is the text translation and annotation page of the Dongba scripture.
[0066] It can be understood that through the above steps, it can be determined that the content projection value corresponding to the current image page does not meet the preset Dongba Sutra homepage conditions and the preset Dongba Sutra introduction page conditions. Therefore, it can be determined that the current image page is the main text translation and annotation page of the Dongba Sutra.
[0067] In a specific embodiment, if the current image page is determined to meet the preset Dongba Sutra first page condition based on the content projection value corresponding to the current image page, the current image page can be determined to be the first page of a single Dongba Sutra. Furthermore, the next document image page corresponding to the current image page can be determined to be the introduction page of the Dongba Sutra. It is understandable that the first page of a single Dongba Sutra is generally a single page, and the introduction page follows the first page, which is mostly a single page, but in rare cases, there may be multiple introduction pages. Therefore, in certain specific embodiments, if the current image page is determined to meet the preset Dongba Sutra first page condition, the document image page following the current image page can also be determined to be the introduction page, and the next document image page can be determined to be the main text translation and annotation page of the Dongba Sutra.
[0068] Step S15: Recognize the Dongba scriptures, the literal translation text, and the paraphrase text in the main text annotation page according to the preset height threshold, the first preset text line spacing, and the second preset text line spacing, and obtain corresponding annotation page recognition results.
[0069] Furthermore, after determining the current image page as the main text annotation page of the Dongba scripture, the content of the main text annotation page can be identified. In a specific embodiment, the identification of Dongba scripture, literal translation text, and paraphrase text in the main text annotation page based on a preset height threshold, a first preset text line spacing, and a second preset text line spacing can include: determining content information in the main text annotation page whose pixel height is not less than the preset height threshold as Dongba scripture based on the first content projection value; determining content information in the main text annotation page whose spacing between adjacent text lines is not greater than the first preset text line spacing as literal translation text based on the first content projection value; and determining content information in the main text annotation page whose spacing between adjacent text lines is greater than the first preset text line spacing and not greater than the second preset text line spacing as paraphrase text based on the first content projection value. Specifically, Dongba scripture, literal translation, and paraphrase text in the main text annotation page can be identified separately based on a preset height threshold, a first preset text line spacing, and a second preset text line spacing. Text blocks (content information) in the main text annotation page whose pixel height is no less than the preset height threshold can be identified as Dongba scripture based on the first content projection value corresponding to the main text annotation page. Accordingly, the content of the main text annotation page can be identified based on the first preset text line spacing, and text blocks whose spacing between adjacent text lines is no greater than the first preset text line spacing can be identified as literal translation text. Furthermore, text blocks whose spacing between adjacent text lines in the main text annotation page is greater than the first preset text line spacing, while those whose spacing is less than the second preset text line spacing, can be identified as paraphrase text. It should be noted that the main text of a Dongba scripture scroll consists of three parts: the original Dongba scripture text, the pronunciation and literal translation of the scripture text, and the paraphrase text. The original scripture text is an image, and statistically, its height is between 550 and 650 pixels, indicating significant characteristics. The pronunciation of the scripture consists of two lines of International Phonetic Alphabet (IPA) and literal Chinese character translations, with a large spacing between each line, almost twice the normal line spacing. The Chinese character paraphrase, on the other hand, is organized as a single line, with a smaller spacing. Therefore, the spacing between different types of text lines can be used to achieve classification. Specifically, when determining the projection results corresponding to the image page of the main text, content with a top-bottom spacing greater than a first spacing can be identified as Dongba scripture; correspondingly, content with a text line spacing greater than a second spacing can be identified as literal scripture translation. Furthermore, for paraphrase text, content with a text line spacing less than the second spacing and greater than a third spacing in the projection results is identified as paraphrase text. It should be noted that the top-bottom spacing of a single line of Chinese character paraphrase is also almost twice the normal line spacing. In the pronunciation of the scripture, the height and width of individual characters in the International Phonetic Alphabet vary significantly, while the height and width of individual characters in the literal Chinese character translation are nearly uniform. This characteristic can be used to distinguish paraphrase lines from the pronunciation of the scripture.In a specific embodiment, for the k (k<m) text lines that constitute the body of the document, if it is assumed that the height of the i-th text line is H(i), a line of text includes n characters, the average width and height of the characters are ww(j) and wh(j) respectively, the upper spacing of the i-th text line is uw(i), and the lower spacing is dw(i), then the type of the text line Type is:
[0070]
[0071] Among them, limiting the average height of a single character in the pronunciation line to 1.5 times the average width can effectively avoid the problem of misjudging the paraphrase line as the pronunciation line. In addition, through the statistics of a large number of segmentation results of Dongba document text lines, it can be seen that the height of a single text line H(i)∈[40,70] pixels. Therefore, when the height of a single text line H(i)>75 pixels, the text line is segmented twice to ensure the effectiveness of text line segmentation. The automatic segmentation and classification results of text lines are shown in Figure 2. Figure 4 As shown in the figure, the red text line is the header line, the blue text line is the pronunciation line and the literal translation line, the green text line is the free translation line, the yellow text line is the footnote line, and the aqua text line is the footer line.
[0072] In another specific embodiment, the obtaining of the corresponding translation and annotation page recognition result may include: using a preset character recognition tool to perform content recognition on the paraphrase text of the scripture to obtain the corresponding content recognition result; the preset character recognition tool is a tool obtained by merging a preset number of character recognition tools using a preset merging formula. It is understandable that since different recognition tools have different recognition rates for Chinese characters, English characters and punctuation marks, the recognition results obtained have different degrees of errors in Chinese and English character recognition, missing punctuation marks and other problems. Therefore, the present application can further correct the recognition effect of text images and improve the recognition rate by merging the recognition results of the three tools. For example, for a text line containing n characters; if the three pointers p, q and r (p, q, r∈[1,n]) are distributed to point to the recognition result strings Ptext, Ttext and Etext of the three tools, such as Paddle OCR, Tesseract OCR and Easy OCR, respectively, then the length and calculation formula of the merged result Rtext are as follows:
[0073] Len(Rtext)=max(Len(Ptext), Len(Ttext), Len(Etext));
[0074]
[0075] Here, i,p,q,r≤Len(Rtext). Furthermore, if the three characters at the same position are not equal, the algorithm first determines whether they are punctuation marks. If so, the pointer moves backward; otherwise, the pointer remains unchanged. This ensures that even if punctuation marks are omitted from the original string, they will not interfere with the string merging process, enhancing the robustness of the recognition algorithm.
[0076] It can be seen that the present application can identify the pages of the Dongba sutra. First, the image page of the Dongba sutra is processed through morphological operations and Hough transform and other operations, which can optimize the display effect of the image page and reduce interference. Then, the type of the current image page is judged according to the projection value. The home page, introduction page and main text translation and annotation page of the Dongba sutra can be quickly separated according to the structural characteristics of the page, and the Dongba sutra can be classified and divided. The content of the relevant main text translation and annotation pages can be identified, and the content information of the main text translation and annotation pages of the Dongba sutra can be accurately identified, which facilitates the automatic reading of the pages of the Dongba sutra and improves the accuracy and reliability of the recognition results.
[0077] like Figure 5 As shown, the embodiment of the present application discloses a Dongba scripture page recognition device, comprising:
[0078] The projection value acquisition module 11 is used to obtain the content projection value corresponding to the current image page of the Dongba scripture;
[0079] The first judgment module 12 is used to judge whether the content projection value meets the preset Dongba scripture homepage condition;
[0080] The second judgment module 13 is used to judge whether the content projection value meets the preset Dongba Sutra introduction page condition when the content projection value does not meet the preset Dongba Sutra homepage condition;
[0081] The third judgment module 14 is configured to determine that the current image page is a text translation and annotation page of the Dongba Sutra when the content projection value does not meet the preset Dongba Sutra introduction page condition;
[0082] The page recognition module 15 is used to recognize the Dongba scriptures, the literal translation text and the free translation text in the main text annotation page according to the preset height threshold, the first preset text line spacing and the second preset text line spacing, and obtain the corresponding annotation page recognition result.
[0083] It can be seen from this that the present application can judge the type of the image page of the Dongba scripture to obtain the main text translation and annotation page of the Dongba scripture, and then perform content recognition on the relevant main text translation and annotation content, so that the homepage, introduction page and main text translation and annotation page of the Dongba scripture can be accurately separated, which facilitates automatic reading of the Dongba scripture and improves the accuracy and reliability of the recognition results.
[0084] In a specific embodiment, the device may further include:
[0085] A morphological operation unit, used for sequentially performing morphological opening operations on the document image of the Dongba scripture to obtain pages after the operations;
[0086] A page adjustment unit, configured to determine the boundary of the calculated page according to a vertical projection segmentation algorithm, and to adjust the angle of the calculated page using a Hough transform to obtain a complete image page;
[0087] The page structure elimination unit is used to eliminate the page structure features of the complete image page to obtain a target image page corresponding to the Dongba scripture.
[0088] Accordingly, the projection value acquisition module 11 may include:
[0089] a page determining unit, configured to determine a current image page from the target image pages;
[0090] The page projection unit is configured to perform projection segmentation processing on the current image page using a horizontal projection algorithm and a vertical projection algorithm respectively, to obtain a first content projection value and a second content projection value corresponding to the current image page.
[0091] In a specific embodiment, the first determining module 12 may include:
[0092] a first determining unit configured to determine whether there is content information in which the blank spacing between the upper and lower parts of the first content projection value is not less than the first spacing in the preset Dongba scripture homepage condition;
[0093] a first determining unit configured to determine the current image page as the first page of the Dongba scripture when the first content projection value contains content information in which the blank spacing between the upper and lower parts is not less than the first spacing;
[0094] a second determining unit configured to determine whether, in the second content projection value, there is content information whose height is not less than a first height of a preset Dongba scripture front page condition when there is no content information in the first content projection value whose upper and lower blank spaces are not less than the first spacing;
[0095] a second determining unit configured to determine the current image page as the front page of the Dongba Sutra scroll when the second content projection value contains content information whose height is not less than a first height of a preset Dongba Sutra scroll front page condition;
[0096] The third determination unit is configured to determine that the content projection value does not meet the preset Dongba Sutra homepage condition when there is no content information in the second content projection value whose height is not less than a first height of the preset Dongba Sutra homepage condition.
[0097] In a specific embodiment, the second determining module 13 may include:
[0098] A text line determination unit, configured to determine a target text line in which the text line spacing in the first content projection value is not greater than the second spacing in the preset Dongba scripture introduction page condition;
[0099] a third judging unit, configured to judge whether a ratio of the target text line to all text lines in the first content projection value is greater than a first threshold in the preset Dongba scripture introduction page;
[0100] a fourth determination unit, configured to determine the current image page as an introduction page of the Dongba Sutra when a ratio of the target text lines to all text lines in the first content projection value is greater than a first threshold value in the preset Dongba Sutra introduction page;
[0101] The fifth determination unit is configured to determine that the current image page does not meet the preset Dongba scripture introduction page condition when the proportion of the target text lines to all text lines in the first content projection value is not greater than a first threshold in the preset Dongba scripture introduction page.
[0102] In another specific embodiment, the page identification module 15 may include:
[0103] A Dongba scripture determination unit is configured to determine, based on the first content projection value, content information in the main text translation and annotation page whose pixel height is not less than a preset height threshold as Dongba scripture;
[0104] a literal translation text determination unit, configured to determine, based on the first content projection value, content information in the main text annotation page where the spacing between adjacent text lines is not greater than a first preset text line spacing as a literal translation text;
[0105] The paraphrase text determination unit is used to determine, based on the first content projection value, content information in which the spacing between adjacent text lines in the main text annotation page is greater than the first preset text line spacing and not greater than the second preset text line spacing as the paraphrase text of the scripture.
[0106] In another specific embodiment, the page identification module 15 may include:
[0107] The content recognition unit is used to perform content recognition on the paraphrase text of the scripture using a preset character recognition tool to obtain a corresponding content recognition result; the preset character recognition tool is a tool obtained by merging a preset number of character recognition tools using a preset merging formula.
[0108] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0109] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the Dongba scripture page identification method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0110] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0111] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0112] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20. The operating system 221 can be Windows Server, NetWare, Unix, Linux, etc. In addition to including a computer program capable of performing the Dongba Sutra page recognition method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.
[0113] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned method for identifying pages of Dongba scriptures. The specific steps of this method can be found in the corresponding contents disclosed in the aforementioned embodiments and will not be further described here.
[0114] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0115] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0116] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0117] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0118] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A Dongba scripture page recognition method, characterized in that: include: Get the content projection value corresponding to the current image page of the Dongba scripture; Determining whether the content projection value meets the preset Dongba scripture homepage condition; If the content projection value does not meet the preset Dongba Sutra homepage condition, determining whether the content projection value meets the preset Dongba Sutra introduction page condition; If the content projection value does not meet the preset Dongba Sutra introduction page condition, then determining that the current image page is the main text translation and annotation page of the Dongba Sutra; According to the preset height threshold, the first preset text line spacing and the second preset text line spacing, the Dongba scriptures, the literal translation text and the free translation text in the main text annotation page are respectively identified to obtain the corresponding annotation page recognition results.
2. The Dongba scripture page recognition method according to claim 1, characterized in that: Before obtaining the content projection value corresponding to the current image page of the Dongba scripture, the method further includes: Performing morphological opening operations on the document images of the Dongba scripture in sequence to obtain pages after the operation; Determining the boundary of the calculated page according to a vertical projection segmentation algorithm, and adjusting the angle of the calculated page using a Hough transform to obtain a complete image page; The page structure features of the complete image page are eliminated to obtain a target image page corresponding to the Dongba scripture.
3. The Dongba scripture page recognition method according to claim 2, characterized in that: The step of obtaining the content projection value corresponding to the current image page of the Dongba scripture includes: determining a current image page from the target image pages; A horizontal projection algorithm and a vertical projection algorithm are respectively used to perform projection segmentation processing on the current image page to obtain a first content projection value and a second content projection value corresponding to the current image page.
4. The Dongba scripture page recognition method according to claim 3, characterized in that: The determining whether the content projection value meets the preset Dongba scripture homepage condition includes: Determine whether there is content information in which the blank spacing between the upper and lower parts of the first content projection value is not less than the first spacing in the preset Dongba scripture homepage condition; If the first content projection value contains content information in which the blank spacing between the upper and lower parts is not less than the first spacing, the current image page is determined to be the first page of the Dongba scripture; If the first content projection value does not contain content information with upper and lower blank spaces not less than the first space, then determining whether the second content projection value contains content information with a height not less than a first height of a preset Dongba scripture homepage condition; If the second content projection value includes content information whose height is not less than a first height of a preset Dongba Sutra homepage condition, the current image page is determined to be the homepage of the Dongba Sutra; If the second content projection value does not contain content information with a height not less than the first height of the preset Dongba Sutra homepage condition, it is determined that the content projection value does not meet the preset Dongba Sutra homepage condition.
5. The Dongba scripture page recognition method according to claim 3, characterized in that: The determining whether the content projection value meets the preset Dongba scripture introduction page condition includes: Determine a target text line in which the text line spacing in the first content projection value is not greater than the second spacing in the preset Dongba scripture introduction page condition; Determining whether a ratio of the target text line to all text lines in the first content projection value is greater than a first threshold in the preset Dongba scripture introduction page; If yes, the current image page is determined to be the introduction page of the Dongba scripture; If not, it is determined that the current image page does not meet the preset Dongba scripture introduction page condition.
6. The Dongba scripture page recognition method according to claim 3, characterized in that: The identifying of the Dongba scriptures, the literal translation text and the paraphrase text in the main text annotation page according to the preset height threshold, the first preset text line spacing and the second preset text line spacing respectively includes: Determining, based on the first content projection value, that content information in the main text translation and annotation page whose pixel height is not less than a preset height threshold is Dongba scripture; Determining, based on the first content projection value, content information in which the spacing between adjacent text lines in the main text annotation page is not greater than a first preset text line spacing as a literal translation text; According to the first content projection value, content information in which the spacing between adjacent text lines in the main text annotation page is greater than the first preset text line spacing and not greater than the second preset text line spacing is determined as the scripture paraphrase text.
7. The Dongba scripture page recognition method according to any one of claims 1 to 6, characterized in that: Obtaining the corresponding translation annotation page recognition result includes: The content of the paraphrase text of the scripture is recognized by using a preset character recognition tool to obtain a corresponding content recognition result; the preset character recognition tool is a tool obtained by merging a preset number of character recognition tools using a preset merging formula.
8. A Dongba scripture page recognition device, characterized in that: include: The projection value acquisition module is used to obtain the content projection value corresponding to the current image page of the Dongba scripture; The first judgment module is used to judge whether the content projection value meets the preset Dongba scripture homepage condition; a second judgment module, configured to judge whether the content projection value meets the preset Dongba Sutra introduction page condition when the content projection value does not meet the preset Dongba Sutra homepage condition; a third judgment module, configured to determine that the current image page is a text translation and annotation page of the Dongba Sutra when the content projection value does not meet the preset Dongba Sutra introduction page condition; The page recognition module is used to recognize the Dongba scriptures, the literal translation text and the free translation text in the main text annotation page according to a preset height threshold, a first preset text line spacing and a second preset text line spacing, and obtain corresponding annotation page recognition results.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the Dongba scripture page recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the Dongba scripture page recognition method according to any one of claims 1 to 7.