Mongolian text line and word alignment corpus construction method based on text line recognition

By combining a text line recognition method with deep learning and image processing technology, the alignment problem of text lines and words in early lead-printed Mongolian newspapers was solved, achieving more efficient corpus construction and recognition accuracy.

CN120612700APending Publication Date: 2025-09-09INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510635692.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional methods have difficulty in accurately segmenting and aligning text lines and words in early lead-printed Mongolian newspapers, resulting in inaccurate corpus construction. In particular, due to the unique writing style of Mongolian and the limitations of printing technology, problems such as under-segmentation and over-segmentation occur.

Method used

A Mongolian text line and word alignment corpus is constructed by adopting a method based on text line recognition, combining vertical projection, OSTU algorithm, binarization, ResNet feature extraction, BiLSTM sequence encoding and CTC decoding. Through rough alignment of text lines and generation of word segmentation lines, combined with connected domain analysis and manual correction, a Mongolian text line and word alignment corpus is constructed.

Benefits of technology

It improves the accuracy of text line and word alignment, reduces recognition errors, improves the accuracy and efficiency of corpus construction, and reduces manual correction costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612700A_ABST
    Figure CN120612700A_ABST
Patent Text Reader

Abstract

The invention discloses a Mongolian text line and word alignment corpus construction method based on text line recognition, and the method comprises the steps: carrying out the image segmentation of an image containing Mongolian text based on a vertical projection method, and obtaining an initial text line linear segmentation result; performing text line coarse alignment, and training a text line recognition network by using a text line corpus subjected to coarse alignment to obtain a recognized text character sequence; predicting word segmentation lines in the text lines, and according to the generated word segmentation lines, combining connected domain analysis to generate a word frame; and connecting the frames of the words in the same line to generate a minimum rectangular frame which completely surrounds the whole text line, so as to construct the line and word alignment corpus of the Mongolian text. According to the method, traditional image processing and a deep learning model are combined, the problems of under-segmentation and over-segmentation caused by printing defects of an early-stage lead printing newspaper image are solved, the accuracy of alignment of lines and words of Mongolian texts is remarkably improved, and meanwhile the manual correction cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing and image processing, and in particular to a method for constructing a Mongolian text line and word alignment corpus based on text line recognition. Background Art

[0002] Lead-type newspaper images feature diverse layout elements and complex, compact layouts. Due to their age, the pages have faded, resulting in blurred text, faded ink, and damaged newspapers. Furthermore, due to the immaturity of early printing technology, words exhibit inconsistent glyphs, ink bleed, missing characters, and distorted text. Furthermore, there are variable gaps between type and between words, making traditional word segmentation methods based on connected component analysis or projection segmentation difficult to accurately apply. These methods are prone to under-segmentation and over-segmentation when processing early lead-type newspaper images, failing to accurately segment text lines and align words, severely impacting the accuracy of corpus construction.

[0003] This problem is even more prominent in Mongolian recognition due to its unique writing style: words are arranged vertically from top to bottom, and paragraphs are then arranged horizontally from left to right, forming a compound directional structure of "vertical columns for words, horizontal rows for paragraphs." This characteristic particularly highlights the challenges in lead type newspaper typesetting: first, vertical words cannot be split (for example, Latin words can wrap after syllables). If horizontal elements such as illustrations and tables are inserted into the layout, it is easy to cause words to break or forced column width to be compressed, destroying semantic integrity; second, manual typesetting requires precise alignment of multiple column baselines, but the difference in character height in different columns (such as high vowels) can easily cause visual dislocation, and the physical limitations of lead type molds increase the difficulty of adjustment. Craftsmen need to repeatedly measure the spacing between type and the position of dividing lines; third, when layout elements are mixed, the vertical dividing line must define the border area (such as separating the title and the text), and must also imply the logical relationship of the content through its position (such as the main news and sidebar comments), but the complex carved border lines often squeeze the text space, forcing the font size to be reduced or sacrificing the beauty of white space. The technical constraints of the lead movable type era (such as the irreversible fixing of movable type and the need for custom zinc plate splicing of illustrations) further magnified the inherent contradiction between vertical writing and horizontal paragraphs. Mongolian newspapers had to maintain a compact grid layout while compromising with the unevenness between columns and the adhesion of images and text caused by manual errors. This made Mongolian newspapers a highly recognizable but difficult to replicate visual paradigm in traditional printing technology. Summary of the Invention

[0004] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method for constructing a Mongolian text line and word alignment corpus based on text line recognition, which is used to improve the accuracy of text line and word alignment, aiming to solve the various difficulties faced in the processing of Mongolian lead movable type newspaper text, and improve the accuracy and efficiency of Mongolian lead movable type newspaper text processing, especially a method and system for constructing a word and line-level alignment corpus for documents such as early lead-printed newspaper images.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A method for constructing a Mongolian text line and word alignment corpus based on text line recognition includes the following steps:

[0007] Step 1: For an image containing Mongolian text, perform image segmentation based on the vertical projection method to obtain an initial text line segmentation result;

[0008] Step 2: perform coarse alignment of text lines, and use the coarsely aligned text line corpus to train the text line recognition network to obtain the recognized text character sequence;

[0009] Step 3: predict the word segmentation lines within the text line, and generate word bounding boxes based on the generated word segmentation lines and connected domain analysis;

[0010] Step 4: Connect the borders of each word in the same line to generate a minimum rectangular box that completely surrounds the entire text line, thereby constructing a Mongolian text line and word alignment corpus.

[0011] Furthermore, in step 1, image segmentation is performed based on the vertical projection method, and the implementation method is as follows:

[0012] Use OSTU algorithm to binarize the image;

[0013] The extreme points are obtained by vertical projection and one-dimensional Gaussian filtering and smoothing, and the image is segmented according to the positions of the extreme points to obtain the initial text line segmentation result.

[0014] Furthermore, in step 2, the method for implementing rough alignment of text lines is as follows:

[0015] The horizontal projection variance is calculated for each image, and the edge text line segmentation lines whose variance is less than the set threshold are deleted.

[0016] Furthermore, 1 / 3 of the average value of the horizontal projection variance of each image is set as the segmentation line determination threshold, d i represents the horizontal projection variance of the i-th text line, and the formula is as follows:

[0017]

[0018] N is the number of text lines, and II() represents a judgment function used to determine whether the condition is met.

[0019] Furthermore, the horizontal dividing line is not processed.

[0020] Furthermore, the text line recognition network includes a ResNet feature extraction layer, a BiLSTM sequence encoding layer and a CTC decoding layer;

[0021] First, a CNN-based ResNET feature extraction layer is used for downsampling feature extraction. The output is passed to the BiLSTM sequence encoding layer in sequence. The BiLSTM sequence encoding layer is followed by a fully connected network. The dimension of the CTC decoding layer is set to the same dimension as the output sequence mapping. An OCR model based on the CTC decoding algorithm is used to decode the text image to obtain the recognized text character sequence.

[0022] Furthermore, the step 3 predicts the word segmentation line within the text line by backtracking the receptive fields of the spaces and non-space characters on the input image, and the implementation method is as follows:

[0023] The word segmentation line of the text line is determined according to the mapping area of ​​the characters before and after all spaces in the character sequence in the original image.

[0024] Furthermore, the word segmentation line determination algorithm is as follows:

[0025] y i =b i +(t i -b i ) / 2

[0026] Where y i represents the i-th dividing line, t i Indicates the lower bound of the rectangular area where the character on the i-th space is mapped to the original image, b i Indicates the upper bound of the area to which the next character is mapped; the center line between the upper bound of the mapping area of ​​the previous character and the lower bound of the mapping area of ​​the previous character is used as the dividing line; when there is an intersection, the center position of the intersection area is the dividing line.

[0027] Furthermore, based on the known number of words and text lines, non-aligned text is determined and manually corrected to correct bounding box errors.

[0028] Compared with the existing technology, the present invention can better and more accurately divide words and text lines, reduce recognition and division errors, and ultimately realize the construction of a text line and word aligned corpus. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is the overall framework diagram of the present invention.

[0030] Figure 2 This is a diagram of the text line rough alignment process.

[0031] Figure 3 It is a word segmentation line generation graph.

[0032] Figure 4 It is manual correction. DETAILED DESCRIPTION

[0033] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.

[0034] In order to effectively solve the problems encountered in word and line level alignment and word alignment corpus construction of early lead-printed newspaper images and other documents, improve the accuracy and efficiency of corpus construction, and reduce labor costs. The present invention provides a Mongolian text line and word alignment corpus construction method based on text line recognition, which integrates text line segmentation, rough alignment algorithm, and backtracking algorithm to detect and delete segmentation lines of Mongolian lead movable type newspaper text. Figure 1 As shown, the main steps are as follows:

[0035] Step 1: Binarize the image and use one-dimensional Gaussian filtering to smooth the extreme points. Then perform image segmentation to obtain an initial text line segmentation result.

[0036] For example, the OSTU algorithm may be used to perform binarization processing on the image, and extreme points may be obtained through vertical projection and one-dimensional Gaussian filtering and smoothing, and the image may be segmented according to the positions of the extreme points.

[0037] Step 2: Since the traditional text line segmentation method will cause the left and right dividing lines to cause mismatch between the image and the text, an algorithm is designed to calculate the horizontal projection variance based on the relatively uniform horizontal distribution of foreground pixels of the dividing line. The edge text line dividing lines with variance less than the set threshold are deleted to perform rough alignment of the text lines.

[0038] In this invention, the horizontal uniform distribution of foreground pixels of the segmentation line is used to detect and delete the segmentation line. The algorithm calculates the horizontal projection variance of each image. The edge text line image with small variance is considered as the segmentation line. 1 / 3 of the average variance is set as the segmentation line judgment threshold. i It represents the horizontal projection variance of text line i. The horizontal dividing line is not processed. The formula is as follows:

[0039]

[0040] N is the number of text lines, and II() represents a judgment function used to determine whether the condition is met.

[0041] Figure 2The figure shows the rough alignment of text lines in this step. The text content in the figure is irrelevant to the present invention and has been fuzzy processed. Figure 2 The blue curve between the image and text in the left image is the smoothed vertical projection curve. The dividing line and decorative line to be deleted are marked with red arrows. The aligned text and image are connected by a blue arrow.

[0042] Step 3: Design a word extraction method guided by a text recognition pre-training model, using the coarsely aligned text line corpus (i.e. Figure 1 The text line recognition network is trained based on a Chinese text line image set. The model mainly consists of a ResNet feature extraction layer, a BiLSTM sequence encoding layer, and a CTC decoding layer. The trained model can be used to obtain the recognized text character sequence.

[0043] Specifically, we first use the CNN-based ResNET to perform downsampling feature extraction, and the output is passed to a bidirectional LSTM in turn. This is analyzed based on a BiLSTM model that combines the forward LSTM and the backward LSTM. The bidirectional LSTM is followed by a fully connected network. The dimension of the CTC decoding is set to the same dimension as the output sequence mapping. The OCR model based on the CTC decoding algorithm is used to decode the text image to obtain the recognized text character sequence.

[0044] Step 4: In the present invention, the word segmentation lines in the text line are predicted by backtracking and the receptive fields of the spaces and non-space characters on the input image, and the word borders are generated based on the generated word segmentation lines in combination with the connected domain analysis.

[0045] Specifically, in the inference phase, the recognized text character sequence is first obtained by pre-training the OCR model (the text line recognition network obtained in step 3), and then the word segmentation line in the text line is predicted by backtracking the receptive field of the space and non-space characters on the input image. Segmentation line generation is a key step. Since it is necessary to accurately backtrack the mapping area of ​​the space and non-space labels in the original image, it is necessary to extract the index value of the relevant character before the CTC decoder. The word segmentation line of the text line is determined based on the mapping area of ​​the characters before and after all spaces in the output character sequence (including the Mongolian space "U+202F") in the original image. Design a word segmentation line determination algorithm:

[0046] y i =b i +(t i -b i ) / 2

[0047] where y i represents the i-th dividing line, t i Indicates the lower bound of the rectangular area where the character on the i-th space is mapped to the original image, b i Indicates the upper bound of the area to which the next character is mapped.

[0048] The center line between the lower bound of the previous character's mapping area and the lower bound of the previous character's mapping area is used as the dividing line. If there is an intersection, the center position of the intersection area is the dividing line.

[0049] Finally, based on the generated word segmentation lines, the word bounding box is generated in combination with the connected domain analysis. Specifically, the connected domains between two adjacent segmentation lines are merged to obtain the final word bounding box, which is the row bounding box in the figure.

[0050] Figure 3 The following figure shows the word segmentation line generation process in this step, in which Mongolian characters have been fuzzy processed. Figure 3 The left picture shows the upper and lower boundaries, and the right picture shows the center of the intersection area as the dividing line, with incomplete characters without actual meaning displayed at the dividing line.

[0051] Step 5: In order to better accurately identify Mongolian text lines, the present invention connects the borders of each word in the same line (i.e. Figure 1 ), generating a minimum rectangular box that completely surrounds the entire text line.

[0052] For example, the bounding boxes of words in the same row are connected to generate a minimum bounding rectangle as the bounding box of the text line image.

[0053] Step 6, such as Figure 4 As shown, manual correction detects and corrects bounding box errors, and based on the known number of words and text lines, determines non-aligned text and performs manual correction.

[0054] Manual correction involves, for example, modifying bounding boxes and correcting text. Bounding box errors are detected and corrected, and based on the known number of words and text lines, misaligned text is identified and manually corrected. Aligned text lines are not processed.

[0055] Figure 4 The Aletheia tool displays the automatically generated word and line borders. In this interface, the word bounding boxes are manually checked and proofread one by one. The text in the figure is irrelevant to the present invention and has been blurred.

[0056] In summary, the present invention can effectively solve the problems of misclassification of neighboring pixels, missed detection of regions, and insufficient contextual semantic modeling in Mongolian recognition. ResNet adapts to complex layouts and low-quality images through its powerful feature extraction capabilities, BiLSTM improves the recognition performance of fuzzy words and characters with similar glyphs through contextual semantic modeling, and CTC avoids word segmentation errors and flexibly processes character sequences through sequence-to-sequence mapping. This combined model forms a unified framework through end-to-end training, significantly improving robustness and accuracy to meet the needs of practical applications; the word segmentation lines within the text line are predicted based on the receptive fields of spaces and non-space characters on the image through the backtracking method; finally, the alignment results are optimized through connected domain analysis and manual correction. The present invention combines traditional image processing with deep learning models to solve the problems of under-segmentation and over-segmentation of early lead-printed newspaper images caused by printing defects, significantly improving the accuracy of Mongolian text line and word alignment, and reducing the cost of manual correction.

Claims

1. A method for constructing a Mongolian text line and word alignment corpus based on text line recognition, characterized in that: The steps include: Step 1: For an image containing Mongolian text, perform image segmentation based on the vertical projection method to obtain an initial text line segmentation result; Step 2: perform coarse alignment of text lines, and use the coarsely aligned text line corpus to train the text line recognition network to obtain the recognized text character sequence; Step 3: predict the word segmentation lines within the text line, and generate word bounding boxes based on the generated word segmentation lines and connected domain analysis; Step 4: Connect the borders of each word in the same line to generate a minimum rectangular box that completely surrounds the entire text line, thereby constructing a Mongolian text line and word alignment corpus.

2. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 1, characterized in that: In step 1, image segmentation is performed based on the vertical projection method, and the implementation method is as follows: Use OSTU algorithm to binarize the image; The extreme points are obtained by vertical projection and one-dimensional Gaussian filtering and smoothing, and the image is segmented according to the positions of the extreme points to obtain the initial text line segmentation result.

3. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 1, characterized in that: In step 2, the method for implementing rough alignment of text lines is as follows: The horizontal projection variance is calculated for each image, and the edge text line segmentation lines whose variance is less than the set threshold are deleted.

4. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 3, characterized in that: The 1 / 3 of the average value of the horizontal projection variance of each image is set as the segmentation line determination threshold, d i represents the horizontal projection variance of the i-th text line, and the formula is as follows: N is the number of text lines, and II() represents a judgment function used to determine whether the condition is met.

5. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 1, characterized in that: The text line recognition network includes a ResNet feature extraction layer, a BiLSTM sequence encoding layer and a CTC decoding layer; First, a CNN-based ResNET feature extraction layer is used for downsampling feature extraction. The output is passed to the BiLSTM sequence encoding layer in sequence. The BiLSTM sequence encoding layer is followed by a fully connected network. The dimension of the CTC decoding layer is set to the same dimension as the output sequence mapping. An OCR model based on the CTC decoding algorithm is used to decode the text image to obtain the recognized text character sequence.

6. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 1, characterized in that: The step 3 predicts the word segmentation line within the text line by backtracking the receptive field of the space and non-space characters on the input image, and the implementation method is as follows: The word segmentation line of the text line is determined according to the mapping area of ​​the characters before and after all spaces in the character sequence in the original image.

7. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 1 or 6, characterized in that: The word segmentation line determination algorithm is as follows: y i =b i +(t i -b i ) / 2 Where y i represents the i-th dividing line, t i Indicates the lower bound of the rectangular area where the character on the i-th space is mapped to the original image, b i Indicates the upper bound of the area to which the next character is mapped; the center line between the upper bound of the mapping area of ​​the previous character and the lower bound of the mapping area of ​​the previous character is used as the dividing line; when there is an intersection, the center position of the intersection area is the dividing line.

8. The method for constructing a Mongolian text line and word alignment corpus based on text line recognition according to claim 1, characterized in that: Based on the known number of words and text lines, misaligned text is identified and manually corrected to correct bounding box errors.

Citation Information

Cited By

  • Text typesetting

    US20250131181A1