Multi-level text correction method and system based on document layout analysis
By employing a multi-level text correction method, combined with self-similarity and frequency domain feature algorithms, efficient document image processing is achieved under conditions of lack of training data and high-performance hardware. This solves the problems of OCR recognition and layout analysis of complex documents, and provides high-precision text correction and structured output.
Patent Information
- Application Number
- CN202511461255.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing technologies suffer from low OCR recognition rates and inaccurate document structure reconstruction when processing complex document images. In particular, without a large amount of training data and high-performance hardware support, it is difficult to achieve efficient text line detection and layout analysis.
A multi-level text correction method based on document layout analysis is adopted. It combines multi-scale self-similarity feature algorithm and directional frequency domain peak feature algorithm for image type discrimination, uses unsupervised clustering technology to extract text connected components, and performs horizontal alignment and skewing correction of text boxes through multi-level processing to form standard rectangular text blocks and output structured JSON data.
It efficiently processes diverse document formats under various hardware environments, maintains high recognition accuracy, adapts to complex layouts and distorted documents, provides accurate structured location information, and supports subsequent OCR and information extraction tasks.
Smart Images

Figure CN120932245A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence application technology in document image processing, specifically to a multi-level text correction method and system based on document layout analysis. Background Technology
[0002] In recent years, with the increasing demand for electronic office work and the widespread use of mobile devices, the need for digitizing paper documents has become increasingly prominent. Especially thanks to the improved performance of mobile phone cameras and advancements in image processing algorithms, individual users can easily acquire document images and apply them to intelligent tasks such as text recognition and structured information extraction. However, in practical applications, original document images often suffer from severe distortion problems, such as page curvature, shadow occlusion, blurred imaging, and unclear text. These problems directly affect the accuracy of subsequent OCR and layout restoration, thus limiting the efficiency and accuracy of automated document processing.
[0003] Several solutions exist for correcting document warping distortion. Early solutions primarily relied on text line detection and reconstructing the text orientation using mathematical transformation models (such as coordinate mapping) to restore it to a horizontal or vertical alignment. However, these traditional methods are highly sensitive to document image quality and page layout structure, often performing poorly when processing documents containing complex charts or non-standard layouts; furthermore, false detection of text lines can distort the correction results.
[0004] To improve processing accuracy, researchers have recently proposed optimization-based methods that train neural network models using iterative loss functions to approximate the ideal displacement mapping. While these methods perform better with high-quality images, they often suffer from excessive computation time in practical applications, making it difficult to meet real-time requirements.
[0005] With the advancement of deep learning technology and the emergence of large-scale labeled datasets, methods based on learned displacement field generation have gradually become a research hotspot. These methods extract features from document images by constructing neural network models and directly predict the pixel-level displacement corresponding to the distorted areas, thereby achieving high-precision image distortion removal. Especially after introducing a self-attention mechanism architecture, this technology has achieved significant results in complex document recognition and layout analysis. Although deep learning-based methods possess strong generalization capabilities and high performance, they still have certain dependencies on data quality, annotation standards, and hardware resources, facing certain adaptation bottlenecks in practical business applications.
[0006] Currently, a relatively mature alternative strategy is to directly utilize large-scale multimodal pre-trained models for end-to-end document structure reconstruction. For example, by designing specific prompts or instruction templates, the system can automatically complete text recognition and information extraction tasks, and to some extent achieve layout analysis. However, this method relies on the processing power and supporting resources of the large model itself, and suffers from "overcapacity" when processing simple document images, making it difficult to demonstrate cost-effectiveness. Summary of the Invention
[0007] This invention aims to overcome numerous challenges currently faced in the field of document image processing. For example, factors such as page curvature, shadows, and blurring often lead to unsatisfactory OCR (Optical Character Recognition) recognition rates, and the restoration of document structure is often inaccurate. To address these challenges, this invention provides a multi-level text correction method and system based on document layout analysis. This invention can accurately detect, align, and rotate text lines in complex document images even without large amounts of training data or high-performance hardware support, ultimately assisting OCR technology in outputting structured information.
[0008] This invention is achieved through the following technical solution:
[0009] A multi-level text correction method based on document layout analysis includes:
[0010] For the image to be corrected, the type of the image to be corrected is determined by combining the multi-scale self-similarity feature algorithm and the directional frequency domain peak feature algorithm, and the image to be corrected is adaptively preprocessed according to the determined image type to obtain a standardized image;
[0011] Extract the text connected components from the standardized image, use unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters, merge word clusters that meet preset conditions to form at least one text block, and obtain the minimum bounding quadrilateral of each text block to obtain the corresponding text box.
[0012] The center point coordinates of each text box are obtained respectively, and it is determined whether two text boxes are the same line of text based on the difference in the ordinates of the center points of adjacent text boxes;
[0013] Perform horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line;
[0014] The characters in the rotated text box are shaped and regularized so that the area where the characters are located forms a standard rectangular text block with neat edges and no misalignment, so as to avoid the character shape being destroyed by anti-aliasing or smoothing.
[0015] Output the corrected image and structured JSON data.
[0016] As an optimization, the multi-scale self-similarity feature algorithm calculates the structural similarity index (SSIM) of adjacent scales of the image to be corrected to determine image scaling stability. When the SSIM is ≥ 0.7, the image to be corrected is determined to be a high-quality digital document image; when the SSIM is < 0.7, the image to be corrected is determined to be a low-quality photograph. The directional frequency domain peak feature algorithm detects the peak values in the horizontal and vertical directions of the image to be corrected to recognize screen-captured images. When the peak values in the low and mid frequencies of the horizontal or vertical directions exceed a threshold, it is determined to be a screen-captured image.
[0017] As an optimization, when performing adaptive preprocessing on the image to be corrected, if it is determined to be a screen capture image or a low-quality photo, a noise removal operation is performed using a hybrid algorithm based on local extrema and frequency domain energy distribution combined with gamma curves. Then, the image to be corrected after the noise removal operation is standardized to obtain a standardized image.
[0018] As an optimization, the process of extracting the text connected components from the standardized image and using unsupervised clustering techniques to cluster each symbol in the text connected components to obtain several word clusters is as follows:
[0019] The unsupervised clustering technique is set as the DBSCAN clustering algorithm. When using the DBSCAN clustering algorithm to cluster each symbol in the connected component of the text, a 4-dimensional feature vector is constructed based on the horizontal distance, vertical distance, area ratio and centroid angle difference between any two adjacent symbols. The clustering neighborhood radius ε and the minimum number of samples MinPts are set. When the feature vector distance between two adjacent symbols is ≤ε and the number of samples in the neighborhood is ≥MinPts, the two symbols are assigned to the same cluster.
[0020] For each cluster, calculate the set of horizontal distances between all symbol pairs within the cluster. With vertical spacing set Take the set of horizontal spacings The median is the horizontal spacing of the symbol pairs. Take the set of vertical spacings The median is the vertical spacing of the symbol pairs. If the number of symbols in the cluster is ≥3, then the horizontal spacing set is removed. With vertical spacing set The deviation from the median is >1.5× The word clusters are obtained by using the symbols of outliers, the median is recalculated, the median is corrected, and the horizontal spacing set is recalculated based on the corrected median. With vertical spacing set IQR stands for interquartile range.
[0021] As an optimization, when merging word clusters that meet preset conditions to form text blocks, the preset conditions include a first merging condition and a second merging condition: the first merging condition is that the horizontal overlap ratio between word clusters is >50% and the vertical center difference is <1 / 2 of the median height of the word clusters. Word clusters that meet the first merging condition are merged into text lines, and the centroid coordinates of each text line are calculated. The second merging condition is used to merge text lines into text blocks, and the specific determination logic is as follows:
[0022] Calculate the vertical distance between the centroids of all adjacent text lines, and take the median of all vertical distances as the standard line spacing. ;
[0023] For any two text lines to be judged, calculate the actual vertical distance d between their centroids, and at the same time calculate the overlap ratio of all pixels in the vertical direction of the two text lines to be judged.
[0024] When the actual vertical distance d between two text lines to be judged satisfies When the vertical overlap ratio of all pixels is greater than 20%, the two text lines are determined to be related text lines, and the two text lines to be determined are merged into the same text block.
[0025] Repeat the previous step until all text lines that meet the association conditions have been merged into independent text blocks.
[0026] As an optimization, the specific method for obtaining the center point coordinates of the text box is as follows: The four vertices of the smallest bounding quadrilateral corresponding to the text box are denoted clockwise from the bottom left corner. Construct area weighting coefficient , ;
[0027] Calculate the coordinates of the center point of the text box. The coordinates of the center point The x and y coordinates are respectively The coordinates of the center point The calculation formula is: , .
[0028] As an optimization, the specific process for determining whether adjacent text boxes belong to the same line of text based on the difference in the ordinates of their center points is as follows:
[0029] Calculate the vertical projection span T of each text box; ; This represents the coordinate value of the i-th vertex in the vertical direction;
[0030] Calculate the dynamic threshold based on two horizontally adjacent text boxes. , , These are empirical parameters. These are the q-th text box and the (q-1)-th text box, respectively.
[0031] When the difference in the y-coordinates of the center points of two adjacent text boxes When the two adjacent text boxes are determined to be on the same line of text, These represent the ordinates of the center points of the q-th and (q-1)-th text boxes, respectively.
[0032] As an optimization, horizontal alignment and skewing correction are performed on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line. The specific process is as follows:
[0033] Calculate the average y-coordinate of all text boxes belonging to the same line of text. And based on the average ordinate Adjust the vertical position of all text boxes belonging to the same line of text so that the bottom left pixel of all text boxes belonging to the same line of text is horizontally aligned. Where N is the number of text boxes. This represents the area of the q-th text box. This represents the y-coordinate of the center point of the q-th text box;
[0034] Calculate the slope of a single text box, and calculate the tilt angle of the single text box based on the slope of the single text box. The formula for calculating the slope of a single text box is: , ,in, Represents the pixel vertex of the bottom left corner of a single text box. The coordinates; ) indicates the relationship with the vertex The vertex with the furthest horizontal distance If the tilt angle of the single frame If the value is greater than the set first threshold, it will be calculated based on the pixel vertex. Center, Angle Rotate the single text box.
[0035] As an optimization, the characters in the rotated text box undergo shape regularization processing to form a standard rectangular text block with neat edges and no misalignment in the area where the characters are located. The specific process to avoid the character shape being destroyed by anti-aliasing or smoothing is as follows:
[0036] Enlarge the rotated text box to a multiple of 4 pixels, then starting from the bottom left corner, create an integer grid of cells with a step size of 4 pixels. Set the bottom left cell as the origin and number it accordingly. ;
[0037] From the origin cell Begin by merging adjacent cells with identical content attributes, following a right-to-top order, and using the set merge rules to obtain the largest cell. The merge rules include:
[0038] Horizontal priority: If the cell With cells With consistent attributes, first select the cells mentioned above. With cells Perform a horizontal merge to obtain a horizontally merged cell group. Let the coordinates be the cell coordinates, and This means shifting one cell to the right. This indicates shifting up one cell;
[0039] Secondly, vertically: if the attribute of the cell preceding each cell in the horizontally merged cell group is consistent with the attribute of the horizontally merged cell group, then each cell in the horizontally merged cell group is merged with its corresponding preceding cell;
[0040] Attribute consistency determination: If the grayscale histogram of the pixels in the cell has only one main peak, then the cell is a solid color block. Conversely, it is a colorless block. ;
[0041] Record each largest cell according to the order generated by the merging rules. Then, connect the largest cell at the origin by the bottom left pixel after rotation correction. Based on this, reverse splicing is performed, prioritizing vertical and then horizontal, to obtain a preliminary combined text block;
[0042] The initial text blocks are then filled with gaps: for each horizontal gap, the outermost column of pixels of the adjacent small blocks on the left and right are filled with Gaussian blur; for vertical gaps, one row of pixels at the top and bottom are filled in the same way; after filling the gaps, a standard rectangular text block with neat edges and no misalignment is finally obtained, completing the layout reorganization of the original pixel information.
[0043] This invention also discloses a multi-level text correction system based on document layout analysis, used to perform the aforementioned multi-level text correction method based on document layout analysis, comprising:
[0044] The image preprocessing module is used to determine the type of the image to be corrected by combining a multi-scale self-similarity feature algorithm and a directional frequency domain peak feature algorithm, and to perform adaptive preprocessing on the image to be corrected according to the determined image type to obtain a standardized image.
[0045] The text layout analysis module is used to extract the text connected components in the standardized image, use unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters, merge word clusters that meet preset conditions to form at least one text block, and obtain the minimum bounding quadrilateral of each text block to obtain the corresponding text box.
[0046] The text box-level processing module is used to obtain the center point coordinates of each text box, and to determine whether two text boxes are the same line of text based on the difference in the ordinates of the center points of adjacent text boxes.
[0047] The out-of-frame line-level processing module is used to perform horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line.
[0048] The in-frame character-level processing module is used to perform shape regularization processing on the characters in the rotated text box, so that the area where the characters are located forms a standard rectangular text block with neat edges and no misalignment, so as to avoid the character shape being destroyed by anti-aliasing or smoothing processing.
[0049] The module processes and integrates image and location information into a structured output module, which outputs corrected images and structured JSON data.
[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0051] This invention does not require large-scale training data. Its core advantage lies in its unique design based on geometric principles and grouping rules, which gives it strong adaptability. It can flexibly handle diverse document formats and content without relying on a large number of training samples for learning.
[0052] For documents with complex layouts and distortions, this invention demonstrates extremely high robustness and can effectively handle document images in various real-world scenarios. Whether it is a scanned document, a photograph, or a printed document, it can accurately identify and analyze it. Even when the document has wrinkles, stains, or uneven lighting, it can maintain high recognition accuracy.
[0053] In terms of computing efficiency, this invention is outstanding. It is not only fast in processing speed, but also easy to deploy. It can adapt to the computing power of edge devices and ordinary PCs, without relying on high-performance GPUs, enabling it to run efficiently in various hardware environments.
[0054] Furthermore, the structured position information output provided by this invention is very accurate, which greatly facilitates character position correction for subsequent downstream tasks such as OCR (Optical Character Recognition) and information extraction. The recognized information can be rearranged according to the results of this method as needed.
[0055] Finally, the present invention is highly scalable, with multi-granularity correction at the character level, line level, and text block level. This means that the present invention can flexibly adjust the processing granularity according to different application scenarios and needs, thereby meeting the needs of various complex document processing. Attached Figure Description
[0056] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:
[0057] Figure 1 This is a flowchart of a multi-level text correction method based on document layout analysis as described in this invention;
[0058] Figure 2 A flowchart for image preprocessing;
[0059] Figure 3 A diagram illustrating the clustering effect of the DBSCAN algorithm.
[0060] Figure 4 This is a schematic diagram of the correction result for a standard rectangular frame;
[0061] Figure 5 This invention defines the module relationships and output content of the system described herein.
[0062] Figure 6 This is a schematic diagram illustrating how unsupervised clustering techniques are used to cluster adjacent characters in an image to obtain clusters in a specific scheme.
[0063] Figure 7 To Figure 6 A schematic diagram illustrating the results of text line relationship analysis using interline character group (word) clustering;
[0064] Figure 8 To Figure 7 A diagram showing the result after processing the text box;
[0065] Figure 9 To Figure 8 A diagram obtained after character-level / line-level processing;
[0066] Figure 10 To Figure 9 A schematic diagram obtained by dividing the text block into internal rectangular blocks;
[0067] Figure 11 To Figure 10 A schematic diagram showing the text block after being corrected and filled with Gaussian blur, etc.
[0068] Figure 12 To extract Figure 11 A schematic diagram of the result after OCR extraction. Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0070] This embodiment 1 provides a multi-level text correction method based on document layout analysis. In the field of image processing, accurate detection and clustering of text lines are key steps in text recognition tasks. For example... Figure 1 As shown, it includes:
[0071] S1. For the image to be corrected, the type of the image to be corrected is determined by combining the multi-scale self-similarity feature algorithm and the directional frequency domain peak feature algorithm, and the image to be corrected is adaptively preprocessed according to the determined image type to obtain a standardized image.
[0072] S2. Extract the text connected components in the standardized image, and use unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters. Merge the word clusters that meet the preset conditions to form at least one text block, and obtain the minimum bounding quadrilateral of each text block to obtain the corresponding text box.
[0073] S3. Obtain the center point coordinates of each text box, and determine whether the two text boxes are on the same line of text based on the difference in the ordinates of the center points of adjacent text boxes.
[0074] S4. Perform horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line.
[0075] S5. Perform shape regularization processing on the characters in the rotated text box so that the area where the characters are located forms a standard rectangular text block with neat edges and no misalignment, so as to avoid the character shape being destroyed by anti-aliasing or smoothing processing.
[0076] S6 outputs the corrected image and structured JSON data.
[0077] Next, each step will be explained in detail.
[0078] Step S1 is responsible for correctly extracting the required content from the image and standardizing it. The main goal of image preprocessing is to improve the quality and accuracy of subsequent feature extraction. The flowchart is as follows: Figure 2 As shown.
[0079] In some embodiments, the multi-scale self-similarity feature algorithm determines image scaling stability by calculating the structural similarity index (SSIM) of adjacent scales of the image to be corrected. When the SSIM is ≥ 0.7, the image to be corrected is determined to be a high-quality digital document image; when the SSIM is < 0.7, the image to be corrected is determined to be a low-quality photograph. The directional frequency domain peak feature algorithm recognizes screen-captured images by detecting the combined peak values in the horizontal and vertical directions of the image to be corrected in the frequency domain. When the combined peak values in the low and mid frequencies of the horizontal or vertical directions exceed a threshold, it is determined to be a screen-captured image.
[0080] It's important to note that moiré patterns are generated by the interference between the sensor grid and the screen pixel grid. The superposition of these two grids creates distinct directional peaks in the low and mid-frequency domains. Whether these peaks are horizontal or vertical depends on the angle between the camera sensor and the screen pixel grid. If one direction has a significant peak, the other will be weaker; therefore, the presence of a directional peak in either direction is sufficient for identification. If the peak doesn't exceed a certain threshold, the image is a clear photograph, screenshot, or other image without moiré patterns or noise.
[0081] This section involves determining the image type of the image to be corrected: by judging the content in the image to be corrected, the subsequent processing method is determined, including but not limited to document photos / scanned copies, screen capture images, and combining multi-scale self-similarity feature algorithms and directional frequency domain peak feature algorithms, after determining the image type, targeted image processing is performed.
[0082] The multi-scale self-similarity feature algorithm captures the stability of an image during scaling by calculating the structural similarity index between adjacent scales. Scanned documents, due to digital acquisition and the enhancement process during scanning, exhibit high self-similarity; paper photographs, due to blurring, have low self-similarity. The algorithm scales the image at k scales and then uses the SSIM algorithm to calculate its structural similarity.
[0083] ;
[0084] in, Let S represent the average structural similarity, used to measure the degree of structural similarity between multiple images or image regions. The higher S is, the higher the stability of the image, and the content is almost unaffected by scaling. In this invention, when S ≥ 0.7, the image content is determined to be a high-quality digital document image such as a scanned copy or screenshot with a relatively high digital acquisition instruction; when S < 0.7, the image content is determined to be a photograph or screen capture image.
[0085] The directional frequency domain peak feature algorithm targets moiré patterns in screen captures. Since screen captures are typically taken with the camera directly facing the viewer, the algorithm primarily detects peaks in the horizontal and vertical directions in the frequency domain. When the peak values of the horizontal or vertical projection at low and mid frequencies exceed those of a normal image, it is identified as a screen capture image. The directional projection is denoted by P, the frequency domain is f (the frequency domain to be evaluated is low and mid frequencies), and the combined peak value is T. h represents the vertical direction, and v represents the horizontal direction. When T exceeds a set threshold, it is identified as a screen capture image.
[0086] The formula for calculating the overall peak value is: ; This represents the combined peak value in the frequency domain when the frequency is f. These represent the projections in the vertical and horizontal directions, respectively, in the frequency domain at f.
[0087] It should be noted that the vertical direction can be understood as the longitudinal direction.
[0088] In some embodiments, when performing adaptive preprocessing on the image to be corrected, if it is determined to be a screen capture image or a low-quality photo, a noise removal operation is performed using a hybrid algorithm based on local extrema and frequency domain energy distribution combined with gamma curves. Then, the image to be corrected after the noise removal operation is standardized to obtain a standardized image.
[0089] More specifically:
[0090] Noise Removal: For screen-captured images or photos of poor quality, a hybrid algorithm based on local extrema and frequency domain energy distribution, combined with gamma curves, is used to remove noise such as screen reflections, moiré patterns, overexposure / underexposure, etc., in the image, enhance the contrast of key information in the image, and improve the accuracy of subsequent text and image feature extraction.
[0091] Image standardization: For document photos / scanned copies, distortion and angle corrections need to be performed on the document in the image to ensure that the geometry of the image is standardized and to eliminate interference caused by factors such as shooting angle and distortion.
[0092] This step provides high-quality, standardized input for downstream steps, ensuring the reliability of subsequent analyses.
[0093] Step S2 mainly involves performing text layout analysis on the image to be corrected in order to determine the original structure of the text.
[0094] In some embodiments, the specific process of extracting the text connected components in the standardized image and using unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters includes three steps: S2.1, S2.2, and S2.3.
[0095] S2.1. Set the unsupervised clustering technique as the DBSCAN clustering algorithm. When using the DBSCAN clustering algorithm to cluster each symbol in the connected components of the text, construct a 4-dimensional feature vector based on the horizontal distance, vertical distance, area ratio, and centroid angle difference between any two adjacent symbols. Set the clustering neighborhood radius ε and the minimum number of samples MinPts. When the feature vector distance between two adjacent symbols is ≤ε and the number of samples in the neighborhood is ≥MinPts, the two symbols are assigned to the same cluster.
[0096] The centroid can be understood as the geometric center, which is the average of the coordinates of all pixels within the set of text lines.
[0097] Geometric center pixel coordinates: ;
[0098] Where N is the number of pixels in the text line set, ( , ) represents the coordinates of the i-th pixel.
[0099] First, the standardized image is binarized to extract the text connected components (character-level symbols).
[0100] Cluster analysis: For each symbol, the unsupervised clustering technique (DBSCAN) is used to group the feature vectors of adjacent symbols into clusters, outputting character group (word) clusters. The main method is to perform clustering on any symbol... According to Manhattan distance in radius Find the nearest neighbor symbol j within the inner circle. Indicates the average character width The median of the dataset is calculated, and then a 4-dimensional eigenvector is calculated:
[0101] ;
[0102] ;
[0103] ;
[0104] in For horizontal spacing, For vertical spacing, For area ratio, This is due to the difference in the center of mass angle. These are the center coordinates in the k (horizontal or vertical) direction associated with the symbols j and i, respectively; These are the dimensions (such as width or height) of symbols i and j in the k direction, respectively. Let i and j represent the areas, respectively. It is the arctangent function; These represent the center x-coordinates of symbols i and j, respectively; These represent the center ordinates of symbols i and j, respectively.
[0105] The next step is to put all ( , , , Using as input, the DBSCAN clustering algorithm is employed. The conceptual diagram of this algorithm is shown below. Figure 3 As shown, it can adapt to the layout analysis requirements of this invention.
[0106] The basic concepts of this algorithm are as follows:
[0107] ε-neighborhood: The neighborhood of an object within a radius of ε. The area within.
[0108] Core object: if symbol of -The neighborhood contains at least One sample, i.e. ,but It is a core object.
[0109] Noise point: An object that does not belong to any cluster is called a noise point.
[0110] Cluster: A density-based cluster is the set of all objects that are connected by the highest density.
[0111] In the application of this invention, an optical recognition method is used to roughly determine the average object radius of the character. Parameters, core objects The parameters are set to 1-4, corresponding to character group (word) clusters, and noise points are marked as single characters. Figure 3 In the diagram, a dashed circle represents a cluster. , , , These are the core objects of the four clusters.
[0112] S2.2 For each cluster, calculate the set of horizontal distances between all symbol pairs within the cluster. With vertical spacing set Take the set of horizontal spacings The median is the horizontal spacing of the symbol pairs. Take the set of vertical spacings The median is the vertical spacing of the symbol pairs. If the number of symbols in the cluster is ≥3, then the horizontal spacing set is removed. With vertical spacing set The deviation from the median is >1.5× The word clusters are obtained by using the symbols of outliers, the median is recalculated, the median is corrected, and the horizontal spacing set is recalculated based on the corrected median. With vertical spacing set IQR stands for interquartile range.
[0113] For each cluster, count the set of horizontal spacings between all symbol pairs within the cluster. With vertical spacing set ;Pick The median is the horizontal spacing between characters / words. ,Pick The median is the vertical spacing between characters / words. If the number of symbols in a cluster is ≥3, then remove it. , The deviation from the median is >1.5× Outliers identified by the interquartile range (IQR) method are recalculated and the median is corrected. This method is more effective for pixel-level processing in character clustering tasks.
[0114] Specifically, for the set of horizontal spacing and vertical spacing set In this dataset, the median is first determined, and then the deviation of each data point from the median is calculated. If a data point deviates from the median by more than 1.5 times the interquartile range, it is considered an outlier and removed. This is to reduce the interference of outliers on subsequent statistical analyses such as median calculation, ensuring a more accurate horizontal spacing. and vertical spacing It better reflects the central tendency of the data, making subsequent processing based on these intervals (such as cluster-related operations) more accurate.
[0115] S2.3 When merging word clusters that meet preset conditions to form text blocks, the preset conditions include a first merging condition and a second merging condition: the first merging condition is that the horizontal overlap ratio between word clusters is >50% and the vertical center difference is <1 / 2 of the median height of the word clusters. Word clusters that meet the first merging condition are merged into text lines, and the centroid coordinates of each text line are calculated. The second merging condition is used to merge text lines into text blocks, and the specific determination logic is as follows:
[0116] A1. Calculate the vertical distance between the centroids of all adjacent text lines, and take the median of all vertical distances as the standard line spacing. The calculation logic is as follows:
[0117] Subtracting the y-coordinate of the geometric center of the next row from the y-coordinate of the geometric center of the previous row, and then taking the absolute value, gives the perpendicular distance between the centroids of the two rows. The formula can be understood as... .
[0118] A2. For any two text lines to be judged, calculate the actual vertical distance d between their centroids, and at the same time calculate the overlap ratio of all pixels of the two text lines in the vertical direction.
[0119] A3. When the actual vertical distance d between two text lines to be judged satisfies When the vertical overlap ratio of all pixels is greater than 20%, the two text lines are determined to be related text lines, and the two text lines to be determined are merged into the same text block.
[0120] A4. Repeat step A3 until all text lines that meet the association conditions have been merged into independent text blocks.
[0121] The above process can be simply understood as follows: if the clustering output is a word-level cluster, then traverse the word clusters and select those with a horizontal overlap ratio > 50% and a vertical central difference < 50%. Neighboring word clusters are merged into text lines; the formula for the vertical overlap ratio of all pixels is: ;
[0122] It is the highest vertical coordinate value in the text line. It is the lowest coordinate value in the vertical direction of the text line.
[0123] The median vertical distance between the centroids of adjacent row clusters is used as the row spacing. Calculate the vertical distance between all lines of text. Taking a 30% tolerance, if satisfy If the horizontal projection overlap ratio is greater than 20%, the corresponding lines will be merged into the same text block. For each text block, the minimum set of vertices of the bounding quadrilateral of all its symbols will be calculated, and the coordinates of the four corner points will be output to form a text box for subsequent processing.
[0124] In some embodiments, step S3 mainly includes S3.1 and S3.2. S3.1 is to obtain the coordinates of the center point of the text box, and S3.2 is to determine whether the two text boxes are the same line of text based on the difference in the vertical coordinates of the center points of adjacent text boxes.
[0125] The specific method of S3.1 is as follows:
[0126] S3.1.1, denot the four vertices of the smallest bounding quadrilateral corresponding to the text box in clockwise order, starting from the bottom left corner. Construct area weighting coefficient , ;
[0127] S3.1.2 Calculate the coordinates of the center point of the text box. The coordinates of the center point The x and y coordinates are respectively The coordinates of the center point The calculation formula is: .
[0128] This represents the coordinate value of the i-th vertex in the k-direction. The meaning is as follows: the subscript k indicates the direction of the coordinate axis, taking the value of x or y. The subscript i indicates the i-th vertex of the quadrilateral (numbered clockwise from the bottom left corner; as mentioned above, i = 1, 2, 3, 4). Therefore, It is the x-coordinate of the i-th vertex, and It is the y-coordinate of the i-th vertex. Simplified to: It is the xy coordinate value of the i-th vertex of the quadrilateral in the k-direction.
[0129] First, each detected text box in the image is analyzed in detail. These text boxes are likely irregular quadrilaterals. The coordinates of their center points are calculated by marking the four vertices of the quadrilateral clockwise from the bottom left corner. Construction area weighting coefficient This weighting makes the long diagonal contribute more, suppressing center drift caused by perspective distortion.
[0130] This step is to determine the relative position of the text boxes in the image, providing basic data for subsequent clustering and sorting.
[0131] In some embodiments, the specific process of S3.2 is as follows:
[0132] S3.2.1 Calculate the vertical projection span T of each text box;
[0133] S3.2.2 Calculate the dynamic threshold based on two horizontally adjacent text boxes. , , These are empirical parameters. These are the q-th text box and the (q-1)-th text box, respectively.
[0134] S3.2.3, When the difference in the ordinates of the center points of two adjacent text boxes When the two adjacent text boxes are determined to be on the same line of text, These represent the ordinates of the center points of the q-th and (q-1)-th text boxes, respectively.
[0135] Calculate the projected span of the quadrilateral in the vertical direction. Based on the center point sequence obtained in the previous step Define dynamic thresholds The decision to group all text boxes onto the same line.
[0136] Determined by the horizontally adjacent text boxes, the difference in the ordinates of the center points of two adjacent detection boxes is less than the vertical length of the text box. times, of which Further fine-tuning based on the statistical variance of the dataset is needed, but the empirical constant was confirmed to be [value missing] when this method was applied. At this time, it can usually meet the resolution robustness requirements of most document images:
[0137] The judgment criteria are: At this point, the two detection boxes will be determined to belong to the same line of text.
[0138] In some embodiments, the specific process of S4 is as follows:
[0139] S4.1 Calculate the average ordinate of all text boxes belonging to the same line of text. And based on the average ordinate Adjust the vertical position of all text boxes belonging to the same line of text so that the bottom left pixel of all text boxes belonging to the same line of text is horizontally aligned. Where N is the number of text boxes. This represents the area of the q-th text box. This represents the y-coordinate of the center point of the q-th text box;
[0140] S4.2 Calculate the slope of a single text box, and calculate the tilt angle of the single text box based on the slope of the single text box. The formula for calculating the slope of a single text box is: , ,in, Represents the pixel vertex of the bottom left corner of a single text box. The coordinates; ) indicates the relationship with the vertex The vertex with the furthest horizontal distance If the tilt angle of the single frame If the value is greater than the set first threshold, it will be calculated based on the pixel vertex. Center, Angle Rotate the single text box.
[0141] After clustering the text lines, the detection boxes for each line are further processed. Their average ordinate is calculated, and since there are N boxes per line, their respective areas are used for further processing. When weight:
[0142] ;
[0143] Based on this, horizontal alignment is performed. This method allows larger text boxes to define a larger portion of the horizontal line, making the entire line more stable. The result of horizontal alignment is to ensure consistency in the horizontal direction of text lines.
[0144] Next, the text content needs to be rotated and corrected to fix text boxes with excessively large tilt angles. The tilt angle... This is achieved by calculating the ratio of the text box's height to its width and then taking its arctangent value, with the reference point being the bottom-left pixel vertex. .
[0145] First calculate and The vertex with the furthest horizontal distance Then calculate the vertex to The vertical drop is calculated, then the slope of the single frame is calculated and the arctangent is obtained to determine the tilt angle. .
[0146] The calculation formula is as follows, where for The ordinate of the point:
[0147] ;
[0148] If the tilt angle If the rotation angle exceeds 3°, we will perform rotation correction on the text content. Center, Angle Rotate the box to be parallel to the edge, prioritizing ensuring that the line segment extending from the lower left corner of the text box is a right angle aligned with the image edge. This initial calibration can minimize the number of subsequent rectangular segments.
[0149] However, due to pixel unevenness at the edges of the text box, simple rotation cannot guarantee that the horizontal / vertical coordinates of the rotated rectangle are consistent, resulting in a misalignment of 1-3 pixels. Therefore, anti-aliasing and edge smoothing processing are required.
[0150] Therefore, anti-aliasing and edge smoothing are achieved through step S5.
[0151] In some embodiments, the specific process of S5 is as follows:
[0152] S5.1. Enlarge the rotated text box to a multiple of 4 pixels, then starting from the bottom left corner, create an integer grid of cells with a step size of 4 pixels; and set the bottom left corner cell as the origin for numbering, marking it as... ;
[0153] S5.2, From the origin cell Begin by merging adjacent cells with identical content attributes, following a right-to-top order, and using the set merge rules to obtain the largest cell. The merge rules include:
[0154] S5.2.1, Horizontal Priority: If the cell With cells With consistent attributes, first select the cells mentioned above. With cells Perform a horizontal merge to obtain a horizontally merged cell group. Let the coordinates be the cell coordinates, and This means shifting one cell to the right. This indicates shifting up one cell;
[0155] S5.2.2, Vertically: If the attribute of the cell preceding each cell in the horizontally merged cell group is consistent with the attribute of the horizontally merged cell group, then each cell in the horizontally merged cell group is merged with its corresponding preceding cell;
[0156] S5.2.3, Attribute Consistency Determination: If the grayscale histogram of the pixels in the cell has only one main peak, then the cell is a solid color block. Conversely, it is a colorless block. ;
[0157] S5.3. Record each largest cell in the order generated according to the merging rules. Then, connect the largest cell at the origin by the bottom left pixel after rotation correction. Based on this, reverse splicing is performed, prioritizing vertical and then horizontal, to obtain a preliminary combined text block;
[0158] S5.4. Perform gap filling processing on the initially combined text blocks: For each horizontal gap, take the outermost column of pixels of the left and right adjacent small blocks and perform Gaussian blur filling; for vertical gaps, take one row of pixels at the top and bottom and perform the same filling; after filling the gaps, a standard rectangular text block with neat edges and no misalignment is finally obtained, completing the layout reorganization of the original pixel information.
[0159] Traditional anti-aliasing and edge smoothing algorithms can destructively repair characters within text boxes, making it difficult to preserve the original character shape information. Therefore, for each rectangle's interior, a maximum rectangular block approach is used to segment the image.
[0160] First, enlarge the rotated text box to a multiple of 4. Then, starting from the bottom left corner, create an integer grid with a step size of 4 pixels. Set the bottom left rectangle as the origin rectangle (0,0) and number it accordingly. .
[0161] Then, the maximum rectangle is grown, that is, the maximum rectangle is cut out, from... Begin by placing adjacent items with the same content attributes in a "right-to-top" order. Merge them into a larger rectangle. The merging rules are as follows:
[0162] 1. Horizontal priority: If and If the attributes are consistent, expand to the right first;
[0163] 2. Vertically: If the entire row can be merged, then expand upwards to... ;
[0164] 3. Attribute consistency determination: The grayscale histogram of an inner pixel has only one main peak, which is a solid color block. Conversely, it is a colorless block. .
[0165] Record each maximum rectangle in the order it was generated. That is, the largest cell, in which The order is horizontal. The order is vertical. Then, a large rectangle is formed by connecting the bottom left pixels after rotation correction. Based on this, reverse the splicing process by prioritizing vertical alignment (aligning with the left side) and then horizontal alignment (aligning with the bottom side), ultimately allowing the largest split rectangle to be reassembled into a new standard rectangular text box.
[0166] However, during the rotation → splitting → baseline stitching operation, because the pixels do not fall exactly within integer grids, and due to image quality factors, a 1-2 pixel gap appears between the largest rectangular blocks during stitching. Although this is almost imperceptible to the naked eye, it needs to be filled for downstream recognition tasks. The system fills each horizontal gap with Gaussian blur by taking the outermost column of pixels from the left and right adjacent small blocks, and the vertical gap by taking one row of pixels from the top and bottom. After filling the gap rows, a standard rectangular text block with neat edges and no misalignment is finally obtained, completing the layout reorganization of the original pixel information.
[0167] The text box after S5, such as Figure 4 As shown.
[0168] Recorded according to the order in which the largest rectangle is generated. The core idea is to assign unique coordinates to each merged largest rectangular block. The meanings of i and j need to be understood in conjunction with the merging order of "right first, then top":
[0169] Generation order: Continues the merging logic of "right first, then top".
[0170] The largest rectangle is generated from the origin cell. Initially, following the rule of "expanding horizontally first, then vertically," the generation order is as follows:
[0171] Step 1: First process the bottom horizontal layer (the mergeable cell group in row j=0), generating the largest horizontal rectangle from left to right (if merging is done first). Form the first block, then process the right side. (Forming a second block).
[0172] Step 2: Process the upper horizontal layer (the mergeable cell group in row j=1), similarly generating the largest horizontal rectangle from left to right (if merging is possible). (forming the third block)
[0173] This process continues until all cells are merged into a block.
[0174] The coordinate definition of ) is as follows:
[0175] i (horizontal order): Within the same vertical height (i.e., the horizontal layer corresponding to the same j value), the block's arrangement number from left to right, starting with the first i=0, and increasing sequentially;
[0176] j (vertical order): The horizontal layer where the block is located, arranged from bottom to top. The bottom horizontal layer j=0, and the numbers increase sequentially.
[0177] For example: If j=0 (bottom layer) generates 2 blocks from left to right, denoted as... (Left), (Right); Layer j=1 (above j=0) generates one block from left to right, denoted as . Then the coordinates of all blocks clearly reflect their horizontal position and vertical height.
[0178] "Reverse splicing" does not mean "sponging in reverse order of generation," but rather using the bottom left corner of the rotated text box as the starting point. To establish a fixed reference point, other blocks are joined around the reference point according to the rule of "vertical priority → horizontal priority" to ultimately form a standard rectangle. The specific steps are as follows:
[0179] Step 1: Determine the baseline. Special status:
[0180] It is "the large rectangle connected by the bottom left corner pixels after rotation correction", meaning it includes the origin cell. The largest rectangle, which is also the block whose bottom left corner coincides with the bottom left corner of the original text box, is therefore selected as the splicing reference. Its function is:
[0181] Fix the "bottom left corner anchor point" of the standard rectangle after splicing to ensure that the final text box position is aligned with the core area of the original rotated text box and avoid overall offset;
[0182] Provides an "alignment reference" for other blocks; the splicing position of all blocks is based on this reference. The calculation is based on the edge.
[0183] Step 2: Perform the splicing, "Vertical priority (align with the left side) → Horizontal priority (align with the bottom side)".
[0184] The splicing logic needs to be combined with the block's coordinates (i,j), following the order of "first splicing vertically in the same column, then splicing horizontally in the same row" to ensure that there are no misalignments or gaps after splicing.
[0185] Vertical priority (aligning left edges): Blocks with the same i value are concatenated. The same i value represents blocks with the same horizontal position (i.e., aligned left and right). These blocks are concatenated from bottom to top according to their j values (j=0→j=1→j=2...), and the left edge of all blocks must be aligned with the left edge of the right edge. Align the left side:
[0186] calculate bottom edge and The vertical distance from the top edge will bottom edge and The top edges are seamlessly joined, while ensuring that the left edges of both are perfectly aligned (with no lateral offset).
[0187] If there are still blocks with i=0 and j=2, continue to connect their bottom edges to... The top edges are joined, while the left edges remain aligned. The left side eventually forms a continuous vertical rectangular strip with "i=0 column".
[0188] Next, horizontally (aligning the bottom edge): Blocks with the same j value are joined. The same j value represents blocks with the same vertical height (i.e., vertically aligned). They must be joined from left to right according to the i value (i=0→i=1→i=2...), and the bottom edge of all blocks must be aligned with the reference bottom edge of the same j value.
[0189] Taking layer j=0 as an example, the baseline is The bottom edge (coinciding with the bottom left corner of the original text box);
[0190] Find other blocks with the same j=0 (e.g.) (i=1), calculate its left side and The horizontal distance on the right side will The left side and The right side of each is seamlessly joined, while ensuring that the bottom edges of both are perfectly aligned (with no vertical offset).
[0191] If there are still blocks with j=0 and i=2, continue to connect their left edges to... The right side is aligned, while the bottom side remains aligned. The bottom edge eventually forms a continuous horizontal rectangular strip with "j=0 row".
[0192] Repeat the splicing process to form a standard rectangle. Following the order of "first splice column i=0 vertically → then splice row j=0 horizontally → then splice column i=1 vertically → then splice row j=1 horizontally", after all the blocks are spliced together, the originally scattered blocks will form a rectangle with "neat left side, neat bottom side, and no protrusions or depressions". This is the final "new standard rectangular text box".
[0193] Assume there are 3 blocks generated: (12 pixels wide, 8 pixels high, bottom left corner coordinates (0,0)) (12 pixels wide, 8 pixels high) (12 pixels wide, 8 pixels high), the stitching steps are as follows:
[0194] Prioritize merging columns i=0: bottom edge and The top edge (y=8) is aligned with the left edge (x=0), and the column i=0 forms a rectangular bar with a width of 12 pixels and a height of 16 pixels (coordinates (0,0)-(12,16)).
[0195] The horizontal row with j=0 is then filled in: The left side and The right side (x=12) is aligned with the bottom edge, and the bottom edge is aligned with y=0. At this time, the row j=0 forms a rectangular strip with a width of 24 pixels and a height of 8 pixels (coordinates (0,0)-(24,8)).
[0196] Final stitching result: The three blocks form a standard rectangle with a width of 24 pixels and a height of 16 pixels (coordinates (0,0)-(24,16)), without any gaps or misalignments.
[0197] The essence of this process is to solve the problem of "irregular edges that may exist in the text box after rotation". By splitting the text box into blocks with uniform attributes and then aligning and splicing them with a fixed reference, the "protruding, concave and tilted edges" of the original text box are forcibly eliminated, and the final output is a standard rectangular text box with straight edges that is suitable for OCR recognition.
[0198] S6 outputs corrected images and structured JSON data, providing a basis for subsequent information positioning calibration. The corrected image undergoes text recognition using an OCR tool, while the integrated structured JSON data stores essential information such as the text box number, the coordinates of the lower left corner reference point, and the text box size. After the text content recognized by the OCR tool is processed, it can be aligned to the correct position recorded in the JSON data. This results in new text that retains both the accurate document location information recorded by this method and the content information recognized by OCR technology, facilitating subsequent digital archiving, storage, and other business needs.
[0199] Example 2 discloses a multi-level text correction system based on document layout analysis, used to execute the multi-level text correction method based on document layout analysis described in Example 1, such as... Figure 5 As shown, it includes:
[0200] The image preprocessing module is used to determine the type of the image to be corrected by combining a multi-scale self-similarity feature algorithm and a directional frequency domain peak feature algorithm, and to perform adaptive preprocessing on the image to be corrected according to the determined image type to obtain a standardized image.
[0201] The text layout analysis module is used to extract the text connected components in the standardized image, use unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters, merge word clusters that meet preset conditions to form at least one text block, and obtain the minimum bounding quadrilateral of each text block to obtain the corresponding text box.
[0202] The text box-level processing module is used to obtain the center point coordinates of each text box, and to determine whether two text boxes are the same line of text based on the difference in the ordinates of the center points of adjacent text boxes.
[0203] The out-of-frame line-level processing module is used to perform horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line.
[0204] The in-frame character-level processing module is used to perform shape regularization processing on the characters in the rotated text box, so that the area where the characters are located forms a standard rectangular text block with neat edges and no misalignment, so as to avoid the character shape being destroyed by anti-aliasing or smoothing processing.
[0205] The module processes and integrates image and location information into a structured output module, which outputs corrected images and structured JSON data.
[0206] The following is an example of using this invention for document layout analysis and text correction of medical information images, illustrating the actual process of the method and system described in this invention. This embodiment aims to demonstrate the key processes and effectiveness of this invention in a real-world scenario.
[0207] 1. Text layout analysis:
[0208] Text layout analysis is performed on the image to determine the original text structure. The text is then converted into characters using binarization, and unsupervised clustering is applied to cluster adjacent characters. The resulting cluster relationships of adjacent character groups (words) are plotted in the original image as shown below. Figure 6 As shown in the figure. The results of subsequent text line relationship analysis on the interline character group (word) clusters are then plotted in the original figure, as shown below. Figure 7 As shown.
[0209] 2. Text box processing:
[0210] First, using the image after removing moiré patterns, draw text boxes according to the text layout analysis. Then, extract all parts of the original image that contain text information, such as... Figure 8 As shown.
[0211] 3. Character-level / line-level processing:
[0212] Using the method proposed in this invention, combined with the character group (word) cluster relationships analyzed from document layout, line-level alignment and rotation correction are performed on text blocks, ensuring that the bottom left corner of the same line frame is aligned with a right angle to a horizontal line, such as... Figure 9 As shown.
[0213] However, due to the small angle and image pixel count, the perceived slope is caused by the jagged edges of the rectangular text block, which cannot be corrected by rotation. Therefore, the method proposed in this invention is used to divide the text block into internal rectangular sections, such as... Figure 10 As shown.
[0214] Subsequently, using the correction method proposed in this invention, the rectangular text block is corrected and filled with Gaussian blur, thus correcting the text block into a aligned rectangle from a character-level perspective, such as... Figure 11As shown.
[0215] The overall image processing result after OCR extraction is as follows: Figure 12 The instructions are clear in structure and complete in content, which greatly facilitates the subsequent storage of persistent electronic documents.
[0216] This invention creatively proposes a set of text line detection and structured reconstruction techniques for complex scenes. Its core innovation lies in solving the systemic problems of traditional methods in image type discrimination, text line clustering, geometric distortion correction, and pixel-level reconstruction through a multi-level adaptive processing mechanism. Specifically, it includes the following innovations:
[0217] 1. Image type adaptive preprocessing mechanism:
[0218] By integrating multi-scale self-similarity features (scaled structural stability) and directional frequency domain peak features (moiré pattern detection), a joint criterion is constructed using quantification indicators S (scale stability mean) and T (directional projection peak). This avoids the limitations of traditional single image processing, enables collaborative perception of noise type and image source, and provides standardized images with optimized type for subsequent modules.
[0219] 2. Unsupervised text layout reconstruction techniques:
[0220] Based on connected component DBSCAN clustering, a three-level architecture of character-word-line is completed, using the median spacing dynamic reduction algorithm: based on the area-weighted centroid ordinate and the projected span T, combined with the empirical threshold λ = 1 / 5. This method achieves unsupervised line aggregation. Based on centroid coordinate sorting, it reconstructs the geometric consistency of slanted text lines by sorting and horizontally aligning the text using the y-coordinate of the text box's center point. This solves problems such as line segmentation and character distortion / drift in non-uniformly typed text using traditional rule-based methods.
[0221] 3. Pixel-preserving text line geometric correction technology:
[0222] The rotated text box is divided into a sub-rectangular grid, and then sequentially translated and stitched together with the lower left corner as the reference to avoid pixel misalignment caused by rotation. For the 1-2 pixel loss caused by segmented stitching, a non-destructive Gaussian blur fill is used to eliminate jagged edges while maintaining the integrity of the character structure. Traditional methods cannot handle pixel-level tilt due to global rotation. This invention achieves compatibility between rotation correction and original pixel preservation through segmented translation and local filling.
[0223] 4. Location-Content Dual Closed-Loop Structured Output System:
[0224] JSON-OCR dual-modal alignment technology: The lower left corner coordinates and the size of the bounding rectangle after text line correction are written into JSON metadata and spatially bound to the OCR recognition result; It breaks through the limitations of OCR structured documents caused by the influence of the original image, and builds an integrated framework of geometric layout and semantic content, providing a foundation for the reconstruction of structured documents such as certificates / contracts.
[0225] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-level text correction method based on document layout analysis, characterized in that, include: For the image to be corrected, the type of the image to be corrected is determined by combining the multi-scale self-similarity feature algorithm and the directional frequency domain peak feature algorithm, and the image to be corrected is adaptively preprocessed according to the determined image type to obtain a standardized image; Extract the text connected components from the standardized image, use unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters, merge word clusters that meet preset conditions to form at least one text block, and obtain the minimum bounding quadrilateral of each text block to obtain the corresponding text box. The center point coordinates of each text box are obtained respectively, and it is determined whether two text boxes are the same line of text based on the difference in the ordinates of the center points of adjacent text boxes; Perform horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line; The characters in the rotated text box are shaped and regularized so that the area where the characters are located forms a standard rectangular text block with neat edges and no misalignment, so as to avoid the character shape being destroyed by anti-aliasing or smoothing. Output the corrected image and structured JSON data.
2. The multi-level text correction method based on document layout analysis according to claim 1, characterized in that, The multi-scale self-similarity feature algorithm calculates the structural similarity index (SSIM) of adjacent scales of the image to be corrected to determine image scaling stability. When the SSIM is ≥ 0.7, the image to be corrected is determined to be a high-quality digital document image; when the SSIM is < 0.7, the image to be corrected is determined to be a low-quality photograph. The directional frequency domain peak feature algorithm detects the peak values in the horizontal and vertical directions of the image to be corrected to recognize screen-captured images. When the peak values in the low and mid frequencies of the horizontal or vertical directions exceed the threshold, it is determined to be a screen-captured image.
3. The multi-level text correction method based on document layout analysis according to claim 1, characterized in that, When performing adaptive preprocessing on the image to be corrected, if it is determined to be a screen capture image or a low-quality photo, a noise removal operation is performed using a hybrid algorithm based on local extrema and frequency domain energy distribution combined with gamma curves. Then, the image to be corrected after the noise removal operation is standardized to obtain a standardized image.
4. The multi-level text correction method based on document layout analysis according to claim 1, characterized in that, The specific process of extracting the text connected components from the standardized image and using unsupervised clustering techniques to cluster each symbol in the text connected components to obtain several word clusters is as follows: The unsupervised clustering technique is set as the DBSCAN clustering algorithm. When using the DBSCAN clustering algorithm to cluster each symbol in the connected component of the text, a 4-dimensional feature vector is constructed based on the horizontal distance, vertical distance, area ratio and centroid angle difference between any two adjacent symbols. The clustering neighborhood radius ε and the minimum number of samples MinPts are set. When the feature vector distance between two adjacent symbols is ≤ε and the number of samples in the neighborhood is ≥MinPts, the two symbols are assigned to the same cluster. For each cluster, calculate the set of horizontal distances between all symbol pairs within the cluster. With vertical spacing set Take the set of horizontal spacings The median is the horizontal spacing of the symbol pairs. Take the set of vertical spacings The median is the vertical spacing of the symbol pairs. If the number of symbols in the cluster is ≥3, then the horizontal spacing set is removed. With vertical spacing set The deviation from the median is >1.5× The word clusters are obtained by using the symbols of outliers, the median is recalculated, the median is corrected, and the horizontal spacing set is recalculated based on the corrected median. With vertical spacing set IQR stands for interquartile range.
5. A multi-level text correction method based on document layout analysis according to claim 1, characterized in that, When merging word clusters that meet preset conditions to form a text block, the preset conditions include a first merging condition and a second merging condition: the first merging condition is that the horizontal overlap ratio between word clusters is >50% and the vertical center difference is <1 / 2 of the median height of the word cluster. Word clusters that meet the first merging condition are merged into text lines, and the centroid coordinates of each text line are calculated. The second merging condition is used to merge text lines into text blocks, and the specific determination logic is as follows: Calculate the vertical distance between the centroids of all adjacent text lines, and take the median of all vertical distances as the standard line spacing. ; For any two text lines to be judged, calculate the actual vertical distance d between their centroids, and at the same time calculate the overlap ratio of all pixels in the vertical direction of the two text lines to be judged. When the actual vertical distance d between two text lines to be judged satisfies When the vertical overlap ratio of all pixels is greater than 20%, the two text lines are determined to be related text lines, and the two text lines to be determined are merged into the same text block. Repeat the previous step until all text lines that meet the association conditions have been merged into independent text blocks.
6. The multi-level text correction method based on document layout analysis according to claim 1, characterized in that, The specific method for obtaining the center point coordinates of the text box is as follows: The four vertices of the smallest bounding quadrilateral corresponding to the text box are denoted clockwise from the bottom left corner as follows: Construct area weighting coefficient , ; Calculate the coordinates of the center point of the text box. The coordinates of the center point The x and y coordinates are respectively The coordinates of the center point The calculation formula is: , This represents the coordinate value of the i-th vertex in the k-direction.
7. A multi-level text correction method based on document layout analysis according to claim 1, characterized in that, The specific process for determining whether adjacent text boxes belong to the same line of text based on the difference in the ordinates of their center points is as follows: Calculate the vertical projection span T of each text box; ; This represents the coordinate value of the i-th vertex in the vertical direction; Calculate the dynamic threshold based on two horizontally adjacent text boxes. , , These are empirical parameters. These are the q-th text box and the (q-1)-th text box, respectively. When the difference in the y-coordinates of the center points of two adjacent text boxes When the two adjacent text boxes are determined to be on the same line of text, These represent the ordinates of the center points of the q-th and (q-1)-th text boxes, respectively.
8. A multi-level text correction method based on document layout analysis according to claim 1, characterized in that, The specific process of performing horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line, is as follows: Calculate the average y-coordinate of all text boxes belonging to the same line of text. And based on the average ordinate Adjust the vertical position of all text boxes belonging to the same line of text so that the bottom left pixel of all text boxes belonging to the same line of text is horizontally aligned. Where N is the number of text boxes. This represents the area of the q-th text box. This represents the y-coordinate of the center point of the q-th text box; Calculate the slope of a single text box, and calculate the tilt angle of the single text box based on the slope of the single text box. The formula for calculating the slope of a single text box is: , ,in, Represents the pixel vertex of the bottom left corner of a single text box. The coordinates; ) indicates the relationship with the vertex The vertex with the furthest horizontal distance The coordinates; if the tilt angle of the single frame If the value is greater than the set first threshold, it will be calculated based on the pixel vertex. Center, Angle Rotate the single text box.
9. A multi-level text correction method based on document layout analysis according to claim 1, characterized in that, The specific process of performing shape regularization on the characters in the rotated text box to form a standard rectangular text block with neat edges and no misalignment, in order to avoid the character shape being destroyed by anti-aliasing or smoothing, is as follows: Enlarge the rotated text box to a multiple of 4 pixels, then starting from the bottom left corner, create an integer grid of cells with a step size of 4 pixels. Set the bottom left cell as the origin and number it accordingly. ; From the origin cell Begin by merging adjacent cells with identical content attributes, following a right-to-top order, and using the set merge rules to obtain the largest cell. The merge rules include: Horizontal priority: If the cell With cells With consistent attributes, first select the cells mentioned above. With cells Perform a horizontal merge to obtain a horizontally merged cell group. Let the coordinates be the cell coordinates, and This means shifting one cell to the right. This indicates shifting up one cell; Secondly, vertically: if the attribute of the cell preceding each cell in the horizontally merged cell group is consistent with the attribute of the horizontally merged cell group, then each cell in the horizontally merged cell group is merged with its corresponding preceding cell; Attribute consistency determination: If the grayscale histogram of the pixels in the cell has only one main peak, then the cell is a solid color block. Conversely, it is a colorless block. ; Record each largest cell according to the order generated by the merging rules. Then, connect the largest cell at the origin by the bottom left pixel after rotation correction. Based on this, reverse splicing is performed, prioritizing vertical and then horizontal, to obtain a preliminary combined text block; The initial text blocks are then filled with gaps: for each horizontal gap, the outermost column of pixels of the adjacent small blocks on the left and right are filled with Gaussian blur; for vertical gaps, one row of pixels at the top and bottom are filled in the same way; after filling the gaps, a standard rectangular text block with neat edges and no misalignment is finally obtained, completing the layout reorganization of the original pixel information.
10. A multi-level text correction system based on document layout analysis, used to execute the multi-level text correction method based on document layout analysis as described in any one of claims 1-9, characterized in that, include: The image preprocessing module is used to determine the type of the image to be corrected by combining a multi-scale self-similarity feature algorithm and a directional frequency domain peak feature algorithm, and to perform adaptive preprocessing on the image to be corrected according to the determined image type to obtain a standardized image. The text layout analysis module is used to extract the text connected components in the standardized image, use unsupervised clustering technology to cluster each symbol in the text connected components to obtain several word clusters, merge word clusters that meet preset conditions to form at least one text block, and obtain the minimum bounding quadrilateral of each text block to obtain the corresponding text box. The text box-level processing module is used to obtain the center point coordinates of each text box, and to determine whether two text boxes are the same line of text based on the difference in the ordinates of the center points of adjacent text boxes. The out-of-frame line-level processing module is used to perform horizontal alignment and skewing correction on all text boxes belonging to the same line of text, so that all text boxes belonging to the same line of text are rotated to the same horizontal line. The in-frame character-level processing module is used to perform shape regularization processing on the characters in the rotated text box, so that the area where the characters are located forms a standard rectangular text block with neat edges and no misalignment, so as to avoid the character shape being destroyed by anti-aliasing or smoothing processing. The module processes and integrates image and location information into a structured output module, which outputs corrected images and structured JSON data.
Citation Information
Patent Citations
Correction method for warped document image
CN108921804A
Text picture correction method and device, electronic equipment and machine readable storage medium
CN111832371A
Certificate identification method based on deep learning OCR and layout structure
CN112926469A
Engineering image text detection and identification method, device and system
CN114049648A
Algorithm for detecting and correcting any inclination angle of text image
CN116363654A