Layout analysis method and system, electronic device, and storage medium
By clustering and sorting text lines, the problem of text order recovery in document image layout analysis is solved, and the accurate sorting and splicing of text content is achieved, which is in line with human reading habits.
Patent Information
- Application Number
- PCT/CN2025/070017
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-10
AI Technical Summary
The prior art is difficult to effectively analyze the layout of document images, and it is impossible to accurately restore the reading order of text.
The clustering algorithm is used to cluster text lines, determine paragraphs, and sort them according to the location information of the paragraphs, and finally restore the reading order of the text.
It realizes the accurate sorting and splicing of text content, conforms to human reading habits, and improves the accuracy of layout analysis.
Smart Images

Figure CN2025070017_10072025_PF_FP_ABST
Abstract
Description
Layout analysis method, system, electronic device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202410011527.8 filed in China on January 3, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] Embodiments of the present invention relate to the field of image processing technology, and in particular to a layout analysis method, system, electronic device, and storage medium. Background Art
[0004] Layout analysis analyzes and identifies text and layout elements within an input document image, linking multiple lines of text and restoring the text order consistent with human reading habits. Analyzing document image layouts to obtain accurate text content with accurate text order is a current research hotspot. Summary of the Invention
[0005] Embodiments of the present invention provide a layout analysis method, system, electronic device, and storage medium for solving the problem of how to perform layout analysis on document images to obtain text content information with accurate text order.
[0006] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0007] In a first aspect, an embodiment of the present invention provides a layout analysis method, comprising:
[0008] Performing text detection and text recognition on the image to obtain the content of the text line in the image and the position information of the text line;
[0009] Clustering the position information of the text lines to obtain paragraphs consisting of the text lines;
[0010] Determining position information of the paragraph according to position information of text lines within the paragraph;
[0011] determining the page area to which the paragraph belongs according to the position information of the paragraph, and sorting the paragraphs in the page area according to the position information of the paragraph;
[0012] The text lines in the paragraph are sorted according to the position information of the text lines.
[0013] Optionally, clustering the position information of the text lines to obtain paragraphs consisting of the text lines includes:
[0014] Obtaining a data set, a domain radius, and a minimum number of elements in a neighborhood, wherein the data set is position information of the text line and the elements are text lines;
[0015] The data set, the domain radius, and the minimum number of elements in the neighborhood are input into a clustering algorithm to obtain the category to which each text line belongs, and the text lines belonging to the same category belong to the same paragraph.
[0016] Optionally, the field radius is determined according to the width of the image in a first direction, the image resolution of the image, and information related to line spacing of the image, where the first direction is the arrangement direction of the text lines in the image.
[0017] Optionally, the position information of the text line includes a minimum bounding rectangle of the text line; and determining the position information of the paragraph according to the position information of the text line in the paragraph includes:
[0018] Obtaining the minimum and maximum x and y coordinates of each vertex of the minimum bounding rectangle of all text lines in the paragraph, and determining the minimum bounding rectangle of the paragraph based on the minimum and maximum x and y coordinates as the position information of the paragraph; or determining the position information of the paragraph using a minimum bounding rectangle algorithm based on the position information of the text lines in the paragraph;
[0019] Calculating the intersection-over-union ratio of the minimum bounding rectangles of different paragraphs;
[0020] If the intersection-over-union ratio between the two paragraphs is less than or equal to a first threshold, it is determined that the two paragraphs belong to the same paragraph, the two paragraphs are merged, and position information of the merged paragraph is determined.
[0021] Optionally, determining the page area to which the paragraph belongs based on the position information of the paragraph, the method further includes:
[0022] Using a line detection model, detecting book page edges and seam lines in the image;
[0023] The image is divided into different areas according to the detected edges and center seams of the book pages, wherein the different areas include the page area.
[0024] Optionally, the text lines in the paragraph are sorted according to position information of the text lines, and the method further includes:
[0025] Get the first text line and the second text line in the paragraph;
[0026] According to an overlap condition between the first text line and the second text line on the y-axis coordinate, it is determined whether the first text line and the second text line belong to the same text line.
[0027] Optionally, determining whether the first text line and the second text line belong to the same text line according to an overlap between the first text line and the second text line on the y-axis coordinate includes:
[0028] Comparing a first y-coordinate value of an upper left corner vertex of a minimum bounding rectangle of the first text line with a second y-coordinate value of an upper left corner vertex of a minimum bounding rectangle of the second text line;
[0029] If the first y-coordinate value is smaller than the second y-coordinate value, and a first distance value between the first y-coordinate value and the second y-coordinate value in the y-axis direction is smaller than a second threshold, it is determined that the first text line and the second text line belong to the same text line, and the first text line and the second text line are merged.
[0030] Optionally, the text lines in the paragraph are sorted according to position information of the text lines, and the method further includes:
[0031] Get the first text line and the second text line in the paragraph;
[0032] Obtaining line angles of the first text line and the second text line;
[0033] If the direction of the line angles of the first text line and the second text line are consistent and both are greater than a preset angle, determining a second distance value, where the second distance value is the vertical distance from the upper left corner vertex of the minimum bounding rectangle of the first text line to the upper longest side of the minimum bounding rectangle of the second text line;
[0034] If the second distance value is less than a second threshold, it is determined that the first text line and the second text line belong to the same text line, and the first text line and the second text line are merged.
[0035] Optionally, determining the second distance value includes:
[0036] When a first y-coordinate value y0 of the upper left corner vertex of the minimum bounding rectangle of the first text line is equal to a second y-coordinate value y1 of the upper left corner vertex of the minimum bounding rectangle of the second text line, the second distance value is determined based on a first formula, the first formula being: box_h_dis=|x0-x1|*sin(θ);
[0037] When the first y-coordinate value y0 is greater than the second y-coordinate value y1, the second distance value is determined based on a second formula, wherein the second formula is: box_h_dis=|x0-x1|*sin(θ)+|y0-y1| / cos(θ);
[0038] Among them, box_h_dis is the second distance value, x0 is the x-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the first text line, x1 is the x-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the second text line, and θ is the angle between the upper longest side of the minimum circumscribed rectangle of the second text line or the extension of the upper longest side and the horizontal line where the upper left corner vertex of the minimum circumscribed matrix of the first text line is located.
[0039] Optionally, sorting the paragraphs in the page area according to the position information of the paragraphs includes: sorting the paragraphs in the page area from top to bottom according to the position information of the paragraphs;
[0040] Sorting the text lines in the paragraph according to the position information of the text lines includes: sorting the text lines in the paragraph from top to bottom according to the position information of the text lines.
[0041] Optionally, the method further includes:
[0042] According to the sorted paragraphs and the text lines within the paragraphs, the contents of the text lines are spliced to obtain text content information of the image.
[0043] In a second aspect, an embodiment of the present invention provides a layout analysis system, comprising:
[0044] A text recognition module is used to perform text detection and text recognition on an image to obtain the content of a text line in the image and the position information of the text line;
[0045] a clustering module, configured to cluster the position information of the text lines to obtain paragraphs consisting of the text lines;
[0046] A first determining module, configured to determine position information of the paragraph based on position information of text lines within the paragraph;
[0047] A first sorting module is configured to determine the page region to which the paragraph belongs based on the position information of the paragraph, and sort the paragraphs in the page region based on the position information of the paragraph;
[0048] The second sorting module is configured to sort the text lines in the paragraph according to the position information of the text lines.
[0049] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the layout analysis method described in the first aspect above are implemented.
[0050] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the layout analysis method as described in the first aspect above are implemented.
[0051] In an embodiment of the present invention, the identified text lines are clustered in a clustering manner to obtain paragraphs composed of text lines, and the paragraphs and the text lines within the paragraphs are sorted. According to the order of the sorted paragraphs and the text lines within the paragraphs, text content that conforms to human reading habits can finally be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0053] FIG1 is a schematic flow chart of a layout analysis method according to an embodiment of the present invention;
[0054] FIG2 is a schematic diagram of a minimum circumscribed matrix of text lines and paragraphs according to an embodiment of the present invention;
[0055] FIG3 is a schematic diagram of merging sections whose intersection-over-union ratio is less than or equal to a first threshold according to an embodiment of the present invention;
[0056] FIG4 is a schematic diagram of a line detection model according to an embodiment of the present invention dividing an image into multiple regions;
[0057] FIG5 is a schematic diagram of a U-Net network structure for line detection according to an embodiment of the present invention;
[0058] FIG6 is a schematic diagram of the structure of Conv_Block according to an embodiment of the present invention;
[0059] FIG7 is a schematic diagram of the structure of UpSampling according to an embodiment of the present invention;
[0060] FIG8 is a schematic diagram of a method for correcting text order by using a method for determining the overlap ratio of two text lines on the y-coordinate axis according to an embodiment of the present invention;
[0061] FIG9 is a schematic diagram of an embodiment of the present invention in which a text line is recognized as two text lines due to an obstructed text line;
[0062] FIG10 is a schematic diagram of an optimized sorting method when text lines are tilted according to an embodiment of the present invention;
[0063] FIG11 is a schematic structural diagram of a layout analysis system according to an embodiment of the present invention;
[0064] FIG12 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0066] Referring to FIG1 , an embodiment of the present invention provides a layout analysis method, including:
[0067] Step 11: Perform text detection and text recognition on the image to obtain the content of the text line in the image and the position information of the text line;
[0068] In the embodiment of the present invention, the image may be a photographed or scanned image. In the embodiment of the present invention, optionally, the image may be an image of a book.
[0069] In an embodiment of the present invention, an image may be input into an Optical Character Recognition (OCR) model, and the OCR model may be used to perform text detection and text recognition on the image to obtain the content of text lines in the image and the position information of the text lines.
[0070] In an embodiment of the present invention, optionally, the position information of the text line may be represented in the form of a minimum circumscribed matrix (box), for example, the coordinates of the four vertices of the minimum circumscribed matrix of the text line are used as the position information of the text line.
[0071] In an embodiment of the present invention, optionally, the correspondence between the content of a text line and the position information of the text line may be recorded in a dictionary, wherein the key in the dictionary records the position information of the text line and the value records the content of the text line.
[0072] Step 12: clustering the position information of the text lines to obtain paragraphs consisting of the text lines;
[0073] Step 13: Determine the position information of the paragraph according to the position information of the text line in the paragraph;
[0074] Step 14: determining the page area to which the paragraph belongs based on the location information of the paragraph, and sorting the paragraphs in the page area according to the location information of the paragraph;
[0075] In an embodiment of the present invention, the position information of the paragraph can be represented in the form of a minimum bounding box. When sorting paragraphs in the same page area, they can be sorted according to the size of the y coordinate of the upper left corner vertex of the minimum bounding rectangle of the paragraph.
[0076] Step 15: Sort the text lines in the paragraph according to the position information of the text lines.
[0077] In an embodiment of the present invention, the position information of the text line can be represented in the form of a minimum circumscribed matrix (box). When sorting text lines in the same paragraph, they can be sorted according to the size of the y coordinate of the upper left corner vertex of the minimum circumscribed rectangle of the text line.
[0078] That is, according to the order of the sorted paragraphs and text lines within the paragraphs, combined with the correspondence between the position information of the text lines and the content of the text lines saved in step 11, the text line contents are spliced together in order, and finally text content that conforms to human reading habits is obtained.
[0079] It should be noted that the serial numbers of the above steps do not represent the order in which they are executed, but are only for the convenience of explanation. For example, the above step 14 can be executed before any one of steps 11-13.
[0080] In an embodiment of the present invention, the identified text lines are clustered in a clustering manner to obtain paragraphs composed of text lines, and the paragraphs and text lines within the paragraphs are sorted. According to the order of the sorted paragraphs and text lines within the paragraphs, the contents of the text lines are spliced together in order, and finally text content that conforms to human reading habits is obtained.
[0081] In step 12 above, a clustering algorithm may optionally be used to cluster the positional information of the text lines in the image to obtain paragraphs consisting of text lines. The clustering algorithm may be, for example, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), a density-based clustering algorithm. Compared to other clustering algorithms, DBSCAN does not require a pre-specified number of cluster categories, can discover clusters of any shape, and can also detect outliers (noise points) during the clustering process.
[0082] Optionally, in step 12 above, clustering the position information of the text lines to obtain paragraphs consisting of the text lines includes:
[0083] Step 121: Obtaining a data set, a domain radius, and a minimum number of elements in a neighborhood, wherein the data set is the position information of the text line, the domain radius is related to the book type, and the element is the text line;
[0084] In the embodiment of the present invention, the selection of the neighborhood radius needs to ensure that text lines belonging to the same paragraph are clustered in one paragraph, and text lines belonging to different paragraphs are clustered in different paragraphs.
[0085] In the embodiment of the present invention, the minimum number of elements in a neighborhood can be set as needed. If the minimum number of elements in a neighborhood is set to 2, but there is only one text line in a paragraph, it is considered as noise.
[0086] In an embodiment of the present invention, optionally, the field radius is determined based on the width of the image in a first direction, the image resolution of the image, and information related to line spacing of the image, where the first direction is the arrangement direction of the text lines of the image.
[0087] Optionally, the field radius is determined using the following formula: r = dis_thr * img width / img_ori width
[0088] Where r is the neighborhood radius, dis_thr is the row spacing related information of the image (it can be the row spacing itself or a value determined based on the row spacing), img width is the image resolution of the image, img_ori width is the width of the image in the first direction. For example, for children's picture books, img_ori width =1280 resolution, dis_thr=500.
[0089] Step 122: Input the data set, the domain radius, and the minimum number of elements in the neighborhood into a clustering algorithm to obtain the category to which each text line belongs. The text lines belonging to the same category belong to the same paragraph.
[0090] The clustering process of the clustering algorithm can be as follows: randomly select an element to be clustered (text line), such as text line 1, text line 2 and text line 3 are within the neighborhood radius of text line 1, then text line 1, text line 2 and text line 3 are considered to be a paragraph; then find other text lines within the neighborhood radius of text line 2, and so on, until all text lines are divided into paragraphs.
[0091] Optionally, the position information of the text line includes a minimum bounding rectangle of the text line; in the above step 13, determining the position information of the paragraph based on the position information of the text line in the paragraph includes:
[0092] Step 131: Obtain the minimum and maximum x and y coordinates of each vertex of the minimum bounding rectangle of all text lines in the paragraph, and determine the minimum bounding rectangle of the paragraph based on the minimum and maximum x and y coordinates as the position information of the paragraph; or, determine the position information of the paragraph using a minimum bounding rectangle algorithm based on the position information of the text lines in the paragraph;
[0093] Please refer to FIG. 2 , in which 21 is the minimum circumscribed matrix of a text line, and 22 is the minimum circumscribed matrix of a paragraph determined based on the minimum circumscribed matrices of all text lines in the paragraph.
[0094] Step 132: Calculate the intersection-over-union ratio of the minimum bounding rectangles of different segments;
[0095] The intersection-to-union ratio is the ratio of the intersection and union of two minimum enclosing rectangles.
[0096] Step 133: If the intersection-over-union ratio between the two paragraphs is less than or equal to the first threshold, determine that the two paragraphs belong to the same paragraph, merge the two paragraphs, and determine the position information of the merged paragraph.
[0097] The first threshold can be set as needed, for example, it can be 0.2.
[0098] In this embodiment of the present invention, steps 132 and 133 are performed for each paragraph, i.e., the IoU ratio of the paragraph is compared with that of each other paragraph to determine whether the paragraphs can be merged. As shown in FIG3 , if the IoU ratio of two paragraphs in FIG3 is less than or equal to the first threshold, the paragraphs are merged.
[0099] Optionally, the above step 14 determines the page area to which the paragraph belongs based on the position information of the paragraph, and further includes:
[0100] Step 141: using a line detection model to detect the edges and center seams of book pages in the image;
[0101] Step 142: Divide the image into different regions based on the detected edges and center seams of the book pages, wherein the different regions include the page region.
[0102] In an embodiment of the present invention, referring to FIG4 , the line detection model may use an image segmentation algorithm (such as U-Net) to detect vertical lines in an image, and the image may be divided into different regions by extending the detected vertical lines to the edge of the image.
[0103] Please refer to Figure 5, which shows the U-Net network structure for line detection according to an embodiment of the present invention. The U-Net includes several Conv_Blocks, one Conv_Middle_Block, and several UpSamplings. Conv_Block is a downsampling module, and UpSampling is an upsampling module. After the image is input into the U-Net, it is downsampled by 2 times after each Conv_Block, and the dimension size remains unchanged after each Conv_Middle_Block. It is upsampled by 2 times after each UpSampling and superimposed with the features output by the Conv_Block of the corresponding layer. The result obtained after the output passes through a convolutional layer is used for classification, that is, to segment the vertical line area in the image. The minimum enclosing rectangle of the vertical line area is obtained to obtain the coordinates of the rectangular box containing the vertical line area, and the coordinates of the midpoint of its short side are obtained to obtain the coordinate values of the endpoints of the vertical line.
[0104] Please refer to Figures 6 and 7. Figure 6 is a schematic diagram of the structure of Conv_Block according to an embodiment of the present invention. The Conv_Block includes: a conv layer, a Normalize layer, a conv layer, a Normalize layer, and a MaxPooling layer. The conv layer is a convolutional layer, the Normalize layer is a normalization layer, and the MaxPooling layer is a maximum pooling layer. Figure 7 is a schematic diagram of the structure of UpSampling according to an embodiment of the present invention. UpSampling includes: a Concat layer, multiple alternating conv layers, and a Normalize layer. The conv layer is a convolutional layer, the Normalize layer is a normalization layer, and the Concat layer is a concatenation layer.
[0105] In the embodiment of the present invention, the line detection model may also use the traditional Hough line detection algorithm.
[0106] In step 15 above, when sorting text lines within a paragraph, if a text line is obscured or the spacing between characters in a line is too large, text that was originally in one line may be detected as two text lines. Furthermore, if the height of the subsequent text line in the same line is higher than the preceding text line, the sorting algorithm using text line height will fail. To address this issue, embodiments of the present invention can use a method to determine the overlap ratio of two text lines on the y-axis to correct the text order.
[0107] That is, the above step 15 sorts the text lines in the paragraph according to the position information of the text lines, and also includes:
[0108] Get the first text line and the second text line in the paragraph;
[0109] According to an overlap condition between the first text line and the second text line on the y-axis coordinate, it is determined whether the first text line and the second text line belong to the same text line.
[0110] Optionally, determining whether the first text line and the second text line belong to the same text line according to an overlap between the first text line and the second text line on the y-axis coordinate includes:
[0111] Step 151a: Compare the first y-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the first text line with the first y-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the first text line;
[0112] Step 152a: If the first y-coordinate value is smaller than the second y-coordinate value, and a first distance value between the first y-coordinate value and the second y-coordinate value in the y-axis direction is smaller than a second threshold, determine that the first text line and the second text line belong to the same text line, and merge the first text line and the second text line.
[0113] Optionally, the second threshold may be determined based on the following formula:
[0114] box_h_dis_thr=a*(box1_h+box2_h) / 2
[0115] Among them, box_h_dis_thr is the second threshold, a is set according to prior knowledge, for example, 0.5, box1_h is the length of the minimum enclosing rectangle box1 of the first text line in the y-axis direction, and box2_h is the length of the minimum enclosing rectangle box2 of the second text line in the y-axis direction. Please refer to Figure 8. When the first y-coordinate value box1_y0 of the upper left corner vertex of box1 of the first text line is less than the first y-coordinate value box2_y0 of the upper left corner vertex of the minimum enclosing rectangle box2 of the second text line and the first distance value box_h_dis between the first y-coordinate value and the second y-coordinate value in the y-axis direction is less than box_h_dis_thr, box1 and box2 are considered to be in the same line, with box1 in front and box2 in the back.
[0116] Please refer to Figure 9. The last text line in Figure 9 is identified as two text lines due to being blocked. The minimum bounding rectangle 3 and the minimum bounding rectangle 4 of the two text lines are compared based on the above overlap and are determined to belong to the same line, so they need to be merged.
[0117] In the embodiment of the present invention, the probability of errors in sorting text lines within a paragraph can be reduced by determining the overlap ratio of two text lines on the y-coordinate axis.
[0118] In the above step 15, when sorting the text lines in the paragraph, if the text lines are tilted, the text lines that were originally two lines above and below may be considered as one line (please refer to box 1 and box 2 in Figure 10), resulting in sorting errors.
[0119] To solve the above problem, optionally, the above step 15 sorts the text lines in the paragraph according to the position information of the text lines, and also includes:
[0120] Step 151b: Obtain the first text line and the second text line in the paragraph;
[0121] Step 152b: Obtain line angles of the first text line and the second text line;
[0122] Step 153b: If the angles of the first text line and the second text line are aligned and both are greater than a predetermined angle, a second distance value is determined. The second distance value is the vertical distance from the upper left corner of the minimum bounding rectangle of the first text line to the upper longest side of the minimum bounding rectangle of the second text line (see the line connecting the two dots in FIG. 10 ).
[0123] Optionally, the preset angle may be 5 degrees, and the specific value may be set as needed.
[0124] Step 154b: If the second distance value is less than a second threshold, determine that the first text line and the second text line belong to the same text line, and merge the first text line and the second text line.
[0125] Optionally, determining the second distance value includes:
[0126] Step 153b1: When the first y-coordinate value y0 of the upper left corner vertex of the minimum bounding rectangle of the first text line is equal to the second y-coordinate value y1 of the upper left corner vertex of the minimum bounding rectangle of the second text line, the second distance value is determined based on a first formula, wherein the first formula is: box_h_dis = |x0-x1|*sin(θ);
[0127] Step 153b2: When the first y-coordinate value y0 is greater than the second y-coordinate value y1, the second distance value is determined based on a second formula, wherein the second formula is: box_h_dis = |x0-x1|*sin(θ)+|y0-y1| / cos(θ);
[0128] In the embodiment of the present invention, when the first y-coordinate value y0 is smaller than the second y-coordinate value y1, it is considered that the angle of the book in the image is too large, and the image is not further processed.
[0129] Among them, box_h_dis is the second distance value, x0 is the x-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the first text line, x1 is the x-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the second text line, and θ is the angle between the upper longest side of the minimum circumscribed rectangle of the second text line or the extension of the upper longest side and the horizontal line where the upper left corner vertex of the minimum circumscribed matrix of the first text line is located.
[0130] The above method in the embodiment of the present invention can reduce the problem of text lines that are originally two lines above and below being considered as one line due to the inclination of the text lines, thereby causing sorting errors.
[0131] Optionally, sorting the paragraphs in the page area according to the position information of the paragraphs includes: sorting the paragraphs in the page area from top to bottom according to the position information of the paragraphs;
[0132] Sorting the text lines in the paragraph according to the position information of the text lines includes: sorting the text lines in the paragraph from top to bottom according to the position information of the text lines;
[0133] Optionally, the method further includes:
[0134] According to the sorted paragraphs and the text lines within the paragraphs, the contents of the text lines are spliced to obtain text content information of the image.
[0135] Optionally, after obtaining the text content information of the image, the method further includes: automatically reading the text content information of the image in order from top to bottom. Optionally, a text-to-speech plug-in can be used to automatically read the text content information of the image to achieve automatic sequential reading of book content.
[0136] Referring to FIG. 11 , an embodiment of the present invention further provides a layout analysis system 110 , including:
[0137] A text recognition module 111 is configured to perform text detection and text recognition on an image to obtain the content of a text line in the image and position information of the text line;
[0138] A clustering module 112, configured to cluster the position information of the text lines to obtain paragraphs consisting of the text lines;
[0139] A first determining module 113, configured to determine position information of the paragraph based on position information of text lines within the paragraph;
[0140] A first sorting module 114 is configured to determine the page region to which the paragraph belongs based on the location information of the paragraph, and sort the paragraphs in the page region based on the location information of the paragraph;
[0141] The second sorting module 115 is configured to sort the text lines in the paragraph according to the position information of the text lines.
[0142] Optionally, the clustering module 112 is used to obtain a data set, a domain radius, and a minimum number of elements in a neighborhood, wherein the data set is the location information of the text line, the domain radius is related to the book type, and the element is the text line; the data set, the domain radius, and the minimum number of elements in the neighborhood are input into the clustering algorithm to obtain the category to which each text line belongs, and the text lines belonging to the same category belong to the same paragraph.
[0143] Optionally, the field radius is determined according to the width of the image in a first direction, the image resolution of the image, and information related to line spacing of the image, where the first direction is the arrangement direction of the text lines in the image.
[0144] Optionally, the position information of the text line includes the minimum enclosing rectangle of the text line; the first determination module 113 is used to obtain the minimum and maximum x, y coordinates of each vertex of the minimum enclosing rectangle of all text lines in the paragraph, and determine the minimum enclosing rectangle of the paragraph according to the minimum and maximum x, y coordinates as the position information of the paragraph; or, based on the position information of the text line in the paragraph, use the minimum enclosing rectangle algorithm to determine the position information of the paragraph; calculate the intersection-and-union ratio of the minimum enclosing rectangles of different paragraphs; if the intersection-and-union ratio between two paragraphs is less than or equal to a first threshold, determine that the two paragraphs belong to the same paragraph, merge the two paragraphs, and determine the position information of the merged paragraph.
[0145] Optionally, the layout analysis system 110 further includes:
[0146] The segmentation module is used to detect the book page edges and center seams in the image using a straight line detection model; and divide the image into different areas according to the detected book page edges and center seams, wherein the different areas include the page area.
[0147] Optionally, the layout analysis system 110 further includes:
[0148] The first processing module is used to obtain the first text line and the second text line in the paragraph; and determine whether the first text line and the second text line belong to the same text line based on the overlap of the first text line and the second text line on the y-axis coordinate.
[0149] Optionally, determining whether the first text line and the second text line belong to the same text line according to an overlap between the first text line and the second text line on the y-axis coordinate includes:
[0150] Comparing a first y-coordinate value of an upper left corner vertex of a minimum bounding rectangle of the first text line with a second y-coordinate value of an upper left corner vertex of a minimum bounding rectangle of the second text line;
[0151] If the first y-coordinate value is smaller than the second y-coordinate value, and a first distance value between the first y-coordinate value and the second y-coordinate value in the y-axis direction is smaller than a second threshold, it is determined that the first text line and the second text line belong to the same text line, and the first text line and the second text line are merged.
[0152] Optionally, the layout analysis system 110 further includes:
[0153] A second processing module is used to obtain the first text line and the second text line in the paragraph; obtain the line angles of the first text line and the second text line; if the directions of the line angles of the first text line and the second text line are consistent and both are greater than a preset angle, determine a second distance value, the second distance value being the vertical distance from the upper left corner vertex of the minimum circumscribed rectangle of the first text line to the upper longest side of the minimum circumscribed rectangle of the second text line; if the second distance value is less than a second threshold, determine that the first text line and the second text line belong to the same text line, and merge the first text line and the second text line.
[0154] Optionally, determining the second distance value includes:
[0155] When a first y-coordinate value y0 of the upper left corner vertex of the minimum bounding rectangle of the first text line is equal to a second y-coordinate value y1 of the upper left corner vertex of the minimum bounding rectangle of the second text line, the second distance value is determined based on a first formula, the first formula being: box_h_dis=|x0-x1|*sin(θ);
[0156] When the first y-coordinate value y0 is greater than the second y-coordinate value y1, the second distance value is determined based on a second formula, wherein the second formula is: box_h_dis=|x0-x1|*sin(θ)+|y0-y1| / cos(θ);
[0157] Among them, box_h_dis is the second distance value, x0 is the x-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the first text line, x1 is the x-coordinate value of the upper left corner vertex of the minimum circumscribed rectangle of the second text line, and θ is the angle between the upper longest side of the minimum circumscribed rectangle of the second text line or the extension of the upper longest side and the horizontal line where the upper left corner vertex of the minimum circumscribed matrix of the first text line is located.
[0158] Optionally, the first sorting module is configured to sort the paragraphs in the page area from top to bottom according to the position information of the paragraphs;
[0159] The second sorting module is configured to sort the text lines in the paragraph from top to bottom according to position information of the text lines.
[0160] Optionally, the layout analysis system 110 further includes:
[0161] The splicing module is used to splice the contents of the text lines according to the sorted paragraphs and the text lines in the paragraphs to obtain the text content information of the image.
[0162] Optionally, the layout analysis system 110 further includes:
[0163] The automatic reading module is used to automatically read the text content information of the image in a top-to-bottom order.
[0164] Please refer to Figure 12. An embodiment of the present invention further provides an electronic device 120, including a processor 121, a memory 122, and a computer program stored in the memory 122 and executable on the processor 121. When the computer program is executed by the processor 121, each process of the above-mentioned layout analysis method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0165] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described layout analysis method embodiment and achieves the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0166] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0168] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A layout analysis method, wherein, Including: Performing text detection and text recognition on the image to obtain the content of the text lines in the image and the position information of the text lines; Clustering the position information of the text lines to obtain paragraphs composed of the text lines; Determining the position information of the paragraph according to the position information of the text lines in the paragraph; Determining the page area to which the paragraph belongs according to the position information of the paragraph, and sorting the paragraphs in the page area according to the position information of the paragraph; Sorting the text lines in the paragraph according to the position information of the text lines.
2. The method according to claim 1, wherein, Clustering the position information of the text lines to obtain paragraphs composed of the text lines, including: Obtaining a data set, a domain radius, and the minimum number of elements in the neighborhood, where the data set is the position information of the text lines and the element is a text line; Inputting the data set, the domain radius, and the minimum number of elements in the neighborhood into a clustering algorithm to obtain the category to which each text line belongs, and the text lines belonging to the same category belong to the same paragraph.
3. The method according to claim 2, wherein, The domain radius is determined according to the width of the image in the first direction, the image resolution of the image, and the line spacing related information of the image, and the first direction is the arrangement direction of the text lines of the image.
4. The method according to claim 1, wherein The position information of the text line includes the minimum bounding rectangle of the text line; determining the position information of the paragraph according to the position information of the text lines in the paragraph includes: Obtaining the minimum and maximum x and y coordinates of the vertices of the minimum bounding rectangles of all text lines in the paragraph, and determining the minimum bounding rectangle of the paragraph according to the minimum and maximum x and y coordinates as the position information of the paragraph; or, determining the position information of the paragraph using a minimum bounding rectangle algorithm according to the position information of the text lines in the paragraph; Calculating the intersection over union of the minimum bounding rectangles of different paragraphs; If the intersection over union between two paragraphs is less than or equal to a first threshold, determining that the two paragraphs belong to the same paragraph, merging the two paragraphs, and determining the position information of the paragraph obtained after merging.
5. The method according to claim 1, wherein, Before determining the page area to which the paragraph belongs according to the position information of the paragraph, further including: Using a line detection model to detect the book page edges and the center seam in the image; Dividing the image into different regions according to the detected book page edges and center seam, and the different regions include the page area.
6. The method according to claim 1, wherein Before sorting the text lines in the paragraph according to the position information of the text lines, further including: Obtaining a first text line and a second text line in the paragraph; Determining whether the first text line and the second text line belong to the same text line according to the overlapping situation of the first text line and the second text line on the y-axis coordinate.
7. The method according to claim 6, wherein, Determining whether the first text line and the second text line belong to the same text line according to the overlapping situation of the first text line and the second text line on the y-axis coordinate, including: Compare the first y - coordinate value of the upper - left vertex of the minimum bounding rectangle of the first text line and the second y - coordinate value of the upper - left vertex of the minimum bounding rectangle of the second text line; If the first y - coordinate value is less than the second y - coordinate value, and the first distance value in the y - axis direction between the first y - coordinate value and the second y - coordinate value is less than a second threshold, determine that the first text line and the second text line belong to the same text line, and merge the first text line and the second text line.
8. The method according to claim 1, wherein Sort the text lines in the paragraph according to the position information of the text lines, and it also includes before: Obtain the first text line and the second text line in the paragraph; Obtain the line angles of the first text line and the second text line; If the directions of the line angles of the first text line and the second text line are the same and both are greater than a preset angle, determine a second distance value, where the second distance value is the vertical distance from the upper - left vertex of the minimum bounding rectangle of the first text line to the upper long side of the minimum bounding rectangle of the second text line; If the second distance value is less than the second threshold, determine that the first text line and the second text line belong to the same text line, and merge the first text line and the second text line.
9. The method according to claim 8, wherein, Determining the second distance value includes: When the first y - coordinate value y0 of the upper - left vertex of the minimum bounding rectangle of the first text line is equal to the second y - coordinate value y1 of the upper - left vertex of the minimum bounding rectangle of the second text line, the second distance value is determined based on a first formula, and the first formula is: box_h_dis = |x0 - x1| * sin(θ); When the first y - coordinate value y0 is greater than the second y - coordinate value y1, the second distance value is determined based on a second formula, and the second formula is: box_h_dis = |x0 - x1| * sin(θ)+|y0 - y1| / cos(θ); Where, box_h_dis is the second distance value, x0 is the x - coordinate value of the upper - left vertex of the minimum bounding rectangle of the first text line, x1 is the x - coordinate value of the upper - left vertex of the minimum bounding rectangle of the second text line, and θ is the angle between the upper long side or the extension of the upper long side of the minimum bounding rectangle of the second text line and the horizontal line where the upper - left vertex of the minimum bounding matrix of the first text line is located.
10. According to the method of claim 1, wherein, Sorting the paragraphs in the page area according to the position information of the paragraphs includes: sorting the paragraphs in the page area from top to bottom according to the position information; Sorting the text lines in the paragraph according to the position information of the text lines includes: sorting the text lines in the paragraph from top to bottom according to the position information.
11. The method according to claim 1, wherein, It also includes: According to the sorted paragraphs and the text lines in the paragraphs, splice the content of the text lines to obtain the text content information of the image.
12. A layout analysis system, wherein, It includes: A text recognition module, used for text detection and text recognition of an image to obtain the content of the text lines in the image and the position information of the text lines; A clustering module, configured to cluster the position information of the text lines to obtain paragraphs composed of the text lines; A first determination module, configured to determine the position information of the paragraph according to the position information of the text lines within the paragraph; A first sorting module, configured to determine the page area to which the paragraph belongs according to the position information of the paragraph, and sort the paragraphs in the page area according to the position information of the paragraph; A second sorting module, configured to sort the text lines within the paragraph according to the position information of the text lines.
13. An electronic device, wherein, Comprising: A processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the layout analysis method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium, wherein, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the layout analysis method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Document paragraph sorting method and device, electronic equipment and storage medium
CN109657221A
OFD format document paragraph identification method and device
CN114359943A
Text recognition method and device, electronic equipment and storage medium
CN116052196A
Paragraph detection method and device, electronic equipment and storage medium
CN116758573A
Layout analysis method and system, electronic equipment and storage medium
CN117576712A