Layout analysis method and device, computer-readable medium, and electronic device

By performing fusion comparison of layout analysis and contour detection of image documents, the problem of not being able to recognize image document layout information in the prior art is solved, automatic editing of electronic documents is realized, and recognition accuracy and stability are improved.

CN113642514BActive Publication Date: 2025-07-18GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111005197.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-30
Publication Date
2025-07-18
Estimated Expiration
2041-08-30

AI Technical Summary

Technical Problem

The prior art cannot effectively identify the layout information of image documents, resulting in manual intervention in editing of electronic documents, which is cumbersome to operate.

Method used

By performing layout analysis and contour detection on the target image, the minimum external rectangle frame and its marking information are obtained, and the results of the two are fused and compared to supplement the missed inspection content to generate complete layout analysis results.

Benefits of technology

Improve the accuracy and stability of layout recognition, ensure the integrity of document data content, and reduce the need for manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113642514B_ABST
    Figure CN113642514B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology, and particularly to a layout analysis method and apparatus, a computer-readable medium, and an electronic device. The method includes: obtaining a target image corresponding to a document to be processed; performing layout analysis on the target image to obtain a first target detection result; wherein the first target detection result includes a plurality of minimum bounding rectangles and corresponding marking information; and performing contour detection on the target image to obtain a second text contour detection result; wherein the second text contour detection result includes a plurality of minimum bounding rectangles and corresponding marking information; fusing and comparing the first target detection result and the second text contour detection result to obtain a supplementary content set; and determining a layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result. This method can effectively improve the accuracy of the page layout recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a layout analysis method, a layout analysis device, a computer-readable medium, and an electronic device. Background Art

[0002] With the rapid development of computer technology, users can store documents and data in the form of electronic documents, which are convenient to carry and consult; they can also directly use electronic devices to create, edit, store, and share documents. For paper-based documents, a scanner or a high-definition camera can be used to take pictures and store the image documents. However, image documents cannot be edited such as adding, deleting, or typesetting the text content.

[0003] In the related art, the OCR (Optical Character Recognition) technology can be used to recognize the text content in the image. However, the OCR technology can only recognize the text content and cannot recognize the layout information of the image document.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The present disclosure provides a layout analysis method, a layout analysis device, a computer-readable medium, and an electronic device, which can effectively improve the accuracy of the recognition result of the page layout.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, a layout analysis method is provided, including:

[0008] Obtaining a target image corresponding to a document to be processed;

[0009] Performing layout analysis on the target image to obtain a first target detection result; wherein, the first target detection result includes a plurality of minimum circumscribed rectangle frames and corresponding marking information; and

[0010] Performing contour detection on the target image to obtain a second text contour detection result; wherein, the second text contour detection result includes a plurality of minimum circumscribed rectangle frames and corresponding marking information;

[0011] Fusing and comparing the first target detection result and the second text contour detection result to obtain a supplementary content set;

[0012] Determine the layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result.

[0013] According to a second aspect of the present disclosure, there is provided a layout analysis device, including:

[0014] A target image acquisition module, configured to acquire a target image corresponding to a document to be processed;

[0015] A layout analysis module, configured to perform layout analysis on the target image to obtain a first target detection result; wherein, the first target detection result includes a plurality of minimum circumscribed rectangle frames and corresponding marking information; and

[0016] A contour analysis module, configured to perform contour detection on the target image to obtain a second text contour detection result; wherein, the second text contour detection result includes a plurality of minimum circumscribed rectangle frames and corresponding marking information;

[0017] A fusion processing module, configured to fuse and compare the first target detection result and the second text contour detection result to obtain a supplementary content set;

[0018] An analysis result output module, configured to determine the layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result.

[0019] According to a third aspect of the present disclosure, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned layout analysis method is implemented.

[0020] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:

[0021] One or more processors;

[0022] A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned layout analysis method.

[0023] An embodiment of the present disclosure provides a layout analysis method. By performing layout analysis on a target image, a corresponding first target detection result is obtained. At the same time, by performing contour detection on the target image, a second text contour detection result is obtained. Thus, the first target detection result and the second text contour detection result can be compared and fused to obtain a supplementary content set. The supplementary content set is used to supplement the undetected content of the first target detection result, thereby improving the accuracy of layout recognition while ensuring the integrity of the data content of the document to be processed. Moreover, by separately executing the layout analysis process and the text contour detection process, over-coupling of the model is avoided, and the stability of the algorithm is improved.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0026] Figure 1 A schematic diagram schematically showing a layout analysis method in an exemplary embodiment of the present disclosure;

[0027] Figure 2 A schematic diagram schematically showing a method for performing contour detection on a target image in an exemplary embodiment of the present disclosure;

[0028] Figure 3 A schematic diagram schematically showing a method for fusion comparison in an exemplary embodiment of the present disclosure;

[0029] Figure 4 A schematic diagram schematically showing the architecture of a layout analysis method in an exemplary embodiment of the present disclosure;

[0030] Figure 5 A schematic diagram schematically showing the composition of a layout analysis device in an exemplary embodiment of the present disclosure;

[0031] Figure 6 A schematic diagram schematically showing the structure of an electronic device in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0033] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0034] In the related art, with the popularization of portable electronic devices such as laptop computers and mobile phones, more and more users choose to store documents in electronic devices in electronic form, that is, directly create, edit, store, and share documents using electronic devices. For existing paper-based documents, many users will choose to use a scanner or a high-definition camera to take pictures, and then store the corresponding document images after obtaining them. Taking pictures is an effective method for electronic conversion of paper documents. However, although the document image contains all the information of the original paper document, it cannot perform editing operations such as adding or deleting text and typesetting. If there is an editing requirement, it is necessary to manually enter information such as text / images into the computer and then perform typesetting, which is very cumbersome. Therefore, after the OCR technology emerged, it was quickly applied to the application of electronic conversion of paper documents. However, the OCR technology can only recognize the text itself and cannot obtain layout information. Therefore, only the manual entry step can be omitted, and typesetting still requires manual intervention. In existing document conversion from images solutions, in some solutions, the layout analysis algorithm is not integrated and an electronic document with a layout cannot be generated; in other solutions, the layout analysis algorithm is highly coupled with other algorithms and the effect is not stable enough.

[0035] In view of the above-mentioned shortcomings and deficiencies of the prior art, in this example embodiment, a layout analysis method is provided, which can be applied to intelligent terminal devices to perform accurate layout analysis on document images. Refer to Figure 1 As shown in, the above-mentioned layout analysis method may include the following steps:

[0036] S11, obtaining a target image corresponding to the document to be processed;

[0037] S12. Perform layout analysis on the target image to obtain a first target detection result. The first target detection result includes a number of minimum bounding rectangles and corresponding marking information. And

[0038] S13. Perform contour detection on the target image to obtain a second text contour detection result. The second text contour detection result includes a number of minimum bounding rectangles and corresponding marking information.

[0039] S14. Fuse and compare the first target detection result and the second text contour detection result to obtain a supplementary content set.

[0040] S15. Determine the layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result.

[0041] In the layout analysis method provided by this exemplary embodiment, on the one hand, by performing layout analysis on the target image, a corresponding first target detection result is obtained. At the same time, by performing contour detection on the target image, a second text contour detection result is obtained, so that the first target detection result and the second text contour detection result can be compared and fused to obtain a supplementary content set. On the other hand, by using the supplementary content set to supplement the undetected content of the first target detection result, the accuracy of layout recognition is improved on the premise of ensuring the integrity of the data content of the document to be processed. And by performing the layout analysis process and the text contour detection process separately, the over-coupling of the model is avoided, and the stability of the algorithm is improved.

[0042] Next, each step of the layout analysis method in this exemplary embodiment will be described in more detail with reference to the drawings and embodiments.

[0043] In step S11, obtain the target image corresponding to the document to be processed.

[0044] In this exemplary embodiment, the above method can be applied to intelligent terminal devices such as mobile phones, tablet computers, and laptop computers. The above document to be processed can be a paper document. The user can obtain the document image corresponding to the document to be processed by means of photographing, scanning, etc., and use it as the target image. Use the target image as the output data and perform layout recognition on it. In addition, after obtaining the document image corresponding to the document to be processed input by the user, it can also be subjected to image conversion to convert the document image into a target image with a specified resolution.

[0045] For example, the above-mentioned document to be processed can be an unstructured document, such as a document data including any one or more of text, images, tables, etc., such as papers, magazines, books, etc. In the document data, different titles or outline levels can also be included. Of course, the above-mentioned document to be processed can also be a document containing structured content; for example, it can be document data such as invoices, receipts, bills, etc.

[0046] In step S12, perform layout analysis on the target image to obtain a first target detection result; wherein, the first target detection result includes a plurality of minimum bounding rectangles and corresponding marking information.

[0047] In this exemplary embodiment, take the document to be processed as an unstructured data document or a document containing partial structured data content as an example.

[0048] Specifically, a layout analysis model based on an object detection algorithm can be pre-trained, and the layout analysis model based on the object detection algorithm is used to identify the target image, identify the position of the layout content corresponding to the target image, that is, the position of the text content; and, identify the type of the layout, such as types like paragraphs, titles, etc.

[0049] In this exemplary embodiment, specifically, the above step S12 may include:

[0050] Step S121, extract features from the target image to obtain feature information;

[0051] Step S122, use a position regression model based on an object detection algorithm to calculate the feature information to obtain a rectangle box recognition result; and

[0052] Step S123, use a label classification model based on an object detection algorithm to calculate the feature information to obtain a classification recognition result of the rectangle box.

[0053] Specifically, after obtaining the target image input by the user, a corresponding first subtask can be created, that is, a first layout detection task, and the layout analysis model is used to perform layout detection on the target image.

[0054] For example, a layout detection model can be pre-trained. When selecting training samples, in the application scenario of image digitization of academic documents, the open-source dataset publaynet can be used. Or, in the case of other types of documents or application scenarios, a training dataset can also be customized. For example, use the python-docx tool to parse existing word documents as the training dataset. For a specific object detection model, a lightweight YOLO V5 model can be used, and a smaller processed image resolution can be set, such as 320 dpi.

[0055] The object detection model may include a feature extraction network, a location regression branch model, and a label classification branch model. For example, the Focus structure of the YOLO V5 model can be used as the feature extraction network to extract features from the target image to obtain the feature map corresponding to the target image. For example, a target image of 608*608*3 is input into the Focus structure, and slicing operations are adopted. First, it becomes a feature map of 304*304*12, and then after a convolution operation with 32 convolutional kernels, it finally becomes a feature map of 304*304*32.

[0056] After obtaining the feature map corresponding to the target image, the feature map can be input into the location regression branch model and the label classification branch model respectively. In the location regression label model, N 5-dimensional vectors (x, y, w, h, c) can be output; where x and y represent the center point coordinates of the rectangular box, w and h represent the length / width of the rectangular box, and c represents the predicted confidence level (between 0 and 1). In the label classification model, N K-dimensional vectors (p1, p2,..., pk) can be output; where K represents the number of possible labels; pi represents the probability that the part inside the rectangular box belongs to the i-th category, 0 <= pi <= 1, and p1 + p2 +... + pk = 1.

[0057] The outputs of the two branch models are concatenated to obtain N (5 + K)-dimensional vectors. For example, if one of the vectors is (1, 2, 3, 4, 0.8, 0.1, 0.2, 0.3, 0.4), it means that the network has a probability of 0.8 to judge that the area with the center point (1, 2) and size (3, 4) has a probability of 0.1 to belong to category 1, a probability of 0.2 to belong to category 2, a probability of 0.3 to belong to category 3, and a probability of 0.4 to belong to category 4. By removing predictions with too low confidence levels according to a certain threshold and merging overlapping rectangular boxes (NMS, non-maximum suppression), the final result can be obtained. For example, the above types can be titles, title levels, text paragraphs, and specific types determined according to the text outline, etc.

[0058] Alternatively, in some exemplary embodiments, other object detection algorithms based on the CNN model can also be used for layout recognition to obtain the text box recognition result corresponding to the target image and the classification result corresponding to each text box.

[0059] In step S13, contour detection is performed on the target image to obtain a second text contour detection result; where the second text contour detection result includes a plurality of minimum circumscribed rectangular boxes and corresponding marking information.

[0060] In this exemplary embodiment, after obtaining the target image, a corresponding second subtask can also be created for the target image to perform contour detection on the target image. For example, the first subtask and the second subtask corresponding to the target image can be executed synchronously; alternatively, the second subtask can also be executed after the first subtask is executed and the text box result is obtained.

[0061] In this exemplary embodiment, specifically, referring to Figure 2 as shown, the above step S13 may include:

[0062] Step S131, converting the target image to obtain a corresponding grayscale image;

[0063] Step S132, performing Gaussian blur processing on the grayscale image to obtain a blurred image;

[0064] Step S133, performing binarization processing on the blurred image to obtain a binary image;

[0065] Step S134, performing morphological dilation operation on the binary image to obtain a dilation result;

[0066] Step S135, performing contour detection on the dilation result to obtain the minimum bounding rectangle of each contour, and configuring corresponding marking information for each minimum bounding rectangle.

[0067] Specifically, the original target image IMG can be converted to a grayscale image GRAY. Then, Gaussian blur can be performed on GARY to obtain a blurred image BLUR; among them, the size of the blur kernel is positively correlated with the image size. When the general image size is about 1280, 7 can be taken. Then BLUR is converted to a binary image BIN. 4. Continuously perform k times of morphological dilation operations on BIN; among them, the horizontal size of the dilation kernel should be greater than the vertical size to ensure that the text in the same row / sentence can be connected together. When the general image size is about 1280, k can be taken as 3, and the dilation kernel can be set to (11, 3). Then, contour detection can be performed on the dilation result, the minimum bounding rectangle of each contour is taken, and the most common "paragraph" label is given to these boxes, which is the contour detection result.

[0068] Due to the training data and network structure, the CNN-based layout detection model will inevitably have missed detection phenomena. In the image document to image task, the integrity priority of the target document is much higher than the aesthetics of the document layout. Therefore, an additional contour detection module needs to be set to ensure that all text regions can be detected. The contour detection model is based on the recognition of text edges, so it is almost impossible to have missed detection, but it cannot obtain the corresponding semantic information (layout type). Therefore, it can only be used as a supplement or "fallback method" for the layout detection model and cannot be used alone.

[0069] In step S14, the first target detection result and the second text contour detection result are fused and compared to obtain a supplementary content set.

[0070] In this exemplary embodiment, specifically, referring to Figure 3 as shown, the above step S14 may include:

[0071] Step S141: Construct a coordinate system for the target image, mark the position of the minimum bounding rectangle in the first target detection result, and sort them according to a preset rule to obtain a first rectangle frame sequence; and, mark the position of the minimum bounding rectangle in the second target detection result, and sort them according to a preset rule to obtain a second rectangle frame sequence;

[0072] Step S142: Based on the position mark of the first rectangle frame in the first rectangle frame sequence, select the minimum bounding rectangle in the second target detection result that is above the position of the first rectangle frame, and add it to the supplementary content set;

[0073] Step S143: Based on the position mark of the last rectangle frame in the above first rectangle frame sequence, select the minimum bounding rectangle in the second target detection result that is below the position of the last rectangle frame, and add it to the supplementary content set; and

[0074] Step S144: Obtain the gap frames between two adjacent rectangle frames in the first rectangle frame sequence, calculate the intersection area between the gap frames and other minimum bounding rectangles in the second target detection result, and add the rectangle frames of the intersection part to the supplementary content set when the intersection area is much larger than a preset threshold.

[0075] Specifically, for the first target detection result and the second text contour detection result corresponding to the target image, they have the same image size. For the first target detection result and the second text contour detection result, construct the same coordinate system and mark the positions of each minimum bounding rectangle. For example, the position can be marked by the two coordinates of the diagonal of the rectangle frame. For the first target detection result, that is, the target detection frame, denoted as D, all the rectangle frames can be sorted according to a preset rule, in the order from top to bottom and from left to right for the text frames, to obtain the corresponding first rectangle frame sequence. Similarly, for the second text contour detection result, that is, the contour detection frame, denoted as C, all the rectangle frames can be sorted in the order from top to bottom and from left to right for the text frames, to obtain the corresponding second rectangle frame sequence.

[0076] When performing rectangle comparison, first determine the first rectangle (the first D box) in the first target detection result, that is, the topmost D box. According to the position markings of each rectangle in the coordinate system, select all C boxes in the second text contour detection result that are above the position of the first D box. Mark the selected C boxes as processed and add them to the supplementary content set E. At the same time, determine the last rectangle in the first target detection result, that is, the bottommost D box, which is also the last rectangle in the first rectangle sequence. According to the position markings of the rectangles, select all C boxes in C that are below the position of the last D box, mark them as processed, and add them to the supplementary content set E.

[0077] After that, for the first rectangle sequence, select the 1st and 2nd boxes. Mark the coordinates of the upper left corner and the lower right corner of the first rectangle as (b1_x1, b1_y1) and (b1_x2, b1_y2) respectively; mark the diagonal coordinates of the second rectangle as (b2_x1, b2_y1) and (b2_x2, b2_y2) respectively. Configure the gap threshold as T1. For example, the gap threshold can be configured according to the font size in the document. For example, half of the font size is configured as the value of T1. Calculate the gap size between the adjacent first rectangle and the second rectangle. For example, calculate based on the lower right corner coordinate of the first rectangle and the upper left corner coordinate of the second rectangle. If (b2_y1 - b1_y2) is less than T1, it means the gap between the two rectangles is small, and the current two rectangles can be skipped; or, if (b2_y1 - b1_y2) is greater than or equal to T1, it indicates that the gap between the two rectangles is large, and the size of the gap box and the coordinates of relevant points can be configured according to the coordinates of the two rectangles. Specifically, the coordinates of the upper left corner and the lower right corner of the gap box space can be configured as (min(b1_x1, b2_x1), b1_y2) and (max(b1_x2, b2_x2), b2_y1) respectively. Here, min is the operation of taking the minimum value, and max is the operation of taking the maximum value. Traverse the unprocessed rectangles in C, calculate the intersection area s between each unprocessed rectangle and the gap box respectively, and compare it with the preset area threshold T2. If the intersection area is greater than T2, mark the corresponding C box as processed and add the intersection part between the C box and the gap box to the set E; or, if the intersection area is less than or equal to T2, ignore it. Repeat the above process, successively select adjacent two rectangles in the first rectangle sequence, such as selecting the second and third rectangles, the third and fourth rectangles, and repeat the process of calculating the gap between the rectangles, calculating the gap box, and calculating and comparing the intersection area between the gap box and the unprocessed boxes in C until all adjacent rectangles in D are traversed to construct the final supplementary content set E.

[0078] In step S15, the layout analysis result corresponding to the target image is determined according to the supplementary content set and the first target detection result.

[0079] In this exemplary embodiment, after obtaining the complete supplementary content set, the first target detection result can be merged with the supplementary content set, and the union is taken as the final layout analysis result. Thus, the layout analysis result is supplemented by using the contour detection result.

[0080] In addition, in some exemplary embodiments, the above method may further include: performing table recognition on the target image and adding the table recognition result to the layout analysis result.

[0081] Specifically, after obtaining the target image, an independent table recognition model can be used to perform table recognition on the target image and generate a corresponding table recognition result. After performing layout analysis on the template image and obtaining the first target detection result, the table recognition result can be cross-compared and fused with the first target detection result first to generate a first target detection result including table recognition labels. That is, in the first target detection result, the recognition results and corresponding labels of tables, different-level headings, and paragraphs can be included.

[0082] Based on the above, in some exemplary embodiments, the above method may further include: obtaining the text recognition result of the target image; generating the document recognition result corresponding to the target according to the text recognition result in combination with the layout analysis result.

[0083] Specifically, for the target image corresponding to the document to be processed, OCR character recognition can be performed synchronously to obtain the text content corresponding to the target image. Then, the text content recognition result, the first target detection result, and the layout analysis result are combined to obtain the document recognition result including the layout analysis result corresponding to the document to be processed. The automatic conversion from an image to an electronic document is realized.

[0084] The layout analysis method provided by the embodiments of the present disclosure refers to Figure 4As shown, the layout detection 402 is combined with the contour detection 403. The layout detection model based on the object detection algorithm is used to obtain the layout detection result, and the accurate position information and category information of each rectangular box are determined. The layout detection result and the contour detection result are fused through the fusion model 404 to obtain a supplementary content set. Using the accurate position information of each rectangular box in the supplementary content set, the first object detection result output by the object detection algorithm is fused and supplemented to obtain the final layout analysis result 405, so as to improve the accuracy as much as possible on the premise of ensuring the integrity of the result in the image-to-document task, that is, all texts can be detected. That is, the semantic information output by the object detection model can be considered very accurate, which can effectively improve the user experience. By regarding the layout analysis task as an object detection task, while obtaining high-accuracy semantic labels by applying the existing object detection model, the defect of high missed detection rate of the object detection algorithm is solved by supplementing the contour detection algorithm and designing the corresponding fusion technology, which can effectively improve the algorithm performance. Moreover, the two parts of the model are relatively independent, with low coupling, ensuring the stability of the algorithm.

[0085] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0086] Further, referring to Figure 5 As shown, in the embodiment of this example, a layout analysis device 50 is further provided, including: a target image acquisition module 501, a layout analysis module 502, a contour analysis module 503, a fusion processing module 504, and a fusion processing module 505. Among them,

[0087] The target image acquisition module 501 can be used to acquire the target image corresponding to the document to be processed.

[0088] The layout analysis module 502 can be used to perform layout analysis on the target image to obtain a first object detection result; among them, the first object detection result includes a plurality of minimum bounding rectangles and corresponding marking information.

[0089] The contour analysis module 503 can be used to perform contour detection on the target image to obtain a second text contour detection result; among them, the second text contour detection result includes a plurality of minimum bounding rectangles and corresponding marking information.

[0090] The fusion processing module 504 can be used to fuse and compare the first object detection result and the second text contour detection result to obtain a supplementary content set.

[0091] The analysis result output module 505 can be used to determine the layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result.

[0092] In an example of the present disclosure, the layout analysis module 502 may include: extracting features from the target image to obtain feature information; calculating the feature information using a position regression model based on an object detection algorithm to obtain a rectangular box recognition result; and calculating the feature information using a label classification model based on an object detection algorithm to obtain a classification recognition result of the rectangular box.

[0093] In an example of the present disclosure, the contour analysis module 503 may include: converting the target image to obtain a corresponding grayscale image; performing Gaussian blur processing on the grayscale image to obtain a blurred image; performing binarization processing on the blurred image to obtain a binary image; performing morphological dilation operation on the binary image to obtain a dilation result; performing contour detection on the dilation result to obtain the minimum circumscribed rectangular boxes of the respective contours, and configuring corresponding marking information for each minimum circumscribed rectangular box.

[0094] In an example of the present disclosure, the fusion processing module 504 may include: constructing a coordinate system for the target image, marking the positions of the minimum circumscribed rectangular boxes in the first target detection result, and sorting them according to a preset rule to obtain a first rectangular box sequence; and marking the positions of the minimum circumscribed rectangular boxes in the second target detection result, and sorting them according to a preset rule to obtain a second rectangular box sequence; based on the position marking of the first rectangular box in the first rectangular box sequence, selecting the minimum circumscribed rectangular box in the second target detection result that is located above the position of the first rectangular box, and adding it to the supplementary content set; based on the position marking of the last rectangular box in the first rectangular box sequence, selecting the minimum circumscribed rectangular box in the second target detection result that is located below the position of the last rectangular box, and adding it to the supplementary content set; and obtaining the gap boxes between two adjacent rectangular boxes in the first rectangular box sequence, calculating the intersection area between the gap boxes and the other minimum circumscribed rectangular boxes in the second target detection result, and adding the rectangular boxes of the intersection part to the supplementary content set when the intersection area is much larger than a preset threshold.

[0095] In an example of the present disclosure, the fusion processing module 504 may further include: selecting two adjacent rectangular frames in the first rectangular frame sequence to form a group to be recognized, calculating a gap value based on the second coordinate of the first rectangular frame and the third coordinate of the second rectangular frame in the group to be recognized; when the gap value is less than a preset gap threshold, determining that there is no gap between the two rectangular frames in the group to be recognized; or, when the gap value is greater than the preset gap threshold, determining that there is a gap between the two rectangular frames in the group to be recognized, and determining the diagonal coordinates of the gap frame based on the diagonal coordinates of the first rectangular frame and the second rectangular frame.

[0096] In an example of the present disclosure, the layout analysis device 50 may further include: a document recognition result output module.

[0097] The document recognition result output module may be configured to obtain the text recognition result of the target image; generate a document recognition result corresponding to the target image according to the text recognition result in combination with the layout analysis result.

[0098] In an example of the present disclosure, the layout analysis device 50 may further include: a table recognition module.

[0099] The table recognition module may be configured to perform table recognition on the target image and add the table recognition result to the layout analysis result.

[0100] The specific details of each module in the above layout analysis device have been described in detail in the corresponding layout analysis method, and thus will not be elaborated here.

[0101] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.

[0102] Figure 6 The schematic diagram of an electronic device suitable for implementing the embodiments of the present invention is shown.

[0103] It should be noted that Figure 6 The electronic device 600 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0104] Such as Figure 6As shown, the electronic device 600 includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 402 or the program loaded from the storage section 908 into the Random Access Memory (RAM) 603. In the RAM 603, various programs and data required for system operations are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The Input / Output (I / O) interface 605 is also connected to the bus 604.

[0105] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.

[0106] Specifically, according to an embodiment of the present invention, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the Central Processing Unit (CPU) 601, various functions defined in the system of the present application are executed.

[0107] Specifically, the above-mentioned electronic device can be a smart mobile terminal device such as a mobile phone, a tablet computer, or a laptop computer. Alternatively, the above-mentioned electronic device can also be a smart terminal device such as a desktop computer.

[0108] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0110] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation on the units themselves in some cases.

[0111] It should be noted that, on the other hand, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. When the one or more programs carried by the above computer-readable medium are executed by an electronic device, the electronic device is caused to implement the methods described in the following embodiments. For example, the electronic device may implement the Figure 1 or Figure 2 steps shown.

[0112] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0113] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0114] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0115] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A layout analysis method, characterized in that, Including: Obtaining a target image corresponding to a document to be processed; Performing layout analysis on the target image to obtain a first target detection result; wherein, the first target detection result includes a plurality of minimum circumscribed rectangle frames and corresponding marking information; and Performing contour detection on the target image to obtain a second text contour detection result; wherein, the second text contour detection result includes a plurality of minimum circumscribed rectangle frames and corresponding marking information; Fusing and comparing the first target detection result and the second text contour detection result to obtain a supplementary content set; Determining a layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result; Wherein, the fusing and comparing the first target detection result and the second text contour detection result to obtain a supplementary content set includes: constructing a coordinate system for the target image, marking the positions of the minimum circumscribed rectangle frames in the first target detection result according to the coordinate system, and sorting them according to a preset rule to obtain a first rectangle frame sequence; and, marking the positions of the minimum circumscribed rectangle frames in the second text contour detection result and sorting them according to a preset rule to obtain a second rectangle frame sequence; based on the position marking of the first rectangle frame in the first rectangle frame sequence, selecting the minimum circumscribed rectangle frame in the second text contour detection result that is above the position of the first rectangle frame and adding it to the supplementary content set; based on the position marking of the last rectangle frame in the first rectangle frame sequence, selecting the minimum circumscribed rectangle frame in the second text contour detection result that is below the position of the last rectangle frame and adding it to the supplementary content set; obtaining the gap frames between two adjacent rectangle frames in the first rectangle frame sequence, calculating the intersection area between the gap frames and other minimum circumscribed rectangle frames in the second text contour detection result, and adding the rectangle frames of the intersection part to the supplementary content set when the intersection area is much larger than a preset threshold.

2. The layout analysis method according to claim 1, wherein The performing layout analysis on the target image to obtain a first target detection result includes: Performing feature extraction on the target image to obtain feature information; Calculating the feature information by using a position regression model based on a target detection algorithm to obtain a rectangle frame recognition result; and Calculating the feature information by using a label classification model based on a target detection algorithm to obtain a classification recognition result of the rectangle frame.

3. The layout analysis method according to claim 1, characterized in that The performing contour detection on the target image to obtain a second text contour detection result includes: Converting the target image to obtain a corresponding grayscale image; Performing Gaussian blur processing on the grayscale image to obtain a blurred image; Performing binarization processing on the blurred image to obtain a binary image; Performing a morphological dilation operation on the binary image to obtain a dilation result; Performing contour detection on the dilation result to obtain the minimum circumscribed rectangle frames of each contour and configuring corresponding marking information for each minimum circumscribed rectangle frame.

4. The layout analysis method according to claim 1, wherein The obtaining the gap frames between two adjacent rectangle frames in the first rectangle frame sequence includes: Select two adjacent rectangular frames from the first sequence of rectangular frames to form a group to be recognized, and calculate the gap value according to the second coordinate of the first rectangular frame and the third coordinate of the second rectangular frame in the group to be recognized; When the gap value is less than the preset gap threshold, it is determined that there is no gap between the two rectangular frames in the group to be recognized; or, When the gap value is greater than the preset gap threshold, it is determined that there is a gap between the two rectangular frames in the group to be recognized, and the diagonal coordinates of the gap frame are determined according to the diagonal coordinates of the first rectangular frame and the second rectangular frame.

5. The layout analysis method according to claim 1, wherein The method further includes: Obtain the text recognition result of the target image; Generate the document recognition result corresponding to the target image according to the text recognition result in combination with the layout analysis result.

6. The layout analysis method according to claim 1 or 5, characterized in that The method further includes: Perform table recognition on the target image, and add the table recognition result to the layout analysis result.

7. A layout analysis device, characterized in that, Including: A target image acquisition module, configured to acquire a target image corresponding to a document to be processed; A layout analysis module, configured to perform layout analysis on the target image to obtain a first target detection result; wherein, the first target detection result includes a plurality of minimum circumscribed rectangular frames and corresponding marking information; and A contour analysis module, configured to perform contour detection on the target image to obtain a second text contour detection result; wherein, the second text contour detection result includes a plurality of minimum circumscribed rectangular frames and corresponding marking information; A fusion processing module, configured to fuse and compare the first target detection result and the second text contour detection result to obtain a supplementary content set; An analysis result output module, configured to determine the layout analysis result corresponding to the target image according to the supplementary content set and the first target detection result; Among them, the fusing and comparing the first target detection result and the second text contour detection result to obtain a supplementary content set includes: constructing a coordinate system for the target image, marking the positions of the minimum circumscribed rectangular frames in the first target detection result according to the coordinate system, and sorting them according to a preset rule to obtain a first sequence of rectangular frames; and, marking the positions of the minimum circumscribed rectangular frames in the second text contour detection result and sorting them according to a preset rule to obtain a second sequence of rectangular frames; based on the position marking of the first rectangular frame in the first sequence of rectangular frames, select the minimum circumscribed rectangular frame in the second text contour detection result that is above the position of the first rectangular frame and add it to the supplementary content set; based on the position marking of the last rectangular frame in the first sequence of rectangular frames, select the minimum circumscribed rectangular frame in the second text contour detection result that is below the position of the last rectangular frame and add it to the supplementary content set; obtain the gap frames between two adjacent rectangular frames in the first sequence of rectangular frames, calculate the intersection area between the gap frames and other minimum circumscribed rectangular frames in the second text contour detection result, and add the rectangular frames of the intersection part to the supplementary content set when the intersection area is much larger than the preset threshold.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the layout analysis method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the layout analysis method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Character detection method and device, apparatus and computer readable storage medium

    CN110097046A

  • Text information structured extraction method and device

    CN112733639A