Method and device for identifying illustration area in document page

By parsing the layout document pages and detecting correlations, identifying and processing the illustration areas, the problems of high resource consumption and high false detection rate in the existing technology are solved, and efficient and low-consumption text recognition is achieved.

CN120635930APending Publication Date: 2025-09-12ZHUHAI KINGSOFT OFFICE SOFTWARE +3
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410273981.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing technology consumes a lot of resources and has a high false positive rate when identifying text-surrounding illustrations in electronic documents, and is prone to losing information such as the writing order of the original text objects.

Method used

By parsing the layout document page, identifying the initial text and non-text objects, determining the area to be identified and detecting the correlation between the target text objects, shielding or deleting the illustration area, and utilizing the object characteristics of the document page for identification rather than large-scale image processing.

Benefits of technology

It reduces resource consumption and false detection rate, retains the native object information of the document page, and improves the accuracy and efficiency of text recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635930A_ABST
    Figure CN120635930A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for identifying an illustration area in a document page, and the method comprises the steps: carrying out the determination of an area to be identified and a detection area and the searching of a target text object through an initial text object and a non-text object which are analyzed from document pages such as a PDF (Portable Document Format) page; according to the embodiment of the invention, the target text object is extracted, and whether the target text object has the correlation is detected, so that the region belonging to the illustration region in the document page can be recognized, and the interference of the illustration region on text recognition can be reduced when the document page is subjected to text recognition; in the identification process of the illustration area, the original data characteristics of the object in the document page are utilized, and the process of firstly converting the whole document page into image data and then carrying out large-scale image processing is not needed, so that the resource consumption level and the false drop rate level are relatively low; and information such as the original object writing sequence in the document page can be reserved. The method is widely applied to the technical field of computer document processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer document processing, and in particular to a method and device for identifying an illustration area in a document page. Background Art

[0002] Electronic documents often contain information in various forms, including text and images. When performing effective information extraction, optical character recognition (OCR), and typesetting on electronic documents, it is often necessary to identify, select, copy, and edit text content within the electronic document page. Taking PDF, a typical electronic document format, as an example, due to its flexible integration, when non-text content such as images is present on a PDF page, the text content is often displayed very close to the image, disrupting the integrity of the text content. This makes it difficult to identify, select, copy, and edit the text content. This difficulty often occurs when text wraps around illustrations.

[0003] To address the difficulty of identifying text content in situations like text-wrapped illustrations, current technologies perform operations such as text pixel clustering and fuzzy pixel clustering based on the resulting image of the document page printout. This generates text and non-text layers, and then performs operations such as pixel correlation clustering based on the texture features of the text and non-text layers to identify text content. Current technologies require computing a large number of texture features and performing image processing such as pixel correlation clustering. This leads to high resource consumption, a high false positive rate, and a tendency to lose information such as the object writing order of the original text in the original document. Summary of the Invention

[0004] In view of the current technical problems faced in recognizing text content in document pages, such as large resource consumption and high false detection rate, the purpose of the present invention is to provide a method and device for recognizing illustration areas in document pages.

[0005] In one aspect, an embodiment of the present invention includes a method for identifying an illustration region in a document page, the method comprising the following steps:

[0006] Parsing a document page of the layout document to obtain at least one initial text object and at least one non-text object;

[0007] Extracting a region of the document page according to at least one non-text object to determine a region to be identified;

[0008] Determining at least one detection area according to the area to be identified;

[0009] determining at least one target text object from the at least one initial text object according to the detection area, wherein a bounding box of the target text object intersects the detection area;

[0010] A correlation detection is performed on the target text objects, and when it is detected that the correlation exists between the target text objects corresponding to the area to be identified, the area to be identified is identified as an illustration area.

[0011] Furthermore, extracting a region of the document page according to at least one non-text object to determine a region to be identified includes:

[0012] Performing connectivity detection on the non-text objects to determine connected areas in the document page;

[0013] The connected area is used as the area to be identified.

[0014] Furthermore, the performing connectivity detection on the non-text objects to determine connected areas in the document page includes:

[0015] Create and initialize a binary memory bitmap;

[0016] Mapping the non-text object from the document page to the binary memory bitmap; bits in the mapped binary memory bitmap have corresponding mark masks, and the mark mask is a first value or a second value;

[0017] Mapping the bit in the binary memory bitmap corresponding to the mark mask having the first value to the document page;

[0018] The mapped portion of the document page is used as the connected area.

[0019] Furthermore, determining at least one detection area based on the area to be identified includes:

[0020] Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified;

[0021] Determine a first dimension and a second dimension; the positive direction of the first dimension is the row direction of the document page, and the negative direction of the second dimension is the block direction of the document page;

[0022] Starting from the edge of the first bounding box and extending in the negative direction of the first dimension, a first detection area is determined;

[0023] Starting from the edge of the first bounding box and extending in the positive direction of the first dimension, a second detection area is determined;

[0024] Starting from the edge of the first bounding box and extending in the positive direction of the second dimension, a third detection area is determined;

[0025] Starting from the edge of the first bounding box and extending in the negative direction of the second dimension, a fourth detection area is determined.

[0026] Furthermore, the performing correlation detection on the target text object includes:

[0027] Classifying each target text object into a first text set, a second text set, a third text set, or a fourth text set; wherein the target text objects in the first text set intersect with the first detection area, the target text objects in the second text set intersect with the second detection area, the target text objects in the third text set intersect with the third detection area, and the target text objects in the fourth text set intersect with the fourth detection area;

[0028] Performing correlation detection on each text set combination respectively; each text set combination is a combination of the first text set and the third text set, a combination of the first text set and the fourth text set, a combination of the second text set and the third text set, and a combination of the second text set and the fourth text set;

[0029] In the case where the correlation is detected from any combination of the text sets, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized areas.

[0030] Furthermore, determining at least one detection area based on the area to be identified includes:

[0031] Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified;

[0032] Extending outward from each side and each vertex of the first bounding box to determine an extended area;

[0033] The epitaxial region is used as the fifth detection region.

[0034] Furthermore, the performing correlation detection on the target text object includes:

[0035] Sort each target text object according to the LRTB sequence corresponding to the document page;

[0036] Traversing the sorted target text objects in sequence, and performing correlation detection on at least two adjacent target text objects;

[0037] In the case where the correlation is detected from any at least two target text objects, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized area.

[0038] Furthermore, the correlation is a contextual semantic correlation, or a positional correlation; wherein the contextual semantic correlation represents the correlation between the contextual semantics of multiple target text objects, and the positional correlation represents the correlation between the positions of multiple target text objects in a document page.

[0039] Furthermore, after the step of identifying the area to be identified as an illustration area, the method for identifying an illustration area in a document page further includes:

[0040] Finding a target text object from the third text set and / or the fourth text set;

[0041] Performing content determination on the found target text object;

[0042] When it is determined that the found target text object constitutes a text illustration, the area within the bounding box corresponding to the found target text object is identified as the illustration area.

[0043] Furthermore, after the step of identifying the area to be identified as an illustration area, the method for identifying an illustration area in a document page further includes:

[0044] Finding a target graphic object in the document page; the target graphic object is a graphic object located in the illustration area;

[0045] The found target graphic object is deleted from the document page.

[0046] Furthermore, after the step of deleting the found target graphic object from the document page, the method for identifying an illustration area in a document page further includes:

[0047] Text recognition is performed on the document page.

[0048] On the other hand, an embodiment of the present invention further includes a device for identifying an illustration region in a document page, the device comprising:

[0049] A document parsing module, configured to parse a document page of a layout document to obtain at least one initial text object and at least one non-text object;

[0050] A region extraction module, configured to extract a region of the document page according to at least one non-text object and determine a region to be identified;

[0051] An area determination module, configured to determine at least one detection area based on the area to be identified;

[0052] an object determination module, configured to determine at least one target text object from the at least one initial text object based on the detection area, wherein a bounding box of the target text object intersects the detection area;

[0053] The first region recognition module is configured to perform correlation detection on the target text objects, and identify the region to be recognized as an illustration region when the correlation is detected between the target text objects corresponding to the region to be recognized.

[0054] Furthermore, the region extraction module is specifically used to:

[0055] Performing connectivity detection on the non-text objects to determine connected areas in the document page;

[0056] The connected area is used as the area to be identified.

[0057] Furthermore, the region extraction module is specifically used to:

[0058] Create and initialize a binary memory bitmap;

[0059] Mapping the non-text object from the document page to the binary memory bitmap; bits in the mapped binary memory bitmap have corresponding mark masks, and the mark mask is a first value or a second value;

[0060] Mapping the bit in the binary memory bitmap corresponding to the mark mask having the first value to the document page;

[0061] The mapped portion of the document page is used as the connected area.

[0062] Furthermore, the region determination module is specifically configured to:

[0063] Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified;

[0064] Determine a first dimension and a second dimension; the positive direction of the first dimension is the row direction of the document page, and the negative direction of the second dimension is the block direction of the document page;

[0065] Starting from the edge of the first bounding box and extending in the negative direction of the first dimension, a first detection area is determined;

[0066] Starting from the edge of the first bounding box and extending in the positive direction of the first dimension, a second detection area is determined;

[0067] Starting from the edge of the first bounding box and extending in the positive direction of the second dimension, a third detection area is determined;

[0068] Starting from the edge of the first bounding box and extending in the negative direction of the second dimension, a fourth detection area is determined.

[0069] Furthermore, the first region identification module is specifically configured to:

[0070] Classifying each target text object into a first text set, a second text set, a third text set, or a fourth text set; wherein the target text objects in the first text set intersect with the first detection area, the target text objects in the second text set intersect with the second detection area, the target text objects in the third text set intersect with the third detection area, and the target text objects in the fourth text set intersect with the fourth detection area;

[0071] Performing correlation detection on each text set combination respectively; each text set combination is a combination of the first text set and the third text set, a combination of the first text set and the fourth text set, a combination of the second text set and the third text set, and a combination of the second text set and the fourth text set;

[0072] In the case where the correlation is detected from any combination of the text sets, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized areas.

[0073] Furthermore, the region determination module is specifically configured to:

[0074] Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified;

[0075] Extending outward from each side and each vertex of the first bounding box to determine an extended area;

[0076] The epitaxial region is used as the fifth detection region.

[0077] Furthermore, the first region identification module is specifically configured to:

[0078] Sort each target text object according to the LRTB sequence corresponding to the document page;

[0079] Traversing the sorted target text objects in sequence, and performing correlation detection on at least two adjacent target text objects;

[0080] In the case where the correlation is detected from any at least two target text objects, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized area.

[0081] Furthermore, the correlation is a contextual semantic correlation, or a positional correlation; wherein the contextual semantic correlation represents the correlation between the contextual semantics of multiple target text objects, and the positional correlation represents the correlation between the positions of multiple target text objects in a document page.

[0082] Furthermore, the illustration area recognition device further includes:

[0083] A first object search module is configured to search for a target text object from the third text set and / or the fourth text set after the area to be identified is identified as an illustration area;

[0084] A content determination module, configured to determine the content of the found target text object;

[0085] The second area recognition module is used to recognize the area within the boundary box corresponding to the found target text object as the illustration area when it is determined that the found target text object constitutes a text illustration.

[0086] Furthermore, the illustration area recognition device further includes:

[0087] A second object search module is configured to search for a target graphic object in the document page after identifying the area to be identified as an illustration area; the target graphic object is a graphic object located in the illustration area;

[0088] The object deletion module is used to delete the found target graphic object from the document page.

[0089] Furthermore, the illustration area recognition device further includes:

[0090] The text recognition module is configured to perform text recognition on the document page after deleting the found target graphic object from the document page.

[0091] The beneficial effects of the present invention are as follows: the method and device for identifying illustration areas in document pages in the embodiment, by determining the areas to be identified and the detection areas and searching for target text objects for the initial text objects and non-text objects parsed from the document pages of the layout document, and then detecting whether there is a correlation between the target text objects, can identify the areas in the document page that belong to the illustration areas, which helps to take operations such as shielding or deleting the illustration areas when performing text recognition on the document page, thereby reducing the interference of the illustration areas on text recognition; since the identification process of the illustration areas utilizes the original data characteristics of the objects in the document page, there is no need to first convert the entire document page into image data and then perform a large-scale image processing process, thereby having a lower resource consumption level and a lower false detection rate level, and can retain information such as the original object writing order in the document page. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 Schematic diagram of the steps of the method for identifying an illustration area in a document page in an embodiment;

[0093] Figure 2 Schematic diagram of a PDF file and a PDF page in an embodiment;

[0094] Figure 3 is a schematic diagram of a graphic object in a PDF page in an embodiment;

[0095] Figure 4 is a schematic diagram of mapping graphic objects in an embodiment;

[0096] Figure 5 is a schematic diagram of a bounding box and its coordinates corresponding to a graphic object in an embodiment;

[0097] Figure 6 Schematic diagram of the relationship between the display effect of a PDF page and a graphic object in an embodiment;

[0098] Figure 7 Schematic diagram of the principle of extracting regions from a PDF page based on non-text objects and determining at least one region to be identified in an embodiment;

[0099] Figure 8 Schematic diagram of the principle of merging multiple initial text objects in an embodiment;

[0100] Figure 9 2 is a schematic diagram showing a first execution mode of the step of determining at least one detection area according to the area to be identified in the embodiment;

[0101] Figure 10 2 is a schematic diagram showing the principle of a second execution method of the step of determining at least one detection area according to the area to be identified in the embodiment;

[0102] Figure 11 2 is a schematic diagram of the principle of the step of determining at least one target text object from all initial text objects according to the detection area in the embodiment;

[0103] Figure 12 Schematic diagram of the principle of performing position correlation detection on target text objects in an embodiment;

[0104] Figure 13 Schematic diagram of the structure of an apparatus for identifying illustration areas in a document page in an embodiment;

[0105] Figure 14 is a schematic structural diagram of a computer device in an embodiment;

[0106] Figure 15 Schematic diagram of the working principle of the computer-readable storage medium in the embodiment. DETAILED DESCRIPTION

[0107] A layout document is an electronic document format with a fixed layout presentation effect. PDF documents and OFD (Open Fixed-layout Document) are both layout documents. In the embodiments of the present application, PDF is used as an example of an electronic document, that is, a PDF page is used as an example of a document page to illustrate the method for identifying illustration areas in a document page. It is understandable that the illustration area identification method of the present application can also be applied to other document formats similar to PDF, such as OFD format, CAJ format, XML format, HTML format, etc., and only adaptive modifications are required.

[0108] Among them, PDF is the abbreviation of Portable Document Format in English, and its Chinese name is "Portable File Format". It is a cross-operating system platform file format that can encapsulate text, fonts, graphics, images, colors, layouts and parameters related to printing equipment in one file, keeping the page elements unchanged during network transmission, printing and plate-making output, and can also include hypertext links, audio, video and other electronic information.

[0109] In this embodiment, refer to Figure 1 ,The method for recognizing illustration regions in a document page includes the following steps:

[0110] S1. Parsing the document page of the layout document to obtain at least one initial text object and at least one non-text object;

[0111] S2. Extract regions from the document page based on non-text objects to determine the regions to be identified;

[0112] S3. Determine at least one detection area based on the area to be identified;

[0113] S4. Determining at least one target text object from at least one initial text object according to the detection area, wherein the bounding box of the target text object intersects with the detection area;

[0114] S5. Perform correlation detection on the target text objects. When a correlation is detected between target text objects corresponding to the area to be identified, identify the area to be identified as an illustration area.

[0115] Steps S1-S5 can be performed by a computer. For example, computer programming techniques can be used to generate computer code, which can be input into a computer device, causing the computer device to execute steps S1-S5. The computer code can be integrated into application software such as a document reader or document editor, and steps S1-S5 can be executed when the computer device runs the application software such as a document reader or document editor. This allows the illustration region recognition effect achieved by the computer device executing steps S1-S5 to become a function of the application software such as a document reader or document editor.

[0116] The computer device can download the PDF data to be processed through the Internet, or connect to a mobile device such as a mobile phone, tablet computer, or removable hard disk to read the PDF data to be processed from the mobile device, thereby reading PDF data such as PDF files and PDF pages for the computer device to execute steps S1-S5.

[0117] Before executing steps S1-S4, you can refer to Figure 2 , obtain the PDF file to be processed (i.e., layout document), the PDF file contains multiple PDF pages (i.e., document pages of the layout document), and extract a single PDF page from the PDF file. Specifically, the file directory (Document Catalog) of the PDF file can be obtained, the page tree (PageTree) can be determined based on the file directory, and the PDF page can be found based on the page tree. Since the illustration area recognition method in the document page can use the PDF page in the PDF file as the processing unit when identifying the illustration area of ​​the PDF file, that is, when processing multiple PDF pages in the PDF file, the same processing process is performed on each of the PDF pages, so the processing of one of the PDF pages is taken as an example for explanation. In this embodiment, when referring to "PDF page", unless otherwise specified, it can refer to a specific PDF page in a PDF file.

[0118] When executing step S1, all graphics objects contained in the PDF page can be obtained from the content stream data of the PDF page. Each graphics object can be specifically classified as any one of a text object, a path object, an external object (XObject), an inline image object, or a shading object. Among the various graphic objects mentioned above, path objects represent graphics through elements such as straight lines, rectangular boxes, and cubic Bezier curves, and their combinations. Complex shapes can be drawn by filling, drawing lines / strokes, and filling and stroking (both) these elements. Text objects consist of one or more characters, and text can be drawn by stroking and filling the characters in the text objects. External objects are defined outside the content stream. As resource-defined objects, external objects can reference resources outside the content stream. External objects specifically include image external objects (Image XObject), form external objects (Form XObject), reference external objects (Reference XObject), and group external objects (Group XObject). Built-in image objects use specific syntax to display data as a small image. Compared with image external objects, built-in image objects do not need to reference external resources. Shading objects can describe a geometric shape and calculate the color of each position in the geometric shape.

[0119] Specifically, a PDF page can be output to an output device such as a display or printer. For example, each graphic object in a PDF page is rendered with a corresponding visual effect. Text objects are rendered as text, while other graphic objects, such as path objects, are rendered as images. Therefore, in this embodiment, object types other than text objects, such as path objects, external objects (XObjects), inline image objects, and shading objects, are referred to as non-text objects.

[0120] In this embodiment, according to the technical specifications of PDF files, each graphic object (specifically a text object or a non-text object) in a PDF page corresponds to a bounding box in the form of a rectangular area (Rectangle). This bounding box can determine the position of the graphic object in the PDF page. This bounding box can be represented by a set of [ll x ,ll y ,ur x ,ur y ] data representation. For example, refer to Figure 3 , PDF page data through a set of data [ll x ,ll y ,ur x ,ur y ] to describe a graphic object. This set of data can represent a bounding box in the form of a rectangular area on the PDF page, that is, the graphic object is located in Figure 3 Specifically, the data in this set are within the bounding box shown in the dotted box. x Indicates the horizontal coordinate of the lower-left corner of the bounding box, ll y Indicates the vertical coordinate of the lower left corner of the bounding box, ur x Indicates the horizontal coordinate of the upper-right corner of the bounding box, ur y Indicates the ordinate of the upper-right corner of the bounding box.

[0121] according to Figure 3 The characteristics of the PDF page shown can be referred to Figure 4 , mapping each graphic object into a plane space. Figure 4 As shown, the plane space can be represented by the X-axis (first dimension) and the Y-axis (second dimension). Each graphic object has corresponding coordinates in the plane space, where the coordinates on the X-axis (first dimension) are represented as value values ​​and the coordinates on the Y-axis (second dimension) are represented as key values. Specifically, the scale of the plane space can be the same as that of the PDF page, so that the bounding box coordinate data of a graphic object contained in the PDF page data can be directly converted to the

[11] x ,ll y ,ur x ,ur y ], as the coordinates of this graphic object in the plane space [value1, key1, value2, key2], that is, value1 = ll x ,value2=ur x , key1 = ll y , key2 = ur y .

[0122] In this embodiment, the directions of the X-axis (first dimension) and the Y-axis (second dimension) can be set according to the reading order of the PDF page. Considering that LRTB (Left Right Top Bottom, i.e., from left to right and from top to bottom) is a more commonly used reading order, the directions of the X-axis (first dimension) and the Y-axis (second dimension) can be set according to the reading order of LRTB. When the reading order of the PDF page is LRTB, the arrangement direction of the characters in the same text object in the PDF page is from left to right, that is, the line movement direction of the PDF page is from left to right; when the characters in the text object are arranged along a line to the end and need to wrap to form a new text object (text block), the new text object is below the original text object, that is, the block movement direction of the PDF page is from top to bottom.

[0123] For the convenience of explanation, the first dimension and the second dimension can be set according to the userspace coordinate space in the PDF technical standard. Figure 4 , set the X-axis (first dimension) as the left-right dimension, and set the Y-axis (second dimension) as the up-down dimension. Specifically, set the positive direction of the X-axis (first dimension) to be the same as the row direction of the PDF page, that is, the positive direction of the X-axis (first dimension) is from left to right; set the negative direction of the Y-axis (second dimension) to be the same as the block direction of the PDF page, that is, the positive direction of the Y-axis (second dimension) is opposite to the block direction of the PDF page, that is, the positive direction of the Y-axis (second dimension) is from bottom to top. The direction settings of the first and second dimensions do not affect the implementation of the illustration area recognition method.

[0124] according to Figure 4 According to the principle shown, when a PDF page contains multiple graphic objects (specifically, they can be text objects or non-text objects), each graphic object can be mapped into a plane space, thereby determining a corresponding rectangular area.

[0125] It should be noted that Figure 4 and Figure 5 The “mapping” shown can be a mathematical correspondence, and for Figure 4 and Figure 5 The "mapping" of the graphic objects shown in the figure from the PDF page to the plane space, since the PDF page itself contains the corresponding [ll x ,ll y ,ur x ,ur y ], which is equivalent to a rectangular area in the plane space. Therefore, when executing steps S1-S5, there is no need to perform mapping operations on the graphics level, but you can directly use the [llx ,ll y ,ur x ,ur y ] coordinate data as the value and key value in the plane space, that is, to complete the "mapping" operation; or the [ll x ,ll y ,ur x ,ur y ] coordinate data is converted to obtain the value and key value of each graphic object in the plane space, that is, to complete the "mapping" operation.

[0126] By executing step S1, referring to Figure 5 , each graphics object in the PDF page has a corresponding bounding box, and each bounding box has a corresponding set of value and key values.

[0127] Next, we will take a PDF page with specific content as an example. Figure 6 , where the upper left corner is the visual effect after the PDF page is output (for example, displayed after being rendered on a monitor, or printed on paper or other media through a printer), including the main text that reads "Long-term exposure to the computer may cause damage such as vision loss, dry eyes, cervical spondylosis, and periarthritis of the shoulder. Because computer screens can irritate the eyes, long-term exposure to blue light may cause visual fatigue. If the discomfort symptoms are not relieved in time, further aggravation may affect vision and even lead to vision loss.", an image of a person facing a computer screen, and a text illustration that reads "white-collar workers facing the screen for a long time", where the text illustration explains the content of the image.

[0128] Reference Figure 6 The arrow between the upper left corner and the upper right corner represents the intuitive display of each graphic object obtained by parsing the PDF page. Figure 6 In the upper right corner of the figure, each dotted box represents a corresponding graphic object. For example, the words "facing the computer for a long time may" correspond to a text object, the words "having decreased vision, dry eyes" correspond to a text object, and the computer screen part of the image "a person facing the computer screen" is a non-text object (specifically, it can be an Inline Image Object), and the person part of the image "a person facing the computer screen" is another non-text object (specifically, it can be an Inline ImageObject).

[0129] Reference Figure 6The arrow between the upper right corner and the lower right corner indicates that multiple text objects and multiple non-text objects can be obtained by executing step S1. In this embodiment, the text objects obtained by executing step S1 are not processed and are called "initial text objects". Therefore, executing step S1 can obtain Figure 6 The initial text objects shown in the lower right corner of the PDF include initial text object 1, initial text object 2, initial text object 13, and non-text objects such as non-text object 1 and non-text object 2. Each graphic object corresponds to a specific value and key in plane space. Based on the value and key of a graphic object, the rectangular area of ​​the graphic object in the PDF page can be mapped. Therefore, in this embodiment, it is not necessary to distinguish between the rectangular area of ​​a graphic object in the PDF page, the bounding box in the PDF page, the coordinates of the diagonal points of the bounding box in the PDF page, and the value and key in plane space.

[0130] In this embodiment, Figure 6 The dotted box in the lower right corner of the figure where each graphic object is located can be regarded as an intuitive representation of the bounding box corresponding to the graphic object, so as to facilitate the understanding of the correspondence between a graphic object, its corresponding bounding box and the rectangular area defined by it. In the process of processing the graphic object and its corresponding bounding box, the computer does not need to perform graphic-level processing on the bounding box. For example, it does not need to render, display or print the bounding box corresponding to the graphic object in the form of a dotted box or other forms. It only needs to calculate the bounding box of the graphic object according to the relevant data of the bounding box (including the graphic object in the Figure 6 The corresponding value and key value in the plane space shown can be processed.

[0131] In this embodiment, focus on Figure 6 The lower right corner of Figure 8 As shown, it is assumed that the coordinate on the X-axis (first dimension) corresponding to the dotted box of the initial text object 1 is value text11 and value text12 , the coordinate on the X-axis (first dimension) corresponding to the dotted box of the initial text object 2 is value text21 and value text22 , define the distance between the dotted box of initial text object 1 and the dotted box of initial text object 2 on the X axis (first dimension) as gap, then gap = |value text12 -value text21 In this embodiment, a threshold value can be set. gap ; When gap < threshold gapIn the case of , the initial text object 1 and the initial text object 2 can be merged into the same initial text object, and marked as initial text object 1, initial text object 2 or other names, and the initial text object obtained after the merger is Figure 8 The corresponding bounding box in can be the union of the bounding box corresponding to the original initial text object 1, the bounding box corresponding to the initial text object 2, and the part between these two bounding boxes, that is, the initial text object obtained after the merger is in Figure 8 The coordinate range on the X-axis (first dimension) corresponding to the bounding box in the corresponding text11 ,value text22 ]; in gap ≥ threshold gap In the case of , the initial text object 1 and the initial text object 2 can still be maintained as two independent objects. Figure 8 The two corresponding bounding boxes in may not be merged.

[0132] Figure 8 The process of judging the size of the gap between the initial text object 1 and the initial text object 2 and determining whether to merge the two initial text objects based on the size judgment result can also be applied to any other two initial text objects. Figure 8 In the embodiment, corresponding judgments may also be performed on the two text objects, initial text object 11 and initial text object 12.

[0133] In this embodiment, Figure 8 The principle of the gap judgment process shown is as follows: for any two initial text objects, the corresponding gap value represents the distance between the bounding boxes corresponding to the two initial text objects, which is equivalent to the distance between the two text blocks corresponding to the two initial text objects after the PDF page is output by display or printing. gap In this case, it can be determined that the distance between the two text blocks corresponding to the two initial text objects is small, that is, they can be visually regarded as one text block, and the two initial text objects can be merged into one text object. The first advantage is that it does not change the text content contained in the PDF page. The second advantage is that the total number of initial text objects in a PDF page can be reduced, which is beneficial for processing the PDF page. The third advantage is that the two initial text objects with a small distance have a close positional relationship and usually have a close semantic association relationship. Merging the two initial text objects with a small distance into one text object is beneficial for the semantic algorithm to identify the semantic features of the text object when performing semantic analysis on the text object, which is beneficial for text recognition and other processing.

[0134] When executing step S2, that is, extracting the region of the PDF page according to the non-text object and determining at least one region to be identified, refer to Figure 6 Since each non-text object has a corresponding value and key value in the plane space, the rectangular area of ​​the non-text object in the PDF page can be mapped. Therefore, the rectangular area corresponding to the non-text object in the PDF page can be determined as the area to be identified.

[0135] Reference Figure 6 The PDF page contains two non-text objects, namely non-text object 1 and non-text object 2. Therefore, two rectangular areas can be determined in the PDF page based on the bounding boxes of non-text object 1 and non-text object 2. These two rectangular areas are both the areas to be identified obtained by executing step S2.

[0136] In this embodiment, when executing step S2, that is, extracting regions from the PDF page based on non-text objects and determining at least one region to be identified, the following steps may be specifically performed:

[0137] S201 performs connectivity detection on non-text objects to determine the connected areas in the PDF page;

[0138] S202. Use the connected area as the area to be identified.

[0139] by Figure 6 Taking the graphic object in the PDF page shown in the lower right corner of FIG as an example, steps S201-S202 are described. Figure 6 The PDF page contains non-text object 1 and non-text object 2. The principle of steps S201-S202 is as follows: Figure 7 shown.

[0140] Reference Figure 7 When executing step S201, the following steps may be performed:

[0141] S20101. First, create and initialize a binary memory bitmap. The binary memory bitmap has the same scale as the PDF page, so that each pixel on the PDF page corresponds to a bit in the binary memory bitmap. The flag mask of each bit in the initialized binary memory bitmap is the second value (e.g., 0).

[0142] S20102. Map all non-text objects (i.e., non-text object 1 and non-text object 2) from the PDF page to a binary memory bitmap, such that the mark mask of the bit corresponding to non-text object 1 or non-text object 2 in the binary memory bitmap becomes a first value (e.g., 1), while the mark mask of the bit not corresponding to any non-text object remains a second value (e.g., 0);

[0143] S20103. Map the bits in the binary memory bitmap corresponding to the first value (e.g., 1) of the mark mask back to the PDF page. The pixels mapped to the bits in the PDF page are connected, forming a connected area in the PDF page, i.e. Figure 7 The bold dashed box part in .

[0144] Steps S20101-S20103 actually map multiple non-text objects in the PDF page to a binary memory bitmap to achieve binarization of the non-text object content. After binarization, the pixel values ​​of each point in the multiple non-text objects are respectively a first value (e.g., 1) or a second value (e.g., 0). If there are continuous points (i.e., connected areas) in the multiple non-text objects after binarization that are all of the first value (e.g., 1), these points (i.e., connected areas) indicate that there is a content connection between the multiple non-text objects, i.e., the multiple non-text objects may correspond to the same illustration, thereby treating the connected area as the area to be identified, that is, treating the multiple non-text objects as one area to be identified. This shows that image binarization algorithms such as image segmentation can also be used instead of binary memory bitmaps to binarize the non-text objects in the PDF page.

[0145] Reference Figure 7 When executing step S202, the connected area in the PDF page determined by executing step S201 is used as the area to be identified.

[0146] In this embodiment, refer to Figure 7 It can be seen that the principle of executing steps S201-S202 is that the multiple non-text objects contained in the PDF page may eventually be rendered into different parts of the same illustration, and by performing connectivity detection on the multiple non-text objects, the detected connected areas represent the association relationship between the multiple non-text objects (that is, the multiple non-text objects are eventually rendered into different parts of the same illustration). By using the detected connected areas as the areas to be identified, compared with using the areas within the bounding boxes corresponding to each non-text object as the areas to be identified, the number and area of ​​the areas to be identified to be traversed in executing steps S1-S5 can be reduced, which is beneficial to improving the execution efficiency of the illustration area identification method in the document page.

[0147] After determining the area to be identified, steps S3-S5 are performed for each area to be identified, thereby determining whether each area to be identified is an illustration area. Figure 6In the figure, the rectangular area boundary box determined by the dotted box corresponding to the non-text object 1 in the PDF page is an area to be identified, and the rectangular area boundary box determined by the dotted box corresponding to the non-text object 2 in the PDF page is also an area to be identified. The principles of executing steps S3-S5 for these two areas to be identified are the same. Next, the execution of steps S3-S5 for the area to be identified corresponding to the non-text object 1 is used as an example for explanation. In this embodiment, unless otherwise specified, when a "area to be identified" is mentioned separately, such as "each target text object corresponding to the area to be identified", it refers to a specific area to be identified. In the case where there are multiple areas to be identified, different areas to be identified each correspond to some target text objects, and the target text objects of different areas to be identified may be completely the same or not completely the same.

[0148] In step S3, for the area to be recognized corresponding to the non-text object 1, at least one detection area is determined based on the area to be recognized, thereby facilitating subsequent determination of whether the target text object exists in the detection area. When at least one detection area is set based on the area to be recognized, a certain target position relationship is satisfied between the set detection area and the area to be recognized (for example, the detection area is located on a fixed side of the area to be recognized, or a vector from the center of the area to be recognized to the center of the detection area is a fixed vector, etc.), so that the target text object detected by the detection area satisfies a certain target position relationship with the area to be recognized corresponding to the non-text object 1.

[0149] In this embodiment, when executing step S3, that is, determining at least one detection area based on the area to be identified, the following steps may be specifically performed:

[0150] S301A. Detect a first bounding box;

[0151] S302A. Determine the first dimension and the second dimension;

[0152] S303A. Starting from the edge of the first bounding box, extending in the negative direction of the first dimension, determine the first detection area;

[0153] S304A. Starting from the edge of the first bounding box, extending in the positive direction of the first dimension, determine the second detection area;

[0154] S305A. Starting from the edge of the first bounding box, extending in the positive direction of the second dimension, determine a third detection area;

[0155] S306A. Starting from the edge of the first bounding box, extend in the negative direction of the second dimension to determine a fourth detection area.

[0156] Steps S301A-S306A are the first execution mode of step S3.

[0157] The principles of steps S301A-S306A are as follows: Figure 9 shown.

[0158] Reference Figure 9 In step S301A, after determining a region to be identified, the bounding box corresponding to the region to be identified is marked as a first bounding box.

[0159] In step S302A, the first dimension is Figure 9 The X-axis in the LRTB is in the positive direction from left to right, that is, the direction of row movement in the LRTB. The second dimension is Figure 9 The positive direction of the Y-axis is from bottom to top, that is, the opposite direction of the block moving direction in LRTB.

[0160] In step S303A, refer to Figure 9 , starting from the leftmost edge of the first bounding box, extend in the negative direction of the first dimension, that is, the X axis, by Δvalue, so as to determine the coordinates as [value other11 -Δvalue,value other11 ]×[key other11 ,key other12 ]’s first detection area.

[0161] In step S304A, refer to Figure 9 , starting from the rightmost edge of the first bounding box, extend in the positive direction of the first dimension, that is, the X axis, by Δvalue, so as to determine the coordinates as [value other12 ,value other12 +Δvalue]×[key other11 ,key other12 ]'s second detection area.

[0162] In step S305A, refer to Figure 9 , starting from the top edge of the first bounding box, extend in the negative direction of the second dimension, that is, the Y axis, by Δkey, so as to determine the coordinates as [value other11 ,value other12 ]×[key other11 ,key other11 -Δkey] third detection area.

[0163] In step S306A, refer to Figure 9 , starting from the bottom edge of the first bounding box, extend in the negative direction of the second dimension, that is, the Y axis, by Δkey, so as to determine the coordinates as [value other11 ,value other12 ]×[key other12,key other12 +Δkey] of the fourth detection area.

[0164] By executing steps S301A-S306A, multiple detection areas such as a first detection area, a second detection area, a third detection area, and a fourth detection area can be determined outside the first bounding box of the area to be identified. Figure 9 The first detection area, the second detection area, the third detection area and the fourth detection area are discrete from each other and are located at different positions outside the area to be identified, which is conducive to finding the initial text objects located at different positions near the periphery of the area to be identified as target text objects.

[0165] In this embodiment, the sizes of the first, second, third, and fourth detection areas can be controlled by controlling the sizes of the extensions Δvalue and Δkey, thereby controlling the number of target text objects that can be detected by these detection areas. Specifically, the size of the extension Δkey can be determined by calculating the average width of all bounding boxes in a PDF page; the size of the extension Δvalue can be determined by calculating the average length of all bounding boxes in a PDF page.

[0166] In this embodiment, based on the execution of steps S301A-S306A, when executing step S4, that is, determining at least one target text object from all the initial text objects based on the detection area, by comparing the coordinates of the bounding box of the initial text object with the coordinates of each detection area, it can be determined whether the bounding box of the initial text object intersects with each detection area. Figure 9 Taking initial text object 1 as an example, at least part of the bounding box of initial text object 1 overlaps with the third detection area, so the bounding box of initial text object 1 intersects with the third detection area. Similarly, the bounding boxes of initial text objects 2, 3, 4, 5, 6, 7, 8, and 9 intersect with the first detection area, the second detection area, the third detection area, and the fourth detection area, respectively. Therefore, initial text objects 1, 2, 3, 4, 5, 6, 7, 8, and 9 are all target text objects. However, initial text objects 10, 11, 12, and 13 are not target text objects.

[0167] In this embodiment, on the basis of executing steps S301A-S306A, when executing step S5, that is, performing correlation detection on the target text object, the following steps may be specifically performed:

[0168] S501. Classify each target text object into the first text set TKL, the second text set TKR, the third text set TVT or the fourth text set TVB;

[0169] S502. Perform correlation detection on each text set combination;

[0170] S503. When a correlation is detected from any text set combination, determine whether the correlation exists between target text objects corresponding to the to-be-recognized area.

[0171] Steps S501-S503 are based on Figure 9 The first detection area, the second detection area, the third detection area and the fourth detection area shown are designed.

[0172] When executing step S501, if a target text object intersects with the first detection area, then the target text object is placed in the first text set TKL; if a target text object intersects with the second detection area, then the target text object is placed in the second text set TKR; if a target text object only intersects with the third detection area, then the target text object is placed in the third text set TVT; if a target text object only intersects with the fourth detection area, then the target text object is placed in the fourth text set TVB.

[0173] Reference Figure 9 The first text set TKL includes the initial text object 3, the initial text object 5 and the initial text object 7, that is, the first text set TKL = {initial text object 3, initial text object 5, initial text object 7}; similarly, the second text set TKR = {initial text object 4, initial text object 6, initial text object 8}, the third text set TVT = {initial text object 1, initial text object 2}, and the fourth text set TVB = {initial text object 9}.

[0174] In step S502, correlation detection is performed on text combination sets such as "the first text set TKL and the third text set TVT", "the first text set TKL and the fourth text set TVB", "the second text set TKR and the third text set TVT", and "the second text set TKR and the fourth text set TVB".

[0175] Specifically, when executing step S502, it is possible to detect whether a contextual semantic relationship exists between two text sets within each text combination set, wherein the contextual semantic correlation relationship represents the correlation between the contextual semantics of multiple target text objects. For example, for the text combination set "a first text set TKL and a third text set TVT", it is possible to detect whether a contextual semantic relationship exists between the first text set TKL and the third text set TVT. Specifically, the entirety of the initial text objects within the first text set TKL, namely "initial text object 3, initial text object 5, initial text object 7", can be considered as a text segment, and the entirety of the initial text objects within the third text set TVT, namely "initial text object 1, initial text object 2", can be considered as another text segment. It is then possible to detect whether a contextual semantic relationship exists between the two text segments. For example, if the first text set TKL is the preceding text in a certain text, and the third text set TVT is the following text in the same text, then it can be determined that a contextual semantic relationship exists between the first text set TKL and the third text set TVT. By treating the first text set TKL and the third text set TVT as a whole text segment for contextual semantic relationship detection, it is beneficial to improve detection efficiency.

[0176] When executing step S502, in addition to performing contextual semantic relationship detection on the first text set TKL and the third text set TVT as a whole text segment, you can also select the whole composed of some initial text objects in the first text set TKL (for example, only select initial text object 5, or select initial text object 5 and initial text object 7) as a text segment, and select the whole composed of some initial text objects in the third text set TVT (for example, only select initial text object 1, or select initial text object 1 and initial text object 2 at the same time) as a text segment to detect whether there is a contextual semantic relationship between the two text segments; since there are multiple ways to select from the first text set TKL There are multiple ways to select text segments from the third text set TVT. Any text segment selected from the first text set TKL can be combined with any text segment selected from the third text set TVT to detect a corresponding context semantic relationship detection result. In this way, multiple context semantic relationship detection results will be obtained. When the context semantic relationship detection results of the first text set TKL and the third text set TVT are "context semantic relationship exists" when a threshold number (or a threshold ratio) of the context semantic relationship detection results are reached, it can be judged that there is a context semantic relationship between the first text set TKL and the third text set TVT. Otherwise, it can be judged that there is no context semantic relationship between the first text set TKL and the third text set TVT. By selecting text segments from the first text set TKL and the third text set TVT in a variety of ways to perform context semantic relationship detection, it is beneficial to reduce the interference of noise in the first text set TKL and the third text set TVT on the context semantic relationship detection and improve the detection accuracy.

[0177] If step S502 is executed, Figure 9 If any text set combination shown in FIG. 1 detects contextual semantics and other related relationships, then in step S503, it can be determined that Figure 9 There is a correlation between the target text objects corresponding to the area to be identified shown in FIG. Figure 9 The area to be identified is marked as the illustration area; on the contrary, if step S502 is executed, Figure 9 If no contextual semantics or other related relationships are detected for all text collection combinations shown, then Figure 9 The area to be identified is marked as the illustration area (or Figure 9 The area to be identified is marked as a non-illustration area).

[0178] In this embodiment, the principle of executing steps S301A-S306A, and steps S501-S503 executed based on steps S301A-S306A is that by setting the first detection area, the second detection area, the third detection area, and the fourth detection area, the target text objects at specific positions of the first text set TKL, the second text set TKR, the third text set TVT, and the fourth text set TVB can be determined, and combined with Figure 9 It can be seen that the possibility of the existence of contextual semantic relationships and other related relationships between target text objects at specific positions such as the first text set TKL, the second text set TKR, the third text set TVT and the fourth text set TVB is greater than the possibility of the existence of contextual semantic relationships and other related relationships between target text objects at other positions. Therefore, detecting these target text objects can reduce the number of target text objects that need to detect related relationships and improve the detection efficiency of the related relationships of the target text objects.

[0179] In this embodiment, when executing step S3, that is, determining at least one detection area based on the area to be identified, the following steps may be further performed:

[0180] S301B. Detect the first bounding box;

[0181] S302B extends outward from each edge and each corner of the first bounding box to determine the extension area;

[0182] S303B. Use the epitaxial region as the fifth detection region.

[0183] Steps S301B-S303B are a second execution method of step S3.

[0184] The principle of steps S301B-S303B is as follows Figure 10 shown.

[0185] Reference Figure 10 , similar to step S301A, in step S301B, after determining a region to be identified, the bounding box corresponding to the region to be identified is marked as a first bounding box.

[0186] In step S302B, refer to Figure 10 , starting from each edge and each vertex of the first bounding box, extend outward in all directions outside the first bounding box by a certain extension amount, where the extension amount in the first dimension (X axis) direction is Δvalue, and the extension amount in the second dimension (Y axis) direction is Δkey, thus forming Figure 10 The epitaxial region shown is used as the fifth detection region in step S303B.

[0187] In this embodiment, through steps S301B-S303B and Figure 10 The fifth detection area shown is used to detect each initial text object, referring to Figure 10 , where the bounding boxes corresponding to initial text object 1, initial text object 2, initial text object 3, initial text object 4, initial text object 5, initial text object 6, initial text object 7, initial text object 8, initial text object 9, etc. all intersect with the fifth detection area, then if Figure 11 As shown, initial text object 1, initial text object 2 ... initial text object 9 and other initial text objects are selected as target text objects, while the remaining initial text object 10, initial text object 11 ... initial text object 13 and other initial text objects are not selected as target text objects.

[0188] Specifically, when the correlation to be judged is the semantic correlation between the text contents of each target text object, the semantic correlation between each target text object can be judged individually, or the semantic correlation between the text paragraphs can be judged after the target text objects are combined into text paragraphs.

[0189] For example, Figure 11 The contents of each target text object determined in are:

[0190] Target text object (original initial text object 1): "Long-term use of computers may";

[0191] Target text object (original initial text object 2): "Has decreased vision, dry eyes", ";

[0192] Target text object (original initial text object 3): "cervical spondylosis and shoulder";

[0193] Target text object (original initial text object 4): "perioditis and other hazards. Cause";

[0194] Target text object (original initial text object 5): "for computer screen meeting";

[0195] Target text object (original initial text object 6): "irritating to the eyes";

[0196] Target text object (original initial text object 7): "excited, long-term blue light";

[0197] Target text object (original initial text object 8): "Irradiation may cause";

[0198] Target text object (original initial text object 9): "white-collar workers who face the screen for a long time."

[0199] Each target text object can be arranged according to the reading order of LRTB, that is, a sequence of target text object (original initial text object 1), target text object (original initial text object 2), target text object (original initial text object 3), target text object (original initial text object 4), target text object (original initial text object 5), target text object (original initial text object 6), target text object (original initial text object 7), target text object (original initial text object 8), and target text object (original initial text object 9) can be formed. For the above sequence, the computer can read two target text objects each time in sequence to perform word-granularity text matching, sentence-granularity text matching, or paragraph-granularity text matching. For example, first read the target text object (original initial text object 1) and the target text object (original initial text object 2), add the number of text characters of each of the two target text objects to obtain the total number of characters, and use the threshold comparison method to judge whether the total number of characters belongs to the level of "large", "medium" or "small". When the total number of characters is at the level of "small", the word granularity text matching algorithm is used to perform text matching on the target text object (original initial text object 1) and the target text object (original initial text object 2). When the total number of characters is at the level of "medium", the sentence granularity text matching algorithm is used to perform text matching on the target text object (original initial text object 1) and the target text object (original initial text object 2). If the number of characters is "large," the target text object (original initial text object 1) and the target text object (original initial text object 2) are matched using a paragraph-granularity text matching algorithm. All three text matching algorithms output corresponding feature similarities to represent the semantic relevance between the target text object (original initial text object 1) and the target text object (original initial text object 2). When the feature similarity is greater than the threshold, it can be determined that a semantic correlation exists between the target text object (original initial text object 1) and the target text object (original initial text object 2). Otherwise, it is determined that no semantic correlation exists between the target text object (original initial text object 1) and the target text object (original initial text object 2). After determining whether a semantic correlation exists between the target text object (original initial text object 1) and the target text object (original initial text object 2), between the target text object (original initial text object 3) and the target text object (original initial text object 4), and so on, it can be determined in sequence whether a semantic correlation exists between the target text object (original initial text object 2) and the target text object (original initial text object 3), between the target text object (original initial text object 3) and the target text object (original initial text object 4), and so on.

[0200] In this embodiment, the following judgment rules can be set: Figure 11 As an example, if there is a semantic correlation between at least two target text objects, Figure 11If there is a semantic correlation between all target text objects shown in , then it is judged that Figure 11 There is no correlation between all target text objects shown in ; it is also possible to set the number of target text objects that meet the semantic correlation relationship to reach a threshold (for example, 4), and then determine Figure 11 There is a correlation between all target text objects shown in , otherwise, that is, less than the threshold number of target text objects meet the semantic correlation, then it is judged Figure 11 There is no correlation between all the target text objects shown in .

[0201] On the other hand, the position correlation relationship represents the correlation between the positions of the plurality of target text objects in the document page. When the correlation relationship to be determined is the position correlation relationship between the bounding boxes corresponding to the respective target text objects, it can be determined whether the bounding boxes corresponding to the respective target text objects have a relationship such as line alignment or end alignment. For example, referring to Figure 12 When there are more than one target text object's bounding boxes with the same coordinate value on the Y axis (the second dimension) (or the difference between the coordinate values ​​of more than one target text object's bounding boxes is less than the threshold), then these target text objects will get a "row alignment" visual effect when rendered or printed out, so that it can be determined that the specific position-related relationship between such target text objects satisfies "row alignment", for example Figure 12 In the example, the target text objects corresponding to the original initial text object 1 and the original initial text object 2 both have coordinate values ​​key on the Y axis (second dimension). align1 , then the target text objects corresponding to the original initial text object 1 and the original initial text object 2 have a specific position correlation relationship of "line alignment". Similarly, the target text objects corresponding to the original initial text object 3 and the original initial text object 4 both have the coordinate value key on the Y axis (the second dimension) align2 , then the target text objects corresponding to the original initial text object 3 and the original initial text object 4 also have a specific position correlation relationship of "line alignment"; refer to Figure 12 When there are more than one target text object's bounding boxes with the same coordinate value on the X-axis (first dimension) (or the difference between the coordinate values ​​of more than one target text object's bounding boxes is less than the threshold), then these target text objects will obtain an "end-aligned" visual effect when rendered or printed out, so that it can be determined that the specific position-related relationship between such target text objects satisfies "end-alignment", for example Figure 12 In the example, the target text objects corresponding to the original initial text object 1, the original initial text object 3, the original initial text object 5, and the original initial text object 7 all have coordinate values ​​on the X-axis (first dimension).align1 , then the target text objects corresponding to original initial text object 1, original initial text object 3, original initial text object 5 and original initial text object 7 have a specific positional correlation relationship of "end alignment". Similarly, the target text objects corresponding to original initial text object 2, original initial text object 4, original initial text object 6 and original initial text object 8 all have a coordinate value on the Y axis (the second dimension) align2 , then the target text objects corresponding to the original initial text object 2, the original initial text object 4, the original initial text object 6 and the original initial text object 8 also have a specific position correlation relationship of "end alignment".

[0202] In this embodiment, the following judgment rules can be set: Figure 12 As an example, if there are at least two target text objects with a specific position correlation relationship such as "line alignment" or "end alignment" between them, it is determined that Figure 12 If there is a correlation between all target text objects shown in , on the contrary, if there is no specific position correlation such as "line alignment" or "end alignment" between any two target text objects, then it is judged that Figure 12 There is no correlation between all the target text objects shown in ; it can also be set that the number of target text objects that meet the specific position correlation relationship such as "line alignment" or "end alignment" reaches a threshold (for example, 4), and then it is determined Figure 12 There is a correlation between all target text objects shown in , otherwise, that is, less than the threshold number of target text objects meet the specific position correlation relationship such as "line alignment" or "end alignment", then it is judged Figure 12 There is no correlation between all the target text objects shown in .

[0203] for Figure 11 The non-text object 1 shown corresponds to the area to be identified and its corresponding multiple target text objects. Step S4 is executed to perform correlation detection on each target text object. Specifically, the purpose of correlation detection is to determine whether there is a correlation between the target text objects corresponding to the area to be identified corresponding to the non-text object 1. The correlation represents a possible correlation between the target text objects to be examined, such as the semantic correlation between the text contents of the target text objects, or the positional correlation between the bounding boxes corresponding to the target text objects.

[0204] When the correlation to be examined is set to a specific correlation, Figure 11The multiple target text objects detected by the fifth detection area as shown execute step S4, and the results may be: ① there is a correlation between the multiple target text objects detected by the fifth detection area; ② there is no correlation between the multiple target text objects detected by the fifth detection area.

[0205] For non-text object 1, if step S4 is executed and it is determined that there is a correlation between the multiple target text objects detected in the fifth detection area, then when step S5 is executed, Figure 11 The area to be identified corresponding to the non-text object 1 shown is Figure 10 The area within the first bounding box in is marked as the illustration area; on the contrary, if step S4 is executed and it is determined that there is no correlation between the multiple target text objects detected in the fifth detection area, then the area to be identified corresponding to the non-text object 1 will not be marked as the illustration area.

[0206] In this embodiment, by Figure 10 and Figure 9 By comparison, it can be seen that the range of the fifth detection area determined by executing steps S301B-S303B is larger than the sum of the first detection area, the second detection area, the third detection area and the fourth detection area determined by executing steps S301A-S306A. When executing step S4, it is beneficial to detect more target text objects, so that when executing step S5, more target text objects can be detected for correlation, which is beneficial to executing step S5 to detect the correlation between the target text objects, thereby identifying the area to be identified as the illustration area, and achieving a higher recognition sensitivity of the area to be identified; by Figure 11 It can be seen from the illustrated principle that, based on steps S301B-S303B, when executing step S5, more diverse methods can be used to determine whether there is a correlation between the multiple target text objects.

[0207] After executing steps S3-S4 for non-text object 1, steps S3-S4 are also executed for non-text object 2 to determine whether the area to be identified corresponding to non-text object 2 is an illustration area. If there are more non-text objects, all non-text objects are traversed and steps S3-S4 are executed for the area to be identified corresponding to each non-text object to determine whether the area to be identified corresponding to each non-text object is an illustration area, thereby identifying the portion of the PDF page that is an illustration area.

[0208] In this embodiment, the principle of executing steps S1-S5 is as follows: by performing region extraction on the PDF page according to the non-text objects, it is possible to determine the potential area on the PDF page where illustrations may exist, that is, the area to be identified; further judgment is performed on any area to be identified, if the target text object is detected through its corresponding detection area, and there is a semantic correlation relationship or position correlation relationship between these target text objects, it indicates that there is a contextual relationship between these target text objects, and the text characters corresponding to these target text objects form a text paragraph, that is, the graphic objects existing at the position corresponding to the area to be identified in the PDF page are rendered and displayed together with each target text object. When the image is identified, the graphic objects at the corresponding positions of the area to be identified will form an illustration surrounded by the texts corresponding to the target text objects (equivalent to the illustration being embedded in the text paragraph), so that the area to be identified can be marked as an illustration area, and the part belonging to the illustration area can be identified from the PDF page; since steps S1-S5 are based on the analysis of the objects on the PDF page to realize the identification of the illustration area, the original data characteristics of the objects in the PDF format are utilized, and there is no need to convert the entire PDF page into image data before performing a large-scale image processing process, thereby having a lower resource consumption level and a lower false detection rate level, and can retain information such as the original object writing order in the PDF page.

[0209] In this embodiment, when executing the method for identifying illustration areas in document pages, in addition to executing steps S1-S5 (and executing steps S501-S503 when executing step S5), the following steps may also be executed:

[0210] S6. In the case where the area to be identified is an illustration area, searching for a target text object from the third text set and / or the fourth text set;

[0211] S7. Determine the content of the target text object found;

[0212] S8. When it is determined that the found target text object constitutes a text illustration, the found target text object is identified as an illustration area according to the position of the bounding box corresponding to the found target text object.

[0213] In step S6, refer to Figure 9 When it is determined that the area to be identified within the first bounding box is an illustration area, the target text object is found from the corresponding third text set TVT = {initial text object 1, initial text object 2} or from the fourth text set TVB = {initial text object 9}, that is, at least one of the initial text object 1, initial text object 2 and initial text object 9 is found.

[0214] In step S7, the target text object found in step S6 is subjected to content determination, thereby determining whether the found target text object belongs to the text illustration description of the illustration area.

[0215] When executing step S7, the key dictionary method can be used to judge the target text object found in step S6. Specifically, an image recognition algorithm can be run to perform image recognition on the non-text object in the first bounding box, thereby obtaining a keyword representing the content of the non-text object in the first bounding box, and then searching for text that is the same as or similar to the keyword in the target text object found in step S6 (that is, at least one of the initial text object 1, initial text object 2, and initial text object 9). For example, if a text that is the same as or similar to the keyword is found in the initial text object 9, then in step S7, the initial text object 9 is determined to be a text illustration description belonging to the illustration area, and in step S8, the position of the bounding box corresponding to the initial text object 9 in the PDF page is also identified as the illustration area.

[0216] When executing step S7, the continuity check method can also be used to judge the target text objects found in step S6. Specifically, for the target text objects found in step S6, that is, initial text object 1, initial text object 2 and initial text object 9, refer to Figure 12 According to the principle shown, each of the initial text objects is checked one by one to see if there is a continuity relationship such as line alignment or end alignment between them and other text objects. For initial text objects 1 and 2, since they have a continuity relationship of line alignment with each other, there is also a continuity relationship of end alignment between initial text objects 1 and 3, and there is also a continuity relationship of end alignment between initial text objects 2 and 4, it can be determined that initial text objects 1 and 2 do not belong to the text illustration description of the illustration area; and for initial text object 9, since initial text object 9 does not have a continuity relationship such as line alignment or end alignment with any other text object, it can be determined that initial text object 9 is a text illustration description belonging to the illustration area. In step S8, the position of the bounding box corresponding to initial text object 9 in the PDF page is also identified as the illustration area.

[0217] In this embodiment, the principle of executing steps S6-S8 is: Figure 9According to the principles of steps S501-S503, when the PDF page is in the reading order of LRTB, the target text object in the third text set TVT or the fourth text set TVB is respectively located above and below the area to be identified as the illustration area. The target text object in the third text set TVT or the fourth text set TVB is more likely to belong to the text illustration description of the illustration area, that is, the target text object in the third text set TVT or the fourth text set TVB is integrated with the illustration. Therefore, the position of the bounding box corresponding to the target text object constituting the text illustration description is also identified as the illustration area. The range of the illustration area can be expanded so that the parts of the PDF page belonging to the illustration can be batch processed.

[0218] In this embodiment, when executing the method for identifying illustration areas in document pages, in addition to executing steps S1-S5 or steps S1-S8, the following steps may also be executed:

[0219] S9. Find the target graphic object in the PDF page;

[0220] S10. Delete the target graphic object found from the PDF page;

[0221] S11. Perform text recognition on the PDF page after deleting the target graphic object.

[0222] In step S9, all illustration areas are determined in the PDF page according to the execution results of steps S1-S5 or steps S1-S8, wherein the graphic objects located in the illustration areas are target graphic objects.

[0223] In step S10, the data corresponding to the target graphic object determined in step S9 is deleted from the original data of the PDF page, and only the data of the graphic objects that do not belong to the target graphic object is retained.

[0224] In step S11, based on the data of the target graphic object deleted in step S10, a text recognition process such as text search or OCR recognition is performed on the data of the graphic object retained in the PDF page, thereby realizing operations such as searching, selecting, editing, and deleting the text content in the PDF page.

[0225] In this embodiment, the principle of executing steps S9-S11 is that the target graphic object in the PDF page is a graphic object located in the illustration area, which is specifically a non-text object that forms the illustration part surrounded by text, or a text object that forms a text illustration description, that is, the target graphic object does not belong to the text content to be recognized in the PDF page. By first deleting the target graphic object from the PDF page and then performing text recognition on the PDF page, the interference of the target graphic object on the text recognition process can be eliminated, thereby achieving correct recognition of the text content in the PDF page when the text is surrounded by the illustration.

[0226] In this embodiment, the method for identifying illustration areas in reference document pages is applied to PDF pages, and can be used to perform illustration area recognition processing on other types of document pages with similar data structures such as CAJ, Html and Xml, thereby achieving the same technical effect of retaining information such as the original object writing order in the document page.

[0227] In this embodiment, the illustration area recognition device in the document page can be run to perform the illustration area recognition method in the document page. Specifically, corresponding to the above method embodiment, as shown in FIG. Figure 13 As shown, the illustration area recognition device in the document page includes:

[0228] The document parsing module 1301 is used to parse the document page of the layout document to obtain at least one initial text object and at least one non-text object;

[0229] A region extraction module 1302 is configured to extract a region of the document page based on at least one non-text object to determine a region to be identified;

[0230] The area determination module 1303 is configured to determine at least one detection area based on the area to be identified;

[0231] an object determination module 1304 for determining, based on the detection area, at least one target text object from the at least one initial text object, wherein a bounding box of the target text object intersects with the detection area;

[0232] The first region identification module 1305 is configured to perform correlation detection on the target text objects, and identify the region to be identified as an illustration region when the correlation is detected between the target text objects corresponding to the region to be identified.

[0233] In an optional embodiment, the region extraction module 1302 is specifically configured to:

[0234] Performing connectivity detection on the non-text objects to determine connected areas in the document page;

[0235] The connected area is used as the area to be identified.

[0236] In an optional embodiment, the region extraction module 1302 is specifically configured to:

[0237] Create and initialize a binary memory bitmap;

[0238] Mapping the non-text object from the document page to the binary memory bitmap; bits in the mapped binary memory bitmap have corresponding mark masks, and the mark mask is a first value or a second value;

[0239] Mapping the bit in the binary memory bitmap corresponding to the mark mask having the first value to the document page;

[0240] The mapped portion of the document page is used as the connected area.

[0241] In an optional implementation, the region determination module 1303 is specifically configured to:

[0242] Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified;

[0243] Determine a first dimension and a second dimension; the positive direction of the first dimension is the row direction of the document page, and the negative direction of the second dimension is the block direction of the document page;

[0244] Starting from the edge of the first bounding box and extending in the negative direction of the first dimension, a first detection area is determined;

[0245] Starting from the edge of the first bounding box and extending in the positive direction of the first dimension, a second detection area is determined;

[0246] Starting from the edge of the first bounding box and extending in the positive direction of the second dimension, a third detection area is determined;

[0247] Starting from the edge of the first bounding box and extending in the negative direction of the second dimension, a fourth detection area is determined.

[0248] In an optional implementation, the first region identification module 1305 is specifically configured to:

[0249] Classifying each target text object into a first text set, a second text set, a third text set, or a fourth text set; wherein the target text objects in the first text set intersect with the first detection area, the target text objects in the second text set intersect with the second detection area, the target text objects in the third text set intersect with the third detection area, and the target text objects in the fourth text set intersect with the fourth detection area;

[0250] Performing correlation detection on each text set combination respectively; each text set combination is a combination of the first text set and the third text set, a combination of the first text set and the fourth text set, a combination of the second text set and the third text set, and a combination of the second text set and the fourth text set;

[0251] In the case where the correlation is detected from any combination of the text sets, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized areas.

[0252] In an optional implementation, the region determination module 1303 is specifically configured to:

[0253] Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified;

[0254] Extending outward from each side and each vertex of the first bounding box to determine an extended area;

[0255] The epitaxial region is used as the fifth detection region.

[0256] In an optional implementation, the first region identification module 1305 is specifically configured to:

[0257] Sort each target text object according to the LRTB sequence corresponding to the document page;

[0258] Traversing the sorted target text objects in sequence, and performing correlation detection on at least two adjacent target text objects;

[0259] In the case where the correlation is detected from any at least two target text objects, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized area.

[0260] In an optional embodiment, the correlation relationship is a contextual semantic correlation relationship, or a positional correlation relationship; wherein, the contextual semantic correlation relationship represents the correlation between the contextual semantics of multiple target text objects, and the positional correlation relationship represents the correlation between the positions of multiple target text objects in a document page.

[0261] In an optional embodiment, the illustration region identification device further includes:

[0262] A first object search module is configured to search for a target text object from the third text set and / or the fourth text set after the area to be identified is identified as an illustration area;

[0263] A content determination module, configured to determine the content of the found target text object;

[0264] The second area recognition module is used to recognize the area within the boundary box corresponding to the found target text object as the illustration area when it is determined that the found target text object constitutes a text illustration.

[0265] In an optional embodiment, the illustration region identification device further includes:

[0266] A second object search module is configured to search for a target graphic object in the document page after identifying the area to be identified as an illustration area; the target graphic object is a graphic object located in the illustration area;

[0267] The object deletion module is used to delete the found target graphic object from the document page.

[0268] In an optional embodiment, the illustration region identification device further includes:

[0269] The text recognition module is configured to perform text recognition on the document page after deleting the found target graphic object from the document page.

[0270] The document parsing module, region extraction module, region determination module, object determination module and first region identification module included in the illustration region identification device are used to execute steps S1-S5 respectively. On the basis of running the above modules, the first object search module, content determination module and second region identification module can also be run to respectively execute steps S6-S8; the second object search module and object deletion module can also be further run to respectively execute steps S9-S10; the text recognition module can also be further run to execute step S11. Among them, the modules such as the document parsing module, region extraction module, region determination module, object determination module, first region identification module, first object search module, content determination module, second region identification module, second object search module, object deletion module and text recognition module can be hardware modules, software modules or a combination of hardware and software that have the function of executing the corresponding steps. For example, the document parsing module, the region extraction module, the region determination module, the object determination module, the first region identification module, the first object search module, the content determination module, the second region identification module, the second object search module, the object deletion module and the text recognition module are respectively computer devices that can run independently; or, the document parsing module, the region extraction module, the region determination module, the object determination module, the first region identification module, the first object search module, the content determination module, the second region identification module, the second object search module, the object deletion module and the text recognition module are respectively different subroutines (or functions) in the same application software run by the same computer device.

[0271] By running the device for identifying illustration areas in a document page, the same technical effect as the method for identifying illustration areas in a document page in the embodiment can be achieved.

[0272] In this embodiment, the computer is programmed to execute Figure 1 and storing such a computer program in Figure 14 The memory 1401 of the computer device shown, or Figure 15 The computer program stored in the memory 1401 or the computer readable storage medium 1501 can be Figure 14 Processor 1402 in or Figure 15 The processor 1502 in the processor reads it and obtains the data for execution through operations such as compilation. Figure 1 The processor executes these instructions and can perform Figure 1 Each step S1-S5 in the method is to execute the illustration area recognition method in the document page.

[0273] because Figure 14 The computer device shown and Figure 15The computer readable storage medium shown can enable the processor to execute the illustration area recognition method in the document page, so Figure 14 The computer device shown and Figure 15 The computer-readable storage media shown can achieve the same technical effect as the method for identifying illustration areas in document pages in the embodiment.

[0274] It should be noted that, unless otherwise specified, the singular forms "a" and "the" used in this disclosure are also intended to include plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by those skilled in the art. The terms used in this embodiment description are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any combination of one or more related listed items.

[0275] It should be understood that, although the present disclosure may adopt the term first, second, third etc. to describe various elements, these elements should not be limited to these terms.These terms are only used to distinguish the elements of the same type from each other.For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.The use of any and all examples or exemplary language ("for example", "such as" etc.) provided by the present embodiment is only intended to better illustrate embodiments of the present invention, and unless otherwise required, the scope of the present invention will not be limited.

[0276] It should be appreciated that embodiments of the present invention can be implemented or practiced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The methods can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and figures described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed application-specific integrated circuit for this purpose.

[0277] In addition, the operations of the processes described in this embodiment may be performed in any suitable order, unless otherwise indicated in this embodiment or otherwise clearly contradicted by the context. The processes described in this embodiment (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. A computer program includes multiple instructions that can be executed by one or more processors.

[0278] Furthermore, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, RAM, ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the above steps in conjunction with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.

[0279] The above are merely preferred embodiments of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods may be made.

Claims

1. A method for identifying illustration areas in a document page, characterized in that: The illustration area recognition method includes: Parsing a document page of the layout document to obtain at least one initial text object and at least one non-text object; Extracting a region of the document page according to at least one non-text object to determine a region to be identified; Determining at least one detection area according to the area to be identified; determining at least one target text object from the at least one initial text object according to the detection area, wherein a bounding box of the target text object intersects the detection area; A correlation detection is performed on the target text objects, and when it is detected that the correlation exists between the target text objects corresponding to the area to be identified, the area to be identified is identified as an illustration area.

2. The illustration area recognition method according to claim 1, characterized in that: The extracting a region of the document page according to the at least one non-text object to determine a region to be identified includes: Performing connectivity detection on the non-text objects to determine connected areas in the document page; The connected area is used as the area to be identified.

3. The illustration area recognition method according to claim 2, characterized in that: The performing connectivity detection on the non-text objects to determine connected areas in the document page includes: Create and initialize a binary memory bitmap; Mapping the non-text object from the document page to the binary memory bitmap; bits in the mapped binary memory bitmap have corresponding mark masks, and the mark mask is a first value or a second value; Mapping the bit in the binary memory bitmap corresponding to the mark mask having the first value to the document page; The mapped portion of the document page is used as the connected area.

4. The illustration region recognition method according to claim 1, characterized in that: The determining of at least one detection area according to the area to be identified includes: Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified; Determine a first dimension and a second dimension; the positive direction of the first dimension is the row direction of the document page, and the negative direction of the second dimension is the block direction of the document page; Starting from the edge of the first bounding box and extending in the negative direction of the first dimension, a first detection area is determined; Starting from the edge of the first bounding box and extending in the positive direction of the first dimension, a second detection area is determined; Starting from the edge of the first bounding box and extending in the positive direction of the second dimension, a third detection area is determined; Starting from the edge of the first bounding box and extending in the negative direction of the second dimension, a fourth detection area is determined.

5. The illustration area recognition method according to claim 4, characterized in that: The performing correlation detection on the target text object includes: Classifying each target text object into a first text set, a second text set, a third text set, or a fourth text set; wherein the target text objects in the first text set intersect with the first detection area, the target text objects in the second text set intersect with the second detection area, the target text objects in the third text set intersect with the third detection area, and the target text objects in the fourth text set intersect with the fourth detection area; Performing correlation detection on each text set combination respectively; each text set combination is a combination of the first text set and the third text set, a combination of the first text set and the fourth text set, a combination of the second text set and the third text set, and a combination of the second text set and the fourth text set; In the case where the correlation is detected from any combination of the text sets, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized areas.

6. The illustration region recognition method according to claim 1, characterized in that: The determining of at least one detection area according to the area to be identified includes: Detecting a first bounding box; the first bounding box is a bounding box corresponding to the area to be identified; Extending outward from each side and each vertex of the first bounding box to determine an extended area; The epitaxial region is used as the fifth detection region.

7. The illustration region recognition method according to claim 5, characterized in that: The performing correlation detection on the target text object includes: Sort each target text object according to the LRTB sequence corresponding to the document page; Traversing the sorted target text objects in sequence, and performing correlation detection on at least two adjacent target text objects; In the case where the correlation is detected from any at least two target text objects, it is determined that the correlation exists between the target text objects corresponding to the to-be-recognized area.

8. The illustration region recognition method according to claim 1, characterized in that: The correlation relationship is a contextual semantic correlation relationship or a positional correlation relationship; wherein, the contextual semantic correlation relationship represents the correlation between the contextual semantics of multiple target text objects, and the positional correlation relationship represents the correlation between the positions of multiple target text objects in a document page.

9. The illustration region recognition method according to claim 6, characterized in that: After the step of identifying the area to be identified as an illustration area, the method for identifying an illustration area in a document page further includes: Finding a target text object from the third text set and / or the fourth text set; Performing content determination on the found target text object; When it is determined that the found target text object constitutes a text illustration, the area within the bounding box corresponding to the found target text object is identified as the illustration area.

10. The illustration region recognition method according to any one of claims 1 to 9, characterized in that: After the step of identifying the area to be identified as an illustration area, the method for identifying an illustration area in a document page further includes: Finding a target graphic object in the document page; the target graphic object is a graphic object located in the illustration area; The found target graphic object is deleted from the document page.

11. The illustration region recognition method according to claim 10, characterized in that: After the step of deleting the found target graphic object from the document page, the method for identifying an illustration area in a document page further includes: Text recognition is performed on the document page.

12. A device for identifying illustration areas in a document page, characterized in that: The illustration area recognition device includes: A document parsing module, configured to parse a document page of a layout document to obtain at least one initial text object and at least one non-text object; A region extraction module, configured to extract a region of the document page according to at least one non-text object and determine a region to be identified; An area determination module, configured to determine at least one detection area based on the area to be identified; an object determination module, configured to determine at least one target text object from the at least one initial text object based on the detection area, wherein a bounding box of the target text object intersects the detection area; The first region recognition module is configured to perform correlation detection on the target text objects, and identify the region to be recognized as an illustration region when the correlation is detected between the target text objects corresponding to the region to be recognized.

Citation Information

Patent Citations

  • Device and method for extracting composite graph in format document

    CN104346615A

  • Method for identifying chart in PDF document

    CN107133566A

  • PDF (Portable Document Format) document analysis method and device, equipment and storage medium

    CN116451677A

  • Detection of numbered captions

    US20140212038A1