Document processing method and device, and image data extraction method and device
By generating and merging bounding boxes and combining pixel analysis and text recognition technology, the problem of high cost and low accuracy of chart data extraction in the existing technology is solved, and efficient and accurate chart data extraction is achieved.
Patent Information
- Application Number
- CN202111156200.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-09-29
AI Technical Summary
In the prior art, manually intercepting charts in documents to extract structured data is costly and has low accuracy, especially in documents with multiple columns of text.
It generates multiple bounding boxes to mark text-sparse areas, merges adjacent bounding boxes to determine the local image, and extracts structured data from the image using pixel analysis and text recognition techniques.
It achieves efficient and accurate extraction of chart data in multi-column text documents, reduces labor costs and improves the accuracy of data extraction.
Smart Images

Figure CN113886582B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to the field of document processing technology. More specifically, the present disclosure provides a document processing method and apparatus, a method and apparatus for extracting data from an image, an electronic device, and a storage medium. Background Art
[0002] A document may contain one or more charts. The data in these charts can be unstructured, such as images or background images. In related technologies, charts in documents can be manually captured and then their characteristic points (e.g., axis origins, scale line endpoints, etc.) and data values can be observed to extract structured data from these charts. Summary of the Invention
[0003] The present disclosure provides a document processing method and device, a data extraction method and device for an image, an electronic device, and a storage medium.
[0004] According to a first aspect, a document processing method is provided, the method comprising: generating a plurality of first bounding boxes based on position information of line text images in a document page; generating a plurality of second bounding boxes based on the position information of the plurality of first bounding boxes, each second bounding box being used to mark a text-sparse area in the document page; performing a merging operation on adjacent second bounding boxes to obtain a plurality of candidate bounding boxes; determining a plurality of partial images of the document page for the plurality of candidate bounding boxes based on the position information of each candidate bounding box; and generating a target image based on the content in the plurality of partial images.
[0005] According to a second aspect, a data extraction method for an image is provided, the method comprising: determining the coordinates of N marking points located on a coordinate axis in the target image based on the pixel value of each pixel in the target image; performing a division operation on the target image based on the coordinates of the N marking points to obtain N+1 sub-regions; performing a text recognition operation on the i-th sub-region in the N+1 sub-regions to obtain an i-th group of data corresponding to the i-th sub-region; i=1,...,N+1; wherein the target image is generated according to the document processing method provided by the present disclosure.
[0006] According to a third aspect, a document processing device is provided, which includes: a first generating module for generating a plurality of first bounding boxes based on position information of line text images in a document page; a second generating module for generating a plurality of second bounding boxes based on the position information of the plurality of first bounding boxes, each second bounding box being used to mark a text-sparse area in the document page; a merging module for performing a merging operation on adjacent second bounding boxes to obtain a plurality of candidate bounding boxes; a first determining module for determining a plurality of partial images of the document page based on the position information of each candidate bounding box for the plurality of candidate bounding boxes; and a third generating module for generating a target image based on the content in the plurality of partial images.
[0007] According to a fourth aspect, a data extraction device for an image is provided, which includes: a second determination module, used to determine the coordinates of N marking points located on the coordinate axis in the target image according to the pixel value of each pixel in the target image; a division module, used to perform a division operation on the target image according to the coordinates of the N marking points, to obtain N+1 sub-regions; a text recognition module, used to perform a text recognition operation on the i-th sub-region among the N+1 sub-regions, to obtain the i-th group of data corresponding to the i-th sub-region; i=1,...,N+1; wherein the target image is generated according to the document processing device provided by the present disclosure.
[0008] According to a fifth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.
[0009] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method provided according to the present disclosure.
[0010] According to a seventh aspect, a computer program product is provided, comprising a computer program, which implements the method provided according to the present disclosure when executed by a processor.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0013] Figure 1is a schematic diagram of an exemplary system architecture to which a document processing method and apparatus and a data extraction method for an image can be applied according to an embodiment of the present disclosure;
[0014] Figure 2 is a flowchart of a document processing method according to an embodiment of the present disclosure;
[0015] Figures 3A to 3B is a schematic diagram of a document processing method according to an embodiment of the present disclosure;
[0016] Figures 4A to 4C is a schematic diagram of a document processing method according to an embodiment of the present disclosure;
[0017] Figure 5 is a flow chart of a method for extracting data from an image according to one embodiment of the present disclosure;
[0018] Figures 6A to 6B is a schematic diagram of a method for extracting data from an image according to an embodiment of the present disclosure;
[0019] Figure 7 is a block diagram of a document processing apparatus according to an embodiment of the present disclosure;
[0020] Figure 8 is a block diagram of a data extraction apparatus for an image according to an embodiment of the present disclosure; and
[0021] Figure 9 is a block diagram of an electronic device for a document processing method and / or a data extraction method for an image according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] In one example method, structured data of charts in documents can be manually extracted, which is costly and has a low accuracy rate in data extraction.
[0024] Figure 1 This is a schematic diagram of an exemplary system architecture to which a document processing method and apparatus and / or a data extraction method and apparatus for a chart can be applied according to an embodiment of the present disclosure. It should be noted that: Figure 1The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0025] like Figure 1 As shown, the system architecture 100 according to this embodiment may include multiple terminal devices 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal devices 101 and the server 103. The network 102 may include various connection types, such as wired and / or wireless communication links, etc.
[0026] A user may use a terminal device 101 to interact with a server 103 via a network 102 to receive or send messages, etc. The terminal device 101 may be any electronic device, including but not limited to a smartphone, a tablet computer, a laptop computer, and the like.
[0027] At least one of the document processing method and the data extraction method for images provided in the embodiments of the present disclosure can generally be executed by the server 103. Accordingly, the document processing apparatus and the data extraction apparatus for images provided in the embodiments of the present disclosure can generally be set in the server 103. The document processing method and the data extraction method for images provided in the embodiments of the present disclosure can also be executed by a server or a server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103. Accordingly, the document processing apparatus and the data extraction apparatus for images provided in the embodiments of the present disclosure can also be set in a server or a server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103.
[0028] Figure 2 is a flowchart of a document processing method according to an embodiment of the present disclosure.
[0029] like Figure 2 As shown, the document processing method 200 may include operations S210 to S250.
[0030] In operation S210 , a plurality of first bounding boxes are generated according to position information of text images in a document page.
[0031] In the embodiment of the present disclosure, the document may be a PDF (Portable Document Format) document.
[0032] For example, the document may be a searchable PDF document or a PDF document created by other editing applications.
[0033] For example, the document can be a PDF document containing only images. In one example, a PDF document containing only images can be created by a scanning operation.
[0034] For example, the document may be a DOC (Document) document. In one example, because data in a DOC document is damaged, each page thereof is displayed in an image format.
[0035] In the embodiment of the present disclosure, the position information of the line text image can be obtained based on the position information of each text image.
[0036] For example, you can use open source software such as XPDFReader to parse a PDF document and obtain the location information of each text image in the PDF document. In one example, you can use open source software such as XPDFReader to parse a searchable PDF document or a PDF document created by other editing applications to obtain the location information of each text image in the PDF document. In another example, you can first perform OCR (optical character recognition) on a PDF document containing only images, create a text layer in the PDF document containing only images, and then use open source software such as XPDFReader to obtain the location information of each text image in the PDF document.
[0037] For example, each text image in a PDF document has position information relative to the top left vertex of the page on which it resides, such as coordinates (e.g., the coordinates of the top left vertex of the text image), height, and width. Based on the vertical coordinate values of the text images, a fourth sub-bounding frame is generated, where the vertical coordinate values of each text image in the fourth sub-bounding frame are the same or similar. Then, based on the vertical coordinate values of the fourth sub-bounding frames (the average of the vertical coordinate values of each text image), the spacing (e.g., the distance between edges) of multiple fourth sub-bounding frames within a preset range is calculated. Based on the spacing of the multiple fourth sub-bounding frames, a fifth sub-bounding frame is generated. The fifth sub-bounding frame includes at least one fourth sub-bounding frame, and the spacing between at least one of the fourth sub-bounding frames is less than a preset spacing threshold. Based on the position information of the fifth sub-bounding frame, position information of a row of text images can be obtained. The position information of a row of text images is obtained based on the position information of each text image in the row. Based on the position information of the row of text images, a first bounding frame is generated. The first bounding frame includes at least one fifth sub-bounding frame. The vertical coordinate values of each text image in the first bounding frame are the same, or the difference in vertical coordinate values is within a preset difference range.
[0038] In operation S220 , a plurality of second bounding boxes are generated according to the position information of the plurality of first bounding boxes, where each second bounding box is used to mark a text-sparse area in the document page.
[0039] In the embodiment of the present disclosure, the text-sparse area may be a blank area without text or an area with the number of texts being less than a preset threshold.
[0040] For example, a text-sparse area may be an area of a document where a chart is located.
[0041] In the embodiment of the present disclosure, the second bounding box includes a first sub-bounding box and a second sub-bounding box.
[0042] In the embodiment of the present disclosure, a first sub-enclosing frame is generated between any two vertically adjacent first enclosing frames. In the embodiment of the present disclosure, the width of the first sub-enclosing frame may be the width of the page.
[0043] For example, the first bounding box is generated based on the position information of a line of text images, and there is a blank area between multiple lines of text images. The blank area may be a sparse text area and can be marked by the first sub-bounding box.
[0044] In the embodiment of the present disclosure, a second sub-enclosing frame is generated on the left and / or right side of each first enclosing frame. In the embodiment of the present disclosure, the width of the second sub-enclosing frame is the length from the edge of the first enclosing frame to the edge of the document page.
[0045] For example, a document may have page margins, and there may be blank areas on the left and right sides of the line text image. These blank areas may be text sparse areas and may be marked by the second sub-bounding box.
[0046] In an embodiment of the present disclosure, a document page may include at least two columns of text.
[0047] For example, some papers contain two columns of text per document page.
[0048] In an embodiment of the present disclosure, for each of the multiple first enclosing frames, at least one overlapping area is determined according to position information of the first enclosing frame and position information of a second enclosing frame, and one overlapping area corresponds to at least one second enclosing frame.
[0049] For example, in a document containing two columns of text, multiple first bounding boxes and multiple second bounding boxes are generated for each line of text image in each column. In a document containing two columns of text, the positional information of the line of text images in the two columns is not completely consistent. Therefore, the first bounding box may overlap with multiple second bounding boxes.
[0050] In the embodiment of the present disclosure, for at least one overlapping region, the overlapping region is removed from at least one second bounding box corresponding to each overlapping region to obtain a plurality of adjusted second bounding boxes.
[0051] For example, the overlapping region may be removed from the second bounding box corresponding to each overlapping region. For another example, the overlapping region and the region above or below the overlapping region may be removed from at least one second bounding box corresponding to each overlapping region.
[0052] In operation S230 , a merging operation is performed on adjacent second bounding boxes to obtain a plurality of candidate bounding boxes.
[0053] In the embodiment of the present disclosure, the candidate bounding box is a rectangle.
[0054] For example, the width of the second sub-enclosing frame generated on the left side of the first enclosing frame may be the length from the left frame line of the first enclosing frame to the left edge of the document page.
[0055] For another example, the width of the second sub-enclosing frame generated on the right side of the first enclosing frame may be the length from the right frame line of the first enclosing frame to the right edge of the document page.
[0056] In an embodiment of the present disclosure, after performing a merging operation on adjacent second bounding boxes, the width of the obtained candidate bounding box may be the same as that of the second sub-bounding box, and the height of the obtained candidate bounding box may be greater than or equal to the height of the merged second bounding box.
[0057] In the embodiment of the present disclosure, a splitting operation is performed on each first sub-bounding frame to obtain a plurality of third sub-bounding frames having the same width as the second sub-bounding frame.
[0058] For example, the width of the first sub-enclosing frame may be greater than the width of the second sub-enclosing frame. The width of the first sub-enclosing frame needs to be adjusted to perform the merge operation. After performing the split operation, multiple third sub-enclosing frames and the remaining portion of the first sub-enclosing frame excluding the third sub-enclosing frames are obtained.
[0059] In the embodiment of the present disclosure, a merging operation is performed on the third sub-bounding frame and the second sub-bounding frame to obtain a candidate bounding frame.
[0060] For example, when there are multiple adjacent second bounding boxes, the width of the candidate bounding box may be the same as that of the second sub-bounding box, and the height of the candidate bounding box may be greater than or equal to the heights of the adjacent second sub-bounding boxes and the third sub-bounding box.
[0061] For example, after performing the merging operation, at least one third sub-bounding box is merged with at least one second sub-bounding box, and the obtained box can be used as a candidate bounding box.
[0062] For another example, after performing the merging operation, the remaining portion of the first sub-bounding box may also be used as a candidate bounding box.
[0063] In operation S240, for the plurality of candidate bounding boxes, a plurality of partial images of the document page are determined according to the position information of each candidate bounding box.
[0064] For example, each candidate bounding box may determine an area on the document page, and based on the area, an image within the area may be determined.
[0065] In operation S250, a target image is generated based on the contents of the plurality of partial images.
[0066] For example, the contents of the plurality of partial images may be a chart, a blank space, or a landscape image. The partial image with the chart content may be used as the target image.
[0067] For example, you can remove white edges from a target image based on the margins of a document page.
[0068] Through the disclosed embodiments, since images can be text-sparse areas within document pages, the above operations can accurately extract charts from document pages. This is particularly true for papers with multiple columns of text. Charts can be quickly and accurately extracted from a large number of papers, saving significant manpower.
[0069] Figures 3A to 3B It is a schematic diagram according to an embodiment of the present disclosure.
[0070] like Figure 3A As shown, the document page 301 includes three line text images, and three first bounding boxes can be generated according to the position information of each line text image. The three first bounding boxes include, for example Figure 3A The first enclosing box 3021 and the first enclosing box 3022, and the first enclosing box corresponding to "this line of text is the second example".
[0071] According to the position information of the three first bounding boxes, multiple second bounding boxes can be generated, for example Figure 3A The four second sub-bounding boxes 3032 in . Figure 3A A first sub-bounding box 30311 and a second sub-bounding box 30312 are also shown.
[0072] like Figure 3B As shown, a merge operation can be performed on the adjacent second bounding boxes to obtain four candidate bounding boxes, such as Figure 3BCandidate bounding boxes 3041, 3042, and 3043 are shown in FIG. The remaining portion of the first sub-bounding box between the first bounding box marked "This line of text is only an example" and first bounding box 3021 can also be used as a candidate bounding box. After the merge operation, candidate bounding box 3043 is the remaining portion of first sub-bounding box 30312. For each of the four candidate bounding boxes, multiple partial images of document page 301 can be determined based on the position information of each candidate bounding box. Based on the content of the multiple partial images, a target image can be determined. For example, a partial image determined based on candidate bounding box 3043 can be determined as the target image.
[0073] Figures 4A to 4C is a schematic diagram according to another embodiment of the present disclosure.
[0074] like Figure 4A As shown, the document page 401 includes two columns of text. The column of text on the left side of the document page 401 includes four line text images, while the column of text on the right side of the document page 401 includes one line text image and a chart 405.
[0075] According to the position information of the text image in the document page, five first bounding boxes can be generated, such as Figure 4A The first bounding box 4021 and the first bounding box 4022 in .
[0076] According to the position information of the five first bounding boxes, multiple second bounding boxes can be generated. The second bounding boxes may include, for example Figure 4A The first sub-enclosing frame 4031 between "Example Title" and "Left Column Text - Example" is shown in FIG. The second enclosing frame may also include, for example Figure 4A The second sub-enclosing boxes 4033 on both sides of the “left column text example 1” and the second sub-enclosing boxes 4032 on both sides of the “right column text example”.
[0077] like Figure 4A As shown, a third enclosing frame 406 may be generated above the first enclosing frame 4021 in the document page 401 , and a third enclosing frame 406 may be generated below the first enclosing frame 4022 in the document page 401 .
[0078] like Figure 4A As shown, the first enclosing frame marked "right column text example" has overlapping areas with the first sub-enclosing frame 4031 and the second sub-enclosing frame 4033. The overlapping areas can be removed from the first sub-enclosing frame 4031 and the second sub-enclosing frame 4033 to obtain, for example Figure 4BIn a similar manner, the overlapping region can be removed from at least one second enclosing frame corresponding to each overlapping region to obtain other adjusted second enclosing frames.
[0079] A merge operation can be performed on adjacent second enclosing frames to obtain multiple candidate enclosing frames. During the merge operation, second sub-enclosing frame 4035 and adjusted second sub-enclosing frame 4033′ are adjacent to first sub-enclosing frame 4034. Adjacent second enclosing frames whose width difference is within a preset difference range can be merged. For example, if the width difference between first sub-enclosing frame 4034 and second sub-enclosing frame 4035 in second sub-enclosing frame 4035 and adjusted second sub-enclosing frame 4033′ is small and within the preset difference range, first sub-enclosing frame 4034 and second sub-enclosing frame 4035 can be merged.
[0080] In one example, a split operation is performed on the first sub-bounding box 4034 to obtain a third sub-bounding box with the same width as the second sub-bounding box 4035. In a similar manner, a split operation is performed on the first sub-bounding box 4036, and then a merge operation is performed to obtain, for example Figure 4C Candidate bounding box 4043 in .
[0081] The multiple candidate bounding boxes may include, for example Figure 4C The candidate bounding boxes 4041, 4042, 4043, and 4044 in FIG. For multiple candidate bounding boxes, multiple partial images of the document page 401 can be determined based on the position information of each candidate bounding box. A target image can be determined based on the contents of the multiple partial images. For example, a partial image determined based on the candidate bounding box 4043 can be determined as the target image. The target image can be, for example, Figure 4C The line chart in 405.
[0082] It should be noted that, for example Figures 3A to 3B 、 Figures 4A to 4C The width of the second bounding box in the figure can be the length from the edge of the first bounding box to the edge of the document page, but in order to clearly distinguish the edge of the first bounding box, the edge of the second bounding box, and the edge of the document page in the figure, the width of the second bounding box is reduced to a certain extent.
[0083] Figure 5 is a flowchart of a method for extracting data from an image according to an embodiment of the present disclosure.
[0084] like Figure 5 As shown, the data extraction method 500 for an image may include operations S510 to S530.
[0085] In operation S510 , coordinates of N marking points located on a coordinate axis in the target image are determined according to a pixel value of each pixel in the target image.
[0086] For example, the target image can be e.g. Figure 4C The line chart in 405.
[0087] For example, the marked points can be the endpoints of the scale lines on the coordinate axis.
[0088] In the embodiment of the present disclosure, pixel analysis may be performed on the target image to obtain the pixel value of each pixel.
[0089] For example, based on the results of pixel analysis, it can be determined that the target image contains multiple consecutive pixels with the same pixel value. Multiple horizontal and vertical line segments in the target image can then be identified. The longest, mutually perpendicular line segments are used as coordinate axes, such as the longest horizontal segment as the horizontal axis and the longest vertical segment as the vertical axis. The origin is then determined based on the coordinate axes. In one example, the intersection of the horizontal and vertical axes can be used as the origin.
[0090] In the embodiment of the present disclosure, the coordinate axis includes M pixels.
[0091] For example, the horizontal axis includes M pixels.
[0092] In the embodiment of the present disclosure, M pixels in each row of K rows of pixels closest to the coordinate axis may be obtained; K≥1.
[0093] For example, the K rows of pixels closest to the horizontal axis may be above or below the horizontal axis. For another example, the K rows of pixels closest to the coordinate axis may be to the left or to the right of the vertical axis.
[0094] For example, if K=2, M pixels in the first row and M pixels in the second row of pixels in the two rows closest to the horizontal axis can be obtained. The first row of pixels is closest to the horizontal axis, followed by the second row of pixels.
[0095] In the embodiment of the present disclosure, the similarity between the jth pixel of the coordinate axis and the jth pixel of the first row of pixels is calculated, and the similarity between the jth pixel of the coordinate axis and the jth pixel of the second row of pixels is calculated.
[0096] For example, in response to the similarity between the jth pixel on the coordinate axis and the jth pixel in each row of pixels being greater than a preset similarity threshold, the similarity between the j-1th pixel on the coordinate axis and the j-1th pixel in each row of pixels being less than a preset similarity threshold, and the similarity between the j+λth pixel on the coordinate axis and the j+λth pixel in each row of pixels being less than a preset similarity threshold, the jth pixel on the coordinate axis is determined to be a marked point; j = 2, ..., M; λ is a preset value, and λ is a natural number. In one example, the coordinate axis can be a horizontal axis. In another example, λ can be 1. In another example, the preset similarity threshold can be 50%. In another example, the scale line is wider and λ can be 3.
[0097] In operation S520 , a division operation is performed on the target image according to the coordinates of the N marking points to obtain N+1 sub-regions.
[0098] For example, N vertical lines passing through N marking points on the horizontal axis are generated to obtain N+1 sub-regions.
[0099] In operation S530 , a text recognition operation is performed on the i-th sub-region among the N+1 sub-regions to obtain an i-th group of data corresponding to the i-th sub-region.
[0100] In the embodiment of the present disclosure, i=1, ..., N+1.
[0101] For example, if there are five tick marks on the horizontal axis, six subregions can be obtained. Each subregion corresponds to a value on the horizontal axis. After performing text recognition, data can be extracted from the subregions, and a corresponding relationship between the values on the horizontal axis and the data extracted from the subregions can be established.
[0102] It should be noted that the text recognition operation may be an OCR (optical character recognition) operation or other operations. For example, pixel analysis may be used to determine the inflection point or endpoint of a line graph in a sub-region, and then determine the distance from the inflection point or endpoint to the horizontal axis to determine the data corresponding to the inflection point or endpoint.
[0103] Through the embodiments of the present disclosure, data can be accurately extracted from a bar chart or a line chart, and in particular, data can be accurately extracted from a bar chart or a line chart in a paper with two columns of text.
[0104] Figures 6A to 6B It is a schematic diagram of a method for extracting data from an image according to an embodiment of the present disclosure.
[0105] like Figure 6AAs shown, based on the pixel value of each pixel in the target image 601, the coordinates of N marking points in the target image 601 can be determined. For example, the coordinates of 5 marking points in the target image can be determined. One of the marking points is the endpoint of the first scale line 605 on the horizontal axis 602.
[0106] For example, it can be found that there are multiple consecutive pixels with the same pixel value in the target image, and then multiple horizontal and vertical line segments in the target image can be determined, and the longer and mutually perpendicular line segments are used as coordinate axes, such as the longest horizontal line segment as the horizontal axis 602 and the longest two vertical line segments as the first vertical axis 603 and the second vertical axis 604. Figure 6A As shown, the horizontal axis 602 is the longest of the horizontal line segments, while the first vertical axis 603 and the second vertical axis 604 are the longest of the vertical line segments. The origin can be determined based on the coordinate axes, and the intersection of the horizontal axis 602 and the first vertical axis 603 can be used as the origin.
[0107] After determining the horizontal axis 602 from the target image 601, the scale lines on the horizontal axis can be determined by referring to the method described above regarding operation S510. The endpoints of the scale lines on the horizontal axis 602 can be used as marking points, resulting in five marking points.
[0108] like Figure 6B As shown in the figure, the target image is divided into six sub-regions based on the coordinates of the five markers. The six sub-regions correspond to the six values on the horizontal axis, namely 2014, 2015, 2016, 2017, 2018, and 2019.
[0109] Then, by performing text recognition operations on the six sub-regions respectively, six data can be obtained, which are 127, 163, 283, 113, 189, and 435 respectively.
[0110] Finally, 6 groups of structured data (2014, 127), (2015, 163), (2016, 283), (2017, 113), (2018, 189), and (2019, 435) can be extracted as the results of text recognition.
[0111] Figure 7 is a block diagram of a document processing apparatus according to an embodiment of the present disclosure.
[0112] like Figure 7 As shown, the apparatus 700 may include a first generating module 710 , a second generating module 720 , a merging module 730 , a first determining module 740 and a third generating module 750 .
[0113] The first generating module 710 is configured to generate a plurality of first bounding boxes according to the position information of the text image in the document page. In some embodiments, the first generating module 710 can be configured to perform the above operation S210, which will not be described in detail in this disclosure.
[0114] The second generating module 720 is configured to generate a plurality of second bounding boxes based on the position information of the plurality of first bounding boxes, each second bounding box being used to mark a text-sparse region in the document page. In some embodiments, the second generating module 720 can be configured to perform the aforementioned operation S220, which will not be further described in this disclosure.
[0115] The merging module 730 is configured to perform a merging operation on adjacent second bounding boxes to obtain multiple candidate bounding boxes. In some embodiments, the merging module 730 may be configured to perform the above-mentioned operation S230, which will not be described in detail herein.
[0116] The first determination module 740 is configured to determine, for the plurality of candidate bounding boxes, a plurality of partial images of the document page according to the position information of each candidate bounding box. In some embodiments, the first determination module 740 may be configured to perform the aforementioned operation S240, which will not be described in detail herein.
[0117] The third generating module 750 is configured to generate a target image based on the contents of the plurality of partial images. In some embodiments, the third generating module 750 may be configured to perform the above operation S250, which will not be described in detail herein.
[0118] In some embodiments, the above-mentioned second enclosing box includes a first sub-enclosing box and a second sub-enclosing box; the above-mentioned second generation module includes: a first generation unit, used to generate a first sub-enclosing box between any two upper and lower adjacent first enclosing boxes; and a second generation unit, used to generate a second sub-enclosing box on the left and / or right side of each of the above-mentioned first enclosing boxes.
[0119] In some embodiments, the candidate bounding box is a rectangle; the width of the second sub-bounding box is the length from the edge of the first bounding box to the edge of the document page; the merging module includes: a division unit, used to perform a division operation on each first sub-bounding box to obtain multiple third sub-bounding boxes with the same width as the second sub-bounding box; and a merging unit, used to perform a merging operation on the third sub-bounding box and the second sub-bounding box to obtain a candidate bounding box.
[0120] In some embodiments, the above-mentioned second generation module includes: a first determination unit, used to determine at least one overlapping area for each first enclosing box among the above-mentioned multiple first enclosing boxes based on the position information of the first enclosing box and the position information of the above-mentioned second enclosing box, and one overlapping area corresponds to at least one of the above-mentioned second enclosing boxes; and a removal unit, used to remove the overlapping area from at least one of the above-mentioned second enclosing boxes corresponding to each overlapping area, respectively, to obtain multiple adjusted second enclosing boxes.
[0121] In some embodiments, the first generating module is further configured to generate a plurality of first bounding boxes for each line of text in the document page according to the position information of each text image and the height of each text image in the document page.
[0122] Figure 8 is a block diagram of a data extraction apparatus for an image according to another embodiment of the present disclosure.
[0123] like Figure 8 As shown, the apparatus 800 may include a second determination module 810 , a division module 820 and a text recognition module 830 .
[0124] The second determining module 810 is configured to determine the coordinates of N marking points located on the coordinate axis in the target image based on the pixel value of each pixel in the target image. In some embodiments, the second determining module 810 can be configured to perform the above operation S510, which will not be described in detail herein.
[0125] The division module 820 is configured to perform a division operation on the target image according to the coordinates of the N marker points to obtain N+1 sub-regions. In some embodiments, the division module 820 may be configured to perform the above operation S820, which will not be described in detail in this disclosure.
[0126] The text recognition module 830 is configured to perform a text recognition operation on the i-th sub-region among the N+1 sub-regions, thereby obtaining an i-th set of data corresponding to the i-th sub-region; i=1, ..., N+1. In some embodiments, the text recognition module 830 may be configured to perform the above-described operation S530, which is not further described in this disclosure.
[0127] The target image is based on, for example Figure 7 Generated by the document processing device in.
[0128] In some embodiments, the above-mentioned coordinate axis includes M pixels; the above-mentioned second determination module includes: an acquisition unit, used to acquire M pixels in each row of K rows of pixels closest to the above-mentioned coordinate axis; K≥1; a second determination unit, used to determine that the jth pixel on the above-mentioned coordinate axis is a marking point in response to the similarity between the jth pixel on the above-mentioned coordinate axis and the jth pixel in each row of pixels being greater than a preset similarity threshold, the similarity between the j-1th pixel on the above-mentioned coordinate axis and the j-1th pixel in each row of pixels being less than the preset similarity threshold, and the similarity between the j+λth pixel point on the above-mentioned coordinate axis and the j+λth pixel in each row of pixels being less than the preset similarity threshold; j=2,...,M; λ is a preset value, and λ is a natural number.
[0129] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0130] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0131] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0132] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0133] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0134] The computing unit 901 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the document processing method or the data extraction method for an image. For example, in some embodiments, the document processing method or the data extraction method for an image can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the document processing method or the data extraction method for an image described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the document processing method or the data extraction method for an image in any other appropriate manner (eg, by means of firmware).
[0135] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0136] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0137] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0138] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0139] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0140] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0141] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0142] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A document processing method, comprising: generating a plurality of first bounding boxes based on position information of a text image in a document page, wherein the document page includes at least two columns of text; generating a plurality of second bounding boxes according to the position information of the plurality of first bounding boxes, each second bounding box being used to mark a text-sparse area in the document page, the second bounding box including a first sub-bounding box and a second sub-bounding box, the width of the second sub-bounding box being the length from the edge of the first bounding box to the edge of the document page; For each first enclosing frame in the plurality of first enclosing frames, determining at least one overlapping region according to position information of the first enclosing frame and position information of the second enclosing frame, where one overlapping region corresponds to at least one second enclosing frame; For at least one overlapping area, remove the overlapping area from at least one second bounding frame corresponding to each overlapping area to obtain a plurality of adjusted second bounding frames; performing a merging operation on a plurality of adjacent second bounding boxes to obtain a plurality of candidate bounding boxes, wherein the plurality of adjacent second bounding boxes include the adjusted second bounding box; For the multiple candidate bounding boxes, determining multiple partial images of the document page according to the position information of each candidate bounding box; as well as generating a target image according to the contents of the plurality of partial images, The step of generating a plurality of second bounding boxes according to the position information of the plurality of first bounding boxes includes: generating a first sub-enclosing frame between any two vertically adjacent first enclosing frames; and A second sub-bounding box is generated on the left and / or right side of each of the first bounding boxes.
2. The method according to claim 1, wherein The candidate bounding box is a rectangle; The performing a merging operation on the plurality of adjacent second bounding boxes to obtain the plurality of candidate bounding boxes comprises: Performing a splitting operation on each first sub-enclosing frame to obtain a plurality of third sub-enclosing frames having the same width as the second sub-enclosing frame; and A merging operation is performed on the third sub-bounding box and the second sub-bounding box to obtain a candidate bounding box.
3. The method according to claim 1, wherein Generating a plurality of first bounding boxes according to position information of the text image in the document page includes: A plurality of first bounding boxes are generated for each line of text in the document page according to the position information of each text image in the document page and the height of each text image.
4. A method for extracting data from an image, comprising: Determine the coordinates of N marked points on the coordinate axis in the target image according to the pixel value of each pixel in the target image; Performing a division operation on the target image according to the coordinates of the N marking points to obtain N+1 sub-regions; Performing a text recognition operation on the i-th sub-region among the N+1 sub-regions to obtain an i-th group of data corresponding to the i-th sub-region; i=1, ..., N+1; The target image is generated according to the document processing method according to any one of claims 1 to 3.
5. The method according to claim 4, wherein The coordinate axis includes M pixels; Determining the coordinates of N marking points located on a coordinate axis in the target image according to the pixel value of each pixel in the target image includes: Obtain M pixels in each row of K rows of pixels closest to the coordinate axis; K ≥ 1; In response to the similarity between the j-th pixel on the coordinate axis and the j-th pixel in each row of pixels being greater than a preset similarity threshold, the similarity between the j-1-th pixel on the coordinate axis and the j-1-th pixel in each row of pixels being less than the preset similarity threshold, and the similarity between the j+λ-th pixel on the coordinate axis and the j+λ-th pixel in each row of pixels being less than the preset similarity threshold, the j-th pixel on the coordinate axis is determined to be a marking point; j=2,…,M; λ is a preset value, and λ is a natural number.
6. A document processing device comprising: A first generating module, configured to generate a plurality of first bounding boxes based on position information of a text image in a document page, wherein the document page includes at least two columns of text; a second generating module, configured to generate a plurality of second bounding boxes based on the position information of the plurality of first bounding boxes, each second bounding box being used to mark a text-sparse area in the document page, the second bounding box comprising a first sub-bounding box and a second sub-bounding box, the width of the second sub-bounding box being the length from the edge of the first bounding box to the edge of the document page; a first determining unit configured to determine, for each first enclosing frame among the plurality of first enclosing frames, at least one overlapping region based on position information of the first enclosing frame and position information of the second enclosing frame, wherein one overlapping region corresponds to at least one second enclosing frame; a removing unit, configured to remove, for at least one overlapping area, the overlapping area from at least one second bounding frame corresponding to each overlapping area, to obtain a plurality of adjusted second bounding frames; a merging module, configured to perform a merging operation on a plurality of adjacent second bounding boxes to obtain a plurality of candidate bounding boxes, wherein the plurality of adjacent second bounding boxes include the adjusted second bounding box; A first determining module is configured to determine, for the plurality of candidate bounding boxes, a plurality of partial images of the document page according to position information of each candidate bounding box; as well as The third generating module is configured to generate a target image based on the contents of the plurality of partial images. The second generation module includes: A first generating unit is configured to generate a first sub-enclosing frame between any two vertically adjacent first enclosing frames; and The second generating unit is configured to generate a second sub-enclosing frame on the left side and / or the right side of each of the first enclosing frames.
7. The device according to claim 6, wherein The candidate bounding box is a rectangle; the width of the second sub-bounding box is the length from the edge of the first bounding box to the edge of the document page; The merging module includes: a dividing unit, configured to perform a dividing operation on each first sub-enclosing frame to obtain a plurality of third sub-enclosing frames having the same width as the second sub-enclosing frame; as well as A merging unit is configured to perform a merging operation on the third sub-bounding frame and the second sub-bounding frame to obtain a candidate bounding frame.
8. The device according to claim 6, wherein The first generating module is further configured to generate a plurality of first bounding boxes for each line of text in the document page according to the position information of each text image in the document page and the height of each text image.
9. A data extraction device for an image, comprising: A second determining module is used to determine the coordinates of N marking points located on the coordinate axis in the target image according to the pixel value of each pixel in the target image; A division module, configured to perform a division operation on the target image according to the coordinates of the N marking points to obtain N+1 sub-regions; A text recognition module is configured to perform a text recognition operation on an i-th sub-region among the N+1 sub-regions to obtain an i-th group of data corresponding to the i-th sub-region; i=1, ..., N+1; The target image is generated by the document processing device according to any one of claims 6 to 8.
10. The device according to claim 9, wherein The coordinate axis includes M pixels; The second determining module includes: An acquisition unit, configured to acquire M pixels in each row of K rows of pixels closest to the coordinate axis; K ≥ 1; A second determination unit is configured to determine that the jth pixel on the coordinate axis is a marking point in response to the similarity between the jth pixel on the coordinate axis and the jth pixel in each row of pixels being greater than a preset similarity threshold, the similarity between the j-1th pixel on the coordinate axis and the j-1th pixel in each row of pixels being less than the preset similarity threshold, and the similarity between the j+λth pixel on the coordinate axis and the j+λth pixel in each row of pixels being less than the preset similarity threshold; j=2,…,M; λ is a preset value, and λ is a natural number.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Document chart extraction method, electronic device and computer readable storage medium
CN107688788A
Columnar diagram recognition method and device, equipment and computer readable storage medium
CN110363092A