Document conversion method, device, computer equipment and computer-readable storage medium

By classifying and formatting the text detection box in the document image, the problems of information loss and format confusion in traditional document conversion are solved, and the complete restoration of document content is achieved.

CN113920510BActive Publication Date: 2025-07-18ZHAOYIN YUNCHUANG INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111136955.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-07-18
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

Traditional document conversion methods cannot accurately detect text categories and formats, resulting in lost information and confusing formats in document conversion.

Method used

By obtaining the text detection box from the document image, classifying it according to the properties of the text detection box, including size and position information, and writing text content into the editable document according to the category and write format information.

Benefits of technology

It realizes complete retention of information and accurate restoration of formats in document conversion, avoiding the problems of information loss and format confusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920510B_ABST
    Figure CN113920510B_ABST
Patent Text Reader

Abstract

The present application relates to a document conversion method, apparatus, computer device, and computer-readable storage medium. The method includes: obtaining a text detection box from a document image corresponding to an original document, where the text detection box includes text content; classifying the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes; writing the text content in the text detection boxes into an editable document according to the position information and the writing format information of the text detection boxes of different categories to generate a first conversion document. The classification of the text detection boxes is realized, and for text detection boxes of different categories, according to the position information and the writing format information of the text detection boxes, the text content in the text detection boxes is written into the editable document. Therefore, adopting this method can effectively solve the problems of information loss and format chaos in document conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet information technology, and particularly to a document conversion method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] At present, with the rapid development of Internet information technology, people have higher and higher requirements for document processing. Among them, document conversion, as a key part of document processing, has received extensive attention.

[0003] Document conversion is a commonly used document operation in daily life. Usually, document conversion is to convert non-editable electronic documents such as PDF format into editable documents. The traditional document conversion method first converts the electronic document into a picture format, then uses related technologies such as text recognition to recognize the content in the picture, and finally writes the recognized text content into an editable document.

[0004] However, the traditional document conversion method can only recognize text content, and cannot accurately detect text categories, pictures, tables in the document and judge the corresponding text formats, which leads to problems of information loss and format chaos in the obtained text content. Summary of the Invention

[0005] Based on this, in view of the problems of information loss and format chaos in document conversion, a document conversion method, apparatus, computer device, and computer-readable storage medium are provided.

[0006] A document conversion method, the method includes:

[0007] Obtain a text detection box from a document image corresponding to an original document, where the text detection box includes text content;

[0008] Classify the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes;

[0009] Write the text content in the text detection box into an editable document to generate a first conversion document according to the position information of the text detection boxes of different categories and the write format information of the text detection boxes.

[0010] In one embodiment, a document conversion method further includes:

[0011] Obtain a non-text detection box from the document image according to the position information of the text detection box;

[0012] Write the content in the non-text detection box into the first conversion document according to the position information of the non-text detection box to generate a second conversion document.

[0013] In one embodiment, different categories of text detection frames include a title text detection frame and a paragraph text detection frame, and the attributes of the text content in the title text detection frame and the paragraph text detection frame are different; classifying the text detection frames according to the attributes of the text detection frames to obtain different categories of text detection frames, including:

[0014] Classifying the text detection frames according to the size and position information of the text detection frames to obtain the title text detection frame;

[0015] Classifying the text detection frames according to the size and position information of the text detection frames to obtain the paragraph text detection frame.

[0016] In one embodiment, the size of the text detection frame includes the height of the text detection frame; classifying the text detection frames according to the size and position information of the text detection frames to obtain the title text detection frame, including:

[0017] Judging whether the height of the text detection frame conforms to the height range of the preset level title text detection frame; the preset level title text detection frame includes at least one level of title text detection frame;

[0018] If so, judging whether the distance between the text detection frame and the edge of the document image conforms to the first preset distance range;

[0019] If so, taking the text detection frame as the preset level title text detection frame.

[0020] In one embodiment, the size of the text detection frame includes the height of the text detection frame; the position information of the text detection frame includes the starting position information of the text detection frame; classifying the text detection frames according to the size and position information of the text detection frames to obtain the paragraph text detection frame, including:

[0021] Judging whether the height of the text detection frame conforms to the height range of the preset paragraph text detection frame;

[0022] If so, judging whether the distance between the text detection frames is less than the first preset threshold according to the position information of the text detection frames, and taking the text detection frames with a distance less than the first preset threshold as the intermediate text detection frames;

[0023] Obtaining, from the intermediate text detection frames, the text detection frames whose distance between the starting position information of the intermediate text detection frames and the edge of the document image conforms to the second preset distance range as the starting text detection frames;

[0024] Obtaining the paragraph text detection frame according to the intermediate text detection frames between two adjacent starting text detection frames.

[0025] In one embodiment, a document conversion method further includes: obtaining, from a document image, writing format information corresponding to the text content in a text detection box for different categories of text detection boxes.

[0026] In one embodiment, obtaining a non-text detection box from a document image according to the position information of a text detection box includes:

[0027] Obtaining the position information of a text detection box and the position information of the next text detection box adjacent to the text detection box;

[0028] Judging whether the distance between the text detection box and the next text detection box is greater than a preset distance threshold according to the position information of the text detection box and the position information of the next text detection box;

[0029] If so, dividing the area between the text detection box and the next text detection box into a non-text detection box.

[0030] In one embodiment, writing the content in a non-text detection box into a first conversion document according to the position information of the non-text detection box to generate a second conversion document includes:

[0031] Obtaining the writing format information of the non-text detection box from the document image;

[0032] Writing the content of the non-text detection box into the first conversion document according to the position information of the non-text detection box and the writing format information of the non-text detection box to generate a second conversion document.

[0033] In one embodiment, the non-text detection box includes at least one of a figure detection box, a table detection box, a figure caption detection box, a table caption detection box, a formula detection box, a header detection box, and a footer detection box.

[0034] In one embodiment, obtaining a text detection box from a document image corresponding to an original document includes

[0035] Converting the format of the original document to generate a document image corresponding to the original document;

[0036] Inputting the document image into a preset text detection model for text box detection to generate a plurality of text detection boxes; the preset text detection model includes a CTPN neural network model.

[0037] A document conversion device, the device includes:

[0038] A text detection box obtaining module, configured to obtain a text detection box from a document image corresponding to an original document, and the text detection box includes text content;

[0039] A text detection box classification module, which is used to classify text detection boxes according to the attributes of the text detection boxes, so as to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes;

[0040] A text detection box writing module, which is used to write the text content in the text detection box into an editable document according to the position information of the text detection boxes of different categories and the writing format information of the text detection boxes to generate a first conversion document.

[0041] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0042] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0043] For the above document conversion method, device, computer device and computer-readable storage medium, first, text detection boxes are obtained from the document image corresponding to the original document, and the text detection boxes include text content. Then, the text detection boxes are classified according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes. Finally, according to the position information of the text detection boxes of different categories and the writing format information of the text detection boxes, the text content in the text detection boxes is written into an editable document to generate a first conversion document. The classification of the text detection boxes is realized, and for text detection boxes of different categories, according to the position information of the text detection box and the writing format information of the text detection box, the text content in the text detection box is written into an editable document. It is possible to restore the text content in the document detection box according to the category of the text detection box, without missing the restoration of any category of text detection box, and the text content restored from text detection boxes of different categories has different writing formats. Therefore, the problems of information loss and format confusion in document conversion are effectively solved, and the text content in the document is more completely retained. Brief Description of the Drawings

[0044] Figure 1 It is an application environment diagram of the document conversion method in an embodiment;

[0045] Figure 2 It is a flowchart of the document conversion method in an embodiment;

[0046] Figure 3 It is a flowchart of the document conversion method in an embodiment;

[0047] Figure 4 It is a flowchart of obtaining text detection boxes of different categories in an embodiment;

[0048] Figure 5 Schematic diagram of the process for obtaining the title text detection frame in an embodiment;

[0049] Figure 6 Schematic diagram of the process for determining the first-level title text detection frame in an embodiment;

[0050] Figure 7 Schematic diagram of the first-level title text detection frame in an embodiment;

[0051] Figure 8 Schematic diagram of the process for obtaining the paragraph text detection frame in an embodiment;

[0052] Figure 9 Schematic diagram of the paragraph text detection frame in an embodiment;

[0053] Figure 10 Schematic diagram of the process for obtaining the non-text detection frame in an embodiment;

[0054] Figure 11 Schematic diagram of the process for generating the second conversion document in an embodiment;

[0055] Figure 12 Schematic diagram of the process for obtaining the figure detection frame and the table detection frame in an embodiment;

[0056] Figure 13 Schematic diagram of the process for obtaining the text detection frame in an embodiment;

[0057] Figure 14 Schematic diagram of the CTPN model in an embodiment;

[0058] Figure 15 Schematic diagram of the training process of the preset text detection model in an embodiment;

[0059] Figure 16 Schematic diagram of the document conversion method in a specific embodiment;

[0060] Figure 17 Structural block diagram of the document conversion device in an embodiment;

[0061] Figure 18 Structural block diagram of the text detection frame classification module in an embodiment;

[0062] Figure 19 Internal structural diagram of the computer device in an embodiment. Detailed implementation manners

[0063] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] Figure 1 FIG. is an application scenario diagram of document conversion in an embodiment. As Figure 1 shown, the application environment includes a computer device 140. The computer device 140 first obtains a text detection box from the document image corresponding to the original document, and the text detection box includes text content. Then, the computer device 140 classifies the text detection box according to the attributes of the text detection box to obtain text detection boxes of different categories; the attributes of the text detection box include the size and position information of the text detection box. Finally, the computer device 140 writes the text content in the text detection box into an editable document according to the position information of the text detection boxes of different categories and the writing format information of the text detection box to generate a first conversion document.

[0065] Figure 2 FIG. is a schematic flowchart of a document conversion method in an embodiment. As Figure 2 shown, a document conversion method is provided, which is applied to a computer device and includes steps 210 to 230.

[0066] S210. Obtain a text detection box from the document image corresponding to the original document, and the text detection box includes text content.

[0067] Specifically, the original document is an uneditable document, such as a document in PDF format. The elements such as text, pictures, and tables in the uneditable document cannot be edited arbitrarily. Further, the original document is converted into a corresponding document image, and the format of the document image includes but is not limited to: PNG, JPG, TIFF, etc. Optionally, the present application converts the original document into a document image in PNG format. The text detection box is the outline of the included text content, and the outline is approximately rectangular. Further, the text content is the text in the document, and at least one text content is included in the text detection box, that is, at least one text is included in the text detection box.

[0068] Further, first convert the original document into a document image in the PNG format, and then obtain text detection boxes from the document image, where the text detection boxes include text content. Specifically, the present application does not limit the method for obtaining the text detection boxes. Further, the method for obtaining the text detection boxes includes, but is not limited to: the EAST (Efficient and Accuracy Scene Text detection pipeline) model, the FTSN (Fused Text Segmentation Networks) model, the RRPN (Rotation Region Proposal Networks) model, the DMPNet (Deep Matching Prior Network) model, the CTPN (Connectionist Text Proposal Network) model, etc.

[0069] Specifically, the EAST model first uses a fully convolutional network (FCN) to generate a multi-scale fused feature map, and then directly performs pixel-level text block prediction on this basis. In the EAST model, two text region annotation forms, namely rotated rectangular boxes and arbitrary quadrilaterals, are supported. Corresponding to the quadrilateral annotation, when the EAST model executes, it predicts the coordinate differences of each pixel in the feature map to the four vertices. Corresponding to the rotated rectangular box annotation, when the EAST model executes, it predicts the distances from each pixel in the feature map to the four sides of the rectangular box and the orientation angle of the rectangular box. Therefore, the EAST model can effectively detect rotated text. Further, the FTSN model uses a segmentation network to support inclined text detection. The FTSN model uses Resnet as the basic network and uses a multi-scale fused feature map. In addition, the annotation data includes the pixel mask and the bounding box of the text instance, and multi-object joint training of pixel prediction and bounding box detection is used. Further, the RRPN model incorporates the rotation factor into the classical region proposal network (such as Faster RCNN). The RRPN model uses a rotated region of interest pooling layer to first divide the region proposals in any direction into sub-regions, and then performs max pooling operations on these sub-regions respectively and projects the results onto a small feature map with a fixed spatial size. Further, the DMPNet model uses a quadrilateral (non-rectangular) to more compactly annotate the boundary of the text region, and the trained model has a better detection effect on inclined text blocks. Further, CTPN is one of the most widely used open-source text detection models. The CTPN model can detect horizontal or slightly inclined text lines. A text line can be regarded as a character sequence, rather than a single independent object in general object detection. The character images on the same text line can be context for each other. Letting the detection model learn the context statistical laws in the image during the training stage can effectively improve the prediction accuracy of text blocks in the prediction stage. In addition, in the image prediction process of the CTPN model, the popular VGG16 at that time is used as the basic network at the front end to extract the local image features of each character, the BLSTM layer is used in the middle to extract the context features of the character sequence, and then through the FC fully connected layer, and the coordinate values and classification result probability values of each text block are output at the end. In the data post-processing stage, adjacent small text blocks are merged into text lines.

[0070] S220. Classify the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes.

[0071] Among them, the size of the text detection box is related to the quantity of the included text content, and the position information of the text detection box includes the coordinate values of the four endpoints of the text detection box in the document image. Since the sizes and position information of text detection boxes of different categories are different, the text detection boxes can be classified according to the sizes and position information of the text detection boxes to obtain text detection boxes of different categories.

[0072] For example, text detection boxes of different categories include title text detection boxes and paragraph text detection boxes. Title text detection boxes generally have specific sizes and position information. Based on this specific size and position information, the text detection boxes can be classified as title text detection boxes. Correspondingly, paragraph text detection boxes generally also have specific size and position information different from those of title text detection boxes. Based on this specific size and position information, the text detection boxes can be classified as paragraph text detection boxes.

[0073] S230. Write the text content in the text detection box into an editable document according to the position information of text detection boxes of different categories and the write format information of the text detection box to generate a first conversion document.

[0074] Specifically, the text content included in text detection boxes of different categories has different write format information. The above write format information includes format information such as the font, font size, color, etc. of the text content. This application does not limit this. Further, the methods for obtaining the write format information corresponding to the text detection box include: open-source font recognition tools, neural network algorithms, etc. This application does not limit this. For example, OCR font recognition can achieve high-precision font recognition and detection. Currently, there are multiple open platforms for OCR character recognition for users to use. In addition, there are various font recognition websites in existing network resources. Only by using the text detection box image can the write format information corresponding to the included text content be detected. Further, based on the position information of text detection boxes of different categories, obtain the coordinate values of text detection boxes of different categories. Based on the coordinate values of text detection boxes of different categories, obtain the text detection box images at the corresponding positions in the document image. Based on the text detection box images, obtain the write format information of the text content in the text detection box. Write the text content in the text detection box into the corresponding position of the editable document according to the coordinate values of the text detection box and the write format information of the text detection box to generate a first conversion document. Further, this application does not limit the method of writing the text content into the editable document. Optionally, this application uses the open-source tool python-docx to write the text content of the text detection box into the editable document. Specifically, the open-source tool python-docx is a module in python. Python-docx can be used to create docx documents, which include almost all the commonly used functions in word documents such as paragraphs, page breaks, tables, pictures, titles, styles, etc.

[0075] In an embodiment of the present application, when converting a document, first obtain a text detection box from the document image corresponding to the original document, and the text detection box includes text content. Then classify the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes. Finally, according to the position information of the text detection boxes of different categories and the writing format information of the text detection boxes, write the text content in the text detection boxes into an editable document to generate a first converted document. By classifying the text detection boxes, text detection boxes of different categories are obtained, and according to the position information of the text detection box and the writing format information of the text detection box, the text content in the text detection box is written into the editable document. Furthermore, the text content in the document detection box can be restored according to the category of the text detection box, and the text content of any category of text detection box will not be missed. Therefore, the problems of information loss and format confusion in document conversion are effectively solved, and the text content in the document is more completely retained.

[0076] In one of the embodiments, Figure 3 is a schematic flowchart of a document conversion method in an embodiment, as Figure 3 shown, a document conversion method is provided, which further includes steps 240 to 250.

[0077] Step 240, obtain a non-text detection box from the document image according to the position information of the text detection box.

[0078] Specifically, the position information of the text detection box includes the coordinate values of the four endpoints of the text detection box. Generally speaking, in addition to paragraph text content and title text content in the document, there are also contents such as figures, tables, figure captions, table captions, formulas, headers, and footers. Further, since the text detection box only includes text content such as paragraph text content and title text content, and does not include contents such as figures, tables, figure captions, table captions, formulas, headers, and footers, in the document image, the area outside the text detection box includes the above-mentioned contents such as figures, tables, figure captions, table captions, formulas, headers, and footers.

[0079] Further, according to the position information of the text detection box, a non-text detection box can be obtained from the document image. Specifically, based on the position information of two adjacent text detection boxes, the coordinate values of the two adjacent text detection boxes are obtained. Based on the coordinate values of the two adjacent text detection boxes, the coordinate values of the area between the two adjacent text detection boxes can be obtained. Further, based on the coordinate values of the area between the two adjacent text detection boxes, the area at the corresponding position in the document image is used as the non-text detection box.

[0080] Step 250: Write the content in the non-text detection box into the first conversion document according to the position information of the non-text detection box to generate a second conversion document.

[0081] Specifically, the position information of the non-text detection box includes the coordinate values of the four endpoints of the non-text detection box. Further, based on the position information of the non-text detection box, obtain the coordinate values of the non-text detection box. Based on the coordinate values of the non-text detection box, intercept the corresponding non-text detection box image in the document image, and then write the content in the non-text detection box image into the first conversion document to generate a second conversion document.

[0082] Further, for the content of the non-text detection box, it can be written into an editable document through existing open-source tools. The present application does not limit the writing method of the content in the non-text detection box. Preferably, if the content of the non-text detection box is a table, first use the OpenCV open-source tool to perform Hough line detection and vertex detection on the table area to obtain the coordinates of the intersections of the lines and the table. Then, intercept the cells according to the coordinate information of the intersections of the lines and the table and use the open-source tool pytesser to recognize the text. Further, obtain the size and text information of each cell according to the coordinate information of the intersections of the lines and the table, and use the open-source tool python-docx to draw a table in the word document. Specifically, the OpenCV open-source tool is a computer vision library developed for application developers. OpenCV includes C / C++ versions and Python versions, which contain more than 300 C / C++ programs and Python programs for image processing and computer vision. In addition, OpenCV contains a large number of functions to handle common problems in the field of computer vision, such as motion analysis and tracking, face recognition, 3D reconstruction, and object recognition. The pytesser open-source tool is a module of the Google OCR open-source project. Importing this module in python can convert the text in the picture into text. The python-docx open-source tool is a module in python. Python-docx can be used to create docx documents, which contain almost all the common functions in word documents, such as paragraphs, page breaks, tables, pictures, headings, styles, etc.

[0083] In the embodiment of the present application, first obtain the non-text detection box from the document image according to the position information of the text detection box, and then write the content in the non-text detection box into the corresponding position of the first conversion document according to the position information of the non-text detection box, and finally obtain a second conversion document. The non-text detection box is obtained, and the content in the non-text detection box is written into the editable document without omission, which can effectively solve the problem of information loss of non-text content during document conversion.

[0084] In one embodiment, as Figure 4 shown, the text detection boxes of different categories include a title text detection box and a paragraph text detection box, and the attributes of the text content in the title text detection box and the paragraph text detection box are different. Classifying the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories includes steps 222 to 224. Among them, S222, classifying the text detection boxes according to the size and position information of the text detection boxes to obtain the title text detection boxes.

[0085] Specifically, the size of the text detection box may specifically include the height and length information of the text detection box, and the position information of the text detection box may include the coordinate information of the four endpoints in the text detection box. Among them, the height and length information of the text detection box can be obtained according to the coordinate information of the four endpoints in the text detection box.

[0086] Further, a document includes a title and paragraphs. Therefore, the text detection boxes of different categories obtained from the document image include a title text detection box and a paragraph text detection box. Among them, the title text detection box is a text detection box including the title, that is, the text content in the title text detection box is the title. The paragraph text detection box is a text detection box including paragraphs, that is, the text content in the paragraph text detection box is paragraphs. And the attributes of the text content in the title text detection box and the paragraph text detection box are different, that is, the title included in the title text detection box and the paragraphs included in the paragraph text detection box have different attributes. Specifically, it is reflected in that the position information of the title and the paragraphs in the document is different, and the font size and typeface of the title and the paragraphs are different.

[0087] Further, since the title in the title text detection box has specific attributes, such as the font size and typeface of the title. Further, since the title text detection box is the outline corresponding to the title, the size and typeface of the title in the title text detection box determine the size of the title text detection box, so the title text detection box has a specific height. Further, the title in the document has specific position information, and this position information includes the distance from the document edge. For example, for a title close to the left edge of the document, the distance between the title and the left edge of the document is a specific value, so the distance between this title text detection box and the left edge of the document should also be a specific value. Another example is that for a title close to the upper edge of the document, the distance between the title and the upper edge of the document is a specific value, so the distance between this title text detection box and the upper edge of the document should also be a specific value.

[0088] Further, it is possible to determine whether the text detection box is a title text detection box by judging whether the size of the text detection box meets the preset title height threshold range. If the size of the text detection box meets the preset title height threshold range, then the text detection box is a title text detection box. Otherwise, the text detection box is not a title text detection box. Among them, the size of the above-mentioned text detection box includes, but is not limited to, the height and length information of the text detection box.

[0089] Further, it is also possible to determine whether the text detection box is a title text detection box by judging whether the distance between the text detection box and the document edge meets the preset title distance threshold range. If the distance between the text detection box and the document edge meets the preset title distance threshold range, then the text detection box is a title text detection box. Otherwise, the text detection box is not a title text detection box. Among them, the distance between the above-mentioned text detection box and the document edge includes, but is not limited to, the distance between the text detection box and the upper edge of the document, and the distance between the text detection box and the left edge of the document.

[0090] The method for determining whether the text detection box is a title text detection box according to the size and position information of the text detection box includes, but is not limited to, the methods mentioned above, and the present application does not limit this.

[0091] S224. Classify the text detection box according to the size and position information of the text detection box to obtain a paragraph text detection box.

[0092] Further, since the text content in the paragraph text detection box has specific attributes, such as the size and font size of the font in the paragraph. Further, since the paragraph text detection box is the outline of the paragraph, the size and font size of the paragraph in the paragraph text detection box determine the size of the paragraph text detection box, so the paragraph text detection box has a specific height. Further, the paragraphs in the document have specific position information, and this position information includes that the starting line of the paragraph is indented by two characters, and there is a specific line spacing between two adjacent lines in the paragraph.

[0093] Further, it is possible to determine whether the text detection box is a paragraph text detection box by judging whether the size of the text detection box meets the preset paragraph height threshold range. If the size of the text detection box meets the preset paragraph height threshold range, then the text detection box is a paragraph text detection box. Otherwise, the text detection box is not a paragraph text detection box. Among them, the size of the above-mentioned text detection box includes, but is not limited to, the height and length information of the text detection box.

[0094] Further, it is also possible to determine whether the text detection frame is a paragraph text detection frame by judging whether the distance between adjacent text detection frames meets the range of the first preset paragraph distance threshold. If the size of the text detection frame meets the range of the first preset paragraph distance threshold, then the text detection frame is a paragraph text detection frame. Otherwise, the text detection frame is not a paragraph text detection frame.

[0095] Further, it is also possible to determine whether the text detection frame is at the beginning or end paragraph by judging whether the distance between the text detection frame and the document edge meets the range of the second preset paragraph distance threshold. If the distance between the text detection frame and the document edge meets the range of the second preset paragraph distance threshold, then the text detection frame is at the beginning or end paragraph. Otherwise, the text detection frame is not at the beginning or end paragraph. Among them, the distance between the text detection frame and the document edge includes, but is not limited to: the distance between the text detection frame and the left edge of the document and the distance between the text detection frame and the right edge of the document.

[0096] Further, it is possible to determine the beginning or end paragraph of the paragraph, and then divide the text detection frame into paragraphs to obtain paragraph text detection frames. The above paragraph text detection frames include a complete paragraph.

[0097] The method for judging whether the text detection frame is a paragraph text detection frame according to the size and position information of the text detection frame includes, but is not limited to, the methods mentioned above. This application does not make any limitations in this regard.

[0098] In the embodiments of the present application, the detection frames are classified according to the size and position information of the text detection frames to obtain title text detection frames and paragraph text detection frames. By classifying the text detection frames in the document, the title text content and paragraph text content in the title text detection frames and paragraph text detection frames can be written into an editable document, and the text content of the title text and paragraph text in the document is more completely retained.

[0099] In one of the embodiments, as Figure 5 shown, the size of the text detection frame includes the height of the text detection frame. Classifying the text detection frames according to the size and position information of the text detection frames to obtain title text detection frames includes steps 320 to 340. Among them,

[0100] S320, judging whether the height of the text detection frame conforms to the height range of the preset level title text detection frame; the preset level title text detection frame includes at least one level of title text detection frame.

[0101] Specifically, the height of the text detection box is obtained according to the position information of the text detection box. Specifically, the size of the text detection box may specifically include the height and length information of the text detection box; the position information of the text detection box may include the coordinate information of the four endpoints in the text detection box. Further, the text detection box is approximately rectangular. Therefore, the coordinate information of the text detection box is the coordinates of the four endpoints, which can be expressed as:

[0102] Ω m,n ={(x i,j ,y i,j ),(x i,j+1 ,y i,j+1 ),(x i+1,j ,y i+1,j ),(x i+1,j+1 ,y i+1,j+1 )} (1)

[0103] Among them, the text detection boxes in the document are arranged in the order from top to bottom and from left to right. Therefore, Ω m,n represents the text detection box corresponding to the m-th row and the n-th column in the document. The coordinates (x i,j ,y i,j ) are the coordinate values of the upper left endpoint in the text detection box, the coordinates (x i,j+1 ,y i,j+1 ) are the coordinate values of the lower left endpoint in the above text detection box, the coordinates (x i+1,j ,y i+1,j ) are the coordinate values of the upper right endpoint in the above text detection box, and the coordinates (x i+1,j+1 ,y i+1,j+1 ) are the coordinate values of the lower right endpoint in the text detection box. Therefore, the height of the text detection box can be obtained through the coordinate information of the text detection box:

[0104] Δy=y i,j+1 -y i,j (2)

[0105] Among them, Δy is the height of the text detection box, y i,j+1 is the larger ordinate value, and y i,j is the smaller ordinate value. Similarly, the length of the text detection box can be obtained through the coordinate information of the text detection box:

[0106] Δx=x i+1,j -x i,j (3)

[0107] Among them, Δx is the length of the text detection box, x i+1,j is the larger abscissa value, and x i,j is the smaller abscissa value.

[0108] Further, the height Δy of the text detection box can be obtained through Equation (1) and Equation (2).

[0109] Further, it is determined whether the height of the text detection box conforms to the height range of the preset level title text detection box. Specifically, the above-mentioned preset level title text detection box includes at least one level of title text detection box. The font sizes corresponding to the title texts of different levels in the document are different. For example, the first-level title, the second-level title, and the third-level title usually adopt different font sizes. Further, the font size affects the size of the title text detection box. Therefore, it can be determined which level of title text detection box the text detection box belongs to by determining whether the height of the text detection box conforms to the height range of the preset level title text detection box. Further, if the height of the text detection box conforms to the height range of the preset level title text detection box, the next step is performed; otherwise, the text detection box does not belong to the preset level title text detection box.

[0110] S340. If so, it is determined whether the distance between the text detection box and the edge of the document image conforms to the first preset distance range. If so, the text detection box is used as the preset level title text detection box.

[0111] Specifically, the distances between the title texts of different levels in the document and the edge of the document are different. For example, for the left-aligned title text, the distance between the title text and the left edge of the document is a fixed value, and the distances between the title texts of different levels and the left edge of the document are different. Therefore, by determining whether the distance between the text detection box and the edge of the document image conforms to the first preset distance range, it is further determined which level of title text detection box the text detection box belongs to. Further, if the distance between the text detection box and the edge of the document image conforms to the first preset distance range, the text detection box belongs to the preset level title text detection box; otherwise, the text detection box does not belong to the preset level title text detection box.

[0112] Specifically, the preset level title text detection box in this application includes a first-level title text detection box and a second-level title text detection box. The height range of the preset level title text detection box includes the height range of the preset first-level title text detection box and the height range of the preset second-level title text detection box. The first preset distance range includes the preset first-level title distance range and the preset second-level title distance range.

[0113] Specifically, Figure 6 is a schematic flowchart for determining the first-level title text detection box in an embodiment. As Figure 6 shown, a method for determining the first-level title text detection box is provided, including steps 420 to 460.

[0114] S420. Obtain the height of the text detection box according to the position information of the text detection box.

[0115] Specifically, the height of the text detection box can be obtained through Equation (1) and Equation (2). Further, Figure 7 is a schematic diagram for judging the text detection box of the first-level title in an embodiment. As Figure 7 shown, the text detection boxes in the document image 510 are arranged in the order from top to bottom and from left to right. Therefore, the position information of the document image edge 520 is determined through the coordinate values of each text detection box. In addition, as Figure 7 shown, the coordinate values of the text detection box are determined based on the coordinate system as Figure 7 shown. That is to say, the upper left end of the document image 510 is the origin of the coordinate system, and the coordinate value of the upper left end of the document image 510 is (0, 0). Further, based on the position information of the text detection box, obtain the maximum abscissa, minimum abscissa, maximum ordinate, and minimum ordinate of all text detection boxes in the document image 510. Further, a rectangular box can be determined from the maximum abscissa, minimum abscissa, maximum ordinate, and minimum ordinate, and this rectangular box is the document image edge 520.

[0116] S440. Judge whether the height of the text detection box conforms to the height range of the preset first-level title text detection box. If so, judge the distance between the text detection box and the document image edge.

[0117] Specifically, as Figure 7 shown, the height Δy1 of the text detection box 530 can be obtained through Equation (1) and Equation (2). The maximum value of the above preset height range of the first-level title text detection box is The minimum value of the above preset height range of the first-level title text detection box is Further, when the height Δy1 of the text detection box 530 conforms to the height range of the preset first-level title text detection box, the height Δy1 of the text detection box 530 satisfies the following conditions:

[0118]

[0119] It can be judged whether the height Δy1 of the text detection box 530 conforms to the height range of the preset first-level title text detection box from Equation (4). Specifically, when is set to 4 cm, When it is set to 3 cm, if the height Δy1 of the text detection frame 530 is 3.5 cm, then the height Δy1 of the text detection frame 530 conforms to the height range of the preset first-level heading text detection frame, and the next judgment is carried out. If the height Δy1 of the text detection frame 530 is 4.5 cm, then the height Δy1 of the text detection frame 530 does not conform to the height range of the preset first-level heading text detection frame, and the text detection frame 530 is not a first-level heading text detection frame.

[0120] S460. Determine whether the distance between the text detection frame and the edge of the document image conforms to the preset first-level heading distance range. If so, regard the text detection frame as the first-level heading text detection frame.

[0121] Specifically, as Figure 7 shown, the distance Δx1 between the text detection frame 530 and the edge 520 of the document image can be obtained through the minimum abscissa of the edge 520 of the document image and the minimum abscissa of the text detection frame. The maximum value of the above preset first-level heading distance range is The minimum value of the above preset first-level heading distance range is Further, when the distance Δx1 between the text detection frame 530 and the edge 520 of the document image conforms to the preset first-level heading distance range, the following conditions are satisfied:

[0122]

[0123] It can be judged from formula (5) whether the distance Δx1 between the text detection frame 530 and the edge 520 of the document image conforms to the preset first-level heading distance range. Specifically, when is set to 1.4 cm, is set to 1.6 cm, if the distance Δx1 between the text detection frame 530 and the edge 520 of the document image is 1.5 cm, then the distance Δx1 between the text detection frame 530 and the edge 520 of the document image conforms to the preset first-level heading distance range, and the text detection frame is the first-level heading text detection frame. If the distance Δx1 between the text detection frame 530 and the edge 520 of the document image is 1.3 cm, then the distance Δx1 between the text detection frame 530 and the edge 520 of the document image does not conform to the preset first-level heading distance range, and the text detection frame is not the first-level heading text detection frame.

[0124] The judgment process for the second-level heading text detection frame is similar to the judgment method for the first-level heading text detection frame, and the judgment method process of the above first-level heading text detection frame can be referred to, and it will not be elaborated in detail here.

[0125] In an embodiment of the present application, a method for determining a title text box is provided. First, it is determined whether the height of the text detection box conforms to the height range of the preset-level title text detection box; wherein the preset-level title text detection box includes at least one level of title text detection box; if so, it is determined whether the distance between the text detection box and the edge of the document image conforms to the first preset distance range; if so, the text detection box is used as the preset-level title text detection box. By judging different levels of title text detection boxes, the text content in different levels of title text detection boxes can be written into the editable document in an accurate format, and the problems of missing title text content and format confusion can be effectively solved.

[0126] In one embodiment, as Figure 8 shown, the size of the text detection box includes the height of the text detection box; the position information of the text detection box includes the starting position information of the text detection box; classifying the text detection box according to the size and position information of the text detection box to obtain a paragraph text detection box, including steps 620 to 680. Among them,

[0127] S620. Determine whether the height of the text detection box conforms to the height range of the preset paragraph text detection box.

[0128] Specifically, the height of the text detection box can be obtained through formulas (1) and (2). Further, Figure 9 is a schematic diagram for determining a paragraph text detection box in one embodiment. As Figure 9 shown, most of the text content in the text detection box in the document image 510 is paragraph text content except for the title text content of the title text box. The height Δy2 of the text detection box 570 can be obtained through formulas (1) and (2). The maximum value of the height range of the above-mentioned preset paragraph text detection box is The minimum value of the height range of the above-mentioned preset paragraph text detection box is Further, when the height Δy2 of the text detection box 570 conforms to the height range of the preset paragraph text detection box, the height Δy2 of the text detection box 570 satisfies the following conditions:

[0129]

[0130] It can be determined whether the height Δy2 of the text detection box 570 conforms to the height range of the preset paragraph text detection box from formula (6). Specifically, when is set to 0.8 cm, When set to 0.6 cm, if the height Δy2 of the text detection box 570 is 0.7 cm, then the height Δy2 of the text detection box 570 meets the height range of the preset paragraph text detection box, and the next judgment is carried out. If the height Δy2 of the text detection box 570 is 0.9 cm, then the height Δy2 of the text detection box 570 does not meet the height range of the preset paragraph text detection box, and the text detection box is not a paragraph text detection box.

[0131] S640. If so, according to the position information of the text detection boxes, determine whether the distance between the text detection boxes is less than the first preset threshold, and regard the text detection boxes with a distance less than the first preset threshold as intermediate text detection boxes.

[0132] Specifically, as Figure 9 shown, there is a fixed line spacing between adjacent two lines of paragraph text in the document. Therefore, by judging the line spacing between each paragraph text, all paragraph text detection boxes can be divided into different regions. Among them, the distance between each text detection box belonging to the same region is approximately the same. Specifically, by obtaining the position information of two adjacent text detection boxes, the line spacing Δh between two adjacent paragraph texts can be obtained. Further, for two adjacent text detection boxes, obtain the maximum ordinate value of the text detection box close to the horizontal axis x, and obtain the minimum ordinate value of the text detection box far from the horizontal axis x, and subtract the two to get the line spacing Δh between the two adjacent text detection boxes. Further, when the distance Δh between the text detection box 570 and the previous text detection box is less than the first preset threshold ξ, the text detection box 570 is an intermediate text detection box; otherwise, the text detection box 570 is not an intermediate text detection box. Specifically, when the first preset threshold ξ is set to 0.3 cm, if the distance Δh between the text detection box 570 and the previous text detection box is 0.6 cm, then the text detection box 570 is not an intermediate text detection box. If the distance Δh between the text detection box 570 and the previous text detection box is 0.26 cm, then the text detection box 570 is an intermediate text detection box, and the next judgment is carried out. As Figure 8 shown, after sequentially judging the distance Δh between two adjacent text detection boxes, the text detection boxes are divided into an intermediate text detection box 540 and an intermediate text detection box 550, where the intermediate text detection box 540 and the intermediate text detection box 550 include at least one text detection box.

[0133] S660. From the intermediate text detection boxes, obtain the text detection boxes whose distance between the starting position information of the intermediate text detection box and the edge of the document image meets the second preset distance range, and regard them as starting text detection boxes.

[0134] Specifically, as Figure 9As shown, for the intermediate text detection box 540, there is continuous paragraph text. Usually, the first line of a complete paragraph is indented by two characters. Therefore, by determining whether the distance Δx2 between the intermediate text detection box and the document edge meets the text detection box within the second preset distance range, the starting text detection box can be further determined. The maximum value of the above-mentioned second preset distance range is The minimum value of the above-mentioned second preset distance range is Furthermore, when the distance Δx2 between the text detection box 570 and the document image edge 520 meets the second preset distance range, the distance Δx2 between the text detection box 570 and the document image edge 520 satisfies the following conditions:

[0135]

[0136] It can be judged from Equation (7) whether the distance Δx2 between the text detection box 570 and the document image edge 520 meets the second preset distance range. Specifically, when is set to 1.5 cm, is set to 1.3 cm, if the distance Δx2 between the text detection box 570 and the document image edge 520 is 1.4 cm, then the distance Δx2 between the text detection box 570 and the document image edge 520 meets the second preset distance range, and the text detection box 570 is the starting detection box. If the distance Δx2 between the text detection box 570 and the document image edge 520 is 0.8 cm, then the distance Δx2 between the text detection box 570 and the document image edge 520 does not meet the second preset distance range, and the text detection box is not the starting detection box.

[0137] S680. Obtain the paragraph text detection box according to the intermediate text detection box between two adjacent starting text detection boxes.

[0138] Specifically, after determining the starting text detection box in the intermediate text detection box 540, first determine two adjacent starting text detection boxes from many starting text detection boxes. Then, determine all the intermediate text detection boxes between the two adjacent starting text detection boxes from the intermediate text detection boxes. Finally, divide all the intermediate text detection boxes and the former of the two adjacent starting text detection boxes into a paragraph text detection box 560.

[0139] In an embodiment of the present application, a method for obtaining a detection frame of paragraph text is provided. First, it is determined whether the height of the text detection frame conforms to the height range of the preset paragraph text detection frame. If so, according to the position information of the text detection frame, it is determined whether the distance between text detection frames is less than a first preset threshold, and the text detection frames with a distance less than the first preset threshold are used as intermediate text detection frames. Then, from the intermediate text detection frames, the text detection frames whose distance between the starting position information of the intermediate text detection frame and the edge of the document image conforms to a second preset distance range are obtained as starting text detection frames. Finally, based on the intermediate text detection frames between two adjacent starting text detection frames, a paragraph text detection frame is obtained. Based on the judgment of the intermediate text detection frame and the starting text detection frame, it can be ensured that the text content in the paragraph text detection frame is a complete paragraph, effectively solving the problem of missing text content in the paragraph text detection frame.

[0140] In one embodiment, for different types of text detection frames, writing format information corresponding to the text content in the text detection frame is obtained from the document image.

[0141] Specifically, the text content in different types of text detection frames has different format information, such as font, size and other format information. The present application does not limit the method for obtaining the writing format information. Further, for different types of text detection frames, according to the position information of the text detection frame, a text detection frame image at the corresponding position in the document image is intercepted, and the format information of the text content in the text detection frame image is recognized through a font recognition tool, so as to obtain the corresponding writing format information. Further, the methods for obtaining the writing format information in the present application include but are not limited to: deep learning algorithms, font recognition tools, etc. For example, OCR font recognition can achieve high-precision font recognition detection, and currently there are multiple open platforms for OCR text recognition for users to use. In addition, there are various font recognition websites in existing network resources, and the writing format information corresponding to the text content included can be detected only through the text detection frame image.

[0142] In an embodiment of the present application, for different types of text detection frames, writing format information corresponding to the text content in the text detection frame is obtained from the document image. Based on the obtained writing format information corresponding to the text content in different types of text detection frames, the text content in different types of text detection frames can be accurately written into an editable document according to the corresponding writing format information, effectively solving the problem of chaotic text content format of the text detection frame.

[0143] In one embodiment, as Figure 10 shown, according to the position information of the text detection frame, non-text detection frames are obtained from the document image, including steps 720 to 760. Among them,

[0144] S720. Obtain the position information of the text detection frame and the position information of the next text detection frame adjacent to the text detection frame.

[0145] Specifically, the position information of the text detection frame includes the coordinate values of the four endpoints of the text detection frame. Further, obtain the position information of the text detection frame, and based on the position information of the text detection frame, obtain the coordinate values of the text detection frame. Obtain the position information of the next text detection frame adjacent to the text detection frame, and based on the position information of the next text detection frame adjacent to the text detection frame, obtain the coordinate values of the next text detection frame adjacent to the text detection frame.

[0146] S740. Determine whether the distance between the text detection frame and the next text detection frame is greater than a preset distance threshold according to the position information of the text detection frame and the position information of the next text detection frame.

[0147] Specifically, based on the coordinate values of the text detection frame and the coordinate values of the next text detection frame adjacent to the text detection frame, obtain the distance between the text detection frame and the next text detection frame adjacent to the text detection frame. The present application does not limit the calculation method of the distance between two adjacent detection frames. Optionally, obtain the maximum ordinate value of the text detection frame based on the coordinate values of the text detection frame, and obtain the minimum ordinate value of the next text detection frame adjacent to the text detection frame based on the coordinate values of the next text detection frame adjacent to the text detection frame, and subtract the two ordinate values to obtain the distance between the text detection frame and the next text detection frame adjacent to the text detection frame. Further, the above preset distance threshold includes a first preset distance threshold and a second preset distance threshold, where the second preset distance threshold is greater than the first preset distance threshold. Further, if the distance between the text detection frame and the next text detection frame adjacent to the text detection frame is within the range between the first preset distance threshold and the second preset distance threshold, then there is a non-text detection frame in the area between the text detection frame and the next text detection frame, otherwise there is no non-text detection frame in the area between the text detection frame and the next text detection frame.

[0148] S760. If so, divide the area between the text detection frame and the next text detection frame into non-text detection frames.

[0149] Specifically, based on the coordinate values of the text detection frame and the coordinate values of the next text detection frame adjacent to the text detection frame, determine the coordinate values of the area between the text detection frame and the next text detection frame adjacent to the text detection frame. Determine the rectangular detection frame of the area between the text detection frame and the next text detection frame adjacent to the text detection frame based on the above coordinate values, and then use the rectangular detection frame as a non-text detection frame.

[0150] As an alternative solution, based on the obtained non-text detection boxes above, rectangular contour detection is performed on the non-text detection boxes. If a rectangular contour is detected, the rectangular contour is used as the non-text detection box. This can reduce the area of the non-text detection box, and thus the content in the non-text detection box can be recognized more accurately. Preferably, in this application, the OpenCV open-source tool is used to perform rectangular detection on the non-text detection box. Specifically, based on the coordinate values of the non-text detection box, the non-text detection box image at the corresponding position in the document image is intercepted. The OpenCV open-source tool is used to perform rectangular contour detection on the non-text detection box image. If a rectangular contour is detected in the non-text detection box image, the above non-text detection box is replaced with the rectangular area of the rectangular contour. If no rectangular contour is detected, the non-text detection box remains unchanged.

[0151] In an embodiment of this application, a method for obtaining a non-text detection box is provided. First, the position information of the text detection box and the position information of the next text detection box adjacent to the text detection box are obtained; then, based on the position information of the text detection box and the position information of the next text detection box, it is determined whether the distance between the text detection box and the next text detection box is greater than a preset distance threshold; if so, the area between the text detection box and the next text detection box is divided into a non-text detection box. By judging the non-text detection box, it can be ensured that the content of the non-text detection box will not be missed, effectively solving the problem of content loss in the non-text detection box and ensuring the integrity of the document content.

[0152] In one embodiment, as Figure 11 shown, according to the position information of the non-text detection box, the content in the non-text detection box is written into a first conversion document to generate a second conversion document, including steps 252 to 254. Among them,

[0153] S252. Obtain the writing format information of the non-text detection box from the document image;

[0154] Specifically, the position information of the non-text detection box includes the coordinate values of the four endpoints of the non-text detection box. Further, in addition to the title and paragraph text in the document, there are usually also contents such as figures, tables, figure captions, table captions, formulas, headers, and footers. The content of the non-text detection box includes the above-mentioned figures, tables, figure captions, table captions, formulas, headers, footers, etc. Similar to the text content, the figures, tables, figure captions, table captions, formulas, headers, footers, etc. also have corresponding format information. The above format information includes but is not limited to: the font size and font of the figure caption, the font size and font of the table caption, the style of the table, etc. This application does not make any limitations on this. Further, the methods for obtaining the writing format information include but are not limited to: deep learning algorithms, existing font recognition tools, etc. This application does not make any limitations on the way of obtaining the writing format information.

[0155] Further, based on the position information of the non-text detection box, obtain the coordinate values of the non-text detection box. Based on the coordinate values of the non-text detection box, intercept the non-text detection box image at the corresponding position in the document image, and use an identification tool to identify the format information of the content in the non-text detection box image, thereby obtaining the corresponding writing format information. For example, there are various font recognition websites in existing network resources. Only by using the non-text detection box image can the writing format information corresponding to the included content be detected, and then writing format information such as figure captions, table captions, headers, and footers can be obtained.

[0156] S254. Write the content of the non-text detection box into the first conversion document according to the position information of the non-text detection box and the writing format information of the non-text detection box to generate a second conversion document.

[0157] Specifically, based on the position information of the non-text detection box, obtain the coordinate values of the non-text detection box. Based on the coordinate values of the non-text detection box, intercept the non-text detection box image in the document image. Then obtain the content in the non-text detection box image, and then write the content in the non-text detection box image into the corresponding position of the first conversion document based on the writing format information of the non-text detection box to generate a second conversion document. Specifically, the content of the non-text detection box can be any one of figures, tables, figure captions, table captions, formulas, headers, footers, etc. Therefore, different tools can be used to write different types of content into the first conversion document, and this application does not make any limitations in this regard.

[0158] Specifically, if the content of the non-text detection box is a table, first use the OpenCV open-source tool to perform Hough line detection and vertex detection on the table area to obtain the coordinates of the intersections of the lines and the table. Then intercept the cells according to the coordinate information of the intersections of the lines and the table and use the open-source tool pytesser to recognize the text. Further, obtain the sizes and text information of each cell according to the coordinate information of the intersections of the lines and the table, and use the open-source tool python-docx to draw a table in the word document. If the content of the non-text detection box is a figure caption or a table caption, use the open-source tool pytesser to recognize the figure caption text or the table caption text, and then use the open-source tool python-docx to write the figure caption text or the table caption text into the word document.

[0159] In the embodiments of the present application, first, the writing format information of the non-text detection frame is obtained from the document image; then, according to the position information of the non-text detection frame and the writing format information of the non-text detection frame, the non-text detection frame is written into the first conversion document to generate a second conversion document. By writing the content in the non-text detection frame into the corresponding position of the first conversion document according to the writing format information of the corresponding non-text detection frame through the position information of the non-text detection frame, the content in the non-text detection frame can be restored more accurately, and the original format of the content in the non-text detection frame is retained, effectively solving the problems of content loss and format confusion in the non-text detection frame.

[0160] In one embodiment, the non-text detection frame includes at least one of a figure detection frame, a table detection frame, a figure caption detection frame, a table caption detection frame, a formula detection frame, a header detection frame, and a footer detection frame.

[0161] Specifically, the non-text detection frame includes at least one of a figure detection frame, a table detection frame, a figure caption detection frame, a table caption detection frame, a formula detection frame, a header detection frame, and a footer detection frame. Specifically, the content in the figure detection frame is an illustration in the document; the content in the table detection frame is a table; the content in the figure caption detection frame is a figure caption, usually located below the figure; the content in the table caption detection frame is a table caption, usually located above the table; the content in the formula detection frame is a formula, usually occupying a single line and having a specific height; the content in the header detection frame is the header in the upper edge of the document, which can be a page number, etc.; the content in the footer detection frame is the footer in the lower edge of the document, which can be a page number and a footnote, etc.

[0162] Specifically, the non-text detection frame in the present application includes a figure detection frame, a table detection frame, a figure caption detection frame, and a table caption detection frame. Further, Figure 12 is a schematic flowchart of obtaining the figure detection frame and the table detection frame in an embodiment, as Figure 12 shown, a method for obtaining the figure detection frame and the table detection frame is provided, including steps 820 to 880.

[0163] S820. Obtain the position information of the text detection frame and the position information of the next text detection frame adjacent to the text detection frame.

[0164] Specifically, the position information of the text detection frame includes the coordinate values of the four endpoints of the text detection frame. Further, to obtain the position information of the text detection frame, based on the position information of the text detection frame, obtain the coordinate values of the text detection frame. Obtain the position information of the next text detection frame adjacent to the text detection frame, and based on the position information of the next text detection frame adjacent to the text detection frame, obtain the coordinate values of the next text detection frame adjacent to the text detection frame.

[0165] S840. Determine whether the distance between the text detection box and the next text detection box is greater than a preset distance threshold according to the position information of the text detection box and the position information of the next text detection box.

[0166] Specifically, the distance between the text detection box and the next text detection box adjacent to it is obtained through the coordinate values of this text detection box and the coordinate values of the next text detection box adjacent to this text detection box. The maximum ordinate value of this text detection box is obtained through the coordinate values of this text detection box, and the minimum ordinate value of the next text detection box adjacent to this text detection box is obtained through the coordinate values of the next text detection box adjacent to this text detection box. The difference between the two ordinate values is used to obtain the distance between the text detection box and the next text detection box adjacent to it. The above preset distance threshold includes a first preset chart distance threshold and a second preset chart distance threshold, where the second preset chart distance threshold is greater than the first preset chart distance threshold. Further, if the distance between the text detection box and the next text detection box adjacent to it is within the range between the first preset chart distance threshold and the second preset chart distance threshold, then the area between the text detection box and the next text detection box is regarded as a non-text detection box, otherwise there is no non-text detection box between the text detection box and the next text detection box, where the above non-text detection box is a figure detection box or a table detection box.

[0167] Further, based on the obtained non-text detection box, rectangular detection is performed on the non-text detection box, which can reduce the area of the non-text detection box, and then the chart content in the non-text detection box can be recognized more accurately. In this application, the OpenCV open-source tool is used to perform rectangular detection on the non-text detection box. If a rectangular contour is detected, the obtained non-text detection box is replaced with the rectangular area of this rectangular contour. If no rectangular contour is detected, the non-text detection box remains unchanged.

[0168] S860. Perform text detection on the content in the non-text detection box according to the position information of the non-text detection box.

[0169] Specifically, the text content inside the table in the table detection box is regularly arranged, and the figure detection box may also include text content. Therefore, it is possible to determine whether the non-text detection box is a figure detection box or a table detection box by judging the arrangement of the text content in the table detection box and the figure detection box. Further, based on the position information of the non-text detection box, the coordinate values of the non-text detection box are obtained. Based on the coordinate values of the non-text detection box, the corresponding non-text detection box image in the document image is intercepted. By performing text detection on the non-text detection box image, the text detection box in the non-text detection box image is obtained. Further, the method for obtaining the text detection box in this application is not limited. The methods for obtaining the text detection box include but are not limited to: EAST model, FTSN model, RRPN model, DMPNet model, CTPN model, etc. Specifically, this application uses the CTPN model to perform text detection on the content in the non-text detection box image.

[0170] S880. Determine whether the text detection box in the non-text detection box is regularly arranged. If so, the non-text detection box is a table detection box; otherwise, the non-text detection box is a figure detection box.

[0171] Specifically, generally speaking, the text content in the table is regularly arranged in the form of a table. Therefore, the text detection box obtained after performing text detection on the table detection box is also regularly arranged. Further, judge the arrangement of the text detection box in the non-text detection box. If the text detection box in the non-text detection box shows a regular arrangement, the non-text detection box is a table detection box; otherwise, the non-text detection box is a figure detection box.

[0172] Further, after obtaining the figure detection box or the table detection box, determine whether there is a figure caption detection box in the figure detection box or a table caption detection box in the table detection box. Specifically, if the non-text detection box is a figure detection box, based on the position information of the figure detection box, the figure detection box image at the corresponding position in the document image is obtained. Then, text detection is performed on the figure detection box image to obtain the text detection box of the figure detection box image. Specifically, this application uses the CTPN model to perform text detection on the figure detection box image. Further, if a line text detection box appears below the figure detection box image, the text content in the line text detection box is used as the figure caption, and the line text detection box is used as the figure caption detection box. Among them, the line text detection box means that the text content in the text detection box is line text.

[0173] Further, if the non-text detection box is a table detection box, according to the position information of the table detection box, obtain the table detection box image at the corresponding position in the document image. Then, perform text detection on the table detection box image to obtain the text detection box of the table detection box. If a line text detection box appears above the table detection box image, then use the text content in the line text detection box as the table note, and the line text detection box as the table note detection box.

[0174] In the embodiments of the present application, the non-text detection box includes at least one of a figure detection box, a table detection box, a figure note detection box, a table note detection box, a formula detection box, a header detection box, and a footer detection box. By judging different types of non-text detection boxes, the content in the non-text detection box can be written into the editable document more accurately, effectively solving the problem of content loss in the non-text detection box.

[0175] In one embodiment, as Figure 13 shown, obtaining the text detection box from the document image corresponding to the original document includes steps 212 to 214. Among them,

[0176] S212. Convert the format of the original document to generate a document image corresponding to the original document.

[0177] Specifically, convert the format of the original document to generate a document image corresponding to the original document. The format of the document image includes but is not limited to: PNG, JPG, TIFF, etc. Optionally, the present application converts the original document into a document image in PNG format.

[0178] S214. Input the document image into a preset text detection model for text box detection to generate multiple text detection boxes; the preset text detection model includes a CTPN neural network model.

[0179] Specifically, CTPN (Connectionist Text Proposal Network) detects text in the image and gives the proposed regions where text may exist in the image, that is, the text detection boxes. By combining CNN and LSTM deep networks, CTPN can effectively detect horizontally distributed text in complex scenes. Figure 14 For a schematic diagram of the CTPN model in one embodiment, as Figure 14 shown, a CTPN model is provided. CTPN first extracts features through VGG16, then performs a convolution, passes through BLSTM, and is input into the FC fully connected layer, and finally obtains the output result of the network. The output result includes the position information of the text detection box.

[0180] Specifically, the CTPN model first obtains a document image, assuming it is a color image with a size of W*H*C, where W represents the width of the image, H represents the height of the image, and C represents the number of channels of the image. The image is input into the VGG16 deep network. Through a series of 3*3 convolution operations and pooling operations, a feature map of the text image is extracted. At this time, the size of the feature map is W*H*512. The extracted feature map is then subjected to a convolution operation through a convolutional layer with a convolution kernel size of 3*3 and a number of 512, so as to obtain a feature map with a size of W*H*512.

[0181] To avoid the deviation brought by the prediction of each small window, which affects the prediction results of the surrounding sliding windows, causing some small windows similar to text to be misdetected as text or some weak texts with unclear features to be discarded. At the same time, considering that text is a content with contextuality, the context information before and after the text line is crucial for the final prediction result. Therefore, the feature obtained after convolution is unfolded into a row vector and input into a bidirectional BLSTM network layer, enabling it to encode the context content in both forward and backward directions, thus achieving end-to-end training without other additional computational overhead. The BLSTM network layer at this time is an LSTM network containing two 128-dimensional hidden nodes, and the obtained feature map is still W*H*512. Then, the vector output by the BLSTM is used as the input of the fully connected layer FC and input into an RPN network similar to Faster R-CNN. This RPN network is divided into two parts. One part is used for detection box regression, and the other part is used for Softmax to classify the detection box. In the RPN network, the RPN network generates a multi-channel feature map with a scale of 1 / 16 of the original image for the input image. These feature maps can reflect the coordinate information of W*H*10 candidate text boxes with a width of 16 and a length of 11-283 on this image. The output is obtained through classification or regression. The output represents the 2k vertical coordinates of the height and the y-axis coordinate of the center of the selected box, the 2k scores of the category information of k anchor points indicating whether it is a character, and the k side-refinement of the horizontal offset of the selected box. Through the above output, the position information of the text box is obtained.

[0182] Further, the present application uses the CTPN neural network model as a preset text detection model. Figure 15 For a flow schematic diagram of the training process of the preset text detection model in an embodiment, as Figure 15 shown, a training method for the preset text detection model is provided, including steps 910 to step 950. The training process of the preset text detection model includes steps 910 to step 950.

[0183] S910. Obtain a document data set, where the document data set includes a number of preset document pictures;

[0184] S920. Set annotation labels for each preset document picture, where the annotation labels include preset text detection frames;

[0185] S930. Input the preset document picture into the initial text detection model to generate predicted text detection frames, where the initial text detection model is a CTPN neural network model;

[0186] S940. Calculate the loss function of the CTPN neural network model according to the predicted text detection frames and the preset text detection frames;

[0187] S950. Adjust the parameters of the initial text detection model according to the loss function to generate a preset text detection model.

[0188] Furthermore, input the document image into the preset text detection model obtained above for text box detection to obtain multiple text detection frames, and each text detection frame includes text content.

[0189] In the embodiment of the present application, a method for obtaining text detection frames is provided. First, the original document is subjected to format conversion to generate a document image corresponding to the original document, and then the document image is input into the preset text detection model for text box detection to generate multiple text detection frames. By using a neural network model to perform text detection on the document image, automatic detection of text content is realized, effectively improving the efficiency of document conversion.

[0190] In a specific embodiment, Figure 16 is a schematic flow chart of a document conversion method in a specific embodiment. As Figure 16 shown, a document conversion method is provided, including steps 1001 to 1012.

[0191] S1001. Perform format conversion on the original document to generate a document image corresponding to the original document;

[0192] S1002. Input the document image into the preset text detection model for text box detection to generate multiple text detection frames; the preset text detection model includes a CTPN neural network model;

[0193] S1003. Classify the text detection frames according to the attributes of the text detection frames to obtain text detection frames of different categories; the attributes of the text detection frames include the size and position information of the text detection frames;

[0194] Specifically, text detection boxes of different categories include title text detection boxes and paragraph text detection boxes, and the attributes of the text content in the title text detection boxes and the paragraph text detection boxes are different. Further, the text detection boxes are classified according to the size and position information of the text detection boxes to obtain title text detection boxes and paragraph text detection boxes.

[0195] S1004. Determine whether the height of the text detection box conforms to the height range of the preset-level title text detection box. If so, proceed to the next step; otherwise, the text detection box is not a title text detection box. Among them, the preset-level title text detection boxes include first-level title text detection boxes and second-level title detection boxes.

[0196] S1005. Determine whether the distance between the text detection box and the edge of the document image conforms to the first preset distance range. If so, determine that the text detection box is a title text detection box and proceed to step 1010; otherwise, the text detection box is not a title text detection box and proceed to step 1006.

[0197] S1006. Determine whether the height of the text detection box conforms to the height range of the preset paragraph text detection box. If so, proceed to the next step; otherwise, the text detection box does not belong to the paragraph text detection box.

[0198] S1007. According to the position information of the text detection boxes, determine whether the distance between the text detection boxes is less than the first preset threshold, and regard the text detection boxes with a distance less than the first preset threshold as intermediate text detection boxes.

[0199] S1008. From the intermediate text detection boxes, obtain the text detection boxes whose distance between the starting position information of the intermediate text detection box and the edge of the document image conforms to the second preset distance range as the starting text detection boxes.

[0200] S1009. Obtain paragraph text detection boxes based on the intermediate text detection boxes between two adjacent starting text detection boxes and proceed to step 1010.

[0201] S1010. According to the position information of the text detection box and the position information of the next text detection box, determine whether the distance between the text detection box and the next text detection box is greater than the preset distance threshold. If so, divide the area between the text detection box and the next text detection box into non-text detection boxes; otherwise, there are no non-text detection boxes between the text detection box and the next text detection box. The non-text detection boxes include figure detection boxes, table detection boxes, figure caption detection boxes, and table caption detection boxes.

[0202] S1011. Write the text content in the text detection box into an editable document according to the position information of the text detection boxes of different categories and the writing format information of the text detection boxes to generate a first conversion document.

[0203] S1012. Write the content of the non-text detection box into the first conversion document according to the position information of the non-text detection box and the writing format information of the non-text detection box to generate a second conversion document.

[0204] In an embodiment of the present application, first, the document image corresponding to the original document is input into the CTPN neural network model to obtain a plurality of text detection boxes, and then the text detection boxes are classified according to the attributes of the text detection boxes to obtain text detection boxes of different categories. Specifically, the text detection box categories include title text detection boxes and paragraph text detection boxes. Further, non-text detection boxes are obtained according to the position information of the text detection boxes. The non-text detection boxes include figure detection boxes, table detection boxes, figure caption detection boxes, and table caption detection boxes. Finally, according to the position information of the text detection boxes and the non-text detection boxes, the text content in the text detection boxes and the content in the non-text detection boxes are written into an editable document according to the corresponding writing format information, so that the original format of the text content can be accurately restored and the integrity of the text content information in the document can be ensured.

[0205] It should be understood that although Figure 2-16 the steps in the flowchart of Figure 2-16 are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover,

[0206] In one embodiment, Figure 17 is a structural block diagram of a document conversion device in an embodiment. As Figure 17 shown, a document conversion device 1700 is provided, including:

[0207] A text detection box acquisition module 1720, configured to acquire text detection boxes from a document image corresponding to an original document, where the text detection boxes include text content;

[0208] A text detection box classification module 1740, configured to classify the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection boxes;

[0209] The text detection box writing module 1760 is used to write the text content in the text detection box into an editable document according to the position information of text detection boxes of different categories and the writing format information of the text detection boxes, so as to generate a first conversion document.

[0210] In one embodiment, as Figure 18 shown, the text detection box classification module 1740 includes:

[0211] The title text detection box classification unit 1742 is used to classify the text detection box according to the size and position information of the text detection box, so as to obtain the title text detection box;

[0212] The paragraph text detection box classification unit 1744 is used to classify the text detection box according to the size and position information of the text detection box, so as to obtain the paragraph text detection box.

[0213] In one embodiment, the title text detection box classification unit 1742 is further used to determine whether the height of the text detection box meets the height range of the preset level title text detection box; the preset level title text detection box includes at least one level of title text detection box; if so, it is determined whether the distance between the text detection box and the edge of the document image meets the first preset distance range; if so, the text detection box is used as the preset level title text detection box.

[0214] In one embodiment, the paragraph text detection box classification unit 1744 is further used to determine whether the height of the text detection box meets the height range of the preset paragraph text detection box; if so, according to the position information of the text detection box, it is determined whether the distance between the text detection boxes is less than the first preset threshold, and the text detection boxes with a distance less than the first preset threshold are used as intermediate text detection boxes; from the intermediate text detection boxes, the text detection boxes whose distance between the starting position information of the intermediate text detection box and the edge of the document image meets the second preset distance range are obtained as the starting text detection boxes; according to the intermediate text detection boxes between two adjacent starting text detection boxes, the paragraph text detection box is obtained.

[0215] In one embodiment, the document conversion device 1700 further includes:

[0216] The non-text detection box classification module is used to obtain the position information of the text detection box and the position information of the next text detection box adjacent to the text detection box; according to the position information of the text detection box and the position information of the next text detection box, it is determined whether the distance between the text detection box and the next text detection box is greater than the preset distance threshold; if so, the area between the text detection box and the next text detection box is divided into non-text detection boxes.

[0217] In one embodiment, the document conversion device 1700 further includes:

[0218] A non-text detection box writing module, configured to obtain the writing format information of the non-text detection box from the document image; and write the content of the non-text detection box into a first converted document according to the position information of the non-text detection box and the writing format information of the non-text detection box, to generate a second converted document.

[0219] The division of each module in the above document conversion device is only for illustrative purposes. In other embodiments, the document conversion device may be divided into different modules as needed to complete all or part of the functions of the above document conversion device.

[0220] In one embodiment, Figure 19 is a schematic internal structure diagram of a computer device in one embodiment. As Figure 19 shown, the computer device may be a server, and the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The database of the computer device is used to store audio data, the network interface of the computer device is used to communicate with an external terminal through a network connection, and the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The computer program can be executed by the processor to implement the document conversion method provided in each of the above embodiments.

[0221] In the embodiments of the present application, the implementation of each module in the provided document conversion device may be in the form of a computer program. The computer program can run on a computer device or a server. The program module constituted by the computer program can be stored on the memory of the computer device or the server. When the computer program is executed by the processor, the steps of the method described in the embodiments of the present application are implemented.

[0222] The embodiments of the present application also provide a computer-readable storage medium. One or more non-volatile computer-readable storage media containing computer-executable instructions, when the computer-executable instructions are executed by one or more processors, cause the processors to execute the steps of the document conversion method.

[0223] A computer program product containing instructions, when it runs on a computer, causes the computer to execute the document conversion method.

[0224] Any reference to memory, storage, database, or other media used in the embodiments of this application may include non-volatile and / or volatile memory. Suitable non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).

[0225] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0226] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. A document conversion method, characterized in that, The method includes: Obtaining a text detection box from a document image corresponding to the original document, where the text detection box includes text content; Classifying the text detection box according to the attributes of the text detection box to obtain text detection boxes of different categories; the attributes of the text detection box include the size and position information of the text detection box; Writing the text content in the text detection box into an editable document according to the position information of the text detection boxes of different categories and the writing format information of the text detection box to generate a first converted document; Obtaining the position information of the text detection box and the position information of the next text detection box adjacent to the text detection box; Judging whether the distance between the text detection box and the next text detection box is greater than a preset distance threshold according to the position information of the text detection box and the position information of the next text detection box, where the position information of the text detection box includes the coordinate values of the four endpoints of the text detection box; If so, determining the coordinate values of the area between two adjacent text detection boxes based on the coordinate values corresponding to the text detection box and the coordinate values corresponding to the next text detection box, and dividing the area at the corresponding position into a non-text detection box based on the coordinate values of the area between the two adjacent text detection boxes, where the position information of the non-text detection box includes the coordinate values of the four endpoints of the non-text detection box; Writing the content in the non-text detection box into the first converted document according to the position information of the non-text detection box to generate a second converted document.

2. The method according to claim 1, wherein The text detection boxes of different categories include a title text detection box and a paragraph text detection box, and the attributes of the text content in the title text detection box and the paragraph text detection box are different; The classifying the text detection box according to the attributes of the text detection box to obtain text detection boxes of different categories includes: Classifying the text detection box according to the size and position information of the text detection box to obtain the title text detection box; Classifying the text detection box according to the size and position information of the text detection box to obtain the paragraph text detection box.

3. The method according to claim 2, wherein The size of the text detection box includes the height of the text detection box; the classifying the text detection box according to the size and position information of the text detection box to obtain the title text detection box includes: Judging whether the height of the text detection box conforms to the height range of a preset level title text detection box; the preset level title text detection box includes at least one level of title text detection box; If so, judging whether the distance between the text detection box and the edge of the document image conforms to a first preset distance range; If so, regarding the text detection box as the preset level title text detection box.

4. The method according to claim 2 or 3, characterized in that, The size of the text detection box includes the height of the text detection box; the position information of the text detection box includes the starting position information of the text detection box; The classifying the text detection box according to the size and position information of the text detection box to obtain the paragraph text detection box includes: Determine whether the height of the text detection box meets the height range of the preset paragraph text detection box; If so, according to the position information of the text detection box, determine whether the distance between the text detection boxes is less than a first preset threshold, and use the text detection boxes with a distance less than the first preset threshold as intermediate text detection boxes; From the intermediate text detection boxes, obtain the text detection boxes whose distance between the starting position information of the intermediate text detection box and the edge of the document image meets a second preset distance range as starting text detection boxes; Obtain the paragraph text detection box according to the intermediate text detection boxes between two adjacent starting text detection boxes.

5. The method according to claim 1, wherein The writing the content in the non-text detection box into the first conversion document according to the position information of the non-text detection box to generate a second conversion document includes: Obtain the writing format information of the non-text detection box from the document image; According to the position information of the non-text detection box and the writing format information of the non-text detection box, write the content of the non-text detection box into the first conversion document to generate the second conversion document.

6. A document conversion device, characterized in that, The device includes: A text detection box acquisition module, configured to acquire a text detection box from a document image corresponding to an original document, where the text detection box includes text content; A text detection box classification module, configured to classify the text detection boxes according to the attributes of the text detection boxes to obtain text detection boxes of different categories; the attributes of the text detection boxes include the size and position information of the text detection box; A text detection box writing module, configured to write the text content in the text detection box into an editable document according to the position information of different categories of text detection boxes and the writing format information of the text detection box to generate a first conversion document; A non-text detection box classification module, configured to obtain the position information of the text detection box and the position information of the next text detection box adjacent to the text detection box; according to the position information of the text detection box and the position information of the next text detection box, determine whether the distance between the text detection box and the next text detection box is greater than a preset distance threshold, where the position information of the text detection box includes the coordinate values of the four endpoints of the text detection box; if so, based on the coordinate values corresponding to the text detection box and the coordinate values corresponding to the next text detection box, determine the coordinate values of the area between two adjacent text detection boxes, and based on the coordinate values of the area between the two adjacent text detection boxes, divide the area at the corresponding position into non-text detection boxes, where the position information of the non-text detection box includes the coordinate values of the four endpoints of the non-text detection box; A non-text detection box writing module, configured to write the content in the non-text detection box into the first conversion document according to the position information of the non-text detection box to generate a second conversion document.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Image processing method and device, terminal and computer readable storage medium

    CN110598566A

  • Text processing method, device and equipment based on artificial intelligence, and medium

    CN111242083A

  • Document picture recognition method and device and computer equipment

    CN113221632A