Document recognition method, device, electronic device and computer-readable storage medium

By extracting and detecting the layout structure features and content features, the problem of inaccurate document recognition under complex layout structures in the prior art is solved, and the effect of accurately identifying the layout content area in the complex layout structure is achieved.

CN115131804BActive Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210425659.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-08-22
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

Existing document recognition methods cannot accurately identify the publication structure information under complex layout structure, resulting in low document recognition accuracy.

Method used

By extracting the layout structure features and layout content features in the document image to be identified, the layout content area and content type are detected, and the layout content is determined based on the text content and content type are determined to generate an editable target document.

Benefits of technology

Accurately identifying the content area of ​​the publication in a complex layout structure improves the accuracy of document recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131804B_ABST
    Figure CN115131804B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a document recognition method, device, electronic device and computer-readable storage medium; after displaying a document recognition page, the embodiment of the present invention responds to a trigger operation on a recognition control in the document recognition page, extracts publication structure features and layout content features from a document image to be recognized in the document recognition page, and then, based on the layout structure features and layout content features, detects at least one layout content area and the content type of the layout content area in the document image to be recognized, recognizes the text content corresponding to the publication content area in the document image to be recognized, and determines the layout content of the layout content area based on the text content and the content type, and then, based on the layout content, generates a target document corresponding to the document image to be recognized, and displays the target document, which is an editable document; this solution can improve the accuracy of document recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a document recognition method, device, electronic device and computer-readable storage medium. Background Art

[0002] In recent years, with the rapid development of Internet technology, the content of images has become increasingly rich, and a growing number of image recognition methods have emerged. In addition to identifying image types, it is also possible to identify documents within images. Existing document recognition methods often use connected component segmentation and semantic segmentation to segment the document image to identify the target document.

[0003] In the research and practice of the existing technology, the inventors of the present invention found that when using connected domain segmentation and semantic segmentation in the document recognition process, the layout of the document image to be recognized is often restored according to the reading order. Under complex layout structures, the publication structure information cannot be accurately identified, resulting in low accuracy of document recognition. Summary of the Invention

[0004] Embodiments of the present invention provide a document recognition method, device, electronic device, and computer-readable storage medium, which can improve the accuracy of document recognition.

[0005] A document recognition method, comprising:

[0006] Displaying a document recognition page, wherein the document recognition page includes a document image to be recognized and a recognition control;

[0007] In response to a triggering operation on the recognition control, extracting the publication page structure features and the layout content features from the document image to be recognized;

[0008] Detecting at least one layout content area and a content type of the layout content area in the document image to be identified based on the layout structure features and the layout content features;

[0009] Identifying text content corresponding to the layout content area in the document image to be identified, and determining layout content of the layout content area based on the text content and content type;

[0010] According to the layout content, a target document corresponding to the document image to be identified is generated and displayed, and the target document is an editable document.

[0011] Accordingly, an embodiment of the present invention provides a document recognition device, comprising:

[0012] A display unit, configured to display a document recognition page, wherein the document recognition page includes an image of a document to be recognized and a recognition control;

[0013] an extraction unit, configured to extract, in response to a triggering operation on the recognition control, publication page structure features and layout content features from the document image to be recognized;

[0014] a detection unit, configured to detect at least one layout content area and a content type of the layout content area in the document image to be identified based on the layout structure features and the layout content features;

[0015] a determination unit, configured to identify text content corresponding to the layout content area in the document image to be identified, and determine layout content of the layout content area based on the text content and content type;

[0016] A generating unit is used to generate a target document corresponding to the document image to be identified according to the layout content, and display the target document, wherein the target document is an editable document.

[0017] Optionally, in some embodiments, the extraction unit can be specifically used to perform layout correction on the document image to be identified to obtain a corrected document image; adjust the image size of the corrected document image to obtain an adjusted document image; and extract publication structure features and layout content features from the adjusted document image.

[0018] Optionally, in some embodiments, the extraction unit can be specifically used to use a trained layout detection model to perform image feature extraction on the adjusted document image to obtain basic image features; perform multi-dimensional layout feature extraction on the basic image features to obtain basic layout features of each dimension; and perform multi-dimensional layout feature extraction on the basic image features based on the basic layout features to obtain the layout structure features and layout content features of the document image to be identified.

[0019] Optionally, in some embodiments, the extraction unit can be specifically used to fuse the basic layout features with the basic image features to obtain fused image features; perform layout feature extraction on the fused image features to obtain initial layout features corresponding to the target dimension; and identify the layout structure features and layout content features of the document image to be identified in the initial layout features.

[0020] Optionally, in some embodiments, the extraction unit can be specifically used to sort the dimensional information of the basic layout features, and based on the sorting information, screen out target basic layout features that exceed the target dimension from the basic layout features; fuse the target basic layout features and the initial layout features to obtain fused layout features; and extract the layout structure features and layout content features of the document image to be identified from the fused layout features.

[0021] Optionally, in some embodiments, the extraction unit can be specifically used to use the fused layout features as the fused image features, and return to execute the step of performing layout feature extraction on the fused image features to obtain the initial layout features corresponding to the target dimension, until the target basic layout features do not exist, and obtain the layout features corresponding to each dimension; obtain the weighting coefficient corresponding to each dimension, and weight the layout features based on the weighting coefficient to obtain the weighted layout features; identify the layout structure features and layout content features of the document image to be identified in the weighted layout features.

[0022] Optionally, in some embodiments, the document recognition device may further include a training unit, which may be specifically used to obtain document image samples, and use a preset layout detection model to extract layout features of the document image samples to obtain basic sample layout features; identify target sample layout features from the basic sample layout features, and determine the main loss information of the document image sample based on the target sample layout features; determine the auxiliary loss information of the document image sample based on the basic sample layout features, and converge the preset layout detection model based on the main loss information and the auxiliary loss information to obtain the trained layout detection model.

[0023] Optionally, in some embodiments, the detection unit can be specifically used to detect at least one layout structure area and the area type corresponding to the layout structure area in the document image to be identified based on the layout structure features; determine the layout structure type of the layout structure area based on the area type; and identify at least one layout content area and the content type of the layout content area in the layout structure area based on the layout content features and the layout structure type.

[0024] Optionally, in some embodiments, the detection unit can be specifically used to identify at least one layout content area and the content type of the layout content area in the layout structure area based on the layout content features when the layout structure type is a column structure area, and the column structure area is an area used for content columnarization in the document contained in the document image to be identified; when the layout structure type is a non-column structure area, the layout structure area is used as the layout content area, and the area type is used as the content type of the layout content area.

[0025] Optionally, in some embodiments, the detection unit can be specifically used to identify at least one layout content area and the initial content type of the layout content area in the layout structure area based on the layout content features; when the initial content type is a formula, obtain the formula position information and formula format information of the layout content area corresponding to the formula, and determine the formula type of the formula based on the formula position information and formula format information to obtain the content type of the layout content area; when the initial content type is non-formula, use the initial content type as the content type of the layout content area.

[0026] Optionally, in some embodiments, the determination unit can be specifically used to identify the image corresponding to the layout content area in the document image to be identified to obtain the layout content when the content type is an image; when the content type is non-image, determine the text type of the layout content according to the content type, and convert the text content into layout content corresponding to the text type.

[0027] Optionally, in some embodiments, the determination unit can be specifically used to obtain the text format of the layout content area when the text type is basic text, and adjust the format of the text content based on the text format to obtain the layout content; when the text type is table text, convert the text content into table content, and use the table content as the layout content; when the text type is formula text, convert the text content into formula content according to the formula type corresponding to the formula text, and use the formula content as the layout content.

[0028] Optionally, in some embodiments, the determination unit can be specifically used to obtain the layout information of the layout content area and extract the basic formula format from the layout information; when the formula type is an inline formula, the basic formula format is used as the formula format of the layout content, and the inline formula is the formula in the text paragraph; when the formula type is an interline formula, the basic formula format is adjusted according to the type of the interline formula to obtain the formula format of the layout content, and the interline formula is the formula between the text paragraphs; the text content is converted into the formula content corresponding to the formula format.

[0029] Optionally, in some embodiments, the determination unit can be specifically used to determine the text alignment of the inter-line formula according to the type of the inter-line formula; and add the text alignment to the basic formula format to obtain the formula format of the layout content.

[0030] Optionally, in some embodiments, the generation unit can be specifically used to obtain the area location information of each of the layout content areas, and sort the layout content according to the area location information; create an initial document in a preset format, and based on the sorting information, write the layout content into the initial document to obtain the target document corresponding to the document image to be identified.

[0031] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to implement the document recognition method provided by the embodiment of the present invention.

[0032] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps in any one of the document recognition methods provided by the embodiments of the present invention.

[0033] After displaying the document recognition page, the embodiment of the present invention responds to the trigger operation of the recognition control in the document recognition page, extracts the publication structure features and layout content features from the document image to be recognized in the document recognition page, and then, based on the layout structure features and layout content features, detects at least one layout content area and the content type of the layout content area in the document image to be recognized, recognizes the text content corresponding to the publication content area in the document image to be recognized, and determines the layout content of the layout content area based on the text content and the content type, and then, based on the layout content, generates the target document corresponding to the document image to be recognized, and displays the target document, which is an editable document; since the scheme can directly extract the publication structure features and layout content features in the document image to be recognized, and then, based on the layout structure features and layout content features, can directly detect the publication content area, so that the publication content area can be accurately identified in a complex layout structure, thereby improving the accuracy of document recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0035] Figure 1 Schematic diagram of a scenario of a document recognition method provided by an embodiment of the present invention;

[0036] Figure 2 This is the process of the document recognition method provided by the embodiment of the present invention;

[0037] Figure 3 1 is a schematic diagram of a document image photographing page provided by an embodiment of the present invention;

[0038] Figure 4 1 is a schematic diagram of a process for extracting layout features of each dimension provided by an embodiment of the present invention;

[0039] Figure 5 is a schematic diagram of a double-column area identified in a document image to be identified provided by an embodiment of the present invention;

[0040] Figure 6 is a schematic diagram of different types of formulas in a document image to be recognized provided by an embodiment of the present invention;

[0041] Figure 7 This is a schematic diagram of a process for restoring the layout of a document image to be recognized provided by an embodiment of the present invention;

[0042] Figure 8Schematic diagram of the main algorithm framework of CBNetV2 provided by an embodiment of the present invention;

[0043] Figure 9 Schematic diagram of the algorithm structure of scaled-yolov4 provided in an embodiment of the present invention;

[0044] Figure 10 It is a schematic structural diagram of an improved layout detection model provided by an embodiment of the present invention;

[0045] Figure 11 is another flowchart of the document recognition method provided by an embodiment of the present invention;

[0046] Figure 12 Schematic diagram of the structure of a document recognition device provided by an embodiment of the present invention;

[0047] Figure 13 is another structural diagram of a document recognition device provided by an embodiment of the present invention;

[0048] Figure 14 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] The embodiments of the present invention provide a document recognition method, device, electronic device, and computer-readable storage medium. The document recognition device can be integrated into an electronic device, which can be a server, a terminal, or other device.

[0051] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and this application does not limit this.

[0052] For example, see Figure 1 Taking the document recognition device integrated in an electronic device as an example, after the electronic device displays the document recognition page, it responds to the trigger operation of the recognition control in the document recognition page, extracts the publication structure features and layout content features from the document image to be recognized in the document recognition page, and then, based on the layout structure features and layout content features, detects at least one layout content area and the content type of the layout content area in the document image to be recognized, recognizes the text content corresponding to the publication content area in the document image to be recognized, and determines the layout content of the layout content area based on the text content and the content type, and then, based on the layout content, generates a target document corresponding to the document image to be recognized, and displays the target document, thereby improving the accuracy of document recognition.

[0053] Among them, the response is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay; unless otherwise specified, there is no restriction on the order of execution of the multiple operations executed.

[0054] It can be understood that in the specific implementation of this application, when the following embodiments of this application are applied to specific products or technologies, relevant data such as the document image to be identified of the object require permission or consent, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0055] Among them, the keyword recognition method provided in the embodiment of the present application involves computer vision technology in the field of artificial intelligence, that is, in this application, artificial intelligence machine vision technology can be used to extract layout features of the document image to be identified, and based on the extracted layout structure features and layout content features, at least one layout content area and the area type of the layout content area can be detected in the document image to be identified, and based on the detected layout content area and area type, the document image to be identified can be restored to the target document.

[0056] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. AI software technologies primarily include computer vision and machine learning / deep learning.

[0057] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, which uses cameras and computers to replace the human eye in identifying and measuring targets, and then further processes images to make them more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0058] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0059] This embodiment will be described from the perspective of a document recognition device, which can be specifically integrated into an electronic device, which can be a server or a terminal; wherein the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a wearable device, a virtual reality device or other smart devices that can perform document recognition.

[0060] A document recognition method, comprising:

[0061] A document recognition page is displayed, which includes a document image to be recognized and a recognition control. In response to a trigger operation on the recognition control, publication structure features and layout content features are extracted from the document image to be recognized. Based on the layout structure features and layout content features, at least one layout content area and the content type of the layout content area are detected in the document image to be recognized. The text content corresponding to the publication content area is recognized in the document image to be recognized, and based on the text content and the content type, the layout content of the layout content area is determined. Based on the layout content, a target document corresponding to the document image to be recognized is generated, and the target document is displayed. The target document is an editable document.

[0062] like Figure 2 As shown, the specific process of the document recognition method is as follows:

[0063] 101. Display the document recognition page.

[0064] Among them, the document recognition page is a page that recognizes the target document in the document image to be recognized. The document recognition page includes the document image to be recognized and the recognition control. The document image to be recognized can be an image containing a document. The so-called document can be understood as an electronic archive file containing at least one type of content. The document can be of various types, for example, it can be a text document, a spreadsheet document, a slide document, etc.

[0065] There are multiple ways to display the document recognition page, which are as follows:

[0066] For example, a document image acquisition page is displayed, which includes an acquisition control. In response to a trigger operation on the acquisition control, the current document image is acquired, and an acquisition image preview page is displayed. The acquisition image preview page includes the current document image and a confirmation control. In response to a trigger operation on the confirmation control, the current document image is used as the document image to be identified, and a document identification page is displayed. Alternatively, a user operation page is displayed, which includes an upload control. In response to a trigger operation on the upload control, a candidate document image list is displayed. In response to a selection operation on the candidate document image list, the candidate document image corresponding to the selection operation is filtered out from the candidate document image list to obtain the document image to be identified, and a document identification page is displayed based on the document image to be identified.

[0067] There are multiple collection controls, and different collection controls can exist for collecting different contents. For example, the collection controls can include a collection control for extracting text, a collection control for taking photos and scanning, a collection control corresponding to scanning a code, and a collection control corresponding to scanning, etc. Taking the collection control as a photo control as an example, the way to display the document recognition page can be that after displaying the document image photo page, the user triggers the photo control and takes a photo of the document or the content containing the document, thereby obtaining the document image to be recognized, and then displays the document recognition page. The document image photo page can be as follows: Figure 3 shown.

[0068] Among them, the way of displaying the document recognition page can be found that there are two sources of the document image to be recognized. One is to select the document image to be recognized from the existing document images, and the other is to directly capture the image containing the document through the image acquisition control to obtain the document image to be recognized.

[0069] 102. In response to a triggering operation on a recognition control, extracting publishing page structure features and layout content features from the document image to be recognized.

[0070] Among them, the layout structure feature can be understood as the characteristic information that characterizes the layout structure of the document in the document image to be identified. The so-called layout structure can be understood as the structure of the layout in the document, for example, it can include the document's layout columns (single column, double column or multiple columns), headers, footers and titles, etc.

[0071] Among them, the layout content feature can be feature information representing the layout content of the document in the document image to be identified. The so-called layout content can be understood as the content of each content area in the document layout, for example, it can include paragraphs, tables, images, formulas, etc.

[0072] There are many ways to extract the publication structure features and layout content features from the document image to be identified in response to the triggering operation on the control to be identified, which can be specifically as follows:

[0073] For example, in response to a trigger operation for a control to be identified, the layout correction is performed on the document image to be identified to obtain a corrected document image, the image size of the corrected document image is adjusted to obtain an adjusted document image, and the publication structure features and layout content features are extracted from the adjusted document image.

[0074] Among them, the method of performing layout correction on the document image to be recognized, for example, correcting the distortion or tilt of the document image to be recognized, thereby obtaining a corrected document image.

[0075] After performing layout correction on the document image with recognition, the image size of the corrected document image can be adjusted in various ways, such as obtaining current image size information of the corrected document image, selecting the longest side and the shortest side from the current image size information, scaling the longest side to a first preset size, and scaling the shortest side to a second preset size, thereby obtaining an adjusted document image. The first preset size and the second preset size can be set according to actual application. For example, the first preset size can be 1280, and the second preset size can be a multiple of 128, etc.

[0076] After adjusting the image size of the corrected document image, there are many ways to extract the publication structure features and layout content features from the adjusted document image. For example, the trained layout detection model can be used to extract image features of the adjusted document image to obtain basic image features, and multi-dimensional layout features can be extracted from the basic image features to obtain basic layout features of each dimension. Based on the basic layout features, multi-dimensional layout features can be extracted from the basic image features to obtain the layout structure features and layout content features of the document image to be identified.

[0077] Among them, there are many ways to perform multi-dimensional layout feature extraction on basic image features based on layout features. For example, the basic layout features can be fused with the basic image features to obtain fused image features, and the fused image features can be subjected to layout feature extraction to obtain the initial layout features corresponding to the target dimension. The layout structure features and layout content features of the document image to be identified can be identified in the initial layout features.

[0078] Among them, there can be many ways to identify the layout structure features and layout content features of the document image to be identified in the initial layout features. For example, the dimensional information of the basic layout features is sorted, and based on the sorting information, the target basic layout features that exceed the target dimension are screened out from the basic layout features, the target basic layout features and the initial layout features are fused to obtain the fused layout features, and the layout structure features and layout content features of the document image to be identified are extracted from the fused layout features.

[0079] Among them, there can be many ways to extract the layout structure features and layout content features of the document image to be identified from the fused layout features. For example, the fused layout features are used as the fused image features, and the step of extracting the layout features from the fused image features is returned to obtain the initial layout features corresponding to the target dimension, until the target basic layout features do not exist, the layout features corresponding to each dimension are obtained, the weighting coefficients corresponding to each dimension are obtained, and the layout features are weighted based on the weighting coefficients to obtain weighted layout features, and the layout structure features and layout content features of the document image to be identified in the weighted layout features.

[0080] Among them, when extracting multi-dimensional layout features from basic image features, taking the feature extraction dimension as 4 dimensions as an example, the process of extracting the layout features corresponding to each dimension from the basic image features can be as follows: Figure 4 As shown, the adjusted document image is convolved to obtain basic image features, the first backbone network (Backbone 1) performs multi-dimensional layout feature extraction on the basic image features to obtain basic layout features corresponding to each dimension, the basic layout features from the first to fourth dimensions are fused with the basic image features to obtain fused image features, the second backbone network (Backbone 2) performs layout feature extraction of the first dimension on the fused image features to obtain initial layout features corresponding to the first dimension, then, the basic layout features from the second to fourth dimensions are fused with the initial layout features corresponding to the first dimension to obtain fused layout features corresponding to the first dimension, the second backbone network performs layout feature extraction of the second dimension on the fused layout features corresponding to the first dimension to obtain initial layout features corresponding to the second dimension, then, the basic layout features from the third to fourth dimensions are fused with the initial layout features corresponding to the second dimension to obtain fused layout features corresponding to the second dimension, and so on, thereby obtaining layout features corresponding to each dimension.

[0081] Optionally, the trained layout detection model can be set according to actual applications. In addition, it should be noted that the trained layout detection model can be pre-set by maintenance personnel or trained by the document recognition device itself. That is, before the step of "using the trained layout detection model to extract image features from the adjusted document image to obtain basic image features", the document recognition method may further include:

[0082] Obtain document image samples, and use a preset layout detection model to extract layout features of the document image samples to obtain basic sample layout features, identify target sample layout features in the basic sample layout features, and determine the main loss information of the document image samples based on the target sample layout features. According to the basic sample layout features, determine the auxiliary loss information of the document image samples, and based on the main loss information and the auxiliary loss information, converge the preset layout detection model to obtain a trained layout detection model.

[0083] Among them, the document image sample includes a document image with annotated layout content area and content type of the layout content area. There are many ways to use the preset layout detection model to extract layout features of the document image sample. For example, the preset layout detection model can be used to extract image features of the document image sample to obtain sample image features, and the first backbone network (Backbone 1) of the preset layout detection model can be used to perform multi-dimensional layout feature extraction on the sample image features to obtain basic sample layout features.

[0084] After extracting the basic sample layout features, the target sample layout features can be identified from the basic layout features. There are many ways to identify the target sample layout features. For example, the second backbone network (Backbone 1) can be used to perform multi-dimensional layout feature extraction on the sample image features and the basic sample layout features to obtain the target sample layout features.

[0085] Among them, the first backbone network and the second backbone network can be compositely connected. In addition, the network structures of the first backbone network and the second backbone network can be the same. In the preset layout detection model, the number of backbone networks can be 2 or more. In the case of multiple networks, composite connections are used between each backbone network.

[0086] After identifying the layout features of the target sample, the trunk loss information of the document image sample can be determined. There are many ways to determine the trunk loss information. For example, based on the layout features of the target sample, at least one layout content area and the content type of the layout content area are predicted in the document image sample to obtain a first predicted layout content area and a first predicted content type. The first predicted layout content area and the first predicted content type are compared with the annotated layout content area and the annotated content type respectively to obtain the trunk loss information of the document image sample.

[0087] Among them, there are many ways to determine the auxiliary information of the document image sample based on the basic sample layout characteristics. For example, based on the basic sample layout characteristics, at least one layout content area and the content type of the layout content area are predicted in the document image sample to obtain a second predicted layout content area and a second predicted content type. The second predicted layout content area and the second predicted content type are respectively compared with the annotated layout content area and the annotated content type to obtain the auxiliary loss information of the document image sample.

[0088] After determining the main loss information and the auxiliary loss information, the preset layout detection model can be converged based on the main loss information and the auxiliary loss information. There are many ways to converge, for example, the main loss information and the auxiliary loss information are fused to obtain the fused loss information, and the fused loss information is used to converge the preset layout detection model to obtain the trained layout detection model. Alternatively, the weighting coefficients corresponding to the main loss information and the auxiliary loss information can be obtained, and the main loss information and the auxiliary loss information can be weighted based on the weighting coefficients to obtain the weighted main loss information and the weighted auxiliary loss information, and the weighted main loss information and the weighted auxiliary loss information can be fused to obtain the fused loss information, and the preset layout detection model can be converged based on the fused loss information to obtain the trained layout detection model.

[0089] 103. Detect at least one layout content area and a content type of the layout content area in the document image to be identified based on the layout structure features and the layout content features.

[0090] The layout content area may be the content area of ​​the document layout in the document image to be recognized. The layout content area may have various types, such as a column structure area (single column or multi-column), a header and footer, a title, a paragraph, etc. The content type of the layout content area may be understood as the type of content contained in the layout content area. The content type may also have various types, such as an image, text, a table, a formula, etc.

[0091] There are various ways to detect at least one layout content area and the content type of the layout content area in the document image to be recognized based on the layout structure features and layout content features, which can be specifically as follows:

[0092] For example, based on the layout structure features, at least one layout structure area and the area type corresponding to the layout structure area can be detected in the document image to be identified; based on the area type, the layout structure type of the layout structure area can be determined; based on the layout content features and the layout structure type, at least one layout content area and the content type of the layout content area can be identified in the layout structure area.

[0093] The layout structure area may be a structure area of ​​the layout of a document contained in the document image to be identified. The layout structure area may have various types of layout structure areas, such as column structure areas, headers and footers, titles, and the like. Based on the layout content features and the layout structure type, there may be various ways to identify at least one layout content area and the content type of the layout content area in the layout structure area. For example, when the layout structure type is a column structure area, at least one layout content area and the content type of the layout content area may be identified in the layout structure area based on the layout content features. When the layout structure type is a non-column structure area, the layout structure area may be used as the layout content area, and the area type may be used as the content type of the layout content area.

[0094] The column structure area is an area used for content column division in the document contained in the document image to be identified. There are many types of column structure areas, for example, they can include single column structure areas, double column structure areas, and multi-column structure areas. Taking the classification structure area as a double column structure area as an example, the two column areas in the double column area can be marked in the document image to be identified. Specifically, Figure 5 As shown. When the layout structure type is a column structure area, there can be multiple ways to identify at least one layout content area and the content type of the layout content area in the layout structure area based on the layout content features. For example, based on the layout content features, at least one layout content area and the initial content type of the layout content area can be identified in the layout structure area. When the initial content type is a formula, the formula position information and formula format information of the layout content area corresponding to the formula are obtained, and the formula type of the formula is determined based on the formula position information and formula format information to obtain the content type of the layout content area. When the initial content type is non-formula, the initial content type is used as the content type of the layout content area.

[0095] The formula position information is used to indicate the position information of the layout content area corresponding to the formula in the document image to be recognized, and the formula format information is used to indicate the format information contained in the layout content area corresponding to the formula. There are multiple ways to determine the formula type of the formula based on the formula position information and the formula format information. For example, the layout content area whose initial content type is a paragraph is screened out from the layout content area to obtain the paragraph content area, the paragraph position information of the paragraph content area is obtained, the paragraph position of the paragraph content area is extracted from the paragraph position information, and the formula position is extracted from the formula position information. The formula position is compared with the paragraph position. When the formula position is within the paragraph position, the formula type of the formula is determined to be an inline formula. When the formula position is outside the paragraph position, the formula type of the formula is determined to be an interline formula. The target formula format information corresponding to the interline formula is screened out from the formula format information. When the target formula format information contains a formula sequence number, the interline formula is determined to be a sequenced interline formula. When the target formula format information does not contain a formula sequence number, the interline formula is determined to be an unsequenced interline formula.

[0096] Among them, there can be multiple types of formulas in this solution, for example, it can include in-line formulas, unnumbered inter-line formulas and numbered inter-line formulas. Compared with the existing layout area detection, it often only considers the segmentation of content areas within the layout, and these areas often do not consider overlapping situations, that is, an area must be one of the above-mentioned layout content areas. However, such a design cannot solve the situation of in-line formulas, that is, if there is a formula inside a paragraph, then the area belongs to both a text paragraph and an in-line formula. Therefore, when designing the layout category, this solution adds in-line formulas, unnumbered inter-line formulas, and numbered inter-line formulas. The reason for considering the formulas with and without number separately is that the alignment of the formulas with number is often special. The formula is generally centered, and the number is generally right-aligned. Specifically, it can be as follows Figure 6 As shown, the inline formula is in the middle of the normal text and appears on the same line. The difference between the unnumbered inline formula and the inline formula is that the inline formula is often on a separate line and has no serial number. The difference between the inline formula with a serial number and the unnumbered inline formula is that the formula often has a formula at the end of the line, such as (2-1), (2), (Formula 2), etc.

[0097] Optionally, after detecting at least one layout content area and the content type of the layout content area in the text image to be identified, a document recognition preview page can also be displayed. The document recognition preview page is used to display the detected layout content area and the area type of the layout content area. For example, a document recognition preview page can be displayed. The document recognition preview page includes a restoration control and the layout content area and the content type of the layout content area marked in the document image to be identified. In response to the triggering operation of the restoration control, the text content corresponding to the layout content area is identified in the document image to be identified, and the layout content of each layout content area is determined based on the text content and the content type.

[0098] 104. Identify text content corresponding to the layout content area in the document image to be identified, and determine the layout content of the layout content area based on the text content and content type.

[0099] The layout content refers to the content within the restored layout content area after the layout is restored in the document image to be identified.

[0100] There are many ways to identify the text content corresponding to the publication page content area in the document image to be identified, which can be specifically as follows:

[0101] For example, text recognition can be performed in the document image to be identified to obtain a text content set and the text position corresponding to each text content in the text content set, match the text position with the area position of each layout content area, and filter out the text content that matches the area position of the layout content area in the text content set, thereby obtaining the text content corresponding to the layout content area.

[0102] Among them, there are many ways to perform text recognition on the document image to be recognized. For example, an OCR (optical character recognition) network can be used to detect and recognize text lines to extract text from the document image to be recognized, thereby obtaining a text content set and the text position corresponding to each text content in the text content set. Alternatively, other text recognition networks can be used to perform text recognition on the document image to be recognized, thereby obtaining a text content set and the text position corresponding to each text content in the text content set.

[0103] After identifying the text content corresponding to the publication page content area, the layout content of each layout content area can be determined based on the text content and content type. There are many ways to determine the layout content. For example, when the content type is an image, the image corresponding to the publication page content area is identified in the document image to be identified to obtain the layout content. When the content type is non-image, the text type of the layout content is determined according to the content type, and the text content is converted into the layout content corresponding to the text type.

[0104] The text type is used to indicate the type of text represented in the layout content. There can be multiple text types, for example, basic text, table text, and formula text. According to the content type, there can be multiple ways to determine the text type of the layout content. For example, when the content type is a table, the text type of the layout content can be determined as table text. When the content type is a formula, the text type of the layout content can be determined as formula text. When the content type is a content type other than a table or formula, such as a title, header, footer, or paragraph, the text type of the layout content can be determined as basic text.

[0105] After determining the text type, the text content can be converted into the layout content corresponding to the text type. There are many ways to convert it. For example, when the text type is basic text, the text format of the layout content area is obtained, and based on the text format, the text content is formatted to obtain the layout content. When the text type is table text, the text content is converted into table content, and the table content is used as the layout content. When the text type is formula text, the text content is converted into formula content according to the formula type corresponding to the formula text, and the formula content is used as the layout content.

[0106] Among them, there are many ways to adjust the format of text content based on the text format. For example, based on the text format, the basic text format such as font and paragraph of the text in the text content can be adjusted, so that the text content after format adjustment can be obtained, and the text content after format adjustment can be used as the layout content.

[0107] There are many ways to convert text content into table content. For example, you can obtain the table format of the layout content area, create a basic table based on the table format, and then add the text content to the basic table to obtain the table content.

[0108] Among them, according to the formula type corresponding to the formula text, there are many ways to convert the text content into formula content. For example, the layout information of the layout content area can be obtained, and the basic formula format can be extracted from the layout information. When the formula type is an in-line formula, the basic formula format is used as the formula format of the layout content. When the formula type is an inter-line formula, the basic formula format is adjusted according to the type of the inter-line formula to obtain the formula format of the layout content, and the text content is converted into the formula content corresponding to the formula format.

[0109] Among them, inline formulas are formulas within text paragraphs, and interline formulas are formulas between text paragraphs. The basic formula format can be some basic formats of the formula, such as font, size, paragraph, formula string format, etc. Depending on the type of interline formula, there are many ways to adjust the basic formula format. For example, you can determine the text alignment of the interline formula based on the type of interline formula, and add the text alignment to the basic formula format to obtain the formula format of the layout content.

[0110] Among them, the text alignment can be used to represent the alignment position of the formula text in the column structure area. Depending on the type of the inter-line formula, there can be multiple ways to determine the text alignment of the inter-line formula. For example, when the inter-line formula is a numbered inter-line formula, the text alignment of the inter-line formula can be determined to be right-aligned. When the inter-line formula is an unnumbered inter-line formula, the text alignment of the inter-line formula can be determined to be center-aligned.

[0111] 105. Generate a target document corresponding to the document image to be identified based on the layout content, and display the target document.

[0112] The target document is an editable document. The so-called editable document can be understood as a document that can be edited. There are many ways to generate the target document corresponding to the document image to be recognized based on the layout content, which can be as follows:

[0113] For example, the regional location information of each layout content area can be obtained, and the layout content can be sorted according to the regional location information to create an initial document in a preset format. Based on the sorting information, the layout content can be written into the initial document to obtain the target document corresponding to the document image to be identified.

[0114] The initial document can be a blank document created in a preset format or an initialized document. Based on the sorting information, there are multiple ways to write the layout content into the initial document. For example, based on the sorting information, the layout content can be directly added to the corresponding position of the initial document, thereby obtaining a target document corresponding to the document image to be recognized. For example, taking the initial document as a Word document in docx format, the paragraph content, table content, formula content, title content, and header and footer content can be written into the Word document in docx format in order from top to bottom and from left to right, thereby obtaining the target document.

[0115] After generating the target document corresponding to the document image to be identified, the target document can be displayed. There are many ways to display the target document. For example, the target document can be displayed directly, and the target document is an editable document. Alternatively, a document recognition result page can be displayed, which includes thumbnail information of the target document and a display control of the target document. The target document is displayed in response to a trigger operation on the display control.

[0116] Optionally, after displaying the target document, the target document can be edited in response to an editing operation on the target document, thereby obtaining an edited document. There are many ways to edit the target document, such as adding or deleting text content, editing formula content or table content, or adjusting the document format of the target document.

[0117] The process of identifying the target document in the document image to be identified can be regarded as restoring the layout of the document image to be identified. Taking the target document as a word document in docx format as an example, the layout restoration process can be as follows: Figure 7 As shown, the user inputs the document image to be identified, and the image correction module corrects the document image to be identified, and the trained layout detection model is used to detect at least one layout content area and the content type of the layout area content in the document image to be identified, and the OCR text line detection and recognition module is used to identify the text content in the document image to be identified, and the layout detection result is matched with the text content, and the paragraphs, text titles, chart titles, and tables are matched with the text content. Then, the layout detection result, such as the table area and the corresponding text content in the table, is sent to the table recognition module to obtain the table recognition result. At the same time, the formula area is segmented and sent to the formula recognition module to obtain the formula recognition result. Finally, the information in the layout box is written into a word document in docx format according to the layout coordinates in the order from top to bottom and from left to right.

[0118] It should be noted that during the entire layout restoration process, this solution proposes a new layout content area definition logic, that is, a layout detection data annotation design scheme. At the same time, an improved layout detection model is used to complete the detection of the layout content area. The specific details are as follows:

[0119] (1) Layout content area definition logic;

[0120] For example, the existing technology often only considers the segmentation of content areas within the layout, such as paragraphs, tables, pictures, formulas, headers and footers, etc., and these areas often do not consider overlapping situations, that is, an area must be one of the above-mentioned layout content areas. However, such a design cannot solve the situation of in-line formulas, that is, if there is a formula inside a paragraph, then the area belongs to both a text paragraph and an in-line formula; in order to solve this problem, this solution adds in-line formulas, unnumbered interline formulas, and numbered interline formulas when designing the categories of layout content areas. The reason for considering the formulas with and without sequence numbers separately is that the alignment of the formulas with sequence numbers is often special. The formulas are generally aligned in the center, and the sequence numbers are generally aligned to the right. In addition, the existing technology often relies on rules and strategies for the restoration of layouts with column structures. This solution directly incorporates the structural information of single columns and multiple columns into the layout detection model, which can greatly improve the accuracy of layout content area detection.

[0121] (2) Improvement of layout detection model

[0122] For example, when detecting layout content areas, layout areas overlap. Pixels in the same image may belong to multiple layout areas, so semantic segmentation methods cannot achieve the effect of extracting layout areas all at once. To achieve layout area extraction, this solution combines scaled-yolov4 (an object detection model) and CBNetV2 (a backbone network-based object detection model). This not only extracts all layout information at once, but also targets images of different sizes and complex backgrounds. Figure 8 It is the main algorithm framework of CBNetV2. Figure 9 is the main algorithm structure of scaled-yolov4, where yolov4-p5, yolov4-p6 and yolov4-p7 correspond to networks of different depths and scales respectively. The improvement of the layout detection model in this solution is mainly to replace the backbone network in scaled-yolov4 with the backbone network of CBNetV2. Figure 10As shown, the backbone network of CBNetV2 is used to extract features of the document image to be identified, and the layout structure features and layout content features of the document image to be identified are obtained. Then, the detection network in scaled-yolov4 is used to detect at least one layout content area and the content type of the layout content area in the document image to be identified based on the layout structure features and layout content features. The layout detection model with structural improvement is helpful for extracting complex layout information, and it can also obtain good results for multi-scale layout areas. Therefore, in the document recognition process, a variety of layout area information and layout information with complex layout structures can be extracted at one time, so that when restoring the word document, the structural information can be perfectly restored, and the in-line formulas and inter-line formulas can also be extracted. At the same time, the formulas with serial numbers can be perfectly aligned and restored, thereby greatly improving the accuracy of document recognition.

[0123] It's important to note that the training method for the layout detection model obtained after the structural improvement is mainly similar to the traditional CBNetv2. Although the improved scaled-yolov4 uses the same two original backbones, the features generated by both backbones need to pass through the neck and head parts of scaled-yolov4 and undergo supervised training, also known as auxiliary supervised training, to obtain the trained layout detection model.

[0124] Optionally, in addition to scaled-yolov4, other object detection models can be used for layout detection, such as cascade-rcnn (an object detection model) and DERT (an object detection model). Furthermore, the backbone network can also be replaced with another feature extraction network.

[0125] Optionally, the layout detection model can be deployed on the terminal or on the server. When deployed on the terminal side, the terminal can directly extract the publication structure features and layout content features from the document image to be identified, and then, based on the layout content features and layout content, detect at least one layout content area and the content type of the layout content area in the document image to be identified, identify the text content corresponding to each layout content area in the document image to be identified, determine the layout content corresponding to each layout content area based on the text content and content type, generate the target document corresponding to the document image to be identified according to the layout content, and then display the target document. When deployed on the server side The terminal can perform layout correction on the document image to be identified, and send the corrected document image to the server so that the server can extract the publication structure features and layout content features from the document image to be identified. Then, based on the layout content features and layout content, at least one layout content area and the content type of the layout content area are detected in the document image to be identified, and the text content corresponding to each layout content area is identified in the document image to be identified. Based on the text content and content type, the layout content corresponding to the layout content area is determined, and the target document corresponding to the document image to be identified is generated according to the layout content. Then, the terminal receives the target document returned by the server and displays the target document.

[0126] From the above, it can be seen that after the document recognition page is displayed, the embodiment of the present application responds to the trigger operation of the recognition control in the document recognition page, extracts the publication structure features and layout content features from the document image to be recognized in the document recognition page, and then, based on the layout structure features and layout content features, detects at least one layout content area and the content type of the layout content area in the document image to be recognized, identifies the text content corresponding to the publication content area in the document image to be recognized, and determines the layout content of the layout content area based on the text content and the content type, and then, based on the layout content, generates an editable target document corresponding to the document image to be recognized, and displays the target document; since this scheme can directly extract the publication structure features and layout content features in the document image to be recognized, and then, based on the layout structure features and layout content features, can directly detect the publication content area, so that the publication content area can be accurately identified in a complex layout structure, thereby improving the accuracy of document recognition.

[0127] The method described in the above embodiment will be further described in detail below with examples.

[0128] In this embodiment, the document recognition device is specifically integrated into an electronic device, the electronic device is a terminal, and the target document is a Word document.

[0129] like Figure 11 As shown, a document recognition method, the specific process is as follows:

[0130] 201. The terminal displays a document recognition page.

[0131] For example, the terminal displays a document image acquisition page, which includes an acquisition control. In response to a trigger operation on the acquisition control, the current document image is acquired and a capture image preview page is displayed. The capture image preview page includes the current document image and a confirmation control. In response to a trigger operation on the confirmation control, the current document image is used as the document image to be identified and a document identification page is displayed. Alternatively, a user operation page is displayed, which includes an upload control. In response to a trigger operation on the upload control, a candidate document image list is displayed. In response to a selection operation on the candidate document image list, the candidate document image corresponding to the selection operation is filtered out from the candidate document image list to obtain the document image to be identified, and the document identification page is displayed based on the document image to be identified.

[0132] 202. The terminal extracts the publication structure features and layout content features from the document image to be recognized in response to the triggering operation on the recognition control.

[0133] For example, the terminal corrects the distortion or tilt of the document image to be identified to obtain a corrected document image, obtains the current image size information of the corrected document image, filters out the longest side and the shortest side in the current image size information, scales the longest side to 1280, and scales the shortest side to a multiple of 128, so that the adjusted document image can be obtained.

[0134] The terminal uses the trained layout detection model to perform convolution processing on the adjusted document image to obtain basic image features. The first backbone network (Backbone 1) performs multi-dimensional layout feature extraction on the basic image features to obtain basic layout features corresponding to each dimension, and fuses the basic layout features from the first dimension to the fourth dimension with the basic image features to obtain fused image features. The second backbone network (Backbone 2) performs layout feature extraction of the first dimension on the fused image features to obtain initial layout features corresponding to the first dimension. Then, the basic layout features from the second dimension to the fourth position are fused with the initial layout features corresponding to the first dimension to obtain fused layout features corresponding to the first dimension. The second backbone network performs layout feature extraction of the second dimension on the fused layout features corresponding to the first dimension to obtain initial layout features corresponding to the second dimension. Then, the basic layout features from the third dimension to the fourth dimension are fused with the initial layout features corresponding to the second dimension to obtain fused layout features corresponding to the second dimension. And so on, thereby obtaining layout features corresponding to each dimension.

[0135] Optionally, the trained layout detection model can be set according to actual applications. In addition, it should be noted that the trained layout detection model can be pre-set by maintenance personnel or trained by the document recognition device itself. The training process can be as follows:

[0136] The terminal obtains a document image sample and uses a preset layout detection model to extract image features from the document image sample to obtain sample image features. The first backbone network (Backbone 1) of the preset layout detection model then performs multidimensional layout feature extraction on the sample image features to obtain basic sample layout features. The second backbone network (Backbone 1) then performs multidimensional layout feature extraction on the sample image features and basic sample layout features to obtain target sample layout features.

[0137] Based on the target sample layout features, the terminal predicts at least one layout content area and the content type of the layout content area in the document image sample, obtaining a first predicted layout content area and a first predicted content type, and compares the first predicted layout content area and the first predicted content type with the annotated layout content area and the annotated content type, thereby obtaining main loss information of the document image sample. Based on the basic sample layout features, the terminal predicts at least one layout content area and the content type of the layout content area in the document image sample, obtaining a second predicted layout content area and a second predicted content type, and compares the second predicted layout content area and the second predicted content type with the annotated layout content area and the annotated content type, thereby obtaining auxiliary loss information of the document image sample.

[0138] The terminal fuses the main loss information and the auxiliary loss information to obtain fused loss information, and uses the fused loss information to converge the preset layout detection model to obtain the trained layout detection model. Alternatively, the terminal can also obtain the weighting coefficients corresponding to the main loss information and the auxiliary loss information, respectively, and weight the main loss information and the auxiliary loss information based on the weighting coefficients to obtain weighted main loss information and weighted auxiliary loss information, and fuse the weighted main loss information and the weighted auxiliary loss information to obtain fused loss information, and converge the preset layout detection model based on the fused loss information to obtain the trained layout detection model.

[0139] 203. The terminal detects at least one layout content area and a content type of the layout content area in the document image to be identified based on the layout structure features and the layout content features.

[0140] For example, the terminal uses the detection network in the trained layout detection model to detect at least one layout structure region and the region type corresponding to the layout structure region in the document image to be recognized based on the layout structure features. Based on the region type, the terminal determines the layout structure type of the layout structure region. If the layout structure type is a non-column structure region, the layout structure region is treated as the layout content region, and the region type is used as the content type of the layout content region. When the layout structure type is a column structure area, based on the layout content features, at least one layout content area and the initial content type of the layout content area are identified in the layout structure area; when the initial content type is non-formula, the initial content type is used as the content type of the layout content area; when the initial content type is a formula, the formula position information and formula format information of the layout content area corresponding to the formula are obtained; the layout content area whose initial content type is a paragraph is screened out in the layout content area to obtain the paragraph content area; the paragraph position information of the paragraph content area is obtained; the paragraph position of the paragraph content area is extracted from the paragraph position information; and the formula position is extracted from the formula position information; the formula position is compared with the paragraph position; when the formula position is within the paragraph position, the formula type of the formula is determined to be an inline formula; when the formula position is outside the paragraph position, the formula type of the formula is determined to be an interline formula. Target formula format information corresponding to the inter-line formula is filtered out in the formula format information. When the target formula format information contains a formula serial number, the inter-line formula is determined to be an inter-line formula with a serial number. When the target formula format information does not contain a formula serial number, the inter-line formula is determined to be an inter-line formula without a serial number, thereby obtaining the content type of each layout content area.

[0141] Optionally, after detecting at least one layout content area and the content type of the layout content area in the text image to be identified, the terminal can display a document recognition preview page, which includes a restoration control and the layout content area and the content type of the layout content area marked in the document image to be identified. In response to the triggering operation of the restoration control, the text content corresponding to the layout content area is identified in the document image to be identified, and the layout content of each layout content area is determined based on the text content and the content type.

[0142] 204. The terminal identifies text content corresponding to the publication content area in the document image to be identified.

[0143] For example, the terminal uses an OCR network to detect and recognize text lines to extract text from the document image to be recognized, thereby obtaining a text content set and the text position corresponding to each text content in the text content set. Alternatively, other text recognition networks can be used to perform text recognition on the document image to be recognized, thereby obtaining a text content set and the text position corresponding to each text content in the text content set. The text position is matched with the regional position of each layout content area, and text content that matches the regional position of the layout content area is screened out from the text content set to obtain the text content corresponding to the layout content area.

[0144] 205. The terminal determines the layout content of the layout content area based on the text content and the content type.

[0145] For example, when the content type is an image, the terminal identifies the image corresponding to the publication content area in the document image to be identified to obtain the layout content. When the content type is non-image, when the content type is a table, it can be determined that the text type of the layout content is table text. When the content type is a formula, it can be determined that the text type of the layout content is formula text. When the content type is a content type other than tables and formulas, such as titles, headers, footers or paragraphs, it can be determined that the text type of the layout content is basic text.

[0146] When the text type is basic text, the terminal adjusts the basic text formats such as font and paragraph of the text content based on the text format, so as to obtain the text content after format adjustment, and uses the text content after format adjustment as the layout content. When the text type is table text, the table format of the layout content area is obtained, and based on the table format, a basic table is created, and then the text content is added to the basic table to obtain the table content. When the text type is formula text, the layout information of the layout content area is obtained, and the basic formula format is extracted from the layout information. When the formula type is an inline formula, the basic formula format is used as the formula format of the layout content. When the formula type is an interline formula, and the interline formula is a numbered interline formula, the text alignment of the interline formula can be determined to be right-aligned. When the interline formula is an unnumbered interline formula, the text alignment of the interline formula can be determined to be center-aligned. The text alignment is added to the basic formula format to obtain the formula format of the layout content, and the text content is converted into the formula content corresponding to the formula format.

[0147] 206. The terminal generates a word document corresponding to the document image to be recognized based on the layout content.

[0148] For example, the terminal can obtain the regional location information of each layout content area, and sort the layout content according to the regional location information, create an initial word document in docx format, and write the paragraph content, table content, formula content, title content, header and footer content in order from top to bottom and from left to right into the initial word document in docx format, so that the word document corresponding to the document image to be recognized can be obtained.

[0149] 207. The terminal displays the word document corresponding to the document image to be recognized.

[0150] For example, the terminal can directly display the word document corresponding to the image of the document to be identified, which is an editable document. Alternatively, it can also display a document recognition result page, which includes thumbnail information of the target document and a display control of the target document. In response to a trigger operation on the display control, the word document corresponding to the image of the document to be identified is displayed.

[0151] Optionally, after displaying the target document, the terminal can also respond to editing operations on the target document. There are many ways to edit the target document, such as adding or deleting text content, or editing formula content or table content, or adjusting the document format of the target document, etc., to obtain the edited document. From the above, it can be seen that after the terminal of this embodiment displays the document recognition page, in response to the trigger operation of the recognition control in the document recognition page, the publication structure features and layout content features are extracted from the document image to be recognized in the document recognition page, and then, based on the layout structure features and layout content features, at least one layout content area and the content type of the layout content area are detected in the document image to be recognized, the text content corresponding to the publication content area in the document image to be recognized is recognized, and based on the text content and the content type, the layout content of the layout content area is determined, and then, based on the layout content, an editable target document corresponding to the document image to be recognized is generated, and the target document is displayed; since this scheme can directly extract the publication structure features and layout content features in the document image to be recognized, and then, based on the layout structure features and layout content features, the publication content area can be directly detected, so that the publication content area can be accurately identified in a complex layout structure, thereby improving the accuracy of document recognition.

[0152] In order to better implement the above method, an embodiment of the present invention also provides a document recognition device, which can be integrated into an electronic device, such as a server or terminal, and the terminal can include a tablet computer, a laptop computer and / or a personal computer.

[0153] For example, Figure 12As shown, the document recognition device may include a display unit 301, an extraction unit 302, a detection unit 303, a determination unit 304, and a generation unit 305, as follows:

[0154] (1) Display unit 301;

[0155] The display unit 301 is used to display a document recognition page, which includes a document image to be recognized and a recognition control.

[0156] For example, the display unit 301 can be specifically used to display a document image acquisition page, which includes an acquisition control. In response to a trigger operation on the acquisition control, the current document image is acquired and a capture image preview page is displayed. The capture image preview page includes the current document image and a confirmation control. In response to a trigger operation on the confirmation control, the current document image is used as the document image to be identified, and a document identification page is displayed. Alternatively, a user operation page is displayed, which includes an upload control. In response to a trigger operation on the upload control, a candidate document image list is displayed. In response to a selection operation on the candidate document image list, the candidate document image corresponding to the selection operation is filtered out from the candidate document image list to obtain the document image to be identified, and the document identification page is displayed based on the document image to be identified.

[0157] (2) extraction unit 302;

[0158] The extraction unit 302 is configured to extract the publication layout structure features and layout content features from the document image to be identified in response to a triggering operation on the identification control.

[0159] For example, the extraction unit 302 can be specifically used to respond to a trigger operation on the control to be identified, perform layout correction on the document image to be identified, obtain a corrected document image, adjust the image size of the corrected document image, obtain an adjusted document image, use the trained layout detection model to extract image features on the adjusted document image to obtain basic image features, perform multi-dimensional layout feature extraction on the basic image features to obtain basic layout features of each dimension, and perform multi-dimensional layout feature extraction on the basic image features based on the basic layout features to obtain layout structure features and layout content features of the document image to be identified.

[0160] (3) Detection unit 303;

[0161] The detection unit 303 is configured to detect at least one layout content area and a content type of the layout content area in the document image to be identified based on the layout structure features and the layout content features.

[0162] For example, the detection unit 303 can be specifically used to detect at least one layout structure area and the area type corresponding to the layout structure area in the document image to be identified based on the layout structure features, determine the layout structure type of the layout structure area based on the area type, and identify at least one layout content area and the content type of the layout content area in the layout structure area based on the layout content features and the layout structure type.

[0163] (4) determining unit 304;

[0164] The determination unit 304 is configured to identify text content corresponding to the page content area in the document image to be identified, and determine page content of the page content area based on the text content and the content type.

[0165] For example, the determination unit 304 can be specifically used to perform text recognition in the document image to be recognized, obtain a text content set and the text position corresponding to each text content in the text content set, match the text position with the regional position of each layout content area, and filter out text content in the text content set that matches the regional position of the layout content area, thereby obtaining the text content corresponding to the layout content area. When the content type is an image, the image corresponding to the layout content area is identified in the document image to be recognized to obtain the layout content. When the content type is non-image, the text type of the layout content is determined based on the content type, and the text content is converted into the layout content corresponding to the text type.

[0166] (5) generating unit 305;

[0167] The generating unit 305 is used to generate a target document corresponding to the document image to be recognized according to the layout content, and display the target document, which is an editable document.

[0168] For example, the generation unit 305 can be specifically used to obtain the regional location information of each layout content area, and sort the layout content according to the regional location information, create an initial document in a preset format, and write the layout content into the initial document based on the sorting information, obtain the target document corresponding to the document image to be identified, and directly display the target document, which is an editable document. Alternatively, a document recognition result page can also be displayed, which includes the thumbnail information of the target document and the display control of the target document, and displays the target document in response to a trigger operation on the display control.

[0169] Optionally, the document recognition device may further include a training unit 306, such as Figure 13 As shown, the specific details can be as follows:

[0170] The training unit 306 is used to train the preset layout detection model to obtain a trained layout detection model.

[0171] For example, the training unit 306 can be specifically used to obtain document image samples, and use a preset layout detection model to extract layout features of the document image samples to obtain basic sample layout features, identify target sample layout features in the basic sample layout features, and determine the main loss information of the document image sample based on the target sample layout features, determine the auxiliary loss information of the document image sample based on the basic sample layout features, and converge the preset layout detection model based on the main loss information and the auxiliary loss information to obtain a trained layout detection model.

[0172] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.

[0173] As can be seen from the above, in this embodiment, after the display unit 301 displays the document recognition page, the extraction unit 302 extracts the publication structure features and layout content features from the document image to be recognized in the document recognition page in response to the trigger operation of the recognition control in the document recognition page. Then, the detection unit 303 detects at least one layout content area and the content type of the layout content area in the document image to be recognized based on the layout structure features and layout content features. The determination unit 304 identifies the text content corresponding to the publication content area in the document image to be recognized, and determines the layout content of the layout content area based on the text content and the content type. Then, the generation unit 305 generates an editable target document corresponding to the document image to be recognized based on the layout content, and displays the target document. Since this scheme can directly extract the publication structure features and layout content features in the document image to be recognized, and then, the publication content area can be directly detected based on the layout structure features and layout content features, so that the publication content area can be accurately identified in a complex layout structure, thereby improving the accuracy of document recognition.

[0174] An embodiment of the present invention further provides an electronic device, such as Figure 14 , which shows a schematic structural diagram of an electronic device involved in an embodiment of the present invention, specifically:

[0175] The electronic device may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will understand that Figure 14 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0176] Processor 401 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and circuits. It performs various functions of the electronic device and processes data by running or executing software programs and / or modules stored in memory 402 and accessing data stored in memory 402. Optionally, processor 401 may include one or more processing cores. Preferably, processor 401 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.

[0177] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0178] The electronic device also includes a power supply 403 for supplying power to various components. Preferably, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0179] The electronic device may further include an input unit 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0180] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0181] A document recognition page is displayed, which includes a document image to be recognized and a recognition control. In response to a trigger operation on the recognition control, publication structure features and layout content features are extracted from the document image to be recognized. Based on the layout structure features and layout content features, at least one layout content area and the content type of the layout content area are detected in the document image to be recognized. The text content corresponding to the publication content area is recognized in the document image to be recognized, and based on the text content and the content type, the layout content of the layout content area is determined. Based on the layout content, a target document corresponding to the document image to be recognized is generated, and the target document is displayed. The target document is an editable document.

[0182] For example, the electronic device displays a document image acquisition page, which includes an acquisition control. In response to a trigger operation on the acquisition control, the current document image is acquired and a capture image preview page is displayed. The capture image preview page includes the current document image and a confirmation control. In response to a trigger operation on the confirmation control, the current document image is used as the document image to be identified and a document identification page is displayed. Alternatively, the electronic device displays a user operation page, which includes an upload control. In response to a trigger operation on the upload control, a candidate document image list is displayed. In response to a selection operation on the candidate document image list, the candidate document image corresponding to the selection operation is filtered out from the candidate document image list to obtain the document image to be identified, and the document identification page is displayed based on the document image to be identified. In response to a trigger operation for a control to be identified, layout correction is performed on the document image to be identified to obtain a corrected document image. The image size of the corrected document image is adjusted to obtain an adjusted document image. Image feature extraction is performed on the adjusted document image using a trained layout detection model to obtain basic image features. Multi-dimensional layout feature extraction is performed on the basic image features to obtain basic layout features of each dimension. Based on the basic layout features, multi-dimensional layout feature extraction is performed on the basic image features to obtain layout structure features and layout content features of the document image to be identified. Based on the layout structure features, at least one layout structure area and the area type corresponding to the layout structure area are detected in the document image to be identified. Based on the area type, the layout structure type of the layout structure area is determined. Based on the layout content features and the layout structure type, at least one layout content area and the content type of the layout content area are identified in the layout structure area. Text recognition is performed in a document image to be recognized, obtaining a text content set and the text position corresponding to each text content in the text content set. The text positions are matched with the regional positions of each layout content area, and text content matching the regional positions of the layout content area is screened out from the text content set to obtain the text content corresponding to the layout content area. When the content type is an image, the image corresponding to the layout content area is recognized in the document image to be recognized to obtain the layout content. When the content type is non-image, the text type of the layout content is determined based on the content type, and the text content is converted into layout content corresponding to the text type. Regional position information for each layout content area is obtained, and the layout content is sorted based on the regional position information to create an initial document in a preset format. Based on the sorting information, the layout content is written to the initial document to obtain a target document corresponding to the document image to be recognized. The target document is directly displayed, and the target document is an editable document. Alternatively, a document recognition result page may be displayed, which includes thumbnail information of the target document and a display control for the target document. The target document is displayed in response to a triggering operation on the display control.When an editing operation on a target document is detected, the target document is edited in response to the editing operation on the target document, thereby obtaining an edited document.

[0183] The specific implementation of the above operations can be found in the previous embodiments and will not be described in detail here.

[0184] From the above, it can be seen that after the document recognition page is displayed, the embodiment of the present application responds to the trigger operation of the recognition control in the document recognition page, extracts the publication structure features and layout content features from the document image to be recognized in the document recognition page, and then, based on the layout structure features and layout content features, detects at least one layout content area and the content type of the layout content area in the document image to be recognized, identifies the text content corresponding to the publication content area in the document image to be recognized, and determines the layout content of the layout content area based on the text content and the content type, and then, based on the layout content, generates an editable target document corresponding to the document image to be recognized, and displays the target document; since this scheme can directly extract the publication structure features and layout content features in the document image to be recognized, and then, based on the layout structure features and layout content features, can directly detect the publication content area, so that the publication content area can be accurately identified in a complex layout structure, thereby improving the accuracy of document recognition.

[0185] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0186] To this end, an embodiment of the present invention provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any document recognition method provided in an embodiment of the present invention. For example, the instructions may execute the following steps:

[0187] A document recognition page is displayed, which includes a document image to be recognized and a recognition control. In response to a trigger operation on the recognition control, publication structure features and layout content features are extracted from the document image to be recognized. Based on the layout structure features and layout content features, at least one layout content area and the content type of the layout content area are detected in the document image to be recognized. The text content corresponding to the publication content area is recognized in the document image to be recognized, and based on the text content and the content type, the layout content of the layout content area is determined. Based on the layout content, a target document corresponding to the document image to be recognized is generated, and the target document is displayed. The target document is an editable document.

[0188] For example, a document image acquisition page is displayed, which includes an acquisition control. In response to a trigger operation on the acquisition control, the current document image is acquired, and an acquisition image preview page is displayed. The acquisition image preview page includes the current document image and a confirmation control. In response to a trigger operation on the confirmation control, the current document image is used as the document image to be identified, and a document identification page is displayed. Alternatively, a user operation page is displayed, which includes an upload control. In response to a trigger operation on the upload control, a candidate document image list is displayed. In response to a selection operation on the candidate document image list, the candidate document image corresponding to the selection operation is filtered out from the candidate document image list to obtain the document image to be identified, and a document identification page is displayed based on the document image to be identified. In response to a trigger operation for a control to be identified, layout correction is performed on the document image to be identified to obtain a corrected document image. The image size of the corrected document image is adjusted to obtain an adjusted document image. Image feature extraction is performed on the adjusted document image using a trained layout detection model to obtain basic image features. Multi-dimensional layout feature extraction is performed on the basic image features to obtain basic layout features of each dimension. Based on the basic layout features, multi-dimensional layout feature extraction is performed on the basic image features to obtain layout structure features and layout content features of the document image to be identified. Based on the layout structure features, at least one layout structure area and the area type corresponding to the layout structure area are detected in the document image to be identified. Based on the area type, the layout structure type of the layout structure area is determined. Based on the layout content features and the layout structure type, at least one layout content area and the content type of the layout content area are identified in the layout structure area. Text recognition is performed in a document image to be recognized, obtaining a text content set and the text position corresponding to each text content in the text content set. The text positions are matched with the regional positions of each layout content area, and text content matching the regional positions of the layout content area is screened out from the text content set to obtain the text content corresponding to the layout content area. When the content type is an image, the image corresponding to the layout content area is recognized in the document image to be recognized to obtain the layout content. When the content type is non-image, the text type of the layout content is determined based on the content type, and the text content is converted into layout content corresponding to the text type. Regional position information for each layout content area is obtained, and the layout content is sorted based on the regional position information to create an initial document in a preset format. Based on the sorting information, the layout content is written to the initial document to obtain a target document corresponding to the document image to be recognized. The target document is directly displayed, and the target document is an editable document. Alternatively, a document recognition result page may be displayed, which includes thumbnail information of the target document and a display control for the target document. The target document is displayed in response to a triggering operation on the display control.When an editing operation on a target document is detected, the target document is edited in response to the editing operation on the target document, thereby obtaining an edited document.

[0189] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0190] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0191] Since the instructions stored in the computer-readable storage medium can execute the steps in any document recognition method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any document recognition method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0192] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the aforementioned document recognition or document restoration aspects.

[0193] The above is a detailed introduction to a document recognition method, device, electronic device and computer-readable storage medium provided in an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A document recognition method, characterized in that: include: Displaying a document recognition page, wherein the document recognition page includes a document image to be recognized and a recognition control; In response to a triggering operation on the recognition control, performing layout correction on the document image to be recognized to obtain a corrected document image; Adjusting the image size of the corrected document image to obtain an adjusted document image; A trained layout detection model is used to extract publication structure features and layout content features from the adjusted document image, wherein the trained layout detection model is obtained by replacing the backbone network in scaled-yolov4 with the backbone network of CBNetV2; detecting, in the document image to be identified, at least one layout structure region and a region type corresponding to the layout structure region according to the layout structure feature; Determining the layout structure type of the layout structure area according to the area type; When the layout structure type is a column structure area, based on the layout content feature, identifying at least one layout content area and an initial content type of the layout content area in the layout structure area, wherein the column structure area is an area for content columnarization in the document contained in the document image to be identified; When the initial content type is a formula, obtaining formula position information and formula format information of the layout content area corresponding to the formula, and determining the formula type of the formula based on the formula position information and formula format information to obtain the content type of the layout content area, wherein the formula type includes inline formulas and interline formulas. The inline formula is a formula in a text paragraph, and the interline formula is a formula between the text paragraphs. Identifying text content corresponding to the layout content area in the document image to be identified, and determining layout content of the layout content area based on the text content and content type; According to the layout content, a target document corresponding to the document image to be identified is generated and displayed, and the target document is an editable document.

2. The document recognition method according to claim 1, wherein: The method of extracting the publication structure features and the layout content features from the adjusted document image using the trained layout detection model includes: Using the trained layout detection model to extract image features from the adjusted document image to obtain basic image features; Performing multi-dimensional layout feature extraction on the basic image features to obtain basic layout features of each dimension; Based on the basic layout features, multi-dimensional layout feature extraction is performed on the basic image features to obtain the layout structure features and layout content features of the document image to be identified.

3. The document recognition method according to claim 2, characterized in that: The step of performing multi-dimensional layout feature extraction on the basic image features based on the basic layout features to obtain layout structure features and layout content features of the document image to be identified includes: Fusing the basic layout features with the basic image features to obtain fused image features; Performing layout feature extraction on the fused image features to obtain initial layout features corresponding to the target dimension; The layout structure features and layout content features of the document image to be identified are identified from the initial layout features.

4. The document recognition method according to claim 3, wherein: The step of identifying the layout structure features and layout content features of the image of the document to be identified from the initial layout features includes: sorting the dimension information of the basic layout features, and screening out target basic layout features exceeding the target dimension from the basic layout features according to the sorting information; Fusing the target basic layout features with the initial layout features to obtain fused layout features; The layout structure features and layout content features of the document image to be identified are extracted from the fused layout features.

5. The document recognition method according to claim 4, characterized in that: Extracting the layout structure features and layout content features of the document image to be identified from the fused layout features includes: The fused layout features are used as the fused image features, and the step of extracting layout features from the fused image features to obtain initial layout features corresponding to the target dimension is returned to the step of obtaining layout features corresponding to each dimension until the target basic layout features no longer exist. Obtaining a weighting coefficient corresponding to each dimension, and weighting the layout features based on the weighting coefficients to obtain weighted layout features; The layout structure features and layout content features of the document image to be identified are identified from the weighted layout features.

6. The document recognition method according to claim 2, characterized in that: Before extracting image features from the adjusted document image using the trained layout detection model to obtain basic image features, the method further includes: Obtaining a document image sample, and extracting layout features from the document image sample using a preset layout detection model to obtain basic sample layout features; Identifying target sample layout features from the basic sample layout features, and determining the backbone loss information of the document image sample based on the target sample layout features; According to the basic sample layout features, the auxiliary loss information of the document image sample is determined, and based on the trunk loss information and the auxiliary loss information, the preset layout detection model is converged to obtain the trained layout detection model.

7. The document recognition method according to claim 1, wherein: After determining the layout structure type of the layout structure area according to the area type, the method further includes: When the layout structure type is a non-column structure area, the layout structure area is used as the layout content area, and the area type is used as the content type of the layout content area.

8. The document recognition method according to claim 1, wherein: After identifying at least one layout content area and an initial content type of the layout content area in the layout structure area based on the layout content feature, the method further includes: When the initial content type is non-formula, the initial content type is used as the content type of the layout content area.

9. The document recognition method according to any one of claims 1 to 6, characterized in that: The determining of the layout content of the layout content area based on the text content and the content type includes: When the content type is an image, identifying an image corresponding to the layout content area in the document image to be identified to obtain the layout content; When the content type is non-image, the text type of the layout content is determined according to the content type, and the text content is converted into layout content corresponding to the text type.

10. The document recognition method according to claim 9, characterized in that: The converting the text content into layout content corresponding to the text type includes: When the text type is basic text, obtaining the text format of the layout content area, and adjusting the format of the text content based on the text format to obtain the layout content; When the text type is table text, converting the text content into table content, and using the table content as the layout content; When the text type is a formula text, the text content is converted into a formula content according to the formula type corresponding to the formula text, and the formula content is used as the layout content.

11. The document recognition method according to claim 10, characterized in that: The converting the text content into formula content according to the formula type corresponding to the formula text includes: Acquiring layout information of the layout content area and extracting a basic formula format from the layout information; When the formula type is an inline formula, the basic formula format is used as the formula format of the layout content; When the formula type is an interline formula, adjusting the basic formula format according to the type of the interline formula to obtain the formula format of the layout content; The text content is converted into formula content corresponding to the formula format.

12. The document recognition method according to claim 11, characterized in that: The step of adjusting the basic formula format according to the type of the interline formula to obtain the formula format of the layout content includes: Determining a text alignment of the inter-line formula according to the type of the inter-line formula; The text alignment is added to the basic formula format to obtain the formula format of the layout content.

13. The document recognition method according to any one of claims 1 to 5, characterized in that: Generating a target document corresponding to the document image to be recognized according to the layout content includes: Obtaining the area position information of each of the layout content areas, and sorting the layout content according to the area position information; An initial document in a preset format is created, and based on the sorting information, the layout content is written into the initial document to obtain a target document corresponding to the document image to be recognized.

14. A document recognition device, characterized in that: include: A display unit, configured to display a document recognition page, wherein the document recognition page includes an image of a document to be recognized and a recognition control; an extraction unit configured to, in response to a triggering operation on the recognition control, perform layout correction on the document image to be recognized to obtain a corrected document image, adjust the image size of the corrected document image to obtain an adjusted document image, and extract publication structure features and layout content features from the adjusted document image using a trained layout detection model, wherein the trained layout detection model is obtained by replacing the backbone network in scaled-yolov4 with the backbone network of CBNetV2; a detection unit configured to detect, in the document image to be identified, at least one layout structure area and an area type corresponding to the layout structure area based on the layout structure feature; determine, based on the area type, the layout structure type of the layout structure area; and, when the layout structure type is a column structure area, identify, based on the layout content feature, at least one layout content area and an initial content type of the layout content area in the layout structure area, wherein the column structure area is an area for performing content columns in the document contained in the document image to be identified; and, when the initial content type is a formula, obtain formula position information and formula format information of the layout content area corresponding to the formula; and determine, based on the formula position information and formula format information, a formula type of the formula to obtain the content type of the layout content area, wherein the formula type includes inline formulas and interline formulas, wherein the inline formulas are formulas in text paragraphs, and the interline formulas are formulas between the text paragraphs; a determination unit, configured to identify text content corresponding to the layout content area in the document image to be identified, and determine layout content of the layout content area based on the text content and content type; A generating unit is used to generate a target document corresponding to the document image to be identified according to the layout content, and display the target document, wherein the target document is an editable document.

15. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps of the document recognition method according to any one of claims 1 to 13.

16. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the document recognition method according to any one of claims 1 to 13 are implemented.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the document recognition method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Image file transfer method, device and equipment based on OCR and readable storage medium

    CN109933756A

  • Image processing method and device, computer equipment and storage medium

    CN111144320A

  • Image recognition method and device, storage medium and terminal

    CN113313066A