A complex document recognition method and medium based on layout analysis and OCR

CN121617115BActive Publication Date: 2026-09-22ZHEJIANG BAORONG MEDIA TECH (ZHEJIANG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511753752.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-09-22
Estimated Expiration
2045-11-26

AI Technical Summary

Technical Problem

[0003]发明人在实现本发明的过程中,发现现有技术存在如下缺陷:目前,传统OCR识别技术主要针对简单版面、单一字体的文档进行字符识别,将图像中的文字转换为可编辑的文本格式,广泛应用于书籍扫描、票据识别等场景,但是对报纸的多栏布局、论文的图文混排、表格的跨页嵌套等复杂结构识别率低,无法将每个板块内容精准分离,无法区分文本区域、标题、图表以及表格等语义单元,导致输出内容逻辑混乱

Benefits of technology

[0019]本发明实施例的技术方案,实时获取待识别文档图像,并通过预处理模块对所述待识别文档图像进行预处理操作,得到待识别标准文档图像;通过版面分析模块对所述待识别标准文档图像进行处理,并结合预先设置的版面分析策略方法,得到至少一个目标区域图像,以及各所述目标区域图像分别对应的目标区域图像类型和区域视觉矩阵特征;根据各目标区域图像类型,通过所述多模态OCR引擎模块来分别每个目标区域图像进行文本分析,得到各文本语义分析结果,以及与各文本语义分析结果分别对应的区域语义矩阵特征;通过所述多模态注意力融合模块,将各所述区域视觉矩阵特征和各所述区域语义矩阵特征进行深度耦合,得到当前视觉文本融合特征;通过后处理与结构化输出模块,根据当前视觉文本融合特征,并结合预设的基于区域坐标与语义标签构建文档对象模型,来生成并反馈复杂文档识别输出结果。解决了无法处理不规则版面,导致准确率差的问题,提高了复杂版面识别的准确率和灵活性,提高了复杂版面识别的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617115B_ABST
    Figure CN121617115B_ABST
Patent Text Reader

Abstract

The application discloses a complex document recognition method and medium based on layout analysis and OCR. A document image to be recognized is acquired in real time and preprocessed by a preprocessing module to obtain a standard document image to be recognized; a layout analysis module is used to process the standard document image to be recognized, and a layout analysis strategy method is combined to obtain at least one target region image, a target region image type and a region visual matrix feature; according to the target region image type, a multi-modal OCR engine module is used to analyze the text of each target region image to obtain a text semantic analysis result and a region semantic matrix feature; a multi-modal attention fusion module is used to obtain a current visual text fusion feature; and a complex document recognition output result is generated and fed back according to the current visual text fusion feature. The application solves the problem that an irregular layout cannot be processed, resulting in poor accuracy, and improves the accuracy and flexibility of complex layout recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a method and medium for recognizing complex documents based on layout analysis and OCR. Background Technology

[0002] As the layouts of academic papers, newspapers, and other media become increasingly complex, the requirements for complex layout recognition technology are becoming more demanding, making it crucial to improve the accuracy of complex layout recognition.

[0003] In the process of developing this invention, the inventors discovered the following shortcomings in existing technologies: Currently, traditional OCR recognition technology mainly targets simple layouts and single-font documents for character recognition, converting text in images into editable text formats. It is widely used in scenarios such as book scanning and invoice recognition. However, it has low recognition rates for complex structures such as multi-column layouts in newspapers, mixed text and images in papers, and nested tables across pages. It cannot accurately separate the content of each section, nor can it distinguish semantic units such as text areas, titles, charts, and tables, leading to logically chaotic output content. Furthermore, based on OCR recognition technology, simple layout analysis cannot handle irregular layouts, resulting in poor accuracy. Summary of the Invention

[0004] This invention provides a method and medium for recognizing complex documents based on layout analysis and OCR, so as to achieve high accuracy and flexibility in recognizing complex layouts.

[0005] According to one aspect of the present invention, a method for recognizing complex documents based on layout analysis and OCR is provided, comprising: acquiring a document image to be recognized in real time, and performing a preprocessing operation on the document image to be recognized through a preprocessing module to obtain a standard document image to be recognized;

[0006] The complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture. The complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module.

[0007] The layout analysis module processes the standard document image to be identified and combines it with a pre-set layout analysis strategy to obtain at least one target region image, as well as the target region image type and region visual matrix features corresponding to each target region image.

[0008] Based on the image type of each target region, the multimodal OCR engine module performs text analysis on each target region image to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result.

[0009] The multimodal attention fusion module deeply couples the visual matrix features and semantic matrix features of each region to obtain the current visual-text fusion features.

[0010] The post-processing and structured output module generates and feeds back complex document recognition output results based on the current visual text fusion features and a pre-defined document object model constructed based on region coordinates and semantic tags.

[0011] According to another aspect of the present invention, a complex document recognition device based on layout analysis and OCR is provided, comprising: a standard document image determination module for real-time acquisition of a document image to be recognized, and a preprocessing module for preprocessing the document image to be recognized to obtain a standard document image to be recognized;

[0012] The complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture. The complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module.

[0013] The target region image type and region visual matrix feature determination module is used to process the standard document image to be identified through the layout analysis module, and combine it with the pre-set layout analysis strategy method to obtain at least one target region image, and the target region image type and region visual matrix features corresponding to each target region image.

[0014] The region semantic matrix feature determination module is used to perform text analysis on each target region image according to the image type of each target region through the multimodal OCR engine module, to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result.

[0015] The current visual-text fusion feature determination module is used to deeply couple the visual matrix features of each region and the semantic matrix features of each region through the multimodal attention fusion module to obtain the current visual-text fusion features;

[0016] The complex document recognition output generation and feedback module is used to generate and feedback complex document recognition output results by combining the current visual text fusion features with a pre-set document object model based on region coordinates and semantic labels, through the post-processing and structured output module.

[0017] According to another aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the complex document recognition method based on layout analysis and OCR as described in any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the complex document recognition method based on layout analysis and OCR as described in any embodiment of the present invention.

[0019] The technical solution of this invention involves acquiring a document image to be recognized in real time, preprocessing it using a preprocessing module to obtain a standard document image, processing it using a layout analysis module, and combining it with a pre-set layout analysis strategy to obtain at least one target region image, as well as the target region image type and region visual matrix features corresponding to each target region image, based on the target region image type, and performing text analysis on each target region image using a multimodal OCR engine module to obtain text semantic analysis results and region semantic matrix features corresponding to each text semantic analysis result, and deeply coupling the region visual matrix features and region semantic matrix features using a multimodal attention fusion module to obtain the current visual-text fusion features, and generating and feeding back complex document recognition output results using a post-processing and structured output module based on the current visual-text fusion features and a pre-set document object model constructed based on region coordinates and semantic tags. This solves the problem of poor accuracy caused by the inability to handle irregular layouts, improves the accuracy and flexibility of complex layout recognition, and increases the efficiency of complex layout recognition.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1This is a flowchart of a complex document recognition method based on layout analysis and OCR provided in Embodiment 1 of the present invention;

[0023] Figure 2 This is a schematic diagram of a complex document recognition device based on layout analysis and OCR according to Embodiment 2 of the present invention;

[0024] Figure 3 This is a schematic diagram of the structure of an electronic device provided according to Embodiment 3 of the present invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "target," "current," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] It is worth noting that the information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse; if the user chooses to refuse, the process will proceed to the expert decision-making process.

[0028] Example 1

[0029] Figure 1The flowchart of a complex document recognition method based on layout analysis and OCR is provided in Embodiment 1 of the present invention. This embodiment is applicable to the case of performing layout analysis on documents with complex layouts. The method can be executed by a complex document recognition device based on layout analysis and OCR, which can be implemented in hardware and / or software.

[0030] Correspondingly, such as Figure 1 As shown, the method includes:

[0031] S110. Acquire the document image to be identified in real time, and perform preprocessing operations on the document image to be identified through the preprocessing module to obtain the standard document image to be identified.

[0032] The complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture. The complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module.

[0033] The document image to be identified can be an image obtained by scanning a document such as a newspaper or a thesis.

[0034] Optionally, the step of preprocessing the document image to be identified through the preprocessing module to obtain a standard document image to be identified includes: performing tilt correction on the document image to be identified using an image tilt correction method in the preprocessing module to obtain a tilt-corrected document image to be identified; acquiring each formula region corresponding to the tilt-corrected document image to be identified, enhancing it using a preset contrast enhancement method, and combining it with a preset noise removal method to obtain a standard document image to be identified.

[0035] Among them, the image tilt correction method can be a method for correcting the image edges of a document.

[0036] Specifically, the image tilt correction method in the preprocessing module is used to correct the tilt of the document image to be recognized, resulting in a tilt-corrected document image. In detail, OpenCV can be used to detect document edges and correct tilt. An angle error can be set, and if the angle error condition is not met, the document image to be recognized needs to be tilt corrected.

[0037] Furthermore, it is necessary to acquire the formula regions corresponding to the tilt-corrected document image to be recognized, and enhance them using contrast enhancement methods. Specifically, CLAHE can be used to enhance the contrast of the formula regions, and then noise removal methods can be combined to denoise the contrast-enhanced image to obtain the standard document image to be recognized. Additionally, blurry parts in the document image to be recognized can be reconstructed using super-resolution methods. Specifically, the ESRGAN model can be used to perform multiple super-resolution on the blurry parts to improve the OCR input quality.

[0038] S120. The standard document image to be identified is processed by the layout analysis module, and combined with the pre-set layout analysis strategy method, at least one target area image is obtained, as well as the target area image type and area visual matrix features corresponding to each target area image.

[0039] Among them, layout analysis strategies can be used to more accurately divide complex layouts.

[0040] Optionally, the step of performing keyword recognition processing on the standard document image to be identified through the layout analysis module, and dividing the standard document image to be identified based on at least one identified keyword to obtain semantic regions for each keyword, includes: determining the document type of the standard document image to be identified through document type recognition and adaptive strategies to determine the target document type; mapping and matching the target document type with a preset document type layout rule library to obtain a set of target keywords; and performing keyword recognition processing on the standard document image to be identified through the layout analysis module based on the set of target keywords, and dividing the standard document image to be identified based on at least one identified keyword to obtain semantic regions for each keyword.

[0041] In this embodiment, the first step is to perform document type identification on the standard document image to be identified, which yields the target document type. Specifically, the target document type can include document types such as newspapers or academic papers. For newspapers, this may include multi-column news, advertisements, or images; for academic papers, it may include a title, author, abstract, and main text. Based on this information, a document type layout rule base can be constructed, allowing for the determination of the keyword set according to the document type. Then, a mapping and matching process can be performed between the target document type and the document type layout rule base to obtain the target keyword set. Correspondingly, each keyword in the target keyword set can be used to perform keyword identification and image segmentation processing on the standard document image to be identified, resulting in a keyword semantic region matching each keyword.

[0042] Optionally, the step of processing the standard document image to be identified through the layout analysis module, and combining it with a pre-set layout analysis strategy to obtain at least one target region image, and the target region image type and region visual matrix features corresponding to each target region image, includes: performing keyword recognition processing on the standard document image to be identified through the layout analysis module, and dividing the standard document image to be identified according to the at least one identified keyword to obtain each keyword semantic region; constructing a relationship graph for each keyword semantic region to obtain a target document element relationship graph G; wherein, N represents the set of neighboring nodes, which includes text block elements, table elements, and image elements; E represents a relation edge, which consists of spatial and semantic relationships; based on the target document element relationship graph, the graph structure reasoning path is optimized using a pre-set node state update rule method combined with the layout analysis strategy method to obtain at least one target region image; wherein, the layout analysis strategy method is based on a Markov decision strategy; the node state update rule method is... , Let n be the state of node n at time t; Represents the edge feature encoding function. Let m be the feature vector of edge nm; m is the set of neighboring nodes. The nodes in the text; among them, the layout analysis strategy method is... ;in, This is a function for layout analysis strategies; in the current state The probability distribution of the next choice To analyze the action, For global context vectors; The trainable weight matrix maps the input features to the action space dimension; The current state of the focused node is determined; based on the images of each target region, the image type of each target region and the visual matrix features of each region are obtained.

[0043] In this embodiment, after determining the semantic regions of each keyword, it is necessary to identify the text block elements, table elements, and image elements in each keyword semantic region as nodes, and construct the relationship graph by forming the relationship edges between the nodes based on spatial and semantic relationships, so as to obtain the target document element relationship graph.

[0044] Furthermore, based on the target document element relationship graph, the graph structure reasoning path can be optimized by using node state update rules and combining layout analysis strategies (which can be strategies based on Markov decision strategies) to obtain at least one target region image. The target region image can be generated by segmenting the standard document image to be identified into multiple region images.

[0045] Correspondingly, features can be extracted for each target region type to obtain the corresponding visual matrix features. Alternatively, each target region image can be identified to determine its corresponding type. Specifically, this can be achieved by combining the multimodal classification model in the region classification module to perform fine-grained classification of segmented regions, i.e., identifying one of the following image types: title region image type, body text region image type, table cell region image type, formula region image type, and special symbol region image type, thus obtaining the target region image type. During the training process of the multimodal classification model, prior knowledge of document types can be incorporated to correct the classification results, further improving the classification accuracy of the model.

[0046] S130. Based on the image type of each target region, the multimodal OCR engine module is used to perform text analysis on each target region image to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result.

[0047] Among them, the multimodal OCR engine module can call different dedicated OCR engine modules to perform text analysis operations for different types of target region images.

[0048] For example, for title area image types or body area image types, the SVTR module in the multimodal OCR engine module can be used for text analysis, and the resolution of the input target area image can be set to... For table cell range images, the multimodal OCR engine module can be used. The module performs text analysis and allows setting the resolution of the input target region image. For image types containing formula regions or special symbols, the multimodal OCR engine module can be used. The module performs text analysis and allows setting the resolution of the input target region image. Additionally, for image types in table cell ranges, the OCR results can be corrected by combining row or column layout information.

[0049] Optionally, the step of performing text analysis on each target region image according to the type of each target region image, using the multimodal OCR engine module to obtain semantic analysis results for each text and region semantic matrix features corresponding to each text semantic analysis result, includes: determining each dedicated OCR engine module in the multimodal OCR engine module according to the type of each target region image; wherein, the target region image type includes title region image type, body text region image type, table cell region image type, formula region image type, and special symbol region image type; performing text analysis on the corresponding target region image according to each dedicated OCR engine module to obtain semantic analysis results for each text; and extracting semantic matrix features from each text semantic analysis result to obtain region semantic matrix features.

[0050] In this embodiment, after determining the target region image type and corresponding dedicated OCR engine module for each target region image, it is necessary to perform text analysis on the target region image using the corresponding dedicated OCR engine module to obtain the semantic analysis results of each text. That is, the SVTR module is used to perform text analysis on the title region image type or the body text region image type; through... This module performs text analysis on image types within table cell ranges; through The module performs text analysis on image types of formula regions or special symbol regions; it can determine the semantic analysis results of each text based on the text analysis results.

[0051] S140. Through the multimodal attention fusion module, the visual matrix features and semantic matrix features of each region are deeply coupled to obtain the current visual text fusion features.

[0052] Among them, the multimodal attention fusion module can be a module that deeply couples the visual matrix features and semantic matrix features of each region to obtain fused features.

[0053] Optionally, the step of deeply coupling the visual matrix features and semantic matrix features of each region through the multimodal attention fusion module to obtain the current visual-text fusion features includes: deeply coupling the visual matrix features and semantic matrix features of each region through the multimodal attention fusion module to calculate the cross-modal attention weights. ;in, ; This represents the attention level of the i-th semantic unit to the j-th visual region, and , Represents the semantic query projection matrix. ; The feature vector of the i-th semantic unit has the shape as follows: ; Represents the visual key projection matrix. ; This represents the feature vector of the j-th visual region, with shape [formula missing]. ; V represents the scaling factor; V represents the region visual matrix feature. H represents the feature map height, W represents the feature map width, C represents the number of channels; S represents the region semantic matrix features. N represents the length of the text sequence, and D represents the dimension of each semantic unit. Based on the cross-modal attention weights, the current visual-text fusion features are calculated using a pre-set fusion feature calculation formula; wherein, the fusion feature calculation formula is... ; Represents the cross-modal fusion feature of the i-th semantic unit; By each visual area pass The eigenvalues ​​obtained by projection.

[0054] In this embodiment, according to the formula By deeply coupling the visual matrix features and semantic matrix features of each region, the cross-modal attention weights can be calculated. This allows for bidirectional alignment of visual and semantic features.

[0055] Furthermore, it can be done according to the formula The current visual text fusion features are obtained by calculating the fusion features using the fusion feature calculation formula.

[0056] S150. Through the post-processing and structured output module, based on the current visual text fusion features and combined with the preset document object model constructed based on region coordinates and semantic tags, the complex document recognition output results are generated and fed back.

[0057] In this embodiment, a document object model based on region coordinates and semantic tags can be constructed in the post-processing and structured output module. Matching and localization can be performed based on the current visual-text fusion features to determine the corresponding document object. Complex document recognition output results are then generated based on each document object, and these results are processed. For example, for a newspaper, the recognition results corresponding to the current visual-text fusion features can be identified, which may include detailed information such as the newspaper's header information, the date in the header, the newspaper's column titles, and the text content of the newspaper's columns.

[0058] Optionally, after generating and feeding back the complex document recognition output result by the post-processing and structured output module based on the current visual text fusion features and combined with a preset document object model constructed based on region coordinates and semantic tags, the method further includes: setting corresponding chart position connection addresses for the chart portion in the complex document recognition output result, and determining the target complex document recognition output result by combining it with the text portion in the complex document recognition output result; and storing the target complex document recognition output result in a document using a preset document saving format.

[0059] In this embodiment, for each identified chart portion, a chart location link address can be set and inserted into the corresponding text portion. After the user clicks the chart location link address, they can be redirected to the corresponding chart location storage location. Based on the text portion containing the link addresses of each chart location, the target complex document recognition output result can be determined. Furthermore, the target complex document recognition output result can be stored as a document using a document saving format. For example, structured data in JSON or XML format can be generated, and the text content can be saved using the .md format.

[0060] In addition, for the generated complex document recognition output, key fields can be validated by using regular expressions. For example, parameters such as the unique identifier of the paper's digital object or the newspaper date can be used for validation.

[0061] The technical solution of this invention involves acquiring a document image to be recognized in real time, preprocessing it using a preprocessing module to obtain a standard document image, processing it using a layout analysis module, and combining it with a pre-set layout analysis strategy to obtain at least one target region image, as well as the target region image type and region visual matrix features corresponding to each target region image, based on the target region image type, and performing text analysis on each target region image using a multimodal OCR engine module to obtain text semantic analysis results and region semantic matrix features corresponding to each text semantic analysis result, and deeply coupling the region visual matrix features and region semantic matrix features using a multimodal attention fusion module to obtain the current visual-text fusion features, and generating and feeding back complex document recognition output results using a post-processing and structured output module based on the current visual-text fusion features and a pre-set document object model constructed based on region coordinates and semantic tags. This solves the problem of poor accuracy caused by the inability to handle irregular layouts, improves the accuracy and flexibility of complex layout recognition, and increases the efficiency of complex layout recognition.

[0062] Example 2

[0063] Figure 2 This is a schematic diagram of a complex document recognition device based on layout analysis and OCR provided in Embodiment 2 of the present invention. The complex document recognition device based on layout analysis and OCR provided in this embodiment can be implemented by software and / or hardware, and can be configured in a terminal device or server to implement a complex document recognition method based on layout analysis and OCR in this embodiment of the present invention. Figure 2 As shown, the device includes: a standard document image determination module 210, a target region image type and region visual matrix feature determination module 220, a region semantic matrix feature determination module 230, a current visual text fusion feature determination module 240, and a complex document recognition output result generation and feedback module 250.

[0064] The standard document image determination module 210 is used to acquire the document image to be identified in real time and perform preprocessing operations on the document image to be identified through the preprocessing module to obtain the standard document image to be identified.

[0065] The complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture. The complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module.

[0066] The target region image type and region visual matrix feature determination module 220 is used to process the standard document image to be identified through the layout analysis module, and combine it with the pre-set layout analysis strategy method to obtain at least one target region image, and the target region image type and region visual matrix features corresponding to each target region image.

[0067] The region semantic matrix feature determination module 230 is used to perform text analysis on each target region image according to the image type of each target region through the multimodal OCR engine module, to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result.

[0068] The current visual-text fusion feature determination module 240 is used to deeply couple the visual matrix features of each region and the semantic matrix features of each region through the multimodal attention fusion module to obtain the current visual-text fusion features;

[0069] The complex document recognition output generation and feedback module 250 is used to generate and feedback the complex document recognition output results by using the post-processing and structured output module, based on the current visual text fusion features and combined with a preset document object model constructed based on region coordinates and semantic labels.

[0070] The technical solution of this invention involves acquiring a document image to be recognized in real time, preprocessing it using a preprocessing module to obtain a standard document image, processing it using a layout analysis module, and combining it with a pre-set layout analysis strategy to obtain at least one target region image, as well as the target region image type and region visual matrix features corresponding to each target region image, based on the target region image type, and performing text analysis on each target region image using a multimodal OCR engine module to obtain text semantic analysis results and region semantic matrix features corresponding to each text semantic analysis result, and deeply coupling the region visual matrix features and region semantic matrix features using a multimodal attention fusion module to obtain the current visual-text fusion features, and generating and feeding back complex document recognition output results using a post-processing and structured output module based on the current visual-text fusion features and a pre-set document object model constructed based on region coordinates and semantic tags. This solves the problem of poor accuracy caused by the inability to handle irregular layouts, improves the accuracy and flexibility of complex layout recognition, and increases the efficiency of complex layout recognition.

[0071] Based on the above embodiments, the standard document image determination module 210 can be specifically used to: perform tilt correction on the document image to be identified using the image tilt correction method in the preprocessing module to obtain a tilt-corrected document image to be identified; acquire each formula region corresponding to the tilt-corrected document image to be identified, perform enhancement processing using a preset contrast enhancement method, and combine it with a preset noise removal method to obtain a standard document image to be identified.

[0072] Based on the above embodiments, the target region image type and region visual matrix feature determination module 220 can be specifically used to: perform keyword recognition processing on the standard document image to be identified through the layout analysis module, and divide the standard document image to be identified according to at least one identified keyword to obtain each keyword semantic region; construct a relationship graph for each keyword semantic region to obtain a target document element relationship graph G; wherein, N represents the set of neighboring nodes, which includes text block elements, table elements, and image elements; E represents a relation edge, which consists of spatial and semantic relationships; based on the target document element relationship graph, the graph structure reasoning path is optimized using a pre-set node state update rule method combined with the layout analysis strategy method to obtain at least one target region image; wherein, the layout analysis strategy method is based on a Markov decision strategy; the node state update rule method is... , Let n be the state of node n at time t; Represents the edge feature encoding function. Let m be the feature vector of edge nm; m is the set of neighboring nodes. The nodes in the text; among them, the layout analysis strategy method is... ;in, This is a function for layout analysis strategies; in the current state The probability distribution of the next choice To analyze the action, For global context vectors; The trainable weight matrix maps the input features to the action space dimension; The current state of the focused node is determined; based on the images of each target region, the image type of each target region and the visual matrix features of each region are obtained.

[0073] Based on the above embodiments, the target region image type and region visual matrix feature determination module 220 can also be specifically used to: determine the target document type by identifying the document type of the standard document image to be identified through document type recognition and adaptive strategies; map and match the target document type with a preset document type layout rule library to obtain a target keyword set; and perform keyword recognition processing on the standard document image to be identified through the layout analysis module based on the target keyword set, and divide the standard document image to be identified based on at least one identified keyword to obtain the semantic regions of each keyword.

[0074] Based on the above embodiments, the region semantic matrix feature determination module 230 can be specifically used to: determine each dedicated OCR engine module in the multimodal OCR engine module according to the target region image type; wherein, the target region image type includes title region image type, body text region image type, table cell region image type, formula region image type, and special symbol region image type; perform text analysis on the corresponding target region image according to each dedicated OCR engine module to obtain each text semantic analysis result; and extract semantic matrix features from each text semantic analysis result to obtain the semantic matrix features of each region.

[0075] Based on the above embodiments, the current visual-text fusion feature determination module 240 can be specifically used to: deeply couple the visual matrix features and semantic matrix features of each region through the multimodal attention fusion module to calculate the cross-modal attention weights. ;in, ; This represents the attention level of the i-th semantic unit to the j-th visual region, and , Represents the semantic query projection matrix. ; The feature vector of the i-th semantic unit has the shape as follows: ; Represents the visual key projection matrix. ; This represents the feature vector of the j-th visual region, with shape [formula missing]. ; V represents the scaling factor; V represents the region visual matrix feature. H represents the feature map height, W represents the feature map width, C represents the number of channels; S represents the region semantic matrix features. N represents the length of the text sequence, and D represents the dimension of each semantic unit. Based on the cross-modal attention weights, the current visual-text fusion features are calculated using a pre-set fusion feature calculation formula; wherein, the fusion feature calculation formula is... ; Represents the cross-modal fusion feature of the i-th semantic unit; By each visual area pass The eigenvalues ​​obtained by projection.

[0076] Based on the above embodiments, a target complex document recognition output result storage module is also included, which can be specifically used to: after the post-processing and structured output module generates and feeds back the complex document recognition output result based on the current visual text fusion features and in combination with a preset document object model constructed based on region coordinates and semantic tags, set the corresponding chart position connection address for the chart part in the complex document recognition output result, and determine the target complex document recognition output result in combination with the text part in the complex document recognition output result; and store the target complex document recognition output result in a document using a preset document saving format.

[0077] The complex document recognition device based on layout analysis and OCR provided in the embodiments of the present invention can execute the complex document recognition method based on layout analysis and OCR provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.

[0078] Example 3

[0079] Figure 3A schematic diagram of an electronic device 10, which can be used to implement Embodiment 3 of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0080] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0081] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0082] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as complex document recognition methods based on layout analysis and OCR.

[0083] In some embodiments, the complex document recognition method based on layout analysis and OCR can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the complex document recognition method based on layout analysis and OCR described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the complex document recognition method based on layout analysis and OCR by any other suitable means (e.g., by means of firmware).

[0084] The method includes: acquiring a document image to be recognized in real time, and performing preprocessing operations on the document image to be recognized through a preprocessing module to obtain a standard document image to be recognized; wherein, the complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture; the complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module; the standard document image to be recognized is processed by the layout analysis module, and combined with a pre-set layout analysis strategy method, at least one target region image is obtained, and each of the target region images is respectively... The system identifies the corresponding target region image type and region visual matrix features. Based on each target region image type, the multimodal OCR engine module performs text analysis on each target region image to obtain semantic analysis results for each text and corresponding region semantic matrix features. The multimodal attention fusion module deeply couples the region visual matrix features and the region semantic matrix features to obtain the current visual-text fusion features. The post-processing and structured output module generates and outputs complex document recognition results based on the current visual-text fusion features and a preset document object model constructed based on region coordinates and semantic labels.

[0085] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0086] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0087] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0088] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0089] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0090] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0091] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0092] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

[0093] Example 4

[0094] Embodiment 4 of the present invention also provides a computer-readable storage medium, wherein the computer-readable instructions, when executed by a computer processor, are used to perform a complex document recognition method based on layout analysis and OCR. The method includes: acquiring a document image to be recognized in real time, and performing preprocessing operations on the document image to be recognized through a preprocessing module to obtain a standard document image to be recognized; wherein the complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture; the complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module; the layout analysis module processes the standard document image to be recognized and combines it with pre-set parameters. The layout analysis strategy method obtains at least one target region image, as well as the target region image type and region visual matrix features corresponding to each target region image. Based on the target region image type, the multimodal OCR engine module performs text analysis on each target region image to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result. Through the multimodal attention fusion module, the region visual matrix features and the region semantic matrix features are deeply coupled to obtain the current visual text fusion features. Through the post-processing and structured output module, based on the current visual text fusion features and combined with a preset document object model constructed based on region coordinates and semantic tags, the complex document recognition output results are generated and fed back.

[0095] Of course, the computer-executable instructions provided in the embodiments of the present invention, which include a computer-readable storage medium, are not limited to the method operations described above, but can also perform related operations in complex document recognition based on layout analysis and OCR provided in any embodiment of the present invention.

[0096] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0097] It is worth noting that in the above embodiments of complex document recognition based on layout analysis and OCR, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0098] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for recognizing complex documents based on layout analysis and OCR, characterized in that, include: The system acquires the image of the document to be identified in real time and performs preprocessing operations on the image of the document to be identified through the preprocessing module to obtain the standard document image to be identified. The complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture. The complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module. The layout analysis module processes the standard document image to be identified and combines it with a pre-set layout analysis strategy to obtain at least one target region image, as well as the target region image type and region visual matrix features corresponding to each target region image. Based on the image type of each target region, the multimodal OCR engine module performs text analysis on each target region image to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result. The multimodal attention fusion module deeply couples the visual matrix features and semantic matrix features of each region to obtain the current visual-text fusion features. The post-processing and structured output module generates and feeds back complex document recognition output results based on the current visual text fusion features and a pre-defined document object model constructed based on region coordinates and semantic tags. Specifically, the layout analysis module performs keyword recognition processing on the standard document image to be identified, and divides the standard document image to be identified based on at least one identified keyword to obtain the semantic region of each keyword; A relational graph is constructed for each of the keyword semantic regions to obtain the target document element relational graph G; in, N represents the set of neighboring nodes, which includes text block elements, table elements, and image elements; E represents the relation edge, which consists of spatial relations and semantic relations. Based on the target document element relationship graph, the graph structure reasoning path is optimized by using a pre-set node state update rule method and combining it with the layout analysis strategy method to obtain at least one target region image. The layout analysis strategy method is based on a Markov decision strategy; the node state update rule method is... , Let n be the state of node n at time t; Represents the edge feature encoding function. Let m be the feature vector of edge nm; m is the set of neighboring nodes. Nodes in; Among them, the layout analysis strategy method is ;in, This is a function for layout analysis strategies; in the current state The probability distribution of the next choice To analyze the action, For global context vectors; The trainable weight matrix maps the input features to the action space dimension; The current state of the focused node; Based on the images of each target region, the image type of each target region and the visual matrix features of each region are obtained.

2. The method according to claim 1, characterized in that, The step of preprocessing the document image to be recognized through the preprocessing module to obtain the standard document image to be recognized includes: The image tilt correction method in the preprocessing module is used to correct the tilt of the document image to be identified, thereby obtaining the tilt-corrected document image to be identified. The formula regions corresponding to the tilt correction document image to be identified are acquired, and enhanced by a preset contrast enhancement method. Combined with a preset noise removal method, the standard document image to be identified is obtained.

3. The method according to claim 2, characterized in that, The step involves using a layout analysis module to perform keyword recognition processing on the standard document image to be identified, and then dividing the standard document image based on at least one identified keyword to obtain semantic regions for each keyword, including: By using document type recognition and adaptive strategies, the document type of the standard document image to be identified is determined, thereby identifying the target document type. Based on the target document type, a mapping and matching process is performed with a preset document type layout rule library to obtain the target keyword set; Based on the target keyword set, the layout analysis module performs keyword recognition processing on the standard document image to be identified, and divides the standard document image to be identified based on at least one identified keyword to obtain the semantic region of each keyword.

4. The method according to claim 3, characterized in that, The step involves performing text analysis on each target region image based on its image type using the multimodal OCR engine module, obtaining semantic analysis results for each text, and corresponding region semantic matrix features, including: Based on the image type of each target region, each dedicated OCR engine module is determined in the multimodal OCR engine module; The target area image types include title area image type, body text area image type, table cell area image type, formula area image type, and special symbol area image type; Each dedicated OCR engine module performs text analysis on the corresponding target region image to obtain the semantic analysis results of each text. Semantic matrix features were extracted from the semantic analysis results of each text to obtain the semantic matrix features of each region.

5. The method according to claim 4, characterized in that, The process involves deeply coupling the visual matrix features and semantic matrix features of each region through the multimodal attention fusion module to obtain the current visual-text fusion features, including: The multimodal attention fusion module deeply couples the visual matrix features and semantic matrix features of each region to calculate the cross-modal attention weights. ; in, ; This represents the attention level of the i-th semantic unit to the j-th visual region, and , Represents the semantic query projection matrix. ; The feature vector of the i-th semantic unit has the shape as follows: ; Represents the visual key projection matrix. ; This represents the feature vector of the j-th visual region, with shape [formula missing]. ; V represents the scaling factor; V represents the region visual matrix feature. H represents the feature map height, W represents the feature map width, C represents the number of channels; S represents the region semantic matrix features. N represents the length of the text sequence, and D represents the dimension of each semantic unit; Based on the attention weights of each cross-modal, the current visual text fusion features are calculated using a pre-set fusion feature calculation formula. The formula for calculating the fusion feature is as follows: ; Represents the cross-modal fusion feature of the i-th semantic unit; By each visual area pass The eigenvalues ​​obtained by projection.

6. The method according to claim 1, characterized in that, After the post-processing and structured output module generates and feeds back complex document recognition output results based on the current visual text fusion features and a preset document object model constructed based on region coordinates and semantic tags, the following steps are also included: For the chart portion of the complex document recognition output, set the corresponding chart location link address, and combine it with the text portion of the complex document recognition output to determine the target complex document recognition output; The results of the identification of the target complex document are stored in a preset document saving format.

7. A complex document recognition device based on layout analysis and OCR, characterized in that, include: The standard document image determination module is used to acquire the document image to be identified in real time, and to perform preprocessing operations on the document image to be identified through the preprocessing module to obtain the standard document image to be identified. The complex document recognition system based on layout analysis and OCR adopts an end-to-end pipeline architecture. The complex document recognition system based on layout analysis and OCR includes a preprocessing module, a layout analysis module, a multimodal OCR engine module, a multimodal attention fusion module, and a post-processing and structured output module. The target region image type and region visual matrix feature determination module is used to process the standard document image to be identified through the layout analysis module, and combine it with the pre-set layout analysis strategy method to obtain at least one target region image, and the target region image type and region visual matrix features corresponding to each target region image. The region semantic matrix feature determination module is used to perform text analysis on each target region image according to the image type of each target region through the multimodal OCR engine module, to obtain the semantic analysis results of each text and the region semantic matrix features corresponding to each text semantic analysis result. The current visual-text fusion feature determination module is used to deeply couple the visual matrix features of each region and the semantic matrix features of each region through the multimodal attention fusion module to obtain the current visual-text fusion features; The complex document recognition output generation and feedback module is used to generate and feedback complex document recognition output results by using the post-processing and structured output module, based on the current visual text fusion features and combined with a preset document object model constructed based on region coordinates and semantic labels. The target region image type and region visual matrix feature determination module is used to: perform keyword recognition processing on the standard document image to be identified through the layout analysis module, and divide the standard document image to be identified according to at least one identified keyword to obtain semantic regions for each keyword; construct a relationship graph for each keyword semantic region to obtain a target document element relationship graph G; wherein, N represents the set of neighboring nodes, which includes text block elements, table elements, and image elements; E represents a relation edge, which consists of spatial and semantic relationships; based on the target document element relationship graph, the graph structure reasoning path is optimized using a pre-set node state update rule method combined with the layout analysis strategy method to obtain at least one target region image; wherein, the layout analysis strategy method is based on a Markov decision strategy; the node state update rule method is... , Let n be the state of node n at time t; Represents the edge feature encoding function. Let m be the feature vector of edge nm; m is the set of neighboring nodes. The nodes in the text; among them, the layout analysis strategy method is... ;in, This is a function for layout analysis strategies; in the current state The probability distribution of the next choice To analyze the action, For global context vectors; The trainable weight matrix maps the input features to the action space dimension; The current state of the focused node is determined; based on the images of each target region, the image type of each target region and the visual matrix features of each region are obtained.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a complex document recognition method based on layout analysis and OCR as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute a complex document recognition method based on layout analysis and OCR as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Document layout analysis method, model training method and device and equipment

    CN113361247A

  • Document layout recognition method and device, electronic equipment and storage medium

    CN113901954A