Multimodal document recognition method, device, equipment and storage medium

Through multimodal fusion feature extraction and entity recognition technology, the problem of poor document recognition accuracy in complex text structure scenarios is solved, and high-precision and high generalization document recognition effect is achieved.

CN115131801BActive Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210386897.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-05-13
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

Existing document recognition technology is difficult to capture long-distance dependencies and correlations in complex text structure scenarios, resulting in poor recognition accuracy.

Method used

A multimodal-based document recognition method is adopted to obtain document images, image segmentation, word segmentation feature extraction, image feature extraction and feature fusion, and multimodal fusion features are generated for entity recognition.

Benefits of technology

It significantly improves the accuracy and accuracy of document recognition, enhances generalization, and realizes the accuracy of fine-grained attribute recognition and position marking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131801B_ABST
    Figure CN115131801B_ABST
Patent Text Reader

Abstract

The present application provides a document recognition method, device, equipment and storage medium based on multimodality, which relates to the field of artificial intelligence and can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving. The method includes: performing image segmentation processing on the document image corresponding to the document to be recognized to obtain text image blocks, non-text image blocks and block position information; performing feature extraction on the text image blocks and non-text image blocks respectively to obtain word segmentation features, word segmentation position information, second image features and word segmentation position features of text word segmentation, and first image features and block position features of non-text image blocks; based on the word segmentation position information and block position information, performing feature fusion processing on the word segmentation features, first image features, second image features, word segmentation position features and block position features, performing entity recognition on the obtained multimodal fusion features, and obtaining document recognition results. The present application significantly improves the recognition accuracy and has strong generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a document recognition method and device based on multimodality. Background Art

[0002] Document recognition is an important basic task in the field of document processing, and can provide valuable input data for high-level document processing tasks such as summary generation and knowledge question answering. At present, in document recognition tasks, prior knowledge (such as expert knowledge or rule base, etc.) is usually used to build a knowledge base, and then document recognition is achieved by matching the similarity between the input text and the objects in the knowledge base, or the transition probability between different text words is identified based on traditional machine learning methods to obtain recognition results. However, the former method is limited by the amount and breadth of knowledge in the knowledge base and has poor generalization, while the latter method is only applicable to simple text recognition scenarios. In complex text structure scenarios, due to the difficulty in capturing long-distance dependencies and related relationships, the document recognition accuracy is poor.

[0003] Therefore, it is necessary to provide a reliable document recognition solution to solve the above existing problems. Summary of the invention

[0004] The present application provides a document recognition method, apparatus, device and storage medium based on multimodality, which significantly improves the precision and accuracy of document recognition and has strong generalization.

[0005] On the one hand, the present application provides a document recognition method based on multimodality, the method comprising:

[0006] Acquire a document image corresponding to a document to be identified, wherein the document to be identified includes at least one document element;

[0007] Performing image segmentation processing on the document image corresponding to the document to be identified to obtain text image blocks, non-text image blocks and segmentation position information corresponding to the document to be identified;

[0008] Extracting word segmentation features from the text image block and the non-text image block respectively to obtain word segmentation features and word segmentation position information of the text word segmentation corresponding to the document to be identified;

[0009] Extracting image features of the non-text image block and the text segmentation from the document image to obtain a first image feature of the non-text image block and a second image feature of the text segmentation;

[0010] Performing feature mapping processing on the word segmentation position information of the text word segmentation and the block position information respectively to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image block;

[0011] Based on the word segmentation position information and the block position information, feature fusion processing is performed on the word segmentation feature, the first image feature, the second image feature, the word segmentation position feature and the block position feature to obtain a multimodal fusion feature of the document to be identified;

[0012] Entity recognition is performed on the multimodal fusion features to obtain a document recognition result of the document to be recognized, wherein the document recognition result includes the text segmentation and entity category of the non-text image block corresponding to the document to be recognized.

[0013] On the other hand, a multimodal document recognition device is provided, the device comprising:

[0014] Document data acquisition module: used to acquire a document image corresponding to a document to be identified, wherein the document to be identified includes at least one document element;

[0015] Image segmentation module: used to perform image segmentation processing on the document image corresponding to the document to be identified, and obtain text image blocks, non-text image blocks and segment position information corresponding to the document to be identified;

[0016] A word segmentation feature extraction module: used to extract word segmentation features from the text image block and the non-text image block respectively, to obtain word segmentation features and word segmentation position information of the text word segmentation corresponding to the document to be identified;

[0017] An image feature extraction module is used to extract the image features of the non-text image block and the text segmentation from the document image to obtain a first image feature of the non-text image block and a second image feature of the text segmentation;

[0018] Position feature mapping module: used to perform feature mapping processing on the word segmentation position information of the text word segmentation and the block position information respectively, to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image block;

[0019] A feature fusion module: configured to perform feature fusion processing on the word segmentation feature, the first image feature, the second image feature, the word segmentation position feature and the block position feature based on the word segmentation position information and the block position information, so as to obtain a multimodal fusion feature of the document to be identified;

[0020] Entity recognition module: used to perform entity recognition on the multimodal fusion features to obtain a document recognition result of the document to be recognized, wherein the document recognition result includes the entity category of the text segmentation and non-text image block corresponding to the document to be recognized.

[0021] On the other hand, a multimodal document recognition device is provided, the device comprising a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the multimodal document recognition method as described above.

[0022] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimodal-based document recognition method as described above.

[0023] On the other hand, a computer-readable storage medium is provided, in which at least one instruction or at least one program is stored, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the multimodal-based document recognition method as described above.

[0024] On the other hand, a terminal is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimodal-based document recognition method as described above.

[0025] On the other hand, a server is provided, the server includes a processor and a memory, the terminal includes a processor and a memory, the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimodal-based document recognition method as described above.

[0026] On the other hand, a computer program product or a computer program is provided, characterized in that the computer program product or the computer program comprises computer instructions, and when the computer instructions are executed by a processor, the multimodal-based document recognition method as described above is implemented.

[0027] The multimodal document recognition method, apparatus, device, terminal, server, storage medium and computer program provided in this application have the following technical effects:

[0028] The present application first obtains a document image corresponding to a document to be identified, wherein the document to be identified includes at least one document element; performs image segmentation processing on the document image corresponding to the document to be identified, and obtains text image blocks, non-text image blocks, and segmentation position information corresponding to the document to be identified; then, word segmentation feature extraction is performed on the text image blocks and non-text image blocks, respectively, to obtain word segmentation features and word segmentation position information of text word segments corresponding to the document to be identified; image feature extraction is performed on the non-text image blocks and text word segments of the document image, to obtain first image features of the non-text image blocks and second image features of the text word segments; feature mapping processing is performed on the word segmentation position information and the segmentation position information of the text word segments, respectively. The word segmentation position features of text word segmentation and the block position features of non-text image blocks are obtained, and then the fine-grained features of multiple modalities of the document to be identified are obtained; further, based on the word segmentation position information and the block position information, feature fusion processing is performed on the word segmentation features, the first image features, the second image features, the word segmentation position features and the block position features to obtain the multimodal fusion features of the document to be identified, and entity recognition is performed on the multimodal fusion features containing multi-level and multimodal document information to obtain the document recognition result of the document to be identified, thereby achieving accurate fine-grained attribute recognition of document elements, significantly improving the accuracy of element attribute recognition and position marking, and being able to provide high-value input for high-order document recognition tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0030] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0031] Figure 2 It is a flowchart of a multimodal document recognition method provided in an embodiment of the present application;

[0032] Figure 3 It is a flowchart of another multimodal document recognition method provided in an embodiment of the present application;

[0033] Figure 4 It is a flowchart of another multimodal document recognition method provided in an embodiment of the present application;

[0034] Figure 5 It is a flowchart of another multimodal document recognition method provided in an embodiment of the present application;

[0035] Figure 6 is a visualization diagram of a document recognition result provided by an embodiment;

[0036] Figure 7 is a schematic diagram of document recognition results before and after correction provided by an embodiment;

[0037] Figure 8 is a structural framework diagram of a document recognition system provided by an embodiment;

[0038] Fig. 9 is a principle flow chart of a document recognition method provided by an embodiment;

[0039] Fig.10 It is a schematic diagram of a framework of a multimodal document recognition device provided in an embodiment of the present application;

[0040] Fig.11 It is a hardware structure block diagram of an electronic device of a multimodal document recognition method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0043] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0044] OCR: Optical Character Recognition, the process of extracting text from an image using image algorithms.

[0045] NER: Named Entity Recognition, named entity recognition, identifies entities with specific meanings in text. In this application scenario, entity types can correspond to document element categories, including ordinary text / title / chapter / caption, etc.

[0046] Bert: Bidirectional Encoder Representations from Transformers, a bidirectional encoder representation technology based on transformers, a pre-training technology for natural language processing.

[0047] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0048] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0049] Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further perform graphic processing so that computer processing becomes an image that is more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0050] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0051] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely used in many fields. The solution provided in the embodiment of the present application involves artificial intelligence machine learning / deep learning and natural language processing and other technologies, which are specifically described by the following embodiments:

[0052] See also Figure 1 , Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application, such as Figure 1 As shown, the application environment may include at least a server 01 and a terminal 02. The server 01 and the terminal 02 may be directly or indirectly connected via wired or wireless communication, which is not limited in the present application.

[0053] In the embodiment of the present application, server 01 may include an independently operated server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In addition, multiple servers can also be composed of a blockchain, and the server is a node on the blockchain. Specifically, server 01 may include a network communication unit, a processor, a memory, and the like.

[0054] Specifically, cloud technology refers to a hosting technology that unifies hardware, software, network and other resources in a wide area network or local area network to achieve data computing, storage, processing and sharing. It distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. The network that provides resources is called "cloud". Among them, artificial intelligence cloud service is generally also called AIaaS (AI as a Service, Chinese for "AI as a Service"). This is a mainstream service mode of artificial intelligence platform. Specifically, AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI theme mall: all developers can access one or more artificial intelligence services provided by the platform through API interfaces. Some senior developers can also use the AI ​​framework and AI infrastructure provided by the platform to deploy and operate their own cloud artificial intelligence services.

[0055] Specifically, the server 01 may be a node in the distributed system 100, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes through network communication. The nodes may form a peer-to-peer (P2P) network, and any form of computer device, such as the server 01, the client 02 and other electronic devices may become a node in the blockchain system by joining the peer-to-peer network, wherein the blockchain includes a series of blocks that are connected to each other in the order of their generation. Once a new block is added to the blockchain, it will not be removed, and the block records the record data submitted by the nodes in the blockchain system.

[0056] In the embodiment of the present application, the terminal 02 may include physical devices such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, and vehicle-mounted terminals, but is not limited thereto, and may also include software running in physical devices, such as applications, etc. The operating system running on the terminal 02 in the embodiment of the present application may include but is not limited to Android, IOS, Linux, Windows, etc.

[0057] In the embodiment of the present application, the server 01 can be used to provide a document recognition service to generate a corresponding document recognition result, and can also provide a pre-training service for an initial recognition network to obtain a pre-trained recognition network, and provide a constraint training service for entity recognition of the pre-trained recognition network to obtain a target entity recognition network. The terminal 02 can be used to send a document recognition instruction and a document to be recognized to the server 01, so that the server 01 performs the corresponding document recognition.

[0058] In addition, it should be noted that Figure 1 What is shown is merely an application environment of a multimodal document recognition method and device. In actual applications, the application environment may include more or fewer nodes, and this application does not impose any limitation thereto.

[0059] The following combination Figure 2 This application introduces a multimodal document recognition method. Figure 2 It is a flowchart of a multimodal document recognition method provided in an embodiment of the present application. The present application provides method operation steps such as the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many orders, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, a parallel processor or a multi-threaded processing environment). The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. Specific examples include Figure 2 As shown, the method may include:

[0060] S201: Obtain a document image corresponding to a document to be recognized.

[0061] In the embodiment of the present application, the document to be identified may be a plain text document or a document including multimodal document elements, such as an official document, a graphic document, or a bill document. The document to be identified includes at least one document element, and the document element may include but is not limited to a picture element and a text element. In some cases, the document element may also include a caption, a header, and a footer. The document image is the image data of the document to be identified, which may be obtained by shooting, scanning, or format conversion.

[0062] S203: Perform image segmentation processing on the document image corresponding to the document to be recognized, and obtain text image blocks, non-text image blocks and segment position information corresponding to the document to be recognized.

[0063] In an embodiment of the present application, the image segmentation of the document image can be achieved by a semantic segmentation method to obtain at least one text image block corresponding to the text object in the document to be identified, at least one non-text image block corresponding to the non-text object, and the block position information. Among them, the non-text objects may include but are not limited to drawings, tables, captions, headers and footers, etc., and correspondingly, the non-text image blocks may include but are not limited to drawing image blocks, table image blocks, caption image blocks, header image blocks and footer image blocks, etc., and the block position information includes the position information of each text image block and the position information of each non-text image block. The layout analysis of the document to be identified is achieved through image segmentation processing, and the image blocks of various document elements such as text, figures, tables, captions, etc. in the document to be identified are extracted, which is conducive to providing fine-grained information for subsequent feature extraction and document recognition, thereby improving the accuracy and generalization of document recognition.

[0064] In practical applications, the semantic segmentation method can be implemented based on a semantic segmentation model. The semantic segmentation model may include but is not limited to U-Net, FCN, ResNet, and SegNet, such as ResNet-101.

[0065] S205: extracting word segmentation features from the text image block and the non-text image block respectively, and obtaining word segmentation features and word segmentation position information of the text word segmentation corresponding to the document to be recognized.

[0066] In the embodiment of the present application, text recognition is performed on the text image block to obtain the first text data of each text image block in the document to be recognized. At the same time, based on the position information of the text image block, the position information of the first text data can be determined in the text recognition process, and the first text data is segmented and feature extracted to obtain the segmentation features of each text segmentation in the first text data, and the segmentation position information of each text segmentation is determined according to the position information of the first text data. For non-text image blocks, the non-text image blocks can be first mapped to the corresponding second text data, and the second text data of each non-text image block can be segmented and feature extracted to obtain the segmentation features of each text segmentation in the second text data, and the position information of the non-text image block is used as the segmentation position information of the corresponding text segmentation. It can be understood that the above-mentioned text segmentation can be word granularity, character granularity or with a preset character length as the segmentation granularity. The segmentation granularity can be set based on actual needs and is not limited here.

[0067] In some embodiments, the image fingerprint of the non-text image block can be calculated by a preset image fingerprint algorithm to obtain the image fingerprint ID of each non-text image block, the image fingerprint ID can be segmented to obtain the corresponding text segmentation, and the text segmentation can be embedded with features to obtain the segmentation features of each non-text image block. It can be understood that the segmentation of the image fingerprint ID can obtain one segmentation or multiple segmentations. In this way, the segmentation features obtained contain the content information of the non-text image block, which is beneficial to the subsequent training of the target entity recognition network, so that it can learn the feature extraction of the non-text object content and the relationship between it and other text elements, further refine the recognition granularity and improve the accuracy.

[0068] In other embodiments, element block identification texts of non-text image blocks of various document element categories may be pre-stored, for example, the element block identification text of caption non-text image blocks is 30, and the element block identification text of figure non-text image blocks is 50. In some embodiments, the semantic features of non-text image blocks of different document element categories are different, and the document element category of the non-text image block is identified based on the semantic features, and then the element block identification text corresponding to the identified document element category is determined according to a preset corresponding relationship, wherein the preset corresponding relationship represents the association relationship between multiple document element categories and multiple element block identification texts. Accordingly, the element block identification text represents the document element category of the non-text image block, please refer to Figure 3 , S205 may include the following steps S2051-S2055.

[0069] S2051: Obtain the element block identification text corresponding to the non-text image block.

[0070] S2052: Perform character recognition on the text image block to obtain the text line corresponding to the text image block and the position information of the text line.

[0071] S2053: Perform word segmentation processing on the text line and the element block identification text respectively to obtain the text word segmentation corresponding to the document to be recognized.

[0072] S2054: Determine the word segmentation position information of the text segmentation based on the text line position information and the block position information.

[0073] S2055: Perform feature embedding processing on the text word segmentation to obtain the word segmentation features of the text word segmentation.

[0074] Specifically, the text data in the text image block can be identified by a character recognition method, such as using optical character recognition (OCR) to obtain the text line in the text image block, and obtain the fine-grained information of the text block. In the character recognition process, the position information of the text line can be determined according to the number of lines of the text line in the text image block to which it belongs and the position information of the text image block to which it belongs. For example, the text image block can be equidistantly segmented based on the total number of lines of the text image block, and the position information of each text line image block obtained by segmentation is determined according to the overall position information of the text image block, that is, the position information of the text line is obtained. In one embodiment, the OCR method is a sequence recognition method based on CRNN (Convolutional Recurrent Neural Network).

[0075] Further, after obtaining each text line of each text image block, each text line is segmented based on a preset segmentation method to obtain the text segmentation of each text line, and each element block identification text is segmented to obtain the text segmentation of each non-text image block, and all text segmentations corresponding to the document to be identified are feature embedded to obtain the segmentation features of each text segmentation. Among them, the preset segmentation method can adopt an existing segmentation algorithm of natural language processing based on actual needs, such as the tokenizer provided by the Bert method, which is not limited in this application. In addition, the position information of the non-text image block is used as the position information of the corresponding text segmentation, and the position information of the text segmentation corresponding to the text line is determined based on the position information of the text line. In some cases, the position information of the text line is directly determined as the position information of the corresponding text segmentation, and in other cases, the coordinate range of the text line can be evenly divided according to the number of text segmentations corresponding to the text line to obtain the position information of each text segmentation. In one embodiment, the word segmentation feature is obtained by using a tokenizer to perform word segmentation processing on a text line or element block identification text and natural language conversion text vector embedding processing. Each text segmentation corresponds to at least one token, each token corresponds to a word segmentation feature, and the feature dimension of the word segmentation feature is 1x512.

[0076] It should be noted that the above-mentioned various location information may include area coordinate information, which can characterize the coordinate range of an area. For example, the location information of a rectangular area may include the coordinates, width and height of the vertices in the rectangular area, or the coordinates, width and height of the center point of the rectangular area.

[0077] Exemplarily, taking the text "if" as an example, after word segmentation, the character-level word segmentation results in the words "ru" and "guo". The position information of "ru" includes the coordinates (x1, y1) of its upper left vertex and the coordinates (x2, y2) of its lower right vertex. The position information of "guo" includes the coordinates (x1 + 10, y1 + 10) of its upper left vertex and the coordinates (x2 + 10, y2 + 10) of its lower right vertex.

[0078] S207: Extract image features of non-text image blocks and text word segmentation from the document image to obtain the first image features of non-text image blocks and the second image features of text word segmentation.

[0079] In the embodiments of the present application, the above image feature extraction can be implemented based on a preset image feature extraction network. Specifically, before using the preset image feature extraction network for feature extraction, preprocess the input document image or regional image, scale its resolution to the preset size (H*W) to obtain a normalized image, and then input the normalized image into the preset image feature extraction network to obtain the corresponding feature map. Furthermore, perform feature extraction on the feature maps corresponding to each non-text image block and each text word segmentation respectively to obtain the first image features and the second image features. Among them, the preset image feature extraction network can be constructed based on a convolutional neural network (CNN), such as ResNet-101. Correspondingly, each input image is processed into a feature map with a resolution of H / 32 * W / 32 and 16 channels. In one embodiment, for the feature map corresponding to each text word segmentation or non-text image block, the RoiAlign algorithm is used for regional feature extraction to obtain the first image features or the second image features. The feature dimensions of the first image features and the second image features are 1x512.

[0080] In practical applications, please refer to Figure 4 , S207 may include the following steps S2071 - S2072.

[0081] S2071: Respectively obtain the feature maps of the image regions corresponding to non-text image blocks and text word segmentation in the document image.

[0082] S2072: Respectively perform feature extraction on the feature map corresponding to the non-text image block and the feature map corresponding to the text word segmentation to obtain the first image features and the second image features.

[0083] In some embodiments, S2071 may specifically include: performing convolutional processing on the document image to obtain the document feature map corresponding to the document image; based on the word segmentation position information and the block position information, determine the feature map corresponding to the text word segmentation and the feature map corresponding to the non-text image block from the document feature map.

[0084] Specifically, after preprocessing the document image, convolution processing is performed on the normalized document image, and then based on the word segmentation position information of the text word segmentation, the first area corresponding to each text word segmentation in the document feature map is determined, and the first area of ​​the document feature map is determined as the feature map corresponding to the text word segmentation; based on the block position information of the non-text image block, the second area corresponding to each non-text image block in the document feature map is determined, and the second area in the document feature map is determined as the feature map corresponding to the non-text image block, in this way, the feature map of each text word segmentation and the feature map of each non-text image block are obtained.

[0085] In some other embodiments, S2071 may specifically include: respectively obtaining the image areas corresponding to the text segmentation and non-text image blocks in the document image; performing convolution processing on the image areas corresponding to the text segmentation and non-text image blocks to obtain feature maps corresponding to the text segmentation and non-text image blocks.

[0086] Specifically, according to the word segmentation position information and the block position information, the image area corresponding to each text segmentation and each non-text block in the document image is determined, and the normalized image corresponding to each image area is convolved to obtain a feature map of each text segmentation and a feature map of each non-text image block.

[0087] Furthermore, feature extraction is performed on the feature graph of each text segmentation word and the feature graph of each non-text image block to obtain a first image feature of each text segmentation word and a second image feature of each non-text image block.

[0088] S209: Perform feature mapping processing on the word segmentation position information and the block position information of the text word segmentation respectively to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image block.

[0089] In an embodiment of the present application, the word segmentation position information of each text word segment is vector-embedded, and the block position information of each non-text image block is vector-embedded to obtain the word segmentation position features of each word segmentation position information and the block position features of each non-text image block. In one embodiment, the feature dimensions of the word segmentation position features and the block position features are 1x512.

[0090] S211: Based on the word segmentation position information and the block position information, feature fusion processing is performed on the word segmentation feature, the first image feature, the second image feature, the word segmentation position feature and the block position feature to obtain a multimodal fusion feature of the document to be identified.

[0091] In the embodiment of the present application, the multimodal fusion feature integrates the feature information of text, image and location, and can provide fine-grained attribute features of the document, which is conducive to realizing fine-grained entity recognition of the document and improving the accuracy of document entity recognition. In addition, when there are fuzzy areas in the document, it is easy to cause character recognition errors. By superimposing image features, character recognition errors can be corrected, and the robustness of the recognition system can be improved.

[0092] In actual application, please refer to the figure, S211 may include the following steps S2111-S2112.

[0093] S2111: Based on the word segmentation position information and the block position information, feature stitching processing of the word segmentation features, feature stitching processing of the first image features and the second image features, and feature stitching processing of the word segmentation position features and the block position features are performed respectively to obtain text stitching features, image stitching features and position stitching features of the document to be identified.

[0094] S2112: Fusing text splicing features, image splicing features, and position splicing features of the document to be recognized to obtain multimodal fusion features.

[0095] Specifically, based on the word segmentation position information of each text word segmentation and the block position information of each non-text image block, the position order of all text word segments and non-text image blocks corresponding to the document to be identified is determined; then, the word segmentation features of the text word segmentation are spliced ​​using the position order as the feature splicing order to obtain text splicing features, each first image feature and each second image feature are spliced ​​to obtain image splicing features, and each word segmentation position feature and each block position feature are spliced ​​to obtain position splicing features; then, a multimodal fusion feature is obtained through feature fusion processing, and the feature fusion processing here can be feature addition, such as simple addition. In some cases, the text splicing feature and the image splicing feature are added in the first direction to obtain a first fusion feature, and then the position splicing feature and the first fusion feature are added in the second direction to obtain a multimodal fusion feature.

[0096] S213: Perform entity recognition on the multimodal fusion features to obtain a document recognition result of the document to be recognized.

[0097] In an embodiment of the present application, the document recognition result includes the entity category of the text segmentation and non-text image block corresponding to the document to be recognized. The entity category may include but is not limited to ordinary text, title, chapter, caption, header, footer and formula, etc. By performing entity recognition on the multimodal fusion features, the entity category of each text segmentation and the entity category of each non-text image block in the document to be recognized are obtained. Different entity categories can use different category labels, such as different colors or shapes, etc., and then the category labels can be visualized on the document image to display the fine-grained category recognition results of the document to be recognized. Please refer to Figure 6 , Figure 6 A schematic diagram of a document recognition result provided by an embodiment is shown in FIG. Figure 6 The type tags M1-M6 represent the entity categories of header / footer, normal text, end of line, beginning of line, figure and caption respectively.

[0098] In summary, by performing entity recognition on multimodal fusion features containing multi-level and multimodal document information, the document recognition result of the document to be identified is obtained, accurate fine-grained attribute recognition of document elements is achieved, and the accuracy of element attribute recognition and position marking is significantly improved, which can provide high-value input for high-order document recognition tasks.

[0099] In practical applications, the target entity recognition network can be called to perform entity recognition on the multimodal fusion features to obtain document recognition results. The target entity recognition network is obtained by constraining the entity recognition of the pre-trained recognition network based on the sample fusion features and entity category labels corresponding to the first sample document image, and the pre-trained recognition network is obtained by jointly training the initial recognition network for feature mask prediction and document classification recognition based on the sample fusion features and document category labels corresponding to the second sample document image.

[0100] Specifically, a training sample set for joint training of the initial recognition network and a training sample set for entity recognition training of the pre-trained recognition network are obtained, and the sample document images in the two training sample sets may be the same, different or partially overlapped.

[0101] Specifically, the first sample document image and the second sample document image can be a plain text document, or a document including multimodal document elements, including but not limited to official documents, graphic documents or bill documents, etc., and the document can include multiple document elements. The method for obtaining the sample fusion features corresponding to the first sample document image and the second sample document image is similar to the method for obtaining the multimodal fusion features described above, and will not be repeated here. The entity category label represents the entity category of each text segmentation and non-text image block corresponding to the first sample document image, and the document category label represents the document category of the second sample document. The document category may include but is not limited to bills, official documents, papers, and web pages, etc.

[0102] In one embodiment, the target entity recognition network may be a NER network, in which a transformer architecture is used as a basic network to perform feature extraction of multimodal fusion features. Exemplarily, the transformer may include a 12-layer encoder and a 12-layer decoder.

[0103] In a specific embodiment, the following method may be used to obtain a pre-trained recognition network.

[0104] S301: Acquire a training data set and an initial recognition network, where the training data set includes a second sample document image and a corresponding document category label.

[0105] S303: Extract features from the second sample document image to obtain sample fusion features corresponding to the second sample document image.

[0106] S305: Perform feature masking processing on the sample fusion features to obtain target sample features.

[0107] Specifically, the sample fusion feature is obtained by fusion processing based on text features, image features and position features. The masking process refers to masking at least one dimension of the image / position / text feature information with a certain probability, and inferring the masked information with the unmasked information, thereby realizing the prediction task of multimodal feature masking. Exemplarily, the text features are partially masked, such as replacing the "feature" in the text line "perform feature masking processing on the sample fusion feature" in the second sample document with "Mask", and obtaining the masked text line "perform feature masking processing on the sample fusion MaskMask", and then extracting features of the masked second sample document image to obtain the target sample features, or directly performing masking processing on the word segmentation features corresponding to the two words "feature" in the sample fusion feature to obtain the target sample features.

[0108] S307: Using the target sample features as inputs of the initial recognition network, and using the mask features and document category labels as expected outputs, the initial recognition network is jointly trained for feature mask prediction and document classification recognition to obtain a pre-trained recognition network.

[0109] In some embodiments, the base networks of the initial recognition network and the pre-trained recognition network are both transformers, the training tasks performed by the initial recognition network and the pre-trained recognition network are different, and the loss calculation methods used are also different. In a specific embodiment, the initial recognition network is jointly trained based on the feature mask prediction task and the document classification recognition task, wherein the document classification recognition task is to identify and classify the category of the document based on the document features, which can be specifically a multi-label document classification task.

[0110] Specifically, the basic network is used to extract features of the sample fusion features, and the mask features are predicted based on the extracted features, and the document category recognition is performed based on the extracted features to obtain the mask feature prediction results and the document category recognition results respectively; the loss function of the feature mask prediction task is used to calculate the loss of the mask feature prediction results and the mask features to obtain the first loss, and the loss function of the document classification and recognition task is used to calculate the loss of the document category recognition results and the document category labels to obtain the second loss; the first loss and the second loss are added to obtain the total model loss. If the total model loss or the current number of iterations meets the model convergence condition, the current initial recognition network is used as the pre-trained recognition network. Otherwise, the network parameters of the initial recognition network are adjusted based on the total model loss to obtain an updated initial recognition network; the sample fusion features of the second sample document image are input into the updated initial recognition network to perform feature extraction, mask feature prediction, document category recognition and loss calculation to perform iterative training of the updated initial recognition network until the total model loss or the number of iterations meets the model convergence condition to obtain the pre-trained recognition network. The model convergence condition may be that the total loss is less than or equal to a preset loss, or the number of iterations reaches a preset number. In one embodiment, the loss functions used in the feature cover prediction task and the document classification and recognition task are both cross entropy functions.

[0111] In this way, based on a large number of different types of sample documents, fusion features including multimodal information are generated, and then the initial recognition network is trained jointly with multiple tasks. This can make full use of the multi-level and multimodal information of the document, so that the recognition network can learn the global features of the input document while fully learning the features and relationships of the document elements, significantly improving the network learning effect, reducing the training cost of the subsequent pre-trained recognition network, and improving the model effect of the final target entity recognition network.

[0112] In practical applications, the sample fusion features corresponding to the first sample document image are used as input to iteratively train the pre-trained recognition network obtained above for entity recognition, and obtain the target entity recognition network. Specifically, the document element objects in the first sample document can be labeled with entity categories, and the document element objects can be marked as ordinary text, titles, chapters or captions, etc. Through the above pre-training process, the pre-trained recognition network has learned the relationship between document elements, which can reduce the amount of sample data required in the entity recognition training process. Only by tuning the model parameters of the pre-trained recognition model based on a small amount of labeled training data, the entity recognition training can be completed, and the target entity recognition network can be obtained to perform complex document fine-grained entity recognition tasks and obtain the fine-grained attribute categories of the document. In addition, the above pre-training process can improve the generalization and transferability of the document recognition method.

[0113] In practical applications, the above-mentioned pre-training and entity recognition training are implemented using a hardware environment equipped with a GPU chip. The GPU chip supports GPU parallel computing, which can improve training efficiency.

[0114] Based on some or all of the above implementations, in an embodiment of the present application, the method may further include the following steps of correcting the document recognition results, specifically including the following S401-S407.

[0115] S401: According to the document recognition result, a target text line is determined from text lines corresponding to the document to be recognized, and the target text line contains text segmentations of at least two entity categories.

[0116] S403: Count the number of word segments for at least two entity categories to obtain the number of text word segments for each entity category in the at least two entity categories.

[0117] S405: The entity category with the largest number of text segmentations is used as the target entity category of the target text line.

[0118] S407: Based on the target entity category, update the entity category of each text segmentation in the target text line.

[0119] In practical applications, there may be recognition errors in the document recognition results. Based on the above correction steps, the results can be corrected to improve the recognition accuracy and system robustness. Specifically, the document recognition results include the entity category of each text segmentation and non-text image block, with the text line and non-text image block as anchor points, the text line including two or more entity categories as the target text line, and the entity category of each text segmentation in each target text line is voted, and the entity category with the most votes, that is, the entity category with the most text segmentations, is determined as the target entity category of the text line, that is, the text segmentations of other entity categories in the corresponding target text line are all updated to the target entity category. In some embodiments, the non-text image block including two or more entity categories can also be determined as the target image block, and the actual entity category of the target image block is determined based on the above voting processing method, and the entity category is updated. In this way, local errors in document recognition can be avoided and the accuracy and robustness of the recognition results can be improved.

[0120] Specifically, the aforementioned voting operation can also be performed on each text line and each non-text image block to obtain respective entity categories, and then each text line and each non-text image block are marked with overall entity categories. Figure 7 , Figure 7 The figure is a schematic diagram of document recognition results before and after correction provided by an embodiment. The left figure is the document recognition result before correction. Figure 7The entity category result of the text segmentation word marked by the arrow is shown in the attached figure, while the entity category results of other text segmentations in the text line are ordinary text. Obviously, there is a recognition error. After voting on the text line, it is determined that the entity category of the text segmentation word marked by the arrow is ordinary text, and it is updated to obtain the entity recognition result in the right figure.

[0121] In the present application, please refer to Figure 8 and Fig. 9 , Figure 8 A structural framework diagram of a document recognition system is shown. Fig. 9 A principle flow chart of a document recognition method provided by an embodiment is shown. The document recognition system includes an object extraction module, a character recognition module, a feature extraction module, a feature fusion module and a target entity recognition network; wherein the object extraction module is used to perform image segmentation processing on a document image to be recognized or a sample document image to realize document layout analysis and obtain text image blocks, non-text image blocks and block position information; the character recognition module is used to perform character recognition on text image blocks to obtain corresponding text lines and position information of text lines; the feature extraction module may include a word segmentation submodule, a word segmentation feature extraction network, an image feature extraction network and a position feature embedding submodule, which are respectively used to perform word segmentation processing and word segmentation feature extraction on text lines and non-text image blocks output by the character recognition module, perform image feature extraction on image areas of text word segments and non-text image blocks, perform position feature embedding processing on position information, and splice the obtained multiple modal features to obtain text splicing features, image splicing features and position splicing features; the feature fusion module is used to perform addition processing on text splicing features, image splicing features and position splicing features to obtain multi-modal fusion features, and input them into the target entity recognition network to perform document recognition processing and obtain document recognition results. In this way, end-to-end result output is achieved through system design, simplifying operating costs.

[0122] Existing document recognition solutions mainly include the following three methods: 1) Based on expert knowledge / rule base, based on the common element categories of general documents, a knowledge base is built, and entity recognition is completed by matching the similarity between the input text and the objects in the knowledge base. For example, if "2.1XXX" is input, "2.1" is a common chapter number, so "2.1XXX" is recognized as "chapter"; 2) Based on traditional machine learning methods, such as hidden Markov models, superimposed conditional random fields, the transition probability between different text words is established through hidden Markov models, and further constraint learning and result optimization are performed through conditional random fields; 3) Based on deep learning methods, such as long short-term memory networks or attention networks, through the powerful modeling capabilities of neural networks, stronger and closer word / word or sentence / sentence relationships are established to complete NER. However, method 1) is limited by the amount of knowledge in the library and has poor generalization. Although method 2) can achieve good results in simple scenarios, it is difficult to achieve satisfactory results in complex text structure scenarios because it is difficult to capture long-distance dependencies and related relationships. Method 3) can achieve good results in specified document scenarios, but it is difficult to achieve good generalization in general document scenarios due to the amount of training data. The above technical solution based on this application does not need to build a knowledge base, reduces the amount of annotated training data required, improves the generalization of the method, and combines multimodal features to achieve fine-grained attribute recognition of documents and improve recognition accuracy.

[0123] The present application also provides a multi-modal document recognition device 700, such as Fig.10 As shown, the device comprises:

[0124] Document data acquisition module 10: used to acquire a document image corresponding to a document to be identified, wherein the document to be identified includes at least one document element;

[0125] Image segmentation module 20: used to perform image segmentation processing on the document image corresponding to the document to be identified, and obtain text image blocks, non-text image blocks and segment position information corresponding to the document to be identified;

[0126] The word segmentation feature extraction module 30 is used to extract word segmentation features from the text image block and the non-text image block respectively, and obtain the word segmentation features and word segmentation position information of the text word corresponding to the document to be identified;

[0127] Image feature extraction module 40: used to extract image features of non-text image blocks and text segmentation from the document image, and obtain first image features of the non-text image blocks and second image features of the text segmentation;

[0128] Position feature mapping module 50: used to perform feature mapping processing on the word segmentation position information and the block position information of the text word segmentation, respectively, to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image block;

[0129] Feature fusion module 60: used to perform feature fusion processing on word segmentation features, first image features, second image features, word segmentation position features and block position features based on word segmentation position information and block position information to obtain multimodal fusion features of the document to be identified;

[0130] Entity recognition module 70: used to perform entity recognition on the multimodal fusion features to obtain a document recognition result of the document to be recognized, wherein the document recognition result includes the text segmentation and entity category of the non-text image block corresponding to the document to be recognized.

[0131] In some embodiments, the entity recognition module 70 may be specifically used to: call a target entity recognition network to perform entity recognition on the multimodal fusion features to obtain a document recognition result;

[0132] Among them, the target entity recognition network is obtained by constraining the entity recognition of the pre-trained recognition network based on the sample fusion features and entity category labels corresponding to the first sample document image, and the pre-trained recognition network is obtained by jointly training the initial recognition network for feature masking prediction and document classification recognition based on the sample fusion features and document category labels corresponding to the second sample document image.

[0133] In some embodiments, the apparatus may further include:

[0134] Training data acquisition module: used to acquire a training data set and an initial recognition network, the training data set includes a second sample document image and a corresponding document category label;

[0135] Sample feature extraction module: used to extract features from the second sample document image to obtain sample fusion features corresponding to the second sample document image;

[0136] Feature masking module: used to perform feature masking on sample fusion features to obtain target sample features;

[0137] Pre-training module: It is used to use the target sample features as the input of the initial recognition network, and the cover features and document category labels as the expected outputs, respectively, to jointly train the initial recognition network for feature cover prediction and document classification recognition, and obtain a pre-trained recognition network.

[0138] In some embodiments, the word segmentation feature extraction module 30 may include:

[0139] The identification text acquisition submodule is used to acquire the element block identification text corresponding to the non-text image block, where the element block identification text represents the document element category of the non-text image block;

[0140] Character recognition submodule: used to perform character recognition on the text image block to obtain the text line and position information of the text line corresponding to the text image block;

[0141] Word segmentation processing submodule: used to perform word segmentation processing on the text line and element block identification text respectively to obtain the text segmentation corresponding to the document to be identified;

[0142] The word segmentation position determination submodule is used to determine the word segmentation position information of the text segmentation based on the position information of the text line and the block position information;

[0143] Word segmentation feature embedding submodule: used to perform feature embedding processing on text word segmentation to obtain word segmentation features of text word segmentation.

[0144] In some embodiments, the image feature extraction module 40 may include:

[0145] Feature map acquisition submodule: used to respectively acquire feature maps of image areas corresponding to non-text image blocks and text segmentation words in document images;

[0146] Image feature extraction submodule: used to extract features from the feature map corresponding to the non-text image block and the feature map corresponding to the text word segmentation, respectively, to obtain the first image feature and the second image feature.

[0147] In some embodiments, the feature map acquisition submodule may include:

[0148] A first convolution processing unit: used for performing convolution processing on the document image to obtain a document feature map corresponding to the document image;

[0149] Feature map determination unit: used to determine the feature map corresponding to the text word segmentation and the feature map corresponding to the non-text image block from the document feature map based on the word segmentation position information and the block position information.

[0150] In some other embodiments, the feature map acquisition submodule may include:

[0151] Image region acquisition unit: used to respectively acquire image regions corresponding to text segmentation words and non-text image blocks in the document image;

[0152] The second convolution processing unit is used to perform convolution processing on the image area corresponding to the text segmentation and the non-text image block to obtain a feature map corresponding to the text segmentation and the non-text image block.

[0153] In some embodiments, the feature fusion module 60 may include:

[0154] Feature splicing submodule: used for performing feature splicing processing of word segmentation features, feature splicing processing of first image features and second image features, and feature splicing processing of word segmentation position features and block position features based on word segmentation position information and block position information, to obtain text splicing features, image splicing features and position splicing features of the document to be identified;

[0155] Feature fusion submodule: used to fuse the text splicing features, image splicing features and position splicing features of the document to be identified to obtain multimodal fusion features.

[0156] In some embodiments, the apparatus may further include:

[0157] Target text line determination module: used to determine the target text line from the text line corresponding to the document to be recognized according to the document recognition result, and the target text line contains text segmentation of at least two entity categories;

[0158] Segment count module: used to count the segment counts of at least two entity categories, and obtain the number of text segment counts of each entity category in at least two entity categories;

[0159] Target entity category determination module: used to take the entity category with the largest number of text segmentation words as the target entity category of the target text line;

[0160] Entity category update module: used to update the entity category of each text segmentation in the target text line based on the target entity category.

[0161] The device and method embodiments in the above device embodiments are based on the same application concept.

[0162] An embodiment of the present application provides a multimodal document recognition device, which includes a processor and a memory, in which at least one instruction or at least one program is stored, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimodal document recognition method provided in the above method embodiment.

[0163] The memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0164] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal, a server or similar electronic devices. Fig.11 1 is a hardware structure block diagram of an electronic device for executing a multi-modal document recognition method provided by an embodiment of the present application. Fig.11 As shown, the electronic device 800 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 810 (the processor 810 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 830 for storing data, and one or more storage media 820 (such as one or more mass storage devices) for storing application programs 823 or data 822. Among them, the memory 830 and the storage medium 820 can be short-term storage or permanent storage. The program stored in the storage medium 820 may include one or more modules, each of which may include a series of instruction operations in the electronic device. Furthermore, the central processing unit 810 can be configured to communicate with the storage medium 820 and execute a series of instruction operations in the storage medium 820 on the electronic device 800. The electronic device 800 may also include one or more power supplies 860, one or more wired or wireless network interfaces 850, one or more input and output interfaces 840, and / or, one or more operating systems 821, such as Windows Server TM , Mac OS X TM , Unix TM ,LinuxTM, FreeBSDTM, etc.

[0165] The input / output interface 840 may be used to receive or send data via a network. The specific example of the network may include a wireless network provided by a communication provider of the electronic device 800. In one example, the input / output interface 840 includes a network adapter (Network Interface Controller, NIC), which may be connected to other network devices via a base station so as to communicate with the Internet. In one example, the input / output interface 840 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0166] It can be understood by those skilled in the art that Fig.11 The structure shown is for illustration only and does not limit the structure of the above electronic device. Fig.11 More or fewer components as shown, or with Fig.11 Different configurations are shown.

[0167] An embodiment of the present application also provides a storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to implementing a noise addition processing method for an image in a method embodiment, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the noise addition processing method for the image provided by the above method embodiment.

[0168] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network electronic devices in a computer network, such as at least one of a plurality of network servers. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0169] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.

[0170] It can be seen from the embodiments of the multimodal document recognition method, device, equipment, terminal, server, storage medium or computer program provided by the present application that the present application first obtains a document image corresponding to a document to be recognized, wherein the document to be recognized includes at least one document element; performs image segmentation processing on the document image corresponding to the document to be recognized, and obtains text image blocks, non-text image blocks and segmentation position information corresponding to the document to be recognized; then, word segmentation feature extraction is performed on the text image blocks and non-text image blocks respectively, and the word segmentation feature and word segmentation position information of the text word corresponding to the document to be recognized are obtained; image feature extraction is performed on the non-text image blocks and the text word segments on the document image, and the first image feature of the non-text image block and the second image feature of the text word segment are obtained; and the document image is segmented and processed respectively. The word segmentation position information and block position information of the word segmentation are subjected to feature mapping processing to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image blocks, thereby obtaining the fine-grained features of multiple modalities of the document to be identified; further, based on the word segmentation position information and the block position information, feature fusion processing is performed on the word segmentation features, the first image features, the second image features, the word segmentation position features and the block position features to obtain the multimodal fusion features of the document to be identified, entity recognition is performed on the multimodal fusion features containing multi-level and multimodal document information to obtain the document recognition result of the document to be identified, accurate fine-grained attribute recognition of document elements is achieved, the accuracy of element attribute recognition and position marking is significantly improved, and high-value input can be provided for high-order document recognition tasks.

[0171] It should be noted that the above-mentioned sequence of the embodiments of the present application is for description only and does not represent the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of the present application are described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0172] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0173] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0174] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A document recognition method based on multimodality, characterized in that: The method comprises: Acquire a document image corresponding to a document to be identified, wherein the document to be identified includes at least one document element; Performing image segmentation processing on the document image corresponding to the document to be identified to obtain text image blocks, non-text image blocks and segmentation position information corresponding to the document to be identified; Extracting word segmentation features from the text image block and the non-text image block respectively to obtain word segmentation features and word segmentation position information of the text word segmentation corresponding to the document to be identified; Extracting image features of the non-text image block and the text segmentation from the document image to obtain a first image feature of the non-text image block and a second image feature of the text segmentation; Performing feature mapping processing on the word segmentation position information of the text word segmentation and the block position information respectively to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image block; Based on the word segmentation position information and the block position information, feature fusion processing is performed on the word segmentation feature, the first image feature, the second image feature, the word segmentation position feature and the block position feature to obtain a multimodal fusion feature of the document to be identified; Entity recognition is performed on the multimodal fusion features to obtain a document recognition result of the document to be recognized, wherein the document recognition result includes the text segmentation and entity category of the non-text image block corresponding to the document to be recognized.

2. The method according to claim 1, characterized in that The performing entity recognition on the multimodal fusion feature to obtain the document recognition result of the document to be recognized includes: Calling a target entity recognition network to perform entity recognition on the multimodal fusion feature to obtain the document recognition result; Among them, the target entity recognition network is obtained by constraining the entity recognition of the pre-trained recognition network based on the sample fusion features and entity category labels corresponding to the first sample document image, and the pre-trained recognition network is obtained by jointly training the initial recognition network for feature masking prediction and document classification recognition based on the sample fusion features and document category labels corresponding to the second sample document image.

3. The method according to claim 1, characterized in that The method further comprises: Acquire a training data set and an initial recognition network, wherein the training data set includes a second sample document image and a corresponding document category label; Performing feature extraction on the second sample document image to obtain a sample fusion feature corresponding to the second sample document image; Performing feature masking processing on the sample fusion features to obtain target sample features; The target sample features are used as inputs of the initial recognition network, and the cover features and document category labels are used as expected outputs respectively. The initial recognition network is jointly trained for feature cover prediction and document classification recognition to obtain a pre-trained recognition network.

4. The method according to any one of claims 1 to 3, characterized in that The extracting word segmentation features of the text image block and the non-text image block respectively to obtain the word segmentation features and word segmentation position information of the text word corresponding to the document to be identified includes: Acquire an element block identification text corresponding to the non-text image block, wherein the element block identification text represents a document element category of the non-text image block; Performing character recognition on the text image block to obtain a text line corresponding to the text image block and position information of the text line; Performing word segmentation processing on the text line and the element block identification text respectively to obtain text word segments corresponding to the document to be identified; Determining word segmentation position information of the text segmentation based on the position information of the text line and the block position information; Perform feature embedding processing on the text word segmentation to obtain the word segmentation features of the text word segmentation.

5. The method according to any one of claims 1 to 3, characterized in that: The extracting the image features of the non-text image block and the text segmentation from the document image to obtain the first image feature of the non-text image block and the second image feature of the text segmentation comprises: Respectively obtaining feature maps of the image area corresponding to the non-text image block and the text segmentation in the document image; Feature extraction is performed on the feature map corresponding to the non-text image block and the feature map corresponding to the text word segmentation respectively to obtain the first image feature and the second image feature.

6. The method according to claim 5, characterized in that The step of respectively acquiring the feature graphs of the image regions corresponding to the non-text image blocks and the text segmentation words in the document image comprises: Performing convolution processing on the document image to obtain a document feature map corresponding to the document image; Based on the word segmentation position information and the block position information, a feature map corresponding to the text word segmentation and a feature map corresponding to the non-text image block are determined from the document feature map.

7. The method according to claim 5, characterized in that The step of respectively acquiring the feature graphs of the image regions corresponding to the non-text image blocks and the text segmentation words in the document image comprises: Respectively obtaining image regions corresponding to the text segmentation and the non-text image block in the document image; Convolution processing is performed on the image area corresponding to the text segmentation and the non-text image block to obtain a feature map corresponding to the text segmentation and a feature map corresponding to the non-text image block.

8. The method according to any one of claims 1 to 3, characterized in that The step of performing feature fusion processing on the word segmentation feature, the first image feature, the second image feature, the word segmentation position feature and the block position feature based on the word segmentation position information and the block position information to obtain the multimodal fusion feature of the document to be identified includes: Based on the word segmentation position information and the block position information, feature splicing processing of the word segmentation feature, feature splicing processing of the first image feature and the second image feature, and feature splicing processing of the word segmentation position feature and the block position feature are respectively performed to obtain text splicing features, image splicing features, and position splicing features of the document to be identified; The text splicing features, image splicing features and position splicing features of the document to be identified are fused to obtain the multimodal fusion features.

9. The method according to claim 4, characterized in that After performing entity recognition on the multimodal fusion feature to obtain a document recognition result, the method further includes: According to the document recognition result, a target text line is determined from the text lines corresponding to the document to be recognized, wherein the target text line contains text segmentations of at least two entity categories; Performing word segmentation statistics on the at least two entity categories to obtain the number of text word segments of each entity category in the at least two entity categories; Taking the entity category with the largest number of text segmentations as the target entity category of the target text line; Based on the target entity category, the entity category of each text word in the target text line is updated.

10. A document recognition device based on multimodality, characterized in that: The device comprises: Document data acquisition module: used to acquire a document image corresponding to a document to be identified, wherein the document to be identified includes at least one document element; Image segmentation module: used to perform image segmentation processing on the document image corresponding to the document to be identified, and obtain text image blocks, non-text image blocks and segment position information corresponding to the document to be identified; A word segmentation feature extraction module: used to extract word segmentation features from the text image block and the non-text image block respectively, to obtain word segmentation features and word segmentation position information of the text word segmentation corresponding to the document to be identified; An image feature extraction module is used to extract the image features of the non-text image block and the text segmentation from the document image to obtain a first image feature of the non-text image block and a second image feature of the text segmentation; Position feature mapping module: used to perform feature mapping processing on the word segmentation position information of the text word segmentation and the block position information respectively, to obtain the word segmentation position features of the text word segmentation and the block position features of the non-text image block; A feature fusion module: configured to perform feature fusion processing on the word segmentation feature, the first image feature, the second image feature, the word segmentation position feature and the block position feature based on the word segmentation position information and the block position information, so as to obtain a multimodal fusion feature of the document to be identified; Entity recognition module: used to perform entity recognition on the multimodal fusion features to obtain a document recognition result of the document to be recognized, wherein the document recognition result includes the entity category of the text segmentation and non-text image block corresponding to the document to be recognized.

11. The device according to claim 10, characterized in that The entity recognition module is specifically used for: Calling a target entity recognition network to perform entity recognition on the multimodal fusion feature to obtain the document recognition result; Among them, the target entity recognition network is obtained by constraining the entity recognition of the pre-trained recognition network based on the sample fusion features and entity category labels corresponding to the first sample document image, and the pre-trained recognition network is obtained by jointly training the initial recognition network for feature masking prediction and document classification recognition based on the sample fusion features and document category labels corresponding to the second sample document image.

12. The device according to claim 10, characterized in that The device also includes: A training data acquisition module: used to acquire a training data set and an initial recognition network, wherein the training data set includes a second sample document image and a corresponding document category label; Sample feature extraction module: used to extract features from the second sample document image to obtain sample fusion features corresponding to the second sample document image; Feature masking module: used to perform feature masking processing on the sample fusion features to obtain target sample features; Pre-training module: used to use the target sample features as the input of the initial recognition network, and the cover features and document category labels as the expected outputs, to jointly train the initial recognition network for feature cover prediction and document classification recognition, and obtain a pre-trained recognition network.

13. The device according to any one of claims 10 to 12, characterized in that The word segmentation feature extraction module includes: The identification text acquisition submodule is used to acquire the element block identification text corresponding to the non-text image block, wherein the element block identification text represents the document element category of the non-text image block; Character recognition submodule: used for performing character recognition on the text image block to obtain the text line corresponding to the text image block and the position information of the text line; The word segmentation processing submodule is used to perform word segmentation processing on the text line and the element block identification text respectively to obtain the text segmentation corresponding to the document to be identified; A word segmentation position determination submodule: used to determine the word segmentation position information of the text segmentation based on the position information of the text line and the block position information; The word segmentation feature embedding submodule is used to perform feature embedding processing on the text word segmentation to obtain the word segmentation features of the text word segmentation.

14. The device according to any one of claims 10 to 12, characterized in that The image feature extraction module comprises: A feature map acquisition submodule: used to respectively acquire feature maps of the image area corresponding to the non-text image block and the text segmentation word in the document image; Image feature extraction submodule: used to extract features from the feature map corresponding to the non-text image block and the feature map corresponding to the text segmentation, respectively, to obtain the first image feature and the second image feature.

15. The device according to claim 14, characterized in that The feature map acquisition submodule includes: A first convolution processing unit: configured to perform convolution processing on the document image to obtain a document feature map corresponding to the document image; A feature map determining unit is used to determine the feature map corresponding to the text word segmentation and the feature map corresponding to the non-text image block from the document feature map based on the word segmentation position information and the block position information.

16. The device according to claim 14, characterized in that The feature map acquisition submodule includes: An image region acquisition unit: used to respectively acquire the image regions corresponding to the text segmentation words and the non-text image blocks in the document image; The second convolution processing unit is used to perform convolution processing on the image area corresponding to the text segmentation and the non-text image block to obtain a feature map corresponding to the text segmentation and a feature map corresponding to the non-text image block.

17. The device according to any one of claims 10 to 12, characterized in that The feature fusion module includes: A feature splicing submodule: used for performing feature splicing processing of the word segmentation feature, feature splicing processing of the first image feature and the second image feature, and feature splicing processing of the word segmentation position feature and the block position feature based on the word segmentation position information and the block position information, respectively, to obtain text splicing features, image splicing features and position splicing features of the document to be identified; Feature fusion submodule: used for fusing the text splicing features, image splicing features and position splicing features of the document to be identified to obtain the multimodal fusion features.

18. The device according to claim 13, characterized in that The device also includes: A target text line determination module is used to determine a target text line from the text line corresponding to the document to be identified according to the document recognition result, wherein the target text line contains text segmentation of at least two entity categories; A word segmentation counting module: used to count the word segmentations of the at least two entity categories to obtain the number of text segmentations of each entity category in the at least two entity categories; A target entity category determination module: used for taking the entity category with the largest number of text segmentations as the target entity category of the target text line; Entity category updating module: used to update the entity category of each text word in the target text line based on the target entity category.

19. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimodal document recognition method as described in any one of claims 1-9.

20. A computer device, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimodal document recognition method as described in any one of claims 1-9.

21. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the multimodal-based document recognition method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Document layout analysis method and device, model training method and device, and equipment

    CN113378580A

  • Document layout recognition method and device, electronic equipment and storage medium

    CN113901954A