A method, apparatus, and system for extracting key-value pair information from document images
Through the multimodal model combining image and text features, the accuracy and generalization of key-value pair information extraction in document images is solved, end-to-end training and prediction are achieved, and the accuracy and adaptability of document information extraction are improved.
Patent Information
- Application Number
- CN202111528389.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-14
AI Technical Summary
In the prior art, when processing non-European spatial data, especially when extracting key-value pair information in document images, it is difficult to effectively utilize image and text features, resulting in poor generalization ability and low accuracy.
A multimodal model is used, combining pre-trained models of images and text, graph neural networks and question-and-answer systems, and end-to-end training and prediction are carried out through neural network structures such as transformers, and key-value information extraction is used for image, text and position features of the document.
It improves the accuracy and generalization ability of key-value information extraction, reduces dependence on different layouts, and improves the learning efficiency and generalization ability of the model.
Smart Images

Figure CN114419642B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method, apparatus, and system for extracting key-value pair information from document images. Background Art
[0002] In reality, it is common to encounter key-value pairs in many documents. For example, Figure 1 In a bank check, "Issue Date (in capital letters)" and "February 19, 2007" form a key-value pair. The former is the keyword, and the latter is the true value. The keyword describes the true value, and together they constitute a useful piece of information. There may be multiple similar key-value pair information structures in a document, and usually, all the corresponding true values need to be extracted.
[0003] The traditional method is to generate a template for each document layout. First, store the positions of each keyword in the template. After finding the keyword, the value behind or below it is the corresponding true value. This method can solve the problem well for fixed templates with high accuracy, but it will go wrong if the layout is slightly different. Therefore, a set of templates needs to be maintained for each layout. When there are many layouts to be processed, it will consume a lot of time and effort to create and maintain a large number of templates, and a new set of templates needs to be created for each new layout, with very poor generalization ability. With the development of deep learning, some neural network-based models have gradually replaced the traditional template method. Such methods do not require manually creating templates for each layout. Instead, a large amount of data with different layouts is input into the model, allowing the neural network to learn the general features hidden in different layouts by itself, thus greatly improving the generalization ability. A representative method is to splice the entire text into a string and send it into the model, and then perform NER to extract the required entities. However, such methods only utilize the text information in the document and completely ignore the special correspondence between the image information of the document and the key-value pairs, which is very helpful for improving the accuracy.
[0004] In order to better utilize the text features and image features of the document, as well as the special position correspondence contained in the key-value pairs, our team innovatively proposed a multi-modal model that combines text, image, and position features. In the model, pre-trained models for images and text, graph neural networks, and question-and-answer systems are mainly used. The backgrounds of these aspects are introduced separately below.
[0005] After entering the big data era, the available data has grown exponentially. However, the vast majority of this data is unlabeled and may have little relevance to the specific tasks that need to be solved. So, how can we learn useful knowledge from this massive amount of data and apply it to specific tasks? This is where pre-trained models come in. The training of pre-trained models usually involves designing some unsupervised training tasks aimed at learning general information in the data, such as knowledge about image classification, grammar, and syntax in language. Pre-trained models initially achieved breakthrough progress on ImageNet in the field of computer vision. With the emergence of BERT and its excellent performance, pre-trained models have rapidly developed in the NLP field and achieved great results. After obtaining a pre-trained model, it can be applied to different downstream tasks by changing its output layer, such as question answering systems, text classification, object detection, named entity recognition, and so on. Compared with models trained from scratch, pre-trained models can provide good preparatory knowledge, and this knowledge is extremely helpful for downstream tasks, enabling the model to converge faster and achieve higher accuracy.
[0006] Although traditional deep learning methods have achieved great success in extracting features from data in Euclidean space, the data in many practical application scenarios is generated from non-Euclidean spaces, and the performance of traditional deep learning methods in dealing with non-Euclidean space data is still far from satisfactory. For example, in e-commerce, a graph-based learning system can use the interactions between users and products to make very accurate recommendations, but the complexity of the graph makes it extremely challenging for existing deep learning algorithms to handle. This is because graphs are irregular, each graph has a variable-sized unordered set of nodes, and each node in the graph has a different number of adjacent nodes, resulting in some important operations (such as convolution) being easy to calculate on images but not suitable for direct use on graphs. In addition, a core assumption of existing deep learning algorithms is that data samples are independent of each other. However, for graphs, this is not the case. Each data sample (node) in a graph is connected by edges to other data samples (nodes) in the graph, and this information can be used to capture the interdependencies between instances. To make full use of this information, researchers have borrowed ideas from convolutional networks, recurrent networks, and deep autoencoders to define and design graph neural networks for processing graph data. Information between nodes is propagated through the edges connecting them. Through information propagation, the information of each node is an aggregation of the information of its adjacent nodes, which reveals the relationships between adjacent nodes, thus enabling the model to make better use of the positional relationships between key-value pairs in the document and helping the model achieve better results.
[0007] As a classic task in natural language processing, the question answering system is an advanced form of information retrieval system, aiming to answer questions raised by users in natural language with accurate and concise natural language. The research on question answering systems can be traced back to the 1960s. At that time, the methods were based on templates and rules, and both the robustness and accuracy of the models were relatively poor. There are many methods and technologies for current question answering systems, which are divided into two types according to different processing methods: knowledge graph-based question answering systems and reading comprehension-based question answering systems. The first type constructs a factual question answering system based on the knowledge graph and finds answers from the knowledge graph, with relatively high accuracy. The disadvantage is that it is overly dependent on the knowledge graph and cannot give answers outside the knowledge graph. This requires sufficient resources to build a relatively large-scale knowledge graph. The second type is to perform reading comprehension on unstructured articles to obtain answers. The data form is to give an article and ask some questions around this article. The task is to directly extract the answers from the article. Relatively common models include FastQAExt, BERT, RoBERTa, etc. Summary of the Invention
[0008] The present invention relates to a method, device and system for extracting key-value pair information in a document image. This method integrates multiple tasks of pre-trained models for images and texts, graph neural networks, and question answering systems into one model. Use neural network structures such as transformers to build the model, achieve end-to-end training and prediction, and finally output all key-value pair information in the document.
[0009] According to the first aspect of the present invention, there is provided a method for extracting key-value pair information in a document image. The document image includes multiple key-value pairs composed of keywords and true values. The input information includes: the document image, the text in each text block in the document image, the position coordinates corresponding to each text block, and the name of the entity type to be extracted. The extraction method includes the following steps:
[0010] Feature encoding step: Encode the input information and output an image + content + coordinate concatenated feature vector and a final entity type name feature vector;
[0011] Image convolution step: Take each piece of text as a node, aggregate the image + content + coordinate concatenated feature vectors of adjacent nodes, and obtain the text feature vector of each piece of text;
[0012] Task inference step: Based on the text feature vector of each piece of text, classify each text block according to the entity type. At the same time, based on the final entity type name feature vector and the text feature vector of each piece of text, output key-value pairs composed of all entity types and the corresponding text blocks through a question answering system.
[0013] Further, the feature encoding step specifically includes:
[0014] Encode the document image, the text in each text block of the document image, the name of the entity type to be extracted, and the position coordinates corresponding to each text block to obtain a document image feature vector, a text block content feature vector, a preliminary entity type name feature vector, and a text block coordinate feature vector;
[0015] Concatenate the document image feature vector, the text block coordinate feature vector, and the text block content feature vector to obtain an image + content + coordinate concatenated feature vector;
[0016] Input the preliminary entity type name feature vector into the Transformer model to output the final entity type name feature vector.
[0017] Further, the encoding of the document image, the text in each text block of the document image, the name of the entity type to be extracted, and the position coordinates corresponding to each text block to obtain a document image feature vector, a text block content feature vector, a preliminary entity type name feature vector, and a text block coordinate feature vector specifically includes:
[0018] Encode the document image to obtain a document image feature vector;
[0019] For the text in each text block of the document image and the name of the entity type to be extracted, input them into the pre-trained Chinese BERT (Bidirectional Encoder Representations from Transformers) model respectively to output a text block content feature vector and a preliminary entity type name feature vector;
[0020] Encode the position coordinates corresponding to each text block to obtain a text block coordinate feature vector.
[0021] Further, the concatenation of the document image feature vector, the text block coordinate feature vector, and the text block content feature vector to obtain an image + content + coordinate concatenated feature vector specifically includes:
[0022] After concatenating the document image feature vector and the text block coordinate feature vector, input them into the ROIAlign model to output a text block image feature vector;
[0023] After concatenating the text block coordinate feature vector and the text block content feature vector, input them into the Transformer model to output a content + coordinate concatenated feature vector;
[0024] Concatenate the content + coordinate concatenated feature vector and the text block image feature vector to obtain an image + content + coordinate concatenated feature vector.
[0025] ROIAlign is a way of regional feature aggregation, which well solves the problem of regional mis - alignment caused by two - stage quantization in traditional ROI Pooling operation. Experiments show that replacing ROI Pooling with ROI Align in the detection task can improve the accuracy of the detection model. The idea of ROI Align is very simple: cancel the quantization operation, and use the method of bilinear interpolation to obtain the image values at the pixel points with floating - point coordinates, thus transforming the whole feature aggregation process into a continuous operation.
[0026] Furthermore, the dimensions of both the image + content + coordinate concatenated feature vector and the final entity type name feature vector are 512.
[0027] Furthermore, the encoding of the document image specifically includes:
[0028] For the document image, a pre - trained deep convolutional neural network is used to encode the text block and the surrounding image features to obtain the sample image feature vector.
[0029] Here, the surrounding image features are obtained through convolution.
[0030] Furthermore, the pre - trained deep convolutional neural network is ResNet - 50 pre - trained on the ImageNet massive images.
[0031] The ImageNet project is a large - scale visualization database for visual object recognition software research. More than 14 million image URLs are manually annotated by ImageNet to indicate the objects in the pictures; bounding boxes are also provided for at least one million images.
[0032] Resnet is the abbreviation of Residual Network. This series of networks is widely used in fields such as object classification and as part of the backbone classic neural network for computer vision tasks. ResNet - 50 is a typical Resnet network, which contains 50 conv2d operations.
[0033] Furthermore, the encoding of the position coordinates corresponding to each text block specifically includes:
[0034] For each text block, the coordinates of the four vertices and the length and width of the text block are put together as the text block coordinate feature vector [x1, y1, x2, y2, x3, y3, x4, y4, w, h] of the text block, where [x1, y1], [x2, y2], [x3, y3], [x4, y4] are the coordinates of the four vertices respectively, and [w, h] are the length and width values of the text block.
[0035] Here, each text block on the sample image can be regarded as a quadrilateral.
[0036] Furthermore, the image convolution step specifically includes:
[0037] Taking words as nodes, the link relationship between words represents the edges of the graph. Calculate the weights of the edges between each node and other nodes according to the Euclidean distance between the image + content + coordinate concatenated feature vectors of each node, and obtain a soft graph adjacency matrix;
[0038] According to the soft graph adjacency matrix, perform weighted aggregation on the image + content + coordinate concatenated feature vectors of adjacent nodes to obtain the aggregated neighbor node features;
[0039] Concatenate the image + content + coordinate concatenated feature vector of a certain node with the aggregated neighbor node features;
[0040] Use a multi-layer perceptron to transform the concatenated features to obtain the word feature vector of each word.
[0041] Furthermore, the number of layers of the graph convolutional neural network is 2.
[0042] Furthermore, the dimension of the word feature vector of each word is 512.
[0043] Furthermore, the task inference step specifically includes two tasks processed in parallel:
[0044] Node classification: Input the word feature vector of each word into a classifier composed of a trained linear neural network, and output N types of entity types to be extracted and the "keyword" category;
[0045] Question answering system: Concatenate the word feature vectors of each word to obtain the article feature vector composed of all text blocks; Use the final entity type name feature vector and the article feature vector composed of all text blocks as the question and the article to input into the question answering system respectively; Output the key-value pairs composed of all entity types and their corresponding text blocks.
[0046] Furthermore, the question answering system is a trained RoBERTa_wwm_ext_large model.
[0047] The RoBERTa_wwm_ext_large model is a representation model based on the BERT model.
[0048] Further, in the RoBERTa_wwm_ext_large model, the next sentence prediction task is removed, dynamic masking is used to replace static masking, and a vocabulary of 50265 characters is used in the text encoding stage, and no additional preprocessing or tokenization is performed on the input.
[0049] According to a second aspect of the present invention, there is provided an apparatus for extracting key-value pair information in a document image. The apparatus operates based on the method provided in any of the foregoing aspects. The apparatus includes:
[0050] A feature encoding module for encoding input information and outputting an image + content + coordinate concatenated feature vector and a final entity type name feature vector;
[0051] An image convolution module for aggregating the image + content + coordinate concatenated feature vectors of adjacent nodes with each character as a node to obtain a character feature vector for each character;
[0052] A task inference module for classifying each text block according to the entity type based on the character feature vector of each character, and at the same time, based on the final entity type name feature vector and the character feature vector of each character, outputting key-value pairs composed of all entity types and their corresponding text blocks through a question answering system.
[0053] According to a third aspect of the present invention, there is provided a system for extracting key-value pair information in a document image. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the method for extracting key-value pair information in a document image as described in any of the foregoing aspects.
[0054] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, characterized in that a computer program is stored thereon, and when the computer program is executed by a processor, it implements the method for extracting key-value pair information in a document image as described in any of the foregoing aspects.
[0055] Advantages of the present invention:
[0056] 1. Taking out the entity type name separately as the question of the question answering system makes the extraction of entities more targeted. At the same time, using the semantic similarity between the entity type name and the keywords in the document, the model can better learn the corresponding relationship between key-value pairs, thereby making the accuracy of entity extraction higher;
[0057] 2. Combining node classification and the question answering system into one model through multi-task learning realizes end-to-end training and prediction. The two tasks have a positive impact on each other, greatly improving the learning efficiency;
[0058] 3. The model has strong generalization ability. The model makes full and efficient use of document features, including grammar and semantics within the text, the relationship between texts within a sentence, the position information of the text on the image, etc., so that the model has little dependence on the layout. There are many cases where key-value pairs are included in the document and the layouts are not the same. A single model can handle these cases well, avoiding the need to train multiple models for different situations;
[0059] 4. Popular and excellent pre-trained models such as Resnet-50, BERT, and RoBERTa are applied, enabling the model to learn richer information, faster learning speed, and higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0061] Figure 1 Show an example of a key-value pair bank check in the prior art.
[0062] Figure 2 Show a flowchart of the extraction algorithm for key-value pair information in a document image according to an embodiment of the present invention.
[0063] Figure 3 Show a structural diagram of the extraction algorithm for key-value pair information in a document image according to an embodiment of the present invention.
[0064] The realization of the object of the present invention, functional features, and advantages will be further described in conjunction with the embodiments and with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0066] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein, for example.
[0067] In addition, the terms "comprising", "having", and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0068] Multiple, including two or more.
[0069] And / or, it should be understood that for the term "and / or" used in this disclosure, it is merely a relational description of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0070] The present invention relates to an accurate method for extracting key information from document images. Aiming at the problem of key-value pair information extraction, the applicant proposes to convert the key-value pair information extraction task into a question-and-answer system task to more specifically extract each entity in the document. The node classification and the question-and-answer system are placed in one model, making full use of the image, text, and position features of the document to achieve end-to-end training and prediction, avoiding the problem of error propagation caused by multiple models. Since there are dependencies between tasks, putting them into one model for training can make full use of the relationship between the two tasks to promote each other and accelerate the learning efficiency, ultimately greatly improving the accuracy of information extraction.
[0071] The applicant combines the node classification and the question-and-answer system in one model to form an organic whole. Here, the nodes are the text boxes recognized by OCR. For ease of understanding, the work within the model of the present invention can be regarded as two tasks: The first task is to classify each node encoded by the neural network. Here, the categories are the entity types to be extracted defined in advance plus the "keyword" entity; the second task is the question-and-answer system. It is carried out simultaneously with the first task. The input of the article features of the question-and-answer system is obtained by splicing the nodes encoded by the neural network together, and the question features are the categories of the entities after encoding. Here, the present invention performs a question-and-answer system task for each entity category to extract that entity.
[0072] Embodiment
[0073] 1. Feature Encoding Module
[0074] The input of the model is the image of the entire sample, the text within the text boxes recognized by OCR, the position coordinates of the text boxes, and the entity names to be extracted. The main task of this module is to encode these inputs to generate feature vectors that can be input into subsequent modules.
[0075] For the input image, the most important thing is to perform aspect ratio-preserving size normalization and padding with zeros at the boundaries, so that the size of the image can support operations such as convolution and downsampling required by the neural network in the encoding module, and maximize the retention of global and local feature information. Image feature encoding mainly uses a deep convolutional neural network to encode the image features of text blocks and their surrounding areas. In this step, ResNet-50 pre-trained on a large number of ImageNet images is used as the feature encoding network. This model has a powerful representation ability for images and can extract and represent the key features of images well. The goal of this step is to output the image feature encodings corresponding to each text box. Therefore, it is necessary to apply ROIAlign at the corresponding positions of the network output feature map in combination with the positions of the text boxes to obtain the corresponding image feature encodings. The dimension of this feature is 512.
[0076] The text input to this model is divided into two parts: one part is the text within the text boxes recognized by OCR; the other part is the predefined entity type names to be extracted, which are equivalent to the keywords in key-value pairs. First, the Chinese BERT model pre-trained on a large number of texts such as Chinese Wikipedia is used to encode the text. This model already has the ability to parse syntax and semantics after pre-training, which can make the encoded features more abundant and enable subsequent modules to learn higher-level features faster. Here, the two parts of the text are separately fed into the BERT model for encoding. The output of the first part of the text is in units of text boxes and is concatenated with the position features and then fed into the subsequent network structure; the output of the second part of the text is in units of each entity type. If there are N entity types, then N features are correspondingly output and fed into the subsequent network structure.
[0077] The text box recognized by OCR can be regarded as a quadrilateral, and the input position coordinates corresponding to each text box are the coordinates of the four vertices. Here, in this embodiment, the four vertices, length, and width of the quadrilateral are concatenated together as the position vector of the text box, that is, [x1, y1, x2, y2, x3, y3, x4, y4, w, h]. Then, the position vector and the text features output after BERT encoding of the text within the text box are concatenated and input into Transformer3 for encoding. In this embodiment, the output of the second part of the text is also fed into the Transformer for re-encoding. Using the powerful multi-layer multi-head self-attention mechanism in the Transformer, the syntax, semantics between words, and the influence of each part on subsequent tasks are learned, which has an important impact on the accuracy of subsequent tasks. The dimension of the feature vector output from the Transformer is 512.
[0078] So far, the image features of each text box, the fused features of the position features and text features of each text box, and the encoded features corresponding to each entity type name have been obtained respectively. In this embodiment, in the dimension of the feature space, the image features and the fused features of the text box are added as the new feature vector of the text box and input into the subsequent module. The dimension of the feature vector after addition is still 512. The features corresponding to each entity type name will be used as question features and input into the question-answering system in the subsequent module. The feature vectors obtained through the feature encoding module not only contain image features, the grammar and semantics of the text itself, and the sentence features unique to such samples, but also the relative position relationships between text boxes. This enables the model in this embodiment to learn a variety of information and better complete the subsequent tasks.
[0079] 2. Graph Convolution Module
[0080] The function of this module is to pass the feature vectors output by the feature encoding module through a multi-layer graph convolutional neural network to fully learn the relative position relationships unique to the document.
[0081] The graph defined by this module is an undirected graph, where the text serves as the nodes of the graph, and the link relationships between the texts represent the edges of the graph. The feature vectors output by the feature encoding module undergo convolutional operations in multiple layers of the graph convolutional network. Each node continuously propagates its own features to neighboring nodes and at the same time fuses the features of adjacent nodes to enhance the representation of this node and learn the inherent local and global graph structures. The graph convolution operation can be divided into three steps. First, calculate the weights of the edges between each node and other nodes according to the Euclidean distance between the features of each node to obtain a soft graph adjacency matrix. According to this adjacency matrix, perform weighted aggregation on the features of adjacent nodes to obtain the aggregated features of neighboring nodes. Second, concatenate the features of this node with the aggregated features of neighboring nodes. Third, use a multi-layer perceptron to transform the concatenated features to obtain the final features of this node. Through experiments, it is found that the effect is relatively good when the number of layers of the graph convolutional neural network is 2. The dimension of the feature vector output by the graph convolutional neural network is 512.
[0082] 3. Task Inference Module
[0083] In the task inference module, this embodiment designs two tasks. One is an auxiliary task, which does not participate in the inference and prediction process but will participate in the training process; the other is the main task, which participates in both inference and training.
[0084] The first task is node classification, which is an auxiliary task. Feature vectors corresponding to each text box are obtained from the graph convolution module, and these feature vectors are fed into a classifier composed of a linear neural network. Each node can be classified into one of N + 1 categories (N is the number of predefined entity types to be extracted, and the extra category is "keyword"). This task enables the model to better learn the feature differences between different types of entities and "keywords", thus making the entity extraction in the main task more accurate.
[0085] The second task is the question-answering system, which is the main task. The input of the question-answering system requires two parts: the question and the article. The question is the output after the entity type names in step 1 pass through the feature encoding module, and each entity type name corresponds to a question. The article is obtained by concatenating the feature vectors of each text box output by the graph convolution module. The task of the question-answering system is to find a certain text box in the article as the answer to each question. Here, in this embodiment, the RoBERTa_wwm_ext_large model trained on a large amount of Chinese data is used. RoBERTa is improved based on BERT, and there are three main improvements: one is to remove the next sentence prediction task, and experiments have shown that the model performance will improve after removing this task; the second is to replace the static mask with a dynamic mask. The advantage is that during the continuous input of a large amount of data, the model will gradually adapt to different mask strategies and learn different language representations; the third is to use a larger vocabulary in the text encoding stage, and no additional preprocessing or word segmentation is performed on the input. This model is retrained on a large number of relevant samples for the task of the Chinese question-answering system on the basis of the original RoBERTa model, so it has better performance for the question-answering system task.
[0086] The two tasks seemingly differ, but actually they both serve one goal, that is, to better find the relationships between entities and other nodes, between entities, and between key-value pairs, so as to extract entities more accurately. Therefore, they can promote each other and improve the learning efficiency.
[0087] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising that element.
[0088] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that the above implementation methods can be realized by means of software plus a necessary general hardware platform. Of course, it can also be realized by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0090] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A method for extracting key-value pair information in a document image, the document image including a plurality of key-value pairs composed of keywords and true values, and the input information including: A document image, the text within each text block in the document image, the position coordinates corresponding to each text block, and the name of the entity type to be extracted, wherein the extraction method comprises the following steps: Feature encoding step: Encode the input information and output an image + content + coordinate concatenated feature vector and a final entity type name feature vector: Encode the document image, the text within each text block in the document image, the name of the entity type to be extracted, and the position coordinates corresponding to each text block to obtain a document image feature vector, a text block content feature vector, a preliminary entity type name feature vector, and a text block coordinate feature vector; Concatenate the document image feature vector, the text block coordinate feature vector, and the text block content feature vector to obtain an image + content + coordinate concatenated feature vector; Input the preliminary entity type name feature vector into a Transformer model and output a final entity type name feature vector; Among them, the encoding of the document image, the text within each text block in the document image, the name of the entity type to be extracted, and the position coordinates corresponding to each text block to obtain a document image feature vector, a text block content feature vector, a preliminary entity type name feature vector, and a text block coordinate feature vector specifically includes: Encode the document image to obtain a document image feature vector; For the text within each text block in the document image and the name of the entity type to be extracted, input them into a pre-trained Chinese BERT model respectively, and output a text block content feature vector and a preliminary entity type name feature vector; Encode the position coordinates corresponding to each text block to obtain a text block coordinate feature vector; Among them, the concatenation of the document image feature vector, the text block coordinate feature vector, and the text block content feature vector to obtain an image + content + coordinate concatenated feature vector specifically includes: After concatenating the document image feature vector and the text block coordinate feature vector, input them into a ROIAlign model and output a text block image feature vector; After concatenating the text block coordinate feature vector and the text block content feature vector, input them into a Transformer model and output a content + coordinate concatenated feature vector; Concatenate the content + coordinate concatenated feature vector and the text block image feature vector to obtain an image + content + coordinate concatenated feature vector; Image convolution step: Take each word as a node, aggregate the image + content + coordinate concatenated feature vectors of adjacent nodes, and obtain a word feature vector for each word; Task inference step: Based on the word feature vector of each word, classify each text block according to the entity type, and at the same time, based on the final entity type name feature vector and the word feature vector of each word, output a key-value pair composed of all entity types and their corresponding text blocks through a question-answering system; Among them, the question-and-answer system is a trained RoBERTa_wwm_ext_large model; among them, in the RoBERTa_wwm_ext_large model, the next sentence prediction task is removed, the static mask is replaced with a dynamic mask, and a vocabulary of 50265 characters is used in the text encoding stage, and no additional preprocessing or tokenization is performed on the input.
2. The extraction method according to claim 1, wherein The specific steps of the image convolution include: Using words as nodes, the link relationships between words represent the edges of the graph. Calculate the weights of the edges between each node and other nodes according to the Euclidean distances between the image + content + coordinate concatenated feature vectors of each node to obtain a soft graph adjacency matrix; According to the soft graph adjacency matrix, perform weighted aggregation on the image + content + coordinate concatenated feature vectors of adjacent nodes to obtain the aggregated neighbor node features; Concatenate the image + content + coordinate concatenated feature vector of a certain node with the aggregated neighbor node features; Use a multi-layer perceptron to transform the concatenated features to obtain the word feature vector of each word.
3. The extraction method according to claim 1, characterized in that, The specific steps of the task reasoning include two tasks processed in parallel: Node classification: Input the word feature vector of each word into a classifier composed of a trained linear neural network, and output N types of entity types to be extracted and the "keyword" category; Question-and-answer system: Concatenate the word feature vectors of each word to obtain the article feature vector composed of all text blocks; Use the final entity type name feature vector and the article feature vector composed of all text blocks as the question and the article respectively and input them into the question-and-answer system; Output the key-value pairs composed of all entity types and the corresponding text blocks.
4. An extraction device for key-value pair information in a document image, characterized in that, The device operates based on the method for extracting key-value pair information in a document image according to any one of claims 1 to 3, and the device includes: A feature encoding module for encoding the input information and outputting an image + content + coordinate concatenated feature vector and a final entity type name feature vector; An image convolution module for using each word as a node to aggregate the image + content + coordinate concatenated feature vectors of adjacent nodes to obtain the word feature vector of each word; A task reasoning module for classifying each text block according to the entity type based on the word feature vector of each word, and at the same time, based on the final entity type name feature vector and the word feature vector of each word, outputting the key-value pairs composed of all entity types and the corresponding text blocks through the question-and-answer system.
5. An extraction system for key-value pair information in a document image, the system comprising: A processor and a memory for storing executable instructions; among them, the processor is configured to execute the executable instructions to perform the method for extracting key-value pair information in a document image according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor, it implements the method for extracting key-value pair information in a document image according to any one of claims 1 to 3.
Citation Information
Patent Citations
Multi-instance document key information extraction method and system
CN113536798A