Document element recognition and extraction methods, systems, devices, and storage media
By combining region selection networks and Transformer encoders, the problem of insufficient generalization ability of OCR and NLP in electronic document recognition is solved, achieving accurate identification and extraction of document elements and improving the efficiency and accuracy of document processing.
Patent Information
- Application Number
- CN202310953577.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-07-31
AI Technical Summary
In existing technologies, the combination of OCR and NLP has insufficient generalization ability when recognizing and extracting elements of electronic documents, and cannot make full use of the structural and visual information contained in the documents, especially in complex document templates and multi-template scenarios where recognition errors occur frequently.
A region selection network is used to recognize electronic documents. A convolutional neural network is used to extract target candidate regions. OCR technology is combined to recognize text content and text box positions. A Transformer encoder is used for feature mapping and fusion to construct a global feature map and generate an online editing page to improve recognition accuracy and generalization ability.
It enables accurate identification and extraction of document elements in electronic documents, improves the generalization ability of identification, reduces identification errors, and improves the efficiency and accuracy of document processing.
Smart Images

Figure CN116959017B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document recognition technology, and in particular to a method, system, device, and storage medium for document element recognition and extraction. Background Technology
[0002] Online processing of electronic documents enables enterprises and institutions to collaboratively process the same electronic document tasks in different regions, thereby reducing the consumption of paper materials, conserving natural resources, and conforming to the global trend of carbon sequestration.
[0003] In related technologies, a combination of OCR (Optical Character Recognition) and NLP (Natural Language Processing) is used to intelligently identify and extract document elements. However, electronic document templates are complex in format, containing not only rich text content but also structural information and visual information based on document layout, such as bolding, underlining, highlighting, and red text. Document element extraction methods implemented using OCR technology have poor generalization ability and cannot fully extract and utilize the rich information contained within the document. Summary of the Invention
[0004] This application provides a document element identification and extraction method, system, device, and storage medium, aiming to improve the generalization ability of document element identification and extraction.
[0005] This application provides a document element identification and extraction method, which includes:
[0006] The electronic document is identified using a region selection network, resulting in multiple target candidate regions;
[0007] Determine the text content and text box position corresponding to each target candidate region, and obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region;
[0008] Spatial distance transformation is performed on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions.
[0009] By fusing the second fusion feature and the second visual feature corresponding to each of the target candidate regions, document elements are obtained.
[0010] Optionally, the step of obtaining the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region includes:
[0011] Map the text content corresponding to each of the target candidate regions to the first real number field, and map the text box position corresponding to each of the target candidate regions to the second real number field;
[0012] The text content mapped to the first real number field and the text box position mapped to the second real number field are concatenated to obtain the first fusion feature corresponding to each of the target candidate regions.
[0013] Optionally, after the step of fusing the second fusion feature and the second visual feature corresponding to each of the target candidate regions to obtain document elements, the method further includes:
[0014] Obtain the global feature map associated with the document elements. The global feature map includes multiple feature items, wherein each feature item includes recommended content and rule requirements.
[0015] Determine the correlation distance between each document element and each feature term, and obtain the target feature term corresponding to the document element based on the correlation distance;
[0016] Based on the document elements and the recommended content and rule requirements corresponding to the target feature items, construct a tree structure diagram;
[0017] An online editing page is generated based on the tree structure diagram. The online editing page includes a template recommendation area and an online editing area.
[0018] Optionally, the step of determining the correlation distance between the document element and each feature term, and obtaining the target feature term corresponding to the document element based on the correlation distance includes:
[0019] The correlation coefficient is determined based on the covariance between the document elements and each feature item, the variance of the document elements, and the variance of each feature item. The larger the absolute value of the correlation coefficient, the higher the correlation between the document elements and the feature items.
[0020] Based on the correlation coefficient, the correlation distance between the document element and each feature item is determined, wherein the larger the correlation coefficient, the shorter the correlation distance;
[0021] The feature term corresponding to the shortest relevant distance is determined as the target feature term.
[0022] Optionally, the step of identifying electronic documents based on a region selection network to obtain multiple target candidate regions includes:
[0023] The electronic document is input into a convolutional neural network to extract target features and generate a shared feature map.
[0024] Based on the region selection network, the shared feature map is slid with anchor boxes of different sizes to obtain candidate regions of different sizes and proportions;
[0025] The target candidate region is determined based on the intersection-union ratio of each candidate region with the anchor frame.
[0026] Optionally, the step of determining the target candidate region based on the intersection-union ratio of each candidate region and the anchor frame includes:
[0027] When the cross-union ratio is greater than a preset value, the candidate region with the cross-union ratio greater than the preset value is determined as the target candidate region.
[0028] Optionally, before the step of inputting the electronic document into a convolutional neural network to extract target features and generate a shared feature map, the method further includes:
[0029] The first document dataset is generated by training a recurrent adversarial network on real and pseudo document datasets.
[0030] The document datasets corresponding to different text types are used as input to the deep fakery adversarial network in the X domain to generate a second document dataset.
[0031] The first and second document datasets are used as the X domain, and then input into the deep fakery adversarial network again to generate the third document dataset;
[0032] The convolutional neural network was trained using the aforementioned third document dataset.
[0033] Furthermore, to achieve the above objectives, the present invention also provides a document element recognition and extraction system, comprising:
[0034] The recognition module is used to recognize electronic documents based on a region selection network to obtain multiple target candidate regions;
[0035] The first fusion module is used to determine the text content and text box position corresponding to each target candidate region, and to obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region.
[0036] The encoding module is used to perform spatial distance transformation on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions.
[0037] The second fusion module is used to fuse the second fusion features and second visual features corresponding to each of the target candidate regions to obtain document elements.
[0038] In addition, to achieve the above objectives, the present invention also provides a document element recognition and extraction device comprising: a memory, a processor, and a document element recognition and extraction program stored in the memory and executable on the processor, wherein the document element recognition and extraction program, when executed by the processor, implements the steps of the document element recognition and extraction method described above.
[0039] In addition, to achieve the above objectives, the present invention also provides a storage medium storing a document element identification and extraction program thereon, which, when executed by a processor, implements the steps of the document element identification and extraction method described above.
[0040] This application provides a technical solution for document element identification and extraction, including a method, system, device, and storage medium. It utilizes a region selection network to identify target candidate regions in electronic documents, obtaining target candidate regions with actual semantics. The text box positions and text content of each target candidate region are determined. The text content and text box positions corresponding to each target candidate region are then fused to obtain a first fused feature. Finally, the first fused feature and the first visual feature corresponding to each target candidate region are subjected to spatial distance transformation to obtain a second fused feature and a second visual feature after the transformation. Since the second fused feature is a local feature and the second visual feature is a global feature, fusing the second fused feature and the second visual feature yields document elements that integrate the document's semantic, structural, and visual information. Document elements extracted using this method are more accurate and have stronger generalization ability. Attached Figure Description
[0041] Figure 1 This is a flowchart illustrating the first embodiment of the document element identification and extraction method of the present invention;
[0042] Figure 2 This is a flowchart illustrating the second embodiment of the document element identification and extraction method of the present invention;
[0043] Figure 3 This is a flowchart illustrating the third embodiment of the document element identification and extraction method of the present invention;
[0044] Figure 4 This is a tree structure diagram of the present invention;
[0045] Figure 5 This is a functional module diagram of the document element recognition and extraction system of the present invention;
[0046] Figure 6 This is a schematic diagram of the document element recognition and extraction device of the present invention.
[0047] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings are only one embodiment and not the entirety of the invention. Detailed Implementation
[0048] Online processing of electronic documents greatly improves the ability of enterprises and institutions to collaboratively process the same electronic document tasks in different regions, saves natural resources, reduces the consumption of paper materials, and also reduces carbon footprint, thus conforming to the global trend of carbon neutrality.
[0049] Currently, research on the multiple elements of electronic documents covers the following aspects:
[0050] First, regarding sample data augmentation, VAE and GAN are used to randomly generate pseudo-samples for training. These electronic documents include printed text, handwritten text, English characters, Arabic numerals, tables, underlines, spaces, and special symbols. Most of the document content is white text on a black background, but some have significant font size variations, and others have watermarks, non-white backgrounds, or highly blurred backgrounds. All these formats require a large amount of sample data. Since production samples do not possess this volume of data, and many production samples are strictly for external use, relying solely on online collection or manual generation of training samples will result in inconsistent training sample quality and other problems.
[0051] To address the aforementioned issues, this application proposes a method for generating indiscriminate data through GAN adversarial dynamic transformation using real document style data and pseudo-style data in the context of electronic document sample data augmentation. This method can generate richer feature images and also generate high-quality samples that are relevant to business needs.
[0052] Specifically, electronic documents typically have a white background, and some may include a text-based watermark. The First Document Dataset uses a recurrent adversarial network algorithm to dynamically generate background images with various styles, including white background, watermark, ripples, tables, underlines, brackets, colons, checkboxes, radio buttons, and horizontal bars. A Chinese Unicode character dataset was constructed using a semi-supervised approach. English dataset and Arabic dataset ,in, , , and The X domain, as input, is used to dynamically generate an electronic document style dataset in the Y domain via a deep forgery adversarial network. The Y domain dataset serves as the second document dataset. Finally, put and As a new X domain, new electronic document style datasets are then dynamically generated using deepfake adversarial networks. That is, the third document dataset, this This refers to the final high-quality sample set with multiple features used in training. This application uses a recurrent adversarial network to generate the first document dataset. This is mainly because recurrent adversarial networks have their limitations. While they guarantee semantic transformation, they cannot guarantee certain details. In contrast, deep forgery adversarial networks are better suited to handling local details, and text style transfer is better suited to adjusting local features, thus making the generated image sets more practically valuable.
[0053] Secondly, regarding multi-feature document recognition, we use a combination of OCR and NLP to intelligently identify and extract document elements. Existing electronic document templates primarily employ OCR recognition methods. For elements that are difficult to recognize, targeted annotation and training are required. For elements prone to recognition errors, we also need to combine NLP contextual semantic understanding and MacBERT error correction. However, NLP and MacBERT also have limitations. For highly complex documents, such as those with many tables, underlines, or handwritten text, the effectiveness of NLP is minimal.
[0054] To address the aforementioned issues, this application proposes a method for intelligent identification and extraction of multiple elements in electronic documents. This method utilizes the inductive approach of knowledge graphs to construct a global feature graph of the electronic document, and then fuses local and global features using the spatial vector distance method of encoding and decoding.
[0055] Specifically, this application plans to first use a region selection network to perform preliminary segmentation of the electronic document. Then, it will analyze the resulting target candidate regions. First, use OCR technology to identify the text content within the area. and text box position ,in It is a quadruple that records the positions of the four vertices of the text box in the document. Constructing a word embedding model. Will Map to a high-dimensional real field for easier subsequent processing Where N is the number of characters in the text content recognized by OCR, and d is the dimension of the embedded word vector. A positional embedding model is constructed to map the text box position to a high-dimensional real number domain. The mapped and By concatenating the features, we obtain new features that include the text content of the target candidate region and the position of the text box. This refers to the first fusion feature. Furthermore, the features The input is fed into the encoder for encoding, yielding the mapped features, i.e., the second fused features. The target candidate region is then... Corresponding first visual features The second visual features are obtained by encoding them into the same feature space using a visual encoder. Finally, by combining the second fusion feature and the second visual feature, a high-dimensional feature that integrates the semantic, structural, and visual information of the document is obtained, thus yielding the document elements. This enables the extraction of document elements.
[0056] Third, regarding online document editing, a single-page document layout should be used. Most existing online electronic document editing methods adopt a waterfall layout. When encountering complex or multi-page electronic documents, it is necessary to manually locate the blank area to edit and drag the scroll bar. Some filling areas are exactly in the middle of the scroll bar between the previous and next pages. This method is inconvenient to use, inefficient, and prone to errors and omissions of some required fields.
[0057] To address the aforementioned issues, this application proposes a two-region segmentation verification method for online document editing data. This method involves associating intelligently identified document elements with a pre-constructed global feature map to identify the recommended content and rule requirements corresponding to each document element. A two-region segmentation verification template capable of real-time intelligent recommendation and processing is then built on the interface. Specifically, based on the content extracted from intelligent multi-element document identification and the association with the global feature map, this application automatically identifies the recommended content and rule requirements corresponding to each document element and constructs a tree-structured recommendation template. The tree-structured template is referenced from... Figure 4 The system comprises three levels: Level 1 represents the document elements of the electronic document; Level 2 represents recommended content; and Level 3 represents rule requirements, such as text size, minimum and maximum text length, required fields, and data type format. An online editing page is generated based on the tree structure diagram. The online editing page consists of a two-area layout in a left-right format, including a template recommendation area and an online editing area. The template recommendation area uses a template validation mode, and the online editing area uses an online editing mode.
[0058] To better understand the above technical solutions, exemplary embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While exemplary embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art.
[0059] First Embodiment
[0060] like Figure 1As shown, in the first embodiment of this application, the document element identification and extraction method of this application includes the following steps:
[0061] Step S110: The electronic document is identified based on the region selection network to obtain multiple target candidate regions.
[0062] In this embodiment, the region proposal network is an RPN (Region Proposal Network), which is used to segment target candidate regions with actual semantics from electronic documents.
[0063] Optionally, a convolutional neural network can be used to extract target features from the electronic document to obtain a shared feature map. This shared feature map is then input into an RPN network to extract a candidate set for document segmentation. Each element in this candidate set represents a target candidate region that may contain complete document elements. The convolutional neural network can be a Faster R-CNN network.
[0064] Step S120: Determine the text content and text box position corresponding to each target candidate region, and obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region.
[0065] In this embodiment, OCR technology can be used in each target candidate region to extract the text content and text box positions in each target candidate region. The text box position can be the position of a single character or the position of a sentence.
[0066] In this embodiment, a pre-trained convolutional neural network is used to perform feature mapping on the document element regions, mapping the two-dimensional image information to a high-dimensional real space. Then, an RPN is used to extract target candidate regions. Finally, the bounding boxes of the retained target candidate regions are fine-tuned to obtain the exact text box positions. In this context, 4 represents the four vertices of the text box, and 2 represents the horizontal coordinate of each vertex.
[0067] In this embodiment, when recognizing text content, the problem of recognizing text content can be transformed into a classification problem based on text content images, that is, finding a mapping. ,in The image represents a two-dimensional image, where H represents the height of the two-dimensional image and W represents the width of the two-dimensional image. This represents the corresponding character tag, where N is the total number of characters in the dictionary. However, in actual processing, since N is relatively large, the character image features are first mapped to... In the space, c is the word vector corresponding to the character y, and then the model is constructed. Make The advantage of first converting character images into word vectors is that it transforms the multi-classification task into a feature mapping from one space to another, where c has a richer representational capability than y. Therefore, for input text images, using the model... The word vector sequence corresponding to the image can be obtained. , which is the text content, where L is the text length.
[0068] The above steps complete the recognition of the text content and text box position corresponding to each target candidate region.
[0069] In this embodiment, after obtaining the text content and text box position corresponding to each target candidate region, a word embedding model is constructed to map the text content to a high-dimensional real number domain, i.e., the first real number domain. A position embedding model is constructed to map the text box position to a high-dimensional real number domain, i.e., the second real number domain.
[0070] Optionally, the step of obtaining the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region includes:
[0071] Step S121: Map the text content corresponding to each of the target candidate regions to the first real number field, and map the text box position corresponding to each of the target candidate regions to the second real number field.
[0072] In this embodiment, a word embedding model is constructed. Will Map to a high-dimensional real field for easier subsequent processing Where N is the number of characters in the text content recognized by OCR, and d is the dimension of the embedded word vector. A positional embedding model is constructed to map the text box position to a high-dimensional real number domain. Where d is the dimension of the embedded position vector.
[0073] Step S122: The text content mapped to the first real number field and the text box position mapped to the second real number field are concatenated to obtain the first fusion feature corresponding to each of the target candidate regions.
[0074] In this embodiment, the mapped and By concatenating the features, we obtain new features that include the text content of the target candidate region and the position of the text box. This refers to the first fusion feature. .
[0075] In summary, by concatenating the mapped text content and the text box position, the first fusion feature containing both text content and text structure is obtained.
[0076] Step S130: Perform spatial distance transformation on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions.
[0077] In this embodiment, the text content can be obtained from the above. Text box position and first visual features Each text box contains structural, content, and visual information about the document elements. After obtaining this information, the text box is flattened and combined with the text content to obtain the first fusion feature, which contains both the document structure and the text element content.
[0078] In this embodiment, the target candidate region is segmented and used as input to the visual information extraction model to obtain the first visual feature. For ease of subsequent processing, the first visual feature of the target candidate region needs to be mapped to a feature space consistent with the word vector space mentioned above. Since the segmented target candidate region is still a two-dimensional image, it needs to be unfolded into a one-dimensional vector before encoding. This application uses a multilayer perceptron structure to map the two-dimensional image corresponding to the target candidate region to... In space.
[0079] In this embodiment, after obtaining the first fused feature and the first visual feature, the first fused feature is input into a Transformer-based encoder for encoding to obtain the second fused feature. The target candidate region is then... Corresponding first visual features The second visual features are obtained by encoding them into the same feature space using a visual encoder. This ensures that the second fusion feature and the second visual feature are in the same feature space.
[0080] In this embodiment, the encoded features should ensure that the intra-class distance is as small as possible, while the inter-class distance is as large as possible, which means that the following formula needs to be optimized:
[0081]
[0082]
[0083] in, for Samples in the class, for The number of samples in the class. For Transformer-based encoders, These are the parameters of the encoder.
[0084] Step S140: Fuse the second fusion feature and the second visual feature corresponding to each of the target candidate regions to obtain document elements.
[0085] In this embodiment, since the encoded second fusion feature and the second visual feature are in the same feature space, the encoded second fusion feature and the second visual feature are concatenated together to obtain a complete feature description of the document element.
[0086] Based on the above technical solution, this application addresses the problem of insufficient model generalization ability caused by the lack of fusion of overall and local features in documents. It utilizes a convolutional neural network to segment feature maps and map them to an RPN network, then segments candidate regions with actual semantics, determines the text box positions and text content of the candidate regions, performs spatial distance transformation on the new features based on a Transformer encoder, and finally fuses the encoded features with global and local features. The document multi-element content extracted by this method is more accurate and has stronger generalization ability.
[0087] Furthermore, the step of identifying electronic documents based on a region selection network to obtain multiple target candidate regions includes:
[0088] Step S111: Input the electronic document into the convolutional neural network to extract target features and generate a shared feature map.
[0089] In this embodiment, the electronic document is input into a convolutional neural network, and a shared feature map is obtained through a series of convolutions. The convolutional neural network can be a Faster R-CNN network.
[0090] Step S112: Based on the region selection network, the shared feature map is slid with anchor boxes of different sizes to obtain candidate regions of different sizes and proportions.
[0091] In this embodiment, different sliding windows are set on the shared feature map, and the RPN network sets one sliding window. The scale and size of the sliding window can be determined according to the actual situation. Setting anchor boxes of different scales and proportions on the sliding window can yield multiple anchor boxes; candidate regions are generated through the sliding window.
[0092] Assuming the shared feature map size is N x 16 x 16, in the RPN stage, it first undergoes a 3 x 3 convolution to obtain a 256 x 16 x 16 feature map, which can also be viewed as 16 x 16 256-dimensional feature vectors. Then, it undergoes two 1 x 1 convolutions to obtain an 18 x 16 x 16 feature map and a 36 x 16 x 16 feature map, resulting in 16 x 16 x 9 results. Each result contains 2 scores and 4 coordinates. Combined with predefined anchor boxes, and after post-processing, candidate regions are obtained.
[0093] Step S113: Determine the target candidate region based on the intersection-union ratio of each candidate region with the anchor frame.
[0094] Further, determining the target candidate region based on the intersection-union ratio of each candidate region and the anchor frame includes the following steps:
[0095] Step S1131: When the intersection-union ratio is greater than a preset value, the candidate region with the intersection-union ratio greater than the preset value is determined as the target candidate region.
[0096] In this embodiment, the preset value can be set to 0.7. RPN mainly performs two different types of predictions: binary classification to determine whether a document element is present, and bounding box regression adjustment. During the model training phase, if the intersection-over-union (IoU) ratio of the candidate region and the anchor box is greater than 0.7, the candidate region is considered a document element region, i.e., the target candidate region; if the IoU ratio is less than 0.1, the candidate region is considered a background region; candidate regions in between are ignored. Here, IoU is the intersection-over-union ratio. It is the set of all pixels in all images in the training set. It represents a set The output of the network on the pixel probability, It is a set If the true value is , then IoU can be expressed as:
[0097]
[0098]
[0099]
[0100] For the retained target candidate regions, calculate the classification loss and the regression loss between the target candidate regions and the anchor boxes respectively.
[0101] Then, based on the obtained target candidate regions, sub-feature maps of the corresponding regions are cropped from the feature maps, and the sizes of these sub-feature maps are standardized. Next, the feature features of each target candidate region are classified using a simple multilayer perceptron network. Finally, two sets of prediction results are obtained: the first set represents the features corresponding to the target candidate region, and the second set represents the offset of the bounding box adjustment for the target candidate region. This operation yields the final document element segmentation and recognition results.
[0102] Second Embodiment
[0103] like Figure 2 As shown, based on the first embodiment, in the second embodiment of this application, before step S111, the following steps are further included:
[0104] Step S210: Perform recurrent adversarial training on the real document dataset and the pseudo document dataset based on the recurrent adversarial network to generate the first document dataset.
[0105] In this embodiment, two input datasets, X1 and X2, are used. X1 and X2 include features such as white background, watermark, wavy lines, tables, underlines, parentheses, colons, checkboxes, radio buttons, and horizontal lines. X1 has a white background without a watermark, while X2 has both a watermarked background and a natural background. X1 is a real document dataset, and X2 is a pseudo-document dataset. Based on the inputs X1 and X2, two generators G are trained using a recurrent adversarial network for X1 and X2. x and G y and two discriminators D x and D y The loss function consists of two parts, including the adversarial loss. GAN and cycle consistency loss cycle That is, Loss = Loss GAN +Loss cycle Loss GAN Guarantee G x G y and D x D y Mutual evolution, thereby ensuring the generator G x and G y It can produce more realistic images, Loss cycle Guarantee generator G x and G y The output image is the same as the input image, only the style is different.
[0106] Loss GAN The specific formula can be expressed as:
[0107]
[0108] in:
[0109]
[0110]
[0111] Loss cycle The specific formula can be expressed as:
[0112] +
[0113] During generator training, discriminator D x and D y The parameters are fixed, only G. x and G y The parameters are adjustable; adjust G. x The parameters make D y For G x The generated image G x (x) score D y (G x (x) should be as high as possible; adjust G y The parameters make D x For G y The generated image G y (y) score D x (G y The higher the value of (y), the better. Therefore, this ensures that the images generated by the generator become increasingly realistic.
[0114] During the training of the discriminator, the generator G x and G y The parameters are fixed, D x and D y Parameters are adjustable, maximize D x The value of (x) allows the discriminator to give the real image x a high score, while minimizing D. x (G y The value of (x) is used by the discriminator to generate the image G. y A lower score improves the discriminator's ability to distinguish. Update D y The process and D x Similar to.
[0115] The recurrent adversarial network is continuously trained using the loss function Loss, and finally the generator G... x and G y For the X1 and X2 datasets, dynamically generate a near-realistic first document dataset. In practical applications, the backgrounds of electronic documents are quite complex, including features such as watermarks, ripples, tables, underlines, brackets, colons, checkboxes, radio buttons, and horizontal lines. This step addresses the problem of insufficient model generalization ability by pre-generating a large number of rich and realistic application production style background datasets.
[0116] Step S220: Input the document datasets corresponding to different text types into the deep forgery adversarial network in the X domain to generate the second document dataset.
[0117] In this embodiment, the document datasets corresponding to different text types include, but are not limited to: Chinese Unicode character datasets, English datasets, and Arabic datasets. The Chinese Unicode character dataset is constructed using a semi-supervised approach. English dataset and Arabic dataset This includes the first document dataset with application production style background. . , , and The DeepFake adversarial network is used as the input to the X domain. The entire network shares a single encoder, which employs fully connected layers to disrupt the spatial relationships within the features extracted by the previous convolutional layers, allowing for full computation between each pixel. The decoder uses a PixelShuffler structure. Using two decoders allows the encoder to learn richer text features, encoding different text features into the same latent space, and then "reconstructing" them using different decoders in different ways. This ensures that the generated results have greater practical application value. The DeepFake adversarial network is used to dynamically generate a Y-domain document dataset, i.e., the second document dataset. .
[0118] In this embodiment, deepfake adversarial networks can be divided into four categories: reproduction, replacement, editing, and synthesis. Due to the many subtle features of electronic documents, they are very suitable for practical applications. The second document dataset... The specific construction process includes three stages: detection, alignment, and mask generation. Details are as follows:
[0119] (1) In the first document dataset In each frame, identify features such as tables, underlines, parentheses, colons, checkboxes, radio buttons, and horizontal bars. Crop these features and then analyze the cropped feature map. Perform eight rotation transformations, such as up / down / left / right, top left, bottom left, top right, and bottom right, and save them in order to form a new feature sequence. .
[0120] (2) For new feature sequences The algorithm performs alignment and classification, attempting to identify potential similar features such as checkboxes, radio buttons, underlines, and horizontal bars. Then, it tries to use this information to align checkboxes, radio buttons, underlines, and horizontal bars.
[0121] (3) Identify aligned features such as underscores, parentheses, colons, checkboxes, radio buttons, and horizontal bars, and mask areas containing backgrounds / obstacles.
[0122] This application extracts features primarily for two purposes: training and transformation. The model is trained using a newly extracted set of data, including underscores, parentheses, colons, checkboxes, radio buttons, and horizontal bars. This training data includes alignment and masks, which are required by the model. This information is stored in the metadata of these extracted features. An alignment file and mask are generated for transforming the final image. The alignment file contains the specific location information of these features, allowing the model to swap new similar features at any given input image location during the transformation process.
[0123] When converting into new features, , , The document dataset is filled with discrete random variables outside the new feature regions to synthesize a second document dataset. Because DeepFake uses MAE as the loss and has meanness, it can cause blurry images with dynamic style transformations. Therefore, two discriminators need to be reintroduced, each corresponding to a different discriminator. and Used to supplement the details of the generated image, without considering and The difference is that the results obtained in this way are visually similar. Firstly, it saves some parameters, and secondly, if... and When the amount of data for either one is small, the training will be more stable.
[0124] Step S230: The first document dataset and the second document dataset are used as the X domain, and then input into the deep forgery adversarial network again to generate the third document dataset.
[0125] In this embodiment, and As a new X domain, a new recurrent adversarial network is formed, and then a new electronic document style dataset, namely the third document dataset, is dynamically generated through the DeepFake adversarial network. ,in, The generation steps and Similar, the only difference is, and There will be instances where the background remains consistent. This is to ensure that when transferring a new style, the transferred style doesn't deviate significantly and remains within a controllable range. Thus, It refers to a sample set that is feature-rich, of high quality, and indistinguishable from real data.
[0126] Step S240: Train a convolutional neural network using the third document dataset.
[0127] In this embodiment, by utilizing a third document dataset A backbone network based on a convolutional neural network (CNN) is trained to extract target features based on the CNN and generate a shared feature map. The region selection network slides the shared feature map with anchor boxes of different sizes to obtain candidate regions of different sizes and proportions. The target candidate region is determined based on the intersection-union ratio (IUU) of each candidate region with the anchor box.
[0128] Based on the above technical solution, this application proposes a method for generating indistinguishable GAN adversarial dynamic transformations from real document style data and pseudo-style data. This method can generate richer feature images and also generate high-quality samples that are relevant to business needs.
[0129] Third Embodiment
[0130] like Figure 3 As shown, based on any embodiment of the first and second embodiments, in the third embodiment of this application, after step S140, the following steps are further included:
[0131] Step S310: Obtain the global feature map associated with the document elements. The global feature map includes multiple feature items, wherein each feature item includes recommended content and rule requirements.
[0132] In this embodiment, one or more document elements are identified in an electronic document. Document elements mainly refer to content information in the electronic document that needs to be filled or edited, such as date and stamp features. Date extraction requires distinguishing between values and null values; only null values need to be extracted. The extracted date information mainly includes semantic information from the context. Based on the semantic context, invalid content can be further filtered out to ensure that the extracted date is editable. Stamp area feature extraction requires further judgment on the rationality of the stamp area based on whether the blank area is near the date. Each extracted feature has three key attributes: type, area range, and contextual content.
[0133] In this embodiment, the global feature map is pre-constructed and mainly includes type, size, length, rule requirements, and recommended content. That is, the global feature map includes multiple feature items, and each feature item includes recommended content and rule requirements. The recommended content is not singular or static; it is dynamically adjusted based on the actual feedback. The feedback entry point is the user's front-end interface. Using this closed-loop, iterative feedback approach, the recommended content is continuously optimized and adjusted, enriching the global feature map. This method improves the quality of the recommended content.
[0134] Step S320: Determine the correlation distance between the document element and each feature item, and obtain the target feature item corresponding to the document element based on the correlation distance.
[0135] In this embodiment, the type of document element is pre-associated with the type of global feature map, which can find the matching rule requirements and recommended content of document element from the global feature map, that is, dynamically find the rule requirements and recommended content corresponding to each document element, that is, find the target feature item corresponding to each document element.
[0136] In this embodiment, the correlation coefficient and correlation distance between document elements and each feature item are calculated. Optionally, the normalized values of the correlation coefficients can be sorted, and the feature item corresponding to the highest correlation coefficient can be taken as the target feature item. Alternatively, the correlation distances between document elements and each feature item can be sorted, and the feature item corresponding to the smallest correlation distance can be taken as the target feature item, thereby obtaining the recommended content and rule requirements.
[0137] Step S330: Construct a tree structure diagram based on the document elements and the recommended content and rule requirements corresponding to the target feature items;
[0138] In this embodiment, after determining the target features, a tree structure diagram is constructed. The tree structure diagram can intuitively display the recommended content and rule requirements corresponding to each document element in the electronic document.
[0139] Optionally, a three-level tree structure diagram is constructed based on the recommended content and rule requirements corresponding to document elements and target feature items, as shown in the figure. Figure 4 As shown, the first level consists of document elements of the electronic document, including coordinate position and contextual relevance; the second level consists of recommended content, which is mainly automatically generated by global feature maps through induction and NLP semantic models; and the third level consists of rule requirements, such as text size, minimum and maximum text length, required fields, data type format, etc.
[0140] Step S340: Generate an online editing page based on the tree structure diagram. The online editing page includes a template recommendation area and an online editing area.
[0141] In this embodiment, an online editing page is automatically constructed based on the content returned by the tree structure diagram. The online editing page can adopt a horizontal layout, for example, the left side is a template recommendation area, and the right side is the online editing area. The left side uses a template validation mode, which distinguishes document elements, i.e., the content to be filled, by color: required fields are red, and optional fields are orange. Rectangular areas are used to separate required and optional fields. The right side is the actual online editing mode. Online editing can directly refer to the intelligent template recommendations on the left. When inputting content in the area to be filled, the system performs strict validation according to rules. This approach standardizes and unifies the content to be filled during multi-person collaboration, and the two-area segmentation intelligent template validation method improves the processing efficiency of online electronic documents.
[0142] Further, the step of determining the correlation distance between the document element and each feature term, and obtaining the target feature term corresponding to the document element based on the correlation distance includes:
[0143] Step S321: Determine the correlation coefficient based on the covariance between the document elements and each feature item, the variance of the document elements, and the variance corresponding to each feature item. The larger the absolute value of the correlation coefficient, the higher the correlation between the document elements and the feature items.
[0144] In this embodiment, the correlation coefficient between a document element and one of its feature terms is calculated as an example. The square root of the variance of the document element is taken, and the square root of the variance of the feature term is taken. The product of the square rooted variance of the document element and the square rooted variance of the feature term is used as the denominator. The covariance between the document element and the feature term is used as the numerator to calculate the correlation coefficient.
[0145] The correlation coefficient can be calculated using the formula for calculating the correlation coefficient. The formula for calculating the correlation coefficient is as follows:
[0146] .
[0147] In the formula, Cov(X, Y) is the covariance of document elements and feature terms, D(X) is the variance of document elements, D(Y) is the variance of feature terms, and the correlation coefficient is a method to measure the degree of correlation between two feature columns. The value range is [-1, 1]. The larger the absolute value of the correlation coefficient, the higher the degree of correlation between document elements and feature terms.
[0148] Step S322: Determine the correlation distance between the document element and each feature item based on the correlation coefficient, wherein the larger the correlation coefficient, the shorter the correlation distance;
[0149] In this embodiment, the formula for calculating the relevant distance is:
[0150] .
[0151] Wherein, the larger the correlation coefficient, the shorter the correlation distance, and the correlation distance D xy The value of is such that the closer it is to 0, the shorter the distance, and vice versa.
[0152] Step S323: The feature term corresponding to the shortest relevant distance is determined as the target feature term.
[0153] In this embodiment, the relevance distances between document elements and each feature item can be sorted, and the feature item with the smallest relevance distance is taken as the target feature item, thereby obtaining the recommended content and rule requirements. The smallest relevance distance indicates that the recommended content and rule requirements corresponding to the feature item are closer to the document element, thus enabling accurate matching of the recommended content and rule requirements corresponding to each document element.
[0154] In this embodiment, the present application uses intelligent extraction of multi-element and global feature maps to calculate spatial distance, finds the most suitable recommended content and rule requirements, constructs a tree-structured recommendation template, and constructs a two-region layout for the online editing page based on the content of the recommendation template. The layout of the two regions is a horizontal mode, with the left side being the system's intelligent template verification mode. This approach can improve the processing efficiency of online electronic documents and standardize and unify the content to be filled in when multiple people are collaborating.
[0155] This invention provides an embodiment of a document element identification and extraction method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0156] like Figure 5 As shown, this application provides a document element recognition and extraction system, which includes:
[0157] The recognition module 10 is used to recognize electronic documents based on a region selection network to obtain multiple target candidate regions.
[0158] Optionally, the recognition module 10 is further configured to input the electronic document into a convolutional neural network to extract target features and generate a shared feature map; slide the shared feature map with anchor boxes of different sizes based on the region selection network to obtain candidate regions of different sizes and proportions; and determine the target candidate region according to the intersection-union ratio of each candidate region with the anchor box.
[0159] Optionally, the identification module 10 is further configured to determine the candidate region with an intersection-union ratio greater than a preset value as the target candidate region when the intersection-union ratio is greater than a preset value.
[0160] The first fusion module 20 is used to determine the text content and text box position corresponding to each target candidate region, and to obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region.
[0161] Optionally, the first fusion module 20 is further configured to map the text content corresponding to each of the target candidate regions to a first real number field, and map the text box position corresponding to each of the target candidate regions to a second real number field; and concatenate the text content mapped to the first real number field and the text box position mapped to the second real number field to obtain the first fusion feature corresponding to each of the target candidate regions.
[0162] Encoding module 30 is used to perform spatial distance transformation on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions;
[0163] The second fusion module 40 is used to fuse the second fusion features and second visual features corresponding to each of the target candidate regions to obtain document elements.
[0164] Optionally, after the second fusion module 40, an online editing page generation module is further connected. The online editing page generation module is used to obtain a global feature map associated with the document elements. The global feature map includes multiple feature items, wherein each feature item includes recommended content and rule requirements; determine the correlation distance between the document elements and each feature item, and obtain the target feature item corresponding to the document elements based on the correlation distance; construct a tree structure diagram according to the document elements and the recommended content and rule requirements corresponding to the target feature items; and generate an online editing page based on the tree structure diagram. The online editing page includes a template recommendation area and an online editing area.
[0165] Optionally, the online editing page generation module is further configured to determine a correlation coefficient based on the covariance between the document element and each feature item, the variance of the document element, and the variance corresponding to each feature item, wherein the larger the absolute value of the correlation coefficient, the higher the correlation between the document element and the feature item; determine the correlation distance between the document element and each feature item based on the correlation coefficient, wherein the larger the correlation coefficient, the shorter the correlation distance; and determine the feature item corresponding to the shortest correlation distance as the target feature item.
[0166] Optionally, a dataset generation module is connected before the recognition module 10. The dataset generation module is used to perform recurrent adversarial training on real document datasets and pseudo document datasets based on a recurrent adversarial network to generate a first document dataset; input document datasets corresponding to different text types as the X domain into a deep forgery adversarial network to generate a second document dataset; input the first document dataset and the second document dataset as the X domain again into the deep forgery adversarial network to generate a third document dataset; and train a convolutional neural network using the third document dataset.
[0167] The specific implementation of the document element recognition and extraction system of the present invention is basically the same as the embodiments of the document element recognition and extraction method described above, and will not be repeated here.
[0168] like Figure 6 As shown, Figure 6 This is a schematic diagram of the document element recognition and extraction device of the present invention. The document element recognition and extraction device may include: a processor 1001, such as a CPU; a memory 1005; a user interface 1003; a network interface 1004; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface. The memory 1005 may be a high-speed RAM memory or a stable memory, such as a disk storage device. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001. For example, the document element recognition and extraction device may be a mobile phone, computer, etc.
[0169] Those skilled in the art will understand that Figure 6 The document feature recognition and extraction device structure shown does not constitute a limitation on the document feature recognition and extraction device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0170] like Figure 6 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a document element recognition and extraction program. The operating system is a program that manages and controls the hardware and software resources of the document element recognition and extraction device, as well as the operation of the document element recognition and extraction program and other software or programs.
[0171] exist Figure 6In the document element recognition and extraction device shown, the user interface 1003 is mainly used to connect to the terminal and communicate data with the terminal; the network interface 1004 is mainly used to communicate data with the backend server; and the processor 1001 can be used to call the document element recognition and extraction program stored in the memory 1005.
[0172] In this embodiment, the document element recognition and extraction device includes: a memory 1005, a processor 1001, and a document element recognition and extraction program stored in the memory and executable on the processor, wherein:
[0173] When processor 1001 calls the document element recognition and extraction program stored in memory 1005, it performs the following operations:
[0174] The electronic document is identified using a region selection network, resulting in multiple target candidate regions;
[0175] Determine the text content and text box position corresponding to each target candidate region, and obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region;
[0176] Spatial distance transformation is performed on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions.
[0177] By fusing the second fusion feature and the second visual feature corresponding to each of the target candidate regions, document elements are obtained.
[0178] When processor 1001 calls the document element recognition and extraction program stored in memory 1005, it also performs the following operations:
[0179] Obtain the global feature map associated with the document elements. The global feature map includes multiple feature items, wherein each feature item includes recommended content and rule requirements.
[0180] Determine the correlation distance between each document element and each feature term, and obtain the target feature term corresponding to the document element based on the correlation distance;
[0181] Based on the document elements and the recommended content and rule requirements corresponding to the target feature items, construct a tree structure diagram;
[0182] An online editing page is generated based on the tree structure diagram. The online editing page includes a template recommendation area and an online editing area.
[0183] When processor 1001 calls the document element recognition and extraction program stored in memory 1005, it also performs the following operations:
[0184] The first document dataset is generated by training a recurrent adversarial network on real and pseudo document datasets.
[0185] The document datasets corresponding to different text types are used as input to the deep fakery adversarial network in the X domain to generate a second document dataset.
[0186] The first and second document datasets are used as the X domain, and then input into the deep fakery adversarial network again to generate the third document dataset;
[0187] The convolutional neural network was trained using the aforementioned third document dataset.
[0188] Based on the same inventive concept, this application also provides a computer-readable storage medium storing a document element identification and extraction program. When the document element identification and extraction program is executed by a processor, it implements the various steps of the document element identification and extraction method described above and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0189] Since the storage medium provided in this application embodiment is the storage medium used to implement the method of this application embodiment, those skilled in the art can understand the specific structure and variations of the storage medium based on the method described in this application embodiment, and therefore will not be repeated here. All storage media used in the method of this application embodiment are within the scope of protection of this application.
[0190] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0191] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0192] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, television, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0193] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for identifying and extracting document elements, characterized in that, The document element identification and extraction method includes: The electronic document is identified using a region selection network, resulting in multiple target candidate regions; Determine the text content and text box position corresponding to each target candidate region, and obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region; Spatial distance transformation is performed on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions. By fusing the second fusion feature and the second visual feature corresponding to each of the target candidate regions, document elements are obtained; Obtain the global feature map associated with the document elements. The global feature map includes multiple feature items, wherein each feature item includes recommended content and rule requirements. Determine the correlation distance between each document element and each feature term, and obtain the target feature term corresponding to the document element based on the correlation distance; Based on the document elements and the recommended content and rule requirements corresponding to the target feature items, construct a tree structure diagram; An online editing page is generated based on the tree structure diagram. The online editing page includes a template recommendation area and an online editing area.
2. The document element recognition and extraction method as described in claim 1, characterized in that, The step of obtaining the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region includes: Map the text content corresponding to each of the target candidate regions to the first real number field, and map the text box position corresponding to each of the target candidate regions to the second real number field; The text content mapped to the first real number field and the text box position mapped to the second real number field are concatenated to obtain the first fusion feature corresponding to each of the target candidate regions.
3. The document element recognition and extraction method as described in claim 1, characterized in that, The step of determining the correlation distance between the document element and each feature term, and obtaining the target feature term corresponding to the document element based on the correlation distance includes: The correlation coefficient is determined based on the covariance between the document elements and each feature item, the variance of the document elements, and the variance of each feature item. The larger the absolute value of the correlation coefficient, the higher the correlation between the document elements and the feature items. Based on the correlation coefficient, the correlation distance between the document element and each feature item is determined, wherein the larger the correlation coefficient, the shorter the correlation distance; The feature term corresponding to the shortest relevant distance is determined as the target feature term.
4. The document element identification and extraction method as described in claim 1, characterized in that, The step of identifying electronic documents based on a region selection network to obtain multiple target candidate regions includes: The electronic document is input into a convolutional neural network to extract target features and generate a shared feature map. Based on the region selection network, the shared feature map is slid with anchor boxes of different sizes to obtain candidate regions of different sizes and proportions; The target candidate region is determined based on the intersection-union ratio of each candidate region with the anchor frame.
5. The document element identification and extraction method as described in claim 4, characterized in that, The step of determining the target candidate region based on the intersection-union ratio of each candidate region and the anchor frame includes: When the cross-union ratio is greater than a preset value, the candidate region with the cross-union ratio greater than the preset value is determined as the target candidate region.
6. The document element identification and extraction method as described in claim 4, characterized in that, Before the step of inputting the electronic document into a convolutional neural network to extract target features and generate a shared feature map, the method further includes: The first document dataset is generated by training a recurrent adversarial network on real and pseudo document datasets. The document datasets corresponding to different text types are used as input to the deep fakery adversarial network in the X domain to generate a second document dataset. The first and second document datasets are used as the X domain, and then input into the deep fakery adversarial network again to generate the third document dataset; The convolutional neural network was trained using the aforementioned third document dataset.
7. A document element recognition and extraction system, characterized in that, The document element recognition and extraction system includes: The recognition module is used to recognize electronic documents based on a region selection network to obtain multiple target candidate regions; The first fusion module is used to determine the text content and text box position corresponding to each target candidate region, and to obtain the first fusion feature corresponding to each target candidate region based on the text content and text box position corresponding to each target candidate region. The encoding module is used to perform spatial distance transformation on the first fusion feature and the first visual feature corresponding to each of the target candidate regions to obtain the second fusion feature and the second visual feature corresponding to each of the target candidate regions. The second fusion module is used to fuse the second fusion features and second visual features corresponding to each of the target candidate regions to obtain document elements; Obtain the global feature map associated with the document elements. The global feature map includes multiple feature items, wherein each feature item includes recommended content and rule requirements. Determine the correlation distance between each document element and each feature term, and obtain the target feature term corresponding to the document element based on the correlation distance; Based on the document elements and the recommended content and rule requirements corresponding to the target feature items, construct a tree structure diagram; An online editing page is generated based on the tree structure diagram. The online editing page includes a template recommendation area and an online editing area.
8. A document element recognition and extraction device, characterized in that, The document element recognition and extraction device includes: a memory, a processor, and a document element recognition and extraction program stored in the memory and running on the processor. When the document element recognition and extraction program is executed by the processor, it implements the steps of the document element recognition and extraction method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a document element recognition and extraction program, which, when executed by a processor, implements the steps of the document element recognition and extraction method according to any one of claims 1-6.
Citation Information
Patent Citations
Bill information extraction method and device, equipment, medium and product
CN115205884A