An ocr recognition and field structuring processing method for complex layout tickets
Patent Information
- Application Number
- CN202611063702.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-17
AI Technical Summary
[0004]在复杂排版票据OCR识别中,现有技术受限于固定模板与关键词锚定机制,难以适应财务票据版式无序、横竖排版混用的非标准化特征,同时,票面红色印鉴对关键字段的像素级遮挡,造成传统图像处理方法无法有效还原被覆盖的文字笔画,严重影响识别完整性,此外,现有方案缺乏对票据空间布局与语义关联的深层理解,导致多行表格及键值对关系解析时频繁出现字段错位或逻辑混淆,难以满足全票面信息高精度结构化的实际需求,因此,如何在不依赖固定版式先验的条件下,克服印鉴干扰、多方向文字混排及复杂空间关系解析等障碍,实现复杂排版票据中关键字段的准确提取与结构化输出,是本发明要解决的技术问题
1、该一种面向复杂排版票据的OCR识别与字段结构化处理方法,通过融合文本语义、空间位置与视觉特征构建综合表征,并基于图神经网络对票据内容的拓扑关系进行推理,能够适应任意陌生版式的票据,即使票据的字段布局发生任意变化,仍能依据语义与空间关联准确锁定关键信息,降低了系统部署与运维成本。
Smart Images

Figure CN122574891B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent document processing technology, and in particular to a method for OCR recognition and field structuring processing of complex typesetting documents. Background Technology
[0002] With the digital transformation of corporate financial shared service centers and cross-border businesses, the demand for automated processing of massive amounts of financial documents is becoming increasingly urgent. In actual business operations, overseas documents, represented by Japanese documents (such as receipts and requests), are extremely difficult to process. There is no unified format standard for these documents, resulting in thousands of unstructured layouts and extremely complex text features, with a common mix of horizontal and vertical text, as well as handwritten and printed text. In addition, the surface of the documents is often covered with red seals (such as personal seals or company corner seals), which often directly obscure key business fields such as amounts and dates, creating high-frequency and difficult-to-remove noise interference.
[0003] For example, a lightweight OCR recognition method and system for invoices, published in Chinese Patent Publication No. CN116612479A, achieves high classification efficiency for automated classification of all mixed invoices. It utilizes open-source text recognition and text detection models, fine-tuning them before entering the optimization algorithm to ensure classification cost while meeting classification accuracy requirements. Finally, comprehensive error correction can improve the accuracy of the returned results.
[0004] In OCR recognition of complex-formatted invoices, existing technologies are limited by fixed templates and keyword anchoring mechanisms, making it difficult to adapt to the non-standardized characteristics of financial invoices, such as disordered layouts and mixed horizontal and vertical layouts. Furthermore, the pixel-level occlusion of key fields by the red seal on the invoice makes it impossible for traditional image processing methods to effectively restore the covered text strokes, severely affecting the integrity of the recognition. In addition, existing solutions lack a deep understanding of the spatial layout and semantic relationships of invoices, leading to frequent field misalignment or logical confusion when parsing multi-line tables and key-value pairs, making it difficult to meet the practical need for high-precision structuring of all invoice information. Therefore, the technical problem this invention aims to solve is how to overcome obstacles such as seal interference, mixed text layouts in multiple directions, and complex spatial relationship parsing without relying on a fixed format prior. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, the present invention provides an OCR recognition and field structuring method for complex typesetting documents, which can effectively solve the problems involved in the prior art.
[0006] The objective of this invention can be achieved through the following technical solution: This invention provides an OCR recognition and field structuring method for complex typesetting documents, the specific steps of which are as follows: Step 1: Image enhancement and seal restoration based on generative adversarial networks. The residual network is used to locate the seal occlusion area and generate a seal mask. The unoccluded text strokes and background textures are used to reconstruct the covered area at the pixel level, remove red noise from the seal and complete the broken strokes, and output a high-definition seal-free ticket image, providing a clean image input for subsequent recognition and significantly improving the recognizability of the text in the occluded area. Step 2: Text detection is performed using an improved differentiable binarization algorithm, which outputs the coordinate information of the horizontal and vertical text boxes. A convolutional recurrent neural network is used in conjunction with a connection-time classification loss function to perform character recognition, so as to output the raw data stream containing text content, confidence score and four-point coordinate position information. Step 3: The acquired text content is transformed into semantic vectors through a pre-trained language model, the normalized coordinates of the text boxes are mapped into positional encoding vectors, and then concatenated with the image visual features to form a comprehensive feature representation, providing a multi-dimensional node attribute foundation for graph network inference; Step 4: Treat each text box as a node, establish edge connections based on Euclidean distance and semantic similarity, construct a topological graph of the ticket content, and perform node classification (key / value discrimination) and edge relationship prediction (key-value pairing) on the graph structure to get rid of the dependence on fixed coordinates; Step 5: Perform logical post-processing and verification based on financial and tax rules. Use the built-in financial and tax logic engine to standardize the data (automatically convert the identified calendar dates to Gregorian calendar date format, and convert full-width numbers and special currency symbols to standard half-width numbers), and perform self-consistency verification. Introduce business rules in the financial and tax field to standardize and verify the consistency of the identification results. Step Six: Based on the self-consistency verification results, determine whether to mark a prompt, and combine the graph network reasoning and rule verification results to generate structured data containing key-value pairs, table row and column relationships, and confidence scores. No preset template is required, it can be directly adapted to any unfamiliar invoice, and outputs a standardized JSON format that can be directly connected to the financial system.
[0007] Preferably, step one specifically includes: The preprocessed ticket image is input into the residual network for multi-scale feature extraction. The texture and edge information under different receptive fields are captured through the feature pyramid structure. Based on the attention mechanism, the unique color saturation and edge sharpness of the seal are enhanced to generate a seal mask map that accurately identifies the occluded pixel area. It is a binary mask map, which provides accurate occlusion area location basis for subsequent mask generation and pixel repair. The original ticket image and the seal mask are input into the generator network. The seal mask is used to locate the area to be repaired, and the stroke direction, gray level gradient and background texture distribution features of the text in the unmasked area are extracted as reference features for pixel-level reconstruction. The prior features of the background and text are fully utilized to guide the repair process and ensure the authenticity of the reconstructed content. Based on the extracted reference features, the generator performs pixel-by-pixel context-aware filling on the mask-covered area, simultaneously eliminating the chromatic interference of the red seal and restoring the morphological continuity of the fragmented characters. The output is a seal-free ticket image with complete text strokes and a clean background. While removing the chromatic interference of the seal, the integrity and continuity of the text strokes are maintained, significantly improving the readability of the text in the occluded area.
[0008] Preferably, the process of performing pixel-by-pixel context-aware fill in step one is as follows: A generator network containing multi-level dilated convolutions is constructed. Convolutional kernels with different dilation rates are used to extract contextual texture information of multiple scales around the mask region in parallel. Based on attention weights, features of each scale are adaptively fused to generate a combined feature vector containing global and local information for each pixel to be repaired. This expands the receptive field to capture long-distance contextual dependencies and achieves complementarity between global and local information. The combined feature vector is input into the pixel generation module. Combined with the gradient continuity constraint at the mask boundary and the consistency constraint of the text stroke direction, the fill content is synthesized pixel by pixel, which is seamlessly connected with the surrounding background in terms of tone, brightness and texture. The constraint boundary transitions smoothly and is consistent with the stroke direction, ensuring that the repaired area blends naturally with the background. A multi-scale discriminator is used to apply adversarial constraints to the restoration results. The generator's filling quality is supervised from two levels: global image semantics and local stroke details. This ensures that the restored text strokes are consistent with the font style of the original document. Multi-scale supervision guarantees overall color harmony and realistic local details, driving continuous optimization of the generator.
[0009] Preferably, step two specifically includes: A direction-sensitive feature extraction branch is added to the differentiable binarization detection network. The axis alignment features of horizontal and vertical text are independently encoded by multi-angle convolution kernels. At the same time, the regression parameters of horizontal and vertical candidate boxes are predicted to output the four-point coordinate detection results containing the direction label. The horizontal and vertical layout features are independently encoded by multi-directional convolution kernels to achieve accurate detection with direction awareness. Based on the direction labels of the four-point coordinate detection results, different sequential reading strategies are adopted for horizontal and vertical text boxes to ensure that vertical text is serialized and spliced in the column direction, avoiding character order disorder caused by horizontal splicing. The reading order is dynamically adjusted according to the direction labels to ensure that characters are arranged in the correct reading direction. The serialized text image blocks are sequentially input into a convolutional recurrent neural network. A bidirectional long short-term memory network is used to extract temporal dependency features. The text content and its confidence score are then output through a connected temporal classification decoder, forming a multi-dimensional raw data stream containing text content, confidence score, and four-point coordinate information. Temporal modeling is used to capture the dependencies between characters, thereby improving the accuracy of complex character recognition.
[0010] Preferably, step three specifically includes: The output text content is input into a language model pre-trained on a large-scale corpus, and high-dimensional semantic embedding vectors corresponding to each text box are extracted to capture the deep semantic commonalities of the same field under different expressions. The pre-trained language model is used to convert the text into semantic vectors to eliminate the ambiguity of understanding caused by synonymous expressions. The normalized center coordinates and aspect ratios of each text box are mapped to a continuous space of the same dimension as the semantic embedding vector to form an absolute position encoding vector. At the same time, the relative displacement between text boxes is encoded into a relative position relationship vector. The spatial coordinates and relative displacement are encoded into a high-dimensional space to quantify the positional dependency between fields. The semantic embedding vector, absolute position encoding vector, relative position relationship vector, and corresponding region visual feature vector of the unsigned ticket image are cross-modal concatenated and linearly transformed to output a comprehensive feature representation matrix of each text box. Cross-modal fusion of text semantics, spatial location and visual appearance is used to construct a multi-dimensional node feature representation.
[0011] Preferably, step four specifically includes: Each text detection box in the ticket is used as an independent node in the topology graph. The comprehensive feature representation of each node is used as the initial node attribute. Based on the Euclidean distance matrix and semantic vector cosine similarity matrix between all nodes, node pairs that meet the preset adjacency conditions are selected to establish edge connections. The graph edge connections are established using the dual constraints of spatial distance and semantic similarity to accurately locate the association between fields. The constructed node attribute matrix and adjacency matrix are input into a multi-layer graph convolutional network. The feature representation of each node is iteratively updated through a neighborhood aggregation mechanism, so that each node gathers the spatial location information and semantic category information of its neighboring nodes, and realizes the full interaction between node features and their context. A linear classifier is applied to the updated node features. Based on the semantic representation pattern of each node and the spatial distribution features of surrounding nodes, it is determined whether each node belongs to the key field category or the value field category. The node-level key / value attribute labels are output. Semantic role determination is performed based on the fused context features to accurately distinguish field types.
[0012] Preferably, step four further includes: After updating the graph convolutional network, all key node features and value node features are aggregated into key feature set and value feature set, respectively. For each key-value node pair, their feature vectors are concatenated and input into the edge prediction multilayer perceptron. The matching confidence score of the pairing relationship is calculated. The reliability of key-value matching is quantified by the pairing relationship score, and low-confidence combinations are filtered out. A global assignment matrix is constructed based on the matching confidence scores of all key-value nodes, and a two-way row and column constraint is introduced. The Hungarian algorithm is used to solve the globally optimal one-to-one key-value pairing scheme, and the reliability confidence score corresponding to each matching edge is output. The globally optimal matching ensures that each key-value field is paired only once, avoiding duplication and conflict. The paired key-value nodes are clustered according to their spatial adjacency in the original document. The block type is automatically determined based on the distribution of node coordinates within each group. The field-level semantic units and table row and column structures are jointly parsed. The pairs are assigned to the corresponding blocks through spatial clustering, the table row and field list structures are identified, and the block layout of the document is restored.
[0013] Preferably, step five specifically includes: A structured knowledge base containing calendar conversion rules, numerical standardization rules, and arithmetic operation constraint rules from the Japanese fiscal and tax system is constructed. The corresponding subset of verification rules is automatically activated based on the invoice type identifier to adapt to the differentiated business requirements of invoices in different industries. The fiscal and tax rules are organized by industry layer to achieve accurate activation and efficient matching of verification logic. The original recognition content of each field in the key-value pairing results output in step four is sent to the financial and tax logic engine. Based on the knowledge base, the engine performs format conversion between the calendar year and the Gregorian calendar year, and normalizes full-width numbers, Chinese numbers and special currency symbols into a unified half-width numerical expression form. Non-standard dates and numerical formats are converted into a unified expression to eliminate ambiguity between measurement and year. Based on arithmetic operation constraint rules, cross-field self-consistency verification is performed on the standardized numerical field group. Specifically, this includes comparing the consistency of the cumulative value of the detailed items and the value of the summary items within the same group, verifying the balance of addition and subtraction operations between component items, performing logical verification on related numerical fields, and verifying the balance relationship between detailed summaries and components.
[0014] Preferably, step five further includes: Based on the business type of the invoice, the corresponding arithmetic logic expression template is loaded from the knowledge base. The standardized numerical fields are substituted into the variable placeholders in the template. The cumulative or operation result on the left side of the expression and the target reference value on the right side are calculated respectively. The corresponding template is loaded according to the business type, and the standardized fields are substituted to perform automatic verification, providing differentiated arithmetic verification logic for different invoices. Set relative error tolerance thresholds and absolute error tolerance thresholds, and compare the calculation results with the target reference value item by item under the dual threshold constraints. If the absolute difference between the calculation result and the target reference value does not exceed the absolute error tolerance threshold, or the relative error does not exceed the relative error tolerance threshold, then the field is determined to pass the verification. If both errors exceed the corresponding tolerance range, it is determined to be a verification anomaly. The comparison is carried out using both relative and absolute tolerances to take into account the verification sensitivity of large and small amounts of data. The fields that are checked for anomalies are marked in multiple levels, and the anomaly type, the difference quantification value and the field identifiers involved are recorded respectively. The marking information is also attached to the corresponding key-value pairing records for classification, display and review guidance in the subsequent output stage. The anomaly fields are marked in layers according to type and degree of deviation, and the anomaly details and related fields are recorded in a complete manner.
[0015] Preferably, step six specifically includes: The key-value pairing results, table row and column parsing results, and verification anomaly marker information output in step four are merged and aligned. Based on the preset standardized data exchange architecture, each field is mapped to the corresponding data level path. The multi-source heterogeneous identification and verification results are integrated according to a unified data architecture to ensure accurate field placement. By combining the confidence score of graph network inference with the results of tax rule verification, a composite quality score is assigned to each key-value pair. Based on the score threshold, the output fields are divided into three levels: high confidence fields, low confidence fields, and verification anomaly fields. The data quality is evaluated by combining the inference confidence and rule verification results, and the data is classified according to the degree of reliability, so that the output data has a quality classification label. The output content is organized into three levels, and a structured data message containing field key names, standardized values, confidence scores, verification status markers and exception details is generated in JSON format. This directly adapts to the data import interface specifications of downstream financial systems, and the output message is organized into a hierarchical structure.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. This method for OCR recognition and field structuring of complex-formatted invoices constructs a comprehensive representation by integrating text semantics, spatial location, and visual features, and infers the topological relationships of the invoice content based on graph neural networks. It can adapt to any unfamiliar invoice format, and even if the field layout of the invoice changes arbitrarily, it can still accurately locate key information based on semantics and spatial association, thereby reducing system deployment and maintenance costs.
[0017] 2. This method for OCR recognition and field structuring of complex-formatted documents addresses the common problem of red seal obscuring in documents by employing generative deep repair technology. Through pixel-by-pixel context-aware reconstruction of the covered area, it can automatically restore the fragmented strokes of the text while completely removing the color noise of the seal. This transforms a large number of documents that cannot be recognized by traditional OCR due to seal obscuring into effective data that can be processed automatically, expanding the coverage of automated processing and improving the overall data utilization rate.
[0018] 3. This OCR recognition and field structuring method for complex layout invoices uses a direction-aware mechanism to simultaneously detect text boxes facing different directions and executes differentiated serialization reading strategies based on the reading direction. This ensures that each text area is recognized in the correct character order, avoiding character errors or missed detections caused by misjudgment of direction, and guaranteeing the complete extraction and accurate output of all text information on the invoice.
[0019] 4. This method for OCR recognition and field structuring of complex-formatted documents constructs a topological graph of the document content. Under the dual constraints of spatial adjacency and semantic similarity, it performs association reasoning on text nodes, accurately determines whether each text belongs to a key field or a value field, and establishes key-value correspondence through a global optimal matching strategy. It fully explores the spatial arrangement rules and semantic dependencies between fields, and can still ensure the accuracy and consistency of key-value pairing even when the field layout is loose or the position is variable. Attached Figure Description
[0020] Figure 1 This is a schematic diagram illustrating the workflow of an OCR recognition and field structuring method for complex layout documents according to the present invention. Figure 2 This is a schematic diagram illustrating the principle of imprint removal and repair based on generative networks in this invention. Figure 3 This is a schematic diagram of the spatial topology construction based on graph neural networks according to the present invention. Detailed Implementation
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0022] Example 1, please refer to Figures 1 to 3 This invention provides a technical solution: a method for OCR recognition and field structuring processing of complex typesetting documents, the specific steps of which are as follows: Step 1: Image enhancement and seal restoration based on generative adversarial networks (GANs). A residual network is used to locate the seal occlusion area and generate a seal mask. Unoccluded text strokes and background texture are used, while pixel-level reconstruction of the covered area is performed. Red noise from the seal is removed, and fragmented strokes are completed, outputting a high-resolution seal-free document image. A deep learning model accurately separates the seal and text regions, effectively eliminating red noise interference and providing clean image input for subsequent recognition, significantly improving the recognizability of text in occluded areas. The preprocessed document image is input into the residual network for multi-scale feature extraction. A feature pyramid structure captures texture and edge information under different receptive fields, and an attention mechanism is used to enhance the seal's unique color saturation and edge sharpness response, generating a seal mask that accurately identifies the occluded pixel areas. This is a binary mask, and multi-scale feature fusion accurately captures the seal's texture and edge responses. Features are used to provide accurate location of occluded areas for subsequent mask generation and pixel restoration. The original ticket image and the seal mask are input into the generator network. The seal mask is used to locate the area to be restored, and the stroke direction, grayscale gradient and background texture distribution features of the unoccluded area are extracted as reference features for pixel-level reconstruction. The prior features of the background and text are fully utilized to guide the restoration process, ensuring the authenticity of the reconstructed content. The generator can perform targeted pixel filling based on the surrounding context. Based on the extracted reference features, the generator performs pixel-by-pixel context-aware filling on the mask-covered area, simultaneously eliminating the color interference of the red seal and restoring the morphological continuity of the fragmented characters. The output is a ticket image without seal with complete text strokes and a clean background. While removing the color interference of the seal, the integrity and continuity of the text strokes are maintained, which significantly improves the readability of the text in the occluded area and provides high-quality input for subsequent recognition. Furthermore, the pixel-by-pixel context-aware filling process in step one involves: constructing a generator network containing multi-level dilated convolutions; extracting contextual texture information at multiple scales around the mask region in parallel using convolutional kernels with different dilation rates; and adaptively fusing features at each scale based on attention weights. This generates a combined feature vector containing global and local information for each pixel to be repaired, expanding the receptive field to capture long-distance contextual dependencies and achieving complementarity between global and local information. This provides a rich feature representation for the pixel generation module, improving the repair quality. The combined feature vector is then input into the pixel generation module, combining gradient continuity constraints at the mask boundary and the direction of character strokes. Consistency constraints are applied to synthesize fill content that seamlessly integrates with the surrounding background in terms of hue, brightness, and texture, pixel by pixel. The constraints ensure smooth transitions at the boundaries and consistency with the stroke direction, guaranteeing a natural blending of the repaired area with the background. This makes the reconstructed content visually indistinguishable from the original area. A multi-scale discriminator is used to apply adversarial constraints to the repair results, supervising the generator's fill quality from both global image semantics and local stroke details. This ensures that the repaired text strokes maintain an intrinsic consistency with the font style of the original document. Multi-scale supervision guarantees overall hue coordination and realistic local details, driving continuous optimization of the generator to ensure that the repaired strokes maintain stylistic consistency with the original document font. Specifically, the acquired grayscale images of tickets with a resolution of at least 300 DPI are scaled down to 512×512 pixels and input into a pre-trained ResNet-50 backbone network for initial feature extraction. A feature pyramid structure is then connected after the backbone network, fusing feature maps of different resolutions from Conv2 to Conv5 layers from top to bottom. These correspond to texture and edge responses with receptive fields of 8×8, 16×16, 32×32, and 64×64 pixels, respectively. Simultaneously, a channel attention module is introduced to enhance the seal sensitivity of the fused feature maps: targeting the seal area... The red channel saturation in the RGB space is concentrated in the range of 0.75 to 0.95, and the edge gradient magnitude is 2 to 3 times greater than that of the background area. The attention module assigns a response gain coefficient of 1.2 to 1.5 times to pixels that meet these statistical characteristics, while suppressing the response values of non-imprint areas. The enhanced feature map is then mapped pixel-by-pixel to the [0, 1] interval using the Sigmoid activation function, and an imprint mask is generated with a binarization threshold of 0.5. Pixels with a value of 1 in the mask precisely correspond to the imprint coverage area, while pixels with a value of 0 correspond to the background and text areas. The original ticket image and the generated binary mask image are concatenated along the channel dimension to form a 4-channel tensor, which is then input into the repair generator network. This generator adopts an encoder-decoder architecture. The encoder part consists of 5 downsampling modules, each containing a convolutional layer with a stride of 2 and a batch normalization layer, with the output feature map size halved step by step to 16×16. The decoder part has 5 corresponding upsampling modules, and skip connections are introduced between the same layers of the encoder and decoder to pass high-resolution detail features. Multi-level dilated convolutional modules are embedded in the last three layers of the decoder, and these modules are set with parallel dilation rates of 2, 3, 4, 5, 6, 7, 8, 9, 10 ... Three sets of 3×3 convolutional kernels (4 and 8) expand the receptive field to 13×13, 21×21, and 29×29 pixels respectively. Contextual information at different scales is extracted layer by layer from the mask boundary outwards. Then, learnable attention weights are used to adaptively weight and sum the features at each scale, generating a 256-dimensional combined feature vector for each pixel to be repaired. After receiving this feature vector, the pixel generation module passes it through two fully connected layers and outputs the RGB three-channel reconstructed value of the pixel. During the reconstruction process, the gradient change rate of pixels on both sides of the mask boundary is forcibly constrained to not exceed 1% of the average gradient change rate of the background region.The accuracy is doubled, and the angular deviation between the main direction of the repaired stroke and the direction of the stroke of the adjacent unoccluded text is controlled within ±5°. A dual-discriminator architecture containing a global discriminator and a local discriminator is constructed for adversarial training to ensure the overall realism of the repaired image and the accuracy of local strokes. The global discriminator receives the complete repaired image, downsamples it through 5 convolutional layers with a stride of 2, and outputs a 256-dimensional feature vector. Then, it outputs a global realism score through a fully connected layer to supervise the overall tonal distribution of the image and the naturalness of the background texture. The local discriminator crops a 96×96 pixel image block centered on the mask area, processes it through 4 convolutional layers, and outputs a local realism score to supervise the sharpness of the repaired text strokes and the consistency of the font style. Both discriminators adopt the WGAN-GP training strategy, with the gradient penalty coefficient set to 10 and the generator learning rate set to 2×10. -4 The discriminator learning rate is set to 1×10. -4 The Adam optimizer is used to update parameters. Four images are processed in each batch, and the output is an RGB three-channel image of a ticket without a seal. The pixel values of the seal-covered area are completely reconstructed by the generator, while the pixel values of the areas not covered by the mask in the original image remain unchanged. Step Two: An improved differentiable binarization algorithm is used for text detection, simultaneously outputting the coordinate information of horizontal and vertical text boxes. A convolutional recurrent neural network combined with a connection-time classification loss function is used for character recognition, outputting a raw data stream containing text content, confidence scores, and four-point coordinate information. A direction-aware detection mechanism is designed for mixed horizontal and vertical layouts to ensure complete detection of text from multiple directions. The text regions in the image are transformed into structured text records with coordinates and confidence scores. A direction-sensitive feature extraction branch is added to the differentiable binarization detection network. Multi-angle convolutional kernels are used to independently encode the axis alignment features of horizontal and vertical text, while simultaneously predicting the regression parameters of horizontal and vertical candidate boxes to output four-point coordinate detection results containing direction labels. By independently encoding horizontal and vertical layout features through multi-directional convolutional kernels, accurate direction-aware detection is achieved, providing targeted detection for text with different layout directions. To enhance the detection recall rate by expressing the characteristics of sexuality, different sequential reading strategies are adopted for horizontal and vertical text boxes based on the direction labels of the four-point coordinate detection results. This ensures that vertical text is serialized and spliced in the column direction, avoiding character order disorder caused by horizontal splicing. The reading order is dynamically adjusted according to the direction labels to ensure that characters are arranged in the correct reading direction, effectively solving the problem of character order inversion caused by horizontal splicing of vertical text. The serialized text image blocks are sequentially input into a convolutional recurrent neural network, and a bidirectional long short-term memory network is used to extract temporal dependency features. The text content and its confidence score are output through a connected temporal classification decoder, forming a multi-dimensional raw data stream containing text content, confidence score and four-point coordinate position information. Temporal modeling is used to capture the dependencies between characters, improve the accuracy of complex character recognition, and provide a complete text record with spatial coordinates and confidence for downstream structured processing. Specifically, the feature extraction branch contains two sets of independent and parallel convolutional channels. The first set uses horizontal strip convolutional kernels of sizes 3×3, 5×5, and 7×7 to enhance the response strength of horizontal text along the horizontal axis. The second set uses vertical strip convolutional kernels of the same size to enhance the response strength of vertical text along the vertical axis. After the feature map output, the two sets of channels are concatenated and fused, and then fed into the region proposal network to simultaneously predict the regression parameters of horizontal and vertical candidate boxes. The regression parameters include the horizontal and vertical coordinates of the candidate box center point, width, height, and rotation angle. The detection network finally outputs the four coordinates of each text box and adds a direction label to each text box. The direction label uses binary identifiers to distinguish between horizontal and vertical text, where horizontal... The four coordinates of the text boxes are arranged in the order of top left, top right, bottom right, and bottom left, while the four coordinates of the vertical text boxes are arranged in the order of top right, bottom right, bottom left, and top left, ensuring that the direction labels correspond to the coordinate arrangement order. Based on the discrimination results of the direction labels, a differentiated serialization reading strategy is applied to the detected text boxes. For text boxes labeled horizontally, pixel rows are extracted in a raster scan order from left to right and from top to bottom, and then concatenated to form a one-dimensional feature sequence. For text boxes labeled vertically, pixel columns are extracted in a column scan order from top to bottom and from right to left. After concatenating the pixels in each column end to end, the columns are then concatenated in a right-to-left order to form a complete one-dimensional feature sequence, ensuring that the characters in the vertical text are arranged in accordance with their coordinate order. The inherent reading direction is used for serialization, avoiding character order reversal or disorder caused by directly using horizontal splicing. During serialization, the relative position index of each character in the original text box is automatically recorded. The serialized text image blocks are adjusted to a uniform height of 32 pixels and the width maintains the original aspect ratio. They are then sequentially input into a convolutional recurrent neural network for feature extraction. The front end of the convolutional recurrent neural network consists of seven convolutional layers, each followed by a batch normalization layer and a ReLU activation function, outputting a 512-dimensional feature sequence. This feature sequence is then fed into a bidirectional long short-term memory network (LSTM). The bidirectional LSM is set to have 256 hidden units, is bidirectional, and has two layers, used to extract the feature sequence in both directions. The temporal dependency is addressed by linearly transforming the output of the bidirectional long short-term memory network and then inputting it into a temporal classification decoder. The decoder employs a dictionary-free, frame-by-frame decoding method, using a character set as the output space for path search. The character set includes hiragana, katakana, kanji, and special currency symbols, totaling approximately 3,000 character categories. The decoder outputs the posterior probability corresponding to each character category and folds the frame-level prediction results into the final character sequence using a prefix beam search algorithm. Simultaneously, it outputs the confidence score of the character sequence, which is the geometric mean of the posterior probabilities of each valid character. In the final raw data stream, each record contains a unique identifier for the text box, four-point horizontal and vertical coordinates, a direction label, the identified text content, and a confidence score. Step 3: The acquired text content is transformed into semantic vectors using a pre-trained language model. The normalized coordinates of the text boxes are mapped to positional encoding vectors and concatenated with image visual features to form a comprehensive feature representation. This integrates textual semantics, spatial location, and visual appearance information to construct a feature expression rich in contextual relationships, providing a multi-dimensional node attribute foundation for graph network inference. The output text content is then input into a language model pre-trained on a large-scale corpus to extract high-dimensional semantic embedding vectors corresponding to each text box. This captures the deep semantic commonalities of the same field under different expressions. The pre-trained language model transforms the text into semantic vectors, eliminating ambiguity caused by synonymous expressions and providing a semantically unified and highly generalizable text representation foundation for subsequent cross-modal fusion. The normalized centers of each text box are then... The coordinates and aspect ratios are mapped to a continuous space of the same dimension as the semantic embedding vector to form an absolute position encoding vector. At the same time, the relative displacement between text boxes is encoded as a relative position relationship vector. The spatial coordinates and relative displacement are encoded into a high-dimensional space to quantify the positional dependencies between fields, providing structured spatial position prior information for graph network construction. The semantic embedding vector, absolute position encoding vector, relative position relationship vector, and corresponding region visual feature vectors of the unsigned ticket image are cross-modal concatenated and linearly transformed to fuse and output a comprehensive feature representation matrix of each text box. Cross-modal fusion of text semantics, spatial position, and visual appearance constructs a multi-dimensional node feature representation, unifying heterogeneous information into the same feature space and providing information-rich initial node attributes for graph network inference. Specifically, the pre-trained language model used is the BERT model based on the Transformer encoder architecture, pre-trained on Wikipedia and news corpora containing Japanese corpora, with the total Japanese corpus being approximately 30GB. The maximum sequence length of the input text is set to 128 tokens; any exceeding this length is truncated, and any insufficient length is padded with [PAD] tokens. The model output is a 768-dimensional semantic embedding vector for each text box. This dimension has been validated through ablation experiments to achieve the optimal balance between semantic representation richness and computational efficiency. During the semantic embedding vector extraction process, the output corresponding to the [CLS] token in the last hidden state of the BERT model is taken as the sentence-level semantic representation of the entire text box. Semantic representation aggregates attention information from the entire text context, effectively capturing the semantic commonality of the same field under different expressions such as "received amount," "accepted amount," and "total." The absolute position encoding vector is generated using a sine and cosine position encoding function. The normalized x-coordinate, normalized y-coordinate, and aspect ratio w / h (the range of values is determined by statistically analyzing the aspect ratio distribution of all text boxes in the training set) of each text box's center point are mapped to subsets of each dimension in a 768-dimensional space. Specifically, the x-coordinate is allocated 256 dimensions, the y-coordinate 256 dimensions, and the aspect ratio 256 dimensions. These are then concatenated to form a complete 768-dimensional absolute position encoding vector. The relative position relationship vector is calculated separately for each pair of text boxes, encoding the difference in x-coordinates, ... The difference in ordinates, the Euclidean distance between the center points, and their orientation angles are also mapped to a 256-dimensional space. During the encoding process, sine functions of different frequencies are used to modulate the positional information of channels in different dimensions, ensuring that text boxes with similar spatial positions have similar encoded representations in the high-dimensional space. A fusion strategy based on a gated cross-modal attention mechanism is employed, adding the semantic embedding vector (768-dimensional) and the absolute position encoding vector (768-dimensional) element-wise to form a 1568-dimensional semantic-position joint representation. For the visual feature vector, affine transformation correction is performed on the four-point coordinates of the text detection boxes from the seal-free document image output in step one. Each text region is cropped and scaled to a uniform size of 48×48 pixels and input into the visual feature extraction network. The network consists of four convolutional layers, each with a kernel size of 3×3, a stride of 1, and 64, 128, 256, and 512 channels, respectively. Each layer is followed by a batch normalization layer and a ReLU activation function, which then outputs a 512-dimensional visual feature vector through global average pooling. The three types of feature vectors are concatenated along the channel dimension to form a 2048-dimensional original fusion vector. After a fully connected linear transformation, a 768-dimensional comprehensive feature representation is output. The fully connected layer uses a trainable weight matrix and bias term. The weight matrix has a size of 768×2048, the bias term has a size of 768, and the comprehensive feature representation matrix has a dimension of N×768, where N is the total number of text boxes in the current ticket. Each row in the matrix corresponds to the comprehensive multimodal features of one text box. Step 4: Treat each text box as a node, establish edge connections based on Euclidean distance and semantic similarity, construct a topology graph of the ticket content, and perform node classification (key / value discrimination) and edge relationship prediction (key-value pairing) on the graph structure. Get rid of the dependence on fixed coordinates, and model the spatial and semantic associations between fields through graph neural networks to achieve accurate identification of field roles, so as to have the ability to understand any format of ticket without the need for preset templates; Step 5: Perform logical post-processing and verification based on financial and tax rules. Use the built-in financial and tax logic engine to standardize the data (automatically convert the identified calendar dates to Gregorian calendar date format, and convert full-width numbers and special currency symbols to standard half-width numbers), and perform self-consistency verification. Introduce business rules in the financial and tax field to standardize and verify the recognition results, ensuring that the output data meets the format specifications and logical requirements of the financial system. Step Six: Based on the self-consistency verification results, determine whether to mark a prompt. Combine the graph network reasoning and rule verification results to generate structured data containing key-value pairs, table row and column relationships, and confidence scores. No preset template is required, and it can be directly adapted to any unfamiliar invoice. The output is a standardized JSON format that can be directly connected to the financial system. The output fields are organized hierarchically according to the confidence score and verification status, which facilitates downstream systems to process them according to priority. The output is in the form of standardized data messages, realizing seamless integration between the recognition results and the financial system.
[0023] Example 2, as Figures 1 to 3As shown, based on Embodiment 1, this invention provides a technical solution: Step four specifically includes: taking each text detection box in the ticket as an independent node in the topology graph, taking the comprehensive feature representation of each node as the initial node attribute, and selecting node pairs that meet the preset adjacency conditions to establish edge connections based on the Euclidean distance matrix and semantic vector cosine similarity matrix between all nodes. The graph edge connections are established using the dual constraints of spatial distance and semantic similarity to accurately locate the correlation between fields, providing a high-quality topology connection that reflects the inherent structure of the ticket for the graph network. The constructed node attribute matrix and adjacency matrix are input into a multi-layer graph convolutional network, and the feature representation of each node is iteratively updated through a neighborhood aggregation mechanism, so that each node... The system aggregates spatial location and semantic category information of neighboring nodes to enable full interaction between node features and their context. It iteratively fuses context information through neighborhood message passing to enhance the expressive richness of node features, enabling each node's features to perceive the distribution of surrounding fields and semantic environment. A linear classifier is applied to the updated node features, and based on the semantic expression pattern carried by each node and the spatial distribution characteristics of surrounding nodes, it determines whether each node belongs to the key field category or the value field category, outputting node-level key / value attribute labels. Semantic role determination is performed based on the fused context features to accurately distinguish field types, classifying nodes into key and value categories, thus establishing a clear role basis for subsequent key-value pairing. Specifically, the adjacency condition is established using a dual judgment mechanism of distance threshold and semantic similarity threshold: when the Euclidean distance between two nodes is less than 12% of the short side size of the ticket image (this threshold is determined after statistically analyzing the distribution of adjacent field spacing in the training set) and the cosine similarity of the semantic vectors is greater than 0.35, an undirected edge is established between the two nodes; for regions with sparse text box density, the distance threshold is relaxed to 18% of the short side size to ensure that isolated text boxes can establish necessary connections with the surrounding context, ensuring the connectivity of the graph structure. Edge weights are assigned based on the semantic similarity between nodes and the reciprocal of the distance, ranging from 0.2 to 0.9. The larger the weight value, the closer the semantic and spatial association between the two nodes. High-weight edges are given higher message passing priority in subsequent graph convolution iterations; the graph convolutional network adopts a three-layer stacked structure. The first graph convolutional layer receives N×768-dimensional node feature input, outputs N×512-dimensional features after neighborhood aggregation; the second layer outputs N... The third layer outputs N×128 dimensional features. In each graph convolution operation, the node's own features are fused with the weighted sum of the features of its neighboring nodes using an average aggregation method. During aggregation, the average of the features of all neighboring nodes in the neighborhood is taken and concatenated with the node's own features, followed by linear transformation and ReLU activation. Residual connections are used between the three layers, and the input and output features of each layer are added element-wise before being fed into the next layer, effectively alleviating the gradient vanishing problem in deep network training. After the three-layer graph convolution, the final 128-dimensional features of each node fuse its own semantics, spatial location, and contextual information of all nodes within its second-order neighborhood. The linear classifier uses a single-layer fully connected network with an input dimension of 128 dimensions and an output dimension of 2 dimensions, corresponding to the unnormalized logistic values of the key and value categories, respectively. After normalization by the Softmax function, the higher probability value is taken as the attribute label of the node. The cross-entropy loss function is used during classifier training, and the learning rate is set to 3×10. -4 Each batch processes 16 tickets, and the classifier front end is connected to the Dropout layer with a dropout rate of 0.3 to ensure classification accuracy and prevent overfitting to specific field expression patterns in the training set. It ensures that common key fields such as "received amount" and their corresponding numerical fields can be accurately distinguished. After the node attribute labels are output, the label results are written back to the feature records of each node, and the semantic category name corresponding to each key node and the numerical type of the value node are marked. Furthermore, step four also includes: aggregating all key node features and value node features after the graph convolutional network update into key feature sets and value feature sets, respectively; for each key-value node pair, concatenating their feature vectors and inputting them into the edge prediction multilayer perceptron; calculating the matching confidence score of the pairing relationship; quantifying the reliability of key-value matching through the pairing relationship score; filtering out low-confidence combinations; providing an accurate basic scoring basis for subsequent global matching; constructing a global assignment matrix based on the matching confidence scores of all key-value nodes; introducing bidirectional row and column constraints; and using the Hungarian algorithm to solve for the globally optimal one-to-one key-value pairing scheme. Simultaneously, the reliability confidence score corresponding to each matching edge is output. Global optimal matching ensures that each key-value field is paired only once, avoiding duplication and conflict, so that the pairing results are optimal overall, improving the accuracy of structured extraction. The paired key-value node pairs are clustered according to their spatial adjacency in the original ticket. The block type is automatically determined based on the distribution of node coordinates within each group. Field-level semantic units and table row and column structures are jointly parsed. Pairs are assigned to corresponding blocks through spatial clustering, and the table row and field list structures are identified to restore the block layout of the ticket, so that the output structured data retains the original page organization information. Specifically, the edge prediction multilayer perceptron employs a three-layer fully connected structure, with hidden layer dimensions set to 256, 128, and 64 dimensions respectively. ReLU activation is used between each layer, and a dropout layer with a dropout rate of 0.2 is added after the first hidden layer to prevent overfitting. The output layer of the multilayer perceptron is a single-neuron structure, outputting a matching confidence score in the range of 0 to 1 after being mapped by a sigmoid function. During the training phase, positive samples are correctly labeled key-value pairs from real tickets, while negative samples are randomly sampled incorrect pairs, maintaining a 1:1 ratio. A binary cross-entropy loss function is used to optimize the network parameters, and the learning rate is fixed at 1×10⁻⁶. -3During the inference phase, each key-value node to be paired independently passes through the multilayer perceptron to obtain its corresponding matching confidence score. The closer the score is to 1, the higher the probability that the two nodes form a correct key-value pair. The rows of the global assignment matrix correspond to all key nodes (number denoted as M), and the columns correspond to all value nodes (number denoted as K). The matrix elements are the matching confidence scores of the corresponding key-value pairs. Since the number of key nodes and value nodes in a ticket is usually unequal, the matrix is completed into an M×M square matrix (when M>K) or a K×K square matrix (when K>M). The completed positions are filled with a fixed value of 0.01, representing invalid pairings. The row and column constraints ensure that each key node matches at most one value node, and each value node matches at most one key node. The Hungarian algorithm solves the problem by maximizing the total global matching score. After obtaining the globally optimal one-to-one matching scheme, the virtual matches corresponding to the completed positions are removed, and only the matching edges between real nodes are retained. The original matching confidence score corresponding to each matching edge is used as the reliability of that matching relationship. The system outputs the confidence level of the key-value pairs. After pairing, the coordinates of the key nodes in each key-value pair are used as the spatial representative position of the pair. A density-based spatial clustering algorithm is used to group all pairs. The cluster radius is set to 8% of the short side size of the ticket image, and the minimum number of pairs in the neighborhood of the core point is set to 2. Key-value pairs in the same cluster are spatially adjacent to each other and together form an independent semantic block. For each cluster, the bounding box coordinates of all key nodes and value nodes in the group are extracted, and the aspect ratio of the bounding box is calculated: if the aspect ratio is greater than 2.5, the block is determined to be a horizontal table row structure, and each pair is arranged in the horizontal direction; if the aspect ratio is less than 0.6, it is determined to be a vertical field list structure, and each pair is stacked in the vertical direction; if it is between the two, it is determined to be an independent field block. After the block type is determined, each block is globally sorted according to the spatial arrangement order of the blocks (from top to bottom, from left to right), and then the block type label, the list of all key-value pairs contained in the block, and their spatial bounding box coordinates are output. Step five specifically includes: constructing a structured knowledge base containing calendar conversion rules, numerical standardization rules, and arithmetic operation constraint rules from the Japanese tax system; automatically activating the corresponding subset of verification rules based on the invoice type identifier to adapt to the differentiated business requirements of invoices in different industries; organizing tax rules by industry layer to achieve accurate activation and efficient matching of verification logic; providing the engine with a multi-industry, scalable rule support system to ensure business adaptability; and sending the original recognition content of each field in the key-value pairing results output in step four into the tax logic engine, performing format conversion between the calendar year and the Gregorian calendar based on the knowledge base, and converting full-width numbers, etc. Chinese numerals and special currency symbols are normalized into a unified half-width numerical representation, and non-standard dates and numerical formats are converted into a unified representation, eliminating ambiguity in measurement and year reckoning, ensuring data consistency and calculation accuracy in subsequent arithmetic verification, and performing cross-field self-consistency verification on the standardized numerical field group based on arithmetic operation constraint rules. Specifically, this includes comparing the consistency of the cumulative value of the detailed items within the same group with the value of the summary item, as well as verifying the balance of addition and subtraction operations between component items, performing logical verification on related numerical fields, verifying the balance relationship between detailed summaries and components, effectively identifying identification errors or abnormal invoices, and ensuring the audit-level accuracy of financial data. Specifically, the knowledge base is constructed using a layered storage architecture. The top layer stores general financial and tax rules applicable to all types of invoices. The middle layer is divided by industry category, specifically covering four main industry branches: supermarket retail, catering services, transportation, and hotel accommodation. Each branch stores a mapping table of account names and tax calculation rules specific to that industry. The bottom layer stores merchant-level customized rules, which are indexed and retrieved through the enterprise's unified social credit code or merchant registration number pre-printed in the invoice image. The entire knowledge base is in JSON format. The system is organized in schema format, with approximately 420 rule entries in total. These include 48 calendar conversion rules, 96 numerical standardization rules, and 276 arithmetic operation constraint rules. All rules are stored as executable logical expressions, facilitating direct loading and interpretation by the tax logic engine. When the recognition confidence level is below 0.75, a subset of alternative rules is automatically activated based on keyword matching results, and the selection criteria are noted in the output record. The tax logic engine runs resident in memory. Upon startup, it loads all rule entries and compiles them into an executable decision tree structure. Input fields are processed by type: date fields are formatted using a regular expression template library and then uniformly converted to the Gregorian calendar format for output. Conversion deviations are corrected by a built-in Gregorian calendar-Chinese calendar conversion table. Numeric fields undergo code point conversion from full-width to half-width numbers, semantic conversion from Chinese numerals to Arabic numerals, and currency symbol standardization. The standardization process preserves the original numbers. The values are rounded to two decimal places without any rounding. Arithmetic operation constraint rules are grouped and organized according to the semantic category of the fields. Interrelated fields within the same semantic block are grouped into the same verification group. The allocation of verification groups is determined by the block type label and spatial clustering group output in step four. Each verification group contains one target field and several source fields. The values of the source fields are combined and calculated according to the predefined operators (addition or subtraction) in the rules. The calculation results are compared with the values of the target fields. An absolute error tolerance threshold is set in the comparison process. That is, when the absolute value of the difference between the combined calculated value of the source fields and the value of the target field is less than or equal to one, the verification is considered to pass. The relative error tolerance threshold is set to five per thousand. The more lenient of the absolute error and relative error is used as the final judgment criterion. During the verification process, the original text before standardization and the value after standardization, as well as the operators used in the calculation process, are recorded for each source field involved in the calculation to form a complete calculation path log. Furthermore, step five also includes: loading the corresponding arithmetic logic expression template from the knowledge base according to the business type of the invoice; substituting the standardized numerical fields into the variable placeholders in the template; calculating the summation or operation result on the left side of the expression and the target reference value on the right side; loading the corresponding template according to the business type; substituting the standardized fields to perform automatic verification; providing differentiated arithmetic verification logic for different invoices; ensuring that the verification rules match the business scenario; setting relative error tolerance thresholds and absolute error tolerance thresholds; comparing the operation result with the target reference value item by item under the dual threshold constraints; if the absolute difference between the operation result and the target reference value does not exceed the absolute error tolerance threshold, or the relative error does not exceed the absolute error tolerance threshold, the operation is considered complete. If the relative error tolerance threshold is exceeded, the field group is deemed to have passed the verification. If both types of errors exceed the corresponding tolerance range, the field group is deemed to be a verification anomaly. This approach balances the sensitivity of verification for both large and small amounts of data, ensuring accuracy while avoiding misjudgment due to minor deviations, thus improving the rationality of verification. The field groups with verification anomalies are marked in multiple levels, recording the anomaly type, the difference quantification value, and the field identifiers involved. The marking information is also synchronously attached to the corresponding key-value pairing records for subsequent output stages for classification, display, and review guidance. The anomaly fields are marked in layers according to type and degree of deviation, and the anomaly details and related fields are fully recorded, providing reviewers with clear anomaly location and traceability information, thus improving the efficiency of manual intervention. Specifically, the financial and tax logic engine determines the business category of the current invoice based on the block type label output in step four, and locates the corresponding expression template file from the middle-level industry rule subset of the knowledge base. The templates are stored in JSON format, and each template contains a template identifier, applicable invoice type, version number, and a set of predefined arithmetic expression definitions. Each expression definition contains a list of variables on the left and the corresponding operator sequence, a target variable name on the right, and a description of the business meaning of the expression. After receiving the block type label, the engine scans all expression definitions in the template, matches and associates the field key names in the standardized field set within the current validation group with the template variable names. For fields that match successfully, their standardized values are arranged in the order of variable placeholders. When the left side of the expression is filled in, the engine performs arithmetic operations from left to right according to the operator sequence to obtain the result. Similarly, the target variable on the right side is located from the standardized field set, corresponding to the field value. The absolute error tolerance threshold is set to 1.00 yuan, based on the smallest unit of legal tender, the yuan. This means that the absolute error constraint is met when the absolute value difference is no greater than 1.00 yuan. The relative error tolerance threshold is set to 0.5%, meaning the ratio of the absolute value of the difference between the result and the target reference value to the target reference value does not exceed 0.005. The dual thresholds use an "OR" logic; if either the absolute error is ≤1.00 yuan or the relative error is ≤0.005, the field group is considered to be within the absolute error tolerance range. In terms of consistency, the engine meets the requirements. After completing the single-item comparison, it sequentially executes the above comparison process for all expressions within the verification group, accumulating and recording all abnormal items that do not meet the dual constraint conditions. During the comparison process, double-precision floating-point numbers are used to store intermediate calculation results. All abnormal items are temporarily stored in the engine's memory buffer in list form, awaiting subsequent multi-level marking processing. After completing the comparison of all arithmetic expressions, the financial and tax logic engine executes a multi-level marking process for each verification group determined to be abnormal. The first level marks the abnormality type, generating a corresponding four-digit abnormality type code based on the predefined type coding system in the expression template, where the first digit identifies the abnormality category. The second level records the difference quantification value, and the absolute difference between the engine's calculation result and the target reference value is recorded. The difference, rounded to two decimal places, is stored in the difference quantification value field in yuan. The third level records the field identifiers involved. The engine extracts the unique identifiers of all source and target fields involved in the calculation in the expression, and assembles them into a string array in the order of source field first and target field last. After completing the three-level marking, the engine serializes the marking information into a JSON object. It locates the corresponding entry in the key-value pairing record output in step four by indexing the field key name, and writes the marking object into the abnormal marking extended attribute field of the entry. The original recognition text and standardized value are still retained in the original fields of the pairing record for reviewers to check. After the marking information is successfully written, the temporary memory buffer for this verification is released, and the verification process of a single ticket is completed. Step six specifically includes: fusing and aligning the key-value pairing results, table row and column parsing results, and verification anomaly marker information output from step four; mapping each field to its corresponding data level path according to a pre-defined standardized data exchange architecture; integrating multi-source heterogeneous identification and verification results according to a unified data architecture to ensure accurate field placement; providing a clear and well-defined data input foundation for downstream systems; eliminating field ambiguity; and assigning a composite quality score to each key-value pair based on the confidence score of graph network inference and the results of tax rule verification. Finally, based on the score threshold, the output fields are divided into high-confidence fields and low-confidence fields. The system employs a three-tiered approach, encompassing segments, verification anomaly fields, and a comprehensive evaluation of data quality based on reasoning confidence and rule verification results. Data is categorized hierarchically according to reliability, giving output data a quality grading identifier. This facilitates differentiated processing by downstream systems based on priority. The output content is organized into three tiers, generating structured data messages in JSON format that include field key names, standardized values, confidence scores, verification status markers, and anomaly details. This directly adapts to the data import interface specifications of downstream financial systems. The hierarchical structured output messages, carrying complete quality identifiers and anomaly details, enable seamless integration of identification results with the financial system, reducing system integration and manual review costs. Specifically, during the fusion and alignment process, the key-value pairing records output from step four are read. Each record contains a unique identifier for the key node, a key semantic category name, a unique identifier for the value node, the value identification text content, a standardized numerical value or date, the value node numerical type, a pairing confidence score, and a block ownership number. Simultaneously, the verification anomaly marker information output from step five is loaded. Each anomaly marker contains a four-digit anomaly type code, a difference quantification value (unit: yuan, rounded to two decimal places), and an array of related field identifiers. Using the unique identifier for the key node as the primary key, a copy of the anomaly marker object is written to the anomaly marker extended attribute field of the corresponding key-value pairing record. Subsequently, based on the predefined field mapping dictionary in the standardized data exchange architecture, the key names of each field (such as "received amount," "tax amount," "total amount," "transaction date," etc.) are mapped one by one to the corresponding JSON data level path. After completion, each paired record obtains a complete data hierarchy path identifier. The fusion and alignment process adopts a transactional write method. After the write is completed, the uniqueness of the index primary key is checked to ensure that no duplicate records are written. For each key-value pair, the matching confidence score (i.e., the real value in the range of 0 to 1 output by the edge prediction multilayer perceptron) and the financial and tax rule verification result (pass or anomalous) are read respectively. The composite quality score is calculated as follows: For key-value pairs that pass the verification, the composite quality score is directly taken as the matching confidence score multiplied by a coefficient of 1.0; for key-value pairs that fail the verification, the composite quality score is taken as the matching confidence score multiplied by a coefficient of 0.7, and then a penalty item of 0.05 multiplied by the first digit of the anomalous type code is subtracted. The score is truncated to the range of 0 to 1. The score thresholds are divided as follows: 0.85 (inclusive) and above are high confidence fields, 0.60 (inclusive) to 0.85 are low confidence fields, and 0.Fields below 60 or with any validation anomalies are grouped into the validation anomaly field level. After all scoring calculations are completed, the level label is written into the composite quality score field of each paired record. The JSON output message is organized in the order of high-confidence fields, low-confidence fields, and validation anomaly fields. The root node of the message contains the invoice metadata (unique identifier for the invoice image, processing timestamp, and format type label) and arrays of the three levels of fields. The high-confidence field array contains all paired records that meet the scoring threshold. Output fields include field key names, normalized values (date in "YYYY-MM-DD" format), and settings. The system outputs a reliability score and a list of empty anomaly details. The low-confidence field array contains all paired records that meet the low threshold range. In addition to the above fields, the original identified text is output for verification and comparison. The validation anomaly field array contains all anomaly paired records, and the system outputs complete field key names, standardized values, confidence scores, validation status flags, and anomaly details (including anomaly type codes, difference quantification values, and the array of involved fields). Before output, all field values are UTF-8 encoded and escaped to ensure error-free transmission of special characters. The generated JSON message directly conforms to the data format and field naming conventions of the downstream financial system's data import interface.
[0024] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for OCR recognition and field structuring processing of complexly formatted documents, characterized in that, The specific steps are as follows: Step 1: Image enhancement and seal restoration based on generative adversarial network. The residual network is used to locate the seal occlusion area and generate a seal mask map. At the same time, the covered area is reconstructed at the pixel level to remove the red noise of the seal and complete the broken strokes, and output a seal-free ticket image. Step 2: Text detection is performed using an improved differentiable binarization algorithm, which outputs the coordinate information of the horizontal and vertical text boxes. Character recognition is then performed using a convolutional recurrent neural network combined with a connection-time classification loss function to output the original data stream. Step 3: The acquired text content is converted into semantic vectors through a pre-trained language model, the normalized coordinates of the text box are mapped into positional encoding vectors, and then concatenated with the image visual features to form a comprehensive feature representation; Step 4: Treat each text box as a node, establish edge connections based on Euclidean distance and semantic similarity, construct a topological graph of the ticket content, and perform node classification and edge relationship prediction on the graph structure, specifically including: Each text detection box in the ticket is used as an independent node in the topology graph. The comprehensive feature representation of each node is used as the initial node attribute. Based on the Euclidean distance matrix and semantic vector cosine similarity matrix between all nodes, edge connections are established for node pairs that meet the preset adjacency conditions. The constructed node attribute matrix and adjacency matrix are input into a multi-layer graph convolutional network. The feature representation of each node is iteratively updated through a neighborhood aggregation mechanism, so that each node gathers the spatial location information and semantic category information of its neighboring nodes. A linear classifier is applied to the updated node features. Based on the semantic representation pattern of each node and the spatial distribution features of surrounding nodes, the classifier determines whether each node belongs to the key field category or the value field category, and outputs the node-level key / value attribute label. Step four also includes: After updating the graph convolutional network, all key node features and value node features are aggregated into a key feature set and a value feature set, respectively. For each key-value node pair, their feature vectors are concatenated and input into the edge prediction multilayer perceptron to calculate the matching confidence score of the pairing relationship. A global assignment matrix is constructed based on the matching confidence scores of all key-value nodes, and a two-way row and column constraint is introduced. The Hungarian algorithm is used to solve the globally optimal one-to-one key-value pairing scheme, and the reliability confidence score corresponding to each matching edge is output. The paired key-value nodes are clustered according to their spatial adjacency in the original ticket, and the block type is automatically determined based on the node coordinate distribution within each group. The field-level semantic units and table row and column structures are jointly parsed. Step 5: Perform logical post-processing and verification based on financial and tax rules, standardize data using the built-in financial and tax logic engine, and perform self-consistency verification; Step 6: Determine whether to mark a hint based on the self-consistency verification results, and combine the graph network reasoning and rule verification results to generate structured data containing key-value pairs, table row and column relationships, and confidence scores.
2. The OCR recognition and field structuring method for complex layout documents according to claim 1, characterized in that: Step one specifically includes: The preprocessed ticket image is input into the residual network for multi-scale feature extraction. The texture and edge information under different receptive fields are captured through the feature pyramid structure. Based on the attention mechanism, the color saturation and edge sharpness of the seal are enhanced to generate a seal mask map that accurately identifies the occluded pixel area. It is a binary mask map. The original ticket image and the seal mask are input into the generator network. The seal mask is used to locate the area to be repaired, and the stroke direction, gray level gradient and background texture distribution features of the text in the unmasked area are extracted as reference features for pixel-level reconstruction. Based on the extracted reference features, the generator performs pixel-by-pixel context-aware filling on the mask-covered area, simultaneously eliminating the chromatic interference of the red seal and restoring the morphological continuity of the fragmented characters, outputting a seal-free ticket image with complete text strokes and a clean background.
3. The OCR recognition and field structuring method for complex layout documents according to claim 2, characterized in that: The process of performing pixel-by-pixel context-aware fill in step one is as follows: A generator network containing multi-level dilated convolutions is constructed. Convolutional kernels with different dilation rates are used to extract contextual texture information of multiple scales around the mask region in parallel. Based on attention weights, features of each scale are adaptively fused to generate a combined feature vector containing global and local information for each pixel to be repaired. The combined feature vector is input into the pixel generation module. Combined with the gradient continuity constraint at the mask boundary and the consistency constraint of the text stroke direction, the fill content is synthesized pixel by pixel, which is seamlessly connected with the surrounding background in terms of tone, brightness and texture. A multi-scale discriminator is used to apply adversarial constraints to the restoration results, supervising the generator's filling quality from both global image semantics and local stroke details levels, so that the restored text strokes maintain an intrinsic consistency with the font style of the original document.
4. The OCR recognition and field structuring method for complex layout documents according to claim 1, characterized in that: Step two specifically includes: A direction-sensitive feature extraction branch is added to the differentiable binarization detection network. The axis alignment features of horizontal and vertical text are independently encoded by multi-angle convolution kernels. At the same time, the regression parameters of horizontal and vertical candidate boxes are predicted to output the four-point coordinate detection results containing the direction label. Based on the direction labels of the four-point coordinate detection results, different sequential reading strategies are used for the horizontal and vertical text boxes to enable the vertical text to be serialized and spliced in the column direction. The serialized text image blocks are sequentially input into a convolutional recurrent neural network. A bidirectional long short-term memory network is used to extract temporal dependency features. The text content and its confidence score are then output through a connected temporal classification decoder, forming a multi-dimensional raw data stream containing text content, confidence score, and four-point coordinate information.
5. The OCR recognition and field structuring method for complex layout documents according to claim 1, characterized in that: Step three specifically includes: The output text content is input into a language model pre-trained on a large-scale corpus, and high-dimensional semantic embedding vectors corresponding to each text box are extracted to capture the deep semantic commonalities of the same field under different expression methods. The normalized center coordinates and aspect ratios of each text box are mapped to a continuous space of the same dimension as the semantic embedding vector to form an absolute position encoding vector. At the same time, the relative displacement between text boxes is encoded as a relative position relationship vector. The semantic embedding vector, absolute position encoding vector, relative position relationship vector, and corresponding region visual feature vector of the unsigned ticket image are cross-modal concatenated and linearly transformed to output a comprehensive feature representation matrix of each text box.
6. The OCR recognition and field structuring method for complex layout documents according to claim 1, characterized in that: Step five specifically includes: Construct a structured knowledge base that includes calendar conversion rules, numerical standardization rules, and arithmetic operation constraint rules in the fiscal and tax system, and automatically activate the corresponding subset of verification rules based on the invoice type identifier; The original recognition content of each field in the key-value pairing results output in step four is sent to the financial and tax logic engine. Based on the knowledge base, the engine performs format conversion between the calendar year and the Gregorian calendar year, and normalizes full-width numbers, Chinese numerals and special currency symbols into a unified half-width numerical representation. Based on arithmetic operation constraint rules, cross-field self-consistency verification is performed on the standardized numerical field group. Specifically, this includes a consistency comparison between the cumulative value of the detailed items and the value of the summary item within the same group, as well as a balance verification of addition and subtraction operations between component items.
7. The OCR recognition and field structuring method for complex layout documents according to claim 6, characterized in that: Step five also includes: Based on the business type of the invoice, the corresponding arithmetic logic expression template is loaded from the knowledge base. The standardized numerical fields are substituted into the variable placeholders in the template, and the cumulative or operation result on the left side of the expression and the target reference value on the right side are calculated respectively. Set relative error tolerance thresholds and absolute error tolerance thresholds, and compare the calculation results with the target reference value item by item under the dual threshold constraints. If the absolute difference between the calculation result and the target reference value does not exceed the absolute error tolerance threshold, or the relative error does not exceed the relative error tolerance threshold, then the field is determined to pass the verification; if both errors exceed the corresponding tolerance range, then the verification is determined to be abnormal. The fields that are checked for anomalies are marked in multiple levels, and the anomaly type, the difference quantification value and the field identifiers involved are recorded respectively. The marking information is then synchronously attached to the corresponding key-value pair records.
8. The OCR recognition and field structuring method for complex layout documents according to claim 7, characterized in that: Step six specifically includes: The key-value pairing results, table row and column parsing results, and verification anomaly marker information output in step four are merged and aligned, and each field is mapped to the corresponding data level path according to the preset standardized data exchange architecture. Combining the confidence scores of graph network reasoning with the results of tax rule verification, a composite quality score is assigned to each key-value pair, and the output fields are divided into three levels: high confidence fields, low confidence fields, and verification anomaly fields according to the scoring threshold. The output content is organized into three levels, and a structured data message containing field key names, standardized values, confidence scores, verification status markers, and exception details is generated in JSON format.
Citation Information
Patent Citations
Lightweight bill OCR (Optical Character Recognition) method and system
CN116612479A
Method for extracting structured information of bill and electronic equipment
CN114694158A
Financial bill automatic identification generation and decision-making method and system
CN120472483A