Finance and accounting document data summarization method based on OCR correction by artificial intelligence

Through the financial perception self-correction model based on artificial intelligence, multi-dimensional quality parameters evaluation and dynamic correction of accounting documents are solved, which solves the problem of low accuracy in traditional OCR when processing poor quality documents, and realizes efficient electronicization of accounting data.

CN120375385APending Publication Date: 2025-07-25BEIHAI FORECASTING CENT OF STATE OCEANIC ADMINISTRATION ((QINGDAO MARINE FORECASTING STATION OF STATE OCEANIC ADMINISTRATION) (QINGDAO MARINE ENVIRONMENT MONITORING CENT OF STATE OCEANIC ADMINISTRATION))
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510796572.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When processing accounting documents of poor quality, the OCR identification accuracy is low, and the adaptive adjustment ability is lacking, and differentiated processing is not possible for differentiated areas, resulting in many identification errors and low efficiency.

Method used

The financial perception self-correction model based on artificial intelligence is adopted, and the gating function is deeply corrected through uniform light compensation, multi-dimensional quality parameter evaluation and quality impact gating function, combined with the financial field knowledge graph, to achieve accurate correction of accounting document images.

Benefits of technology

It significantly improves the accuracy of accounting documents identification under complex quality conditions, reduces manual intervention, and improves the efficiency and accuracy of electronic accounting data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375385A_ABST
    Figure CN120375385A_ABST
Patent Text Reader

Abstract

The invention provides a financial and accounting document data summarization method based on artificial intelligence correction OCR, and belongs to the technical field of financial and accounting document data summarization, and the method comprises the steps: firstly carrying out the illumination equalization processing, calculating the parameters such as wrinkle degree and stain readability, carrying out the preprocessing of a document image, and then carrying out the preliminary recognition through an OCR engine, and obtaining original text data; the core innovation lies in that a financial perception self-correction model based on a Vision Transform architecture is constructed, the model integrates visual and text features through a double-flow attention network, and a quality influence gating function is introduced to dynamically adjust attention distribution, so that differential correction of different quality regions is realized. And the scanning quality compensation function and the financial domain knowledge graph verification are combined, so that the accuracy and efficiency of financial document data identification under the complex quality condition are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of financial document data summarization. Specifically, it relates to a method for summarizing financial document data based on artificial intelligence to correct OCR. Background Art

[0002] The electronic processing of financial document data is the basis of modern enterprise financial management. Traditional technologies mainly convert document images into editable digital texts through optical character recognition (OCR). These systems usually rely on general OCR engines for text recognition and combine some basic image preprocessing technologies, such as binarization, skew correction, etc., to improve the recognition accuracy. In actual applications, financial documents often have quality problems such as wrinkles, stains, uneven illumination, etc. after multiple transmissions and storage processes. However, traditional OCR technologies perform poorly when dealing with financial documents under such non-ideal conditions. The existing technologies mainly use image enhancement algorithms with fixed parameters to preprocess the documents, lacking the ability of adaptive adjustment for different quality problems. At the same time, the correction of recognition errors mainly relies on general text proofreading algorithms or manual intervention, and fails to make full use of professional knowledge in the financial field for targeted error correction, resulting in a large number of recognition errors that need to be corrected manually. Especially when financial documents have multiple quality problems such as wrinkles, stains, and uneven illumination simultaneously, the existing technologies are difficult to accurately evaluate the influence degree of each quality factor on the OCR recognition result, lack a differential processing mechanism for different quality regions, and cannot achieve precise correction of the text, ultimately leading to low accuracy and inefficiency in the process of electronic financial data. That is to say, there is a technical problem in the existing technologies that the poor quality of financial document images leads to low OCR recognition accuracy. Summary of the Invention

[0003] In view of this, the present invention provides a method for summarizing financial document data based on artificial intelligence to correct OCR, which can solve the technical problem in the existing technologies that the poor quality of financial document images leads to low OCR recognition accuracy.

[0004] The present invention is implemented as follows: The present invention provides a method for summarizing financial document data based on artificial intelligence to correct OCR, including: collecting financial document images and performing preprocessing through a light uniformity compensation function, and simultaneously calculating the wrinkle degree parameter and the stain readability parameter; inputting the preprocessed financial document images into an OCR engine for preliminary recognition to obtain the original recognition text data and the recognition confidence matrix; using a financial perception self-correction model to correct the original recognition text data, the financial perception self-correction model includes a visual feature extraction stream and a text feature extraction stream, and performing information integration through an adaptive fusion module; using a scanning quality compensation function to calculate the degree of influence of each character region in the original recognition text data on the quality of the financial document image, and generating a quality influence weight vector; analyzing the deformation influence of the financial document image according to the wrinkle change degree parameter and the wrinkle jump degree parameter, and adjusting the attention distribution parameter through a quality influence gating function; combining the quality influence weight vector with the corrected text data for multi-layer correction, and simultaneously verifying the data rationality using a financial domain knowledge graph; performing structured integration and summarization on the corrected text data to generate a standard financial data format for output.

[0005] Among them, the light uniformity compensation function refers to a mathematical model that analyzes the brightness distribution histogram of the financial document image, calculates the degree of uneven illumination in each region of the financial document image, and adjusts the pixel value distribution through an adaptive gamma correction and a multi-scale histogram equalization algorithm to make the overall illumination of the financial document image tend to be uniform.

[0006] Among them, the scanning quality compensation function refers to a function model that calculates the quality level of the financial document image based on an image sharpness evaluation index and assigns different credibility weight coefficients to the original recognition text data for different quality levels.

[0007] Among them, the wrinkle degree parameter refers to the total amount of wrinkle features in the financial document image. By detecting the linear edge density and direction discontinuity in the financial document image, and combining high-frequency filtering to extract texture abnormal regions, it is quantitatively represented as an integer value from 0 to 100.

[0008] Among them, the wrinkle change degree parameter refers to the degree of drastic change in the wrinkle shape in the financial document image. By calculating the included angle between adjacent region wrinkle feature vectors and the wrinkle line density gradient, it quantitatively represents the complexity of the deformation of the financial document image.

[0009] Among them, the wrinkle jump degree parameter refers to the degree of discontinuity in the spatial distribution of wrinkle features in the financial document image. By detecting the number of breakpoints of wrinkle line segments and the jump frequency of the wrinkle region boundary, it reflects the degree of influence of local deformation of the financial document image on text recognition.

[0010] Among them, the stain readability parameter refers to the degree to which the text information in the area covered by the stain in the financial document image can still be recognized, and is calculated by analyzing features such as the integrity of the text strokes, contrast, and edge sharpness in the stain area.

[0011] Among them, the specific structure of the financial perception self-correction model is a two-stream attention network based on the Vision Transformer architecture. The visual feature extraction stream adopts a resolution adaptive blockization mechanism, and dynamically adjusts the block size parameter according to the wrinkle degree parameter. The text feature extraction stream adopts a hierarchical attention network structure, which can process character-level features, word-level features, and sentence-level features simultaneously.

[0012] Among them, the information of the visual feature extraction stream and the text feature extraction stream is interactively fused through the multi-head self-attention mechanism. During the fusion process, financial domain knowledge constraints are introduced, and the original information is retained through the skip connection mechanism and financial term representations are introduced.

[0013] Among them, the core of the financial perception self-correction model is the adaptive gating attention mechanism, which realizes selective extraction of information in different quality regions through dynamically adjusted attention distribution parameters. The attention distribution parameters are controlled by the quality impact gating function.

[0014] The present invention deeply corrects the preliminary OCR recognition result by constructing a financial perception self-correction model. This method first calculates multi-dimensional quality parameters such as the wrinkle degree and stain readability of the document through refined image analysis, and uses the light uniform compensation function to achieve light equalization processing, significantly improving the preprocessing effect. In the core correction link, the financial perception self-correction model of the present invention can receive the original recognition text data and multi-dimensional quality parameters, adaptively adjust the attention distribution mechanism, and implement differential correction strategies for different quality regions. By dynamically controlling the attention distribution parameters of the model through the quality impact gating function, the present invention can accurately identify the text information in the wrinkled deformation and stain-covered areas, greatly improving the recognition accuracy under complex quality conditions. The present invention solves the problem of low accuracy of traditional technologies in processing financial documents with poor quality, realizes precise evaluation and targeted correction of multiple quality problems in financial document images, and makes the OCR recognition result more in line with financial data specifications by introducing financial domain knowledge constraints and an adaptive quality compensation mechanism, reducing the need for manual intervention and improving the efficiency and accuracy of financial data digitization. Description of the Drawings

[0015] Figure 1 It is a flowchart of the method of the present invention.

[0016] Figure 2 It is a schematic structural diagram of the financial perception self-correction model related to the present invention.

[0017] Figure 3 Schematic diagram of the quality impact gating function mechanism related to the present invention. Detailed implementation manners

[0018] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0019] As Figure 1 shown, it is a flowchart of a method for summarizing financial document data based on artificial intelligence to correct OCR provided by the present invention. The method includes the following steps:

[0020] S01. Collect a financial document image and perform preprocessing. Perform illumination equalization processing on the financial document image through a light uniformity compensation function, and calculate the wrinkle degree parameter and stain readability parameter of the financial document image at the same time;

[0021] S02. Input the preprocessed financial document image into an OCR engine for preliminary recognition to obtain original recognition text data and a corresponding recognition confidence matrix;

[0022] S03. Use a financial perception self-correction model to correct the original recognition text data. The financial perception self-correction model receives the original recognition text data, the recognition confidence matrix, the wrinkle degree parameter and the stain readability parameter, and outputs corrected text data and a credibility index;

[0023] S04. Use a scanning quality compensation function to calculate the degree of influence of the quality of the financial document image on each character area in the original recognition text data, and generate a quality impact weight vector;

[0024] S05. Analyze the influence of the deformation of the financial document image on the original recognition text data according to the wrinkle change degree parameter and the wrinkle jump degree parameter, and adjust the attention distribution parameter in the financial perception self-correction model through a quality impact gating function;

[0025] S06. Combine the quality impact weight vector with the corrected text data to perform multi-layer correction on the original recognition text data, and verify the data rationality of the corrected text data by using a financial domain knowledge graph;

[0026] S07. Structurally integrate and summarize the corrected text data to generate a standard financial data format for output, and continuously optimize the parameters of the financial perception self-correction model through regression testing.

[0027] Among them, the light uniform compensation function refers to a mathematical model that analyzes the brightness distribution histogram of the financial document image, calculates the uneven illumination degree of each area of the financial document image, and adjusts the pixel value distribution through the adaptive gamma correction and multi-scale histogram equalization algorithms to make the overall illumination of the financial document image tend to be uniform; the input of the light uniform compensation function includes the original pixel data of the financial document image and the statistical values of the brightness distribution of each area, and the output of the light uniform compensation function is the financial document image after the illumination equalization process.

[0028] Among them, the scanning quality compensation function refers to a function model that calculates the quality level of the financial document image based on the image sharpness evaluation index and assigns different credibility weight coefficients to the original recognition text data for different quality levels; the input of the scanning quality compensation function includes multi-dimensional parameters such as the variance of the gradient amplitude, the frequency domain energy distribution, and the texture features of the financial document image, and the output of the scanning quality compensation function is the quality influence weight vector, which is used to adjust the credibility weights of each character in the original recognition text data.

[0029] Among them, the wrinkle degree parameter refers to the total amount of wrinkle features in the financial document image. By detecting the linear edge density and direction discontinuity in the financial document image, and combining high-frequency filtering to extract the texture abnormal area, it is quantitatively represented as an integer value from 0 to 100; the wrinkle degree parameter is calculated by the preprocessing in step S01 and is used to adjust the block size parameter of the financial perception self-correction model in step S03.

[0030] Among them, the wrinkle change degree parameter refers to the degree of drastic change in the wrinkle shape in the financial document image. By calculating the included angle of the wrinkle feature vectors in adjacent areas and the gradient of the wrinkle line density, it quantitatively represents the deformation complexity of the financial document image; the wrinkle change degree parameter is calculated based on the financial document image during the preprocessing in step S01 and is used as the input parameter of the quality influence gating function in step S05.

[0031] Among them, the wrinkle jump degree parameter refers to the degree of discontinuity in the spatial distribution of the wrinkle features in the financial document image. By detecting the number of break points of the wrinkle line segments and the jump frequency of the wrinkle area boundary, it reflects the influence degree of the local deformation of the financial document image on the character recognition; the wrinkle jump degree parameter is calculated based on the financial document image during the preprocessing in step S01 and is used as the input parameter of the quality influence gating function in step S05.

[0032] Among them, the stain readability parameter refers to the degree to which the text information in the area covered by the stain in the financial document image can still be recognized, and is calculated by analyzing features such as the integrity of text strokes, contrast, and edge sharpness in the stain area; the stain readability parameter is calculated based on the financial document image during the preprocessing process in step S01, and is used as an input parameter for the financial perception self-correction model in step S03 to adjust the hierarchical feature fusion weights in the financial perception self-correction model.

[0033] As Figure 2 shown, the specific structure of the financial perception self-correction model is a two-stream attention network based on the Vision Transformer architecture, including a visual feature extraction stream and a text feature extraction stream, and the visual feature extraction stream and the text feature extraction stream perform information integration through an adaptive fusion module; the visual feature extraction stream adopts a resolution adaptive blockization mechanism, and dynamically adjusts the block size parameter according to the wrinkle degree parameter. When the wrinkle degree parameter is high, a finer-grained block is used to retain detailed features; the text feature extraction stream adopts a hierarchical attention network structure, which can simultaneously process character-level features, word-level features, and sentence-level features, and adjusts the fusion weights of features at each level according to the stain readability parameter; the information of the visual feature extraction stream and the text feature extraction stream is interactively fused through a multi-head self-attention mechanism, and financial domain knowledge constraints are introduced during the fusion process, and the original information is retained through a skip connection mechanism and financial term representations are introduced; the core of the financial perception self-correction model is an adaptive gating attention mechanism, which realizes selective extraction of information in different quality regions through dynamically adjusted attention distribution parameters, and the attention distribution parameters are controlled by the quality impact gating function.

[0034] Among them, the steps for establishing the training data set of the financial perception self-correction model specifically include collecting multi-type financial document image samples from the real enterprise financial system, classifying and labeling different types of documents to ensure that the data set includes common financial document types such as VAT invoices, general invoices, receipts, and bank statements; grading the quality of the collected document images and constructing a document sample library with different quality levels, including high-quality scanned samples, low-quality scanned samples, wrinkled samples, stained samples, and samples with multiple problems; performing field structure annotation for each type of document, and the annotation content includes document type, field position, field content, and field relationship; using data augmentation technology to artificially construct various quality problem samples, including adding different degrees of uneven illumination, wrinkles, stains, and blurring effects, and calculating the corresponding illumination uniformity parameter, the wrinkle degree parameter, the wrinkle change degree parameter, the wrinkle jump degree parameter, and the stain readability parameter for each constructed sample; establishing a knowledge base of the logical relationship between document fields as a constraint condition for model training; performing preliminary OCR recognition on all samples and recording the differences between the recognition results and the true annotation content as the training objective of the financial perception self-correction model; constructing a balanced training data batch according to the document type, the illumination uniformity parameter, the wrinkle degree parameter, the wrinkle change degree parameter, the wrinkle jump degree parameter, the stain readability parameter, and the recognition error type to ensure that the financial perception self-correction model has sufficient learning for various problem scenarios.

[0035] Among them, the steps of training the financial perception self-correction model specifically include adopting a two-stage training strategy of pre-training and fine-tuning. First, pre-train the basic parameters of the financial perception self-correction model on a large-scale general text recognition dataset to enable the financial perception self-correction model to have basic text recognition capabilities; then use a partial financial document dataset to be summarized to fine-tune the financial perception self-correction model to make the financial perception self-correction model adapt to the text format and field relationships in the financial field; in the training process, adopt a multi-task joint learning framework to optimize three objective functions of text recognition accuracy, field localization accuracy, and document relationship reasoning accuracy at the same time; introduce an adversarial learning mechanism to enhance the robustness of the financial perception self-correction model to various quality problems by constructing adversarial samples; for samples with different illumination uniformity parameters, wrinkle degree parameters, wrinkle change degree parameters, wrinkle jump degree parameters, and stain readability parameters, adopt a dynamic weight adjustment mechanism to assign higher training weights to difficult samples; in the later stage of training, introduce financial domain knowledge distillation to integrate financial expert knowledge into the parameters of the financial perception self-correction model; adopt a gradient accumulation strategy to process large-batch data training to ensure that the financial perception self-correction model fully learns various document styles; prevent overfitting through periodic learning rate adjustment and early stopping strategies; after training, perform quantization and pruning optimization on the financial perception self-correction model to reduce the computational overhead during the inference of the financial perception self-correction model.

[0036] Such as Figure 3As shown, the quality impact gating function refers to a multi-parameter based dynamic adjustment mechanism for controlling the dynamic changes of the attention distribution parameters in the financial perception self-correcting model; the input of the quality impact gating function includes five quality parameters: the illumination uniformity parameter, the wrinkling degree parameter, the wrinkling change degree parameter, the wrinkling jump degree parameter, and the stain readability parameter; the calculation process of the quality impact gating function first calculates the comprehensive quality balance value, and the calculation formula of the comprehensive quality balance value is the weighted average of the illumination uniformity parameter, the wrinkling degree parameter, the wrinkling change degree parameter, the wrinkling jump degree parameter, and the stain readability parameter, and the weight coefficients are determined through historical data analysis; then different adjustment strategies are selected according to the range of the comprehensive quality balance value. When the comprehensive quality balance value is in the high-quality interval, a relatively conservative adjustment strategy is adopted, focusing on retaining the original recognition text data; when the comprehensive quality balance value is in the medium-quality interval, a moderate adjustment strategy is adopted, and the low-confidence regions are mainly corrected on the basis of retaining the original recognition text data; when the comprehensive quality balance value is in the low-quality interval, an aggressive adjustment strategy is adopted, greatly enhancing the attention of the attention mechanism of the financial perception self-correcting model to the low-quality regions; the quality impact gating function realizes the precise control of the behavior of the financial perception self-correcting model by dynamically adjusting three key parameters: the number of attention heads, the attention distribution parameters, and the feature fusion weights; the output of the quality impact gating function is the adjusted attention distribution parameters, which are used to control the behavior of the adaptive gating attention mechanism in the financial perception self-correcting model.

[0037] Among them, the illumination uniformity parameter refers to an index for measuring the uniform degree of the illumination distribution in each region of the financial accounting document image, and is obtained by calculating the brightness standard deviation of different regions of the financial accounting document image; the illumination uniformity parameter is calculated based on the financial accounting document image by the preprocessing process in step S01 and is used as an input parameter of the quality impact gating function in step S05.

[0038] The specific implementation manners of the above steps are described in detail below. The specific implementation manner of step S01 is to comprehensively preprocess the collected financial document images. First, the image is subjected to illumination equalization processing through a light uniformity compensation function. This function is based on the adaptive histogram equalization algorithm, and the image is processed in blocks, with each block having a size of 32×32 pixels, and bilinear interpolation is used for transition between adjacent blocks. The luminance distribution histograms are calculated separately for the RGB channels of the image. When the ratio of the histogram peak to the mean exceeds 3.5, it is determined that there is uneven illumination. When using the adaptive gamma correction algorithm, the gamma value γ is dynamically calculated according to the regional luminance mean. In the dark area, the γ value in the range of 0.5 to 0.8 is used, and in the bright area, the γ value in the range of 1.2 to 1.5 is used. At the same time, the wrinkle degree parameter is calculated. The linear edge features are extracted through the Canny edge detection algorithm, and the double thresholds are set to 50 and 150 respectively. The wrinkle degree parameter is quantified as an integer value from 0 to 100, and is calculated based on the detected linear edge density and direction discontinuity. In addition, the stain readability parameter is calculated. The stain area is identified through the local contrast enhancement and adaptive threshold segmentation algorithms, and then the integrity, contrast, and edge sharpness features of the text strokes in the stain area are analyzed. The purpose of this step is to improve the image quality, provide a clearer image input for subsequent OCR recognition, and at the same time obtain image quality parameters for subsequent processing.

[0039] The specific implementation manner of step S02 is to input the preprocessed financial document images into the OCR engine for preliminary recognition. In this step, the deep learning model CRNN (Convolutional Recurrent Neural Network) is used as the basic recognition engine, and the recognition effect is optimized by combining the attention mechanism. The image first extracts features through the VGG16 network, then is input into the bidirectional LSTM network for sequence modeling, and finally the text sequence and the recognition confidence of each character are output through the CTC (Connectionist Temporal Classification) loss function. For Chinese character recognition, a dictionary of 5000 common Chinese characters is used as the recognition basis, and an additional dictionary is established for optimization of financial special vocabulary. The confidence matrix is in units of characters, recording the recognition credibility of each character. The confidence threshold is set to 0.75, and the recognition results below this threshold are marked as suspicious areas. The role of this step is to obtain the original recognition text data and the corresponding recognition confidence matrix, providing the basic data and confidence basis for subsequent correction.

[0040] The specific implementation of step S03 is to use a financial perception self-correction model to correct the original recognized text data. This model receives the original recognized text data, recognition confidence matrix, wrinkle degree parameter, and stain readability parameter as inputs. First, the model encodes the original recognized text data into character-level feature vectors, with the feature vector dimension of each character being 128. Then, different weights are assigned to the feature vectors according to the recognition confidence matrix. When the recognition confidence is lower than 0.6, the model increases the constraint weight of the financial domain knowledge graph to 1.5 times the original weight; when the recognition confidence is between 0.6 and 0.75, the constraint weight of the financial domain knowledge graph is increased to 1.2 times the original weight. When the wrinkle degree parameter value exceeds 70, the model dynamically adjusts the block size, changing from the default 16×16 to 8×8 to more precisely capture local features. When the stain readability parameter is lower than 30, the model strengthens the weight of the visual feature extraction stream, increasing it to 1.8 times the original weight. By fusing multi-source information, corrected text data is generated, and a credibility index is calculated, which comprehensively considers the original confidence and the confidence change during the correction process. The purpose of this step is to use financial domain knowledge to perform preliminary intelligent correction on the original recognition result and improve the text recognition accuracy.

[0041] The specific implementation of step S04 is to use a scanning quality compensation function to calculate the degree of influence of the quality of the accounting document image on each character region in the original recognized text data. This function uses a multi-scale image quality assessment method, comprehensively considering the variance of the image gradient amplitude, frequency domain energy distribution, and texture features. The image is decomposed at multiple scales through wavelet transform, and the decomposition level is set to 3 layers to extract the energy distribution features of the high-frequency sub-bands. The standard deviation of the image gradient amplitude is calculated, and when the standard deviation is less than 15, the region is determined to be a low-definition region. The texture features are calculated through the gray-level co-occurrence matrix, including four statistics: energy, contrast, correlation, and entropy. Based on the above features, a quality influence weight vector is calculated, and the weight value range is 0 to 1. The smaller the value, the more serious the quality influence. When the standard deviation of the gradient amplitude is lower than 10 or the high-frequency energy ratio is lower than 0.15, the weight of the corresponding region is set below 0.3; when the Shannon entropy value of the local region of the image is lower than 4.5, the weight of the corresponding region is set below 0.5. The purpose of this step is to quantify the degree of influence of image quality on the recognition result and provide a targeted quality compensation basis for subsequent correction.

[0042] The specific implementation of step S05 is to analyze the impact of the deformation of the financial document image on the original recognized text data according to the fold parameters. First, the fold change degree parameter is analyzed. This parameter is obtained by calculating the included angle of the fold feature vectors in adjacent regions and the fold line density gradient. When the included angle of the feature vectors in adjacent regions exceeds 45 degrees, it is regarded as a high change degree region. The fold jump degree parameter is calculated by detecting the number of breakpoints of the fold line segments and the jump frequency of the fold region boundary. When the number of breakpoints per unit area exceeds 5, it is determined as a high jump degree region. Based on these parameters, the attention distribution parameters in the financial perception self-correction model are adjusted through the quality impact gating function. When the fold change degree parameter exceeds 75, the number of attention heads increases from the default 8 to 12; when the fold jump degree parameter exceeds 60, the adaptive gating attention threshold decreases from the default 0.65 to 0.45, increasing the model's attention to low-quality regions. This step aims to dynamically adjust the correction model parameters according to the document deformation characteristics and improve the accuracy of text recognition in the fold region.

[0043] The specific implementation of step S06 is to combine the quality impact weight vector with the corrected text data and perform multi-layer correction on the original recognized text data. First, a hierarchical correction framework is constructed, including three levels: character-level correction, word-level correction, and sentence-level correction. Character-level correction is based on the quality impact weight vector and the corrected text data. When the weight is lower than 0.4, the corrected result is preferentially adopted; otherwise, the original result is retained. Word-level correction uses the N-gram model for context analysis. The 3-gram model is used to analyze the collocation probability of words. When the collocation probability is lower than 0.25, correction is triggered. Sentence-level correction uses the financial domain knowledge graph to verify the rationality of the corrected text data. The knowledge graph contains 5,000 financial core concepts and 30,000 concept relationships. The recognized text is matched with the knowledge graph entities through the entity alignment algorithm, and the matching threshold is set to 0.7. When there is a logical conflict between the recognition result and the knowledge graph, the maximum likelihood estimation method is used to select the most reasonable correction scheme. The purpose of this step is to maximize the accuracy of text correction through a multi-level correction strategy combined with domain knowledge constraints.

[0044] The specific implementation of step S07 is to structurally integrate and summarize the corrected text data. First, a spatial relationship model of the data is established based on the document layout analysis algorithm. The YOLO v5 object detection network is used to identify the keyword field areas, and the detection confidence threshold is set to 0.8. The recognized text is logically grouped by the hierarchical clustering algorithm, and the clustering distance threshold is set to 15 pixels. According to the pre-defined financial document template library, the recognized text is mapped to the standard fields, and the template matching similarity threshold is set to 0.85. A semantic consistency check is performed on the recognized data to verify whether the calculation relationship of the digital fields conforms to the financial rules, and the allowable error range is 0.01. Next, cross-document data summarization processing is performed. First, an index tree structure is established according to the document type, and the same type of documents are organized into a tree structure for quick retrieval, and the tree depth is set to 3 layers. The adaptive data fusion algorithm is used to intelligently merge the same type of data from different documents, and the fusion granularity is divided into three levels: document level, field level, and numerical level. When there are multiple related documents in the same business scenario, a conflict resolution strategy is set, and the data selection is determined according to the document priority and timestamp. The priority configuration matrix is preset according to the importance of the documents. The sliding time window technology is used to achieve dynamic data summarization, and the window size can be configured as daily, weekly, monthly, quarterly, and annually, with the default setting being monthly summarization. For the data that needs to be summarized by subject, multi-dimensional aggregation calculations are implemented, and automatic balance verification of the debit and credit directions is supported, and the balance difference threshold is set to 0.001. The system can automatically generate standard financial statements including balance sheets, income statements, and cash flow statements according to the preset rules. Standard financial data formats are generated for output, including XML, JSON, and CSV, and at the same time, the data source and confidence information are recorded. The performance of the summarization model is evaluated through regression testing, and the summarization accuracy and consistency indicators are introduced as evaluation criteria. When the accuracy is lower than 95% or the consistency indicator is lower than 0.93, the model parameter optimization is triggered. This step aims to convert the corrected document text data into structured summary information, realize the conversion from scattered document data to centralized financial information, and support enterprise financial analysis and decision-making.

[0045] The detailed structure of the financial perception self-correction model is a two-stream attention network based on the Vision Transformer architecture, which includes a visual feature extraction stream and a text feature extraction stream. The visual feature extraction stream adopts a resolution adaptive blockification mechanism. The input image first extracts low-level features through a convolutional neural network and then is segmented into image blocks of variable sizes, with a default block size of 16×16 pixels. When the wrinkle degree parameter exceeds 75, the block size is reduced to 8×8 pixels to improve the feature extraction accuracy. Each image block is input into a 12-layer Transformer encoder after position encoding, with a hidden layer dimension of 768, a feed-forward network dimension of 3072, and 12 attention heads. The text feature extraction stream adopts a hierarchical attention network structure, including a character-level feature extraction layer, a word-level feature extraction layer, and a sentence-level feature extraction layer. The character-level features are obtained through character embedding, with an embedding dimension of 128; the word-level features are extracted through a bidirectional GRU network, with a hidden layer dimension of 256; the sentence-level features are integrated through a self-attention mechanism, with 8 attention heads. The model dynamically adjusts the feature fusion weights at each level according to the stain readability parameter. When the stain readability parameter is lower than 30, the weight of the character-level features is increased to 0.6, the weight of the word-level features is reduced to 0.3, and the weight of the sentence-level features is set to 0.1. The two feature streams integrate information through an adaptive fusion module. The fusion module uses a gating mechanism to control the information flow, and the gating function is the Sigmoid function. The input is the concatenated vector of the features of the two streams. The financial domain knowledge constraints are introduced in the fusion process, and the original information is retained through a skip connection mechanism. The core of the model is an adaptive gating attention mechanism, which selectively extracts information from different quality regions by dynamically adjusting the attention distribution parameters.

[0046] The detailed steps for establishing the training dataset of the financial perception self-correction model include collecting more than 100,000 financial document image samples from the financial systems of 20 enterprises of different scales, covering 7 common types of financial documents such as value-added tax invoices, ordinary invoices, receipts, and bank statements. The collected document images are graded in terms of quality into 5 levels: Level A is high-quality scanned samples with a resolution of not less than 300 DPI; Level B is standard-quality scanned samples with a resolution of 200 - 300 DPI; Level C is low-quality scanned samples with a resolution lower than 200 DPI or with minor quality problems; Level D is samples with obvious quality problems, including single problems such as wrinkles, stains, and blurs; Level E is samples with multiple problems, having multiple quality defects simultaneously. For each type of document, field structure annotation is carried out, and the annotation content includes document type, field position, field content, and field relationship. The annotation accuracy requires that the framing error of the field position does not exceed 5 pixels. Data augmentation techniques are used to construct various quality problem samples, and by adding different degrees of uneven illumination, wrinkles, stains, and blur effects, the scale of the training set is expanded to 3 times the original data. For each constructed sample, parameters such as illumination uniformity parameter, wrinkling degree parameter, wrinkling change degree parameter, wrinkling jump parameter, and stain readability parameter are calculated to form a quality parameter annotation set. A knowledge base of the logical relationships between document fields is established, including the core logical constraints such as the relationships among quantity, unit price, and amount, and the relationships among tax rate, tax-exclusive amount, and tax amount commonly found in financial documents. OCR preliminary recognition is carried out on all samples, and the edit distance between the recognition result and the true annotation content is recorded as the error correction target for model training. Balanced training data batches are constructed according to document type and quality parameters, and each batch contains samples of different types and different quality levels to ensure that the model has sufficient learning for various problem scenarios.

[0047] The training of the financial perception self-correction model adopts a two-stage strategy of pre-training and fine-tuning. First, the basic parameters of the model are pre-trained on a general dataset containing 5 million text images, using the Adam optimizer with an initial learning rate set to 10 -4 , and the weight decay is 10 -5 , and it is trained for 500 epochs. Then, fine-tuning is carried out using a partial financial document dataset to be summarized, and the initial learning rate is reduced to 10 -5, the cosine annealing learning rate scheduling strategy is used. The training process adopts a multi-task joint learning framework, and simultaneously optimizes three objective functions: the text recognition accuracy, the field localization accuracy, and the document relationship reasoning accuracy. The weight ratio of the three is 5:3:2. An adversarial learning mechanism is introduced, and adversarial samples are constructed by adding perturbations to the training samples, with the perturbation amplitude controlled within 3% of the original image pixel value. A dynamic weight adjustment mechanism is adopted for samples with different quality parameters, and difficult samples with a quality parameter lower than 30 are given a training weight twice as large. Financial domain knowledge distillation is introduced in the later stage of training, and the knowledge of the pre-trained financial expert model is incorporated into the self-correcting model parameters, with the distillation temperature parameter set to 2.0. The gradient accumulation strategy is used to handle large-batch data training, and the accumulation step is set to 4. Early stopping is prevented by periodic learning rate adjustment and early stopping strategy. Early stopping is triggered when the performance of the validation set does not improve for 5 consecutive rounds. After training, the model is quantized and pruned for optimization. The weight precision is reduced from 32-bit floating-point numbers to 8-bit integers, and the pruning ratio is set to 30% to ensure the computational efficiency of the model during inference.

[0048] It should be noted that there are mainly three core technical ideas in the present invention: one is the multi-dimensional quality parameter evaluation system, the second is the financial perception self-correcting model based on the Vision Transformer architecture, and the third is the dynamic adjustment mechanism of the quality impact gating function.

[0049] The multi-dimensional quality parameter evaluation system realizes the refined quantification of the quality problems of financial accounting document images by calculating the wrinkle degree parameter, the wrinkle change degree parameter, the wrinkle jump degree parameter, the stain readability parameter, and the illumination uniformity parameter. Different from the traditional method that only relies on simple global image quality evaluation, the present invention can perform regional and differential analysis on the quality problems unique to financial accounting documents, providing an accurate quality evaluation basis for subsequent correction. This multi-dimensional and fine-grained quality evaluation mechanism enables the system to identify the influence degree of different types of quality problems on text recognition, so as to perform targeted correction, avoiding the problems of insufficient correction or over-correction caused by the "one-size-fits-all" treatment in the traditional method.

[0050] The financial perception self-correction model based on the Vision Transformer architecture adopts a two-stream attention network structure, which extracts visual features and text features respectively and performs adaptive fusion. Compared with traditional OCR error correction methods that only rely on text features or simple image features, this model can make full use of the complementary information of images and texts. Among them, the visual feature extraction stream can dynamically adjust the block size according to the fold degree parameter through a resolution adaptive block mechanism, ensuring that finer-grained blocks are used in high-fold regions to retain more details; the text feature extraction stream adopts a hierarchical attention structure, processes character-level, word-level, and sentence-level features simultaneously, and introduces financial domain knowledge constraints to make the correction results more in line with financial data specifications. This deep learning architecture that integrates visual and text features and introduces domain knowledge enables the model to understand the complex relationship between text content and image quality, thus achieving more accurate text correction.

[0051] As a regulation mechanism, the quality impact gating function can calculate the comprehensive quality balance value according to multi-dimensional quality parameters, and select different adjustment strategies accordingly to dynamically control the change of the attention distribution parameters in the financial perception self-correction model. This dynamic adjustment mechanism enables the model to adopt different processing strategies for different quality regions: retaining the original recognition results for high-quality regions, making moderate adjustments for medium-quality regions, and adopting an aggressive correction strategy for low-quality regions. Compared with traditional methods that use a unified correction strategy, this adaptive adjustment mechanism can accurately intervene in the quality problems of different regions in the document, avoiding unnecessary interference with the correctly recognized regions.

[0052] The synergistic effect of these three core technical ideas forms a closed-loop optimization system: multi-dimensional quality parameters provide an accurate quality assessment basis, the financial perception self-correction model performs deep feature extraction and fusion based on these parameters, and the quality impact gating function dynamically adjusts the attention distribution of the model according to the quality parameters, making the correction intensity match the degree of quality problems. This multi-level, closed-loop synergistic mechanism enables the system to accurately identify the text information in the wrinkled deformation and stain-covered areas, achieve high-precision recognition and correction of the text of accounting documents under complex quality conditions, thus significantly improving the accuracy and efficiency of the computerization of accounting data, reducing the need for manual intervention, and providing a more reliable automated solution for enterprise financial data processing.

[0053] Specifically, the principle of the present invention is as follows: The core of the technical principle of the present invention lies in constructing a two-stream attention network architecture that integrates visual features and text features, and realizing targeted correction of different quality regions through an adaptive adjustment mechanism driven by quality parameters. First, the invention proposes a complete document quality evaluation system, including illumination uniformity parameters, wrinkling degree parameters, wrinkling change degree parameters, wrinkling jump degree parameters, and stain readability parameters. These parameters quantify the quality characteristics of the document image from different dimensions, providing a scientific basis for subsequent differential processing.

[0054] Secondly, the financial perception self-correction model of the present invention adopts the Vision Transformer architecture. Through the resolution adaptive blockification mechanism, the model can dynamically adjust the block size according to the wrinkling degree parameters, and adopt a finer-grained block in the high-wrinkling degree area to retain more detailed features. At the same time, the text feature extraction stream adopts a hierarchical attention network structure, which can process character-level, word-level, and sentence-level features simultaneously, and adjust the fusion weights of features at each level according to the stain readability parameters, realizing refined processing of texts with different damage degrees.

[0055] Most innovatively, the present invention introduces a quality impact gating function. This function takes multi-dimensional quality parameters as inputs, calculates the comprehensive quality balance value, and accordingly selects different adjustment strategies to achieve precise control of the model's attention distribution parameters. When the document quality is high, the model tends to retain the original recognition result; while when the document quality is low, the model increases the correction intensity for the quality problem areas. This adaptive adjustment mechanism enables the model to accurately correct the recognition errors in the low-quality areas while maintaining accurate recognition in the high-quality areas. In addition, by integrating financial domain knowledge constraints and using financial term representations and logical relationships between fields, the model further improves the professionalism and rationality of the correction results.

[0056] The present invention adopts a two-stage training strategy of pre-training plus fine-tuning, and simultaneously optimizes the text recognition accuracy, field localization accuracy, and document relationship reasoning accuracy through a multi-task joint learning framework. By introducing an adversarial learning mechanism and dynamic weight adjustment, the robustness of the model to various quality problems is enhanced. The organic combination of these technical means enables the present invention to effectively solve multiple quality problems in financial document OCR recognition and achieve precise correction of the recognition results.

[0057] The following provides a specific Embodiment 1 of the present invention, and the specific implementation manners of each step in this Embodiment 1 are described in detail as follows.

[0058] The specific implementation of step S01 is to comprehensively preprocess the collected financial document images. First, the image is subjected to illumination equalization processing through a light uniformity compensation function. This function is based on the adaptive histogram equalization algorithm and processes the image in blocks, with each block having a size of 32×32 pixels. Bilinear interpolation is used for transition between adjacent blocks. The luminance distribution histograms are calculated separately for the RGB channels of the image. When the ratio of the histogram peak to the mean exceeds 3.5, it is determined that there is uneven illumination. The mathematical expression of the light uniformity compensation function is as follows:

[0059] I eq (x, y) = I orig (x, y)·γ(x, y);

[0060] In the formula, I eq (x, y) is the pixel value of the equalized image at the coordinate (x, y); I orig (x, y) is the pixel value of the original image at the coordinate (x, y); γ(x, y) is the adaptive gamma correction coefficient, which is dynamically calculated according to the regional luminance mean μ(x, y). For the dark area, the calculation formula of γ(x, y) is:

[0061]

[0062] In the formula, μ(x, y) is the local regional luminance mean centered on the coordinate (x, y); μ max is the maximum luminance value of the image (usually 255). For the bright area, the calculation formula of γ(x, y) is:

[0063]

[0064] At the same time, the wrinkle degree parameter W d is calculated. Linear edge features are extracted through the Canny edge detection algorithm, and double thresholds are set to 50 and 150 respectively. The wrinkle degree parameter is quantized to an integer value from 0 to 100, and the calculation formula is:

[0065]

[0066] In the formula, L i is the length of the i-th detected linear edge; D i is the direction discontinuity coefficient of the i-th linear edge; A total is the total area of the image (number of pixels); α is the normalization coefficient, with a value of 0.05; n is the total number of detected edge lines. The calculation formula of the direction discontinuity coefficient D i is:

[0067]

[0068] In the formula, θj is the direction angle of the j-th segment on the i-th edge line; m is the number of sampling points on this edge line. Additionally, calculate the stain readability parameter R s , identify the stain area through local contrast enhancement and adaptive threshold segmentation algorithms, and the calculation formula is:

[0069]

[0070] In the formula, S represents the set of pixels in the detected stain area; |S| represents the total number of pixels in the stain area; C(x, y) represents the text stroke integrity coefficient at the coordinate (x, y) (range 0 - 1); E(x, y) represents the edge sharpness coefficient at the coordinate (x, y) (range 0 - 1); I(x, y) represents the contrast coefficient at the coordinate (x, y) (range 0 - 1). The purpose of this step is to improve the image quality, provide a clearer image input for subsequent OCR recognition, and obtain image quality parameters for subsequent processing.

[0071] The specific implementation of step S02 is to input the preprocessed financial document image into the OCR engine for preliminary recognition. This step uses the deep learning model CRNN (Convolutional Recurrent Neural Network) as the basic recognition engine and combines the attention mechanism to optimize the recognition effect. The image first extracts features through the VGG16 network, then inputs into the bidirectional LSTM network for sequence modeling, and finally outputs the text sequence and the recognition confidence of each character through the CTC (Connectionist Temporal Classification) loss function. For Chinese character recognition, a dictionary of 5000 common Chinese characters is used as the recognition basis, and an additional dictionary is established for optimization of financial special vocabulary. The confidence matrix records the recognition credibility of each character in units of characters, and the confidence threshold is set to 0.75. Recognition results below this threshold are marked as suspicious areas. The role of this step is to obtain the original recognition text data and the corresponding recognition confidence matrix, providing basic data and confidence basis for subsequent correction.

[0072] The specific implementation of step S03 is to use the financial perception self-correction model to correct the original recognition text data. This model receives the original recognition text data, the recognition confidence matrix, the wrinkle degree parameter, and the stain readability parameter as inputs. The model first encodes the original recognition text data into character-level feature vectors, with the feature vector dimension of each character being 128, and then assigns different weights to the feature vectors according to the recognition confidence matrix. For each character c i , the weighted calculation formula of its feature vector is:

[0073]

[0074] In the formula, is the weighted feature vector of character c i ; is the original feature vector for character c i ; p i is the recognition confidence of character c i ; K(c i ) is the constraint weight for character c i in the financial domain knowledge graph; β(p i ) is the confidence-based weight adjustment function, and its calculation formula is:

[0075]

[0076] When the fold degree parameter W d value exceeds 70, the model dynamically adjusts the block size, which is adjusted from the default 16×16 to 8×8 to more finely capture local features. When the stain readability parameter R s is lower than 30, the model strengthens the weight of the visual feature extraction stream and increases it to 1.8 times the original weight. By fusing multi-source information, corrected text data is generated, and the credibility index CI is calculated. Its calculation formula is:

[0077] CI = λ1·p orig +λ2·p corr +λ3·p know ;

[0078] In the formula, CI is the comprehensive credibility index; p orig is the original confidence; p corr is the change in confidence during the correction process; p know is the confidence contribution of the knowledge graph constraint; λ1, λ2, and λ3 are weight coefficients, and the default values are 0.3, 0.4, and 0.3 respectively, and λ1 + λ2 + λ3 = 1. This step aims to use financial domain knowledge to perform preliminary intelligent correction on the original recognition result and improve the text recognition accuracy.

[0079] The specific implementation of step S04 is to calculate the degree of influence of the quality of the accounting document image on each character area in the original recognition text data using the scanning quality compensation function. This function uses a multi-scale image quality assessment method, comprehensively considering the variance of the image gradient amplitude, the frequency domain energy distribution, and the texture features. The image is decomposed into multiple scales through wavelet transform, and the decomposition level is set to 3 layers to extract the energy distribution features of the high-frequency subbands. The quality influence weight vector is calculated as follows:

[0080]

[0081] In the formula, N is the total number of characters in the text data; q i is the quality influence weight of the i-th character area, with a range of 0 to 1. The smaller the value, the more serious the quality influence. The quality influence weight q iThe calculation formula is as follows:

[0082] q i = σ(ω1·G i + ω2·F i + ω3·T i );

[0083] In the formula, σ is the Sigmoid function; G i is the standard deviation of the normalized gradient amplitude of the i-th character region; F i is the normalized high-frequency energy ratio of the i-th character region; T i is the normalized texture complexity of the i-th character region; ω1, ω2, and ω3 are weight coefficients, with default values of 0.4, 0.3, and 0.3 respectively, and satisfy ω1 + ω2 + ω3 = 1. Among them, the calculation formula for the standard deviation of the gradient amplitude G i is as follows:

[0084]

[0085] In the formula, is the image gradient of the i-th character region; is the standard deviation of the gradient amplitude; σ max is the normalization coefficient, with a value of 25. When is less than 10, G i is set to 0.2. The calculation formula for the high-frequency energy ratio F i is as follows:

[0086]

[0087] In the formula, E high is the high-frequency subband energy; E total is the total energy. When F i is lower than 0.15, F i is set to 0.25.

[0088] The texture complexity T i is calculated through the gray-level co-occurrence matrix, and the calculation formula is:

[0089]

[0090] In the formula, H i is the Shannon entropy value of the i-th character region; H max is the maximum possible entropy value (8 for 8-bit images). When H i is lower than 4.5, T i is set to 0.3. The purpose of this step is to quantify the influence degree of image quality on the recognition result and provide a targeted quality compensation basis for subsequent correction.

[0091] The specific implementation of step S05 is to analyze the impact of the deformation of the financial document image on the original recognized text data according to the fold parameters. First, analyze the fold change degree parameter W v , which is obtained by calculating the included angle of the fold feature vectors in adjacent regions and the fold line density gradient. The calculation formula is:

[0092]

[0093] In the formula, M is the number of regions into which the image is divided; is the fold feature vector of the j-th region; is the included angle between the feature vectors of adjacent regions; ρ j is the fold line density of the j-th region; ρ max is the normalization coefficient, and its value is 20% of the area of the region. When the included angle between the feature vectors of adjacent regions exceeds 45 degrees (i.e., π / 4), it is regarded as a high change degree region. The fold jump degree parameter W j is calculated by detecting the number of break points of the fold line segments and the jump frequency of the fold region boundaries. The calculation formula is:

[0094]

[0095] In the formula, P is the total number of detected fold line segments; J k is the number of break points of the k-th fold line segment; A total is the total area of the image (number of pixels); γ is the normalization coefficient, and its value is 0.005. When the number of break points per unit area exceeds 5, it is determined as a high jump degree region. Based on these parameters, adjust the attention distribution parameter A in the financial perception self-correction model through the quality impact gating function d , and the calculation formula is:

[0096] A d = A0·(1 + δ1·f(W v ) + δ2·g(W j ));

[0097] In the formula, A d is the adjusted attention distribution parameter; A0 is the basic attention distribution parameter; f(W v ) and g(W j ) are the influence functions of the fold change degree and the fold jump degree respectively; δ1 and δ2 are weight coefficients, and their default values are both 0.3. When the fold change degree parameter W v exceeds 75, the number of attention heads increases from the default 8 to 12; when the fold jump degree parameter W jWhen it exceeds 60, the adaptive gating attention threshold is reduced from the default 0.65 to 0.45, increasing the model's attention to low-quality regions. This step aims to dynamically adjust the calibration model parameters according to the document deformation characteristics and improve the accuracy of text recognition in the wrinkled areas.

[0098] The specific implementation of step S06 is to combine the quality influence weight vector with the corrected text data to perform multi-layer calibration on the original recognized text data. First, a hierarchical calibration framework is constructed, including three levels: character-level calibration, word-level calibration, and sentence-level calibration. The decision function D of character-level calibration c The calculation formula is:

[0099]

[0100] In the formula, c i is the originally recognized character; c i ′ is the corrected character; q i is the quality influence weight of the i-th character. Word-level calibration uses the N-gram model for context analysis and uses the 3-gram model to analyze the collocation probability P(w i |w i-2 , w i-1 ). When the collocation probability is lower than 0.25, calibration is triggered. The decision function D of word correction w The calculation formula is:

[0101]

[0102] In the formula, w i is the originally recognized word; w i ′ is the corrected word; P(w i |w i-2 , w i-1 ) is the conditional probability of the word w i under the condition of the previous two words w i-2 and w i-1 . Sentence-level calibration uses the financial domain knowledge graph to verify the rationality of the corrected text data. The knowledge graph contains 5,000 financial core concepts and 30,000 concept relationships. The recognized text is matched with the knowledge graph entities through the entity alignment algorithm, and the matching threshold is set to 0.7. When there is a logical conflict between the recognition result and the knowledge graph, the maximum likelihood estimation method is used to select the most reasonable calibration scheme. The calculation formula is:

[0103] s * = argmax s∈S P(s|G);

[0104] In the formula, s *is the optimal correction scheme; S is the set of all possible correction schemes; P(s|G) is the conditional probability of scheme s under the condition of knowledge graph G. The purpose of this step is to maximize the accuracy of text correction through a multi-level correction strategy and combined with domain knowledge constraints.

[0105] The specific implementation of step S07 is to structurally integrate and summarize the corrected text data. First, based on the document layout analysis algorithm, a spatial relationship model is built for the data. The YOLO v5 object detection network is used to identify the keyword field areas, and the detection confidence threshold is set to 0.8. The recognized text is logically grouped by the hierarchical clustering algorithm, and the clustering distance threshold is set to 15 pixels. The distance calculation formula for hierarchical clustering is:

[0106]

[0107] In the formula, C i and C j are two different text regions; p and q are the points in regions C i and C j respectively; ||p - q|| is the Euclidean distance between two points. According to the predefined financial document template library, the recognized text is mapped to the standard fields, and the template matching similarity threshold is set to 0.85. A semantic consistency check is performed on the recognized data to verify whether the calculation relationship of the digital fields conforms to the financial rules, and the allowable error range is 0.01. Next, cross-document data summarization processing is performed. First, an index tree structure is built according to the document type, and the same type of documents are organized into a tree structure for quick retrieval. The tree depth is set to 3 layers. The adaptive data fusion algorithm is used to intelligently merge the same type of data from different documents. The fusion granularity is divided into three levels: document level, field level, and numerical level. When there are multiple associated documents in the same business scenario, a conflict resolution strategy is set, and the data selection is determined according to the document priority and timestamp. The priority configuration matrix is preset according to the importance of the documents. The sliding time window technology is used to achieve dynamic data summarization, and the window size can be configured as daily, weekly, monthly, quarterly, and annual. The default setting is monthly summarization. For the data that needs to be summarized by subject, multi-dimensional aggregation calculation is implemented, and automatic balance verification of debit and credit directions is supported. The balance difference threshold is set to 0.001. The summarization accuracy Acc sum is calculated by the formula:

[0108]

[0109] In the formula, N is the total number of summarized data items; V i is the actual value of the i-th item; V i ′ is the predicted value of the i-th item; V max is the maximum value of the data. The data consistency index CI sum is calculated by the formula:

[0110]

[0111] In the formula, N consist is the number of data items for maintaining consistency; N total is the total number of data items. When the accuracy rate is lower than 95% or the consistency index is lower than 0.93, the optimization of model parameters is triggered. This step aims to convert the corrected document text data into structured summary information, realizing the conversion from scattered document data to centralized financial information, and supporting corporate financial analysis and decision-making.

[0112] The detailed structure of the financial perception self-correction model is a two-stream attention network based on the Vision Transformer architecture, including a visual feature extraction stream and a text feature extraction stream. The visual feature extraction stream adopts a resolution adaptive blockification mechanism. The input image first extracts low-level features through a convolutional neural network, and then is segmented into image blocks with variable sizes, and the default block size is 16×16 pixels. When the wrinkle degree parameter W d exceeds 75, the block size is reduced to 8×8 pixels to improve the feature extraction accuracy. Each image block is input into a 12-layer Transformer encoder after position encoding, with a hidden layer dimension of 768, a feed-forward network dimension of 3072, and 12 attention heads. The text feature extraction stream adopts a hierarchical attention network structure, including a character-level feature extraction layer, a word-level feature extraction layer, and a sentence-level feature extraction layer. The two feature streams integrate information through an adaptive fusion module. The fusion module uses a gating mechanism to control the information flow, and the gating function is the Sigmoid function, and the input is the concatenated vector of the features of the two streams. The mathematical expression of the adaptive fusion module is:

[0113] F fused = G(F v , F t )·F v +(1 - G(F v , F t ))·F t ;

[0114] In the formula, F fused is the fused feature; F v is the visual feature; F t is the text feature; G(F v , F t ) is the gating function, and the calculation formula is:

[0115] G(F v , F t ) = σ(W g ·[F v ; F t +b g );

[0116] where σ is the Sigmoid function; W g is the gating weight matrix; b g is the bias vector; [F v ; F t represents the concatenation of visual features and text features. The fusion process introduces financial domain knowledge constraints and preserves the original information through a skip connection mechanism. The core of the model is an adaptive gating attention mechanism, which selectively extracts information from different quality regions by dynamically adjusting the attention distribution parameters.

[0117] The detailed steps for establishing the training dataset of the financial perception self-correction model include collecting more than 100,000 financial document image samples from the financial systems of 20 enterprises of different scales, covering 7 common types of financial documents such as value-added tax invoices, ordinary invoices, receipts, and bank statements. The collected document images are graded for quality into 5 levels: Level A is high-quality scanned samples with a resolution of not less than 300 DPI; Level B is standard-quality scanned samples with a resolution of 200 - 300 DPI; Level C is low-quality scanned samples with a resolution below 200 DPI or with minor quality problems; Level D is samples with obvious quality problems, including single problems such as wrinkles, stains, and blurs; Level E is samples with multiple problems, with multiple quality defects. Each type of document is structurally annotated for fields, and the annotation content includes document type, field location, field content, and field relationships. The annotation accuracy requires that the framing error of the field location does not exceed 5 pixels. Data augmentation techniques are used to construct various quality problem samples, expanding the scale of the training set to 3 times the original data. Illumination uniformity parameters, wrinkle degree parameters, wrinkle change degree parameters, wrinkle jump degree parameters, and stain readability parameters are calculated for each constructed sample to form a quality parameter annotation set.

[0118] The training of the financial perception self-correction model adopts a two-stage strategy of pre-training plus fine-tuning. First, the basic parameters of the model are pre-trained on a general dataset containing 5 million text images, using the Adam optimizer with an initial learning rate set to 10 -4 , and a weight decay of 10 -5 , and trained for 500 epochs. Then, fine-tuning is performed using a partial financial document dataset to be summarized, with the initial learning rate reduced to 10 -5 , and the cosine annealing learning rate scheduling strategy is used. The training process adopts a multi-task joint learning framework, simultaneously optimizing three objective functions: text recognition accuracy, field localization accuracy, and document relationship reasoning accuracy, with a weight ratio of 5:3:2 for the three. The multi-task joint loss function L joint is calculated as follows:

[0119] L joint = 0.5·L text + 0.3·L field + 0.2·Lrel ;

[0120] In the formula, L text is the text recognition loss; L field is the field localization loss; L rel is the document relationship reasoning loss. The text recognition loss L text adopts the cross-entropy loss function; the field localization loss L field adopts a combination of smooth L1 loss and cross-entropy loss; the document relationship reasoning loss L rel adopts binary cross-entropy loss. An adversarial learning mechanism is introduced. By adding perturbations to the training samples to construct adversarial samples, the perturbation amplitude is controlled within 3% of the original image pixel value. A dynamic weight adjustment mechanism is adopted for samples with different quality parameters. Difficult samples with a quality parameter lower than 30 are given a training weight of 2 times. The weight adjustment function W adj has the following calculation formula:

[0121]

[0122] In the formula, Q is the comprehensive quality parameter of the sample, which is the weighted average of the illumination uniformity parameter, the wrinkling degree parameter, the wrinkling change degree parameter, the wrinkling jump degree parameter, and the stain readability parameter. After training, the model is quantized and pruned for optimization. The weight precision is reduced from 32-bit floating-point numbers to 8-bit integers, and the pruning ratio is set to 30% to ensure the calculation efficiency of the model during inference.

[0123] The quality impact gating function is a dynamic adjustment mechanism based on multiple parameters, used to control the dynamic changes of the attention distribution parameters in the financial perception self-correction model. The inputs of this function include the illumination uniformity parameter L u , the wrinkling degree parameter W d , the wrinkling change degree parameter W v , the wrinkling jump degree parameter W j , and the stain readability parameter R s , a total of five quality parameters. Its calculation process first calculates the comprehensive quality balance value Q bal , and the calculation formula is:

[0124] Q bal = α1·L u + α2·(100 - W d ) + α3·(100 - W v ) + α4·(100 - W j ) + α5·R s ;

[0125] In the formula, α1, α2, α3, α4, and α5 are weight coefficients, and the default values are 0.2, 0.25, 0.2, 0.15, and 0.2 respectively, and satisfy Select different adjustment strategies according to the range of the comprehensive quality balance value, and the interval division is as follows:

[0126]

[0127] In the formula, S conservative represents the conservative adjustment strategy, and S moderate represents the moderate adjustment strategy, and S aggressive represents the aggressive adjustment strategy. When Q bal is in the high-quality interval (≥75), the conservative adjustment strategy is adopted, focusing on retaining the original recognized text data; when Q bal is in the medium-quality interval (40 - 75), the moderate adjustment strategy is adopted, and the low-confidence regions are mainly corrected on the basis of retaining the original recognized text data; when Q bal is in the low-quality interval (<40), the aggressive adjustment strategy is adopted, and the attention mechanism of the financial perception self-correction model is greatly enhanced to pay attention to the low-quality regions.

[0128] The quality influence gating function precisely controls the behavior of the financial perception self-correction model by dynamically adjusting three key parameters: the number of attention heads H, the attention distribution parameter A d and the feature fusion weight W f . The adjustment formula for the number of attention heads H is:

[0129]

[0130] In the formula, H0 is the basic number of attention heads, and the default value is 8; δ H is the adjustment coefficient, and the default value is 0.5. After rounding, the actual number of attention heads is obtained. The adjustment formula for the attention distribution parameter A d is:

[0131]

[0132] In the formula, A0 is the basic attention distribution parameter, and the default value is 0.65; γ A is the adjustment coefficient, and the default value is 0.4.

[0133] The adjustment formula for the feature fusion weight W f is:

[0134]

[0135] In the formula, w1, w2, and w3 are the fusion weights of character-level, word-level, and sentence-level features respectively, and satisfy w1 + w2 + w3 = 1. The output of this function is the adjusted attention distribution parameter, which is used to control the behavior of the adaptive gating attention mechanism in the financial perception self-correction model.

[0136] The light uniformity parameter Lu It refers to an index for measuring the evenness of the illumination distribution in each area of the financial document image, and the calculation formula is:

[0137]

[0138] In the formula, σ I is the standard deviation of the brightness in different areas of the image; μ I is the average value of the overall brightness of the image. This parameter is calculated based on the financial document image through the preprocessing process in step S01 and is used as an input parameter for the quality impact gating function in step S05.

[0139] To sum up, this method constructs an OCR recognition and data summarization system for financial documents that can effectively handle various quality problems by integrating multiple image processing technologies, deep learning models, and financial domain knowledge. The key feature of this method is that it designs a special parametric model for common quality problems such as wrinkles and stains in financial documents. Through precise quantitative analysis of image quality, it dynamically adjusts the recognition and correction strategies, achieving high-accuracy text recognition and data structuring. At the same time, this method makes full use of the constraints of financial domain knowledge and introduces financial rules and logical relationship verification in the multi-level correction framework, further improving data accuracy and consistency. The financial perception self-correction model used in the method adopts a two-stream attention network structure, which can process visual features and text features simultaneously and achieve optimal feature integration through an adaptive fusion mechanism. During the model training process, a multi-task learning framework and an adversarial learning mechanism are adopted to enhance the robustness of the model to various quality problems. The structured summarization link adopts a hierarchical processing strategy, from the character level to the document level and then to the business level, gradually realizing data integration and summarization, and finally generating a standardized financial data format output, providing reliable data support for enterprise financial analysis and decision-making.

[0140] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: Researchers summarized a large number of financial documents, including about 2,000 various financial documents such as value-added tax invoices, ordinary invoices, receipts, and bank statements. Due to the diverse sources and uneven quality of the documents, the recognition accuracy of traditional OCR systems is not high, and the data summarization efficiency is low. The researchers randomly selected 500 financial documents of different types and quality grades within a quarter as test samples, including problem documents with different degrees of wrinkles, stains, uneven illumination, etc. The sample distribution is shown in Table 1:

[0141] Table 1 Test sample distribution table

[0142]

[0143] In the implementation of step S01, the researchers preprocessed all the test samples. After the image preprocessing, the calculated average quality parameters are shown in Table 2:

[0144] Table 2 Average Quality Parameter Table after Preprocessing

[0145]

[0146] In the implementation of step S02, the researchers input the preprocessed images into the OCR engine for preliminary recognition, obtaining the original recognition text data and the recognition confidence matrix. The OCR preliminary recognition results of documents with different quality levels are shown in Table 3:

[0147] Table 3 OCR Preliminary Recognition Accuracy Table

[0148] Document quality grade Character accuracy rate (%) Field accuracy rate (%) Average confidence level Proportion of suspicious areas (%) Grade A 97.8 95.3 0.92 3.5 Grade B 93.6 90.2 0.87 8.2 Grade C 84.5 79.8 0.76 18.4 Grade D 72.3 65.7 0.67 32.6 Grade E 58.9 48.3 0.52 45.8

[0149] In the implementation of steps S03 to S06, the researchers applied the financial perception self-correction model to correct the original recognition text data. For documents with different quality levels, the model automatically adjusted the relevant parameters, as shown in Table 4:

[0150] Table 4 Model Automatic Parameter Adjustment Table

[0151]

[0152] After multiple corrections, the recognition accuracy was significantly improved, and the correction effect is shown in Table 5:

[0153] Table 5 Recognition Accuracy Table after Correction

[0154]

[0155] In the implementation of step S07, the researchers structured and summarized the corrected text data to generate the standard financial data format. For the 500 tested documents, the summary processing efficiency and accuracy are shown in Table 6:

[0156] Table 6 Data Summary Efficiency and Accuracy Table

[0157]

[0158] Traditional OCR financial document recognition methods usually adopt recognition models with fixed parameters and simple post - processing rules, and are unable to effectively process document images of different quality levels. Especially for low - quality documents with problems such as wrinkles and stains, the recognition accuracy rate is usually no more than 60%. Traditional methods also lack the support of financial domain knowledge in the data aggregation link, resulting in relatively low accuracy and consistency of the aggregated data, and the aggregation accuracy rate is usually around 85%. By introducing a financial perception self - calibration model and combining it with a quality - impact gating function, the present invention can dynamically adjust the recognition strategy according to the quality of the document image, significantly improving the recognition accuracy rate of low - quality documents. The test results show that for class - E seriously problematic samples, the recognition accuracy rate of the method of the present invention reaches 82.4%, which is about 18.7% higher than that of traditional methods. In the data aggregation link, by introducing financial domain knowledge constraints and a multi - level calibration framework, the present invention enables the aggregation accuracy rate to reach 97.6%, which is about 12.6% higher than that of traditional methods. At the same time, the method of the present invention has stronger adaptability to different types of documents, faster processing speed, can meet the actual needs of enterprise financial work, and significantly improves the efficiency and accuracy of financial data processing.

[0159] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 7 and 8 below.

[0160] Table 7 Variable Explanation Table (Part 1)

[0161]

[0162]

[0163] Table 8 Variable Explanation Table (Part 2)

[0164]

[0165]

[0166] As described above, the above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for summarizing financial document data based on artificial intelligence to correct OCR, characterized in that, Including: Collecting the images of financial accounting documents and preprocessing them through a light uniformity compensation function, while calculating the wrinkle degree parameter and the stain readability parameter; Inputting the preprocessed images of financial accounting documents into an OCR engine for preliminary recognition to obtain the original recognition text data and the recognition confidence matrix; correcting the original recognition text data by using a financial perception self-correction model, the financial perception self-correction model includes a visual feature extraction stream and a text feature extraction stream, and integrating information through an adaptive fusion module; calculating the degree of influence of each character region in the original recognition text data on the quality of the financial accounting document image by using a scanning quality compensation function to generate a quality influence weight vector; analyzing the deformation influence of the financial accounting document image according to the wrinkle change degree parameter and the wrinkle jump degree parameter, and adjusting the attention distribution parameter through a quality influence gating function; combining the quality influence weight vector with the corrected text data for multi-layer correction, and at the same time verifying the data rationality by using a financial domain knowledge graph; structurally integrating and summarizing the corrected text data to generate a standard financial data format for output.

2. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 1, wherein The light uniformity compensation function refers to a mathematical model that analyzes the brightness distribution histogram of the financial accounting document image, calculates the unevenness degree of illumination in each region of the financial accounting document image, and adjusts the pixel value distribution through an adaptive gamma correction and a multi-scale histogram equalization algorithm to make the overall illumination of the financial accounting document image tend to be uniform.

3. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 2, wherein The scanning quality compensation function refers to a function model that calculates the quality level of the financial accounting document image based on an image sharpness evaluation index and assigns different credibility weight coefficients to the original recognition text data for different quality levels.

4. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 3, wherein, The wrinkle degree parameter refers to the total amount of wrinkle features in the financial accounting document image. By detecting the linear edge density and direction discontinuity in the financial accounting document image, and combining high-frequency filtering to extract the texture abnormal region, it is quantitatively represented as an integer value from 0 to 100.

5. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 4, wherein The wrinkle change degree parameter refers to the degree of drastic change in the wrinkle form in the financial accounting document image. By calculating the included angle between the wrinkle feature vectors in adjacent regions and the wrinkle line density gradient, it quantitatively represents the complexity of the deformation of the financial accounting document image.

6. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 5, wherein, The wrinkle jump degree parameter refers to the degree of discontinuity in the spatial distribution of the wrinkle features in the financial accounting document image. By detecting the number of break points of the wrinkle line segments and the jump frequency of the wrinkle region boundaries, it reflects the degree of influence of the local deformation of the financial accounting document image on character recognition.

7. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 6, wherein, The stain readability parameter refers to the degree to which the text information in the area covered by the stain in the financial accounting document image can still be recognized. It is calculated by analyzing the characteristics such as the integrity of the text strokes, the contrast, and the edge sharpness in the stain area.

8. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 7, wherein, The specific structure of the financial perception self-correction model is a two-stream attention network based on the Vision Transformer architecture. The visual feature extraction stream adopts a resolution adaptive blockization mechanism and dynamically adjusts the block size parameter according to the wrinkle degree parameter. The text feature extraction stream adopts a hierarchical attention network structure and can process character-level features, word-level features, and sentence-level features simultaneously.

9. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 8, characterized in that The information of the visual feature extraction stream and the text feature extraction stream is interactively fused through the multi-head self-attention mechanism. During the fusion process, financial domain knowledge constraints are introduced, and the original information is retained through the skip connection mechanism and financial term representations are introduced.

10. The method for summarizing financial document data based on artificial intelligence to correct OCR according to claim 9, wherein The core of the financial perception self-correction model is the adaptive gating attention mechanism, which realizes the selective extraction of information in different quality regions through dynamically adjusted attention distribution parameters, and the attention distribution parameters are controlled by the quality impact gating function.

Citation Information

Cited By

  • Document table extraction method and device, equipment and medium

    CN120877323A