Financial robot invoice identification and matching method based on deep reinforcement learning
Through deep reinforcement learning methods, text semantics, image structure and position information features are integrated to realize adaptive matching of invoice fields, solving the problems of low recognition accuracy and poor adaptability in the prior art, and improving the accuracy and efficiency of invoice recognition and matching.
Patent Information
- Application Number
- CN202510407180.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing invoice recognition and matching systems have low recognition accuracy in the face of non-standard expression, position misalignment or image blur, and lack of intelligent learning and adaptive adjustment, resulting in inaccurate matching results and difficult to meet the needs of high-frequency trading and fast settlement.
Using a deep reinforcement learning method, through image enhancement, feature extraction and multi-objective optimization model, combined with Dueling DQN reinforcement learning module, the adaptive matching strategy of invoice fields is realized, text semantics, image structure and location information features are integrated, and real-time dynamic adjustments are performed.
Maintaining high matching accuracy in non-standard expression, positional misalignment or image blurring significantly improves processing efficiency and robustness, and reduces matching error rates and manual intervention needs.
Smart Images

Figure CN120340044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of invoice recognition, and in particular to a financial robot invoice recognition and matching method based on deep reinforcement learning. Background Art
[0002] With the continuous improvement of the level of information technology in corporate financial management, invoices, as important documents for financial accounting and tax compliance, are increasingly in need of digital processing and automated matching. In the daily operations of enterprises, a large amount of invoice data collection, identification and matching processes are involved, which often require a lot of manual effort for verification and processing, resulting in low efficiency and high error rate, making it difficult to meet the business needs of high-frequency transactions and rapid settlement. Therefore, the development of an intelligent system that can automatically identify and match invoice fields has become an important research direction in the field of financial information processing.
[0003] At present, mainstream invoice recognition and matching systems mostly use methods based on OCR technology and template matching to convert invoice information in paper or image format into structured data for comparison and verification. They have achieved certain results in the initial realization of automation, but there are still many technical bottlenecks. For example, traditional OCR technology has low recognition accuracy when facing blurred invoice images, information occlusion or complex layouts; fixed template matching methods have poor adaptability to changes in invoice styles and have difficulty in handling the deep connection between structured information and unstructured images. In addition, in the field matching link, existing methods are mostly based on static rules or manually set similarity functions for judgment, lacking intelligent learning and adaptive adjustment mechanisms, resulting in matching results that are difficult to cover the changing actual scenarios. In the presence of multiple invoices, field misalignment, and content variation, inaccurate matching, omissions or duplications are prone to occur.
[0004] In addition, the existing invoice recognition and matching process is often separated into different steps and lacks a unified optimization framework. It is unable to comprehensively balance recognition accuracy, processing efficiency and result robustness at the global level. When processing large-scale invoice data, traditional methods cannot ensure matching accuracy while taking into account processing efficiency and stability, which in turn affects the automation level of financial processes and risk control capabilities. Summary of the invention
[0005] One object of the present invention is to propose a financial robot invoice recognition and matching method based on deep reinforcement learning. The present invention can maintain a high matching accuracy rate even when the field content has non-standard expressions, position misalignment or image blur.
[0006] A financial robot invoice recognition and matching method based on deep reinforcement learning according to an embodiment of the present invention comprises the following steps:
[0007] S1. Collect the original invoice data containing structured invoice data and unstructured invoice image data, and perform image enhancement, layout recognition, and text extraction on the original invoice data to form an initially recognized invoice dataset;
[0008] S3. Preprocess the initially recognized invoice dataset to generate a preprocessed invoice dataset;
[0009] S6. Extract features for each invoice field in the preprocessed invoice dataset, convert each invoice field into a unified high-dimensional feature vector, and construct an invoice field feature vector set;
[0010] S9. Establish a multi-objective matching optimization model for invoice fields based on the invoice field feature vector set;
[0011] S12. Use the differential evolution algorithm to perform a global search on the multi-objective matching optimization model for invoice fields, obtain a set of candidate invoice field matching solutions, and output a preliminary candidate matching solution;
[0012] S15. Take the set of candidate invoice field matching solutions as the state input, and use the Dueling DQN deep reinforcement learning module to perform real-time dynamic evaluation on the candidate matching solutions, make action decisions through state value and advantage evaluation, and generate an adaptively adjusted matching strategy;
[0013] S18. Perform fine-grained matching on the invoice fields in the preprocessed invoice dataset according to the adaptive matching strategy, realize real-time dynamic adjustment of the matching relationship of each invoice field, and output the final invoice field matching result.
[0014] Optionally, S1 includes the following steps:
[0015] S11. Obtain the original invoice dataset, where the original invoice dataset includes structured invoice field information D struct and unstructured invoice image information D image , and define the original invoice dataset D raw :
[0016]
[0017] where d i represents the original data of the i-th invoice, represents the structured invoice field, represents the corresponding unstructured invoice image, and N is the total number of invoices;
[0018] S12. Perform image enhancement processing on the unstructured invoice image information D image , process the unstructured invoice image, and obtain an enhanced invoice image dataset;
[0019] S13. Perform layout structure recognition on the enhanced invoice image dataset, identify the key areas in the invoice image, including the invoice number area, invoice date area, amount area, and commodity item area, and extract the structural layout mapping set of each invoice image;
[0020] S14. Perform region-level text extraction on the enhanced image data based on the structural layout mapping set, and extract the corresponding candidate field text T from each layout region i,j ;
[0021] S15. Perform preliminary combination matching on the structured invoice field information D struct and the corresponding candidate field text T i,j to construct the preliminarily recognized invoice data D pre represented by the combined invoice field of text and image:
[0022]
[0023] wherein, represents the preliminary recognition result of the i-th invoice, and K i is the total number of field regions extracted from the i-th invoice.
[0024] Optionally, the S2 includes the following steps:
[0025] S21. Perform image denoising on the unstructured image information in the preliminarily recognized invoice data D pre , apply an image filtering operator to the candidate field regions in each invoice image to remove image noise, and generate a set of invoice images after image denoising;
[0026] S22. Perform text standardization processing on the preliminarily recognized candidate field text set to convert the candidate field text into standard candidate field text
[0027] S23. Perform fusion processing on the structured invoice field information and the image data after image denoising and the standard candidate field text set to construct the preprocessed invoice dataset D prep :
[0028]
[0029] wherein, represents the data of the i-th invoice after image denoising and text standardization processing.
[0030] Optionally, the S3 includes the following steps:
[0031] S31. Based on the preprocessed invoice dataset D prep Perform text vectorization on the standard candidate field texts, and use a word embedding model to convert each candidate field text into a text feature vector of a fixed length
[0032]
[0033] Among them, represents the feature vector of the j-th candidate field text of the i-th invoice, and Embedding(·) is the word embedding mapping function;
[0034] S32. For the invoice image data after image denoising Perform image feature encoding, and use a convolutional neural network to encode each candidate field area image to generate an image feature vector
[0035]
[0036] Among them, represents the image feature vector of the j-th candidate field area of the i-th invoice, represents the j-th candidate field area of the i-th invoice after image denoising processing, and CNN(·) is the convolutional neural network encoding function;
[0037] S33. Extract the location information features of each candidate field area in the preprocessed invoice dataset, and generate the corresponding location information feature vector according to the coordinate information of the candidate field area in the invoice layout structure Generate the corresponding location information feature vector
[0038] S34. Through the vector splicing function, fuse the text feature vector The image feature vector And the location information feature vector Construct a unified high-dimensional feature vector for each invoice field
[0039] S35. Based on the unified high-dimensional feature vectors of each field of all invoices Construct an invoice field feature vector set F for the multi-object matching optimization model feature .
[0040] Optionally, the S4 includes the following steps:
[0041] S41. Based on the invoice field feature vector set F feature Define the input space of the multi-object matching optimization model as the set of feature combinations between all candidate fields:
[0042]
[0043] Among them, represents all field matching pairs in the invoice, and respectively represent the high-dimensional feature vectors of the jth and kth fields in the ith invoice;
[0044] S42. For the multi-objective characteristics in the field matching problem, construct a multi-objective optimization function and define the overall optimization objective as:
[0045] F = {f1, f2, f3};
[0046]
[0047] Among them, f1(x) represents the matching accuracy objective function, measuring the semantic and structural consistency of field pairs, f2(x) represents the matching efficiency objective function, measuring the computational and inference complexity, and f3(x) represents the matching robustness objective function, measuring the consistent matching performance of field pairs under noise perturbation, is any candidate field matching strategy, and w1, w2, w3 are the weighting coefficients of the multi-objective function;
[0048] S43. Model the matching accuracy objective function f1(x), define the semantic distance using the cosine similarity between field pairs, and combine the field position deviation penalty term:
[0049]
[0050] Among them, cos(·) represents the cosine similarity function, and Dist pos (j, k) represents the coordinate position distance between the jth and kth fields, and λ1 is the position deviation penalty factor;
[0051] S44. Model the matching efficiency objective function f2(x) and estimate the function based on the complexity and processing time of the matching strategy:
[0052] f2(x) = λ2·T(x) + λ3·C(x);
[0053] Among them, T(x) represents the average inference time of the current matching strategy, C(x) represents the computational complexity of this strategy, and λ2, λ3 are the corresponding weighting coefficients;
[0054] S45. Model the matching robustness objective function f3(x) and measure the output stability of the matching strategy under noise interference:
[0055]
[0056] Among them, x mDenote the output of the matching strategy after the m-th perturbation, denote the reference output before perturbation, M is the number of perturbation samplings, used to measure the stability performance of the strategy;
[0057] S46. Construct a multi-objective matching optimization model M for invoice fields based on a multi-objective optimization function match , the input of the multi-objective matching optimization model for invoice fields is the invoice field feature vector set F feature , the output of the multi-objective matching optimization model for invoice fields is the optimal solution x in all field matching pairs in the invoice ; * .
[0058] Optionally, the S5 includes the following steps:
[0059] S51. Based on the multi-objective matching optimization model M for invoice fields match and its set of optimization objective functions F, initialize the population of the differential evolution algorithm where M is the population size, denote the m-th individual, and each individual corresponds to an invoice field matching strategy;
[0060] S52. Perform a mutation operation on each generation of the population of the differential evolution algorithm to generate a mutant vector
[0061] S53. Perform a crossover operation based on the mutant vector, and use a hybrid crossover strategy to perform feature recombination between the original individual and the mutant vector to construct a trial individual
[0062]
[0063] where, denote the values of the k-th dimension of the trial vector, mutant vector, and original vector respectively, CR is the crossover probability, rand k is the k-th dimensional random value, k rand is the randomly selected dimension index;
[0064] S54. Based on the multi-objective fitness function, perform a multi-objective fitness comparison on the trial individual and the original individual , and retain the individual with better fitness to enter the next generation:
[0065]
[0066] S55. Repeat steps S52 to S54 until the maximum number of iterations G max or the population convergence condition is reached, and finally output a preliminary candidate matching scheme:
[0067]
[0068] Optionally, S6 includes the following steps:
[0069] S61. Use the preliminary candidate matching solution set X * as the state input of reinforcement learning, and define each candidate matching solution as a state The state space includes multi-dimensional features such as the text similarity of the matching pair, the image structure similarity, the position information deviation, and the objective function value;
[0070] S62. Construct a Dueling DQN reinforcement learning network structure, which is split into a state value function V(s t ) and an advantage function A(s t , a t ), and obtain the Q-value function through fusion:
[0071]
[0072] where Q(s t , a t ) represents the evaluation value of taking action a t in state s t , is the action space, including field replacement, field retention, and field reordering matching adjustment operations;
[0073] S63. Set the reward function r t , which is used to measure the impact on the matching performance after taking action a t in the current state. The reward function is defined according to the improvement amplitude of the matching accuracy, efficiency, and robustness metrics;
[0074] S64. Use the experience replay mechanism to record the state transition triple (s t , a t , r t , s t+1 ), construct a training sample set for the iterative training of the Dueling DQN network, and at the same time adopt a combination of a target network and an online network to stabilize the policy update;
[0075] S65. During the training process, select the optimal action in the current state based on the greedy policy:
[0076]
[0077] And update the current matching solution according to the selected action to form a new state s t+1 ;
[0078] Repeat steps S61 to S65 until the Dueling DQN network converges or reaches the maximum number of training rounds, and finally output the optimally matched policy set after adaptive adjustment:
[0079]
[0080] where, π θ represents the trained Dueling DQN policy network, s m is the state, θ is the network parameter, represents the finally matched invoice field policy after adaptive adjustment.
[0081] The beneficial effects of the present invention are:
[0082] (1) By constructing a unified high-dimensional feature vector of invoice fields, the present invention integrates text semantic features, image structure features, and field position information features, realizes the deep coupling expression of structured and unstructured information, and the multi-modal vector fusion method effectively breaks through the problems of incomplete information and semantic ambiguity caused by traditional methods that only rely on text matching or OCR structured results, significantly improves the modeling ability of semantic consistency between fields, and can still maintain a high matching accuracy in the case of non-standard expressions, position misalignment, or image blur of field contents.
[0083] (2) The present invention designs an optimization model including a three-dimensional objective function of matching accuracy, inference efficiency, and matching robustness, and uses the differential evolution algorithm for global search, so as to effectively jump out of the local optimum and obtain a more stable field matching policy. In dealing with large batches of high-complexity invoice scenarios, the high-adaptability population search mechanism of the differential evolution algorithm can effectively reduce the matching error rate and execution time.
[0084] (3) The present invention introduces the Dueling DQN reinforcement learning mechanism into the field of invoice field matching. By taking the candidate matching scheme as the environmental state and designing a structure that separates the state value function and the advantage function, the sensitivity of key variables in the matching policy is strengthened. During the training process, by dynamically evaluating the impact of the current matching policy on the matching performance and continuously optimizing the matching action based on the reinforcement learning feedback mechanism, the fine matching and real-time adjustment of field relationships are realized, and it has stronger generalization ability and significantly improved robustness in the face of field overlap, redundancy, or information missing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0086] Figure 1Flowchart of a method for invoice recognition and matching of a financial robot based on deep reinforcement learning proposed by the present invention. Detailed implementation manners
[0087] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0088] Refer to Figure 1 , a method for invoice recognition and matching of a financial robot based on deep reinforcement learning, comprising the following steps:
[0089] S1. Collect the original invoice data including structured invoice data and unstructured invoice image data, and perform image enhancement, layout recognition, and text extraction on the original invoice data to form an initially recognized invoice data set;
[0090] S2. Preprocess the initially recognized invoice data set to generate a preprocessed invoice data set;
[0091] S3. Extract features from each invoice field in the preprocessed invoice data set, convert each invoice field into a unified high-dimensional feature vector, and construct an invoice field feature vector set;
[0092] S4. Establish a multi-objective matching optimization model for invoice fields based on the invoice field feature vector set;
[0093] S5. Use the differential evolution algorithm to perform a global search on the multi-objective matching optimization model for invoice fields, obtain a set of candidate invoice field matching solution sets, and output a preliminary candidate matching solution;
[0094] S6. Use the candidate invoice field matching solution set as the state input, and adopt the Dueling DQN deep reinforcement learning module to perform real-time dynamic evaluation on the candidate matching solution, make action decisions through state value and advantage evaluation, and generate an adaptively adjusted matching strategy;
[0095] S7. Perform fine matching on the invoice fields in the preprocessed invoice data set according to the adaptive matching strategy, realize real-time dynamic adjustment of the matching relationship of each invoice field, and output the final invoice field matching result.
[0096] In this embodiment, S1 includes the following steps:
[0097] S11. Obtain the original invoice data set, where the original invoice data set includes structured invoice field information D struct and unstructured invoice image information D image , and define the original invoice data set D raw :
[0098]
[0099] Among them, d i represents the original data of the i-th invoice, represents the structured invoice fields, represents the corresponding unstructured invoice image, and N is the total number of invoices;
[0100] S12. Perform image enhancement processing on the unstructured invoice image information D image to obtain an enhanced invoice image dataset after processing the unstructured invoice image;
[0101] S13. Perform layout structure recognition on the enhanced invoice image dataset, identify the key areas in the invoice image, including the invoice number area, invoice date area, amount area, and commodity item area, and extract the structure layout mapping set of each invoice image;
[0102] S14. Perform region-level text extraction on the enhanced image data based on the structure layout mapping set, and extract the corresponding candidate field text T i,j ;
[0103] S15. Perform preliminary combination matching on the structured invoice field information D struct and the corresponding candidate field text T i,j to construct the preliminary recognized invoice data D pre represented by the combined graphic and text invoice fields:
[0104]
[0105] Among them, represents the preliminary recognition result of the i-th invoice, and K i is the total number of field areas extracted from the i-th invoice.
[0106] In this embodiment, S2 includes the following steps:
[0107] S21. Perform image denoising processing on the unstructured image information in the preliminary recognized invoice data D pre Apply an image filtering operator to the candidate field area in each invoice image to remove image noise and generate a set of invoice images after image denoising;
[0108] S22. Perform text standardization processing on the preliminary recognized candidate field text set to convert the candidate field text into standard candidate field text
[0109] S23. The structured invoice field information and the image data after image denoising and the standard candidate field text set perform a fusion process to construct the preprocessed invoice dataset D prep :
[0110]
[0111] wherein, represents the data of the i-th invoice after image denoising and text standardization processing.
[0112] In this embodiment, S3 includes the following steps:
[0113] S31. Based on the preprocessed invoice dataset D prep perform text vectorization processing on the standard candidate field text, and use a word embedding model to convert each candidate field text into a fixed-length text feature vector
[0114]
[0115] wherein, represents the feature vector of the j-th candidate field text of the i-th invoice, and Embedding(·) is a word embedding mapping function;
[0116] S32. Perform image feature encoding on the invoice image data after image denoising and use a convolutional neural network to encode each candidate field area image to generate an image feature vector
[0117]
[0118] wherein, represents the image feature vector of the j-th candidate field area of the i-th invoice, represents the j-th candidate field area of the i-th invoice after image denoising processing, and CNN(·) is a convolutional neural network encoding function;
[0119] S33. Extract location information features for each candidate field area in the preprocessed invoice dataset, and generate corresponding location information feature vectors according to the coordinate information of the candidate field area in the invoice layout structure
[0120] S34. Through a vector splicing function, fuse the text feature vector the image feature vector and the location information feature vector to construct a unified high-dimensional feature vector for each invoice field
[0121] S35. High-dimensional feature vectors unified for each field of all invoices Construct an invoice field feature vector set F for the multi-objective matching optimization model feature 。
[0122] In this embodiment, S4 includes the following steps:
[0123] S41. Based on the invoice field feature vector set F feature Define the input space of the multi-objective matching optimization model as the set of feature combinations between all candidate fields:
[0124]
[0125] where, represents all field matching pairs in the invoice, and respectively represent the high-dimensional feature vectors of the jth and kth fields in the ith invoice;
[0126] S42. For the multi-objective characteristics in the field matching problem, construct a multi-objective optimization function and define the overall optimization objective as:
[0127] F = {f1, f2, f3};
[0128]
[0129] where, f1(x) represents the matching accuracy objective function, measuring the semantic and structural consistency of the field pair, f2(x) represents the matching efficiency objective function, measuring the computational and inference complexity, and f3(x) represents the matching robustness objective function, measuring the consistent matching performance of the field pair under noise perturbation, is any candidate field matching strategy, and w1, w2, w3 are the weighted coefficients of the multi-objective function;
[0130] S43. Model the matching accuracy objective function f1(x), define the semantic distance using the cosine similarity between field pairs, and combine the field position deviation penalty term:
[0131]
[0132] where, cos(·) represents the cosine similarity function, Dist pos (j, k) represents the coordinate position distance between the jth and kth fields, and λ1 is the position deviation penalty factor;
[0133] S44. Model the matching efficiency objective function f2(x) and estimate the function based on the complexity and processing time of the matching strategy:
[0134] f2(x) = λ2·T(x) + λ3·C(x);
[0135] Wherein, T(x) represents the average inference time of the current matching strategy, C(x) represents the computational complexity of the strategy, and λ2, λ3 are corresponding weighting coefficients;
[0136] S45. Model the matching robustness objective function f3(x) to measure the output stability of the matching strategy under noise interference:
[0137]
[0138] where x m represents the output of the matching strategy after the m-th perturbation, represents the reference output before the perturbation, and M is the number of perturbation samplings, which is used to measure the stability performance of the strategy;
[0139] S46. Construct an invoice field multi-objective matching optimization model M based on the multi-objective optimization function match , the input of the invoice field multi-objective matching optimization model is the invoice field feature vector set F feature , and the output of the invoice field multi-objective matching optimization model is the optimal solution x in * .
[0140] In this embodiment, S5 includes the following steps:
[0141] S51. Based on the invoice field multi-objective matching optimization model M match and its optimization objective function set F, initialize the population of the differential evolution algorithm where M is the population size, represents the m-th individual, and each individual corresponds to an invoice field matching strategy;
[0142] S52. Perform a mutation operation on each generation of the population of the differential evolution algorithm to generate a mutant vector
[0143] S53. Perform a crossover operation based on the mutant vector, and use a hybrid crossover strategy to perform feature recombination between the original individual and the mutant vector to construct a trial individual
[0144]
[0145] where, respectively represent the values of the k-th dimension of the trial vector, the mutant vector, and the original vector, CR is the crossover probability, and rand k is the k-th dimensional random value, krand is a randomly selected dimension index;
[0146] S54. Based on the multi-objective fitness function, compare the fitness of the trial individuals and the original individuals in terms of multi-objective fitness, and retain the individuals with better fitness to enter the next generation:
[0147]
[0148] S55. Repeat steps S52 to S54 until the maximum number of iterations G max or the population convergence condition is reached, and finally output the preliminary candidate matching scheme:
[0149]
[0150] In this embodiment, S6 includes the following steps:
[0151] S61. Take the preliminary candidate matching scheme set X * as the state input of reinforcement learning, and define each candidate matching scheme as the state The state space includes multi-dimensional features such as the text similarity, image structure similarity, position information deviation, and objective function value of the matching pair;
[0152] S62. Construct a Dueling DQN reinforcement learning network structure, which is split into a state value function V(s t ) and an advantage function A(s t , a t ), and obtain the Q-value function through fusion:
[0153]
[0154] Among them, Q(s t , a t ) represents the evaluation value of taking action a t in state s t , is the action space, which includes operations such as field replacement, field retention, and field reordering matching adjustment;
[0155] S63. Set the reward function r t , which is used to measure the impact on the matching performance after taking action a t in the current state. The reward function is defined according to the improvement amplitude of the matching accuracy, efficiency, and robustness indicators;
[0156] S64. Use the experience replay mechanism to record the state transition triple (s t , a t,r t ,s t+1 ), construct a training sample set for iterative training of the DuelingDQN network, and use a combination of the target network and the online network to stabilize the strategy update;
[0157] S65. During the training process, the optimal action in the current state is selected based on the greedy strategy:
[0158]
[0159] And update the current matching scheme according to the selected action to form a new state s t+1 ;
[0160] S66. Repeat steps S61 to S65 until the DuelingDQN network converges or reaches the maximum number of training rounds, and finally output the optimal matching strategy set after adaptive adjustment:
[0161]
[0162] Among them, π θ Represents the DuelingDQN policy network after training, s m is the state, θ is the network parameter, Represents the final invoice field matching strategy after adaptive adjustment.
[0163] Embodiment 1:
[0164] On October 15, 2024, a cross-border e-commerce company located in the Science and Technology Park of Zone A discovered during its daily financial processing that the invoices entered in batches in early October had many matching anomalies and field duplication problems. The company's traditional invoice processing method was based on OCR recognition and manual secondary review. However, since most of the invoices involved that month were scanned and photographed copies provided by external logistics companies, the image quality was generally poor. At the same time, some invoice content was missing or handwritten, the matching accuracy of the traditional system dropped sharply, resulting in delays in financial processing.
[0165] In this context, the enterprise's financial information department decided to use the present invention to test the processing of some invoices to verify the system's processing capabilities in real scenarios.
[0166] The test data was selected from a total of 500 customs declaration and VAT general invoices from September 25 to October 5, 2024. All invoices were scanned and uploaded to the internal ERP system for archiving by three financial personnel every day. The image sources include: scanners (60%), mobile phone photos (35%), and fax scanned images (5%). Among them, 72 invoices had abnormal features such as tilt, shadow coverage, stamp pressure, and handwritten amounts.
[0167] In the method of the present invention, the system first performs image enhancement and denoising processing on each invoice. For example, in the invoice image numbered FP20241001 - 0342, the lower commodity item area is severely shaded and blocked, and the traditional OCR method cannot recognize the content of the "amount" field. However, through the combination strategy of image brightness enhancement + structural edge enhancement, the image clarity of this system is improved by 21%, and the OCR character confidence is increased from 0.42 to 0.88.
[0168] Subsequently, the system performs structural layout recognition on the invoice image and accurately divides 8 core areas including invoice number, invoice date, amount, capital amount, and purchaser information. In FP20241001 - 0342, the recognized "invoice date" is "September 29, 2024", and the OCR text extraction result is "September 2g, 2024". The traditional method misidentifies "9" as "g", resulting in a matching failure; while the present invention corrects this field to "September 29, 2024" adaptively through field - level semantic models and historical data learning, achieving a successful match.
[0169] In the process of field vectorization, the system performs triple - modal feature fusion on the "purchaser name" field: the text feature is "A Certain Technology Co., Ltd. in City A"; the image feature comes from the character - area image encoded by convolution, with a positioning confidence as high as 94%; the position - information feature is the rectangular coordinates (165, 323, 545, 357). Finally, the constructed field feature vector has a length of 256 dimensions, and the fusion accuracy is significantly higher than that of a single modality.
[0170] In the field matching stage, the system performs differential evolution optimization matching on a total of 3960 groups of fields in 500 invoices. For example, invoices numbered FP20241003 - 0115 and FP20241003 - 0116 are two invoices from the same supplier, with similar times and amounts. The traditional system repeatedly cross - matches the "amount" field between them. However, the present invention dynamically evaluates the semantic, position, and image - feature consistency of the two through DQN. Finally, it is determined that the amount of 0115 is "¥4,200.00" and that of 0116 is "¥2,100.00", achieving a successful match.
[0171] Among 10 Spanish - language invoices, for the invoice numbered FP20241004 - 0082, the traditional method cannot parse the fields of "Fecha de emisión" (invoice date) and "Importe total" (total amount), and the system field loss rate reaches 40%; while the present invention successfully recognizes that "Fecha de emisión" is "2024 / 09 / 30" through the support of multi - language embedding model training, and the matching - field confidence reaches 0.91.
[0172] The system processing results show that the total processing time for 500 invoices is 982 seconds, with an average of 1.96 seconds per invoice. The total number of accurately matched fields is 3,932, and the average accuracy rate reaches 96.4%. The comparison data is as follows:
[0173] Item Results of traditional method Results of the method of the present invention Total number of fields 3960 3960 Number of successfully matched fields 3213 3932 Matching accuracy rate 81.1% 96.4% Average processing time per single sheet 3.12 seconds 1.96 seconds Number of failed processes for abnormally imaged invoices 47 3 Loss rate of matching fields for foreign language invoices 42% 6% Number of duplicate matches / mismatched fields 88 11
[0174] Finally, in the invoice audit report system, the system generates a detailed report based on the matching confidence level and automatically marks the invoice items that require manual review (confidence level < 0.6 or field matching conflicts). On October 20, 2024, the system output 5 invoices for review. After verification by the financial staff, it was confirmed that 4 were due to OCR blur caused by unclear handwritten amounts from suppliers, and 1 was due to mistransmission of duplicate invoice numbers. The system prompted that all were valid.
[0175] This Example 1 truly demonstrates the recognition and matching capabilities of the system of the present invention in a complex, low-quality, mixed-language, and multi-template invoice environment. It not only significantly improves the automatic processing efficiency but also greatly reduces the matching error rate and the need for manual intervention, laying a solid foundation for large-scale financial automation.
[0176] The present invention constructs a unified high-dimensional feature vector for invoice fields, integrating text semantic features, image structure features, and field position information features, realizing the deep coupling expression of structured and unstructured information. The multi-modal vector fusion method effectively breaks through the problems of incomplete information and semantic ambiguity brought about by traditional methods that only rely on text matching or OCR structured results, significantly improving the modeling ability of semantic consistency between fields, and still maintaining a high matching accuracy rate in cases where the field content has non-standard expressions, position misalignments, or image blurs.
[0177] The present invention designs an optimization model that includes a three-dimensional objective function of matching accuracy rate, inference efficiency, and matching robustness, and uses the differential evolution algorithm for global search, thereby effectively jumping out of the local optimum and obtaining a more stable field matching strategy. In the scenario of processing a large number of high-complexity invoices, the high-adaptability population search mechanism of the differential evolution algorithm can effectively reduce the matching error rate and execution time.
[0178] The present invention introduces the Dueling DQN reinforcement learning mechanism into the field of invoice field matching. By taking the candidate matching scheme as the environmental state and designing a structure that separates the state value function and the advantage function, it strengthens the sensitivity to key variables in the matching strategy. During the training process, by dynamically evaluating the impact of the current matching strategy on the matching performance and continuously optimizing the matching actions based on the reinforcement learning feedback mechanism, it realizes the fine matching and real-time adjustment of field relationships, and has stronger generalization ability and significantly improved robustness in the face of scenarios of field overlap, redundancy, or information loss.
[0179] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A method for invoice recognition and matching of a financial robot based on deep reinforcement learning, characterized in that, It includes the following steps: S1. Collect the original invoice data including structured invoice data and unstructured invoice image data, and perform image enhancement, layout recognition, and text extraction on the original invoice data to form an initially recognized invoice data set; S2. Preprocess the initially recognized invoice data set to generate a preprocessed invoice data set; S3. Extract features for each invoice field in the preprocessed invoice data set, convert each invoice field into a unified high-dimensional feature vector, and construct an invoice field feature vector set; S4. Establish a multi-objective matching optimization model for invoice fields based on the invoice field feature vector set; S5. Use the differential evolution algorithm to perform a global search on the multi-objective matching optimization model for invoice fields, obtain a candidate invoice field matching solution set, and output a preliminary candidate matching solution; S6. Take the candidate invoice field matching solution set as the state input, use the Dueling DQN deep reinforcement learning module to perform real-time dynamic evaluation on the candidate matching solutions, make action decisions through state value and advantage evaluation, and generate an adaptively adjusted matching strategy; S7. Perform fine matching on the invoice fields in the preprocessed invoice data set according to the adaptive matching strategy, realize real-time dynamic adjustment of the matching relationship of each invoice field, and output the final invoice field matching result.
2. The financial robot invoice recognition and matching method based on deep reinforcement learning according to claim 1, characterized in that The S1 includes the following steps: S11. Obtain an original invoice dataset, where the original invoice dataset includes structured invoice field information D struct and unstructured invoice image information D image , and define the original invoice dataset D raw : Among them, d i represents the original data of the i-th invoice, represents the structured invoice fields, represents the corresponding unstructured invoice image, and N is the total number of invoices; S12. Process the unstructured invoice image information D image Perform image enhancement processing on the unstructured invoice image to obtain an enhanced invoice image dataset; S13. Perform layout structure recognition on the invoice image data set after image enhancement, identify the key areas in the invoice image, including the invoice number area, invoice date area, amount area, and commodity item area, and extract the structural layout mapping set of each invoice image; S14. Perform region-level text extraction on the enhanced image data based on the structural layout mapping set, and extract the corresponding candidate field text T from each layout region i,j ; S15. Combine the structured invoice field information D struct with the corresponding candidate field text T i,j for preliminary combined matching to construct the initially recognized invoice data D represented by the combined graphic and text invoice fields pre : Among them, represents the preliminary recognition result of the i-th invoice, and K i is the total number of field regions extracted from the i-th invoice.
3. A financial robot invoice recognition and matching method based on deep reinforcement learning according to claim 1, characterized in that The S2 includes the following steps: S21. Perform image denoising on the unstructured image information in the initially identified invoice data D pre Apply an image filtering operator to the candidate field areas in each invoice image to remove image noise and generate a set of invoice images after image denoising; S22. Perform text standardization processing on the candidate field text set after preliminary recognition, and convert the candidate field text into standard candidate field text S23. Combine the structured invoice field information with the image data after image denoising and the standard candidate field text set to perform a fusion process to construct the preprocessed invoice dataset D prep : Among them, represents the data of the i-th invoice after image denoising and text normalization processing.
4. A method for invoice recognition and matching of a financial robot based on deep reinforcement learning according to claim 1, characterized in that, The S3 includes the following steps: S31. Based on the preprocessed invoice dataset D prep Perform text vectorization on the standard candidate field texts, and use a word embedding model to convert each candidate field text into a text feature vector of a fixed length Among them, represents the feature vector of the j-th candidate field text of the i-th invoice, and Embedding(·) is the word embedding mapping function; The invoice image data after image denoising Perform image feature encoding, and use a convolutional neural network to encode each candidate field area image to generate an image feature vector Among them, represents the image feature vector of the j-th candidate field region of the i-th invoice, represents the j-th candidate field region of the i-th invoice after image denoising processing, and CNN(·) is a convolutional neural network encoding function; S33. Extract the location information features of each candidate field area in the preprocessed invoice dataset, and generate the corresponding location information feature vectors according to the coordinate information of the candidate field area in the invoice layout structure Generate the corresponding location information feature vectors S34. Fuse the text feature vectors through a vector concatenation function Image feature vector And the position information feature vector Construct a unified high-dimensional feature vector for each invoice field S35. High-dimensional feature vectors unified for each field of all invoices Construct an invoice field feature vector set F for the multi-objective matching optimization model feature .
5. A method for invoice recognition and matching of a financial robot based on deep reinforcement learning according to claim 1, characterized in that, The S4 includes the following steps: S41. Based on the invoice field feature vector set F feature Define the input space of the multi-objective matching optimization model as the set of feature combinations between all candidate fields: Among them, represents all field matching pairs in the invoice, and respectively represent the high-dimensional feature vectors of the j-th and k-th fields in the i-th invoice; S42. For the multi-objective characteristics in the field matching problem, construct a multi-objective optimization function, and define the overall optimization objective as: F = {f1, f2, f3}; Among them, f1(x) represents the matching accuracy objective function, which measures the semantic and structural consistency of field pairs, f2(x) represents the matching efficiency objective function, which measures the computing and inference complexity, and f3(x) represents the matching robustness objective function, which measures the consistent matching performance of field pairs under noise perturbation. is an arbitrary candidate field matching strategy, and w1, w2, and w3 are the weighting coefficients of the multi-objective function; S43. Model the matching accuracy objective function f1(x), use the cosine similarity between field pairs to define the semantic distance, and combine the field position deviation penalty term: where cos(·) represents the cosine similarity function, and Dist pos (j, k) represents the coordinate position distance between the j-th and k-th fields, and λ1 is the position deviation penalty factor; S44. Model the matching efficiency objective function f2(x), and estimate the function based on the complexity and processing time of the matching strategy: f2(x) = λ2·T(x) + λ3·C(x); where, T(x) represents the average inference time of the current matching strategy, C(x) represents the computational complexity of this strategy, and λ2, λ3 are the corresponding weighting coefficients; S45. Model the matching robustness objective function f3(x) to measure the output stability of the matching strategy under noise interference: Among them, x m represents the output of the matching strategy after the m-th perturbation, represents the reference output before perturbation, and M is the number of perturbation samplings, which is used to measure the stability performance of this strategy; S46. Construct an invoice field multi-objective matching optimization model M based on a multi-objective optimization function match , the input of the invoice field multi-objective matching optimization model is the invoice field feature vector set F feature , the output of the invoice field multi-objective matching optimization model is the optimal solution x in all field matching pairs in the invoice * . 6. The method for invoice recognition and matching of a financial robot based on deep reinforcement learning according to claim 1, wherein The S5 includes the following steps: S51. Optimize model M for multi-objective matching of invoice fields match And its set of optimization objective functions F, initialize the population of the differential evolution algorithm Where M is the population size, represents the m-th individual, and each individual corresponds to an invoice field matching strategy; S52. Perform a mutation operation on each generation of the population of the differential evolution algorithm to generate a mutant vector. S53. Perform crossover operations based on the mutation vector, and use a hybrid crossover strategy to perform feature recombination between the original individual and the mutation vector to construct a trial individual Among them, respectively represent the values of the k-th dimension of the test vector, the mutant vector and the original vector. CR is the crossover probability, and rand k is the random value of the k-th dimension, and k rand is the randomly selected dimension index; S54. Compare the multi-objective fitness of the trial individuals and the original individuals to perform multi-objective fitness comparison, and retain the individuals with better fitness to enter the next generation: Repeat steps S52 to S54 until the maximum number of iterations G max or the population convergence condition is reached, and finally output the preliminary candidate matching scheme:
7. A method for invoice recognition and matching of a financial robot based on deep reinforcement learning according to claim 1, characterized in that, The S6 includes the following steps: S61. Take the preliminary candidate matching solution set X * as the state input of reinforcement learning, and define each candidate matching solution as a state The state space includes multi-dimensional features such as the text similarity of matching pairs, the image structure similarity, the position information deviation, and the objective function value; S62. Construct the Dueling DQN reinforcement learning network structure, which is split into a state value function V(s t ) and an advantage function A(s t , a t ). The Q-value function is obtained by fusion: Among them, Q(s t , a t ) represents the evaluation value of taking action a t in state s t . is the action space, including operations such as field replacement, field retention, and field reordering matching adjustment; S63. Set the reward function r t , which is used to measure the impact on the matching performance when taking action a t after that, and the reward function is defined according to the improvement amplitude of the matching accuracy, efficiency, and robustness metrics; S64. Record state transition triples (s t , a t , r t , s t+1 ) using the experience replay mechanism, construct a training sample set for the iterative training of the Dueling DQN network, and at the same time adopt a combination of a target network and an online network to stabilize policy updates; S65. During the training process, select the optimal action in the current state based on the greedy strategy: And update the current matching scheme according to the selected action to form a new state s t+1 ; S66. Repeat steps S61 to S65 until the Dueling DQN network converges or reaches the maximum number of training rounds, and finally output the optimally adjusted adaptive matching strategy set: where, π θ represents the trained Dueling DQN policy network, s m is the state, and θ is the network parameter, represents the final invoice field matching policy after adaptive adjustment.
Citation Information
Cited By
Dynamic fault-tolerant invoice duplicate checking system and method based on multi-mode OCR and mixed index
CN121074899A