Financial robot invoice element identification method based on semantic extraction

Through semantic potential modeling and context encoding technology, combined with the DeBERTa model, the accuracy and stability problems of invoice element recognition in complex layout scenarios in the existing technology are solved, and efficient structured processing of invoice images is achieved.

CN120708239AInactive Publication Date: 2025-09-26LIANYUNGANG GUOTU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510915119.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When faced with complex layouts and diverse invoice scenarios, existing invoice element recognition methods lack methods to effectively integrate graphic and text spatial information with contextual semantic features, resulting in low recognition accuracy and poor stability, making it difficult to achieve accurate extraction.

Method used

A financial robot invoice element recognition method based on semantic extraction is adopted. Through semantic potential modeling and context encoding technology, a semantic core generation module, a potential field construction module, a semantic aggregation module and an element decoding module are constructed. Combined with the DeBERTa model, a deep understanding and element aggregation of invoice images are achieved.

Benefits of technology

It improves the accuracy and stability of structured recognition in complex scenarios, enhances the ability to model the semantic relationship between character units and invoice element categories, improves the robustness and adaptability of the recognition process, and can effectively deal with non-standard invoices such as multi-field mixing, text overlapping, and image and text interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708239A_ABST
    Figure CN120708239A_ABST
Patent Text Reader

Abstract

The invention discloses a financial robot invoice element identification method based on semantic extraction, and the method comprises the following steps: S1, obtaining an original invoice image, and carrying out the image preprocessing; s2, executing optical character recognition operation, and extracting invoice text information; s3, inputting a semantic potential model; s4, constructing a semantic kernel vector set according to preset invoice element categories; s5, generating a potential tensor field based on the semantic kernel vector set; s6, performing iterative semantic migration operation on the character units in the potential tensor field to form a semantic clustering region; s7, calculating a comprehensive confidence score, and outputting an invoice element recognition result; and S8, performing field legality verification on the invoice element identification result, and submitting the invoice element identification result to a financial robot system after verification is passed to drive related business processes. According to the method, semantic potential modeling and context coding technologies are fused, invoice elements are accurately extracted, and the method has the advantages of being clear in structure, high in robustness and high in adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent financial technology, and in particular to a financial robot invoice element recognition method based on semantic extraction. Background Art

[0002] With the continuous improvement of the level of financial informatization in enterprises, automatic recognition and structured processing of invoice images are playing an increasingly important role in processes such as financial reimbursement, account management, and tax comparison. Existing invoice element recognition methods mainly rely on optical character recognition technology to extract the text on the invoice, and then combine regular expressions, template matching, or rule-based text parsing methods to try to extract key information such as invoice codes, amounts, and dates from unstructured text. Such methods are applicable to certain scenarios in which invoices have clear structures and standardized layouts. However, when the invoice format is complex, the text layout is diverse, or there are problems such as character adhesion and unclear printing, the recognition accuracy drops significantly, and the stability of the structured output results is poor.

[0003] To overcome the problem of traditional methods' strong dependence on layout structure, some studies have introduced deep learning models to identify invoice element fields from optical character recognition results through natural language processing technologies such as named entity recognition. Although this type of method enhances the modeling ability of semantic information, most models still only take linear text sequences as input, ignoring the spatial arrangement relationship of text in the image.

[0004] Existing technologies generally lack a method to effectively integrate image and text spatial information with contextual semantic features, and are unable to accurately extract feature fields in complex layout scenarios in invoice images. At the same time, traditional models lack an explicit semantic aggregation mechanism, making it difficult to simulate the clustering judgment process of "semantic belonging" in manual recognition, resulting in severe discretization of extraction results and low structural quality.

[0005] Therefore, it is urgent to propose a new invoice element recognition method that can simultaneously consider the character context semantics, image spatial layout and field category prior knowledge, so as to improve the accuracy and stability of semantic understanding and meet the needs of intelligent financial systems for high-quality invoice structured data. Summary of the Invention

[0006] One purpose of the present invention is to propose a method for invoice element recognition of a financial robot based on semantic extraction. The present invention integrates semantic potential modeling and context coding technology, fully utilizes the semantic relationship and spatial distribution characteristics between characters, and realizes a deep understanding and element aggregation of unstructured invoice text. It has the advantages of clear structure, strong robustness and high adaptability, can effectively improve the automated processing capability and recognition accuracy of the financial robot, and is suitable for a variety of complex invoice scenarios.

[0007] According to an embodiment of the present invention, a method for recognizing invoice elements by a financial robot based on semantic extraction includes the following steps:

[0008] S1. Obtaining the original invoice image and performing image preprocessing on the invoice image;

[0009] S2. Perform optical character recognition on the preprocessed invoice image to extract the corresponding invoice text information, generate an original text sequence, and perform denoising to obtain a cleaned text sequence;

[0010] S3, inputting the cleaned text sequence into a semantic potential model, wherein the semantic potential model includes a semantic core generation module, a potential field construction module, a semantic aggregation module, and an element decoding module;

[0011] S4. Constructing a semantic core vector set according to preset invoice element categories in a semantic core generation module;

[0012] S5. In the potential field construction module, based on the semantic kernel vector set and the cleaned text sequence, a potential tensor field containing semantic gradient information is generated;

[0013] S6. Performing an iterative semantic shift operation on the character units in the potential tensor field through the semantic aggregation module, so that the character units gradually aggregate toward the most matching semantic core along the semantic gradient to form a semantic clustering area;

[0014] S7. Calculate the comprehensive confidence score of the semantic clustering area through the element decoding module and output the structured invoice element recognition result;

[0015] S8. Perform field legitimacy verification on the structured invoice element recognition results. After verification, submit them to the financial robot system to drive related business processes.

[0016] Optionally, the image preprocessing includes image denoising, rotation correction, edge enhancement and size normalization.

[0017] Optionally, the S2 specifically includes:

[0018] S21, receiving the invoice image after image preprocessing as the input image of the OCR recognition model;

[0019] S22, performing text region detection on the input image, identifying text segments in the image, and outputting a set of candidate text regions including position information;

[0020] S23, cropping the image of each candidate text region and uniformly adjusting it to the standard input size required by the OCR recognition model;

[0021] S24, inputting the processed candidate text region image into the OCR recognition model, extracting the character sequence corresponding to each region, and generating the recognition confidence of each character;

[0022] S25. Sort the recognition results in row and column order and align their positions according to the spatial layout of the candidate text regions in the image to generate a complete original text sequence, where the original text sequence represents the text content recognized in the entire invoice image;

[0023] S26. Filter the recognition confidence of each character in the original text sequence, remove characters with recognition confidence lower than a set threshold, and output a denoised cleaned text sequence.

[0024] Optionally, the semantic core generation module generates a semantic core vector corresponding to each invoice element category and constructs a semantic core vector set. The potential field construction module models the semantic association relationship between each character unit and the semantic core into a potential tensor field. The semantic aggregation module drives the character units to gradually move closer to the most matching semantic core to form a semantic clustering area. The element decoding module parses the structured invoice element recognition results from the semantic clustering area and calculates the comprehensive confidence score.

[0025] Optionally, the S4 specifically includes:

[0026] S41. Receive preset invoice element categories, including invoice code, invoice number, invoice date, amount, tax rate, buyer information, and seller information;

[0027] S42. Extracting a representative set of field samples corresponding to the target invoice element category from a historical set of structured invoice samples that have been annotated with invoice semantic categories. The representative set of field samples is a set of text with expressive features selected from historical real invoices and manually annotated data for each preset invoice element category. This set of text is used to enhance the semantic core generation model's ability to understand the context of the target field category.

[0028] S43. Input each sample in the representative field sample set into the pre-trained DeBERTa model to extract a contextual semantic vector for each sample, wherein the contextual semantic vector has the ability to simultaneously express character semantics and position dependency;

[0029] S44. Calculate the semantic core vector of the invoice element category by using the mean aggregation method of all contextual semantic vectors:

[0030]

[0031] Among them, k j The semantic core vector representing the jth invoice element category, n jrepresents the number of samples under the jth invoice element category, v j,m Represents the context semantic vector of the mth sample;

[0032] S45. All semantic core vectors are formed into a set, and provided as a semantic core vector set to the potential field construction module, so as to guide the character units to aggregate toward the semantic core direction.

[0033] Optionally, the S5 specifically includes:

[0034] S51, receiving a semantic kernel vector set and a cleaned text sequence in a potential field construction module;

[0035] S52, using the pre-trained DeBERTa model to semantically encode the cleaned text sequence, mapping each character unit in the cleaned text sequence into a contextual semantic vector;

[0036] S53. Based on the Euclidean distance between the context semantic vector and the semantic core vector, a semantic attraction strength function is constructed:

[0037]

[0038] Among them, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, h i Represents the contextual semantic vector of the character unit, k j represents the semantic kernel vector of the j-th invoice element, ‖·‖ represents the Euclidean norm operation, and σ represents the potential attenuation control parameter, which is used to adjust the distribution shape of the attraction range and gradient response;

[0039] S54, constructing a potential tensor field by combining the semantic attraction strengths between all character units and all semantic core vectors;

[0040] The potential tensor field represents the semantic attraction strength of each character unit relative to different semantic core vectors.

[0041] Optionally, the S6 specifically includes:

[0042] S61, receiving a potential tensor field in a semantic aggregation module;

[0043] S62, obtaining the coordinate position of each character unit based on the OCR recognition model, and constructing a coordinate position set;

[0044] S63. After the potential tensor field is generated, the response center of each semantic kernel vector in the image space is calculated as the reference coordinate position. The reference coordinate position and the coordinate position of the character unit are both two-dimensional vectors, and the horizontal coordinate and the vertical coordinate are calculated independently:

[0045]

[0046] Among them, c j Represents the reference coordinate position of the semantic kernel vector in the image space, Z j represents the normalization factor, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, p i Represents the two-dimensional coordinate position of the character unit, and N represents the total number of character units;

[0047] S64. Constructing a semantic core response map based on the cosine similarity between semantic core vectors, wherein the semantic core response map describes the semantic relevance between different invoice element categories and enhances the cooperative attraction ability between cores in the potential field;

[0048] S65. Based on the potential tensor field, the semantic kernel response map, and the two-dimensional coordinate position set, the semantic offset direction under coupling drive is calculated for each character unit. The semantic offset direction is a two-dimensional vector:

[0049]

[0050] Among them, g i Indicates the semantic offset direction of the character unit, c l represents the reference coordinate position of the semantic kernel vector in the image space, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, Ω j,l represents the semantic kernel response graph, describing the semantic relevance and coupling relationship between semantic kernels, and M represents the number of semantic kernel vectors;

[0051] S66. Based on the context window of the character unit in the cleaned text sequence, the pre-trained DeBERTa model is used to calculate the semantic similarity between the character unit and its context character unit, and a context consistency gating factor is generated. The context consistency gating factor is used to adjust the context credibility of the semantic shift:

[0052]

[0053] Among them, β i represents the context consistency gating factor, W represents the context window size, cos(e i ,e j ) represents the semantic similarity between character units, h i and h j Contextual semantic vectors representing different character units, generated by the pre-trained DeBERTa model;

[0054] S67. Set the semantic offset step size, and update the character unit position based on the semantic offset step size and the context consistency gating factor:

[0055] p i ' =p i +α·β i ·g i ;

[0056] Among them, p i ' represents the updated position of the character unit after the current round of iteration, α represents the semantic offset step parameter, β i represents the context consistency gating factor;

[0057] S68, repeating the semantic shift operation from S65 to S67, and recalculating the reference coordinate position after each round of iteration, dynamically updating the semantic core response map and the context consistency gating factor, until the position change of all character units is less than the set position change threshold or the number of iterations reaches the upper limit;

[0058] S69, taking the final stabilized character unit position set as input, performing a spatial clustering operation based on position distribution, and dividing into a number of semantic clustering areas;

[0059] The spatial clustering operation includes:

[0060] After multiple rounds of iterative semantic shifting, the two-dimensional position of each character unit will gradually approach the most matching semantic core position, forming multiple semantic direction aggregation flows;

[0061] After the final round of iteration, the final semantic core is determined for each character unit, and a semantic label set is generated;

[0062] All character units belonging to the same semantic core are clustered according to their spatial positions to form a set of two-dimensional points.

[0063] Then the minimum bounding rectangle method is used to generate closed semantic clustering areas for the two-dimensional point set.

[0064] Optionally, it is characterized in that the S7 specifically includes:

[0065] S71. For each semantic cluster region, extract all character units contained therein, and splice the contents in the order of the characters from left to right and from top to bottom in the image to generate candidate field contents. Simultaneously, determine the field category based on the semantic labels corresponding to the semantic cluster regions, and use the position set of the character units as the field position. The three together constitute a candidate field triple, which includes field content, field category, and field position.

[0066] S72. Based on the character unit information contained in the candidate field triples, the recognition confidence of the OCR recognition model of the character and the semantic attraction strength of the character unit relative to the semantic core are respectively counted, and a weighted sum is performed. The score is normalized by the number of character units to obtain a comprehensive confidence score.

[0067] S73. Screen the candidate field triples based on a preset confidence threshold, and eliminate recognition results whose comprehensive confidence scores are less than the preset confidence threshold.

[0068] S74. Integrate all the field triplets that pass the screening into a structured invoice element recognition result.

[0069] The beneficial effects of the present invention are:

[0070] The present invention provides a method for invoice element recognition of a financial robot based on semantic extraction. In response to the problems of insufficient semantic modeling, missing spatial information and poor robustness of field recognition in the existing technology, a new recognition framework that integrates semantic kernel generation, potential tensor modeling and semantic offset aggregation is proposed, which improves the accuracy and stability of structured recognition in complex scenarios. By introducing a semantic potential model, the semantic relationship between character units and invoice element categories is explicitly modeled, so that the recognition process not only relies on the semantic expression of the characters themselves, but also can use the representative semantic kernels of the fields within the class for semantic attraction, thereby enhancing the ability of semantic attribution judgment.

[0071] In addition, the present invention constructs a semantic gradient-driven potential tensor field and designs an iterative semantic shift mechanism to simulate the aggregation process of character units toward their most matching field categories, thereby realizing self-organizing clustering of spatial semantics and improving the boundary accuracy of character classification. Compared with traditional NER methods, this method does not rely on fixed formats and rule bases, and still maintains high robustness and adaptability when facing non-standard invoices such as multi-field mixing, text overlap, and image and text interference. Combined with the contextual semantic encoding capability provided by the DeBERTa pre-training model, the model's ability to distinguish fine-grained semantic differences between fields is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0073] Figure 1 This is a flowchart of a method for recognizing invoice elements of a financial robot based on semantic extraction proposed by the present invention;

[0074] Figure 2This is a schematic diagram of the offset iteration process performed by character units in the semantic aggregation module of the financial robot invoice element recognition method based on semantic extraction proposed by the present invention;

[0075] Figure 3 This is a schematic diagram of the spatial clustering and semantic clustering area generation of a financial robot invoice element recognition method based on semantic extraction proposed in the present invention. DETAILED DESCRIPTION

[0076] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0077] refer to Figure 1-3 , a financial robot invoice element recognition method based on semantic extraction, including the following steps:

[0078] S1. Obtaining the original invoice image and performing image preprocessing on the invoice image;

[0079] S2. Perform optical character recognition on the preprocessed invoice image to extract the corresponding invoice text information, generate an original text sequence, and perform denoising to obtain a cleaned text sequence;

[0080] S3, inputting the cleaned text sequence into a semantic potential model, wherein the semantic potential model includes a semantic core generation module, a potential field construction module, a semantic aggregation module, and an element decoding module;

[0081] S4. Constructing a semantic core vector set according to preset invoice element categories in a semantic core generation module;

[0082] S5. In the potential field construction module, based on the semantic kernel vector set and the cleaned text sequence, a potential tensor field containing semantic gradient information is generated;

[0083] S6. Performing an iterative semantic shift operation on the character units in the potential tensor field through the semantic aggregation module, so that the character units gradually aggregate toward the most matching semantic core along the semantic gradient to form a semantic clustering area;

[0084] S7. Calculate the comprehensive confidence score of the semantic clustering area through the element decoding module and output the structured invoice element recognition result;

[0085] S8. Perform field legitimacy verification on the structured invoice element recognition results. After verification, submit them to the financial robot system to drive related business processes.

[0086] The present invention proposes a financial robot invoice element recognition method based on semantic extraction. By constructing a semantic potential model, it realizes the complete processing flow from original image acquisition to structured invoice element results. This method innovatively integrates OCR recognition and semantic vector-guided clustering mechanism, effectively overcoming the shortcomings of traditional methods in field mis-entry, omission and ambiguous recognition, and can achieve highly robust analysis of invoices with complex formats and diverse layouts, providing accurate invoice element input for financial automation processing systems.

[0087] In this embodiment, the image preprocessing includes image denoising, rotation correction, edge enhancement and size normalization.

[0088] The present invention designs a multi-step operation of denoising, rotation correction, size normalization and edge enhancement in the image preprocessing stage, which improves the character recognition accuracy and stability of the OCR model. By normalizing the input image quality, it eliminates interference factors such as scanning skew and shooting shadows, laying a unified and clear basic environment for subsequent character recognition and semantic analysis.

[0089] In this embodiment, S2 specifically includes:

[0090] S21, receiving the invoice image after image preprocessing as the input image of the OCR recognition model;

[0091] S22, performing text region detection on the input image, identifying text segments in the image, and outputting a set of candidate text regions including position information;

[0092] S23, cropping the image of each candidate text region and uniformly adjusting it to the standard input size required by the OCR recognition model;

[0093] S24, inputting the processed candidate text region image into the OCR recognition model, extracting the character sequence corresponding to each region, and generating the recognition confidence of each character;

[0094] S25. Sort the recognition results in row and column order and align their positions according to the spatial layout of the candidate text regions in the image to generate a complete original text sequence, where the original text sequence represents the text content recognized in the entire invoice image;

[0095] S26. Filter the recognition confidence of each character in the original text sequence, remove characters with recognition confidence lower than a set threshold, and output a denoised cleaned text sequence.

[0096] The present invention improves the accuracy and anti-interference ability of text information extraction by refining the OCR processing flow, covering sub-steps such as text area detection, area standardization, and character confidence screening. It generates clean text sequences through spatial layout sorting and confidence filtering, ensuring that the data input to the semantic model is structured and highly confident, providing high-quality input for subsequent semantic potential modeling, and effectively reducing misrecognition and semantic offset problems.

[0097] In this embodiment, the semantic core generation module generates a semantic core vector corresponding to each invoice element category and constructs a semantic core vector set. The potential field construction module models the semantic association relationship between each character unit and the semantic core into a potential tensor field. The semantic aggregation module drives the character unit to gradually move closer to the most matching semantic core to form a semantic clustering area. The element decoding module parses the structured invoice element recognition results from the semantic clustering area and calculates the comprehensive confidence score.

[0098] In this embodiment, the S4 specifically includes:

[0099] S41. Receive preset invoice element categories, including invoice code, invoice number, invoice date, amount, tax rate, buyer information, and seller information;

[0100] S42. Extracting a representative set of field samples corresponding to the target invoice element category from a historical set of structured invoice samples that have been annotated with invoice semantic categories. The representative set of field samples is a set of text with expressive features selected from historical real invoices and manually annotated data for each preset invoice element category. This set of text is used to enhance the semantic core generation model's ability to understand the context of the target field category.

[0101] S43. Input each sample in the representative field sample set into the pre-trained DeBERTa model to extract a contextual semantic vector for each sample, wherein the contextual semantic vector has the ability to simultaneously express character semantics and position dependency;

[0102] S44. Calculate the semantic core vector of the invoice element category by using the mean aggregation method of all contextual semantic vectors:

[0103]

[0104] Among them, k j The semantic core vector representing the jth invoice element category, n j represents the number of samples under the jth invoice element category, v j,m Represents the context semantic vector of the mth sample;

[0105] S45. All semantic core vectors are formed into a set, and provided as a semantic core vector set to the potential field construction module, so as to guide the character units to aggregate toward the semantic core direction.

[0106] In the semantic core generation stage, the present invention makes full use of the DeBERTa model to extract contextual semantic vectors, constructs representative field semantic cores for each type of invoice element, and introduces a mean aggregation mechanism to improve semantic representativeness. By introducing manually labeled samples to enhance semantic expression capabilities, each semantic core has category discriminability and context generalization capabilities, providing a solid semantic coordinate basis for potential modeling and semantic shift, and improving the accuracy and generalization robustness of element aggregation.

[0107] In this embodiment, the S5 specifically includes:

[0108] S51, receiving a semantic kernel vector set and a cleaned text sequence in a potential field construction module;

[0109] S52, using the pre-trained DeBERTa model to semantically encode the cleaned text sequence, mapping each character unit in the cleaned text sequence into a contextual semantic vector;

[0110] S53. Based on the Euclidean distance between the context semantic vector and the semantic core vector, a semantic attraction strength function is constructed:

[0111]

[0112] Among them, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, h i Represents the contextual semantic vector of the character unit, k j represents the semantic kernel vector of the j-th invoice element, ‖·‖ represents the Euclidean norm operation, and σ represents the potential attenuation control parameter, which is used to adjust the distribution shape of the attraction range and gradient response;

[0113] S54, constructing a potential tensor field by combining the semantic attraction strengths between all character units and all semantic core vectors;

[0114] The potential tensor field represents the semantic attraction strength of each character unit relative to different semantic core vectors.

[0115] The present invention designs an attraction function based on the DeBERTa semantic vector and the Euclidean distance of the semantic kernel vector in the potential field construction module, and constructs a semantic tensor field, realizing fine-grained semantic modeling between character units and multiple semantic categories. This method realizes the continuous expression of semantic intensity. Compared with traditional classification methods, it can more accurately capture the association strength between characters and target elements, provide gradient guidance for subsequent aggregation, and improve the precision and interpretability of the recognition process.

[0116] In this embodiment, S6 specifically includes:

[0117] S61, receiving a potential tensor field in a semantic aggregation module;

[0118] S62, obtaining the coordinate position of each character unit based on the OCR recognition model, and constructing a coordinate position set;

[0119] S63. After the potential tensor field is generated, the response center of each semantic kernel vector in the image space is calculated as the reference coordinate position. The reference coordinate position and the coordinate position of the character unit are both two-dimensional vectors, and the horizontal coordinate and the vertical coordinate are calculated independently:

[0120]

[0121] Among them, c j Represents the reference coordinate position of the semantic kernel vector in the image space, Z j represents the normalization factor, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, p i Represents the two-dimensional coordinate position of the character unit, and N represents the total number of character units;

[0122] S64. Constructing a semantic core response map based on the cosine similarity between semantic core vectors, wherein the semantic core response map describes the semantic relevance between different invoice element categories and enhances the cooperative attraction ability between cores in the potential field;

[0123] S65. Based on the potential tensor field, the semantic kernel response map, and the two-dimensional coordinate position set, the semantic offset direction under coupling drive is calculated for each character unit. The semantic offset direction is a two-dimensional vector:

[0124]

[0125] Among them, g i Indicates the semantic offset direction of the character unit, c l represents the reference coordinate position of the semantic kernel vector in the image space, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, Ω j,l represents the semantic kernel response graph, describing the semantic relevance and coupling relationship between semantic kernels, and M represents the number of semantic kernel vectors;

[0126] S66. Based on the context window of the character unit in the cleaned text sequence, the pre-trained DeBERTa model is used to calculate the semantic similarity between the character unit and its context character unit, and a context consistency gating factor is generated. The context consistency gating factor is used to adjust the context credibility of the semantic shift:

[0127]

[0128] Among them, β i represents the context consistency gating factor, W represents the context window size, cos(e i ,e j ) represents the semantic similarity between character units, h i and h j Contextual semantic vectors representing different character units, generated by the pre-trained DeBERTa model;

[0129] S67. Set the semantic offset step size, and update the character unit position based on the semantic offset step size and the context consistency gating factor:

[0130] p i ' =p i +α·β i ·g i ;

[0131] Among them, p i ' represents the updated position of the character unit after the current round of iteration, α represents the semantic offset step parameter, β i represents the context consistency gating factor;

[0132] S68, repeating the semantic shift operation from S65 to S67, and recalculating the reference coordinate position after each round of iteration, dynamically updating the semantic core response map and the context consistency gating factor, until the position change of all character units is less than the set position change threshold or the number of iterations reaches the upper limit;

[0133] S69, taking the final stabilized character unit position set as input, performing a spatial clustering operation based on position distribution, and dividing into a number of semantic clustering areas;

[0134] The spatial clustering operation includes:

[0135] After multiple rounds of iterative semantic shifting, the two-dimensional position of each character unit will gradually approach the most matching semantic core position, forming multiple semantic direction aggregation flows;

[0136] After the final round of iteration, the final semantic core is determined for each character unit, and a semantic label set is generated;

[0137] All character units belonging to the same semantic core are clustered according to their spatial positions to form a set of two-dimensional points.

[0138] Then the minimum bounding rectangle method is used to generate closed semantic clustering areas for the two-dimensional point set.

[0139] The present invention designs a semantic shift mechanism, which iteratively promotes the aggregation of character units towards semantic cores in the potential field, and dynamically updates their positions in combination with the context consistency gating factor, effectively realizing semantic aggregation. At the same time, it introduces reference coordinates and semantic core response maps to strengthen the coupling between cores, ensuring that the semantic migration path has directionality and contextual coherence, thereby improving the stability, accuracy and boundary clarity of field clustering.

[0140] Through spatial clustering operations, the character positions after iterative semantic shifting are integrated to form semantic clustering areas, and the minimum enclosing rectangle is used to construct closed boundaries to improve the geometric stability and computational efficiency of the area. This operation makes full use of the relationship between character distribution and semantic trend, making the field division more in line with the actual visual layout. It can effectively deal with invoice images with uneven character spacing and mixed layout, and improve the accuracy and compatibility of structured field area positioning.

[0141] In this embodiment, it is characterized in that the S7 specifically includes:

[0142] S71. For each semantic cluster region, extract all character units contained therein, and splice the contents in the order of the characters from left to right and from top to bottom in the image to generate candidate field contents. Simultaneously, determine the field category based on the semantic labels corresponding to the semantic cluster regions, and use the position set of the character units as the field position. The three together constitute a candidate field triple, which includes field content, field category, and field position.

[0143] S72. Based on the character unit information contained in the candidate field triples, the recognition confidence of the OCR recognition model of the character and the semantic attraction strength of the character unit relative to the semantic core are respectively counted, and a weighted sum is performed. The score is normalized by the number of character units to obtain a comprehensive confidence score.

[0144] S73. Screen the candidate field triples based on a preset confidence threshold, and eliminate recognition results whose comprehensive confidence scores are less than the preset confidence threshold.

[0145] S74. Integrate all the field triplets that pass the screening into a structured invoice element recognition result.

[0146] The present invention constructs field triplets and combines recognition confidence with semantic attraction strength for weighted fusion, outputs information structure fields with clear semantic categories and positions, and sets confidence thresholds to implement candidate field screening. This mechanism realizes fine decoding from region to field, avoids content redundancy and label errors, improves the reliability and availability of structured recognition results, and enhances the data accuracy and business flow automation capabilities of the financial system.

[0147] Example 1:

[0148] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the financial shared service center of a certain enterprise group, and application verification is carried out on the value-added tax special invoices, ordinary invoices and electronic invoices that are centrally processed in its routine operations. The enterprise processes an average of more than 300 invoices per day, covering a variety of bill formats issued by suppliers in different industries. The text layout, font style and information expression are extremely diverse, and there are problems such as complex layout structure, irregular field distribution, high OCR recognition error, inaccurate element extraction, etc., which result in a huge workload for manual verification, low efficiency and a certain risk of accounting errors.

[0149] In actual deployment, the financial robot invoice element recognition method based on semantic extraction proposed in the present invention is embedded between the company's automatic reimbursement system and tax control system. First, by deploying the image acquisition terminal, the paper invoice is scanned or photographed and uploaded to the server. The image preprocessing module completes image denoising, rotation correction, edge enhancement, and size normalization to ensure image quality. Subsequently, the image content is recognized through a deep learning OCR model, and the character information is organized into an original text sequence. The system further performs denoising to generate a clean text sequence, excludes low-confidence characters, and improves the accuracy of subsequent processing.

[0150] The core recognition stage is completed through the semantic potential model. The system constructs a set of semantic core vectors corresponding to the invoice element category based on historical annotation data, encodes the character context semantics with the help of the pre-trained DeBERTa model, constructs the potential tensor field between the character unit and the semantic core, and performs multiple rounds of iterative semantic shift in the semantic field, so that the character unit converges to the optimal semantic clustering area under the dual guidance of semantic attraction and context consistency gating factors. The system then uses a spatial clustering method to form a closed semantic boundary, and completes the field content splicing, category judgment and confidence score output through the element decoding module to form a structured invoice field result.

[0151] To verify the performance of the present invention in practice, a comparison is made with the existing mainstream invoice element recognition methods. The comparison results are shown in Table 1.

[0152] Table 1 Overview of performance comparison of intelligent invoice element recognition methods

[0153]

[0154]

[0155] In this recognition performance comparison experiment, three current mainstream invoice element recognition methods were selected, namely the structural analysis method based on the traditional rule engine, the semantic recognition method based on the bidirectional encoding model, and the "financial robot invoice element recognition method based on semantic extraction" proposed in this invention. The comparison dimensions cover eight indicators including recognition accuracy, field recall rate, field position deviation, character error recognition rate, semantic consistency score, scenario adaptability, structured output integrity and processing delay.

[0156] From the comparison of accuracy and recall, it can be seen that the semantic potential model adopted in the present invention has advantages in feature positioning and category recognition. Compared with the traditional method relying on fixed templates and character regular inference, the semantic attraction mechanism constructs a more stable matching relationship between character units and semantic cores, and the recognition accuracy performance is better. At the same time, the DeBERTa contextual semantic encoder enhances the model's ability to capture contextual semantics, thereby improving the field recall rate and maintaining a high recognition quality even in invoices with interference factors such as blur, occlusion, and line breaks.

[0157] In terms of field position deviation and character error recognition rate, the present invention drives character aggregation through semantic offset operations, and performs minimum circumscribed rectangle fitting in combination with spatial position distribution, effectively alleviating the problems of character recognition position drift and semantic mismatch. The field position deviation control is more precise, and the character error recognition rate is significantly reduced.

[0158] In addition, in the dimension of semantic consistency scoring, the present invention constructs a contextual consistency gating mechanism to dynamically adjust the semantic credibility during the semantic aggregation process, thereby improving the accuracy and consistency of semantic clustering; traditional methods are usually unable to handle ambiguous information across fields, resulting in weak performance in this item.

[0159] This invention also demonstrates excellent adaptability to various scenarios, supporting various formats such as VAT invoices, general invoices, and electronic invoices without requiring pre-set templates. Regarding structured output integrity, a candidate field triple and comprehensive confidence screening mechanism effectively ensures the integrity and stability of output data.

[0160] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A financial robot invoice element recognition method based on semantic extraction, characterized by: The steps include: S1. Obtaining the original invoice image and performing image preprocessing on the invoice image; S2. Perform optical character recognition on the preprocessed invoice image to extract the corresponding invoice text information, generate an original text sequence, and perform denoising to obtain a cleaned text sequence; S3, inputting the cleaned text sequence into a semantic potential model, wherein the semantic potential model includes a semantic core generation module, a potential field construction module, a semantic aggregation module, and an element decoding module; S4. Constructing a semantic core vector set according to preset invoice element categories in a semantic core generation module; S5. In the potential field construction module, based on the semantic kernel vector set and the cleaned text sequence, a potential tensor field containing semantic gradient information is generated; S6. Performing an iterative semantic shift operation on the character units in the potential tensor field through the semantic aggregation module, so that the character units gradually aggregate toward the most matching semantic core along the semantic gradient to form a semantic clustering area; S7. Calculate the comprehensive confidence score of the semantic clustering area through the element decoding module and output the structured invoice element recognition result; S8. Perform field legitimacy verification on the structured invoice element recognition results. After verification, submit them to the financial robot system to drive related business processes.

2. A financial robot invoice element recognition method based on semantic extraction according to claim 1, characterized in that: The image preprocessing includes image denoising, rotation correction, edge enhancement and size normalization.

3. The method for recognizing invoice elements of a financial robot based on semantic extraction according to claim 1 is characterized in that: The S2 specifically includes: S21, receiving the invoice image after image preprocessing as the input image of the OCR recognition model; S22, performing text region detection on the input image, identifying text segments in the image, and outputting a set of candidate text regions including position information; S23, cropping the image of each candidate text region and uniformly adjusting it to the standard input size required by the OCR recognition model; S24, inputting the processed candidate text region image into the OCR recognition model, extracting the character sequence corresponding to each region, and generating the recognition confidence of each character; S25. Sort the recognition results in row and column order and align their positions according to the spatial layout of the candidate text regions in the image to generate a complete original text sequence, where the original text sequence represents the text content recognized in the entire invoice image; S26. Filter the recognition confidence of each character in the original text sequence, remove characters with recognition confidence lower than a set threshold, and output a denoised cleaned text sequence.

4. The method for recognizing invoice elements of a financial robot based on semantic extraction according to claim 1 is characterized in that: The semantic core generation module generates a semantic core vector corresponding to each invoice element category and constructs a set of semantic core vectors. The potential field construction module models the semantic association relationship between each character unit and the semantic core into a potential tensor field. The semantic aggregation module drives the character units to gradually move closer to the most matching semantic core to form a semantic clustering area. The element decoding module parses the structured invoice element recognition results from the semantic clustering area and calculates the comprehensive confidence score.

5. The method for recognizing invoice elements of a financial robot based on semantic extraction according to claim 1 is characterized in that: The S4 specifically includes: S41. Receive preset invoice element categories, including invoice code, invoice number, invoice date, amount, tax rate, buyer information, and seller information; S42. Extracting a representative set of field samples corresponding to the target invoice element category from a historical set of structured invoice samples that have been annotated with invoice semantic categories. The representative set of field samples is a set of text with expressive features selected from historical real invoices and manually annotated data for each preset invoice element category. This set of text is used to enhance the semantic core generation model's ability to understand the context of the target field category. S43. Input each sample in the representative field sample set into the pre-trained DeBERTa model to extract a contextual semantic vector for each sample, wherein the contextual semantic vector has the ability to simultaneously express character semantics and position dependency; S44. Calculate the semantic core vector of the invoice element category by using the mean aggregation method of all context semantic vectors; S45. All semantic core vectors are formed into a set, and provided as a semantic core vector set to the potential field construction module.

6. The method for recognizing invoice elements of a financial robot based on semantic extraction according to claim 1 is characterized in that: The S5 specifically includes: S51, receiving a semantic kernel vector set and a cleaned text sequence in a potential field construction module; S52, using the pre-trained DeBERTa model to semantically encode the cleaned text sequence, mapping each character unit in the cleaned text sequence into a contextual semantic vector; S53. Constructing a semantic attraction strength function based on the Euclidean distance between the context semantic vector and the semantic core vector; S54, constructing a potential tensor field by combining the semantic attraction strengths between all character units and all semantic core vectors; The potential tensor field represents the semantic attraction strength of each character unit relative to different semantic core vectors.

7. The method for recognizing invoice elements of a financial robot based on semantic extraction according to claim 1 is characterized in that: The S6 specifically includes: S61, receiving a potential tensor field in a semantic aggregation module; S62, obtaining the coordinate position of each character unit based on the OCR recognition model, and constructing a coordinate position set; S63. After the potential tensor field is generated, the response center of each semantic kernel vector in the image space is calculated as the reference coordinate position; S64. Constructing a semantic core response map based on the cosine similarity between semantic core vectors, wherein the semantic core response map describes the semantic relevance between different invoice element categories and enhances the cooperative attraction ability between cores in the potential field; S65. Based on the potential tensor field, the semantic kernel response map, and the two-dimensional coordinate position set, the semantic offset direction under coupling drive is calculated for each character unit. The semantic offset direction is a two-dimensional vector: Among them, g i Indicates the semantic offset direction of the character unit, c l represents the reference coordinate position of the semantic kernel vector in the image space, Φ i,j represents the semantic attraction strength of the character unit relative to the semantic kernel vector, Ω j,l represents the semantic kernel response graph, describing the semantic relevance and coupling relationship between semantic kernels, and M represents the number of semantic kernel vectors; S66. Calculate the semantic similarity between the character unit and its context character unit based on the context window of the character unit in the cleaned text sequence using the pre-trained DeBERTa model to generate a context consistency gating factor, wherein the context consistency gating factor is used to adjust the context credibility of the semantic shift; S67. Set the semantic offset step size, and update the character unit position based on the semantic offset step size and the context consistency gating factor: p i ' =p i +a·b i ·g i ; Among them, p i ' represents the updated position of the character unit after the current round of iteration, α represents the semantic offset step parameter, β i represents the context consistency gating factor; S68, repeating the semantic shift operation from S65 to S67, and recalculating the reference coordinate position after each round of iteration, dynamically updating the semantic core response map and the context consistency gating factor, until the position change of all character units is less than the set position change threshold or the number of iterations reaches the upper limit; S69, taking the final stabilized character unit position set as input, performing a spatial clustering operation based on position distribution, and dividing into a number of semantic clustering areas; The spatial clustering operation includes: After multiple rounds of iterative semantic shifting, the two-dimensional position of each character unit will gradually approach the most matching semantic core position, forming multiple semantic direction aggregation flows; After the final round of iteration, the final semantic core is determined for each character unit, and a semantic label set is generated; All character units belonging to the same semantic core are clustered according to their spatial positions to form a set of two-dimensional points. Then the minimum bounding rectangle method is used to generate closed semantic clustering areas for the two-dimensional point set.

8. The method for recognizing invoice elements of a financial robot based on semantic extraction according to claim 1 is characterized in that: The S7 specifically includes: S71. For each semantic cluster region, extract all character units contained therein, and splice the contents in the order of the characters from left to right and from top to bottom in the image to generate candidate field contents. Simultaneously, determine the field category based on the semantic labels corresponding to the semantic cluster regions, and use the position set of the character units as the field position. The three together constitute a candidate field triplet. S72. Based on the character unit information contained in the candidate field triples, the recognition confidence of the OCR recognition model of the character and the semantic attraction strength of the character unit relative to the semantic core are respectively counted, and a weighted sum is performed. The score is normalized by the number of character units to obtain a comprehensive confidence score. S73. Screen the candidate field triples based on a preset confidence threshold, and eliminate recognition results whose comprehensive confidence scores are less than the preset confidence threshold. S74. Integrate all the field triplets that pass the screening into a structured invoice element recognition result.

Citation Information

Cited By

  • Digital instrument character image analysis method

    CN121392815A

  • Electronic bill block chain data analysis system and method based on big data

    CN121526619A

  • Image correction effect evaluation method and device and electronic equipment

    CN121725320A

  • Image correction effect evaluation method, device and electronic equipment

    CN121725320B