Electronic report bill generation method and device and storage medium

By using deep learning network models and semantic analysis technology, the problems of inaccurate image recognition of invoices and insufficient risk identification in the financial reimbursement system have been solved, realizing efficient and accurate electronic reimbursement generation and intelligent risk management.

CN120876124APending Publication Date: 2025-10-31CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510908867.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing financial reimbursement systems suffer from poor resolution when recognizing blurry or obscured document images, resulting in inaccurate electronic reimbursements. Furthermore, they lack in-depth analysis of text data and risk identification capabilities, leading to high labor costs, low efficiency, and insufficient risk control.

Method used

A deep learning network model based on attention convolutional layers is adopted, combined with multi-layer encoders and decoders. Key features of invoice images are extracted through attention mechanism and semantic analysis is performed to generate electronic invoices. At the same time, dynamic risk rules are designed for risk warning.

Benefits of technology

It improved the accuracy of invoice image recognition, generated accurate electronic expense reports, reduced labor costs, and enhanced the intelligent management and control capabilities of financial risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876124A_ABST
    Figure CN120876124A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an electronic report bill generation method and device and a storage medium. The method comprises the steps that a bill image and a bill template are input to a first network model, the first network model comprises an input layer, a plurality of first convolution layers, an attention convolution layer, a plurality of second convolution layers and an output layer, the input layer is used for receiving the bill image, the plurality of first convolution layers are used for extracting first features of the bill image, and the attention convolution layer is used for extracting second features of the bill image; the attention convolution layer is used for receiving the bill image and the bill template, obtaining a first position feature and a second position feature according to the bill image and the bill template, and obtaining a second feature according to the first feature, the first position feature and the second position feature, and the plurality of second convolution layers are used for obtaining text data of the bill image according to the second feature, the output layer is used for outputting text data; and generating an electronic bill according to the text data. According to the method, the bill image can be accurately identified to obtain accurate text data, so that an accurate electronic bill is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to methods, apparatus and storage media for generating electronic bills. Background Technology

[0002] With the development of information technology, financial reimbursement systems have emerged. On the one hand, these systems use optical character recognition (OCR) technology to convert image data such as invoice images into editable and processable text data. On the other hand, based on the recognition results of OCR technology, these systems can automatically generate electronic reimbursement forms.

[0003] However, when using OCR technology to recognize invoice images, there are problems such as poor recognition resolution and recognition errors when recognizing partially blurred or occluded invoice images, which in turn leads to inaccurate electronic bills generated from the recognized text data. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for generating electronic expense reports, which can accurately identify invoice images to obtain precise text data, thereby generating accurate electronic expense reports. This enables intelligent generation of electronic expense reports and reduces financial processing costs. The following describes this disclosure from different aspects; it should be understood that the implementation methods and beneficial effects of the different aspects described below can be referenced interchangeably.

[0005] In a first aspect, this application provides a method for generating electronic bills. The method includes: inputting a bill image and a bill template into a first network model, the first network model comprising an input layer, multiple first convolutional layers, an attention convolutional layer, multiple second convolutional layers, and an output layer. The input layer receives the bill image; the multiple first convolutional layers extract a first feature from the bill image; the attention convolutional layer receives the bill image and the bill template; the attention convolutional layer includes a first branch and a second branch; the first branch obtains a first positional feature based on the bill image and the bill template, the first positional feature being a positional feature where the bill image and the bill template match; the second branch obtains a second positional feature based on the bill image and the bill template, the second positional feature being a positional feature where the bill image and the bill template do not match; the attention convolutional layer further obtains a second feature based on the first feature, the first positional feature, and the second positional feature; the multiple second convolutional layers obtain text data of the bill image based on the second feature; and the output layer outputs the text data. Based on the text data, an electronic bill is generated.

[0006] According to the above technical means, the invoice image is input into a first network model, and the first feature is extracted through multiple first convolutional layers. Then, the invoice image and invoice template are input into an attention convolutional layer. Based on the first and second branches of the attention convolutional layer, first and second positional features are extracted, representing the matching and non-matching cases of the invoice image and invoice template, respectively. The attention convolutional layer fuses the first feature, the first positional feature, and the second positional feature to obtain the second feature, which is then input into multiple second convolutional layers to obtain the text data of the invoice image. Compared to single-mode localization, extracting the key region positions in the invoice image through the first and second branches allows for case-by-case identification of key regions in the invoice image, improving the accuracy of text data recognition and generating accurate electronic invoices.

[0007] In one possible implementation, the attention convolutional layer further includes a first activation layer and a second activation layer. The first activation layer is connected to the output of the first branch and is used to obtain a first attention feature based on a first positional feature and a first weight function. The second activation layer is connected to the output of the second branch and is used to obtain a second attention feature based on a second positional feature and a second weight function. The attention convolutional layer is also used to fuse the first attention feature, the second attention feature, and the first feature to obtain a second feature, and input the second feature into multiple second convolutional layers to obtain text data.

[0008] Based on the aforementioned techniques, the attention convolutional layer, combined with an attention mechanism, transforms the first positional feature into a first attention feature and the second positional feature into a second attention feature. The first attention feature, the second attention feature, and the first feature are then fused to obtain the second feature, from which text data is derived. By introducing an attention mechanism, the first network model can identify key features of the ticket image and mask non-key features, thereby improving the model's recognition accuracy.

[0009] In one possible implementation, the above-mentioned "generating an electronic bill of exchange based on text data" includes: generating a vector representation of the text data; and inputting the vector representation into a second network model, wherein the second network model includes a multi-layer encoder and a multi-layer decoder, the multi-layer encoder being used to extract features from the vector representation, and the multi-layer decoder being used to generate an electronic bill of exchange based on the features output by the multi-layer encoder.

[0010] Based on the aforementioned technical methods, text data is transformed into computable semantic vectors by analyzing the semantic relationships between text data to generate vector representations. A second network model is then used to train these semantic vectors. This second network model accurately understands the semantics within the invoice image and automatically fills in the expense report content based on semantic relationships, generating an electronic expense report. This electronic expense report generation method provides users with an intelligent and personalized expense report experience.

[0011] In one possible implementation, the method further includes: displaying the risk type of the electronic expense report, wherein the risk type includes one or more of the following: duplicate expense report, expense report exceeding the standard, expense report spanning multiple periods, or sensitive supplier.

[0012] Based on the aforementioned technical means, the second network model automatically identifies risk types, accurately locates risks by setting risk thresholds, and enhances financial risk management capabilities.

[0013] In one possible implementation, the method further includes: acquiring business information, which includes a first business, a second business, or a third business, wherein the first business is a business in which the invoice image matches the invoice template, the second business is a business in which the invoice image does not match the invoice template and the standardization result of the invoice image is a first value, and the third business is a business in which the invoice image does not match the invoice template and the standardization result of the invoice image is a second value, wherein the first value indicates that the invoice image has been standardized and the second value indicates that the invoice image cannot be standardized; and preprocessing the invoice image according to the business information.

[0014] Based on the aforementioned technical means, the invoice images are preprocessed in a targeted manner for three different business processes, enabling the preprocessed invoice images to more accurately identify text data.

[0015] In one possible implementation, the aforementioned "preprocessing of the ticket image according to business information" includes: in a first business scenario, performing tilt correction and smoothing denoising processing on the ticket image; in a second business scenario, dividing the ticket image to obtain multiple first sub-regions, generating a first local grayscale histogram for each first sub-region, and correcting the brightness of the ticket image based on each first local grayscale histogram; in a third business scenario, performing tilt correction and smoothing denoising processing on the ticket image, dividing the ticket image after tilt correction and smoothing denoising processing to obtain multiple second sub-regions, generating a second local grayscale histogram for each second sub-region, correcting the brightness of the ticket image after tilt correction and smoothing denoising processing based on each second local grayscale histogram, and performing image binarization processing on the brightness-corrected ticket image.

[0016] Based on the aforementioned technical means, local adjustments are made to bill images that do not match the bill template, achieving precise optimization of the brightness and contrast of various local areas of such bill images, thus giving the bill images stronger adaptability and correction effects.

[0017] Secondly, this application provides an apparatus for generating electronic bills. The apparatus includes: an input unit for inputting a bill image into a first network model, the first network model including an input layer, multiple first convolutional layers, an attention convolutional layer, multiple second convolutional layers, and an output layer, wherein the input layer is used to receive the bill image, the multiple first convolutional layers are used to extract a first feature from the bill image, the attention convolutional layer is used to receive the bill image and a bill template, the attention convolutional layer includes a first branch and a second branch, the first branch is used to obtain a first positional feature based on the bill image and the bill template, the first positional feature being a positional feature that matches the bill image and the bill template, the second branch is used to obtain a second positional feature based on the bill image and the bill template, the second positional feature being a positional feature that does not match the bill image and the bill template, the attention convolutional layer is also used to obtain a second feature based on the first feature, the first positional feature, and the second positional feature, the multiple second convolutional layers are used to obtain text data of the bill image based on the second feature, and the output layer is used to output the text data; and a generation unit for generating an electronic bill based on the text data.

[0018] In one possible implementation, the attention convolutional layer further includes a first activation layer and a second activation layer. The first activation layer is connected to the output of the first branch and is used to obtain a first attention feature based on a first positional feature and a first weight function. The second activation layer is connected to the output of the second branch and is used to obtain a second attention feature based on a second positional feature and a second weight function. The attention convolutional layer is also used to fuse the first attention feature, the second attention feature, and the first feature to obtain a second feature. The input unit is also used to input the second feature into multiple second convolutional layers to obtain text data.

[0019] In one possible implementation, the generation unit is also used to generate a vector representation of the text data; the vector representation is input into a second network model, wherein the second network model includes a multi-layer encoder and a multi-layer decoder, the multi-layer encoder is used to extract features from the vector representation, and the multi-layer decoder is used to generate an electronic bill based on the features output by the multi-layer encoder.

[0020] In one possible implementation, the generation unit is also used to display the risk type of the electronic expense report, wherein the risk type includes one or more of the following: duplicate expense report, expense report exceeding the standard, expense report across periods, or sensitive supplier.

[0021] In one possible implementation, the input unit is further configured to acquire business information, which includes a first business, a second business, or a third business. The first business is a business where the invoice image matches the invoice template; the second business is a business where the invoice image does not match the invoice template and the standardization result of the invoice image is a first value; and the third business is a business where the invoice image does not match the invoice template and the standardization result of the invoice image is a second value. The first value indicates that the invoice image has been standardized, and the second value indicates that the invoice image cannot be standardized. Based on the business information, the invoice image is preprocessed.

[0022] In one possible implementation, the input unit is further configured to: perform tilt correction and smoothing denoising processing on the ticket image in a first business scenario; divide the ticket image into multiple first sub-regions in a second business scenario, generate a first local grayscale histogram for each first sub-region, and correct the brightness of the ticket image based on each first local grayscale histogram; and in a third business scenario, perform tilt correction and smoothing denoising processing on the ticket image, divide the ticket image after tilt correction and smoothing denoising processing into multiple second sub-regions, generate a second local grayscale histogram for each second sub-region, correct the brightness of the ticket image after tilt correction and smoothing denoising processing based on each second local grayscale histogram, and perform image binarization processing on the brightness-corrected ticket image.

[0023] Thirdly, this application provides an electronic billing statement generation apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement the method described in the first aspect and any of its possible embodiments.

[0024] Fourthly, this application provides a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by the processor of an electronic billing generation apparatus, enables the electronic billing generation apparatus to perform the methods described in the first aspect and any of its possible embodiments.

[0025] Fifthly, this application provides a computer program product containing instructions that, when run on an electronic billing generation apparatus, enables the electronic billing generation apparatus to perform the methods described in the first aspect and any of its possible implementations. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of a system applied in an embodiment of this application;

[0027] Figure 2 This is an exemplary flowchart of a method for generating an electronic bill of exchange, as applied in an embodiment of this application.

[0028] Figure 3This is a schematic flowchart illustrating the key area localization performed by the first and second branches in the embodiments of this application;

[0029] Figure 4 This is a schematic diagram illustrating a feature extraction method applied in an embodiment of this application;

[0030] Figure 5 This is a schematic diagram illustrating the document image preprocessing method for three business scenarios applied in the embodiments of this application;

[0031] Figure 6 This is a schematic diagram of the structure of an electronic billing system used in an embodiment of this application;

[0032] Figure 7 This is a schematic diagram of another electronic billing generation device used in the embodiments of this application. Detailed Implementation

[0033] To facilitate understanding of the embodiments of this disclosure, the following points will be explained before introducing the embodiments of this disclosure.

[0034] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Furthermore, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0035] The terms "first" and "second," etc., used in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.

[0036] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0037] It should be noted that in the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0038] In financial management, expense reimbursement is fundamental to financial work, with companies processing a large number of paper and electronic invoices daily. In the information age, as companies increasingly prioritize their IT infrastructure and development, various management information systems have emerged to address complex and ever-changing business scenarios. Financial reimbursement systems are one such type of management information system. To shorten settlement time for business personnel and improve approval efficiency, automated identification of paper and electronic invoices and intelligent filling of reimbursement forms have become key aspects of platform development.

[0039] With increasing market demand, many financial systems have adopted OCR technology for recognizing invoice images and text information. OCR technology is widely used across various fields, demonstrating significant advantages in convenience and accuracy. Examples include license plate recognition in smart parking lots, and the recognition and input of information such as ID cards, bank cards, financial instruments, and financial statements. Compared to manual data entry, OCR technology not only significantly improves the structuring of invoice information and reimbursement efficiency, reducing the error rate that may occur during manual data entry, but also effectively saves labor costs. With the continuous growth of diversified market demands and the widespread adoption of smart mobile terminals, the application of OCR technology in financial reimbursement has also ushered in new development opportunities. The industry has made customized improvements and applications to OCR technology, which has, to some extent, promoted the intelligentization and efficiency of workflows.

[0040] The original images used in OCR technology typically come from specialized acquisition devices such as scanners and document cameras, which capture images of very high quality. However, because these devices are fixed to office desks, their convenience and flexibility are limited, and their application scenarios are restricted to desktops. In addition, OCR technology mostly uses image processing methods such as binarization, connected component analysis, and projection analysis, as well as statistical machine learning methods like support vector machines (SVM) to recognize images. While these methods can recognize text data from images, they still suffer from problems such as text misalignment, blurring, and poor resolution when recognizing certain key information or occluded / blurred data, resulting in highly unstable recognition performance. For example, external factors such as shooting angle and lighting can cause a series of problems in the original image, including shadow occlusion, angular distortion, and out-of-focus blurring. Correspondingly, some engineering methods to solve these problems have emerged. In some embodiments, in addition to using OCR technology to recognize the text data of the invoice, Fourier transform, Radon transform, Hough transform, scale-invariant feature transform (SIFT), speed-up robust features (SURF), and various convolution operators can be used for feature extraction and mapping to enhance the recognition effect of the invoice image. However, these methods still have problems such as poor recognition resolution and recognition errors, and do not solve the risk of rework by business personnel due to incorrectly recognized invoice images.

[0041] In financial systems, besides recognizing invoice images to obtain invoice text data, it's also necessary to draft expense reports based on known invoice information. Currently, invoice information is uploaded as image attachments to the front-end financial business system. The invoice verification process in this type of expense report is largely manual. Data entry involves manually inputting the paper text of the invoices into the business system, or directly uploading electronic invoices as image attachments. The invoice information in the system still needs to be manually filled in and supplemented. This method suffers from high labor costs, slow speed, and a high risk of errors, and it cannot utilize intelligent technology to improve the efficiency of expense report drafting and processing.

[0042] The existing reimbursement system relies more on users manually filling out reimbursement forms. Due to the large amount of content in the reimbursement forms and the transcription of a large amount of invoice information, manual operation is not only time-consuming and labor-intensive, but also prone to input errors, resulting in a slow reimbursement process and increasing financial processing costs and time costs. At present, some reimbursement systems adopt intelligent reimbursement form drafting to solve the problems of manual reimbursement form drafting. In some embodiments, intelligent reimbursement form drafting adopts the following technical means: (1) In terms of data processing, the processing of invoice information is only simple text recognition and structured storage. For example, basic OCR technology is used to convert the text on the invoice image into the text data of the invoice, and then the text data of the invoice is directly stored in the database according to a fixed format. There is a lack of deeper analysis and processing of the text data of the invoice, and no multimodal information fusion and semantic mining are involved. In terms of entity recognition and relationship processing, rule-based matching is used to identify a small number of key entities. For example, only simple string matching is used to extract information such as amount and date, which makes it difficult to build complex semantic relationships between entities. This method is limited to keyword search and does not convert text into semantic vectors for in-depth analysis. (2) Regarding the implementation of the bill drafting function, a static template filling method is adopted, and users manually fill in the various contents of the bill. The business system only provides basic format verification and does not have the ability to automatically predict and intelligently fill.

[0043] During the drafting of expense reports, business systems need to identify potential risks in textual data, promptly and accurately identify these risks, and alert users. In some embodiments, the financial expense reporting system uses a rule engine to identify risks in its risk rule verification method. The rule engine uses a static setting approach, pre-setting simple rule sets, such as amount limits and date ranges, to verify expense report data. These rule sets are relatively fixed and singular, making it difficult to update them in a timely manner according to policy changes and the diversity of business scenarios. When reimbursement standards are adjusted or new business types emerge, static rules cannot accurately identify the corresponding risks; for example, they cannot identify the compliance of reimbursement for special expenses in new businesses, leading to loopholes in risk control. In some embodiments, the business system's rule engine can only perform superficial data comparisons, lacking an understanding of data semantics and unable to effectively identify hidden risks, such as semantic-level risks of duplicate reimbursements. Furthermore, when existing business systems detect risks, they can only provide general error messages, failing to offer specific and targeted guidance. That is, they cannot combine specific business scenarios and data semantics to explain in detail the causes of the risks and provide accurate remedial suggestions. Users need to spend additional time communicating with finance personnel for confirmation, affecting expense reporting efficiency and accuracy.

[0044] Therefore, this application aims to solve the problem of how to automatically and accurately recognize user-entered invoice images and other information in the financial reimbursement system, and then fill in the recognized text data to generate an electronic reimbursement form to simplify the process. Additionally, the verification of the generated electronic reimbursement form is also a problem that this application needs to address.

[0045] To address the aforementioned issues, this application proposes a method for generating electronic expense reports. This method provides a more accurate invoice image recognition technology and performs semantic analysis on the recognized text data to achieve intelligent prediction and filling of expense reports. Furthermore, the electronic expense report generation method proposed in this application also designs reasonable risk rules to achieve risk warnings and provide remedial suggestions.

[0046] For ease of understanding, the method for generating electronic expense reports provided in this application will be described in detail below with reference to the accompanying drawings.

[0047] Figure 1 This is a schematic diagram of a system applied in an embodiment of this application. The electronic billing method provided in this embodiment can be applied to this system. Figure 1 As shown, the system 100 includes terminals 101 and 102, a network 103, a database server 104, and a server 105. The network 103 serves as the medium providing a communication link between terminals 101 and 102 and the database servers 104 and 105. The network 103 can connect terminals 101 and 102 and the database servers 104 and 105 using various connection methods, such as wired or wireless communication links or fiber optic cables.

[0048] User 110 can use terminals 101 and 102 to interact with server 105 via network 103 to receive or send messages, etc. Various client applications or systems can be installed on terminals 101 and 102, such as model training applications, invoice recognition applications, expense report generation applications, financial reimbursement systems, web browsers, and instant messaging tools, etc.

[0049] The terminals 101 and 102 here can be either hardware or software. When terminals 101 and 102 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), laptops, and desktop computers. When terminals 101 and 102 are software, they can be installed in the electronic devices listed above. Terminals 101 and 102 can be multiple software programs or software modules (e.g., for providing distributed services) or a single software program or software module. No specific limitations are imposed here.

[0050] When terminals 101 and 102 are hardware devices, image acquisition devices can also be installed on them. These image acquisition devices can be various devices capable of capturing images, such as cameras and sensors. User 110 can use the image acquisition devices on terminals 101 and 102 to capture images of various documents, such as medical receipts.

[0051] Database server 104 can be a database server that provides various services. For example, database server 104 can store a sample set. The sample set contains a large number of samples. The samples may include sample ticket images and text data corresponding to the sample ticket images. In this way, user 110 can also select samples from the sample set stored in database server 104 through terminals 101 and 102.

[0052] Server 105 can be a server providing various services, such as a backend server supporting various applications displayed on terminals 101 and 102. The backend server can use samples from the sample set sent by terminals 101 and 102 to train an initial model, and can send the training results (such as a generated invoice recognition model and an electronic bill generation model) back to terminals 101 and 102. In this way, users can apply the generated models to perform invoice recognition and electronic bill generation. Terminals 101 and 102 can also directly access the training results of server 105 via network 103.

[0053] The database server 104 and server 105 here can be either hardware or software. When they are hardware, they can be a distributed server cluster consisting of multiple servers or a single server. When they are software, they can be multiple software programs or software modules (e.g., for providing distributed services) or a single software program or software module. No specific limitations are made here.

[0054] It should be noted that the electronic billing statement generation method provided in this application embodiment can be executed by terminals 101 and 102, or by server 105. Correspondingly, the electronic billing statement generation device can also be located in terminals 101 and 102 or server 105.

[0055] It should be noted that if server 105 can perform the relevant functions of database server 104, database server 104 may not need to be set up in system 100.

[0056] It should be understood that Figure 1 The number of terminals, networks, database servers, and servers shown is merely illustrative. Depending on implementation needs, the system can have any number of terminals, networks, database servers, and servers.

[0057] Figure 2 This is an exemplary flowchart of an electronic bill of exchange generation method 200 applied in an embodiment of this application. The electronic bill of exchange generation method 200 includes steps 220 and 230. Below, in conjunction with... Figure 2 The method described in the embodiments of this application is explained.

[0058] 220. Input the ticket image and ticket template into a first network model. The first network model includes an input layer, multiple first convolutional layers, an attention convolutional layer, multiple second convolutional layers, and an output layer. The input layer receives the ticket image, the multiple first convolutional layers extract first features from the ticket image, and the attention convolutional layers receive the ticket image and ticket template. The attention convolutional layer includes a first branch and a second branch. The first branch obtains a first positional feature based on the ticket image and ticket template; the first positional feature is the positional feature where the ticket image and ticket template match. The second branch obtains a second positional feature based on the ticket image and ticket template; the second positional feature is the positional feature where the ticket image and ticket template do not match. The attention convolutional layer also obtains a second feature based on the first feature, the first positional feature, and the second positional feature. The multiple second convolutional layers obtain text data from the ticket image based on the second feature, and the output layer outputs the text data.

[0059] For example, terminals 101 and 102 or server 105 input the ticket image and ticket template into the first network model. If terminals 101 and 102 input the ticket image and ticket template into the first network model, the first network model can be deployed on terminals 101 and 102. If server 105 inputs the ticket image and ticket template into the first network model, the first network model can be deployed on server 105 or on other servers. Furthermore, server 105 can obtain the ticket image through terminals 101 and 102, that is, terminals 101 and 102 send the ticket image to server 105.

[0060] For example, in the first network model, the input layer, multiple first convolutional layers, attention convolutional layers, multiple second convolutional layers, and output layer are sequentially connected, and the number of multiple first convolutional layers and multiple second convolutional layers is not limited. In this embodiment, a VGG16 (visual geometry group network) neural network model is used as an example to perform image recognition processing on a ticket image. The ticket image may or may not undergo preprocessing. The ticket image is input into the VGG16 neural network model, first passing through the input layer, then sequentially through multiple first convolutional layers for feature extraction to obtain the first feature. Next, the ticket image and ticket template are input into the attention convolutional layer, and the first feature is also input into the attention convolutional layer. This attention convolutional layer is divided into two branches, a first branch and a second branch, used to locate the key regions in the ticket image, respectively. After calculation, the first branch and the second branch obtain the first positional feature and the second positional feature, respectively. The attention convolutional layer fuses the first feature, the first positional feature, and the second positional feature to obtain the second feature. The second feature is then input into multiple second convolutional layers. The second feature is processed sequentially by multiple second convolutional layers to finally obtain the text data of the ticket image.

[0061] In one embodiment, the fifth layer of the VGG16 neural network is used as an attention convolutional layer. In another embodiment, the eighth layer of the VGG16 neural network is used as an attention convolutional layer. Those skilled in the art can also set other layers of the VGG16 neural network as attention convolutional layers, and this application does not limit this.

[0062] Figure 3 This is a schematic flowchart illustrating the location of key areas using the first and second branches in the embodiments of this application. Figure 3 As shown in Figure 310, the ticket image and ticket template are input into the attention convolutional layer, or the attention convolutional layer receives the ticket image and ticket template. Figure 3 As shown in Figure 320, the attention convolutional layer determines whether the ticket image matches the ticket template. If the ticket image matches the ticket template, the ticket image enters the first branch for processing. The first branch is used to locate the key regions. In the first branch, as shown... Figure 3 As shown in Figure 331, one or more regions are fixed using prior information from the document image. This prior information can be a standardized document template, a fixed document template, or known format information from the document image. Then, the key regions are located based on these one or more regions. Figure 3 As shown in Figure 332, in one embodiment, the Euclidean distance formula is used as the loss function to calculate and locate the key regions of the ticket image. The Euclidean distance formula can be expressed as:

[0063]

[0064] Where ρ is the Euclidean distance between points (x2, y2) and (x1, y1). Points (x1, y1) and (x2, y2) can be the coordinates of any two points in the document image. In one embodiment, point (x1, y1) is the coordinate of the upper left corner of the supplier's area location box, and point (x2, y2) is the coordinate of the upper left corner of the seller's area location box. In another embodiment, point (x1, y1) is the coordinate of the lower right corner of the supplier's area location box, and point (x2, y2) is the coordinate of the lower right corner of the seller's area location box. The positions of points (x1, y1) and (x2, y2) can be various, as long as they are in the same position in both area location boxes; this application does not impose any limitations on this.

[0065] like Figure 3 As shown in Figure 333, the first branch can roughly locate the key area based on the Euclidean distance.

[0066] If the document image does not match the document template, the document image enters the second branch for processing to locate the key areas. For example... Figure 3 As shown in Figure 341, the format of the ticket image is a known format. (For example...) Figure 3 As shown in Figure 342, for a document image of a known format, template matching is used to quickly locate the region in the document image and extract the text data from that region, such as amount, date, and payee information. Then, as... Figure 3 As shown in Figure 343, key regions are marked for text detection and recognition. By defining the target region, text data, such as text and numbers, can be accurately extracted from the image.

[0067] For example, the first branch and the second branch can be two parallel methods for processing the ticket image, or the ticket image can be first input into the first branch for image processing, and then the ticket image and feature data processed by the first branch can be input into the second branch for further image processing.

[0068] For feature data processed by the first or second branch, such as Figure 3 As shown in Figure 350, calculate the probability distribution of the characters within it, as follows: Figure 3 As shown in Figure 360, the output is the feature data result after recognition, namely the first position feature and the second position feature.

[0069] For example, the attention convolutional layer further includes a first activation layer and a second activation layer. The first activation layer is connected to the output of the first branch and is used to obtain a first attention feature based on a first positional feature and a first weight function. The second activation layer is connected to the output of the second branch and is used to obtain a second attention feature based on a second positional feature and a second weight function. The attention convolutional layer is also used to fuse the first attention feature, the second attention feature, and the first feature to obtain a second feature, and input the second feature into multiple second convolutional layers to obtain text data.

[0070] Figure 4 This is a schematic diagram illustrating a feature extraction method applied in an embodiment of this application. For example... Figure 4 As shown, the ticket image is input into the first network model 403 and the attention convolutional layer combined with the attention mechanism 401, respectively. Using the attention mechanism 401, the attention convolutional layer generates first positional features and second positional features based on the first branch and the second branch. The first positional features and the second positional features can constitute a feature map, which can be represented as f∈H×W×C, where H, W, and C represent the height, width, and number of channels of the feature map, respectively. i,j Let ∈C represent the feature at a spatial location (i,j) on the feature map. The attentional convolutional layer then uses the attention mechanism to learn a weight function a(f) for each spatial location's feature. i,j θ), where θ represents the parameters of the weight function. The first and second positional features are respectively processed by the activation function G(x) to obtain a weight map with invariant size, namely attention map 402. Attention map 402 is an image generated by combining the attention convolutional layer with the attention mechanism 401 and performing weight calculations on the feature maps of the first and second positional features. This image includes the first and second attention features. The formula for calculating attention map 402 is as follows:

[0071]

[0072] in, Let represent the attention feature at a spatial location (i,j) on attention map 402. The first location feature is used to obtain the first attention feature through an activation function, and the second location feature is used to obtain the second attention feature through an activation function. The weight function a(·) consists of two 1×1 convolutional layers and a Softplus activation function.

[0073] The ticket image is processed through multiple first convolutional layers of the first network model 403 to generate a feature map 404, which includes a first feature. The feature map 404 is fused with the attention map 402 to form a second feature, namely a feature vector 405. This second feature is then fed into multiple second convolutional layers to extract the semantic features 406 of the feature vector 405, thus obtaining the text data of the ticket image.

[0074] For the recognition of complex document images, compared to single-method localization, using the first and second branches of the attention convolutional layer in the first network model to locate key regions can identify key and important features of the document image. By utilizing distance and labeled positional information, non-critical background can be masked. Furthermore, this method is more robust and efficient in enhancing "local" features, improving the accuracy of recognition and detection by 2-3 percentage points.

[0075] 230. Generate an electronic bill based on the text data.

[0076] Existing expense reimbursement systems suffer from drawbacks such as low efficiency due to manual entry, static risk rules, and reliance on experience for review. While existing rule engines can perform basic risk checks, they lack semantic understanding of unstructured data and cannot proactively predict the content to be filled in. Therefore, this application proposes an intelligent method to generate electronic expense reimbursement forms based on the identified text data, building upon the recognition of the aforementioned invoice images.

[0077] For example, terminals 101, 102, or server 105 generate electronic bills based on text data.

[0078] The following details the process of intelligently generating electronic expense reports.

[0079] First, generate a vector representation of the text data.

[0080] For example, multimodal information is fused. Multimodal information includes data in various modes such as images, text, video, and audio. In one embodiment, the text data identified by the above method undergoes structured fusion processing. A spaCy pre-trained model is used for word segmentation and part-of-speech tagging, combined with a custom financial domain dictionary to expand entity recognition capabilities. Then, an entity relationship graph is constructed. A Flair-based named entity recognition (NER) model is used to extract key entities such as expense type, amount, date, and supplier. Semantic relationships between entities are constructed through dependency parsing (e.g., amount-expense type, date-business occurrence time). Flair is a PyTorch-based natural language processing (NLP) library that supports NER functionality for multiple languages. Finally, the semantic relationship vectors are spatially mapped. Vector representations are generated using the Word2Vec continuous bag of words (CBOW) model. The weight matrix is ​​calculated using term frequency-inverse document frequency (TF-IDF), and then dimensionality is reduced to a 300-dimensional feature space using principal component analysis (PCA). It should be understood that the above process can be performed on a word-by-word or character-by-character basis. Vector representations obtained by processing on a word-by-word basis are called word vector representations. Vector representations obtained by processing on a character-by-character basis are called character vector representations.

[0081] Next, the vector representation is input into a second network model, which includes a multi-layer encoder and a multi-layer decoder. The multi-layer encoder is used to extract features from the vector representation, and the multi-layer decoder is used to generate the electronic bill based on the features output by the multi-layer encoder.

[0082] For example, terminals 101, 102, or server 105 input vector representations into the second network model. If terminals 101 and 102 input vector representations into the second network model, the second network model can be deployed on terminals 101 and 102. If server 105 inputs vector representations into the second network model, the second network model can be deployed on server 105 or on other servers.

[0083] For example, the second network model uses the Transformer-XL model architecture based on a multi-head attention mechanism, and designs a 12-layer encoder and a 6-layer decoder structure. Those skilled in the art can also design encoder and decoder structures with other numbers of layers, and this application does not limit this.

[0084] During the pre-training phase, a masked language model (MLM) and next sentence prediction (NSP) tasks are employed. A dynamic masking mechanism, combined with the AdamW optimizer, is used to implement a dynamic learning rate adjustment strategy. In one embodiment, the dynamic masking mechanism performs 100% masking of financial technical terms (e.g., "value-added tax invoice") and 15% masking of common words. In another embodiment, the dynamic masking mechanism performs 90% masking of financial technical terms and 30% masking of common words. Those skilled in the art can set other masking ratios, and this application does not limit this.

[0085] Unsupervised pre-training is performed on a pre-defined knowledge base, such as a policy knowledge base (PKB) and a corpus for invoice recognition. Information is classified and distinguished according to the semantic relationships between entities and pre-defined prefix rules, so as to realize context-aware prediction and filling of invoices and generate electronic invoices.

[0086] Finally, the risk type of the electronic expense report is displayed, which includes one or more of the following: duplicate expense reports, expense reports exceeding the standard, expense reports spanning multiple periods, or sensitive suppliers.

[0087] For example, the risk type of the electronic expense report can be obtained based on a preset knowledge base, which can be a corporate policy knowledge base (PKB). Based on the PKB and the electronic expense report generated through the above process, combined with similar compliant expense reporting history, risk anomalies (such as duplicate reimbursements or reimbursements exceeding standards) are identified. Taking into account factors such as risk type, probability of occurrence, and degree of impact, a reasonable risk threshold is set. When data in the expense report triggers the threshold, the business system automatically performs an interception operation, marks high-risk fields, and simultaneously matches the expense report text format using regular expressions. It also uses an N-gram language model to verify the grammatical validity of invoice headers, supplier names, etc., and provides accurate remediation suggestions. Table 1 introduces the risk types and their corresponding feature inputs and warning thresholds.

[0088] Table 1

[0089] Risk type Feature Input Warning threshold Duplicate reimbursement Invoice summary, amount, date, buyer and seller >0.95 Reimbursement exceeding the standard Cost type, project budget, historical average >120% Reimbursement for expenses incurred beyond the specified period Business occurrence time, expense reimbursement submission time >6 months Sensitive Suppliers Supplier name, transaction frequency, amount Risk score > 70

[0090] Duplicate reimbursement refers to the same invoice being reimbursed again on top of an already reimbursed expense. The model verifies that if the similarity between the reimbursement's characteristic data (e.g., invoice summary, amount, date, buyer and seller) and existing characteristic data in the database is above 95%, the reimbursement is considered duplicate. Over-standard reimbursement refers to the reimbursement amount exceeding the project budget. The model verifies the expense type field and considers both the project budget and historical average amounts. If the reimbursement amount is 120% of the project budget or historical average, it is considered over-standard. Cross-period reimbursement refers to the reimbursement exceeding the time limit. The model verifies that if the time between the transaction and the reimbursement submission exceeds 6 months, the reimbursement is considered cross-period. Sensitive supplier refers to the supplier in this reimbursement being a sensitive supplier. The model determines whether the supplier in this reimbursement is a sensitive supplier based on the supplier list in the database or illegal order transactions. If the supplier is a sensitive supplier, the reimbursement is included in the risk warning, and a risk score is calculated according to preset rules. If the risk score exceeds 70, the supplier is considered a high-risk supplier.

[0091] Intelligent electronic expense reimbursement generation utilizes multimodal information fusion processing to perform deep structured analysis of the text data of invoices. Combined with a custom financial domain dictionary and advanced natural language processing models, it can quickly and accurately extract key information. By constructing an entity relationship graph and mapping a semantic vector space, it deeply mines the semantic relationships between text data, transforming the text data into computable semantic vectors. Simultaneously, based on the context-aware function of the Transformer-XL model, the training strategy is optimized for financial domain data during the pre-training stage, enabling the model to accurately understand the semantics of the invoice text data and automatically fill in the expense reimbursement content based on semantic relationships, generating pre-filled electronic expense reimbursement suggestions. This method eliminates static template filling, reduces manual data entry workload, improves expense reimbursement drafting efficiency by approximately 70%, significantly shortens the expense reimbursement process time, reduces financial processing costs, and provides users with an intelligent and personalized expense reimbursement experience.

[0092] In addition, existing technologies employ static rule engines for risk verification, which are ill-suited to complex and ever-changing business scenarios. This application's embodiments construct a dynamic risk identification system based on the enterprise policy knowledge base (PKB) and historical compliance data. This system not only identifies anomalies such as duplicate reimbursements and reimbursements exceeding standards based on real-time policies and similar reimbursement behaviors, but also accurately locates risks by setting comprehensive risk thresholds. By combining regular expressions and N-gram language models, further verification is performed on text format and grammatical rationality, increasing the risk identification accuracy to over 85%. This effectively compensates for the shortcomings of traditional static rules and enhances financial risk management capabilities.

[0093] In one implementation of the embodiments of this application, such as Figure 2 As shown, the method for generating electronic bills 200 may further include step 210.

[0094] 210. Obtain business information, which includes a first business, a second business, or a third business. The first business is the matching of the invoice image with the invoice template; the second business is the mismatch between the invoice image and the invoice template, but the standardized result of the invoice image is a first value; the third business is the mismatch between the invoice image and the invoice template, but the standardized result of the invoice image is a second value. The first value indicates that the invoice image has been standardized, and the second value indicates that the invoice image cannot be standardized. Based on the business information, preprocess the invoice image.

[0095] For example, terminals 101 and 102 or server 105 acquire business information and preprocess the document image based on the business information. If terminals 101 and 102 acquire business information, the preprocessing operation on the document image based on the business information can be performed on terminals 101 and 102. If server 105 acquires business information, the preprocessing operation on the document image based on the business information can be performed on server 105 or on other servers.

[0096] The document image used for document recognition can be either pre-processed or unprocessed. In one embodiment, if the document image is unprocessed, preprocessing is required. Since the document reporting standards differ across business scenarios, the locations of key information in the document image may also vary. Therefore, different methods need to be used to preprocess the document image for different business scenarios. Figure 5 This is a schematic diagram illustrating the document image preprocessing method for three business scenarios applied in the embodiments of this application.

[0097] For example, in the scenario of the first business, the invoice image matches the invoice template, meaning the format of the invoice image is consistent with the format of the invoice template. In some embodiments, the financial reimbursement system is controlled by a front-end business system. The invoice images uploaded to the business system are reviewed and approved by business personnel, and the image attachments are confirmed, thus achieving standardized and regulated uploading of invoice images. Figure 5 The first line of the business system or pre-application form is used for preliminary review of invoice images as business attachments. This preliminary review clarifies the business type and the scope of review for financial attachments, defines the business review process and the reimbursement approval process, and ensures that the province's existing business system or pre-reimbursement capabilities support it. Pre-approval is loaded into the reimbursement process to ensure that the format of image attachments is consistent with the invoice template. After financial review and voucher import, the invoice image is preprocessed. In the first business process, the invoice image only needs to undergo steps such as tilt correction and smoothing denoising to improve the accuracy of semantic matching of text data in the invoice image.

[0098] For example, in the scenario of the second business, the invoice image does not match the invoice template and the standardization result of the invoice image is the first value. That is, the format of the invoice image is inconsistent with the format of the invoice template, but a format matching the invoice template can be obtained based on the meaning of the text data in the invoice image, i.e., the invoice image is standardized. In some embodiments, the financial reimbursement system does not have front-end business system control, the format requirements of the invoice images uploaded by the business system are different, the text data of the invoices is irregularly arranged, or the invoice images contain interference such as seals. Figure 5 The second line of input does not include the pre-application form from the business system, but it sorts out the attached image documents and identifies those that can be formatted into text (Word) or table (Excel) formats. After financial review and voucher import, the attached image documents are confirmed, the image document data is extracted and applied, and regularization improves the approval rate. In the second business process, preprocessing methods such as tilt correction and smoothing are insufficient to correct text angles and resolve background interference. The following details how to preprocess the image documents to address these issues.

[0099] First, the ticket image is divided into multiple sub-regions.

[0100] For example, a hybrid partitioning strategy using two partitioning methods is adopted to balance the structural characteristics of the invoices and computational efficiency.

[0101] As an optional embodiment, this application employs a grid-based partitioning method to divide the invoice image. The invoice image is initially divided according to a fixed-size grid to obtain fixed-size sub-regions. In one embodiment, the image is divided according to a 32×32 pixel grid to obtain 32×32 pixel sub-regions. This partitioning method can quickly perform initial segmentation of the invoice image, is applicable to most invoice images, and ensures that the subsequent computational complexity remains within a controllable range.

[0102] As an optional embodiment, this application employs an adaptive segmentation method based on contour detection to segment the document image. For documents with obvious tabular or graphic structures (e.g., invoices, expense reports), the Canny edge detection algorithm is first used to extract the contours of the document image. Then, the Hough transform is combined to identify straight lines and rectangular structures, thereby accurately segmenting regions with independent significance in the document image (e.g., amount fields, date fields, or title fields). In one embodiment, when a rectangular frame is detected, the area within the frame is treated as an independent sub-region. This segmentation method can selectively segment key information regions according to the actual structure of the document.

[0103] For example, the segmented sub-regions are not processed independently; subsequent calculations merge the sub-regions based on the feature similarity of adjacent sub-regions. In one embodiment, the merging rule is as follows: if the difference in the mean brightness of adjacent sub-regions is less than a set threshold (e.g., 10 grayscale values) and the difference in variance is less than another set threshold (e.g., 5 squared grayscale values), then the two sub-regions are merged to reduce computational redundancy caused by over-segmentation.

[0104] Secondly, a local grayscale histogram is generated for each sub-region.

[0105] For example, after obtaining the sub-regions, the brightness distribution characteristics of each sub-region are analyzed using local grayscale histograms, and the grayscale value distribution of each sub-region is statistically analyzed to generate a local grayscale histogram. In one embodiment, the grayscale range of 0-255 is divided into 16 intervals, and the number of pixels in each interval of each sub-region is counted to obtain the local histogram data of each sub-region.

[0106] Next, the brightness of the ticket image is corrected based on each local grayscale histogram.

[0107] For example, the Gamma correction method is used to dynamically calculate and adjust the ticket image. The Gamma correction method aims to enhance the contrast of the ticket image by introducing a Gamma transformation, thereby correcting the brightness response of irregular ticket images. The principle expression of the Gamma correction method is as follows:

[0108]

[0109] Among them, V in Input grayscale value, the value range is 0. <V in <1, V out The output is the grayscale value, and the value of the parameter γ (Gamma) determines the grayscale mapping between the input and output images. In one embodiment, this application uses a multi-parameter regression model to dynamically calculate the Gamma value. The dynamic calculation process of the Gamma value is described in detail below.

[0110] First, the model is trained. A large number of ticket images of different types and qualities are collected as the training set, and the Gamma value of each sub-region under the optimal correction effect is manually labeled as a tag. The feature data of the sub-regions (e.g., mean brightness, variance, skewness, neighborhood information, ticket type identifier) ​​are used as input to the model, and the model is trained using the support vector regression (SVR) algorithm or random forest regression algorithm to establish the mapping relationship between the feature data and the optimal Gamma value.

[0111] Secondly, real-time calculation is performed. For a newly input ticket image, feature data of each sub-region of the ticket image is extracted, and the feature data is input into the pre-trained regression model to predict the optimal Gamma value of each sub-region in real time.

[0112] For example, using the maximization of information entropy of the sub-region image before and after applying the Gamma correction method as the criterion for determining the optimal Gamma value, the information entropy formula can be expressed as:

[0113]

[0114] Where p(i) is the probability of grayscale value i in the image. The Gamma value that maximizes the information entropy of the sub-region is selected to retain more image detail. For different ticket images or different sub-regions of the same ticket image, the final Gamma value is dynamically calculated based on its feature data. The process of determining the Gamma value is as follows: the feature data of the sub-region is input into the trained regression model to obtain the predicted Gamma value. If the value does not meet the optimal judgment criterion, that is, the Gamma value cannot maximize the information entropy of the sub-region image, then the Gamma value is iterated and adjusted within a certain range with a fixed step size. The information entropy of the sub-region is calculated under each adjusted value, and the Gamma value that maximizes the information entropy is selected as the final value. In one embodiment, within a numerical range greater than 0.5 and less than 1.5, the current Gamma value is iterated and adjusted with a step size of 0.01. It is understood that the numerical range and step size can be other values, and this application does not limit them.

[0115] Secondly, dynamic adjustments are made for each sub-region. Each sub-region calculates an independent Gamma value based on its own feature data to achieve precise local correction. In one embodiment, for a darker sub-region with low contrast, the calculated Gamma value may be less than 1 to enhance the contrast of its low grayscale areas; in another embodiment, for a brighter sub-region, the calculated Gamma value may be greater than 1 to reduce brightness.

[0116] Finally, global collaborative adjustment is performed. After calculating the Gamma values ​​of all sub-regions, to avoid excessive differences in the correction effects between adjacent sub-regions, a smoothing filtering algorithm (e.g., Gaussian filtering) is used to smooth the Gamma values ​​of adjacent sub-regions, making the transition of Gamma values ​​between adjacent regions natural and ensuring the visual consistency of the overall ticket image.

[0117] For example, in the scenario of the third business, the invoice image does not match the invoice template and the standardized result of the invoice image is the second value. That is, the format of the invoice image is inconsistent with the format of the invoice template, and it is impossible to obtain a format that matches the invoice template based on the meaning of the text data in the invoice image. In other words, the invoice image cannot be standardized. In some embodiments, the types of attachments uploaded to the business system are different, or the position of key information in the invoice image is not fixed. Figure 5 The third line of input consists of attachments and external documents that cannot be standardized by the internal business system. After financial review and voucher import of such attachments, a combination of OCR, NLP, and rule engine is used to process the attachment images independently and step-by-step, according to different scenarios. In the third business process, the document images need to be processed using the methods of the first and second business processes. On this basis, image binarization is performed on the document images, and adaptive threshold binarization is used to visualize the grayscale distribution of different regions of the image. That is, in the third business process, the document images are first subjected to tilt correction and smoothing denoising. Then, the document images after tilt correction and smoothing denoising are divided into sub-regions. A local grayscale histogram is generated for each sub-region. The brightness of the document images is corrected based on each local grayscale histogram. Finally, the brightness-corrected document images are subjected to image binarization, and adaptive threshold binarization is used to visualize the grayscale distribution of different regions of the image.

[0118] For example, for a ticket image preprocessed through one of the three business scenarios mentioned above, Principal Component Analysis (PCA) is used to detect and correct the tilt angle of the ticket image. The PCA algorithm maps the ticket image data to a low-dimensional space and then back to a high-dimensional space, achieving noise reduction. Compared with methods using Fourier, Radon, and Hough algorithms, the embodiments of this application show a 30% improvement in recognition accuracy for classification data preprocessing.

[0119] Based on the design of basic regularized image attachment calibration, this application preprocesses three types of data for document image attachments to calibrate key information in the attachments. Furthermore, unlike general images which use a globally fixed Gamma value for contrast and brightness correction, this application proposes a dynamic Gamma correction method based on local feature analysis. This avoids over-enhancement or under-enhancement issues that occur in global correction, effectively improving the clarity of text and patterns in document images. This application divides the document image into regions and uses local histogram analysis to determine the brightness distribution characteristics of each sub-region, dynamically adjusting the Gamma value. Unlike Gamma correction using fixed parameters, this application can precisely optimize the brightness and contrast of various local regions of irregular document images, demonstrating stronger adaptability and correction effectiveness.

[0120] The foregoing section detailed examples of the methods provided in this application. It is important to understand that the methods provided in this application are not limited to document image attachment recognition or intelligent financial reimbursement form filling. The purpose of the methods provided in this application is to simplify the document submission process for claimants and settlement specialists, streamline the reimbursement form filling process, reduce manual comparison of financial information, improve document circulation efficiency, and effectively solve the problems of high error rates, low efficiency, and long approval times associated with existing manual form filling, batch invoices, and attachment information uploading and recognition. This provides information technology support for financial reimbursement and settlement work. Similar recognition and data entry methods in other industries can also be used as a reference, demonstrating broad application prospects.

[0121] The following is combined Figures 6 to 7 This application describes an apparatus for generating electronic bills according to embodiments of the present application.

[0122] Figure 6 and Figure 7 This is a schematic diagram of the structure of the electronic bill of exchange generation apparatus used in the embodiments of this application. These electronic bill of exchange generation apparatuses are used to implement any embodiment of the above-described electronic bill of exchange generation method. For example... Figure 6 As shown, the electronic billing generation device 600 includes an input unit 610, a processing unit 620, and an output unit 630.

[0123] When the electronic billing generation device 600 is used to realize Figure 2 The function of the method embodiment shown is as follows: the input unit 610 is used to input the invoice image and the invoice template to the processing unit 620; the processing unit 620 is equipped with a first network model, and the processing unit 620 obtains the text data of the invoice image through the first network module, and obtains the electronic bill of lading based on the text data; the output unit 630 is used to output the electronic bill of lading.

[0124] In one implementation of this application, a first network model includes an input layer, multiple first convolutional layers, an attention convolutional layer, multiple second convolutional layers, and an output layer. The input layer receives a ticket image, the multiple first convolutional layers extract a first feature from the ticket image, the attention convolutional layers receive the ticket image and a ticket template, and the attention convolutional layers include a first branch and a second branch. The first branch obtains a first positional feature based on the ticket image and the ticket template. The first positional feature is a positional feature that matches the ticket image and the ticket template. The second branch obtains a second positional feature based on the ticket image and the ticket template. The second positional feature is a positional feature that does not match the ticket image and the ticket template. The attention convolutional layers also obtain a second feature based on the first feature, the first positional feature, and the second positional feature. The multiple second convolutional layers obtain text data of the ticket image based on the second feature. The output layer outputs the text data.

[0125] In one implementation of this application, the attention convolutional layer further includes a first activation layer and a second activation layer. The first activation layer is connected to the output of the first branch and is used to obtain a first attention feature based on a first positional feature and a first weight function. The second activation layer is connected to the output of the second branch and is used to obtain a second attention feature based on a second positional feature and a second weight function. The attention convolutional layer is also used to fuse the first attention feature, the second attention feature, and the first feature to obtain a second feature, and input the second feature into multiple second convolutional layers to obtain text data.

[0126] In one implementation of this application, the processing unit 620 further includes a second network model. The processing unit 620 is further configured to: generate a vector representation of the text data; and input the vector representation into the second network model, wherein the second network model includes a multi-layer encoder and a multi-layer decoder, the multi-layer encoder being used to extract features from the vector representation, and the multi-layer decoder being used to generate an electronic bill based on the features output by the multi-layer encoder.

[0127] In one implementation of this application, the output unit 630 is further configured to: display the risk type of the electronic expense report, wherein the risk type includes one or more of the following: duplicate reimbursement, over-standard reimbursement, cross-period reimbursement, or sensitive supplier.

[0128] In one implementation of this application, the processing unit 620 is further configured to: acquire business information, which includes a first business, a second business, or a third business, wherein the first business is a business in which the invoice image matches the invoice template, the second business is a business in which the invoice image does not match the invoice template and the standardization result of the invoice image is a first value, and the third business is a business in which the invoice image does not match the invoice template and the standardization result of the invoice image is a second value, wherein the first value indicates that the invoice image is standardized and the second value indicates that the invoice image cannot be standardized; and preprocess the invoice image according to the business information.

[0129] In one implementation of this application, the processing unit 620 is further configured to: perform tilt correction and smoothing denoising processing on the ticket image in a first business scenario; in a second business scenario, divide the ticket image to obtain multiple first sub-regions, generate a first local grayscale histogram for each first sub-region, and correct the brightness of the ticket image based on each first local grayscale histogram; in a third business scenario, perform tilt correction and smoothing denoising processing on the ticket image, divide the ticket image after tilt correction and smoothing denoising processing to obtain multiple second sub-regions, generate a second local grayscale histogram for each second sub-region, correct the brightness of the ticket image after tilt correction and smoothing denoising processing based on each second local grayscale histogram, and perform image binarization processing on the brightness-corrected ticket image.

[0130] like Figure 7 As shown, the electronic billing statement generation apparatus 700 includes a processor 710, an interface circuit 720, and a memory 740. The processor 710, interface circuit 720, and memory 740 are coupled to each other via a bus 730. It is understood that the interface circuit 720 can be an input / output interface. The memory 740 is used to store instructions executed by the processor 710, or to store input data required by the processor 710 to execute instructions, or to store data generated after the processor 710 executes instructions. Sometimes, the interface circuit 720 can also be understood as part of the processor 710, in which case the electronic billing statement generation apparatus 700 includes the processor 710 and the memory 740. The processor 710 is used to control the device to implement any embodiment of the above-described electronic billing statement generation method.

[0131] The electronic billing generation device in this application embodiment can be a digital computer of various forms, such as a laptop computer, desktop computer, workbench, personal digital assistant, server, blade server, mainframe computer, and other suitable computers. The electronic billing generation device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0132] It is understood that the processor mentioned above can be a central processing unit (CPU), or it can be one or more combinations of other general-purpose processors, digital signal processors (DSPs), microprocessor units (MPUs), microcontroller units (MCUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), artificial intelligence processors (AI processors), such as tensor processing units (TPUs) or neural processing units (NPUs), discrete gate or transistor logic devices, or discrete hardware components.

[0133] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0134] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the electronic billing method shown in the method flow of the above method embodiments.

[0135] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires; a portable computer disk drive; a hard disk drive; a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM); a register; a hard disk drive; an optical fiber; a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination thereof; or any other form of computer-readable storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read a product from the storage medium and to write a product to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). In the embodiments of this application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0136] Embodiments of this application provide a computer program product containing instructions that, when executed on a computer, cause the computer to perform the electronic billing generation method of embodiments of this application.

[0137] Since the electronic billing device, computer-readable storage medium, and computer program product in the embodiments of this application can be applied to the above method, the technical effects that can be obtained can also be referred to the above method embodiments. The embodiments of this application will not be repeated here.

[0138] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0140] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0141] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.

Claims

1. A method for generating an electronic bill, characterized in that, include: A ticket image and a ticket template are input into a first network model. The first network model includes an input layer, multiple first convolutional layers, an attention convolutional layer, multiple second convolutional layers, and an output layer. The input layer receives the ticket image. The multiple first convolutional layers extract a first feature from the ticket image. The attention convolutional layers receive the ticket image and the ticket template. The attention convolutional layer includes a first branch and a second branch. The first branch obtains a first positional feature based on the ticket image and the ticket template. The first positional feature is a positional feature where the ticket image and the ticket template match. The second branch obtains a second positional feature based on the ticket image and the ticket template. The second positional feature is a positional feature where the ticket image and the ticket template do not match. The attention convolutional layer also obtains a second feature based on the first feature, the first positional feature, and the second positional feature. The multiple second convolutional layers obtain text data of the ticket image based on the second feature. The output layer outputs the text data. The electronic bill is generated based on the text data.

2. The method according to claim 1, characterized in that, The attention convolutional layer further includes a first activation layer and a second activation layer. The first activation layer is connected to the output of the first branch and is used to obtain a first attention feature based on the first position feature and the first weight function. The second activation layer is connected to the output of the second branch and is used to obtain a second attention feature based on the second position feature and the second weight function. The attention convolutional layer is also used to fuse the first attention feature, the second attention feature, and the first feature to obtain the second feature, and input the second feature into the plurality of second convolutional layers to obtain the text data.

3. The method according to claim 1 or 2, characterized in that, The step of generating the electronic bill based on the text data includes: Generate a vector representation of the text data; The vector representation is input into a second network model, which includes a multi-layer encoder and a multi-layer decoder. The multi-layer encoder is used to extract features from the vector representation, and the multi-layer decoder is used to generate the electronic bill based on the features output by the multi-layer encoder.

4. The method according to any one of claims 1 to 3, characterized in that, Also includes: The electronic expense report will display the risk type, which includes one or more of the following: duplicate expense report, expense report exceeding the standard, expense report across periods, or sensitive supplier.

5. The method according to any one of claims 1 to 4, characterized in that, Also includes: Obtain business information, which includes a first business, a second business, or a third business. The first business is the business where the invoice image matches the invoice template. The second business is the business where the invoice image does not match the invoice template and the standardization result of the invoice image is a first value. The third business is the business where the invoice image does not match the invoice template and the standardization result of the invoice image is a second value. The first value indicates that the invoice image has been standardized, and the second value indicates that the invoice image cannot be standardized. The ticket image is preprocessed based on the business information.

6. The method according to claim 5, characterized in that, The step of preprocessing the ticket image based on the business information includes: In the first business scenario, the ticket image is subjected to tilt correction and smoothing / denoising processing; In the second business scenario, the ticket image is divided into multiple first sub-regions, a first local grayscale histogram is generated for each first sub-region, and the brightness of the ticket image is corrected based on each first local grayscale histogram. In the third business scenario, the ticket image is subjected to tilt correction and smoothing denoising processing. The ticket image after tilt correction and smoothing denoising processing is divided into multiple second sub-regions. A second local grayscale histogram is generated for each second sub-region. The brightness of the ticket image after tilt correction and smoothing denoising processing is corrected according to each second local grayscale histogram. Finally, the ticket image after brightness correction is subjected to image binarization processing.

7. An apparatus for generating electronic expense reports, characterized in that, include: A module that performs the method according to any one of claims 1 to 6.

8. An apparatus for generating electronic expense reports, characterized in that, It includes a processor and a memory, the processor and the memory being coupled, the processor being used to control the device to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed by the ticket recognition device, implement the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a computer, implement the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Invoice risk control management method and system based on multi-source data linkage

    CN121120289A