A method for training a return single identification model, a method for identifying a return single, and related devices

CN121354130BActive Publication Date: 2026-09-29KINGDEE DEEKING CLOUDCOMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511519080.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-09-29
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

[0004]但是,现有的回单识别方法依赖人工收集和定义大量规则模板,工作量大且难以覆盖所有回单格式,导致新格式的回单出现时识别失败

Benefits of technology

[0030]从以上技术方案可以看出,本申请实施例具有以下优点:本申请通过两次训练集的迭代,利用数据抽样技术构建涵盖各种完整程度类别以及各个识别信息类别的应用场景的基础数据集,并通过筛选和优化得到更高质量的目标回单训练集,再基于目标回单训练集单独微调每个识别模型。这样,有效利用了有限数据资源,使识别模型学习更多样化特征,从而提升对不同回单格式的适应性和识别准确性、鲁棒性与泛化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121354130B_ABST
    Figure CN121354130B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a return single identification model training method, a return single identification method and related equipment, which are used for training a return single identification model while improving the accuracy of return single identification. The method of the embodiments of the present application comprises: obtaining a plurality of first original return single samples and an initial return single training set, identifying the plurality of first original return single samples by using a plurality of pre-trained identification models respectively, and determining a plurality of second original return single samples; classifying the plurality of second original return single samples based on each identification information category, and determining a plurality of target return single samples corresponding to each identification information category; identifying the plurality of target return single samples by using the plurality of pre-trained identification models respectively; and obtaining a plurality of identification models after individual fine-tuning when the loss between the predicted target identification information and the labeled target identification information corresponding to each of the plurality of pre-trained identification models reaches a preset convergence condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of receipt recognition, and more specifically, to a receipt recognition model training method, a receipt recognition method, a receipt recognition model training device, a receipt recognition device, an electronic device, a computer-readable storage medium, and a computer program product containing instructions. Background Technology

[0002] With the development of financial services, the number and diversity of bank receipts are constantly increasing, and the demand for receipt identification is also growing.

[0003] The existing method for recognizing receipts involves first converting the receipt into an image, then extracting structured text information using OCR technology. Next, the system sets corresponding rule templates for receipts from different banks, and the program reads these rule templates to segment the receipt and extract key information.

[0004] However, existing receipt recognition methods rely on manually collecting and defining a large number of rule templates, which is labor-intensive and difficult to cover all receipt formats, leading to recognition failures when new receipt formats appear. Therefore, the accuracy of receipt recognition is low. Summary of the Invention

[0005] This application provides a method for training a receipt recognition model, a receipt recognition method, a receipt recognition model training device, a receipt recognition device, an electronic device, a computer-readable storage medium, and a computer program product containing instructions, which can train a receipt recognition model while improving the accuracy of receipt recognition.

[0006] In a first aspect, embodiments of this application provide a method for training a receipt recognition model, including:

[0007] Multiple first original receipt samples and an initial receipt training set are obtained. The first original receipt samples are original receipt samples whose recognition information is consistent across multiple recognition models. The original receipt samples are labeled with corresponding recognition information. The initial receipt training set includes multiple initial receipt samples corresponding to each category of completeness of recognition information. The initial receipt samples are labeled with initial recognition information.

[0008] The multiple first original receipt samples are identified using multiple pre-trained recognition models, and multiple second original receipt samples with consistent recognition information corresponding to the multiple pre-trained recognition models are identified. The multiple pre-trained recognition models are trained from the initial receipt training set.

[0009] The multiple second original return receipt samples are classified based on each identification information category, and multiple target return receipt samples corresponding to each identification information category are determined from the multiple second original return receipt samples corresponding to each identification information category to construct a target return receipt training set, wherein the target return receipt samples are labeled with target identification information.

[0010] The multiple target return receipt samples are identified by the multiple pre-trained recognition models respectively, and the predicted target recognition information of the multiple target return receipt samples corresponding to each of the multiple pre-trained recognition models is obtained. When the loss between the predicted target recognition information and the labeled target recognition information corresponding to each of the multiple pre-trained recognition models reaches a preset convergence condition, multiple individually fine-tuned recognition models are obtained. The multiple individually fine-tuned recognition models are used to identify return receipts.

[0011] Secondly, embodiments of this application provide a receipt identification method, including:

[0012] Obtain at least one receipt;

[0013] The at least one receipt is identified using multiple individually fine-tuned recognition models as described in the first aspect, and the initial recognition results of the at least one receipt output by each of the multiple individually fine-tuned recognition models are obtained.

[0014] Based on the determination of whether the initial recognition results of the at least one receipt output by the multiple individually fine-tuned recognition models are consistent, the target recognition result of the at least one receipt is determined.

[0015] Thirdly, embodiments of this application provide a receipt recognition model training device, including:

[0016] The acquisition unit is used to acquire multiple first original receipt samples and an initial receipt training set. The first original receipt samples are original receipt samples whose recognition information is consistent across multiple recognition models. The original receipt samples are labeled with corresponding recognition information. The initial receipt training set includes multiple initial receipt samples corresponding to each category of completeness of recognition information. The initial receipt samples are labeled with initial recognition information.

[0017] The determining unit is used to identify the multiple first original receipt samples using multiple pre-trained recognition models, and to determine multiple second original receipt samples whose recognition information is consistent with the multiple pre-trained recognition models. The multiple pre-trained recognition models are trained from the initial receipt training set.

[0018] The determining unit is further configured to classify the plurality of second original return receipt samples based on each identification information category, and determine the plurality of target return receipt samples corresponding to each identification information category from the plurality of second original return receipt samples corresponding to each identification information category, so as to construct a target return receipt training set, wherein the target return receipt samples are labeled with target identification information;

[0019] The acquisition unit is further configured to use the pre-trained multiple recognition models to identify the multiple target return receipt samples respectively, obtain the predicted target recognition information of the multiple target return receipt samples corresponding to each of the pre-trained multiple recognition models, and when the loss between the predicted target recognition information and the labeled target recognition information corresponding to each of the pre-trained multiple recognition models reaches a preset convergence condition, obtain the multiple recognition models after individual fine-tuning, and the multiple recognition models after individual fine-tuning are used to identify return receipts.

[0020] Fourthly, embodiments of this application provide a receipt identification device, including:

[0021] The acquisition unit is used to acquire at least one receipt;

[0022] The identification unit is used to identify the at least one receipt using the multiple identification models that have been individually fine-tuned as described in the first aspect, and to obtain the initial identification results of the at least one receipt output by each of the multiple identification models that have been individually fine-tuned.

[0023] The determining unit is used to determine the target recognition result of the at least one receipt based on the determination result of whether the initial recognition results of the at least one receipt output by the multiple individually fine-tuned recognition models are consistent.

[0024] Fifthly, embodiments of this application provide an electronic device, including:

[0025] Central processing unit, memory, input / output interfaces, wired or wireless network interfaces, and power supply;

[0026] The memory is either a short-term storage memory or a persistent storage memory;

[0027] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the methods of the first or second aspect described above.

[0028] In a sixth aspect, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the methods described in the first or second aspect.

[0029] In a seventh aspect, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the first or second aspect.

[0030] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: This application constructs a basic dataset covering various completeness categories and various identification information categories through two iterations of the training set, using data sampling techniques. A higher-quality target receipt training set is obtained through filtering and optimization, and then each identification model is fine-tuned individually based on the target receipt training set. In this way, limited data resources are effectively utilized, enabling the identification model to learn more diverse features, thereby improving its adaptability to different receipt formats and its identification accuracy, robustness, and generalization ability.

[0031] Accordingly, the receipt recognition model training device, receipt recognition device, electronic device, computer-readable storage medium, and computer program product containing instructions provided in this application also have the above-mentioned technical effects. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the architecture of a receipt recognition model training system disclosed in an embodiment of this application;

[0033] Figure 2 This is a flowchart illustrating a method for training a receipt recognition model disclosed in an embodiment of this application;

[0034] Figure 2-1-1 This is a schematic diagram of a bank receipt disclosed in an embodiment of this application;

[0035] Figure 2-1 This is a schematic diagram of a receipt image disclosed in an embodiment of this application, which "explicitly shows the payee and payer".

[0036] Figure 2-2 This is a schematic diagram of a receipt image of a type that "only explicitly shows the payee" as disclosed in an embodiment of this application;

[0037] Figure 2-3 This is a schematic diagram of a receipt image of a data type "only explicitly showing the payer" disclosed in an embodiment of this application;

[0038] Figure 2-4 This is a schematic diagram of a receipt image of a type "neither the payee nor the payer is explicitly given" as disclosed in an embodiment of this application.

[0039] Figure 2-5 This is a schematic diagram of a receipt image with the data type "Bank A_Fee Receipt" disclosed in an embodiment of this application;

[0040] Figure 2-6 This is a schematic diagram of a receipt image with the data type "Bank B_Tax Payment Receipt" disclosed in an embodiment of this application;

[0041] Figure 2-7 This is a schematic diagram of a receipt image with the data type "C Bank_Transaction Receipt" disclosed in an embodiment of this application;

[0042] Figure 2-8 This is a schematic diagram of a receipt image with the data type "D Bank_Deduction Receipt" disclosed in an embodiment of this application;

[0043] Figure 2-9 This is a schematic diagram of a receipt pagination method disclosed in an embodiment of this application;

[0044] Figure 3 This is a flowchart illustrating a receipt identification method disclosed in an embodiment of this application;

[0045] Figure 3-1 This is a flowchart illustrating another receipt identification method disclosed in an embodiment of this application;

[0046] Figure 4 This is a schematic diagram of the structure of a receipt recognition model training device disclosed in an embodiment of this application;

[0047] Figure 5 This is a schematic diagram of the structure of a receipt recognition device disclosed in an embodiment of this application;

[0048] Figure 6 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. Detailed Implementation

[0049] This application provides a method for training a receipt recognition model, a receipt recognition method, a receipt recognition model training device, a receipt recognition device, an electronic device, a computer-readable storage medium, and a computer program product containing instructions, which can train a receipt recognition model while improving the accuracy of receipt recognition.

[0050] Currently, there are two main methods for receipt recognition: The first involves converting the receipt into an image and then extracting structured text information using OCR technology. Then, corresponding rule templates are set for receipts from different banks within the system, and the program reads these templates to segment the receipt and extract key information. However, existing receipt recognition methods rely on manually collecting and defining a large number of rule templates, which is labor-intensive and difficult to cover all receipt formats, leading to recognition failures when new formats appear. Therefore, the accuracy of receipt recognition is low. The second method converts the receipt into an image, extracts structured text information using OCR technology, and then calls a large model, utilizing its general recognition capabilities to return key information for subsequent processing. However, this method is costly and its accuracy is uncontrollable (due to issues like "large model illusion"). In financial applications, data accuracy is crucial. If the returned data contains uncertain errors, the credibility of the entire business data will be affected, requiring additional manual verification, which limits the actual effectiveness and application of the solution, resulting in low customer acceptance. Based on this, this application provides a method for training and recognizing a receipt recognition model. The target dataset is obtained through two iterations of multiple recognition models on different datasets, thereby training multiple recognition models. It is evident that this application achieves the following effects: First, by iterating through multiple training sets and using data augmentation techniques, limited data resources are fully utilized, enabling the model to learn more diverse features. Second, the model's adaptability to different receipt formats and its recognition accuracy, robustness, and generalization ability are improved. Third, reliance on manually defined templates is reduced, development and maintenance costs are lowered, and the system's usability and customer acceptance are improved.

[0051] Please see Figure 1 The architecture of the return receipt recognition model training system in this application embodiment includes:

[0052] Server 101 and client 102. When training the receipt recognition model, server 101 can acquire multiple first original receipt samples and an initial receipt training set. It then uses multiple pre-trained recognition models to identify these first original receipt samples and determines multiple second original receipt samples whose recognition information is consistent across the pre-trained models. Based on each recognition information category, the server 101 categorizes these second original receipt samples and identifies multiple target receipt samples corresponding to each category from these samples to construct a target receipt training set. The server then uses the pre-trained recognition models to identify these target receipt samples. When the loss between the predicted target recognition information and the labeled target recognition information of each pre-trained recognition model reaches a preset convergence condition, individually fine-tuned recognition models are obtained. Client 102 can call the individually fine-tuned recognition models from server 101 to identify new receipts.

[0053] based on Figure 1 Please refer to the order recognition model training system shown below. Figure 2 , Figure 2 This is a flowchart illustrating a receipt recognition model training method disclosed in an embodiment of this application. The method includes:

[0054] 201. Obtain multiple first original receipt samples and initial receipt training set. The first original receipt samples are original receipt samples for multiple recognition models with consistent recognition information. The original receipt samples are labeled with corresponding recognition information. The initial receipt training set includes multiple initial receipt samples corresponding to each category of completeness of recognition information. The initial receipt samples are labeled with initial recognition information.

[0055] In one optional implementation, receipts may include, but are not limited to, bank receipts (such as transfer receipts, payment receipts, etc.), tax receipts, or internal company documents. Bank receipts are documents provided by banks to customers when processing customer settlement transactions (including various payment and receipt transactions conducted by customers through banks) to prove that a transaction occurred and funds were received or paid. They play an important role in accounting and business verification. Please refer to [link / reference needed] for specific bank receipts. Figure 2-1-1 , Figure 2-1-1 This is a schematic diagram of a bank receipt disclosed in an embodiment of this application. Figure 2-1-1 It is known that the information includes the payee's name, payer's name, payee's account number, payer's account number, receipt number, transaction amount, receipt summary, transaction time, currency, remarks, and bank seal. The first original receipt sample is a sample whose identification information is consistent after cross-validation by multiple recognition models, used to ensure the high quality and reliability of the training data. The completeness category of identification information refers to the classification according to the completeness of key information on the receipt (including but not limited to payee's name, payer's name, payee's account number, payer's account number, receipt number, transaction amount, receipt summary, transaction time, etc.), such as whether the payee and payer information is complete, or whether other key information is complete, etc., aiming to optimize the training data structure and improve the model's adaptability to different data completeness levels.

[0056] 202. Use multiple pre-trained recognition models to identify multiple first original receipt samples, and identify multiple second original receipt samples whose recognition information is consistent with the multiple pre-trained recognition models. The multiple pre-trained recognition models are trained from the initial receipt training set.

[0057] In one optional implementation, multiple first original receipt samples are identified using multiple pre-trained recognition models, and multiple second original receipt samples with consistent recognition information corresponding to the multiple pre-trained recognition models are identified. The aim is to screen out high-confidence samples through multi-model cross-validation, thereby further improving the quality of training data.

[0058] 203. Based on each identification information category, classify multiple second original receipt samples, and determine multiple target receipt samples corresponding to each identification information category from the multiple second original receipt samples corresponding to each identification information category, so as to construct a target receipt training set. The target receipt samples are labeled with target identification information.

[0059] In one alternative implementation, the various identification information categories may include, but are not limited to, categories based on business type (e.g., "business type is receipt", "business type is payment", "business type is transfer") or by combination of bank name and business type (e.g., "Bank A - receipt", "Bank A - payment", "Bank B - receipt", etc.) to perform fine-grained classification of receipts, helping the model to accurately identify different categories of receipts in multi-task learning and enhancing its adaptability to complex receipt formats.

[0060] 204. Use multiple pre-trained recognition models to identify multiple target receipt samples, obtain the predicted target recognition information of each target receipt sample corresponding to each of the multiple pre-trained recognition models, and when the loss between the predicted target recognition information and the labeled target recognition information of each of the multiple pre-trained recognition models reaches the preset convergence condition, obtain multiple individually fine-tuned recognition models, which are then used to identify receipts.

[0061] In one alternative implementation, individual fine-tuning of the representations involves individually fine-tuning each recognition model during training to optimize its performance. The recognition model can be a large model or other machine learning models. For example, assuming Qwen2.5vl-7B is used as recognition model A and InternVL3-8B as recognition model B, both trained using the LoRa method for a total of 10 epochs.

[0062] Thus, this application constructs a basic dataset covering various application scenarios (such as scenarios with different completeness levels and different categories of recognition information) through two iterations of the training set using the BootStrapping method (a data sampling and augmentation technique). A higher-quality target training set is obtained through filtering and optimization. Combined with LoRa fine-tuning (an efficient fine-tuning method), limited data resources are effectively utilized, enabling the model to learn more diverse features, thereby improving its adaptability to different receipt formats and its recognition accuracy, robustness, and generalization ability.

[0063] In one optional implementation, obtaining multiple first original receipt samples and an initial receipt training set includes: obtaining an original receipt training set, which includes multiple original receipt samples from different data sources, each original receipt sample being labeled with corresponding identification information; using multiple identification models to identify the multiple original receipt samples, obtaining identification information for each of the multiple original receipt samples corresponding to each identification model; determining multiple first original receipt samples whose identification information is consistent across the multiple identification models; classifying the multiple original receipt samples based on the completeness category of the identification information labeled on the multiple original receipt samples; and determining multiple initial receipt samples corresponding to each completeness category from the multiple original receipt samples corresponding to each completeness category, thereby constructing an initial receipt training set, wherein the initial receipt samples are labeled with initial identification information.

[0064] Specifically, different data sources refer to the original receipt samples coming from different banks or other financial institutions, ensuring data diversity and breadth. "Determining multiple initial receipt samples corresponding to each completeness category from multiple original receipt samples corresponding to each completeness category" signifies classifying the original receipt samples based on the completeness of information in the receipt samples (such as whether the payee and payer information is complete). Samples meeting the requirements are selected from each category and used as part of the initial receipt training set for subsequent model training. This ensures the diversity and representativeness of the training set, helping the model better adapt to receipt recognition tasks in various situations.

[0065] In this way, obtaining original receipt samples from different data sources such as different banks or other financial institutions ensures the breadth and diversity of the data. This allows the trained model to adapt to various receipt formats and contents, improving the model's versatility and generalization ability. Secondly, classifying the receipt samples based on the completeness of information within them and selecting qualified samples from each category to construct the initial receipt training set ensures the quality and representativeness of the training data. This helps the model maintain high recognition accuracy and robustness when processing receipts with varying levels of completeness.

[0066] In one optional implementation, an original receipt training set is obtained, which includes multiple original receipt samples from different data sources. The original receipt samples are labeled with corresponding recognition results. This includes: obtaining original receipt file samples; performing text pagination on the original receipt file samples to obtain multiple original receipt pagination samples, which include multiple original receipt samples from different data sources; using a pre-trained segmentation model to segment multiple original receipt samples from the multiple original receipt pagination samples; parsing the multiple original receipt samples using a preset rule template; and using the parsed information of the multiple original receipt samples as the recognition information labeled with the multiple original receipt samples to construct the original receipt training set.

[0067] Specifically, a receipt file refers to a document provided by a bank or other financial institution to record transaction information. This file includes various receipts such as transfer receipts and payment receipts. These receipts typically contain key information such as the account information of both parties, the transaction amount, and the transaction time. Text pagination refers to the process of dividing a long text file into multiple independent pages (receipt pagination). Each page (receipt pagination) contains one or more specific receipt records or portions of content. A preset rule template is a set of predefined extraction rules based on the fixed format, keywords, or data location characteristics of the receipt, used to identify and extract key information from the receipt. During parsing, the original receipt sample can be processed according to the preset rule template to identify key information that conforms to the rules and extract it as the identification result. For example, if the amount field on the receipt is always located in a fixed position, the rule template can define this position, and then the amount information can be directly extracted from that position during parsing.

[0068] In this way, by integrating receipt files from different data sources and performing text pagination, combined with a pre-trained segmentation model and preset rule templates, key information can be efficiently extracted to construct a high-quality original receipt training set. This achieves the following technical effects: First, it integrates diverse data sources to improve the model's generalization ability; second, it improves data processing efficiency and annotation quality through structured processing and automated information extraction; and third, it lays a solid foundation for subsequent model training, enhancing the accuracy and adaptability of receipt recognition.

[0069] In one alternative implementation, the general logic in the preset rule template can be extracted and organized into verification rules or prompt word rules for the recognition model to improve the recognition capability of the recognition model.

[0070] Specifically, the preset rule templates are formulated based on features such as fixed receipt formats, keywords, and data locations. General logic refers to the common parts of these preset rule templates, such as the amount being in the upper right corner of the receipt, the date format being "YYYY-MM-DD," and the payee information starting with "Payee." Validation rules are used to verify the correctness of the information extracted by the model. For example, if the preset rule template specifies the date format as "YYYY-MM-DD," a validation rule can be formulated as "Check if the extracted date conforms to this format. If it does not, the information may be incorrect." Hint rules are used to guide the recognition model to extract information more accurately. For example, if payee information often begins with "Payee," a hint rule can be formulated as "Prompt the model to look for content starting with 'Payee' when extracting payee information," making the model's extraction more accurate.

[0071] In this way, by extracting the general logic from the preset rule template and organizing it into verification rules or prompt word rules, additional guidance and verification mechanisms can be provided to the recognition model. This helps the recognition model to more accurately identify and extract key information from receipts, thereby improving the overall recognition capability.

[0072] To facilitate understanding of the training schemes for the various recognition models in the embodiments of this application, a specific example is given below:

[0073] I. Dataset Construction (Constructing the dataset twice):

[0074] This embodiment of the application can construct the dataset twice using the BootStrapping method. Each data entry in the dataset is formatted as (img, label). `img` represents a single receipt image, and `label` represents the label for that receipt. The label includes the payee's name, payer's name, payee's account number, payer's account number, receipt number, transaction amount, receipt summary, and transaction time. The dataset construction process is as follows:

[0075] 1. Collect 10,000 receipt PDF files and use a pre-trained segmentation model to segment these PDF files into 200,000 receipt images.

[0076] 2. Initial dataset construction (initial order training set):

[0077] a) Develop a rule engine (preset rule template): Develop a rule engine that supports recognizing the receipt formats of the top 30 banks.

[0078] b) Use the rule engine for identification: Use the engine to identify 200,000 receipt images and extract the receipt images that can be matched by the rule engine and their corresponding tags (i.e., the labeled identification information, such as the name of the payee, the name of the payer, etc.).

[0079] c) Classification and Sampling: Based on the OCR results of the receipt images, the data obtained in the previous step is divided into four data types ("explicitly showing the payee and payer", "explicitly showing only the payee", "explicitly showing only the payer", and "neither explicitly showing the payee nor the payer"). 25,000 data points are randomly selected from each type to form the initial receipt training set, which serves as the initial dataset. For receipt images of the "explicitly showing the payee and payer" data type, please refer to [link / reference]. Figure 2-1 Please refer to the receipt image for the data type that "only explicitly shows the payee". Figure 2-2 Please refer to the receipt image for the data type that "only explicitly shows the payer". Figure 2-3 Please refer to the receipt image for the data type "Neither the payee nor the payer is explicitly given". Figure 2-4 .

[0080] 3. Train the initial recognition model:

[0081] The two recognition models, Qwen2.5vl-7B and InternVL3-8B, were trained using the initially constructed dataset (initial return order training set).

[0082] 4. Reconstruct the training dataset (target order training set):

[0083] a) Multi-model recognition: Use two large models (Qwen2.5vl-72B and InternVL3-78B) to identify the bank name and business type of all 200,000 receipt images, and take 50,000 data points with consistent results.

[0084] b) Further screening: Use the pre-trained model to re-identify these 50,000 data points, and select the 20,000 data points with consistent results.

[0085] c) Classification and Sampling: Data with consistent identification results from the previous step are classified according to "Bank Name_Business Type," resulting in over 200 categories. A maximum of 120 data entries are taken from each category, resulting in a total of 11,200 data entries. These 11,200 data entries will be used as the target receipt training set. For receipt images with the data type "Bank A_Fee Receipt," please refer to [link / reference]. Figure 2-5 Please refer to the image of the receipt with data type "Bank B_Tax Payment Receipt" Figure 2-6 Please refer to the receipt image with data type "Bank C_Transaction Receipt" for details. Figure 2-7 Please refer to the receipt image with data type "Bank D_Deduction Receipt" for details. Figure 2-8 .

[0086] II. Model Training:

[0087] In this embodiment, the Qwen2.5vl-7B model is used as recognition model A, and the InternVL3-8B model is used as recognition model B. Both are trained using the LoRa method for a total of 10 epochs. By training on the target receipt training set, the recognition model can better learn the features of the receipt, thereby improving recognition accuracy and adaptability.

[0088] In an optional implementation, before segmenting multiple original order samples from multiple original order pagination samples using a pre-trained segmentation model, the method further includes: obtaining a target order pagination training set, which includes multiple target order pagination samples corresponding to various layouts, each target order pagination sample including at least one target order sample of various order sizes and layouts, each target order sample labeled with a corresponding order detection box; using the segmentation model to determine a predicted order detection box for at least one target order sample from the multiple target order pagination samples, the predicted order detection box being used to segment at least one target order sample from the multiple target order pagination samples to obtain at least one target order sample; and obtaining a trained segmentation model when the predicted order detection box for at least one target order sample passes the verification, and when the loss between the predicted order detection box for at least one target order sample and the labeled order detection box for at least one target order sample reaches a preset convergence condition.

[0089] Specifically, order size refers to the aspect ratio of the order, used to distinguish order sizes. Various order sizes can include, but are not limited to, "large order: aspect ratio less than 1," "medium order: aspect ratio greater than 1 and less than or equal to 3," and "small order: aspect ratio greater than 3." Order layout characterizes how orders are arranged on the page. Various order layouts can include, but are not limited to, "three orders evenly arranged at the top, middle, and bottom of an image," "two orders at the top and middle of an image," "two orders at the top and bottom of an image," "one order appearing above an image," "one order randomly appearing at either the top or bottom of an image," and "the entire order pagination is one order," etc. This layout information helps the recognition model understand the possible distribution of orders on the page, thus enabling more accurate segmentation. This application can use the YOLOv5 model architecture for training, with a training period of 100 epochs.

[0090] In this way, by using a target training set with diverse order sizes and layouts, and training the YOLOv5 model for 100 epochs, the segmentation model can efficiently adapt to different order formats, improve the generalization ability of the segmentation model, ensure accurate segmentation of orders with various complex layouts and sizes, and thus improve the overall efficiency and accuracy of the order recognition system.

[0091] In one optional implementation, obtaining the target order pagination training set includes: obtaining an initial order pagination training set, the initial order pagination training set including multiple initial order pagination samples corresponding to various layouts, the multiple initial order pagination samples corresponding to various layouts including at least one initial order sample of various order sizes and various order layouts, and determining multiple target order pagination samples corresponding to various layouts from the multiple initial order pagination samples corresponding to various layouts to construct the target order pagination training set.

[0092] Specifically, the target receipt pagination training set can be an artificially constructed dataset used to train the segmentation model. The dataset construction process is as follows: the training data format is (img, boxes), where img is the image format of a single-page PDF (receipt pagination), and boxes are the detection boxes for bank receipts in the image (receipt pagination). By collecting and analyzing a large number of single-page PDFs, three different receipt sizes can be identified, including but not limited to "large receipt: aspect ratio less than 1", "medium receipt: aspect ratio greater than 1 and less than or equal to 3", and "small receipt: aspect ratio greater than 3". Multiple receipts contained in a single-page PDF image can have the same or different sizes; this application does not impose any limitations. Receipt layouts can include, but are not limited to, the following seven categories:

[0093] 1. An image contains three small receipts, which are evenly arranged at the top, middle, and bottom of the image.

[0094] 2. An image contains two small receipts, located at the top and center of the image.

[0095] 3. An image contains two small receipts, located at the top and bottom of the image.

[0096] 4. A small receipt is contained within an image, and the small receipt appears above the image.

[0097] 5. An image contains two return slips, located at the top and bottom of the image.

[0098] 6. An image contains a return ticket, which appears randomly at either the top or bottom of the image.

[0099] 7. If an image contains a large order, the entire page becomes a large order.

[0100] Understandably, we can collect single-page PDF images of different formats, gathering 10 images for each format, for a total of 100 formats and 1000 images in total (200 large spreadsheets, 300 medium spreadsheets, and 500 small spreadsheets). Then, based on the summarized patterns, we can generate 100,000 training images by randomly combining these spreadsheet images. Please refer to the generated single-page PDF images (split spreadsheets) for details. Figure 2-9 , Figure 2-9 This is a schematic diagram of a receipt pagination method disclosed in an embodiment of this application. Figure 2-9 As can be seen, the three small receipts are evenly arranged in the top, middle, and bottom positions of the image.

[0101] In this way, constructing training datasets through artificial synthesis can reduce annotation costs while improving dataset quality.

[0102] In one optional implementation, obtaining the target order pagination training set includes: obtaining an initial order pagination training set, which includes multiple initial order pagination samples corresponding to various layouts, each initial order pagination sample including at least one initial order sample of various order sizes and layouts; obtaining order pagination samples of new layouts; and adding the new layout order pagination samples to the initial order pagination training set to obtain the target order pagination training set.

[0103] Specifically, once the system is online, if a new receipt format is encountered that is not currently supported, support for that new version can be quickly implemented through simple steps. The specific steps could be: first, collect a small number of samples, add them to the existing dataset, regenerate the dataset, and then retrain the model. For example, you only need to collect 10 receipt images of the new version, then add these 10 images to the existing receipt image set, regenerate the training dataset using the updated receipt image set, and then retrain the segmentation model using the new dataset.

[0104] In this way, the entire process takes only two days, requires minimal development work, and can be quickly deployed to support the new receipt format. Through simple data updates and model retraining, it can quickly adapt to the new receipt format, ensuring that the system can handle diverse receipt formats, thereby improving the system's usability and scalability.

[0105] In one optional implementation, determining that the predicted order detection box of at least one target order sample passes the verification includes at least one of the following: if there are multiple target order samples, and it is determined that there is no overlap between the order content within the predicted order detection boxes of the multiple target order samples, then the predicted order detection box of the multiple target order samples passes the verification; if it is determined that the aspect ratio of the predicted order detection box of at least one target order sample meets a preset aspect ratio range, then the predicted order detection box of at least one target order sample passes the verification; if it is determined that the richness of the text content within the predicted order detection box of at least one target order sample meets a preset richness range, then the predicted order detection box of at least one target order sample passes the verification; if it is determined that the amount of text outside the predicted order detection box of at least one target order sample in the multiple target order pagination samples is less than a preset amount of text, then the predicted order detection box of at least one target order sample passes the verification.

[0106] Specifically, an automatic verification module for the segmentation results can use heuristic rules to determine whether the segmentation results are correct, thereby improving the efficiency of automation. The specific rules are as follows:

[0107] 1. Check if there is overlap in the segmented receipts: If the contents of the prediction detection boxes of multiple receipt samples do not overlap, then these detection boxes are considered to be qualified.

[0108] 2. Check if the aspect ratio of the return receipt is reasonable: If the aspect ratio of the predicted return receipt detection box is within the preset range, the detection box is considered to be qualified, so as to avoid incorrect identification of extreme aspect ratios.

[0109] 3. Check the richness of the text content: If the richness of the text content in the detection box is appropriate, neither too little nor too much, the detection box is considered to have passed the verification, in order to prevent the receipt content from being missing or containing too much irrelevant text.

[0110] 4. Check if the amount of text outside the receipt is appropriate: If the amount of text outside the check box is within a reasonable range, the check box is considered to have passed the verification, in order to avoid incomplete receipt content or the inclusion of too much external text.

[0111] In this way, through these rules, the automatic verification module can determine whether the segmentation is successful, reducing manual intervention and improving the level of automation.

[0112] It's worth noting that the segmentation model can be trained using manually constructed datasets, and this method offers the following advantages and features: First, it can quickly support new receipt formats: If a receipt format not supported by the current segmentation model is encountered, support for that new version can be quickly achieved within 2-3 days by constructing a manually constructed dataset. This provides good adaptability and scalability, enabling rapid response to new business needs or changes. Second, it automatically determines the correctness of the segmentation: The segmentation model can automatically determine whether the segmentation result is correct. If the segmentation is incorrect, the model will throw a "segmentation failed" exception, prompting the user for manual confirmation. If the segmentation is correct, the receipt will automatically proceed to the next step of large-scale model recognition, without manual intervention. Third, it improves automation: This ability to automatically determine the correctness of the segmentation significantly improves the automation level of the entire processing flow. Traditional processing methods may require customers to confirm the segmentation of each receipt individually, while the method in this application reduces this workload of manual confirmation and improves work efficiency.

[0113] Please see Figure 3 , Figure 3 This is a flowchart illustrating a receipt identification method disclosed in an embodiment of this application. The method includes:

[0114] 301. Obtain at least one receipt.

[0115] In one alternative implementation, the receipt may include, but is not limited to, bank receipts (such as transfer receipts, payment receipts, etc.), tax receipts, or internal company documents.

[0116] 302. Utilize respectively, such as Figure 2 Each of the individually fine-tuned recognition models identifies at least one receipt, resulting in the initial recognition result of at least one receipt output by each of the individually fine-tuned recognition models.

[0117] In one alternative implementation, the same receipt is identified using multiple individually fine-tuned recognition models to obtain preliminary results for each model, which are then used to determine the final recognition result.

[0118] 303. Based on the determination of whether the initial recognition results of at least one receipt output by each of the multiple recognition models after individual fine-tuning are consistent, determine the target recognition result of at least one receipt.

[0119] In one optional implementation, if the initial recognition results of at least one receipt output by each of the multiple individually fine-tuned recognition models are consistent, then the consistent initial recognition results are taken as the target recognition results. If the initial recognition results of at least one receipt output by each of the multiple individually fine-tuned recognition models are inconsistent, then the target recognition result of at least one receipt is determined based on the specific inconsistency. The method for determining the target recognition result of at least one receipt may include, but is not limited to, manual intervention.

[0120] In this way, by employing a separately fine-tuned recognition model, the system can intelligently handle some unknown changes in receipts. The recognition model can automatically adapt to new receipt formats and changes, making the system more flexible and adaptable. Secondly, the recognition model can learn to automatically adapt to new receipt formats, reducing the workload and related costs of manually maintaining a large number of receipt template rules.

[0121] In one optional implementation, obtaining at least one receipt includes: obtaining a receipt file and paginating the receipt file to obtain receipt pages, wherein each receipt page includes at least one receipt, using methods such as... Figure 2 The pre-trained segmentation model segments at least one order from the order page.

[0122] Specifically, a transaction receipt is a document provided by a bank or other financial institution to record transaction information. This includes various receipts such as transfer receipts and payment receipts, which typically contain key information such as the account information of both parties, the transaction amount, and the transaction time. Text pagination refers to the process of dividing a long text file into multiple independent pages (receipt paginations). Each page (receipt pagination) contains one or more specific receipt records or portions of content.

[0123] In this way, by paginating the receipt file and using a pre-trained segmentation model to extract individual receipts, the efficiency and accuracy of receipt recognition are effectively improved.

[0124] In one alternative implementation, using, for example Figure 2 The pre-trained segmentation model segments at least one order from the order pagination, including: parsing at least one order using a preset rule template to obtain the parsing result of at least one order; if the parsing result of at least one order does not meet the preset parsing standard, then executing the following steps: Figure 2 The pre-trained segmentation model segments at least one order from the order page.

[0125] Specifically, the "preset rule template" is a set of extraction rules predefined based on the fixed format, keywords, or data positions of the receipt, used to identify and extract key information from the receipt. During the parsing process, the original receipt sample can be processed according to the preset rule template to identify key information that conforms to the rules and extract it as the identification result. For example, if the amount field of the receipt is always located in a fixed position, the rule template can define this position, and then the amount information can be directly extracted from that position during parsing. "If the parsing result of at least one receipt does not meet the preset parsing standard" means that after parsing the receipt using the rule template, the obtained parsing result does not meet the pre-set completeness and accuracy requirements. This may be because the receipt format is new, and the new receipt format does not completely match the existing rule template, resulting in missing or incorrect key information.

[0126] In this way, the dual checks of preset rule templates and recognition models effectively ensure the accuracy of receipt recognition results, reducing the workload of customer verification. Secondly, common receipts are processed using preset rule templates, while the recognition model is only invoked for infrequent receipts. This reduces the frequency of recognition model calls, saves resources, and improves system response speed. Furthermore, this combined approach adapts to new receipt formats while reducing the cost of continuously maintaining rule templates, avoiding the high workload caused by frequent rule template updates. Finally, using multiple recognition models to check the results further confirms the accuracy of the results, enhancing the stability and reliability of the system.

[0127] In one optional implementation, the preset rule template and the recognition model scheme can be separated into two independent schemes for use separately. This achieves the following technical effects: First, high flexibility: the preset rule template or the recognition model scheme can be selected based on the specific circumstances of the receipt. Receipts with fixed formats can be quickly extracted using the preset rule template scheme, while complex or uncommon receipts can be accurately identified using the recognition model scheme. Second, strong targeting: running the two schemes separately can better handle different scenarios. The preset rule template scheme is suitable for common receipts with clear rules, resulting in low computational costs; the recognition model scheme is suitable for complex and variable receipts, adapting to special situations through training to improve recognition accuracy. Third, easy maintenance and optimization: the separated schemes are easier to maintain and optimize independently. The preset rule template can be updated independently to adapt to minor changes; the recognition model can also be trained and tuned independently to improve the ability to handle complex receipts without affecting the stable operation of the other scheme. Fourth, reasonable resource allocation: independent operation allows for on-demand allocation of system resources. When processing large batches of common receipts, the preset rule template scheme is mainly used to save resources; for a small number of complex receipts, more resources are allocated to the recognition model scheme to ensure overall processing efficiency and effectiveness.

[0128] In one optional implementation, after determining the target recognition result of at least one receipt based on whether the initial recognition results of at least one receipt corresponding to the output of each of the multiple recognition models after individual fine-tuning are consistent, the method further includes: determining at least one target receipt whose initial recognition results corresponding to each of the multiple recognition models are inconsistent, and obtaining the target recognition information of at least one target receipt obtained after user intervention; constructing a new receipt training set based on at least one target receipt and the target recognition result of at least one target receipt; and retraining the multiple recognition models individually based on the new receipt training set to obtain the retrained multiple recognition models.

[0129] Specifically, when multiple recognition models produce inconsistent initial recognition results for the same receipt, the system filters out these receipts and obtains the correct recognition information confirmed by the user. These user-intervened receipts and their correct recognition results are added to the training set, forming new training samples. Then, the recognition model is retrained using this new training set to improve its ability to recognize receipt content that was previously difficult to accurately identify, thereby improving the overall accuracy and reliability of the model.

[0130] To facilitate understanding of the receipt recognition scheme provided in this application, a specific example is given below:

[0131] Please see Figure 3-1 , Figure 3-1 This is a flowchart illustrating another receipt identification method disclosed in an embodiment of this application. Figure 3-1 The process of the receipt identification method is as follows:

[0132] Step 1: Determine the feasibility of text parsing: First, check if the incoming receipt file can be directly parsed. Text parsing provides more accurate text information and avoids potential errors associated with using Optical Character Recognition (OCR).

[0133] Step 2, Text Parsing and Pagination: If the receipt file can be directly parsed as text, then paginate it and save the text content of each page.

[0134] Step 3, OCR Recognition: For receipts that cannot be directly parsed, first convert the file to image format page by page, then use OCR technology to recognize the text content of each page and save the text.

[0135] Step 4: Template Feature Retrieval and Information Extraction: Using the paginated text content, retrieve whether it matches the preset template features. If a match is successful, segment the order and extract key information according to the preset template rules.

[0136] Step 5: Call the segmentation model: If the receipt format cannot be confirmed through template features, call the pre-trained segmentation model to segment the receipt.

[0137] Step 6, Multimodal Large Model Recognition: Input the segmented single receipts into two pre-trained multimodal large models (two pre-trained recognition models) to obtain key information.

[0138] Step 7: Model Result Comparison: Compare whether the key information returned by the two large models is consistent. If they are consistent, the recognition result is considered correct, and the result is returned.

[0139] Step 8: Manual intervention: If the results of the two models are inconsistent, the identification is considered to have failed, and a failure message is returned for the customer to perform manual intervention.

[0140] Step 9, Model Optimization: Use the correct recognition results obtained after manual intervention as new training data to further train and optimize the large model in order to improve its recognition ability.

[0141] In this way, by combining text parsing, OCR technology, template matching, machine learning model recognition, and human intervention, the efficiency and accuracy of receipt recognition can be ensured. At the same time, through continuous model training and optimization, the intelligence level and adaptability of the system are improved.

[0142] For further details, please refer to Figure 4 One embodiment of the receipt recognition model training device in this application includes:

[0143] The acquisition unit is used to acquire multiple first original receipt samples and an initial receipt training set. The first original receipt samples are original receipt samples whose recognition information is consistent across multiple recognition models. The original receipt samples are labeled with corresponding recognition information. The initial receipt training set includes multiple initial receipt samples corresponding to each category of completeness of recognition information. The initial receipt samples are labeled with initial recognition information.

[0144] The determining unit is used to identify the multiple first original receipt samples using multiple pre-trained recognition models, and to determine multiple second original receipt samples whose recognition information is consistent with the multiple pre-trained recognition models. The multiple pre-trained recognition models are trained from the initial receipt training set.

[0145] The determining unit is further configured to classify the plurality of second original return receipt samples based on each identification information category, and determine the plurality of target return receipt samples corresponding to each identification information category from the plurality of second original return receipt samples corresponding to each identification information category, so as to construct a target return receipt training set, wherein the target return receipt samples are labeled with target identification information;

[0146] The acquisition unit is further configured to use the pre-trained multiple recognition models to identify the multiple target return receipt samples respectively, obtain the predicted target recognition information of the multiple target return receipt samples corresponding to each of the pre-trained multiple recognition models, and when the loss between the predicted target recognition information and the labeled target recognition information corresponding to each of the pre-trained multiple recognition models reaches a preset convergence condition, obtain the multiple recognition models after individual fine-tuning, and the multiple recognition models after individual fine-tuning are used to identify return receipts.

[0147] In one alternative implementation, the acquisition unit may be used for:

[0148] A training set of original receipts is obtained, comprising multiple original receipt samples from different data sources. Each original receipt sample is labeled with corresponding identification information. Multiple identification models are used to identify these original receipt samples, obtaining identification information for each of the multiple identification models. Multiple first original receipt samples with identical identification information for all multiple identification models are identified. Based on the completeness category of the identification information labeled on the multiple original receipt samples, the multiple original receipt samples are classified. Multiple initial receipt samples corresponding to each completeness category are determined from the original receipt samples corresponding to each completeness category to construct the initial receipt training set. These initial receipt samples are labeled with initial identification information.

[0149] In one alternative implementation, the acquisition unit may be used for:

[0150] Obtain original receipt file samples, perform text pagination on the original receipt file samples to obtain multiple original receipt pagination samples. The original receipt pagination samples include multiple original receipt samples from different data sources. Use a pre-trained segmentation model to segment the multiple original receipt pagination samples into multiple original receipt samples. Use a preset rule template to parse the multiple original receipt samples, and use the parsed information of the multiple original receipt samples as the identification information for the annotation of the multiple original receipt samples to construct the original receipt training set.

[0151] In one alternative implementation, the acquisition unit may be used for:

[0152] Obtain a target order pagination training set, which includes multiple target order pagination samples corresponding to various layouts. The multiple target order pagination samples corresponding to various layouts include at least one target order sample with various order sizes and various order layouts. The target order samples are labeled with corresponding order detection boxes.

[0153] Determining a unit can be used for:

[0154] A segmentation model is used to determine the predicted return detection box of at least one target return sample from the plurality of target return pagination samples. The predicted return detection box is used to segment the at least one target return sample from the plurality of target return pagination samples to obtain at least one target return sample.

[0155] The acquisition unit can be used for:

[0156] The trained segmentation model is obtained when the predicted return detection box of the at least one target return sample is verified as qualified, and when the loss between the predicted return detection box of the at least one target return sample and the labeled return detection box of the at least one target return sample reaches the preset convergence condition.

[0157] In one alternative implementation, the acquisition unit may be used for:

[0158] Obtain an initial order pagination training set, which includes multiple initial order pagination samples corresponding to various layouts. The multiple initial order pagination samples corresponding to various layouts include at least one initial order sample of various order sizes and various order layouts. Determine multiple target order pagination samples corresponding to various layouts from the multiple initial order pagination samples corresponding to various layouts to construct the target order pagination training set.

[0159] In one alternative implementation, the acquisition unit may be used for:

[0160] Obtain an initial order pagination training set, which includes multiple initial order pagination samples corresponding to various layouts. The multiple initial order pagination samples corresponding to various layouts include at least one initial order sample of various order sizes and various order layouts. Obtain order pagination samples of new layouts and add the order pagination samples of new layouts to the initial order pagination training set to obtain the target order pagination training set.

[0161] In one alternative implementation, the determining unit may be used for:

[0162] If there are multiple target order samples, and it is determined that there is no overlap between the order content within the predicted order detection boxes of the multiple target order samples, then the predicted order detection boxes of the multiple target order samples are deemed to be qualified. If it is determined that the aspect ratio of the predicted order detection box of at least one target order sample meets the preset aspect ratio range, then the predicted order detection box of at least one target order sample is deemed to be qualified. If it is determined that the richness of the text content within the predicted order detection box of at least one target order sample meets the preset richness range, then the predicted order detection box of at least one target order sample is deemed to be qualified. If it is determined that the amount of text outside the predicted order detection box of at least one target order sample in the multiple target order pagination samples is less than the preset amount of text, then the predicted order detection box of at least one target order sample is deemed to be qualified.

[0163] For further details, please refer to Figure 5 One embodiment of the receipt identification device in this application includes:

[0164] The acquisition unit is used to acquire at least one receipt;

[0165] Identification unit, used to respectively utilize such as Figure 2 The individually fine-tuned multiple recognition models identify the at least one receipt, and obtain the initial recognition results of the at least one receipt output by each of the individually fine-tuned multiple recognition models;

[0166] The determining unit is used to determine the target recognition result of the at least one receipt based on the determination result of whether the initial recognition results of the at least one receipt output by the multiple individually fine-tuned recognition models are consistent.

[0167] In one alternative implementation, the acquisition unit may be used for:

[0168] Obtain the receipt file and paginate the receipt file to obtain receipt pages, each receipt page including at least one receipt, using methods such as... Figure 2 The pre-trained segmentation model segments at least one order from the order page.

[0169] In one alternative implementation, the acquisition unit may be used for:

[0170] The at least one receipt is parsed using a preset rule template to obtain the parsing result of the at least one receipt. If the parsing result of the at least one receipt does not meet the preset parsing standard, then the process of using the preset rule template is executed. Figure 2 The step of the pre-trained segmentation model segmenting at least one order from the order page.

[0171] In one alternative implementation, the acquisition unit may also be used for:

[0172] Identify at least one target order whose initial recognition results for each of the multiple recognition models are inconsistent, and obtain the target recognition result of the at least one target order obtained after user intervention. Based on the at least one target order and its target recognition result, construct a new order training set. Based on the new order training set, retrain the multiple recognition models individually to obtain multiple recognition models that have been retrained.

[0173] For further details, please refer to Figure 6 One embodiment of the electronic device in this application includes:

[0174] Central processing unit 601, memory 605, input / output interface 604, wired or wireless network interface 603, and power supply 602;

[0175] Memory 605 is either a short-term storage memory or a persistent storage memory;

[0176] The central processing unit 601 is configured to communicate with the memory 605 and execute instructions stored in the memory 605 to perform the aforementioned operations. Figure 2 or Figure 3 The method in the illustrated embodiment.

[0177] Furthermore, embodiments of this application also provide a computer-readable storage medium, which includes instructions that, when executed on a computer, cause the computer to perform the aforementioned... Figure 2 or Figure 3 The method in the illustrated embodiment.

[0178] Furthermore, embodiments of this application also provide a computer program product containing instructions, which, when run on a computer, causes the computer to perform the aforementioned... Figure 2 or Figure 3 The method in the illustrated embodiment.

[0179] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0180] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0181] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0182] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0183] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0184] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for training a receipt recognition model, characterized in that, include: Multiple first original receipt samples and an initial receipt training set are obtained. The first original receipt samples are original receipt samples whose recognition information is consistent across multiple recognition models. The original receipt samples are labeled with corresponding recognition information. The initial receipt training set includes multiple initial receipt samples corresponding to each category of completeness of recognition information. The initial receipt samples are labeled with initial recognition information. The multiple first original receipt samples are identified using multiple pre-trained recognition models, and multiple second original receipt samples with consistent recognition information corresponding to the multiple pre-trained recognition models are identified. The multiple pre-trained recognition models are trained from the initial receipt training set. The multiple second original return receipt samples are classified based on each identification information category, and multiple target return receipt samples corresponding to each identification information category are determined from the multiple second original return receipt samples corresponding to each identification information category to construct a target return receipt training set, wherein the target return receipt samples are labeled with target identification information. The multiple target return receipt samples are identified by the multiple pre-trained recognition models respectively, and the predicted target recognition information of the multiple target return receipt samples corresponding to each of the multiple pre-trained recognition models is obtained. When the loss between the predicted target recognition information and the labeled target recognition information corresponding to each of the multiple pre-trained recognition models reaches a preset convergence condition, multiple individually fine-tuned recognition models are obtained. The multiple individually fine-tuned recognition models are used to identify return receipts.

2. The method according to claim 1, characterized in that, The process of obtaining multiple first original return receipt samples and the initial return receipt training set includes: Obtain the original receipt training set, which includes multiple original receipt samples from different data sources, and the original receipt samples are labeled with corresponding identification information; The multiple original receipt samples are identified using multiple recognition models respectively, and the recognition information of the multiple original receipt samples corresponding to each of the multiple recognition models is obtained. The multiple first original receipt samples whose recognition information corresponding to the multiple recognition models is consistent are then identified. The original return receipts are classified based on the completeness categories of the identification information labeled on the original return receipts. Then, multiple initial return receipts corresponding to each completeness category are determined from the original return receipts corresponding to each completeness category to construct the initial return receipt training set. The initial return receipts are labeled with initial identification information.

3. The method according to claim 2, characterized in that, The process of obtaining the original receipt training set includes multiple original receipt samples from different data sources. Each original receipt sample is labeled with a corresponding recognition result, including: Obtain a sample of the original receipt file; The original receipt file sample is paginated to obtain multiple original receipt pagination samples, which include multiple original receipt samples from different data sources. The multiple original order pagination samples are segmented from the multiple original order pagination samples using a pre-trained segmentation model; The multiple original return receipt samples are parsed using a preset rule template, and the parsed information of the multiple original return receipt samples is used as the identification information for the annotation of the multiple original return receipt samples to construct the original return receipt training set.

4. The method according to claim 3, characterized in that, Before segmenting the multiple original order pagination samples from the multiple original order pagination samples using the pre-trained segmentation model, the method further includes: Obtain a target order pagination training set, which includes multiple target order pagination samples corresponding to various layouts. The multiple target order pagination samples corresponding to various layouts include at least one target order sample with various order sizes and various order layouts. The target order samples are labeled with corresponding order detection boxes. A segmentation model is used to determine the predicted return detection box of at least one target return sample from the plurality of target return pagination samples. The predicted return detection box is used to segment the at least one target return sample from the plurality of target return pagination samples to obtain at least one target return sample. The trained segmentation model is obtained when the predicted return detection box of the at least one target return sample is verified as qualified, and when the loss between the predicted return detection box of the at least one target return sample and the labeled return detection box of the at least one target return sample reaches the preset convergence condition.

5. The method according to claim 4, characterized in that, The process of obtaining the target order pagination training set includes: Obtain an initial order pagination training set, which includes multiple initial order pagination samples corresponding to various layouts, and the multiple initial order pagination samples corresponding to various layouts include at least one initial order sample of various order sizes and various order layouts. Multiple target order pagination samples corresponding to the various formats are determined from the multiple initial order pagination samples corresponding to the various formats, so as to construct the target order pagination training set.

6. The method according to claim 4, characterized in that, The process of obtaining the target order pagination training set includes: Obtain an initial order pagination training set, which includes multiple initial order pagination samples corresponding to various layouts, and the multiple initial order pagination samples corresponding to various layouts include at least one initial order sample of various order sizes and various order layouts. Get a sample of the new version of the order pagination; The new version of the order pagination sample is added to the initial order pagination training set to obtain the target order pagination training set.

7. The method according to claim 4, characterized in that, The determination that the predicted receipt detection box of the at least one target receipt sample is qualified includes at least one of the following situations: If there are multiple target return receipt samples, and it is determined that there is no overlap between the return receipt content in the predicted return receipt detection box of the multiple target return receipt samples, then the predicted return receipt detection box of the multiple target return receipt samples is determined to be qualified. If the aspect ratio of the predicted receipt detection frame of the at least one target receipt sample meets the preset aspect ratio range, then the predicted receipt detection frame of the at least one target receipt sample is deemed to be qualified. If it is determined that the richness of the text content within the predicted return detection box of the at least one target return sample meets the preset richness range, then the predicted return detection box of the at least one target return sample is determined to be qualified. If it is determined that the number of characters outside the predicted return detection box of at least one of the multiple target return pagination samples is less than a preset range, then the predicted return detection box of the at least one target return sample is verified as qualified.

8. A method for recognizing receipts, characterized in that, include: Obtain at least one receipt; The at least one receipt is identified by using the individually fine-tuned recognition models as described in any one of claims 1-7, and the initial recognition results of the at least one receipt output by each of the individually fine-tuned recognition models are obtained. Based on the determination of whether the initial recognition results of the at least one receipt output by the multiple individually fine-tuned recognition models are consistent, the target recognition result of the at least one receipt is determined.

9. The method according to claim 8, characterized in that, Obtaining at least one receipt includes: Obtain the receipt file and perform file pagination on the receipt file to obtain receipt pages, wherein the receipt pages include at least one receipt; The at least one order is segmented from the order page using a pre-trained segmentation model as described in any one of claims 3-7.

10. The method according to claim 9, characterized in that, The step of segmenting the at least one order from the order page using the pre-trained segmentation model as described in any one of claims 3-7 includes: The at least one receipt is parsed using a preset rule template to obtain the parsing result of the at least one receipt; If the parsing result of at least one order does not meet the preset parsing result standard, then the step of segmenting the at least one order from the order page using the pre-trained segmentation model as described in any one of claims 3-7 is executed.

11. The method according to claim 8, characterized in that, After determining the target recognition result of the at least one receipt based on whether the initial recognition results of the at least one receipt output by the multiple individually fine-tuned recognition models are consistent, the method further includes: Identify at least one target order where the initial identification results corresponding to the multiple identification models are inconsistent, and obtain the target identification result of the at least one target order after user intervention; Based on the at least one target receipt and the target identification results of the at least one target receipt, a new receipt training set is constructed; Based on the new return order training set, the multiple recognition models are trained separately again to obtain multiple recognition models that have been retrained.

12. A receipt recognition model training device, characterized in that, include: The acquisition unit is used to acquire multiple first original receipt samples and an initial receipt training set. The first original receipt samples are original receipt samples whose recognition information is consistent across multiple recognition models. The original receipt samples are labeled with corresponding recognition information. The initial receipt training set includes multiple initial receipt samples corresponding to each category of completeness of recognition information. The initial receipt samples are labeled with initial recognition information. The determining unit is used to identify the multiple first original receipt samples using multiple pre-trained recognition models, and to determine multiple second original receipt samples whose recognition information is consistent with the multiple pre-trained recognition models. The multiple pre-trained recognition models are trained from the initial receipt training set. The determining unit is further configured to classify the plurality of second original return receipt samples based on each identification information category, and determine the plurality of target return receipt samples corresponding to each identification information category from the plurality of second original return receipt samples corresponding to each identification information category, so as to construct a target return receipt training set, wherein the target return receipt samples are labeled with target identification information; The acquisition unit is further configured to use the pre-trained multiple recognition models to identify the multiple target return receipt samples respectively, obtain the predicted target recognition information of the multiple target return receipt samples corresponding to each of the pre-trained multiple recognition models, and when the loss between the predicted target recognition information and the labeled target recognition information corresponding to each of the pre-trained multiple recognition models reaches a preset convergence condition, obtain the multiple recognition models after individual fine-tuning, and the multiple recognition models after individual fine-tuning are used to identify return receipts.

13. A receipt recognition device, characterized in that, include: The acquisition unit is used to acquire at least one receipt; The identification unit is used to identify the at least one receipt using the individually fine-tuned identification models as described in any one of claims 1-7, and to obtain the initial identification results of the at least one receipt output by each of the individually fine-tuned identification models. The determining unit is used to determine the target recognition result of the at least one receipt based on the determination result of whether the initial recognition results of the at least one receipt output by the multiple individually fine-tuned recognition models are consistent.

14. An electronic device, characterized in that, include: Central processing unit and memory; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 7 or any one of claims 8 to 11.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 7 or any one of claims 8 to 11.

16. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 7 or any one of claims 8 to 11.

Citation Information

Patent Citations

  • Bank electronic receipt identification method and device

    CN115909348A

  • Defect classification method based on fine-grained recognition and attention mechanism

    CN116109629A