Methods, apparatus, equipment and storage media for classifying vouchers

By performing initial classification of each page in the document and determining the classification results of the target page, invalid pages are automatically identified and processed, solving the problem that the voucher classification model cannot identify, and achieving efficient voucher classification.

CN122489818APending Publication Date: 2026-07-31CLP JINXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CLP JINXIN TECH CO LTD
Filing Date
2026-06-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, voucher classification models cannot effectively identify invalid pages in documents, resulting in low efficiency due to reliance on manual annotation and failing to meet batch processing requirements.

Method used

By classifying each page in the document to be classified using credentials, an initial classification result is obtained, invalid pages are identified, and the classification result of the target page in the document is determined. The page classification is then automatically performed using the initial and target classification results of the model.

Benefits of technology

It enables automatic and efficient credential classification of all pages in a document, reducing manual intervention and improving processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489818A_ABST
    Figure CN122489818A_ABST
Patent Text Reader

Abstract

This application belongs to the field of data processing technology and discloses a method, apparatus, device, and storage medium for classifying vouchers. This application classifies vouchers by classifying each page of a document to be classified, obtaining an initial classification result. Then, based on the initial classification result, invalid pages in the document to be classified are identified. A target page preceding the invalid page is determined in the document to be classified, and the page classification result corresponding to the invalid page is determined based on the target page's target classification result. This application identifies invalid pages (pages that cannot be classified) based on the initial classification result of each page in the document to be classified. It identifies the target page preceding the invalid page in the document to be classified and determines the page classification result of the invalid page based on the target page's target classification result. Compared to existing methods that rely on manual annotation of invalid page types, the above method of this application can automatically and effectively classify vouchers for all pages in the document to be classified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device and storage medium for classifying vouchers. Background Technology

[0002] Voucher classification is a preliminary step in intelligent verification systems. Its purpose is to automatically categorize each page of uploaded documents into its corresponding voucher type, such as contracts, invoices, and bank statements. Typically, voucher classification is done using models. However, for pages in the document that the model cannot recognize, manual labeling of the voucher type is required, which is inefficient and cannot meet the needs of batch processing. Summary of the Invention

[0003] The main objective of this application is to provide a method, apparatus, device, and storage medium for classifying vouchers, aiming to solve the technical problem of how to automatically and effectively classify vouchers across all pages in a document.

[0004] To achieve the above objectives, this application provides a method for classifying vouchers, which includes the following steps: Classify each page in the document to be classified using vouchers to obtain the initial classification results; Based on the initial classification results, invalid pages in the document to be classified are determined. In the document to be classified, determine the target page preceding the invalid page, and determine the page classification result corresponding to the invalid page based on the target page's target classification result.

[0005] Optionally, the step of determining the target page before identifying the invalid page in the document to be classified, and determining the page classification result corresponding to the invalid page based on the target classification result of the target page, includes: In the document to be classified, determine the target page preceding the invalid page, and select the page preceding the invalid page from the target page; Determine whether the previous page is a valid page, and obtain the determination result; The page classification result corresponding to the invalid page is determined based on the judgment result and the target classification result of the target page.

[0006] Optionally, determining the page classification result corresponding to the invalid page based on the judgment result and the target classification result of the target page includes: If the determination result indicates that the previous page is a valid page, the previous category result of the previous page is selected from the target category results of the target page; The previous classification result is used as the page classification result corresponding to the invalid page; If the determination result is that the previous page is not a valid page, the target page is traversed according to a preset order based on the invalid page until a valid page is obtained; Select the valid classification result corresponding to the valid page from the target classification results, and use the valid classification result as the page classification result corresponding to the invalid page.

[0007] Optionally, after determining the target page before the invalid page in the document to be classified, and determining the page classification result corresponding to the invalid page based on the target classification result of the target page, the method further includes: Based on the initial classification results and the page classification results, determine the final classification results for each page in the document to be classified; Based on the final classification results, determine the target group corresponding to each page in the document to be classified; The number of vouchers corresponding to the document to be classified is determined based on the target group, and the voucher classification results are displayed based on the number of vouchers.

[0008] Optionally, determining the target group corresponding to each page in the document to be classified based on the final classification result includes: Based on the final classification result, each page in the document to be classified is grouped according to the voucher type to obtain the first group corresponding to each page; The filename of the page corresponding to the same voucher type is determined based on the first group; Based on the file name, the pages corresponding to the same voucher type are grouped to obtain a second group of pages corresponding to the same voucher type; The first group is updated according to the second group to obtain the target group corresponding to each page in the document to be classified.

[0009] Optionally, after determining the final classification result corresponding to each page in the document to be classified based on the initial classification result and the page classification result, the method further includes: The confidence level of the final classification result is determined based on the classification method corresponding to each page in the document to be classified; If the confidence level is less than a preset threshold, the corresponding page needs to be pushed to manual review.

[0010] Optionally, before performing credential classification on each page of the document to be classified to obtain the initial classification result, the method further includes: Obtain the target document in the first format and determine the first path corresponding to the target document; Convert the first path into a second path list, and use the documents corresponding to the second path list as documents to be classified. Accordingly, after determining the number of vouchers corresponding to the document to be classified based on the target group, and displaying the voucher classification results based on the number of vouchers, the method further includes: The credential classification results of the second path list are converted into the classification results of the first path, and the classification results of the first path are displayed.

[0011] In addition, to achieve the above objectives, this application also provides a voucher sorting device, the voucher sorting device comprising: The voucher classification module is used to classify vouchers in each page of the document to be classified, and obtain the initial classification results. The page determination module is used to determine invalid pages in the document to be classified based on the initial classification results. The result determination module is used to determine the target page preceding the invalid page in the document to be classified, and to determine the page classification result corresponding to the invalid page based on the target classification result of the target page.

[0012] In addition, to achieve the above objectives, this application also proposes a voucher classification device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voucher classification method as described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the credential classification method described above.

[0014] This application categorizes each page of a document to be classified into voucher categories to obtain an initial classification result. Then, based on the initial classification result, it identifies invalid pages within the document. Next, it identifies the target page preceding the invalid page in the document and determines the page classification result corresponding to the invalid page based on the target page's target classification result. This application identifies invalid pages (pages that cannot be classified) based on the initial classification result of each page in the document. It identifies the target page preceding the invalid page in the document and determines the page classification result of the invalid page based on the target page's target classification result. Compared to existing methods that rely on manual annotation of invalid page voucher types, this application's method can automatically and effectively classify all pages in the document to be classified. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the first embodiment of the document classification method of this application; Figure 2 This is a flowchart illustrating the second embodiment of the document classification method of this application; Figure 3 This is a flowchart illustrating the third embodiment of the document classification method of this application; Figure 4 This is a schematic diagram of the overall process of one embodiment of the document classification method of this application; Figure 5 This is a structural block diagram of the first embodiment of the document sorting device of this application; Figure 6 This is a schematic diagram of the structure of the credential classification device in the hardware operating environment involved in the embodiments of this application.

[0018] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] The main solution of this application embodiment is: to classify each page in the document to be classified by credentials to obtain an initial classification result; to determine invalid pages in the document to be classified based on the initial classification result; to determine the target page in the document to be classified before the invalid page, and to determine the page classification result corresponding to the invalid page based on the target classification result of the target page.

[0022] Voucher classification is a preliminary step in intelligent verification systems. Its purpose is to automatically categorize each page of uploaded documents into its corresponding voucher type, such as contracts, invoices, and bank statements. Typically, voucher classification is done using models. However, for pages in the document that the model cannot recognize, manual labeling of the voucher type is required, which is inefficient and cannot meet the needs of batch processing.

[0023] This application categorizes each page of a document to be classified into voucher categories to obtain an initial classification result. Then, based on the initial classification result, it identifies invalid pages within the document. Next, it identifies the target page preceding the invalid page in the document and determines the page classification result corresponding to the invalid page based on the target page's target classification result. This application identifies invalid pages (pages that cannot be classified) based on the initial classification result of each page in the document. It identifies the target page preceding the invalid page in the document and determines the page classification result of the invalid page based on the target page's target classification result. Compared to existing methods that rely on manual annotation of invalid page voucher types, this application's method can automatically and effectively classify all pages in the document to be classified.

[0024] It should be noted that the executing entity of this application can be a computing service device with data processing, network communication, and program execution functions, such as a computer, or an electronic device or voucher classification device capable of performing the above functions. The following description uses a voucher classification device as an example to illustrate this embodiment and the subsequent embodiments.

[0025] Based on this, embodiments of this application provide a method for classifying credentials, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the document classification method of this application.

[0026] In this embodiment, the voucher classification method includes the following steps: Step S10: Classify each page in the document to be classified using vouchers to obtain the initial classification result.

[0027] Understandably, the document to be classified can be a document that requires credential classification. Each page in this document needs to be classified as a credential. In one feasible embodiment, credential classification can be performed using a model, such as BERT or Qwen. This yields the initial classification results for each page in the document to be classified, for example, contracts, invoices, bank statements, etc.

[0028] Furthermore, in this embodiment, before step S10, the method further includes: obtaining a target document in a first format and determining a first path corresponding to the target document; converting the first path into a second path list and using the documents corresponding to the second path list as documents to be classified.

[0029] It should be understood that the target document can be in a primary format, which is PDF, and the target document is also a document that needs to be classified as evidence. The primary path corresponding to the target document, i.e., the PDF path, should be determined.

[0030] Understandably, the first path can be converted into a second path list, which can be a Markdown path list. In one feasible embodiment, the database can be queried using the instruction `etStructuredTextMinIOPaths(documentId)`. The input is the `documentId` of the target document, i.e., the unique identifier of the document. The action is to query the mapping table between PDF paths and Markdown path lists in the database, and the output is a Markdown path list. Passing the Markdown path list into the model, replacing the original PDF images with plain text, that is, using the documents corresponding to the Markdown path list as the documents to be classified, can reduce the computational cost of image processing by the model. Extracting directly from the text can improve the accuracy of the model in extracting structured information.

[0031] In a specific implementation, a mapping relationship can be constructed between the original PDF path and the Markdown path list, which can be represented as origPdfPath→[markdownPage1Path, markdownPage2Path, ...].

[0032] Step S20: Determine invalid pages in the document to be classified based on the initial classification results.

[0033] Understandably, invalid pages in the document to be classified can be determined based on the initial classification results. Invalid pages can be blank pages, transition pages containing only headers and footers, etc. If the model returns an initial classification result of UNKNOW for a certain page, then that page is considered invalid.

[0034] Step S30: Determine the target page preceding the invalid page in the document to be classified, and determine the page classification result corresponding to the invalid page based on the target classification result of the target page.

[0035] It should be understood that if there are invalid pages in the document to be classified, the target page before the invalid page should be selected from the document to be classified. If the invalid page is on the tenth page, pages from the first to the ninth page can be used as the target page.

[0036] In practice, the target classification result corresponding to each target page can be selected from the initial classification result, and the page classification result corresponding to the invalid page can be determined based on the target classification result, thereby obtaining the credential type of all pages in the document to be classified.

[0037] This embodiment classifies each page of a document to be classified into voucher categories to obtain an initial classification result. Then, based on the initial classification result, invalid pages in the document to be classified are identified. A target page preceding the invalid page is determined within the document, and the page classification result corresponding to the invalid page is determined based on the target page's target classification result. This embodiment identifies invalid pages (pages that cannot be classified into vouchers) based on the initial classification result of each page in the document to be classified. It identifies the target page preceding the invalid page in the document and determines the page classification result of the invalid page based on the target page's target classification result. Compared to existing methods that rely on manual annotation of invalid page types, this embodiment can automatically and effectively classify all pages in the document to be classified into voucher categories.

[0038] refer to Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the document classification method of this application.

[0039] Based on the first embodiment described above, in this embodiment, step S30 includes: Step S301: Determine the target page preceding the invalid page in the document to be classified, and select the page preceding the invalid page from the target page.

[0040] Understandably, you can select target pages between invalid pages from the documents to be categorized, and then select the page preceding the invalid page from the target pages. For example, if the invalid page is page 10, the preceding page is page 9.

[0041] Step S302: Determine whether the previous page is a valid page, and obtain the determination result.

[0042] It should be understood that it is possible to determine whether the previous page is a valid page. In one feasible embodiment, it is possible to determine whether the previous page is a valid page based on the initial classification result. For example, if the classification result of the previous page is UNKNOW, then the previous page is an invalid page; if the classification result of the previous page is a contract, invoice or other voucher type, then the previous page is a valid page.

[0043] Step S303: Determine the page classification result corresponding to the invalid page based on the judgment result and the target classification result of the target page.

[0044] Furthermore, in order to automatically obtain the page classification result corresponding to the invalid page, in this embodiment, step S303 includes: if the judgment result is that the previous page is a valid page, selecting the previous classification result of the previous page from the target classification results of the target page; using the previous classification result as the page classification result corresponding to the invalid page; if the judgment result is that the previous page is not a valid page, traversing the target page according to a preset order based on the invalid page until a valid page is obtained; selecting the valid classification result corresponding to the valid page from the target classification results, and using the valid classification result as the page classification result corresponding to the invalid page.

[0045] Understandably, if the previous page is a valid page, since the target page includes valid pages, the classification result of the previous page can be selected from the target classification results of the target page as the previous classification result, and the previous classification result can be used as the page classification result corresponding to the invalid page.

[0046] It should be understood that when the previous page is invalid, the invalid page can be used as the starting page, and the target pages can be traversed in a preset order. The preset order can be from back to front, i.e., backtracking, until a valid page is found. For example, if the invalid page is page 10, pages 9, 8, etc., can be traversed sequentially until a valid page is found. Then, the valid category result corresponding to the valid page is selected from the target category results, and this valid category result is used as the page category result corresponding to the invalid page. For example, P1 is a valid page, and the category result is "Contract"; P2's category result is "UNKNOW", i.e., an invalid page, and it can inherit the category result of P1, so P2's category result is also "Contract"; P3's category result is "UNKNOW", and it can backtrack to the valid page P1, so P3's category result is also "Contract"; P4's category result is "UNKNOW", and it can backtrack to the valid page P1, so P4's category result is also "Contract"; P5 is a valid page, and the category result is "Attachment"; P6's category result is "UNKNOW", and it can inherit the category result of P5, so P6's category result is also "Attachment". The backtracking algorithm has a time complexity of O(n²), which is acceptable in real-world scenarios (a single voucher typically does not exceed 100 pages).

[0047] In a concrete implementation, the code for the above inheritance and backtracking can be represented as follows: FOR i = 0 TO sortedResults.size() - 1: current = sortedResults[i] IF current.voucherNo is empty OR current.voucherNo == "UNKNOW": If the current file is a PDF file AND i>0: prev = sortedResults[i - 1] If prev.voucherNo is valid: / / Inherit directly from the previous page current.voucherNo = prev.voucherNo current.voucherName = prev.voucherName ELSE: / / Backtrack forward to find the first valid page FOR j = i - 1 DOWNTO 0: IF sortedResults[j].voucherNo is valid: current.voucherNo= sortedResults[j].voucherNo current.voucherName=sortedResults[j].voucherName BREAK In the code above, the inheritance constraints include: 1. Inheritance is only performed on PDF files (non-PDF files are categorized independently for each page); 2. Only backward inheritance is performed, not backward inheritance (to ensure the deterministic processing order); 3. If all pages of the entire PDF are not recognized, the UNKNOW state is retained.

[0048] This embodiment identifies the target page preceding the invalid page in the document to be classified, selects the page preceding the invalid page from the target page, determines whether the preceding page is a valid page, obtains the determination result, and then determines the page classification result corresponding to the invalid page based on the determination result and the target page's target classification result. This embodiment, by determining whether the preceding page is a valid page and then determining the page classification result corresponding to the invalid page based on the determination result and the target page's target classification result, can automatically inherit or backtrack the classification results of valid pages preceding the invalid page, thereby obtaining the credential classification results for all pages in the document to be classified.

[0049] refer to Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the document classification method of this application.

[0050] Based on the above embodiments, in this embodiment, after step S30, the method further includes: Step S40: Determine the final classification result for each page in the document to be classified based on the initial classification result and the page classification result.

[0051] Understandably, the initial classification results may include the classification results of all valid pages in the document to be classified, and the page classification results may include the classification results of all invalid pages, thereby enabling the determination of the final classification results of all pages in the document to be classified based on the initial classification results and the page classification results.

[0052] Furthermore, in this embodiment, after step S01, the method further includes: determining the confidence level of the final classification result based on the classification method corresponding to each page in the document to be classified; if the confidence level is less than a preset threshold, the corresponding page needs to be pushed to manual review.

[0053] It should be understood that the classification method for each page in the document to be classified can be determined. Classification methods may include direct classification, inheritance of the previous page's classification result, or backtracking of the classification result of a previous page. Then, the confidence level of the final classification result for each page is determined based on the classification method. For directly classified pages: confidence level = original model confidence level; for pages inheriting from the previous page: confidence level = previous page confidence level × 0.9 (attenuation coefficient); for backtracking inherited pages: confidence level = inherited page confidence level × (0.9^backtracking distance). For example, P1 is a directly classified page, confidence level = original model confidence level = 0.9; P2 is a page inheriting from the previous page, confidence level = 0.9 × 0.9 = 0.81; P3 is a page backtracking to P1, confidence level = 0.9 × (0.9^2) = 0.729.

[0054] In practice, if the confidence level is less than a preset threshold (which is a pre-set threshold), the corresponding page can be pushed to manual review, achieving a balance between accuracy and efficiency.

[0055] Step S50: Determine the target group corresponding to each page in the document to be classified based on the final classification result.

[0056] Further, in this embodiment, step S02 includes: grouping each page in the document to be classified according to the voucher type based on the final classification result to obtain a first group corresponding to each page; determining the file name of the page corresponding to the same voucher type based on the first group; grouping the pages corresponding to the same voucher type based on the file name to obtain a second group of the pages corresponding to the same voucher type; updating the first group based on the second group to obtain the target group corresponding to each page in the document to be classified.

[0057] Understandably, pages in the document to be classified can be grouped according to the voucher type (voucherNo). This means grouping pages based on the final classification result, grouping pages of the same voucher type together to obtain the first group for each page. Then, based on the first group, the filenames of the pages corresponding to the same voucher type are determined, resulting in the second group. Within the same voucher type group, pages are grouped according to their filename (origFileName), and pages with the same filename are considered the same voucher. Updating the first group based on the second group yields the target group for each page in the document to be classified. For example, when the voucher type is a contract, Contract A.pdf (the first copy) → groupingTag=1, including P1, P2, and P3 in Contract A.pdf; Contract B.pdf (the second copy) → groupingTag=2, including P1 and P2 in A.pdf; when the voucher type is an invoice, Invoice X.pdf (the first copy) → groupingTag=1, including P1 in Invoice X.pdf, where groupingTag is the group.

[0058] Step S60: Determine the number of vouchers corresponding to the document to be classified according to the target group, and display the voucher classification result based on the number of vouchers.

[0059] It should be understood that the number of vouchers corresponding to the documents to be classified, groupingTagCounter, can be determined based on the target group. The number of vouchers is: groupingTagCounter = 1 FOR EACH voucherNo IN First-level group: FOR EACH origFileName IN Second-level group: The groupingTag of all pages under this file name is set to groupingTagCounter. groupingTagCounter++ In the specific implementation, the voucher classification results can be displayed based on the number of vouchers. This embodiment employs a two-level strategy: first grouping by voucher type, then by filename. This automatically identifies multiple copies of the same voucher type and enables automatic counting of voucher copies. The grouping tags are persisted to the database, supporting the front-end display of voucher classification results by copy and subsequent review processes by copy.

[0060] Furthermore, in this embodiment, after step S60, the method further includes: converting the credential classification result of the second path list into the classification result of the first path, and displaying the classification result of the first path.

[0061] Understandably, the model returns a list of credential classification results from the second path list, i.e., a list of Markdown file paths, such as {docId} / markdown / page_3.md. Through a reverse mapping table, the Markdown paths are converted back to the classification results from the first path, i.e., the original PDF paths and page numbers. The business system uses the original PDF paths for subsequent display processing, remaining unaware of the Markdown optimization process. This embodiment automatically converts PDF paths to Markdown pagination paths upon request and automatically reverse maps them upon return, making it transparent to upper-layer business logic while supporting Markdown transmission optimization.

[0062] In the specific implementation, refer to Figure 4 , Figure 4 This is a schematic diagram of the overall process of one embodiment of the document classification method of this application, as shown below. Figure 4 As shown, when classifying documents into vouchers, the documents to be classified can be entered into the voucher classification entry point. Then, an intelligent agent data assistance class is used to construct the intelligent agent request data by using the PDF path → Markdown path list. The classification results are then output page by page through the AI ​​classification engine (Agent / LLM). The classification result processor performs the following: ① inheritance or backward tracing of unrecognized pages; ② two-level grouping; ③ count of copies; ④ reverse mapping of Markdown paths. Finally, the data is persisted in the database.

[0063] This embodiment determines the final classification result for each page in the document to be classified based on the initial classification result and the page classification result. Then, it determines the target group for each page in the document to be classified based on the final classification result, and then determines the number of vouchers corresponding to the document to be classified based on the target group. Finally, it displays the voucher classification result based on the number of vouchers. This embodiment determines the target group for each page in the document to be classified based on the final classification result, using a two-level strategy of grouping by voucher type and then by filename. This automatically identifies multiple copies under the same voucher type and automatically counts the number of vouchers.

[0064] Reference Figure 5 , Figure 5 This is a structural block diagram of the first embodiment of the document classification device of this application.

[0065] like Figure 5 As shown, the credential classification device proposed in this application includes: The voucher classification module 10 is used to classify vouchers in each page of the document to be classified and obtain the initial classification result. Page determination module 20 is used to determine invalid pages in the document to be classified based on the initial classification result; The result determination module 30 is used to determine the target page before the invalid page in the document to be classified, and to determine the page classification result corresponding to the invalid page based on the target classification result of the target page.

[0066] This embodiment classifies each page of a document to be classified into voucher categories to obtain an initial classification result. Then, based on the initial classification result, invalid pages in the document to be classified are identified. A target page preceding the invalid page is determined within the document, and the page classification result corresponding to the invalid page is determined based on the target page's target classification result. This embodiment identifies invalid pages (pages that cannot be classified into vouchers) based on the initial classification result of each page in the document to be classified. It identifies the target page preceding the invalid page in the document and determines the page classification result of the invalid page based on the target page's target classification result. Compared to existing methods that rely on manual annotation of invalid page types, this embodiment can automatically and effectively classify all pages in the document to be classified into voucher categories.

[0067] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this application. In practical applications, those skilled in the art can select some or all of it to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.

[0068] In addition, for technical details not described in detail in this embodiment, please refer to the credential classification method provided in any embodiment of this application, which will not be repeated here.

[0069] Based on the first embodiment of the document classification device described in this application, a second embodiment of the document classification device of this application is proposed.

[0070] In this embodiment, the result determination module 30 is further configured to determine the target page preceding the invalid page in the document to be classified, and select the page preceding the invalid page from the target page; determine whether the preceding page is a valid page to obtain a determination result; and determine the page classification result corresponding to the invalid page based on the determination result and the target classification result of the target page.

[0071] Furthermore, the result determination module 30 is also configured to: when the determination result indicates that the previous page is a valid page, select the previous category result of the previous page from the target category results of the target page; use the previous category result as the page category result corresponding to the invalid page; when the determination result indicates that the previous page is not a valid page, traverse the target page according to a preset order based on the invalid page until a valid page is obtained; select the valid category result corresponding to the valid page from the target category results, and use the valid category result as the page category result corresponding to the invalid page.

[0072] Furthermore, the result determination module 30 is also used to determine the final classification result corresponding to each page in the document to be classified based on the initial classification result and the page classification result; determine the target group corresponding to each page in the document to be classified based on the final classification result; determine the number of vouchers corresponding to the document to be classified based on the target group; and display the voucher classification result based on the number of vouchers.

[0073] Furthermore, the result determination module 30 is also used to group each page in the document to be classified according to the voucher type based on the final classification result, to obtain a first group corresponding to each page; determine the file name of the page corresponding to the same voucher type based on the first group; group the pages corresponding to the same voucher type based on the file name, to obtain a second group of the pages corresponding to the same voucher type; update the first group based on the second group, to obtain the target group corresponding to each page in the document to be classified.

[0074] Furthermore, the result determination module 30 is also used to determine the confidence level of the final classification result based on the classification method corresponding to each page in the document to be classified; if the confidence level is less than a preset threshold, the corresponding page needs to be pushed to manual review.

[0075] Furthermore, the voucher classification module 10 is also used to obtain a target document in a first format and determine a first path corresponding to the target document; convert the first path into a second path list and use the document corresponding to the second path list as the document to be classified; the result determination module 30 is also used to convert the voucher classification result of the second path list into the classification result of the first path and display the classification result of the first path.

[0076] Other embodiments or specific implementations of the document classification device of this application can be found in the above-described method embodiments, and will not be repeated here.

[0077] This application provides a credential classification device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the credential classification method in Embodiment 1 above.

[0078] The following is for reference. Figure 6 The diagram illustrates a structural schematic of a credential classification device suitable for implementing embodiments of this application. The credential classification device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The credential sorting device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0079] like Figure 6 As shown, the credential sorting device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the credential sorting device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the credential sorting device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows credential sorting devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0080] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0081] The voucher classification device provided in this application, employing the voucher classification method in the above embodiments, can solve the technical problem of how to automatically and effectively classify vouchers across all pages in a document. Compared with the prior art, the beneficial effects of the voucher classification device provided in this application are the same as those of the voucher classification method provided in the above embodiments, and other technical features of this voucher classification device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0082] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0084] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the credential classification method in the above embodiments.

[0085] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0086] The aforementioned computer-readable storage medium may be included in the credential classification device; or it may exist independently and not be assembled into the credential classification device.

[0087] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the credential classification device, cause the credential classification device to: classify each page in the document to be classified to obtain an initial classification result; determine invalid pages in the document to be classified based on the initial classification result; determine target pages in the document to be classified preceding the invalid pages, and determine the page classification result corresponding to the invalid pages based on the target classification result of the target pages.

[0088] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Python, Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0090] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0091] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described credential classification method, thereby solving the technical problem of how to automatically and effectively classify credentials across all pages in a document. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the credential classification method provided in the above embodiments, and will not be repeated here.

[0092] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.

Claims

1. A method for classifying vouchers, characterized in that, The voucher classification method includes the following steps: Classify each page in the document to be classified using vouchers to obtain the initial classification results; Based on the initial classification results, invalid pages in the document to be classified are determined. In the document to be classified, determine the target page preceding the invalid page, and determine the page classification result corresponding to the invalid page based on the target page's target classification result.

2. The voucher classification method as described in claim 1, characterized in that, The step of determining the target page prior to identifying the invalid page in the document to be classified, and determining the page classification result corresponding to the invalid page based on the target page's target classification result, includes: In the document to be classified, determine the target page preceding the invalid page, and select the page preceding the invalid page from the target page; Determine whether the previous page is a valid page, and obtain the determination result; The page classification result corresponding to the invalid page is determined based on the judgment result and the target classification result of the target page.

3. The voucher classification method as described in claim 2, characterized in that, The step of determining the page classification result corresponding to the invalid page based on the judgment result and the target page's target classification result includes: If the determination result indicates that the previous page is a valid page, the previous category result of the previous page is selected from the target category results of the target page; The previous classification result is used as the page classification result corresponding to the invalid page; If the determination result is that the previous page is not a valid page, the target page is traversed according to a preset order based on the invalid page until a valid page is obtained; Select the valid classification result corresponding to the valid page from the target classification results, and use the valid classification result as the page classification result corresponding to the invalid page.

4. The voucher classification method as described in claim 1, characterized in that, After determining the target page prior to identifying the invalid page in the document to be classified, and determining the page classification result corresponding to the invalid page based on the target page's target classification result, the method further includes: Based on the initial classification results and the page classification results, determine the final classification results for each page in the document to be classified; Based on the final classification results, determine the target group corresponding to each page in the document to be classified; The number of vouchers corresponding to the document to be classified is determined based on the target group, and the voucher classification results are displayed based on the number of vouchers.

5. The voucher classification method as described in claim 4, characterized in that, Determining the target group corresponding to each page in the document to be classified based on the final classification result includes: Based on the final classification result, each page in the document to be classified is grouped according to the voucher type to obtain the first group corresponding to each page; The filename of the page corresponding to the same voucher type is determined based on the first group; Based on the file name, the pages corresponding to the same voucher type are grouped to obtain a second group of pages corresponding to the same voucher type; The first group is updated according to the second group to obtain the target group corresponding to each page in the document to be classified.

6. The voucher classification method as described in claim 4, characterized in that, After determining the final classification result for each page in the document to be classified based on the initial classification result and the page classification result, the method further includes: The confidence level of the final classification result is determined based on the classification method corresponding to each page in the document to be classified; If the confidence level is less than a preset threshold, the corresponding page needs to be pushed to manual review.

7. The voucher classification method as described in claim 4, characterized in that, Before performing credential classification on each page of the document to be classified and obtaining the initial classification result, the process also includes: Obtain the target document in the first format and determine the first path corresponding to the target document; Convert the first path into a second path list, and use the documents corresponding to the second path list as documents to be classified. Accordingly, after determining the number of vouchers corresponding to the document to be classified based on the target group, and displaying the voucher classification results based on the number of vouchers, the method further includes: The credential classification results of the second path list are converted into the classification results of the first path, and the classification results of the first path are displayed.

8. A voucher sorting device, characterized in that, The voucher sorting device includes: The voucher classification module is used to classify vouchers in each page of the document to be classified, and obtain the initial classification results. The page determination module is used to determine invalid pages in the document to be classified based on the initial classification results. The result determination module is used to determine the target page preceding the invalid page in the document to be classified, and to determine the page classification result corresponding to the invalid page based on the target classification result of the target page.

9. A voucher sorting device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the credential classification method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the credential classification method as described in any one of claims 1 to 7.