Accounting Voucher Generation Method, Device and Medium Based on Document Block Traversal
By chunking the accounting voucher rules document and using the key information extraction instruction templates of the big model, the problems of low efficiency and insufficient accuracy of traditional accounting voucher generation are solved, and efficient and accurate automatic generation of accounting vouchers is achieved.
Patent Information
- Application Number
- CN202510630957.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Traditional accounting voucher generation depends on the voucher rule mapping table, which is inefficient and difficult to guarantee the accuracy. The workload is high when the voucher rule documents are changed. Directly calling a large model to generate accounting vouchers has problems such as slow response speed, high resource consumption and difficult information extraction.
The accounting voucher rules document is processed in blocks, and the key information is extracted and generated using a large model. The information extraction and generation process is optimized through pre-generated instruction templates and chunking standards.
It improves the efficiency and accuracy of accounting voucher generation, reduces manual intervention, reduces error rate and resource consumption, and shortens response time.
Smart Images

Figure CN120147043B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and specifically to an accounting voucher generation method, device, and medium based on document block traversal. Background Art
[0002] In traditional solutions, the generation of accounting vouchers relies on a voucher rule mapping table, which is formulated by professional finance and tax personnel according to voucher rule documents. This method has low efficiency, difficult accuracy guarantee, and when the voucher rule document changes, it is necessary to re-formulate the voucher rule mapping table, resulting in a large workload.
[0003] With the development of artificial intelligence, the intelligentization of finance and tax has been promoted. As an important technology in the field of artificial intelligence, large language models (abbreviated as large models) have promoted the intelligent generation of accounting vouchers with their outstanding capabilities in understanding, generation, reasoning, translation, code, etc.
[0004] However, due to the large size of voucher rule documents, usually reaching tens of thousands of words or even more than one hundred thousand words, directly calling a large model to generate accounting vouchers is prone to problems such as slow response speed, large resource consumption, and difficult information extraction. Using the Retrieval-augmented Generation (RAG) method will have problems with the accuracy of document recall, thus affecting the accuracy of accounting voucher generation. Summary of the Invention
[0005] To solve the above problems, this application proposes an accounting voucher generation method based on document block traversal, including:
[0006] Performing block processing on the accounting voucher rule document to obtain multiple document blocks;
[0007] Based on a pre-generated key information extraction instruction template, by calling a large model, corresponding key information is extracted according to the document blocks;
[0008] Based on a pre-generated accounting voucher generation instruction template, by calling a large model, an accounting voucher is generated according to the obtained reimbursement form information and the key information.
[0009] In one example, performing block processing on the accounting voucher rule document to obtain multiple document blocks specifically includes:
[0010] For the smallest block unit obtained from the accounting voucher rule document, determining whether it meets the preset word count requirement;
[0011] If it meets the requirement, the smallest block unit is retained as a document block;
[0012] If not, determine the current chunking round, and based on the chunking round, determine the corresponding current chunking criteria; the chunking criteria include at least one of keywords, natural delimiters, and regular expressions;
[0013] Based on the chunking criteria, perform chunking on the non - compliant minimum chunking unit to obtain multiple new minimum chunking units.
[0014] In one example, the method further includes:
[0015] Chunk a number of the earliest - ordered documents as the first specified document chunks;
[0016] In the first specified document chunks, take the document chunks that have preset keywords and are obtained in the first several preset chunking rounds as the second specified document chunks;
[0017] For the second specified document chunks, by invoking a large - model, extract the corresponding summary information of the second specified document chunks;
[0018] For the other document chunks except the second specified document chunks, add the summary information thereto.
[0019] In one example, the method further includes:
[0020] Determine that the minimum chunking unit does not meet the preset word - count requirement and cannot continue to be chunked by the chunking criteria;
[0021] Divide the minimum chunking unit evenly by word - count to obtain several sub - units that meet the preset word - count requirement;
[0022] For each sub - unit, perform two - way supplementation in accordance with the character order in the accounting voucher rule document until the sub - unit is supplemented to meet the preset word - count requirement;
[0023] Delete the content that does not meet the preset minimum text area from the sub - units that meet the word - count requirement to obtain new document chunks.
[0024] In one example, based on a pre - generated key - information extraction instruction template, by invoking a large - model, according to the document chunks, extract the corresponding key information, specifically including:
[0025] Obtain the pre - generated key - information extraction instruction template, and based on the document chunks and the current scenario requirements, update the key - information extraction instruction template to obtain key - information extraction prompt words;
[0026] Invoke the large - model, and through the key - information extraction prompt words, extract the sub - key information corresponding to each document chunk;
[0027] Integrate the sub - key information to obtain the key information corresponding to the accounting voucher rule document.
[0028] In one example, call a large - model, and extract the sub - key information corresponding to each document chunk through the key - information extraction prompt words, specifically including:
[0029] Call the large - model. In the first conversation for sub - key information extraction, extract the first sub - key information corresponding to the current document chunk through the key - information extraction prompt words.
[0030] According to the sequence relationship corresponding to each document chunk, obtain the historical document chunks before the current document chunk corresponding to the accounting voucher rule document, and obtain the second sub - key information corresponding to the historical document chunks.
[0031] Call the large - model. In a new second conversation, based on the accounting voucher rule document, determine the memory weights between the first sub - key information and each second sub - key information.
[0032] Call the large - model. In the first conversation, based on the memory weights, extract the first sub - key information corresponding to the current document chunk again through the key - information extraction prompt words.
[0033] In one example, based on a pre - generated accounting voucher generation instruction template, by calling a large - model, generate accounting vouchers according to the obtained reimbursement form information and the key information, specifically including:
[0034] Obtain the pre - generated accounting voucher generation instruction template, and update the accounting voucher generation instruction template based on the key information and the obtained reimbursement form information to obtain an accounting voucher generation prompt word.
[0035] Call the large - model, and generate accounting vouchers according to the accounting voucher rule document through the accounting voucher generation prompt word.
[0036] In one example, the method further includes:
[0037] Obtain the feedback opinions of the current department and downstream departments on the accounting vouchers.
[0038] Based on the feedback opinions, modify the large - model architecture and the content of the prompt words.
[0039] On the other hand, the present application also proposes an accounting voucher generation device based on document - chunk traversal, including:
[0040] At least one processor; and,
[0041] A memory communicatively connected to the at least one processor; wherein,
[0042] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute, for example, the accounting voucher generation method based on document block traversal described in any of the above examples.
[0043] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as: the accounting voucher generation method based on document block traversal described in any of the above examples.
[0044] The accounting voucher generation method based on document block traversal proposed by the present application can bring the following beneficial effects:
[0045] Through document block division, efficiency breakthroughs are achieved. Overall, the automated process reduces manual intervention, shortens the time-consuming of document preprocessing, and improves the speed of key information extraction. When it is detected that the content of a document block changes, only this document block needs to be processed, improving the response speed.
[0046] Through document block division, the video memory occupancy ratio is reduced, the reuse rate of computing resources is increased, and the error handling cost is reduced. Through the dual-stage instruction guarantee of key information extraction instructions and voucher generation instructions, and the execution one by one based on the large model, the accuracy of voucher generation can be improved, and the error rate of accounting voucher generation can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0048] Figure 1 It is a flowchart of the accounting voucher generation method based on document block traversal in an embodiment of the present application;
[0049] Figure 2 It is a flowchart of document block division in a case of an embodiment of the present application;
[0050] Figure 3 It is a flowchart of key information extraction in a case of an embodiment of the present application;
[0051] Figure 4 It is a partial schematic diagram of an accounting voucher in a case of an embodiment of the present application;
[0052] Figure 5 It is a schematic diagram of the accounting voucher generation device based on document block traversal in an embodiment of the present application. Detailed implementation manners
[0053] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0054] The technical solutions provided by each embodiment of the present application will be described in detail below with reference to the drawings.
[0055] As Figure 1 shown, the embodiment of the present application provides an accounting voucher generation method based on document block traversal, including:
[0056] S101: Perform block processing on the accounting voucher rule document to obtain a plurality of document blocks.
[0057] The accounting voucher rule document is the basis for generating accounting vouchers. Usually, the document is large, containing tens of thousands of words, or even more than one hundred thousand words. Directly invoking the large model will have problems such as slow response speed, large resource consumption, and low information accuracy. By adopting the method of retrieval-augmented generation (RAG), however, the information accuracy of accounting voucher generation will be affected by the recall accuracy of the RAG retrieved documents.
[0058] Based on this, the method of long document block traversal is adopted to accurately obtain the key information of accounting vouchers by utilizing the intelligence of the large model.
[0059] For the smallest block unit obtained from the accounting voucher rule document, determine whether it meets the preset word count requirement. The smallest block unit refers to the smallest document block obtained in the current block processing. If the accounting voucher rule document has not been block-processed yet, that is, the first round of block processing is being performed, then the current smallest block unit is the accounting voucher rule document itself. If the block processing has been performed, then the currently obtained smallest document block is used as the smallest block unit. For example, if only one round of block processing has been experienced, and the accounting voucher rule document is block-processed to obtain document block A and document block B, then both document block A and document block B are the smallest block units. If two rounds of block processing have been experienced, and document block B is further divided into document block C and document block D, and document block A is not further divided, then the current smallest block units are document block A, document block C, and document block D.
[0060] The word count requirement is pre-set and can be set based on user needs, document length, large model capabilities, etc. For example, it can be set to 8192.
[0061] At this time, if the current document meets the pre-set word count requirement, it is considered that this document chunk can be retained, that is, retain this smallest chunk unit as the document chunk.
[0062] If it does not meet the pre-set word count requirement, then further chunking processing is performed on this document chunk. Based on the chunking criteria, the non-compliant smallest chunk unit is chunked to obtain multiple new smallest chunk units.
[0063] As Figure 2 shown, when performing document chunking processing, multiple chunking rounds can be set. In each chunking round, determine the current chunking round, and based on the chunking round, determine the corresponding current chunking criteria; the chunking criteria include at least one of keywords, natural delimiters, and regular expressions. For example, in the first chunking round, the selected keyword is the first-level heading, and in the second chunking round, the selected keyword is the second-level heading, etc. As Figure 2 shown, up to three levels of headings can be set. If, after three chunking rounds, there are still smallest chunk units that do not meet the pre-set word count requirement, then they are directly chunked according to the word count.
[0064] Among them, keywords can be set through headings of different levels, and natural delimiters refer to the delimiters set in the accounting voucher rule document for distinguishing different regions (such as different paragraphs, different lines, etc.). After identifying this delimiter, the region between the two delimiters can be regarded as an overall region and no chunking processing is performed. And regular expressions are set according to keywords, natural delimiters, etc., and are used to write script code according to this regular expression to automatically execute the corresponding chunking actions.
[0065] S102: Based on the pre-generated key information extraction instruction template, by calling a large model, corresponding key information is extracted according to the document chunk.
[0066] After the accounting voucher rule document is chunked, it is necessary to call a large model to obtain key information from the accounting voucher rule document. At this time, it can be achieved by setting an instruction template. It should be noted that the large model called here mainly refers to the Large Language Model (LLM). This large language model can generate natural language text or understand the meaning of language text according to the input text data. In some scenarios, this large language model can also be a multi-modal large language model that can input and output other modal data in addition to the text modality (such as the image modality, the audio modality, etc.).
[0067] The large language model can be deployed locally or in the cloud. At this time, the large model can be called through the local large model interface API or the cloud service API.
[0068] The large language model can be obtained by the user's own training or be an open-source large language model obtained through the network. It can be called by deploying it locally or using the cloud service of the open-source large language model. It can also be based on the open-source large language model deployed locally and fine-tune the model parameters of the large language model according to its own needs to obtain the corresponding large model.
[0069] Specifically, obtain the pre-generated key information extraction instruction template, and update the key information extraction instruction template based on document chunking and current scenario requirements to obtain the key information extraction prompt. Based on different scenario requirements, the required key information is also different. For example, for different business departments, counterparties, accounting departments, etc., the required key information may be different. For example, the key information may include debit / credit direction, account name, auxiliary accounting category, etc. At this time, a corresponding key information extraction instruction template can be set separately for each key information, or a key information extraction instruction template can be set as a whole for all key information.
[0070] In the key information extraction instruction template, different instructions are constructed by integrating instruction templates with variable parameters and passing different parameters. For example, the key information extraction instruction template can be: "You are a software engineer and need to generate accounting vouchers based on the {{reimbursement document}} information, including voucher headers, voucher entries, auxiliary accounting details, etc. The current task is to find the content about {{key information}} from the {{accounting voucher rule document}}. If found, output the relevant information. If not found, please feedback not found."
[0071] Among them, the content in "{{*}}" is variable content that needs to be filled in. {{reimbursement document}} refers to the original reimbursement document information for generating accounting vouchers, {{accounting voucher rule document}} refers to each document chunk obtained by chunking the accounting voucher rule document, and {{key information}} is the key information for generating accounting vouchers, including debit / credit direction, account name, auxiliary accounting category, etc.
[0072] In this way, after the variable content is updated, the key information extraction prompt can be obtained. At this time, call the large model and extract the sub-key information corresponding to each document chunk through the key information extraction prompt.
[0073] Such as Figure 3As shown, after inputting the key information, generate a key information extraction prompt, and call the large model to extract information from the rule document of the document chunks. If there is relevant information, output the relevant information; if not found, output "not found". Then, in the way of adding 1 to the chunk number of the document chunks, find the next document chunk, continue to call the large model to extract key information, and perform the same operation until the rule document of the last chunk. Integrate all the sub-key information found by the large model into the key information and output it.
[0074] S103: Based on the pre-generated accounting voucher generation instruction template, by calling the large model, generate accounting vouchers according to the obtained reimbursement form information and the key information.
[0075] After obtaining the key information of the accounting voucher, it is necessary to generate accounting vouchers according to the key information of the accounting voucher and the original reimbursement form information. At this time, obtain the pre-generated accounting voucher generation instruction template, and update the accounting voucher generation instruction template based on the key information and the obtained reimbursement form information to obtain the accounting voucher generation prompt.
[0076] The accounting voucher generation instruction template is divided into three parts: accounting subject reasoning, accounting voucher preparation, and amount calculation. For example, in an example, it can be: "You are a software engineer. According to the {{reimbursement form}} information and the {{key information}} in the voucher generation rule document, correctly process the software function of accounting matters. Your task is to execute in sequence: 1. You need to intelligently reason about the accounting subjects, display all debit and credit subjects, and the display content includes the debit and credit directions, subject names, whether there is auxiliary accounting, and the auxiliary accounting category. The auxiliary accounting category for those without auxiliary accounting is empty, and it is presented in tabular form; 2. You need to independently complete the preparation of the accounting voucher. The accounting voucher includes three parts: voucher header, voucher entry, and auxiliary accounting details. You need to complete the voucher header, voucher entry, and auxiliary accounting details in sequence, and generate them strictly according to the format, and display them in tabular form; 3. You need to perform amount calculation, merge and calculate the same items in debit and credit, and verify that the total of debit and credit in the voucher entry is equal. Do not randomly fabricate subject data."
[0077] Among them, the content in "{{*}}" is variable content that needs to be filled in to obtain the accounting voucher generation prompt. At this time, call the large model, and based on the accounting voucher generation prompt, generate accounting vouchers according to the accounting voucher rule document.
[0078] As Figure 4 shown, the finally obtained accounting vouchers can include forms such as tables and texts, and can also include forms such as bar charts and pie charts based on requirements to reflect the corresponding values therein. Among them, for privacy protection, Figure 4 part of the content in is blurred.
[0079] In addition, it is also possible to obtain the feedback opinions of the current department and downstream departments on accounting vouchers, and based on these feedback opinions, modify the large model architecture and prompt content. Here, the current department is the department that generates accounting vouchers, and the downstream department is the department that uses accounting vouchers. Generally speaking, the feedback opinions of the downstream department have a higher weight. When making modifications, the control variable method can be used for adjustment.
[0080] Efficiency breakthroughs are achieved through document chunking. Overall, the automated process reduces manual intervention, shortens the time-consuming of document preprocessing, and improves the speed of key information extraction. When changes in the chunk content are detected, only the relevant document chunk needs to be processed, improving the response speed.
[0081] Document chunking reduces the proportion of video memory occupied and improves the reuse rate of computing resources, reducing the cost of error handling. Through the dual-stage instruction guarantee of key information extraction instructions and voucher generation instructions, and the execution one by one based on the large model, the accuracy of voucher generation can be improved, and the error rate of accounting voucher generation can be reduced.
[0082] In one embodiment, dividing the accounting voucher rule document into multiple document chunks can solve the problem of slow processing speed of the large model. However, since the accounting voucher rule document is usually a coherent whole and there is actually a connection between the document chunks, when extracting key information for each document chunk separately, it is easy to cause omission or misjudgment of key information. And if key information is extracted for the current document chunk by considering context information, it is easy to cause the above-mentioned problem of too slow processing speed.
[0083] Based on this, the first several document chunks with the earliest order are used as the first designated document chunks. The number of these several chunks can be set according to the actual situation. Generally speaking, the earliest document chunks represent the rule summary and other content of the accounting voucher rule document, which can extract relatively important information from the entire accounting voucher rule document.
[0084] However, due to the differences in actual accounting voucher rule documents, not all the first designated document chunks are summary content. Based on this, screening is carried out among the first designated document chunks, and the document chunks that contain preset keywords and are obtained in the first several preset chunk rounds are used as the second designated document chunks. The preset keywords can be set according to actual needs. Since the content of the summary document chunks is usually not too long, generally, the document chunks obtained in the first few chunk rounds are used as the second designated document chunks. In most cases, it can be set as the document chunk obtained in the first chunk round. If the text quantity of the accounting voucher rule document is too long, it can also be set as the document chunks obtained in the first two chunk rounds.
[0085] At this time, it can be considered that the second specified document block is the summary content of the document. For the second specified document block, by calling the large model, the summary information corresponding to the second specified document block can be extracted. The summary information can only contain the corresponding key information, or can be simply expanded based on requirements.
[0086] At this time, for other document blocks except the second specified document block, add summary information to them. When subsequently analyzing through the content of the document block by calling the large model, it is also possible to quickly locate the document block without relying on context information and understand the overall situation of the accounting voucher rule document.
[0087] In one embodiment, as mentioned above, when performing document chunking, first perform chunking processing through first-level headings, second-level headings, etc. As the number of chunking rounds increases, it may occur that there is no corresponding heading in the latest smallest chunk unit obtained, and it is no longer possible to continue chunking in this way, yet its word count still exceeds the preset word count. At this time, it is determined that the smallest chunk unit does not meet the preset word count requirement and cannot continue to be chunked according to the chunking standard.
[0088] Divide the smallest chunk unit evenly according to the word count to obtain several sub-units that meet the preset word count requirement. Generally speaking, to ensure the enrichment of the content in each document block, it is necessary to ensure that it has a certain word count requirement while meeting the preset word count requirement. Therefore, the number of equal divisions can be minimized as much as possible to make the word count in the obtained sub-units as large as possible.
[0089] However, there may be a problem at this time, that is, since the smallest chunk unit is not chunked by keywords, headings, etc., there is a relatively deep connection between the contents of each document block. Directly using the brute-force chunking method is likely to destroy this connection.
[0090] Based on this, for each sub-unit, perform two-way supplementation according to the character order in the accounting voucher rule document until the sub-unit is supplemented to meet the preset word count requirement. Among them, for the first sub-unit in the smallest chunk unit, its supplementation direction only needs to be supplemented backward, and for the last sub-unit, its supplementation direction only needs to be supplemented forward. For other sub-units, when performing two-way supplementation, it can be supplemented simultaneously in both the front and the back. For each character supplemented forward, a character is supplemented backward at the same time, so as to ensure that in the finally supplemented sub-unit, the original sub-unit is located in the middle position, thereby increasing the connection between this sub-unit and other adjacent sub-units.
[0091] Of course, there may be cases where only half a sentence or half a paragraph is included. In this case, in the sub-units that meet the word count requirements, the content that does not meet the preset minimum text area is deleted to obtain new document chunks. The minimum text area can be set to a sentence, a paragraph, etc. (usually set to a paragraph). At this time, the earliest and the last incomplete content can be deleted to make the sub-unit complete, and this sub-unit is used as a new document chunk.
[0092] In one embodiment, when extracting the sub-key information corresponding to each document chunk through a large model, although the word count of each document chunk is restricted to a certain extent, due to the excessive word count of the accounting voucher rule document itself, it is difficult for the large model to consider the actually useful context information specifically when extracting sub-key information.
[0093] Based on this, when extracting sub-key information, a large model is called. In the first conversation for sub-key information extraction, the first sub-key information corresponding to the current document chunk is extracted through key information extraction prompt words. The first conversation includes the extraction process of each document chunk in the current accounting voucher rule document. At this time, the first sub-key information is obtained according to the default memory weight.
[0094] Of course, if there is no previous document chunk before the current document chunk, or the number of previous document chunks is very small, less than the preset number (for example, less than 3), it can be considered that the current context information is less and the setting of the memory weight has less influence. The already obtained first sub-key information can be directly used as the sub-key information of the current document chunk without performing the subsequent steps.
[0095] If the number of previous document chunks is large, then according to the corresponding sequence relationship of each document chunk, the historical document chunks before the current document chunk corresponding to the accounting voucher rule document are obtained, and the second sub-key information corresponding to the historical document chunks is obtained. The second sub-key information is for the document chunks before the accounting voucher rule document in the first conversation.
[0096] At this time, a large model is called. In the new second conversation, based on the accounting voucher rule document, the memory weights between the first sub-key information and each second sub-key information are determined. Using the new second conversation can be not affected by the context information in the first conversation. The memory weight refers to the consideration weight of each context information when the large model considers context information. And since the key information usually has less content, the memory weights between each second sub-key information and the first sub-key information can be determined quickly. The second conversation can be a single one for recording the generation of all memory weights, or it can be multiple ones, and each time the memory weight is generated, a blank second conversation is used.
[0097] Re - call the large - model. In the first conversation, based on the memory weights, extract the first sub - key information corresponding to the current document chunk through key - information extraction prompts. The obtained memory weights can effectively help the large - model consider various context information with appropriate weights and reasonably extract the corresponding key information in the accounting voucher rule document.
[0098] As Figure 5 shown, the embodiment of the present application also proposes an accounting voucher generation device based on document - chunk traversal, including:
[0099] At least one processor; and,
[0100] A memory communicatively connected to the at least one processor; wherein,
[0101] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the accounting voucher generation method based on document - chunk traversal as described in any of the above embodiments.
[0102] The embodiment of the present application also proposes a non - volatile computer storage medium storing computer - executable instructions, and the computer - executable instructions are set to: the accounting voucher generation method based on document - chunk traversal as described in any of the above embodiments.
[0103] The various embodiments in the present application are described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0104] The device and medium provided by the embodiment of the present application correspond one - to - one with the method. Therefore, the device and medium also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium are not elaborated here.
[0105] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An accounting voucher generation method based on document block traversal, characterized in that Including: Chunk the accounting voucher rule document to obtain multiple document chunks; Based on a pre-generated key information extraction instruction template, by invoking a large model, extract the corresponding key information according to the document chunks; Based on a pre-generated accounting voucher generation instruction template, by invoking a large model, generate accounting vouchers according to the obtained reimbursement form information and the key information.
2. The accounting voucher generation method based on document block traversal according to claim 1, wherein Chunk the accounting voucher rule document to obtain multiple document chunks, specifically including: For the smallest chunk unit obtained from the accounting voucher rule document, determine whether it meets the preset word count requirement; If it meets the requirement, retain the smallest chunk unit as a document chunk; If it does not meet the requirement, determine the current chunking round, and based on the chunking round, determine the current corresponding chunking criterion; the chunking criterion includes at least one of keywords, natural delimiters, and regular expressions; Based on the chunking criterion, perform chunking processing on the non-conforming smallest chunk unit to obtain multiple new smallest chunk units.
3. The accounting voucher generation method based on document block traversal according to claim 2, wherein The method further includes: Take several document chunks with the earliest order as the first specified document chunks; In the first specified document chunks, take the document chunks that have preset keywords and are obtained in the first several preset chunking rounds as the second specified document chunks; For the second specified document chunks, extract the corresponding summary information of the second specified document chunks by invoking a large model; For other document chunks except the second specified document chunks, add the summary information to them.
4. The accounting voucher generation method based on document block traversal according to claim 2, wherein, The method further includes: Determine that the smallest chunk unit does not meet the preset word count requirement and cannot continue to be chunked by the chunking criterion; Divide the smallest chunk unit evenly according to the word count to obtain several sub-units that meet the preset word count requirement; For each sub-unit, perform two-way supplementation in the character order in the accounting voucher rule document until the sub-unit is supplemented to meet the preset word count requirement; Delete the content that does not meet the preset minimum text area in the sub-units that meet the word count requirement to obtain new document chunks.
5. The accounting voucher generation method based on document block traversal according to claim 1, wherein, Based on a pre-generated key information extraction instruction template, by invoking a large model, extract the corresponding key information according to the document chunks, specifically including: Obtain a pre-generated key information extraction instruction template, and update the key information extraction instruction template based on the document chunks and the current scenario requirements to obtain key information extraction prompt words; Invoke a large model, and extract the sub-key information corresponding to each document chunk through the key information extraction prompt words; Integrate the sub-key information to obtain the key information corresponding to the accounting voucher rule document.
6. The accounting voucher generation method based on document block traversal according to claim 5, wherein, Invoke a large model, and extract the sub-key information corresponding to each document chunk through the key information extraction prompt words, specifically including: Invoke a large model, and in the first conversation for extracting sub-key information, extract the first sub-key information corresponding to the current document chunk through the key information extraction prompt words; According to the sequence relationship corresponding to each document block, obtain the historical document block before the current document block corresponding to the accounting voucher rule document, and obtain the second sub-critical information corresponding to the historical document block; Call the large model. In the new second conversation, based on the accounting voucher rule document, determine the memory weights between the first sub-critical information and each second sub-critical information; Call the large model. In the first conversation, based on the memory weights, re-extract the first sub-critical information corresponding to the current document block through the critical information extraction prompt words.
7. The accounting voucher generation method based on document block traversal according to claim 1, wherein Based on the pre-generated accounting voucher generation instruction template, by calling the large model, generate accounting vouchers according to the obtained reimbursement form information and the critical information, specifically including: Obtain the pre-generated accounting voucher generation instruction template, and update the accounting voucher generation instruction template based on the critical information and the obtained reimbursement form information to obtain the accounting voucher generation prompt words; Call the large model, and generate accounting vouchers according to the accounting voucher rule document through the accounting voucher generation prompt words.
8. The accounting voucher generation method based on document block traversal according to claim 7, wherein The method further includes: Obtain the feedback opinions of the current department and the downstream department on the accounting vouchers; Modify the large model architecture and the content of the prompt words based on the feedback opinions.
9. An accounting voucher generation device based on document block traversal, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the accounting voucher generation method based on document block traversal as described in any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: the accounting voucher generation method based on document block traversal as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method and system for generating accounting voucher
CN110458674A
Document segmentation method, device and equipment and readable storage medium
CN117520549A