Document Data Processing Method, Device and Equipment

By automatically processing document data and rule expressions, the rule engine is used to realize the automated process of letter of credit review, solving the problem of accuracy and inefficiency in the existing technology, and significantly improving the efficiency and accuracy of review.

CN114219443BActive Publication Date: 2025-06-20CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111547859.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-06-20
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

In the prior art, the document review operation of letters of credit is prone to omissions or errors due to the long time required for manual review, resulting in low accuracy and efficiency of document review.

Method used

By obtaining pending document data and rule expressions, the rule engine is used to automatically process the document data to generate audit result information. This method includes named entity recognition, regular matching, calling the rule engine and generating the parser to implement an automated review process.

Benefits of technology

Through the automated review process, the time required for review is significantly reduced, the efficiency and accuracy of review is improved, and the accuracy and inefficiency of review business is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114219443B_ABST
    Figure CN114219443B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, and device for processing document data, relating to data processing technologies. The method includes: obtaining a plurality of document data to be processed; the plurality of document data to be processed includes first document data of a first node and second document data of a second node, and the second node is a superior node of the first node; in a preset rule library, using a preset rule engine to obtain a plurality of rule expressions; the rule library includes a plurality of rule expressions, and the rule engine is a program for executing the rule expressions; based on the correspondence between the document data to be processed and the rule expressions, performing data processing on each piece of document data to be processed through the rule expression corresponding to each piece of document data to be processed, and obtaining audit result information. The method of the present application automatically calls the rule expressions in the rule library through the rule engine, and performs data processing on the document data to be processed through the rule expressions, forming an automated document auditing process, greatly improving the efficiency and accuracy of document auditing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to data processing technologies, and in particular to a method, apparatus, and device for processing document data. Background Art

[0002] Currently, with the development of trade, the operation of examining documents against letters of credit is extremely important.

[0003] In the prior art, during the operation of examining documents against letters of credit, usually the information in the paper-based documents is manually entered into the system, and then manually checked item by item against the examination rules to obtain the examination results. Finally, all the examination results are summarized to generate an examination conclusion.

[0004] However, in the prior art, since the time required for manual examination is relatively long and it is prone to omissions or errors, the accuracy and efficiency of document examination are relatively low. Summary of the Invention

[0005] This application provides a method, apparatus, and device for processing document data to solve the technical problems of relatively low accuracy and efficiency in document examination services.

[0006] In a first aspect, this application provides a method for processing document data, including:

[0007] Obtain multiple pieces of document data to be processed; wherein, the multiple pieces of document data to be processed include the first document data of the first node and the second document data of the second node, and the second node is the upper-level node of the first node;

[0008] In a preset rule library, use a preset rule engine to obtain multiple rule expressions; wherein, the rule library includes multiple rule expressions, and the rule engine is a program for executing rule expressions;

[0009] Based on the correspondence between the document data to be processed and the rule expressions, perform data processing on each piece of document data to be processed through the rule expression corresponding to each piece of document data to be processed to obtain audit result information.

[0010] Further, obtaining multiple pieces of document data to be processed includes:

[0011] Determine the data information of the first document data of the first node through a named entity recognition method and a regular matching method, where the data information includes a document name and a document value;

[0012] Based on the correspondence between the preset document name and the upper-level node, determine in a preset database that the upper-level node corresponding to the document name in the data information of the first document data is the second node; wherein, the second node has the data information of the second document data;

[0013] Replace the document value of the second document data of the second node with the document value of the first document data of the first node to obtain the updated second document data of the second node.

[0014] Further, the document value includes audit elements and the corresponding element values of the audit elements.

[0015] Further, the method further includes:

[0016] Obtain multiple document names;

[0017] Determine the upper-level node corresponding to each document name, generate the correspondence between the document name and the upper-level node according to each document name and the upper-level node corresponding to each document name, and store the correspondence in the database.

[0018] Further, the method further includes:

[0019] Obtain multiple rule texts; wherein, the rule text includes a document name, a document value, and an operator name;

[0020] Determine the upper-level node corresponding to each document name according to the correspondence between the document name and the upper-level node;

[0021] Generate a rule expression corresponding to the upper-level node according to the document name, the audit element, and the operator name.

[0022] Further, generating a rule expression corresponding to the upper-level node according to the document name, the audit element, and the operator name includes:

[0023] Based on the document name, the audit element, and the operator name, train the initial pre-trained language T5 model through a machine translation method to obtain a conversion model for generating rule expressions;

[0024] Through the conversion model, convert to obtain a rule expression corresponding to the combination of the document name, the audit element, and the operator name.

[0025] Further, after obtaining multiple rule expressions by using a preset rule engine in a preset rule library, it further includes:

[0026] Convert the rule expression into a rule expression in a preset language; wherein, the rule expression in the preset language is used to run through the rule engine.

[0027] Further, the method further includes:

[0028] Based on standard grammar information and operator names, a parser corresponding to a regular expression is generated through a preset language, and a rule engine for executing the parser is generated; wherein, the operator names are used to represent calculation logic and parsing logic, and the standard grammar information is used to represent grammar information about the operator names.

[0029] In a second aspect, the present application provides a document data processing device, including:

[0030] A data acquisition unit, configured to acquire a plurality of to-be-processed document data; wherein, the plurality of to-be-processed document data includes first document data of a first node and second document data of a second node, and the second node is a superior node of the first node;

[0031] A rule acquisition unit, configured to acquire a plurality of rule expressions by using a preset rule engine in a preset rule library; wherein, the rule library includes a plurality of rule expressions, and the rule engine is a program for executing rule expressions;

[0032] A processing unit, configured to perform data processing on each of the to-be-processed document data through a rule expression corresponding to each to-be-processed document data based on the corresponding relationship between the to-be-processed document data and the rule expressions, so as to obtain audit result information.

[0033] Further, the data acquisition unit includes:

[0034] A first document data determination module, configured to determine data information of the first document data of the first node by using a named entity recognition method and a regular matching method, wherein the data information includes a document name and a document value;

[0035] A second node determination module, configured to determine, based on a corresponding relationship between a preset document name and a superior node, that a superior node corresponding to the document name in the data information of the first document data in a preset database is the second node; wherein, the second node has data information of the second document data;

[0036] A replacement module, configured to replace the document value of the second document data of the second node with the document value of the first document data of the first node, so as to obtain the updated second document data of the second node.

[0037] Further, the document value includes audit elements and element values corresponding to the audit elements.

[0038] Further, the device further includes:

[0039] A document name acquisition unit, configured to acquire a plurality of document names;

[0040] A storage relation unit, which is used to determine the upper-level node corresponding to each document name, generate the corresponding relationship between the document name and the upper-level node according to each document name and the upper-level node corresponding to each document name, and store the corresponding relationship in a database.

[0041] Further, the device further includes:

[0042] A rule text acquisition unit, which is used to acquire a plurality of rule texts; wherein, the rule text includes a document name, a document value, and an operator name.

[0043] An upper-level node determination unit, which is used to determine the upper-level node corresponding to each document name according to the corresponding relationship between the document name and the upper-level node.

[0044] A rule generation unit, which is used to generate a rule expression corresponding to the upper-level node according to the document name, the audit element, and the operator name.

[0045] Further, the rule generation unit includes:

[0046] A training module, which is used to train an initial pre-trained language T5 model based on the document name, the audit element, and the operator name through a machine translation method to obtain a conversion model for generating rule expressions.

[0047] A generation module, which is used to convert through the conversion model to obtain a rule expression corresponding to the combination of the document name, the audit element, and the operator name.

[0048] Further, the device further includes:

[0049] A conversion unit, which is used to convert the rule expression into a rule expression in a preset language after obtaining a plurality of rule expressions by using a preset rule engine in a preset rule library; wherein, the rule expression in the preset language is used to run through the rule engine.

[0050] Further, the device further includes:

[0051] A rule engine generation unit, which is used to generate a parser corresponding to the rule expression and a rule engine for executing the parser through a preset language based on standard grammar information and an operator name; wherein, the operator name is used to represent calculation logic and parsing logic, and the standard grammar information is used to represent grammar information about the operator name.

[0052] In a third aspect, the present application provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the method described in the first aspect is implemented.

[0053] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect.

[0054] In a fifth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0055] A method, device, and equipment for processing document data provided by the present application obtain a plurality of document data to be processed. Among them, the plurality of document data to be processed includes first document data of a first node and second document data of a second node, and the second node is a superior node of the first node. In a preset rule library, a plurality of rule expressions are obtained by using a preset rule engine. The rule library includes a plurality of rule expressions, and the rule engine is a program for executing rule expressions. Based on the correspondence between the document data to be processed and the rule expressions, each document data to be processed is processed through the rule expression corresponding to each document data to be processed, and audit result information is obtained. In this solution, since the preset rule library includes a plurality of rule expressions and the rule engine is a program for executing rule expressions, a plurality of rule expressions can be obtained by using the preset rule engine in the preset rule library. Then, based on the correspondence between the document data to be processed and the rule expressions, the rule expression corresponding to each document data to be processed is first determined, and then each document data to be processed is processed through the determined rule expression to obtain audit result information. Therefore, the rule engine automatically calls the rule expressions in the rule library, and then the document data to be processed is processed through the rule expressions, forming an automated document auditing process, reducing the time required for document auditing, greatly improving the efficiency and accuracy of document auditing, and solving the technical problems of low accuracy and low efficiency in the document auditing business. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0057] Figure 1 It is a schematic flowchart of a method for processing document data provided by an embodiment of the present application;

[0058] Figure 2 It is a schematic flowchart of another method for processing document data provided by an embodiment of the present application;

[0059] Figure 3 A schematic structural diagram of a document data processing device provided by an embodiment of the present application;

[0060] Figure 4 A schematic structural diagram of another document data processing device provided by an embodiment of the present application;

[0061] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0062] Figure 6 A block diagram of an electronic device provided by an embodiment of the present application.

[0063] Through the above-mentioned drawings, specific embodiments of the present disclosure have been shown, and there will be more detailed descriptions hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0064] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure.

[0065] In one example, with the development of trade, the document review operation for letters of credit is extremely important. In the prior art, during the document review operation for letters of credit, usually, the information in the paper-based documents is manually entered into the system, and then manually checked item by item against the review rules to obtain the review results. Finally, all the review results are summarized to generate a review conclusion. However, in the prior art, since the time required for manual review is relatively long and it is easy to miss or make mistakes, the accuracy and efficiency of document review are relatively low.

[0066] A document data processing method, device, and equipment provided by the present application aim to solve the above-mentioned technical problems in the prior art.

[0067] The following uses specific embodiments to describe in detail the technical solutions of the present application and how the technical solutions of the present application solve the above-mentioned technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the drawings.

[0068] Figure 1 A flowchart of a document data processing method provided by an embodiment of the present application, as Figure 1 shown, the method includes:

[0069] 101. Obtain multiple pieces of data of bills to be processed; among them, the multiple pieces of data of bills to be processed include the first bill data of the first node and the second bill data of the second node, and the second node is the upper-level node of the first node.

[0070] Exemplarily, the execution subject of this embodiment may be an electronic device, or a terminal device, or a bill data processing device or equipment, or other devices or equipment that can execute this embodiment, and there is no limitation thereto. In this embodiment, the execution subject is introduced as an electronic device.

[0071] First, it is necessary to obtain multiple pieces of data of bills to be processed. The data of bills to be processed can be obtained from the memory; or, the data of bills to be processed can be obtained from a web page, or the data of bills to be processed transmitted by other devices can be received. The multiple pieces of data of bills to be processed include the first bill data of the first node and the second bill data of the second node, and the second node is the upper-level node of the first node, that is, the second node is the parent node of the first node.

[0072] 102. In a preset rule library, use a preset rule engine to obtain multiple rule expressions; among them, the rule library includes multiple rule expressions, and the rule engine is a program for executing the rule expressions.

[0073] Exemplarily, the preset rule library includes multiple rule expressions. The rule expressions are expressions preset for processing the data of bills to be processed, and the rule engine is a program preset for executing the rule expressions. Therefore, multiple rule expressions can be obtained in the preset rule library by using the preset rule engine.

[0074] 103. Based on the correspondence between the data of bills to be processed and the rule expressions, perform data processing on each piece of data of bills to be processed through the rule expressions corresponding to each piece of data of bills to be processed to obtain audit result information.

[0075] Exemplarily, for each piece of data of bills to be processed, first determine whether there is a corresponding rule expression. If there is a corresponding rule expression, perform data processing on each piece of data of bills to be processed through the rule expression corresponding to each piece of data of bills to be processed to obtain audit result information, where the audit result information includes: the number of the involved rule expression, the audit elements involved in each operator name in the rule expression, and the calculation result; if there is no corresponding rule expression, skip this rule expression.

[0076] For example, the audit result information includes the data stream number, information on whether the audit process is successful or not, details of the audit result, rule number, rule result, each audit point, operator name, operator result, and audit elements involved in the operator, etc. Among them, the data stream number represents the number of the document data to be processed. The information on whether the audit process is successful or not includes error information, correct information, etc. The details of the audit result are a list, and each element in the list records the situation of each audit rule. Therefore, through key information such as the data stream number and data number, it is very convenient to review and trace each audit business, facilitating the development of audit and management work.

[0077] In the embodiment of the present application, multiple pieces of document data to be processed are obtained; among them, the multiple pieces of document data to be processed include the first document data of the first node and the second document data of the second node, and the second node is the upper-level node of the first node. In the preset rule library, a plurality of rule expressions are obtained by using the preset rule engine; among them, the rule library includes a plurality of rule expressions, and the rule engine is a program for executing the rule expressions. Based on the correspondence between the document data to be processed and the rule expressions, each piece of document data to be processed is processed through the rule expression corresponding to each piece of document data to be processed to obtain the audit result information. In this solution, since the preset rule library includes a plurality of rule expressions and the rule engine is a program for executing the rule expressions, a plurality of rule expressions can be obtained by using the preset rule engine in the preset rule library, and then based on the correspondence between the document data to be processed and the rule expressions, first determine the rule expression corresponding to each piece of document data to be processed, and then process each piece of document data to be processed through the determined rule expression to obtain the audit result information. Therefore, by automatically invoking the rule expressions in the rule library through the rule engine and then processing the document data to be processed through the rule expressions, an automated document audit process is formed, reducing the time required for document auditing and greatly improving the efficiency and accuracy of document auditing, and solving the technical problems of low accuracy and low efficiency in the document audit business.

[0078] Figure 2 It is a schematic flowchart of another method for processing document data provided by the embodiment of the present application, as Figure 2 shown, and the method includes:

[0079] 201. Obtain multiple document names.

[0080] Exemplarily, in a single bill review operation, the involved document data D1, D2, …, Dn are scanned and input into the preprocessing module in sequence. At this time, each piece of document data is converted into picture formats P1, P2, …, Pn. Then, the documents P1, P2, …, Pn are input into the Intelligent Character Recognition (ICR) module. This ICR module first uses Optical Character Recognition (OCR) technology to convert the text in the images into text formats, and then through deep learning algorithms, combined with the context statement information and semantic network knowledge base of each character, conducts semantic reasoning and semantic analysis to achieve the purpose of error correction and improving the recognition accuracy rate. At this time, each piece of document data is converted into text formats T1, T2, …, Tn, and then multiple document names are obtained.

[0081] 202. Determine the upper-level node corresponding to each document name. According to each document name and the upper-level node corresponding to each document name, generate the corresponding relationship between the document name and the upper-level node, and store the corresponding relationship in the database.

[0082] Exemplarily, in the field of international settlement, because there are a wide variety of document types, a set of upper and lower level attribution relationships can be sorted out and stored in the graph database. All kinds of documents under the same major category apply the same set of rules, which can reduce the workload when establishing the rule library. Therefore, the electronic device can determine the upper-level node corresponding to each document name. According to each document name and the upper-level node corresponding to each document name, generate the corresponding relationship between the document name and the upper-level node, and store the corresponding relationship in the database. The database includes non-relational databases such as graph databases. It is also necessary to store the document names and review elements involved in all business scenarios, as well as the English terms corresponding to the document names and the English terms corresponding to the review elements, in a relational database.

[0083] 203. Obtain multiple rule texts; wherein, the rule texts include document names, document values, and operator names.

[0084] Exemplarily, in order to generate a rule expression, the electronic device needs to first obtain multiple rule texts, and the rule texts include document names, document values, and operator names.

[0085] 204. According to the corresponding relationship between the document name and the upper-level node, determine the upper-level node corresponding to each document name.

[0086] Exemplarily, the electronic device can determine the upper-level node corresponding to each document name according to the correspondence between the document name and the upper-level node, which is equivalent to determining the parent node corresponding to each document name, facilitating the generation of a rule expression for the upper-level node in the next step. Compared with generating rule expressions for each first node, the number of rule expressions can be greatly reduced, thereby improving the efficiency of building the rule library.

[0087] 205. Train the initial pre-trained language T5 model by machine translation method based on the document name, audit elements, and operator name to obtain a conversion model for generating rule expressions.

[0088] Exemplarily, the ways of generating rule expressions include three ways: manual, semi-automatic, and full-automatic for rule configuration. When performing full-automatic rule configuration, rule expressions can be entered in batches. To improve efficiency, an automatic conversion toolbox can be used. The automatic conversion process includes: based on the document name, audit elements, and operator name, input the rule text and the corresponding rule expression into the initial pre-trained language T5 model of Google, train the initial pre-trained language T5 model by machine translation method to complete transfer learning, and obtain a conversion model for generating rule expressions. This conversion model is a model that can adapt to the rule conversion task and can be applied to multiple scenarios.

[0089] Exemplarily, when performing manual rule configuration, the document review personnel can perform operations such as adding, deleting, modifying, and querying the database. For example, for the rule "The payer of the contract should be consistent with the consuming enterprise of the Bill of Entry":

[0090] a) If the person is very familiar with domain terms and operators, the rule expression can be written directly.

[0091] b) If the person is not sure about some domain terms or operators, the database toolbox can be used to select the required document name, element name, and operator name in the form of a drop-down box query, and splice them to form a rule expression.

[0092] c) The person can also perform operations such as deleting and modifying the rule expression.

[0093] 206. Through the conversion model, convert to obtain a rule expression corresponding to the document name, audit elements, and operator name.

[0094] Exemplarily, in the scenario of processing document data, inputting a rule text including the document name, audit elements, and operator name into a conversion model can obtain a rule expression corresponding to the document name, audit elements, and operator name. Finally, the obtained rule expression needs to be input into a rule verification tool generated by a preset rule engine. If the grammar is correct, it is stored in the rule library; otherwise, the rule is revised according to the error information. In other scenarios, for example, in the scenario of translating Chinese into English, inputting Chinese and the rule text for converting Chinese into English into the conversion model can obtain a conversion expression corresponding to the Chinese.

[0095] Exemplarily, after the rule expression is stored in the rule library, if the rule expression in the rule library changes, the original rule expression is directly updated to the latest rule expression. Or, during the process of processing the document data to be processed according to the rule expression, in the preset rule library, after obtaining multiple rule expressions by using the preset rule engine, calculate the first hash value of the rule expression through the preset hash function, and cache the obtained rule expressions. Then, process the document data to be processed according to the rule expression to complete the current document review service. When processing the next document review service, it is necessary to calculate the second hash value of the rule expression in the rule library through the preset hash function, and compare the second hash value with the first hash value. If the second hash value is inconsistent with the first hash value, it means that the rule expression in the rule library has been updated. At this time, it is necessary to re-obtain the updated rule expression in the rule library. If the second hash value is consistent with the first hash value, it means that the rule expression in the rule library has not been updated. At this time, the cached rule expression can be directly used without having to obtain the rule expression from the rule library again. Because the number of rule expressions is large and the change cycle is long, through this step of processing, when the rule library has not changed, using the cached rule expression to process document data can effectively save the conversion time of the rule expression and improve the overall review efficiency.

[0096] 207. Based on the standard grammar information and operator name, generate a parser corresponding to the rule expression through a preset language, and generate a rule engine for executing the parser; wherein, the operator name is used to represent the calculation logic and parsing logic, and the standard grammar information is used to represent the grammar information about the operator name.

[0097] Exemplarily, as shown in Table 1 below, the operator name is used to characterize the calculation logic and parsing logic. The operator name includes multiple operators, operator types, and descriptions (i.e., calculation logic and parsing logic). The standard grammar information characterizes the grammar information about the operator name. The Antlr (Another Tool for Language Recognition) tool is used to write the g4 grammar file to define the standard grammar information of the above operators. For example, when multiple operator names are involved, the priority processing order of multiple operator names, etc. After verifying that the g4 grammar is correct, the electronic device generates a parser corresponding to the regular expression based on the standard grammar information and the operator name. The parser includes a lexer lexical analyzer, a parser syntax analyzer, and also generates a ParseTreeVisitor visitor, a ParseTreeListener listener, and a rule verification tool for the syntax analyzer. The lexical analyzer, syntax analyzer, visitor, listener, and rule verification tool are all Java code, and a rule engine for executing the lexical analyzer, syntax analyzer, visitor, listener, and rule verification tool is generated.

[0098] During the process of generating the lexical analyzer, syntax analyzer, visitor, listener, and rule verification tool, the listener method has no return value (the return type is void), and it will automatically complete the traversal of the syntax analyzer. In the present invention, in order to control the behavior of each child node in the syntax analysis tree, a custom visitor class is implemented by inheriting the visitor base class method, and the code for specifically accessing the node is completed therein. Considering that there are already preset arithmetic operators, logical operators, comparison operators, etc. in javascript, and lambda expressions are supported, which is convenient for expanding the usage of operators, the specific logic of the operator is implemented through javascript functions to achieve the purpose of simplifying the development workload. The return value of the above custom visitor class is set to a javascript statement. When the rule engine runs, the Java program calls node.js to complete the calculation process of javascript.

[0099] Table 1

[0100]

[0101] 208. Determine the data information of the first document data of the first node through the named entity recognition method and the regular matching method, where the data information includes the document name and the document value.

[0102] In one example, the document value includes the audit elements and the corresponding element values of the audit elements.

[0103] Exemplarily, the electronic device determines the data information of the first document data of the first node through the named entity recognition method (NER) and the regular matching method. Among them, the data information includes the document name and the document value. The document value includes the review elements and the corresponding element values of the review elements, and key-value pair Json texts J1, J2,..., Jn with the review elements as keywords can be generated.

[0104] For example, taking the bill of lading as an example, a Json text J1 includes a parent node, a first child node, a second child node, a grandchild node, the unique identification number of the review element, and the value of the review element. Among them, the parent node is the document name, the first child node is the document category, the second child node is the review element extracted from the document, and the grandchild node is the data type of the review element.

[0105] After obtaining multiple Json texts J1, J2,..., the documents J1, J2,..., Jn can also be input into the data integration module. This module splices all the document Json texts involved in the same bill review business to obtain the spliced result information. Furthermore, all the to-be-processed document data of the same bill review business can be quickly spliced using the characteristics of the Json format, achieving the purpose of batch review.

[0106] 209. Based on the correspondence relationship between the preset document name and the upper node, determine the upper node corresponding to the document name in the data information of the first document data as the second node in the preset database; among them, the second node has the data information of the second document data.

[0107] Exemplarily, based on the correspondence relationship between the preset document name and the upper node, the electronic device can first query in the preset database the upper node corresponding to the document name in the data information of the first document data, and then determine this upper node as the second node of the document name in the data information of the first document data. The second node has the data information of the second document data.

[0108] For example, query all the upper nodes of the document name in the data information of the first document data in the database and add these upper nodes as new keys, or query all the upper nodes of the document name in the data information of each first document data in the spliced result information in the database and add these upper nodes as new keys.

[0109] 210. Replace the document value of the second document data of the second node with the document value of the first document data of the first node to obtain the updated second document data of the second node.

[0110] Exemplarily, the electronic device assigns the document value of the first document data of the first node to the document value of the second document data of the second node, and obtains the updated second document data of the second node.

[0111] 211. In a preset rule library, use a preset rule engine to obtain multiple rule expressions; wherein, the rule library includes multiple rule expressions, and the rule engine is a program for executing rule expressions.

[0112] Exemplarily, for this step, reference can be made to Figure 1 step 102 in it, which will not be elaborated here.

[0113] 212. Convert the rule expressions into rule expressions in a preset language; wherein, the rule expressions in the preset language are used to be run through the rule engine.

[0114] Exemplarily, the preset language is the Java language, which is the same as the language used by the program of the rule engine. The electronic device converts the rule expressions into rule expressions of JavaScript statements. Therefore, the rule expressions in the preset language can be run through the rule engine.

[0115] 213. Based on the correspondence between the document data to be processed and the rule expressions, perform data processing on each piece of document data to be processed through the rule expressions corresponding to each piece of document data to be processed, and obtain the audit result information.

[0116] Exemplarily, for each piece of document data to be processed, first determine whether there is a corresponding rule expression. If there is a corresponding rule expression, perform data processing on each piece of document data to be processed through the rule expressions of JavaScript statements corresponding to each piece of document data to be processed, and obtain the audit result information.

[0117] In the embodiments of the present application, multiple document names are obtained. The upper-level node corresponding to each document name is determined. According to each document name and the upper-level node corresponding to each document name, the corresponding relationship between the document name and the upper-level node is generated and stored in the database. Multiple rule texts are obtained; wherein, the rule text includes a document name, a document value, and an operator name. According to the corresponding relationship between the document name and the upper-level node, the upper-level node corresponding to each document name is determined. Based on the document name, the audit element, and the operator name, the initial pre-trained language T5 model is trained by a machine translation method to obtain a conversion model for generating a rule expression. Through the conversion model, a rule expression corresponding to the document name, the audit element, and the operator name is obtained. Based on the standard grammar information and the operator name, a parser corresponding to the rule expression is generated by a preset language, and a rule engine for executing the parser is generated; wherein, the operator name is used to represent the calculation logic and the parsing logic, and the standard grammar information is used to represent the grammar information about the operator name. By a named entity recognition method and a regular matching method, the data information of the first document data of the first node is determined, wherein the data information includes a document name and a document value. Based on the preset corresponding relationship between the document name and the upper-level node, the upper-level node corresponding to the document name in the data information of the first document data is determined in the preset database as the second node; wherein, the second node has the data information of the second document data. The document value of the second document data of the second node is replaced with the document value of the first document data of the first node to obtain the updated second document data of the second node. In the preset rule library, multiple rule expressions are obtained by using the preset rule engine; wherein, the rule library includes multiple rule expressions, and the rule engine is a program for executing the rule expression. The rule expression is converted into a rule expression in a preset language; wherein, the rule expression in the preset language is used to be run by the rule engine. Based on the corresponding relationship between the to-be-processed document data and the rule expression, each to-be-processed document data is processed through the rule expression corresponding to each to-be-processed document data to obtain the audit result information. Therefore, the rule engine automatically calls the rule expressions in the rule library, and then the to-be-processed document data is processed through the rule expressions, forming an automated document audit process, reducing the time required for document auditing, greatly improving the efficiency and accuracy of document auditing, and solving the technical problems of low accuracy and low efficiency in the document audit business.

[0118] Figure 3 FIG. is a schematic structural diagram of a document data processing device provided by an embodiment of the present application, as Figure 3 shown, the device includes:

[0119] Obtain a data unit 31 for obtaining multiple pieces of bill data to be processed; among them, the multiple pieces of bill data to be processed include the first bill data of the first node and the second bill data of the second node, and the second node is the upper-level node of the first node;

[0120] Obtain a rule unit 32 for obtaining multiple rule expressions by using a preset rule engine in a preset rule library; among them, the rule library includes multiple rule expressions, and the rule engine is a program for executing rule expressions;

[0121] A processing unit 33 for performing data processing on each piece of bill data to be processed through the rule expression corresponding to each piece of bill data to be processed based on the corresponding relationship between the bill data to be processed and the rule expression, and obtaining audit result information.

[0122] The device in this embodiment can execute the technical solutions in the above method, and its specific implementation process and technical principle are the same, which will not be elaborated here.

[0123] Figure 4 It is a structural schematic diagram of another bill data processing device provided by an embodiment of the present application. On the basis of the embodiment shown in Figure 3 as shown in Figure 4 as shown, the data acquisition unit 31 includes:

[0124] A first bill data determination module 311 for determining the data information of the first bill data of the first node through a named entity recognition method and a regular matching method, where the data information includes a bill name and a bill value.

[0125] A second node determination module 312 for determining, based on the corresponding relationship between a preset bill name and an upper-level node, that the upper-level node corresponding to the bill name in the data information of the first bill data in a preset database is the second node; among them, the second node has the data information of the second bill data.

[0126] A replacement module 313 for replacing the bill value of the second bill data of the second node with the bill value of the first bill data of the first node to obtain the updated second bill data of the second node.

[0127] In one example, the bill value includes audit elements and the corresponding element values of the audit elements.

[0128] In one example, the device further includes:

[0129] A bill name acquisition unit 41 for acquiring multiple bill names.

[0130] A storage relationship unit 42 is used to determine the upper-level node corresponding to each document name, generate the corresponding relationship between the document name and the upper-level node according to each document name and the upper-level node corresponding to each document name, and store the corresponding relationship in the database.

[0131] In one example, the device further includes:

[0132] A rule text acquisition unit 43 is used to acquire multiple rule texts; wherein, the rule text includes a document name, a document value, and an operator name.

[0133] An upper-level node determination unit 44 is used to determine the upper-level node corresponding to each document name according to the corresponding relationship between the document name and the upper-level node.

[0134] A rule generation unit 45 is used to generate a rule expression corresponding to the upper-level node according to the document name, the audit element, and the operator name.

[0135] In one example, the rule generation unit 45 includes:

[0136] A training module 451 is used to train an initial pre-trained language T5 model based on the document name, the audit element, and the operator name through a machine translation method to obtain a conversion model for generating rule expressions.

[0137] A generation module 452 is used to convert through the conversion model to obtain a rule expression corresponding to the combination of the document name, the audit element, and the operator name.

[0138] In one example, the device further includes:

[0139] A conversion unit 46 is used to, in a preset rule library, after obtaining multiple rule expressions by using a preset rule engine, convert the rule expressions into rule expressions in a preset language; wherein, the rule expressions in the preset language are used to be run by the rule engine.

[0140] In one example, the device further includes:

[0141] A rule engine generation unit 47 is used to generate a parser corresponding to the rule expression and a rule engine for executing the parser in a preset language based on standard syntax information and the operator name; wherein, the operator name is used to represent the calculation logic and the parsing logic, and the standard syntax information is used to represent the syntax information about the operator name.

[0142] The device in this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.

[0143] Figure 5A schematic structural diagram of an electronic device provided by an embodiment of the present application is as follows Figure 5 As shown, the electronic device includes: a memory 51 and a processor 52.

[0144] The memory 51 stores a computer program that can run on the processor 52.

[0145] The processor 52 is configured to execute the method provided by the above embodiment.

[0146] The electronic device further includes a receiver 53 and a transmitter 54. The receiver 53 is used to receive instructions and data sent by an external device, and the transmitter 54 is used to send instructions and data to the external device.

[0147] Figure 6 It is a block diagram of an electronic device provided by an embodiment of the present application. The electronic device can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0148] The device 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.

[0149] The processing component 602 generally controls the overall operation of the device 600, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.

[0150] The memory 604 is configured to store various types of data to support the operation of the device 600. Examples of these data include instructions for any application or method operating on the device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0151] The power supply component 606 provides power for various components of the device 600. The power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 600.

[0152] The multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0153] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive external audio signals when the device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.

[0154] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0155] The sensor assembly 614 includes one or more sensors for providing an assessment of the status of the device 600 in various aspects. For example, the sensor assembly 614 can detect the on / off state of the device 600, the relative positioning of components, such as the display and keypad of the device 600. The sensor assembly 614 can also detect a change in the position of the device 600 or a component of the device 600, the presence or absence of user contact with the device 600, the orientation or acceleration / deceleration of the device 600, and the temperature change of the device 600. The sensor assembly 614 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0156] The communication component 616 is configured to facilitate communication between the device 600 and other devices in a wired or wireless manner. The device 600 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0157] In an exemplary embodiment, the device 600 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0158] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 604 including instructions, and the above instructions can be executed by the processor 620 of the device 600 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0159] An embodiment of the present application also provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device can execute the method provided in the above embodiment.

[0160] An embodiment of the present application further provides a computer program product, which includes: a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the execution of the computer program by the at least one processor causes the electronic device to execute the solution provided in any of the above embodiments.

[0161] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0162] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for processing document data, characterized in that, Including: Obtain multiple pieces of data of documents to be processed; among them, the multiple pieces of data of documents to be processed include the first document data of the first node and the second document data of the second node, and the second node is the upper-level node of the first node; the data of the document to be processed is the data of the documents of the letter of credit. In a preset rule library, use a preset rule engine to obtain multiple rule expressions; among them, the rule library includes multiple rule expressions, and the rule engine is a program for executing rule expressions. Based on the correspondence between the data of the document to be processed and the rule expressions, perform data processing on each piece of data of the document to be processed through the rule expression corresponding to each piece of data of the document to be processed to obtain audit result information. Obtain multiple pieces of data of documents to be processed, including: Determine the data information of the first document data of the first node through a named entity recognition method and a regular matching method, where the data information includes the document name and the document value; the document value includes the audit elements and the corresponding element values of the audit elements. Based on the correspondence between the preset document name and the upper-level node, determine in the preset database that the upper-level node corresponding to the document name in the data information of the first document data is the second node; among them, the second node has the data information of the second document data. Replace the document value of the second document data of the second node with the document value of the first document data of the first node to obtain the updated second document data of the second node. The method further includes: Obtain multiple rule texts; among them, the rule text includes the document name, the document value, and the operator name. Determine the upper-level node corresponding to each document name according to the correspondence between the document name and the upper-level node. Generate a rule expression corresponding to the upper-level node according to the document name, the audit element, and the operator name.

2. The method according to claim 1, characterized in that, The method further includes: Obtain multiple document names. Determine the upper-level node corresponding to each document name, generate the correspondence between the document name and the upper-level node according to each document name and the upper-level node corresponding to each document name, and store the correspondence in the database.

3. The method according to claim 1, characterized in that, Generating a rule expression corresponding to the upper-level node according to the document name, the audit element, and the operator name includes: Based on the document name, the audit element, and the operator name, train the initial pre-trained language T5 model through a machine translation method to obtain a conversion model for generating rule expressions. Through the conversion model, convert to obtain a rule expression corresponding to the three of the document name, the audit element, and the operator name.

4. The method according to any one of claims 1 - 3, characterized in that, After obtaining multiple rule expressions by using a preset rule engine in a preset rule library, it further includes: Convert the rule expression into a rule expression in a preset language; where the rule expression in the preset language is used to run through the rule engine.

5. The method according to any one of claims 1 - 3, characterized in that, The method further includes: Based on standard grammar information and operator names, a parser corresponding to a regular expression is generated through a preset language, and a rule engine for executing the parser is generated; wherein, the operator names are used to represent calculation logic and parsing logic, and the standard grammar information is used to represent grammar information about the operator names.

6. A device for processing document data, characterized in that, It includes: A data acquisition unit for acquiring a plurality of data of documents to be processed; wherein, the plurality of data of documents to be processed include the first data of the first document of the first node and the second data of the second document of the second node, and the second node is the upper node of the first node; the data of the document to be processed is the data of the document of the letter of credit. A rule acquisition unit for acquiring a plurality of rule expressions by using a preset rule engine in a preset rule library; wherein, the rule library includes a plurality of rule expressions, and the rule engine is a program for executing rule expressions. A processing unit for performing data processing on each data of the document to be processed through a rule expression corresponding to each data of the document to be processed based on the corresponding relationship between the data of the document to be processed and the rule expression, and obtaining audit result information. The data acquisition unit includes: A first document data determination module for determining the data information of the first data of the first document of the first node through a named entity recognition method and a regular matching method, wherein the data information includes a document name and a document value; the document value includes audit elements and element values corresponding to the audit elements. A second node determination module for determining, based on the corresponding relationship between the preset document name and the upper node, that the upper node corresponding to the document name in the data information of the first data of the first document in a preset database is the second node; wherein, the second node has the data information of the second data of the second document. A replacement module for replacing the document value of the second data of the second document of the second node with the document value of the first data of the first document of the first node to obtain the updated second data of the second document of the second node. The device further includes: A rule text acquisition unit for acquiring a plurality of rule texts; wherein, the rule texts include document names, document values, and operator names. An upper node determination unit for determining the upper node corresponding to each document name according to the corresponding relationship between the document name and the upper node. A rule generation unit for generating a rule expression corresponding to the upper node according to the document name, the audit elements, and the operator names.

7. The device according to claim 6, wherein, The device further includes: A document name acquisition unit for acquiring a plurality of document names. A storage relationship unit for determining the upper node corresponding to each document name, generating a corresponding relationship between the document name and the upper node according to each document name and the upper node corresponding to each document name, and storing the corresponding relationship in a database.

8. The device according to claim 6, wherein, The rule generation unit includes: A training module for training an initial pre-trained language T5 model through a machine translation method based on the document name, the audit elements, and the operator names to obtain a conversion model for generating rule expressions. A generation module, configured to convert, through the conversion model, to obtain a rule expression corresponding to the bill name, the audit element, and the operator name.

9. The device according to any one of claims 6 - 8, wherein, The device further includes: A conversion unit, configured to, in a preset rule library, after obtaining a plurality of rule expressions by using a preset rule engine, convert the rule expressions into rule expressions in a preset language; wherein the rule expressions in the preset language are used to be run by the rule engine.

10. The device according to any one of claims 6 - 8, wherein, The device further includes: A rule engine generation unit, configured to generate, based on standard grammar information and an operator name, a parser corresponding to the rule expression in a preset language, and generate a rule engine for executing the parser; wherein the operator name is used to represent calculation logic and parsing logic, and the standard grammar information is used to represent grammar information about the operator name.

11. An electronic device, wherein, Comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, and when the processor executes the computer program, the method described in any one of claims 1-5 above is implemented.

12. A computer-readable storage medium, wherein, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by the processor, they are used to implement the method described in any one of claims 1-5.

13. A computer program product, wherein, Comprising a computer program, and when the computer program is executed by the processor, the method described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Method and device for generating credit card order auditing inspection key point list

    CN111783432A

  • Credit 46 domain analysis method and device

    CN112991037A