Data processing method and device based on rule engine, equipment and storage medium
By obtaining and comparing the content data of different versions of financial report files, determining and integrating data format conversion rules to the rule engine, the problem of inaccurate identification results caused by inconsistent financial report files is solved, and a higher document recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510138624.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
AI Technical Summary
During the process of editing and publishing corporate financial reports, the document format is prone to change, resulting in inconsistent formats of financial reports in different versions, and the identification results of existing document recognition services are not accurate enough.
By obtaining the content data of different versions of financial report files, comparing and determining the data format conversion rules, integrating them into the preset rule engine, and data conversion processing is carried out to improve identification accuracy.
When the data format of different versions of documents changes, target conversion rules can be determined and integrated into the rules engine for processing, thereby improving the accuracy of document recognition results.
Smart Images

Figure CN120046596A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a data processing method, apparatus, device, and storage medium based on a rule engine. Background Art
[0002] With the continuous development of information technology and the Internet, the application of using artificial intelligence services to perform document recognition on files is becoming more and more widespread. In the field of enterprise financial report auditing, it is necessary for users to upload financial report files (word, pdf, or pictures), and the uploaded financial report files are audited through a financial report auditing system. However, during the process from the editing to the release of the financial report file, many intermediate documents will be generated, and their document formats often change, resulting in inconsistent formats of different versions of the financial report file. At the same time, the financial report file also involves different time lengths or different time descriptions, such as 2023 - 06, the first half of 2023, or as of the middle of this year, etc. Therefore, the recognition results of the current document recognition service are not accurate enough. Summary of the Invention
[0003] The main purpose of this application is to provide a data processing method, apparatus, device, and storage medium based on a rule engine, aiming to improve the accuracy of document recognition results by updating the data conversion rules in the preset rule engine.
[0004] In the first aspect, this application provides a data processing method based on a rule engine, including:
[0005] Obtain the first content data of the first document and obtain the second content data of the second document, where the first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different;
[0006] Compare the first content data with the second content data to obtain the content difference data between the first document and the second document;
[0007] According to the content difference data, determine the target conversion rule for converting the data format of the first content data to the data format of the second content data;
[0008] After integrating the target conversion rule into the preset rule engine, use the preset rule engine to perform data conversion processing on the second content data of the second document.
[0009] In the second aspect, this application also provides a data processing apparatus, and the data processing apparatus includes:
[0010] A data acquisition module, configured to acquire first content data of a first document and second content data of a second document, where the first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different;
[0011] A data comparison module, configured to compare the first content data with the second content data to obtain content difference data between the first document and the second document;
[0012] A data determination module, configured to determine a target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data;
[0013] A data conversion module, configured to integrate the target conversion rule into a preset rule engine, and then use the preset rule engine to perform data conversion processing on the second content data of the second document.
[0014] In a third aspect, the present application further provides a computer device, which includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the data processing method described above are implemented.
[0015] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data processing method described above are implemented.
[0016] The present application provides a data processing method, apparatus, device, and storage medium based on a rule engine. The present application acquires first content data of a first document and second content data of a second document, where the first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different; compares the first content data with the second content data to obtain content difference data between the first document and the second document; determines a target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data; integrates the target conversion rule into a preset rule engine, and then uses the preset rule engine to perform data conversion processing on the second content data of the second document. The embodiments of the present application can determine the target conversion rule between different changed data formats when the data formats of different version documents change, integrate the target conversion rule into a preset rule engine, and use the preset rule engine to perform data conversion processing, so as to effectively identify the second content data of the second document, and thus can improve the recognition rate of the second content data of the second document, thereby effectively improving the accuracy of the document recognition result. Brief Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of the steps of a data processing method based on a rule engine provided by an embodiment of the present application;
[0019] Figure 2 For Figure 1 It is a schematic sub-step flowchart of the data processing method based on the rule engine in
[0020] Figure 3 It is a schematic diagram of a scenario for implementing the data processing method based on the rule engine provided by this embodiment;
[0021] Figure 4 It is a schematic block diagram of a data processing device provided by an embodiment of the present application;
[0022] Figure 5 For Figure 4 It is a schematic block diagram of a sub-module of the data processing device in
[0023] Figure 6 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application.
[0024] The realization of the purpose of the present application, functional features and advantages will be further described with reference to the embodiments and the drawings. Detailed Description of the Embodiments
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0026] The flowchart shown in the drawings is only an example, and does not necessarily include all the content and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0027] To improve the efficiency of enterprises in editing corporate financial reports and ensure the accuracy of the indicators involved in the financial reports, a financial report intelligent review system has been developed. Its design is as follows: The user uploads a financial report file (word, pdf, or picture), and the financial report intelligent review system calls a front-end model (such as an AI model service) to perform content recognition on the financial report file.
[0028] For the review of a certain corporate financial report, the financial report intelligent review system will call back multiple AI model services, such as OCR services, document review services, and document extraction index services, etc. The back-end program needs to integrate the data of multiple AI model services. From the editing to the final release of the financial report, many intermediate manuscripts will be generated, and their formats may not necessarily be in a consistent document format. The financial report also involves different quarters. For example, there are various text descriptions regarding time (such as 2023-06, the first half of 2023, or as of the middle of this year, etc.). The AI model service is not accurate enough in processing documents in new formats, and it is prone to false alarms in document recognition, such as false alarms in time, false alarms in amount formats, and false alarms in text information ambiguity, etc. Therefore, how to improve the accuracy of document recognition results has become an urgent problem to be solved.
[0029] To solve the above problems, the existing design is that business personnel annotate the displayed results, and developers summarize the rules and encode filtering rules at the code level. However, when annotating the results of large files, it will cause problems such as a large amount of manual annotation work and easy omission.
[0030] Based on this, the embodiments of the present application provide a data processing method, device, equipment, and storage medium based on a rule engine, which can improve the accuracy of document recognition results through the integration of data conversion rules. At the same time, it solves the problems mentioned above, such as the long problem-solving cycle when new data appears, the large amount of manual annotation work for large documents, and easy omission.
[0031] Among them, this data processing method can be applied to a terminal device or a server. The terminal device can be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, and a wearable device, etc.; the server can be a single server or a server cluster composed of multiple servers. The following takes the application of this data processing method to a server as an example for explanation.
[0032] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0033] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of the steps of a data processing method based on a rule engine provided by an embodiment of the present application.
[0034] As shown Figure 1 below, the data processing method includes steps S101 to S104.
[0035] Step S101: Obtain the first content data of the first document and obtain the second content data of the second document.
[0036] Among them, the first document and the second document are different versions of the same business document. For example, the first document can be the prior version of a certain business document, and the second document can be the subsequent version of a certain business document. Business documents can include documents or files in different business fields. Specifically, business fields can include the financial field, the medical field, the smart city field, the technology field, etc., and can also include fields such as banks, insurance, hospitals, design institutes, the Internet, biotechnology, etc. For example, the first document and the second document can be two different versions among multiple versions involved in the process from the editing to the final release of an enterprise financial report.
[0037] Among them, the first content data can be the content data in the first document, the second content data can be the content data in the second document, and the data formats of the first content data and the second content data are different. Specifically, the first content data and the second content data can be obtained by respectively performing content recognition on the first document and the second document based on a document recognition model.
[0038] In one embodiment, obtain the first content data of the first document and obtain the second content data of the second document, where the data formats of the first content data and the second content data are different. The difference in data format can refer to different data types and / or different data structures. Among them, data types can include multiple data types in the word, pdf, Excel, and picture formats, and data structures can include data composition structures under different data types, such as the number of rows and columns of a table, a time description structure, etc.
[0039] In one embodiment, when receiving new format data returned by a document recognition service model, obtain the first content data of the first document and obtain the second content data of the second document, thereby starting the subsequent rule update process. That is to say, when receiving new format data, it indicates that there will be a false alarm problem in document recognition, and the data processing method provided by the embodiments of the present application can effectively improve the accuracy of document recognition results.
[0040] For ease of understanding, the following uses the first document and the second document as examples of corporate financial reports for illustration. In corporate financial reports, for the first document of the previous version, the system will call back multiple AI model services to review the first document and use a backend program to integrate the output data of multiple AI model services. However, for the second document of the later version, the data format of some of its data will change compared to the first document. After the system calls back multiple AI model services to review the second document, its output data cannot be recognized and integrated by the backend program, so there will be a false alarm problem in document recognition. Therefore, new format data will be received from the document recognition service model.
[0041] In one embodiment, obtaining the first content data of the first document and obtaining the second content data of the second document includes: obtaining the first document and the second document, where the first document and the second document are different versions of the same financial document; calling a preset financial document recognition model to perform content recognition on the first document and the second document respectively to obtain the first content data of the first document and the second content data of the second document. Among them, the financial document recognition model is used to recognize different versions of the financial document to obtain the content data of the financial document. Through the financial document recognition model, different version documents of the financial document can be recognized quickly and accurately.
[0042] It should be noted that to further ensure the privacy and security of the above-mentioned relevant information such as the first document and the second document, the above-mentioned relevant information such as the first document and the second document can also be stored in a node of a blockchain. The technical solution of the present application can also be applied to adding other data files stored on the blockchain. The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.
[0043] Step S102: Compare the first content data with the second content data to obtain the content difference data between the first document and the second document.
[0044] Among them, the method of comparing the first content data with the second content data can be implemented based on methods such as text comparison based on word vectors, text classification based on machine learning, and cosine similarity. The embodiments of the present application do not make specific explanations for this.
[0045] It should be noted that since the data formats of the first content data and the second content data are different, there must be content difference data between the first document and the second document. If the data formats of the first content data and the second content data are the same, the content difference data between the first document and the second document is an empty set.
[0046] Exemplarily, if a certain segment of the first content data is text and the corresponding segment of the second content data is a picture, the content difference data between the first document and the second document includes the text of the first content data and the picture of the second content data. If the format of a certain date and time in the first content data is year-month-day and the format of the corresponding date and time in the second content data is month-day-year, the content difference data between the first document and the second document includes the date and time of the first content data and the date and time of the second content data. If a certain table of date and time in the first content data has 5 columns and the corresponding table in the second content data has 7 columns, and the first 5 columns of data in the 7-column table are the same as the 5 columns of the first content data, the content difference data between the first document and the second document includes the last 2 columns of data in the second content data.
[0047] Step S103: Determine a target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data.
[0048] Among them, the content difference data may include the first difference data in the first document and / or the second difference data in the second document. In order for the backend program to recognize the new format data in the second content data, it is necessary to determine a target conversion rule for converting the data format of the first content data into the data format of the second content data.
[0049] In one embodiment, as Figure 2 shown, step S103 includes: sub-steps S1031 to S1032.
[0050] Sub-step S1031: Identify the data format of the first difference data in the document difference data, and the first difference data corresponds to the first document.
[0051] Among them, the first difference data is part of the first content data in the first document. Therefore, the first difference data corresponds to the first document and the first content data. The data format of the first difference data can be identified by a data adapter, which can identify various data types and data structures, including identifying newly emerged data formats. For example, the data format of the first difference data can be identified as a format that can be understood by a preset rule engine through the data adapter.
[0052] Sub-step S1032: Identify the data format of the second difference data in the document difference data, and the second difference data corresponds to the second document.
[0053] Among them, the second difference data is a part of the second content data in the second document. Therefore, the second difference data corresponds to the second document and the second content data. The data format of the second difference data can also be recognized by the data adapter, so as to recognize the newly emerged data format. The data format of the second difference data and the data format of the first difference data can be recognized simultaneously.
[0054] Sub-step S1033: Determine the target conversion rule according to the data format of the first difference data and the data format of the second difference data.
[0055] Among them, the target conversion rule can be determined according to the rule management function of the preset rule engine, so that rules can be dynamically added or modified to adapt to the new data format. Of course, the target conversion rule can also be understood as the loading rule of an external source (such as a database, a middleware file service system, and a configuration file) to realize the dynamic update of the conversion rule.
[0056] In one embodiment, there are multiple first difference data and multiple second difference data; determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data includes: determining the matching relationship between the data format of each first difference data and the data format of each second difference data according to the correspondence between the multiple first difference data and the multiple second difference data; determining multiple target conversion rules according to the matching data formats of the first difference data and the second difference data.
[0057] Among them, the multiple first difference data and the multiple second difference data may be in one-to-one correspondence, or may be partially corresponding and partially non-corresponding. Therefore, there is a mutual correspondence between the multiple first difference data and the multiple second difference data. Furthermore, according to the correspondence between the multiple first difference data and the multiple second difference data, the matching relationship between the data format of each first difference data and the data format of each second difference data can be determined, so that the target conversion rule for converting the data format of the first content data into the data format of the second content data can be accurately determined, and thus the accuracy of obtaining multiple target conversion rules can be improved.
[0058] In one embodiment, the data format of the first difference data is the first data structure, and the data format of the second difference data is the second data structure; determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data includes: analyzing the data structure relationship between the first data structure of the first difference data and the second data structure of the second difference data; determining the data conversion rule corresponding to the data structure relationship from the preset data conversion rule library as the target conversion rule.
[0059] Among them, the first data structure is different from the second data structure. For example, the data composition structures of the first data structure and the second data structure are different. For example, the first data structure is a 5-column table structure, and the second data structure is a 7-column table structure. It should be noted that the data conversion rule library can include the data structure relationships between multiple data structures. By parsing the data structure relationship between the first data structure of the first differential data and the second data structure of the second differential data, the corresponding data conversion rule can be accurately determined from the data conversion rule library as the target conversion rule. Therefore, the accuracy and convenience of obtaining the target conversion rule can be improved.
[0060] In one embodiment, the data format of the first differential data is the first data type, and the data format of the second differential data is the second data type; determining the target conversion rule according to the data format of the first differential data and the data format of the second differential data includes: parsing the data type relationship between the first data type of the first differential data and the second data type of the second differential data; determining the data conversion rule corresponding to the data type relationship from the preset data conversion rule library as the target conversion rule.
[0061] Among them, the first data type is different from the second data type. For example, the first data type is text data, and the second data type is picture data. It should be noted that the data conversion rule library can include the data type relationships between multiple data types. By parsing the data type relationship between the first data type of the first differential data and the second data type of the second differential data, the corresponding data conversion rule can be accurately determined from the data conversion rule library as the target conversion rule. Therefore, the accuracy and convenience of obtaining the target conversion rule can be improved.
[0062] Step S104: After integrating the target conversion rule into the preset rule engine, use the preset rule engine to perform data conversion processing on the second content data of the second document.
[0063] Among them, integrating the target conversion rule into the preset rule engine can adopt the hot update method. Integrating the target conversion rule into the preset rule engine to update the preset rule engine, so that the updated preset rule engine can be used to perform data conversion processing on the second content data of the second document, which can solve the problem that the backend program cannot recognize the second content data of the second document, and thus can recognize the new format data in the second content data, improve the recognition rate of the second content data of the second document, and thus effectively improve the accuracy of the document recognition result.
[0064] In one embodiment, the preset rule engine includes a Drools rule engine, and the target transformation rules include data transformation rules that can be understood by the Drools rule engine. It should be noted that a rule model can be defined in the Drools rule engine, and this rule model can adapt to and contain all possible data returned by the front-end model. The Drools rule engine is used to adapt to the new data format returned by the front-end model and maintain the hot update of the rules to ensure the flexibility and responsiveness of the entire system. At the same time, it is necessary to ensure that the Drools rule engine maintains a high degree of automation and low latency so that it can quickly adapt to data changes.
[0065] In one embodiment, the preset rule engine is used to perform data transformation processing on the second content data of the second document, including: performing data transformation processing on the second content data of the second document through the complex event processing (CEP) component of the Drools rule engine, so as to implement the rule execution for real-time data streams, which can effectively improve the accuracy of the document recognition result.
[0066] In one embodiment, the method further includes: using the Drools rule engine to formulate identification rules to monitor the data that needs to be identified and recognized. Among them, the data to be identified can include data with new formats such as the second content data of the second document, thereby reducing the need for manual identification and marking of each item in the entire document.
[0067] In one embodiment, a feedback loop mechanism can also be established to monitor the execution effect of the preset rule engine on the data transformation processing of the second content data of the second document. If it is found that the target transformation rules are not applicable to the new data format or the labeled data, the target transformation rules can be automatically adjusted, or the developer can be notified to manually adjust the target transformation rules for the new data and push the labeled data to the front-end model.
[0068] Please refer to Figure 3 , Figure 3 for a schematic diagram of a scenario for implementing the data processing method of this embodiment.
[0069] As Figure 3As shown in the figure, the front-end model 10 performs content recognition on the first document and the second document to obtain the first content data of the first document and the second content data of the second document. The first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different. The server 20 compares the first content data with the second content data to obtain the content difference data between the first document and the second document. The server 20 determines the target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data. After integrating the target conversion rule into the preset rule engine, the server 20 uses the preset rule engine to perform data conversion processing on the second content data of the second document, and sends the second content data after the data conversion processing to the back-end program 30 for the user to review and confirm.
[0070] The data processing method based on the rule engine provided in the above embodiment obtains the first content data of the first document and the second content data of the second document by obtaining the first content data of the first document and the second content data of the second document. The first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different. The first content data and the second content data are compared to obtain the content difference data between the first document and the second document. According to the content difference data, the target conversion rule for converting the data format of the first content data into the data format of the second content data is determined. After integrating the target conversion rule into the preset rule engine, the preset rule engine is used to perform data conversion processing on the second content data of the second document. The embodiment of the present application can determine the target conversion rule between different changed data formats when the data formats of different version documents change, integrate the target conversion rule into the preset rule engine, and use the preset rule engine to perform data conversion processing, so as to effectively identify the second content data of the second document, and thus improve the recognition rate of the second content data of the second document, thereby effectively improving the accuracy of the document recognition result.
[0071] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a data processing device provided by an embodiment of the present application.
[0072] As Figure 4 shown, the data processing device 200 includes:
[0073] A data acquisition module 210, configured to acquire the first content data of the first document and the second content data of the second document, where the first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different;
[0074] A data comparison module 220, configured to compare the first content data with the second content data to obtain content difference data between the first document and the second document;
[0075] A data determination module 230, configured to determine a target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data;
[0076] A data conversion module 240, configured to integrate the target conversion rule into a preset rule engine, and then use the preset rule engine to perform data conversion processing on the second content data of the second document.
[0077] In one embodiment, as Figure 5 shown, the data determination module 230 includes:
[0078] A first identification sub-module 2031, configured to identify the data format of the first difference data in the document difference data, where the first difference data corresponds to the first document.
[0079] A second identification sub-module 2032, configured to identify the data format of the second difference data in the document difference data, where the second difference data corresponds to the second document.
[0080] A determination sub-module 2033, configured to determine the target conversion rule according to the data format of the first difference data and the data format of the second difference data.
[0081] In one embodiment, there are multiple first difference data and second difference data; the data determination module 230 is further configured to:
[0082] Determine a matching relationship between the data formats of each first difference data and each second difference data according to the corresponding relationship between the multiple first difference data and the multiple second difference data;
[0083] Determine multiple target conversion rules according to the data formats of the first difference data and the second difference data that match each other.
[0084] In one embodiment, the data format of the first difference data is a first data structure, and the data format of the second difference data is a second data structure; the data determination module 230 is further configured to:
[0085] Analyze the data structure relationship between the first data structure of the first difference data and the second data structure of the second difference data;
[0086] Determine, from a preset data conversion rule library, a data conversion rule corresponding to the data structure relationship as the target conversion rule.
[0087] In one embodiment, the data format of the first difference data is a first data type, and the data format of the second difference data is a second data type; the data determination module 230 is further configured to:
[0088] Analyze the data type relationship between the first data type of the first difference data and the second data type of the second difference data;
[0089] Determine, from a preset data conversion rule library, a data conversion rule corresponding to the data type relationship as the target conversion rule.
[0090] In one embodiment, analyze the data type relationship between the first data type of the first difference data and the second data type of the second difference data;
[0091] Determine, from a preset data conversion rule library, a data conversion rule corresponding to the data type relationship as the target conversion rule.
[0092] In one embodiment, the data acquisition module 210 is further configured to:
[0093] Acquire a first document and a second document, where the first document and the second document are different versions of the same financial document;
[0094] Call a preset financial document recognition model to respectively perform content recognition on the first document and the second document, and obtain first content data of the first document and second content data of the second document.
[0095] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing data processing method embodiments, and will not be elaborated herein.
[0096] The device provided in the above embodiment can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 6 shown.
[0097] Please refer to Figure 6 , Figure 6 , which is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application.
[0098] As shown in Figure 6As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory may include a storage medium and an internal memory. The storage medium may be non-volatile or volatile.
[0099] The storage medium can store an operating system and computer programs. The computer programs include program instructions that, when executed, can cause the processor to execute any one of the data processing methods.
[0100] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0101] The internal memory provides an environment for the operation of the computer programs in the storage medium. When the computer programs are executed by the processor, the processor can be caused to execute any one of the data processing methods.
[0102] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0103] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0104] Among them, in one embodiment, the processor is used to run the computer programs stored in the memory to implement the following steps:
[0105] Obtain the first content data of the first document and obtain the second content data of the second document, where the first document and the second document are different versions of the same business document, and the data formats of the first content data and the second content data are different;
[0106] Compare the first content data with the second content data to obtain content difference data between the first document and the second document;
[0107] Determine a target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data;
[0108] After integrating the target conversion rule into a preset rule engine, use the preset rule engine to perform data conversion processing on the second content data of the second document.
[0109] In one embodiment, when the processor implements determining the target conversion rule for converting the data format of the first content data into the data format of the second content data according to the content difference data, it is used to implement:
[0110] Identify the data format of the first difference data in the document difference data, where the first difference data corresponds to the first document; and
[0111] Identify the data format of the second difference data in the document difference data, where the second difference data corresponds to the second document;
[0112] Determine the target conversion rule according to the data format of the first difference data and the data format of the second difference data.
[0113] In one embodiment, there are multiple first difference data and second difference data; when the processor implements determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data, it is used to implement:
[0114] Determine the matching relationship between the data formats of each first difference data and each second difference data according to the corresponding relationship between multiple first difference data and multiple second difference data;
[0115] Determine multiple target conversion rules according to the matching data formats of the first difference data and the second difference data.
[0116] In one embodiment, the data format of the first difference data is a first data structure, and the data format of the second difference data is a second data structure; when the processor implements determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data, it is used to implement:
[0117] Analyze the data structure relationship between the first data structure of the first difference data and the second data structure of the second difference data;
[0118] Determine, from a preset data conversion rule library, a data conversion rule corresponding to the data structure relationship as the target conversion rule.
[0119] In one embodiment, the data format of the first difference data is a first data type, and the data format of the second difference data is a second data type; when the processor implements determining the target conversion rule according to the data formats of the first difference data and the second difference data, it is used to implement:
[0120] Parse the data type relationship between the first data type of the first difference data and the second data type of the second difference data;
[0121] Determine, from a preset data conversion rule library, a data conversion rule corresponding to the data type relationship as the target conversion rule.
[0122] In one embodiment, the preset rule engine includes a Drools rule engine, and the target conversion rule includes a data conversion rule that the Drools rule engine can understand.
[0123] In one embodiment, when the processor implements obtaining the first content data of the first document and obtaining the second content data of the second document, it is used to implement:
[0124] Obtain a first document and a second document, where the first document and the second document are different versions of the same financial document;
[0125] Call a preset financial document recognition model to perform content recognition on the first document and the second document respectively, to obtain the first content data of the first document and the second content data of the second document.
[0126] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described computer device can refer to the corresponding process in the foregoing data processing method embodiment, and will not be elaborated herein.
[0127] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0128] An embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed can refer to the various embodiments of the data processing method of this application.
[0129] Among them, the computer-readable storage medium can be the internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.
[0130] Furthermore, the computer-usable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of blockchain nodes. The blockchain referred to in this application is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0131] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0132] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or system comprising that element.
[0133] The serial numbers of the embodiments of this application above are only for description and do not represent the superiority or inferiority of the embodiments. The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data processing method based on a rule engine, characterized in that: include: Acquire first content data of a first document, and acquire second content data of a second document, wherein the first document and the second document are different versions of the same business document, and the first content data and the second content data have different data formats; Comparing the first content data with the second content data to obtain content difference data between the first document and the second document; Determining, based on the content difference data, a target conversion rule for converting a data format of the first content data into a data format of the second content data; After the target conversion rule is integrated into a preset rule engine, the preset rule engine is used to perform data conversion processing on the second content data of the second document.
2. The data processing method according to claim 1, characterized in that: The determining, based on the content difference data, a target conversion rule for converting the data format of the first content data into the data format of the second content data comprises: identifying a data format of first difference data in the document difference data, the first difference data corresponding to the first document; and identifying a data format of second difference data in the document difference data, the second difference data corresponding to the second document; The target conversion rule is determined according to a data format of the first difference data and a data format of the second difference data.
3. The data processing method according to claim 2, characterized in that: There are a plurality of the first difference data and the second difference data; The determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data includes: Determine, according to the correspondence between the plurality of first difference data and the plurality of second difference data, a matching relationship between a data format of each of the first difference data and a data format of each of the second difference data; A plurality of target conversion rules are determined according to the matching data format of the first difference data and the data format of the second difference data.
4. The data processing method according to claim 2, characterized in that: The data format of the first difference data is a first data structure, and the data format of the second difference data is a second data structure; The determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data includes: parsing a data structure relationship between a first data structure of the first difference data and a second data structure of the second difference data; A data transformation rule corresponding to the data structure relationship is determined from a preset data transformation rule library as the target transformation rule.
5. The data processing method according to claim 2, characterized in that: The data format of the first difference data is a first data type, and the data format of the second difference data is a second data type; The determining the target conversion rule according to the data format of the first difference data and the data format of the second difference data includes: parsing a data type relationship between a first data type of the first difference data and a second data type of the second difference data; A data conversion rule corresponding to the data type relationship is determined from a preset data conversion rule library as the target conversion rule.
6. The data processing method according to any one of claims 1 to 5, characterized in that: The preset rule engine includes a Drools rule engine, and the target conversion rule includes a data conversion rule that can be understood by the Drools rule engine.
7. The data processing method according to any one of claims 1 to 5, characterized in that: The step of acquiring the first content data of the first document and acquiring the second content data of the second document includes: Acquire a first document and a second document, where the first document and the second document are different versions of the same financial document; A preset financial document recognition model is called to perform content recognition on the first document and the second document respectively to obtain first content data of the first document and second content data of the second document.
8. A data processing device, characterized in that: The data processing device comprises: A data acquisition module, configured to acquire first content data of a first document and acquire second content data of a second document, wherein the first document and the second document are different versions of the same business document, and the first content data and the second content data have different data formats; A data comparison module, used to compare the first content data with the second content data to obtain content difference data between the first document and the second document; A data determination module, configured to determine a target conversion rule for converting a data format of the first content data into a data format of the second content data according to the content difference data; The data conversion module is used to integrate the target conversion rule into a preset rule engine, and then use the preset rule engine to perform data conversion processing on the second content data of the second document.
9. A computer device, characterized in that: The computer device comprises a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the data processing method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.