Business data auditing method and related device

Through the data audit method combining deep learning and reinforcement learning, the flexibility and security problems of multi-source heterogeneous data processing are solved, efficient and secure data audit and differential report generation are achieved, and user experience is improved.

CN120336592AInactive Publication Date: 2025-07-18CHINA TOWER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510390366.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate multi-source heterogeneous data, which lacks flexibility and security, resulting in low data audit efficiency, poor accuracy and poor user experience.

Method used

The deep learning model is used to analyze unstructured data, combine adaptive data cleaning algorithms and reinforcement learning algorithms to dynamically adjust the rule base, use homomorphic encryption technology to ensure data security, and generate differences reports through multi-level comparisons through graph databases.

Benefits of technology

Improves the accuracy and efficiency of data processing, enhances the flexibility and security of the system, provides detailed differences reporting to support decision-making, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_6
    Figure SMS_6
Patent Text Reader

Abstract

The invention provides a business data auditing method and a related device. The method comprises the following steps: acquiring various data from a plurality of heterogeneous business systems; analyzing the unstructured data by using a deep learning model, and cleaning the data by using an adaptive algorithm; dynamically adjusting the data analysis rule base based on reinforcement learning; the data security is ensured by adopting a homomorphic encryption technology and a zero-knowledge proof mechanism; constructing an expected result model of the graph database, and performing multi-level comparison; generating a difference report by using a decision tree algorithm and proposing an optimization suggestion; and automatically sending a difference report and providing an API (Application Program Interface) for query. The device comprises a data integration module, an intelligent analysis module, a rule management module, an encryption and security verification module, a multi-level comparison module, a difference report generation module and a notification and response module. According to the method, the data processing efficiency and accuracy are improved, the flexibility and safety of the system are enhanced, and the method is suitable for data auditing in a complex service environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology and relates to a method and related device for auditing business data. Background Art

[0002] With the rapid development of information technology, enterprises and organizations have accumulated a large amount of business data in the daily operation process. These data not only cover key contents such as transaction records, user information, and financial data, but also are of great significance for enterprise decision-making, risk control, and compliance inspection. However, due to the wide range of data sources and diverse forms (including structured, unstructured, and semi-structured data), how to effectively audit these data has become a complex and urgent problem.

[0003] Traditional data auditing methods usually rely on manual review or simple automated scripts, and this method has many limitations. First, in the face of a large amount of data sets, manual review is not only time-consuming and laborious, but also prone to human errors. Second, traditional automated tools often lack flexibility and are difficult to adapt to different types of data formats and ever-changing business requirements. In addition, data security and integrity are also important aspects that cannot be ignored in the data auditing process. Especially when it comes to sensitive information, ensuring the security of data during transmission and processing is crucial.

[0004] In recent years, with the development of artificial intelligence, machine learning, and encryption technology, new ideas and technical means have been provided to solve the above problems. For example, natural language processing (NLP) models based on deep learning can effectively parse unstructured data and extract valuable information from it; reinforcement learning algorithms can dynamically adjust the data parsing rule base according to the historical data analysis results to improve the parsing accuracy; homomorphic encryption technology allows operations on encrypted data without decryption, thus ensuring data security and privacy protection. Nevertheless, the existing technologies still fail to fully integrate these advanced technologies to form a set of efficient and comprehensive business data auditing solutions.

[0005] Although some solutions on the current market perform well in certain specific aspects, such as the application of data cleaning or encryption technology, they often lack systematic integrated design and cannot meet the requirements of multiple links such as multi-source data integration, intelligent parsing and preprocessing, rule dynamic adjustment, multi-level comparison, and difference report generation at the same time. In addition, most of the existing data auditing systems ignore the importance of user experience, resulting in the generated difference reports being difficult to understand and use, which affects the actual application effect. Summary of the Invention

[0006] The objective of the present invention is to provide a method and related device for auditing business data, which can obtain data in real time or at regular intervals from multiple heterogeneous business systems, effectively process the data using advanced parsing techniques and algorithms, and achieve precise comparison by constructing a multi-level expected result model.

[0007] To solve the above technical problems, the present invention provides a method for auditing business data, including the following steps:

[0008] S1. Multi-source data integration step: Obtain business data in real time or at regular intervals from at least two heterogeneous business systems through a distributed message queue. The business data includes structured data D s , unstructured data D u and semi-structured data D h , and conduct preliminary classification on them;

[0009] S2. Intelligent parsing and preprocessing step: Apply a natural language processing model based on deep learning to parse unstructured data, and use an adaptive data cleaning algorithm to preprocess all types of data;

[0010] S3. Rule dynamic adjustment step: Dynamically update the data parsing rule library using a reinforcement learning algorithm according to the analysis results of historical data and current data;

[0011] S4. Encryption and security verification step: Use homomorphic encryption technology to encrypt the parsed data to ensure operations can be performed without decryption, and use a zero-knowledge proof mechanism to verify data integrity;

[0012] S5. Multi-level comparison step: By constructing an expected result model based on a graph database, perform multi-level comparison between the encrypted parsing results and the expected results in the model;

[0013] S6. Difference report generation and optimization step: When differences are detected, use a decision tree algorithm to automatically generate a detailed difference report, and analyze the reasons for the differences in combination with historical data and propose optimization suggestions;

[0014] S7. Notification and response step: Automatically send the difference report to relevant stakeholders through a configurable notification policy, and provide an API interface for third-party systems to query and receive the difference report.

[0015] Further preferably, in step 2, the intelligent parsing and preprocessing step further includes:

[0016] Apply text mining technology to extract key information from unstructured data and convert it into a structured format; The probability distribution model for key information extraction is defined as:

[0017]

[0018] where ω represents the keyword, represents the matching score between the keyword ω and the unstructured data D u and V represents the vocabulary;

[0019] The data cleaning algorithm is defined as:

[0020]

[0021] where represents the cleaned data, represents the regularization term, and λ is the regularization coefficient.

[0022] The present invention also discloses an auditing device for business data, including:

[0023] A data integration module: configured to collect business data from multiple heterogeneous business systems;

[0024] An intelligent parsing module: including a deep learning NLP model and an adaptive data cleaning algorithm, used to parse and preprocess different types of data;

[0025] A rule management module: using a reinforcement learning algorithm to dynamically adjust the data parsing rule library;

[0026] An encryption and security verification module: applying homomorphic encryption technology and zero-knowledge proof mechanism to ensure the security and integrity of data;

[0027] A multi-level comparison module: constructing an expected result model based on a graph database and performing multi-level data comparison;

[0028] A difference report generation module: using a decision tree algorithm to generate a difference report and analyzing the reasons for differences in combination with historical data;

[0029] A notification and response module: configured to send a difference report according to a configurable notification policy and provide an API interface for external systems to call.

[0030] Further preferably, the intelligent parsing module further includes:

[0031] A text mining unit, used to extract key information from unstructured data and convert it into a structured format;

[0032] A data cleaning unit, whose algorithm is defined as:

[0033]

[0034] where represents the cleaned data, represents the regularization term, and λ is the regularization coefficient.

[0035] Further preferably, the rule management module dynamically updates the data parsing rule library using a reinforcement learning algorithm; the objective function of reinforcement learning is expressed as:

[0036]

[0037] where s represents the current state, a represents the action taken, r represents the immediate reward, and r ∈ [0, 1] represents the discount factor.

[0038] Further preferably, the field-level comparison similarity calculation formula in the multi-level comparison module is:

[0039]

[0040] where f1 and f2 respectively represent two fields, ω i represents the weight of the i-th sub-field, and δ(f1, i, f2, i) represents the similarity function between sub-fields.

[0041] Further preferably, the difference report generation module uses a decision tree algorithm to generate a difference report, and the splitting criterion adopts the principle of maximizing information gain, defined as:

[0042]

[0043] where T represents the dataset of the current node, A represents the candidate splitting attribute, and H(T) represents the entropy of the dataset T.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] First, the present invention uses a deep learning model to parse unstructured data and uses an adaptive data cleaning algorithm to process all types of data, significantly improving the accuracy and efficiency of data processing. Compared with traditional methods, it reduces manual intervention and error rates, ensures high-quality data output, and provides a reliable basis for subsequent analysis.

[0046] Second, the present invention uses a reinforcement learning algorithm to dynamically adjust the data parsing rule library, and the system can automatically optimize the rules according to the historical data analysis results to adapt to the changing business needs. This adaptive mechanism enhances the flexibility and intelligence level of the system, reduces manual maintenance work, and improves the overall operation efficiency and accuracy.

[0047] Third, the present invention uses homomorphic encryption technology and zero-knowledge proof mechanism to operate on data and verify its integrity without decrypting, effectively protecting the security of sensitive information. This method not only ensures the security of data during transmission and processing, but also prevents data tampering or loss, and is particularly suitable for scenarios with high security requirements.

[0048] Fourthly, the present invention constructs an expected result model based on a graph database, supports multi-level comparison at the field level, record level, and transaction level, and can comprehensively detect data differences. When differences are found, the system automatically generates a detailed difference report, analyzes the reasons in combination with historical data, provides optimization suggestions, helps users quickly locate problems, and improves decision-making support capabilities. Specific Embodiments

[0049] The following further elaborates on a business data auditing method and related devices proposed by the present invention in conjunction with specific embodiments. According to the following description, the advantages and features of the present invention will be clearer.

[0050] Embodiment 1, a business data auditing method, includes the following steps:

[0051] S1. Multi-source data integration step: Obtain business data in real time or at regular intervals from at least two heterogeneous business systems through a distributed message queue. The business data includes structured data D s , unstructured data D u and semi-structured data D h , and conduct preliminary classification on them.

[0052] S2. Intelligent parsing and preprocessing step: Apply a natural language processing model based on deep learning to parse unstructured data, and use an adaptive data cleaning algorithm to preprocess all types of data.

[0053] The intelligent parsing and preprocessing step further includes:

[0054] Apply text mining technology to extract key information from unstructured data and convert it into a structured format; the probability distribution model for key information extraction is defined as:

[0055]

[0056] where ω represents a keyword, represents the matching score between the keyword ω and the unstructured data D u , and V represents the vocabulary;

[0057] The data cleaning algorithm is defined as:

[0058]

[0059] where represents the data after cleaning, represents the regularization term, and λ is the regularization coefficient.

[0060] S3. Rule dynamic adjustment step: Dynamically update the data parsing rule library using a reinforcement learning algorithm based on the analysis results of historical data and current data.

[0061] S4. Encryption and Security Verification Step: Use homomorphic encryption technology to encrypt the parsed data to ensure operations can be performed without decryption, and use a zero-knowledge proof mechanism to verify data integrity.

[0062] S5. Multi-level Comparison Step: By constructing an expected result model based on a graph database, perform multi-level comparison of the encrypted parsed results with the expected results in the model.

[0063] S6. Difference Report Generation and Optimization Step: When differences are detected, use a decision tree algorithm to automatically generate a detailed difference report, and analyze the reasons for the differences in combination with historical data and propose optimization suggestions.

[0064] S7. Notification and Response Step: Through a configurable notification policy, automatically send the difference report to relevant stakeholders, and provide an API interface for third-party systems to query and receive the difference report.

[0065] Embodiment 2, An auditing device for business data, comprising:

[0066] Data Integration Module: Configured to collect business data from multiple heterogeneous business systems.

[0067] Intelligent Parsing Module: Includes a deep learning NLP model and an adaptive data cleaning algorithm for parsing and preprocessing different types of data.

[0068] The intelligent parsing module further includes:

[0069] Text Mining Unit: Used to extract key information from unstructured data and convert it into a structured format;

[0070] Data Cleaning Unit, whose algorithm is defined as:

[0071]

[0072] where represents the cleaned data, represents the regularization term, and λ is the regularization coefficient.

[0073] Rule Management Module: Uses a reinforcement learning algorithm to dynamically adjust the data parsing rule library.

[0074] The rule management module uses a reinforcement learning algorithm to dynamically update the data parsing rule library; the objective function of reinforcement learning is expressed as:

[0075]

[0076] Where s represents the current state, a represents the action taken, r represents the immediate reward, and r ∈ [0, 1] represents the discount factor.

[0077] Encryption and security verification module: Applies homomorphic encryption technology and zero-knowledge proof mechanism to ensure the security and integrity of data.

[0078] Multi-level comparison module: Builds an expected result model based on a graph database and performs multi-level data comparison.

[0079] The formula for calculating the similarity of field-level comparison in the multi-level comparison module is:

[0080]

[0081] Where f1 and f2 represent two fields respectively, ω i represents the weight of the i-th sub-field, and δ(f1, i, f2, i) represents the similarity function between sub-fields.

[0082] Difference report generation module: Uses a decision tree algorithm to generate a difference report and analyzes the reasons for differences in combination with historical data. The difference report generation module uses a decision tree algorithm to generate a difference report, and the splitting criterion adopts the principle of maximizing information gain, defined as:

[0083]

[0084] Where T represents the data set of the current node, A represents the candidate splitting attribute, and H(T) represents the entropy of the data set T.

[0085] Notification and response module: Configured to send a difference report according to a configurable notification policy and provide an API interface for external systems to call.

[0086] It should also be supplemented and explained that all "settings" and similar descriptive words in this application (especially in the specification) express that there is a connection relationship between two structures, but the specific means of connection between the two are not limited too much, and usually are conventional connection means, that is, it should be understood that this means is the prior art and does not need to be elaborated too much. For example, "n is provided on m" only expresses that there is an n structure on the m structure, and whether the two are connected by welding, riveting, adhesive bonding or integrally formed is within the protection scope of this application; another example is "y is rotatably provided on x", which only expresses that y and x can rotate relative to each other, and as for whether the two are rotatably connected by a bearing, or y directly passes through x and is rotatably connected to x, or other achievable ways, they are all within the protection scope of this application.

[0087] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure fall within the scope of protection of the claims.

Claims

1. A method for auditing service data, characterized in that, It includes the following steps: S1. Multi-source data integration step: Obtain business data in real time or at regular intervals from at least two heterogeneous business systems through a distributed message queue. The business data includes structured data D s , unstructured data D u and semi-structured data D h , and conduct a preliminary classification on them; S2. Intelligent parsing and preprocessing step: Apply a natural language processing model based on deep learning to parse unstructured data, and use an adaptive data cleaning algorithm to preprocess all types of data; S3. Rule dynamic adjustment step: Dynamically update the data parsing rule library using a reinforcement learning algorithm based on the analysis results of historical data and current data; S4. Encryption and security verification step: Use homomorphic encryption technology to encrypt the parsed data to ensure operations without decryption, and use a zero-knowledge proof mechanism to verify data integrity; S5. Multi-level comparison step: By constructing an expected result model based on a graph database, perform multi-level comparison of the encrypted parsing results with the expected results in the model; S6. Difference report generation and optimization step: When differences are detected, use a decision tree algorithm to automatically generate a detailed difference report, and analyze the reasons for the differences in combination with historical data and propose optimization suggestions; S7. Notification and response step: Automatically send the difference report to relevant stakeholders through a configurable notification policy, and provide an API interface for third-party systems to query and receive the difference report.

2. The auditing method for service data according to claim 1, wherein In step 2, the intelligent parsing and preprocessing step further includes: Apply text mining technology to extract key information from unstructured data and convert it into a structured format; The probability distribution model for key information extraction is defined as: where ω represents a keyword, indicating the matching score between the keyword ω and the unstructured data D u and V represents the vocabulary; The data cleaning algorithm is defined as: Among them represents the data after cleaning, represents the regularization term, and λ is the regularization coefficient.

3. An auditing device for service data, based on the auditing method for service data according to any one of claims 1-2, characterized in that, It includes: Data integration module: Configured to collect business data from multiple heterogeneous business systems; Intelligent parsing module: Includes a deep learning NLP model and an adaptive data cleaning algorithm for parsing and preprocessing different types of data; Rule management module: Use a reinforcement learning algorithm to dynamically adjust the data parsing rule library; Encryption and security verification module: Apply homomorphic encryption technology and zero-knowledge proof mechanism to ensure data security and integrity; Multi-level comparison module: Construct an expected result model based on a graph database and perform multi-level data comparison; Difference report generation module: Use a decision tree algorithm to generate a difference report and analyze the reasons for the differences in combination with historical data; Notification and response module: Configured to send difference reports according to a configurable notification policy and provide an API interface for external systems to call.

4. The auditing device for service data according to claim 3, characterized in that, The intelligent parsing module further includes: Text mining unit, used to extract key information from unstructured data and convert it into a structured format; Data cleaning unit, whose algorithm is defined as: Among them represents the data after cleaning, represents the regularization term, and λ is the regularization coefficient.

5. The auditing device for service data according to claim 3, wherein The rule management module uses a reinforcement learning algorithm to dynamically update the data parsing rule library; The objective function of reinforcement learning is expressed as: Where s represents the current state, a represents the action taken, r represents the immediate reward, and r ∈ [0,1] represents the discount factor.

6. The auditing device for service data according to claim 3, wherein The field-level comparison similarity calculation formula in the multi-level comparison module is: where f1 and f2 respectively represent two fields, ω i represents the weight of the i-th sub-field, and δ(f1, i, f2, i) represents the similarity function between sub-fields.

7. The auditing device for service data according to claim 3, wherein The difference report generation module uses a decision tree algorithm to generate a difference report, and the splitting criterion adopts the principle of maximizing information gain, which is defined as: Where T represents the dataset of the current node, A represents the candidate splitting attribute, and H(T) represents the entropy of the dataset T.

Citation Information

Cited By

  • Automatic quality auditing method for heterogeneous data management

    CN120780699A