Data auditing method, system and equipment
By identifying data types and adopting targeted audit strategies and streaming field processing, the existing audit methods are solved, and efficient and accurate data audits and automatic repairs are achieved to adapt to a variety of business scenarios.
Patent Information
- Application Number
- CN202510358266.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing audit methods are inefficient and poorly accurate in multi-data sources and large-scale data flows, and cannot complete audit tasks in a timely and accurate manner, and cannot meet the needs of complex audit tasks.
By obtaining the data type of the data source, analyzing metadata to determine the index, adopt different audit strategies for relational and analytical data, use streaming fields to audit, generate audit reports, and support customized audit rules and repair rules.
It significantly reduces the database performance burden and memory consumption, improves audit quality and efficiency, ensures that tasks are completed on time and accurately, supports multiple audit modes to adapt to different business scenarios, and provides detailed audit reports and automatic repair functions.
Smart Images

Figure CN120296369A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a data auditing method, system and device. Background Art
[0002] With the wide application of technologies such as big data, cloud computing, and artificial intelligence, data has become the core resource for the development of enterprises and society. At the same time, the characteristics of real-time, concurrency, and massiveness in data processing are becoming increasingly prominent, and ensuring data consistency, integrity, and accuracy faces unprecedented challenges. Once data inconsistency occurs, it may not only lead to business logic errors and decision-making mistakes, but also trigger serious compliance risks and user trust crises. Therefore, data auditing is the primary means to avoid this risk and plays a crucial role in ensuring data accuracy, integrity, and timeliness.
[0003] However, when the existing auditing methods are applied to multi-data sources and large-scale data streams, it is often difficult to complete the auditing tasks in a timely and accurate manner due to the excessive size of the data stream, resulting in difficulties in timely discovery and handling of data quality problems and inability to meet the requirements of complex auditing tasks. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, this application discloses a data auditing method, system and device for solving the problems of low efficiency and poor accuracy of the existing data auditing methods.
[0005] The first aspect of this application discloses a data auditing method, including: obtaining the data source of the data to be audited, determining the data type of the data source, where the data type includes relational data and analytical data; parsing the metadata in the data source to determine the index representing the unique identifier of the metadata; if the data source is relational data, auditing the data source based on the index to determine the target storage location of the target data; auditing the target data at the target storage location according to the preset streaming fields to determine the first auditing result; if the data source is analytical data, auditing the quantity of the data source according to the index to determine the second auditing result; generating an auditing report based on the first auditing result or / and the second auditing result to complete data auditing.
[0006] In some embodiments of the first aspect of this application, before obtaining the data source, it further includes: setting the configuration parameters of the data source to determine the auditing table of the data source, where the auditing table includes whether to audit, the auditing time period, and the auditing mode, and the auditing mode corresponds one-to-one with the data type.
[0007] In some embodiments of the first aspect of the present application, before auditing the data source based on the index, it further includes: if the data source is configured to be audited and the auditing time period exceeds a preset period, the data source is segmented according to the time conditions configured for the auditing time period and split into multiple data packets; each data packet includes the data to be audited and the target data corresponding to each data to be audited, and an auditing subtask corresponding to the data packet is generated;
[0008] Each auditing subtask corresponding to each data packet is sequentially executed, and the auditing subtask is used as the auditing task of the data source to generate an auditing instance.
[0009] In some embodiments of the first aspect of the present application, sequentially executing each auditing subtask corresponding to each data packet further includes: if it is detected that the auditing subtask fails in auditing, an error log is generated, the current auditing task is blocked from continuing to execute and an error is reported; if it is detected that the auditing subtask succeeds in auditing, an auditing log is generated, and the next auditing subtask is executed until the auditing subtask corresponding to the preset data packet is completed to generate the first auditing result.
[0010] In some embodiments of the first aspect of the present application, sequentially executing each auditing subtask corresponding to each data packet includes: allocating the auditing subtasks corresponding to the split multiple data packets to a preset number of processing nodes, so that each processing node audits each data packet through the index to determine the target storage location of each data packet; after determining the target storage location, each processing node audits and compares each data packet according to a preset streaming field to obtain the auditing result of each data packet; the auditing results of each data packet are summarized to determine the first auditing result to which the data source belongs.
[0011] In some embodiments of the first aspect of the present application, the preset streaming field includes at least one of the following: a parameter for defining auditing information, an auditing service class for configuring an entry, an implementation management class for configuring the mapping relationship between different data sources and the corresponding auditing implementation classes of subtasks, an interface for defining an auditing subtask, an auditing subtask implementation class for passing streaming processing parameters, a data source parsing tool class, a task splitting tool class, and a single auditing return result.
[0012] In some embodiments of the first aspect of the present application, an audit report is generated based on the first audit result or / and the second audit result, including: if it is detected based on the audit result that the data source is inconsistent with the preset data, a preset data audit rule is called to analyze the first audit result or the second audit result to obtain an analysis result; feature extraction is performed on the analysis result to obtain abnormal data features and key information features; semantic understanding and information extraction are performed on the abnormal data features and the key information features based on natural language processing technology to obtain each of the index parameters; each of the index parameters is input into a preset text generation model to determine the audit report, and each of the index parameters includes a consistency index, an integrity index, a uniqueness index, and an accuracy index.
[0013] In some embodiments of the first aspect of the present application, if it is determined that there are abnormal index parameters in the audit report, a preset repair rule is called to repair the abnormal index parameters to obtain a correction result, and the correction result is used as a supplementary solution for the audit report.
[0014] The second aspect of the present application discloses a data audit system, including: an acquisition module for acquiring the data source of the data to be audited and determining the data type of the data source, where the data type includes relational data and analytical data; a parsing module for parsing the metadata in the data source to determine an index representing the unique identifier of the metadata; a data audit module for, if the data source is relational data, auditing the data source based on the index to determine the target storage location of the target data, and auditing the target data at the target storage location according to preset streaming fields to determine a first audit result; if the data source is analytical data, auditing the quantity of the data source according to the index to determine a second audit result; a report generation module for generating an audit report based on the first audit result or / and the second audit result to complete data audit.
[0015] The third aspect of the present application discloses an electronic device, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the above method.
[0016] Advantages of the present application: Obtain a data source, and determine the data type of the data source; Parse the metadata in the data source to determine an index representing the unique identifier of the metadata; If the data source is relational data, audit the data source based on the index to determine the target storage location of the target data; Audit the target data at the target storage location according to a preset streaming field to determine a first audit result; In this way, the performance burden and memory consumption of the database are significantly reduced, and the database performance overhead and memory overhead are reduced; If the data source is analytical data, perform a quantity audit on the data source according to the index to determine a second audit result; On the one hand, different audit strategies are adopted according to different data types, improving the audit quality and audit performance; On the other hand, real-time big data is split, improving the processing efficiency and accuracy, and making full use of system resources to ensure that tasks can be completed on time and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an application environment diagram of a method for implementing data auditing in an embodiment of the present application;
[0018] Figure 2 is a flowchart of a method for data auditing in an embodiment of the present application;
[0019] Figure 3 is a complete flowchart of a method for data auditing in an embodiment of the present application;
[0020] Figure 4 is a structural diagram of a data auditing system in an embodiment of the present application;
[0021] Figure 5 is a structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and sub-samples in the embodiments can be combined with each other.
[0023] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the actual components in implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0024] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0025] In the description of the embodiments of the present disclosure, terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0026] Unless otherwise specified, the term "plurality" means two or more.
[0027] In the embodiments of the present disclosure, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0028] The term "and / or" is an associative relationship describing objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B these three relationships.
[0029] Combined with Figure 1 As shown, it is an implementation environment diagram of a data auditing method provided by the embodiments of the present disclosure, including a user side and a server side. Among them, the user side and the server side are network-connected and communicate with each other. In this application, the user side is the front end and the server side is the back end.
[0030] It should be understood that the data auditing method is applied to the user side or / and the server side. In some embodiments, the user side can be at least one of a computer device, a desktop device, a smart phone, and a tablet device; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms; the software can be an application program of the data auditing method, etc., but is not limited to the above forms.
[0031] Combined with Figure 2 As shown, the embodiments of the present disclosure provide a flowchart of a data auditing method, including:
[0032] Step S201: Obtain the data source of the data to be audited, and determine the data type of the data source. The data type includes relational data and analytical data;
[0033] Among them, the data source refers to any system or location that provides data, which is the source or origin of the data. Based on the connection protocol and metadata structure of the data source, the data type is identified. This data type includes relational data and analytical data. For example, relational data usually has a table structure and supports SQL queries; while analytical data exists in the form of large-scale distributed storage and is suitable for batch processing and complex analysis.
[0034] Exemplarily, connect a data type identification system or identification tool to the data source, use methods such as database connection strings, API calls, or file paths for connection, and determine whether the type of the data source is relational data (such as MySQL, Oracle) or analytical data (such as Hadoop, Spark) by querying the metadata of the data source or using specific identification algorithms.
[0035] Step S202: Parse the metadata in the data source to determine the index representing the unique identifier of the metadata;
[0036] Among them, use the metadata query function of the database system or data analysis tools to identify and extract the unique identifier, which involves SQL queries, data dictionary queries, or API calls, parse the metadata of the data source, and extract the fields (such as primary keys, universally unique identifiers, etc.) that can uniquely identify each record or data set as the index.
[0037] For example, start the data source and SQL parsing, automatically select different audit executors according to the type of the data source, perform table metadata queries, analyze the primary keys / unique indexes existing in the table, and obtain the primary key / unique index fields as the index representing the unique identifier of the metadata.
[0038] Step S203: If the data source is relational data, audit the data source based on the index to determine the target storage location of the target data; audit the target data at the target storage location according to the preset streaming fields to determine the first audit result;
[0039] Among them, for the target storage location of the located target data, according to the preset streaming field audit, the target data at the target storage location is checked one by one according to the preset streaming fields to check the correctness and integrity of the data; for example, the preset streaming fields include at least one of the following: the input parameters for defining audit information, the audit service class for configuring the entry, the implementation management class for configuring the mapping relationship between different data sources and the corresponding sub-task audit implementation classes, the interface for defining audit sub-tasks, the audit sub-task implementation class for passing streaming processing parameters, the data source parsing tool class, the task splitting tool class, and the single audit return result.
[0040] Exemplarily, Audit DTO, the base class for audit input parameters, defines the information required for audit; AuditService, the audit service class, the entry class for audit; Base Audit, the interface definition for sub-task audit; AuditConnect Instance Manager, the audit implementation management class, configures the mapping between different data sources and the corresponding sub-task audit implementation classes; MySQL Audit, the sub-task audit implementation class for MySQL, passes streaming processing parameters; Oracle Audit, the sub-task audit implementation class for Oracle, passes streaming processing parameters; Doris Audit, the sub-task audit implementation class for Doris; Maxcompute Audit, the sub-task audit implementation class for Maxcompute; Single Audit Result DTO, returns a single audit result, including the number of audit items, audit time consumption, and error logs; Audit Datasource Util, the data source parsing tool class; Date Split Util, the date splitting and replacement tool class, splits time according to the day / month dimension.
[0041] In this way, the performance burden and memory consumption of the database are significantly reduced, the database performance overhead and memory overhead are reduced, the audit efficiency of relational data is improved, and the accuracy and consistency of the data are ensured.
[0042] Step S204, if the data source is analytical data, perform a quantity audit on the data source according to the index to determine the second audit result;
[0043] Among them, use a big data processing framework (such as Hadoop Map Reduce, Spark) to perform distributed computing and perform efficient statistical auditing on large-scale data sets. For example, for analytical data, the focus is on checking the data volume, such as recording the number of audit items, sum, average value, etc. Quickly locate the data set through the index, perform aggregation queries or statistical analyses, and compare with the expected values.
[0044] Step S205: Generate an audit report based on the first audit result or / and the second audit result to complete data auditing.
[0045] Among them, by calling a preset audit report template and compiling it based on the first audit result or / and the second audit result, an audit report is automatically generated to complete data auditing.
[0046] In this embodiment, by optimizing the data processing flow and adopting streaming data processing technology, the performance burden and memory consumption of the system on the database are significantly reduced, and large-scale data streams can be processed more efficiently and stably, reducing the database performance overhead and memory overhead. At the application deployment level, the memory overhead of deployment is also reduced, making the deployment cost of the audit service itself relatively low; in terms of the diversity of audit schemes, a variety of audit rules and algorithms are provided, and users are supported to customize audit rules; it can meet the requirements under different business scenarios and more accurately discover and locate data quality problems. Advanced algorithms and models are used to achieve automatic task splitting and parallel processing, improving the processing efficiency and accuracy, and making full use of system resources; at the same time, functions such as task priority scheduling and load balancing are also supported to ensure that tasks can be completed on time and accurately.
[0047] In the related art, in big data or information systems, there are numerous data sources, and poor management easily leads to data errors, omissions, or non-compliance; different data types may require different auditing methods, and it is difficult for traditional methods to respond flexibly.
[0048] Based on the above embodiments, before obtaining the data source, the present application further includes:
[0049] Set the configuration parameters of the data source, determine the audit table of the data source, and the audit table includes whether to audit, the audit time period, and the audit mode. Among them, the audit mode corresponds one-to-one with the data type.
[0050] It should be understood that through the configuration management system, personalized settings are made for the data source to meet the requirements of different business scenarios. The audit table is used as the instruction set for audit tasks, clarifying the objectives, time, and methods of auditing. Auditing tasks are managed through a structured table, improving the standardization and traceability of auditing. The audit status is represented by a boolean value (yes / no), simplifying the decision-making process of audit tasks; in this way, flexible control of the audit scope is allowed, avoiding unnecessary audit overhead; in addition, the audit tasks are restricted by the time range, ensuring the pertinence and timeliness of auditing, allowing adjustment of the audit time period according to actual needs, and improving the flexibility of auditing.
[0051] In the above manner, by setting configuration parameters, the requirements of different business scenarios are met; by designing a structured audit form, the audit objectives, time, and methods are clarified; by setting the "whether to audit" field and the audit time period, flexible control of the audit task is achieved; by associating the data type with the audit mode through the mapping relationship, the most suitable audit method is selected, improving the accuracy and effectiveness of the audit; allowing the selection of the most suitable audit method according to the characteristics of the data type also improves the flexibility of the audit.
[0052] When the data volume of the data source is huge and the audit time period is long, traditional audit methods may lead to low processing efficiency, affecting the timeliness and accuracy of the audit. Or, for long-term audit tasks, resource allocation may be uneven, resulting in insufficient or over-auditing of data in some periods. The management of audit tasks for large data sources is complex and it is difficult to effectively track and execute.
[0053] In view of the above technical problems, in some embodiments, before auditing the data source based on the index, the present application further includes:
[0054] If the data source is configured to be audited and the audit time period exceeds the preset period, the data source is split into multiple data packets according to the time conditions configured for the audit time period;
[0055] Each data packet includes the data to be audited and the target data corresponding to each data to be audited, and an audit sub-task corresponding to the data packet is generated;
[0056] The audit sub-tasks corresponding to each data packet are executed in sequence, and the audit sub-tasks are used as the audit tasks of the data source to generate an audit instance.
[0057] It should be understood that by comparing the time to determine whether the audit time period exceeds the preset period, using a data splitting algorithm to logically or physically split the data source according to the time conditions; ensuring that the data volume and audit complexity of each data packet are within a controllable range; using a data identification algorithm to extract the data to be audited and the target data from the data packet, and generating an audit sub-task according to the audit rules and requirements, ensuring that each sub-task has a clear audit objective and scope; using a task scheduling algorithm to reasonably arrange the execution order and priority of the audit sub-tasks, and through the audit execution engine, sequentially execute the audit sub-tasks, collect the execution results, and integrate the execution results to generate an audit instance for subsequent analysis and reporting.
[0058] Exemplarily, according to the audit time period, determine the split time points or time periods, and split the data source into multiple data packets according to the time points or time periods. Each data packet contains the data within a specific time period. For each data packet, identify the data to be audited and the corresponding target data therein, and generate corresponding audit subtasks according to the data to be audited and the target data, clarifying the audit objectives and scope; execute the audit subtasks corresponding to each data packet in a predetermined order or priority, and integrate the execution results of the audit subtasks to generate an audit instance of the data source.
[0059] Through the above method, by data splitting and generation of audit subtasks, the complexity of auditing is reduced and the auditing efficiency is improved; by reasonably arranging the execution order and priority of audit subtasks, the full utilization of audit resources is ensured; by generating audit subtasks and audit instances, the management and tracking process of audit tasks is simplified; by sequentially executing audit subtasks and integrating the execution results, the comprehensiveness and accuracy of auditing are ensured.
[0060] In the auditing process, if a subtask fails to execute and there is a lack of an effective error handling mechanism, it may cause the entire auditing process to be interrupted or the results to be inaccurate. Due to the lack of effective logging and error handling, the reliability of the audit results is difficult to guarantee.
[0061] To address the above technical problems, in some embodiments, sequentially executing the audit subtasks corresponding to each data packet further includes:
[0062] If it is detected that the audit of the audit subtask fails, an error log is generated, the current audit task is blocked from continuing to execute and an error is reported.
[0063] If it is detected that the audit of the audit subtask is successful, an audit log is generated, and the next audit subtask is executed until the audit subtasks corresponding to the preset data packets are completed to generate a first audit result.
[0064] It should be understood that using an exception capture mechanism or a status detection algorithm to monitor the execution process of the audit subtask in real time. When an exception is captured or a failure status is detected, the error handling process is triggered. Using a log generation algorithm, according to the execution status and relevant information of the audit subtask, an error log is generated, which is beneficial for error analysis and repair. Using a task control mechanism, when it is detected that an audit subtask fails, the execution of subsequent tasks is blocked, and through an error reporting mechanism, the error information is passed to the user or system administrator.
[0065] Use a status detection algorithm or a success identification mechanism to monitor the execution process of the auditing subtasks in real time. When the success status is detected, trigger the success handling process; use a log generation algorithm to generate auditing logs based on the execution status and relevant information of the auditing subtasks; use a task scheduling algorithm to execute the next auditing subtask according to the preset order and the current task status; use a task scheduling and control mechanism to ensure that all auditing subtasks are executed in order; use a result integration algorithm to integrate the results of each auditing subtask into the first auditing result.
[0066] Through the above methods, through real-time monitoring, error handling, and success handling mechanisms, the correct execution of the auditing subtasks and the accuracy of the results are ensured; through the task scheduling and control mechanism, the auditing subtasks are ensured to be executed in order, and detailed log records are generated; by integrating the results of each auditing subtask, the first auditing result is generated, providing a reliable data basis for subsequent analysis and decision-making.
[0067] When dealing with large-scale data, the traditional single-node processing method may lead to low processing efficiency and fail to meet the real-time requirements; at the same time, due to the lack of effective auditing methods and tools, the auditing results of the data may be inaccurate, affecting the data quality.
[0068] To solve the above problems, in some embodiments, the auditing subtasks corresponding to each data packet are executed in sequence, including:
[0069] Allocate the auditing subtasks corresponding to the multiple split data packets to a preset number of processing nodes, so that each processing node audits each data packet through an index to determine the target storage location of each data packet;
[0070] After determining the target storage location, each processing node audits and compares each data packet according to the preset streaming fields to obtain the auditing result of each data packet;
[0071] Summarize the auditing results of each of the data packets to determine the first auditing result to which the data source belongs.
[0072] It should be understood that through the task allocation algorithm, the audit subtasks are evenly distributed to each processing node to improve the parallel processing efficiency. Each processing node independently processes the assigned audit subtasks, reducing the dependencies and waiting times between tasks. Indexing techniques (such as hash index, B-tree index, etc.) are used to accelerate the retrieval process of data packets; by matching the index with the data packets, the target storage location of the data packets is determined, improving the efficiency of data storage and retrieval. The streaming fields are used to identify the key information and features of the data packets. By comparing the streaming fields, anomalies or errors in the data packets can be quickly discovered, improving the accuracy of the audit. According to the results of the audit comparison, an audit report or log is automatically generated. Data summarization and analysis algorithms are used to merge and process multiple audit results. According to the summary results, the overall data quality of the data source is evaluated, and a first audit report is generated.
[0073] Exemplarily, the original data is split into multiple data packets. According to the preset number of processing nodes, the audit subtasks corresponding to each data packet are assigned to the corresponding processing nodes; after each processing node receives the assigned audit subtasks, it uses the index to quickly retrieve the data packets; according to the retrieval results, the target storage location of each data packet is determined; after each processing node determines the target storage location of the data packet, it performs an audit comparison on the data packet according to the preset streaming fields (such as timestamp, service identifier, etc.). During the comparison process, it checks whether the content, format, integrity, etc. of the data packet meet the expected requirements; after each processing node completes the audit comparison, it generates the audit results of the corresponding data packet. The audit results include information such as whether the data packet passes the audit and the errors or anomalies found during the audit process; the audit results generated by all processing nodes are collected; the audit results are summarized and analyzed to determine the first audit result to which the data source belongs.
[0074] Through the above methods, first, through task allocation and parallel processing, the data processing time is significantly shortened; second, by using index retrieval and streaming field comparison, the accuracy and reliability of the audit process are improved; third, by determining the target storage location of the data packets, the efficiency of data storage and retrieval is improved; finally, by summarizing and analyzing the audit results of each data packet, the data quality of the data source is comprehensively evaluated.
[0075] Regarding how to automatically generate an audit report containing detailed metric parameters to improve the efficiency and accuracy of report generation.
[0076] To solve the above problems, in some embodiments, an audit report is generated based on the first audit result or / and the second audit result, including:
[0077] If it is detected based on the audit result that the data source is inconsistent with the preset data, then the preset data audit rules are called to analyze the first audit result or the second audit result to obtain an analysis result;
[0078] Extract features from the analysis results to obtain abnormal data features and key information features;
[0079] Based on natural language processing technology, perform semantic understanding and information extraction on the abnormal data features and key information features to obtain various index parameters;
[0080] Input the various index parameters into a preset text generation model to determine the audit report. The various index parameters include consistency index, integrity index, uniqueness index, and accuracy index.
[0081] It should be understood that the preset data audit rules are formulated based on business logic, data specifications, or industry standards, and are used to perform logical verification, numerical comparison, or pattern matching on the audit results to identify inconsistencies. The feature extraction technology is based on machine learning algorithms, and by analyzing the distribution, frequency, correlation, etc. of the data, extracts features that can represent data anomalies or key information. NLP technology includes multiple levels such as lexical analysis, syntactic analysis, and semantic understanding. Through these technologies, the meaning of text data is deeply understood, and useful information is extracted. The text generation model is based on natural language generation algorithms, and according to the input index parameters, generates a compliant audit report through a preset template or algorithm.
[0082] Exemplarily, through auditing, a first audit result or a second audit result is obtained. When it is detected that the data source is inconsistent with the preset data, the preset data audit rules are triggered. Through in-depth analysis of the audit results, a detailed analysis result is obtained. Feature extraction is performed on the results obtained from the analysis of the audit rules to identify and extract abnormal data features and key information features; natural language processing technology is used to perform semantic analysis on the extracted abnormal data features and key information features. Through semantic understanding, parameters such as consistency index, integrity index, and compliance index are extracted; the extracted various index parameters are input into the preset text generation model, and an audit report is automatically generated according to these parameters. The report includes evaluations of indicators such as consistency index, integrity index, uniqueness index, and accuracy index.
[0083] Through the above methods, first, through automated audit rules and feature extraction technology, data inconsistencies and abnormal features are quickly and accurately identified; second, through natural language processing technology and text generation model, an audit report containing detailed index parameters is automatically generated, and a correction plan is provided; third, the automated process reduces the workload of manual review, report writing, and anomaly repair, and reduces the risk of human error.
[0084] Regarding how to discover abnormal indicator parameters in an audit report and repair them automatically or semi - automatically, the present application adopts the following technical solutions: If it is determined that there are abnormal indicator parameters in the audit report, a preset repair rule is called to repair the abnormal indicator parameters, obtaining a corrected result, and using the corrected result as a supplementary solution for the audit report.
[0085] It should be understood that the preset repair rule is formulated based on business logic, data specifications or repair experience, and is used to automatically or semi - automatically repair abnormal indicator parameters, providing correction suggestions or solutions.
[0086] In the above - mentioned manner, the generated audit report is reviewed to identify abnormal indicator parameters; a preset repair rule is called to repair the abnormal indicator parameters, and the repair result is used as a supplementary solution for the audit report and provided to the user for reference. The automated repair rule improves the efficiency and accuracy of abnormal indicator repair.
[0087] Please refer to Figure 3 , which is a complete flowchart of a data auditing method in an embodiment of the present application, and is described in detail as follows:
[0088] Select the data source, that is, the data source, through the page, and configure each parameter using the audit table of the data source. For example, select whether to audit and the audit mode, whether to use the number of rows for auditing or field matching for auditing. Click "Go live", and the program starts to parse the data source and SQL. According to the type of the data source, different audit executors are automatically selected to perform table metadata query, analyze the primary key / unique index existing in the table, obtain the primary key / unique index field, and use it as the query return field for streaming query. Based on this scenario, data query is performed, and the SQL of "select * from table" is optimized to "select ID from table", where ID is the index. On the one hand, it greatly reduces the number of field queries and the data access volume; on the other hand, by positioning the storage location of the table and ID in advance, it is conducive to fast and accurate query. In this way, it greatly reduces the consumption of IO (input / output) traffic. Especially when there are many table fields, it can reduce the IO consumption by 90% or even more. Since the primary key / unique index is automatically recognized and parsed, when performing database SQL execution, based on the characteristics of the database, since the primary key index is sorted and organized according to ID, the database can directly quickly locate the value of ID through the index without the need for a table return operation (since the leaf node of the primary key index stores the physical address of the entire row of data, here, only the primary key itself is queried, and there is no need to search for other data). Therefore, the query speed is usually very fast, especially when the data volume is large, the advantage is more obvious, and the pressure on the database is significantly reduced.
[0089] Split the time field, which can be split by day or by month. By configuring the streaming processing parameters, set the specific result set modes Result Set.TYPE_FORWARD_ONLY, Result Set.CONCUR_READ_ONLY, and fetchSize. The role of Result Set.TYPE_FORWARD_ONLY is to specify that the cursor can only move forward. The database returns data in a streaming manner. Since there is no need to maintain the ability to randomly access the cursor, the database and JDBC driver can process data more efficiently, reducing the use of memory and system resources. This is particularly important for scenarios involving large amounts of data because it avoids the additional resource overhead for supporting random access. In cases where only sequential processing of the result set is required, this single-direction cursor movement method can improve query performance because the database can adopt a simpler streaming processing mechanism to return data.
[0090] Among them, the role of Result Set.CONCUR_READ_ONLY is to ensure that during the query execution, the data in the result set will not be accidentally modified by the application, thus ensuring data consistency and integrity. Since data modification operations are not supported, the database and JDBC driver do not need to consider the issue of concurrent data modification when processing the result set, thus simplifying the processing logic and improving performance.
[0091] Only a small amount of data enters the memory each time, avoiding loading a large amount of data at once, effectively reducing memory occupancy. This advantage is more obvious in the case of a large amount of data. Taking a relational database like MySQL as an example, pass the split SQL to the subtask, and the subtask performs data query and streaming processing, and accumulates the number of rows in the memory. For a database like Doris with an MPP (Massively Parallel Processing Architecture), the number of rows is queried through the Count(1) syntax. The reasons are the characteristics of Doris:
[0092] ① Columnar storage. When counting the number of rows, there is no need to read all column data, reducing I / O overhead. Combining partitioning and bucketing can accurately locate data blocks;
[0093] ② Sparse indexes can quickly locate data-containing blocks, and the pre-aggregation function enables Doris to directly use the pre-computed count results;
[0094] ③ The query optimizer selects the optimal execution plan for Count(1), and parallel computing can distribute tasks to multiple nodes or threads and summarize the results;
[0095] ④ The metadata it maintains includes information such as the number of rows in the table. In some scenarios, approximate results can be directly obtained, accelerating the query.
[0096] For data sources like Maxcompute, using Count(1) to count the number of records can complete the audit of this task at the millisecond level. That is, for different data source types, it can determine and identify the most suitable query method for the current data source, and select the optimal solution in terms of time, memory, performance, and database consumption as the execution strategy with the goal of successful query. Only when the audit of this batch of subtasks is successful will the next batch of subtasks be executed. If there is a failure, it will be blocked and not continue, and the audit fails.
[0097] Differences from other data audits: Usually, data audits require a large amount of memory resources and CPU resources, resulting in high deployment costs, and support fewer data source types. At the same time, it will occupy a large amount of resources of the business database, with long-term occupation, and also consume database computing resources. Moreover, after a long time, it may not be able to audit successfully, that is, the query at the database level will fail, such as timeouts. The data audit method adopted in this application has the following technical effects:
[0098] First, it uses metadata parsing to determine indexes, query field cropping, and streaming data processing technology to avoid performance bottlenecks at the database level. For different data sources, it intelligently selects streaming processing or Count summation counting methods, reducing database performance overhead and memory overhead, and improving the audit efficiency of relational data. By optimizing the data processing process and adopting streaming data processing technology, it reduces the performance burden and memory consumption of the system on the database, and processes large-scale data streams more efficiently and stably.
[0099] Second, the entire system configuration is simple and easy to use. At the application deployment level, it reduces the memory overhead of deployment, making the deployment cost of the audit service itself relatively low.
[0100] Third, this application has built-in multiple audit rules, provides different audit modes, and adapts to different business scenarios. Users can choose appropriate audit solutions according to actual needs. At the same time, it supports users to customize audit rules to meet more complex and personalized business needs. Through diverse audit solutions, it can more accurately discover and locate data quality problems.
[0101] Fourth, intelligent task splitting greatly reduces memory overhead, as well as the occupation and consumption of database resources. Among them, according to the characteristics and actual needs of the audit task, the task is automatically split into multiple subtasks for parallel processing. Through intelligent task splitting, it can make full use of system resources, improve processing efficiency and accuracy. At the same time, it also supports functions such as task priority scheduling and load balancing to ensure that tasks can be completed on time and accurately.
[0102] Fifth, it has strong scalability, meets different heterogeneous data sources, and adds implementation classes for data source auditing.
[0103] Please refer to Figure 4 , which is a schematic structural diagram of a data auditing system in an embodiment of the present application, including:
[0104] An acquisition module 401, configured to acquire the data source of the data to be audited and determine the data type of the data source, where the data type includes relational data and analytical data;
[0105] A parsing module 402, configured to parse the metadata in the data source to determine the index representing the unique identifier of the metadata;
[0106] A data auditing module 403, configured to, if the data source is relational data, audit the data source based on the index to determine the target storage location of the target data, and audit the target data at the target storage location according to the preset streaming fields to determine the first audit result; if the data source is analytical data, audit the quantity of the data source according to the index to determine the second audit result;
[0107] A report generation module 404, which generates an audit report based on the first audit result or / and the second audit result to complete data auditing.
[0108] It should be noted that the data auditing system of the present application corresponds one-to-one with the data auditing method, and the corresponding technical details and technical solutions refer to the above data auditing method, which will not be elaborated herein one by one.
[0109] By the above method, the present application acquires the data source and determines the data type of the data source; parses the metadata in the data source to determine the index representing the unique identifier of the metadata; if the data source is relational data, audits the data source based on the index to determine the target storage location of the target data; audits the target data at the target storage location according to the preset streaming fields to determine the first audit result; in this way, the performance burden and memory consumption of the database are significantly reduced, and the database performance overhead and memory overhead are reduced; if the data source is analytical data, audits the quantity of the data source according to the index to determine the second audit result; on the one hand, different auditing strategies are adopted according to different data types, improving the auditing quality and auditing performance; on the other hand, real-time big data is split, improving the processing efficiency and accuracy, and making full use of system resources to ensure that tasks can be completed on time and accurately.
[0110] In some other embodiments, the embodiments of the present disclosure further provide an electronic device, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the above method.
[0111] Figure 5The figure shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example and should not impose any limitations on the functions and application scope of the embodiments of the present application.
[0112] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503, such as executing the methods in the above embodiments. In the random access memory 503, various programs and data required for system operation are also stored. The central processing unit 501, the read-only memory 502, and the random access memory 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0113] The following components are connected to the input / output interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that the computer program read from it can be installed into the storage section 508 as needed.
[0114] Specifically, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, various functions defined in the method of the present application are executed.
[0115] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and sub-samples of some embodiments may be included in or replace parts and sub-samples of other embodiments. Moreover, the terms used in this application are only for describing embodiments and do not limit the claims. As used in the description of embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, the term "and / or" used in this application refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated sub-samples, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other sub-samples, wholes, steps, operations, elements, components, and / or groupings of these. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, or apparatus including the element. Herein, what each embodiment focuses on may be the differences from other embodiments, and the same or similar parts among various embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method parts disclosed in the embodiments, the relevant parts may refer to the description of the method parts.
[0116] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the methods, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0117] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the various functional units can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit.
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods, systems, and computer program products according to the embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks can also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based method for performing the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A data auditing method, characterized in that, Including: Obtain a data source of data to be audited, and determine the data type of the data source, where the data type includes relational data and analytical data; Parse the metadata in the data source to determine an index representing the unique identifier of the metadata; If the data source is relational data, audit the data source based on the index to determine the target storage location of the target data; Audit the target data at the target storage location according to a preset streaming field to determine a first audit result; If the data source is analytical data, perform a quantity audit on the data source according to the index to determine a second audit result; Generate an audit report based on the first audit result or / and the second audit result to complete the data audit.
2. The method according to claim 1, wherein Before obtaining the data source, it further includes: Set the configuration parameters of the data source to determine the audit table of the data source, where the audit table includes whether to audit, the audit time period, and the audit mode, and the audit mode corresponds one-to-one with the data type.
3. The method according to claim 2, wherein Before auditing the data source based on the index, it further includes: If the data source is configured to be audited and the audit time period exceeds a preset period, split the data source according to the time condition configured for the audit time period into multiple data packets; Each data packet includes data to be audited and the target data corresponding to each data to be audited, and generate an audit sub-task corresponding to the data packet; Sequentially execute the audit sub-tasks corresponding to each data packet, and use the audit sub-tasks as the audit tasks of the data source to generate an audit instance.
4. The method according to claim 3, characterized in that, Sequentially executing the audit sub-tasks corresponding to each data packet further includes: If it is detected that the audit sub-task fails in the audit, generate an error log, block the continuation of the current audit task and report an error; If it is detected that the audit sub-task is successfully audited, generate an audit log, and execute the next audit sub-task until the audit sub-tasks corresponding to the preset data packets are completed to generate the first audit result.
5. The method according to claim 3, wherein Sequentially executing the audit sub-tasks corresponding to each data packet includes: Allocate the audit sub-tasks corresponding to the split multiple data packets to a preset number of processing nodes, so that each processing node audits each data packet through the index to determine the target storage location of each data packet; After determining the target storage location, each processing node audits and compares each data packet according to a preset streaming field to obtain the audit result of each data packet; Summarize the audit results of each data packet to determine the first audit result to which the data source belongs.
6. The method according to claim 1, wherein The preset streaming field includes at least one of the following: input parameters for defining audit information, audit service classes for configuring entrances, implementation management classes for configuring the mapping relationship between different data sources and corresponding sub-task audit implementation classes, interfaces for defining audit sub-tasks, audit sub-task implementation classes for passing streaming processing parameters, data source parsing tool classes, task splitting tool classes, and single audit return results.
7. The method according to any one of claims 1 to 6, characterized in that Generate an audit report based on the first audit result and / or the second audit result, including: If it is detected based on the audit result that the data source is inconsistent with the preset data, call the preset data audit rule to analyze the first audit result or the second audit result to obtain an analysis result; Extract features from the analysis result to obtain abnormal data features and key information features; Based on natural language processing technology, perform semantic understanding and information extraction on the abnormal data features and the key information features to obtain each of the index parameters; Input each of the index parameters into a preset text generation model to determine the audit report, and each of the index parameters includes a consistency index, a completeness index, a uniqueness index, and an accuracy index.
8. The method according to claim 7, wherein If it is determined that there are abnormal index parameters in the audit report, call the preset repair rule to repair the abnormal index parameters to obtain a correction result, and use the correction result as a supplementary solution for the audit report.
9. A data auditing system, characterized in that, Including: An acquisition module for acquiring the data source of the data to be audited and determining the data type of the data source, where the data type includes relational data and analytical data; A parsing module for parsing the metadata in the data source to determine the index representing the unique identifier of the metadata; A data audit module for, if the data source is relational data, auditing the data source based on the index to determine the target storage location of the target data, and auditing the target data at the target storage location according to the preset streaming fields to determine the first audit result; If the data source is analytical data, perform a quantity audit on the data source according to the index to determine the second audit result; A report generation module for generating an audit report based on the first audit result and / or the second audit result to complete the data audit.
10. An electronic device, characterized in that, Including: A processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method according to claims 1 to 8.