Data auditing method, device and equipment and computer readable storage medium

By pre-embedding the audit SDK in Flink real-time tasks, collecting and processing indicator data, and generating audit evaluation data, the problem of difficulty in quickly identifying task failure links in existing technologies is solved, and real-time and rapid fault perception is achieved.

CN120811931AActive Publication Date: 2025-10-17KINCHENG BANK OF TIANJIN CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511309922.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

The existing Flink real-time task auditing solution has difficulty in quickly identifying failure links in real time, and cannot meet the second-level fault detection requirements of scenarios such as financial transactions and real-time risk control.

Method used

By pre-embedding the audit SDK at the output and input ends of each operator in the Flink real-time task, indicator data is collected and subjected to first aggregation processing, deduplication processing, and second aggregation processing to generate audit evaluation data and ultimately determine the audit result of the task.

Benefits of technology

It realizes real-time and rapid identification of failure links in Flink real-time tasks, meeting the second-level fault perception requirements of financial transactions and real-time risk control scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811931A_ABST
    Figure CN120811931A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data, and discloses a data auditing method, device and equipment and a computer readable storage medium. The method comprises the steps that in the running process of an Flink real-time task, index data of each operator in the Flink real-time task are collected through auditing SDKs, and the auditing SDKs are embedded in the output end and the input end of each operator; performing first aggregation processing on the index data through an auditing SDK to obtain a to-be-audited message; performing deduplication processing and second aggregation processing on the to-be-audited message to obtain audit evaluation data; and determining an audit result of the Flink real-time task based on the audit evaluation data. The index data of the operators are collected in real time through the auditing SDKs pre-embedded in the output end and the input end of each operator, the auditing result of the Flink real-time task is determined based on the index data, the failure link in the real-time task can be rapidly recognized, and the second-level fault sensing requirements of financial transactions, real-time risk control and other scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, and in particular to a data auditing method and device, equipment and a computer readable storage medium. BACKGROUND

[0002] In a Flink real-time task, auditing is a core monitoring and verification mechanism for ensuring the stability of task operation, data accuracy and business compliance. The existing solution mainly adopts an end-to-end offline auditing mode, that is, data at the source end and the target end is periodically collected for comparison and verification.

[0003] However, the existing Flink real-time task auditing is a post-remedial type of auditing solution, which is difficult to quickly identify the failure link in the Flink real-time task in real time, and cannot meet the second-level fault perception needs of financial transactions, real-time risk control and other scenarios. SUMMARY

[0004] Therefore, the purpose of the present application is to overcome the deficiencies in the prior art and provide a data auditing method, which comprises: In the process of running a Flink real-time task, the audit SDK is used to collect the index data of each operator in the Flink real-time task, and the audit SDK is pre-embedded at the output end and the input end of each operator; The audit SDK is used to perform first aggregation processing on the index data to obtain a to-be-audited message; The to-be-audited message is subjected to a de-duplication process and second aggregation processing to obtain auditing evaluation data; Based on the auditing evaluation data, an auditing result of the Flink real-time task is determined.

[0005] In an embodiment, the step of performing first aggregation processing on the index data by the audit SDK to obtain a to-be-audited message comprises: The audit SDK determines an aggregation strategy corresponding to each index data based on the data type of each index data; Each index data is subjected to first aggregation processing based on the aggregation strategy to obtain a to-be-audited message.

[0006] In an embodiment, the step of performing first aggregation processing on the index data based on the aggregation strategy to obtain a to-be-audited message comprises: Each index data is subjected to first aggregation processing based on the aggregation strategy to obtain an aggregated data message corresponding to each index data; Identify the dimension field of each of the aggregated data messages, assemble the aggregated data messages with the same dimension field, and obtain the message to be audited corresponding to each of the dimension fields.

[0007] In one embodiment, the step of performing deduplication processing and second aggregation processing on the to-be-audited messages to obtain audit evaluation data includes: Performing dimension and data deduplication processing on the message to be audited; Deserialize the audited message after dimension and data deduplication processing to obtain detailed audit data; Perform a second aggregation process on the detailed audit data to obtain audit evaluation data.

[0008] In one embodiment, the step of performing a second aggregation process on the detailed audit data to obtain audit evaluation data includes: Identify the dimension fields of each of the detailed audit data, cache the detailed audit data with the same dimension fields and perform local aggregation to obtain detailed audit aggregated data corresponding to each of the dimension fields; Based on a preset aggregation time interval, the detailed audit aggregation data corresponding to each dimension field is globally aggregated to obtain audit evaluation data.

[0009] In one embodiment, the step of determining the audit result of the Flink real-time task based on the audit evaluation data includes: Based on the audit evaluation data, determine the data volume audit result, data delay audit result, and task heartbeat audit result of the Flink real-time task; Based on the data volume audit result, the data delay audit result and the task heartbeat audit result, if it is determined that an abnormality exists, an abnormality alarm is issued.

[0010] In one embodiment, after the step of performing a first aggregation process on the indicator data by the audit SDK to obtain the message to be audited, the following steps are included: Adding the to-be-audited message to a to-be-reported queue, and sending the to-be-audited message to an intermediate storage device based on the to-be-reported queue; The step of performing deduplication processing and second aggregation processing on the audited messages to obtain audit evaluation data includes: The to-be-audited messages are read from the intermediate storage device, and deduplication processing and second aggregation processing are performed on the to-be-audited messages to obtain audit evaluation data.

[0011] The present application also provides a data auditing device, comprising: The collection module is used to collect the indicator data of each operator in the Flink real-time task through the audit SDK during the execution of the Flink real-time task. The audit SDK is pre-embedded in the output and input of each operator; A first processing module, configured to perform a first aggregation process on the indicator data through the audit SDK to obtain a message to be audited; A second processing module is used to perform deduplication processing and second aggregation processing on the audited message to obtain audit evaluation data; A determination module is used to determine an audit result of the Flink real-time task based on the audit evaluation data.

[0012] The present application also provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the above-mentioned data auditing method.

[0013] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is run on a processor, the data auditing method described above is executed.

[0014] The embodiments of the present application have the following beneficial effects: During the operation of a Flink real-time task, the embodiment of the present application collects the indicator data of each operator in the Flink real-time task through the audit SDK, which is pre-embedded in the output and input of each operator; performs a first aggregation process on the indicator data through the audit SDK to obtain a message to be audited; performs a deduplication process and a second aggregation process on the message to be audited to obtain audit evaluation data; and determines the audit result of the Flink real-time task based on the audit evaluation data. By using the audit SDK pre-embedded in the output and input of each operator to collect the indicator data of each operator in real time, and then processing the indicator data to determine the audit result of the Flink real-time task, the failure link in the Flink real-time task can be identified in real time and quickly, meeting the second-level fault perception requirements of scenarios such as financial transactions and real-time risk control. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the technical solution of this application, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of this application and should not be considered as limiting the scope of protection of this application. Those skilled in the art can also derive other relevant drawings based on these drawings without inventive effort.

[0016] Figure 1 A flowchart of the first embodiment of the data audit method provided by this application; Figure 2 Flowchart of a second embodiment of the data auditing method provided in the present application; Figure 3 Flowchart of a third embodiment of the data auditing method provided in the present application; Figure 4 Flowchart of a fourth embodiment of the data auditing method provided in the present application; Figure 5 Flowchart of a fifth embodiment of the data auditing method provided in the present application; Figure 6 Interaction between the auditing SDK and the intermediate storage device provided in the present application; Figure 7 Auditing flowchart provided in the present application; Figure 8 Structure diagram of the data auditing device provided in the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application.

[0018] The components of the embodiments of the present application generally described and shown in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] Hereinafter, the terms "include", "have", and their synonymous words used in various embodiments of the present application are only intended to indicate that specific features, numbers, steps, processes, elements, components, or combinations of the foregoing are present, and should not be understood as excluding or adding the possibility of existence or addition of one or more features, numbers, steps, processes, elements, components, or combinations of the foregoing.

[0020] In addition, the terms "first", "second", "third", etc. are only used for differentiation in description, and should not be understood as indicating or implying relative importance.

[0021] Unless specifically defined, all other technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application belong. The terminology used herein (e.g., the terminology used in the description of the figures) is for the purpose of describing particular embodiments only and is not intended to be limiting. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application belong. The terminology used herein, such as the terminology defined in a generally used dictionary, will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized or overly formal meaning, unless clearly defined in various embodiments of the present application.

[0022] It can be understood that the method of the present application is applied to a data auditing device, which can be a smart terminal, a PC terminal, a mobile terminal, etc., which is not limited herein.

[0023] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0024] Please refer to Figure 1 , Figure 1 The flowchart of a first embodiment of the data auditing method provided by the present application is shown, and the method comprises: Step S101, in the process of running the Flink real-time task, collecting the index data of each operator in the Flink real-time task through an auditing SDK, the auditing SDK being pre-embedded at the output end and the input end of each operator.

[0025] In this embodiment, the data auditing device collects the index data of each operator in the Flink real-time task through the auditing SDK in the process of running the Flink real-time task; wherein the auditing SDK (Software Development Kit) is a "standardized tool set" provided by the developer for developing a specific software, interfacing with a specific system or implementing a specific function, and is specially used for collecting the index data of each operator in the Flink real-time task and other functions. The auditing SDK is pre-embedded at the output end and the input end of each operator, and the index data includes: operator name, task name, data table name, operator input record number, operator output record number, operator processing delay, serialization time consumption, deserialization time consumption, timestamp, etc.

[0026] Step S102, performing first aggregation processing on the index data through the auditing SDK to obtain a to-be-audited message.

[0027] In this embodiment, after the data auditing device collects the index data of each operator through the auditing SDK, it performs first aggregation processing on the index data with the same timestamp according to the timestamp in the index data through the auditing SDK to obtain a to-be-audited message. The first aggregation processing is to process and integrate the collected index data according to certain rules.

[0028] Step S103, the de-duplication processing and the second aggregation processing are performed on the to-be-audited message to obtain audit evaluation data.

[0029] In this embodiment, the data auditing device performs de-duplication processing and second aggregation processing on the to-be-audited message to obtain audit evaluation data. The de-duplication processing includes de-duplication of the dimension field of the to-be-audited message and de-duplication of the to-be-audited message and data in the to-be-audited message. De-duplication of the dimension field of the to-be-audited message enables to-be-audited messages with the same dimension field to be assigned to the same processing node, thereby avoiding cross-node synchronization of the de-duplication state. De-duplication of the to-be-audited message and data in the to-be-audited message avoids the influence of repeated to-be-audited messages and data on the accuracy of the audit. The second aggregation processing includes local cache aggregation of data in the to-be-audited message after de-duplication and audit aggregation of the locally cached data at a timing to obtain audit evaluation data.

[0030] Step S104, the audit result of the Flink real-time task is determined based on the audit evaluation data.

[0031] In this embodiment, the data auditing device analyzes the audit evaluation data to determine the audit result of the Flink real-time task, and displays the audit result in real time.

[0032] The data auditing device of this embodiment collects the index data of each operator in the Flink real-time task through the audit SDK during the running of the Flink real-time task, and the audit SDK is embedded in the output end and the input end of each operator. The index data is processed through the audit SDK to obtain to-be-audited messages, and the to-be-audited messages are de-duplicated and aggregated to obtain audit evaluation data. The audit result of the Flink real-time task is determined based on the audit evaluation data. The index data of each operator is collected in real time through the audit SDK embedded in the output end and the input end of each operator, and the index data is processed to determine the audit result of the Flink real-time task. This can quickly identify the failure link in the Flink real-time task in real time, and meet the second-level fault perception needs of financial transactions, real-time risk control and other scenarios.

[0033] Please refer to Figure 2 , Figure 2 The flowchart of the second embodiment of the data auditing method provided in this application is different from the first embodiment in that the step of performing first aggregation processing on the index data through the audit SDK to obtain to-be-audited messages includes: Step S201, determining the aggregation strategy corresponding to each index data based on the data type of each index data through the audit SDK.

[0034] In the embodiment, the data auditing device determines the aggregation policy corresponding to each metric data based on the data type of each metric data through the auditing SDK. It can be understood that different data types correspond to different aggregation policies, and the aggregation policies corresponding to various data types are pre-stored in the auditing SDK. The auditing SDK queries the pre-set aggregation policy table based on the data type of each metric data, so as to determine the aggregation policy corresponding to the data type of each metric data. The aggregation policy includes migagg, sample and full. The migagg is a micro-aggregation policy of a data volume statistical index, mainly reduces the influence of index reporting frequency on network io, and reports the aggregated index after micro-aggregation. The sample is a data stream sampling aggregation policy, which samples data through a pre-set sampling frequency, and is used to judge the running state of the task. The full is a full-amount collection aggregation policy of table change in a data stream, that is, all data are used for evaluation, and the table structure change of the upstream data table can be monitored in real time, which is convenient for subsequent problem positioning and corresponding processing.

[0035] In step S202, first aggregation processing is performed on each metric data based on the aggregation policy, and a to-be-audited message is obtained.

[0036] In the embodiment, after determining the aggregation policy corresponding to each metric data, the data auditing device performs first aggregation processing on each metric data based on the aggregation policy, and obtains a to-be-audited message.

[0037] In an embodiment, the step of performing first aggregation processing on each metric data based on the aggregation policy to obtain a to-be-audited message includes: In step S2021, first aggregation processing is performed on each metric data based on the aggregation policy, and an aggregated data message corresponding to each metric data is obtained.

[0038] In the embodiment, the data auditing device performs first aggregation processing on each metric data based on the aggregation policy, and obtains an aggregated data message corresponding to each metric data. It can be understood that for some data types of metric data, the data auditing device performs aggregation through the aggregation policy migagg to obtain an aggregated data message containing the result of data volume statistics. For some data types of metric data, the data auditing device performs aggregation through the aggregation policy sample to obtain an aggregated data message containing the sampling result of the metric data. For some data types of metric data, the data auditing device performs aggregation through the aggregation policy full to obtain an aggregated data message of all metric data.

[0039] In step S2022, the dimension field of each of the aggregated data messages is identified, and the aggregated data messages with the same dimension field are assembled to obtain the to-be-audited messages corresponding to each of the dimension fields.

[0040] In this embodiment, the data auditing device identifies the dimension field of each of the aggregated data messages, and assembles the aggregated data messages with the same dimension field to obtain the to-be-audited messages corresponding to each of the dimension fields. It should be noted that the dimension field is a combination of fields used to uniquely identify a certain type of auditing data dimension in a data flow conversion link, and each aggregated data message corresponds to a dimension field, which usually consists of multiple fields. The specific field combination can be configured according to actual business requirements. In this solution, the core fields include: jobName, operatorName, batchSeq, clientIp, sourceDb, and sourceTable. Exemplarily, the dimension field is: {"jobName": "realtime_risk_control_job", "operatorName": "kafka_source", "sourceDb": "risk_db", "sourceTable": "user_login_log", "clientIp": "192.168.1.10"}. This dimension field represents: the data of the "risk_db.user_login_log" table from the "192.168.1.10" node processed by the "kafka_source" operator in the "realtime_risk_control_job" task. The data auditing device assembles the aggregated data messages with the same dimension field to obtain the to-be-audited messages corresponding to each of the dimension fields, so as to facilitate subsequent unified processing of the to-be-audited messages corresponding to the same dimension field and improve the auditing efficiency.

[0041] The data auditing device of this embodiment determines the aggregation strategy corresponding to each of the index data based on the data type of each of the index data through the auditing SDK, performs first aggregation processing on each of the index data based on the aggregation strategy, obtains the aggregated data messages corresponding to each of the index data, identifies the dimension field of each of the aggregated data messages, assembles the aggregated data messages with the same dimension field, and obtains the to-be-audited messages corresponding to each of the dimension fields. Through aggregation and assembly of the aggregated data messages with the same dimension field, the data amount is reduced, and the to-be-audited messages corresponding to the same dimension field are uniformly integrated, which helps to improve the auditing efficiency.

[0042] For reference Figure 3 , Figure 3A flowchart of a third embodiment of the data auditing method provided in the present application is shown in FIG. 3. The third embodiment is different from the first embodiment to the second embodiment in that the step of performing the deduplication processing and the second aggregation processing on the to-be-audited message to obtain the audit evaluation data comprises: In step S301, the dimension and data deduplication processing is performed on the to-be-audited message.

[0043] In this embodiment, the data auditing device acquires all the to-be-audited messages and performs the dimension and data deduplication processing on the to-be-audited messages. Specifically, all the to-be-audited messages are first stored in a preset storage device, and the data auditing device acquires all the to-be-audited messages within a certain time range according to the timestamp information of the to-be-audited messages. The data auditing device identifies the dimension field of each to-be-audited message, performs the dimension deduplication processing on all the to-be-audited messages based on the dimension field, so that the to-be-audited messages with the same dimension field are collected for processing, thereby avoiding that the to-be-audited messages with the same dimension field are divided into multiple groups and are inconvenient to process. The data auditing device performs the bitmap interval deduplication on all the to-be-audited messages. The bitmap interval deduplication is a "uniqueness marking tool" that solves the problem of how to efficiently judge whether the data is duplicated. The bitmap interval deduplication uses a compact data structure to realize low memory and high throughput deduplication, and ensures the uniqueness of the data.

[0044] In step S302, the to-be-audited message subjected to the dimension and data deduplication processing is deserialized to obtain the detailed audit data.

[0045] In this embodiment, the data auditing device deserializes each to-be-audited message subjected to the dimension and data deduplication processing, and reads the detailed audit data by expanding the to-be-audited message through the flatMap operator. The flatMap operator is a very flexible operator, and its core function is "one-to-many" data conversion. It can convert one input message into multiple output data.

[0046] In step S303, the second aggregation processing is performed on the detailed audit data to obtain the audit evaluation data.

[0047] In this embodiment, after obtaining the detailed audit data in all the to-be-audited messages, the data auditing device performs the second aggregation processing on the detailed audit data to obtain the audit evaluation data. It can be understood that the second aggregation processing is to analyze and integrate all the accurate and non-duplicated detailed audit data obtained after the collection, the first aggregation, the deduplication processing and the deserialization, so as to obtain the audit evaluation data corresponding to the data output by the Flink real-time task execution process within a period of time.

[0048] In an embodiment, the step of performing the second aggregation processing on the detailed audit data to obtain the audit evaluation data comprises: Step S3031: Identify the dimension fields of each of the detailed audit data, cache the detailed audit data with the same dimension fields, and perform local aggregation to obtain detailed audit aggregated data corresponding to each of the dimension fields.

[0049] In this embodiment, the data audit device identifies the dimension fields of each detailed audit data item, caches and locally aggregates the detailed audit data items with the same dimension fields, and obtains detailed audit aggregate data corresponding to each dimension field. For example, detailed audit data items with the same dimension fields are cached and locally aggregated according to the local cache structure, and the resulting audit aggregate data includes the dimension fields (task name, operator name, batch number, client IP address, source library name, source table name), timestamp, delay time, and data volume (the aggregated detailed audit data sequence under the dimension field).

[0050] Step S3032: Based on a preset aggregation time interval, globally aggregate the detailed audit aggregation data corresponding to each dimension field to obtain audit evaluation data.

[0051] In this embodiment, since the cached audit aggregate data will be automatically deleted after a preset time, the data audit device sets a preset aggregation time interval, which is less than the preset time for the audit aggregate data to be automatically deleted. That is, before the cached audit aggregate data is deleted, the data audit device globally aggregates the detailed audit aggregate data corresponding to each dimension field in the cache to obtain audit evaluation data. For example, the purpose of global aggregation is to calculate the delay fluctuation over a period of time based on the detailed audit aggregate data corresponding to each dimension field, and obtain the audit evaluation data as shown in Table 1 below.

[0052] Table 1

[0053] The interval is the data delay time interval that needs to be audited, and m represents minutes. The quantity is the total amount of data in the corresponding interval. The proportion is the ratio of the total amount of data in each interval to the total amount of all data. The average delay is the average of the sum of the delay times of all data in the interval.

[0054] The data auditing device of the embodiment performs dimension and data deduplication processing on the to-be-audited message; performs deserialization on the to-be-audited message after dimension and data deduplication processing, to obtain detailed auditing data; identifies the dimension field of each piece of detailed auditing data, buffers and locally aggregates the detailed auditing data with the same dimension field, to obtain detailed auditing aggregated data corresponding to each dimension field; and performs global aggregation on the detailed auditing aggregated data corresponding to each dimension field based on a preset aggregation time interval, to obtain auditing evaluation data. The accurate and non-repeated all detailed auditing data obtained after deduplication processing and deserialization are analyzed and integrated, to obtain auditing evaluation data corresponding to the data output by the Flink real-time task in a time period, thereby improving the accuracy of the auditing evaluation data.

[0055] Please refer to Figure 4 , Figure 4 The flowchart of the fourth embodiment of the data auditing method provided in the present application is provided, and the fourth embodiment is different from the first embodiment to the third embodiment in that the step of determining the auditing result of the Flink real-time task based on the auditing evaluation data comprises: Step S401, determining the data volume auditing result, the data delay auditing result and the task heartbeat auditing result of the Flink real-time task based on the auditing evaluation data.

[0056] Step S402, based on the data volume auditing result, the data delay auditing result and the task heartbeat auditing result, if it is determined that there is an exception, performing exception alarm.

[0057] In the embodiment, the data auditing device determines the data volume auditing result, the data delay auditing result and the task heartbeat auditing result of the Flink real-time task based on the auditing evaluation data. Based on the data volume auditing result, the data delay auditing result and the task heartbeat auditing result, if it is determined that there is an exception, performing exception alarm. In an embodiment, the real-time monitoring dashboard constructed by the data auditing device comprises three core modules: 1) data volume auditing dashboard: the growth trend of table-level data is displayed through a line chart, and a threshold is set to alarm abnormal fluctuations; 2) data delay auditing dashboard: a heat map and percentile statistics are used to monitor link delay; and 3) task heartbeat dashboard: a state matrix is used to display the survival state of each task node in real time, to realize minute-level fault detection. Although the auditing evaluation data mainly includes time interval, quantity, proportion and average time delay, these basic indexes can support the data required by the three core monitoring dashboards after reasonable combination and calculation; the three core monitoring dashboards support 10-second-level refresh, and the alarm is connected to the enterprise alarm interface to meet the production-level operation and maintenance requirements.

[0058] The data audit device of the embodiment determines a data volume audit result, a data delay audit result and a task heartbeat audit result of the Flink real-time task based on the audit evaluation data, and performs abnormality alarm if it is determined that there is an abnormality based on the data volume audit result, the data delay audit result and the task heartbeat audit result. The Flink real-time task full-link real-time monitoring of the source task-operator-target source can be constructed, the coarse-grained mode of the traditional end-to-end audit is broken through, the accurate collection of the input / output data of a single operator is realized through SDK burying, and the fault of each operator in the Flink real-time task link can be dynamically tracked and located.

[0059] Please refer to Figure 5 , Figure 5 The flowchart of the fifth embodiment of the data audit method provided in the present application, the fifth embodiment is different from the first embodiment to the fourth embodiment in that, after the step of performing first aggregation processing on the index data through the audit SDK to obtain the to-be-audited message, the following steps are included: In step S501, the to-be-audited message is added to a to-be-reported queue, and the to-be-audited message is sent to an intermediate storage device based on the to-be-reported queue.

[0060] In an embodiment, the step of performing deduplication processing and second aggregation processing on the to-be-audited message to obtain audit evaluation data includes: In step S502, the to-be-audited message is read from the intermediate storage device, and the to-be-audited message is subjected to deduplication processing and second aggregation processing to obtain audit evaluation data.

[0061] In the embodiment, the data audit device adds the to-be-audited message to the to-be-reported queue, and sends the to-be-audited message to the intermediate storage device based on the to-be-reported queue; then, the data audit device reads the to-be-audited message from the intermediate storage device, and performs deduplication processing and second aggregation processing on the to-be-audited message to obtain audit evaluation data.

[0062] In an embodiment, as shown in Figure 6 , Figure 6The interaction schematic diagram between the audit SDK provided in the application and the intermediate storage device is shown. The intermediate storage device is Kafka, and Kafka is a distributed stream processing platform. The core positioning of Kafka is a message queue and stream data storage system with high throughput, persistence, and horizontal scalability. The audit SDK architecture is implemented based on a producer-consumer mode. A producer Reporter is provided as an outermost API to call and report audit information by a monitoring end. The audit SDK collects index data of each operator in a Flink real-time task through the producer Reporter, such as data volume statistics, table structure changes, heartbeat sampling, and other timestamp latitude information. The collected index data is encapsulated into a transmissible format, such as JSON or Protobuf. A capacity timer is responsible for the caching and pre-aggregation of the index data, obtains aggregated data messages, assembles the aggregated data messages, obtains to-be-audited messages, and sends the assembled to-be-audited messages to a Disruptor queue group. The to-be-audited messages are sent to Kafka by a consumer sender through a network protocol in an asynchronous manner. The capacity timer adopts a lock-free design. The pre-aggregation of messages with the same key (dimension field) multiplexes the same atomic variable object, thereby reducing unnecessary garbage collection (gc). A message assembler generates audit messages for all keys in batches and resets the atomic variable value, and then sends the audit messages to a to-be-reported queue by a disruptor producer.

[0063] In an embodiment, as shown in Figure 7 Figure 7 ​An audit process schematic diagram is provided for the present application. The specific process is as follows: read the to-be-audited message reported by the audit SDK from Kafka; use keyby to process the deduplication dimension, further filter the data, and then perform bitmap interval deduplication to ensure data uniqueness; through the flatMap operator, the to-be-audited message after the deduplication operation is expanded to read the detail audit data; based on keyby and keyProcess, the aggregation operation is performed, the detail audit data of the dimension key (dimension field) in the fixed time range is aggregated into a record, and then all the records are integrated to calculate the interval delay fluctuation, obtain the audit evaluation data, and then send the audit evaluation data to the doris sink (Doris Sink refers to writing the calculation results (or upstream data) to the Apache Doris data warehouse "output component", which is the key link between the stream computing engine and Doris), and finally the audit evaluation data is read from the doris sink for analysis to determine the data volume audit result, data delay audit result and task heartbeat audit result of the Flink real-time task. Among them, the aggregation operation based on keyby and keyProcess adopts a double-layer aggregation algorithm, the double-layer aggregation algorithm: 1, local cache for pre-aggregation, cache cache (dimension key, batchSeqs), set the cache expiration time T min, remove the key after expiration. The data in the fixed time range of the dimension key is aggregated into a record (n->1), 2, Timer handles the final aggregation result to sink. The detailed execution steps are as follows: the input stream data is grouped according to the dimension key (dimension field) and stored in the local cache structure (such as GuavaCache), the cache key: dimension key (task name, operator name, batch number, client IP, source library name, source table name), collection timestamp, value: delay time, data volume (to-be-aggregated data sequence under the key), set the cache TTL to T minutes, automatically trigger the removal logic when it expires. Implement the flink timer timer, scan the cache at a fixed interval (≤T) to perform aggregation calculation (match delay time interval and accumulate interval delay data volume) on each key that is about to expire to obtain audit evaluation data.

[0064] The data audit device of the embodiment realizes asynchronous transmission of the to-be-audited message based on the intermediate storage device by sending the to-be-audited message to the intermediate storage device, realizes decoupling between data collection and data audit, avoids affecting the Flink real-time task itself in high-concurrency real-time scenarios such as finance and risk control, and realizes real-time and rapid identification of invalid links in the Flink real-time task.

[0065] Reference Figure 8 , Figure 8is a structural schematic diagram of a data auditing apparatus provided in the application. The data auditing apparatus comprises: The collection module 10 is configured to collect, by an auditing SDK, index data of each operator in a Flink real-time task during running of the Flink real-time task, the auditing SDK being pre-embedded at an output end and an input end of each operator. The first processing module 20 is configured to perform first aggregation processing on the index data by the auditing SDK to obtain a to-be-audited message. The second processing module 30 is configured to perform deduplication processing and second aggregation processing on the to-be-audited message to obtain auditing evaluation data. The determination module 40 is configured to determine an auditing result of the Flink real-time task based on the auditing evaluation data.

[0066] It can be understood that the data auditing apparatus of the embodiment corresponds to the data auditing method of the above-described embodiment, and the optional items in the above-described embodiment are also applicable to the embodiment, and thus are not described herein again.

[0067] The application further provides a computer device. Illustratively, the computer device comprises a processor and a memory. The memory stores a computer program. The processor runs the computer program, so that the computer device performs the functions of the data auditing method or each module of the data auditing apparatus.

[0068] The processor can be an integrated circuit chip with a signal processing capability. The processor can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), and a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or at least one of the above. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the application.

[0069] The memory can be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), and the like. Among them, the memory is used to store a computer program, and the processor can execute the computer program correspondingly after receiving an execution instruction.

[0070] The application further provides a computer storage medium for storing the computer program used in the computer device. The computer storage medium can be a readable storage medium, a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium can include, but is not limited to, a U disk, a mobile hard disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk and various program code storage media.

[0071] In several embodiments provided in the application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are only schematic, for example, the flow charts and structural diagrams in the drawings show the possible implementation architectures, functions and processes of the apparatus, method and computer program product according to the embodiments of the application. In this regard, each block in the flow chart or structural diagram can represent a module, a program segment or a part of code containing one or more executable instructions for implementing the specified logic function. It should also be noted that, in alternative implementation manners, the functions annotated in the blocks can also occur in different orders from those annotated in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the structural diagram and / or flow chart, and the combination of blocks in the structural diagram and / or flow chart, can be implemented by a dedicated hardware-based system for executing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0072] In addition, each functional module or unit in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0073] The functions, if implemented in the form of software functional modules and sold or used as independent products, can be stored in a readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.

[0074] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application.

Claims

1. A data audit method, characterized in that: The method comprises: During the running of the Flink real-time task, the indicator data of each operator in the Flink real-time task is collected through the audit SDK. The audit SDK is pre-embedded in the output and input of each operator. Performing a first aggregation process on the indicator data through the audit SDK to obtain a message to be audited; Performing deduplication processing and second aggregation processing on the messages to be audited to obtain audit evaluation data; Based on the audit evaluation data, an audit result of the Flink real-time task is determined.

2. The data audit method according to claim 1, characterized in that: The step of performing a first aggregation process on the indicator data by the audit SDK to obtain a message to be audited includes: Determining, by the audit SDK, an aggregation strategy corresponding to each indicator data based on the data type of each indicator data; Based on the aggregation strategy, a first aggregation process is performed on each indicator data to obtain a message to be audited.

3. The data audit method according to claim 2, characterized in that: The step of performing a first aggregation process on each indicator data based on the aggregation strategy to obtain a message to be audited includes: Based on the aggregation strategy, performing a first aggregation process on each indicator data to obtain an aggregated data message corresponding to each indicator data; Identify the dimension field of each of the aggregated data messages, assemble the aggregated data messages with the same dimension field, and obtain the message to be audited corresponding to each of the dimension fields.

4. The data audit method according to claim 1, characterized in that: The step of performing deduplication processing and second aggregation processing on the audited messages to obtain audit evaluation data includes: Performing dimension and data deduplication processing on the message to be audited; Deserialize the audited message after dimension and data deduplication processing to obtain detailed audit data; Perform a second aggregation process on the detailed audit data to obtain audit evaluation data.

5. The data audit method according to claim 4, characterized in that: The step of performing a second aggregation process on the detailed audit data to obtain audit evaluation data includes: Identify the dimension fields of each of the detailed audit data, cache the detailed audit data with the same dimension fields and perform local aggregation to obtain detailed audit aggregated data corresponding to each of the dimension fields; Based on a preset aggregation time interval, the detailed audit aggregation data corresponding to each dimension field is globally aggregated to obtain audit evaluation data.

6. The data audit method according to claim 1, characterized in that: The step of determining the audit result of the Flink real-time task based on the audit evaluation data includes: Based on the audit evaluation data, determine the data volume audit result, data delay audit result, and task heartbeat audit result of the Flink real-time task; Based on the data volume audit result, the data delay audit result and the task heartbeat audit result, if it is determined that an abnormality exists, an abnormality alarm is issued.

7. The data audit method according to any one of claims 1 to 6, characterized in that: After the step of performing a first aggregation process on the indicator data through the audit SDK to obtain the message to be audited, the method includes: Adding the to-be-audited message to a to-be-reported queue, and sending the to-be-audited message to an intermediate storage device based on the to-be-reported queue; The step of performing deduplication processing and second aggregation processing on the audited messages to obtain audit evaluation data includes: The to-be-audited messages are read from the intermediate storage device, and deduplication processing and second aggregation processing are performed on the to-be-audited messages to obtain audit evaluation data.

8. A data auditing device, characterized in that: The data auditing device comprises: The collection module is used to collect the indicator data of each operator in the Flink real-time task through the audit SDK during the execution of the Flink real-time task. The audit SDK is pre-embedded in the output and input of each operator; A first processing module, configured to perform a first aggregation process on the indicator data through the audit SDK to obtain a message to be audited; A second processing module is used to perform deduplication processing and second aggregation processing on the audited message to obtain audit evaluation data; A determination module is used to determine an audit result of the Flink real-time task based on the audit evaluation data.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the data auditing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is run on a processor, the data auditing method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Method for monitoring health condition of flink real-time operation in full-link mode

    CN112506737A

  • Application service index monitoring method and device, computer equipment and storage medium

    CN112749056A

  • Data extraction method and device, equipment and storage medium

    CN112949763A

  • Audit data processing system and method, edge server and computer medium

    CN117793093A

  • Data reporting method, data reporting toolkit and data auditing platform

    CN120541111A