Data reporting method, data reporting toolkit and data auditing platform

Through the lock-free design data reporting method, the high performance and flexibility of the data audit system are achieved, and the problems of single index reporting, low performance and high coupling in the existing system are solved, and full-link audit support is provided.

CN120541111AActive Publication Date: 2025-08-26KINCHENG BANK OF TIANJIN CO LTD

Patent Information

Application Number
CN202511038499.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-08-26
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

The existing data audit system lacks full-link audit function in real-time data computing scenarios, with single indicator reporting, low performance, high coupling and poor flexibility.

Method used

The data reporting method with lock-free design is adopted, and audit data is collected through the timer interface, pre-aggregation and secondary aggregation processing is performed. The audit messages are cached using cache middleware, and asynchronously written to distributed message middleware, supporting customized indicators and high concurrency environments.

Benefits of technology

It improves the performance of the data audit system, decouples the process, realizes flexible data reporting and monitoring, and ensures the stability and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541111A_ABST
    Figure CN120541111A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data reporting method, a data reporting toolkit and a data auditing platform, and the method comprises the following steps: calling a preset timer interface in response to an external client, so as to collect auditing data from a preset database by using the timer interface; performing pre-aggregation processing on the audit data according to a preset rule and a preset field to obtain aggregated data; based on a lock-free design mode, performing secondary aggregation processing on the aggregated data according to a preset data structure to obtain an audit message, and caching the audit message in a preset cache middleware; writing the audit message in the cache middleware into preset distributed message middleware in an asynchronous mode; wherein the distributed message middleware is used for asynchronously consuming auditing data from the distributed message middleware by an external data auditing system. Therefore, the problems that an existing data auditing system is low in performance, high in coupling performance, poor in flexibility and the like can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data reporting method, a data reporting toolkit, and a data auditing platform. Background Art

[0002] In real-time data calculation scenarios, it is necessary to audit indicators such as the delay of each branch flow in the entire end-to-end link, the overall link delay, and the link health status to verify the accuracy and timeliness of the data and monitor the operating status of the link.

[0003] However, the industry's monitoring and data auditing systems generally lack support for full-link audit-related functions. For example, the reported indicators are relatively simple and lack support for custom indicators. In addition, there are problems such as low performance, high coupling, and poor flexibility. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a data reporting method, a data reporting toolkit and a data auditing platform, which can effectively solve the problems of low performance, high coupling and poor flexibility of the existing data auditing system.

[0005] In a first aspect, an embodiment of the present application provides a data reporting method, the method comprising: In response to an external client calling a preset timer interface, collecting audit data from a preset database using the timer interface; According to preset rules, the audit data is pre-aggregated according to preset fields to obtain aggregated data; Based on a lock-free design, the aggregated data is subjected to secondary aggregation processing according to a preset data structure to obtain an audit message, which is then cached in a preset cache middleware; Writing the audit message in the cache middleware into the preset distributed message middleware in an asynchronous manner; The distributed message middleware is used for an external data audit system to asynchronously consume the audit data.

[0006] In some embodiments, the preset audit fields in the audit data include audit data type, data source table, and data source library; The pre-aggregating the audit data according to preset rules and preset fields to obtain aggregated data includes: Based on a preset time interval, determining a unique pre-aggregation key corresponding to each piece of the audit data according to a preset field; Determine a unique corresponding atomic variable object for each of the pre-aggregation keys; Aggregate all audit data with the same pre-aggregation key, and assign the obtained aggregated data to the corresponding atomic variable object.

[0007] In some embodiments, determining a unique corresponding atomic variable object for each of the pre-aggregated bonds includes: Check whether there is a corresponding atomic variable object for the pre-aggregation key; if not, generate a corresponding atomic variable object; if so, obtain the atomic variable object.

[0008] In some embodiments, the secondary aggregation processing includes packaging and serialization; the secondary aggregation processing of the aggregated data according to a preset data structure based on a lock-free design approach to obtain an audit message includes: Based on the preset data structure, the corresponding atomic variable values ​​are packaged and serialized in batches through each atomic variable object to obtain serialized data corresponding to each pre-aggregation key as the audit message corresponding to each pre-aggregation key; After the grouping, each of the atomic variables is initialized.

[0009] In some embodiments, based on the preset data structure, periodically and batch-packaging and serializing the corresponding atomic variable values ​​through each atomic variable object to obtain serialized data corresponding to each pre-aggregated key includes: According to the preset data structure, periodically batch-combine each of the atomic variable values ​​into a corresponding protocol buffer object; Based on a preset time interval, a corresponding first producer is obtained according to the audit data type of the audit message, and a serialization method built into the protocol buffer is called through the first producer, and each of the protocol buffer objects is serialized respectively through the serialization method.

[0010] In some embodiments, the cache middleware further includes a first queue; The caching into the preset cache middleware includes: The first producer corresponding to each audit message type caches the audit message of the corresponding type into the first queue.

[0011] In some embodiments, the cache middleware includes a first consumer; the distributed message middleware includes a producer pool, a second queue, and a second consumer; Writing the audit message in the cache middleware into the preset distributed message middleware in an asynchronous manner includes: After consuming the audit message event in the first queue, the first consumer obtains the corresponding second producer instance from the pre-built producer pool according to the message event label; uses the asynchronous interface of the second producer instance to obtain the audit message and stores it in the second queue.

[0012] In some embodiments, the method further comprises one or more of the following: Item 1: Based on at least one preset audit message type, a second producer instance corresponding to each audit message type is generated in the producer pool using a hungry initialization method; Item 2: During the process of sending the audit message to the second queue, if a sending anomaly is detected, recording the anomaly log and asynchronously retrying to send the audit message until a preset retry limit is reached; wherein, the write-at-least-once semantics of the audit message are guaranteed; Item 3: The preset data structure includes a first message class and a second message class; the first message class and the second message class each include a plurality of indicator fields; the protocol buffer object includes the first message class and the second message class, and field values ​​of the plurality of indicator fields in the first message class and the second message class; The performing secondary aggregation processing on the aggregated data according to the preset data structure to obtain an audit message includes: Determining a message header of the audit message according to a plurality of indicator fields and corresponding field values ​​in the first message class; The message data portion of the audit message is determined based on a plurality of indicator fields and corresponding field values ​​in the second message class.

[0013] In a second aspect, an embodiment of the present application provides a data reporting toolkit, the data reporting toolkit comprising: a timer interface, a capacity timer, and a distributed message middleware; The timer interface is used for being called by an external client to collect audit data from a preset database; The capacity timer is used to perform pre-aggregation processing on the audit data according to preset fields in accordance with preset rules to obtain aggregated data; based on a lock-free design method, perform secondary aggregation processing on the aggregated data according to a preset data structure to obtain an audit message, and cache it in a preset cache middleware; The distributed messaging middleware is used to read the audit message from the cache middleware in an asynchronous manner; wherein, the distributed messaging middleware is also used for an external data audit system to asynchronously consume the audit data.

[0014] In a third aspect, an embodiment of the present application provides a data audit platform, the data audit platform comprising: a data reporting toolkit and a data audit system; The data reporting toolkit is used to implement a data reporting method provided in the first aspect of this application; The data audit system is used to asynchronously consume audit messages from the data reporting toolkit to audit the audit messages.

[0015] The embodiments of the present application have the following beneficial effects: The method of the present application includes: responding to an external client calling a preset timer interface to collect audit data from a preset database using the timer interface; pre-aggregating the audit data according to preset fields according to preset rules to obtain aggregated data; based on a lock-free design method, performing secondary aggregation processing on the aggregated data according to a preset data structure to obtain an audit message, and caching it in a preset cache middleware; writing the audit message in the cache middleware into a preset distributed message middleware in an asynchronous manner; wherein the distributed message middleware is used for an external data audit system to asynchronously consume audit data from it. The use of a lock-free design in the present application can improve system performance, and the use of a cache middleware to cache data can decouple processes and implement process control of reported data, and finally the audit message is stored in the distributed message middleware to wait for asynchronous consumption by the external data audit system. In this way, the problems of low performance, high coupling and poor flexibility of the existing data audit system can be effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A framework diagram of the data audit platform according to an embodiment of the present application is shown; Figure 2 A flow chart of the data reporting method according to an embodiment of the present application is shown; Figure 3 A second flow chart of the data reporting method according to an embodiment of the present application is shown; Figure 4 A third flow chart of the data reporting method according to an embodiment of the present application is shown.

[0018] Description of main component symbols: 100-Data reporting toolkit; 200-Data audit system; 300-Database. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0020] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0021] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.

[0022] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.

[0023] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0024] In order to solve the problems of single reporting indicators, low performance, high coupling and poor flexibility of existing data audit systems / monitoring systems, this application provides a data reporting method, a data reporting toolkit and a data audit platform.

[0025] First, the embodiment of the present application provides a data audit platform, which is exemplary as follows: Figure 1As shown, the data audit platform includes: a data reporting toolkit 100 and a data audit system 200. The data audit system 200 is connected to a database 300 through the data reporting toolkit 100; the database 300 is used to store audit data.

[0026] A data reporting toolkit 100 (also referred to as SDK), which is used to read audit data from a database and implement the data reporting method according to the embodiment of the present application; The data audit system consumes audit messages from the data reporting toolkit to audit them. Typically, audit data includes log files. Binlogs are database files used to record all write operations (such as INSERT, UPDATE, and DELETE). The monitoring system analyzes the Binlog to track data changes for business auditing or link monitoring.

[0027] The data reporting method is described below with reference to some specific embodiments.

[0028] Figure 2 A flow chart of the data reporting method according to an embodiment of the present application is shown. The data reporting method according to the embodiment of the present application is suitable for reporting audit data to a data audit system or a monitoring system.

[0029] Exemplarily, the data reporting method includes the following steps: S100 responds to an external client calling a preset timer interface to collect audit data from a preset database using the timer interface.

[0030] The embodiment of the present application provides a complete and unified set of interfaces for external data auditing systems to call. Exemplarily, the timer interface includes a data collection interface and a web application reporting interface; the data collection interface is used to collect audit data from a preset database when called. For example, the database path is pre-configured in the configuration file, the data collection interface reads the database path from the configuration file, and obtains audit data based on the database path. The web application is used to report the click status of each section of the current page, the application response speed, record user behavior, etc. Among them, the current page is the client's page.

[0031] This application only requires the client to call a unified interface, which is less coupled with the client code, and can easily integrate the SDK of this application into the real-time computing chain. The client includes but is not limited to the data audit system.

[0032] S200 , pre-aggregating the audit data according to preset rules and preset fields to obtain aggregated data.

[0033] In order to reduce the amount of network IO transmission data, in the embodiment of the present application, the audit data is pre-aggregated. Pre-aggregation includes but is not limited to merging the audit data according to preset rules based on preset fields contained in the audit data.

[0034] Exemplarily, the preset audit fields in the audit data include a data type, a data source table, and a data source library. The audit data is merged based on any one of the data type, the data source table, and the data source library.

[0035] In one embodiment, in step S200, pre-aggregating the audit data according to preset fields according to preset rules to obtain aggregated data includes: S210 , based on a preset time interval, determine a unique pre-aggregation key corresponding to each piece of audit data according to preset fields.

[0036] Exemplarily, the audit fields in the audit data include: audit data type, audit data set source table, audit data source library, audit data collection time (Binlog time when audit data is collected), and data full-link delay, as shown in Table 1.

[0037] Table 1 Audit data

[0038] The preset fields in audit data include the audit data type, audit data set source table, and audit data source library. Use any of these as a pre-aggregate key. For example, use the audit data type as a pre-aggregate key.

[0039] S220 , determining a unique corresponding atomic variable object for each pre-aggregation key.

[0040] It can be understood that when a pre-aggregate key is first obtained, an atomic variable object is generated for the pre-aggregate key, and the same atomic variable object is reused for subsequent pre-aggregate keys. In other words, before determining a unique atomic variable object for each pre-aggregate key, it is necessary to check whether a corresponding atomic variable object exists for the pre-aggregate key; if not, a corresponding atomic variable object is generated; if so, the atomic variable object is retrieved.

[0041] The method of this embodiment adopts a lock-free design, creating a unique atomic variable object for each pre-aggregated key and assigning the audit data corresponding to that key to the atomic variable. Later in the process, audit data for the same pre-aggregated key is pre-aggregated and reused within the same atomic variable object, thereby reducing unnecessary JVM (Java Virtual Machine) garbage collection.

[0042] In this application, the threads of atomic variables can be updated safely, and all operations on atomic variables (such as increment, add, etc.) use `compareAndSet` or the atomic methods provided by JDK to ensure concurrency safety, and high-concurrency writing can be achieved without explicit locking. In addition, this application can also read and reset the values ​​of atomic variables in batches at regular intervals, and package all currently existing keys and their atomic variable values ​​into Protobuf objects within a fixed time interval (such as once per second). After packaging, the atomic variable values ​​corresponding to these keys are reset to the initial values ​​(such as 0). Among them, Protobuf is an efficient and flexible data serialization protocol developed by Google, the full name of which is Protocol Buffers. Protocol Buffers is a structured data serialization mechanism used in communication protocols, data storage and other scenarios.

[0043] Furthermore, in order to facilitate the search for pre-aggregation keys and corresponding atomic variable objects, the present application uses a dictionary data structure to store pre-aggregation keys and corresponding atomic variable objects to obtain a pre-aggregation dictionary.

[0044] S230: Aggregate all audit data items with the same pre-aggregation key and assign the resulting aggregated data to the corresponding atomic variable object. In other words, add the audit data corresponding to the pre-aggregation key to the same atomic variable. Aggregation includes, but is not limited to, data merging.

[0045] For example, the audit data in Table 1 are pre-aggregated, and the audit data of the same audit data type are merged. That is, audit data 1 and audit data 3 corresponding to audit data type a in Table 1 are merged, and audit data 4 and audit data 5 corresponding to audit data type c in Table 1 are merged, to obtain the pre-aggregated Map (pre-aggregated dictionary) shown in Table 2.

[0046] Table 2 Pre-aggregation Map

[0047] Exemplarily, audit data 1 includes: audit data type 1, audit data set source table 1, audit data source library 1, audit data collection time 1, and data full-link delay 1; audit data 3 includes: audit data type 3, audit data set source table 3, audit data source library 3, audit data collection time 3, and data full-link delay 3; then the pre-aggregation key obtained after pre-aggregation processing is audit data type a, and the atomic variable value corresponding to audit data type a is audit data type 1, audit data set source table 1, audit data source library 1, audit data collection time 1, data full-link delay 1, audit data type 3, audit data set source table 3, audit data source library 3, audit data collection time 3, and data full-link delay 3. Exemplarily, the value of the atomic variable is represented by Java's atomic reference type (AtomicReference). The above data is first packaged to generate a Java object, and then the Java object is managed using Java's atomic reference; the atomic object will be shared by multiple threads when reported, so using atomic references can ensure concurrency safety and performance when performing aggregate calculations.

[0048] The metrics mentioned above can be calculated. For example, using Binlog time as the starting point, the SDK records the processing time at each subsequent node, ultimately calculating the latency from the source to the current node. Data latency = current time - Binlog time.

[0049] The SDK of this application is suitable for high-concurrency environments, where audit data is reported simultaneously by multiple threads. Therefore, the cached values ​​in the SDK use atomic variables, and multiple threads share one SDK to ensure thread safety and performance in high-concurrency environments.

[0050] S300, based on a lock-free design, performs secondary aggregation processing on the aggregated data according to a preset data structure to obtain an audit message, and caches it in a preset cache middleware.

[0051] Exemplarily, the secondary aggregation process includes, but is not limited to, packaging and serialization.

[0052] This application achieves process decoupling and reporting flow control by caching audit messages in the cache middleware, ensuring the stability of the SDK and the subsequent scalability of the code.

[0053] For example, lock-free design includes using mechanisms such as atomic operations and CAS (Compare-And-Swap) instructions to enable concurrent multi-threaded access to shared resources without relying on traditional mutex locks (such as synchronized and ReentrantLock), thereby avoiding performance bottlenecks caused by lock contention.

[0054] In one embodiment, based on a lock-free design approach, secondary aggregation processing is performed on aggregated data according to a preset data structure to obtain an audit message, including: S310, based on the preset data structure, regularly and batch-package and serialize the corresponding atomic variable values ​​through each atomic variable object to obtain serialized data corresponding to each pre-aggregation key as the audit message corresponding to each pre-aggregation key.

[0055] S320, after packaging, initializing each atomic variable value. After packaging, resetting the atomic variable value, that is, resetting the specific index value of each atomic variable object to the initial value.

[0056] In one embodiment, the cache middleware is a lock-free concurrent framework designed based on the first producer-first consumer (also known as the first queue consumer) model, and is mainly used to efficiently process the transmission of events or messages.

[0057] Based on the preset data structure, the corresponding atomic variable values ​​are packaged and serialized through each atomic variable object in a scheduled batch, and the serialized data corresponding to each pre-aggregated key is obtained, including: S311: According to the preset data structure, each atomic variable value is periodically batched and combined into a corresponding protocol buffer object. In other words, the key value is converted into the attribute value of the protocol buffer object. For example, the attribute value of the Protobuf object.

[0058] Exemplarily, the preset data structure includes a first message class and a second message class. The first message class and the second message class each include multiple indicator fields. The protocol buffer object includes the first message class and the second message class, as well as the field values ​​of the multiple indicator fields in the first message class and the second message class. The preset data structure of the present application allows for custom indicator fields, allowing for free configuration of the reported indicator fields.

[0059] Furthermore, the aggregated data is subjected to secondary aggregation processing according to the preset data structure to obtain an audit message, including: Determine a message header of the audit message according to the multiple indicator fields and corresponding field values ​​in the first message class; The message data portion of the audit message is determined according to the multiple indicator fields and corresponding field values ​​in the second message class.

[0060] S312: Based on a preset time interval, obtain the corresponding first producer according to the audit data type of the audit message, call the serialization method built into the protocol buffer through the first producer, and serialize each protocol buffer object respectively through the serialization method.

[0061] For example, the caching middleware uses the Disruptor. Developed by the UK-based foreign exchange trading firm LMAX, the Disruptor is a high-performance queue designed to address latency issues associated with in-memory queues (performance testing revealed latency comparable to I / O operations). A single-threaded system based on the Disruptor can support 6 million orders per second. In this application's audit SDK, it primarily serves as an internal caching middleware, assuming decoupling and flow control responsibilities.

[0062] The Disruptor is a lock-free concurrency framework designed based on the Disruptor producer-Disruptor consumer model. Protocol buffer objects are not limited to Google Protobuf objects.

[0063] Exemplarily, based on the preset data structure, the corresponding atomic variable values ​​are packaged and serialized through each atomic variable object in a scheduled batch, and serialized data corresponding to each pre-aggregated key is obtained, including: (1) The atomic variable values ​​corresponding to each pre-aggregated key are periodically batched and combined into corresponding Google Protobuf objects according to the preset data structure.

[0064] Exemplarily, the Protobuf object includes two message classes; each field in the first message class is used to generate a message header; each field in the second message class is used to generate a message data portion.

[0065] (2) Based on the data type of the audit message, the Disruptor producer of the corresponding data type is used to call the Protobuf built-in serialization method, and each Protobuf object is serialized using the serialization method. A Protobuf object is used to generate an audit message. Each Protobuf object includes an audit data form with a preset structure; the audit data form includes multiple preset audit indicators.

[0066] For example, the types of preset audit indicators include: Specific fields (used to generate the message header of a single piece of data), data delay, data link processing delay, data source type, audit data set source table, audit data source library, target type for data writing, target library for data writing, target table for data writing, audit data set source table, and audit data source library.

[0067] The audit message consists of a message header and message data. The message header contains fields such as data source type, audit data set source table, audit data source repository, data write target type, data write target repository, data write target table, audit data set source table, and audit data source repository. The message data contains fields such as data latency and data link processing latency.

[0068] Understandably, specific fields in the Protobuf object (the header of a single data entry), the data delay, and the data link processing delay are used as the audit message header. Fields such as the data source type, audit dataset source table, audit data source repository, data write target type, data write target repository, data write target table, audit dataset source table, and audit data source repository are used as the audit message data portion.

[0069] In one embodiment, the cache middleware further includes a first queue; and caching into a preset cache middleware includes: The first producer corresponding to each audit message type caches the audit message of the corresponding type into the first queue.

[0070] According to the type of audit message, multiple types of Disruptor producers are set. For example, the audit message types include delay audit message types, data volume audit message types and DDL audit message types, etc. Then the Disruptor producers include delay audit message producers, data volume audit message producers and DDL audit message producers. Among them, the audit message type information is included in the audit message. Exemplarily, the cache middleware is Disruptor, the first queue is the Disruptor queue; the first queue consumer is the Disruptor consumer; the Disruptor producer corresponding to each audit message type is used to cache the audit message to the Disruptor queue. Disruptor producers of different audit message types obtain audit messages of corresponding types and cache them to the Disruptor queue. The embodiment of the present application is based on the high-performance queue of the Disruptor, uses the LMAX Disruptor to build a ring buffer, realizes lock-free, multi-threaded safe message passing, and supports high-concurrency writing and consumption.

[0071] In this application, the audit data is cached and pre-aggregated, and the audit message is assembled, and the assembled audit message is sent to the cache middleware, waiting for the consumer to send it asynchronously to the distributed message middleware.

[0072] S400, writing the audit message in the cache middleware into the preset distributed message middleware in an asynchronous manner; Among them, the distributed message middleware is used for the external data audit system to asynchronously consume audit data.

[0073] Furthermore, the cache middleware includes a first queue consumer; the distributed message middleware includes a producer pool, a second queue and a second consumer.

[0074] The audit message in the cache middleware is written to the preset distributed message middleware in an asynchronous manner, including: The first consumer parses the message event tag in the audit message and obtains the corresponding processing logic from the event processor (also known as the consumer) based on the message event tag; the producer pool is used to provide asynchronous sending capabilities for each processing logic; the processing logic calls the second producer in the producer pool to realize asynchronous sending of the audit message to the second queue.

[0075] For example, this embodiment uses a dictionary Map (such as `HashMap` in Java or `std::map` in C++) to register the mapping between event types and corresponding handlers. Users call the registration method to bind events and handlers to the Map. When the event distribution logic receives an event, it searches the Map for the corresponding handler based on the event type and executes the processing logic in the handler.

[0076] Exemplarily, after consuming an audit message event in the first queue, the first consumer retrieves the corresponding second producer instance from a pre-built producer pool based on the message event tag. The first consumer then uses the second producer instance's asynchronous interface to retrieve the audit message and store it in the second queue. In one embodiment, the cache middleware is a Disruptor, which includes a Disruptor consumer. The distributed messaging middleware is not limited to Kafka, the producer pool is a Kafka producer pool, the second queue is a Kafka queue, and the second consumer is a Kafka consumer.

[0077] The Disruptor consumer obtains the corresponding processing logic from the preset event processing function Map based on the message event tag of the audit message; When processing audit data through processing logic, the corresponding Kafka producer will be obtained from the Kafka producer pool; the Kafka producer is used to send audit messages to the Kafka queue in an asynchronous manner.

[0078] like Figure 3 、 Figure 4As shown, in the embodiment of the present application, specifically, the Disruptor queue is used to receive audit messages; the Disruptor consumer is used to parse the message event tag in the message, and obtain the corresponding processing logic (for example, the functional logic provided by the asynchronous interface) from the event processor mapping table according to the message event tag; the Kafka producer pool is used to provide asynchronous sending capabilities for each processing logic; the processing logic calls the Kafka producer to asynchronously send the audit message to the specified Topic (Kafka queue). Among them, the Kafka producer pool adopts an object pool management strategy, including initialization, reuse, and recycling mechanisms. Event tags include but are not limited to: data volume audit, DDL operation, permission change, login and logout. It can be understood that the Disruptor consumer includes delayed message consumers, data volume message consumers, and DDL message consumers.

[0079] This application is based on an event-driven architecture and dynamic routing. Specifically, audit messages contain tags (message event tags), such as "data_volume_audit" and "ddl_operation." Disruptor consumers dynamically search for corresponding processing functions in a map based on these tags. This enables decoupling of event distribution and logical plug-in expansion.

[0080] Furthermore, the method also includes: based on at least one preset audit message type, a second producer instance corresponding to each audit message type is generated in the producer pool using a hungry initialization method. Exemplarily, based on at least one preset audit message type, a Kafka producer instance corresponding to each audit message type is generated in the Kafka producer pool using a hungry initialization method. The Kafka producer pool uniformly manages multiple Kafka producer instances. In particular, the audit message type is configured in advance when using the SDK in the embodiment of the present application.

[0081] After consuming the audit message events in the Disruptor queue, the Disruptor queue consumer obtains the corresponding Kafka producer instance from the pre-built Kafka producer pool according to the message event label; uses the asynchronous interface of the Kafka producer instance to obtain the audit message and stores it in the Kafka queue. In the real-time computing link, the SDK, as the core component for audit data collection and reporting, must have high availability, low intrusion and strong fault tolerance. The distributed message middleware Kafka is used to achieve data decoupling and asynchronous transmission between systems. However, in actual use, the Kafka producer may fail to send data due to network failures, service unavailability, partition inaccessibility, etc. Therefore, the SDK needs to have a mechanism to deal with these failures to ensure that even in the case of partial failure, the data can be delivered to the target system (Kafka in the embodiment of this application) as much as possible, thereby improving the reliability and robustness of the overall SDK.

[0082] In one embodiment, to build a highly reliable, low-latency audit SDK, if a sending anomaly is detected while sending an audit message to a second queue, the anomaly is logged and asynchronous retries are performed until a preset retry limit is reached, thereby ensuring at-least-once write semantics for audit messages. Exemplarily, if a sending anomaly is detected while sending an audit message to a Kafka queue, the anomaly is logged and asynchronous retries are performed until a preset retry limit is reached.

[0083] Asynchronous retry: When an audit message fails to be reported to Kafka, it will not block the main thread or the current task, but will retry sending it through an independent thread or event-driven mechanism. This prevents a single failure from affecting the performance of the main process and increases the likelihood that the data will eventually be delivered.

[0084] For example, when the SDK calls the send() method of a Kafka producer, if an exception (such as a TimeoutException, NetworkException, or SerializationException) occurs, it captures the exception and logs it. The log includes the exception type, failed packet information (such as key, topic, and timestamp), number of attempts, and stack trace. This facilitates subsequent troubleshooting and evaluation of the SDK's robustness and stability.

[0085] Furthermore, the SDK internally configures a maximum number of retries, for example, three. An interval (e.g., an exponential backoff strategy) is set between each retry to reduce transient pressure. If the maximum number of retries is reached without success, the request is abandoned and an alarm mechanism is triggered. This embodiment of the present application prevents resource exhaustion caused by infinite retries while also ensuring that failed requests are recovered as much as possible within a controllable range.

[0086] Furthermore, the write-at-least-once semantics of audit messages are guaranteed, including: The original data is not discarded before Kafka successfully sends it; a copy of the audit message data to be sent is retained; and after successful sending, the cached audit message copy data is cleared. This ensures that the audit message data is sent, but duplication may occur (i.e., the same audit message may be written multiple times). However, data loss is prevented (assuming Kafka itself is configured for persistence and has a reasonable number of replicas).

[0087] This application includes Kafka producer pool management, where all Kafka producers are managed by a unified object pool to avoid frequent creation and destruction. The corresponding Kafka producer is retrieved based on the audit message type, improving performance and preventing resource leaks. This application also includes an asynchronous send mechanism, where processing logic calls Kafka producers for asynchronous send, improving overall response speed. A callback or retry mechanism is supported to ensure data reliability.

[0088] The audit SDK proposed in this application features ease of use, asynchrony, low coupling (minimizing intrusion into the caller's code), and high flexibility (allowing for customizable metrics). This application provides a set of APIs, meaning that when using the SDK, the caller only needs to use a unified set of APIs and perform simple configuration (for example, configurations that map audit messages to Kafka producer logic and database metadata that the SDK needs to collect). The SDK service provider then matches the implementation corresponding to the configuration and reports the data in a purely asynchronous manner, minimizing impact on ETL performance and improving performance. ETL performance refers to the efficiency and stability of ETL (Extract, Transform, Load) tools during data processing. Key performance indicators for ETL tools include processing speed, resource consumption, error rate, and stability.

[0089] The present application provides a data reporting toolkit. Exemplarily, the data reporting toolkit includes: a timer interface, a capacity timer, and a distributed message middleware.

[0090] Timer interface, used for external clients to call to collect audit data from the preset database; The capacity timer is used to pre-aggregate the audit data according to preset rules and fields to obtain aggregated data; based on a lock-free design, the aggregated data is secondary aggregated according to the preset data structure to obtain the audit message, which is then cached in the preset cache middleware; The distributed message middleware is used to read audit messages from the cache middleware in an asynchronous manner; the distributed message middleware is also used for the external data audit system to consume audit data asynchronously.

[0091] In some embodiments, the capacity timer includes a pre-aggregator.

[0092] Pre-aggregators for: Based on the preset time interval, determine the unique pre-aggregation key corresponding to each audit data according to the preset fields; Determine a unique corresponding atomic variable object for each pre-aggregation key; Aggregate all audit data with the same pre-aggregation key and assign the obtained aggregated data to the corresponding atomic variable object.

[0093] Furthermore, the pre-aggregator is specifically used to: find out whether there is a corresponding atomic variable object for the pre-aggregation key; if not, generate a corresponding atomic variable object; if it exists, obtain the atomic variable object.

[0094] Furthermore, the secondary aggregation process includes packet grouping and serialization; and the capacity timer includes a message assembler.

[0095] Message assembler, used to: Based on the preset data structure, the corresponding atomic variable values ​​are packaged and serialized through each atomic variable object in batches at regular intervals to obtain the serialized data corresponding to each pre-aggregation key as the audit message corresponding to each pre-aggregation key; After packing, initialize the value of each atomic variable.

[0096] Furthermore, the message combiner is specifically used to: According to the preset data structure, each atomic variable value is combined into the corresponding protocol buffer object in batches at regular intervals; Based on the preset time interval, the corresponding first producer is obtained according to the audit data type of the audit message, the serialization method built into the protocol buffer is called through the first producer, and each protocol buffer object is serialized respectively through the serialization method.

[0097] Furthermore, the capacity timer includes a first producer, and the cache middleware further includes a first queue.

[0098] First producer for: According to each audit message type, the audit message of the corresponding type is cached in the first queue.

[0099] Furthermore, the data reporting toolkit also includes distributed messaging middleware, and the cache middleware also includes a first consumer. The distributed messaging middleware includes a producer pool, a second queue, and a second consumer. After consuming the audit message event in the first queue, the first consumer obtains the corresponding second producer instance from the pre-built producer pool based on the message event tag. The second producer instance uses an asynchronous interface to obtain the audit message and stores it in the second queue.

[0100] The data reporting toolkit also includes an initialization module. The initialization module is used to generate a second producer instance corresponding to each audit message type in the producer pool using a hungry initialization method based on at least one preset audit message type; The data reporting toolkit also includes a retry module, which is used to record an exception log if a sending exception is detected during the process of sending the audit message to the second queue, and to asynchronously retry sending the audit message until the preset retry limit is reached; among which, the semantics of writing the audit message at least once are guaranteed.

[0101] Furthermore, the preset data structure includes a first message class and a second message class; the first message class and the second message class each include a plurality of indicator fields; the protocol buffer object includes the first message class and the second message class, and field values ​​of the plurality of indicator fields in the first message class and the second message class; Capacity timer, also used for: Determine a message header of the audit message according to the multiple indicator fields and corresponding field values ​​in the first message class; The message data portion of the audit message is determined according to the multiple indicator fields and corresponding field values ​​in the second message class.

[0102] This application has the following advantages: (1) Key-based atomic variable caching mechanism; supports key aggregation of any dimension; dynamically expands the key set and allocates atomic variables on demand; avoids global locks and improves concurrency performance.

[0103] (2) Lock-free micro-aggregation + timed packaging + asynchronous flushing pipeline; there is no synchronization block or wait / notify logic in the overall process; the packaging and reporting logic are decoupled through the Disruptor; and end-to-end asynchronous, non-blocking, high-performance audit collection is achieved.

[0104] (3) Lightweight access + automatic reset mechanism; the caller only needs to call the unified API to update the atomic variable corresponding to the key; the capacity timer automatically manages packaging and reset without interfering with business logic.

[0105] It can be understood that the device of this embodiment corresponds to the data reporting method of the above embodiment, and the optional items in the above embodiment are also applicable to this embodiment, so they will not be described again here.

[0106] The present application also provides a terminal device. Exemplarily, the terminal device includes a processor and a memory, wherein the memory stores a computer program, and the processor runs the computer program to enable the terminal device to execute the above-mentioned data reporting method or the functions of each module in the above-mentioned data reporting toolkit.

[0107] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0108] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving an execution instruction.

[0109] This application also provides a computer-readable storage medium for storing the computer program used in the terminal device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0110] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0111] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0112] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0113] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A data reporting method, characterized in that: The method comprises: In response to an external client calling a preset timer interface, collecting audit data from a preset database using the timer interface; According to preset rules, the audit data is pre-aggregated according to preset fields to obtain aggregated data; Based on a lock-free design, the aggregated data is subjected to secondary aggregation processing according to a preset data structure to obtain an audit message, which is then cached in a preset cache middleware; The audit message in the cache middleware is written into the preset distributed message middleware in an asynchronous manner; the distributed message middleware is used for an external data audit system to asynchronously consume the audit data.

2. The data reporting method according to claim 1, characterized in that: The pre-aggregating the audit data according to preset rules and preset fields to obtain aggregated data includes: Based on a preset time interval, determining a unique pre-aggregation key corresponding to each piece of the audit data according to a preset field; Determine a unique corresponding atomic variable object for each of the pre-aggregation keys; Aggregate all audit data with the same pre-aggregation key, and assign the obtained aggregated data to the corresponding atomic variable object.

3. The data reporting method according to claim 2, characterized in that: Determining a unique corresponding atomic variable object for each of the pre-aggregated keys includes: Check whether there is a corresponding atomic variable object for the pre-aggregation key; if not, generate a corresponding atomic variable object; if so, obtain the atomic variable object.

4. The data reporting method according to claim 3, characterized in that: The secondary aggregation processing includes packaging and serialization; the lock-free design method is based on the preset data structure to perform secondary aggregation processing on the aggregated data to obtain an audit message, including: Based on the preset data structure, the corresponding atomic variable values ​​are packaged and serialized in batches through each atomic variable object to obtain serialized data corresponding to each pre-aggregation key as the audit message corresponding to each pre-aggregation key; After the grouping, each of the atomic variables is initialized.

5. The data reporting method according to claim 4, characterized in that: The method of periodically and batch-packaging the corresponding atomic variable values ​​through each atomic variable object based on the preset data structure and serializing them to obtain serialized data corresponding to each pre-aggregated key includes: According to the preset data structure, periodically batch-combine each of the atomic variable values ​​into a corresponding protocol buffer object; Based on a preset time interval, a corresponding first producer is obtained according to the audit data type of the audit message, and a serialization method built into the protocol buffer is called through the first producer, and each of the protocol buffer objects is serialized respectively through the serialization method.

6. The data reporting method according to claim 5, characterized in that: The cache middleware further includes a first queue; The caching into the preset cache middleware includes: The first producer corresponding to each audit message type caches the audit message of the corresponding type into the first queue.

7. The data reporting method according to claim 6, characterized in that: The cache middleware includes a first consumer; the distributed message middleware includes a producer pool, a second queue and a second consumer; Writing the audit message in the cache middleware into the preset distributed message middleware in an asynchronous manner includes: The first consumer parses the message event tag in the audit message and obtains corresponding processing logic from the event processor according to the message event tag; wherein the producer pool is used to provide asynchronous sending capabilities for each processing logic; the processing logic calls the second producer in the producer pool to implement asynchronous sending of the audit message to the second queue; Specifically, after consuming the audit message event in the first queue, the first consumer obtains the corresponding second producer instance from the pre-built producer pool according to the message event label; uses the asynchronous interface of the second producer instance to obtain the audit message and stores it in the second queue.

8. The data reporting method according to claim 7, characterized in that: The method may further comprise one or more of the following: Item 1: Based on at least one preset audit message type, a second producer instance corresponding to each audit message type is generated in the producer pool using a hungry initialization method; Item 2: During the process of sending the audit message to the second queue, if a sending anomaly is detected, recording the anomaly log and asynchronously retrying to send the audit message until a preset retry limit is reached; wherein, the write-at-least-once semantics of the audit message are guaranteed; Item 3: The preset data structure includes a first message class and a second message class; the first message class and the second message class each include a plurality of indicator fields; the protocol buffer object includes the first message class and the second message class, and field values ​​of the plurality of indicator fields in the first message class and the second message class; The performing secondary aggregation processing on the aggregated data according to the preset data structure to obtain an audit message includes: Determining a message header of the audit message according to a plurality of indicator fields and corresponding field values ​​in the first message class; The message data portion of the audit message is determined based on a plurality of indicator fields and corresponding field values ​​in the second message class.

9. A data reporting toolkit, characterized in that: The data reporting toolkit includes: a timer interface, a capacity timer and a distributed message middleware; The timer interface is used for being called by an external client to collect audit data from a preset database; The capacity timer is used to perform pre-aggregation processing on the audit data according to preset fields in accordance with preset rules to obtain aggregated data; based on a lock-free design method, perform secondary aggregation processing on the aggregated data according to a preset data structure to obtain an audit message, and cache it in a preset cache middleware; The distributed messaging middleware is used to read the audit message from the cache middleware in an asynchronous manner; wherein, the distributed messaging middleware is also used for an external data audit system to asynchronously consume the audit data.

10. A data audit platform, characterized in that: The data audit platform includes: a data reporting toolkit and a data audit system; The data reporting toolkit is used to implement the data reporting method according to any one of claims 1 to 8; The data audit system is used to asynchronously consume audit messages from the data reporting toolkit to audit the audit messages.

Citation Information

Patent Citations

  • Log processing method and device, server and computer readable storage medium

    CN112347165A

  • User behavior data auditing method and system based on block chain

    CN114372296A

  • Data auditing monitoring system

    CN114584600A

  • Configuration method for describing streaming statistical operation mode

    CN116561196A

  • Calculation method and device of index data, equipment, storage medium and program product

    CN118569733A

Cited By

  • Data auditing method, device and equipment and computer readable storage medium

    CN120811931A

  • Data auditing method, device and apparatus, and computer-readable storage medium

    CN120811931B