A data accuracy monitoring system and method

By obtaining data information from the database cluster and real-time data warehouse in the data monitoring system, and combining the log information comparison of the log information of the log collection module, the problem of data loss and positioning difficulties in the existing technology is solved, and the accuracy of data monitoring is improved.

CN115718690BActive Publication Date: 2025-07-18ZHOUPU DATA TECH NANJING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211512387.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-07-18
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

The prior art cannot effectively locate data loss problems in data monitoring, resulting in a decrease in the accuracy of data monitoring.

Method used

Data information is determined from the database cluster and real-time data warehouse through the data statistics module, and log information is collected in real time with the log collection module. Data analysis module is used to compare data and log information, and data statistics and log statistics results are determined to improve monitoring accuracy.

Benefits of technology

It realizes the identification of data loss positioning and data inconsistency problems, and improves the accuracy of data monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115718690B_ABST
    Figure CN115718690B_ABST
Patent Text Reader

Abstract

The present invention discloses a data accuracy monitoring system and method, including: a data statistics module, configured to determine first data information of a first table to be counted in a tenant database to be counted from a database cluster included in a production system; and determine second data information of a second table to be counted corresponding to the first table to be counted from a real-time data warehouse included in the production system; a log collection module, configured to collect in real time log information included in each node on a data link; and a data analysis module, configured to determine data statistics result information by comparing the first data information and the second data information determined by the data statistics module; and determine log statistics result information by comparing the log information of each node collected by the log collection module. By comparing the first data information and the second data information through the data analysis module to determine data statistics result information, and by comparing the log information of each node to determine log statistics result information, the accuracy of data monitoring is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of data processing, and in particular, to a data accuracy monitoring system and method. Background Art

[0002] With the development of the data processing technology field, the need for processing a large amount of data is increasing continuously. When processing a large amount of data, it is necessary to monitor the large amount of data in order to discover problems such as data loss or data inconsistency in the large amount of data.

[0003] In the prior art, the accuracy monitoring of data usually pays more attention to the accuracy of real-time data and the stability of the data link, etc., and adopts methods such as automatic expansion and primary and standby links to improve the stability of the data link. However, when the primary data link fails and the standby data link is started, data loss may occur. However, the prior art pays more attention to the accuracy of real-time data and cannot locate the lost data, resulting in a decrease in the accuracy of data monitoring. Therefore, how to improve the accuracy of data monitoring is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0004] The present invention provides a data accuracy monitoring system and method, which improves the accuracy of data monitoring.

[0005] In a first aspect, an embodiment of the present invention provides a data accuracy monitoring system, including:

[0006] A data statistics module, configured to determine first data information of a first table to be statistically analyzed in a tenant database to be statistically analyzed from a database cluster included in a production system; and determine second data information of a second table to be statistically analyzed corresponding to the first table to be statistically analyzed from a real-time data warehouse included in the production system;

[0007] A log collection module, configured to collect in real time log information included in each node on a data link;

[0008] A data analysis module, configured to be respectively connected to the data statistics module and the log collection module; determine data statistics result information by comparing the first data information and the second data information determined by the data statistics module; and determine log statistics result information by comparing the log information of each node collected by the log collection module;

[0009] Wherein, the number of the database clusters is one or more, each of the database clusters includes one or more tenant databases, and each tenant database includes one or more first tables to be statistically analyzed; the second table to be statistically analyzed is a table obtained by transmitting the first table to be statistically analyzed through the data link to the real-time data warehouse.

[0010] Second aspect, an embodiment of the present invention provides a method for monitoring data accuracy, which is applied to the data accuracy monitoring system described in the first aspect, and includes:

[0011] Determine the first data information of the first table to be statistically analyzed in the tenant database to be statistically analyzed from the database cluster included in the production system through the data statistics module;

[0012] Determine the second data information of the second table to be statistically analyzed corresponding to the first table to be statistically analyzed from the real-time data warehouse included in the production system through the data statistics module;

[0013] Collect the log information included in each node on the data link in real time through the log collection module;

[0014] Compare the first data information and the second data information determined by the data statistics module through the data analysis module to determine the data statistics result information;

[0015] Compare the log information of each node collected by the log collection module through the data analysis module to determine the log statistics result information.

[0016] In the technical solution of the embodiment of the present invention, the first data information of the first table to be statistically analyzed and the second data information of the second table to be statistically analyzed corresponding to the first table to be statistically analyzed are determined through the data statistics module; the log information included in each node on the data link is collected in real time through the log collection module; the data statistics result information is determined by comparing the first data information and the second data information through the data analysis module, and the log statistics result information is determined by comparing the log information of each node through the data analysis module, which improves the accuracy of monitoring data.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a schematic structural diagram of a data accuracy monitoring system provided in Embodiment 1 of the present invention;

[0020] Figure 2 is a schematic structural diagram of a production system provided in Embodiment 1 of the present invention;

[0021] Figure 3 is a schematic structural diagram of a data accuracy monitoring system provided in Embodiment 2 of the present invention;

[0022] Figure 4 is a schematic structural diagram of a multi-tenant database data accuracy monitoring system provided in Embodiment 3 of the present invention;

[0023] Figure 5 is a schematic structural diagram of a data statistics sub-module provided in Embodiment 3 of the present invention;

[0024] Figure 6 is a flowchart of a data accuracy monitoring method provided in Embodiment 4 of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] It can be understood that before using the technical solutions disclosed in the embodiments of the present invention, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to users and user authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0028] Embodiment 1

[0029] Figure 1 is a schematic structural diagram of a data accuracy monitoring system provided in Embodiment 1 of the present invention, and this embodiment is applicable to the situation of realizing data accuracy monitoring. As Figure 1As shown, a data accuracy monitoring system 10 includes:

[0030] A data statistics module 11, configured to determine first data information of a first table to be statistically analyzed in a tenant database to be statistically analyzed from a database cluster included in a production system; and determine second data information of a second table to be statistically analyzed corresponding to the first table to be statistically analyzed from a real-time data warehouse included in the production system;

[0031] A log collection module 12, configured to collect in real time log information included in each node on a data link;

[0032] A data analysis module 13, configured to be respectively connected to the data statistics module 11 and the log collection module 12; determine data statistics result information by comparing the first data information and the second data information determined by the data statistics module 11; and determine log statistics result information by comparing the log information of each node collected by the log collection module 12;

[0033] Wherein, the number of database clusters is one or more, each database cluster includes one or more tenant databases, and each tenant database includes one or more first tables to be statistically analyzed; the second table to be statistically analyzed is a table obtained by transmitting the first table to be statistically analyzed through a data link to the real-time data warehouse.

[0034] Wherein, the production system may refer to an information system that supports daily business operations under normal circumstances. Figure 2 FIG. 17 is a schematic structural diagram of a production system provided in Embodiment 1 of the present invention. The data accuracy monitoring system 10 in the embodiments of the present invention can monitor the accuracy of data in the production system, so as to discover tenant databases or tables with data loss problems in the production system.

[0035] As Figure 2 shown, the production system may include one or more database clusters (such as Figure 2 each rectangular box on the far left in FIG. 22 may represent a database cluster), each database cluster may include one or more tenant databases (such as Figure 2 each cylindrical box in any one of the rectangular boxes on the far left in FIG. 22 may represent a tenant database), and each tenant database may include one or more database tables (Table).

[0036] The production system may also include a data link. The data link may include a relational database management system binary log consumer (Mysql Binlog Consumer), a message queue middleware (Message Queue) for storing binary logs, and a message consumer (Message Consumer) for processing binary logs. The nodes corresponding to these components can be equipped with corresponding logs and record the data records processed by themselves. Among them, the Mysql Binlog Consumer can be used to indicate the Binlog corresponding to the tenant database, and each tenant database can have a corresponding Binlog; the Message Queue can be used to store Binlogs, and by configuring the Message Queue, the tenant databases and / or tables that need to perform data accuracy analysis can be selected; the Message Consumer can be used to process Binlogs, such as performing format conversion on Binlogs, and the Message Consumer can indicate information such as the time when the Binlog is generated or the message type.

[0037] The production system may also include a real-time data warehouse. The real-time data warehouse pays more attention to the real-time nature of data. The tenant databases and / or tables in the original database cluster (i.e., Figure 2 the tenant databases and / or tables shown by any of the leftmost rectangular boxes) can be transmitted to the real-time data warehouse through the data link. The real-time data warehouse may have the tenant databases and / or tables in the original database cluster, and may also have tenant databases and / or tables corresponding to the tenant databases and / or tables in the original database cluster after being processed by the nodes in the data link.

[0038] The data statistics module 11 is used to determine the first data information of the first table to be statistically analyzed in the tenant database to be statistically analyzed from the database cluster included in the production system.

[0039] Among them, the tenant database to be statistically analyzed may refer to one or more tenant databases to be statistically analyzed in the database cluster of the production system. The first table to be statistically analyzed may refer to any database table to be statistically analyzed in any tenant database, and each tenant database may include one or more first tables to be statistically analyzed. There are no restrictions on the quantity of the first tables to be statistically analyzed and the tenant databases where the first tables to be statistically analyzed are located. Specifically, the first tables to be statistically analyzed can be selected in the tenant databases to be statistically analyzed according to actual application requirements.

[0040] The first data information may refer to the information indicated by the data in the first table to be statistically analyzed. For example, the first data information may include the data volume and business metrics. The data volume may be the total amount of data in the first table to be statistically analyzed, such as the total number of rows of data recorded in the first table to be statistically analyzed. The business metrics may be metrics related to the business scenario selected according to actual application requirements, such as sales amount or sales quantity and other metrics.

[0041] The manner in which the data statistics module 11 determines the first data information of the first table to be statistically processed in the tenant database to be statistically processed is not limited, as long as it can determine the first data information of the first table to be statistically processed in the tenant database to be statistically processed. For any first table to be statistically processed, the data statistics module 11 may provide an interface for the first table to be statistically processed to access the data accuracy monitoring system 10, so that the first table to be statistically processed can access the data accuracy monitoring system 10 through the interface provided by the data statistics module 11; when the first table to be statistically processed accesses the data accuracy monitoring system 10, the data statistics module 11 may abstract the total number of rows of the data included in the first table to be statistically processed into a data volume through the count() function, and abstract the cumulative value of the data included in the first table to be statistically processed into a business indicator through the sum() function; then, the abstracted data is processed through programming statements, and the data statistics module 11 executes the programming statements to determine the data volume and business indicator corresponding to the first table to be statistically processed, that is, to determine the first data information of the first table to be statistically processed. Among them, the programming statements can be set according to actual application needs, such as Structured Query Language (SQL).

[0042] The data statistics module 11 is used to determine the second data information of the second table to be statistically processed corresponding to the first table to be statistically processed from the real-time data warehouse included in the production system.

[0043] The second table to be statistically processed is in the real-time data warehouse included in the production system, and the second table to be statistically processed may refer to the table obtained by transmitting the first table to be statistically processed through the data link to the real-time data warehouse. During the process of transmitting the first table to be statistically processed through the data link to the real-time data warehouse, the nodes corresponding to the components in the data link may be equipped with corresponding logs and record the data records processed by themselves.

[0044] When the user performs operations such as modifying the first table to be statistically processed, the first table to be statistically processed can be processed by the nodes in the data link, such as recording the user's modification operations on the first table to be statistically processed, and then transmitting the processed first table to be statistically processed to the real-time data warehouse, which is the second table to be statistically processed in the real-time data warehouse, and the second table to be statistically processed is the database table after the modification operation on the first table to be statistically processed at the current moment.

[0045] The second data information may refer to the information indicated by the data in the second table to be statistically processed. For example, the second data information may include a data volume and a business indicator. The data volume may be the total amount of data in the second table to be statistically processed, such as the total number of rows of the data recorded in the second table to be statistically processed. The business indicator may be an indicator related to the business scenario selected according to actual application needs, such as indicators such as sales amount or sales quantity.

[0046] The manner in which the data statistics module 11 determines the second data information of the second pending statistical table corresponding to the first pending statistical table from the real-time data warehouse included in the production system is not limited, as long as the second data information of the second pending statistical table can be determined. For example, the manner in which the data statistics module 11 determines the second data information of the second pending statistical table can be the same as the manner of determining the first data information of the first pending statistical table, which will not be elaborated here.

[0047] The log collection module 12 is used to collect in real time the log information included in each node on the data link.

[0048] Among them, each node on the data link may refer to the nodes corresponding to components such as Mysql Binlog Consumer, Message Queue, and Message Consumer. For example, the first node corresponding to the binary log consumer, the second node corresponding to the message queue middleware storing the binary log, and the third node of the message consumer processing the binary log. These nodes can be equipped with corresponding logs and record the data records processed by themselves.

[0049] The log information included in each node may refer to the information indicated by the logs equipped by each node. The log information may include key information of the data processed or stored by each node, such as the primary key value recorded in the table. The logs indicated by the log information can maintain a unified format and directory structure.

[0050] Optionally, the log collection module 12 can also collect in real time the log information of the fourth node corresponding to the real-time data warehouse.

[0051] The manner in which the log collection module 12 collects in real time the log information included in each node on the data link, or collects in real time the log information of the fourth node corresponding to the real-time data warehouse is not limited. For example, the log information can be collected through the log file shipping tool Filebeat. Filebeat can monitor the logs of each node and track and read the files corresponding to the logs of each node.

[0052] The data analysis module 13 is used to connect to the data statistics module 11 and the log collection module 12 respectively.

[0053] The manner in which the data analysis module 13 connects to the data statistics module 11 and the log collection module 12 respectively is not limited, as long as the first data information and the second data information determined by the data statistics module 11, and the log information of each node collected by the log collection module 12 can be transmitted to the data analysis module 13. For example, it can be through electrical connection or through network connection, etc.

[0054] The data statistical result information may refer to the information indicated by the result of comparing the first data information and the second data information. The data statistical result information may indicate that no data loss occurs or data loss occurs from the first table to be statistically analyzed to the second table to be statistically analyzed.

[0055] The data analysis module 13 is used to determine the data statistical result information by comparing the first data information and the second data information determined by the data statistics module 11. For example, by using the Structured Query Language (SQL) to compare the data volume included in the first data information determined at the current moment with the data volume included in the second data information, or by comparing the business metrics included in the first data information determined at the current moment with the business metrics included in the second data information. When the data volume and business metrics included in the first data information are the same as those included in the second data information, it is determined that the data statistical result information indicates that no data loss occurs; otherwise, it is determined that the data statistical result information indicates that data loss occurs.

[0056] Another example is that the data analysis module 13 determines the data statistical result information by comparing the first data information and the second data information determined by the data statistics module 11 through the trend analysis method. The trend analysis method may refer to a method of analyzing the change trends of each time segment of the first data information or the second data information to discover data loss problems. Through the trend analysis method, the analysis problem of statistical errors caused by the inability of the real-time system to remain stationary can be processed.

[0057] The trend analysis method divides the data into segments according to the time length as needed. For example, the data of the most recent week can be divided into 7 time segments by days, and the duration of each time segment is 1 day. The first data information and the second data information are determined once within each time segment. When the statistics are performed on the 8th day, the data statistical result information corresponding to the 8th day can be determined based on the first data information and the second data information determined on the 8th day, combined with the first data information and the second data information determined in the previous 7-day time segments. And when the data accuracy monitoring system 10 is started for the first time, if a consistent baseline is determined to make the first data information and the second data information consistent, the time segment with data anomalies can be traced back through the trend analysis method, and the data statistical result information of this time segment indicates that data loss occurs.

[0058] The log statistical result information may refer to the information indicated by the result of comparing the log information of each node. The log statistical result information may indicate the nodes where data inconsistency problems occur in the data link.

[0059] The data analysis module 13 determines the log statistical result information by comparing the log information of each node collected by the log collection module 12. For example, the data analysis module 13 compares the business primary key, the business occurrence time, the binlog type, etc. to find the nodes with data inconsistency problems in the data link, and determines the nodes with data inconsistency problems as the log statistical result information.

[0060] It should be noted that the frequency of the data analysis module 13 to determine the data statistical result information and the log statistical result information is not limited. For example, the data analysis module 13 determines the data statistical result information once a day, or the data analysis module 13 determines the log statistical result information once a day, or the data analysis module 13 determines the log statistical result information once an hour.

[0061] The technical solution of the embodiment of the present invention determines the first data information of the first table to be counted and the second data information of the second table to be counted corresponding to the first table to be counted through the data statistics module; the log collection module collects the log information included in each node on the data link in real time; the data analysis module compares the first data information and the second data information to determine the data statistical result information, and the data analysis module compares the log information of each node to determine the log statistical result information. Through the data statistical result information, it can indicate whether data loss has occurred or data loss has occurred. When data loss occurs, the time segment of data loss can be determined through the trend analysis method; through the log statistical result information, it can indicate the nodes with data inconsistency problems in the data link; that is, when data loss occurs, the lost data can be located, improving the accuracy of data monitoring.

[0062] Furthermore, the data accuracy monitoring system 10 further includes:

[0063] A visualization processing module, which is used to connect to the data analysis module 13; and perform visualization processing on the data statistical result information and / or the log statistical result information determined by the data analysis module 13.

[0064] The visualization processing module can be connected to the data analysis module 13 by electrical connection or network connection, etc. The data analysis module 13 can transmit the data statistical result information and / or the log statistical result information to the visualization processing module, so that the visualization processing module performs visualization processing on the data statistical result information and / or the log statistical result information.

[0065] The visualization processing module is not limited to the way of visualizing the data statistics result information and / or the log statistics result information. For example, the visualization processing module generates an analysis report according to the data statistics result information and / or the log statistics result information, and displays the analysis report on the display screen according to the database cluster, the tenant database, and the database table. Among them, the analysis report can indicate the database cluster, the tenant database, and the database table where the data loss problem occurs. The display screen can be a monitoring dashboard.

[0066] At the same time, the visualization processing module can combine the data statistics result information and / or the log statistics result information determined in different time segments, and draw a data trend curve according to the time dimension to display the data accuracy. Among them, the data trend curve can refer to the curve formed by the data corresponding to different time segments. For example, the data trend curve can indicate the number of database clusters, tenant databases, and database tables with data loss problems in each time segment, etc.

[0067] Furthermore, the data accuracy monitoring system 10 may further include:

[0068] An exception alarm module, which can be used to send the database cluster, tenant database, and database table with data loss problems to the operation and maintenance and development personnel for processing through alarm.

[0069] Embodiment 2

[0070] Figure 3 FIG. is a schematic structural diagram of a data accuracy monitoring system provided according to Embodiment 2 of the present invention. This embodiment is based on the above Embodiment 1 to illustrate the specific structures of the data statistics module 11 and the log collection module 12. For technical details not described in detail in this embodiment, reference can be made to the above Embodiment 1.

[0071] In the embodiment of the present invention, the data statistics module 11 includes:

[0072] A data statistics sub-module 111 and a data storage sub-module 112;

[0073] The data statistics sub-module 111 is used to determine the first data information of the first table to be counted in the tenant database to be counted from the database cluster included in the production system; and determine the second data information of the second table to be counted corresponding to the first table to be counted from the real-time data warehouse included in the production system;

[0074] The data storage sub-module 112 is used to connect to the data statistics sub-module 111 and the data analysis module 13 respectively; store the first data information and the second data information determined by the data statistics sub-module 111; and send the first data information and the second data information to the data analysis module 13.

[0075] The connection manner between the data storage sub-module 112 and the data statistics sub-module 111 and the data analysis module 13 is not limited. For example, it can be electrically connected or connected through a network, etc.

[0076] The data storage sub-module 112 is connected to the data statistics sub-module 111. After the data statistics sub-module 111 determines the first data information and the second data information, the data statistics sub-module 111 can transmit the first data information and the second data information to the data storage sub-module 112, and the data storage sub-module 112 receives and stores the first data information and the second data information.

[0077] The data storage sub-module 112 can be an Online Analytical Processing Database (OLAP DB). Among them, OLAP is Online Analytical Processing, and OLAP can be used for data analysis. Through OLAP, information from multiple database clusters can be analyzed simultaneously, that is, the required data can be extracted and the data can be queried through OLAP so as to perform data analysis from different perspectives.

[0078] The data storage sub-module 112 is connected to the data analysis module 13. The data storage sub-module 112 can send the first data information and the second data information to the data analysis module 13 so that the data analysis module 13 can determine the data statistics result information according to the first data information and the second data information.

[0079] Further, the data statistics sub-module 111 includes:

[0080] The service interface unit 1111 is used to provide an interface for the first table to be statistically analyzed and the second table to be statistically analyzed to access the system.

[0081] The service interface unit 1111 can provide an interface for the first table to be statistically analyzed and the second table to be statistically analyzed to access the system, that is, the first table to be statistically analyzed and the second table to be statistically analyzed included in the production system can access the data accuracy monitoring system 10 through the service interface unit 1111, so that the data accuracy monitoring system 10 can obtain the first table to be statistically analyzed and the second table to be statistically analyzed, and perform subsequent processing on the first table to be statistically analyzed and the second table to be statistically analyzed, such as determining the first data information of the first table to be statistically analyzed, determining the second data information of the second table to be statistically analyzed, or comparing the first data information and the second data information, etc. Correspondingly, the tenant database or database cluster where the first table to be statistically analyzed or the second table to be statistically analyzed is located can also access the data accuracy monitoring system 10 through the interface provided by the service interface unit 1111.

[0082] The task queue unit 1112 is used to store the statistical requirement tasks waiting to be processed. The statistical requirement tasks are to determine the first data information of the first table to be statistically analyzed and / or determine the second data information of the second table to be statistically analyzed.

[0083] A statistical requirement task may refer to a task that needs to perform statistics on the data in any database table. For example, a statistical requirement task may be to determine the first data information of the first table to be statistically analyzed and / or determine the second data information of the second table to be statistically analyzed. The number of statistical requirement tasks is not limited and can be specifically determined according to actual application needs.

[0084] When the first table to be statistically analyzed or the second table to be statistically analyzed accesses the data accuracy monitoring system 10 through the service interface unit 1111, it indicates that it is necessary to determine the first data information of the first table to be statistically analyzed or the second data information of the second table to be statistically analyzed. Each first table to be statistically analyzed can correspond to a statistical requirement task, and each second table to be statistically analyzed can also correspond to a statistical requirement task. Each statistical requirement task can be stored in the task queue unit 1112 and waits to be processed in the task queue unit 1112.

[0085] The tenant database management unit 1113 is used to synchronize the configurations of each tenant database in each database cluster from the production system and group the tenant databases in the same database cluster into the same group.

[0086] The tenant databases in the production system are physically organized and are continuously adjusted according to the development of the Software as a Service (SaaS) business. The tenant database management unit 1113 can synchronize the configurations of each tenant database in each database cluster from the production system and group the tenant databases in the same database cluster into the same group. That is, the tenant database management unit 1113 can identify the tenant databases in each database cluster in the production system and group all the tenant databases belonging to the same database cluster that access the data accuracy monitoring system 10 into the same group.

[0087] The business abstraction unit 1114 is used to perform abstraction processing on the data in the first table to be statistically analyzed and / or the second table to be statistically analyzed.

[0088] The method of performing abstraction processing on the data in the first table to be statistically analyzed and / or the second table to be statistically analyzed is not limited. For example, if it is necessary to determine the first data information and / or the second data information through the task execution unit 1116 later, and the first data information and / or the second data information include the data volume and business metrics, then the business abstraction unit 1114 can abstract the data in the first table to be statistically analyzed and / or the second table to be statistically analyzed into the data volume and business metrics.

[0089] In the first table to be counted and / or the second table to be counted, the total number of rows of data included in the first table to be counted and / or the second table to be counted can be used as the data volume through the count() function; the accumulated value of the data included in the first table to be counted and / or the second table to be counted can be used as a business indicator through the sum() function. For example, the accumulated value of all sales amounts recorded in the first table to be counted and / or the second table to be counted can be used as a business indicator.

[0090] Through the service abstraction unit 1114, the symbolization of the data in the first table to be counted and / or the second table to be counted can be realized. When subsequently providing the first table to be counted and / or the second table to be counted and the functions required for statistical data, the statistics of the data volume and business indicators can be automatically realized.

[0091] The task execution management unit 1115 is used to manage the statistical requirement tasks for concurrent processing and / or manage the number of tenant databases connected to the system.

[0092] The task execution management unit 1115 is used to manage the statistical requirement tasks for concurrent processing, that is, the task execution management unit 1115 can manage the number of statistical requirement tasks for concurrent processing. When the number of tenant databases that need to be counted increases, the number of corresponding threads is expanded, so that the data statistics of more tenant databases can be processed concurrently, that is, more statistical requirement tasks can be processed concurrently. Furthermore, when the data in the production system is constantly changing, the statistics of the data will not be inaccurate due to too long statistical time, ensuring the timeliness of the statistical data.

[0093] The task execution management unit 1115 is used to manage the number of tenant databases connected to the system. When it is necessary to count the data in the first table to be counted or the second table to be counted in the tenant database, it indicates that there is a corresponding statistical requirement task for this tenant database, and the tenant database where the first table to be counted or the second table to be counted is located is connected to the data accuracy monitoring system 10. Correspondingly, when there is no first table to be counted or second table to be counted in the tenant database, it indicates that there is no database table that needs to be counted in this tenant database, that is, there is no corresponding statistical requirement task, and then this tenant database is disconnected from the data accuracy monitoring system 10. This can save the environmental resources operated by the data accuracy monitoring system 10 and at the same time save the connection overhead of the tenant database.

[0094] The task execution unit 1116 is used to concretize the data abstracted by the service abstraction unit 1114 into executable statements and determine the first data information and / or the second data information.

[0095] The task execution unit 1116 is configured to concretize the data abstracted by the service abstraction unit 1114 into executable statements, and determine the first data information and / or the second data information. It can be understood that the task execution unit 1116 can process the data abstracted by the service abstraction unit 1114 through programming statements, and the first data information and / or the second data information can be determined by executing the programming statements in the task execution unit 1116. Among them, the programming statements can be set according to actual application needs, such as Structured Query Language (SQL).

[0096] The result output unit 1117 is configured to send the first data information and / or the second data information to the data storage sub-module 112.

[0097] In the result output unit 1117, the first data information and / or the second data information can be stored in a buffer queue, and the first data information and / or the second data information stored in the buffer queue can be sent to the data storage sub-module 112 when needed. For example, the first data information and / or the second data information stored in the buffer queue can be batch-written to the data storage sub-module 112 through a timing script. Among them, the timing script can be set according to actual application needs, and is not specifically limited as long as the first data information and / or the second data information can be sent to the data storage sub-module 112.

[0098] Furthermore, the task execution management unit 1115 is specifically configured to:[[]]

[0099] Centrally manage the statistical requirement tasks for concurrent processing through a thread manager, and the number of statistical requirement tasks for concurrent processing has a linear relationship with the number of tenant databases;

[0100] Manage the number of tenant databases connected to the system through a database connection manager in combination with the tenant database management unit 1113.

[0101] Centrally manage the statistical requirement tasks for concurrent processing through a thread manager, and the number of statistical requirement tasks for concurrent processing has a linear relationship with the number of tenant databases. It can be understood that when the number of tenant databases increases, the corresponding number of threads is extended through the thread manager, so that more statistical requirement tasks can be processed concurrently. Among them, concurrent processing can adapt to the linear expansion of tenant databases in the production system, and the corresponding extended number of threads can ensure the timeliness of statistical data and will not cause errors due to the continuous change of data in the production system when the statistical time is too long.

[0102] Correspondingly, when the number of tenant databases decreases, the thread manager reduces the number of corresponding threads, enabling fewer statistical requirement tasks to be processed concurrently, thereby saving the environmental resources on which the data accuracy monitoring system 10 runs and also saving the connection overhead of the tenant databases.

[0103] The database connection manager combines with the tenant database management unit 1113 to manage the number of tenant databases in the connection system. It can be understood that flow limiting processing is performed on the number of tenant databases in the connection system. The tenant databases that need to execute statistical requirement tasks can be recorded through a HashMap data structure, and based on the configurations of the tenant databases in each database cluster in the production system recorded by the tenant database management unit 1113, the tenant databases that need to execute statistical requirement tasks in the same database cluster are identified as the same group. For any database cluster, the key in the HashMap data structure can be used as the identifier of the database cluster, and the value can be used as the number of tenant databases in the database cluster that are connected to the data accuracy monitoring system 10. The number of tenant databases in the database cluster that are connected to the data accuracy monitoring system 10 can be determined through the identifier of the database cluster. After determining the identifier of the database cluster, the database connection manager can connect the number of tenant databases corresponding to the value in the database cluster to the data accuracy monitoring system 10, thus realizing the management of the number of tenant databases in the connection system by the database connection manager combining with the tenant database management unit 1113.

[0104] Among them, there is no limitation on the identifier of the database cluster. For example, it can be a digital number, and each database cluster can have a unique corresponding digital number to facilitate the distinction of different database clusters. Flow limiting processing can ensure that the execution of statistical requirement tasks does not affect the performance of the production system.

[0105] Furthermore, the task execution management unit 1115 is specifically used for:

[0106] Releasing the connection between the tenant database without statistical requirement tasks and the system through the database connection manager;

[0107] Releasing threads by using the thread manager when there are no statistical requirement tasks.

[0108] When there are no statistical requirement tasks in the tenant database, release threads by using the thread manager and release the connection between the tenant database and the data accuracy monitoring system 10 through the database connection manager; when there are statistical requirement tasks in the tenant database, keep the threads active by using the thread manager and keep the connection between the tenant database and the data accuracy monitoring system 10 through the database connection manager, thereby saving the environmental resources on which the data accuracy monitoring system 10 runs and also saving the connection overhead of the tenant databases.

[0109] Further, the data analysis module 13 is specifically configured to:

[0110] Within a historical time period, determine the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for each historical time segment within the historical time period as the first error corresponding to the historical time segment;

[0111] Determine the first mean value and the first mean square deviation corresponding to the historical time period according to the first errors corresponding to each historical time segment;

[0112] Determine the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for the current time segment as the second error corresponding to the current time segment;

[0113] Determine the absolute value of the difference between the second error and the first mean value as the third error of the current time segment relative to the historical time period;

[0114] If the third error does not exceed the first mean square deviation, determine that the data statistical result information indicates that no data loss has occurred; if the third error exceeds the first mean square deviation, determine that the data statistical result information indicates that data loss has occurred;

[0115] Wherein, the historical time period includes multiple historical time segments of the same duration, and the duration of the current time segment is the same as that of the historical time segment.

[0116] When the data analysis module 13 compares the first data information and the second data information to determine the data statistical result information, it can be achieved through the trend analysis method. By using the first data information and the second data information determined for the current time segment, combined with the first data information and the second data information determined for each historical time segment within the historical time period, determine the data statistical result information corresponding to the current time segment.

[0117] Taking the historical time period as 1 week as an example, a week can be divided into 7 historical time segments according to the number of days, the duration of each historical time segment is 1 day, and the historical time period includes 7 historical time segments of the same duration. The duration of the current time segment is also 1 day, and the current time segment can be a time segment after the historical time period.

[0118] Specifically, within the historical time period, determine the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for each historical time segment within the historical time period as the first error corresponding to the historical time segment.

[0119] The values corresponding to the first data information for each historical time segment within the historical time period can be successively represented in chronological order as:

[0120] r1, r2, r3, r4, r5, r6, r7

[0121] The values corresponding to the second data information for each historical time segment within the historical time period can be successively represented in chronological order as follows:

[0122] w1, w3, w3, w4, w5, w6, w7

[0123] Then the first error corresponding to each historical time segment can be successively represented in chronological order as follows:

[0124] δ1 = |w1 - r1|, δ2 = |w2 - r2|, δ3 = |w3 - r3|,..., δ7 = |w7 - r7|

[0125] Based on the first error corresponding to each historical time segment, determine the first mean value and the first mean square deviation corresponding to the historical time period.

[0126] The calculation formula for the first mean value corresponding to the historical time period is:

[0127]

[0128] The calculation formula for the first mean square deviation corresponding to the historical time period is:

[0129]

[0130] Determine the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for the current time segment as the second error corresponding to the current time segment. The value corresponding to the first data information for the current time segment can be represented as r8, and the value corresponding to the second data information for the current time segment can be represented as w8. Then the calculation formula for the second error is:

[0131] δ8 = |w8 - r8|

[0132] Determine the absolute value of the difference between the second error and the first mean value as the third error of the current time segment relative to the historical time period. Then the calculation formula for the third error is:

[0133]

[0134] If the third error does not exceed the first mean square deviation, it is determined that the data statistical result information indicates that no data loss has occurred. Among them, the third error not exceeding the first mean square deviation can be expressed as:

[0135] σ8 ≤ σ

[0136] If the third error exceeds the first mean square deviation, it is determined that the data statistical result information indicates that data loss has occurred. Among them, the third error exceeding the first mean square deviation can be expressed as:

[0137] σ8 > σ

[0138] In an embodiment of the present invention, the log collection module 12 includes:

[0139] A log collection tool 121 for collecting log information of the first node corresponding to the binary log consumer; collecting log information of the second node corresponding to the message queue middleware storing the binary log; collecting log information of the third node corresponding to the message consumer processing the binary log; collecting log information of the fourth node corresponding to the real-time data warehouse;

[0140] A log storage tool 122 for respectively connecting to the log collection tool 121 and the data analysis module 13; storing the log information of each node collected by the log collection tool 121; and sending the log information to the data analysis module 13.

[0141] The log collection tool 121 may be a log file shipper tool Filebeat. Through Filebeat, the log information of the first node, the second node, the third node, and the fourth node can be collected.

[0142] The way that the log storage tool 122 connects to the log collection tool 121 and the data analysis module 13 may be electrical connection or network connection, etc.

[0143] By connecting to the log collection tool 121, the log storage tool 122 can obtain and store the log information of each node transmitted by the log collection tool 121. The log storage tool 122 may be a distributed search and analysis engine, such as ElasticSearch, to support both accuracy comparison and problem location capabilities. The storage of log information can be automatically managed in terms of lifecycle, such as the index template and lifecycle function of Elasticsearch, thereby reducing storage costs and operation and maintenance costs.

[0144] By connecting to the data analysis module 13, the log storage tool 122 can send the log information of each node to the data analysis module 13 so that the data analysis module 13 can compare the log information of each node to determine the log statistical result information.

[0145] Furthermore, the data analysis module 13 is specifically configured to:

[0146] Compare the log information of each node collected by the log collection module 12 at a set frequency to determine the problem node where data inconsistency occurs in the data link;

[0147] Determine the problem node as the log statistical result information.

[0148] The set frequency may refer to the frequency of determining the log statistical result information, and the set frequency can be determined according to actual application requirements. For example, the set frequency can be every hour. The problem node may refer to the node where data inconsistency problems occur on the data link.

[0149] The data analysis module 13 can compare the log information included in each node once every hour. For example, by comparing the business primary key, business occurrence time, binlog type, etc. indicated by the log information, it can find the problem nodes where data inconsistency problems occur on the data link, and determine the problem nodes with data inconsistency problems as the log statistical result information.

[0150] The technical solution of the embodiment of the present invention describes the specific structures of the data statistics module and the log collection module, and the specific functions of the data analysis module. The data analysis module can determine the data statistical result information corresponding to the current time segment through the first data information and the second data information corresponding to each historical time segment within the historical time period, and the first data information and the second data information corresponding to the current time segment; by selecting different current time segments and historical time segments, it can further determine the time segments where data loss occurs; the data analysis module can compare the log information included in each node at the set frequency to determine the log statistical result information, and the log statistical result information indicates the nodes where data inconsistency problems occur on the data link; that is, when data loss occurs, the lost data can be located, improving the accuracy of data monitoring.

[0151] Embodiment III

[0152] The embodiment of the present invention is an exemplary illustration of the above embodiments. The embodiment of the present invention proposes a multi-tenant database data accuracy monitoring system, which can monitor the data in a production system including multiple servers (i.e., multiple database clusters) and multiple tenants, and supports OLAP and real-time analysis simultaneously.

[0153] In a real-time production system, when facing TB-level data, it is very difficult to detect data loss and locate problems. Considering from the perspectives of cost and efficiency, this system uses statistical methods and data link tracing methods to conduct automated inspection, comprehensive analysis, and alarm of data accuracy from macroscopic, microscopic, and time dimensions, combined with analysis charts, so as to solve the problem of real-time data accuracy inspection in a multi-tenant scenario. It can find data accuracy problems in the T+1 cycle, and find the running status of tens of thousands of data sources within the H+1 cycle. It has low running costs, high efficiency, and is easy to use.

[0154] Among them, the period is the frequency of data monitoring; there are two strategies for monitoring: full volume and sampling. In the execution of the full-volume strategy, it is executed once a day on average, and correspondingly, we define this period as T+1. In the execution of the sampling strategy, it is executed once an hour on average, and correspondingly, we define this period as H+1.

[0155] Figure 4 It is a schematic structural diagram of a multi-tenant database data accuracy monitoring system provided in Embodiment 3 of the present invention. Figure 4 The monitoring system shown in is the multi-tenant database data accuracy monitoring system provided in Embodiment 3 of the present invention. Figure 4 The production system is also shown in, and the monitoring system can monitor the data accuracy of the production system.

[0156] Such as Figure 4 As shown, the multi-tenant database data accuracy monitoring system includes: statistics (i.e., the data statistics sub-module), log collection (i.e., the log collection tool), distributed search and analysis engine (i.e., the log storage tool), online analytical processing database (i.e., the data storage sub-module), analysis (i.e., the data analysis module), monitoring dashboard (i.e., the visualization processing module), and exception alarm (i.e., the exception alarm module).

[0157] Among them, statistics (i.e., the data statistics sub-module) is mainly used to meet various requirements of the data statistics step; the data is distributed in different database servers (i.e., the database cluster), and the SaaS service is running on it, and it is necessary to solve the problems of many tenant databases and tight production resources.

[0158] Statistics (i.e., the data statistics sub-module) is designed as a service-type task, which can meet various statistical requirements, thus providing the possibility for automated integration; in implementation, multi-threading is adopted to improve concurrency, and the concurrency ability depends on the size of the requirements, that is, how many database servers there are, so as to achieve a linear expansion ability equivalent to the SaaS service; a strict database operation framework is implemented at the bottom layer to strictly control the management of database connections and requests; the number of database connections is evaluated by the production system operation and maintenance party according to the tenant database capacity and business volume and gives the value allowed for the monitoring system to use; the database does not maintain long connections and only connects when needed, and releases immediately after processing the request; the database request only supports query requests to avoid affecting the production data; all requests only support sequential execution to avoid occupying too much production resources and causing production requests to block.

[0159] Figure 5 It is a schematic structural diagram of a data statistics sub-module provided in Embodiment 3 of the present invention. Specifically, the data statistics sub-module includes:

[0160] A service interface (i.e., service interface unit) for providing applications for statistical requirements, which can specify the clusters (i.e., database clusters) or tables (the first table to be statistically analyzed or the second table to be statistically analyzed) to be statistically analyzed on time and flexibly.

[0161] A task queue (i.e., task queue unit), where statistical requirements (i.e., statistical requirement tasks) enter the system and are uniformly queued for processing.

[0162] A tenant database management module (i.e., tenant database management unit). The tenant databases of the production system are physically organized and continuously adjusted according to the development of the SaaS business. The tenant database management module is responsible for synchronizing the tenant database configurations, identifying the configurations of the tenant databases, and grouping the tenant databases of the same database cluster (i.e., dividing the tenant databases in the same database cluster into the same group).

[0163] A business abstraction module (i.e., business abstraction unit). The business includes data volume statistics and business metric statistics. The business abstraction module abstracts the two types of statistics into count and sum, realizing the symbolization of statistical requirement descriptions, that is, providing the tables for storing corresponding business data and the statistical types. The system can automatically implement the statistics.

[0164] A task execution management module (i.e., task execution management unit). The task execution management module sets rules for ensuring both concurrency and flow control for task execution:

[0165] Concurrency ensures that the number of threads can be correspondingly extended in line with the linear expansion of the tenant databases in the production system, thus ensuring the timeliness of statistical data and preventing errors caused by continuous changes in production data when the statistical time is too long. A thread manager is established in the program to centrally manage concurrent tasks. The number of concurrent tasks is linearly related to the number of tenant databases (i.e., the statistical requirement tasks processed concurrently are centrally managed by the thread manager, and the number of statistical requirement tasks processed concurrently is linearly related to the number of tenant databases).

[0166] Flow control ensures that statistical tasks will not affect the performance of the production system. A HashMap data structure is managed in the program to record each tenant database during statistical operation (i.e., when executing statistical requirement tasks), and the tenant databases on the same database cluster are identified in combination with the tenant management module. The key value of the HashMap is the name of the cluster server (i.e., the identifier of the database cluster), and the value is the number of connections of the current tenant database to the database cluster. At the same time, a database connection manager is implemented to provide the database connections required for task execution, thus managing the number of connections of tenant databases (i.e., the number of tenant databases connecting to the system is managed by the database connection manager in combination with the tenant database management unit).

[0167] When tasks (i.e., statistical requirement tasks) keep coming in continuously, both the threads and database connections remain active. Once it is found that no more tasks are coming in, the thread manager actively releases the threads, and the database connection manager actively releases the database connections (i.e., the connection between the tenant database without statistical requirement tasks and the system is released through the database connection manager; when there are no statistical requirement tasks, the thread manager is used to release the threads). This can save the environmental resources used by the statistical system and also save the connection overhead of the tenant databases in the production system;

[0168] The task execution module (i.e., the task execution unit) is used to abstract the business into executable SQL and execute the SQL to obtain results (i.e., concretize the data processed by the business abstraction unit into executable statements and determine the first data information and / or the second data information);

[0169] The result output module (i.e., the result output unit). The statistical results executed by concurrent threads will be stored in a buffer queue, and a cyclic storage thread job will periodically batch-write the statistical results in the buffer queue to external storage (i.e., send the first data information and / or the second data information to the data storage sub-module).

[0170] Through analysis (i.e., the data analysis module), first compare the data volume of the tenant database (i.e., the data source) with the data volume of the real-time data warehouse (i.e., the target database). This is a data processing process at the ten-million level. Using a database with OLAP capabilities, the algorithm for this step is implemented in SQL, which is simple, fast, and has high execution efficiency; the comparison of business metrics is also implemented in the same way;

[0171] The data source is usually a MySQL-type database cluster, which has multiple servers and each server contains multiple tenant databases. The data in the tables of each tenant database is statistically analyzed, including the total data volume and business metrics, and written into a database that supports OLAP analysis. It is necessary to compare each table of multiple tenant databases to find out the tenant databases and tables with data loss problems;

[0172] When comparing business metrics, the metrics of core businesses can be selected, such as sales amount and sales quantity. By comparing the totals of business metrics, the comparison can be made by statistically comparing each tenant database and the storage tables corresponding to the business metrics, so as to find out the tenant databases and tables with inaccurate data problems.

[0173] Based on the results of the above processing, which reflect static data differences, it is necessary to further analyze using the trend analysis method in combination with historical data; the data accuracy problems caused by errors in the real-time system are reflected in the data as the reasons for analyzing the difference results; by comparing historical data, it can be distinguished whether this difference is a problem left over from history or a recent problem, which requires starting from statistics.

[0174] The trend analysis method is superimposed after the statistical analysis method to deal with the analysis problem of statistical errors caused by the inability of the real-time system to remain static; the trend analysis method divides the data into segments according to the time length as needed. For example, the data of the most recent week (i.e., the historical time period) can be divided into 7 time segments (i.e., historical time segments) by day. Data statistics will be performed once within each time segment, and the statistical and analysis results (i.e., the first data information, the second data information, and the data statistical result information) will be saved. When the statistics are performed on the 8th day, the results of the previous 7-day time segments can be combined to determine the data statistical result information corresponding to the 8th day.

[0175] It should be noted that the trend analysis method compares the difference values (i.e., the first error) at that time in each time period. For example, if about 100 data differences are found each time during the statistics from the 1st to the 7th day, then there are still about 100 differences during the statistics on the 8th day, which means no data is lost today. There are two sources of errors: one is the time difference in statistics. Because the data is constantly changing, there will be errors; the other is the dirty data accumulated in history, which requires backtracking to earlier time segments. If we ensure a consistent baseline when the monitoring system is first started, that is, the data in the tenant database is consistent with the data in the real-time data warehouse, then we can always backtrack to the time point when the anomaly occurred. Or rather, the problem of data loss can actually be discovered at that time point.

[0176] The log collection (i.e., the log collection tool) is designed as a general collection framework, and the logs maintain a unified format and directory structure; the logs are collected uniformly on a storage engine that supports full-text retrieval (i.e., the log storage tool), such as ElasticSearch, to support both accuracy comparison and problem location capabilities at the same time; open-source tools, such as Filebeat, are used for log collection; the storage of logs is managed automatically according to the lifecycle, such as the index template and lifecycle function of Elasticsearch, so as to reduce the storage cost and operation and maintenance cost.

[0177] By collecting logs (i.e., the log collection tool), the log information of each node in the data link is collected in real time. These nodes include the Mysql Binlog Consumer (i.e., the first node), the MessageQueue storing Binlog (i.e., the second node), the Message Consumer processing binlog (i.e., the third node), and the final data warehouse node (i.e., the fourth node). The content indicated by the log information is the key information of the data processed or stored by each node, that is, the primary key value recorded in the table.

[0178] By analyzing (i.e., the data analysis module) and comparing the log information, nodes with data inconsistency problems on the link (i.e., the log statistical result information) can be found. Among them, sampling is targeted at the data domain; usually, the corresponding data domain of the core business is selected, and the real-time acquisition link is tracked within this data domain. Using the logs collected by Elasticsearch, the data logs between each node on the link can be compared. In terms of algorithm implementation, the main comparisons are the business primary key, the business occurrence time, and the binlog type; through sampling inspection, it can be found whether the real-time data acquisition link is running stably and whether the data processing is complete. Especially for the changed data, problems that cannot be discovered by the supplementary statistical method can be found.

[0179] Generate report data from the analysis results and display them on the dashboard according to the database instance, tenant database, and table. At the same time, draw a data trend curve according to the historical statistical results in the time dimension to display the data accuracy report (i.e., the visualization processing module performs visualization processing on the data statistical result information and / or log statistical result information determined by the data analysis module); for the tenant databases and tables with data anomalies, they are sent to the operation and maintenance and development personnel for processing through the method of anomaly alarm of the anomaly alarm module.

[0180] In the automated process, the main thing is to discover abnormal problems, lacking a general description of the overall data. This involves the description complexity when considering multiple tenant databases, multiple tables, and the time dimension together, which is not what the monitoring system is good at handling; at this time, analysis charts can be combined, and analysis can be carried out through visualization chart tools such as tables and curve charts. Configurable analysis of each dimension and data range can be carried out as needed, such as on a certain MySQL Server (i.e., the database cluster), a certain table in a certain tenant database, or the historical data trend chart of a certain table, etc.; in the specific implementation, an open-source visualization analysis tool can be used.

[0181] A multi-tenant database data accuracy monitoring system proposed in an embodiment of the present invention requires, in terms of hardware costs, a machine responsible for collecting data to run the collection task (i.e., the statistical requirement task), a database system supporting tens of millions of data, and an Elasticsearch system; all of these are conventional requirements, and usually 2 small-scale machines can meet them; in terms of running resources, the occupation of production can be ignored, and there is no need to expand production.

[0182] In terms of architecture design, statistics and analysis are separated, and most of them can be implemented at the SQL level. The system complexity is low and the development cost is low; in terms of log collection, it is implemented through an open-source framework, and there are currently very mature various components, without the need to invest additional costs; compared with the method of re-storing the business data details in the database and then calculating the index logic, this system directly utilizes the production resources and avoids the complexity of business data detail storage in the database and the hardware maintenance cost.

[0183] A multi-tenant database data accuracy monitoring system proposed in an embodiment of the present invention can complete a full data verification at the hourly level for a large-scale (tens of billions level), multi-tenant (tens of thousands level) real-time system, improving the efficiency of data accuracy monitoring. Through system integration, it can meet both regular checks and on-demand calls; the entire process is fully automated, greatly reducing the usage complexity and the requirements for users.

[0184] Embodiment 4

[0185] Figure 6 is a flowchart of a data accuracy monitoring method provided according to Embodiment 4 of the present invention. This embodiment is applicable to the situation of realizing data accuracy monitoring, and the data accuracy monitoring method can be applied to the data accuracy monitoring system provided in any embodiment of the present invention.

[0186] As Figure 6 shown, the method includes:

[0187] S110. Determine the first data information of the first table to be counted in the tenant database to be counted from the database cluster included in the production system through the data statistics module.

[0188] S120. Determine the second data information of the second table to be counted corresponding to the first table to be counted from the real-time data warehouse included in the production system through the data statistics module.

[0189] S130. Collect the log information included in each node on the data link in real time through the log collection module.

[0190] S140. Compare the first data information and the second data information determined by the data statistics module through the data analysis module to determine the data statistics result information.

[0191] S150. Compare the log information of each node collected by the log collection module through the data analysis module to determine the log statistics result information.

[0192] This method determines the data statistics result information by comparing the first data information and the second data information through the data analysis module, and determines the log statistics result information by comparing the log information of each node of the data analysis module, improving the accuracy of data monitoring.

[0193] Further, determining the first data information of the first table to be counted in the tenant database to be counted from the database cluster included in the production system through the data statistics module, and determining the second data information of the second table to be counted corresponding to the first table to be counted from the real-time data warehouse included in the production system through the data statistics module includes:

[0194] The data statistics sub-module included in the data statistics module determines the first data information of the first table to be statistically analyzed in the tenant database to be statistically analyzed from the database cluster included in the production system, and determines the second data information of the second table to be statistically analyzed corresponding to the first table to be statistically analyzed from the real-time data warehouse included in the production system;

[0195] The data storage sub-module included in the data statistics module stores the first data information and the second data information determined by the data statistics sub-module; and sends the first data information and the second data information to the data analysis module.

[0196] Furthermore, the data statistics sub-module included in the data statistics module determines the first data information of the first table to be statistically analyzed in the tenant database to be statistically analyzed from the database cluster included in the production system, and determines the second data information of the second table to be statistically analyzed corresponding to the first table to be statistically analyzed from the real-time data warehouse included in the production system, including:

[0197] The service interface unit included in the data statistics sub-module provides an interface for the first table to be statistically analyzed and the second table to be statistically analyzed to access the system;

[0198] The task queue unit included in the data statistics sub-module stores the statistical requirement tasks waiting to be processed, and the statistical requirement tasks are to determine the first data information of the first table to be statistically analyzed, and / or to determine the second data information of the second table to be statistically analyzed;

[0199] The tenant database management unit included in the data statistics sub-module synchronizes the configurations of each tenant database in each database cluster from the production system, and classifies the tenant databases in the same database cluster into the same group;

[0200] The business abstraction unit included in the data statistics sub-module performs abstraction processing on the data in the first table to be statistically analyzed and / or the second table to be statistically analyzed;

[0201] The task execution management unit included in the data statistics sub-module manages the statistical requirement tasks processed concurrently, and / or manages the number of tenant databases connected to the system;

[0202] The task execution unit included in the data statistics sub-module concretizes the data after abstraction processing by the business abstraction unit into executable statements, and determines the first data information and / or the second data information;

[0203] The result output unit included in the data statistics sub-module sends the first data information and / or the second data information to the data storage sub-module.

[0204] Furthermore, the task execution management unit included in the data statistics sub-module manages the statistical requirement tasks processed concurrently, and / or manages the number of tenant databases connected to the system, including:

[0205] The statistical requirement tasks for concurrent processing are centrally managed by a thread manager, and the number of statistical requirement tasks for concurrent processing has a linear relationship with the number of tenant databases.

[0206] The database connection manager combines with the tenant database management unit to manage the number of tenant databases connected to the system.

[0207] Furthermore, the database connection manager combines with the tenant database management unit to manage the number of tenant databases connected to the system, including:

[0208] The database connection manager releases the connection between the tenant database without statistical requirement tasks and the system.

[0209] When there are no statistical requirement tasks, the thread manager releases the threads.

[0210] Furthermore, the data analysis module compares the first data information and the second data information determined by the data statistics module to determine the data statistics result information, including:

[0211] Within the historical time period, the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for each historical time segment within the historical time period is determined as the first error corresponding to the historical time segment.

[0212] Based on the first errors corresponding to each historical time segment, the first mean value and the first mean square deviation corresponding to the historical time period are determined.

[0213] The absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for the current time segment is determined as the second error corresponding to the current time segment.

[0214] The absolute value of the difference between the second error and the first mean value is determined as the third error of the current time segment relative to the historical time period.

[0215] If the third error does not exceed the first mean square deviation, it is determined that the data statistics result information indicates that no data loss has occurred; if the third error exceeds the first mean square deviation, it is determined that the data statistics result information indicates that data loss has occurred.

[0216] Among them, the historical time period includes multiple historical time segments of the same duration, and the duration of the current time segment is the same as that of the historical time segment.

[0217] Furthermore, the log collection module collects in real time the log information included in each node on the data link, including:

[0218] Collect the log information of the first node corresponding to the binary log consumer through the log collection tool included in the log collection module; collect the log information of the second node corresponding to the message queue middleware storing the binary log; collect the log information of the third node corresponding to the message consumer processing the binary log; collect the log information of the fourth node corresponding to the real-time data warehouse.

[0219] Store the log information included in each node collected by the log collection tool through the log storage tool included in the log collection module; send the log information to the data analysis module.

[0220] Furthermore, compare the log information of each node collected by the log collection module through the data analysis module to determine the log statistical result information, including:

[0221] Compare the log information included in each node collected by the log collection module at a set frequency to determine the problem node where data inconsistency occurs in the data link;

[0222] Determine the problem node as the log statistical result information.

[0223] Furthermore, the data accuracy monitoring method further includes:

[0224] Through the visualization processing module, perform visualization processing on the data statistical result information and / or log statistical result information determined by the data analysis module.

[0225] The data accuracy monitoring method provided by the embodiments of the present invention can be implemented based on the data accuracy monitoring system provided by any of the above embodiments, belonging to the same inventive concept, and having corresponding functions and beneficial effects.

[0226] It should be understood that various forms of the processes shown above can be used, reordering, adding or deleting steps. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitations are imposed herein.

[0227] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data accuracy monitoring system, characterized in that, Including: A data statistics module, configured to determine first data information of a first table to be counted in a tenant database to be counted from a database cluster included in a production system; Determine second data information of a second table to be counted corresponding to the first table to be counted from a real-time data warehouse included in the production system; A log collection module, configured to collect in real time log information included in each node on a data link; A data analysis module, configured to connect to the data statistics module and the log collection module respectively; determine data statistics result information by comparing the first data information and the second data information determined by the data statistics module; Determine log statistics result information by comparing the log information of each node collected by the log collection module; Wherein, the data statistics result information indicates whether data loss occurs or not from the first table to be counted to the second table to be counted; the log statistics result information indicates the nodes where data inconsistency problems occur on the data link; Wherein, the number of the database clusters is one or more, each of the database clusters includes one or more tenant databases, and each tenant database includes one or more first tables to be counted; the second table to be counted is a table obtained by transmitting the first table to be counted through the data link to the real-time data warehouse.

2. The system according to claim 1, wherein The data statistics module includes: A data statistics sub-module and a data storage sub-module; The data statistics sub-module is configured to determine first data information of a first table to be counted in a tenant database to be counted from a database cluster included in a production system; determine second data information of a second table to be counted corresponding to the first table to be counted from a real-time data warehouse included in the production system; The data storage sub-module is configured to connect to the data statistics sub-module and the data analysis module respectively; store the first data information and the second data information determined by the data statistics sub-module; send the first data information and the second data information to the data analysis module.

3. The system according to claim 2, wherein The data statistics sub-module includes: A service interface unit, configured to provide an interface for accessing the system for the first table to be counted and the second table to be counted; A task queue unit, configured to store statistical requirement tasks waiting to be processed, where the statistical requirement tasks are to determine the first data information of the first table to be counted and / or determine the second data information of the second table to be counted; A tenant database management unit, configured to synchronize the configurations of each tenant database in each database cluster from the production system, and divide the tenant databases in the same database cluster into the same group; A business abstraction unit, configured to perform abstraction processing on the data in the first table to be counted and / or the second table to be counted; A task execution management unit, configured to manage the statistical requirement tasks processed concurrently and / or manage the number of tenant databases connected to the system; A task execution unit, configured to concretize the data after being abstracted by the business abstraction unit into executable statements, and determine the first data information and / or the second data information; A result output unit, configured to send the first data information and / or the second data information to the data storage sub-module.

4. The system according to claim 3, wherein The task execution management unit is specifically configured to: Centrally manage the statistical requirement tasks processed concurrently through a thread manager, and the number of the statistical requirement tasks processed concurrently has a linear relationship with the number of tenant databases; Manage the number of tenant databases connected to the system in combination with the tenant database management unit through a database connection manager.

5. The system according to claim 4, wherein The task execution management unit is specifically configured to: Release the connection between the tenant database without statistical requirement tasks and the system through a database connection manager; When there are no statistical requirement tasks, release threads using a thread manager.

6. The system according to claim 1, wherein The data analysis module is specifically configured to: Within a historical time period, determine the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for each historical time segment within the historical time period as the first error corresponding to the historical time segment; Determine the first mean value and the first mean square deviation corresponding to the historical time period according to the first errors corresponding to the respective historical time segments; Determine the absolute value of the difference between the value corresponding to the first data information and the value corresponding to the second data information for the current time segment as the second error corresponding to the current time segment; Determine the absolute value of the difference between the second error and the first mean value as the third error of the current time segment relative to the historical time period; If the third error does not exceed the first mean square deviation, determine that the data statistical result information indicates that no data loss has occurred; If the third error exceeds the first mean square deviation, determine that the data statistical result information indicates that data loss has occurred; Wherein, the historical time period includes a plurality of historical time segments of the same duration, and the duration of the current time segment is the same as that of the historical time segment.

7. The system according to claim 1, wherein The log collection module includes: A log collection tool for collecting log information of a first node corresponding to a binary log consumer; collecting log information of a second node corresponding to a message queue middleware storing binary logs; collecting log information of a third node corresponding to a message consumer processing binary logs; collecting log information of a fourth node corresponding to a real-time data warehouse; A log storage tool for respectively connecting to the log collection tool and the data analysis module; storing the log information of each node collected by the log collection tool; and sending the log information to the data analysis module.

8. The system according to claim 1, wherein The data analysis module is specifically configured to: Compare the log information of each node collected by the log collection module at a set frequency to determine a problem node where a data inconsistency problem occurs in the data link; Determine the problem node as the log statistical result information.

9. The system according to claim 1, wherein It further includes: A visualization processing module for connecting to the data analysis module; Performing visualization processing on the data statistical result information and / or the log statistical result information determined by the data analysis module.

10. A method for monitoring data accuracy, characterized in that, Applied to the data accuracy monitoring system according to any one of claims 1-9, including: Determining the first data information of the first table to be statistically analyzed in the tenant database to be statistically analyzed from the database cluster included in the production system through a data statistics module; Determine the second data information of the second table to be counted corresponding to the first table to be counted from the real-time data warehouse included in the production system through the data statistics module; Collect the log information included in each node on the data link in real time through the log collection module; Compare the first data information and the second data information determined by the data statistics module through the data analysis module to determine the data statistics result information; Compare the log information of each node collected by the log collection module through the data analysis module to determine the log statistics result information; Wherein, the data statistics result information indicates that no data loss or data loss occurs from the first table to be counted to the second table to be counted; the log statistics result information indicates the nodes where data inconsistency problems occur on the data link.

Citation Information

Patent Citations

  • Metadata management method based on binglog, and method and device used for providing metadata

    CN105447014A

  • Method for application system log monitoring management in cloud computing environment

    CN105589791A