Dirty data detection method and device for distributed database and program product
By traversing and calculating the set of shard identifiers, combined with streaming reading strategies and forced route access, dirty data in distributed databases is automatically detected, solving the problem of low detection efficiency in existing technologies and achieving efficient and accurate dirty data detection.
Patent Information
- Application Number
- CN202510882390.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies for detecting dirty data in sharded tables of distributed databases are inefficient and prone to omissions, leading to production problems.
By traversing each shard database in the distributed database, reading data from the target shard table, calculating the shard identifier set, and determining the dirty data detection result based on the shard identifier set and the identifier of the target shard database, a streaming read strategy and forced routing access to the shard table are adopted to generate a dirty data detection report.
It improves the efficiency of dirty data detection, avoids the inefficiency of manual detection, ensures the comprehensiveness and accuracy of detection, and promptly detects and locates dirty data.
Smart Images

Figure CN120804071A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of distributed databases, and in particular, to a dirty data detection method, device and program product for a distributed database. BACKGROUND
[0002] In a distributed database, if a dirty data exists in a sharded table, the data cannot be accessed, especially in the fields of finance and accounting, the shard location of the dirty data is often incorrect, resulting in production problems such as unaligned accounts.
[0003] In related technologies, when a dirty data in a sharded table needs to be processed, a manual comparison method is usually used. When a problem occurs, it is manually checked whether the dirty data exists in the sharded table in the database. However, the manual comparison method is inefficient and prone to omissions. Moreover, the problem is not addressed until it occurs, which is a lagging problem.
[0004] Currently, there is no effective solution to the problem of low detection efficiency of the dirty data in the sharded table in the distributed database by using the manual detection method in related technologies. SUMMARY
[0005] The main purpose of the present application is to provide a dirty data detection method, device and program product for a distributed database, to solve the problem of low detection efficiency of the dirty data in the sharded table in the distributed database by using the manual detection method in related technologies.
[0006] To achieve the above purpose, according to one aspect of the present application, a dirty data detection method for a distributed database is provided. The method comprises: based on the identifier of each shard database, traversing each shard database in the distributed database, and in the case of traversing to a target shard database, reading the data of a target sharded table in the target shard database to obtain a target data set, wherein the distributed database comprises N shard databases, the target shard database is one of the N shard databases, the target sharded table comprises a shard of a target data table stored in the distributed database, and N is a positive integer; calculating the shard identifier of each data in the target data set to obtain a shard identifier set, wherein the shard identifier set includes the data shard identifier of each data in the target data set, and the data shard identifier of each data includes the identifier of the shard database logically storing the data; based on the shard identifier set and the identifier of the target shard database, determining a dirty data detection result of the target sharded table, wherein the dirty data detection result includes whether there is dirty data in the target data set.
[0007] Further, based on the shard identifier set and the identifier of the target shard database, a dirty data detection result of the target shard table is determined, including: comparing the identifier of the target shard database and the data shard identifier of each data in the target data set to obtain a first comparison result; in a case where the first comparison result indicates that the identifier of the target shard database and the data shard identifier of a piece of data in the target data set are inconsistent, determining that the piece of data is dirty data; or in a case where the first comparison result indicates that the identifier of the target shard database and the data shard identifier of a piece of data in the target data set are consistent, determining that the piece of data is not dirty data; based on the dirty data determination result of each data in the target data set, generating the dirty data detection result of the target shard table.
[0008] Further, the dirty data detection result of the target shard table includes: the number of dirty data in the target data set, and the data amount of data in the target shard table that has undergone dirty data detection; based on the dirty data determination result of each data in the target data set, the dirty data detection result of the target shard table is generated, including: based on whether each data in the target data set is dirty data, the number of dirty data in the target data set is counted; the data amount of data in the target data set is obtained to obtain a detection data amount, wherein the detection data amount is used to indicate the data amount of data in the target shard table that has undergone dirty data detection; based on the number of dirty data in the target data set and the detection data amount, the dirty data detection result of the target shard table is determined.
[0009] Further, after determining the dirty data detection result of the target shard table based on the number of dirty data in the target data set and the detection data amount, it further includes: obtaining the data amount of data in the target data table to obtain a target data amount; comparing the target data amount and the detection data amount to obtain a second comparison result; in a case where the second comparison result indicates that the target data amount and the detection data amount are consistent, it is determined that each data in the target shard table has undergone dirty data detection; or in a case where the second comparison result indicates that the target data amount and the detection data amount are inconsistent, it is determined that there is part of data in the target shard table that has not undergone dirty data detection.
[0010] Further, reading data of a target shard table in the target shard database to obtain a target data set includes: obtaining an identifier of the target shard table, and based on the identifier of the target shard database and the identifier of the target shard table, a target query statement is determined; the target query statement is executed, and the data in the target shard table is read by using a streaming reading strategy to obtain the target data set.
[0011] Further, the computing the shard identifier of each data in the target data set obtains a shard identifier set, comprising: computing the hash value of each data in the target data set; determining the shard identifier set based on the hash value of each data in the target data set.
[0012] Further, after determining the dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database, it further comprises: obtaining the dirty data detection result of the shard table of the target data table in each shard database, obtaining N dirty data detection results; visualizing the N dirty data detection results.
[0013] In order to achieve the above purpose, according to another aspect of the present application, a dirty data detection device of a distributed database is provided, the device comprises: a processing unit, configured to traverse each shard database in a distributed database based on the identifier of each shard database, and in the case of traversing to a target shard database, read the data of a target shard table in the target shard database to obtain a target data set, wherein the distributed database comprises N shard databases, the target shard database is one of the N shard databases, the target shard table comprises a shard of a target data table stored in the distributed database, and N is a positive integer; a computing unit, configured to compute the shard identifier of each data in the target data set to obtain a shard identifier set, wherein the shard identifier set comprises the data shard identifier of each data in the target data set, and the data shard identifier of each data comprises the identifier of the shard database logically storing the data; a first determining unit, configured to determine the dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database, wherein the dirty data detection result comprises whether there is dirty data in the target data set.
[0014] Further, the first determining unit comprises: a comparison subunit, configured to compare the identifier of the target shard database with the data shard identifier of each data in the target data set to obtain a first comparison result; a first determination subunit, configured to determine that a data is dirty data in the case that the first comparison result indicates that the identifier of the target shard database is inconsistent with the data shard identifier of the data in the target data set; or a second determination subunit, configured to determine that a data is not dirty data in the case that the first comparison result indicates that the identifier of the target shard database is consistent with the data shard identifier of the data in the target data set; a generating subunit, configured to generate the dirty data detection result of the target shard table based on the dirty data determination result of each data in the target data set.
[0015] Further, the dirty data detection result of the target shard table includes: the number of dirty data in the target data set, and the data amount of data in the target shard table that has undergone dirty data detection, and the generating subunit includes: a statistics module configured to count the number of dirty data in the target data set based on whether each piece of data in the target data set is dirty data; an acquisition module configured to acquire the data amount of data in the target data set to obtain a detection data amount, where the detection data amount is used to indicate the data amount of data in the target shard table that has undergone dirty data detection; and a determination module configured to determine the dirty data detection result of the target shard table based on the number of dirty data in the target data set and the detection data amount.
[0016] Further, the dirty data detection device of the distributed database further includes: a first acquisition unit configured to acquire the data amount of data in the target data table to obtain a target data amount after determining the dirty data detection result of the target shard table based on the number of dirty data in the target data set and the detection data amount; a comparison unit configured to compare the target data amount and the detection data amount to obtain a second comparison result; and a second determination unit configured to determine that each piece of data in the target shard table has undergone dirty data detection in a case where the second comparison result indicates that the target data amount is consistent with the detection data amount, or configured to determine that there is part of data in the target shard table that has not undergone dirty data detection in a case where the second comparison result indicates that the target data amount is inconsistent with the detection data amount.
[0017] Further, the processing unit includes: a processing subunit configured to acquire an identifier of a target shard table and determine a target query statement based on the identifier of the target shard database and the identifier of the target shard table; and a reading subunit configured to execute the target query statement and read data in the target shard table using a streaming reading strategy to obtain the target data set.
[0018] Further, the computing unit includes: a computing subunit configured to calculate a hash value of each piece of data in the target data set; and a determination subunit configured to determine the shard identifier set based on the hash value of each piece of data in the target data set.
[0019] Further, the dirty data detection device of the distributed database further includes: a second acquisition unit configured to acquire dirty data detection results of shard tables of the target data table in each shard database to obtain N dirty data detection results after determining the dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database; and a display unit configured to visually display the N dirty data detection results.
[0020] According to another aspect of the present application, a computer readable storage medium is provided, comprising a stored executable program, wherein the computer readable storage medium is caused to perform the dirty data detection method of the distributed database when the executable program is run.
[0021] According to another aspect of the present application, an electronic device is provided, comprising a memory storing an executable program, and a processor configured to run the program, wherein the program is configured to perform the dirty data detection method of the distributed database when run.
[0022] According to another aspect of the present application, a computer program product is provided, comprising computer instructions configured to perform the steps of the dirty data detection method of the distributed database when executed by a processor.
[0023] In the embodiments of the present application, based on the identifier of each shard database, each shard database in the distributed database is traversed, and in the case of traversing to a target shard database, data of a target shard table in the target shard database is read to obtain a target data set, wherein the distributed database comprises N shard databases, the target shard database is one of the N shard databases, the target shard table comprises a shard of a target data table stored in the distributed database, and N is a positive integer; a shard identifier of each data in the target data set is calculated to obtain a shard identifier set, wherein the shard identifier set comprises a data shard identifier of each data in the target data set, and the data shard identifier of each data comprises an identifier of a shard database logically storing the data; and based on the shard identifier set and the identifier of the target shard database, a dirty data detection result of the target shard table is determined, wherein the dirty data detection result comprises whether there is dirty data in the target data set, thereby solving the technical problem of low detection efficiency in the related art that the dirty data in the shard table of the distributed database is detected by using a manual detection method.
[0024] In the present application, the data in the shard table is read according to the identifier of the shard database and the identifier of the shard table, and the dirty data detection result is determined based on the shard identifier of the data in the shard table and the actual identifier of the shard database, thereby avoiding the low efficiency of manually monitoring the dirty data in the shard table in the related art, and achieving the technical effect of improving the dirty data detection efficiency of the distributed database. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and are used to interpret the illustrative embodiments of the present application and their descriptions, and do not constitute improper limitations to the present application. In the drawings:
[0026] Figure 1A hardware structure block diagram of a computer terminal for implementing a dirty data detection method of a distributed database is shown;
[0027] Figure 2 A flow chart of a dirty data detection method of a distributed database is provided according to an embodiment of the present application;
[0028] Figure 3 A schematic diagram of a dirty data detection device of a distributed database is provided according to an embodiment of the present application;
[0029] Figure 4 A structure block diagram of an electronic device is provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0031] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0033] Distributed database: a logically unified database composed of a plurality of physically dispersed database units connected by a computer network. Each connected database unit is called a site or node. The distributed database has a unified database management system for management, which is called a distributed database management system.
[0034] Sharding: Sharding database and sharding table, generally refers to the physical server division, that is, the corresponding database becomes multiple sharding databases, the data structure of the sharding table (that is, the sharding table) is consistent, exists in different sharding databases, and the data of each sharding table has no intersection, and the data of each sharding table adds up to the complete data.
[0035] Distributed database middleware: The application node accesses the distributed database through the middleware.
[0036] Dirty data of the sharding table: refers to the sharding error data, which appears in the sharding table that should not appear.
[0037] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user selection authorization or refusal. For example, the system and the interface between the related users or institutions provide the corresponding operation portal for the user to select to agree or refuse the automatic decision result; if the user chooses to refuse, the expert decision process is entered.
[0038] The present application can be applied to various software products, control systems, and client terminals (including but not limited to mobile client terminals, PC terminals, etc.) corresponding to distributed data of financial institutions, and can store the distributed database of the data of the business content (including but not limited to transfer, financial management, fund, payment, account checking, advertising, recommendation, etc.) of the financial institutions.
[0039] Embodiment one
[0040] According to the embodiments of the present application, a method embodiment of a dirty data detection method of a distributed database is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] The method embodiment provided by the embodiment one of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a dirty data detection method of a distributed database is shown. As shown in Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0042] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0043] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the dirty data detection method of the distributed database in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned dirty data detection method of the distributed database. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0044] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module that is configured to communicate with the Internet wirelessly.
[0045] The display can be a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0046] In the above operating environment, the present application provides a method for detecting dirty data of a distributed database as shown in Figure 2 Figure 2 is a flowchart of a method for detecting dirty data of a distributed database according to an embodiment of the present application.
[0047] In step S201, based on the identifier of each shard database, each shard database in the distributed database is traversed. In the case of traversing to a target shard database, the data of a target shard table in the target shard database is read to obtain a target data set. The distributed database includes N shard databases, the target shard database is one of the N shard databases, the target shard table includes a shard of a target data table stored in the distributed database, and N is a positive integer.
[0048] The distributed data can be composed of N shard databases, the target data table can be composed of multiple shard tables, and the multiple shard tables can be stored in the N shard databases. For example, if the distributed database is composed of four shard databases, the target data table can also be divided into four shard tables and stored in the four shard databases. Each shard table can determine the shard database to which the data in the table is stored according to the hash value of the data in the table. For example, if the hash value is the same as the identifier of one of the shard databases, the data corresponding to the hash value can be stored in the shard database.
[0049] In order to avoid the case that the related art needs to manually check the database and analyze the dirty data after a problem occurs, in the embodiment, the dirty data in the distributed database can be scanned and batched, for example, based on the identifier of each shard database, each shard database in the distributed database can be traversed, when a certain shard database (i.e., a target shard database) is traversed, the data of the target shard table in the target shard database can be read into a target data set, for example, based on the identifier of the target shard database and the identifier of a single shard data table, the shard table of the single shard database in the target shard database can be accessed by forced routing, all data of the single shard table can be read, the data can be read into a result set by streaming reading, and the target data set is obtained. Based on the identifier of each shard database, each shard database in the distributed database can be periodically traversed.
[0050] In step S202, the shard identifier of each data in the target data set is calculated to obtain a shard identifier set, wherein the shard identifier set includes the data shard identifier of each data in the target data set, and the data shard identifier of each data includes the identifier of the shard database that logically stores the data.
[0051] The shard identifier set can include the data shard identifier of each data in the target data set, and the data shard identifier of each data includes the identifier of the shard database that logically stores the data, i.e., the identifier of the shard database that should originally store the data.
[0052] In the embodiment, the shard identifier of each data in the target data set can be calculated in the same way as the shard identifier of the shard database where the data is stored, i.e., the shard identifier of the shard database where the data is stored is the same as the shard identifier of the shard database where the data is stored.
[0053] In step S203, based on the shard identifier set and the identifier of the target shard database, the dirty data detection result of the target shard table is determined, wherein the dirty data detection result includes whether there is dirty data in the target data set.
[0054] In the embodiment, the data shard identifier of each data in the shard identifier set can be compared with the identifier of the target shard database, and whether they are consistent can be compared, if they are not consistent, it means that the data is dirty data, i.e., the data appears in a position where it should not appear, if they are consistent, it means that the data is not dirty data.
[0055] It should be noted that in the process of traversing each shard table in each shard database in the distributed database, whether each data is dirty data can be determined, and in this way, all dirty data in the distributed database can be determined.
[0056] After obtaining the detection result of all dirty data in the distributed database, printing can be performed, and the number of dirty data in each shard table in each shard database in the distributed database and the total amount of scanned data of each shard table can be printed, wherein the total amount of scanned data can be used to determine whether to scan all data of the shard table.
[0057] In the embodiment, by the above steps, the data in the shard table is read by forced routing according to the identifier of the shard database and the identifier of the shard table, and the dirty data detection result is determined based on the shard identifier of the data in the shard table and the actual identifier of the shard database, thereby avoiding the low efficiency of manually monitoring the dirty data in the shard table in the related art, and achieving the technical effect of improving the dirty data detection efficiency of the distributed database. Further, the technical problem of low detection efficiency of detecting dirty data in the shard table in the distributed database by using manual detection in the related art is solved.
[0058] Optionally, in the dirty data detection method of the distributed database provided in the embodiment of the present application, the dirty data detection result of the target shard table is determined based on the shard identifier set and the identifier of the target shard database, including: comparing the identifier of the target shard database and the data shard identifier of each data in the target data set to obtain a first comparison result; in the case that the first comparison result indicates that the identifier of the target shard database and the data shard identifier of a data in the target data set are inconsistent, determining that the data is dirty data; or in the case that the first comparison result indicates that the identifier of the target shard database and the data shard identifier of a data in the target data set are consistent, determining that the data is not dirty data; and generating the dirty data detection result of the target shard table based on the dirty data determination result of each data in the target data set.
[0059] If the first comparison result indicates that the shard identifier of a data is inconsistent with the shard database identifier currently checked, the data can be marked as dirty data, indicating that the data is incorrectly stored in a shard database that it does not belong to. When detecting dirty data, a warning prompt message can also be sent.
[0060] If the first comparison result indicates that the shard identifier of a data is consistent with the shard database identifier currently checked, the data can be confirmed as not being dirty data, indicating that the data is correctly stored in a shard database that it belongs to.
[0061] According to the dirty data determination result of each data, the batch job generates a dirty data detection report of the target shard table, which can include how many dirty data there are in the target shard table, and can also include the primary key value of the target shard table and the identifier of the shard database where the target shard table is located. Through batch job, dirty data in the distributed database can be efficiently and accurately detected.
[0062] Optionally, in the dirty data detection method for the distributed database provided in the embodiments of the present application, the dirty data detection result of the target shard table includes: the number of dirty data in the target data set, and the data amount of data in the target shard table that has been subjected to dirty data detection; the dirty data detection result of the target shard table is generated based on the dirty data determination result of each piece of data in the target data set, including: based on whether each piece of data in the target data set is dirty data, the number of dirty data in the target data set is counted; the data amount of data in the target data set is obtained to obtain a detection data amount, where the detection data amount is used to indicate the data amount of data in the target shard table that has been subjected to dirty data detection; and the dirty data detection result of the target shard table is determined based on the number of dirty data in the target data set and the detection data amount.
[0063] In the embodiment, the dirty data detection result of the target shard table can include: the number of dirty data in the target data set, and the data amount of data in the target shard table that has been subjected to dirty data detection, and can also include: the primary key of the target shard table and the identifier of the target shard database. In the embodiment, based on whether each piece of data in the target data set is dirty data, the number of dirty data in the target data set is counted, and the data amount (for example, the number of data) in the target data table is also counted. Based on the number of dirty data in the target data set and the detection data amount, the dirty data detection result of the target shard table is composed, and the technical effect of automatically detecting dirty data in the distributed database is achieved.
[0064] Optionally, in the dirty data detection method for the distributed database provided in the embodiments of the present application, after the dirty data detection result of the target shard table is determined based on the number of dirty data in the target data set and the detection data amount, the method further includes: obtaining the data amount of data in the target data table to obtain a target data amount; comparing the target data amount and the detection data amount to obtain a second comparison result; in a case where the second comparison result indicates that the target data amount and the detection data amount are consistent, it is determined that each piece of data in the target shard table has been subjected to dirty data detection; or in a case where the second comparison result indicates that the target data amount and the detection data amount are inconsistent, it is determined that there is part of data in the target shard table that has not been subjected to dirty data detection.
[0065] The target data amount described above can be the total number of all data in the target shard table, and the obtaining process can be: the total number of all data in the target shard table is obtained by querying the metadata of the database or directly executing a counting query.
[0066] In the embodiment, the target data volume and the detected data volume are compared. By comparing the target data volume and the detected data volume, it can be determined whether the detection job covers all data in the target sharded table. If they are equal, it means that all data has been checked. If they are different, it means that there is data that has not been detected. When the target data volume and the detected data volume are consistent, it is confirmed that the detection process completely covers all data of the target sharded table, ensuring the comprehensiveness of the dirty data detection. The number of detected dirty data is accurate and can be used as a basis for further analysis and decision-making. When the target data volume and the detected data volume are inconsistent, it indicates that the detection does not cover all data, and there may be a part of data that has not been checked. At this time, the detected dirty data may not be complete, so dirty data detection can be performed again on the target sharded table to ensure that all data is checked to avoid missing possible dirty data problems.
[0067] In a distributed database environment, through the above comparison process, not only the completeness of the detection job can be confirmed, but also a verifiable connection between the detection result and the actual data volume can be established, improving the efficiency of data management and problem tracking.
[0068] Optionally, in the dirty data detection method for a distributed database provided in the embodiment of the application, the data of the target sharded table in the target sharded database is read to obtain a target data set, including: obtaining an identifier of the target sharded table, and determining a target query statement based on the identifier of the target sharded database and the identifier of the target sharded table; executing the target query statement, and reading data in the target sharded table by using a streaming reading strategy to obtain the target data set.
[0069] The identifier of the target sharded table can be an identifier of the target data table, which can be used to determine the specific data table to be checked. In the environment of a distributed database, since data is scattered in multiple physical databases (sharded databases), it is necessary to explicitly specify the sharded table to be read. The identifier of the target sharded table can be pre-configured in the dirty data detection program, ensuring that the program can accurately locate the specific sharded table to be checked.
[0070] The identifier of the target sharded database can refer to the identifier of the sharded database from which data is currently read, which is used to help the program determine which specific sharded database to access.
[0071] Since the identifiers of all sharded tables of the target data table are the same, the identifier of the target sharded database, the identifier of the target sharded table, and the related query statement can be spliced to obtain the target query statement, so as to ensure that the data query is performed on the target sharded table in the target sharded database.
[0072] In this embodiment, the distributed database middleware can be used to ensure that the query statement can directly access the target sharding table in the target sharding database.
[0073] The stream reading strategy described above can be an efficient data reading method, allowing the program to process data one by one during data reading, rather than loading all data into memory at once, reducing memory usage and avoiding the impact of insufficient memory on program running.
[0074] After executing the query statement, stream reading of data in the target sharding table can be started. After reading each piece of data, it can be put into the result set to obtain the target data set.
[0075] The stream reading strategy is used to read data step by step to build the target data set, providing a data basis for subsequent dirty data detection and statistics. This ensures efficiency and effective use of resources in large-scale data detection scenarios.
[0076] Optionally, in the dirty data detection method of the distributed database provided in the embodiments of the present application, the shard identifier of each piece of data in the target data set is calculated to obtain a shard identifier set, including: calculating the hash value of each piece of data in the target data set; based on the hash value of each piece of data in the target data set, the shard identifier set is determined.
[0077] Since in the distributed database, data is stored in different sharding databases according to certain rules, and this rule is generally based on the output of the hash function. Therefore, in this embodiment, the hash value of each piece of data in the target data set can be calculated, and the hash value of all data in the target data set can be used to form the shard identifier set described above.
[0078] In this embodiment, calculating the shard identifier of each piece of data in the target data set to obtain the shard identifier set can achieve accurate positioning of the actual storage location of the data.
[0079] Optionally, in the dirty data detection method of the distributed database provided in the embodiments of the present application, after determining the dirty data detection result of the target sharding table based on the shard identifier set and the identifier of the target sharding database, the method further includes: obtaining the dirty data detection result of the sharding table in each sharding database of the target data table, obtaining N dirty data detection results; and visualizing the N dirty data detection results.
[0080] The dirty data detection results of the target data table in each shard database can be obtained. If there are N shard databases in the distributed database, the target data table can be divided into N shard tables. After the dirty data detection is completed, N detection results of the shard tables are generated. In this embodiment, the N dirty data detection results can be converted into charts, dashboards or other visual elements to display the data in a more intuitive manner, and improve the readability and understanding speed of the information. In an optional example, bar charts, pie charts, tables or other suitable visualization tools can be used to display the dirty data quantity, dirty data proportion and detection state of each shard database, so as to quickly identify which shard database has the most serious problem and which shard has higher data quality, and facilitate the priority processing of the problem area and the improvement of efficiency.
[0081] In this embodiment, the dirty data detection results of all the shard tables of the distributed database can also be printed, as follows.
[0082] Whether the dirty data exists in all the shard tables of the distributed database can be detected. After the dirty data appears, a dirty data post-printing mechanism can be designed to print all the problems at one time. The printing mechanism is as follows.
[0083] Step 1: Custom configuration of primary key information of each shard table. For example:
[0084] The primary key of A table is callID.
[0085] The primary key of B table is serino.
[0086] Step 2: The dirty data scanning batch job iterates each shard table one by one. When dirty data appears, go to step 3.
[0087] Step 3: According to the primary key information configured for each shard table, write into a print file, and the format is as follows:
[0088] A table callID: 1643091200004300900;
[0089] A table name A table shard number (0) dirtyDataCount is: 1;
[0090] A table name A table shard number (0) totalCount is: 10;
[0091] A table name A table shard number (1) dirtyDataCount is: 0;
[0092] A table name A table shard number (1) totalCount is: 10;
[0093] B table serino: 1643091200004300920;
[0094] B table name B table shard number (0) dirtyDataCount is: 1;
[0095] B table name B table shard number (0) totalCount is: 10;
[0096] B table name B table shard number (1) dirtyDataCount is: 0;
[0097] B table name B table shard number (1) totalCount is: 10;
[0098] dirtyDataCount is indicates the total number of dirty data, and totalCount is indicates the amount of data scanned.
[0099] Step 4: The output directory of the error file is / approot1 / ccis / datacheck / .
[0100] In this embodiment, the efficiency and accuracy of manual query can be improved by periodically performing batch automatic processing. The dirty data post-printing mechanism can clearly indicate which shard of which shard table has dirty data, and the printed information is clear and unambiguous, which can speed up the problem positioning speed. When dirty data occurs, the risk can be timely alarmed, and the problem response efficiency can be effectively improved.
[0101] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.
[0102] Embodiment two
[0103] The embodiments of the present application also provide another dirty data detection method for a distributed database, which comprises: a dirty data scanning batch job, a dirty data event monitoring alarm mechanism, and a dirty data post-printing mechanism.
[0104] The dirty data scanning batch job iterates through the data of each shard database of the distributed database one by one, calculates the shard number C (i.e., the shard identifier) through a self-defined routing algorithm, compares whether the shard number C is consistent with the actual shard database sequence number N (i.e., the shard database identifier) of the data, and if not, reports the problem through the dirty data event monitoring alarm mechanism. After the batch ends, the running result is printed to the output directory through the dirty data post-printing mechanism, and the details are as follows:
[0105] 1 "Dirty data scan batch job" module:
[0106] The application program can access the distributed database using distributed database middleware. Since the related art batch program cannot access only a single shard database, but the embodiment needs to traverse each single shard database one by one to calculate the shard number. Therefore, in the embodiment, each single shard database can be traversed one by one by the batch program through forced routing, and the following is the pseudo code of the "dirty data scan batch job" module.
[0107] Mainprocess(): main entry of the batch program;
[0108] Init(input: shard field (i.e. identification of shard database) N, query mapper): the shard table of a single shard database can be accessed by forced routing through the writing method of / *!dble:dataNode=db${setid}* / , all data of the single shard table is read, and a result set (i.e. target data set) is read into the DBF (streaming read) through the DBF (streaming read) mode.
[0109] Process(input: shard field N): total process of batch job processing, responsible for traversing data one by one, calculating the shard field, and printing error files.
[0110] Calculate(input: shard field N): calculate using a custom algorithm to obtain the calculated shard C.
[0111] For example:
[0112] Int M (where M represents the number of shards of the database, which can generally be 4);
[0113] Mainprocess()
[0114] {
[0115] Loop M times (i<=M);
[0116] {
[0117] Call init(): return data of each shard table; (input: current shard, query mapper);
[0118] Call process(): traverse shard table data one by one, recalculate the shard value, and judge whether it is consistent with the current shard. If not, print to the error file; (input: shard table data).
[0119] }
[0120] }
[0121] Init(input: shard field, query mapper)
[0122] {
[0123] First step: find the SQL statement (i.e. target query statement) through the query mapper, and concatenate the conditions to form the statement as follows:
[0124] " / *!dble:dataNode=db${setid}* / SELECT*FROM${tableName}WHERE${condition}";
[0125] Second step: the result of the query is read into a result set through the DBF streaming reading mode.
[0126] }
[0127] Process(input: shard field)
[0128] {
[0129] First step: iterate through the data one by one, call the calculate() method, and calculate the shard field.
[0130] Second step: compare the calculated shard field with the input shard field. If they are inconsistent, it is dirty data.
[0131] }
[0132] calculate(input: shard field value)
[0133] {
[0134] Use a hash algorithm to calculate the belonging shard; (This hash algorithm should be consistent with the logic of the algorithm).
[0135] }
[0136] 2 "Dirty data event monitoring alarm mechanism" module:
[0137] The existence of dirty data in the shard table is a very serious production problem that needs to be notified to the developer in the first time. For this feature, a dirty data event monitoring alarm mechanism can be set. When the shard number C calculated by the custom routing algorithm is inconsistent with the actual shard library sequence number N of the data, an abnormal branch is triggered for processing, an alarm message is generated, and the message is sent to the master platform. After receiving the message, the platform can notify the developer through SMS, email, etc.
[0138] 3 "Dirty data post-printing mechanism" module:
[0139] The distributed database can detect whether all the shard tables have dirty data. After dirty data appears, a dirty data printing mechanism can be designed to print all the problems at once to locate the problem faster. The printing mechanism is as follows:
[0140] Step 1: Custom configuration of primary key information of each shard table. For example:
[0141] The primary key of table A is callID.
[0142] The primary key of table B is serino.
[0143] Step 2: The dirty data scanning batch job iterates through each shard table one by one. When dirty data appears, go to step 3.
[0144] Step 3: According to the primary key information configured for each shard table, write to the print file in the following format:
[0145] A table callID: 1643091200004300900;
[0146] A table name A table shard number (0) dirtyDataCount is:1;
[0147] A table name A table shard number (0) totalCount is:10;
[0148] A table name A table shard number (1) dirtyDataCount is:0;
[0149] A table name A table shard number (1) totalCount is:10;
[0150] B table serino: 1643091200004300920;
[0151] B table name B table shard number (0) dirtyDataCount is:1;
[0152] B table name B table shard number (0) totalCount is:10;
[0153] B table name B table shard number (1) dirtyDataCount is:0;
[0154] B table name B table shard number (1) totalCount is:10;
[0155] Among them, "dirtyDataCount is" represents the total number of dirty data / total number of dirty data, and "totalCount is" represents the amount of scanned data.
[0156] Step 4: The output directory of the error file is / approot1 / ccis / datacheck / .
[0157] In this embodiment, the efficiency and accuracy of manual query can be improved by periodic execution in a batch automatic processing manner. The dirty data post-printing mechanism can clearly indicate which shard table of which shard has dirty data, and the printed information is clear and understandable, which can speed up the problem positioning speed. When dirty data occurs, the risk can be timely alarmed, and the problem response efficiency can be effectively improved.
[0158] Embodiment three
[0159] The application embodiment further provides a dirty data detection device of a distributed database. It should be noted that the dirty data detection device of the distributed database in the application embodiment can be used to execute the dirty data detection method for the distributed database provided in the application embodiment. The dirty data detection device of the distributed database provided in the application embodiment is introduced as follows.
[0160] According to the application embodiment, a device for implementing the dirty data detection method of the distributed database is further provided, as shown in Figure 3 The device comprises a processing unit 31, a calculation unit 32, and a first determination unit 33.
[0161] The processing unit 31 is configured to traverse each shard database in the distributed database based on the identifier of each shard database, read the data of the target shard table in the target shard database in the case of traversing to the target shard database, and obtain a target data set, wherein the distributed database comprises N shard databases, the target shard database is one of the N shard databases, the target shard table comprises a shard of a target data table stored in the distributed database, and N is a positive integer.
[0162] The calculation unit 32 is configured to calculate the shard identifier of each data in the target data set to obtain a shard identifier set, wherein the shard identifier set comprises the data shard identifier of each data in the target data set, and the data shard identifier of each data comprises the identifier of the shard database logically storing the data.
[0163] The first determination unit 33 is configured to determine a dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database, wherein the dirty data detection result comprises whether there is dirty data in the target data set.
[0164] In the dirty data detection apparatus for a distributed database provided in the embodiment of the present application, the processing unit 31 can traverse each shard database in the distributed database based on the identifier of each shard database, read data of a target shard table in a target shard database in a case of traversing to the target shard database, and obtain a target data set, wherein the distributed database includes N shard databases, the target shard database is one of the N shard databases, the target shard table includes a shard of a target data table stored in the distributed database, N is a positive integer, the computing unit 32 calculates the shard identifier of each data in the target data set to obtain a shard identifier set, wherein the shard identifier set includes the data shard identifier of each data in the target data set, and the data shard identifier of each data includes the identifier of the shard database logically storing the data, the first determining unit 33 determines a dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database, wherein the dirty data detection result includes whether there is dirty data in the target data set. Thus, the technical problem of low detection efficiency in the related art that the dirty data in the shard table of the distributed database is detected by using the manual detection method is solved. In the embodiment, the data in the shard table is forcibly routed according to the identifier of the shard database and the identifier of the shard table, the dirty data detection result is determined based on the shard identifier of the data in the shard table and the actual identifier of the shard database, the condition that the dirty data in the shard table is manually monitored in the related art and the efficiency is low is avoided, and thus the technical effect of improving the dirty data detection efficiency of the distributed database is achieved.
[0165] Optionally, in the dirty data detection apparatus for a distributed database provided in the embodiment of the present application, the first determining unit includes: a comparison subunit, configured to compare the identifier of the target shard database and the data shard identifier of each data in the target data set to obtain a first comparison result; a first determination subunit, configured to determine that the data is dirty data in a case that the first comparison result indicates that the identifier of the target shard database and the data shard identifier of the data in the target data set are inconsistent; or a second determination subunit, configured to determine that the data is not dirty data in a case that the first comparison result indicates that the identifier of the target shard database and the data shard identifier of the data in the target data set are consistent; and a generation subunit, configured to generate the dirty data detection result of the target shard table based on the dirty data determination result of each data in the target data set.
[0166] Optionally, in the dirty data detection apparatus for the distributed database provided in the embodiments of the present application, the dirty data detection result of the target shard table comprises: the number of dirty data in the target data set, and the data amount of data in the target shard table that has been subjected to dirty data detection; the generating subunit comprises: a statistics module configured to count the number of dirty data in the target data set based on whether each piece of data in the target data set is dirty data; an acquisition module configured to acquire the data amount of data in the target data set to obtain a detection data amount, wherein the detection data amount is used to indicate the data amount of data in the target shard table that has been subjected to dirty data detection; and a determination module configured to determine the dirty data detection result of the target shard table based on the number of dirty data in the target data set and the detection data amount.
[0167] Optionally, in the dirty data detection apparatus for the distributed database provided in the embodiments of the present application, the dirty data detection apparatus for the distributed database further comprises: a first acquisition unit configured to acquire the data amount of data in the target data table to obtain a target data amount after determining the dirty data detection result of the target shard table based on the number of dirty data in the target data set and the detection data amount; a comparison unit configured to compare the target data amount and the detection data amount to obtain a second comparison result; and a second determination unit configured to determine that each piece of data in the target shard table has been subjected to dirty data detection in a case where the second comparison result indicates that the target data amount and the detection data amount are consistent; or the second determination unit is configured to determine that there is part of data in the target shard table that has not been subjected to dirty data detection in a case where the second comparison result indicates that the target data amount and the detection data amount are inconsistent.
[0168] Optionally, in the dirty data detection apparatus for the distributed database provided in the embodiments of the present application, the processing unit comprises: a processing subunit configured to acquire the identifier of the target shard table, and determine the target query statement based on the identifier of the target shard database and the identifier of the target shard table; and a reading subunit configured to execute the target query statement, and read data in the target shard table by using a streaming reading strategy to obtain the target data set.
[0169] Optionally, in the dirty data detection apparatus for the distributed database provided in the embodiments of the present application, the computing unit comprises: a computing subunit configured to compute the hash value of each piece of data in the target data set; and a determination subunit configured to determine the shard identifier set based on the hash value of each piece of data in the target data set.
[0170] Optionally, in the dirty data detection device for a distributed database provided in an embodiment of the present application, the dirty data detection device for a distributed database further includes: a second acquisition unit, for obtaining the dirty data detection results of the shard tables of the target data table in each shard database after determining the dirty data detection results of the target shard table based on the shard identification set and the identification of the target shard database, to obtain N dirty data detection results; and a display unit, for visually displaying the N dirty data detection results.
[0171] It should be noted that the processing unit 31, the calculation unit 32, and the first determination unit 33 described above correspond to steps S201 to S203 in the first embodiment. The examples and application scenarios implemented by each unit and the corresponding steps are the same, but are not limited to the contents disclosed in the first embodiment. It should be noted that the modules or units described above can be hardware components or software components stored in a memory (e.g., the memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be run as part of a device in the computer terminal 10 provided in the first embodiment.
[0172] Example 4
[0173] An embodiment of the present application may provide an electronic device, Figure 4 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 4 As shown, the electronic device may include: one or more ( Figure 4 Only one is shown) processor 402, memory 404, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0174] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0175] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: based on the identification of each shard database, traversing each shard database in the distributed database, in the case of traversing to the target shard database, reading the data of the target shard table in the target shard database to obtain a target data set, wherein the distributed database includes N shard databases, the target shard database is one of the N shard databases, the target shard table includes a shard of a target data table stored in the distributed database, and N is a positive integer; calculating the shard identification of each data in the target data set to obtain a shard identification set, wherein the shard identification set includes the data shard identification of each data in the target data set, and the data shard identification of each data includes the identification of the shard database logically storing the data; based on the shard identification set and the identification of the target shard database, determining the dirty data detection result of the target shard table, wherein the dirty data detection result includes whether there is dirty data in the target data set.
[0176] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: based on the shard identification set and the identification of the target shard database, determining the dirty data detection result of the target shard table, including: comparing the identification of the target shard database and the data shard identification of each data in the target data set to obtain a first comparison result; in the case that the first comparison result indicates that the identification of the target shard database and the data shard identification of a data in the target data set are inconsistent, determining that the data is dirty data; or in the case that the first comparison result indicates that the identification of the target shard database and the data shard identification of a data in the target data set are consistent, determining that the data is not dirty data; based on the dirty data determination result of each data in the target data set, generating the dirty data detection result of the target shard table.
[0177] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: the dirty data detection result of the target shard table includes the number of dirty data in the target data set and the data amount of data in the target shard table that has been detected for dirty data, based on the dirty data determination result of each data in the target data set, generating the dirty data detection result of the target shard table, including: based on whether each data in the target data set is dirty data, counting the number of dirty data in the target data set; obtaining the data amount of data in the target data set to obtain a detection data amount, wherein the detection data amount is used to indicate the data amount of data in the target shard table that has been detected for dirty data; based on the number of dirty data in the target data set and the detection data amount, determining the dirty data detection result of the target shard table.
[0178] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: after determining the dirty data detection result of the target shard table based on the number of dirty data in the target data set and the detection data amount, the steps further include: obtaining the data amount of the data in the target data table to obtain a target data amount; comparing the target data amount and the detection data amount to obtain a second comparison result; in a case where the second comparison result indicates that the target data amount and the detection data amount are consistent, determining that each piece of data in the target shard table has been subjected to dirty data detection; or in a case where the second comparison result indicates that the target data amount and the detection data amount are inconsistent, determining that there is part of the data in the target shard table that has not been subjected to dirty data detection.
[0179] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: reading the data of the target shard table in the target shard database to obtain a target data set, including: obtaining the identifier of the target shard table, and determining a target query statement based on the identifier of the target shard database and the identifier of the target shard table; executing the target query statement, and reading the data in the target shard table using a streaming reading strategy to obtain the target data set.
[0180] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: calculating the shard identifier of each piece of data in the target data set to obtain a shard identifier set, including: calculating the hash value of each piece of data in the target data set; determining the shard identifier set based on the hash value of each piece of data in the target data set.
[0181] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: after determining the dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database, the steps further include: obtaining the dirty data detection result of the shard table in each shard database of the target data table to obtain N dirty data detection results; and visually displaying the N dirty data detection results.
[0182] According to the identifier of the shard database and the identifier of the shard table, the data in the shard table is read by forced routing, and the dirty data detection result is determined based on the shard identifier of the data in the shard table and the actual identifier of the shard database, thereby avoiding the low efficiency of manually monitoring the dirty data in the shard table in the related art, and achieving the technical effect of improving the dirty data detection efficiency of the distributed database.
[0183] Those skilled in the art can understand that Figure 4 The structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, etc.Figure 4 It does not cause limitation to the structure of the electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) or have a different configuration from that shown in the drawings. Figure 4 Figure 4
[0184] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0185] Embodiment Five
[0186] The embodiments of the present application further provide a storage medium. Optionally, in the embodiments, the storage medium can be used to save the program code executed by the dirty data detection method of the distributed database provided in the embodiment one.
[0187] Optionally, in the embodiments, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0188] The present application further provides a computer program product, which, when executed on a data processing device, is adapted to execute the steps of the dirty data detection method of the distributed database.
[0189] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0190] In the above-mentioned embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0191] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the unit division in the above-mentioned device embodiment is only a logical function division, and there can be another division manner during actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.
[0192] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0193] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0194] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various program code storage media.
[0195] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method for detecting dirty data in a distributed database, characterized in that: include: Based on the identifiers of each shard database, traverse each shard database in the distributed database, and when traversing to a target shard database, read data of a target shard table in the target shard database to obtain a target data set, wherein the distributed database includes N shard databases, the target shard database is one of the N shard databases, and the target shard table includes: a shard of the target data table stored in the distributed database, where N is a positive integer; Calculating a shard identifier for each piece of data in the target data set to obtain a shard identifier set, wherein the shard identifier set includes: a data shard identifier for each piece of data in the target data set, and the data shard identifier for each piece of data includes: an identifier of a shard database that logically stores the piece of data; Based on the shard identifier set and the identifier of the target shard database, a dirty data detection result of the target shard table is determined, wherein the dirty data detection result includes: whether dirty data exists in the target data set.
2. The dirty data detection method according to claim 1, characterized in that: Determining a dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database includes: Comparing the identifier of the target shard database with the data shard identifier of each piece of data in the target data set to obtain a first comparison result; If the first comparison result indicates that the identifier of the target shard database is inconsistent with the data shard identifier of a certain data in the target data set, the data is determined to be dirty data; or If the first comparison result indicates that the identifier of the target shard database is consistent with the data shard identifier of a piece of data in the target data set, determining that the piece of data is not dirty data; Based on the dirty data determination result of each data in the target data set, a dirty data detection result of the target shard table is generated.
3. The dirty data detection method according to claim 2, characterized in that: The dirty data detection result of the target shard table includes: the number of dirty data in the target data set, and the amount of data in the target shard table that has undergone dirty data detection. Based on the dirty data determination result of each data in the target data set, the dirty data detection result of the target shard table is generated, including: Based on whether each piece of data in the target data set is dirty data, counting the amount of dirty data in the target data set; Acquire the data volume of the data in the target data set to obtain the detection data volume, wherein the detection data volume is used to indicate the data volume of the data in the target shard table that has been subjected to dirty data detection; Based on the amount of dirty data in the target data set and the amount of detected data, a dirty data detection result of the target shard table is determined.
4. The dirty data detection method according to claim 3, characterized in that: After determining the dirty data detection result of the target shard table based on the amount of dirty data in the target data set and the amount of detected data, the method further includes: Obtaining the data volume of the data in the target data table to obtain the target data volume; Comparing the target data volume with the detected data volume to obtain a second comparison result; If the second comparison result indicates that the target data amount is consistent with the detected data amount, determining that each data in the target shard table has been subjected to dirty data detection; or When the second comparison result indicates that the target data amount and the detected data amount are inconsistent, it is determined that some data in the target shard table has not been subjected to dirty data detection.
5. The dirty data detection method according to claim 1, characterized in that: Read the data of the target shard table in the target shard database to obtain the target data set, including: Obtaining an identifier of a target sharding table, and determining a target query statement based on the identifier of the target sharding database and the identifier of the target sharding table; The target query statement is executed, and the data in the target shard table is read using a streaming read strategy to obtain the target data set.
6. The dirty data detection method according to claim 1, characterized in that: Calculate the shard identifier of each data item in the target data set to obtain a shard identifier set, including: Calculate the hash value of each data item in the target data set; The shard identifier set is determined based on the hash value of each data piece in the target data set.
7. The dirty data detection method according to claim 1, characterized in that: After determining the dirty data detection result of the target shard table based on the shard identifier set and the identifier of the target shard database, the method further includes: Obtain dirty data detection results of the shard tables of the target data table in each of the shard databases to obtain N dirty data detection results; Visually display the N dirty data detection results.
8. A dirty data detection device for a distributed database, characterized in that: include: a processing unit, configured to traverse each of the shard databases in the distributed database based on identifiers of the respective shard databases, and when traversing to a target shard database, read data of a target shard table in the target shard database to obtain a target data set, wherein the distributed database includes N shard databases, the target shard database is one of the N shard databases, and the target shard table includes: a shard of the target data table stored in the distributed database, where N is a positive integer; a calculation unit, configured to calculate a shard identifier for each piece of data in the target data set to obtain a shard identifier set, wherein the shard identifier set includes: a data shard identifier for each piece of data in the target data set, and the data shard identifier for each piece of data includes: an identifier of a shard database that logically stores the piece of data; A first determining unit is configured to determine a dirty data detection result of the target shard table based on the shard identifier set and an identifier of the target shard database, wherein the dirty data detection result includes: whether dirty data exists in the target data set.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the dirty data detection method for a distributed database according to any one of claims 1 to 7.
10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the dirty data detection method for a distributed database according to any one of claims 1 to 7 are implemented.