Data quality monitoring method, computer device and storage medium

CN117909763BActive Publication Date: 2026-09-18TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410128167.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2026-09-18
Estimated Expiration
2044-01-30

AI Technical Summary

Technical Problem

但是,这种通过预测模型监控数据质量的方式,需要大量准确的历史数据进行训练,如果用于训练预测模型的数据量不足,则可能导致预测模型的精确性下降,无法准确的预测数据质量

Benefits of technology

[0059]The aforementioned data quality monitoring method, apparatus, computer equipment, storage medium, and computer program product acquire a first comparison data group from a target monitoring data group, and acquire a second comparison data group from a sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group; determine a first target data group corresponding to the first comparison data group, and determine a second target data group corresponding to the second comparison data group; the first target data group is obtained based on the frequency of each sub-data in the first comparison data group; the second target data group is obtained based on the frequency of each sub-data in the second comparison data group; based on the first target data group, obtain a first probability distribution information of the first comparison data group, and based on the second target data group, obtain a second probability distribution information of the second comparison data group; based on the difference in distribution information between the first probability distribution information and the second probability distribution information, obtain the quality monitoring result of the target monitoring data group. By comparing the data distribution of the target monitoring data group with that of a sample data group with the same data format, this method can quickly and accurately analyze the quality monitoring results of the target monitoring data group, effectively improving the efficiency of data quality monitoring of the target monitoring data group. Compared with the traditional technical solution of training prediction models for data quality prediction, the data quality monitoring method of this application is simpler, more efficient and reliable in the overall process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117909763B_ABST
    Figure CN117909763B_ABST
Patent Text Reader

Abstract

The application relates to a data quality monitoring method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a first to-be-compared data group in a target monitoring data group, and acquiring a second to-be-compared data group in a sample data group corresponding to the target monitoring data group; determining a first target data group corresponding to the first to-be-compared data group, and determining a second target data group corresponding to the second to-be-compared data group; obtaining first probability distribution information of the first to-be-compared data group according to the first target data group, and obtaining second probability distribution information of the second to-be-compared data group according to the second target data group; and obtaining a quality monitoring result of the target monitoring data group according to a distribution information difference between the first probability distribution information and the second probability distribution information. The method can improve the data quality monitoring efficiency of the target monitoring data group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular to a data quality monitoring method, apparatus, computer equipment, storage medium, and computer program product. Background Technology

[0002] Data quality refers to the degree to which data meets the requirements of accuracy, completeness, and consistency. Data quality is crucial for the effective use of data and data-driven decision-making; low-quality data can lead to erroneous decision outcomes.

[0003] Traditional techniques typically use historical data to build predictive models of data quality, and then use these models to predict the quality of target data, monitoring the data quality based on the prediction results. However, this method of monitoring data quality through predictive models requires a large amount of accurate historical data for training. If the amount of data used to train the predictive model is insufficient, the accuracy of the predictive model may decrease, making it unable to accurately predict data quality. Summary of the Invention

[0004] Therefore, it is necessary to provide a data quality monitoring method, device, computer equipment, computer-readable storage medium, and computer program product that can improve the efficiency of data monitoring in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a data quality monitoring method. The method includes:

[0006] Obtain the first comparison data group from the target monitoring data group, and obtain the second comparison data group from the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0007] A first target data group corresponding to the first data group to be compared is determined, and a second target data group corresponding to the second data group to be compared is determined; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0008] Based on the first target data group, a first probability distribution information of the first comparison data group is obtained, and based on the second target data group, a second probability distribution information of the second comparison data group is obtained.

[0009] Based on the difference in distribution information between the first probability distribution information and the second probability distribution information, the quality monitoring result of the target monitoring data group is obtained.

[0010] In one embodiment, determining the first target data group corresponding to the first data group to be compared includes:

[0011] The amplification factor is obtained based on the data type of the first or second set of data to be compared.

[0012] Based on the amplification factor, the first initial data group is constructed; each sub-data in the first initial data group is used to characterize the preset initial frequency for each sub-data in the first comparison data group.

[0013] Based on the amplification factor and the first comparison data group, each sub-data in the first initial data group is updated to obtain the first target data group corresponding to the first comparison data group.

[0014] In one embodiment, based on the amplification factor and the first comparison data group, each sub-data in the first initial data group is updated to obtain a first target data group of the first comparison data group, including:

[0015] The amplification factor, the minimum value in the first data group to be compared, and each sub-data in the first data group to be compared are input into the data group index prediction model to obtain the data index to be updated in the first initial data group.

[0016] The sub-data in the first initial data group corresponding to the index of the data to be updated is updated to obtain the first target data group of the first initial data group.

[0017] In one embodiment, a first initial data set is constructed based on the amplification factor, including:

[0018] Determine the numerical difference between the maximum and minimum values ​​in the first set of data to be compared;

[0019] The numerical difference and the amplification factor are input into the data set size evaluation model to obtain the first data set size.

[0020] The first initial data group is constructed based on the size of the first data group.

[0021] In one embodiment, determining the second target data group corresponding to the second data group to be compared includes:

[0022] Based on the amplification factor, a second initial data set is constructed; each sub-data in the second initial data set is used to characterize the preset initial frequency for each sub-data in the second comparison data set;

[0023] Based on the amplification factor and the second set of data to be compared, each sub-data in the second initial data set is updated to obtain the second target data set corresponding to the second set of data to be compared.

[0024] In one embodiment, obtaining a first probability distribution information of the first comparison data group based on the first target data group, and obtaining a second probability distribution information of the second comparison data group based on the second target data group, includes:

[0025] Obtain the sum of all sub-data in the first target data group, and obtain the sum of all sub-data in the second target data group;

[0026] Each sub-data in the first target data group and all sub-data in the first target data group are input into the probability distribution detection model to obtain the first probability distribution information of the first data group to be compared.

[0027] Each sub-data point in the second target data group and all sub-data points in the second target data group are input into the probability distribution detection model to obtain the second probability distribution information of the second comparison data group.

[0028] In one embodiment, obtaining a first comparison data group from the target monitoring data group and obtaining a second comparison data group from the sample data group corresponding to the target monitoring data group includes:

[0029] A first candidate data group is obtained from the target monitoring data group, and a second candidate data group corresponding to the first candidate data group is obtained from the sample data group; wherein the data in the first candidate data group and the second candidate data group have the same meaning;

[0030] Determine the candidate numerical difference between the maximum and minimum values ​​in the first candidate data group;

[0031] If the difference between the candidate values ​​does not exceed a preset difference threshold, then the first candidate data group is determined as the first data group to be compared, and the second candidate data group is determined as the second data group to be compared.

[0032] If the difference between the candidate values ​​exceeds the preset difference threshold, then a first candidate data group is re-acquired from the target monitoring data group, and a second candidate data group corresponding to the first candidate data group is re-acquired from the sample data group.

[0033] In one embodiment, the quality monitoring result of the target monitoring data group is obtained based on the difference in distribution information between the first probability distribution information and the second probability distribution information, including:

[0034] Obtain the statistical characteristic indicators of the target monitoring data group;

[0035] Based on the differences in distribution information and the statistical characteristic indicators of the target monitoring data group, the quality monitoring results of the target monitoring data group are obtained.

[0036] In one embodiment, the quality monitoring result of the target monitoring data group is obtained based on the distribution information differences and the statistical characteristic indicators of the target monitoring data group, including:

[0037] If the statistical characteristic indicators of the target monitoring data group exceed the preset indicator threshold, and / or the difference in the distribution information exceeds the preset distribution threshold, then the quality monitoring result of the target monitoring data group is confirmed to be abnormal, and the processing tasks associated with the target monitoring data group are blocked.

[0038] If the statistical characteristic indicators of the target monitoring data group do not exceed the preset indicator threshold, and the difference in distribution information does not exceed the preset difference threshold, then the quality monitoring result of the target monitoring data group is confirmed to be normal, and a prompt message for the quality monitoring result is generated.

[0039] Secondly, this application also provides a data quality monitoring device. The device includes:

[0040] The comparison data group filtering module is used to obtain a first comparison data group in the target monitoring data group and a second comparison data group in the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0041] The distributed data group determination module is used to determine a first target data group corresponding to the first data group to be compared, and to determine a second target data group corresponding to the second data group to be compared; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0042] The probability distribution acquisition module is used to obtain the first probability distribution information of the first comparison data group based on the first target data group, and to obtain the second probability distribution information of the second comparison data group based on the second target data group.

[0043] The monitoring result acquisition module is used to obtain the quality monitoring result of the target monitoring data group based on the distribution information difference between the first probability distribution information and the second probability distribution information.

[0044] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0045] Obtain the first comparison data group from the target monitoring data group, and obtain the second comparison data group from the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0046] A first target data group corresponding to the first data group to be compared is determined, and a second target data group corresponding to the second data group to be compared is determined; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0047] Based on the first target data group, a first probability distribution information of the first comparison data group is obtained, and based on the second target data group, a second probability distribution information of the second comparison data group is obtained.

[0048] Based on the difference in distribution information between the first probability distribution information and the second probability distribution information, the quality monitoring result of the target monitoring data group is obtained.

[0049] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0050] Obtain the first comparison data group from the target monitoring data group, and obtain the second comparison data group from the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0051] A first target data group corresponding to the first data group to be compared is determined, and a second target data group corresponding to the second data group to be compared is determined; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0052] Based on the first target data group, a first probability distribution information of the first comparison data group is obtained, and based on the second target data group, a second probability distribution information of the second comparison data group is obtained.

[0053] Based on the difference in distribution information between the first probability distribution information and the second probability distribution information, the quality monitoring result of the target monitoring data group is obtained.

[0054] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0055] Obtain the first comparison data group from the target monitoring data group, and obtain the second comparison data group from the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0056] A first target data group corresponding to the first data group to be compared is determined, and a second target data group corresponding to the second data group to be compared is determined; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0057] Based on the first target data group, a first probability distribution information of the first comparison data group is obtained, and based on the second target data group, a second probability distribution information of the second comparison data group is obtained.

[0058] Based on the difference in distribution information between the first probability distribution information and the second probability distribution information, the quality monitoring result of the target monitoring data group is obtained.

[0059] The aforementioned data quality monitoring method, apparatus, computer equipment, storage medium, and computer program product acquire a first comparison data group from a target monitoring data group, and acquire a second comparison data group from a sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group; determine a first target data group corresponding to the first comparison data group, and determine a second target data group corresponding to the second comparison data group; the first target data group is obtained based on the frequency of each sub-data in the first comparison data group; the second target data group is obtained based on the frequency of each sub-data in the second comparison data group; based on the first target data group, obtain a first probability distribution information of the first comparison data group, and based on the second target data group, obtain a second probability distribution information of the second comparison data group; based on the difference in distribution information between the first probability distribution information and the second probability distribution information, obtain the quality monitoring result of the target monitoring data group. By comparing the data distribution of the target monitoring data group with that of a sample data group with the same data format, this method can quickly and accurately analyze the quality monitoring results of the target monitoring data group, effectively improving the efficiency of data quality monitoring of the target monitoring data group. Compared with the traditional technical solution of training prediction models for data quality prediction, the data quality monitoring method of this application is simpler, more efficient and reliable in the overall process. Attached Figure Description

[0060] Figure 1 This is a flowchart illustrating a data quality monitoring method in one embodiment;

[0061] Figure 2 This is a flowchart illustrating the steps of determining the first target data group corresponding to the first data group to be compared in one embodiment;

[0062] Figure 3 This is a flowchart illustrating the steps of determining the second target data group corresponding to the second data group to be compared in one embodiment;

[0063] Figure 4 This is a flowchart illustrating the steps of obtaining a first comparison data group in a target monitoring data group and a second comparison data group in a sample data group corresponding to the target monitoring data group in one embodiment.

[0064] Figure 5 This is a flowchart illustrating a data quality monitoring method in another embodiment;

[0065] Figure 6 This is a timing diagram of data scheduling interactions performed by the server in one embodiment;

[0066] Figure 7 This is a structural block diagram of a data quality monitoring device in one embodiment;

[0067] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0069] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0070] In one embodiment, such as Figure 1 As shown, a data quality monitoring method is provided. This embodiment illustrates the method applied to a server, but it is understood that the method can also be applied to a terminal, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0071] Step S101: Obtain the first data group to be compared in the target monitoring data group, and obtain the second data group to be compared in the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0072] The target monitoring data set refers to the task data whose quality needs to be monitored. The sample data set refers to verified and accurate data. It should be noted that the data format and meaning of the sample data set are the same as those of the target monitoring data set, but the data volume of the target monitoring data set and the sample data set may differ within a certain range.

[0073] In this context, data format refers to the storage format of the data, such as the number of columns. Data meaning refers to the physical meaning of the data. For example, a target monitoring data set can be stored in the form of a data table. If the target monitoring data set contains 5 columns of data, then the sample data set must also contain 5 columns of data. However, the amount of data in the corresponding column in the target monitoring data set and the sample data set can be different. For example, the amount of data in the first column of the target monitoring data set may be 1 million, while the amount of data in the first column of the sample data set may be 1.01 million.

[0074] Specifically, the server, through a configured interface, retrieves the target monitoring data group (for data quality monitoring) from data sources such as relational databases, file systems (including single-machine or distributed file systems), and distributed data warehouses. It also retrieves a sample data group that has been verified, confirmed to be accurate, and whose data format and meaning are identical to the target monitoring data group. The server compares the data volume of the target monitoring data group with that of the sample data group. If the difference in data volume is too large (e.g., exceeding a preset threshold), further comparison steps are unnecessary. If the target monitoring data group is too large, data sampling can be performed, extracting a portion of the data as the monitoring data. Similarly, if the sample data group is too large, data sampling can also be performed, extracting a portion of the data as a new sample data group.

[0075] Furthermore, a column of data for comparative analysis is determined from the target monitoring data group (or the data to be monitored) and set as the first data group to be compared, and a second data group to be compared is obtained from the sample data group (or the sampled sample data group) at the position corresponding to the first data group to be compared.

[0076] Step S102: Determine the first target data group corresponding to the first data group to be compared, and determine the second target data group corresponding to the second data group to be compared; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0077] Here, the first target data set refers to the data used to analyze the probability distribution of the first comparison data set. The second target data set refers to the data used to analyze the probability distribution of the second comparison data set. Both the first and second target data sets can be in array format.

[0078] Specifically, the server can construct an initial data group for the first data group to be compared and the second data group to be compared, i.e., construct a first initial data group and a second initial data group. Then, it can update the data in the first initial data group according to the first data group to be compared. This can be done by determining the data that needs to be updated in the first initial data group based on the first data group to be compared, calculating the data that needs to be updated in the first initial data group to obtain new data, and then replacing the data that needs to be updated with the new data. In this way, the server obtains the updated first target data group. Similarly, it can update the data in the second initial data group according to the second data group to be compared to obtain the updated second target data group.

[0079] Step S103: Based on the first target data group, obtain the first probability distribution information of the first comparison data group, and based on the second target data group, obtain the second probability distribution information of the second comparison data group.

[0080] The first probability distribution information describes the probability distribution of the first set of data to be compared. The second probability distribution information is similar.

[0081] The server calculates the probability distribution information of the first and second data groups to be compared, respectively. Specifically, the server can pre-build a probability distribution detection model, and then input the first and second target data groups into the probability distribution detection model to output the first and second probability distribution information.

[0082] Step S104: Based on the difference in distribution information between the first probability distribution information and the second probability distribution information, obtain the quality monitoring results of the target monitoring data group.

[0083] Among them, quality monitoring results refer to information used to reflect the data quality of the target monitoring data.

[0084] Specifically, the server calculates the difference between the first probability distribution information and the second probability distribution information to obtain the distribution information difference. This can be achieved by inputting the first and second probability distribution information into a distribution difference evaluation model, thus obtaining the distribution information difference between the first and second probability distribution information. After obtaining the distribution information difference between the extracted first and second data groups to be compared in this round, the server jumps to step S101 above to proceed to the next round, continuing to determine the next column of data for comparative analysis from the target monitoring data group (or the data to be monitored), and resetting it as the first data group to be compared. Additionally, the server obtains the second data group to be compared from the sample data group (or the sampled sample data group) at the position corresponding to the reset first data group to be compared. Ultimately, the server can obtain multiple distribution information differences.

[0085] Furthermore, the server compares multiple distribution information differences with a preset distribution threshold; if none of the multiple distribution information differences exceed the preset difference threshold, the quality monitoring result of the target monitoring data group is confirmed to be normal, and a prompt message for the quality monitoring result is generated; if at least one of the multiple distribution information differences exceeds the preset distribution threshold, the quality monitoring result of the target monitoring data group is confirmed to be abnormal, and the processing tasks associated with the target monitoring data group are blocked.

[0086] In practical applications, a distribution difference assessment model based on Jensen-Shannon divergence (JS divergence) can be used to process the first and second probability distribution information using JS divergence to obtain the distribution difference between them. The distribution difference assessment model based on JS divergence can be expressed by the following formula:

[0087]

[0088] In the formula, JSD(P||Q) represents the difference in distribution information between the first probability distribution information P and the second probability distribution information Q.

[0089] In the aforementioned data quality monitoring method, a first comparison data group is selected from the target monitoring data group, and a second comparison data group is obtained from the corresponding sample data group. A first target data group corresponding to the first comparison data group is determined, and a second target data group corresponding to the second comparison data group is determined. Based on the first target data group, a first probability distribution information of the first comparison data group is obtained, and based on the second target data group, a second probability distribution information of the second comparison data group is obtained. Based on the difference in distribution information between the first and second probability distribution information, the quality monitoring result of the target monitoring data group is obtained. Using this method, by comparing the data distribution between the target monitoring data group and a sample data group with the same data format, the quality monitoring result of the target monitoring data group is quickly and accurately analyzed, effectively improving the efficiency of data quality monitoring of the target monitoring data group. Compared with traditional technical solutions for training prediction models to predict data quality, the data quality monitoring method of this application is simpler, more efficient, and more reliable overall.

[0090] In one embodiment, such as Figure 2 As shown, step S102 above, which determines the first target data group corresponding to the first data group to be compared, specifically includes the following:

[0091] Step S201: Obtain the amplification factor according to the data type of the first or second data group to be compared.

[0092] The data types include, but are not limited to, byte (the smallest data type), short (short integer), int (integer), long (long integer), float (floating-point), double (double-precision floating-point), char (character), and boolean (boolean).

[0093] Specifically, the server can pre-set the correspondence between different data types and amplification factors. Since the first and second data groups to be compared have the same data meaning, the data types of the first and second data groups to be compared are the same. When the server obtains the data type of the first or second data group to be compared, it can query the correspondence based on the data type of the first or second data group to be compared to obtain the amplification factor.

[0094] The magnification factor indicates the level of detail in the detection of data distribution; the larger the magnification factor, the more detailed the examination of the data distribution, meaning a more comprehensive examination of the data distribution. In essence, the magnification factor amplifies the data scale, allowing the server to more clearly observe the distribution characteristics of the data, thereby enabling more accurate identification of potential distributions, outliers, or data trends.

[0095] In practical applications, the magnification factor for integers can be set to 1, and the magnification factor for other data types can be set to 10, or other values ​​can be used.

[0096] Step S202: Construct a first initial data group based on the amplification factor; each sub-data in the first initial data group is used to characterize the preset initial frequency for each sub-data in the first comparison data group.

[0097] Specifically, the server determines the size of the initial data group to be constructed based on the amplification factor, and then constructs an array with all identical data (e.g., 0, i.e., the initial frequency is set to 0) according to the size, and sets it as the first initial data group.

[0098] Step S203: Update each sub-data in the first initial data group according to the amplification factor and the first comparison data group to obtain the first target data group corresponding to the first comparison data group.

[0099] Specifically, the server determines the index of the data that needs to be updated in the first initial data group based on the amplification factor and the first data group to be compared. Then, it updates the data at the corresponding index in the first initial data group. After all updates are completed, the first target data group is obtained.

[0100] In this embodiment, the amplification factor is first obtained based on the data type of the first data group to be compared or the data type of the second data group to be compared; then, a first initial data group is constructed based on the amplification factor; the first initial data group is updated based on the amplification factor and the first data group to be compared to obtain the first target data group corresponding to the first data group to be compared to. This realizes the detection process of the data distribution of the first data group to be compared to be compared, so that the data distribution of the first data group to be compared to be compared can be calculated based on the first target data group in subsequent steps, and the data quality analysis of the target monitoring data group can be completed by comparing the differences in the data distribution.

[0101] In one embodiment, step S203 above, which updates the first initial data group according to the amplification factor and the first data group to be compared, to obtain the first target data group corresponding to the first data group to be compared, specifically includes the following: inputting the amplification factor, the minimum value in the first data group to be compared, and each sub-data in the first data group to be compared into the data group index prediction model to obtain the data index to be updated in the first initial data group; updating the sub-data in the first initial data group corresponding to the data index to be updated to obtain the first target data group of the first initial data group.

[0102] The data set index prediction model is used to predict the index of the data that needs to be updated in the first initial data set. In practical applications, the data set index prediction model can be expressed by the following formula: Index of data to be updated = Sub-data in the comparison data set * Magnification factor - Minimum value in the comparison data set; where the comparison data set can be either the first comparison data set or the second comparison data set. This formula is a linear transformation mathematical formula, which can scale the data to a certain range to facilitate subsequent calculation of the probability distribution.

[0103] Specifically, the server iterates through the sub-data in the first data group to be compared, and then inputs each sub-data, the amplification factor, and the minimum value in the first data group to the data group indicator prediction model to calculate the data index to be updated in the first initial data group. The server queries the data in the first initial data group corresponding to the data index to be updated and increments the value of that data by 1. After all the sub-data in the first data group to be compared has been traversed and all the corresponding data indices to be updated in the first initial data group have been updated, the server obtains the first target data group. The updated data in the target data group is used to represent the frequency of occurrence of the sub-data in the data group to be compared.

[0104] In this embodiment, by inputting the amplification factor, the minimum value in the first data group to be compared, and each sub-data in the first data group to be compared into the data group index prediction model, the index of the data to be updated in the first initial data group is calculated; then, the sub-data in the first initial data group corresponding to the index of the data to be updated is updated to obtain the first target data group of the first initial data group, thus realizing the reasonable construction of the first target data group. This allows subsequent steps to determine the data distribution of the first data to be compared based on the first target data group and to use the data distribution to analyze the data quality of the target monitoring data group.

[0105] In one embodiment, step S202 above, constructing a first initial data set based on the amplification factor, specifically includes the following: determining the numerical difference between the maximum and minimum values ​​in the first data set to be compared; inputting the numerical difference and the amplification factor into the data set size evaluation model to obtain the size of the first data set; and constructing the first initial data set based on the size of the first data set.

[0106] The data set size evaluation model is used to assess the size of data sets used for data distribution analysis. In practical applications, the data set size evaluation model can be expressed by the following formula: Data set size = (maximum value in the comparison data set - minimum value in the comparison data set) * magnification factor; where the comparison data set can be either the first comparison data set or the second comparison data set.

[0107] Specifically, the server selects the maximum and minimum values ​​from the first set of data to be compared, and then calculates the numerical difference between the maximum and minimum values. The numerical difference and the magnification factor are processed using a data set size evaluation model; this can be done by multiplying the numerical difference and the magnification factor to obtain the size of the first data set. The server constructs an array with the size of the first data set to obtain the first initial data set. Default values ​​can also be set for the first initial data set, such as setting it to an array of all zeros.

[0108] In this embodiment, the size of the first data group is calculated by the numerical difference between the maximum and minimum values ​​in the first data group to be compared and the amplification factor; then, based on the size of the first data group, the first initial data group is constructed, thus realizing the reasonable construction of the initial data group and setting a default value for it.

[0109] In one embodiment, such as Figure 3 As shown, step S102 above, which determines the second target data group corresponding to the second data group to be compared, specifically includes the following:

[0110] Step S301: Construct a second initial data set based on the amplification factor; each sub-data in the second initial data set is used to characterize the preset initial frequency for each sub-data in the second comparison data set.

[0111] Step S302: Update each sub-data in the second initial data group according to the amplification factor and the second comparison data group to obtain the second target data group corresponding to the second comparison data group.

[0112] It is understood that the process of obtaining the amplification factor is the same as that of obtaining the amplification factor in the above embodiments, the process of constructing the second initial data group is also the same as that of constructing the first initial data group in the above embodiments, and the process of updating the second initial data group to obtain the second target data group is also the same as that of obtaining the first target data group in the above embodiments, and will not be repeated here.

[0113] In this embodiment, the amplification factor is first obtained based on the data type of the first or second data group to be compared; then, a second initial data group is constructed based on the amplification factor; the second initial data group is updated based on the amplification factor and the second data group to be compared, to obtain the second target data group corresponding to the second data group to be compared. This realizes the detection process of the data distribution of the second data group to be compared, so that the data distribution of the second data group to be compared can be calculated based on the second target data group in subsequent steps, and the data quality monitoring of the target monitoring data group can be completed by comparing the differences in the data distribution.

[0114] In one embodiment, step S103, which obtains the first probability distribution information of the first comparison data group based on the first target data group and the second probability distribution information of the second comparison data group based on the second target data group, specifically includes the following: obtaining the sum of all sub-data in the first target data group and obtaining the sum of all sub-data in the second target data group; inputting each sub-data in the first target data group and the sum of all sub-data in the first target data group into the probability distribution detection model to obtain the first probability distribution information of the first comparison data group; inputting each sub-data in the second target data group and the sum of all sub-data in the second target data group into the probability distribution detection model to obtain the second probability distribution information of the second comparison data group.

[0115] The probability distribution detection model is used to calculate the probability distribution information of the data set. It can be understood that each sub-data point in the target data set represents the frequency of occurrence of the corresponding data in the data set to be compared.

[0116] Specifically, the server adds up all the sub-data in the first target data group to obtain the sum of all the sub-data in the first target data group. Similarly, the server obtains the sum of all the sub-data in the second target data group. The server iterates through each sub-data in the first target data group, and sequentially inputs each sub-data and the sum of all the sub-data in the first target data group into the probability distribution detection model to calculate the first probability distribution information of the first comparison data group. The server obtains the second probability distribution information in the same way as it obtains the first probability distribution information, and will not be described again here.

[0117] In practical applications, the probability distribution detection model can be represented by the following formula: Probability distribution information = Sub-data in the target data group / Sum of all sub-data in the target data group. In this formula, the probability distribution information can be either a first probability distribution or a second probability distribution; the target data group can be either a first target data group or a second target data group; and the sum of the data groups can be either the sum of the first target data group or the sum of the second target data group.

[0118] For example, suppose the target data set includes five sub-data points: a, b, c, d, and e. a, b, c, d, and e represent the frequencies of occurrence of the five different sub-data points in the corresponding comparison data set. Then, the probability distribution information of the comparison data set corresponding to the target data set is {a / (a+b+c+d+e), b / (a+b+c+d+e), c / (a+b+c+d+e), d / (a+b+c+d+e), e / (a+b+c+d+e)}.

[0119] In this embodiment, the server obtains the first probability distribution information of the first comparison data group by inputting each sub-data in the first target data group and the sum of the data in the first target data group into the probability distribution detection model; and obtains the second probability distribution information of the second comparison data group by inputting each sub-data in the second target data group and the sum of the data in the second target data group into the probability distribution detection model. This achieves rapid acquisition of the probability distribution information of the first and second comparison data groups. The acquisition method of probability distribution is simple and has low computational complexity. Subsequent steps can also use the first and second probability distribution information as the basis for data quality analysis, which greatly improves the data quality monitoring efficiency of the target monitoring data group.

[0120] In one embodiment, such as Figure 4 As shown, step S101 above, which involves obtaining the first comparison data group from the target monitoring data group and the second comparison data group from the corresponding sample data group, specifically includes the following:

[0121] Step S401: Obtain a first candidate data group from the target monitoring data group, and obtain a second candidate data group corresponding to the first candidate data group from the sample data group; wherein, the data in the first candidate data group and the second candidate data group have the same meaning.

[0122] Step S402: Determine the candidate numerical difference between the maximum and minimum values ​​in the first candidate data group.

[0123] In step S403, if the difference between the candidate values ​​does not exceed the preset difference threshold, the first candidate data group is determined as the first data group to be compared, and the second candidate data group is determined as the second data group to be compared.

[0124] Step S404: If the difference between candidate values ​​exceeds a preset difference threshold, then the first candidate data group is retrieved again from the target monitoring data group, and the second candidate data group corresponding to the first candidate data group is retrieved again from the sample data group.

[0125] Specifically, the server extracts a column of data to be compared from the target monitoring data group as the first candidate data group, and obtains a second candidate data group from the sample data group whose corresponding column is the first candidate data group. It is understandable that since the target monitoring data group and the sample data group have the same data format and meaning, and the first and second candidate data groups contain the same column data from the target monitoring data group and the sample data group respectively, the meanings of the data in the first and second candidate data groups obtained here are also the same. For example, if the first column of data in the target monitoring data group represents height, then the first column of data in the sample data group also represents height; if the second column of data in the target monitoring data group represents age, then the second column of data in the sample data group also represents age.

[0126] Further, the server determines the difference between the maximum and minimum values ​​in the current first candidate data group, setting it as the candidate numerical difference. Since the sample data group is pre-verified and accurate, there is no need to verify the second candidate data group here. The candidate numerical difference is compared with a preset difference threshold. If the candidate numerical difference does not exceed the preset difference threshold, the first candidate data group is determined as the first comparison data group, and the second candidate data group is determined as the second comparison data group. In addition, the server can also filter out null values ​​in the first and second comparison data groups. If the candidate numerical difference exceeds the preset difference threshold, the server retrieves the next column of data to be compared from the target monitoring data group and sets it as the first candidate data group, and obtains the second candidate data group corresponding to the new first candidate data group from the sample data group, then proceeds to the step of comparing the first and second numerical differences with the preset difference thresholds respectively.

[0127] In this embodiment, a first candidate data group is obtained from the target monitoring data group, and a second candidate data group corresponding to the first candidate data group is obtained from the sample data group; the candidate numerical difference between the maximum and minimum values ​​in the first candidate data group is determined; then the candidate numerical difference is compared with a preset difference threshold to determine the first and second data groups to be compared. By using the numerical difference of the candidate data groups, the data groups used for comparative analysis in the target monitoring data group and the sample data group are selected, saving the data quality analysis time of the data groups with large numerical differences and further improving the efficiency of data quality analysis of the target monitoring data group.

[0128] In one embodiment, step S104, obtaining the quality monitoring result of the target monitoring data group based on the distribution information difference between the first probability distribution information and the second probability distribution information, specifically includes the following: obtaining the statistical characteristic indicators of the target monitoring data group; obtaining the quality monitoring result of the target monitoring data group based on the distribution information difference and the statistical characteristic indicators of the target monitoring data group.

[0129] Among them, statistical characteristic indicators refer to indicators that can reflect the statistical characteristics of data. For example, statistical characteristic indicators include indicators such as maximum value, minimum value, average value, median, and variance.

[0130] To further improve the accuracy of data quality analysis, the server can also analyze the data quality of the target monitoring data group from other perspectives. Specifically, if the target monitoring data group is scalar data, the server calculates multiple statistical characteristic indicators for both the target monitoring data group and the sample data group. Then, based on the differences in distribution information, the statistical characteristic indicators of the target monitoring data group, and the statistical characteristic indicators of the sample data group, the server obtains the quality monitoring results for the target monitoring data group. Since the sample data group is pre-verified and accurate, there is no need to verify the statistical characteristic indicators of the sample data group here.

[0131] In this embodiment, when the target monitoring data group is scalar data, the server can also obtain the statistical characteristic indicators of the target monitoring data group and combine them with the previously obtained distribution information differences to comprehensively analyze the quality monitoring results of the target monitoring data group, which expands the data quality analysis perspective of the target monitoring data group and helps to improve the accuracy of the quality monitoring results.

[0132] In one embodiment, the quality monitoring result of the target monitoring data group is obtained based on the distribution information difference and the statistical characteristic indicators of the target monitoring data group. Specifically, this includes the following: if the statistical characteristic indicators of the target monitoring data group exceed a preset indicator threshold, and / or the distribution information difference exceeds a preset distribution threshold, then the quality monitoring result of the target monitoring data group is confirmed as abnormal, and the processing tasks associated with the target monitoring data group are blocked; if the statistical characteristic indicators of the target monitoring data group do not exceed the preset indicator threshold, and the distribution information difference does not exceed a preset difference threshold, then the quality monitoring result of the target monitoring data group is confirmed as normal, and a prompt message for the quality monitoring result is generated.

[0133] Specifically, the server can pre-set threshold values ​​for statistical indicator characteristics and distribution threshold values ​​for differences in distribution information. After acquiring the statistical indicator values ​​and multiple distribution information differences of the target monitoring data group, the server compares the statistical indicator values ​​of the target monitoring data group with the preset threshold values, and compares each distribution information difference with a preset distribution threshold value. If the statistical indicator values ​​of the target monitoring data group exceed the preset threshold values, and / or at least one of the multiple distribution information differences exceeds the preset distribution threshold value, the quality monitoring result of the target monitoring data group is confirmed as abnormal, and the server can block the processing tasks associated with the target monitoring data group. If the statistical indicator values ​​of the target monitoring data group do not exceed the preset threshold values, and the multiple distribution information differences do not exceed the preset distribution threshold values, the quality monitoring result of the target monitoring data group is confirmed as normal, and the server can generate a quality monitoring result prompt message for security alerts.

[0134] In this embodiment, the server determines whether the quality monitoring result of the target monitoring data group is normal or abnormal by using statistical characteristic indicators and differences in multiple distribution information of the target monitoring data group, and takes corresponding processing measures based on the quality monitoring result. This can promptly block the execution of downstream processing tasks for abnormal data, prevent downstream processing tasks from using incorrect data to obtain incorrect processing results, and improve the reliability of the target monitoring data group.

[0135] In one embodiment, such as Figure 5 As shown, another data quality monitoring method is provided. Taking the application of this method to a server as an example, the steps include:

[0136] Step S501: Obtain the first comparison data group in the target monitoring data group, and obtain the second comparison data group in the sample data group corresponding to the target monitoring data group.

[0137] Step S502: Determine the first target data group corresponding to the first data group to be compared, and determine the second target data group corresponding to the second data group to be compared.

[0138] The first target data group is obtained based on the frequency of each sub-data in the first comparison data group; the second target data group is obtained based on the frequency of each sub-data in the second comparison data group.

[0139] Step S503: Based on the first target data group, obtain the first probability distribution information of the first comparison data group, and based on the second target data group, obtain the second probability distribution information of the second comparison data group.

[0140] Step S504: Determine the difference in distribution information between the first probability distribution information and the second probability distribution information.

[0141] Step S505: Obtain the statistical characteristic indicators of the target monitoring data group.

[0142] It should be noted that step S505 is not limited to being executed after step S504, but can also be executed in parallel with any of the steps S501 to S504.

[0143] Step S506: If the statistical characteristic indicators of the target monitoring data group exceed the preset indicator threshold, and / or the difference in distribution information exceeds the preset distribution threshold, then the quality monitoring result of the target monitoring data group is confirmed to be abnormal, and the processing tasks associated with the target monitoring data group are blocked.

[0144] Step S507: If the statistical characteristic indicators of the target monitoring data group do not exceed the preset indicator threshold and the distribution information difference does not exceed the preset difference threshold, then the quality monitoring result of the target monitoring data group is confirmed to be normal, and a prompt message for the quality monitoring result is generated.

[0145] The above-mentioned data quality monitoring method can achieve the following beneficial effects: by comparing the data distribution between the target monitoring data group and the sample data group with the same data format, the quality monitoring results of the target monitoring data group can be analyzed quickly and accurately, which effectively improves the efficiency of data quality monitoring of the target monitoring data group. Compared with the traditional technical solution of training prediction models for data quality prediction, the data quality monitoring method of this application is simpler, more efficient and reliable in the whole process.

[0146] To more clearly illustrate the data quality monitoring method provided in this disclosure, a specific embodiment is given below to describe the above-mentioned data quality monitoring method. Another data quality monitoring method is also provided, which can be applied to a server, and specifically includes the following:

[0147] (1) Collect target monitoring data sets with a data volume of approximately 1 million to 10 million records, collect verified and accurate sample data sets, and ensure that the data meaning and data format of the sample data sets are the same as those of the target monitoring data sets.

[0148] (2) Compare the data volume of the sample data group with the data volume of the target monitoring data group. If they are the same, continue with the subsequent steps.

[0149] (3) If the target monitoring data group is scalar data, obtain the maximum, minimum, average, median and variance of the target monitoring data group and the sample data group.

[0150] (4) Obtain the first probability distribution information of each first comparison data group in the target monitoring data group, and obtain the second probability distribution information of each second comparison data in the sample data group, and determine the distribution information difference between each first probability distribution information and each second probability distribution information.

[0151] (5) Compare the differences in each distribution information with the preset distribution threshold, and compare each statistical characteristic indicator with the preset indicator threshold. If any one of them exceeds the threshold, the target monitoring data group is considered abnormal and the downstream tasks of the target monitoring data group can be blocked. If all information does not exceed the threshold, a safety prompt for the quality monitoring results is issued.

[0152] This embodiment consumes fewer computational resources and is faster. Experiments have verified that a target monitoring data set of 1 billion records can be processed within one minute using 40 cores and 200GB of memory. Compared to traditional methods that use machine learning-based predictive models for data quality prediction, this embodiment can save over 90% of the data quality analysis time.

[0153] In one embodiment, the above data quality monitoring method is specifically described from the server's perspective. Figure 6 The sequence diagram for data scheduling interactions with the server includes the following:

[0154] Configuration and front-end modules: These are used to configure the data to be monitored and related thresholds, and provide query interfaces and methods for other modules to use.

[0155] Data Acquisition Module: Used to collect data that needs to be monitored, including target monitoring data groups and sample data groups. The data acquisition module supports multiple data sources, such as relational databases, file systems (single-machine file systems or distributed file systems), distributed data warehouses, etc.

[0156] The data comparison module compares the target monitoring data set and the sample data set, and saves the differences in distribution information and quality monitoring results to a specified data storage. The storage of quality monitoring results includes, but is not limited to, relational databases, files, and distributed storage systems.

[0157] Data blocking and alert module: Determines whether to block or alert downstream tasks based on quality monitoring results.

[0158] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0159] Based on the same inventive concept, this application also provides a data quality monitoring device for implementing the data quality monitoring method described above. The solution provided by this device is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more data quality monitoring device embodiments provided below can be found in the limitations of the data quality monitoring method described above, and will not be repeated here.

[0160] In one embodiment, such as Figure 7 As shown, a data quality monitoring device 700 is provided, comprising:

[0161] The comparison data group filtering module 701 is used to obtain the first comparison data group in the target monitoring data group and the second comparison data group in the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group.

[0162] The distributed data group determination module 702 is used to determine the first target data group corresponding to the first data group to be compared, and to determine the second target data group corresponding to the second data group to be compared; the first target data group is obtained based on the frequency of each sub-data in the first data group to be compared; the second target data group is obtained based on the frequency of each sub-data in the second data group to be compared.

[0163] The probability distribution acquisition module 703 is used to obtain the first probability distribution information of the first comparison data group based on the first target data group, and to obtain the second probability distribution information of the second comparison data group based on the second target data group.

[0164] The monitoring result acquisition module 704 is used to obtain the quality monitoring result of the target monitoring data group based on the difference in distribution information between the first probability distribution information and the second probability distribution information.

[0165] In one embodiment, the distributed data group determination module 702 is further configured to obtain an amplification factor based on the data type of the first data group to be compared or the second data group to be compared; each sub-data in the first initial data group is used to characterize the preset initial frequency for each sub-data in the first data group to be compared; a first initial data group is constructed based on the amplification factor; and each sub-data in the first initial data group is updated based on the amplification factor and the first data group to be compared to obtain a first target data group corresponding to the first data group to be compared.

[0166] In one embodiment, the data quality monitoring device 700 further includes a distributed data group update module, which is used to input the amplification coefficient, the minimum value in the first data group to be compared, and each sub-data in the first data group to be compared into the data group index prediction model to obtain the data index to be updated in the first initial data group; and to update the sub-data in the first initial data group corresponding to the data index to be updated to obtain the first target data group of the first initial data group.

[0167] In one embodiment, the data quality monitoring device 700 further includes a distributed data group construction module, used to determine the numerical difference between the maximum and minimum values ​​in the first data group to be compared; input the numerical difference and the amplification factor into the data group size evaluation model to obtain the first data group size; and construct a first initial data group based on the first data group size.

[0168] In one embodiment, the distributed data group determination module 702 is further configured to construct a second initial data group based on the amplification factor; each sub-data in the second initial data group is used to characterize the preset initial frequency for each sub-data in the second comparison data group; and each sub-data in the second initial data group is updated based on the amplification factor and the second comparison data group to obtain the second target data group corresponding to the second comparison data group.

[0169] In one embodiment, the probability distribution acquisition module 703 is further configured to acquire the sum of all sub-data in the first target data group and the sum of all sub-data in the second target data group; input each sub-data in the first target data group and the sum of all sub-data in the first target data group into the probability distribution detection model to obtain the first probability distribution information of the first data group to be compared; input each sub-data in the second target data group and the sum of all sub-data in the second target data group into the probability distribution detection model to obtain the second probability distribution information of the second data group to be compared.

[0170] In one embodiment, the comparison data group filtering module 701 is further configured to obtain a first candidate data group from the target monitoring data group and a second candidate data group corresponding to the first candidate data group from the sample data group; wherein the data meaning of the first candidate data group and the second candidate data group is the same; determine a first numerical difference between the maximum and minimum values ​​in the first candidate data group and determine a second numerical difference between the maximum and minimum values ​​in the second candidate data group; if neither the first numerical difference nor the second numerical difference exceeds a preset difference threshold, then the first candidate data group is determined as the first data group to be compared and the second candidate data group is determined as the second data group to be compared; if at least one of the first numerical difference and the second numerical difference exceeds the preset difference threshold, then the first candidate data group is re-obtained from the target monitoring data group and the second candidate data group corresponding to the first candidate data group is re-obtained from the sample data group.

[0171] In one embodiment, the monitoring result acquisition module 704 is further configured to acquire statistical characteristic indicators of the target monitoring data group; and obtain the quality monitoring results of the target monitoring data group based on the differences in distribution information and the statistical characteristic indicators of the target monitoring data group.

[0172] In one embodiment, the data quality monitoring device 700 further includes a data quality monitoring module, configured to: if the statistical characteristic indicators of the target monitoring data group exceed a preset indicator threshold, and / or the distribution information difference exceeds a preset distribution threshold, then confirm that the quality monitoring result of the target monitoring data group is abnormal and block the processing tasks associated with the target monitoring data group; if the statistical characteristic indicators of the target monitoring data group do not exceed the preset indicator threshold, and the distribution information difference does not exceed a preset difference threshold, then confirm that the quality monitoring result of the target monitoring data group is normal and generate a prompt message for the quality monitoring result.

[0173] Each module in the aforementioned data quality monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0174] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores target monitoring data sets, sample data sets, and other data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a data quality monitoring method.

[0175] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0176] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0177] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0178] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0179] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0180] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0181] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data quality monitoring method, characterized in that, The method includes: Obtain the first comparison data group from the target monitoring data group, and obtain the second comparison data group from the sample data group corresponding to the target monitoring data group; the data format of the target monitoring data group is the same as the data format of the sample data group. A magnification factor is obtained based on the data type of the first or second data group to be compared; a first initial data group is constructed based on the magnification factor; each sub-data in the first initial data group is used to characterize the preset initial frequency for each sub-data in the first data group to be compared; the magnification factor is used to indicate the level of detail in the detection of the data distribution. Based on the amplification factor and the first comparison data group, each sub-data in the first initial data group is updated to obtain a first target data group corresponding to the first comparison data group, and a second target data group corresponding to the second comparison data group is determined; the first target data group is obtained based on the frequency of each sub-data in the first comparison data group; the second target data group is obtained based on the frequency of each sub-data in the second comparison data group. Based on the first target data group, a first probability distribution information of the first comparison data group is obtained, and based on the second target data group, a second probability distribution information of the second comparison data group is obtained. Based on the difference in distribution information between the first probability distribution information and the second probability distribution information, the quality monitoring result of the target monitoring data group is obtained.

2. The method according to claim 1, characterized in that, The step of updating each sub-data in the first initial data group according to the amplification factor and the first comparison data group to obtain the first target data group of the first comparison data group includes: The amplification factor, the minimum value in the first data group to be compared, and each sub-data in the first data group to be compared are input into the data group index prediction model to obtain the data index to be updated in the first initial data group. The sub-data in the first initial data group corresponding to the index of the data to be updated is updated to obtain the first target data group of the first initial data group.

3. The method according to claim 1, characterized in that, The step of constructing the first initial data set based on the amplification factor includes: Determine the numerical difference between the maximum and minimum values ​​in the first set of data to be compared; The numerical difference and the amplification factor are input into the data set size evaluation model to obtain the first data set size. The first initial data group is constructed based on the size of the first data group.

4. The method according to claim 1, characterized in that, Determining the second target data group corresponding to the second data group to be compared includes: Based on the amplification factor, a second initial data set is constructed; each sub-data in the second initial data set is used to characterize the preset initial frequency for each sub-data in the second comparison data set; Based on the amplification factor and the second set of data to be compared, each sub-data in the second initial data set is updated to obtain the second target data set corresponding to the second set of data to be compared.

5. The method according to claim 1, characterized in that, The step of obtaining the first probability distribution information of the first comparison data group based on the first target data group, and obtaining the second probability distribution information of the second comparison data group based on the second target data group, includes: Obtain the sum of all sub-data in the first target data group, and obtain the sum of all sub-data in the second target data group; Input each sub-data in the first target data group and the sum of all sub-data in the first target data group into the probability distribution detection model to obtain the first probability distribution information of the first data group to be compared. The probability distribution detection model is input into the probability distribution detection model by taking each sub-data in the second target data group and the sum of all sub-data in the second target data group, and then obtaining the second probability distribution information of the second comparison data group.

6. The method according to claim 1, characterized in that, The steps of acquiring the first comparison data group in the target monitoring data group and acquiring the second comparison data group in the sample data group corresponding to the target monitoring data group include: A first candidate data group is obtained from the target monitoring data group, and a second candidate data group corresponding to the first candidate data group is obtained from the sample data group; wherein the data in the first candidate data group and the second candidate data group have the same meaning; Determine the candidate numerical difference between the maximum and minimum values ​​in the first candidate data group; If the difference between the candidate values ​​does not exceed a preset difference threshold, then the first candidate data group is determined as the first data group to be compared, and the second candidate data group is determined as the second data group to be compared. If the difference between the candidate values ​​exceeds the preset difference threshold, then a first candidate data group is re-acquired from the target monitoring data group, and a second candidate data group corresponding to the first candidate data group is re-acquired from the sample data group.

7. The method according to claim 1, characterized in that, The step of obtaining the quality monitoring result of the target monitoring data group based on the distribution information difference between the first probability distribution information and the second probability distribution information includes: Obtain the statistical characteristic indicators of the target monitoring data group; Based on the differences in distribution information and the statistical characteristic indicators of the target monitoring data group, the quality monitoring results of the target monitoring data group are obtained.

8. The method according to claim 7, characterized in that, The step of obtaining the quality monitoring result of the target monitoring data group based on the distribution information differences and the statistical characteristic indicators of the target monitoring data group includes: If the statistical characteristic indicators of the target monitoring data group exceed the preset indicator threshold, and / or the difference in the distribution information exceeds the preset distribution threshold, then the quality monitoring result of the target monitoring data group is confirmed to be abnormal, and the processing tasks associated with the target monitoring data group are blocked. If the statistical characteristic indicators of the target monitoring data group do not exceed the preset indicator threshold, and the difference in distribution information does not exceed the preset difference threshold, then the quality monitoring result of the target monitoring data group is confirmed to be normal, and a prompt message for the quality monitoring result is generated.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Out-of-distribution data detection method and device, computer equipment and storage medium

    CN117113182A