Multidimensional data quality detection method based on data center
By utilizing time-series prediction and classification models within a data platform environment, and combining the correlations between data, multidimensional data quality is evaluated. This addresses the issue of insufficient accuracy in data quality detection in existing technologies, enabling more precise data quality assessment and decision support.
Patent Information
- Application Number
- CN202411801934.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-09
AI Technical Summary
Existing data quality inspection methods cannot effectively utilize the relevant information between multidimensional data in a data platform environment, resulting in poor accuracy of data quality inspection results and failing to meet the needs of high-quality data management and utilization.
A multi-dimensional data quality detection method based on a data platform is adopted. By combining the correlation of initial time series data with time series prediction and classification models, the degree of prediction deviation and correlation probability values are obtained. The data quality is comprehensively evaluated, and stable data is selected as a reference and data with obvious deviations are selected as targets for multi-dimensional data quality assessment.
It improves the accuracy of data quality detection by considering the correlation between data and their own quality status, providing more accurate data quality scores, reducing the risk of erroneous decisions, and improving enterprise operational efficiency and competitiveness.
Smart Images

Figure CN119645984B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a multi-dimensional data quality detection method based on a data platform. Background Technology
[0002] In today's digital age, data has become a key element in enterprise decision-making and business operations. With the explosive growth of data volume and the diversification of data sources, data quality issues are becoming increasingly prominent. Users often build data platforms to integrate and manage data, but in a data platform environment, the complexity and scale of data pose significant challenges to the management and utilization of high-quality data.
[0003] Traditional data quality inspection methods typically utilize statistical principles to analyze data, identifying outliers or data that deviates from normal distributions, or comparing the data to be inspected with sample or historical data to detect data changes inconsistent with anomalous patterns. However, these methods focus on the data itself across individual dimensions, failing to leverage the interrelationships between multidimensional data points. This makes it difficult to comprehensively assess data quality, resulting in less accurate data quality inspection results that cannot meet the needs of data platforms for managing and utilizing high-quality data.
[0004] Therefore, in the field of data processing technology, how to improve the accuracy of data quality detection results has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a multi-dimensional data quality detection method based on a data platform to solve the problem of low accuracy of existing data quality detection methods.
[0006] This invention provides a multi-dimensional data quality detection method based on a data platform. This method is applied to a data platform and includes:
[0007] S1: Obtain the trained time series prediction model, the trained classification model, N sets of initial time series data, and the correlation between the initial time series data. The correlation includes both correlated and uncorrelated relationships, and N is an integer greater than 1.
[0008] S2, input each set of initial time series data into the trained time series prediction model to obtain the predicted time series data corresponding to each set of initial time series data.
[0009] S3. Based on each set of initial time series data and the corresponding predicted time series data, obtain the deviation threshold for N sets of initial time series data and the prediction deviation for each set of initial time series data.
[0010] S4. Based on the degree of prediction deviation and the deviation threshold, classify the N sets of initial time series data into target data and initial reference data, and obtain the first data quality score of the corresponding initial reference data according to the degree of prediction deviation of each initial reference data.
[0011] S5. For any target data, based on the correlation between the initial time series data, obtain the target reference data corresponding to the current target data from all the initial reference data, and obtain the data group to be classified consisting of the current target data and each target reference data.
[0012] S6. Input each group of data to be classified corresponding to the current target data into the trained classification model to obtain the correlation probability value between the current target data and the corresponding target reference data.
[0013] S7. Based on the correlation probability value between the current target data and each corresponding target reference data, the prediction deviation degree of each target reference data corresponding to the current target data, and the prediction deviation degree of the current target data, the second data quality score corresponding to the current target data is obtained.
[0014] The beneficial effects of this invention compared to existing technologies are as follows: By learning the patterns and regularities of time series data during the training process of the time series prediction model, a connection is established between actual data and predicted data. This facilitates the examination of data quality from a dynamic perspective of time series. Based on each set of initial time series data and the corresponding predicted time series data, a deviation threshold and a prediction deviation degree corresponding to each set of initial time series data are obtained. The prediction deviation degree is used to accurately determine whether the initial time series data conforms to the expected time series change pattern and the magnitude of the deviation. Based on the comparison between the prediction deviation degree and the deviation threshold, relatively stable data with small deviations from expectations are selected as initial reference data, while data with significant deviations from expectations are selected as target data. The initial reference data is used as a reference benchmark for subsequent evaluation of target data quality. By obtaining relevant probability values, the strength of the correlation between target data and target reference data is quantified. Furthermore, considering the correlation between target data and each target reference data, as well as the impact of the quality of the target reference data itself on the target data quality evaluation, a second data quality score corresponding to the current target data is calculated according to the set correlation rules, thereby improving the accuracy of data quality analysis. Attached Figure Description
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 This is a schematic diagram of an application environment for a multi-dimensional data quality detection method based on a data platform provided in Embodiment 1 of the present invention;
[0017] Figure 2 This is a flowchart illustrating a multi-dimensional data quality detection method based on a data platform, provided in Embodiment 1 of the present invention. Detailed Implementation
[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0019] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0020] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0021] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0022] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0023] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0024] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0025] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0026] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0027] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0028] The multi-dimensional data quality detection method based on a data platform provided in Embodiment 1 of this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0029] Among them, the multi-dimensional data quality detection method based on the data middle platform can be applied to the data middle platform, which is a data processing and service platform located between the business front end and the data back end. It can extract, clean, transform, and store data generated by various business systems within the enterprise (such as sales system, financial system, customer relationship management system, etc.), and then build a unified data asset system based on the processed data. This provides standardized and reusable data services for different departments and business scenarios within the enterprise, making it easier for the enterprise to fully and accurately understand the overall business picture and achieve data interconnection.
[0030] Given the diverse types and wide range of sources of enterprise data, the primary challenge facing data platforms is how to conduct reasonable and accurate data quality testing. This testing results will provide reliable input for users' decision support systems and data analysis applications, reducing the risk of erroneous decisions caused by data quality issues and ultimately improving users' overall operational efficiency and competitiveness.
[0031] See also Figure 2 This is a flowchart illustrating a multi-dimensional data quality detection method based on a data platform, as provided in Embodiment 1 of the present invention. The aforementioned multi-dimensional data quality detection method based on a data platform can be applied to... Figure 1 For clients within a data platform, this multi-dimensional data quality inspection method may include the following steps:
[0032] S1: Obtain the trained time series prediction model, the trained classification model, N sets of initial time series data, and the correlation between the initial time series data. The correlation includes both correlated and uncorrelated relationships, and N is an integer greater than 1.
[0033] Initial time-series data, representing a series of data records generated over time within a specific scenario, is the object to be quality-tested. Analyzing and processing this initial time-series data allows for a comprehensive understanding of its quality over time, identifying potential problems such as abnormal data fluctuations and missing data. This provides reliable input for user decision support systems and data analysis applications, reducing the risk of erroneous decisions due to data quality issues and ultimately improving overall operational efficiency and competitiveness.
[0034] In practical applications, many data points are not isolated but rather interconnected. For example, product sales may be related to marketing activities, seasonal changes, and other factors. Therefore, by clearly defining the correlations between initial time-series data, this embodiment can more accurately assess data quality in subsequent steps by combining these correlations, avoiding misjudgments of data quality due to viewing data in isolation.
[0035] The correlation between the initial time series data can be determined by the implementer through statistical analysis methods, such as calculating the correlation coefficient between the initial time series data (e.g., Pearson correlation coefficient, Spearman correlation coefficient, etc.).
[0036] Time series prediction models are tools used for predictive analysis of time series data. In the multi-dimensional data quality detection task of a data platform, the trained time series prediction model leverages its ability to learn patterns in time series data to predict the possible values of initial time series data at a future point in time or within a time period. The data quality is then assessed by comparing the predicted and actual values. In this embodiment, the trained time series prediction model can be a Long-Short Term Memory (LSTM) artificial neural network, a Recurrent Neural Network (RNN), or other models used for predictive analysis of time series data. Those skilled in the art will recognize that any existing time series prediction model falls within the scope of this invention, and will not be elaborated upon further here.
[0037] The classification model is mainly used to classify and judge data. In this embodiment, the trained classification model is used to classify data based on factors such as the relationship between data, and outputs the confidence level of the classification result as the probability value of the correlation between data, so as to further evaluate the data quality.
[0038] The steps described above for obtaining the trained time-series prediction model, the trained classification model, N sets of initial time-series data, and the correlation between the initial time-series data provide a data foundation and model support for the subsequent calculation of the quality score of the initial time-series data.
[0039] S2, input each set of initial time series data into the trained time series prediction model to obtain the predicted time series data corresponding to each set of initial time series data.
[0040] In this embodiment, a pre-trained time series prediction model is used to perform predictive analysis on each set of initial time series data, thereby obtaining the predicted time series data corresponding to each set of initial time series data. By comparing the actually acquired initial time series data with the predicted time series data obtained by the model, a basis can be provided for subsequent operations such as evaluating the degree of data quality deviation, so as to accurately determine whether the initial time series data conforms to the expected time series change pattern, and thus discover potential data quality problems.
[0041] Optionally, the initial time series data includes the actual data corresponding to the M1th preset time point to the M2th preset time point, and S2 includes the following steps:
[0042] S21, for any initial time series data, input the current initial time series data into the trained time series prediction model to obtain the predicted time series data corresponding to the current initial time series data, wherein the predicted time series data includes the predicted data corresponding to the M1+1th preset time point to the M2+1th preset time point.
[0043] S22, iterate through all the initial time series data to obtain the predicted time series data corresponding to each set of initial time series data.
[0044] The initial time series data covers the actual data from the M1th preset time point to the M2th preset time point. When each set of initial time series data is input into the time series prediction model, the time series prediction model will predict the data for the next time point based on the rules and patterns of the time series data learned during the training process, and thus output the predicted time series data corresponding to each set of initial time series data, that is, the predicted data from the M1+1th preset time point to the M2+1th preset time point, providing comparative reference data for subsequent analysis of the quality of the initial time series data.
[0045] The above steps, which input each set of initial time series data into the trained time series prediction model to obtain the predicted time series data corresponding to each set of initial time series data, establish a connection between actual data and predicted data based on the rules and patterns of time series data learned by the time series prediction model during training. This facilitates the examination of data quality from the dynamic perspective of time series and provides a foundation for data quality detection.
[0046] S3. Based on each set of initial time series data and the corresponding predicted time series data, obtain the deviation threshold for N sets of initial time series data and the prediction deviation for each set of initial time series data.
[0047] Specifically, by comparing and analyzing each set of initial time series data and its corresponding predicted time series data, the deviation threshold corresponding to N sets of initial time series data and the prediction deviation corresponding to each set of initial time series data are quantified. The prediction deviation is used to accurately determine whether the initial time series data conforms to the expected time series change pattern and the degree of deviation, thus providing a data foundation for subsequent operations such as accurately judging the data quality status, classifying the data, and finally giving a data quality score.
[0048] Optionally, S3 includes the following steps:
[0049] S31. Based on the actual data and predicted data corresponding to the Kth preset time point of each set of initial time series data, obtain the Kth prediction deviation value corresponding to each set of initial time series data, where K = M1+1, M1+2, ..., M2.
[0050] S32, iterate through K = M1+1, M1+2, ..., M2, and obtain the M2-M1 prediction bias values corresponding to each set of initial time series data.
[0051] S33, the sum of the M2-M1 prediction deviation values corresponding to each group of initial time series data is determined as the total prediction deviation value corresponding to each group of initial time series data.
[0052] S34, the average value of the actual data corresponding to each group of initial time series data from the M1 preset time point to the M2 preset time point is determined as the average data value corresponding to each group of initial time series data.
[0053] S35, the difference between the total prediction deviation and the corresponding average data value for each set of initial time series data is determined as the prediction deviation level for each set of initial time series data.
[0054] Specifically, the deviation of each set of initial time series data at each specific time point is quantitatively analyzed, and the deviation of each set of data at each time point is summarized by summation to obtain the total predicted deviation value, which is a quantitative indicator reflecting the degree of deviation of the data set over the entire time period. This yields the degree of predicted deviation, which characterizes the deviation of each set of initial time series data from its own average level, providing a key analytical dimension for subsequent data classification and quality scoring operations.
[0055] Optionally, S3 also includes the following steps:
[0056] S36. Based on the prediction deviation degree corresponding to the N sets of initial time series data, obtain the probability distribution function of the prediction deviation degree corresponding to the N sets of initial time series data.
[0057] S37, the third quartile corresponding to the probability distribution function of the prediction deviation is determined as the deviation threshold corresponding to the N initial time series data.
[0058] Specifically, based on the prediction deviation levels corresponding to N sets of initial time series data, statistical methods are used to construct and obtain the probability distribution function of the prediction deviation levels corresponding to N sets of initial time series data. This function describes the probability of different prediction deviation levels occurring in the N sets of initial time series data, which helps to gain a deeper understanding of the overall distribution characteristics of the data deviation levels and provides a more scientific basis for subsequently determining the deviation level threshold.
[0059] The quartile is a statistical measure that divides the data distribution into four equal parts. By determining the third quartile as the deviation threshold, N initial time series data can be divided into data with small deviation and data with large deviation based on the deviation threshold, which facilitates further data quality assessment and processing.
[0060] The above steps, based on each set of initial time series data and the corresponding predicted time series data, obtain the deviation threshold for N sets of initial time series data and the predicted deviation for each set of initial time series data. By using the predicted deviation, the initial time series data can be accurately judged whether it conforms to the expected time series change pattern and the degree of deviation. This provides a data foundation for subsequent operations such as accurately judging the data quality, classifying the data, and finally giving a data quality score.
[0061] S4. Based on the degree of prediction deviation and the deviation threshold, classify the N sets of initial time series data into target data and initial reference data, and obtain the first data quality score of the corresponding initial reference data according to the degree of prediction deviation of each initial reference data.
[0062] Based on a comparison of the degree of prediction deviation and the deviation threshold, the N initial time series data sets are classified into target data and initial reference data. Target data consists of data that deviates significantly from expectations and requires further investigation into its quality. Initial reference data, on the other hand, is closer to expectations and can serve as a reference standard for evaluating the quality of the target data. The first data quality score is determined based on the degree of prediction deviation of the initial reference data, thus providing more detailed differentiation and quantitative indicators for a comprehensive assessment of data quality.
[0063] Optionally, S4 includes the following steps:
[0064] S41, for any initial time series data, if the prediction deviation corresponding to the current initial time series data is less than the deviation threshold, then the current initial time series data is classified as initial reference data.
[0065] S42, if the prediction deviation corresponding to the current initial time series data is greater than or equal to the deviation threshold, then the current initial time series data is classified as target data.
[0066] S43. Based on the prediction deviation degree corresponding to each initial reference data, obtain the first data quality score of the corresponding initial reference data, wherein the first data quality score is negatively correlated with the corresponding prediction deviation degree.
[0067] In this process, relatively stable data with small deviations from expectations are selected as initial reference data. The smaller the prediction deviation of the initial reference data, the higher the corresponding first data quality score. Conversely, the larger the prediction deviation of the initial reference data, the lower the corresponding first data quality score.
[0068] Data that deviates significantly from expectations is selected as target data, and the initial reference data is used as a benchmark for subsequent evaluation of the target data quality. By combining the correlation between data, the data quality of the target data can be analyzed from more dimensions, thereby improving the accuracy of data quality analysis.
[0069] The above steps, based on the degree of prediction deviation and the deviation threshold, classify N sets of initial time series data into target data and initial reference data, and obtain the first data quality score of the corresponding initial reference data according to the degree of prediction deviation. The steps then select relatively stable data with small deviations from the expectation as initial reference data, select data with significant deviations from the expectation as target data, and use the initial reference data as a reference benchmark for subsequent evaluation of the target data quality. This allows for a more multi-dimensional analysis of the target data quality by combining the correlation between data, thereby improving the accuracy of data quality analysis.
[0070] S5. For any target data, based on the correlation between the initial time series data, obtain the target reference data corresponding to the current target data from all the initial reference data, and obtain the data group to be classified consisting of the current target data and each target reference data.
[0071] Specifically, for target data that may have significant quality issues or deviate greatly from expectations, target reference data related to it are selected from the relatively high-quality initial reference data based on the pre-determined correlation between the initial time series data. The target data and each target reference data are then combined into data groups to be classified, providing a more targeted and relevant analytical basis for further evaluation of the quality of the target data.
[0072] Optionally, S5 includes the following steps:
[0073] S51, for any target data, obtain the correlation between the current target data and each initial reference data based on the correlation between the initial time series data.
[0074] S52, for any initial reference data, if the correlation between the current target data and the current initial reference data is positive, then the current initial reference data is determined as the target reference data corresponding to the current target data.
[0075] S53, combine the current target data and the current target reference data to obtain a data group to be classified consisting of the current target data and the current target reference data.
[0076] S54, iterate through all the initial reference data to obtain all the data groups to be classified corresponding to the current target data.
[0077] S55, iterate through all the target data to obtain all the data groups to be classified for each target data.
[0078] The above steps, for any target data, involve obtaining the target reference data corresponding to the current target data from all the initial reference data based on the correlation between the initial time series data, and obtaining the data group to be classified composed of the current target data and each target reference data. Combining the target data with each target reference data into data groups to be classified provides a more targeted and relevant analytical basis for further evaluation of the quality of the target data. This approach can fully utilize the inherent correlation between data, avoid viewing the target data in isolation, and thus improve the accuracy of data quality analysis.
[0079] S6. Input each group of data to be classified corresponding to the current target data into the trained classification model to obtain the correlation probability value between the current target data and the corresponding target reference data.
[0080] Among them, the correlation probability value between the target data and the corresponding target reference data is obtained based on the trained classification model. This is used to quantitatively represent the strength of the correlation between the target data and the target reference data, providing an important basis for more accurate evaluation of the quality of the target data.
[0081] The correlation probability value is a value between 0 and 1, where 0 indicates that there is almost no correlation between the two and 1 indicates that there is a strong correlation between them.
[0082] The steps described above involve inputting each group of data to be classified corresponding to the current target data into the trained classification model to obtain the correlation probability value between the current target data and the corresponding target reference data. Obtaining the correlation probability value quantifies the strength of the correlation between the target data and the target reference data, providing an important basis for more accurately evaluating the quality of the target data in the future.
[0083] S7. Based on the correlation probability value between the current target data and each corresponding target reference data, the prediction deviation degree of each target reference data corresponding to the current target data, and the prediction deviation degree of the current target data, the second data quality score corresponding to the current target data is obtained.
[0084] Optionally, S7 includes the following steps:
[0085] S71. Based on any target reference data corresponding to the current target data, based on the correlation probability value between the current reference data and the current target reference data, and the prediction deviation degree corresponding to the current target reference data, the first score influence parameter corresponding to the current target data is obtained. The first score influence parameter is positively correlated with the corresponding correlation probability value and negatively correlated with the corresponding prediction deviation degree.
[0086] S72, iterate through all target reference data corresponding to the current target data to obtain all first score influence parameters corresponding to the current target data, and determine the average value of all first score influence parameters corresponding to the current target data as the second score influence parameter corresponding to the current target data.
[0087] S73. Based on the second score influence parameter and prediction deviation degree corresponding to the current target data, obtain the second data quality score corresponding to the current target data. The second data quality score is negatively correlated with the corresponding second score influence parameter and prediction deviation degree.
[0088] In this context, for target data and target reference data that are correlated, if the correlation probability value given by the classification model indicates a weak correlation between the two, it can characterize that the target data has a large data error and poor data quality. Therefore, the first scoring influence parameter is positively correlated with the correlation probability value, which is also negatively correlated with the second data quality score. Furthermore, the smaller the prediction deviation of the target reference data, the higher its reliability. Thus, the higher the reliability of the data quality characterized by the correlation probability value, the higher the reliability of the target data. Therefore, the first scoring influence parameter is negatively correlated with the prediction deviation of the corresponding target reference data, which is also positively correlated with the second data quality score.
[0089] The above steps, which derive the second data quality score for the current target data based on the correlation probability value between the current target data and each corresponding target reference data, the prediction deviation degree of each target reference data corresponding to the current target data, and the prediction deviation degree of the current target data, comprehensively consider the correlation between the target data and each target reference data, as well as the impact of the quality status of the target reference data itself on the target data quality assessment. The second data quality score for the current target data is calculated according to the set correlation rules, thereby improving the accuracy of data quality analysis.
[0090] This invention establishes a connection between actual and predicted data based on the patterns and regularities of time series data learned during the training of a time series prediction model. This facilitates the examination of data quality from a dynamic perspective of time series. Based on each set of initial time series data and its corresponding predicted time series data, it obtains a threshold for the degree of deviation for N sets of initial time series data and the degree of prediction deviation for each set. This allows for precise judgment of whether the initial time series data conforms to the expected time series change pattern and the magnitude of deviation. By comparing the degree of prediction deviation with the deviation threshold, relatively stable data with small deviations from expectations are selected as initial reference data, while data with significant deviations from expectations are selected as target data. The initial reference data serves as a benchmark for subsequent evaluation of target data quality. Relevant probability values are obtained to quantify the strength of the correlation between target data and target reference data. Furthermore, considering the correlation between target data and each target reference data, as well as the impact of the quality of the target reference data itself on the target data quality assessment, a second data quality score is calculated according to the established correlation rules, thus improving the accuracy of data quality analysis.
[0091] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0092] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0093] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A multi-dimensional data quality detection method based on a data platform, wherein the multi-dimensional data quality detection method based on a data platform is applied to a data platform, characterized in that, The multi-dimensional data quality detection method based on a data platform includes: S1, obtain the trained time series prediction model, the trained classification model, N sets of initial time series data, and the correlation between the initial time series data, wherein the correlation includes correlated and uncorrelated, and N is an integer greater than 1; S2, input each set of initial time series data into the trained time series prediction model to obtain the predicted time series data corresponding to each set of initial time series data; S3. Based on each set of initial time series data and the corresponding predicted time series data, obtain the deviation threshold for N sets of initial time series data and the prediction deviation for each set of initial time series data. S4. Based on the degree of prediction deviation and the deviation degree threshold, the N sets of initial time series data are classified into target data and initial reference data, and a first data quality score for the corresponding initial reference data is obtained based on the degree of prediction deviation for each initial reference data. S5, for any target data, based on the correlation between the initial time series data, obtain the target reference data corresponding to the current target data from all the initial reference data, and obtain the data group to be classified consisting of the current target data and each target reference data; S6, input each data group to be classified corresponding to the current target data into the trained classification model to obtain the correlation probability value between the current target data and the corresponding target reference data; S7. Based on the correlation probability value between the current target data and each corresponding target reference data, the prediction deviation degree of each target reference data corresponding to the current target data, and the prediction deviation degree of the current target data, the second data quality score corresponding to the current target data is obtained.
2. The multi-dimensional data quality detection method based on a data platform according to claim 1, characterized in that, The initial time series data includes the actual data corresponding to the M1st preset time point to the M2th preset time point. S2 includes the following steps: S21, for any initial time series data, input the current initial time series data into the trained time series prediction model to obtain the predicted time series data corresponding to the current initial time series data, wherein the predicted time series data includes the predicted data corresponding to the M1+1th preset time point to the M2+1th preset time point. S22, iterate through all the initial time series data to obtain the predicted time series data corresponding to each set of initial time series data.
3. The multi-dimensional data quality detection method based on a data platform according to claim 2, characterized in that, S3 includes the following steps: S31, based on the actual data and predicted data corresponding to the Kth preset time point of each set of initial time series data, obtain the Kth prediction deviation value corresponding to each set of initial time series data, where K = M1+1, M1+2, ..., M2; S32, iterate through K = M1+1, M1+2, ..., M2, and obtain the M2-M1 prediction bias values corresponding to each set of initial time series data; S33, the sum of the M2-M1 prediction deviation values corresponding to each group of initial time series data is determined as the total prediction deviation value corresponding to each group of initial time series data; S34, the average value of the actual data corresponding to each group of initial time series data from the M1th preset time point to the M2th preset time point is determined as the average data value corresponding to each group of initial time series data; S35, the difference between the total prediction deviation and the corresponding average data value for each set of initial time series data is determined as the prediction deviation level for each set of initial time series data.
4. The multi-dimensional data quality detection method based on a data platform according to claim 3, characterized in that, S3 also includes the following steps: S36. Based on the prediction deviation degree corresponding to the N sets of initial time series data, obtain the probability distribution function of the prediction deviation degree corresponding to the N sets of initial time series data. S37, the third quartile corresponding to the probability distribution function of the prediction deviation degree is determined as the deviation degree threshold corresponding to the N sets of initial time series data.
5. The multi-dimensional data quality detection method based on a data platform according to claim 1, characterized in that, S4 includes the following steps: S41, for any initial time series data, if the prediction deviation corresponding to the current initial time series data is less than the deviation threshold, then the current initial time series data is classified as initial reference data; S42, if the prediction deviation corresponding to the current initial time series data is greater than or equal to the deviation threshold, then the current initial time series data is classified as target data; S43, based on the prediction deviation degree corresponding to each initial reference data, obtain the first data quality score of the corresponding initial reference data, wherein the first data quality score is negatively correlated with the corresponding prediction deviation degree.
6. The multi-dimensional data quality detection method based on a data platform according to claim 1, characterized in that, S5 includes the following steps: S51, for any target data, obtain the correlation between the current target data and each initial reference data based on the correlation between the initial time series data; S52, for any initial reference data, if the correlation between the current target data and the current initial reference data is positive, then the current initial reference data is determined as the target reference data corresponding to the current target data; S53, combine the current target data and the current target reference data to obtain a data group to be classified consisting of the current target data and the current target reference data; S54, traverse all the initial reference data to obtain all the data groups to be classified corresponding to the current target data; S55, iterate through all the target data to obtain all the data groups to be classified for each target data.
7. The multi-dimensional data quality detection method based on a data platform according to claim 6, characterized in that, S7 includes the following steps: S71, based on any target reference data corresponding to the current target data, based on the correlation probability value between the current reference data and the current target reference data, and the prediction deviation degree corresponding to the current target reference data, a first scoring influence parameter corresponding to the current target data is obtained, wherein the first scoring influence parameter is positively correlated with the corresponding correlation probability value, and the first scoring influence parameter is negatively correlated with the corresponding prediction deviation degree; S72, traverse all target reference data corresponding to the current target data, obtain all first score influence parameters corresponding to the current target data, and determine the average value of all first score influence parameters corresponding to the current target data as the second score influence parameter corresponding to the current target data; S73. Based on the second score influence parameter and the degree of prediction deviation corresponding to the current target data, obtain the second data quality score corresponding to the current target data, wherein the second data quality score is negatively correlated with the corresponding second score influence parameter and the degree of prediction deviation.
Citation Information
Patent Citations
Data quality evaluation method and device for time series data and electronic equipment
CN116204563A
New energy station monitoring data quality evaluation method and system based on multi-source data
CN119066541A