A data quality management method and system applied to a big data system
By constructing a data quality system, based on the concentrated areas of advantageous data and past optimization events, the data in the big data system is accurately divided and managed, which solves the problem of inaccurate data quality levels and improves the accuracy of data quality management.
Patent Information
- Application Number
- CN202511168473.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-08-20
AI Technical Summary
The lack of optimized data segmentation in existing big data systems results in poor accuracy of data quality levels, impacting data quality management.
By identifying the concentrated areas of advantageous data and past data optimization events, a data quality system is constructed. The data to be processed is divided, and the quality level is determined according to the location, priority, and quality score of the sub-data. Data management events are then matched in the data acquisition path.
This improved the accuracy of data quality levels and the precision of event management, enabling comprehensive data quality management.
Smart Images

Figure CN121070905B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of data quality management methods, and more particularly to a data quality management method and system applied to big data systems. Background Technology
[0002] With the development of technology, big data systems are highly complex technical systems mainly used to process, store, and analyze massive, heterogeneous, and multi-type data. They can also be applied to online platforms such as e-commerce. Big data systems help enterprises and organizations extract value from data and support decision-making and innovation through multiple stages such as data collection, storage, processing, analysis, and visualization. In existing technologies, big data systems contain multiple datasets and distribute them in corresponding data spaces. However, the data partitioning is not optimized, which affects the perfection of the data quality system and results in poor accuracy of the quality level of the data to be processed, thus affecting data quality management. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a data quality management method and system applicable to big data systems.
[0004] This invention provides a data quality management method for big data systems, comprising:
[0005] Multiple data spaces are identified based on the detection of the big data system, and multiple advantageous data are determined based on the data volume, spatial form and current state of the big data system.
[0006] The concentration area of the advantageous data is determined based on the distribution of multiple advantageous data, and the data quality system is determined based on the concentration area of the advantageous data and the past data optimization events of the big data system.
[0007] The data quality system divides the data to be processed into multiple sub-data. The quality level of each sub-data is determined based on its data location, priority, and quality score. The quality level of the data to be processed is then determined based on the quality level of each sub-data.
[0008] If the quality level of the data to be processed is lower than the preset quality level, the data collection path is determined based on the traceability of the data to be processed, and the corresponding data control events are matched to each data control node of the data collection path.
[0009] Based on the data control events of each data control node, determine the events to be optimized, and based on the data types and application scenarios corresponding to the events to be optimized and the data to be processed, determine the quality management events of the data to be processed.
[0010] This invention provides a data quality management system for big data systems, which applies the aforementioned data quality management method for big data systems. The data quality management system for big data systems includes:
[0011] The Advantage Data Module is used to identify multiple data spaces based on the detection of the big data system, and to determine multiple advantageous data based on the data volume, spatial form and current state of the big data system.
[0012] The data quality system module is used to determine the concentrated area of advantageous data based on the distribution location of multiple advantageous data, and to determine the data quality system based on the concentrated area of advantageous data and the past data optimization events of the big data system.
[0013] The quality level module is used to divide the data to be processed in this data quality system to output multiple sub-data. The quality level of each sub-data is determined according to its data position, priority and quality score. The quality level of the data to be processed is determined based on the quality level of each sub-data.
[0014] The data management event module is used to determine the data collection path based on the tracing of the data to be processed if the quality level of the data to be processed is lower than the preset quality level, and to match the corresponding data management events to each data management node of the data collection path.
[0015] The quality management event module is used to determine the events to be optimized based on the data control events of each data control node, and to determine the quality management events of the data to be processed based on the data type and application scenario of the events to be optimized and the data to be processed.
[0016] Compared with the prior art, the beneficial effects of the present invention are:
[0017] In this embodiment of the invention, a data quality system is determined based on the concentrated area of advantageous data and the previous data optimization events of the big data system. This data quality system divides the data to be processed into multiple sub-data. The quality level of each sub-data is determined according to its data location, priority, and quality score. The quality level of the data to be processed is determined based on the quality level of each sub-data. This data quality system incorporates the overall consideration of the quality levels of each sub-data, thereby improving the accuracy of the quality level of the data to be processed.
[0018] Therefore, if the quality level of the data to be processed is lower than the preset quality level, the data collection path is determined based on the traceability of the data to be processed, and corresponding data control events are matched to each data control node of the data collection path; the events to be optimized are determined based on the data control events of each data control node, and the quality management events of the data to be processed are determined based on the data types and application scenarios corresponding to the events to be optimized and the data to be processed. The introduction of data control events realizes the overall consideration of the data types and functions corresponding to the events to be optimized and the data to be processed, improves the accuracy of the quality management events of the data to be processed, and realizes the quality management of data. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating a data quality management method applied to a big data system according to an embodiment of the present invention;
[0020] Figure 2 This is a flowchart illustrating step S11 of the data quality management method applied to a big data system in an embodiment of the present invention;
[0021] Figure 3 This is a flowchart illustrating step S12 in the data quality management method applied to a big data system according to an embodiment of the present invention.
[0022] Figure 4 This is a flowchart illustrating step S13 in the data quality management method applied to a big data system according to an embodiment of the present invention.
[0023] Figure 5 This is a flowchart illustrating step S14 of the data quality management method applied to a big data system in an embodiment of the present invention.
[0024] Figure 6 This is a flowchart illustrating step S15 of the data quality management method applied to a big data system in an embodiment of the present invention.
[0025] Figure 7 This is a schematic diagram of the structural composition of a data quality management system applied to a big data system in an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0027] Please see Figures 1 to 7 A data quality management method applied to big data systems, and the scenarios in which it is applied; the data quality management method applied to big data systems includes:
[0028] Step S11: Based on the detection of the big data system, determine multiple data spaces, and determine multiple advantageous data based on the data volume, spatial form and current state of the big data system of the multiple data spaces;
[0029] Step S12: Determine the concentrated area of the advantageous data based on the distribution location of multiple advantageous data, and determine the data quality system based on the concentrated area of the advantageous data and the previous data optimization events of the big data system;
[0030] Step S13: The data quality system divides the data to be processed into multiple sub-data. The quality level of each sub-data is determined based on its data position, priority, and quality score. The quality level of the data to be processed is determined based on the quality level of each sub-data.
[0031] Step S14: If the quality level of the data to be processed is lower than the preset quality level, the data collection path is determined based on the tracing of the data to be processed, and the corresponding data control events are matched to each data control node of the data collection path.
[0032] Step S15: Determine the events to be optimized based on the data management events of each data management node, and determine the quality management events of the data to be processed based on the data types and application scenarios corresponding to the events to be optimized and the data to be processed;
[0033] refer to Figure 2 In step S11, the specific steps are as follows:
[0034] S111: Collect data from the big data system, determine the corresponding detection method based on the system parameters and data volume of the big data system, and trigger the detection of the big data system along the detection method to determine multiple data spaces. At this time, each data space contains data of the corresponding type.
[0035] S112: In each data space, the data volume and spatial form of the data space are determined based on the identification of the data space. At the same time, multiple working parameters of the big data system are collected, and the current state of the big data system is determined according to the big data system.
[0036] S113: Determine a first advantageous data set based on the data volume and spatial form of multiple data spaces; determine a second advantageous data set based on the data volume of multiple data spaces and the current state of the big data system; and determine multiple advantageous data sets based on the first and second advantageous data sets.
[0037] In the embodiments of this application, a big data system is collected, and system parameters (such as cluster size, number of nodes, and storage capacity) are obtained through system APIs or monitoring tools (such as Prometheus and KafkaManager); data metadata (such as data table structure and data source) is obtained using data catalog tools (such as Apache Atlas and HiveMetastore); and data flow and data processing logs are obtained by combining log analysis tools (such as ELKStack), and system parameters (such as cluster size and storage capacity), data metadata (such as data table structure), and data flow logs are output.
[0038] Based on parameters such as cluster size, number of nodes, and storage capacity, assess the system's processing capacity and resource limitations; select the detection method based on the total data volume (e.g., PB level, TB level) and data distribution (e.g., uniform distribution, skewed distribution); detection method selection: full detection: suitable for scenarios with small data volume and high value; scan the entire dataset to identify data spaces (e.g., divided by data tables or data themes); through the selected detection method, identify multiple data spaces in the big data system, each space containing specific types of data; each space contains corresponding types of data (e.g., transaction data space, user data space, log data space).
[0039] Furthermore, scan the entire dataset to identify data spaces (e.g., divided by data tables or data themes); data volume: obtain the size of the data space (e.g., storage capacity, number of records) through data cataloging tools (e.g., HiveMetastore, ApacheAtlas) or file system tools (e.g., HDFS commands); space form: analyze data structure (e.g., columnar storage / row storage), data distribution (e.g., partitioning, bucketing), and data integrity (e.g., null value ratio, duplicate value ratio), and output the data volume (e.g., TB, GB) and space form (e.g., structured, semi-structured, data quality indicators) for each data space.
[0040] Acquire real-time system status to assess data processing performance and stability; obtain metrics through monitoring systems (such as Prometheus and Grafana), including: resource utilization (CPU, memory, disk I / O); task execution status (such as the number of MapReduce jobs, Spark task latency); network bandwidth and latency, and output system operating parameters (such as CPU utilization, task latency). Based on these operating parameters, determine whether the system is under high load, low load, or an abnormal state; Status assessment: High load: CPU / memory utilization > 80%, task latency > 10 minutes; Low load: CPU / memory utilization < 50%, task latency < 1 minute; Abnormal state: Disk I / O errors, task failure rate > 5%.
[0041] Therefore, a first advantageous data set is determined based on the data volume and spatial form of multiple data spaces, and a second advantageous data set is determined based on the data volume of multiple data spaces and the current state of the big data system. Multiple advantageous data sets are determined based on the first and second advantageous data sets, which takes into account the overall consideration of the first and second advantageous data sets and ensures the accuracy of multiple advantageous data sets.
[0042] At this point, from all the data spaces, those that perform "excellently" in terms of "data volume" and "spatial form" (i.e., data structure, quality, and other intrinsic characteristics) are initially screened out. Here, "excellent" usually means that the data volume is large enough (meaning high business value or high analytical value) and the spatial form is good (clear data structure and high quality).
[0043] Simultaneously, set a threshold for data volume or use relative comparisons (e.g., select the top-ranked spaces by data volume); large data volume usually means more information and potential value, but also higher processing costs; a trade-off is needed; analyze the previously determined morphological characteristics; for example: degree of structuring: structured data (such as relational database tables) is generally easier to manage and analyze than unstructured data (such as raw logs), and has a better morphology; data quality indicators: low null value rate, low duplication rate, high integrity, high consistency, etc., indicate a better spatial morphology; organization method: good partitioning and bucketing strategies help query efficiency, and have a better morphology; combine the evaluation results of data volume and spatial morphology; a data space may have a large data volume but poor quality (poor morphology), or a small data volume but extremely high quality and perfect structure (good morphology); the priority of "advantages" needs to be defined according to business needs; for example, prioritize spaces with "large data volume and good morphology", or spaces with "excellent morphology but moderate data volume"; include data spaces that meet the above "advantage" criteria into the first-advantage data set.
[0044] From all data spaces, select those that perform "excellently" in terms of "data volume" and "current state of the big data system" (such as load and performance). Here, "excellent" means that the data volume is moderate and the system currently has sufficient resources to process the data efficiently, or that the data volume is large but the system is in good condition to support it. At the same time, while still considering the data volume, we pay more attention to the matching degree with the system's processing capacity. For example, if the system load is very high, we tend to choose a space with a smaller data volume for fast processing. High load state: system resources are tight (high CPU, high memory, high latency).
[0045] At this point, it is preferable to choose a space with a smaller data volume to avoid further burdening the system; or, choose spaces with low computational complexity for processing; Low load state: the system has sufficient resources; it can handle spaces with a larger data volume or perform more complex data processing tasks; Abnormal state: the system has a failure or performance bottleneck; it is necessary to prioritize processing monitoring data spaces related to the anomaly, or temporarily avoid spaces with large processing volumes that may exacerbate the problem; Including data spaces that meet the criteria of "good matching between data volume and system state" into the second advantageous data set is more like an advantage in terms of "processability" or "timeliness".
[0046] Based on the results of the first two steps, the final determination is made of which data spaces are truly "advantageous data." This means that these data spaces are either of high quality and value (first advantageous data set), or are easy to process and valuable under the current system conditions (second advantageous data set), or both. Simultaneously, the business logic determines how to combine these two sets. The union is taken: data is considered advantageous as long as it appears in either set, thus broadening the scope of advantageous data. The intersection is taken: only data spaces appearing in both sets are considered advantageous data, which is more stringent, ensuring that the data is both of high quality and processable under current conditions. Different weights are assigned to members of the first and second advantageous data sets, and only spaces with high overall scores are considered advantageous data. The selected set operations are then performed to obtain the final list of advantageous data, which will be prioritized for subsequent data quality analysis and management.
[0047] refer to Figure 3 In step S12, the specific steps are as follows:
[0048] S121: Collect multiple dominant data and mark the distribution location of multiple dominant data. Determine multiple sub-dominant regions based on the distribution location, data type and data volume of multiple dominant data. Determine the concentrated area of dominant data based on the cross combination of multiple sub-dominant regions.
[0049] S122: Collect data from the big data system, and determine past data optimization events of the big data system based on the traceability of the big data system, and determine multiple data optimization projects based on the analysis of past data optimization events;
[0050] S123: A data quality system is determined based on cross-training of multiple data optimization projects, concentrated areas of advantageous data, and distribution locations of multiple advantageous data. At this time, the data quality system is used to control the data quality of data in the big data system or external data to be processed.
[0051] In the embodiments of this application, after determining the advantageous data (step S11), the distribution of these data in the system is further analyzed to identify areas with high data density or strong correlation, providing spatial focus for subsequent quality management strategy formulation. Simultaneously, detailed information about these data is obtained from the advantageous data set obtained in step S11, including data source, storage location (e.g., HDFS path, database table name), data format, etc., and each advantageous data is labeled with a geographical or logical location tag. This could be a physical server cluster, a logical database / data warehouse partition, or even a specific subject domain (e.g., user domain, transaction domain). Based on the distribution location, data type (e.g., transaction data, user profile data), and data volume of the advantageous data, similarly distributed or closely related data are clustered into smaller regions. For example, all user-related advantageous data (user basic information, user behavior logs, user profiles) constitute a "user sub-advantageous region." The relationships between these sub-advantageous regions are analyzed, especially their intersections (i.e., areas where multiple sub-regions overlap or are closely adjacent). These intersection areas or core areas are the concentrated areas of advantageous data.
[0052] By reviewing history and understanding how big data systems have handled data quality issues in the past, we can summarize lessons learned and provide a reference for building a new data quality system. Simultaneously, we can collect information about past data quality improvement activities by accessing system logs, version control systems (such as Git), project management tools (such as Jira), knowledge bases, or directly interviewing relevant personnel. Through system tracing (such as querying logs and code commit records), we can identify specific data quality issues that occurred and their corresponding actions. For example, "discovering a large number of duplicate IDs in the user table," "addressing the out-of-order timestamp issue in the transaction log table," and "optimizing the ETL process for log data to reduce missing values" are all specific events. Related optimization events can be categorized and summarized to form more macro-level data optimization projects. For example, multiple events addressing the duplicate ID issue in the user table can be summarized into a single "User Table Deduplication Optimization Project." Projects typically have clearer goals and scope.
[0053] Furthermore, understanding how the big data system handled data quality issues in the past provides a reference for building a new data quality system. Simultaneously, accessing system logs, version control systems (such as Git), project management tools (such as Jira), knowledge bases, or directly interviewing relevant personnel gathers information about past data quality improvement activities. Through system tracing (such as querying logs and code commit records), identify specific data quality issues that occurred and their corresponding actions; for example, "discovering a large number of duplicate IDs in the user table," "addressing the out-of-order timestamp issue in the transaction log table," and "optimizing the ETL process for log data to reduce missing values" are all specific events. Related optimization events are categorized and summarized to form more macro-level data optimization projects; for example, multiple events addressing the duplicate ID issue in the user table can be summarized into a "User Table Deduplication Optimization Project." Projects typically have clearer goals and scope.
[0054] Therefore, a data quality system is determined based on the cross-training of multiple data optimization projects, concentrated areas of advantageous data, and the distribution locations of multiple advantageous data. At this time, the data quality system is used to control the data quality of data in the big data system or external data to be processed. This system takes into account the overall consideration of the cross-training of multiple data optimization projects, concentrated areas of advantageous data, and the distribution locations of multiple advantageous data, thus ensuring the accuracy of the data quality system.
[0055] At this point, by combining historical experience (data optimization projects), the important areas of current data (areas where advantageous data is concentrated), and the distribution of data (locations where advantageous data is distributed), a systematic and targeted data quality control strategy and method can be formulated, namely, a data quality system.
[0056] Meanwhile, "cross-training" here can be understood as a comprehensive analysis and mapping process; it involves correlation analysis between data optimization projects (knowing where problems were likely to occur in the past and how to solve them) and areas of concentrated advantageous data (knowing which areas of data are most important and active now) and the distribution location of advantageous data (knowing where the data is specifically located); for example, it involves correlation analysis between the "user data quality improvement project" and the "user core information concentration area" and the distribution location of the user table, analyzing whether the reasons for the project's failure / success are related to the data characteristics, processing flow, or system environment of that area; and analyzing whether the "transaction data time series accuracy assurance project" is applicable to the current "user-transaction concentration area," or whether that area has introduced new time series problem risks.
[0057] Based on the results of cross-training, general data quality rules, monitoring indicators, inspection methods, processing procedures, and responsibility allocation are extracted. This system can: set up monitoring in areas with concentrated advantageous data and key distribution locations to promptly identify quality issues; define quality measurement standards applicable to different data types and regions; formulate standardized data cleaning, repair, and rejection processes based on historical project experience; and establish a mechanism for continuous improvement to regularly review the effectiveness of the quality system.
[0058] refer to Figure 4 In step S13, the specific steps are as follows:
[0059] S131: Collect the data to be processed, determine the basic information set of the data to be processed based on the traceability of the data to be processed, and at the same time, collect the data quality system, determine the corresponding data partitioning mode based on the data quality system, the data volume of the data to be processed and the basic information set, and partition the data to be processed along the data partitioning mode, and output multiple sub-data.
[0060] S132: Determine the data location of multiple sub-data based on the location detection of multiple sub-data, and determine the priority of multiple sub-data according to the data type, data usage frequency and data application scenario of multiple sub-data. At the same time, determine the corresponding quality score based on the quality detection of multiple sub-data.
[0061] S133: Determine the first-level coefficient based on the data position and priority of multiple sub-data; determine the second-level coefficient based on the priority and quality score of multiple sub-data; determine the quality level of each sub-data based on the mapping relationship between the first-level coefficient, the second-level coefficient, and the quality level; determine the quality level of the data to be processed based on the quality level and data position of each sub-data.
[0062] In the embodiments of this application, data to be processed is collected, and the physical or logical location of the data is determined. This could be a file on a file system, a table in a database, a message in a message queue, a dataset returned by an API, etc. Depending on the data source, appropriate technical means are used for collection. For example, file transfer tools, database connectors (such as JDBC, ODBC), message queue consumers, API calls, etc. are used. After collection, a basic integrity check is performed to ensure that the data has been successfully acquired and there are no obvious transmission errors.
[0063] Gain a deep understanding of the background, structure, and basic characteristics of the data to be processed, as this information is crucial for subsequent quality assessment and processing; trace the data's generation process, source system, collection time, transformation steps, etc., through metadata management systems, data catalogs, data lineage tools, or manual records; extract key information from the data itself or its metadata to construct a basic information set, which typically includes: data identifier, data source, collection time, data format, data structure, data volume, intended use, and known issues (if any).
[0064] The data quality system is collected, and its specific contents are read and loaded, including: Data quality rules: defining the constraints that data must meet (such as non-empty, uniqueness, format validation, range validation, referential integrity, etc.); Evaluation indicators: defining how to quantify data quality (such as integrity score, accuracy score, consistency score); Priority definitions: containing priority rules for different data types or application scenarios, and defining the steps to be taken after a quality problem is discovered.
[0065] Decide how to divide the collected, raw data into smaller, more manageable, and evaluable units (subdata). The choice of partitioning mode will affect the efficiency and accuracy of subsequent quality assessment. The following factors should be considered when deciding on the partitioning mode: Data quality system: Are the rules for the entire dataset or specific fields / records? Some rules need to be evaluated by field or record, while others need to be evaluated by time window or user grouping. For example, debouncing rules need to be grouped by user and time. Data volume: When the data volume is huge, partitioning is necessary to avoid insufficient memory or processing timeouts. The granularity of partitioning (e.g., by file, by time block, by user sharding) needs to balance parallel processing capabilities and subdata size. Basic information set: Data structure (which fields are suitable as partitioning keys), intended use (which dimensions the analysis needs to aggregate), known problems (whether known problem areas need to be partitioned separately).
[0066] Common partitioning patterns include: Partitioning by time: If the data has obvious time attributes and its quality changes over time, it can be partitioned by day, hour, etc.; Partitioning by field value (hash / range): If the data volume is huge and needs to be processed in parallel, it can be partitioned by the hash value or range of a key field (such as user_id, product_id), often used in distributed computing; Partitioning by file / data block: The simplest partitioning method, directly partitioning by physical storage files or data blocks automatically split by big data frameworks (such as Spark, Hadoop); Hybrid partitioning: Combining multiple methods, such as first partitioning by time, and then partitioning by user hash within each time block; Defining partitioning rules, clarifying the specific partitioning operations, such as what the partitioning key is, and how to name or identify the sub-data after partitioning.
[0067] Based on the partitioning pattern determined in the previous step, the data partitioning operation is actually performed to generate multiple sub-data. The partitioning logic is implemented using programming languages (Python, Java), big data processing frameworks (Spark, Flink), or database queries (such as SQL's PARTITIONBY). The partitioned sub-data is stored in appropriate locations, such as distributed file systems (HDFS, S3), different partitions of the database, or tables. It is ensured that the sub-data can be accessed by subsequent steps (S132). Information such as the source, partitioning method, and data range (such as time range, hash range) of each sub-data is recorded so that subsequent steps can correctly associate and process it.
[0068] The raw, unstructured data to be processed is transformed into a series of sub-data with a clear scope and structure, and combined with the data background and quality assessment criteria, laying the foundation for subsequent refined quality inspection and assessment (step S132). The choice of partitioning mode directly affects the efficiency and accuracy of subsequent processing and requires careful consideration.
[0069] Furthermore, clarify the specific storage location or logical location of each sub-data in the system; determine the physical path of the sub-data storage; for example, which directory the file is in on HDFS, which table and which partition in a relational database, and which bucket and which path in object storage (such as S3); determine the logical affiliation of the sub-data in the data architecture; for example, which data domain (such as user domain, transaction domain), which data theme (such as user profile, order details), and which data model it belongs to.
[0070] Assess the importance or urgency of each data sub-data point and assign it a priority. Higher priority data sub-data usually means they have a greater impact on the business, are used more frequently, or are used in critical scenarios. Assess data type: different types of data have different importance; for example, core transaction data (such as order amount, payment status) is usually more important than user preference data (such as browsing history). Assess data usage frequency: data that is frequently queried, analyzed, or used for real-time decision-making has higher priority; this can be obtained through monitoring systems, data warehouse query logs, ETL task scheduling frequency, etc. Assess data application scenarios: analyze which business scenarios the data is used for; for example, data used for core transaction processing, risk control, and key reports has higher priority than data used for internal research and non-core operations; data in critical scenarios (such as fraud detection, inventory management) has higher priority. Based on these three dimensions, design a scoring model or rule engine to calculate a priority score or level (such as high, medium, low) for each data sub-data point; for example, a weighted summation can be used: Priority Score = w1 * Type Weight + w2 * Frequency Weight + w3 * Scenario Weight.
[0071] Perform quality checks on each sub-data set to quantify its quality status. Based on the data quality system determined in step S11, select applicable quality rules for each sub-data set. Rules include completeness (e.g., NOT NULL checks), accuracy (e.g., value range checks, pattern matching), consistency (e.g., key field matching), timeliness (e.g., timestamp validity), and uniqueness (e.g., primary key uniqueness). Use automated tools (e.g., GreatExpectations, Deepu, custom scripts) to perform the selected rule checks on each sub-data set. The tools will record the pass / fail status of each rule on each sub-data set, as well as the specific number or proportion of failed records. Calculate the quality score for each sub-data set based on the severity and impact of the check results. For example, the percentage of rules that pass, a weighted average score (different rules have different weights), or a score based on error type and quantity can be calculated. The score range can be from 0 to 100, or divided into several levels (e.g., Excellent, Good, Average, Poor).
[0072] Each sub-data item is labeled with three key tags: location, priority, and quality score. This information forms the basis for the subsequent step (S133) to perform refined quality level classification, making the data quality assessment more specific and targeted, and able to reflect the actual status of the data in the system and its business importance. In practical applications, the selection of quality detection rules, the design of priority calculation models, and the calculation method of quality scores all need to be carefully designed to ensure the accuracy and effectiveness of the assessment results.
[0073] Therefore, a first-level coefficient is determined based on the data location and priority of multiple sub-data points, and a second-level coefficient is determined based on the priority and quality score of multiple sub-data points. The quality level of each sub-data point is determined based on the mapping relationship between the first-level coefficient, the second-level coefficient, and the quality level. The quality level of the data to be processed is determined based on the quality level and data location of each sub-data point, which takes into account the overall consideration of the quality level and data location of each sub-data point, ensuring the accuracy of the quality level of the data to be processed. At the same time, a data quality system is introduced, which takes into account the overall consideration of the quality level of each sub-data point, improving the accuracy of the quality level of the data to be processed.
[0074] At this point, the first-level coefficient is determined based on the data location and priority of multiple sub-data. This coefficient reflects the importance of the sub-data in the system (determined by priority) and its management convenience or risk (determined by data location).
[0075] Priority (from S132): The importance of data type, frequency of use, and application scenario has been quantified; Data location: Introduces additional weights or adjustment factors; For example: Easily accessible locations (such as the master data warehouse, commonly used data lake directories): Give priority a small positive adjustment because it is easy to manage and monitor; Difficult-to-access or isolated locations (such as temporary tablespaces, archives, third-party system data): Give priority a small negative adjustment or penalty because management and repair are difficult and the risk is higher; First-level coefficient calculation: It can be a simple weighted average or a more complex function; For example: First-level coefficient = α * priority + (1-α) * location weight, where α is a balancing factor between 0 and 1, and the location weight is a value pre-set based on location characteristics (accessibility, risk, etc.); If the location has no special influence, it can be simplified to the first-level coefficient being mainly equal to the priority.
[0076] The second-level coefficient is determined based on the priority and quality score of multiple sub-data points. This coefficient reflects the degree of match or conflict between the importance of the sub-data points (determined by priority) and their actual quality status (determined by quality score).
[0077] Priority (from S132): This again reflects the importance of the data; Quality Score (from S132): This represents the degree to which the data meets the quality requirements; Second-Level Coefficient Calculation: This coefficient is designed to highlight either "important data with poor quality" or "secondary data with good quality"; Highlighting Important Data with Poor Quality: This can be designed as a combination of priority and quality score, so that when the priority is high but the quality score is low, the coefficient value is significantly reduced; for example: Second-Level Coefficient = Priority * (Quality Score / 100) (assuming the quality score is a percentage of 0-100), thus, a high priority but a low quality score will result in a lower second-level coefficient; Highlighting Secondary Data with Good Quality: This can also be designed so that when the priority is low but the quality score is high, the coefficient value is relatively high (but still lower than the case of high priority and high score); for example, using the above formula, low priority multiplied by high score results in a moderate result.
[0078] The two coefficients calculated earlier are transformed into a standardized quality level label (e.g., Excellent, Good, Average, Poor, or 1-5 levels) through a predefined rule (mapping relationship). The quality level mapping relationship is a crucial configuration or rule set; it defines how the combination of the first and second level coefficients maps to the final quality level. This can be a complex scoring formula. A comprehensive scoring formula can be designed, for example: Comprehensive Score = w1 * First Level Coefficient + w2 * Second Level Coefficient; then, levels are assigned based on the comprehensive score (e.g., score > 0.8 is Excellent, 0.6-0.8 is Good, etc.). w1 and w2 are weights, reflecting the relative importance of positional factors and priority-quality matching factors. Determining the quality level of sub-data: For each sub-data, its calculated first and second level coefficients are substituted into the mapping relationship to find or calculate the corresponding quality level.
[0079] The overall quality level of the entire dataset is obtained by summing the quality levels of all sub-data. How to derive an overall level from multiple sub-levels depends on business requirements and the definition of overall data quality. A weighted average method is used: the quality level of each sub-data is weighted according to its size (data volume) or priority. For example, levels are mapped to scores (Excellent=5, Good=4, Average=3, Poor=2, Inferior=1), a weighted average score is calculated, and then mapped back to a level. Overall quality score = Σ(sub-data quality score * sub-data weight) / Σ(sub-data weight); the quality level of the data to be processed = the level mapped from the overall quality score.
[0080] During aggregation, data location comes into play again; for example, if the data location (sub-data) corresponding to a critical business system generally has a low quality, even if the data quality in other locations is good, the overall quality will be dragged down; or, rules can be set such that if the data quality in any "core" location is below a certain threshold, the overall quality cannot be higher than that threshold; the selected aggregation strategy is executed to obtain the final quality level of the data to be processed.
[0081] refer to Figure 5 In step S14, the specific steps are as follows:
[0082] S141: Determine a preset quality level based on the matching between the data to be processed and the big data system, and compare the quality level of the data to be processed with the preset quality level; if the quality level of the data to be processed is higher than the preset quality level, the big data system directly receives the data to be processed and transmits the data to be processed to the corresponding data space;
[0083] S142: If the quality level of the data to be processed is lower than the preset quality level, the data to be processed will be controlled accordingly. At this time, the data to be processed will be traced, and the data to be processed will be marked at each data control node in the data collection process during the tracing process. The data collection path will be determined based on the location of each data control node and the quality level of the data to be processed.
[0084] S143: Multiple data control areas are determined based on the detection of the data acquisition path. Each data control area is matched with at least one data control node. Multiple data control coefficients are determined based on the detection of the data control areas. The corresponding data control events are determined based on the mapping relationship between the multiple data control coefficients and control events. Each data control node is matched with the corresponding data control event.
[0085] In the embodiments of this application, a preset quality level is determined based on the matching between the data to be processed and the big data system. The quality level of the data to be processed is compared with the preset quality level. If the quality level of the data to be processed is higher than the preset quality level, the big data system directly receives the data to be processed and transmits the data to be processed to the corresponding data space. This approach takes into account the overall matching between the data to be processed and the big data system, ensuring the accuracy of the preset quality level.
[0086] At this point, a "threshold" quality level is determined for the data to be processed to measure whether it can be directly accepted by the system. Here, "matching" refers to analyzing whether the characteristics of the data to be processed are consistent with the needs of the big data system. This includes: data type / category: Is the data user behavior logs, transaction records, device sensor data, or others? Different types of data have different uses in the system and different quality requirements; application scenario: Will the data be used for real-time recommendations, financial statements, operational analysis, or model training? Core business decisions usually require higher quality data; data space attribution: According to the data space defined in step S11, which space does the data belong to? Different spaces have different quality baselines.
[0087] Based on the results of the above "matching" analysis, a suitable level is selected from the multiple quality level standards preset by the system as the benchmark for this comparison. The preset quality level is usually defined in the system design stage and is a globally unified minimum standard, as well as different standards set for different data types, application scenarios, or data spaces. For example, the system presets are: Level A (High): Data used for core transaction systems and key decision support, requiring extremely high accuracy, completeness, and timeliness; Level B (Medium): Data used for routine business analysis and reports, allowing for a small number of errors or omissions; Level C (Low): Data used for exploratory analysis and log monitoring, with relatively relaxed quality requirements.
[0088] The process involves several steps: First, determining whether the quality of the data to be processed meets the system's set acceptance criteria. Second, obtaining the quality level of the data to be processed: This level is calculated in step S13 (specifically S133) and comprehensively reflects the overall quality level of the batch of data (e.g., high, medium, low, or a specific score). Third, obtaining the preset quality level: As described in S141, the "threshold" level determined based on data characteristics and system requirements. Fourth, directly comparing the two levels: The comparison logic typically involves determining whether the quality level of the data to be processed "meets" or "is higher than" the preset level. Here, "higher than" can be understood as the quality not being lower than the preset standard. For example, if the preset level is "Level A (High)," and the level calculated for the data to be processed is also "Level A (High)," then the condition is met. If the preset level is "Level B (Medium)," and the level calculated for the data to be processed is "Level A (High)," then the condition is met. If the preset level is "Level B (Medium)," and the level calculated for the data to be processed is "Level C (Low)," then the condition is not met.
[0089] When the data quality meets the standards, it is successfully incorporated into the big data system for storage and management. Based on the comparison results of S141, if the quality level of the data to be processed reaches or exceeds the preset level, a decision to "directly accept" is made. The receiving module of the big data system (which is a data lake, data warehouse, or specific subject database) formally accepts this batch of data, which involves data format conversion, preliminary storage confirmation, and other operations. According to the division results of step S11, it is determined which (or which) "data space" this batch of data belongs to. Then, the received data is physically or logically moved to the storage area or partition corresponding to the data space. This ensures that the data is organized according to its type, source, or purpose, which facilitates subsequent querying, analysis, and application.
[0090] Furthermore, if the quality level of the data to be processed is lower than the preset quality level, the data to be processed is subject to corresponding control. At this time, the data to be processed is traced, and the data control nodes of the data to be processed in the collection process are marked during the tracing process. The data collection path is determined based on the location of each data control node and the quality level of the data to be processed. This takes into account the overall consideration of the location of each data control node and the quality level of the data to be processed, and ensures the accuracy of the data collection path.
[0091] At this point, a special processing procedure for low-quality data is initiated, instead of directly releasing or discarding it. First, the system needs to stop executing the direct receiving procedure defined in step S141 for this batch of low-quality data; mark the status of this batch of data as "pending control" or "quality anomaly" for subsequent tracking and processing; trigger a processing branch specifically for low-quality data, ready to perform subsequent traceability and analysis operations.
[0092] Specifically, assuming step S13 evaluates a batch of data from the "User Feedback Form" with a quality level of C (low) and a preset level of B (medium); the system will not directly receive this batch of data; instead, it will: suspend the flow of this batch of data to the User Feedback Analysis Module; change the label of this batch of data to "Pending Control: Low Quality" in the data management platform; and start a "Low Quality Data Traceability" task.
[0093] Tracing the entire lifecycle of data from its source to its current pending state helps identify the steps that introduce quality issues. Using data lineage tracing technology, the origin of this data (e.g., which business system, table, field, and time range of data) is traced. The ETL (Extract, Transform, Load) processes, cleaning rules, and transformation logic that this data underwent before entering the big data system are reviewed. These processes are checked to see if errors, missing values, or outliers were introduced. Based on the specific manifestations of the quality issues (if known), the point of introduction is inferred. For example, if data is severely missing, it may be due to an unstable interface in the source system, or a cleaning step in the ETL process being skipped or misconfigured.
[0094] Specifically, the system traced the data from the "User Feedback Form" with a quality level of C. The system began backtracking and discovered that the data originated from the "feedback" table in the "Customer Service System," with fields including "User ID," "Feedback Type," "Feedback Content," and "Submission Time." The ETL process was examined: the data was extracted and transformed by a Spark job named "feedback_etl" and then loaded into the big data warehouse. The job's code and configuration were checked, and a problem was found: during the transformation phase, a rule was set to uniformly convert "Suggestion" in the "Feedback Type" to "Suggestion," but this rule was recently temporarily commented out. This caused "Suggestion" type data to directly enter the subsequent process, resulting in inconsistent data types and affecting data consistency.
[0095] Along the data traceability path, identify and record the key points or links responsible for data quality control. These nodes are usually the links in the data lifecycle where specific quality assurance measures are implemented; for example: data verification points in the source system (such as database constraints, application layer verification); data cleaning steps and verification rule execution points in the ETL process; format checks and uniqueness checks when data is loaded into the big data platform; quality audits before data enters the data warehouse / data mart; during the traceability process, whenever such a control node is passed, it is recorded and associated with the current batch of data to be processed.
[0096] Specifically, during the process of tracing the "User Feedback Form" data, the system marked the following control nodes on the backtracking path: Node 1: The "feedback" table in the customer service system database - field non-null constraints (checking whether "User ID" and "Submission Time" are null); Node 2: The "Type Conversion" step in the "feedback_etl" Spark job (originally planned to verify and convert the "feedback type"); Node 3: The "Data Loading" step in the "feedback_etl" Spark job - checking the field format loaded into the Hive table; Marking result: The system recorded that for this batch of low-quality data, they passed through Node 1, Node 2, and Node 3.
[0097] By combining the location information of the control nodes and the low quality of the data itself, we can infer the most problematic data collection path, taking into account the order of the control nodes in the data flow. The problem occurs before the downstream nodes in the data flow. The lower the quality level, the more serious the problem and the more links involved. A low level means that multiple control nodes have failed, or that the failure of a key node has caused a chain reaction. Based on the node location and the severity of the problem, we can draw or infer a data flow path from the source to the current state, and specifically mark the marked control nodes, especially those that have failed or have problems. This path is the "data collection path" that needs to be focused on for subsequent control.
[0098] Specifically, the quality level is known to be C, and nodes 1, 2 (failed), and 3 have been marked. System analysis: Node 2 (type conversion) is located after node 1 (source verification) and before node 3 (load verification). Quality level C indicates a serious problem, with issues in more than one step. However, based on tracing, the main problem lies in the commenting out of the rules at node 2. Determined path: The system determines the data acquisition path as: source system (node 1) > ETL job (node 2 - problem point) > load to Hive (node 3). This path specifically emphasizes that node 2 is the key link leading to low data quality.
[0099] Therefore, multiple data control areas are determined based on the detection of the data acquisition path, and each data control area is matched with at least one data control node. Multiple data control coefficients are determined based on the detection of the data control areas, and the corresponding data control events are determined based on the mapping relationship between the multiple data control coefficients and control events. Each data control node is matched with the corresponding data control event. This approach takes into account the overall consideration of multiple data control coefficients and control event mapping relationships in ABC, ensuring the accuracy of the corresponding data control events.
[0100] At this point, the problematic data acquisition paths identified in step S142 are divided into several logically or functionally related areas for centralized management; the identified data acquisition paths are reviewed (e.g., node 1 > node 2 > node 3); based on the location, function, and processing logic of the nodes on the path, adjacent or functionally similar nodes are grouped together; an area typically contains one or more data control nodes (the nodes marked in S142), and clear boundaries and included nodes are defined for each area; for example, data source connection and preliminary verification can be considered as one area, and complex transformation logic can be considered as another area.
[0101] Specifically, the data acquisition path is: source system (node 1) > ETL job (node 2 - problem point) > loading into Hive (node 3); system analysis of this path: Region 1: defined as the region consisting of node 1 (source system) and node 2 (ETL job), named "Data Extraction and Transformation Region", because node 2 is the problem point and it is directly related to the source data, so it is more effective to analyze it together; Region 2: defined as the region where node 3 (loading into Hive) is located, named "Data Loading Region"; although node 3 itself is not the problem, it is the final stage where the data is implemented, so it also needs to be monitored.
[0102] Assign a numerical value to each defined data control area, representing the intensity or priority of control over that area; evaluate each control area, considering factors including: the importance of the control nodes within the area (e.g., source verification nodes are more important than log recording nodes); the severity of known problems within the area (e.g., node 2 is a critical point leading to low quality, so the area coefficient should be higher); the area's position in the data flow (source areas are generally more important than downstream areas); the frequency and volume of data processed by the area (high-frequency, large-volume areas require higher control); and convert the evaluation results into a specific numerical value (control coefficient) according to preset rules or algorithms; the higher the coefficient, the more attention or stronger control measures are needed for that area.
[0103] Specifically, there are Region 1 (data extraction and transformation region) and Region 2 (data loading region). The system evaluates and scores these two regions: Region 1: Contains problem node 2 and involves data extraction and transformation logic, making it the core region for data quality issues; the evaluation result considers this region to be of high risk and high importance; the calculated control coefficient is 0.8 (assuming the coefficient ranges from 0 to 1, with higher coefficients indicating stronger control); Region 2: Although node 3 is the final data storage point, according to tracing, the problem mainly lies in Region 1; Region 2 itself is simply loading normally; the evaluation result considers this region to be of medium risk; the calculated control coefficient is 0.4.
[0104] Based on the control coefficient of each control area and the preset mapping rules, determine the specific control actions that need to be performed in that area; predefine the correspondence between the control coefficient range and the specific control events; for example: coefficient > 0.7: trigger high-intensity events such as "deep diagnosis", "manual review", "pause processing and alarm"; 0.4 < coefficient < 0.7: trigger medium-intensity events such as "add verification rules", "record detailed logs", "send early warning notifications"; coefficient < 0.4: trigger low-intensity events such as "standard log recording" and "minor alarms"; substitute the control coefficient of each control area into the mapping relationship to find the corresponding control events.
[0105] Specifically, the coefficient for Region 1 is 0.8, and the coefficient for Region 2 is 0.4. The system queries the preset mapping relationship: Region 1 (coefficient 0.8): Based on the mapping relationship (assuming a high-intensity event is triggered if the coefficient > 0.7), the corresponding data control events are determined as follows: Event A: Suspend the data processing flow in this region and send an emergency alert email to the administrator; Event B: Conduct in-depth code review and testing on the transformation logic of Node 2 (ETL job); Region 2 (coefficient 0.4): Based on the mapping relationship (assuming a medium-intensity event is triggered if 0.4 < coefficient < 0.7), the corresponding data control events are determined as follows: Event C: Add a verification rule for the integrity of the target table fields at Node 3 (loaded into Hive); Event D: Record detailed logs of the Node 3 loading process, including loading time, data volume, number of errors, etc.
[0106] The events identified in step S143 for the entire control area are further assigned to specific control nodes within that area for execution; each control event is analyzed to see which one or more specific nodes within the area it affects; an event may affect only one node or multiple nodes within the area; the events to be executed are then explicitly assigned to the corresponding control nodes, so that when subsequent control is executed, it is known which action to perform on which node.
[0107] Specifically, Region 1 contains Node 1 and Node 2, and events A and B; Region 2 contains Node 3, and events C and D. The system assigns events to specific nodes: Region 1: Event A (Pause process, send alarm): mainly affects the entire region, but the specific execution is triggered by the region's entry node (such as after Node 1), or by a dedicated monitoring node; let's assume it's triggered by Node 1; Event B (deep review of ETL logic): explicitly assigned to Node 2 (ETL job); Region 2: Event C (add validation rules): explicitly assigned to Node 3 (loaded into Hive); Event D (record detailed logs): explicitly assigned to Node 3 (loaded into Hive).
[0108] The management of low-quality data is no longer general, but precisely targets the area and process where the problem occurs, and specifies concrete handling actions. This greatly improves the pertinence and efficiency of problem repair. In practical applications, the rationality of the mapping relationship, the effectiveness of the event design, and the accuracy of the node and region division are key.
[0109] refer to Figure 6 In step S15, the specific steps are as follows:
[0110] S151: Collect data control events from each data control node, determine multiple sub-data control items based on the detection of each data control event, determine abnormal events based on the data control content, data control level and data function corresponding to each data control node of each sub-data control item, and determine events to be optimized based on the abnormal events and the data to be processed;
[0111] S152: Collect the data types corresponding to the data to be processed, and determine the application scenarios of the data to be processed based on the data types and data usage requirements. Determine the first quality management coefficient based on the events to be optimized and the data types corresponding to the data to be processed.
[0112] S153: Determine the second quality management coefficient based on the application scenario of the event to be optimized and the data to be processed, and determine the quality management event of the data to be processed based on the first quality management coefficient, the second quality management coefficient and the quality control mapping relationship.
[0113] In the embodiments of this application, data management events of each data management node are collected, multiple sub-data management items are determined based on the detection of each data management event, abnormal events are determined according to the data management content, data management level and data function corresponding to each data management node of each sub-data management item, and events to be optimized are determined based on the abnormal events and the data to be processed. This approach takes into account both abnormal events and the data to be processed as a whole, ensuring the accuracy of the events to be optimized.
[0114] At this point, raw information about data quality issues is collected from data quality monitoring tools, rule engines, data probing tools, etc. (i.e., data control nodes) deployed in the system. These nodes continuously monitor data streams, database tables, and data warehouses. When data violates preset quality rules (such as uniqueness, non-nullability, format, range, consistency, etc.), the node generates a "data control event." A centralized mechanism (such as a log collection system, event bus, or data quality platform) needs to be established to aggregate these raw events from different nodes with varying formats, and output a batch of raw, unprocessed "data control event" records. Each record typically includes: event ID, occurrence time, data source / path, data entity (table / file name), field name, specific issues detected (such as "duplicate value", "null value", "format mismatch"), and the original data value.
[0115] Original, repetitive, or related data governance events are categorized and aggregated to form higher-level, more semantic "sub-data governance projects," which helps to understand the nature and scope of the problem. The collected raw events are analyzed and clustered, typically based on data entities (tables / files), field names, and detected problem types (such as "null value issues," "formatting issues," and "consistency conflicts"). Multiple events for the same data entity, field, and problem type can be merged, and the frequency or severity of the event can be recorded. Each clustering result is defined as a "sub-data governance project." For example, "user table - email field - null value issue" can be considered a sub-project, outputting a set of "sub-data governance projects." Each project describes a specific data quality problem pattern and includes information such as the frequency of the pattern's occurrence and the scope of data involved.
[0116] From the "sub-data governance projects," select those events that truly constitute "abnormalities" and assign them a certain severity level. This requires considering the specific content and severity of the problem, as well as the judgment of the data function department responsible for that data. Conduct in-depth analysis of the "data governance content" (i.e., the specific problem type and impact) represented by each sub-project; apply predefined "data governance level" rules; for example, specify "null values in critical fields (such as primary keys and unique identifiers)" as the highest level, and "format errors in non-critical fields" as a lower level; apply the level rules to each sub-project; and consider the "data function corresponding to the data governance node" (i.e., which department or team is responsible for that data). (Maintenance and management); Sometimes, even if the problem itself is not serious, if it involves key data under the responsibility of core business departments, it is also raised to an anomaly. This can be based on preset rules or manually configured weights. Based on the above analysis, determine which sub-projects meet the criteria for "anomaly" (e.g., the level reaches a certain threshold, or it involves key data of a specific functional department); mark the sub-projects that meet the criteria as "anomaly events" and output a batch of records that have been confirmed as "anomaly events"; each record should include: anomaly event ID, information on the associated sub-data control project, the determined anomaly level (e.g., "serious", "medium", "general"), and information on the associated data function, etc.
[0117] The identified "abnormal events" are associated with the data currently requiring processing (such as data batches about to enter the production environment or a specific dataset to be analyzed) to determine whether these anomalies directly affect the data to be processed, thereby identifying specific optimization tasks. The scope and identifier of the current "data to be processed" are determined (e.g., specific data batch ID, dataset name, time range, etc.). It is then checked whether the "sub-data control projects" associated with each "abnormal event" overlap with the "data to be processed." For example, if the abnormal event is "user table - email field - null value issue," and the data to be processed is "annual report dataset containing user table data," then these two overlap. For anomalies with overlap... The event is defined as follows: determine its specific impact on the "data to be processed"; for example, will empty mailboxes cause annual report calculation errors or prevent generation? For those abnormal events that do affect the data to be processed and require action to resolve, transform them into one or more specific "events to be optimized"; each event to be optimized should clearly indicate: the specific data problem to be optimized (originating from the abnormal event), the scope of data involved (originating from the data to be processed), and the suggested optimization actions (such as "filling in empty values", "correcting the format", "removing problematic data", "investigating and fixing the source problem"); a set of "events to be optimized"; each event describes a specific quality improvement task that needs to be performed on the current data to be processed.
[0118] Furthermore, the data types corresponding to the data to be processed are collected, and the application scenarios of the data to be processed are determined based on the data types and data usage requirements. The first quality management coefficient is determined based on the data types corresponding to the events to be optimized and the data to be processed, taking into account the overall consideration of the data types corresponding to the events to be optimized and the data to be processed, thus ensuring the accuracy of the first quality management coefficient.
[0119] At this point, it's crucial to identify the type (or categories) of data to be processed. Data classification can be based on various dimensions such as data structure, source, and business meaning. Extract metadata information from the data source or data directory, and label the data to be processed with the appropriate "data type" according to predefined classification rules or standards. For example, data types may include "customer information," "transaction records," "equipment sensor data," and "product attributes." Different types of data have different business values and sensitivities, and they also differ significantly in data quality requirements, processing procedures, and control strategies. Identifying the data type is fundamental to subsequent analysis.
[0120] Specifically, suppose we have a batch of data exported from an online shopping mall order system. This batch of data contains information such as the user's order number, user ID, product ID, purchase time, amount, and shipping address. Based on our business understanding, we can classify this batch of data into two main types: "transaction records" and "customer information" (or further subdivided into "order data," "user data," etc.).
[0121] Describe the business environment or usage scenario of the data by combining the data type and how it will be used; the application scenario describes the role of the data in a specific business process or decision support; understand what the data will ultimately be used for, such as financial statement statistics, user profiling analysis, inventory management, precision marketing, or after-sales service. These requirements usually come from the business department's requirements documents, project plans, or system function designs, considering the role of different data types in specific application scenarios; for example, in the "precision marketing" scenario, "customer information" data is used to identify target customers, and "transaction record" data is used to analyze purchasing behavior; combine the data types and usage requirements to form a specific and identifiable business scenario description, which can be a brief text description, a scenario code, or a label.
[0122] Different application scenarios have different requirements for data quality. For example, the "financial statement statistics" scenario has extremely high requirements for the accuracy and completeness of data, while the "user profile analysis" scenario has higher requirements for the comprehensiveness of data and a slightly higher tolerance for the missing of individual fields. Clearly defining the application scenario helps to more accurately assess and manage data quality in the future.
[0123] Specifically, assuming this batch of order data will be used for the "Annual Customer Value Analysis" project, which aims to identify high-value customers and maintain them in a targeted manner; then, combining the data type (transaction records, customer information) and the data usage requirements (customer value analysis), we can determine that the application scenario for this batch of data is "Annual Customer Value Analysis".
[0124] Calculate a preliminary quality management coefficient, which reflects the importance and urgency of the events to be optimized under the current data type and application scenario; review the "events to be optimized" identified in step S151, including events such as "a field has a large number of null values", "a field has inconsistent format", and "a field has logical errors"; analyze the impact of each event to be optimized on the "data type" and "application scenario"; for example, for the "annual customer value analysis" scenario, the event "the customer ID field has null values" has a very large impact because it directly leads to the inability to associate customer information; while the event "the product description field has a few spelling errors" has a smaller impact unless the spelling errors cause the products to be incorrectly classified.
[0125] Based on predefined rules or models, calculate the "first quality management coefficient". This coefficient can be a comprehensive score or a weight for a specific event. The calculation usually considers: event severity: the degree of impact caused by the event itself; data type sensitivity / importance: the criticality of the data type in the business; scenario requirements: the specific requirements of the application scenario for data quality; event frequency / scope: the frequency of the event or the amount of data affected.
[0126] A simple weighted formula can be set: First quality management coefficient = Σ(Severity score of event i * Impact weight of event i on the current scenario * Frequency factor of event i); or, a more complex machine learning model can be built, with input features including event type, data type, scenario code, historical processing results, etc., and output quality management coefficient. This coefficient provides a quantitative basis for subsequent steps, making quality management decisions more scientific and objective; it helps to distinguish which data quality issues are more worthy of priority treatment.
[0127] Step S152 completes the process from understanding data types and defining application scenarios to quantifying preliminary quality risks (first quality management coefficient), laying the foundation for more refined quality management decisions (such as step S153). In practical applications, the definition of data types, the identification of application scenarios, and the calculation method of quality management coefficients all need to be customized and optimized according to specific business and data characteristics.
[0128] Therefore, a second quality management coefficient is determined based on the application scenario of the event to be optimized and the data to be processed. The quality management event of the data to be processed is determined based on the first quality management coefficient, the second quality management coefficient, and the quality control mapping relationship. This approach is compatible with the overall consideration of the first quality management coefficient, the second quality management coefficient, and the quality control mapping relationship, ensuring the accuracy of the quality management event of the data to be processed. At the same time, data control events are introduced, realizing the overall consideration of the data types and functions corresponding to the event to be optimized and the data to be processed, improving the accuracy of the quality management event of the data to be processed, and achieving data quality management.
[0129] At this point, assess the severity and urgency of the events to be optimized in a specific application scenario, and generate a scenario-related quality risk quantification value; analyze the specific problems caused by each event to be optimized (such as missing values, duplicate records, format errors, etc.) in the current application scenario; for example, a missing user email address is fatal in a user registration scenario, but has little impact in a simple product browsing scenario; based on factors such as the severity, occurrence, and scope of the scenario-specific impact, assign a scenario impact score to each event to be optimized. This score can be an absolute value (such as 1-10 points) or a relative weight.
[0130] The second quality management coefficient is calculated by combining the scenario impact score and event frequency. The calculation method can be a simple weighted average or a more complex model. This coefficient reflects the overall quality risk level of the event in a specific scenario. The scenario impact score is adjusted based on the frequency or proportion of the event to be optimized in the data. High-frequency minor issues require more attention than low-frequency serious issues. The first quality management coefficient focuses on the general quality importance of the data itself, while the second quality management coefficient focuses on the actual value and risk of the data in a specific application scenario, making quality management decisions more aligned with business needs.
[0131] The two quality management coefficients are combined and a specific, actionable quality management action plan is finally output through a predefined rule base (quality control mapping relationship). The first quality management coefficient calculated in step S152 and the second quality management coefficient calculated in step S153 are integrated. The integration method can be simple addition, weighted summation, or more complex logical judgment (e.g., taking the higher value, or judging based on the combined range of the two coefficients).
[0132] The quality control mapping relationship is a predefined rule table or decision tree; it defines which quality management events correspond to different combinations of coefficients; for example: total coefficients < 5: quality management event = "record the problem and monitor it regularly"; 5 ≤ total coefficients < 15: quality management event = "perform data cleaning and prioritize high-frequency issues"; total coefficients ≥ 15: quality management event = "immediately perform data cleaning, suspend related business applications, trace the source and fix it"; based on the merged coefficient values, the most appropriate quality management event is found in the mapping relationship. This event is a specific action instruction that tells the data team what to do. This step transforms the results of quantitative analysis into actual action; the design of the quality control mapping relationship directly determines the effectiveness and timeliness of quality management measures.
[0133] Specifically, the first quality management coefficient is 26; the second quality management coefficient is 4.25. Assuming we combine these two coefficients using a weighted summation method, and set the weight of the first coefficient to 0.6 and the weight of the second coefficient to 0.4: the combined coefficient = (26 * 0.6)(4.25 * 0.4) = 15.6 * 1.7 = 17.3. Now, we query the quality control mapping relationship: assuming the mapping relationship is as follows: coefficient < 10: event = "low priority, record and monitor"; 10 ≤ coefficient < 20: event = "medium priority, schedule data cleaning task"; coefficient ≥ 20: event = "high priority". "Level 1, immediately clean, suspend related applications, and investigate the source"; the merged coefficient is 17.3, falling within the range of 10 ≤ coefficient < 20. Therefore, the determined quality management event is: "Medium priority, arrange data cleaning tasks." This means the data team needs to: prioritize handling the missing phone number issue (Event A), as it contributes significantly to the first coefficient and has a high impact on the scenario; the product name spelling error issue (Event B) can be handled later, or treated as a secondary task; execute the data cleaning script to fill in or correct missing / erroneous data; and use the cleaned data for annual customer value analysis.
[0134] Please see Figure 7 , Figure 7 This is a schematic diagram of the structural composition of a data quality management system applied to a big data system in an embodiment of the present invention; the data quality management system applied to the big data system includes:
[0135] Advantage data module 21 is used to determine multiple data spaces based on the detection of the big data system, and to determine multiple advantageous data based on the data volume, spatial form and current state of the big data system of the multiple data spaces;
[0136] Data quality system module 22 is used to determine the concentrated area of advantageous data based on the distribution location of multiple advantageous data, and to determine the data quality system based on the concentrated area of advantageous data and the previous data optimization events of the big data system;
[0137] The quality level module 23 is used to divide the data to be processed in the data quality system to output multiple sub-data. The quality level of each sub-data is determined according to the data position, priority and quality score of the multiple sub-data. The quality level of the data to be processed is determined based on the quality level of each sub-data.
[0138] The data control event module 24 is used to determine the data collection path based on the tracing of the data to be processed if the quality level of the data to be processed is lower than the preset quality level, and to match the corresponding data control events to each data control node of the data collection path.
[0139] The quality management event module 25 is used to determine the events to be optimized based on the data control events of each data control node, and to determine the quality management events of the data to be processed based on the data type and application scenario corresponding to the events to be optimized and the data to be processed.
[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A data quality management method applied to big data systems, characterized in that, include: Multiple data spaces are identified based on the detection of the big data system, and multiple advantageous data are determined based on the data volume, spatial form and current state of the big data system. The concentration area of the advantageous data is determined based on the distribution of multiple advantageous data, and the data quality system is determined based on the concentration area of the advantageous data and the previous data optimization events of the big data system. The data quality system divides the data to be processed into multiple sub-data. The quality level of each sub-data is determined based on its data location, priority, and quality score. The quality level of the data to be processed is then determined based on the quality level of each sub-data. If the quality level of the data to be processed is lower than the preset quality level, the data collection path is determined based on the traceability of the data to be processed, and the corresponding data control events are matched to each data control node of the data collection path. Based on the data control events of each data control node, determine the events to be optimized, and based on the data types and application scenarios corresponding to the events to be optimized and the data to be processed, determine the quality management events of the data to be processed.
2. The data quality management method applied to big data systems according to claim 1, characterized in that, The detection based on the big data system identifies multiple data spaces, and based on the data volume, spatial form, and current state of the big data system, identifies multiple advantageous data, including: The big data system is used to collect data. Based on the system parameters and data volume of the big data system, the corresponding detection method is determined, and the detection of the big data system is triggered along the detection method to identify multiple data spaces. At this time, each data space contains data of the corresponding type. In each data space, the data volume and spatial form of the data space are determined based on the identification of the data space. At the same time, multiple working parameters of the big data system are collected, and the current state of the big data system is determined based on the big data system. A first advantageous data set is determined based on the data volume and spatial form of multiple data spaces. A second advantageous data set is determined based on the data volume of multiple data spaces and the current state of the big data system. Multiple advantageous data sets are then determined based on the first and second advantageous data sets.
3. The data quality management method applied to big data systems according to claim 1, characterized in that, The process of determining the concentrated area of advantageous data based on the distribution location of multiple advantageous data points, and determining the data quality system based on the concentrated area of advantageous data and past data optimization events of the big data system, includes: Collect multiple advantageous data and mark the distribution location of multiple advantageous data. Determine multiple sub-advantage regions based on the distribution location, data type and data volume of multiple advantageous data. Determine the concentrated area of advantageous data based on the cross combination of multiple sub-advantage regions. Collect data from the big data system, and determine past data optimization events based on the traceability of the big data system. Then, determine multiple data optimization projects based on the analysis of past data optimization events. A data quality system is determined by cross-training based on multiple data optimization projects, concentrated areas of advantageous data, and the distribution locations of multiple advantageous data. At this time, the data quality system is used to control the data quality of data in the big data system or external data to be processed.
4. The data quality management method applied to big data systems according to claim 1, characterized in that, The data quality system divides the data to be processed into multiple sub-data sets. Based on the data location, priority, and quality score of each sub-data set, the quality level of each sub-data set is determined. Then, based on the quality levels of each sub-data set, the quality level of the data to be processed is determined, including: Collect the data to be processed, determine the basic information set of the data to be processed based on the traceability of the data to be processed, and at the same time, collect the data quality system, determine the corresponding data partitioning pattern based on the data quality system, the data volume of the data to be processed and the basic information set, and divide the data to be processed according to the data partitioning pattern, and output multiple sub-data.
5. The data quality management method applied to a big data system according to claim 4, characterized in that, The data quality system divides the data to be processed into multiple sub-data sets. It determines the quality level of each sub-data set based on its data location, priority, and quality score. Based on the quality levels of each sub-data set, it determines the quality level of the data to be processed. The system also includes: The data location of multiple sub-data is determined based on the location detection of multiple sub-data, and the priority of multiple sub-data is determined according to the data type, data usage frequency and data application scenario of multiple sub-data. At the same time, the corresponding quality score is determined based on the quality detection of multiple sub-data. The first-level coefficient is determined based on the data location and priority of multiple sub-data points. The second-level coefficient is determined based on the priority and quality score of multiple sub-data points. The quality level of each sub-data point is determined based on the mapping relationship between the first-level coefficient, the second-level coefficient, and the quality level. The quality level of the data to be processed is determined based on the quality level of each sub-data point and its data location.
6. The data quality management method applied to a big data system according to claim 1, characterized in that, If the quality level of the data to be processed is lower than the preset quality level, then a data acquisition path is determined based on the tracing of the data to be processed, and corresponding data control events are matched to each data control node of the data acquisition path, including: A preset quality level is determined based on the matching between the data to be processed and the big data system. The quality level of the data to be processed is compared with the preset quality level. If the quality level of the data to be processed is higher than the preset quality level, the big data system directly receives the data to be processed and transmits the data to the corresponding data space.
7. The data quality management method applied to a big data system according to claim 6, characterized in that, If the quality level of the data to be processed is lower than the preset quality level, then the data acquisition path is determined based on the tracing of the data to be processed, and corresponding data control events are matched to each data control node of the data acquisition path, which also includes: If the quality level of the data to be processed is lower than the preset quality level, the data to be processed will be subject to corresponding control. At this time, the data to be processed will be traced, and the data to be processed will be marked at each data control node in the data collection process during the tracing process. The data collection path will be determined based on the location of each data control node and the quality level of the data to be processed. Multiple data control areas are determined based on the detection of the data acquisition path. Each data control area is matched with at least one data control node. Multiple data control coefficients are determined based on the detection of the data control areas. The corresponding data control events are determined based on the mapping relationship between the multiple data control coefficients and control events. Each data control node is matched with the corresponding data control event.
8. The data quality management method applied to a big data system according to claim 1, characterized in that, The process involves determining the events to be optimized based on the data management events at each data management node, and determining the quality management events for the data to be processed based on the events to be optimized, the data types corresponding to the data to be processed, and the application scenarios. This includes: Data control events are collected from each data control node. Based on the detection of each data control event, multiple sub-data control items are determined. Abnormal events are determined according to the data control content, data control level, and data function corresponding to each data control node of each sub-data control item. Based on the abnormal events and the data to be processed, events to be optimized are determined.
9. The data quality management method applied to a big data system according to claim 8, characterized in that, The process of determining the events to be optimized based on the data management events of each data management node, and determining the quality management events of the data to be processed based on the events to be optimized, the data types corresponding to the data to be processed, and the application scenarios, also includes: Collect the data types corresponding to the data to be processed, and determine the application scenarios of the data to be processed based on the data types and data usage requirements. Determine the first quality management coefficient based on the events to be optimized and the data types corresponding to the data to be processed. The second quality management coefficient is determined based on the application scenario of the events to be optimized and the data to be processed. The quality management events of the data to be processed are determined based on the first quality management coefficient, the second quality management coefficient, and the quality control mapping relationship.
10. A data quality management system applied to big data systems, characterized in that, The data quality management system applied to the big data system is applied to the data quality management method applied to the big data system as described in any one of claims 1-9, and the data quality management system applied to the big data system includes: The Advantage Data Module is used to identify multiple data spaces based on the detection of the big data system, and to determine multiple advantageous data based on the data volume, spatial form and current state of the big data system. The data quality system module is used to determine the concentrated area of advantageous data based on the distribution location of multiple advantageous data, and to determine the data quality system based on the concentrated area of advantageous data and the past data optimization events of the big data system. The quality level module is used to divide the data to be processed in this data quality system to output multiple sub-data. The quality level of each sub-data is determined according to its data position, priority and quality score. The quality level of the data to be processed is determined based on the quality level of each sub-data. The data management event module is used to determine the data collection path based on the tracing of the data to be processed if the quality level of the data to be processed is lower than the preset quality level, and to match the corresponding data management events to each data management node of the data collection path. The quality management event module is used to determine the events to be optimized based on the data control events of each data control node, and to determine the quality management events of the data to be processed based on the data type and application scenario of the events to be optimized and the data to be processed.
Citation Information
Patent Citations
Intelligent guiding system and guiding method for digital exhibition hall
CN117115402A
Cloud-based steel structure big data management system and method
CN120337065A