Data quality evaluation method and device, electronic equipment and computer storage medium

By extracting the full amount of data from the data source within the time window and combining the data field richness and business value scoring, the problem of inaccurate and incomplete data quality assessment in the existing technology is solved, the comprehensiveness and accuracy of data quality assessment is achieved, and real-time quality feedback is provided.

CN120672208APending Publication Date: 2025-09-19SANGFOR TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510828822.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies fail to fully consider the business value of data in data quality assessment, resulting in inaccurate and incomplete assessment results.

Method used

By extracting the full data of the target data source within the preset time window, combined with the data field richness and business value score, a comprehensive weighted assessment of data quality is performed, including the data field richness quality score, business value quality score and data resolution rate score.

Benefits of technology

It achieves comprehensiveness and accuracy in data quality assessment, can reflect the comprehensive assessment results of data in the process dimension and business value dimension in real time, and provide accurate quality feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672208A_ABST
    Figure CN120672208A_ABST
Patent Text Reader

Abstract

The invention provides a security service-oriented data quality evaluation method and apparatus, an electronic device and a computer storage medium, and relates to the technical field of computers, and the method comprises the steps of dynamically extracting total data reported by a target data source within a preset time period by using a preset time evaluation window; according to the data quality evaluation method, the whole data quality evaluation process is established, the reported total data is comprehensively evaluated from the process dimension and the service value dimension, the whole data quality evaluation process is high in real-time performance, the evaluation dimension is also comprehensive, and finally the purpose of obtaining an accurate and comprehensive data quality evaluation result is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data quality assessment method, device, electronic device, and computer storage medium for security services. Background Art

[0002] When conducting data quality assessments, vendors in traditional big data services, SIEM (Security Information and Event Management), SOC (Security Operations Center), and other fields currently focus on evaluating whether the data source reporting process meets expectations based on the four evaluation dimensions pre-defined by the vendor: data accuracy, data completeness, data timeliness, and data compliance.

[0003] Current data quality assessment methods focus more on the reliability and privacy protection of data transmission or storage processes, and rarely or not at all consider the business value of data. As a result, the final data quality assessment results are not completely accurate and comprehensive. Summary of the Invention

[0004] In view of this, the present application provides a data quality assessment method, device, electronic device and computer storage medium for security business to solve the problem that the data quality assessment results obtained by the data quality assessment methods in the prior art are not completely accurate and comprehensive.

[0005] To solve the above problems, this application provides the following technical solutions:

[0006] A first aspect of an embodiment of the present application provides a data quality assessment method, the method comprising:

[0007] When the time evaluation window is reached, the full amount of data reported by the target data source within the preset time period is extracted; the sliding step size of the time evaluation window is set according to the data quality evaluation requirements of the business involved in the full amount of data;

[0008] Determining a predefined full data field corresponding to the full data, and determining a data field richness quality score based on the data field in the full data and the predefined full data field;

[0009] Determine the business scenarios involved in the full data, evaluate the data fields in the full data based on the field requirements corresponding to the business scenarios, and obtain a business value quality score;

[0010] The data richness quality score and the business value quality score are comprehensively weighted to obtain a data quality score for data quality assessment within the time assessment window.

[0011] Preferably, it also includes:

[0012] Counting the data resolution rate of the target data source within a preset time period, and determining a data resolution rate score corresponding to the data resolution rate based on a correspondence between the preset data resolution rate and the data resolution rate dimension score;

[0013] After obtaining the data richness quality score and the business value quality score, the data richness quality score, the business value quality score and the data resolution rate score are comprehensively weighted to obtain a data quality score for data quality evaluation within the time evaluation window.

[0014] Preferably, the data richness quality score, the business value quality score, and the data resolution rate score are comprehensively weighted to obtain the data quality score within the time evaluation window, including:

[0015] Determining a first weight corresponding to the data richness quality score, a second weight corresponding to the business value quality score, and a third weight corresponding to the data resolution rate score; wherein the sum of the first weight, the second weight, and the third weight is 1;

[0016] Based on the first weight, the second weight and the third weight, the data richness quality score, the business value quality score and the data resolution rate score are weighted and summed respectively to obtain a data quality score for data quality evaluation within the time evaluation window.

[0017] Preferably, determining the data field richness quality score based on the data fields in the full data and the predefined full data fields includes:

[0018] Comparing the semantics of the data fields in the full data with the semantics of the corresponding predefined full data fields for completeness, and assigning a value to the full data based on the comparison result, and using the value as the data field richness quality score;

[0019] Alternatively, the proportion of data fields in the full data that conform to the predefined full data fields is calculated, and based on the correspondence between the preset proportion and the data field richness quality score, the proportion is converted into a corresponding data field richness quality score.

[0020] Preferably, the determining of the data field richness quality score based on the data fields in the full data and the predefined full data fields includes:

[0021] Counting the number of data fields in the full amount of data that meet a predefined specification, where the predefined specification is a field that has a value and is not empty;

[0022] Based on the total number of predefined full data fields and the number of data fields, a ratio that conforms to the predefined full data fields is obtained;

[0023] The ratio is rounded down, and the resulting value is used as the data field richness quality score of the full amount of data reported within the preset time period.

[0024] Preferably, the field requirements corresponding to the business scenario are evaluated on the data fields in the full amount of data to obtain a business value quality score, including:

[0025] If the business scenario includes multiple sub-business scenarios, the data fields in the full data are evaluated for the fields corresponding to each sub-business scenario to obtain the business data quality score of the sub-business scenario;

[0026] The quality score of each of the business data is counted, and the obtained sum is used as the business value quality score of the full amount of data reported within the preset time period.

[0027] Preferably, it also includes:

[0028] Recording the data quality scores obtained within the time evaluation window in a data quality table, wherein the data quality table records the data quality scores of data quality evaluations performed within multiple time evaluation windows sliding with the sliding step size;

[0029] Dynamically display the data quality scores recorded in the data quality table.

[0030] A second aspect of an embodiment of the present application provides a data quality assessment device, the data quality assessment device comprising:

[0031] An extraction module is configured to extract the full amount of data reported by the target data source within a preset time period when a time evaluation window is reached; the sliding step size of the time evaluation window is set according to the data quality assessment requirements of the business indicated by the full amount of data;

[0032] a data field richness quality scoring module, configured to determine a predefined full data field corresponding to the full data, and determine a data field richness quality score based on the data field in the full data and the predefined full data field;

[0033] A business value quality scoring module is used to determine the business scenarios involved in the full data, evaluate the data fields in the full data according to the field requirements corresponding to the business scenarios, and obtain a business value quality score;

[0034] The comprehensive weighted scoring module is used to comprehensively weight the data richness quality score and the business value quality score to obtain a data quality score for data quality evaluation within the time evaluation window.

[0035] A third aspect of the embodiments of the present application provides an electronic device, including:

[0036] one or more processors;

[0037] a storage device having one or more programs stored thereon;

[0038] When the one or more programs are executed by the one or more processors, the one or more processors implement the data quality assessment method provided in the first aspect of the embodiment of the present application.

[0039] A fourth aspect of an embodiment of the present application provides a computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the data quality assessment method provided in the first aspect of the embodiment of the present application is implemented.

[0040] It can be seen from the above technical solutions that the data quality assessment method, device, electronic device and computer storage medium for security business provided by this application use a pre-set time assessment window to dynamically extract the full amount of data reported by the target data source within a preset time period, and conduct a comprehensive assessment of the reported full amount of data from the process dimension and business value dimension. The entire data quality assessment process is not only highly real-time, but also has comprehensive assessment dimensions, ultimately achieving the purpose of obtaining accurate and comprehensive data quality assessment results. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0042] Figure 1 This is a system architecture diagram for data quality assessment disclosed in an embodiment of the present application;

[0043] Figure 2 A schematic diagram of all fields defined by the field benchmark disclosed in the embodiments of this application;

[0044] Figure 3 The structure disclosed in the embodiment of this application Figure 2 Schematic diagram of the four sub-parts;

[0045] Figure 4 A schematic diagram of a reported data situation disclosed in an embodiment of the present application;

[0046] Figure 5 A schematic diagram of another reported data situation disclosed in an embodiment of the present application;

[0047] Figure 6 A flow chart of a data quality assessment method disclosed in an embodiment of the present application;

[0048] Figure 7 A flow chart of another data quality assessment method disclosed in an embodiment of the present application;

[0049] Figure 8 This is a schematic diagram of analyzing the richness index of data fields based on the reporting of specific fields in log data disclosed in an embodiment of the present application;

[0050] Figure 9 A visualization diagram of the relationship between data quality and business value of the data fields disclosed in the embodiments of this application;

[0051] Figure 10 A schematic diagram of the structure of a data quality assessment device disclosed in an embodiment of the present application;

[0052] Figure 11 This is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0055] Furthermore, in this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0056] The professional terms involved in this application are as follows:

[0057] Security Operations Platform (SOP): A SOP is an integrated technology framework designed to help organizations monitor, detect, and respond to cybersecurity threats. It integrates various security tools and data sources to provide real-time security situational awareness and incident response capabilities. This platform typically includes security information and event management (SIEM), threat intelligence, automated response, and analytics capabilities, helping security teams improve efficiency and quickly identify and respond to potential security incidents, thereby protecting the organization's digital assets and information security.

[0058] Security Detection Platform (SDP): An integrated system for identifying, assessing, and managing security risks within information systems. It leverages automated tools and technologies to conduct real-time monitoring and vulnerability scanning of networks, applications, and devices, ensuring compliance with security standards and policies. The platform generates security reports and provides risk analysis, helping enterprises promptly identify and remediate security risks and enhance overall security protection. Through centralized management and intelligent analysis, the SDP provides organizations with effective security assurance, mitigating the risk of potential cyberattacks and data breaches.

[0059] Extended Detection and Response (XDR): This platform integrates SIEM log access, a security detection engine, a security operations module, and coordinated response. It generates security alerts based on secondary analysis of access logs. Based on the security capabilities of reported security devices, it further identifies potential threats and collaborates with response devices to complete the detection and response loop. This brings semi-automated operational capabilities to network security within the covered environment, saving manpower while improving the quality of network security. XDR is generally considered a mainstream product implementation of security operations platforms and security detection platforms.

[0060] Chief Security Officer (CSO): The CSO is responsible for coordinating the development and implementation of security strategies, covering network, information, physical and other aspects of security protection, responding to various security threats and risks, and ensuring the safe and stable operation of the organization.

[0061] Data Governance Module (Security Information and Event Management, SIEM): This SIEM accesses various external raw logs and outputs logs in a unified standard format defined within the platform. Organizations with network security construction requirements usually purchase security devices from multiple manufacturers and categories, and security operations personnel find it difficult to frequently check the monitored security logs on each device and analyze the security events they detect. SIEM relies on accessing logs from various devices to the same platform to build a centralized operating environment for customers, saving manpower and improving network security operation efficiency.

[0062] Security Operations Center (SOC): A SOC is a department dedicated to enterprise information security, with the primary responsibility of monitoring, detecting, and responding to security incidents within enterprise networks and systems. An SOC is typically comprised of a team of professional security analysts and engineers who use a variety of security tools and technologies to protect an enterprise's information assets, including real-time monitoring of network traffic, analysis of security logs, and detection of malware and cyberattacks. SIEMs tend to be technical solutions for data collection, analysis, and storage of security data from diverse sources, while SOCs tend to be centralized teams and facilities for monitoring, detecting, and responding to security incidents. As a tool for data collection and analysis, SIEMs provide the SOC with the necessary information support, while the SOC uses this information for real-time monitoring and response. The effective combination of the two can significantly enhance an organization's security capabilities.

[0063] GDPR (General Data Protection Regulation): This law is a data protection regulation that came into effect on May 25, 2018, in the European Union. It aims to strengthen and unify the protection of personal data. GDPR requires data processors to obtain explicit consent when collecting and processing personal data and to ensure the security and privacy of that data.

[0064] HIPAA (Health Insurance Portability and Accountability Act): This act is a U.S. law passed in 1996 to protect the privacy and security of personal health information. The act requires these entities to take appropriate security measures to prevent unauthorized access and disclosure.

[0065] Data baseline: A specific data standard or specification that defines the complete set of fields required for data, including their name, type, format, source, priority, mandatory field status, and examples. Defining a clear data baseline within a business system ensures that business operations avoid conflicts and diverging understandings caused by inconsistent data standards when using data.

[0066] As can be seen from the background technology, the current data quality assessment method mainly starts from the four evaluation dimensions of data accuracy, data integrity, data timeliness and data compliance. The data quality assessment process focuses more on the reliability and privacy protection of data transmission or storage processes, and does not consider or rarely considers the business value of data. As a result, the final data quality assessment is not completely accurate and comprehensive.

[0067] For example, the data returned by an upstream component fully complies with data standards and specifications, including keys for all fields, but all corresponding values ​​are filled with default values. Based on existing technology, all keys have corresponding values, which, from a traditional process perspective, fully complies with specifications. However, the default values ​​(generally because actual values ​​are not available in the production environment) are, to a certain extent, useless data with low or no value to the business, and fail to reflect the content of the data fields required by the business.

[0068] For example, in a security analysis scenario involving qualitative threat analysis based on security alerts, analysis and judgment is performed based on the attack results, HTTP request headers, and HTTP response header fields provided by network security logs. Suppose both data sources A and B detect network threat behavior and report a series of security logs. Data source A reports all attack results as "Unknown." The HTTP request and response headers only contain the first line, not the complete HTTP protocol header content, but an empty string placeholder. However, the security logs reported by data source B contain not only "Unknown" attack results, but also log data indicating both successful and failed attacks.

[0069] In this security analysis scenario for security alert threat characterization, the data reported by Data Sources A and B were evaluated based on four evaluation dimensions: accuracy, completeness, timeliness, and compliance. The evaluation results showed no significant differences between the data reported by Data Sources A and B in terms of accuracy, completeness, timeliness, and compliance. However, from the perspective of attack results, Data Source B's field richness far surpasses that of Data Source A. Data Source B reports complete values ​​for both the HTTP request and response header fields, including all header fields specified by the HTTP protocol. This means that, considering the data fields' support for actual business operations (e.g., alert threat characterization), Data Source B's data quality assessment results in the business dimension are significantly better than those of Data Source A. Overall, Data Source B's overall data quality is superior to that of Data Source A.

[0070] Based on this, this application discloses a new data quality assessment scheme, especially a business-oriented data quality assessment method, starting from at least the above-mentioned process dimensions (data accuracy, data integrity, data timeliness and data compliance) and business value dimensions (business adaptability, data completeness and dynamic responsiveness).

[0071] In the specific data quality assessment process, the data is evaluated to see whether it meets the predefined specifications in the process dimension. That is, during the upstream and downstream processing and flow of data, all predefined data field content, data format, data transmission protocol, and data resolution should be checked to see whether they meet the expected results. Specifically:

[0072] Data accuracy: The data recipient strictly standardizes the data according to predefined data fields and formats, maps the original data to the data form, sets unified field names and field data types, and verifies the accuracy of the data by comparing it with known data benchmarks.

[0073] Data integrity: This involves collecting log and event data from various data sources (such as network devices, servers, and applications) and ensuring that the data has not been tampered with during transmission through technologies such as checksums and hashes. Data integrity is considered complete if its format and content are consistent at both the input and output ends, and vice versa. By identifying and verifying data integrity, the integrity and reliability of data transmitted downstream can be guaranteed.

[0074] Data timeliness: This provides data source monitoring dashboards and reports to help users monitor data quality indicators such as data loss rate and latency. Alerts are triggered when data source indicators such as data loss rate or latency exceed preset standards.

[0075] Data compliance: This means ensuring that data processing complies with industry standards and regulatory requirements, such as GDPR and HIPAA, to maintain the legality and compliance of data.

[0076] Evaluate whether the data meets the actual business needs in terms of business value, and fully consider the dependence of downstream businesses on the data fields and content of upstream businesses.

[0077] The business value dimension can be broken down into:

[0078] Business adaptability: refers to the coverage of scenario-specific fields, the matching degree of business rules, the key indicator mapping rate, etc., which is used to evaluate the support capabilities of data fields for specific business scenarios.

[0079] Data completeness: refers to the coverage of all business scenarios (normal and abnormal) patterns, the representativeness of data distribution, etc., which is used to evaluate the sample diversity of data in business scenarios.

[0080] Dynamic responsiveness: refers to the timeliness of a field, state change latency, and context relevance, and is used to assess the timeliness of synchronization between data changes and business dynamics.

[0081] Based on this, this application combines specific business value dimensions and process dimensions to conduct hierarchical dynamic evaluation, and conducts real-time data quality evaluation from both the process dimension and the business value dimension of the data, thereby obtaining completely accurate and comprehensive data evaluation results.

[0082] like Figure 1 FIG2 is a diagram showing a system architecture of a data quality assessment system disclosed in an embodiment of the present application. The system architecture includes a quality assessment platform 1 and a third-party device 2.

[0083] The quality assessment platform 1 includes at least an assessment module 11 , a SIEM and a visualization tool 12 .

[0084] It should be noted that the quality assessment platform 1 can be any platform with data quality requirements, preferably a security operations platform, a security analysis platform, a security testing platform, an application analysis platform, and XDR. It can be applied to all application scenarios with data quality requirements, as well as other production analysis application scenarios where data quality can actually affect business analysis.

[0085] The third-party device 2 at least includes a data source 21 for reporting data.

[0086] The third-party device 2 is connected to the quality assessment platform 1 and reports data required for data quality assessment to the quality assessment platform 1 through the data source 21 , and the data includes but is not limited to logs and event data.

[0087] The quality assessment platform 1 dynamically assesses the quality of the reported data based on the process dimension and the business value dimension. Specifically:

[0088] The quality assessment platform 1 sets a sliding time assessment window, dynamically extracts data reported by the third-party device 2 within a preset time period based on the time assessment window, performs corresponding assessment, and dynamically displays the assessment results through a visualization tool.

[0089] For example, the sliding step of the time evaluation window is 10 minutes, and the preset time period is 24 hours. The quality evaluation platform 1 extracts the data reported by the data source 21 in the past 24 hours every 10 minutes for corresponding evaluation, and dynamically displays the evaluation results through a visualization tool.

[0090] It should be noted that the sliding step size and window duration of the time evaluation window can be adjusted according to specific data quality evaluation requirements.

[0091] When sliding to the time evaluation window, the evaluation module 11 in the quality evaluation platform 1 performs data quality evaluation in the time evaluation window. The main process includes:

[0092] 1. Extracting data: In the time evaluation window, extract the data reported by the data source 111 within a preset time period; wherein, the data reported by the data source 21 can be analyzed to determine the business type indicated by it.

[0093] 2. Data field richness dimension assessment: Based on the predefined full data fields that the data needs to contain, calculate the ratio of the data field richness in the reported data to the predefined full data fields, and convert this ratio into a data field richness quality score that represents the process dimension assessment results.

[0094] Specifically, data field richness refers to the proportion of fields with values ​​and non-empty values ​​in all the data reported by data source 21 during the preset time period within the actual time evaluation window, accounting for all the fields in the data benchmark. The closer this ratio is to 100%, the higher the data richness and the higher the evaluation value, that is, all predefined data fields are reported, and default values ​​are not filled in to meet data standards. Conversely, the further this ratio is from 100%, the lower the data richness and the lower the evaluation value, that is, no predefined data fields are reported, or default values ​​may be filled in to meet data standards.

[0095] 3. Data business value dimension assessment: Determine the business scenarios involved in the data, and evaluate the data fields in the reported data based on the fields that each business scenario relies on. That is, evaluate the quality of the above data from the business data usage dimension to obtain a business value quality score.

[0096] For example, in a security analysis scenario involving qualitative threat analysis of security alerts, the fields relied upon include attack results, HTTP request headers, and HTTP response headers provided by network security logs. By analyzing these fields, we can assess data quality in terms of the business value of the data in this security analysis scenario.

[0097] Specifically, the data business value dimension refers to the data quality assessment from the dimension of fields used by several businesses.

[0098] In this application, the data field richness dimension is evaluated as a surface, and the data business value dimension is evaluated as a region on the surface. Figures 2 to 5 As shown, it is assumed that all fields defined by the field base are as follows Figure 2 As shown, marked as B. This B consists of four sub-parts (such as Figure 3 The square, circle, parallelogram, and rhombus shown represent the entire data field of a specific business, and it is assumed that the business value of each part is equal to the weight.

[0099] Figure 4 and Figure 5 Represent two actual reported data situations. Figure 4 The area marked C in the figure represents the reported fields, and the reported area accounts for exactly 50%, while the rest represents the unreported fields. Figure 5 The D area marked in the figure represents the reported field, and the area of ​​the reported part accounts for exactly 50%, and the rest represents the unreported part.

[0100] Execute 2 and evaluate scenarios C and D from the "Data Field Richness Dimension". The resulting data quality is the same.

[0101] Execution 3 evaluates scenarios C and D from the "Data Business Value Dimension," resulting in different data quality. In scenario C, 3 / 4 (three sub-parts) of the business-required fields are reported, while in scenario D, only 1 / 4 (one sub-part) of the business-required fields are reported. The data quality assessment results show that scenario C has better data quality than scenario D.

[0102] In other words, conducting data quality assessment from different dimensions will obtain targeted assessment results, which can reflect the real situation more comprehensively and accurately.

[0103] 4. Data Resolution Rate Statistics: Calculates the data resolution rate of data reported by data source 111 for a preset time period and converts it into a data resolution rate score. This verifies whether there are any temporary data reports or network transmission anomalies that result in data non-compliance with data specifications. Specifically, the data resolution rate is obtained through SIEM.

[0104] 5. Comprehensive weighted scoring: The data field enrichment quality score, business value quality score, and data resolution rate score obtained in 2, 3, and 4 are comprehensively weighted to obtain the final data quality score of the third-party device 11 in the time evaluation window.

[0105] 6. Data quality visualization tabulation: The data quality score (including the evaluation results of the process dimension and the business dimension) obtained by the third-party device 11 in each time evaluation window is recorded in the data quality table, and displayed in real time on the data quality dynamic monitoring screen using the visualization tool 102, so that the CSO to which the third-party device 11 belongs can accurately see the real-time overview and changes of the data quality of its own business system.

[0106] When the next time evaluation window arrives, steps 1 to 6 above are continued. A dynamic data quality score for the third-party device 11 is obtained for each time evaluation window. The data quality score obtained for each time evaluation window is recorded in the data quality table. If the sliding step size of the time evaluation window is 10 minutes, the data quality score corresponding to the 10th minute, the 20th minute, the 30th minute, and so on is recorded in the data quality table.

[0107] By visualizing the data quality table (by plotting points as lines) using visualization tool 102, the data quality changes of third-party devices 11 currently connected to security operation platform 10 can be dynamically visualized. Furthermore, visualization tool 102 can also be used to display the visual relationships and chart interpretation of data fields during the evaluation process.

[0108] Based on the system architecture of data quality assessment disclosed in this application, a pre-set time assessment window is used to analyze the data reported by the access device in real time, and the duration of the time assessment window is dynamically adjusted according to actual needs, and the data classification standard is dynamically adjusted to ensure that the data quality assessment can adapt to changes in different time periods and environments, and provide more accurate quality feedback. In the specific data quality assessment process, both the process dimension and the business value dimension are taken into account, and a comprehensive assessment is conducted on the data reported by the access device from multiple dimensions to ensure the comprehensiveness and reliability of the data quality. The data quality assessment results are displayed through visualization tools, so that users can intuitively understand the current status and changing trends of the device data quality, and facilitate timely improvement measures to improve data quality. In particular, by visually displaying the real-time data quality assessment results, when data quality anomalies occur, data quality changes can be discovered in real time, which is convenient for diagnosis and giving possible causes and improvement suggestions, so that the CSO can quickly follow up and investigate and locate them.

[0109] like Figure 6The figure is a flow chart of a data quality assessment method disclosed in an embodiment of the present application, which is applicable to the above Figure 1 The quality assessment platform 10 in the public data quality assessment system architecture can be applied to all application scenarios with high data quality requirements, as well as other production analysis scenarios where data quality can actually affect business analysis, including but not limited to various security analysis, application analysis, and generated data analysis scenarios that rely on the richness and quality of raw log field content.

[0110] The data quality assessment method includes:

[0111] S601: When the time evaluation window is reached, extract all data reported by the target data source within a preset time period.

[0112] In S601, the sliding step size of the time evaluation window is set according to the data quality evaluation requirements of the business indicated by the full amount of data. The full amount of data is specifically logs.

[0113] In one embodiment of the present application, the sliding step size ranges from 5 to 20 minutes. Preferably, when data quality assessment is required for security services, the sliding step size can be set to 10 minutes.

[0114] It should be noted that the sliding step of the time evaluation window can be dynamically adjusted according to the actual business situation and is not only linked to the business type. That is, when it comes to data quality assessment needs for security-oriented business, the sliding step can be set to 10 minutes, but based on the current actual situation of the security-oriented business, the sliding step can also be set to 15 minutes.

[0115] In the specific process of executing S601 , every time the time evaluation window is slid to, the full amount of data reported by the currently connected device (target data source) within the preset time period is extracted based on the time evaluation window.

[0116] It should be noted that this application performs evaluation based on a time evaluation window. The extracted full data is based on the reported time axis. The full data changes dynamically. Therefore, the full data can be dynamically presented based on the time evaluation window.

[0117] S602: Determine the predefined full data fields corresponding to the full data, and determine a data field richness quality score based on the data fields in the full data and the predefined full data fields.

[0118] During the specific execution of S602, the data type corresponding to the reported full data is determined based on the analysis, and the predefined full data fields required for the data type are determined. A richness assessment is performed based on the data fields in the full data and the predefined full data fields, and a data field richness quality score is determined based on the assessment results.

[0119] For example, if the full data is parsed as a full log of a network security log, then predefined full data fields required for the full log of the network security log are determined.

[0120] It should be noted that the predefined full data fields required under this data type can be increased or decreased according to the actual problem scenarios, industry standards, and evaluation rules.

[0121] In the present application, one or more calculation rules may be provided. Optionally, each calculation rule is adapted to correspond to one or more data types. After determining the corresponding data type based on the parsed and reported full data and determining the predefined full data fields required under the data type, when determining the data field richness quality score, the calculation rule adapted to the current response data type is called to determine the data field richness quality score based on the data fields in the full data and the predefined full data fields.

[0122] In one embodiment of the present application, the process of determining the data field richness quality score based on the data fields in the full data and the predefined full data fields includes:

[0123] The semantics of the data fields in the full data are compared for completeness with the semantics of the corresponding predefined full data fields, and a value is assigned to the full data based on the comparison result, with the value serving as the data field richness quality score. The comparison result is a difference value, and a pre-established correspondence between the difference value and the data field richness quality score is used. Based on the currently obtained difference value and the correspondence between the difference value and the data field richness quality score, a value is assigned to the full data to obtain a data field richness quality score for the full data reported within a preset time period.

[0124] For example, the semantic integrity of a field value is used as a criterion for determining field richness. The semantic integrity of ruleName: Generic OS Command Injection Attempt in HTTP POST Parameter Windows is better than that of ruleName: CMD Injection. Therefore, the data quality score for this type of data is higher.

[0125] In one embodiment of the present application, the process of determining the data field richness quality score based on the data fields in the full data and the predefined full data fields includes:

[0126] Calculate the proportion of data fields in the full data that conform to the predefined full data fields, and determine the data field richness quality score based on the proportion. Specifically:

[0127] Calculate the ratio of the richness of the data fields in the full data to the predefined full data fields, and convert the ratio into the corresponding data field richness quality score based on the preset correspondence between the ratio and the data field richness quality score.

[0128] In one embodiment of the present application, the process of determining the data field richness quality score based on the data fields in the full data and the predefined full data fields includes:

[0129] S10: Count the number of data fields in the full amount of data that meet the predefined specifications.

[0130] The predefined specification is a field that has a value and is not empty; S1 is executed to count the number of data fields that meet the requirements of having a value and being non-empty in the full data.

[0131] S11: Based on the total number of predefined full data fields and the number of the data fields, a ratio that meets the predefined full data fields is obtained.

[0132] For example, the total number of predefined full data fields is 42, and 30 data fields that meet the predefined specifications have been reported. The proportion of predefined full data fields that meet the predefined specifications is: 30 / 42=71.4%.

[0133] S12: Rounding down the ratio, and using the obtained value as the data field richness quality score of the full amount of data reported within the preset time period.

[0134] Execute S12 to round down the ratio obtained by executing S11, and use the obtained value as the data field richness quality score of the full amount of data reported within the preset time period.

[0135] The principle of rounding down is: ratio*10.

[0136] Based on the above example, the ratio 71.4% is rounded down to: 71.4%*10=7 (rounded down).

[0137] S603: Determine the business scenarios involved in the full data, evaluate the data fields in the full data according to the field requirements corresponding to the business scenarios, and obtain a business value quality score.

[0138] During the specific execution of S603, the specific business scenario to which the full data is applied, as well as the sub-business scenarios contained in the business scenario, are determined. Based on the field requirements required for the business scenario (excluding the sub-business scenario) or the sub-business scenario, the data fields in the full data are evaluated from the business data usage dimension to obtain the final business value quality score.

[0139] If a business scenario contains several sub-business scenarios, the number of the sub-business scenarios can also be adjusted according to actual conditions.

[0140] It should be noted that each business scenario or sub-business scenario has different requirements for data fields, that is, different requirements for the fields to be utilized. Based on this, in one embodiment of the present application, a unified scoring rule is set for each business scenario or sub-business scenario with different requirements for data fields when evaluating the business value dimension. According to the set scoring rule, the fields that meet the requirements are assigned values, and the corresponding business value quality score is finally obtained. Alternatively, the assigned values ​​in all sub-business scenarios are summed up to obtain the corresponding business value quality score.

[0141] Preferably, the scoring rules include but are not limited to:

[0142] 1. If all basic fields of a dimension are met and more than half of the multiple evidence scenarios meet the field requirements, 1 point will be awarded.

[0143] 2. If all basic fields of a dimension are met but more than half of the multiple evidence scenarios do not meet the field requirements, 0.5 points will be awarded.

[0144] In one embodiment of the present application, taking a business scenario including multiple sub-business scenarios as an example, the process of evaluating the data fields in the full data according to the field requirements corresponding to the business scenario and obtaining the business value quality score includes:

[0145] S20: Evaluate the data fields in the full data according to the field requirements corresponding to each sub-business scenario to obtain the business data quality score of the sub-business scenario.

[0146] Execute S20 to score the corresponding data fields in the full data for each sub-business scenario based on the preset scoring rules to obtain the business data quality score of the full data in each sub-business scenario.

[0147] S21: Count the quality scores of each business data, and use the obtained sum as the business value quality score of the full amount of data reported within a preset time period.

[0148] Execute S21 to accumulate the quality score of each business data, and use the obtained sum as the business value quality score of the full amount of data reported within the preset time period.

[0149] S604: Comprehensively weight the data richness quality score and the business value quality score to obtain a data quality score for data quality assessment within the time assessment window.

[0150] In the specific process of executing S604, the data richness quality score and the business value quality score are comprehensively weighted according to the preset weights to obtain the data quality score for data quality evaluation within the time evaluation window.

[0151] In one embodiment of the present application, the data quality score for data quality assessment within the time assessment window is obtained by comprehensively weighting the data richness quality score and the business value quality score, including:

[0152] S30: Determine a first weight corresponding to the data richness quality score and a second weight corresponding to the business value quality score.

[0153] The sum of the first weight and the second weight is 1. Different weights can be set according to different data quality assessment requirements.

[0154] For example, in a security application scenario, considering that the business scenario has a greater impact on data quality, the first weight is set to 20% and the second weight is set to 80%.

[0155] S31: Based on the first weight and the second weight, respectively perform weighted summation on the data richness quality score and the business value quality score to obtain a data quality score for data quality assessment within the time assessment window.

[0156] Based on the example of S30 above, if the data richness quality score is 5 points and the business value quality score is 7 points, execute S31 for weighted summation, and the final data quality score for data quality assessment within the time assessment window is = 5*20%+7*80%=6.6 points.

[0157] It should be noted that, in the process of executing the above S601 to S604, a visualization tool can be used to analyze the richness dimension of the data field based on the reporting of specific fields in the full data, as well as to visualize the relationship between the data quality and business value of the data field.

[0158] Based on the data quality assessment method disclosed in the embodiment of this application, a pre-set time assessment window is used to dynamically extract the full amount of data reported by the target data source within a preset time period, and a comprehensive assessment of the reported full amount of data is performed from the process dimension and the business value dimension. The entire data quality assessment process is not only highly real-time, but also has a comprehensive assessment dimension, ultimately achieving the goal of obtaining accurate and comprehensive data quality assessment results. Furthermore, the duration of the time assessment window can be dynamically adjusted according to actual needs, and the data classification standards can be dynamically adjusted to ensure that the data quality assessment can adapt to changes in different time periods and environments, and provide more accurate quality feedback.

[0159] like Figure 7 FIG. 1 is a flow chart of another data quality assessment method disclosed in an embodiment of the present application, and FIG. Figure 6 The difference of the disclosed data quality assessment method is that it adds a data analysis dimension.

[0160] The data quality assessment method includes:

[0161] S701: When the time evaluation window is reached, extract the full amount of data reported by the target data source within the preset time period.

[0162] S702: Determine the predefined full data fields corresponding to the full data, and determine a data field richness quality score based on the data fields in the full data and the predefined full data fields.

[0163] S703: Determine the business scenarios involved in the full data, evaluate the data fields in the full data according to the field requirements corresponding to the business scenarios, and obtain a business value quality score.

[0164] The specific execution process and principle of S701 to S702 are consistent with those of S601 to S603 above, and will not be repeated here.

[0165] S704: Counting the data resolution rate of the target data source within a preset time period, and determining a data resolution rate score corresponding to the data resolution rate based on a correspondence between the preset data resolution rate and the data resolution rate dimension score.

[0166] In S704 , the corresponding relationship between the data resolution rate and the data resolution rate dimension score is as follows: the closer the data resolution rate is to 100%, the closer the data resolution rate dimension score is to the full score.

[0167] Preferably, the maximum score for the data resolution dimension is 10 points.

[0168] S705: Comprehensively weight the data richness quality score, the business value quality score, and the data resolution rate score to obtain a data quality score for data quality assessment within the time assessment window.

[0169] In the specific execution of S705, the data richness quality score, the business value quality score and the data resolution rate score are comprehensively weighted according to the preset weights to obtain the data quality score for data quality evaluation within the time evaluation window.

[0170] In one embodiment of the present application, the process of comprehensively weighting the data richness quality score, the business value quality score, and the data resolution rate score to obtain the data quality score within the time evaluation window includes:

[0171] S40: Determine a first weight corresponding to the data richness quality score, a second weight corresponding to the business value quality score, and a third weight corresponding to the data resolution rate score.

[0172] The sum of the first weight, the second weight, and the third weight is 1; the first weight, the second weight, and the third weight can be set according to different data quality assessment requirements.

[0173] For example, in a security application scenario, considering that the business scenario and data resolution rate have a greater impact on data quality, the first weight is set to 20%, the second weight is set to 50%, and the third weight is set to 30%.

[0174] S41: Based on the first weight, the second weight and the third weight, the data richness quality score, the business value quality score and the data resolution rate score are weighted and summed respectively to obtain the data quality score for data quality evaluation within the time evaluation window.

[0175] Based on the example of S30 above, if the data richness quality score is 7 points, the business value quality score is 7 points, and the data resolution rate score is 9 points, execute S41 for weighted summation, and the final data quality score for data quality assessment within the time assessment window is = 7*20%+7*80%+9*30%=9.7 points.

[0176] It should be noted that, in the process of executing the above S701 to S705, a visualization tool can be used to analyze the richness dimension of the data field based on the reporting of specific fields in the full data, as well as to visualize the relationship between the data quality and business value of the data field.

[0177] Based on the data quality assessment method disclosed in the embodiment of this application, a pre-set time assessment window is used to dynamically extract the full amount of data reported by the target data source within a preset time period, and a comprehensive assessment is performed on the reported full amount of data from multiple dimensions such as the process dimension, business value dimension, and data analysis dimension. The entire data quality assessment process is not only highly real-time, but also has comprehensive assessment dimensions, ultimately achieving the goal of obtaining accurate and comprehensive data quality assessment results. Furthermore, the duration of the time assessment window can be dynamically adjusted according to actual needs, and the data classification standards can be dynamically adjusted to ensure that the data quality assessment can adapt to changes in different time periods and environments, and provide more accurate quality feedback.

[0178] In the above Figure 6 and Figure 7 Based on the disclosed data quality assessment method, this application also discloses another data quality assessment method. Figure 6 After S604 shown, or when executing Figure 7 After S705 shown, the following steps are specifically performed:

[0179] S50: Record the data quality score obtained within the time evaluation window in a data quality table.

[0180] In S50 , the data quality table records the data quality scores of the data quality evaluations performed within a plurality of time evaluation windows sliding with a sliding step size.

[0181] In the specific process of executing S50 , after each execution of S604 or S705 , the data quality score obtained for the data quality evaluation within the current time evaluation window is recorded in a preset data quality table.

[0182] S51: Dynamically display the data quality scores recorded in the data quality table.

[0183] During the specific execution of S51 , the data quality scores recorded in the data quality table are plotted and presented, visually and dynamically displaying the changes in data quality after the device is connected.

[0184] At the same time, the data quality assessment results can be analyzed by combining the visualization of the reporting status of specific fields in the full data to analyze the richness dimension of the data fields and the relationship between the data quality and business value of the data fields.

[0185] Based on the data quality assessment method disclosed in the embodiments of this application, during the specific data quality assessment process, the process dimension, business value dimension, and data resolution dimension are taken into account, and a comprehensive assessment is conducted on the data reported by the access device from multiple dimensions. This can ensure the comprehensiveness and reliability of data quality, and display the data quality assessment results through visualization tools, so that users can intuitively understand the current status and changing trends of device data quality, facilitating timely implementation of improvement measures to improve data quality. In particular, by visually displaying real-time data quality assessment results, when data quality anomalies occur, data quality changes can be discovered in real time, facilitating diagnosis and providing possible causes and improvement suggestions, so that the CSO can quickly follow up and investigate and locate the problem.

[0186] In combination with the data quality assessment method disclosed in the above-mentioned embodiment of this application, this application is explained using an actual device accessing network security log data to perform security threat analysis as an application example.

[0187] For example, a cloud WAF device will report the logs detected by the component. When the data quality assessment method disclosed in the above embodiment of the present application is specifically executed, it is confirmed that the cloud WAF device has been correctly connected and the data fields carried in the log have been correctly parsed.

[0188] When the time evaluation window arrives, the extracted reported data is as follows:

[0189] "{"method":"POST","domain":"mss.muxx.com.cn","host":"mss.muxx.com.cn:8043","request_length":1207,"http_x_forwarded_for":"-","msec":1709714974747,"query":"g=obj_app_upfile","upstream_response_time":0,"scheme":"http","request_time":0.004,"http_user_agent":"Mozilla / 5.0 (compatible; MSIE 6.0;Windows NT 5.0; Trident / 4.0)","cookie":"","ip_info":{"city":"Shanghai","province":"Shanghai","state":"CN","nation":"China","operator":"chinatele.com.cn","detail":"chinatele.com.cn","longitude":"121.5447","dimensionality":"31.22249"},"headers":"stgw-orgservername: mss.muxx.com.cn\n stgw-orgreq:POST / ?g=obj_app_upfileHTTP / 1.1\n stgw_request_id: cd3822d842ac5c40b841e998\nwaf-customize-lbid: lb-ewr7b3xn\n accept-encoding:gzip,deflate\n x-waf-uuid:f70e81c1b4ad4bf8d3f8b52469e7779s88\n", "remote_addr":"116.233.74.246","bytes_sent":2307,"content_type":"multipart / form-data; boundary=----WebKitFormBoundary2dJ8bMB66toy58a8xRJCCu0uzZA","language":"-","uuid" :"f70e81c1b4ad4bf8d3f8b52469e77739-6a98201f68f19c6bc128dc163bd9e935","referer":"- ","instance":"waf_2kzd1yfu00g88tbe","upstream_addr":"-","create_time":"2024-03-06 T16:49:34.000000747+08:00","connection":"close","time_local":"06 / Mar / 2024:16:49:34 +0800", "url":" / ", "edition":"clb-waf","upstream_connect_time":0,"encoding":"gzip,deflate", "appid":1315593449, "upstream_status":0,"status":615}".

[0190] S1: Analyzing the data field richness dimension reveals that this log is a network security log. Based on the type of network security log, a full table of predefined data fields is predefined to participate in the data field richness dimension assessment. Table 1 shows the partial fields.

[0191] Table 1:

[0192]

[0193] Among them, enriched fields are enriched from other fields reported in the original log (for example, srcIpTag is enriched by querying asset information using srcIp, indicating whether the source IP is an intranet asset). Except for enriched fields, all other field types are mapped through original log parsing.

[0194] S2: Extract the logs reported by the cloud WAF device in the past 24 hours (preset time period). The logs are parsed into the full log of the network security log. The proportion of data fields with values ​​and non-empty fields contained in the full log to the predefined full data fields is counted and rounded down to obtain the data field richness quality score.

[0195] For example, there are 18 data fields in the reported full log that meet the requirements of having values ​​and being non-empty, and 29 predefined full data fields. The resulting ratio is: 18 / 29 = 62.1%, which is rounded down to obtain the data field richness quality score (rounded down) = 62.1% * 10 = 6.

[0196] like Figure 8 As shown in FIG, the richness index of the data field is analyzed based on the reporting of specific fields in the log data. Figure 8 It mainly includes the reporting status, including reported and unreported status; the function points of each dimension, and whether there is a display area for the corresponding fields of the function points; the log fields of different evidence scenarios, Figure 8 It shows that "The following are HTTP-related evidence information fields" include requestHead (request header), requestBody (request body) and reqMethod (request method) and other fields; "The following are intelligence evidence fields" include tilntelligence (IOC), tiType (threat intelligence type), dinsQueries (DNS query domain name) and other fields; "The following are virus evidence fields" include vinusName (virus name) and other fields; and whether the corresponding fields of the function points are reported, not reported, and no display area is required.

[0197] S3: Identify which fields in the full log need to be used for detection and analysis for specific businesses and applications. That is, determine the business scenarios involved in the full log and the field requirements corresponding to the business scenarios.

[0198] Taking the security analysis scenario of reporting network security logs as an example, this scenario includes several sub-scenarios, such as multi-source alarm fusion, similarity correlation noise reduction, business false alarm characterization, and attack threat characterization. Here, the number of sub-scenarios is set to 10. Because each sub-scenarios has different requirements for log fields in the full log, a unified scoring rule is set here.

[0199] The scoring rules include but are not limited to:

[0200] 1. If all basic fields of a dimension are met and more than half of the multiple evidence scenarios meet the field requirements, 1 point will be awarded.

[0201] 2. If all basic fields of a dimension are met but more than half of the multiple evidence scenarios do not meet the field requirements, 0.5 points will be awarded.

[0202] Combine Figure 8 ,like Figure 9 The following is a visualization of the relationship between data quality and business value of data fields. Figure 9It includes the functional points of each dimension, log fields of different evidence scenarios, and instructions on whether the data fields required by each dimension should be reported in different evidence scenarios ( Figure 9 In the table, √ (bold) indicates that it has been reported; √ (not bold) indicates that it has not been reported; — indicates that it is not required).

[0203] Based on the scoring rules, the data quality scores of all sub-business application scenarios are calculated in sequence, and the sum of the business data quality scores of all sub-business application scenarios is calculated, and the sum is used as the data quality score of the overall business value dimension.

[0204] For example, the data quality scores of the current 10 sub-business application scenarios are distributed as follows: 1, 1, 0.5, 1, 1, 0.5, 0.5, 0.5, 0.5, 1, for a total of 7.5 points. This means that the data quality score for the overall business value dimension is 7.5 points.

[0205] S4: The log resolution rate of the cloud WAF device in the past 24 hours is 100%, which is converted into a data resolution rate score of 10 points.

[0206] S5: Comprehensively weight the data richness quality score, business value quality score, and data resolution score to obtain a data quality score for data quality assessment within the time assessment window.

[0207] In this specific application example, different weights are set for the data richness quality score, business value quality score, and data resolution rate score. Specifically, the data richness quality score is weighted at 20%, the business value quality score is weighted at 70%, and the data resolution rate score is weighted at 10%. The final data quality score for data quality assessment within the time assessment window is 6*0.2+7.5*0.7+10*0.1=7.45 points.

[0208] S6: The data quality score of the data quality assessment performed within the time assessment window is recorded in the data quality table based on the end time of the time assessment window, and the points are drawn into a line. Combined with the data quality scores of the data quality assessment performed in other time assessment windows recorded in the data quality table, the data quality changes after the cloud WAF device is connected to the platform are visualized.

[0209] In the application scenario of using security log data for security threat analysis shown in this application example, data quality assessment takes into account the process dimension, business value dimension and data resolution rate dimension, and comprehensively evaluates the data reported by the access device from multiple dimensions to ensure the comprehensiveness and reliability of data quality. The data quality assessment results are displayed through visualization tools, allowing users to intuitively understand the current status and changing trends of device data quality. At the same time, the relationship between data quality and business value of data fields is displayed, which facilitates the analysis of single items with low data quality scores (including specific sub-business scenarios), thereby obtaining improvement suggestions for the data quality construction of the cloud WAF device.

[0210] In summary, the data quality assessment method proposed in this application has the advantages of strong real-time performance, comprehensive assessment dimensions, and high visibility of the correlation between data quality and business results.

[0211] Based on the data quality assessment device disclosed in the above embodiment of the present application, the embodiment of the present application also discloses a data quality assessment device, such as Figure 10 As shown, the data quality assessment device includes: an extraction module 101, a data field richness quality scoring module 102, a business value quality scoring module 103 and a comprehensive weighted scoring module 104.

[0212] The extraction module 101 is used to extract the full amount of data reported by the target data source within a preset time period when the time evaluation window is reached; the sliding step of the time evaluation window is set according to the data quality evaluation requirements of the business indicated by the full amount of data.

[0213] The data field richness quality scoring module 102 is configured to determine the predefined full data fields corresponding to the full data, and determine the data field richness quality score based on the data fields in the full data and the predefined full data fields.

[0214] The business value quality scoring module 103 is used to determine the business scenarios involved in the full data, evaluate the data fields in the full data according to the field requirements corresponding to the business scenarios, and obtain a business value quality score;

[0215] The comprehensive weighted scoring module 104 is used to comprehensively weight the data richness quality score and the business value quality score to obtain a data quality score for data quality assessment within a time assessment window.

[0216] In one embodiment of the present application, the data quality assessment device further includes: a data resolution scoring module 105 .

[0217] The data resolution rate scoring module 105 is used to count the data resolution rates of the target data source within a preset time period, and determine a data resolution rate score corresponding to the data resolution rate based on a correspondence between the preset data resolution rates and the data resolution rate dimension scores.

[0218] Accordingly, the comprehensive weighted scoring module 104 is used to comprehensively weight the data richness quality score, the business value quality score and the data resolution rate score after obtaining the data richness quality score and the business value quality score, so as to obtain a data quality score for data quality evaluation within a time evaluation window.

[0219] In one embodiment of the present application, the data resolution scoring module 105 is specifically used to determine a first weight corresponding to the data richness quality score, a second weight corresponding to the business value quality score, and a third weight corresponding to the data resolution score; wherein the sum of the first weight, the second weight, and the third weight is 1; based on the first weight, the second weight, and the third weight, the data richness quality score, the business value quality score, and the data resolution score are weighted and summed respectively to obtain a data quality score for data quality evaluation within the time evaluation window.

[0220] In one embodiment of the present application, the data field richness quality scoring module 102 that determines the data field richness quality score based on the data fields in the full data and the predefined full data fields is specifically configured to:

[0221] Compare the semantics of the data fields in the full data with the semantics of the corresponding predefined full data fields for completeness, and assign a value to the full data based on the comparison result, using the value as the data field richness quality score;

[0222] Alternatively, the proportion of data fields in the full data that conform to the predefined full data fields is calculated, and based on the correspondence between the preset proportion and the data field richness quality score, the proportion is converted into a corresponding data field richness quality score.

[0223] In one embodiment of the present application, the data field richness quality scoring module 102 that determines the data field richness quality score based on the data fields in the full data and the predefined full data fields is specifically configured to:

[0224] Count the number of data fields in the full data that meet the predefined specifications, where the predefined specifications are fields with values ​​and are not empty; based on the total number of predefined full data fields and the number of data fields, obtain the proportion of full data fields that meet the predefined specifications; round down the proportion and use the obtained value as the data field richness quality score for the full data reported within the preset time period.

[0225] In one embodiment of the present application, the business value quality scoring module 103 evaluates the data fields in the full amount of data according to the field requirements corresponding to the business scenario to obtain a business value quality score, which is specifically used to:

[0226] If a business scenario contains multiple sub-business scenarios, the data fields in the full data are evaluated for the field requirements corresponding to each sub-business scenario to obtain the business data quality score of the sub-business scenario; the quality score of each business data is counted, and the sum of the obtained values ​​is used as the business value quality score of the full data reported within the preset time period.

[0227] In an embodiment of the present application, the data quality assessment device further includes: a visualization module 106 .

[0228] The visualization module 106 is configured to record the data quality scores obtained within the time evaluation window in a data quality table and dynamically display the data quality scores recorded in the data quality table. The data quality table records the data quality scores of data quality evaluations performed within multiple time evaluation windows sliding with a sliding step size.

[0229] Based on the data quality assessment device disclosed in the embodiment of the present application, data quality assessment is performed taking into account the process dimension, business value dimension and data resolution dimension, and a comprehensive assessment is performed on the data reported by the access device from multiple dimensions, which can ensure the comprehensiveness and reliability of data quality, and display the data quality assessment results through visualization tools, so that users can intuitively understand the current status and changing trends of device data quality, which has the beneficial effects of strong real-time performance, comprehensive assessment dimensions, and high visibility of the correlation between data quality and business effects.

[0230] Based on the data quality assessment device disclosed in the above embodiment of the present disclosure, each of the above modules can be implemented by a hardware device composed of a processor and a memory. Specifically, each of the above modules is stored in the memory as a program unit, and the processor executes the program unit stored in the memory to implement thread control.

[0231] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and database expansion can be achieved by adjusting kernel parameters.

[0232] An embodiment of the present application discloses a computer storage medium having a program stored thereon. When the program is executed by a processor, the data quality assessment method disclosed in the above-mentioned embodiment of the present application is implemented.

[0233] Computer-readable media include permanent and non-permanent, removable and non-removable media that can store information using any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.

[0234] An embodiment of the present application discloses an electronic device, comprising one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by one or more processors, the one or more processors implement the data quality assessment method disclosed in the above-mentioned embodiment of the present application.

[0235] Specifically, such as Figure 11 FIG2 is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. The electronic device 110 includes at least one processor 111 , at least one memory 112 connected to the processor, and a bus 113 .

[0236] The processor 111 and the memory 112 communicate with each other via the bus 113 .

[0237] The processor 111 is configured to execute the program stored in the memory.

[0238] The memory 112 is used to store a program, which is at least used to implement the data quality assessment method disclosed in the above-mentioned embodiment of the present application.

[0239] In a typical configuration, the device includes one or more processors (CPUs), memory, and a bus. The device may also include input / output interfaces, network interfaces, and the like.

[0240] Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip. Memory is an example of a computer-readable medium.

[0241] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0242] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0243] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data quality assessment method, characterized in that: The method comprises: When the time evaluation window is reached, the full amount of data reported by the target data source within the preset time period is extracted; the sliding step size of the time evaluation window is set according to the data quality evaluation requirements of the business involved in the full amount of data; Determining a predefined full data field corresponding to the full data, and determining a data field richness quality score based on the data field in the full data and the predefined full data field; Determine the business scenarios involved in the full data, evaluate the data fields in the full data according to the field requirements corresponding to the business scenarios, and obtain a business value quality score; The data richness quality score and the business value quality score are comprehensively weighted to obtain a data quality score for data quality assessment within the time assessment window.

2. The method according to claim 1, characterized in that Also includes: Counting the data resolution rate of the target data source within a preset time period, and determining a data resolution rate score corresponding to the data resolution rate based on a correspondence between the preset data resolution rate and the data resolution rate dimension score; After obtaining the data richness quality score and the business value quality score, the data richness quality score, the business value quality score and the data resolution rate score are comprehensively weighted to obtain a data quality score for data quality evaluation within the time evaluation window.

3. The method according to claim 2, characterized in that The data richness quality score, the business value quality score, and the data resolution rate score are comprehensively weighted to obtain a data quality score within the time evaluation window, including: Determining a first weight corresponding to the data richness quality score, a second weight corresponding to the business value quality score, and a third weight corresponding to the data resolution rate score; wherein the sum of the first weight, the second weight, and the third weight is 1; Based on the first weight, the second weight and the third weight, the data richness quality score, the business value quality score and the data resolution rate score are weighted and summed respectively to obtain a data quality score for data quality evaluation within the time evaluation window.

4. The method according to claim 1, wherein The determining of the data field richness quality score based on the data fields in the full data and the predefined full data fields includes: Comparing the semantics of the data fields in the full data with the semantics of the corresponding predefined full data fields for completeness, and assigning a value to the full data based on the comparison result, and using the value as the data field richness quality score; Alternatively, the proportion of data fields in the full data that conform to the predefined full data fields is calculated, and based on the correspondence between the preset proportion and the data field richness quality score, the proportion is converted into a corresponding data field richness quality score.

5. The method according to claim 1, wherein The determining of the data field richness quality score based on the data fields in the full data and the predefined full data fields includes: Counting the number of data fields in the full amount of data that meet a predefined specification, where the predefined specification is a field that has a value and is not empty; Based on the total number of predefined full data fields and the number of data fields, a ratio that conforms to the predefined full data fields is obtained; The ratio is rounded down, and the resulting value is used as the data field richness quality score of the full amount of data reported within the preset time period.

6. The method according to claim 1, characterized in that The field requirements corresponding to the business scenario are evaluated on the data fields in the full amount of data to obtain a business value quality score, including: If the business scenario includes multiple sub-business scenarios, the data fields in the full data are evaluated for the fields corresponding to each sub-business scenario to obtain the business data quality score of the sub-business scenario; The quality score of each of the business data is counted, and the obtained sum is used as the business value quality score of the full amount of data reported within the preset time period.

7. The method according to any one of claims 1 to 6, characterized in that Also includes: Recording the data quality scores obtained within the time evaluation window in a data quality table, wherein the data quality table records the data quality scores of data quality evaluations performed within multiple time evaluation windows sliding with the sliding step size; Dynamically display the data quality scores recorded in the data quality table.

8. A data quality assessment device, characterized in that: The data quality assessment device comprises: An extraction module is configured to extract the full amount of data reported by the target data source within a preset time period when a time evaluation window is reached; the sliding step size of the time evaluation window is set according to the data quality assessment requirements of the business indicated by the full amount of data; a data field richness quality scoring module, configured to determine a predefined full data field corresponding to the full data, and determine a data field richness quality score based on the data field in the full data and the predefined full data field; A business value quality scoring module is used to determine the business scenarios involved in the full data, evaluate the data fields in the full data according to the field requirements corresponding to the business scenarios, and obtain a business value quality score; The comprehensive weighted scoring module is used to comprehensively weight the data richness quality score and the business value quality score to obtain a data quality score for data quality evaluation within the time evaluation window.

9. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the data quality assessment method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that A computer program is stored thereon, wherein when the computer program is executed by a processor, the data quality assessment method according to any one of claims 1 to 7 is implemented.