A method, apparatus, medium, and product for quality assessment of network resource data.

By quantifying the hierarchical attributes and automation level of network resource data, and combining clustering fitting models and linear regression algorithms, the problem of inaccurate evaluation of network resource data in existing technologies is solved, and data services with higher accuracy and reliability are achieved.

CN118964340BActive Publication Date: 2025-10-31CHINA MOBILE GROUP DESIGN INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410967015.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-10-31
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

Existing methods for assessing the quality of network resource data are not accurate enough, resulting in high data verification indicators that are unusable in certain scenarios and cannot represent the accuracy and reliability of the data.

Method used

By quantifying the hierarchical attributes and automation level of network resource data, and combining clustering fitting models and linear regression algorithms, the reliability assessment value of the data is calculated, freeing it from the constraints of human experience and the assumption of data isolation, and providing more accurate and reliable data services.

Benefits of technology

It improves the accuracy and reliability of credibility assessment of network resource data in various business scenarios, eliminates the influence of human subjective factors, and ensures the accuracy and reliability of data services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964340B_ABST
    Figure CN118964340B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, medium, and product for quality assessment of network resource data. The method quantifies the importance of the data to be assessed based on the hierarchical attributes of the resource objects, obtaining a quantified importance value. It also quantifies the automation level of data maintenance based on the maintenance methods of each attribute of the resource objects, obtaining a quantified automation value. Furthermore, it quantifies the data quality of the data to be assessed, determining quality quantification values ​​for different indicators. Finally, it calculates the reliability assessment value of the data based on the quantified importance value, the quantified automation value, and the quality quantification values ​​of different indicators. By utilizing the hierarchical attributes of network resource data and the maintenance methods of each attribute of the data to be assessed, the method quantifies the data based on both importance and automation level, freeing it from the constraints of human experience and the assumption of isolated data, thus providing more accurate and reliable data services for business management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data quality assessment technology, and more specifically, to a method, apparatus, medium, and product for assessing the quality of network resource data. Background Technology

[0002] In the operation support system, network resource data is primarily managed centrally within the resource management system. This involves a wide variety of basic network types and data categories, resulting in a massive data volume. It provides data services for various aspects of operations, including customer service activation, network planning and design, fault monitoring, asset inventory, and maintenance cost settlement. Therefore, the accuracy of resource data is one of the key priorities for major operators.

[0003] Existing methods for assessing the quality of network resource data are still not accurate enough. Often, although data verification indicators are high, they do not necessarily mean that the data is accurate and reliable, and in some scenarios, the data may still be unusable. Summary of the Invention

[0004] Compared with existing technologies, this invention proposes a method, apparatus, medium, and product for quality assessment of network resource data, providing more accurate and reliable data services for business management.

[0005] This invention provides a method for quality assessment of network resource data, the method comprising:

[0006] The importance of the data to be evaluated is quantified based on the hierarchical attributes of the resource objects to be evaluated, and the quantified value of the importance of the data to be evaluated is obtained.

[0007] Based on the maintenance method of each attribute of the resource object of the data to be evaluated, the degree of automation of data maintenance is quantified to obtain the quantified value of the degree of automation of the data to be evaluated;

[0008] The data quality of the data to be evaluated is quantified to determine the quality quantification values ​​of different indicators;

[0009] The credibility assessment value of the data to be evaluated is calculated based on the quantitative value of importance, the quantitative value of automation, and the quantitative value of quality of different indicators.

[0010] Preferably, the reliability assessment value of the data to be evaluated is calculated based on the importance quantification value, the automation degree quantification value, and the quality quantification values ​​of different indicators, including:

[0011] The importance quantification value, the automation quantification value, and the quality quantification value of different indicators are used as evaluation items for the data to be evaluated.

[0012] The evaluation items are fitted using a pre-trained clustering fitting model, and the weights of different evaluation items are calculated.

[0013] The credibility assessment value is obtained by weighting and summing the weights of different assessment items with the corresponding quantitative values ​​of the evaluation items.

[0014] Preferably, the importance of the data to be evaluated is quantified according to the hierarchical attributes of the resource objects to be evaluated, to obtain a quantified value of the importance of the data to be evaluated, including:

[0015] Obtain the hierarchical attributes of the resource objects in the data to be evaluated;

[0016] Based on the hierarchical attributes of the resource objects in the data to be evaluated, the corresponding importance quantification value is matched from a preset importance matching library;

[0017] The hierarchical attributes include at least one of level, type, or hierarchy.

[0018] Preferably, the degree of automation in data maintenance is quantified based on the maintenance method of each attribute of the resource object of the data to be evaluated, to obtain a quantified value of the degree of automation of the data to be evaluated, including:

[0019] Obtain the maintenance method of each attribute of the resource object of the data to be evaluated;

[0020] Calculate the proportion of attributes whose preset essential attribute centralized maintenance method is automatic collection or automatic system maintenance.

[0021] The degree of automation is quantified based on the calculated proportion.

[0022] Preferably, the data quality of the data to be evaluated is quantified to determine the quality quantification values ​​of different indicators, including:

[0023] A preset number of fields are extracted from the data to be evaluated for verification, and the extracted fields are checked to see if they meet the requirements of the preset verification rules for indicators of different dimensions.

[0024] The pass rate of the verification for different dimensions of indicators is calculated separately and used as the direct data quality quantification value for the corresponding indicators.

[0025] Preferably, the method further includes:

[0026] The parent and child resources corresponding to the indicators of the data to be evaluated are determined based on the dependency relationships of the network resource data.

[0027] Calculate the quality quantification values ​​of different indicators of the sub-resources in the data to be evaluated;

[0028] The direct data quality quantification values ​​of different indicators of the data to be evaluated, the number of sub-resources, and the direct data quality quantification values ​​of the corresponding indicators of the sub-resources are input into a preset comprehensive evaluation model to calculate the comprehensive evaluation value, which is used as the indirect data quality quantification value of the indicator.

[0029] As a preferred embodiment, the process of constructing the clustering fitting model includes:

[0030] The acquired training dataset is labeled with its credibility level.

[0031] Based on the importance quantification value and automation quantification value in the training dataset, the K-means algorithm is introduced to perform cluster analysis on the data, dividing the training dataset into multiple intervals;

[0032] The labeled training dataset is divided into multiple intervals, and a preset linear regression algorithm is used to fit the clustering fitting model.

[0033] As a preferred embodiment, the method further includes:

[0034] Calculate the credibility assessment value of different data to be evaluated in the acquired network resource data;

[0035] The network resource data is filtered based on the credibility assessment values ​​of different data to be evaluated and the preset credibility range.

[0036] This invention also provides a network resource data quality assessment device, the device comprising:

[0037] The importance quantification module is used to quantify the importance of the data to be evaluated based on the hierarchical attributes of the resource objects of the data to be evaluated, and to obtain the importance quantification value of the data to be evaluated.

[0038] The automation quantification module is used to quantify the degree of automation of data maintenance based on the maintenance method of each attribute of the resource object of the data to be evaluated, and obtain the quantified value of the degree of automation of the data to be evaluated.

[0039] The quality quantification module is used to quantify the data quality of the data to be evaluated and determine the quality quantification values ​​of different indicators.

[0040] The credibility assessment module is used to calculate the credibility assessment value of the data to be assessed based on the importance quantification value, the automation quantification value, and the quality quantification value of different indicators.

[0041] This invention also provides a network resource data quality assessment apparatus, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a network resource data quality assessment method as described in any of the above embodiments.

[0042] This invention also provides a computer-readable storage medium, which includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a network resource data quality assessment method as described in any of the above embodiments.

[0043] This invention also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any of the above embodiments.

[0044] Compared with existing technologies, this invention provides a method, apparatus, medium, and product for quality assessment of network resource data. It quantifies the importance of the data to be assessed based on the hierarchical attributes of the resource objects, obtaining a quantified importance value; it quantifies the automation level of data maintenance based on the maintenance methods of each attribute of the resource objects, obtaining a quantified automation value; it quantifies the data quality of the data to be assessed, determining quality quantified values ​​for different indicators; and it calculates the reliability assessment value of the data to be assessed based on the quantified importance value, the quantified automation value, and the quality quantified values ​​of different indicators. This application's solution quantifies data based on its hierarchical attributes and the maintenance methods of each attribute of the data to be assessed, quantifying both importance and automation levels, thus overcoming the constraints of human experience and the assumption of isolated data, and providing more accurate and reliable data services for business management. Attached Figure Description

[0045] Figure 1 This is a flowchart illustrating the quality assessment method for network resource data provided in an embodiment of the present invention;

[0046] Figure 2 This is a schematic diagram illustrating the construction process of the clustering fitting model provided in an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram illustrating the principle of the clustering analysis algorithm provided in this embodiment of the invention;

[0048] Figure 4 This is a schematic diagram of the structure of a network resource data quality assessment device provided in an embodiment of the present invention;

[0049] Figure 5This is another structural schematic diagram of a network resource data quality assessment device provided in an embodiment of the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] To address the aforementioned technical problems, this application proposes a method for quality assessment of network resource data, see [link to relevant documentation]. Figure 1 This is a flowchart illustrating a method for quality assessment of network resource data provided in an embodiment of the present invention. The method includes the following steps:

[0052] Step S1: Quantify the importance of the data to be evaluated based on the hierarchical attributes of the network resource data to obtain the quantified value of the importance of the data to be evaluated;

[0053] Step S2: Based on the maintenance method of each attribute of the data to be evaluated, quantify the degree of automation of data maintenance to obtain the quantified value of the degree of automation of the data to be evaluated;

[0054] Step S3: Quantify the data quality of the data to be evaluated and determine the quality quantification values ​​of different indicators;

[0055] Step S4: Calculate the credibility assessment value of the data to be evaluated based on the importance quantification value, the automation quantification value, and the quality quantification values ​​of different indicators.

[0056] In this specific implementation, when assessing the quality of network resource data, verification rules are established, the verification results are quantified, and different weights are assigned to different verification dimensions to evaluate data quality. However, it has been proven that this assessment method is still not accurate enough. Often, although data verification indicators are high, the data may still be unusable in certain scenarios. High data verification indicators do not necessarily indicate data reliability.

[0057] Therefore, in response to the aforementioned technical problems, this case provides a method for quality assessment of network resource data, which improves the quantification of data in various dimensions, specifically including:

[0058] The hierarchical attributes of the network provide a basis for quantifying the importance of data. That is, the importance of the data to be evaluated is quantified according to the hierarchical attributes of network resource data, so that the closer the resources are to the core layer and the backbone layer, the higher the data credibility evaluation requirements are, and thus the quantitative value of the importance of the data to be evaluated is obtained.

[0059] When quantifying the degree of automation in data maintenance, the degree of automation in data maintenance is quantified according to the maintenance method of each attribute of each resource object, so as to obtain the quantified value of the degree of automation of the data to be evaluated.

[0060] The data quality of the data to be evaluated is quantified according to a preset data quality quantification method to determine the quality quantification values ​​of different indicators.

[0061] The importance quantification value, the automation quantification value, and the quality quantification values ​​of different indicators are input into the credibility assessment model to calculate the credibility assessment value of the data to be assessed.

[0062] Based on the above quantitative results, the overall calculation model for the credibility assessment model of network resource data can be expressed as: (S,F)=f(x,y,z) 1.1 ,z 1.2 ,...,z m.n )

[0063] Input: x is the quantitative value of the importance of the data to be evaluated, y is the quantitative value of the automation level of the data to be evaluated, and z is... n.m The data quality measures for each dimension are given. The model output S represents whether the data is trustworthy, and F represents the comprehensive trustworthiness assessment measure.

[0064] This solution quantifies data based on its importance and automation level by using hierarchical attributes of network resource data and the maintenance methods for each attribute of the data to be evaluated. Automation level indicates whether the data was automatically acquired; automatically acquired data is considered superior to manually entered or manually set data, as it has higher reliability. All indicator evaluation weights are based on real data, thus eliminating subjective human factors to a certain extent. This avoids the influence of existing technologies where indicator weights are derived from probabilistic data relying on human experience, which contains significant subjective factors and affects accuracy.

[0065] It should be noted that the above embodiments only disclose the credibility assessment method for each resource to be evaluated in the network resource data. In specific implementation, each piece of data in the network resource data can be used as the resource to be evaluated for credibility assessment.

[0066] The importance of each piece of data is quantified based on the hierarchical attributes of the resource objects in the network resource data, and the quantified value of the importance of each piece of data is obtained.

[0067] Based on the maintenance methods of each attribute of each data resource object, the degree of automation of data maintenance is quantified to obtain the quantified value of the degree of automation of each data.

[0068] Quantify the quality of each data point and determine the quality quantification values ​​for different indicators;

[0069] The reliability assessment value of each data point is calculated based on the quantitative values ​​of importance, automation, and quality of different indicators.

[0070] Network resource data is filtered based on the credibility assessment values ​​of each data point, and the filtered data is used to provide data services for management operations.

[0071] This solution improves the quantification of data across various dimensions, intelligently calculates the impact of each dimension on the data, breaks free from the constraints of human experience and isolated data assumptions, and constructs a more accurate credibility assessment model for individual network resource data. The resulting credibility assessment value has high credibility. After determining the credibility of the network resource data, the data of the objects to be assessed in the network resource data is filtered according to the credibility assessment value, thereby improving the reliability of network resource data in providing data services in various business scenarios.

[0072] In another embodiment provided by the present invention, step S4 specifically includes:

[0073] The quantitative values ​​of the importance of the data to be evaluated, the quantitative values ​​of the degree of automation, and the quantitative values ​​of the quality of different indicators are used as evaluation items for the overall calculation model of the credibility assessment model.

[0074] After fitting all the labeled data, the weight of each evaluation item in different intervals is determined based on the cluster fitting model.

[0075] In practice, based on DS evidence theory, AHP and other algorithms, the credibility weight of each dimension is calculated by using a single algorithm or in combination with the EWM algorithm. When using DS evidence theory, it is necessary to assess the basic probability of each data point's distribution in different credibility intervals in advance based on expert experience; when using the AHP algorithm, it is necessary to assess the relative importance between any two dimensions based on expert experience.

[0076] The overall credibility is calculated by weighting and summing the quantified values ​​of different evaluation items according to their respective weights. Specifically, the overall credibility is calculated by weighted summing of the quantified results and credibility weights for each dimension. The calculation formula is expressed as follows: Where s i w represents the quantitative value of each evaluation item in the verification dimension. i This indicates the weight of each evaluation item in the verification dimension.

[0077] By comprehensively calculating the weights of various evaluation items and incorporating a regression algorithm, the credibility assessment value is obtained. This avoids the subjective limitations of human experience, constructs a more accurate credibility assessment model for single network resource data, and improves the accuracy of credibility assessment.

[0078] In another embodiment provided by the present invention, step S1 specifically includes:

[0079] The hierarchical attributes of the resource objects to be evaluated are obtained, and these attributes provide the primary basis for quantifying the importance of the data. Resources closer to the core or backbone layer require higher data credibility for evaluation.

[0080] The network's hierarchical attributes are the provincial core layer and the cross-provincial trunk resources. The importance of its data is extremely important, and its importance is quantified as 4.

[0081] For the network's hierarchical attributes of provincial aggregation layer and secondary backbone resources, the importance of its data is very important, and its importance is quantified as 3.

[0082] For the network's layered attributes, namely the metropolitan area access layer and the area between the central office access point and the customer-side access point, the importance of its data is relatively important, and its importance is quantified as 2.

[0083] For network layered attributes, the data of the end access layer is of general importance, and its importance is quantified as 1.

[0084] Most resource data already maintains attributes such as level / category / hierarchy directly or indirectly through association / inheritance relationships. This proposal quantifies these attributes by pre-constructing an importance matching library based on the network's hierarchical attributes. The importance matching library for resources in the space and transport specialties is shown in Table 1:

[0085] Table 1, Importance Matching Library

[0086]

[0087]

[0088] Based on the hierarchical attributes of the resource objects in the data to be evaluated, the corresponding importance quantification value is matched from a preset importance matching library. The hierarchical attributes include at least one of level, type, or hierarchy.

[0089] In this case, the interdependence of network resource data was taken into account when quantifying the data. The data was no longer evaluated based on the assumption that they were independent of each other. The correlation between network resource data was respected. At the same time, the impact of different levels on the accuracy of the data was also taken into account. To a certain extent, the contradiction of high-quality data being unusable was eliminated, making the evaluation results more accurate.

[0090] In another embodiment provided by the present invention, step S2 specifically includes:

[0091] Obtain the maintenance method of each attribute of the resource object of the data to be evaluated;

[0092] The method of maintaining attributes, that is, the way of maintaining data for different attributes, includes at least automatic collection or system automation, and may also include other methods, such as manual collection and manual input.

[0093] Based on the maintenance method of each attribute of each resource object, the proportion of automatically collected and automatically maintained attributes in the necessary attribute set is statistically analyzed, and the degree of automation is determined based on the calculated proportion.

[0094] The proportion of automatically collected data in the necessary attribute set is used as a quantitative basis for the degree of automation in data maintenance.

[0095] Taking IP bearer network element as a resource object as an example, the required and conditionally required attributes are defined as necessary attributes. There are 39 necessary attributes in total, and 21 attributes can be automatically maintained. Then the proportion of IP network element is 21 / 39*100=53.8. The automation level is determined according to the proportion of attributes.

[0096] It should be noted that the automation level quantification value is determined based on the calculated proportion. Specifically, the proportion can be used directly as the automation level quantification value, or the automation level quantification value can be matched according to the proportion based on a preset matching relationship.

[0097] Taking into account the impact of different maintenance methods on the accuracy of network resource data, the contradiction of high data quality but unusable data has been eliminated to a certain extent, making the evaluation results more accurate.

[0098] In another embodiment provided by the present invention, step S3 specifically includes a direct quality scoring quantification process, specifically including:

[0099] A preset number of fields are extracted from the data to be evaluated for verification, and the extracted fields are checked to see if they meet the requirements of the preset verification rules for indicators of different dimensions.

[0100] The pass rate of the verification for different dimensions of indicators is calculated separately and used as the direct data quality quantification value for the corresponding indicators.

[0101] The obtained direct data quality quantification index is used as the quality quantification value of the index.

[0102] In practice, the direct data quality score will be quantified by consisting of 11 indicators across 6 dimensions: completeness, compliance, standardization, uniqueness, consistency, and activity.

[0103] For the completeness indicator in the completeness dimension, when quantifying the indicator, it is based on the number of completeness checks that pass for the required and conditionally required fields of a single data entry. The completeness rate is calculated as: (Number of fields that pass the completeness check / Total number of required and conditionally required fields) * 100.

[0104] For the cross-professional correlation rate indicator in the compliance dimension, when quantifying the indicator, it is based on the number of cross-professional correlation fields that have passed the verification for a single data point. Cross-professional correlation rate = (Number of cross-professional correlation fields that have passed the verification / Total number of cross-professional correlation fields that have passed the verification) * 100.

[0105] For the business compliance rate indicator in the compliance dimension, when quantifying the indicator, it is quantified based on the number of business compliance verification rules passed for a single data point. Business compliance rate = number of rules passed / total number of business compliance rules * 100.

[0106] For the end-to-end completeness rate metric in the compliance dimension, when quantifying the metric, it is quantified based on the number of end-to-end verification rules passed for a single data point. The end-to-end completeness metric quantification result = number of end-to-end rules that passed verification / total number of end-to-end rules * 100.

[0107] For the professional internal correlation rate indicator in the standardization dimension, when quantifying the indicator, it is based on the number of professional internal correlation verification passes for the correlation fields of a single data point. Professional internal correlation rate = number of fields that pass professional internal correlation verification / total number of correlation verification fields * 100.

[0108] For the enumeration standardization rate indicator in the standardization dimension, when quantifying the indicator, it is based on the number of enumeration fields that pass the verification of a single data point. The enumeration standardization rate = number of fields that pass the enumeration verification / total number of enumeration fields * 100.

[0109] For the naming compliance rate indicator in the standardization dimension, when quantifying the indicator, it is quantified based on the number of naming compliance verification rules for a single data point. The naming compliance rate = number of rules that have passed verification / total number of naming compliance verification rules * 100.

[0110] For uniqueness metrics, quantification is based on whether the unique primary key of a single data entry is duplicated, and is not limited to UID, IP, or name. If any primary key is duplicated, the quantification is 0; otherwise, it is quantized to 100.

[0111] For the consistency rate metric, when quantifying the metric, consistency refers to whether the data is consistent with external systems such as fault, signaling, and professional workbench. The consistency quantification result for each system is the percentage of fields that are consistent out of the total number of fields involved in the comparison, and the consistency quantification result for all systems is the average of all systems. That is, consistency rate = avg(number of fields that are consistent in each system / number of fields involved in the comparison * 100).

[0112] For the dormancy time metric in the activity dimension, when quantifying the metric, dormancy time refers to the time elapsed since the last data change. The longer the dormancy time, the lower the reliability. The quantified dormancy time result is calculated as (current time - last update time) / (current time - network entry time) * 100.

[0113] For the external usage time metric of the activity dimension, when quantifying the metric, the data is judged based on the feedback of work orders such as business activation, fault dispatch, and inspection of the supporting external system to determine whether the data has been used externally and the usage time. The external usage time quantification result = (current time - the most recent usage time of the external system) / (current time - network access time) * 100.

[0114] When quantifying indicators, the quantification of data quality is accurate enough. The verification of atomic-level evaluation items in each dimension is carried out through the proportion of fields and rules. Intermediate results are allowed, which can objectively reflect the overall situation of each dimension and avoid information loss caused by one-size-fits-all quantification.

[0115] In another embodiment of the present invention, based on the direct quality score quantification in step S3, the specific implementation of this case further includes an indirect quality score quantification process, specifically including:

[0116] The dependency relationships of network resource data include the correspondence between parent resources and child resources.

[0117] For example, the sub-resources of a network element are ports, the sub-resources of a computer room are network elements, and the sub-resources of an optical cable are fiber cores.

[0118] Then, based on the dependency relationship of network resource data, the parent resource of the indicator of the data to be evaluated that has child resources, and the child resources corresponding to the parent resource can be determined.

[0119] In practice, the method for quantifying the quality of resource objects can adopt the scheme in the above embodiments to calculate the quality quantification values ​​of different indicators of resource objects as sub-resources.

[0120] The direct data quality quantification values ​​of different indicators of the data to be evaluated, the number of sub-resources, and the direct data quality quantification values ​​of the corresponding indicators of the sub-resources are input into a preset comprehensive evaluation model to calculate the comprehensive evaluation value and obtain the indirect data quality quantification value of the indicator.

[0121] The obtained indirect data quality quantification index is used as the quality quantification value of the index.

[0122] For example, if the direct evaluation value of the data quality of a parent resource is Z1, and there are N child resources, each with an evaluation value of S... n Then the final assessed value of the parent resource

[0123] When N is 0, meaning there are no child resources under the parent resource, then Z2 is directly 0.

[0124] That is, for two or more resources that are dependent on each other, if the quality of the child resource is not high, the quality of the parent resource will be reduced accordingly, and the quality quantification value of the indicator will be updated.

[0125] Based on the fact that network resource data are interconnected, data that is independent and does not have sub-resources indicates that its data quality is problematic and its quantitative value of network data quality is low. Taking into account the interdependence of network resource data, the data is no longer evaluated based on the assumption of mutual independence, thereby improving the reliability of data quality assessment.

[0126] In another embodiment of the present invention, the process of constructing a clustering fitting model specifically includes a labeling stage, a training stage, and a service publishing stage.

[0127] During annotation, the obtained training dataset is labeled with its credibility level. See Table 2, which is a data annotation format table provided in this embodiment of the invention:

[0128] Table 2, Data Labeling Format Table

[0129]

[0130] The obtained training dataset is labeled with confidence level according to the format in Table 2.

[0131] During the training phase, this proposal no longer relies on expert experience to determine the weighting of each evaluation metric. Instead, it introduces a linear regression algorithm for intelligent training and optimization. The entire process is mainly divided into three parts: annotation, training, and service deployment. See also... Figure 2 This is a schematic diagram illustrating the construction process of the clustering fitting model provided in this embodiment of the invention.

[0132] Based on the formula derivation of the subjective weighting algorithm, the data credibility is ultimately calculated by weighted summation of the evaluation results of each dimension and their credibility measurement weights. This is the main basis for choosing the linear regression algorithm in this proposal. Ridge regression or LASSO regression algorithms are then introduced simultaneously for bias correction. The purpose of this stage is to calculate the weights of each indicator based on the labeled data.

[0133] Before performing the fitting calculation, the K-means algorithm is first introduced to perform cluster analysis on the data based on the importance of the data and the degree of automation in data maintenance. See [link to relevant documentation]. Figure 3 This is a schematic diagram of the clustering analysis algorithm provided in this embodiment of the invention. The labeled training dataset is divided into multiple intervals, and the weights within each interval are determined to ensure the consistency of the data evaluation criteria.

[0134] That is, after fitting all the labeled data, the weight of each evaluation item in different intervals and the threshold of different confidence levels are determined, the data confidence model is completed, and it is packaged into a service to obtain the cluster fitting model for subsequent use.

[0135] This solution improves upon existing data quality assessment methods by quantifying data across various dimensions. It proposes a hierarchical assessment method based on linear regression algorithms, which intelligently calculates the weights of each influencing factor, freeing it from the constraints of human experience and isolated data assumptions, thereby enhancing the accuracy of credibility assessment.

[0136] In another embodiment provided by the present invention, the method further includes:

[0137] Calculate the credibility assessment value of different data to be evaluated in the acquired network resource data;

[0138] The network resource data is filtered based on the credibility assessment values ​​of different data to be evaluated and the preset credibility range.

[0139] Based on the credibility assessment calculation of the data to be evaluated, this solution calculates the credibility assessment value of each data point according to the quantitative value of importance, the quantitative value of automation, and the quantitative value of quality of different indicators.

[0140] The network resource data is filtered based on the credibility assessment value of each data point and the preset credibility range, and the data that falls within the credibility range is selected.

[0141] The filtered data is used to provide data services for business management, providing reliable data services for business management.

[0142] See Figure 4 This is a schematic diagram of a network resource data quality assessment device provided in an embodiment of the present invention. The device includes:

[0143] The importance quantification module is used to quantify the importance of the data to be evaluated based on the hierarchical attributes of the resource objects of the data to be evaluated, and to obtain the importance quantification value of the data to be evaluated.

[0144] The automation quantification module is used to quantify the degree of automation of data maintenance based on the maintenance method of each attribute of the resource object of the data to be evaluated, and obtain the quantified value of the degree of automation of the data to be evaluated.

[0145] The quality quantification module is used to quantify the data quality of the data to be evaluated and determine the quality quantification values ​​of different indicators.

[0146] The credibility assessment module is used to calculate the credibility assessment value of the data to be assessed based on the importance quantification value, the automation quantification value, and the quality quantification value of different indicators.

[0147] It should be noted that the network resource data quality assessment device provided in this embodiment can perform all the steps and functions of the network resource data quality assessment method provided in any of the above embodiments, and the specific functions of the device will not be described in detail here.

[0148] See Figure 5 This is another structural schematic diagram of a network resource data quality assessment device provided in an embodiment of the present invention. The network resource data quality assessment device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a network resource data quality assessment program. When the processor executes the computer program, it implements the steps in the various embodiments of the network resource data quality assessment method described above, for example... Figure 1 The steps S1 to S4 are shown. Alternatively, when the processor executes the computer program, it implements the functions of each module in the above-described device embodiments.

[0149] For example, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the network resource data quality assessment device. For example, the computer program can be divided into several modules, the specific functions of which have been described in detail in the network resource data quality assessment method provided in any of the above embodiments; therefore, the specific functions of the device will not be repeated here.

[0150] The network resource data quality assessment device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The network resource data quality assessment device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the schematic diagram is merely an example of a network resource data quality assessment device and does not constitute a limitation on such a device. It may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the network resource data quality assessment device may also include input / output devices, network access devices, buses, etc.

[0151] The processor referred to can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the network resource data quality assessment device, connecting various parts of the device via various interfaces and lines.

[0152] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the network resource data quality assessment device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0153] If the module integrated into the network resource data quality assessment device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0154] This invention also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any of the above embodiments.

[0155] It should be noted that the computer program product provided in this embodiment can execute all the steps and functions of a network resource data quality assessment method provided in any of the above embodiments, and the specific functions of the device will not be described in detail here.

[0156] It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered to be within the scope of protection of this invention.

Claims

1. A method for quality assessment of network resource data, characterized in that, The method includes: The importance of the data to be evaluated is quantified based on the hierarchical attributes of the resource objects to be evaluated, and the quantified value of the importance of the data to be evaluated is obtained. Based on the maintenance method of each attribute of the resource object of the data to be evaluated, the degree of automation of data maintenance is quantified to obtain the quantified value of the degree of automation of the data to be evaluated; The data quality of the data to be evaluated is quantified to determine the quality quantification values ​​of different indicators; The credibility assessment value of the data to be evaluated is calculated based on the quantitative value of importance, the quantitative value of automation, and the quantitative value of quality of different indicators. The data quality of the data to be evaluated is quantified to determine the quality quantification values ​​of different indicators, including: A preset number of fields are extracted from the data to be evaluated for verification, and the extracted fields are checked to see if they meet the requirements of the preset verification rules for indicators of different dimensions. The pass rate of the verification for different dimensions of indicators is calculated separately and used as the direct data quality quantification value for the corresponding indicators; The method further includes: The parent and child resources corresponding to the indicators of the data to be evaluated are determined based on the dependency relationships of the network resource data. The direct data quality quantification values ​​of different indicators of the data to be evaluated, the number of sub-resources, and the direct data quality quantification values ​​of the corresponding indicators of the sub-resources are input into a preset comprehensive evaluation model to calculate the comprehensive evaluation value, which is used as the indirect data quality quantification value of the indicator; the obtained indirect data quality quantification index is used as the quality quantification value of the indicator.

2. The method for quality assessment of network resource data according to claim 1, characterized in that, The reliability assessment value of the data to be evaluated is calculated based on the quantitative values ​​of importance, automation, and quality of different indicators, including: The importance quantification value, the automation quantification value, and the quality quantification value of different indicators are used as evaluation items for the data to be evaluated. The evaluation items are fitted using a pre-trained clustering fitting model, and the weights of different evaluation items are calculated. The credibility assessment value is obtained by weighting and summing the weights of different assessment items with the corresponding quantitative values ​​of the evaluation items.

3. The method for quality assessment of network resource data according to claim 1, characterized in that, The importance of the data to be evaluated is quantified based on the hierarchical attributes of the resource objects to be evaluated, resulting in a quantified value of the importance of the data to be evaluated, including: Obtain the hierarchical attributes of the resource objects in the data to be evaluated; Based on the hierarchical attributes of the resource objects in the data to be evaluated, the corresponding importance quantification value is matched from a preset importance matching library; The hierarchical attributes include at least one of level, type, or hierarchy.

4. The method for quality assessment of network resource data according to claim 1, characterized in that, Based on the maintenance methods of each attribute of the resource object in the data to be evaluated, the degree of automation in data maintenance is quantified to obtain a quantitative value of the degree of automation of the data to be evaluated, including: Obtain the maintenance method of each attribute of the resource object of the data to be evaluated; Calculate the proportion of attributes whose preset essential attribute centralized maintenance method is automatic collection or automatic system maintenance. The degree of automation is quantified based on the calculated proportion.

5. The method for quality assessment of network resource data according to claim 2, characterized in that, The process of constructing the clustering fitting model includes: The acquired training dataset is labeled with its credibility level. Based on the importance quantification value and automation quantification value in the training dataset, the K-means algorithm is introduced to perform cluster analysis on the data, dividing the training dataset into multiple intervals; The labeled training dataset is divided into multiple intervals, and a preset linear regression algorithm is used to fit the clustering fitting model.

6. The method for quality assessment of network resource data according to claim 1, characterized in that, The method further includes: Calculate the credibility assessment value of different data to be evaluated in the acquired network resource data; The network resource data is filtered based on the credibility assessment values ​​of different data to be evaluated and the preset credibility range.

7. A quality assessment device for network resource data, characterized in that, The device includes: The importance quantification module is used to quantify the importance of the data to be evaluated based on the hierarchical attributes of the resource objects of the data to be evaluated, and to obtain the importance quantification value of the data to be evaluated. The automation quantification module is used to quantify the degree of automation of data maintenance based on the maintenance method of each attribute of the resource object of the data to be evaluated, and obtain the quantified value of the degree of automation of the data to be evaluated. The quality quantification module is used to quantify the data quality of the data to be evaluated and determine the quality quantification values ​​of different indicators. The credibility assessment module is used to calculate the credibility assessment value of the data to be assessed based on the importance quantification value, the automation quantification value, and the quality quantification value of different indicators. The quality quantification module is used for: A preset number of fields are extracted from the data to be evaluated for verification, and the extracted fields are checked to see if they meet the requirements of the preset verification rules for indicators of different dimensions. The pass rate of the verification for different dimensions of indicators is calculated separately and used as the direct data quality quantification value for the corresponding indicators; The device is also used for: The parent and child resources corresponding to the indicators of the data to be evaluated are determined based on the dependency relationships of the network resource data. The direct data quality quantification values ​​of different indicators of the data to be evaluated, the number of sub-resources, and the direct data quality quantification values ​​of the corresponding indicators of the sub-resources are input into a preset comprehensive evaluation model to calculate the comprehensive evaluation value, which is used as the indirect data quality quantification value of the indicator; the obtained indirect data quality quantification index is used as the quality quantification value of the indicator.

8. A quality assessment device for network resource data, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements a method for quality assessment of network resource data as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a method for quality assessment of network resource data as described in any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for evaluating reliability of data assets

    CN105023119A

  • Network asset management method and device, equipment and medium

    CN113408948A