Data Value Evaluation and Analysis Method and Device Based on Data Attribute Analysis

Through the method based on data attribute analysis, the value of a single data attribute and multiple data attributes in different application scenarios is calculated, which solves the problem of inaccurate data value assessment in the existing technology, and realizes accurate value assessment of data assets and data set mining that maximizes value.

CN113901106BActive Publication Date: 2025-06-13FUJIAN ZHONGXIN NET SAFETY INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111175477.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-09
Publication Date
2025-06-13
Estimated Expiration
2041-10-09

AI Technical Summary

Technical Problem

The existing technology lacks a complete data value assessment analysis method, making it difficult to accurately evaluate data assets.

Method used

Using a method based on data attribute analysis, we use a method to calculate the value of a single data attribute in multiple application scenarios and evaluate whether the combination of multiple data attributes can generate higher data asset value, and finally determine the data set with the highest value of data assets in the current application scenario.

Benefits of technology

Accurate value evaluation of data assets is achieved, not only considering the value of a single data attribute, but also evaluating the combined effect between multiple data attributes, thereby mining data sets with higher value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901106B_ABST
    Figure CN113901106B_ABST
Patent Text Reader

Abstract

The present invention discloses a data value evaluation and analysis method and device based on data attribute analysis. The first data asset value of a single-attribute data set of a single data attribute is calculated and obtained according to a data value evaluation and analysis algorithm in multiple different application scenarios; the data attribute sets required in different application scenarios are obtained, and all the data attributes in the data attribute sets are combined into multiple data attribute subsets; the data sets of all the data attributes under the same data attribute subset are used as a multi-attribute data set, and the second data asset value of each multi-attribute data set in the associated application scenario is calculated and obtained according to the data value evaluation and analysis algorithm; the data set with the highest data asset value in the current application scenario is obtained according to the first data asset value and the second data asset value in the current application scenario. The present invention not only indirectly improves the value of data assets, but also realizes the accurate value evaluation of data assets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and particularly relates to a data value evaluation and analysis method and device based on data attribute analysis. Background Art

[0002] A large amount of data generated by all walks of life has increasingly become a digital asset comparable to tangible assets, and the data of certain key industries has even become a strategic resource that the country urgently needs to protect. The digital economy with data as the key element has become a new engine for industrial development and a new driving force for national development.

[0003] As a new type of intangible asset, data assets have the characteristics of non-consumability, value-added, dependence, and value volatility. They are controlled by enterprise entities and depend on tangible assets. The value of this application is also affected by many variable factors, such as the quality of data and the application value of data in different scenarios. Therefore, the existing value evaluation of data assets is still very lacking, and there is currently no complete data value evaluation and analysis method. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to provide a data value evaluation and analysis method and device based on data attribute analysis to accurately evaluate the value of data assets.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is:

[0006] A data value evaluation and analysis method based on data attribute analysis, comprising:

[0007] Step S1, calculating and obtaining the first data asset value of a single-attribute data set of a single data attribute in multiple different application scenarios according to a data value evaluation and analysis algorithm;

[0008] Step S2, obtaining a set of data attributes required in different application scenarios, combining all data attributes in the set of data attributes into multiple data attribute subsets, and each data attribute subset is associated with the corresponding application scenario and includes at least two of the data attributes;

[0009] Step S3, taking the data sets of all data attributes under the same data attribute subset as a multi-attribute data set, and calculating and obtaining the second data asset value of each multi-attribute data set in the associated application scenario according to a data value evaluation and analysis algorithm;

[0010] Step S4, obtaining the data set with the highest data asset value in the current application scenario according to the first data asset value and the second data asset value in the current application scenario.

[0011] In order to solve the above technical problems, another technical solution adopted by the present invention is:

[0012] A data value assessment and analysis device based on data attribute analysis includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned data value assessment and analysis method based on data attribute analysis are implemented.

[0013] The beneficial effects of the present invention are: a data value assessment analysis method and device based on data attribute analysis, which calculates and obtains the first data asset value of a single attribute data set of a single data attribute in multiple different application scenarios according to a data value assessment analysis algorithm, and calculates and obtains the second data asset value of each multi-attribute data set in the associated application scenario, and obtains the data set with the highest data asset value in the current application scenario according to the first data asset value and the second data asset value in the current application scenario. Thus, not only the data asset value of each single data attribute is taken into account, but also whether the combination of multiple data attributes can produce a higher data asset value is evaluated, thereby mining a data set with a higher data asset value to reflect the true value of the data assets of each data attribute, which not only indirectly increases the value of the data assets, but also realizes the accurate value assessment of the data assets. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A schematic diagram of a flow chart of a data value assessment and analysis method based on data attribute analysis according to an embodiment of the present invention;

[0015] Figure 2 It is a structural schematic diagram of a data value assessment and analysis device based on data attribute analysis according to an embodiment of the present invention.

[0016] Description of labels:

[0017] 1. A data value assessment and analysis device based on data attribute analysis; 2. A processor; 3. A memory. DETAILED DESCRIPTION

[0018] In order to explain the technical content, achieved objectives and effects of the present invention in detail, the following is an explanation in conjunction with the implementation modes and the accompanying drawings.

[0019] Please refer to Figure 1 , a data value assessment and analysis method based on data attribute analysis, comprising:

[0020] Step S1: Calculate and obtain the first data asset value of a single attribute data set of a single data attribute in multiple different application scenarios according to a data value evaluation and analysis algorithm;

[0021] Step S2: obtaining a set of data attributes required in different application scenarios, and combining all data attributes in the set of data attributes into a plurality of data attribute subsets, each of which is associated with a corresponding application scenario and includes at least two of the data attributes;

[0022] Step S3: taking the data sets of all data attributes under the same data attribute subset as a multi-attribute data set, and calculating and obtaining the second data asset value of each multi-attribute data set in the associated application scenario according to the data value evaluation and analysis algorithm;

[0023] Step S4: Obtain a data set with the highest data asset value in the current application scenario according to the first data asset value and the second data asset value in the current application scenario.

[0024] From the above description, it can be seen that the beneficial effects of the present invention are: according to the data value evaluation analysis algorithm, the first data asset value of a single attribute data set of a single data attribute in multiple different application scenarios is calculated and obtained, and the second data asset value of each multi-attribute data set in the associated application scenario is calculated and obtained, and the data set with the highest data asset value in the current application scenario is obtained according to the first data asset value and the second data asset value in the current application scenario. Therefore, not only the data asset value of each single data attribute is taken into account, but also whether the combination of multiple data attributes can produce a higher data asset value is evaluated, so as to mine a data set with a higher data asset value to reflect the true value of the data assets of each data attribute, which not only indirectly increases the value of the data assets, but also realizes the accurate value evaluation of the data assets.

[0025] Furthermore, the calculation process of the data value assessment analysis algorithm specifically includes:

[0026] Step S11, traversing all data of the first data set, obtaining the number of missing data fields, the number of data fields that do not comply with the corresponding data attribute regulations, and whether the data field values ​​on the matching associated items of all data tables are consistent, thereby obtaining the integrity value, the validity value, and the consistency value in sequence, wherein the first data set is a single-attribute data set or a multi-attribute data set;

[0027] Step S12: Obtain all professional data related to data value evaluation and analysis in academic papers, academic journals, and published patents. Screen out the first professional data containing integrity, validity, and consistency from all the professional data. Uniformly convert them into relative proportion relationships with a sum of 1 according to the specific proportion relationships of integrity, validity, and consistency in the first professional data. After accumulating all the relative proportion relationships corresponding to integrity, validity, and consistency, obtain the quality weight ratio among the three. Calculate the data quality score of the first data set based on the integrity value, validity value, and consistency value of the first data set according to the quality weight ratio.

[0028] Step S13: According to the number of data sources and data update timeliness of different data attributes in the first data set, the consumption data of the first application scenario, and the ratio between the types of data attributes in the first data set and the data attributes required by the first application scenario, sequentially obtain the rarity value, timeliness value, consumption value, and feasibility value. The first application scenario is any one of multiple different application scenarios.

[0029] Step S14: Screen out the second professional data containing rarity, timeliness, consumption, and feasibility from all the professional data. Uniformly convert them into relative proportion relationships with a sum of 1 according to the specific proportion relationships of rarity, timeliness, consumption, and feasibility in the second professional data. After accumulating all the relative proportion relationships corresponding to rarity, timeliness, consumption, and feasibility, obtain the scenario weight ratio among the four. Calculate the data scenario score of the first data set based on the rarity value, timeliness value, consumption value, and feasibility value of the first data set according to the scenario weight ratio.

[0030] Step S15: Take the product of the data quality score and the data scenario score as the data asset value.

[0031] As can be seen from the above description, different dimensional data in terms of data quality and application scenarios are calculated. By retrieving and analyzing the weight ratios of each dimensional data from all professional data related to data value evaluation and analysis in academic papers, academic journals, and published patents, the setting of the weight ratios is made more accurate, so as to achieve accurate value evaluation of data assets.

[0032] Further, the specific method for screening out the first professional data containing the integrity, the validity, and the consistency from all the professional data is as follows: screening out the first professional data that contains at least two of the integrity, the validity, and the consistency from all the professional data;

[0033] The specific method for screening out the second professional data containing rarity, timeliness, consumability, and feasibility from all the professional data is as follows: screening out the second professional data that contains at least two of rarity, timeliness, consumability, and feasibility from all the professional data;

[0034] In the conversion of the specific proportional relationship to a relative proportional relationship with a sum of 1, if a certain property does not exist, it is counted as 0 for calculation.

[0035] As can be seen from the above description, more than half inclusion has a certain reference value, so as to increase the data volume to ensure the more accurate setting of the weight ratio.

[0036] Further, the quality weight ratios of the integrity, the validity, and the consistency, and the scenario weight ratios of the rarity, the timeliness, the consumability, and the feasibility are respectively given a value range in advance by the expert side;

[0037] If each weight ratio obtained in step S12 is within the corresponding value range, then calculate the data quality score of the first data set according to the quality weight ratio for the integrity value, the validity value, and the consistency value of the first data set, otherwise send each generated weight ratio to the expert side;

[0038] If each weight ratio obtained in step S14 is within the corresponding value range, then calculate the data scenario score of the first data set according to the scenario weight ratio for the rarity value, the timeliness value, the consumability value, and the feasibility value of the first data set, otherwise send each generated weight ratio to the expert side.

[0039] As can be seen from the above description, by setting a value range to restrict the results of big data analysis, artificial means are used to avoid all possible "deviations" in machine learning and ensure the accuracy of the weight ratio.

[0040] Further, there are multiple expert sides, and the value range is obtained through discussion and negotiation by multiple expert sides.

[0041] Further, the step of sending each generated weight ratio to the expert side in step S14 specifically includes the following steps:

[0042] Each generated weight ratio and a plurality of professional data closest to the generated weight ratio are sent to the expert end.

[0043] From the above description, it can be seen that when the weight ratio exceeds the set range of values, that is, there is a dispute between manual constraints and machine learning, multiple professional data in machine learning that are closest to the weight ratio generated will be sent to the expert side. After reading the relevant professional data, the expert side will determine whether the weight ratio is reasonable and reliable, so that the weight ratio is set jointly by humans and machines to ensure the accuracy of the weight ratio.

[0044] Furthermore, the step S4 specifically includes:

[0045] According to the first data asset value and the second data asset value that are greater than the data cost in the current application scenario, a data set with the highest data asset value in the current application scenario is obtained.

[0046] Furthermore, the sum of the quality weight ratios of the completeness, the effectiveness and the consistency is 1, and the sum of the scenario weight ratios of the rarity, the timeliness, the consumability and the feasibility is 1.

[0047] Furthermore, before step S1, the following steps are also included:

[0048] Perform metadata management on the original data and use the obtained metadata as a data set.

[0049] Please refer to Figure 2 A data value assessment and analysis device based on data attribute analysis includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned data value assessment and analysis method based on data attribute analysis are implemented.

[0050] From the above description, it can be seen that the beneficial effects of the present invention are: according to the data value evaluation analysis algorithm, the first data asset value of a single attribute data set of a single data attribute in multiple different application scenarios is calculated and obtained, and the second data asset value of each multi-attribute data set in the associated application scenario is calculated and obtained, and the data set with the highest data asset value in the current application scenario is obtained according to the first data asset value and the second data asset value in the current application scenario. Therefore, not only the data asset value of each single data attribute is taken into account, but also whether the combination of multiple data attributes can produce a higher data asset value is evaluated, so as to mine a data set with a higher data asset value to reflect the true value of the data assets of each data attribute, which not only indirectly increases the value of the data assets, but also realizes the accurate value evaluation of the data assets.

[0051] Please refer to Figure 1 , Embodiment 1 of the present invention is:

[0052] Data value assessment and analysis methods based on data attribute analysis include:

[0053] Step S0: Perform metadata management on the original data and use the obtained metadata as a data set.

[0054] That is, all collected data are converted into metadata to facilitate subsequent statistical analysis.

[0055] Step S1: Calculate and obtain the first data asset value of a single attribute data set of a single data attribute in multiple different application scenarios according to a data value evaluation and analysis algorithm;

[0056] In this embodiment, the data attribute refers to what the data represents, for example, the data attribute is purchased items, consumption amount, location information, etc. In step S1, only a single data attribute is evaluated for assets.

[0057] Step S2: obtaining a data attribute set required in different application scenarios, and combining all data attributes in the data attribute set into multiple data attribute subsets, each data attribute subset being associated with a corresponding application scenario and including at least two data attributes;

[0058] In this embodiment, there are two application scenarios, A and B, and the data attributes include purchased items, consumption amount, and location information. Application scenario A requires two data attributes, purchased items and consumption amount, and thus has only one data attribute subset, while application scenario B requires three data attributes, purchased items, consumption amount, and location information, and thus has four data attribute subsets.

[0059] Step S3: taking the data sets of all data attributes under the same data attribute subset as a multi-attribute data set, and calculating and obtaining the second data asset value of each multi-attribute data set in the associated application scenario according to the data value evaluation and analysis algorithm;

[0060] Therefore, we evaluate whether the combination of multiple data attributes can produce higher data asset value, thereby mining data sets with higher data asset value to reflect the true value of data assets of each data attribute, which not only indirectly increases the value of data assets, but also realizes accurate value assessment of data assets.

[0061] Step S4: Obtain the data set with the highest data asset value in the current application scenario according to the first data asset value and the second data asset value that are greater than the data cost in the current application scenario.

[0062] That is, the data cost is mainly the hardware cost during data storage. If the value of the data asset is less than the data cost, there is no need to store it and it can be directly discarded.

[0063] Please refer to Figure 1 , Example 1 of the present invention is:

[0064] Based on the data value evaluation and analysis method of data attribute analysis, on the basis of the above Example 1, the calculation process of the data value evaluation and analysis algorithm specifically includes:

[0065] Step S11: Traverse all the data in the first data set to obtain the number of missing data fields, the number of data fields that do not conform to the corresponding data attribute regulations, and whether the data field values on the matching association items of all data tables are consistent, so as to obtain the integrity value, validity value, and consistency value in sequence. The first data set is a single-attribute data set or a multi-attribute data set;

[0066] In this embodiment, if a data field is missing, it is incomplete. Therefore, the ratio of the quantity of this part of the data subtracted to the quantity of all data is used as the integrity value, which is 95% in this embodiment. At the same time, the validity value and the consistency value are 90% and 96% respectively.

[0067] Step S12: Obtain all the professional data related to data value evaluation and analysis in academic papers, academic journals, and published patents. Screen out the first professional data containing integrity, validity, and consistency from all the professional data. According to the specific proportional relationship of integrity, validity, and consistency in the first professional data, uniformly convert them into relative proportional relationships with the sum being 1, and accumulate all the relative proportional relationships corresponding to integrity, validity, and consistency to obtain the quality weight ratio among the three. Calculate the data quality score of the first data set based on the quality weight ratio for the integrity value, validity value, and consistency value of the first data set;

[0068] Among them, the sum of the quality weight ratios of integrity, validity, and consistency is 1.

[0069] In this embodiment, first professional data that includes at least two of integrity, validity, and consistency is screened out from all professional data. That is, any two properties existing in the professional data are regarded as the first professional data. In the subsequent specific proportional relationship, when converting to a relative proportional relationship with a sum of 1, if a certain property does not exist, it is counted as 0 for calculation. For example, in a certain professional data, there are two properties, integrity and validity, which are 100 and 50 respectively in a certain professional data. After converting to 1, they become 2 / 3 and 1 / 3, and the consistency is 0 because consistency is not included in this professional data. This indicates that for the author of this professional data, the weight ratio of consistency is not important enough to be equivalent to 0.

[0070] In this embodiment, the quality weight ratios of integrity, validity, and consistency, as well as the scenario weight ratios of rarity, timeliness, consumability, and feasibility, are respectively given a value range by the expert side in advance. That is, if each weight ratio obtained in step S12 is within the corresponding value range, the integrity value, validity value, and consistency value of the first data set are calculated according to the quality weight ratio to obtain the data quality score of the first data set. Otherwise, each generated weight ratio and multiple pieces of professional data closest to the generated weight ratio are sent to the expert side together.

[0071] Among them, there are at least three expert sides, and the value range is obtained through joint discussion by multiple expert sides. After at least three expert sides each give a value range, they send the value ranges of different expert sides to all expert sides. After all expert sides receive the value ranges of different expert sides and then conduct joint discussion to redefine the value range, and so on, until at least all expert sides jointly discuss and come up with a unified value range.

[0072] In this embodiment, the value ranges of integrity, validity, and consistency are 30%-45%, 30%-45%, and 10%-25% respectively. The weight ratios of the three obtained from the perceived professional data are 39%, 41%, and 20% respectively. Since all three are within the value range, the above weight ratios are used for subsequent calculations. At this time, the integrity value, validity value, and consistency value are 95%, 90%, and 96% respectively. Then, the calculated data quality score = 95% * 39% + 90% * 41% + 96% * 20% = 37.05% + 36.9% + 19.2% = 93.15%.

[0073] Thus, by retrieving and analyzing the weight ratios of data in each dimension from all professional data related to data value evaluation and analysis in academic papers, academic journals, and publicly disclosed patents, the setting of the weight ratio can be made more accurate.

[0074] Step S13: Based on the number of data sources with different data attributes and the data update timeliness in the first dataset, the consumption data of the first application scenario, and the ratio between the types of data attributes in the first dataset and the data attributes required by the first application scenario, the rarity value, timeliness value, consumption value, and feasibility value are obtained in sequence. The first application scenario is any one of multiple different application scenarios;

[0075] Thus, referring to the above Step S11, the rarity value, timeliness value, consumption value, and feasibility value obtained in Step S13 are 60%, 40%, 25%, and 50% respectively.

[0076] Step S14: Screen out the second professional data containing rarity, timeliness, consumption, and feasibility from all professional data. According to the specific proportional relationship of rarity, timeliness, consumption, and feasibility in the second professional data, they are uniformly converted into a relative proportional relationship with the sum being 1, and the cumulative sum of all relative proportional relationships corresponding to rarity, timeliness, consumption, and feasibility is obtained to get the scenario weight ratio among the four. Based on the scenario weight ratio, calculate the data scenario score of the first dataset for the rarity value, timeliness value, consumption value, and feasibility value of the first dataset;

[0077] Among them, the sum of the scenario weight ratios of rarity, timeliness, consumption, and feasibility is 1.

[0078] Among them, if each weight ratio obtained in Step S14 is within the corresponding value range, calculate the data scenario score of the first dataset for the rarity value, timeliness value, consumption value, and feasibility value of the first dataset based on the scenario weight ratio, otherwise send each generated weight ratio to the expert side.

[0079] In this embodiment, screen out the second professional data containing at least two of rarity, timeliness, consumption, and feasibility from all professional data, specifically referring to the detailed introduction in Step S12.

[0080] Thus, referring to the above Step S12, the weight ratios of rarity, timeliness, consumption, and feasibility obtained in Step S14 are 15%, 20%, 40%, and 25% respectively. Then, calculate the data scenario score = 60% * 15% + 40% * 20% + 25% * 40% + 50% * 25% = 9% + 8% + 10% + 12.5% = 39.5%.

[0081] Step S15: Take the product of the data quality score and the data scenario score as the data asset value.

[0082] Thus, the data asset value of the above specific implementation is 39.5% * 93.15% = 36.8%.

[0083] Please refer to Figure 1 , Embodiment 3 of the present invention is as follows:

[0084] The data value evaluation and analysis method based on data attribute analysis. On the basis of the above Embodiment 2, step S4 specifically includes the following:

[0085] Step S41: Obtain the first single-attribute data set corresponding to the first data asset value greater than the data cost and the first multi-attribute data set corresponding to the second data asset value greater than the data cost in the current application scenario;

[0086] Among them, the data cost includes the data original cost required for each data attribute from acquisition to storage and the data sales cost for each sold data set. Among them, the data cost corresponding to the multi-attribute data set is naturally the sum of the data original costs of all data attributes included plus one data sales cost.

[0087] Step S42: After deducting the data cost from the first data asset value of all the first single-attribute data sets, obtain the first data asset net profit. After deducting the data cost from the second data asset value of the first multi-attribute data set, obtain the second data asset net profit. Combine all the first single-attribute data sets and the first multi-attribute data sets according to the principle of at most non-repetition to obtain a data asset sale combination and calculate the total profit of the data asset sale combination based on the corresponding data asset net profit. Among them, the principle of at most non-repetition means that the number of data attributes included in the data asset sale combination is the theoretical maximum attribute value and all data attributes in the data asset sale combination are uniquely stored;

[0088] Among them, since a part of the data set below the data cost is filtered in step S41, not all data asset sale combinations can include all data attributes. We assume that the data attributes include purchased items, consumption amount, and location information. Among them, the first data asset value of the consumption amount is less than the data cost, and the first data asset values of the remaining purchased items and location information are both greater than the data cost. At the same time, the second data asset values of the 4 multi-attribute data sets of the combination of purchased items, consumption amount, and location information are also greater than the data cost. That is, there are purchased items, location information, purchased items + consumption amount, purchased items + location information, consumption amount + location information, and purchased items + consumption amount + location information. Then, the purchased items + consumption amount + location information itself is used as a data asset sale combination. In addition, purchased items and consumption amount + location information, location information and purchased items + consumption amount, and purchased items + location information are also three data asset sale combinations respectively, totaling four groups. Among them, the purchased items + location information does not include the three data attributes because the first data asset value of the consumption amount is less than the data cost, but two are also the theoretical maximum attribute values that can be combined from this data set.

[0089] Step S43: Use the data asset sale combination with the highest total profit as the data set collection with the highest data asset value in the current application scenario.

[0090] Therefore, among the above four groups, the total profit of location information and purchased items + consumption amount is the highest. That is, the location information is sold individually, and the purchased items + consumption amount are sold bundled to ensure the maximization of the data asset value.

[0091] In other embodiments, considering the negative impact brought by multiple separate sales in the same application scenario, the number of data sets in the data asset sale combination can be limited to at most two or three.

[0092] Please refer to Figure 2 , Embodiment 4 of the present invention is:

[0093] The data value evaluation and analysis device 1 based on data attribute analysis includes a memory 3, a processor 2, and a computer program stored on the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, it implements the steps of the above Embodiment 1 or 2 or 3.

[0094] In summary, the data value evaluation and analysis method and device based on data attribute analysis provided by the present invention set different dimensional data on data quality and application scenarios for calculation, retrieve and analyze the weight ratio of each dimensional data in academic papers, academic journals and all professional data related to data value evaluation and analysis in public patents, and artificially constrain the value range provided by experts to obtain a more accurate data value evaluation and analysis algorithm, calculate and obtain the first data asset value of a single attribute data set of a single data attribute in multiple different application scenarios according to the data value evaluation and analysis algorithm, and calculate and obtain the second data asset value of each multi-attribute data set in the associated application scenario, and obtain the data set with the highest data asset value in the current application scenario according to the first data asset value and the second data asset value in the current application scenario. Thus, not only the data asset value of each single data attribute is taken into account, but also the combination of multiple data attributes is evaluated to see whether it can produce a higher data asset value, so as to mine a data set with a higher data asset value to reflect the true value of the data asset of each data attribute, which not only indirectly improves the value of the data asset, but also realizes the accurate value evaluation of the data asset.

[0095] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's specification and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A data value evaluation and analysis method based on data attribute analysis, characterized in that, it includes: Step S1: Calculate and obtain the first data asset value of the single-attribute data set of a single data attribute in multiple different application scenarios according to the data value evaluation and analysis algorithm; The specific calculation process of the data value evaluation and analysis algorithm includes: Step S11: Traverse all the data in the first data set, obtain the number of missing data fields, the number of data fields that do not conform to the corresponding data attribute regulations, and whether the data field values on the matching association items of all data tables are consistent, so as to obtain the integrity value, the validity value, and the consistency value in sequence. The first data set is a single-attribute data set or a multi-attribute data set; Step S12: Obtain all the professional data related to data value evaluation and analysis in academic papers, academic journals, and published patents. Screen out the first professional data containing integrity, validity, and consistency from all the professional data. According to the specific proportional relationship of integrity, validity, and consistency in the first professional data, uniformly convert them with a sum of 1 into relative proportional relationships, and accumulate all the relative proportional relationships corresponding to integrity, validity, and consistency to obtain the quality weight ratio among the three. Calculate the data quality score of the first data set according to the quality weight ratio for the integrity value, the validity value, and the consistency value of the first data set; Step S13: According to the number of data sources and data update timeliness of different data attributes in the first data set, the consumption data of the first application scenario, and the ratio between the types of data attributes in the first data set and the data attributes required by the first application scenario, obtain the rarity value, the timeliness value, the consumption value, and the feasibility value in sequence. The first application scenario is any one of multiple different application scenarios; Step S14: Screen out the second professional data containing rarity, timeliness, consumption, and feasibility from all the professional data. According to the specific proportional relationship of rarity, timeliness, consumption, and feasibility in the second professional data, uniformly convert them with a sum of 1 into relative proportional relationships, and accumulate all the relative proportional relationships corresponding to rarity, timeliness, consumption, and feasibility to obtain the scenario weight ratio among the four. Calculate the data scenario score of the first data set according to the scenario weight ratio for the rarity value, the timeliness value, the consumption value, and the feasibility value of the first data set; Step S15: Take the product of the data quality score and the data scenario score as the data asset value; Step S2: Obtain the data attribute set required in different application scenarios, combine all the data attributes in the data attribute set into multiple data attribute subsets, and each data attribute subset is associated with the corresponding application scenario and includes at least two data attributes; Step S3: Use the datasets of all data attributes under the same data attribute subset as a multi-attribute dataset, and calculate and obtain the second data asset value of each multi-attribute dataset in the associated application scenario according to the data value evaluation and analysis algorithm; Step S4: Obtain the dataset with the highest data asset value in the current application scenario based on the first data asset value and the second data asset value in the current application scenario.

2. The data value evaluation and analysis method based on data attribute analysis according to claim 1, characterized in that, The specific method of screening out the first professional data containing the integrity, the validity and the consistency from all the professional data is: screening out the first professional data containing at least two of the integrity, the validity and the consistency from all the professional data; The specific method of screening out the second professional data containing rarity, timeliness, consumability and feasibility from all the professional data is: screening out the second professional data containing at least two of rarity, timeliness, consumability and feasibility from all the professional data; In the specific proportional relationship unified with the sum being 1 for conversion into the relative proportional relationship, if a certain property does not exist, it is counted as 0 for calculation.

3. The data value evaluation and analysis method based on data attribute analysis according to claim 1, characterized in that, The quality weight ratios of the integrity, the validity and the consistency, and the scenario weight ratios of the rarity, the timeliness, the consumability and the feasibility are respectively given a value range by the expert side in advance; If each weight ratio obtained in step S12 is within the corresponding value range, calculate the data quality score of the first dataset according to the quality weight ratios for the integrity value, the validity value and the consistency value of the first dataset, otherwise send each generated weight ratio to the expert side; If each weight ratio obtained in step S14 is within the corresponding value range, calculate the data scenario score of the first dataset according to the scenario weight ratios for the rarity value, the timeliness value, the consumability value and the feasibility value of the first dataset, otherwise send each generated weight ratio to the expert side.

4. The data value evaluation and analysis method based on data attribute analysis according to claim 3, characterized in that, There are multiple expert sides, and the value range is obtained through discussion and deliberation by multiple expert sides.

5. The data value evaluation and analysis method based on data attribute analysis according to claim 4, characterized in that, The specific steps of sending each generated weight ratio to the expert side in step S14 specifically include the following steps: Send each generated weight ratio and multiple pieces of professional data closest to the generated weight ratio to the expert side.

6. The data value evaluation and analysis method based on data attribute analysis according to claim 1, characterized in that, Step S4 specifically includes: Obtain the data set with the highest data asset value in the current application scenario based on the first data asset value and the second data asset value that are greater than the data cost in the current application scenario.

7. The data value evaluation and analysis method based on data attribute analysis according to claim 1, wherein, the sum of the quality weight ratios of the integrity, the validity, and the consistency is 1, and the sum of the scenario weight ratios of the rarity, the timeliness, the consumability, and the feasibility is 1.

8. The data value evaluation and analysis method based on data attribute analysis according to claim 1, wherein, before the step S1, the following steps are further included: Perform metadata management on the original data, and use the obtained metadata as the data set.

9. A data value evaluation and analysis device based on data attribute analysis, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the steps in the data value evaluation and analysis method based on data attribute analysis according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Risk assessment method and system based on correlation analysis

    CN110401625A

  • Intelligence use value evaluation method

    CN111667072A