Data cost assessment method, electronic device, storage medium and program product

By obtaining the metadata information and blood relationship diagram of the data node, combining physical cost and target scoring information, the problem of inaccurate data cost assessment in the prior art is solved, and a more accurate data cost assessment is achieved.

CN120011340APending Publication Date: 2025-05-16CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510073024.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art only considers the metadata information of the data itself in data cost assessment, resulting in inaccurate cost assessment results.

Method used

By obtaining the metadata information and blood relationship diagram of each data node in the data to be evaluated, the physical cost information and target scoring information of each data node are determined, and cost evaluation is carried out in combination with multiple target scoring information.

Benefits of technology

Improve the accuracy of data cost assessment, and comprehensively evaluate the importance and cost of data by taking into account the physical cost and blood relationship of the data node.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011340A_ABST
    Figure CN120011340A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data cost evaluation method, electronic equipment, a storage medium and a program product. The method comprises the steps of obtaining metadata information and a blood relationship graph of each data node in to-be-evaluated data; determining physical cost information of each data node according to the metadata information of each data node; determining target score information of each data node according to the physical cost information of each data node and the blood relationship graph, wherein the target score information of each data node is used for representing the importance degree of each data node; and according to the multiple pieces of target scoring information, determining cost evaluation information of the to-be-evaluated data. The method is used for achieving the technical effect of improving the data cost evaluation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data cost assessment, and in particular to a data cost assessment method, electronic equipment, storage medium and program product. Background Art

[0002] In the field of data governance and data cost accounting, with the rapid increase in data volume and the increase in data management complexity, how to accurately calculate and manage data storage and computing costs has become a highly concerned issue.

[0003] The storage and cost calculation methods for data in the prior art mainly calculate the asset value of the data based on the metadata information of the data and a predefined calculation model.

[0004] However, in the prior art, when calculating the cost of data, only the data itself is considered. Therefore, the prior art has a technical problem that the data cost evaluation result is inaccurate. Summary of the invention

[0005] The embodiments of the present application provide a data cost assessment method, an electronic device, a storage medium, and a program product to achieve the technical effect of improving the accuracy of data cost assessment.

[0006] In a first aspect, an embodiment of the present application provides a data cost assessment method, including:

[0007] Obtain metadata information and lineage relationship diagram for each data node in the data to be evaluated;

[0008] Determine physical cost information of each data node according to metadata information of each data node;

[0009] Determine the target score information of each data node according to the physical cost information and the blood relationship diagram of each data node. The target score information of each data node is used to indicate the importance of each data node.

[0010] Cost assessment information of the data to be assessed is determined based on multiple target scoring information.

[0011] In a possible implementation, determining the physical cost information of each data node according to the metadata information of each data node includes:

[0012] Determine computing cost information and storage cost information of each data node according to metadata information of each data node;

[0013] The physical cost information of each data node is determined according to the computational cost information and the storage cost information.

[0014] In a possible implementation, determining the computing cost information and storage cost information of each data node according to the metadata information of each data node includes:

[0015] The metadata information of each data node is input into the calculation cost formula to obtain the calculation cost information of each data node. The expression of the calculation cost formula is:

[0016]

[0017] Among them, C 1 To calculate cost information; n cpu is the number of CPU cores occupied by the data node during the computing process; C cpu is the unit price of CPU core; mem The amount of memory occupied by the data node during the calculation process; C mem is the unit price of memory; b read b is the network bandwidth occupied by the read operation of the data node during the calculation process; write C is the network bandwidth occupied by the write operation of the data node during the computing process; b Unit price of network bandwidth; c The time that the data node takes in the computing operation;

[0018] The metadata information of each data node is input into the storage cost formula to obtain the storage cost information of each data node. The expression of the storage cost formula is:

[0019]

[0020] Among them, C 2 To calculate cost information; S mem_s It is the memory value occupied by the storage data node; ts is the duration of storing the data node.

[0021] In a possible implementation, determining target score information of each data node according to the physical cost information and the lineage relationship diagram of each data node includes:

[0022] Step a, setting the physical cost information of each data node as the initial scoring information of each data node;

[0023] Step b, determining the output data node of each data node according to the blood relationship diagram of each data node, and determining the node information between each data node and each output data node; wherein the number of output data nodes is at least 1;

[0024] Step c: inputting the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain first score information;

[0025] Step d, calculating the difference information between the initial scoring information and the first scoring information, and determining the first scoring information as new initial scoring information;

[0026] Step e: If the difference information does not meet the preset threshold, redetermine the new node information between each data node and each output data node, determine the new node information as the node information, repeat steps c to d until the difference information meets the preset threshold, and determine the obtained first scoring information as the target scoring information.

[0027] In a possible implementation, the score update formula is expressed as:

[0028]

[0029] Among them, PR is the first scoring information of each data node; k is the sum of the physical cost information of all data nodes; w is the coefficient, and the value of w is less than or equal to 1; d is the damping factor, d = 0.85; A is the first node information of each data node and each output node; B is the second node information of each data node and all output nodes; A / B is the third node information of each output data node; among them, the number of output data nodes is j.

[0030] In a possible implementation, the method further includes:

[0031] Determine the output edges between each data node and each output data node; wherein the number of output edges is at least 1;

[0032] According to the blood relationship graph of each data node, determine the operation type and frequency information of each output edge;

[0033] Determine the weight information of each output edge according to the operation type of each output edge;

[0034] Determine output edge information of each output edge according to weight information and frequency information of each output edge;

[0035] Determine first node information of each data node and each output data node according to the plurality of output edge information;

[0036] Determine fourth node information of each data node and each output data node according to the weight information of each output edge;

[0037] According to the plurality of fourth node information, each data node information and the second node information of the output data node are determined.

[0038] In a second aspect, an embodiment of the present application provides a data cost assessment device, including:

[0039] The acquisition module is used to obtain the metadata information and blood relationship diagram of each data node in the data to be evaluated;

[0040] A first processing module, configured to determine physical cost information of each data node according to metadata information of each data node;

[0041] A second processing module is used to determine target score information of each data node according to the physical cost information and the blood relationship diagram of each data node, and each data node target score information is used to indicate the importance of each data node;

[0042] The third processing module is used to determine the cost evaluation information of the data to be evaluated according to the multiple target scoring information.

[0043] The first processing module is also used for:

[0044] Determine computing cost information and storage cost information of each data node according to metadata information of each data node;

[0045] The physical cost information of each data node is determined according to the computational cost information and the storage cost information.

[0046] The first processing module is also used for:

[0047] The metadata information of each data node is input into the calculation cost formula to obtain the calculation cost information of each data node. The expression of the calculation cost formula is:

[0048]

[0049] Among them, C 1 To calculate cost information; n cpu is the number of CPU cores occupied by the data node during the computing process; C cpu is the unit price of CPU core; mem The amount of memory occupied by the data node during the calculation process; C mem is the unit price of memory; b read b is the network bandwidth occupied by the read operation of the data node during the calculation process; write C is the network bandwidth occupied by the write operation of the data node during the computing process; b Unit price of network bandwidth; c The time that the data node takes in the computing operation;

[0050] The metadata information of each data node is input into the storage cost formula to obtain the storage cost information of each data node. The expression of the storage cost formula is:

[0051]

[0052] Among them, C 2 To calculate cost information; S mem_s It is the memory value occupied by the storage data node; ts is the duration of storing the data node.

[0053] The second processing module is also used for:

[0054] Step a, setting the physical cost information of each data node as the initial scoring information of each data node;

[0055] Step b, determining the output data node of each data node according to the blood relationship diagram of each data node, and determining the node information between each data node and each output data node; wherein the number of output data nodes is at least 1;

[0056] Step c: inputting the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain first score information;

[0057] Step d, calculating the difference information between the initial scoring information and the first scoring information, and determining the first scoring information as new initial scoring information;

[0058] Step e: If the difference information does not meet the preset threshold, redetermine the new node information between each data node and each output data node, determine the new node information as the node information, repeat steps c to d until the difference information meets the preset threshold, and determine the obtained first scoring information as the target scoring information.

[0059] The second processing module is also used for:

[0060]

[0061] Among them, PR is the first scoring information of each data node; k is the sum of the physical cost information of all data nodes; w is the coefficient, and the value of w is less than or equal to 1; d is the damping factor, d = 0.85; A is the first node information of each data node and each output node; B is the second node information of each data node and all output nodes; A / B is the third node information of each output data node; among them, the number of output data nodes is j.

[0062] The second processing module is also used for:

[0063] Determine the output edges between each data node and each output data node; wherein the number of output edges is at least 1;

[0064] According to the blood relationship graph of each data node, determine the operation type and frequency information of each output edge;

[0065] Determine the weight information of each output edge according to the operation type of each output edge;

[0066] Determine output edge information of each output edge according to weight information and frequency information of each output edge;

[0067] Determine first node information of each data node and each output data node according to the plurality of output edge information;

[0068] Determine fourth node information of each data node and each output data node according to the weight information of each output edge;

[0069] According to the plurality of fourth node information, each data node information and the second node information of the output data node are determined.

[0070] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0071] The memory stores computer-executable instructions;

[0072] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0073] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.

[0074] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0075] The data cost assessment method, electronic device, storage medium and program product provided by the embodiments of the present application obtain metadata information and a lineage relationship diagram of each data node in the data to be evaluated, determine the physical cost information of each data node through the metadata information, and determine the target score information of each data node based on the physical cost information and the lineage relationship diagram, which information represents the importance of each data node. The target score information of each data node is added to obtain the cost assessment information of the data to be evaluated. This method is not only based on the metadata information of the data to be evaluated, but also combines the lineage relationship between the data to be evaluated and other data to obtain the cost assessment information of the data to be evaluated, thereby making up for the technical defect that data cost assessment only considers metadata, and achieves the technical effect of improving the accuracy of data cost assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0077] Figure 1 Schematic diagram of the data cost evaluation process provided by this application Figure 1 ;

[0078] Figure 2 Schematic diagram of the data cost evaluation process provided by this application Figure 2 ;

[0079] Figure 3 Schematic diagram of the data cost evaluation process provided by this application Figure 3 ;

[0080] Figure 4 A schematic diagram of the structure of the data cost assessment provided for this application;

[0081] Figure 5 Hardware diagram for data cost evaluation provided for this application.

[0082] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0083] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of methods and methods consistent with some aspects of the present application as detailed in the appended claims.

[0084] First, let me explain the nouns:

[0085] Metadata: Metadata refers to data about data. It contains information that describes the storage, calculation, and use of data, such as the data's storage location, the amount of space it occupies, and the usage of computing resources.

[0086] Bloodline diagram: also known as data lineage diagram, data origin diagram or data genealogy diagram, refers to a graphical representation of the relationship similar to human bloodline that is naturally formed between data during the entire life cycle of data, from data generation, processing, processing, fusion, circulation to final extinction.

[0087] Data cost assessment: This application refers to the cost assessment of resources such as storage space and computing space occupied by data during operation. The assessment results obtained can reflect the cost of resources occupied by the data to be evaluated.

[0088] With the rapid development of information technology, especially the widespread application of technologies such as big data and artificial intelligence, the cost assessment of data operation, storage and computing has become increasingly important. Data has become an important asset for enterprises, and effective cost assessment can help enterprises optimize data management, improve decision-making efficiency, and gain an advantage in the fierce market competition.

[0089] In the existing technology, the value assessment and cost accounting of data assets can be achieved through different technical means and methods. These methods usually rely on technical means such as metadata collection, static evaluation models or replacement cost calculation. However, in actual applications, the storage cost and computing cost of data in operation are not only related to the attributes of the data itself, but also closely related to the flow and transformation of data in the entire business process and the blood relationship between data and other data.

[0090] In order to solve the technical problems of the prior art, the data cost evaluation method, electronic device, storage medium and program product provided by the present application obtain the metadata information and lineage relationship diagram of each data node in the data to be evaluated, determine the physical cost information of each data node through the metadata information, and determine the target score information of each data node based on the physical cost information and the lineage relationship diagram. This information represents the importance of each data node. The target score information of each data node is added to obtain the cost evaluation information of the data to be evaluated. This method not only relies on the metadata information of the data to be evaluated, but also combines the lineage relationship between the data to be evaluated and other data, thereby comprehensively improving the accuracy of data cost evaluation.

[0091] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0092] Figure 1 Schematic diagram of the data cost assessment process provided for this application Figure 1 ,like Figure 1 As shown, the method includes:

[0093] S101, obtaining metadata information and a blood relationship diagram of each data node in the data to be evaluated;

[0094] In this embodiment, according to a pre-established metadata management platform, the data to be evaluated includes multiple data nodes, and metadata information of each data node in the data to be evaluated is obtained. At the same time, according to a pre-established blood relationship management platform, the corresponding blood relationship graph of each data node in the data to be evaluated is obtained.

[0095] In one possible implementation, a pre-established metadata management platform is responsible for metadata collection, management, storage, and calculation of metadata information. The metadata management platform in this application mainly includes three types of metadata information. The first type is metadata obtained by directly accessing external systems, namely direct metadata, which includes the Hive table schema, the Hive table life cycle, and the Hive data portrait; the second type of metadata information is obtained by using external systems and then obtaining it through complex calculations, namely indirect metadata, which includes the storage information of the Hive table, the usage of the Hive table, and the partition metadata of the Hive table; the third type of metadata information is obtained through collection tools, and the third type of metadata is obtained through information processing.

[0096] In one possible implementation, the blood relationship management platform can implement fine-grained data pricing strategies for metadata query, impact analysis, traceability analysis, data node importance analysis, data field importance analysis, and data cost calculation steps. The blood relationship management platform of the present application is classified according to the operations of data nodes, with a total of 21 operations; including new operations, modification operations, deletion operations, query operations and other operations. New operations include: creating a new database (CreateDatabase), creating a new data table (CreateTable), copying existing data to a new data table (CreateTableAsSelect) and creating a new view (CreateView); modification operations include: modifying database properties (AlterDatabase), adding partitions to an existing table (AlterTableAddParts), modifying an existing view (AlterViewAs), deleting partitions in an existing table (AlterTableDropParts), adding new columns to an existing table (AlterTableAddCols), replace columns in an existing table (AlterTableReplaceCols), change the data type of a partition column in a table (AlterTablePartColType), perform any type of modification on a table (AlterTable), rename an existing table (AlterTableRename), and rename columns in an existing table (AlterTableRenameCol); deletion operations include deleting the entire database (DropDatabase) and deleting a table in a database (DropTable); query operations include query (Query); other operations include loading (Load), exporting (Export), importing (Import), and clearing a table (TruncateTable).

[0097] In a possible implementation, the blood relationship diagram of data is encapsulated in the form of points and edges. The nodes and edges also contain their own attribute information. The data table node includes 11 attribute information, including table name, table type, creation time, update time, test or generation type, person in charge, whether it is a view, and whether it is deleted; the database node includes 5 attributes, including library name, description statement, storage location, parameters, and whether it is deleted; the HDFS file node includes two attributes, storage location and file name; the Hbase table node includes four attributes, including table name, namespace, URI, and whether it is deleted; the field node includes three attributes, including field type, comment, and whether it is deleted; the partition node includes four attributes, including partition field type, comment, storage location, and whether it is deleted; the storage location node includes two attributes, path and file name; the edge from the data table node to the partition node also has the edge attributes of operation creation time and operation ID; each operation node includes five attributes, including execution type, execution time, number of machine cores used, memory occupied, and user.

[0098] S102, determining physical cost information of each data node according to metadata information of each data node;

[0099] In this embodiment, data production involves the generation, collection, storage, and processing of a large amount of information. Although data is generally regarded as a digital asset, its production also has physical costs. In data production, there are many physical costs associated with the storage, processing, and transmission of data. Calculating the physical cost of each data node through the metadata information of the data to be evaluated is an effective means of evaluating the cost of the data to be evaluated.

[0100] S103, determining target score information of each data node according to the physical cost information and the blood relationship diagram of each data node, where the target score information of each data node is used to indicate the importance of each data node;

[0101] In this embodiment, the target score information of each data node is determined based on the physical cost information and the blood relationship diagram, and this information represents the importance of each data node. This method not only obtains the physical cost of the data based on the metadata information of the data to be evaluated, but also explores the cost of resources occupied by the data in storage, computing, etc. from the perspective of metadata information. At the same time, the blood relationship between the data to be evaluated and other data is considered, and the dependency relationship between the data to be evaluated and other data as well as different data operations is effectively explored, thereby obtaining accurate target score information for each data node, which represents the importance of each data node in the data to be evaluated.

[0102] S104: Determine cost assessment information of the data to be assessed based on multiple target scoring information.

[0103] In this embodiment, the target scoring information corresponding to each data node is added to obtain the cost assessment information of the data to be evaluated. The cost assessment information not only takes into account the resource cost occupied by the data operation process where the metadata information is located, but also fully considers the blood relationship between the data to be evaluated and other data and other data operations, thereby achieving multi-dimensional improvement in the accuracy of the data cost assessment results and achieving the technical effect of accurate and multi-faceted evaluation of data costs.

[0104] The data cost evaluation method provided in the embodiment of the present application obtains metadata information of each data node in the data to be evaluated and the corresponding blood relationship diagram according to a pre-established metadata management platform and a pre-established blood relationship management platform, further calculates the physical cost information of each data node according to the metadata information of the data to be evaluated, and determines the target score information of each data node according to the physical cost information and the blood relationship diagram, which information represents the importance of each data node, and adds the target score information corresponding to each data node to obtain the cost evaluation information of the data to be evaluated. This method not only obtains the physical cost of the data based on the metadata information of the data to be evaluated, but also explores the cost of resources occupied by the data in storage, calculation, etc. from the perspective of metadata information, while also considering the blood relationship between the data to be evaluated and other data, and effectively explores the dependency between the data to be evaluated and other data and different data operations, so as to achieve multi-dimensional improvement of the accuracy of the data cost evaluation results and achieve the technical effect of accurate and multi-faceted evaluation of data costs.

[0105] Figure 2 Schematic diagram of the data cost assessment process provided for this application Figure 2 In this embodiment, Figure 1 Based on the embodiment, the data cost evaluation method is described in detail, wherein the physical cost information of each data node determined according to the metadata information can be implemented through step S202, and the target score information of each data node determined according to the physical cost information and the blood relationship diagram can be implemented through steps S203 to S207. Figure 2 As shown, the data cost evaluation method provided in this embodiment includes:

[0106] S201, obtaining metadata information and a blood relationship diagram of each data node in the data to be evaluated;

[0107] S202, determining the computing cost information and storage cost information of each data node according to the metadata information of each data node; determining the physical cost information of each data node according to the computing cost information and the storage cost information;

[0108] In this embodiment, storage cost information refers to the space and resource consumption occupied by data on the storage device, including but not limited to disk space and maintenance costs of the storage server; the system calculates the storage cost information of each data table based on the first and second types of metadata information of the metadata platform; computing cost information refers to the CPU, memory and bandwidth resources consumed by the data during the computing process; the system calculates the computing cost information of each data node based on the third type of metadata information in the metadata platform; CPU cost refers to the cost of using CPU resources in the data production process, which is usually measured by the number of CPU cores and usage time. Each core represents a processing unit, and the more cores used, the higher the cost. The cost of each core can be defined based on the enterprise's infrastructure costs, licensing fees or other relevant considerations; memory cost refers to the cost associated with using memory resources for data processing, and the memory cost can be based on the amount of memory used and the duration of use; network bandwidth cost refers to the cost associated with data transmission when data is transmitted between different systems, networks or regions. The present invention uses the read and write bandwidth and usage time occupied by the data production process to represent it.

[0109] In a possible implementation, the physical cost information of each data node includes computing cost information and storage cost information. The computing cost information and storage cost information can be obtained respectively by the following formulas. The metadata information of each data node is input into the computing cost formula to obtain the computing cost information of each data node. The expression of the computing cost formula is:

[0110]

[0111] Among them, C 1 To calculate cost information; n cpu is the number of CPU cores occupied by the data node during the computing process; C cpu is the unit price of CPU core; mem The amount of memory occupied by the data node during the calculation process; C mem is the unit price of memory; b read b is the network bandwidth occupied by the read operation of the data node during the calculation process; write C is the network bandwidth occupied by the write operation of the data node during the computing process; b Unit price of network bandwidth; c The time that the data node takes in the calculation operation.

[0112] The metadata information of each data node is input into the storage cost formula to obtain the storage cost information of each data node. The expression of the storage cost formula is:

[0113]

[0114] Among them, C 2 To calculate cost information; S mem_s It is the memory value occupied by the storage data node; ts is the duration of storing the data node.

[0115] S203, setting the physical cost information of each data node as the initial scoring information of each data node;

[0116] In this embodiment, the physical cost information of each data node is used as the initial scoring information of each data node, and this initial scoring information is the data basis for the subsequent calculation of the target scoring information.

[0117] S204, determining the output data node of each data node according to the blood relationship diagram of each data node, and determining the node information between each data node and each output data node;

[0118] In this embodiment, each data node is regarded as a vertex in the corresponding blood relationship graph. The data node has multiple output data nodes, which are connected by edges. The edges represent the dependency relationship between two data nodes. The number of output data nodes is at least 1. Further, the node information between each data node and each output data node is determined.

[0119] Optionally, if the number of output data nodes is 0, it means that the data node has no corresponding output data node, and the target score information of the data node is determined by the physical cost information.

[0120] S205, inputting the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain first score information;

[0121] In this embodiment, the PageRank algorithm is an iterative calculation algorithm based on a graph structure, which is generally used to calculate the importance of a web page. The present application inputs the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain the calculation result of the first iteration, that is, the first score information.

[0122] In a possible implementation, the score update formula is expressed as:

[0123]

[0124] Among them, PR is the first scoring information of each data node; k is the sum of the physical cost information of all data nodes; w is the coefficient, and the value of w is less than or equal to 1; d is the damping factor, d = 0.85; A is the first node information of each data node and each output node; B is the second node information of each data node and all output nodes; A / B is the third node information of each output data node; among them, the number of output data nodes is j.

[0125] S206, calculating the difference information between the initial scoring information and the first scoring information, and determining the first scoring information as new initial scoring information;

[0126] In this embodiment, the convergence condition of the PageRank algorithm is that the PageRank value change of a certain node when the algorithm iterates is less than a preset threshold, so the difference information between the initial scoring information and the first scoring information needs to be calculated, and the first scoring information is determined as the new initial scoring information.

[0127] S207, if the difference information does not meet the preset threshold, then re-determine the new node information between each data node and each output data node, determine the new node information as the node information, repeat steps S205 to S206 until the difference information meets the preset threshold, and determine the obtained first scoring information as the target scoring information;

[0128] In this embodiment, if the difference information does not meet the preset threshold, the new node information between each data node and each output data node is re-determined. This method can realize dynamic adjustment of the node information, abandoning the cost of using fixed node information to calculate data in the prior art, and further determining the new node information as the node information. Steps S205 to S206 are repeated until the difference information satisfies the preset threshold, and the obtained first scoring information is determined as the target scoring information.

[0129] S208. Determine cost assessment information of the data to be assessed based on multiple target scoring information.

[0130] Figure 3 Schematic diagram of the data cost assessment process provided for this application Figure 3 In this embodiment, Figure 2 Based on the embodiment, step S205 is described in detail, wherein obtaining the first node information can be implemented through steps S301 to S304, and obtaining the second node information can be implemented through step S305, such as Figure 3 As shown, the data cost evaluation method provided in this embodiment includes:

[0131] S301, determining an output edge between each data node and each output data node;

[0132] In this embodiment, there is an output edge between each data node and each output data node, and the number of the output edges is at least one.

[0133] Optionally, if the number of output edges is 0, it means that the data node has no corresponding output data node, and the target score information of the data node is determined by the physical cost information.

[0134] S302, determining the operation type and frequency information of each output edge according to the blood relationship diagram of each data node;

[0135] In this embodiment, the operation type and frequency information corresponding to each output edge can be obtained according to the blood relationship graph.

[0136] S303, determining weight information of each output edge according to the operation type of each output edge; determining output edge information of each output edge according to the weight information and frequency information of each output edge;

[0137] In this embodiment, different operation types correspond to different weight information. After determining the weight information of each output edge by determining the operation type of each output edge, the weight information of each output edge and the frequency information are multiplied to obtain the output edge information of each output edge. Optionally, when executing step S207, if the difference information does not meet the preset threshold, the weight information of the output edge is dynamically adjusted according to the operation type, thereby obtaining new output edge information and performing further iterative calculation.

[0138] In one possible implementation, the operation types include the 21 operations in the blood relationship management platform mentioned above. By pre-setting the weight information of each operation, the operation type corresponding to each output edge is determined, thereby determining the weight information of the output edge.

[0139] In a possible implementation, the weight information of the following operations can be set, namely: the weight information of copying existing data to a new data table (CreateTableAsSelect) is 3; the weight information of creating a new view (CreateView) is 2; the weight information of modifying an existing view (AlterViewAs) is 1.5; the weight information of query (Query) is 2; the weight information of renaming an existing table (AlterTableRename) is 1.5; the weight information of renaming a column in an existing table (AlterTableRenameCol) is 1.5; the weight information of the other 15 operations can be 1.

[0140] S304, determining first node information of each data node and each output data node according to the multiple output edge information;

[0141] In this embodiment, multiple output edge information is added to obtain the first node information of each data node and each output data node, that is, the parameter A in the score update formula.

[0142] S305. Determine fourth node information of each data node and each output data node according to the weight information of each output edge; and determine second node information of each data node information and output data node according to the plurality of fourth node information.

[0143] In this embodiment, the weight information of each output edge is added to obtain the fourth node information of each data node and each output data node. Further, all the fourth node information is added to obtain the second node information of each data node information and all output data nodes, which is the parameter B in the score update formula.

[0144] The present application provides a data cost evaluation method, which obtains metadata information of each data node in the data to be evaluated and the corresponding blood relationship graph according to a pre-established metadata management platform and a pre-established blood relationship management platform, and further determines the computing cost information and storage cost information of each data node according to the metadata information of each data node; determines the physical cost information of each data node according to the computing cost information and the storage cost information, and sets the physical cost information of each data node as the initial scoring information of each data node; determines the output data node of each data node according to the blood relationship graph of each data node, and determines the output data node of each data node and each output data node. The method comprises the following steps: first, calculating the node information between each data node and each output data node; inputting the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain the first score information; calculating the difference information between the initial score information and the first score information, and determining the first score information as the new initial score information; if the difference information does not meet the preset threshold, re-determining the new node information between each data node and each output data node, determining the new node information as the node information, and recalculating the new difference information until the difference information meets the preset threshold, determining the obtained first score information as the target score information, and then determining the cost evaluation information of the data to be evaluated based on the multiple target score information. The method can comprehensively consider the direct storage and calculation costs of the data, as well as the indirect costs of the data in different operation scenarios, and more accurately evaluate the data cost by comprehensively integrating the metadata information and blood relationship information of the data. At the same time, through the iterative calculation of the PageRank algorithm, the weight can be dynamically adjusted according to the actual situation of the data node, so that the changes in the data relationship and the operation frequency will directly affect the calculation of the weight, so that the evaluation result is more in line with the actual situation, reducing the subjectivity and uncertainty of human intervention, thereby achieving the technical effect of comprehensively improving the accuracy of data cost evaluation.

[0145] Figure 4 The structural diagram of the data cost evaluation provided for this application is as follows: Figure 4 As shown, the data cost evaluation 40 provided in this embodiment includes:

[0146] The acquisition module 401 is used to acquire metadata information and a blood relationship diagram of each data node in the data to be evaluated;

[0147] A first processing module 402 is used to determine the physical cost information of each data node according to the metadata information of each data node;

[0148] The second processing module 403 is used to determine the target score information of each data node according to the physical cost information and the blood relationship diagram of each data node, and the target score information of each data node is used to indicate the importance of each data node;

[0149] The third processing module 404 is used to determine cost evaluation information of the data to be evaluated according to multiple target scoring information.

[0150] The first processing module 402 is further configured to:

[0151] Determine computing cost information and storage cost information of each data node according to metadata information of each data node;

[0152] The physical cost information of each data node is determined according to the computational cost information and the storage cost information.

[0153] The first processing module 402 is further configured to:

[0154] The metadata information of each data node is input into the calculation cost formula to obtain the calculation cost information of each data node. The expression of the calculation cost formula is:

[0155]

[0156] Among them, C 1 To calculate cost information; n cpu is the number of CPU cores occupied by the data node during the computing process; C cpu is the unit price of CPU core; mem The amount of memory occupied by the data node during the calculation process; C mem is the unit price of memory; b read b is the network bandwidth occupied by the read operation of the data node during the calculation process; write C is the network bandwidth occupied by the write operation of the data node during the computing process; b Unit price of network bandwidth; c The time that the data node takes in the computing operation;

[0157] The metadata information of each data node is input into the storage cost formula to obtain the storage cost information of each data node. The expression of the storage cost formula is:

[0158]

[0159] Among them, C 2 To calculate cost information; S mem_s It is the memory value occupied by the storage data node; ts is the duration of storing the data node.

[0160] The second processing module 403 is further used for:

[0161] Step a, setting the physical cost information of each data node as the initial scoring information of each data node;

[0162] Step b, determining the output data node of each data node according to the blood relationship diagram of each data node, and determining the node information between each data node and each output data node; wherein the number of output data nodes is at least 1;

[0163] Step c: inputting the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain first score information;

[0164] Step d, calculating the difference information between the initial scoring information and the first scoring information, and determining the first scoring information as new initial scoring information;

[0165] Step e: If the difference information does not meet the preset threshold, redetermine the new node information between each data node and each output data node, determine the new node information as the node information, repeat steps c to d until the difference information meets the preset threshold, and determine the obtained first scoring information as the target scoring information.

[0166] The second processing module 403 is further used for:

[0167]

[0168] Among them, PR is the first scoring information of each data node; k is the sum of the physical cost information of all data nodes; w is the coefficient, and the value of w is less than or equal to 1; d is the damping factor, d = 0.85; A is the first node information of each data node and each output node; B is the second node information of each data node and all output nodes; A / B is the third node information of each output data node; among them, the number of output data nodes is j.

[0169] The second processing module 403 is further used for:

[0170] Determine the output edges between each data node and each output data node; wherein the number of output edges is at least 1;

[0171] According to the blood relationship graph of each data node, determine the operation type and frequency information of each output edge;

[0172] Determine the weight information of each output edge according to the operation type of each output edge;

[0173] Determine output edge information of each output edge according to weight information and frequency information of each output edge;

[0174] Determine first node information of each data node and each output data node according to the plurality of output edge information;

[0175] Determine fourth node information of each data node and each output data node according to the weight information of each output edge;

[0176] According to the plurality of fourth node information, each data node information and the second node information of the output data node are determined.

[0177] The data cost evaluation device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail in this embodiment.

[0178] Figure 5 Hardware diagram for data cost evaluation provided for this application. Figure 5 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.

[0179] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor 501 executes the above method.

[0180] The specific implementation process of the processor 501 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0181] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0182] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0183] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0184] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0185] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0186] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0187] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0188] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, method or unit, which can be electrical, mechanical or other forms.

[0189] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0191] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0192] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0193] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A data cost assessment method, characterized in that: include: Obtain metadata information and lineage relationship diagram for each data node in the data to be evaluated; Determine physical cost information of each data node according to metadata information of each data node; Determine the target score information of each data node according to the physical cost information and the blood relationship diagram of each data node. The target score information of each data node is used to indicate the importance of each data node. Cost assessment information of the data to be assessed is determined according to the multiple target scoring information.

2. The method according to claim 1, characterized in that Determining the physical cost information of each data node according to the metadata information of each data node includes: Determine computing cost information and storage cost information of each data node according to metadata information of each data node; The physical cost information of each data node is determined according to the computing cost information and the storage cost information.

3. The method according to claim 2, characterized in that Determining the computing cost information and storage cost information of each data node according to the metadata information of each data node includes: The metadata information of each data node is input into the calculation cost formula to obtain the calculation cost information of each data node. The expression of the calculation cost formula is: Among them, C1 is the calculation cost information; n cpu is the number of CPU cores occupied by the data node during the computing process; C cpu is the unit price of CPU core; mem The amount of memory occupied by the data node during the calculation process; C mem is the unit price of memory; b read b is the network bandwidth occupied by the read operation of the data node during the calculation process; write C is the network bandwidth occupied by the write operation of the data node during the computing process; b Unit price of network bandwidth; c The time that the data node takes in the computing operation; The metadata information of each data node is input into the storage cost formula to obtain the storage cost information of each data node. The expression of the storage cost formula is: Among them, C2 is the calculation cost information; S mem_s It is the memory value occupied by the storage data node; ts is the duration of storing the data node.

4. The method according to claim 1, characterized in that: Determining target score information of each data node according to the physical cost information and the blood relationship diagram of each data node includes: Step a, setting the physical cost information of each data node as the initial scoring information of each data node; Step b, determining the output data node of each data node according to the blood relationship diagram of each data node, and determining the node information between each data node and each output data node; wherein the number of output data nodes is at least 1; Step c: inputting the node information between each data node and each output data node into the score update formula of the PageRank algorithm to obtain first score information; Step d, calculating the difference information between the initial scoring information and the first scoring information, and determining the first scoring information as new initial scoring information; Step e: if the difference information does not meet the preset threshold, then re-determine the new node information between each data node and each output data node, determine the new node information as the node information, repeat steps c to d until the difference information meets the preset threshold, and determine the obtained first scoring information as the target scoring information.

5. The method according to claim 4, characterized in that The expression of the score update formula is: Among them, PR is the first scoring information of each data node; k is the sum of the physical cost information of all data nodes; w is the coefficient, and the value of w is less than or equal to 1; d is the damping factor, d = 0.85; A is the first node information of each data node and each output node; B is the second node information of each data node and all output nodes; A / B is the third node information of each output data node; among them, the number of output data nodes is j.

6. The method according to claim 5, characterized in that The method further comprises: Determine the output edges between each data node and each output data node; wherein the number of output edges is at least 1; According to the blood relationship graph of each data node, determine the operation type and frequency information of each output edge; Determine the weight information of each output edge according to the operation type of each output edge; Determine output edge information of each output edge according to weight information and frequency information of each output edge; Determine first node information of each data node and each output data node according to the plurality of output edge information; Determine fourth node information of each data node and each output data node according to the weight information of each output edge; According to the plurality of fourth node information, each data node information and the second node information of the output data node are determined.

7. A data cost assessment device, characterized in that: include: The acquisition module is used to obtain the metadata information and blood relationship diagram of each data node in the data to be evaluated; A first processing module, configured to determine physical cost information of each data node according to metadata information of each data node; A second processing module is used to determine target score information of each data node according to the physical cost information and the blood relationship diagram of each data node, and each data node target score information is used to indicate the importance of each data node; The third processing module is used to determine the cost evaluation information of the data to be evaluated according to multiple target scoring information.

8. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.