A cluster data resource cost evaluation method, device and equipment
By dividing the cluster data resource cost into storage and computing components and obtaining corresponding evaluation metrics, the problem of unreasonable use of cluster hardware resources is solved, and more accurate and efficient resource evaluation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2026-04-07
AI Technical Summary
In big data scenarios, existing technologies lack effective methods to assess whether the use of cluster hardware resources is reasonable, resulting in low resource utilization efficiency.
By dividing the cluster data resource cost into data storage resource cost and data computing resource cost, and obtaining corresponding evaluation indicators from storage evaluation dimensions and computing evaluation dimensions, including the proportion of invalid and inefficient storage and the proportion of invalid and inefficient computing, a precise evaluation can be carried out.
It enables a more accurate assessment of cluster data resource costs, reflects resource usage more intuitively, improves assessment efficiency, and reduces labor costs.
Smart Images

Figure CN117407261B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and specifically to a method, apparatus, and equipment for evaluating the cost of cluster data resources. Background Technology
[0002] In big data scenarios, data-related businesses are rapidly developing, and various business scenarios rely on cluster resources, such as cluster hardware resources. Hardware can be used to build services, undertake computing tasks, and meet the needs of various business scenarios.
[0003] Cluster data resource cost is used to characterize the usage of cluster hardware resources. In order to determine whether the use of cluster hardware resources is reasonable, there is an urgent need for a method to evaluate cluster data resource cost. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus and equipment for evaluating the cost of cluster data resources, used to assess whether the use of cluster data resources is reasonable.
[0005] To solve the above problems, the technical solution provided in this application is as follows:
[0006] Firstly, this application provides a method for evaluating the cost of cluster data resources, the method comprising:
[0007] Obtain assessment dimensions for evaluating cluster data resource costs; the cluster data resource costs include data storage resource costs and data computing resource costs, and the assessment dimensions include storage assessment dimensions and computing assessment dimensions.
[0008] Obtain storage evaluation metrics under the storage evaluation dimension and computing evaluation metrics under the computing evaluation dimension; the storage evaluation metrics include the proportion of invalid storage resource cost in the total data storage resource cost and the proportion of inefficient storage resource cost in the total data storage resource cost; the computing evaluation metrics include the proportion of invalid computing resource cost in the total data computing resource cost and the proportion of inefficient computing resource cost in the total data computing resource cost.
[0009] Obtain the index values of the storage evaluation index and the index values of the calculated evaluation index;
[0010] The cost of data storage resources is evaluated based on the value of the storage evaluation index, and the cost of data computing resources is evaluated based on the value of the computing evaluation index.
[0011] Wherein, the invalid data storage resource cost is the data storage resource cost that meets the first storage condition, the inefficient data storage resource cost is the data storage resource cost that meets the second storage condition, the invalid data computing resource cost is the data computing resource cost that meets the first computing condition, and the inefficient data computing resource cost is the data computing resource cost that meets the second computing condition.
[0012] Secondly, this application provides an apparatus for evaluating the cost of cluster data resources, the apparatus comprising:
[0013] The first acquisition unit is used to acquire assessment dimensions for evaluating the cost of cluster data resources; the cost of cluster data resources includes data storage resource costs and data computing resource costs, and the assessment dimensions include storage assessment dimensions and computing assessment dimensions.
[0014] The second acquisition unit is used to acquire storage evaluation indicators under the storage evaluation dimension and computing evaluation indicators under the computing evaluation dimension; the storage evaluation indicators include the proportion of invalid storage resource cost in the total data storage resource cost and the proportion of inefficient storage resource cost in the total data storage resource cost; the computing evaluation indicators include the proportion of invalid computing resource cost in the total data computing resource cost and the proportion of inefficient computing resource cost in the total data computing resource cost.
[0015] The third acquisition unit is used to acquire the index values of the stored evaluation index and the index values of the calculated evaluation index.
[0016] An evaluation unit is used to evaluate the data storage resource cost based on the index value of the storage evaluation index, and to evaluate the data computing resource cost based on the index value of the computing evaluation index.
[0017] Wherein, the invalid data storage resource cost is the data storage resource cost that meets the first storage condition, the inefficient data storage resource cost is the data storage resource cost that meets the second storage condition, the invalid data computing resource cost is the data computing resource cost that meets the first computing condition, and the inefficient data computing resource cost is the data computing resource cost that meets the second computing condition.
[0018] Thirdly, this application provides an electronic device, comprising:
[0019] One or more processors;
[0020] Storage device, on which one or more programs are stored,
[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for evaluating the cost of cluster data resources.
[0022] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for evaluating the cost of cluster data resources.
[0023] Therefore, this application has the following beneficial effects:
[0024] This application provides a method, apparatus, and device for evaluating cluster data resource costs. It divides cluster data resource costs into data storage resource costs and data computing resource costs, and evaluates the usage of data storage resources from a storage evaluation perspective and the usage of data computing resources from a computing evaluation perspective, making the evaluation of data resource costs more reasonable and closer to reality. Specifically, it obtains storage evaluation indicators under the storage evaluation dimension and computing evaluation indicators under the computing evaluation dimension. The storage evaluation indicators include the proportion of invalid storage resource costs in the total invalid storage resource costs and the proportion of inefficient storage resource costs in the total inefficient storage resource costs. Invalid storage resource costs are the data storage resource costs that meet a first storage condition, and inefficient data storage resource costs are the data storage resource costs that meet a second storage condition. The computing evaluation indicators include the proportion of invalid computing resource costs in the total invalid computing resource costs and the proportion of inefficient computing resource costs in the total inefficient computing resource costs. Invalid computing resource costs are the computing resource costs that meet the first computing condition, and inefficient computing resource costs are the computing resource costs that meet the second computing condition.
[0025] Based on this, we obtain and calculate the values of storage evaluation metrics. The usage of data storage resources is assessed based on these storage evaluation metrics (i.e., the ratio of invalid storage to inefficient storage), which is equivalent to assessing the cost of data storage resources. Simultaneously, the usage of data computing resources is assessed based on the values of computational evaluation metrics (i.e., the ratio of invalid computation to inefficient computation), which is equivalent to assessing the cost of data computing resources. In this way, distinguishing between invalid and inefficient storage in data storage resource costs, and between invalid and inefficient computation in data computing resource costs, makes the assessment dimensions of data storage resource costs more complete and comprehensive. This leads to a more accurate assessment of cluster data resource costs and allows users to more intuitively perceive the usage of data resources. Attached Figure Description
[0026] Figure 1A flowchart illustrating a method for evaluating the cost of cluster data resources provided in this application embodiment;
[0027] Figure 2 A schematic diagram illustrating an evaluation dimension of cluster data resource cost provided in an embodiment of this application;
[0028] Figure 3 A schematic diagram illustrating an indicator for evaluating the cost of data storage resources, provided as an embodiment of this application;
[0029] Figure 4 A schematic diagram illustrating an indicator for evaluating the cost of data computing resources, provided as an embodiment of this application;
[0030] Figure 5 A schematic diagram of a cluster data resource cost assessment device provided in an embodiment of this application;
[0031] Figure 6 This is a schematic diagram of the basic structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0032] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0033] To facilitate understanding of this application, a method for evaluating the cost of cluster data resources provided in an embodiment of this application will be described below with reference to the accompanying drawings. For example, this method for evaluating the cost of cluster data resources can be performed by terminal devices and / or servers, etc., and is not limited here.
[0034] See Figure 1 As shown, this figure is a flowchart of a method for evaluating the cost of cluster data resources provided in an embodiment of this application. Figure 1 As shown, the method may include S101-S104:
[0035] S101: Obtain the evaluation dimensions used to assess the cost of cluster data resources; the cost of cluster data resources includes the cost of data storage resources and the cost of data computing resources, and the evaluation dimensions include the storage evaluation dimension and the computing evaluation dimension.
[0036] Cluster hardware resources can include various resources used to build services and undertake computing tasks, such as memory, central processing units (CPUs), disks, and hard drives. Among these, CPUs and other processing hardware devices are used to implement various computing tasks, while memory, disks, and hard drives are used to store data in databases, partitions, and data tables. Therefore, the implementation of computing tasks in a cluster depends on the computation of devices such as CPUs, and the output of these tasks (such as generated data tables) depends on the storage of data on devices such as memory and disks.
[0037] Cluster data resources are quantified values corresponding to cluster hardware resources (referred to as cluster hardware). Cluster data resources include data storage resources and data computing resources. Data storage resources are the data resources used to store data, and can be represented by the data storage space (quantified value corresponding to cluster hardware) of cluster hardware such as memory, disks, and hard drives. For example, if a disk has 80GB of data storage space, then 80GB can be considered as the disk's data storage resource. Data computing resources are the usable data resources used to produce data, and can be represented by the time (quantified value corresponding to cluster hardware) that cluster hardware such as the central processing unit can use to produce data.
[0038] Cluster data resource cost refers to the cluster data resources that have already been consumed or used. This cost includes both data storage resource cost and data computation resource cost. Specifically, data storage resource cost refers to the data storage resources that have already been consumed or used, and data computation resource cost refers to the data computation resources that have already been consumed or used. For example, if the disk's data storage resource is 80GB, and the data table produced by the current computation task is stored on the disk, using 5GB of data storage space, then the data storage resource cost for that data table is 5GB. As another example, if the current computation task produced data over 10 hours, then 10 hours represents the data computation resource cost used by that computation task.
[0039] Based on this, the assessment of cluster data resource costs can be achieved through two dimensions: the assessment of data storage resource costs and the assessment of data computing resource costs. Therefore, the assessment dimensions of cluster data resource costs can be divided into two dimensions: storage assessment and computing assessment. Terminal devices and / or servers will first obtain these two assessment dimensions. The storage assessment dimension is used to assess the cost of data storage resources, i.e., to assess the usage of storage data resources; the computing assessment dimension is used to assess the cost of data computing resources, i.e., to assess the usage of computing data resources. The assessment of cluster data resource costs by terminal devices and / or servers can be understood as assessing whether the use of cluster data resources is reasonable; if the use of cluster data resources is reasonable, then the resource usage requirements are met. Furthermore, in this application embodiment, cluster hardware resources, cluster data resources, and cluster data resource costs can be determined based on the dimension of data resource users. See also... Figure 2 , Figure 2 This diagram illustrates one dimension for evaluating the cost of cluster data resources, as provided in an embodiment of this application. Figure 2As shown, data resource users can be organizations, which include individuals, groups, departments, and business lines. That is, data resource users can be individuals, teams, departments, business lines, etc. Therefore, when the data resource user is a department, the aforementioned cluster hardware resources and cluster data resources represent the total data resources allocated to the department, and the cluster data resource cost is the data resources consumed or used by the entire department after receiving the allocated total data resources. Thus, the cluster data resource cost for each department can be evaluated. The same applies to other data resource users, and will not be elaborated upon here.
[0040] like Figure 2 As shown, data storage resource costs can include data resource costs at the dimensions of partitions, columns, tables, topics, datasets, databases, and data resource groups, and are represented by the used data storage space of partitions, columns, tables, topics, datasets, databases, and data resource groups. Data computing resource costs can include data resource costs at the dimensions of stages, applications, instances, tasks, and queues, and are represented by the time taken to produce data at each stage, application, instance, task, and queue. Details are provided below.
[0041] S102: Obtain storage evaluation metrics under the storage evaluation dimension and computing evaluation metrics under the computing evaluation dimension; storage evaluation metrics include the proportion of invalid storage resource cost in the total data storage resource cost and the proportion of inefficient storage resource cost in the total data storage resource cost; computing evaluation metrics include the proportion of invalid computing resource cost in the total data computing resource cost and the proportion of inefficient computing resource cost in the total data computing resource cost.
[0042] After obtaining the two evaluation dimensions—storage evaluation dimension and computing evaluation dimension—terminal devices and / or servers can further obtain evaluation metrics for directly evaluating the cost of cluster data resources. Specifically, storage evaluation metrics for evaluating data storage resource costs under the storage evaluation dimension are determined, and computing evaluation metrics for evaluating data computing resource costs under the computing evaluation dimension are determined.
[0043] Evaluation metrics (including storage and computing evaluation metrics) are core metrics and are outcome-based. The values of these metrics (including the values corresponding to storage and computing evaluation metrics) directly indicate the evaluation results, providing a clear picture of whether the use of cluster data resources is reasonable or if any problems exist.
[0044] For example, under the two dimensions of storage evaluation and computation evaluation, the cost of cluster data resources is categorized into effective use, inefficient use, and ineffective use based on the actual usage of cluster data resources, making the evaluation process of cluster data resource costs more refined. Specifically, the data storage resource cost under the storage evaluation dimension can be divided into effective data storage resource cost, inefficient data storage resource cost, and ineffective data storage resource cost. Effective data storage resource cost, inefficient data storage resource cost, and ineffective data storage resource cost respectively reflect three levels of data storage resource costs: effective use (which can be called effective storage, indicating normal and reasonable data resource use), inefficient use (which can be called inefficient storage, indicating low utilization of data resources or outputs, which can be improved later), and ineffective use (which can be called ineffective storage, indicating that the output of data resources is completely unutilized, and data resource use can be stopped later). Data stored using effective data storage resource costs can be called effectively stored data, and effectively stored data can be called hot data. Similarly, inefficiently stored data can be called cold data, and ineffectively stored data can be called useless data. Additionally, the data computation resource cost under the computation evaluation dimension can be divided into effective data computation resource cost, inefficient data computation resource cost, and ineffective data computation resource cost. Effective data computing resource cost, inefficient data computing resource cost, and invalid data computing resource cost respectively reflect three levels of data computing resource usage: effective use (also known as effective computing), inefficient use (also known as inefficient computing), and invalid use (also known as invalid computing).
[0045] See Figure 3 , Figure 3 This is a schematic diagram illustrating an indicator for evaluating the cost of data storage resources, provided as an embodiment of this application. Based on the above, combined with... Figure 3 Storage evaluation metrics may include the percentage of invalid storage and the percentage of inefficient storage. The percentage of invalid storage is the proportion of invalid data storage resource costs within the total data storage resource costs, while the percentage of inefficient storage resources is the proportion of inefficient data storage resource costs within the total data storage resource costs. The data storage resource costs can be determined based on the data resource users; if the data resource users are departments, then the data storage resource costs here refer to the department's overall data storage resource costs.
[0046] For example, the cost of invalid data storage resources is the cost of data storage resources that meet the first storage condition, and the cost of inefficient data storage resources is the cost of data storage resources that meet the second storage condition. For instance, the first storage condition is that the frequency of access to data stored using the data storage resources is 0, and the second storage condition is that the frequency of access to data stored using the data storage resources is greater than 0 but less than a target frequency, where the target frequency is specified. It should be understood that the specific content of the first and second storage conditions is not limited here and can be determined according to the actual situation.
[0047] It is understandable that the effective data storage resource cost can be considered as the reasonable data storage resource cost, and the effective data storage resource cost does not need to be considered when evaluating the data storage resource cost.
[0048] See Figure 4 , Figure 4 This is a schematic diagram illustrating an indicator for evaluating the cost of data computing resources, provided as an embodiment of this application. Based on the above, combined with... Figure 4 The evaluation indicators can include the percentage of invalid computation and the percentage of inefficient computation. The percentage of invalid computation is the proportion of invalid data computation resource costs within the total data computation resource costs, and the percentage of inefficient computation is the proportion of inefficient data computation resource costs within the total data computation resource costs. When calculating the percentage of invalid and inefficient storage, the data storage resource costs specifically refer to the overall data storage resource costs of the department.
[0049] For example, the cost of computing resources for invalid data is the cost of computing resources that meets the first computing condition, and the cost of computing resources for inefficient data is the cost of computing resources that meets the second computing condition. For instance, the first computing condition is that the frequency of data generated using computing resources is 0, and the second computing condition is that the efficiency of computing resource utilization is lower than the target efficiency (the target efficiency is not limited here). The efficiency of computing resource utilization can be reflected through various indicators, as detailed below. It should be understood that the specific content of the first and second computing conditions is not limited here and can be determined according to the actual situation.
[0050] It is understandable that when the data resource user is a department, the stored data and the data produced are all data stored and produced within that department. The same applies to other data resource users, and will not be elaborated upon here.
[0051] S103: Obtain the value of the storage evaluation index and calculate the value of the evaluation index.
[0052] After obtaining and calculating the storage evaluation metrics, the terminal devices and / or servers can further obtain the values of the storage evaluation metrics and the calculated evaluation metrics. The value of the storage evaluation metric is the ratio of invalid storage to inefficient storage, and the value of the calculated evaluation metric is the ratio of invalid computing to inefficient computing.
[0053] Therefore, the cost of invalid data storage resources and the cost of data storage resources must be determined before the percentage of invalid storage can be calculated. Specifically, the cost of invalid data storage resources and the cost of data storage resources can be defined as the cost of invalid data storage resources within a department. Other indicators are similar.
[0054] S104: Evaluate the cost of data storage resources based on the value of storage evaluation indicators, and evaluate the cost of data computing resources based on the value of computing evaluation indicators.
[0055] When the ratio of invalid storage to inefficient storage is large, it can be assumed that a significant portion of data storage resources are being used effectively. This means that the proportion of invalid and inefficient storage is high, indicating frequent instances of invalid and inefficient data storage, and that data storage resources are not being used efficiently. Similarly, when the ratio of invalid computation to inefficient computation is large, it can be assumed that a significant portion of data computation resources are being used effectively. This also means that the proportion of invalid and inefficient computation is high, indicating frequent instances of invalid and inefficient data computation, and that data computation resources are not being used efficiently.
[0056] In practical applications, an evaluation score for cluster data resource costs can be calculated. This score reflects the assessment of cluster data resource costs, indicating whether they are being used reasonably. For example, weights can be assigned to the percentage of invalid storage and the percentage of inefficient storage. The weighted sum of these ratios is then used to calculate the evaluation score for data storage resource costs. Similarly, weights can be assigned to the percentage of invalid computing and the percentage of inefficient computing. The weighted sum of these ratios is then used to calculate the evaluation score for data computing resource costs. Finally, the evaluation scores for data storage resource costs and data computing resource costs are summed to obtain the overall evaluation score for cluster data resource costs.
[0057] Based on the relevant content in S101-S104, it is known that terminal devices and / or servers evaluate the usage of cluster data resource costs from both storage and computational evaluation dimensions, making the assessment of data resource costs more reasonable and closer to reality. Distinguishing between invalid and inefficient storage in data storage resource costs, and between invalid and inefficient computation in data computational resource costs, makes the evaluation dimensions of data storage and computational resource costs more complete and comprehensive, resulting in a more accurate assessment of cluster data resource costs and allowing users to more intuitively perceive the usage of data resource costs. Furthermore, the cluster data resource cost evaluation method provided in this application embodiment can be automatically executed by terminal devices and / or servers, eliminating the need for significant manual labor and improving evaluation efficiency.
[0058] The following sections will detail the storage diagnostic metrics under the storage evaluation dimension and the computing diagnostic metrics under the computing evaluation dimension, explaining how to obtain the values of the storage evaluation metrics and the computing evaluation metrics based on these metrics.
[0059] In one possible implementation, the method for evaluating the cost of cluster data resources provided in this application embodiment further includes the following steps:
[0060] Obtain storage diagnostic indicators under the storage evaluation dimension and computing diagnostic indicators under the computing evaluation dimension; storage diagnostic indicators include ineffective storage diagnostic indicators and inefficient storage diagnostic indicators, and computing diagnostic indicators include ineffective computing diagnostic indicators and inefficient computing diagnostic indicators.
[0061] The diagnostic indicators (including stored diagnostic indicators and calculated diagnostic indicators) in this embodiment are used to make detailed judgments on the evaluation indicators (stored evaluation indicators and calculated evaluation indicators) to obtain the indicator values of the evaluation indicators, and are process-related indicators. Accordingly, problem attribution can be achieved based on the diagnostic indicators.
[0062] For example, the first storage condition described above is meeting the invalid storage diagnostic indicators, and the second storage condition is meeting the inefficient storage diagnostic indicators. That is, the data storage resource cost that meets the invalid storage diagnostic indicators is the invalid data storage resource cost, and the data storage resource cost that meets the inefficient storage diagnostic indicators is the invalid data storage resource cost. For example, the first calculation condition described above is meeting the invalid calculation diagnostic indicators, and the second calculation condition is meeting the inefficient calculation diagnostic indicators. That is, the data calculation resource cost that meets the invalid calculation diagnostic indicators is the invalid data calculation resource cost, and the data calculation resource cost that meets the inefficient calculation diagnostic indicators is the inefficient data calculation resource cost.
[0063] The diagnostic indicators for invalid storage include indicators indicating that the utilization rate of stored data is 0, and / or indicators indicating that the stored data does not meet data storage rules. The diagnostic indicators for inefficient storage include indicators indicating that the utilization rate of stored data is lower than a utilization threshold, and / or indicators indicating that the storage type of the data does not meet the target storage type. The diagnostic indicators for invalid computation include indicators indicating that the efficiency of task-generated data is lower than a first preset efficiency, and / or indicators indicating that the utilization rate of task-generated data is 0. The diagnostic indicators for inefficient computation include indicators indicating that the efficiency of task-generated data is lower than a second preset efficiency; the second preset efficiency is better than the first preset efficiency.
[0064] Based on the above, combined with Figure 3 Invalid-level storage diagnostic indicators may include one or more of the following:
[0065] The data table has no lifecycle set; the number of accesses to the data table within the preset time period is 0; the number of accesses to the data in the database within the preset time period is 0; the number of accesses to the data in the directory within the preset time period is 0; and the lifecycle of the data table is greater than the recommended value.
[0066] The Time To Live (TTL) value refers to the allowed storage time of data in the data warehouse, which can also be understood as the allowed storage time of data on the hard drive. In practical applications, data tables are used to store data. A smaller TTL value for a data table leads to more frequent data updates, increasing the burden on the device. A larger TTL value results in longer data storage time, potentially leading to outdated data. Therefore, it is necessary to set the TTL for data tables, and the TTL value must be set appropriately. Recommended TTL values are typically provided, and the TTL setting should follow these recommendations. Based on this, in this embodiment, if a data table does not have a TTL value set or the set TTL value is greater than the recommended value, it indicates that the data table's settings do not comply with the TTL setting regulations, and the data in the data table can be considered invalidly stored data. The data storage resources used by this data table are then considered the data storage space of the data table, and the cost of this used data storage space is considered invalid data storage resource cost. In addition, if the TTL value follows the recommended value setting, it can be considered that the data storage rules are met. Therefore, "the data table has not set a lifespan" and "the lifespan of the data table is greater than the recommended value" can be considered as indicators used to characterize that the stored data does not meet the data storage rules.
[0067] In practical applications, the terminal device and / or server will first obtain the data table and the corresponding lifecycle of the data table. When the obtained lifecycle of the data table is empty, the terminal device and / or server will determine that the data table has not been set with a lifecycle.
[0068] In this embodiment, when stored data remains unused for a preset time period, it is determined to be invalid data, lacking value, and the data storage space used by this data represents the cost of invalid data storage resources. As an optional example, the preset time period can be 30 days, and the preset time periods in different invalid storage diagnostic indicators can be the same or different. The stored data can be categorized as data stored in data tables, data stored in databases, and data stored in directories (e.g., directories in a file storage system like HDFS). The fact that the data has remained unused for the preset time period can be represented by the number of accesses to the data tables being 0, the number of accesses to the database being 0, and the number of accesses to the directory being 0. Therefore, "the number of accesses to the data tables being 0, the number of accesses to the database being 0, and the number of accesses to the directory being 0" can be considered indicators representing a zero utilization rate for the stored data.
[0069] Taking a data table as an example, in practical applications, the terminal device and / or server will first obtain the data table and obtain (can be counted) the number of times the data table is accessed within a preset time period, and compare this number of accesses with 0 to determine whether the number of times the data table is accessed within the preset time period is 0.
[0070] It is understandable that if the stored data meets any of the above invalid storage diagnostic indicators, the data storage resource cost consumed by the data is considered to be invalid data storage resource cost.
[0071] Based on the above, combined with Figure 3 Inefficient storage diagnostic indicators may include one or more of the following:
[0072] The number of small files exceeds the target number, the access frequency of partition data is less than the target frequency, and the stored data is not compressed.
[0073] Specifically, when the access frequency of partition data is non-zero and less than the target frequency, the partition data is considered to be used, but at a low frequency. Thus, the partition data is determined to be inefficiently stored, and the data storage space used by this partition data represents the inefficient data storage resource cost. It can be understood that "partition data access frequency less than the target frequency" can be considered an indicator used to characterize that the utilization rate of stored data is lower than a utilization rate threshold. Access frequency is used to characterize data utilization, and the target frequency is used to characterize the utilization rate threshold.
[0074] In practical applications, terminal devices and / or servers first determine the partition, then obtain the access frequency of the data in that partition, and compare the access frequency of the partition data with the target frequency to determine whether the access frequency of the partition data is less than the target frequency.
[0075] Additionally, small files refer to files with relatively small file sizes, such as 1KB. It's understandable that when the number of small files is excessive (e.g., the number exceeds the target number, regardless of the target number), the data storage is considered scattered, and this storage method is deemed inefficient. The data stored in small files is considered inefficiently stored data. Therefore, the data storage space used by small files represents the inefficient data storage resource cost. Furthermore, when some stored data should be compressed to save storage space but is not, the stored data is considered to occupy excessive data space. Thus, the data storage space used by uncompressed stored data represents the inefficient data storage resource cost. It's understandable that "the number of small files exceeds the target number, and the stored data is uncompressed" can be considered an indicator that the data storage type does not meet the target storage type. A storage type where the number of small files is less than or equal to the target number and the stored data is compressed can be considered the target storage type; this is merely an example and not a limitation.
[0076] In practical applications, the terminal device and / or server first obtains the number of small files, then compares this number with the target number to determine if the number of small files exceeds the target number. Additionally, the terminal device and / or server directly obtains the stored data for which compression is to be determined, and then checks whether the stored data is compressed to determine if it has been compressed.
[0077] It is understandable that if the stored data meets any of the above inefficient storage diagnostic indicators, the data storage resource cost consumed by the data is considered to be an inefficient data storage resource cost.
[0078] It is also understandable that the above-mentioned diagnostic indicators for ineffective and low-efficiency storage list some indicators related to data tables, databases, partitions, etc. Figure 2 As shown, under the storage evaluation dimension, there are also some ineffective and inefficient storage diagnostic indicators related to columns, topics, datasets, and data resource groups, which can be set according to the actual situation.
[0079] Based on the above, combined with Figure 4 The diagnostic indicators for invalid level calculations may include one or more of the following:
[0080] The number of accesses to the task output within the preset time period is 0; the task ends normally within the preset time period but the output is empty; the number of accesses to the dashboard and / or interface within the preset time period is 0; the task fails within the preset time period; and the task queue has zero utilization rate on the day.
[0081] The preset time periods for each invalid level's diagnostic indicators can be the same or different, and can be determined according to the actual situation.
[0082] The example of "the number of accesses generated by the task within the preset time period is 0" is as follows: Figure 4 The phrase "no access to output in the past 30 days" means that the preset time period for this metric is 30 days. Task output can be understood as output data, which can be represented as a data table containing the output data. This is just an example and the data storage method is not limited. It can be seen that when the access count of the data generated by the task in the past 30 days is 0, it means that the task output is invalid and has no value. Therefore, the task output (i.e., the data generated by the task) can be considered invalid computational data, and the data computation resources used for the task output are invalid data computational resource costs.
[0083] In practical applications, the terminal device and / or server will first obtain the task output (such as the output data), then obtain the number of times the task output is accessed within a preset time period, and compare the number of accesses with 0 to determine whether the number of accesses of the task output within the preset time period is 0.
[0084] An example of "the task completing normally within a preset time period but producing no output" can be found here. Figure 4 The phrase "output empty for 3 consecutive days" means that the preset time period for this indicator is 3 days. It should be understood that if a task completes normally but outputs nothing for 3 consecutive days, then the execution of the task for those 3 consecutive days is invalid. In this case, the data computing resources used for task execution are considered invalid data computing resource costs. For example, the task completes normally, producing data for August 20th, but the partition name for storing the output data is incorrectly written as August 8th, resulting in the partition storing August 20th data not containing the data produced by this task; the storage is empty.
[0085] In practical applications, the terminal device and / or server first obtain the status of the task within a preset time period. This status includes whether the task has ended normally and whether the task output is empty. In this way, it can be determined whether the scenario of "the task ending normally within the preset time period but the output being empty" exists.
[0086] "The number of accesses to the dashboard and / or interface within the preset time period is 0" can be exemplified as follows: Figure 4 The phrase "application access zero for 30 consecutive days" means that the preset time period for this metric is 30 days. "Application" can refer to a visual business dashboard within an application and / or an interface within an application. The dashboard displays metric values, and the interface accesses data; this is not limited to a single application. It should be understood that if the dashboard and / or interface have no access for 30 consecutive days, it indicates a potential problem with the data across the entire data chain. The data displayed on the dashboard and the data accessible through the interface are considered invalid calculations, and the computational resources used to obtain this data can be considered invalid data computational resource costs.
[0087] In practical applications, the terminal device and / or server will first obtain the number of accesses to the dashboard and / or interface within a preset time period, and then compare the number of accesses with 0 to determine whether the number of accesses to the dashboard and / or interface within the preset time period is 0.
[0088] Example of "task failure within a preset time period" is: Figure 4 The phrase "three consecutive days of task failure" refers to a preset time period of three days for this metric. It should be understood that if a task fails for three consecutive days, then the execution of the task for those three days is invalid. For example, a misspelled table name could cause the task to fail. In this case, the data computation resources used for task execution are considered invalid data computation resource costs.
[0089] In practical applications, the terminal device and / or server will first obtain the running results of the task within a preset time period, and determine whether the task within the preset time period has failed based on the running results.
[0090] like Figure 2As shown, the task queue is one dimension of computational evaluation. The data computation resources configured for the task queue can be exemplified by a 1000-core CPU and 10TB of memory. "Zero utilization of the task queue on a given day" indicates that the task queue is being used ineffectively, and the corresponding data computation resources represent the cost of ineffective data computation. In practical applications, terminal devices and / or servers first obtain the daily task queue utilization rate and compare it to 0 to determine whether the daily task queue utilization rate is truly zero.
[0091] It is understandable that zero accesses to task outputs within a preset time period, and zero accesses to the dashboard and / or interface within a preset time period, can all be indicators used to represent a zero utilization rate of task output data. Task failures within a preset time period, zero utilization of the task queue on that day, and tasks that complete normally within a preset time period but produce nothing can all be indicators used to represent task output data efficiency being lower than a first preset efficiency, where the first preset efficiency is not limited.
[0092] Based on the above, combined with Figure 4 Inefficient diagnostic indicators may include one or more of the following:
[0093] Data resource utilization rate is lower than the target data resource utilization rate, data skew occurs in tasks, tasks are duplicated, task execution time is longer than the target time, the proportion of task queue blocking time exceeds the target proportion, and the proportion of task queue over-issuance time exceeds the target proportion.
[0094] The "data resource utilization rate" is the ratio of the difference between the requested data resource amount and the used data resource amount to the requested data resource amount. Both the requested and used data resource amounts can refer to the data resource amounts requested and used by departments that are data resource users. Here, data resources refer to data computing resources, including memory data resources and CPU data resources. Therefore, the requested data resource amount includes the requested amount of memory data resources and CPU data resources, while the used data resource amount includes the used amount of memory data resources and CPU data resources. For example, the requested amount of memory data resources is 10GB, and the requested amount of CPU data resources is 8 CPUs. The used amount of memory data resources is 5GB, and the used amount of CPU data resources is 4 CPUs. Low data resource utilization rate includes low memory data resource utilization rate and low CPU data resource utilization rate. It should be understood that when the data resource utilization rate is lower than the target data resource utilization rate, it indicates that data resources are being over-requested. In this case, the requested data computing resources can be considered inefficient data computing resource costs. The specific value of the target data resource utilization rate is not limited and can be determined based on the actual scenario.
[0095] In practical applications, terminal devices and / or servers will first obtain the data resource utilization rate in the manner described above, and then compare the data resource utilization rate with the target data resource utilization rate to determine whether the data resource utilization rate is lower than the target data resource utilization rate.
[0096] "Data skew" indicates that the time taken for data output from subtasks of a task is uneven, with some subtasks consuming data resources for extended periods. For example, if a task has 100 subtasks, and the task is considered complete once all 100 subtasks have finished running, but 99 subtasks complete within one minute, while the remaining subtask takes two hours to complete, then data skew has occurred. Data skew leads to low task efficiency. In this case, it's determined that the entire task is using inefficient data computing resources.
[0097] In practical applications, the terminal device and / or server will first obtain the time taken for the data produced by each subtask of the task, and analyze the time taken for the data produced by each subtask to determine whether the time taken for the data produced by any subtask exceeds the target time (e.g., 1 hour), so as to determine whether the task has data skew.
[0098] "Task duplication" indicates that the constructed task contains duplicate parts, resulting in repeated processing and wasted data computing and data storage resources. For example, the first task is to retrieve three fields (fields a, b, and c) from table A and place them into table B. The second task is to retrieve four fields (fields a, b, c, and d) from table A and place them into table C. It's clear that there's no need to construct the second task; by using an instruction to also place field d into table B, fields a, b, c, and d can all be retrieved from table B. In this example, the first and second tasks are considered duplicate tasks, leading to additional waste of data computing and data storage resources. Therefore, when tasks are duplicated, the data computing resources used by the multiple duplicate tasks are all considered inefficient data computing resource costs.
[0099] In practical applications, terminal devices and / or servers first determine the task content of multiple tasks, and then analyze the task content of multiple tasks to determine whether the task content of multiple tasks can be merged in order to determine whether there is any task duplication.
[0100] "Task execution time exceeds target time" indicates that the task runs for a long time, consuming significant data computing resources. In this case, the data computing resources used by the task are defined as inefficient data computing resource costs. The specific value of the target time is not limited and can be determined based on the actual scenario. For example, ... Figure 4 As shown, the target duration is 10 hours.
[0101] In practical applications, the terminal device and / or server will first obtain the task execution time, and then compare the task execution time with the target time to determine whether the task execution time is greater than the target time.
[0102] The "Task Queue Blocking Time Ratio" is the ratio of the total suspended time of instances in the queue to the total runtime of instances in the queue. It should be understood that when the Task Queue Blocking Time Ratio exceeds the target ratio, it indicates severe queue blocking, and there may be problems with task execution. In this case, the data computing resources corresponding to the task queue are considered inefficient data computing resource costs. The specific value of the target ratio is not limited and can be determined based on the actual scenario. For example, ... Figure 4 As shown, the target ratio is 30%.
[0103] In practical applications, the terminal device and / or server first obtains the total suspended duration and total runtime of instances in the queue, and calculates the task queue blocking time ratio. Then, the task queue blocking time ratio is compared with a target to determine whether the task queue blocking time ratio exceeds the target ratio.
[0104] The "Task Queue Over-issuance Duration Ratio" is the ratio of the number of times the total data resources requested in the task queue exceed the minimum guaranteed data resources to the total number of times data is requested during the tracking period. It should be understood that some task queues may use data resources exceeding the configured data resource limits. For example, a task queue may be configured with 1000 CPU cores and 10TB of memory, allowing it to use data resources exceeding the configured limits. The timeout period for exceeding these limits is the over-issuance duration. When the task queue over-issuance duration ratio is greater than the target ratio, it indicates that the over-issuance duration is too long, and the current data computing resource requests for the task queue are unreasonable. The data computing resources for the task queue can be increased, or the tasks in the queue can be adjusted. In this case, the data computing resources corresponding to the task queue are considered inefficient data computing resource costs. The specific value of the target ratio is not limited and can be determined based on the actual scenario. For example, ... Figure 4 As shown, the target ratio is 30%.
[0105] In practical applications, the terminal device and / or server first obtains the number of times the total data resource requests in the task queue exceed the minimum guaranteed data resource, as well as the total number of data points, and calculates the over-issuance time ratio of the task queue. Then, the over-issuance time ratio of the task queue is compared with the target ratio to determine whether the over-issuance time ratio of the task queue exceeds the target ratio.
[0106] It is understandable that "data resource utilization rate lower than the target data resource utilization rate, data skew in tasks, task duplication, task execution time longer than the target time, task queue blocking time exceeding the target proportion, and task queue over-issuance time exceeding the target proportion" can all be considered indicators used to characterize the efficiency of task output data being lower than the second preset efficiency. Task output data efficiency can be used to characterize the efficiency of data computing resource utilization.
[0107] It is also understandable that the above-mentioned diagnostic indicators for ineffective and low-efficiency computing list some indicators related to tasks, queues, etc. Figure 2 As shown, under the computational evaluation dimension, there can also be some ineffective and inefficient computational diagnostic indicators related to stages, applications, and instances, which can be set according to the actual situation.
[0108] Based on the above, this application provides a specific implementation method for obtaining the index value of the stored evaluation index and calculating the index value of the evaluation index in S103, including:
[0109] A1: Based on the storage diagnostic indicators and their corresponding values, determine the values of the storage evaluation indicators.
[0110] In other words, it determines whether the stored data meets the storage diagnostic indicators and obtains the corresponding indicator values. Since the indicator values can characterize the types of stored data, namely invalid data (data stored ineffectively) and inefficient data (data stored inefficiently), the resource costs for storing invalid data and inefficient data can be determined based on the corresponding indicator values. Furthermore, the percentage of invalid storage and the percentage of inefficient storage can be calculated.
[0111] A2: Based on the calculated diagnostic indicators and their corresponding values, determine the values of the calculated evaluation indicators.
[0112] In other words, it determines whether tasks / queues, etc., meet the computational diagnostic indicators and obtains the corresponding indicator values. Since the indicator values can characterize the type of data computing resource usage, namely invalid computing and inefficient computing, the cost of invalid data computing resources and the cost of inefficient data computing resources can be determined based on the corresponding indicator values. Furthermore, the proportion of invalid computing and the proportion of inefficient computing can be calculated.
[0113] Specifically, in one possible implementation, this application provides a specific implementation method for determining the index values of storage evaluation indicators based on storage diagnostic indicators and their corresponding index values in A1, including:
[0114] A11: Based on the invalid-level storage diagnostic indicators and their corresponding values, determine the amount of invalid data stored in the cluster data resources, and based on the inefficient-level storage diagnostic indicators, determine the amount of inefficient data stored in the cluster data resources; the amount of invalid data stored is used to represent the cost of invalid data storage resources, and the amount of inefficient data stored is used to represent the cost of inefficient data storage resources.
[0115] Based on the above, and using the listed invalid storage diagnostic indicators, determine whether the stored data meets these indicators and obtain the corresponding indicator values (indicator values can be yes or no). If any one of the invalid storage diagnostic indicators is met, the data is determined to be invalid storage, and the stored data is invalid data. After all invalid storage diagnostic indicators have been determined, calculate the data storage volume of the invalid data, i.e., the invalid data storage resource cost. Similarly, calculate the data storage volume of inefficient data, i.e., the inefficient data storage resource cost.
[0116] A12: Determine the total storage usage of data in the cluster data resources; the total storage usage is used to represent the cost of data storage resources.
[0117] A13: The ratio of invalid data storage to total data storage is determined as the ratio of invalid storage percentage, and the ratio of inefficient data storage to total data storage is determined as the ratio of inefficient storage percentage.
[0118] It should be understood that cluster data resources can refer to the total data resources applied for by departments whose data resource users are departments. Total data storage usage can refer to the data storage resources used by the department as a whole, representing the cost of data storage resources. The stored data mentioned in A11 can refer to various types of data stored within the department.
[0119] For example, a department uses data tables to store data. The department applied for 100 data tables (i.e., cluster data resources), each with a data storage capacity of 1GB. The department used 90 of these data tables (i.e., the data storage space corresponding to these 90 data tables represents the total used storage space, i.e., the data storage resource cost). After assessment, 10 data tables meet the listed invalid storage diagnostic indicators; therefore, the data stored in these 10 data tables is invalid, and the data storage capacity of invalid data is 10GB. Additionally, after assessment, 30 data tables meet the listed inefficient storage diagnostic indicators; therefore, the data stored in these 30 data tables is inefficient, and the data storage capacity of inefficient data is 30GB. The department's total used storage capacity is 90GB. Therefore, 10GB / 90GB represents the percentage of invalid storage, and 30GB / 90GB represents the percentage of inefficient storage.
[0120] It is understandable that the percentage of invalid storage indicates the waste of data storage resources, while the percentage of inefficient storage indicates the inefficient use of data storage resources. Thus, we can understand the cost of data storage resources. The less data storage resources are wasted and the less data storage resources are used inefficiently, the lower the cost of data storage resources.
[0121] Furthermore, storage diagnostic metrics can also be used to attribute the causes of data storage resource cost issues. As can be seen in the above judgment process, knowing the specific storage diagnostic metrics triggered allows for the identification of detailed locations where data storage resources are not being used properly, thus achieving problem attribution. For example, if table A triggers the invalid-level storage diagnostic metric "Table does not have a lifecycle set," with a value of "Yes," then it can be determined that the data stored in table A is invalid, table A is the storage location of invalid data, and "Table does not have a lifecycle set" is the cause of the invalid data storage resource cost issue. In this way, after problem attribution, it becomes easier for users to subsequently manage data storage resources.
[0122] In one possible implementation, this application provides a specific implementation method for determining the value of the calculated evaluation index based on the calculated diagnostic index and its corresponding value in A2, including:
[0123] A21: Based on the invalid computing diagnostic indicators and their corresponding values, determine the invalid computing time in the cluster data resources, and based on the inefficient computing diagnostic indicators, determine the inefficient computing time in the cluster data resources; invalid computing time is used to represent the cost of computing resources for invalid data, and inefficient computing time is used to represent the cost of computing resources for inefficient data; computing time is obtained from the computing time used by the central processing unit and the computing time used by memory.
[0124] Based on the above, and using the listed diagnostic indicators for ineffective computing, we determine whether tasks, queues, etc., meet these indicators and obtain the corresponding indicator values (the values can be either yes or no). If any one of the diagnostic indicators for ineffective computing is met, the data computing resources used are determined to be ineffective. After all diagnostic indicators for ineffective computing have been determined, the cost of ineffective data computing resources is calculated. Similarly, the cost of inefficient data computing resources is calculated.
[0125] For example, the cost of data computing resources can be represented in computation time. Computation time can be understood as a unit for evaluating the amount of data computing resources used by a task / queue. Computation time is obtained by calculating the computation time used by the CPU and the computation time used by memory. Specifically, the computation time is the greater of the CPU computation time and the memory computation time. One hour of use of one CPU core is counted as one computation time, and one hour of use of 4GB of memory is counted as one computation time. For example, if a task uses one CPU core and 8GB of memory, and runs for one hour, then the CPU computation time is 1 computation time, and the memory computation time is 2 computation times. Therefore, the corresponding computation time for this task is the greater of 1 and 2 computation times, i.e., 2 computation times.
[0126] Thus, when the data computing resources are invalid, the data computing resource cost corresponds to invalid computing; when the data computing resources are inefficient, the data computing resource cost corresponds to inefficient computing.
[0127] A22: Determine the total data usage computation time in the cluster data resources; the total data usage computation time is used to represent the data computing resource cost.
[0128] A23: The ratio of invalid calculation time to total data usage calculation time is determined as the ratio of invalid calculation percentage, and the ratio of inefficient calculation time to total data usage calculation time is determined as the ratio of inefficient calculation percentage.
[0129] It should be understood that cluster data resources can refer to the total data resources applied for by departments whose data resource users are departments. Total data usage calculations can refer to the data computing resources used by the entire department, representing the cost of data computing resources. The tasks and queues mentioned in A21 can refer to the tasks and queues constructed within the department.
[0130] For example, if a department has only 100 tasks designed to generate data, then it's necessary to determine whether each task triggers the listed ineffective and inefficient computation diagnostic indicators. Tasks triggering ineffective computation diagnostic indicators are counted, and the cumulative computation time used by multiple tasks triggering these indicators is defined as ineffective computation time. Similarly, tasks triggering inefficient computation diagnostic indicators are counted, and the cumulative computation time used by multiple tasks triggering these indicators is defined as inefficient computation time. The cumulative computation time used by all 100 tasks is then determined as the total computation time used for data processing. From this, the ratio of ineffective computation to inefficient computation can be calculated.
[0131] The percentage of invalid computations indicates instances where computation was used ineffectively, while the percentage of inefficient computations indicates instances where computation was used inefficiently. Therefore, we can understand the cost of data computing resources. The less data computing resources are used ineffectively or inefficiently, the more rationally the data computing resources are used.
[0132] Similarly, diagnostic metrics can also be used to attribute the causes of data computing resource cost problems, which will not be elaborated here. In this way, attributing the causes of data computing resource cost problems facilitates users' subsequent governance of data computing resources.
[0133] Understandably, if the data resource users are various departments, then not only can we know the total data storage usage and total data computation time of each department, but we can also identify the ineffective, inefficiently used data storage resources, ineffective data computation resources, and inefficiently used data computation resources within each department. Furthermore, after attributing the problems, we can also manage the data resources to save data storage and data computation resources and reduce ineffective and inefficient use of data resources.
[0134] Based on the above, the cluster data resource cost assessment method provided in this application is a measurable and interpretable cluster data resource assessment system. This method can clearly identify effectively used cluster data resources as well as ineffective and / or inefficiently used cluster data resources. Furthermore, it provides a process and outcome indicators for the assessment, accurately measuring whether data storage resource costs and data computing resource costs are being used reasonably.
[0135] To enable a more granular assessment of data storage and computing resource costs, embodiments of this application provide observation metrics, also known as benefit metrics, including indicators that assist in judgment or decision-making, and indicators that contribute to governance benefits. These observation metrics can be visualized on a dashboard.
[0136] Based on this, in one possible implementation, the cluster data resource cost evaluation method provided in this application embodiment further includes the following steps:
[0137] Obtain storage observation metrics under the storage evaluation dimension and computing observation metrics under the computing evaluation dimension; storage observation metrics include auxiliary judgment metrics for data storage resource costs and / or data storage resource governance metrics; computing observation metrics include auxiliary judgment metrics for data computing resource costs and / or data computing resource governance metrics;
[0138] By combining storage and computation metrics, we can assess the cost of data storage resources and the cost of data computation resources.
[0139] As can be seen, the observation metrics include storage observation metrics under the storage evaluation dimension and computing observation metrics under the computing evaluation dimension. The scope of the observation metrics is relatively broad, and they are used in combination to assist in evaluating the data storage resource cost and in combination to assist in evaluating the data computing resource cost. The evaluation results depend on the value of the observation metrics.
[0140] Storage metrics include auxiliary indicators for assessing data storage resource costs and / or indicators for data storage resource governance. Based on this, as an optional example, combined with... Figure 3 Storage observation metrics include one or more of the following:
[0141] The data storage volume increase within adjacent time periods, the month-on-month increase in data storage volume, the ranking results of the data storage volume increase within a preset time period, the data storage volume of invalid and / or inefficient data that has been addressed within the preset time period, the monetary value corresponding to the data storage volume of invalid and / or inefficient data that has been addressed, the total data storage volume of invalid data within the preset time period, and the total data storage volume of inefficient data within the preset time period.
[0142] The "increase in data storage volume within adjacent time periods" can be exemplified as follows: Figure 3 The "Yesterday's New Storage" indicator refers to the adjacent time periods of the previous day and yesterday. "Yesterday's New Storage" represents the cost of data storage resources added yesterday, specifically the total data storage volume of yesterday minus the total data storage volume of the previous day. When the data resource user is a department, "Yesterday's New Storage" specifically represents the department's cost of data storage resources added yesterday. By analyzing changes in the "Yesterday's New Storage" indicator, we can determine whether the daily increase in data storage volume is normal, thus helping to assess whether data storage resource costs are being used rationally.
[0143] "Data Storage Value-Added Month-on-Month" is the ratio of the difference between yesterday's new storage and the previous day's new storage, to the previous day's new storage. "Data Storage Value-Added Month-on-Month" can be used to represent a horizontal comparison of daily data storage value-added, helping to determine whether data storage resources are being used reasonably.
[0144] The sorting results of "data storage volume increment within a preset time period" can be exemplified as follows: Figure 3 The daily / weekly growth rate TOP10 in the data means that the preset time period for this indicator can be daily or weekly, and the ranking result can be the TOP10. It can be understood that if the data resource users are departments, then the ranking result can specifically be the data storage value-added of the top 10 departments.
[0145] "The amount of invalid and / or inefficient data stored within a preset time period that has been managed" is a data storage resource management indicator, which can be exemplified as follows: Figure 3The "yesterday's managed storage" in this indicator refers to the preset time period of yesterday. For example, if a data table meets the invalid storage diagnostic criteria, it is considered invalid and can be deleted. In this case, the invalid data in the table can be considered managed, and the data storage size of the table is the managed invalid data storage size. Another example: there are 100 small files, each with a data storage size of 1KB, for a total data storage size of 100KB. Since 100 files is a large number, the storage of these 100 small files is considered inefficient, and the corresponding inefficient data storage size is 100KB. Therefore, these 100 small files can be merged into one file to manage the inefficient data. After management, the data storage size of the single file is 100KB, which can be considered effective storage. Therefore, the managed inefficient data storage size can be considered 100KB.
[0146] "The monetary value corresponding to the amount of data stored that has been remediated (invalid and / or inefficient data)" is a data storage resource governance indicator, exemplified by: Figure 3 The "yesterday's governance revenue" corresponds to the "data storage volume of invalid and / or inefficient data that has been governed within the preset time period." It can be considered that the data storage volume corresponds to a certain monetary amount; therefore, after determining the amount of data storage that has been governed, the corresponding monetary amount can be calculated. For example, 1GB of data storage corresponds to a unit price of 100 yuan. The monetary amount corresponding to the governed data storage volume can be considered as the revenue generated from the governance of data storage resources for invalid and / or inefficient data.
[0147] "The total amount of invalid data stored within the preset time period" Figure 3 The "invalid storage amount" shown refers to the total amount of storage that is not being used effectively. The "total amount of inefficient data stored within the preset time period" is... Figure 3 The "Inefficient Storage" shown refers to the total amount of storage that is being used inefficiently. The preset time period can be yesterday; no specific time limit is specified here. Additionally, storage monitoring metrics may include... Figure 3 The dashboard displays the amounts for invalid and inefficient storage. Invalid storage amount is the monetary value corresponding to the amount of invalid storage, and inefficient storage amount is the monetary value corresponding to the amount of inefficient storage. These metrics are displayed in the dashboard, allowing users to understand the relevant storage situation in detail.
[0148] The computational observation indicators include auxiliary judgment indicators of data computing resource costs and / or data computing resource governance indicators. Based on this, as an optional example, combined with... Figure 4 The observation indicators include one or more of the following:
[0149] The ranking results include: total number of tasks, increase in task volume within adjacent time periods, task execution duration, computation time governance results within a preset time period, unresolved invalid computation time within a preset time period, unresolved inefficient computation time within a preset time period, the corresponding monetary value of unresolved invalid computation time within a preset time period, and the corresponding monetary value of unresolved inefficient computation time within a preset time period. The computation time governance results include resolved invalid and / or inefficient computation time, and the corresponding monetary values.
[0150] "Total number of tasks" refers to the total number of tasks from the perspective of data resource users, serving as an auxiliary indicator for judging the cost of data computing resources. For example, when the data resource user is a department, "total number of tasks" specifically refers to the total number of tasks constructed by the department.
[0151] An example of "increased task volume within adjacent time periods" is... Figure 4 The "Number of New Tasks Added Yesterday" is an auxiliary indicator for judging the cost of data computing resources. The adjacent time periods in this indicator are the day before yesterday and yesterday. "Number of New Tasks Added Yesterday" represents the total number of tasks under the data resource user dimension. For example, when the data resource user is a department, "Number of New Tasks Added Yesterday" specifically refers to the number of tasks added by the department yesterday.
[0152] An example of the sorting results for "task execution time" is shown below. Figure 4 The "Execution Time TOP100" list displays the execution times of the top 100 tasks. The execution time is the difference between the task's start and end times. It's understood that longer execution times likely indicate potential issues with data computing resource costs. Displaying the execution times of the TOP100 tasks provides guidance on addressing data computing resource costs.
[0153] The data computation resource governance indicator is the data computation resource cost calculated during the computation time. "Governed invalid computation time and / or inefficient computation time" can be exemplified as "Yesterday's governed data computation resource volume," representing the total data computation resource cost governed yesterday. "The corresponding monetary value for governed invalid computation time and / or inefficient computation time" can be exemplified as... Figure 4 The “Yesterday’s governance revenue” refers to the total revenue corresponding to the data calculation resource cost of yesterday’s governance.
[0154] "When invalid calculations are to be addressed" can be exemplified as follows: Figure 4 The "invalid computational load" refers to the total amount of computational resources currently remaining for unresolved invalid data. The "amount corresponding to invalid computations pending remediation" can be exemplified as follows: Figure 4 The "invalid calculation amount" refers to the monetary value converted from the total amount of remaining unmanaged invalid data calculation resources. An example of "inefficient calculations awaiting remediation" is... Figure 4 The term "inefficient computational load" refers to the total amount of computational resources currently available for the remaining unmanaged inefficient data. The monetary value corresponding to the inefficient computation to be addressed can be exemplified as follows: Figure 4 The "Inefficient Computation Amount" represents the monetary value converted from the total amount of unmanaged inefficient data computing resources. The "Amount Values Corresponding to Invalid Computation to be Managed, Inefficient Computation to be Managed, Invalid Computation to be Managed, and Inefficient Computation to be Managed" displays the current manageable situation and are data computing resource governance indicators related to the cost of data computing resources.
[0155] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0156] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
[0157] Based on the method for evaluating cluster data resource costs provided in the above embodiments, this application also provides a device for evaluating cluster data resource costs. The device will be described below with reference to the accompanying drawings. Since the principle by which the device solves the problem in this disclosure is similar to the method for evaluating cluster data resource costs described in this application, the implementation of the device can refer to the implementation of the method, and repeated details will not be elaborated further.
[0158] See Figure 5 As shown in the figure, this is a schematic diagram of the structure of a cluster data resource cost evaluation device provided in an embodiment of this application. Figure 5 As shown, the cluster data resource cost assessment device 500 includes:
[0159] The first acquisition unit 501 is used to acquire the assessment dimensions for evaluating the cluster data resource cost; the cluster data resource cost includes data storage resource cost and data computing resource cost, and the assessment dimensions include storage assessment dimension and computing assessment dimension.
[0160] The second acquisition unit 502 is used to acquire storage evaluation indicators under the storage evaluation dimension and computing evaluation indicators under the computing evaluation dimension; the storage evaluation indicators include the proportion of invalid storage resource cost in the total data storage resource cost and the proportion of inefficient storage resource cost in the total data storage resource cost; the computing evaluation indicators include the proportion of invalid computing resource cost in the total data computing resource cost and the proportion of inefficient computing resource cost in the total data computing resource cost.
[0161] The third acquisition unit 503 is used to acquire the index value of the stored evaluation index and the index value of the calculated evaluation index.
[0162] The first evaluation unit 504 is used to evaluate the data storage resource cost based on the index value of the storage evaluation index, and to evaluate the data computing resource cost based on the index value of the computing evaluation index.
[0163] Wherein, the invalid data storage resource cost is the data storage resource cost that meets the first storage condition, the inefficient data storage resource cost is the data storage resource cost that meets the second storage condition, the invalid data computing resource cost is the data computing resource cost that meets the first computing condition, and the inefficient data computing resource cost is the data computing resource cost that meets the second computing condition.
[0164] In one possible implementation, the device further includes:
[0165] The fourth acquisition unit is used to acquire storage diagnostic indicators under the storage evaluation dimension and computing diagnostic indicators under the computing evaluation dimension; the storage diagnostic indicators include invalid level storage diagnostic indicators and inefficient level storage diagnostic indicators, and the computing diagnostic indicators include invalid level computing diagnostic indicators and inefficient level computing diagnostic indicators.
[0166] The first storage condition is to meet the invalid level storage diagnostic index, the second storage condition is to meet the inefficient level storage diagnostic index, the first calculation condition is to meet the invalid level calculation diagnostic index, and the second calculation condition is to meet the inefficient level calculation diagnostic index.
[0167] The invalid-level storage diagnostic indicators include indicators that characterize the utilization rate of stored data as 0, and / or indicators that characterize the stored data as not meeting data storage rules; the inefficient-level storage diagnostic indicators include indicators that characterize the utilization rate of stored data as lower than a utilization rate threshold, and / or indicators that characterize the storage type of data as not meeting the target storage type; the invalid-level computation diagnostic indicators include indicators that characterize the efficiency of task-generated data as lower than a first preset efficiency, and / or indicators that characterize the utilization rate of task-generated data as 0; the inefficient-level computation diagnostic indicators include indicators that characterize the efficiency of task-generated data as lower than a second preset efficiency; the second preset efficiency is better than the first preset efficiency.
[0168] The third acquisition unit 503 includes:
[0169] The first determining subunit is used to determine the index value of the storage evaluation index based on the storage diagnostic index and the corresponding index value.
[0170] The second determining subunit is used to determine the index value of the calculated evaluation index based on the calculated diagnostic index and the corresponding index value.
[0171] In one possible implementation, the invalid-level storage diagnostic indicators include one or more of the following:
[0172] The data table has no lifecycle set; the number of accesses to the data table within the preset time period is 0; the number of accesses to the data in the database within the preset time period is 0; the number of accesses to the data in the directory within the preset time period is 0; and the lifecycle of the data table is greater than the recommended value.
[0173] The inefficient storage diagnostic indicators include one or more of the following:
[0174] The number of small files exceeds the target number, the access frequency of partition data is less than the target frequency, and the stored data is not compressed;
[0175] The invalidity level calculation diagnostic indicators include one or more of the following:
[0176] The number of accesses to the task output within the preset time period is 0; the task ends normally within the preset time period but the output is empty; the number of accesses to the dashboard and / or interface within the preset time period is 0; the task fails within the preset time period; and the task queue has zero utilization rate on the day.
[0177] The inefficient computational diagnostic indicators include one or more of the following:
[0178] Data resource utilization rate is lower than the target data resource utilization rate, data skew occurs in tasks, tasks are duplicated, task execution time is longer than the target time, the proportion of task queue blocking time exceeds the target proportion, and the proportion of task queue over-issuance time exceeds the target proportion.
[0179] In one possible implementation, the first determining subunit includes:
[0180] The third determining subunit is used to determine the data storage volume of invalid data in the cluster data resources based on the invalid-level storage diagnostic indicators and the corresponding indicator values, and to determine the data storage volume of inefficient data in the cluster data resources based on the inefficient-level storage diagnostic indicators; the data storage volume of invalid data is used to represent the data storage resource cost of invalid data, and the data storage volume of inefficient data is used to represent the data storage resource cost of inefficient data.
[0181] The fourth determining subunit is used to determine the total storage usage of data in the cluster data resources; the total storage usage is used to represent the cost of the data storage resources.
[0182] The fifth determining subunit is used to determine the ratio of the data storage amount of invalid data to the total data storage amount as the ratio of invalid storage proportion, and to determine the ratio of the data storage amount of inefficient data to the total data storage amount as the ratio of inefficient storage proportion.
[0183] In one possible implementation, the second determining subunit includes:
[0184] The sixth determining subunit is used to determine the invalid computing time in the cluster data resources based on the invalid level calculation diagnostic indicators and the corresponding indicator values, and to determine the inefficient computing time in the cluster data resources based on the inefficient level calculation diagnostic indicators; the invalid computing time is used to represent the cost of invalid data computing resources, and the inefficient computing time is used to represent the cost of inefficient data computing resources; the computing time is obtained by the central processing unit's computing time and the memory's computing time.
[0185] The seventh determining subunit is used to determine the total data usage calculation time in the cluster data resources; the total data usage calculation time is used to represent the data computing resource cost;
[0186] The eighth determining subunit is used to determine the ratio of invalid calculation time to the total data usage calculation time as the ratio of invalid calculation percentage, and to determine the ratio of inefficient calculation time to the total data usage calculation time as the ratio of inefficient calculation percentage.
[0187] In one possible implementation, the device further includes:
[0188] The fifth acquisition unit is used to acquire storage observation indicators under the storage evaluation dimension and computing observation indicators under the computing evaluation dimension; the storage observation indicators include auxiliary judgment indicators of data storage resource cost and / or data storage resource governance indicators; the computing observation indicators include auxiliary judgment indicators of data computing resource cost and / or data computing resource governance indicators.
[0189] The second evaluation unit is used to evaluate the data storage resource cost and the data computing resource cost by combining the storage observation metrics and the computing observation metrics.
[0190] In one possible implementation, the storage observation metrics include one or more of the following:
[0191] The data storage volume increase within adjacent time periods, the month-on-month increase in data storage volume, the sorting result of the data storage volume increase within a preset time period, the data storage volume of invalid and / or inefficient data that has been addressed within the preset time period, the monetary value corresponding to the data storage volume of the addressed invalid and / or inefficient data, the total data storage volume of invalid data within the preset time period, and the total data storage volume of inefficient data within the preset time period.
[0192] The calculated observation indicators include one or more of the following:
[0193] The ranking results include: total number of tasks, increase in task volume within adjacent time periods, task execution time, calculation time management results within a preset time period, invalid calculation time to be managed within a preset time period, the corresponding amount of invalid calculation time to be managed within a preset time period, inefficient calculation time to be managed within a preset time period, and the corresponding amount of inefficient calculation time to be managed within a preset time period.
[0194] The governance results during calculation include the governed invalid and / or inefficient calculation times, and the corresponding monetary values for the governed invalid and / or inefficient calculation times.
[0195] It should be noted that the specific implementation of each unit in this embodiment can be found in the relevant descriptions in the above method embodiments. The division of units in this application embodiment is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. The functional units in this application embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. For example, in the above embodiments, the processing unit and the sending unit can be the same unit or different units. The integrated unit can be implemented in hardware or as a software functional unit.
[0196] Based on the method for evaluating cluster data resource costs provided in the above embodiments, this application also provides an electronic device, including: one or more processors; a storage device storing one or more programs thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for evaluating cluster data resource costs described in any of the above embodiments.
[0197] The following is for reference. Figure 6This document illustrates a structural schematic diagram of an electronic device 600 suitable for implementing embodiments of this application. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Android Devices), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (televisions), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0198] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0199] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0200] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of the embodiments of this application.
[0201] The electronic device provided in this application embodiment and the cluster data resource cost assessment method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0202] Based on the method embodiment provided above, this application provides a computer-readable medium storing a computer program thereon, wherein the program, when executed by a processor, implements the method for evaluating cluster data resource costs as described in any of the above embodiments.
[0203] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0204] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0205] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0206] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the aforementioned method for assessing the cost of cluster data resources.
[0207] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0208] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0209] The units described in the embodiments of this application can be implemented in software or in hardware. The name of the unit / module does not necessarily limit the unit itself; for example, a voice data acquisition module can also be described as a "data acquisition module".
[0210] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0211] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0212] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
[0213] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0214] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0215] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0216] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for evaluating the cost of cluster data resources, characterized in that, The method includes: Obtain assessment dimensions for evaluating cluster data resource costs; the cluster data resource costs include data storage resource costs and data computing resource costs, and the assessment dimensions include storage assessment dimensions and computing assessment dimensions. Obtain storage evaluation metrics under the storage evaluation dimension and computing evaluation metrics under the computing evaluation dimension; the storage evaluation metrics include the proportion of invalid storage resource cost in the total data storage resource cost and the proportion of inefficient storage resource cost in the total data storage resource cost; the computing evaluation metrics include the proportion of invalid computing resource cost in the total data computing resource cost and the proportion of inefficient computing resource cost in the total data computing resource cost. Obtain the index values of the storage evaluation index and the calculation evaluation index; wherein, the index values of the storage evaluation index include the ratio of invalid storage ratio and the ratio of inefficient storage ratio, and the index values of the calculation evaluation index include the ratio of invalid computing ratio and the ratio of inefficient computing ratio. The cost of data storage resources is evaluated based on the value of the storage evaluation index, and the cost of data computing resources is evaluated based on the value of the computing evaluation index. The step of evaluating the data storage resource cost based on the value of the storage evaluation index and evaluating the data computing resource cost based on the value of the computing evaluation index includes: The data storage resource cost assessment score is calculated based on the ratio of the invalid storage ratio to the ratio of the inefficient storage ratio, and the data storage resource cost is assessed based on the assessment score. The data computing resource cost is evaluated based on the ratio of invalid computing percentage to the ratio of inefficient computing percentage. Wherein, the invalid data storage resource cost is the data storage resource cost that meets the first storage condition, the inefficient data storage resource cost is the data storage resource cost that meets the second storage condition, the invalid data computing resource cost is the data computing resource cost that meets the first computing condition, and the inefficient data computing resource cost is the data computing resource cost that meets the second computing condition.
2. The method according to claim 1, characterized in that, The method further includes: Obtain storage diagnostic indicators under the storage evaluation dimension and computing diagnostic indicators under the computing evaluation dimension; the storage diagnostic indicators include ineffective storage diagnostic indicators and inefficient storage diagnostic indicators, and the computing diagnostic indicators include ineffective computing diagnostic indicators and inefficient computing diagnostic indicators. The first storage condition is to meet the invalid level storage diagnostic index, the second storage condition is to meet the inefficient level storage diagnostic index, the first calculation condition is to meet the invalid level calculation diagnostic index, and the second calculation condition is to meet the inefficient level calculation diagnostic index. The invalid-level storage diagnostic indicators include indicators that characterize the utilization rate of stored data as 0, and / or indicators that characterize the stored data as not meeting data storage rules; the inefficient-level storage diagnostic indicators include indicators that characterize the utilization rate of stored data as lower than a utilization rate threshold, and / or indicators that characterize the storage type of data as not meeting the target storage type; the invalid-level computation diagnostic indicators include indicators that characterize the efficiency of task-generated data as lower than a first preset efficiency, and / or indicators that characterize the utilization rate of task-generated data as 0; the inefficient-level computation diagnostic indicators include indicators that characterize the efficiency of task-generated data as lower than a second preset efficiency; the second preset efficiency is better than the first preset efficiency. The steps of obtaining the index value of the storage evaluation index and calculating the index value of the evaluation index include: Based on the storage diagnostic indicators and their corresponding values, the values of the storage evaluation indicators are determined. Based on the calculated diagnostic indicators and their corresponding values, the values of the calculated evaluation indicators are determined.
3. The method according to claim 2, characterized in that, The invalid level storage diagnostic indicators include one or more of the following: The data table has no lifecycle set; the number of accesses to the data table within the preset time period is 0; the number of accesses to the data in the database within the preset time period is 0; the number of accesses to the data in the directory within the preset time period is 0; and the lifecycle of the data table is greater than the recommended value. The inefficient storage diagnostic indicators include one or more of the following: The number of small files exceeds the target number, the access frequency of partition data is less than the target frequency, and the stored data is not compressed; The invalidity level calculation diagnostic indicators include one or more of the following: The number of accesses to the task output within the preset time period is 0; the task ends normally within the preset time period but the output is empty; the number of accesses to the dashboard and / or interface within the preset time period is 0; the task fails within the preset time period; and the task queue has zero utilization rate on the day. The inefficient computational diagnostic indicators include one or more of the following: Data resource utilization rate is lower than the target data resource utilization rate, data skew occurs in tasks, tasks are duplicated, task execution time is longer than the target time, the proportion of task queue blocking time exceeds the target proportion, and the proportion of task queue over-issuance time exceeds the target proportion.
4. The method according to claim 2 or 3, characterized in that, The step of determining the value of the storage evaluation index based on the storage diagnostic index and its corresponding value includes: Based on the invalid storage diagnostic indicators and their corresponding values, the data storage volume of invalid data in the cluster data resources is determined, and based on the inefficient storage diagnostic indicators, the data storage volume of inefficient data in the cluster data resources is determined; the data storage volume of invalid data is used to represent the cost of invalid data storage resources, and the data storage volume of inefficient data is used to represent the cost of inefficient data storage resources. Determine the total storage usage of the data in the cluster data resources; the total storage usage is used to represent the cost of the data storage resources. The ratio of the amount of invalid data stored to the total amount of data used is determined as the ratio of invalid storage percentage, and the ratio of the amount of inefficient data stored to the total amount of data used is determined as the ratio of inefficient storage percentage.
5. The method according to claim 2 or 3, characterized in that, The step of determining the value of the calculated evaluation index based on the calculated diagnostic index and its corresponding value includes: Based on the invalid computing diagnostic indicators and their corresponding values, invalid computing time in the cluster data resources is determined, and inefficient computing time in the cluster data resources is determined based on the inefficient computing diagnostic indicators; the invalid computing time is used to represent the cost of invalid data computing resources, and the inefficient computing time is used to represent the cost of inefficient data computing resources; the computing time is obtained from the computing time used by the central processing unit and the computing time used by memory. When determining the total data usage calculation time in the cluster data resources; the total data usage calculation time is used to represent the data computing resource cost. The ratio of invalid calculation time to the total data usage calculation time is determined as the ratio of invalid calculation percentage, and the ratio of inefficient calculation time to the total data usage calculation time is determined as the ratio of inefficient calculation percentage.
6. The method according to claim 1, characterized in that, The method further includes: Obtain storage observation metrics under the storage evaluation dimension and computing observation metrics under the computing evaluation dimension; the storage observation metrics include auxiliary judgment metrics for data storage resource costs and / or data storage resource governance metrics; the computing observation metrics include auxiliary judgment metrics for data computing resource costs and / or data computing resource governance metrics. By combining the storage observation metrics and the computation observation metrics, the data storage resource cost and the data computation resource cost are evaluated.
7. The method according to claim 6, characterized in that, The stored observation metrics include one or more of the following: The data storage volume increase within adjacent time periods, the month-on-month increase in data storage volume, the sorting result of the data storage volume increase within a preset time period, the data storage volume of invalid and / or inefficient data that has been addressed within the preset time period, the monetary value corresponding to the data storage volume of the addressed invalid and / or inefficient data, the total data storage volume of invalid data within the preset time period, and the total data storage volume of inefficient data within the preset time period. The calculated observation indicators include one or more of the following: The ranking results include: total number of tasks, increase in task volume within adjacent time periods, task execution time, calculation time management results within a preset time period, invalid calculation time to be managed within a preset time period, the corresponding amount of invalid calculation time to be managed within a preset time period, inefficient calculation time to be managed within a preset time period, and the corresponding amount of inefficient calculation time to be managed within a preset time period. The governance results during calculation include the governed invalid and / or inefficient calculation times, and the corresponding monetary values for the governed invalid and / or inefficient calculation times.
8. A device for evaluating the cost of cluster data resources, characterized in that, The device includes: The first acquisition unit is used to acquire assessment dimensions for evaluating the cost of cluster data resources; the cost of cluster data resources includes data storage resource costs and data computing resource costs, and the assessment dimensions include storage assessment dimensions and computing assessment dimensions. The second acquisition unit is used to acquire storage evaluation indicators under the storage evaluation dimension and computing evaluation indicators under the computing evaluation dimension; the storage evaluation indicators include the proportion of invalid storage resource cost in the total data storage resource cost and the proportion of inefficient storage resource cost in the total data storage resource cost; the computing evaluation indicators include the proportion of invalid computing resource cost in the total data computing resource cost and the proportion of inefficient computing resource cost in the total data computing resource cost. The third acquisition unit is used to acquire the index values of the storage evaluation index and the calculation evaluation index; wherein, the index values of the storage evaluation index include the ratio of invalid storage ratio and the ratio of inefficient storage ratio, and the index values of the calculation evaluation index include the ratio of invalid calculation ratio and the ratio of inefficient calculation ratio. The first evaluation unit is used to evaluate the data storage resource cost based on the index value of the storage evaluation index, and to evaluate the data computing resource cost based on the index value of the computing evaluation index. The first evaluation unit is further configured to: The data storage resource cost assessment score is calculated based on the ratio of the invalid storage ratio to the ratio of the inefficient storage ratio, and the data storage resource cost is assessed based on the assessment score. The data computing resource cost is evaluated based on the ratio of invalid computing percentage to the ratio of inefficient computing percentage. Wherein, the invalid data storage resource cost is the data storage resource cost that meets the first storage condition, the inefficient data storage resource cost is the data storage resource cost that meets the second storage condition, the invalid data computing resource cost is the data computing resource cost that meets the first computing condition, and the inefficient data computing resource cost is the data computing resource cost that meets the second computing condition.
9. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the cluster data resource cost assessment method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method for evaluating the cost of cluster data resources as described in any one of claims 1-7.
Citation Information
Patent Citations
Big data asset value evaluation system and method
CN113361980A
Systems, apparatus and methods for cost and performance-based management of resources in a cloud environment
US20200314175A1