A method and device for assessing risks caused by improper use of cluster data resources
By analyzing the division and evaluation dimensions of cluster data resources, the risk assessment problem caused by improper use of cluster hardware resources was solved, enabling more accurate risk identification and prevention, and avoiding business incidents.
Patent Information
- Application Number
- CN202311475501.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-11-07
AI Technical Summary
In big data scenarios, improper use of cluster hardware resources can lead to serious business incidents, and existing technologies are insufficient to effectively assess and prevent these risks.
Cluster data resources are divided into data storage resources and data computing resources. Risks are assessed through storage evaluation dimensions and computing evaluation dimensions, respectively. High-risk diagnostic indicators are obtained and trigger quantities are counted to evaluate the usage of data storage and computing resources.
This enables more accurate risk assessment during the use of cluster data resources, allowing for early identification and prevention of potential business impacts and reducing the occurrence of incidents.
Smart Images

Figure CN117407246B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a method and device for evaluating risks caused by improper use of cluster data resources. BACKGROUND
[0002] In a big data scenario, data-related businesses are rapidly developing, and various business scenarios are implemented by relying on cluster resources, such as cluster hardware resources. Hardware can be used to build services, undertake computing tasks, and meet the needs of various business scenarios. After the investment of cluster hardware resources, improper use of cluster hardware resources in the use process of the cluster hardware resources may cause use risks, and after the risk escalates, it may cause serious business accidents. Cluster data resources are quantitative values corresponding to cluster hardware resources, and in order to avoid business accidents as much as possible, it is also necessary to estimate the risks in the use process of cluster data resources in advance. SUMMARY
[0003] Therefore, the present application provides a method and device for evaluating risks caused by improper use of cluster data resources, which is used to evaluate the risk situation of cluster hardware resources in the use process, so as to perceive the risk in advance.
[0004] To solve the above problems, the technical scheme provided by the present application is as follows:
[0005] In a first aspect, the present application provides a method for evaluating risks caused by improper use of cluster data resources, comprising:
[0006] obtaining an evaluation dimension for evaluating the risk of cluster data resources; the cluster data resources include data storage resources and data computing resources, and the evaluation dimension includes a storage evaluation dimension and a computing evaluation dimension;
[0007] obtaining at least one high-risk diagnostic index under the storage evaluation dimension and at least one high-risk diagnostic index under the computing evaluation dimension; the high-risk diagnostic index under the storage evaluation dimension includes an index for representing that the data resource usage of the data storage resources exceeds a first data storage resource limit value and / or an index for representing that the data storage resources do not meet a first data storage resource usage rule, and the high-risk diagnostic index under the computing evaluation dimension includes an index for representing that the data resource usage of the data computing resources exceeds a first data computing resource limit value;
[0008] count the high-risk indicator trigger quantity in the storage evaluation dimension and the high-risk indicator trigger quantity in the calculation evaluation dimension; the high-risk indicator trigger quantity in the storage evaluation dimension is the number of high-risk diagnostic indicators in the storage evaluation dimension met by the data storage resource in use, and the high-risk indicator trigger quantity in the calculation evaluation dimension is the number of high-risk diagnostic indicators in the calculation evaluation dimension met by the data calculation resource in use;
[0009] evaluate the risk caused by improper use of the data storage resource based on the high-risk indicator trigger quantity in the storage evaluation dimension, and evaluate the risk caused by improper use of the data calculation resource based on the high-risk indicator trigger quantity in the calculation evaluation dimension.
[0010] In a second aspect, the present application provides a device for evaluating the risk of improper use of cluster data resources, and the device comprises:
[0011] an acquisition unit configured to acquire evaluation dimensions for evaluating the risk of cluster data resources; the cluster data resources comprise data storage resources and data calculation resources, and the evaluation dimensions comprise a storage evaluation dimension and a calculation evaluation dimension;
[0012] a first determination unit configured to determine at least one high-risk diagnostic indicator in the storage evaluation dimension and at least one high-risk diagnostic indicator in the calculation evaluation dimension; the high-risk diagnostic indicator in the storage evaluation dimension comprises an indicator for representing that the data resource usage of the data storage resource exceeds a first data storage resource limit value and / or an indicator for representing that the data storage resource does not meet a first data storage resource usage rule, and the high-risk diagnostic indicator in the calculation evaluation dimension comprises an indicator for representing that the data resource usage of the data calculation resource exceeds a first data calculation resource limit value;
[0013] a first counting unit configured to count the high-risk indicator trigger quantity in the storage evaluation dimension and the high-risk indicator trigger quantity in the calculation evaluation dimension; the high-risk indicator trigger quantity in the storage evaluation dimension is the number of high-risk diagnostic indicators in the storage evaluation dimension met by the data storage resource in use, and the high-risk indicator trigger quantity in the calculation evaluation dimension is the number of high-risk diagnostic indicators in the calculation evaluation dimension met by the data calculation resource in use;
[0014] a first evaluation unit configured to evaluate the risk caused by improper use of the data storage resource based on the high-risk indicator trigger quantity in the storage evaluation dimension, and evaluate the risk caused by improper use of the data calculation resource based on the high-risk indicator trigger quantity in the calculation evaluation dimension.
[0015] In a third aspect, the present application provides an electronic device, comprising:
[0016] one or more processors;
[0017] a storage device having stored thereon one or more programs,
[0018] when the one or more programs are executed by the one or more processors, the one or more processors implement any of the cluster data resource misusing induced risk assessment methods.
[0019] In a fourth aspect, the present application provides a computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements any of the cluster data resource misusing induced risk assessment methods.
[0020] Therefore, the present application has the following beneficial effects:
[0021] The application provides a risk evaluation method and device caused by improper use of cluster data resources, the cluster data resources being quantitative values corresponding to cluster hardware resources, including data storage resources and data computing resources. An evaluation dimension for evaluating the risk of the cluster data resources is obtained, and the evaluation dimension includes a storage evaluation dimension and a computing evaluation dimension. That is, under the storage evaluation dimension, the risk generated in the use process of the data storage resources is evaluated, and under the computing evaluation dimension, the risk generated in the use process of the data computing resources is evaluated. Then, at least one high-risk diagnosis index under the storage evaluation dimension and at least one high-risk diagnosis index under the computing evaluation dimension are obtained. The high-risk diagnosis index under the storage evaluation dimension is an index for representing that the data resource usage of the data storage resources exceeds a first data storage resource limit value and / or the data storage resources do not meet a first data storage resource usage rule, and the high-risk diagnosis index under the computing evaluation dimension is an index for representing that the data resource usage of the data computing resources exceeds a first data computing resource limit value. If the high-risk diagnosis index is triggered in the use process of the cluster data resources, one or more of the following occurs: the data resource usage of the data storage resources exceeds the first data storage resource limit value, the data storage resources do not meet the first data storage resource usage rule, and the data resource usage of the data computing resources exceeds the first data computing resource limit value, thereby determining that the use of the cluster data resources has a high risk and may cause an accident. On this basis, the number of high-risk diagnosis indexes under the storage evaluation dimension that are met in the use process of the data storage resources is counted and recorded as the high-risk index trigger quantity under the storage evaluation dimension, and the number of high-risk diagnosis indexes under the computing evaluation dimension that are met in the use process of the data computing resources is counted and recorded as the high-risk index trigger quantity under the computing evaluation dimension. The high-risk index trigger quantity is used to represent the degree of high risk, thereby evaluating the risk caused by improper use of the data storage resources by using the high-risk index trigger quantity under the storage evaluation dimension and evaluating the risk caused by improper use of the data computing resources by using the high-risk index trigger quantity under the computing evaluation dimension, and completing the risk evaluation of the cluster data resources. The greater the high-risk index trigger quantity, the more the number of risk events, and the higher the degree of risk of improper use of the cluster data resources.
[0022] In the application, the cluster data resources are divided into data storage resources and data computing resources, so that the division of the cluster data resources is more reasonable and closer to the actual situation. On this basis, two evaluation dimensions, the storage evaluation dimension and the computing evaluation dimension, are determined, and the risks caused by improper use of the data storage resources and the data computing resources are evaluated from the storage evaluation dimension and the computing evaluation dimension respectively, so that the risk evaluation of the cluster data resources in the use process is more accurate. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1A flowchart illustrating a method for assessing risks arising from improper use of cluster data resources, provided in an embodiment of this application;
[0024] Figure 2 A schematic diagram illustrating risk assessment levels provided for an embodiment of this application;
[0025] Figure 3 A schematic diagram illustrating an indicator for assessing the risk of data storage resources, provided as an embodiment of this application;
[0026] Figure 4 A schematic diagram illustrating an indicator for assessing the risk of data computing resources, provided as an embodiment of this application;
[0027] Figure 5 A schematic diagram of the structure of an assessment device for risks caused by improper use of cluster data resources, provided in an embodiment of this application;
[0028] Figure 6 This is a schematic diagram of the basic structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0030] To facilitate understanding and explanation of the technical solutions provided in the embodiments of this application, the background technology of this application will be described first.
[0031] To facilitate understanding of this application, a method for assessing the risks caused by improper use of cluster data resources, provided in an embodiment of this application, will be described below with reference to the accompanying drawings. For example, this method for assessing the risks caused by improper use of cluster data resources can be performed by terminal devices and / or servers, etc., and is not limited here.
[0032] See Figure 1 As shown, this figure is a flowchart of a method for assessing the risks caused by improper use of cluster data resources according to an embodiment of this application. Figure 1 As shown, the method may include S101-S104:
[0033] S101: Obtain assessment dimensions for evaluating cluster data resource risks; cluster data resources include data storage resources and data computing resources, and the assessment dimensions include storage assessment dimensions and computing assessment dimensions.
[0034] The cluster hardware resources can include memory, central processing unit, disk, hard disk, and various resources for building services and undertaking computing tasks. Among them, the central processing unit and other processing hardware devices can be used to implement various computing tasks, and the memory, disk, hard disk, and other storage hardware devices can be used to implement the storage of data in databases, partitions, data tables, and the like. It can be known that the implementation of computing tasks in the cluster depends on the calculation of central processing unit and other devices, and the output of computing tasks (such as output data table) depends on the storage of memory, disk, and other devices.
[0035] The cluster data resources are quantitative values corresponding to the cluster hardware resources (referred to as cluster hardware) and are used to quantify the cluster hardware resources. The cluster data resources include data storage resources and data computing resources. Among them, the data storage resources are data resources for storing data, which can be represented by the data storage space (quantitative value corresponding to the cluster hardware) of the memory, disk, hard disk, and other cluster hardware. For example, the disk has a data storage space of 80 GB, and 80 GB can be regarded as the data storage resource of the disk. The data computing resources are data resources that can be used to output data, which can be quantitatively represented by the time of the central processing unit and other cluster hardware that can be used to output data.
[0036] In addition, the cluster hardware resources, cluster data resources, and the like in the embodiments of the present application can be determined based on the dimension of the data resource user. For example, the data resource user can be an organization, which includes people, groups, departments, and business lines. That is, the data resource user can be an individual, a group, a department, a business line, and the like. Therefore, when the data resource user is a department, the cluster hardware resources and the cluster data resources described above are the total data resources allocated to the department, and the use of the cluster data resources is also the use of the cluster data resources within the department by the department. Therefore, the evaluation of the cluster data resource risk provided by the embodiments of the present application can specifically be the evaluation of the cluster data resource risk of each department. The rest of the data resource users are similar, and will not be described here.
[0037] The data storage resources can include data resources of dimensions such as partition, column, table, Topic, data set, library, and resource group, and are represented by the data storage space of the partition, column, table, Topic, data set, library, and resource group. The data computing resources can include data resources of dimensions such as stage, application, instance, task, and queue, and are represented by the time used for output data of the stage, application, instance, task, and queue. Details can be found below.
[0038] Based on the above, the cluster data resources are divided into data storage resource usage and data computing resource usage in the usage process. For example, a computing task needs to consume several data computing resources, and the storage of the data output by the computing task needs to consume several data storage resources. The cluster data resources in the usage process may be affected by improper use. For example, when all data storage resources have been used, new data cannot be written, resulting in the failure of new data output, which affects the business. Therefore, it is necessary to determine the risks caused by improper use of cluster data resources in the usage process and evaluate them to intervene in advance based on the evaluation results, so as to avoid the impact of improper use of cluster data resources on the business as much as possible. Specifically, when evaluating the risks in the usage process of cluster data resources, a terminal device and / or a server can obtain two evaluation dimensions, including a storage evaluation dimension and a computing evaluation dimension. In the storage evaluation dimension, the risks of data storage resources in the usage process are determined and evaluated; in the computing evaluation dimension, the risks of data computing resources in the usage process are determined and evaluated.
[0039] Referring to Figure 2 , Figure 2 A risk evaluation level diagram is provided for the embodiments of the present application. As shown in Figure 2 , according to the impact of the risks caused by improper use of data resources on the business, the risk levels of cluster data resources in the usage process are divided into three levels, including a low risk level, a high risk level, and an accident level. Among them, the low risk level indicates a low risk degree, which has not yet affected the business and needs to be concerned and managed, otherwise the low risk level will be upgraded to the high risk level, resulting in the occurrence of unstable conditions. The high risk level indicates a high risk degree, which is likely to affect the business and needs to be controlled and intervened in time, otherwise the high risk level will be upgraded to the accident level, resulting in the occurrence of business accidents. The accident level indicates that a business accident has occurred and has affected the business. The accident level can be divided into L0, L1, L2, L3, and Notice levels (risk gradually increases). In the embodiments of the present application, it is necessary to determine the risks caused by improper use of cluster data resources in the usage process and avoid affecting the business as much as possible, so it is necessary to pay attention to the risks of the low risk level and the high risk level. In addition, since the risk of the high risk level is high, it may affect the business soon, therefore, the risk of the high risk level with a high risk degree can be focused on.
[0040] S102: Obtain at least one high-risk diagnostic indicator in the storage evaluation dimension and at least one high-risk diagnostic indicator in the calculation evaluation dimension. The high-risk diagnostic indicator in the storage evaluation dimension includes an indicator for representing that the data resource usage of the data storage resource exceeds the first data storage resource limit value and / or an indicator for representing that the data storage resource does not meet the first data storage resource usage rule, and the high-risk diagnostic indicator in the calculation evaluation dimension includes an indicator for representing that the data resource usage of the data calculation resource exceeds the first data calculation resource limit value.
[0041] After the terminal device and / or the server obtain two evaluation dimensions such as the storage evaluation dimension and the calculation evaluation dimension, the terminal device and / or the server further obtain at least one high-risk diagnostic indicator in the storage evaluation dimension and at least one high-risk diagnostic indicator in the calculation evaluation dimension.
[0042] It can be known that the risk in the process of using the cluster data resource is generated by a specific event, and the event is related to the use of the cluster data resource. The high-risk diagnostic indicator is used to determine whether a high-risk event occurs. When the event related to the cluster data resource meets the high-risk diagnostic indicator, it indicates that the event is a high-risk event, which may cause the occurrence of a high-risk problem, and the risk level of the cluster data resource related to the high-risk event is a high-risk level.
[0043] The high-risk diagnostic indicator in the storage evaluation dimension includes an indicator for representing that the data resource usage of the data storage resource exceeds the first data storage resource limit value and / or an indicator for representing that the data storage resource does not meet the first data storage resource usage rule, and the number of indicators can be at least one. The data resource usage of the data storage resource exceeding the first data storage resource limit value, the data storage resource not meeting the first data storage resource usage rule, and the like are all high-risk events in the storage evaluation dimension. The high-risk diagnostic indicator in the calculation evaluation dimension is an indicator for representing that the data resource usage of the data calculation resource exceeds the first data calculation resource limit value, and the number of indicators can be at least one. The data resource usage of the data calculation resource exceeding the first data calculation resource limit value is a high-risk event in the calculation evaluation dimension.
[0044] It can be understood that as long as an event related to the data storage resource triggers any high-risk diagnostic indicator in the storage evaluation dimension, it can be determined that the event related to the data storage resource is a high-risk event in the storage evaluation dimension. As long as an event related to the data calculation resource triggers any high-risk diagnostic indicator in the calculation evaluation dimension, it can be determined that the event related to the data calculation resource is a high-risk event in the calculation evaluation dimension.
[0045] Referring to Figure 3 , Figure 3This is a schematic diagram of an indicator for assessing the risk of data storage resources, provided as an embodiment of this application.
[0046] As an optional example, combined Figure 3 High-risk diagnostic indicators under the storage assessment dimension include one or more of the following:
[0047] The following conditions must be met: remaining days of storage usage are less than or equal to the first day; the increase in storage volume is greater than the first percentage; the proportion of small files is greater than the first percentage; the proportion of bad blocks on the disk is greater than the first bad block proportion; the write-prohibition time is greater than the target time; the storage utilization rate exceeds the target utilization rate; and the set lifecycle is less than the recommended value.
[0048] Among these parameters, the remaining days of storage usage, the month-on-month increase in storage volume, the proportion of small files, the proportion of bad blocks on the disk, the trigger write-ban time, the storage utilization rate, and the set lifecycle are all specific to the data resource user. For example, if the data resource user is a department, then "remaining days of storage usage" specifically refers to the department's remaining days of storage usage. The other parameters are similar and will not be elaborated here.
[0049] It is understandable that parameters such as remaining days of storage use, month-on-month increase in storage volume, proportion of small files, proportion of bad blocks on disk, trigger write-prohibition time, storage utilization rate, and set lifecycle are all related events in the use of data storage resources, and all involve data storage resources.
[0050] Data storage resources include used data storage resources and remaining data storage resources. "Remaining days of storage use" refers to the number of days the remaining data storage resources are available for use by data resource users. "Remaining days of storage use less than or equal to the first day" can be exemplified as follows: Figure 3 The indicator shown is "Remaining storage days less than or equal to 10 days". By comparing the remaining storage days with the first day, it can be determined whether the data resource user has sufficient remaining data storage resources. When the remaining storage days are less than or equal to the first day, the "Remaining storage days less than or equal to the first day" indicator is triggered. This indicates that the remaining data storage resources are insufficient to meet the daily increase in data storage volume, indicating a high level of risk in data storage resources. It is necessary to expand the available data storage resources as soon as possible to avoid new data being unable to be written.
[0051] In the embodiments of the present application, the specific value of the first number of days is not limited, and can be determined according to actual scenarios. For example, the first number of days is set to 10 days, and the department has 1 PB of data storage capacity added every day. If the department currently has 8 PB of remaining data storage resources, the corresponding remaining storage usage days are 8 days, and the remaining storage usage days (8 days) of the department are less than the first number of days (10 days), it is determined that the risk level of the data storage resources of the department is high, and the available data storage resources of the department need to be expanded as soon as possible.
[0052] In actual application, the terminal device and / or the server first obtain the remaining storage usage days based on the above method, and then the terminal device and / or the server compare the remaining storage usage days with the first number of days to determine whether the remaining storage usage days are less than or equal to the first number of days. If yes, the terminal device and / or the server determines that the data storage resources meet the high-risk diagnostic index in the use process.
[0053] The "storage value increase year-on-year" is used to determine whether the use of data storage resources has increased sharply. When the storage value increase year-on-year is greater than the first proportion, the "storage value increase year-on-year greater than the first proportion" index is triggered, which indicates that the use of data storage resources has increased sharply, and the situation of insufficient data storage resources may occur, and at this time it is determined that the risk level of data storage resources is high. The "storage value increase year-on-year greater than the first proportion" can be exemplified as "data storage resource use daily increase year-on-year greater than 400%" in the above table, that is, the first proportion is 400%. The "data storage resource use daily increase year-on-year" is specifically the ratio of the difference between the daily added data storage capacity and the yesterday added data storage capacity to the yesterday added data storage capacity. Figure 3
[0054] In actual application, the terminal device and / or the server first calculates the storage value increase year-on-year based on the above method, and then the terminal device and / or the server compares the storage value increase year-on-year with the first proportion to determine whether the storage value increase year-on-year is greater than the first proportion. If yes, the terminal device and / or the server determines that the data storage resources meet the high-risk diagnostic index in the use process.
[0055] The "small file proportion" is specifically the ratio of the number of small files to the total number of small files (after removing the folders), and is used to determine whether the number of small files is a reasonable value. When the small file proportion is greater than the first proportion, the "small file proportion greater than the first proportion" index is triggered, which indicates that the number of small files is large, and the number of small files is not a reasonable value. At this time, the device is prone to freezing problems, which may affect the business running in the device, and it is determined that the risk level of data storage resources is high. Here, the first proportion is not limited and can be set according to actual scenarios. For example, the first proportion can be 70% as shown in the above table. Figure 3
[0056] In actual application, the terminal device and / or the server first calculates the small file ratio based on the above manner, and then compares the small file ratio with the first ratio to determine whether the small file ratio is greater than the first ratio. If yes, the terminal device and / or the server determines that the data storage resource meets the high risk diagnosis index in the use process.
[0057] The "disk bad block ratio" is specifically the ratio of the number of disks with bad blocks to the total number of disks, and is used to determine whether the remaining data storage resource is sufficient. When the disk bad block ratio is greater than the first bad block ratio, the "disk bad block ratio greater than the first bad block ratio" index is triggered, which indicates that the number of disks with bad blocks is relatively large, and the situation that the remaining data storage resource is insufficient may occur, and it is determined that the risk degree of the data storage resource is high. Here, the first bad block ratio is not limited, and can be set according to the actual scene. For example, the first bad block ratio can be 10% as shown. Figure 3
[0058] In actual application, the terminal device and / or the server first calculates the disk bad block ratio based on the above manner, and then compares the disk bad block ratio with the first bad block ratio to determine whether the disk bad block ratio is greater than the first bad block ratio. If yes, the terminal device and / or the server determines that the data storage resource meets the high risk diagnosis index in the use process.
[0059] In the use process of the data storage resource, if two tasks write to the same partition or the storage component appears abnormal at the same time, it may cause the write data to be blocked or prohibited. The time when the data is prohibited to write is the trigger prohibition time, and the trigger prohibition time is used to determine whether the data is successfully written. When the trigger prohibition time of the data is greater than the target time, the "trigger prohibition time greater than the target time" index is triggered, which indicates that the time when the data is prohibited to write is too long, and it may affect the running business. Since the "trigger prohibition time" parameter is also related to the use of the data storage resource, it is determined that the risk degree of the data storage resource is high. Here, the target time is not limited, and can be set according to the actual scene.
[0060] In actual application, the terminal device and / or the server first acquires the trigger prohibition time, and then compares the trigger prohibition time with the target time to determine whether the trigger prohibition time is greater than the target time. If yes, the terminal device and / or the server determines that the data storage resource meets the high risk diagnosis index in the use process.
[0061] The "storage usage rate" is the ratio of the used data storage resources to the total data storage resources, and is used to determine whether the remaining data storage resources are sufficient. When the storage usage rate exceeds the target usage rate, the "storage usage rate exceeds the target usage rate" indicator is triggered, which indicates that the remaining data storage resources may not be sufficient, and the risk level of the data storage resources is determined to be high. Here, the target usage rate is not limited and can be set according to the actual scene.
[0062] In actual application, the terminal device and / or the server first acquire the storage usage rate based on the above manner, and then the terminal device and / or the server compare the storage usage rate with the target usage rate to determine whether the storage usage rate exceeds the target usage rate. If yes, the terminal device and / or the server determine that the data storage resources meet the high-risk diagnostic indicator in the use process.
[0063] The "lifecycle" is specifically the lifecycle of a data table and the like. The lifecycle (Time To Live, TTL) is the time allowed for data to be stored in the data warehouse, and can also be understood as the time allowed for data to be stored in the hard disk. In actual application, data can be stored using a data table. The smaller the TTL value of the data table is set, the more frequent the data update will be, and the more likely it is to increase the device burden. The larger the TTL value of the data table is set, the longer the data storage time will be, and the stored data can be outdated. Therefore, it is necessary to set the TTL of the data table and to set the TTL to an appropriate value. Generally, a recommended value of the TTL can be given, which is an appropriate TTL value, and the setting of the TTL value should follow the recommended value of the TTL. Therefore, it is necessary to determine whether the lifecycle of the data table and the like is appropriate. When the set lifecycle TTL is less than the target time, the "set lifecycle is less than the recommended value" indicator is triggered, indicating that the currently set lifecycle is unreasonable, which can cause data errors to be deleted and affect the current running business. Since the "lifecycle" parameter is also related to the use of data storage resources, the risk level of the data storage resources is determined to be high.
[0064] In actual application, the terminal device and / or the server first acquires the lifecycle corresponding to the data table and the like, and then the terminal device and / or the server compare the acquired lifecycle with the recommended value to determine whether the set lifecycle is less than the recommended value. If yes, the terminal device and / or the server determine that the data storage resources meet the high-risk diagnostic indicator in the use process.
[0065] It can be understood that the high-risk diagnostic indicators listed above under the storage evaluation dimension belong to indicators for indicating that the data resource usage of the data storage resource exceeds the first data storage resource limit value, or belong to indicators for indicating that the data storage resource does not meet the first data storage resource usage rule. The data storage resource includes used data storage resource and remaining data storage resource. Among them, the indicators such as storage remaining usage days less than or equal to the first day, storage amount increase value compared to the last period greater than the first proportion, disk bad block proportion greater than the first bad block proportion, and storage usage rate exceeding the target usage rate are related to the used data storage resource or the remaining data storage resource, so these indicators belong to indicators for indicating that the data resource usage of the data storage resource exceeds the first data storage resource limit value. Taking "storage remaining usage days less than or equal to the first day" as an example, the storage remaining usage days less than or equal to the first day correspondingly indicates that the used data storage resource exceeds the first data storage resource limit value. In addition, the indicators such as the set life cycle being less than the recommended value, the small file proportion being greater than the first proportion, and the trigger write-prohibited time being greater than the target time involve the data storage resource usage rule, and belong to indicators for indicating that the data storage resource does not meet the first data storage resource usage rule. Taking "small file proportion greater than first proportion" as an example, the first data storage resource usage rule can specify that the small file proportion needs to be less than or equal to the first proportion, and the rest is similar and will not be repeated here.
[0066] Referring to Figure 4 , Figure 4 An index diagram for evaluating data computing resource risk provided by an embodiment of the present application. As an optional example, in combination with Figure 4 The high-risk diagnostic indicators under the evaluation dimension include one or more of the following:
[0067] Queue blocking time length exceeds the first time length, scheduling task failure rate is greater than the first failure rate, scheduling task running time length is greater than the first running time length, queue usage computing time proportion exceeds the first computing time proportion, SLA task breaks line.
[0068] The high-risk diagnostic indicators listed above under the computing evaluation dimension can be divided according to the "task" and "queue" dimensions under the computing evaluation dimension. For example, the queue blocking time length exceeds the first time length, and the queue usage computing time proportion exceeds the first computing time proportion are high-risk diagnostic indicators related to "queue". The scheduling task failure rate is greater than the first failure rate, the scheduling task running time length is greater than the first running time length, and the SLA task breaks line are high-risk diagnostic indicators related to "task". In addition, the high-risk diagnostic indicators under the "task" and "queue" dimensions are not limited, and corresponding high-risk diagnostic indicators can also be constructed from the dimensions of stage, application, instance, etc.
[0069] It can be understood that the queue blocking duration, the scheduling task failure rate, the scheduling task running duration, the queue usage computing time proportion, the SLA task, and the like are related events of the data computing resource in the use process. For example, taking the “queue” related indicators (such as the queue blocking duration and the queue usage computing time proportion) as an example, the data resource configured for the queue is a central processing unit of 1000 cores and a memory of 10 TB. When the queue is used, the data computing resource thereof is consumed, and the consumed data computing resource can be expressed as computing time. The computing time can be understood as a unit for evaluating the size of the data computing resource used for a task / queue. The computing time is obtained from the central processing unit usage computing time and the memory usage computing time, specifically, the computing time takes the larger one of the central processing unit usage computing time and the memory usage computing time. Among them, 1 core of the central processing unit is determined to use 1 hour as 1 computing time, and 4 GB of the memory is determined to use 1 hour as 1 computing time. For example, a task uses a central processing unit core and uses 8 GB, and the task runs for 1 hour, then the central processing unit usage computing time is 1 computing time, the memory usage computing time is 2 computing times, and the computing time corresponding to the task is the larger one of 1 computing time and 2 computing times, that is, 2 computing times. The “task” is similar and will not be described here.
[0070] In addition, the queue blocking duration, the scheduling task failure rate, the scheduling task running duration, the queue usage computing time proportion, the SLA task, and the like are parameters for the data resource user. For example, the data resource user is a department, and the “SLA task” is specifically an SLA task constructed in the department. The remaining parameters are similar and will not be described here.
[0071] The “queue blocking duration” is the total suspension duration of the instances in the queue, and is used to determine whether the data computing resource is normally used. When the “queue blocking duration” exceeds the first duration, it is considered that the total suspension duration of the instances in the queue is long, the queue is blocked for a long time, and the data computing resource is continuously consumed. At this time, it is determined that the data computing resource is not normally used, and the risk degree of the data computing resource corresponding to the queue is high. The first duration is not limited here and can be set according to the actual scene. For example, it can be 30 minutes as shown in the following table. Figure 4
[0072] In actual application, the terminal device and / or the server first acquire the queue blocking duration, and then the terminal device and / or the server compare the queue blocking duration with the first duration to determine whether the queue blocking duration exceeds the first duration. If yes, the terminal device and / or the server determines that the data computing resource in the use process meets the high risk diagnosis indicator.
[0073] The scheduling task failure rate is a ratio of the number of scheduling task failures to the total number of scheduling tasks, and is used to determine the waste of data computing resources. When the scheduling task failure rate is greater than a first failure rate, it is determined that the scheduling task failure rate is high, and the data computing resources used by the failed task are wasted. At this time, it is determined that the waste of data computing resources is high, and the risk level of data computing resources is high. The first failure rate is not limited here and can be set according to the actual scene, for example, it can be 10% as shown in the following formula (1). Figure 4
[0074] In actual application, the terminal device and / or the server first acquires the scheduling task failure rate, and then compares the scheduling task failure rate with the first failure rate to determine whether the scheduling task failure rate is greater than the first failure rate. If yes, the terminal device and / or the server determines that the data computing resources meet the high-risk diagnostic index in the use process.
[0075] The scheduling task running time length is the total consumption time length of the scheduling task, and is used to determine whether the scheduling task runs for a long time. When the scheduling task running time length is greater than a first running time length, it indicates that the scheduling task runs for a long time. The scheduling task running for a long time is a high-time-consumption task, which will consume data computing resources all the time, causing the data computing resources not to be released, and possibly causing the task to hang. At this time, it is determined that the risk level of data computing resources is high. The first running time length is not limited here and can be set according to the actual scene, for example, it can be 24 hours as shown in the following formula (2). Figure 4
[0076] In actual application, the terminal device and / or the server first acquires the scheduling task running time length, and then compares the scheduling task running time length with the first running time length to determine whether the scheduling task running time length is greater than the first running time length. If yes, the terminal device and / or the server determines that the data computing resources meet the high-risk diagnostic index in the use process.
[0077] The queue usage computing time length ratio is specifically a ratio of the queue usage computing time length to the total computing time length, and is used to determine the queue load. The total computing time length can be the total computing time length of the data resource user. When the queue usage computing time length ratio exceeds a first computing time length ratio, it indicates that the queue is overfull, the data computing resources consumed by the queue are more, and the queue may have problems in running at this time. At this time, it is determined that the risk level of data computing resources is high. The first computing time length ratio is not limited here and can be set according to the actual scene.
[0078] In actual application, the terminal device and / or the server first obtains the queue usage calculation time proportion based on the above manner, and then compares the queue usage calculation time proportion with the first calculation time proportion to determine whether the queue usage calculation time proportion exceeds the first calculation time proportion. If yes, the terminal device and / or the server determines that the data calculation resource meets the high-risk diagnosis index in the usage process.
[0079] The SLA task is a task of signing an SLA agreement. Usually, the time required for completing the SLA task is agreed. When the time is exceeded, the SLA task is not completed, which indicates that the SLA task is broken. For example, the time required for completing the SLA task is agreed to be 2:30. Before 2:30, the SLA task needs to be normally output. However, if the SLA task is not completed at 2:30, it indicates that the SLA task is broken. The SLA task being broken may cause the subsequent related task to be delayed. Through the judgment of whether the SLA task is broken, the timeliness of the output data can be known. That is, the SLA task being broken indicates that the output data is not timely, and at this time, the data calculation resource consumed by the SLA task may have a problem, and it is determined that the risk degree of the data calculation resource is high.
[0080] In actual application, the terminal device and / or the server first determines the SLA task, and then judges whether the SLA task is completed within the corresponding required time to determine whether the SLA task is broken. If yes, the terminal device and / or the server determines that the data calculation resource meets the high-risk diagnosis index in the usage process.
[0081] It can be understood that the high-risk diagnosis indexes in the above listed calculation evaluation dimensions can all belong to indexes for indicating that the data resource usage amount of the data calculation resource exceeds the first data calculation resource limit value. Taking the "queue blocking time length exceeding the first time length" as an example, when the queue blocking time length exceeds the first time length, it is considered that the data resource usage amount of the data calculation resource is too large and has exceeded the first data calculation resource limit value. The rest of the indexes are similar, which will not be described here.
[0082] S103: statistics of high-risk index trigger amount in the storage evaluation dimension and high-risk index trigger amount in the calculation evaluation dimension; the high-risk index trigger amount in the storage evaluation dimension is the number of high-risk diagnosis indexes in the storage evaluation dimension met by the data storage resource in the usage process, and the high-risk index trigger amount in the calculation evaluation dimension is the number of high-risk diagnosis indexes in the calculation evaluation dimension met by the data calculation resource in the usage process.
[0083] After determining the at least one high-risk indicator in the storage evaluation dimension, it can be judged whether the data storage resource triggers / satisfies the high-risk indicator in the storage evaluation dimension in the use process. Specifically, it is judged whether some related events of the data storage resource in the use process trigger or satisfy the high-risk indicator in the storage evaluation dimension. Further, the number of high-risk diagnostic indicators in the storage evaluation dimension triggered or satisfied by the data storage resource in the use process is counted.
[0084] For example, there are 10 high-risk diagnostic indicators in the storage evaluation dimension, and the data storage resource triggers 5 of them in the use process. Therefore, the high-risk indicator triggering quantity in the storage evaluation dimension is 5. The event (related to the use of the data storage resource) triggering the high-risk diagnostic indicator in the storage evaluation dimension is a high-risk event.
[0085] In addition, the calculation of the high-risk indicator triggering quantity in the evaluation dimension is similar, which will not be described here.
[0086] Based on the above, in a possible implementation, the embodiment of the application provides a specific implementation of counting the high-risk indicator triggering quantity in the storage evaluation dimension and calculating the high-risk indicator triggering quantity in the evaluation dimension, which includes:
[0087] A1: In combination with the use of the data storage resource, the index value corresponding to the high-risk diagnostic indicator in the storage evaluation dimension is determined, and the high-risk indicator triggering quantity in the storage evaluation dimension is determined according to the index value corresponding to the high-risk diagnostic indicator in the storage evaluation dimension.
[0088] Specifically, when the use of the data storage resource is known, the remaining situation of the data storage resource can also be obtained. In the use process of the data storage resource, the index value corresponding to each high-risk diagnostic indicator in the above-listed storage evaluation dimension can be obtained. For example, according to the remaining situation of the data storage resource, it is determined that the storage remaining use days are 5 days, and if the first day is 10 days, the index value of “the storage remaining use days are less than or equal to the first day” is “yes”. The rest of the indexes are similar, which will not be described here.
[0089] It can be understood that as long as the index value is “yes”, it indicates that the index is triggered or satisfied. Therefore, the high-risk diagnostic indicators in the storage evaluation dimension with the index value of “yes” are counted to obtain the high-risk indicator triggering quantity in the storage evaluation dimension.
[0090] A2: In combination with the use of the data computing resource, the index value corresponding to the high-risk diagnostic indicator in the computing evaluation dimension is determined, and the high-risk indicator triggering quantity in the computing evaluation dimension is determined according to the index value corresponding to the high-risk diagnostic indicator in the computing evaluation dimension.
[0091] Specifically, the usage of the data computing resource is learned, and in the process of using the data computing resource, the index value corresponding to the high-risk diagnostic index in each computing evaluation dimension listed above can be obtained. For example, whether the SLA task is broken is judged according to the completion time of the SLA task. If the SLA task is not broken, the index value of the "SLA task broken" index is "no". The rest of the indexes are similar and will not be repeated here.
[0092] Therefore, the high-risk diagnostic indexes in each computing evaluation dimension with the index value of "yes" are counted to obtain the high-risk index trigger quantity in the computing evaluation dimension.
[0093] S104: Based on the high-risk index trigger quantity in the storage evaluation dimension, the risk caused by improper use of the data storage resource is evaluated, and based on the high-risk index trigger quantity in the computing evaluation dimension, the risk caused by improper use of the data computing resource is evaluated.
[0094] The high-risk index trigger quantity in the storage evaluation dimension and the high-risk index trigger quantity in the computing evaluation dimension are result class indexes, which can directly evaluate the risk of the data storage resource and the data computing resource, and directly determine whether there is a problem in the current data resource usage. When the high-risk index trigger quantity reaches a certain number, it is determined that the risk caused by improper use of the cluster data resource is high. For example, the number of high-risk diagnostic indexes in the storage evaluation dimension is 3, and the high-risk index trigger quantity in the storage evaluation dimension is 2, so it is determined that the risk degree is high.
[0095] As an optional example, when it is determined that the risk caused by improper use of the data storage resource is high, the risk level thereof is determined to be a high-risk level. Similarly, when it is determined that the risk caused by improper use of the data computing resource is high, the risk level thereof is determined to be a high-risk level. At this time, the risk level can be pushed to the user, so that the user can timely perceive the risk and timely intervene in the processing.
[0096] In addition, the high-risk diagnostic index used to obtain the high-risk index trigger quantity is a process class index, and the process class index can be used for risk attribution. For example, the high-risk event can be learned through the triggered high-risk diagnostic index, and the high-risk event is the cause of the high risk, thereby realizing risk attribution.
[0097] In addition, the number of high-risk events in the storage evaluation dimension can be counted, and the number of high-risk events in the calculation evaluation dimension can be calculated. The risk caused by improper use of data storage resources is evaluated based on the number of high-risk events in the storage evaluation dimension, and the risk caused by improper use of data calculation resources is evaluated based on the number of high-risk events in the calculation evaluation dimension. When the number of high-risk events in the storage evaluation dimension is high, it is determined that the data storage resources have caused high risk due to improper use during use, and the risk level is high. When the number of high-risk events in the calculation evaluation dimension is high, it is determined that the data calculation resources have caused high risk due to improper use during use, and the risk level is high. The number of high-risk events in the storage evaluation dimension is the number of events that trigger a high-risk diagnostic indicator in the storage evaluation dimension (any indicator can be used), and the event is related to the use of data storage resources. The number of high-risk events in the calculation evaluation dimension is the number of events that trigger a high-risk diagnostic indicator in the calculation evaluation dimension (any indicator can be used), and the event is related to the use of data calculation resources. For example, there are 10 events related to the use of data storage resources, and the events are matched with high-risk diagnostic indicators in the storage evaluation dimension. As long as any high-risk diagnostic indicator in the storage evaluation dimension is triggered, the event is determined to be a high-risk event in the storage evaluation dimension. Therefore, the number of high-risk events in the storage evaluation dimension can be counted, for example, 4.
[0098] Based on the related content of S101-S104, the cluster data resources are divided into data storage resources and data calculation resources in the embodiments of the present application, so that the division of the cluster data resources is more reasonable and closer to the actual situation. On this basis, two evaluation dimensions such as the storage evaluation dimension and the calculation evaluation dimension are obtained, and the risk of the data storage resources during use and the risk of the data calculation resources during use are evaluated from the storage evaluation dimension and the calculation evaluation dimension respectively, which can make the risk evaluation of the cluster data resources during use more accurate. The high-risk indicator trigger quantity is used as a high-risk evaluation indicator to directly evaluate the risk of the cluster data resources during use. The greater the high-risk indicator trigger quantity, the more the number of risk events, and the higher the risk level of improper use of the cluster data resources.
[0099] It can be seen that the above content introduces how to evaluate whether the risk level of the cluster data resources during use is high. Before the cluster data resources have high-risk problems, it can also be evaluated whether the risk level of the cluster data resources during use is low, and low-risk problems of the cluster data resources during use are determined, so that the risk of the cluster data resources can be perceived earlier. For details, see the following.
[0100] In a possible implementation, the method for evaluating the risk caused by improper use of the cluster data resources provided by the embodiments of the present application can further include the following steps B1-B3:
[0101] B1: obtaining and calculating at least one low-risk diagnosis indicator in the evaluation dimension of storage; the low-risk diagnosis indicator in the evaluation dimension of storage includes an indicator for representing that the data resource usage of the data storage resource exceeds the second data storage resource limit value and / or an indicator for representing that the data storage resource does not meet the second data storage resource usage rule, and the low-risk diagnosis indicator in the evaluation dimension of calculation includes an indicator for representing that the data resource usage of the data calculation resource exceeds the second data calculation resource limit value; the second data storage resource limit value is less than the first data storage resource limit value, and the second data calculation resource limit value is less than the first data calculation resource limit value.
[0102] It can be known that, in the low-risk judgment process of the data storage resource, the second data storage resource limit value is set to be less than the first data storage resource limit value to represent that the risk degree of the data storage resource when triggering the low-risk diagnosis indicator in the evaluation dimension of storage is relatively low. In the low-risk judgment process of the data calculation resource, the second data calculation resource limit value is set to be less than the first data calculation resource limit value to represent that the risk degree of the data calculation resource when triggering the low-risk diagnosis indicator in the evaluation dimension of calculation is relatively low.
[0103] Again referring to Figure 3 As shown in Figure 3 , the low-risk diagnosis indicator in the evaluation dimension of storage can include one or more of the following:
[0104] the storage remaining usage days are less than or equal to the second number of days, the storage amount increase value is greater than the second proportion, the small file proportion is greater than the second proportion, and the disk bad block proportion is greater than the second bad block proportion; the second number of days is greater than the first number of days, the second proportion is less than the first proportion, the second proportion is less than the first proportion, and the second bad block proportion is less than the first bad block proportion.
[0105] It can be understood that the storage remaining usage days, the storage amount increase value, the small file proportion, and the disk bad block proportion are all related events of the data storage resource in the use process.
[0106] It can also be understood that, by setting the second number of days to be greater than the first number of days, the second proportion to be less than the first proportion, the second proportion to be less than the first proportion, and the second bad block proportion to be less than the first bad block proportion, the risk degree of the related events in the low-risk diagnosis indicator in the evaluation dimension of storage is relatively low. For example, the second number of days can be 30 days as shown in Figure 3 , the second proportion can be 200% as shown in Figure 3 , the second proportion can be 50% as shown in Figure 3 , and the second bad block proportion can be 5% as shown in Figure 3 .
[0107] The low-risk diagnosis indicators listed above under the storage evaluation dimension belong to indicators for representing that the data resource usage of the data storage resource exceeds the second data storage resource limit value, and / or indicators for representing that the data storage resource does not meet the second data storage resource usage rule. Among them, the storage remaining usage days less than or equal to the second number of days, the storage amount increase value compared to the previous period greater than the second proportion, the disk bad block proportion greater than the second bad block proportion, etc. are related to the used data storage resource or the remaining data storage resource, so these indicators belong to indicators for representing that the data resource usage of the data storage resource exceeds the second data storage resource limit value. The indicators such as the small file proportion greater than the second proportion involve the data storage resource usage rule, and belong to indicators for representing that the data storage resource does not meet the second data storage resource usage rule. The first data storage resource usage rule can specify that the small file proportion needs to be less than or equal to the second proportion.
[0108] Referring again to Figure 4 As shown in Figure 4 , the low-risk diagnosis indicators under the computing evaluation dimension include one or more of the following:
[0109] The queue blocking time length exceeds the second time length, the queue over-issuing time length exceeds the target over-issuing time length, the scheduling task failure rate is greater than the second failure rate, the scheduling task running time length is greater than the second running time length, the waiting through concurrent task proportion exceeds the target proportion, the queue usage computing time proportion exceeds the second computing time proportion, and the day-level task is completed across days; the second time length is greater than the first time length, the second failure rate is less than the first failure rate, the second running time length is less than the first running time length, and the second computing time proportion is less than the first computing time proportion.
[0110] It can be understood that the queue blocking time length, the queue over-issuing time length, the scheduling task failure rate, the scheduling task running time length, the waiting through concurrent task proportion, the queue usage computing time proportion, and the day-level task are all related events of the data computing resource in the use process.
[0111] It can also be understood that through the setting of the conditions that the second time length is greater than the first time length, the second failure rate is less than the first failure rate, the second running time length is less than the first running time length, and the second computing time proportion is less than the first computing time proportion, the risk degree of the related events in the low-risk diagnosis indicators under the computing evaluation dimension is low. For example, the second time length can be 10 minutes as shown in Figure 4 , the second failure rate can be 5% as shown in Figure 4 , and the second running time length can be 8 hours as shown in Figure 4 .
[0112] In addition, "queue overuse duration exceeds target overuse duration" is used to indicate whether the data output is timely. It can be known that the data resources used by some queues can exceed the configured data resource limit value. For example, the data resource limit value of the task queue is configured as 1000 cores of central processing unit and 10 TB of memory. The data resource usage of the queue exceeds the configured data resource limit value, and the overtime of the data resource usage of the queue exceeding the configured data resource limit value is the overuse duration. When the queue overuse duration exceeds the target overuse duration is greater than the target overuse duration, it indicates that the overuse duration of the task queue has exceeded the normal value, and it is determined that the data computing resource has a risk but the risk degree is low. Here, the target overuse duration is not limited and can be set according to the actual scene.
[0113] In actual application, the terminal device and / or the server first acquires the queue overuse duration, and then compares the queue overuse duration and the target overuse duration to determine whether the queue overuse duration exceeds the target overuse duration. If yes, the terminal device and / or the server determines that the data computing resource meets the low-risk diagnostic index in the use process.
[0114] In actual application, some tasks are executed concurrently, and when the number of concurrent tasks waiting for execution is too large, it can cause the data computing resource to be unable to be used for a long time, resulting in the situation of data computing resource vacancy. The "waiting-through-concurrent-task proportion" is the ratio of the number of concurrent tasks waiting for execution to the total amount of tasks. When the waiting-through-concurrent-task proportion exceeds the target proportion, it is determined that the data computing resource has a risk but the risk degree is low. Here, the target proportion is not limited and can be set according to the actual scene.
[0115] In actual application, the terminal device and / or the server first acquires the waiting-through-concurrent-task proportion based on the above-mentioned manner, and then compares the waiting-through-concurrent-task proportion and the target proportion to determine whether the waiting-through-concurrent-task proportion exceeds the target proportion. If yes, the terminal device and / or the server determines that the data computing resource meets the low-risk diagnostic index in the use process.
[0116] "Day-level task is completed across days" indicates that the task required to be completed on the same day is not completed on the same day but is completed across days, thereby determining that the timeliness of data output is poor and the data computing resource used by the task has a risk but the risk degree is low.
[0117] In actual application, the terminal device and / or the server first acquires the completion time of the day-level task based on the above-mentioned manner, and then judges whether the completion time of the day-level task is within the time range of the same day to determine whether the day-level task is completed across days. If yes, the terminal device and / or the server determines that the data computing resource meets the low-risk diagnostic index in the use process.
[0118] It can be understood that the low-risk diagnostic indicators listed above in the calculation evaluation dimension can all belong to indicators for indicating that the data resource usage of the data computing resource exceeds the second data computing resource limit value. Taking the "queue blocking time exceeding the second time length" as an example, when the queue blocking time exceeds the second time length, it is considered that the data resource usage of the data computing resource exceeds the normal value, indicating that "the data resource usage of the data computing resource exceeds the second data computing resource limit value". The rest of the indicators are similar, and will not be repeated here.
[0119] B2: Obtain the low-risk indicator trigger quantity in the storage evaluation dimension and the low-risk indicator trigger quantity in the calculation evaluation dimension; the low-risk indicator trigger quantity in the storage evaluation dimension is the number of low-risk diagnostic indicators in the storage evaluation dimension that are met by the data storage resource in the use process, and the low-risk indicator trigger quantity in the calculation evaluation dimension is the number of low-risk diagnostic indicators in the calculation evaluation dimension that are met by the data computing resource in the use process.
[0120] In a possible implementation, the embodiment of the present application provides a specific implementation of obtaining the low-risk indicator trigger quantity in the storage evaluation dimension and the low-risk indicator trigger quantity in the calculation evaluation dimension, including C1-C2:
[0121] C1: In combination with the use of the data storage resource, determine the indicator value corresponding to the low-risk diagnostic indicator in the storage evaluation dimension, and according to the indicator value corresponding to the low-risk diagnostic indicator in the storage evaluation dimension, determine the low-risk indicator trigger quantity in the storage evaluation dimension.
[0122] C2: In combination with the use of the data computing resource, determine the indicator value corresponding to the low-risk diagnostic indicator in the calculation evaluation dimension, and according to the indicator value corresponding to the low-risk diagnostic indicator in the calculation evaluation dimension, determine the low-risk indicator trigger quantity in the calculation evaluation dimension.
[0123] It can be known that the technical implementation of C1-C2 is similar to that of A1-A2, which will not be repeated here, and the detailed content can be referred to A1-A2.
[0124] B3: Based on the low-risk indicator trigger quantity in the storage evaluation dimension, evaluate the risk caused by improper use of the data storage resource, and based on the low-risk indicator trigger quantity in the calculation evaluation dimension, evaluate the risk caused by improper use of the data computing resource.
[0125] It can be known that the remaining technical implementation of B1-B3 is similar to that of S102-S104, which will not be repeated here, and the detailed content can be referred to S102-S104.
[0126] In order to make the risk assessment of data storage resources and data computing resources more fine-grained, the embodiments of the present application provide observation indexes, which can also be referred to as benefit indexes, for assisting in assessing the risk degree of data storage resources and data computing resources. Details are described below.
[0127] In a possible implementation, the method for assessing the risk caused by improper use of cluster data resources provided by the embodiments of the present application can further include the following steps:
[0128] D1: Obtain the risk observation indexes in the storage evaluation dimension and the risk observation indexes in the computing evaluation dimension.
[0129] As an optional example, in combination with Figure 3 , the risk observation indexes in the storage evaluation dimension include one or more of the following:
[0130] The small file sorting result, the remaining storage amount, the average storage increment in a preset time period, the remaining storage available days, the daily new storage amount, and the total number of small files.
[0131] It should be understood that the risk observation indexes in the storage evaluation dimension described above are all for data resource users. For example, the data resource user is a department, and the small file sorting result is the small file sorting result in the department. The rest of the parameters are similar, and will not be described here.
[0132] The small file can be understood as a file with a storage usage space less than 256 MB. The small file sorting result can be exemplified as the "small file Top100" shown in Figure 3 , which can be sorted according to the size of the storage usage space of the file. The total number of small files is the total number of files with a storage usage space less than 256 MB. The small file sorting result and the total number of small files can be used to observe the stability of small files to observe whether the number of small files is too large.
[0133] The remaining storage amount represents the total amount of the current data storage resources, which is the difference between the total storage amount of the data storage resources and the storage amount of the used data storage resources, and the unit can be GB. The remaining storage available days is the ratio of the remaining storage amount to the average storage increment in the last 7 days, and the unit is days. The indexes such as the remaining storage amount and the remaining storage available days are used to observe the use risk of the data storage resources.
[0134] The average storage increment in a preset time period is the ratio of the total storage increment in a preset time period to the preset time period. For example, the preset time period is the last 7 days, and the average storage increment in the last 7 days is the ratio of the total storage increment in the last 7 days to 7. The average storage increment in a preset time period is used to observe the use risk of the data storage resources.
[0135] The daily newly added storage amount is the daily data storage resource increment, which can be used to observe the use risk of the data storage resource.
[0136] As an optional example, in combination with Figure 4 The risk observation index under the calculation evaluation dimension includes one or more of the following:
[0137] The daily use calculation time, the daily use calculation time proportion, the daily use calculation time comparison, the scheduling task failure amount, the proportion of instances not completed on the same day, the T+1 incomplete rate, the SLA broken line amount, the task scheduling one-time success rate, the total of the task suspension time in the period, and the number of queue tasks in the period.
[0138] It should be understood that the risk observation indexes under the calculation evaluation dimension described above are all for data resource users. For example, the data resource user is a department, and the "SLA broken line amount" is the SLA broken line amount in the department. The rest of the parameters are similar and will not be described here.
[0139] The "daily use calculation time" is the data calculation resource used on the same day, which is represented by calculation time. The "daily use calculation time proportion" is the ratio of the daily use calculation time to the total calculation time, and the total calculation time can be the total calculation time in the past 7 days, which is not limited here. The "daily use calculation time comparison" is the ratio of the daily use calculation time to the yesterday use calculation time. The indexes such as the daily use calculation time, the daily use calculation time proportion, and the daily use calculation time comparison are used to observe the stability of the data calculation resource.
[0140] The "scheduling task failure amount" is the number of scheduling task failures, and the time scale can be the past 7 days, which is not limited here. The "scheduling task failure amount" is used to observe the stability of the data calculation resource. In addition, the risk observation index under the calculation evaluation dimension can also include the "retry still failed task proportion", which is the ratio of the retry still failed task to the retry task, and is used to observe the stability of the data calculation resource.
[0141] There is a task in the data warehouse that delays the statistics data by one day, which is represented as T+1 task. For example, today is October 16, and the data produced today is the data of October 15. If the data of October 15 is not produced on October 16, it is determined that the task is not completed. The "T+1 incomplete rate" represents the proportion of the T+1 task not completed in all tasks, which is used to observe the stability of the data calculation resource.
[0142] The "SLA breaking quantity" is the number of SLA tasks that break. The "task scheduling one-time success rate" is the ratio of task instances that are successfully scheduled once to the total number of task scheduling. The "total suspension time of tasks in a period" is the total suspension time of all task instances running in a period, which can be a day. The "number of queued tasks in a period" is the number of tasks running in the queue in a period, which can be an hour. The "SLA breaking quantity, task scheduling one-time success rate, total suspension time of tasks in a period, and number of queued tasks in a period" are used to observe the stability of data computing resources.
[0143] D2: In combination with the risk observation indexes in the storage evaluation dimension and the risk observation indexes in the computing evaluation dimension, the data storage resources and the data computing resources are evaluated.
[0144] The risk observation indexes in the storage evaluation dimension are used to assist in evaluating the risk of data storage resources, and the risk observation indexes in the computing evaluation dimension are used to assist in evaluating the risk of data computing resources, so that the risk evaluation of cluster data resources is more accurate, the range of evaluation indexes is wider, and the evaluation dimension is richer. For example, the smaller the "remaining storage", the greater the risk of data storage resources, the greater the "SLA breaking quantity", the greater the risk of data computing resources, and the rest are similar, which will not be repeated here.
[0145] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0146] Based on the above-mentioned method embodiment of the cluster data resource usage improper risk evaluation method, the embodiment of the present application also provides a cluster data resource usage improper risk evaluation device. The cluster data resource usage improper risk evaluation device will be described below in conjunction with the drawings. Since the principle of solving the problem in the device of the embodiment of the present application is similar to the above-mentioned cluster data resource usage improper risk evaluation method of the embodiment of the present application, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be repeated.
[0147] Referring to Figure 5 As shown in the figure, the figure is a structure schematic diagram of a cluster data resource usage improper risk evaluation device provided by an embodiment of the present application. As Figure 5 As shown in the figure, the cluster data resource usage improper risk evaluation device 500 comprises:
[0148] The first obtaining unit 501 is configured to obtain an evaluation dimension for evaluating a risk of a cluster data resource; the cluster data resource comprises a data storage resource and a data computing resource, and the evaluation dimension comprises a storage evaluation dimension and a computing evaluation dimension;
[0149] The second obtaining unit 502 is configured to obtain at least one high-risk diagnosis index under the storage evaluation dimension and at least one high-risk diagnosis index under the computing evaluation dimension; the high-risk diagnosis index under the storage evaluation dimension comprises an index for representing that a data resource usage of the data storage resource exceeds a first data storage resource limit value and / or an index for representing that the data storage resource does not meet a first data storage resource usage rule, and the high-risk diagnosis index under the computing evaluation dimension comprises an index for representing that a data resource usage of the data computing resource exceeds a first data computing resource limit value;
[0150] The first statistical unit 503 is configured to count a high-risk index trigger quantity under the storage evaluation dimension and a high-risk index trigger quantity under the computing evaluation dimension; the high-risk index trigger quantity under the storage evaluation dimension is a number of the high-risk diagnosis index under the storage evaluation dimension that is met by the data storage resource in a usage process, and the high-risk index trigger quantity under the computing evaluation dimension is a number of the high-risk diagnosis index under the computing evaluation dimension that is met by the data computing resource in a usage process.
[0151] The first evaluation unit 504 is configured to evaluate a risk caused by improper use of the data storage resource based on the high-risk index trigger quantity under the storage evaluation dimension, and evaluate a risk caused by improper use of the data computing resource based on the high-risk index trigger quantity under the computing evaluation dimension.
[0152] In a possible implementation, the first statistical unit 503 comprises:
[0153] The first determination sub-unit is configured to determine an index value corresponding to the high-risk diagnosis index under the storage evaluation dimension in combination with a usage situation of the data storage resource, and determine the high-risk index trigger quantity under the storage evaluation dimension according to the index value corresponding to the high-risk diagnosis index under the storage evaluation dimension.
[0154] The second determination sub-unit is configured to determine an index value corresponding to the high-risk diagnosis index under the computing evaluation dimension in combination with a usage situation of the data computing resource, and determine the high-risk index trigger quantity under the computing evaluation dimension according to the index value corresponding to the high-risk diagnosis index under the computing evaluation dimension.
[0155] In a possible implementation, the high-risk diagnosis index under the storage evaluation dimension comprises one or more of the following:
[0156] the remaining use days of the storage are less than or equal to the first day number, the storage amount increment ratio is greater than the first ratio, the small file proportion is greater than the first proportion, the disk bad block ratio is greater than the first bad block ratio, the trigger write-prohibited time is greater than the target time, the storage usage rate exceeds the target usage rate, or the set life cycle is less than the recommended value;
[0157] The high-risk diagnosis index in the computing evaluation dimension includes one or more of the following:
[0158] the queue blocking time exceeds the first time length, the scheduling task failure rate is greater than the first failure rate, the scheduling task running time is greater than the first running time, the queue usage proportion during calculation exceeds the first calculation time proportion, or the SLA task breaks the line.
[0159] In a possible implementation, the apparatus further includes:
[0160] The third obtaining unit is configured to obtain at least one low-risk diagnosis index in the storage evaluation dimension and at least one low-risk diagnosis index in the computing evaluation dimension; the low-risk diagnosis index in the storage evaluation dimension includes an index for representing that the data resource usage of the data storage resource exceeds a second data storage resource limit value and / or an index for representing that the data storage resource does not meet a second data storage resource usage rule, and the low-risk diagnosis index in the computing evaluation dimension includes an index for representing that the data resource usage of the data computing resource exceeds a second data computing resource limit value; the second data storage resource limit value is less than the first data storage resource limit value, and the second data computing resource limit value is less than the first data computing resource limit value.
[0161] The second statistical unit is configured to obtain a low-risk index trigger amount in the storage evaluation dimension and a low-risk index trigger amount in the computing evaluation dimension; the low-risk index trigger amount in the storage evaluation dimension is a number of low-risk diagnosis indexes in the storage evaluation dimension that are met by the data storage resource in the use process, and the low-risk index trigger amount in the computing evaluation dimension is a number of low-risk diagnosis indexes in the computing evaluation dimension that are met by the data computing resource in the use process.
[0162] The second evaluation unit is configured to evaluate a risk caused by improper use of the data storage resource based on the low-risk index trigger amount in the storage evaluation dimension, and evaluate a risk caused by improper use of the data computing resource based on the low-risk index trigger amount in the computing evaluation dimension.
[0163] In a possible implementation, the second statistical unit includes:
[0164] The third determining subunit is configured to determine an index value corresponding to a low-risk diagnosis index in the storage evaluation dimension in combination with a usage of the data storage resource, and determine a low-risk index trigger quantity in the storage evaluation dimension according to the index value corresponding to the low-risk diagnosis index in the storage evaluation dimension.
[0165] The fourth determining subunit is configured to determine an index value corresponding to a low-risk diagnosis index in the calculation evaluation dimension in combination with a usage of the data calculation resource, and determine a low-risk index trigger quantity in the calculation evaluation dimension according to the index value corresponding to the low-risk diagnosis index in the calculation evaluation dimension.
[0166] In a possible implementation, the low-risk diagnosis index in the storage evaluation dimension includes one or more of the following:
[0167] The number of remaining storage usage days is less than or equal to a second number of days, the storage value increment is greater than a second ratio, the proportion of small files is greater than a second proportion, and the proportion of disk bad blocks is greater than a second bad block proportion; the second number of days is greater than a first number of days, the second ratio is less than a first ratio, the second proportion is less than a first proportion, and the second bad block proportion is less than a first bad block proportion.
[0168] The low-risk diagnosis index in the calculation evaluation dimension includes one or more of the following:
[0169] The queue blocking time exceeds a second time, the queue over-issuing time exceeds a target over-issuing time, the scheduling task failure rate is greater than a second failure rate, the scheduling task running time is greater than a second running time, the proportion of waiting through concurrent tasks exceeds a target proportion, the queue usage calculation time proportion exceeds a second calculation time proportion, and the day-level task is completed across days; the second time is greater than a first time, the second failure rate is less than a first failure rate, the second running time is less than a first running time, and the second calculation time proportion is less than a first calculation time proportion.
[0170] In a possible implementation, the apparatus further includes:
[0171] The fourth obtaining unit is configured to obtain a risk observation index in the storage evaluation dimension and a risk observation index in the calculation evaluation dimension.
[0172] The third evaluation unit is configured to perform risk evaluation on the data storage resource and the data calculation resource in combination with the risk observation index in the storage evaluation dimension and the risk observation index in the calculation evaluation dimension.
[0173] The risk observation index in the storage evaluation dimension includes one or more of the following:
[0174] Small file sorting result, remaining storage capacity, average storage increment in a preset time period, remaining storage available days, daily new storage, total number of small files;
[0175] The risk observation index under the calculation evaluation dimension includes one or more of the following:
[0176] Daily use calculation time, daily use calculation time proportion, daily use calculation time comparison, scheduling task failure quantity, proportion of instances not completed on the day, T+1 incomplete rate, SLA broken line quantity, task scheduling one-time success rate, total of task suspension time in the period, and queue task quantity in the period.
[0177] On the basis of the implementation manners provided by the above aspects, the application can be further combined to provide more implementation manners.
[0178] It should be noted that the specific implementation of each unit in the present embodiment can be referred to the related description in the above method embodiments. The division of units in the present embodiment is illustrative, and is only a logical function division. In actual implementation, another division manner can be used. Each functional unit in the present embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. For example, in the above embodiments, the processing unit and the sending unit can be the same unit, or can be different units. The integrated unit can be realized in the form of hardware, or in the form of a software functional unit.
[0179] Based on the risk evaluation method of cluster data resource misuse provided in the above method embodiments, the application further provides an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the risk evaluation method of cluster data resource misuse according to any of the above embodiments.
[0180] Reference is made below to Figure 6 which shows a structural schematic diagram of an electronic device 600 suitable for implementing the embodiments of the present application. The terminal device in the embodiments of the present application can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable android devices), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like.Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0181] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0182] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0183] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of the embodiments of this application.
[0184] The electronic device provided in this application embodiment and the risk assessment method for improper use of cluster data resources provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0185] Based on the cluster data resource usage improper risk assessment method provided in the above method embodiments, an embodiment of the present application provides a computer readable medium having a computer program stored thereon, wherein the program is executed by a processor to implement the cluster data resource usage improper risk assessment method according to any of the above embodiments.
[0186] It should be noted that the computer readable medium of the present application described above can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, a RF (radio frequency) or the like, or any suitable combination of the above.
[0187] In some embodiments, the client, server can communicate using any currently known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication (e.g., communication network) of any form or medium, such as a local area network ("LAN"), a wide area network ("WAN"), an internetwork (e.g., the Internet), and an end-to-end network (e.g., an ad hoc end-to-end network), as well as any currently known or future developed network.
[0188] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0189] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method for evaluating the risk caused by improper use of cluster data resources.
[0190] Computer program code for carrying out operations of the present application can be written in any of one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0191] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0192] The units involved in the embodiments of the present application can be implemented in a software manner, or can be implemented in a hardware manner. In some cases, the name of the unit / module does not constitute a limitation on the unit itself, for example, the voice data acquisition module can also be described as a "data acquisition module".
[0193] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program- specific Integrated Circuits (ASICs), Program- specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0194] In the context of this application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0195] It should be noted that the various embodiments described in the specification employ a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between embodiments can be referred to each other.
[0196] It should be understood that in this application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including single or multiple combinations of items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0197] It is also to be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless otherwise indicated. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or "contains" are used in either the detailed description and the claims, such terms are intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or other elements.
[0198] The embodiments disclosed herein can each be implemented as a method, apparatus, or article of manufacture using programming instructions. The embodiments disclosed herein can be implemented using software, firmware, hardware, or a combination thereof. The various elements of the disclosed embodiments, as well as the embodiments themselves, can be constructed from any combination of hardware, software, and / or firmware. The software implementation can be implemented by one or more software modules using object-oriented design methodology, among other techniques. The software modules can be stored on any computer-readable medium, including RAM, ROM, EEPROM, flash memory, or a hard disk, to name a few. The software modules can include one or more routines.
[0199] The above description of disclosed embodiments provides information sufficient to understand how to make and use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of assessing risk of misusing a clustered data resource, characterized by, The method comprises: acquiring an evaluation dimension for evaluating the risk of a cluster data resource; the cluster data resource comprises a data storage resource and a data computing resource, and the evaluation dimension comprises a storage evaluation dimension and a computing evaluation dimension; acquiring at least one high-risk diagnostic index under the storage evaluation dimension and at least one high-risk diagnostic index under the computing evaluation dimension; the high-risk diagnostic index under the storage evaluation dimension comprises an index for representing that the data resource usage of the data storage resource exceeds a first data storage resource limit value and / or an index for representing that the data storage resource does not meet a first data storage resource usage rule, and the high-risk diagnostic index under the computing evaluation dimension comprises an index for representing that the data resource usage of the data computing resource exceeds a first data computing resource limit value; counting the high-risk index trigger quantity under the storage evaluation dimension and the high-risk index trigger quantity under the computing evaluation dimension; the high-risk index trigger quantity under the storage evaluation dimension is the number of high-risk diagnostic indexes under the storage evaluation dimension that are met by the data storage resource during use, and the high-risk index trigger quantity under the computing evaluation dimension is the number of high-risk diagnostic indexes under the computing evaluation dimension that are met by the data computing resource during use; based on the high-risk index trigger quantity under the storage evaluation dimension, evaluating the risk caused by improper use of the data storage resource, and based on the high-risk index trigger quantity under the computing evaluation dimension, evaluating the risk caused by improper use of the data computing resource.
2. The method of claim 1, wherein, The counting of the high-risk index trigger quantity under the storage evaluation dimension and the high-risk index trigger quantity under the computing evaluation dimension comprises: in combination with the use of the data storage resource, determining the index value corresponding to the high-risk diagnostic index under the storage evaluation dimension, and according to the index value corresponding to the high-risk diagnostic index under the storage evaluation dimension, determining the high-risk index trigger quantity under the storage evaluation dimension; in combination with the use of the data computing resource, determining the index value corresponding to the high-risk diagnostic index under the computing evaluation dimension, and according to the index value corresponding to the high-risk diagnostic index under the computing evaluation dimension, determining the high-risk index trigger quantity under the computing evaluation dimension.
3. The method of claim 2, wherein, The high-risk diagnostic index under the storage evaluation dimension comprises one or more of the following: the storage remaining usage days are less than or equal to a first day number, the storage amount increment year-on-year is greater than a first proportion, the small file proportion is greater than a first proportion, the disk bad block proportion is greater than a first bad block proportion, the trigger write-prohibition time is greater than a target time, the storage usage rate exceeds a target usage rate, and the set life cycle is less than a recommended value; The high-risk diagnostic index under the computing evaluation dimension comprises one or more of the following: the queue blocking time exceeds a first time length, the scheduling task failure rate is greater than a first failure rate, the scheduling task running time is greater than a first running time, the queue usage computing time proportion exceeds a first computing time proportion, and the SLA task breaks the line.
4. The method of claim 1, wherein, The method further comprises: obtaining at least one low-risk diagnosis indicator in the storage evaluation dimension and at least one low-risk diagnosis indicator in the computing evaluation dimension; the low-risk diagnosis indicator in the storage evaluation dimension includes an indicator for representing that the data resource usage of the data storage resource exceeds a second data storage resource limit value and / or an indicator for representing that the data storage resource does not meet a second data storage resource usage rule, and the low-risk diagnosis indicator in the computing evaluation dimension includes an indicator for representing that the data resource usage of the data computing resource exceeds a second data computing resource limit value; the second data storage resource limit value is less than the first data storage resource limit value, and the second data computing resource limit value is less than the first data computing resource limit value; obtaining a low-risk indicator trigger quantity in the storage evaluation dimension and a low-risk indicator trigger quantity in the computing evaluation dimension; the low-risk indicator trigger quantity in the storage evaluation dimension is the number of low-risk diagnosis indicators in the storage evaluation dimension met by the data storage resource during use, and the low-risk indicator trigger quantity in the computing evaluation dimension is the number of low-risk diagnosis indicators in the computing evaluation dimension met by the data computing resource during use; based on the low-risk indicator trigger quantity in the storage evaluation dimension, evaluating the risk caused by improper use of the data storage resource, and based on the low-risk indicator trigger quantity in the computing evaluation dimension, evaluating the risk caused by improper use of the data computing resource.
5. The method of claim 4, wherein, The obtaining of the low-risk indicator trigger quantity in the storage evaluation dimension and the low-risk indicator trigger quantity in the computing evaluation dimension comprises: in combination with the use of the data storage resource, determining the indicator value corresponding to the low-risk diagnosis indicator in the storage evaluation dimension, and determining the low-risk indicator trigger quantity in the storage evaluation dimension according to the indicator value corresponding to the low-risk diagnosis indicator in the storage evaluation dimension; in combination with the use of the data computing resource, determining the indicator value corresponding to the low-risk diagnosis indicator in the computing evaluation dimension, and determining the low-risk indicator trigger quantity in the computing evaluation dimension according to the indicator value corresponding to the low-risk diagnosis indicator in the computing evaluation dimension.
6. The method of claim 5, wherein, The low-risk diagnosis indicator in the storage evaluation dimension includes one or more of the following: the storage remaining usage days are less than or equal to a second number of days, the storage amount increment year-on-year is greater than a second proportion, the proportion of small files is greater than a second proportion, and the proportion of disk bad blocks is greater than a second bad block proportion; the second number of days is greater than a first number of days, the second proportion is less than a first proportion, the second proportion is less than a first proportion, and the second bad block proportion is less than a first bad block proportion; The low-risk diagnosis indicator in the computing evaluation dimension includes one or more of the following: The queue blocking duration exceeds a second duration, the queue over-issuing duration exceeds a target over-issuing duration, the scheduling task failure rate is greater than a second failure rate, the scheduling task running duration is greater than a second running duration, the waiting through concurrent task proportion exceeds a target proportion, the queue usage proportion exceeds a second calculation time proportion, and the day-level task is completed across days.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: obtaining risk observation indicators under the storage evaluation dimension and risk observation indicators under the calculation evaluation dimension; combining the risk observation indicators under the storage evaluation dimension and the risk observation indicators under the calculation evaluation dimension, performing risk evaluation on the data storage resource and the data calculation resource; The risk observation indicators under the storage evaluation dimension include one or more of the following: small file sorting results, remaining storage capacity, average storage increment in a preset time period, remaining storage available days, daily new storage capacity, and total number of small files; The risk observation indicators under the calculation evaluation dimension include one or more of the following: daily usage calculation time, daily usage calculation time proportion, daily usage calculation time comparison, scheduling task failure quantity, proportion of instances not completed on the same day, T+1 incomplete rate, SLA broken line quantity, task scheduling one-time success rate, total of task suspension duration within a period, and queue task quantity within a period.
8. An assessment device for risks arising from improper use of cluster data resources, characterized in that, The device includes: an acquisition unit configured to acquire evaluation dimensions for evaluating cluster data resource risks; the cluster data resource includes a data storage resource and a data calculation resource, and the evaluation dimensions include a storage evaluation dimension and a calculation evaluation dimension; a first determination unit configured to acquire at least one high-risk diagnosis indicator under the storage evaluation dimension and at least one high-risk diagnosis indicator under the calculation evaluation dimension; the high-risk diagnosis indicator under the storage evaluation dimension includes an indicator for representing that a data resource usage amount of the data storage resource exceeds a first data storage resource limit value and / or an indicator for representing that the data storage resource does not meet a first data storage resource usage rule, and the high-risk diagnosis indicator under the calculation evaluation dimension includes an indicator for representing that a data resource usage amount of the data calculation resource exceeds a first data calculation resource limit value; a first statistical unit configured to statistically acquire a high-risk indicator trigger amount under the storage evaluation dimension and a high-risk indicator trigger amount under the calculation evaluation dimension; the high-risk indicator trigger amount under the storage evaluation dimension is a number of high-risk diagnosis indicators under the storage evaluation dimension that are met by the data storage resource in a usage process, and the high-risk indicator trigger amount under the calculation evaluation dimension is a number of high-risk diagnosis indicators under the calculation evaluation dimension that are met by the data calculation resource in a usage process; and a second statistical unit configured to statistically acquire a high-risk indicator trigger amount under the storage evaluation dimension and a high-risk indicator trigger amount under the calculation evaluation dimension. The first evaluation unit is configured to evaluate the risk caused by improper use of the data storage resource based on the high-risk indicator trigger quantity in the storage evaluation dimension, and evaluate the risk caused by improper use of the data calculation resource based on the high-risk indicator trigger quantity in the calculation evaluation dimension.
9. An electronic device, comprising: The method comprises the following steps: one or more processors; a storage device having stored thereon one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the cluster data resource misuse risk evaluation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, a computer program is stored thereon, and the computer program is executed by a processor to implement the cluster data resource misuse risk evaluation method according to any one of claims 1-7.
Citation Information
Patent Citations
Software reliability acceptance risk assessment method based on multi-dimensional layering criterion
CN112486790A
Risk assessment method and device
CN113407950A