Storage system health optimization method and system, electronic device and medium

CN116521415BActive Publication Date: 2026-08-11INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]通常情况下,存储系统的运行情况将会影响业务的运行状态和用户的使用感受;但随着存储系统使用时间增长自然而然的会出现软件运行卡顿、访问效率低等各种运行效率低的问题;此时如果不能及时发现,将会直接影响业务的处理速度,严重的将会导致数据的丢失等问题

Benefits of technology

[0054]本申请提供了一种存储系统健康度优化方法,包括监测存储系统中多个硬件部件,并获取所述多个硬件部件的当前硬件效率采集值;监测存储系统上运行的软件,并获取所述软件的当前软件服务状态;根据所述当前硬件效率采集值确定硬件效率值,并根据所述当前软件服务状态确定软件效率值;基于所述硬件效率值及软件效率值,确定异常点并生成健康度报告以实现对所述存储系统健康度的优化。通过定时对关键硬件和软件的监测并输出对应的硬件效率值和软件效率值,来反映存储系统的健康度并及时同步到客户,提高健康度通知的时效性,避免因异常导致客户业务中断等现象,提高系统运行效率。进一步实现将存储系统健康度稳定在较优状态,此外通过自动进行异常分析并解决,存储系统可自动进行对应的优化,减少人力介入,进而提高存储系统的健壮性,最终提升产品的竞争力。此外,通过对存储系统健康度和优化操作出具健康报告,使得客户更容易掌握存储系统的健康度,简明扼要;更改了传统的获取存储系统健康度的方式,提高产品竞争力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116521415B_ABST
    Figure CN116521415B_ABST
Patent Text Reader

Abstract

This application provides a storage system health optimization method, system, electronic device, and medium, including: monitoring multiple hardware components in the storage system and acquiring current hardware efficiency values ​​of these components; monitoring software running on the storage system and acquiring its current software service status; determining hardware efficiency values ​​based on the current hardware efficiency values ​​and software efficiency values ​​based on the current software service status; and identifying anomalies and generating a health report based on the hardware and software efficiency values ​​for health optimization. By monitoring key hardware and software and outputting corresponding hardware and software efficiency values, the health of the storage system is reflected and promptly synchronized with customers, avoiding issues such as business interruptions and improving system operating efficiency. Furthermore, it stabilizes the storage system health at an optimal level, automatically analyzes and resolves anomalies, reduces human intervention, improves system robustness, and ultimately enhances system competitiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, system, electronic device, and storage medium for optimizing the health of a storage system. Background Technology

[0002] Typically, the performance of a storage system affects the operational status of business operations and the user experience. However, as storage systems are used for longer periods, various problems such as software lag and low access efficiency will naturally arise. If these issues are not detected in time, they will directly impact business processing speed and, in severe cases, lead to data loss.

[0003] Based on the above issues, the industry typically employs two methods to address these problems: one is for the storage system to generate an alarm to alert the user, requiring the user to contact customer service to locate the issue; the other is for the system to malfunction, such as a function becoming unusable, necessitating customer service intervention for problem analysis. However, both of these commonly used solutions require manual intervention after a problem occurs. During this process, the customer's business is likely to be interrupted and can only resume operation after the problem is resolved. These common solutions primarily rely on alerts and cannot truly prevent or monitor problems, thus failing to fundamentally and efficiently avoid issues and resulting in a waste of human and material resources.

[0004] Therefore, there is an urgent need for a method that can promptly repair anomalies within the storage system to improve the health of the storage system and solve the above-mentioned technical problems. Summary of the Invention

[0005] Therefore, it is necessary to provide a storage system health optimization method, system, electronic device, and storage medium to address the above-mentioned technical problems, so as to improve the health of the storage system and enable the storage system to always maintain a high level of health.

[0006] In a first aspect, this application provides a method for optimizing the health of a storage system, the method comprising:

[0007] Monitor multiple hardware components in the storage system and obtain the current hardware efficiency values ​​of the multiple hardware components;

[0008] Monitor the software running on the storage system and obtain the current software service status of the software;

[0009] The hardware efficiency value is determined based on the current hardware efficiency acquisition value, and the software efficiency value is determined based on the current software service status.

[0010] Based on the hardware efficiency values ​​and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system.

[0011] In some embodiments, determining the hardware efficiency value based on the current hardware efficiency acquisition value includes:

[0012] The current hardware efficiency values ​​of multiple hardware components are compared with the corresponding theoretical hardware efficiency values ​​of multiple hardware components set in advance to determine the hardware efficiency value.

[0013] The hardware components include a CPU, memory, hard drive, and chassis. The current hardware efficiency acquisition values ​​include the corresponding current CPU efficiency acquisition values, current memory efficiency acquisition values, current hard drive efficiency acquisition values, and current chassis temperature acquisition values. The theoretical hardware efficiency values ​​include the corresponding theoretical CPU efficiency values, theoretical memory efficiency values, theoretical hard drive efficiency values, and theoretical chassis temperature values.

[0014] In some embodiments, determining the hardware efficiency value by comparing the current hardware efficiency acquisition values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple hardware components preset in advance includes:

[0015] If the current hardware efficiency values ​​of the multiple hardware components are all less than the corresponding theoretical hardware efficiency values, then the hardware efficiency value of the storage system is determined to be the first hardware efficiency value.

[0016] If the current hardware efficiency values ​​of the multiple hardware components are all greater than or equal to the corresponding theoretical hardware efficiency values, then the hardware efficiency value is determined to be the second hardware efficiency value.

[0017] The first hardware efficiency value is greater than the second hardware efficiency value.

[0018] In some embodiments, the step of comparing the current hardware efficiency acquisition values ​​of a plurality of hardware components with the corresponding theoretical hardware efficiency values ​​of a plurality of hardware components, and determining the hardware efficiency value, further includes:

[0019] If the current chassis temperature acquisition value is greater than the theoretical chassis temperature value, and the difference between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than the first preset threshold, then the hardware efficiency value is determined to be the third hardware efficiency value.

[0020] If the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold, and the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is greater than or equal to the second preset threshold, then the hardware efficiency value is determined to be the fourth hardware efficiency value.

[0021] If the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold, and the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than the second preset threshold, then the hardware efficiency value is determined to be the fifth hardware efficiency value.

[0022] Among them, the third hardware efficiency value is greater than the second hardware efficiency value, the fourth hardware efficiency value is greater than the third hardware efficiency value, the fifth hardware efficiency value is greater than the fourth hardware efficiency value, and the first hardware efficiency value is greater than the fifth hardware efficiency value.

[0023] In some embodiments, the current software service status includes the current service running status, the current log status, and the current thread running status of the software. Determining the software efficiency value based on the current software service status includes:

[0024] If the current service running status is normal, the current log status is normal, and there are idle threads, then the software efficiency value is determined to be the first software efficiency value.

[0025] If the current service running status is normal, the current log status is normal, and there are no idle threads, then the software efficiency value is determined to be the second software efficiency value.

[0026] If the current service running status is normal but the current log status is abnormal, then the software efficiency value is determined to be the third software efficiency value.

[0027] If the detected service operation status is abnormal, the software efficiency value is determined to be the fourth software efficiency value.

[0028] Among them, the first software efficiency value is greater than the second software efficiency value, the second software efficiency value is greater than the third software efficiency value, and the third software efficiency value is greater than the fourth software efficiency value.

[0029] In some embodiments, the anomalies include hardware anomalies and software anomalies, and determining the anomalies based on the hardware efficiency values ​​and software efficiency values ​​includes:

[0030] The hardware efficiency value is obtained and the hardware anomaly point is determined based on a pre-set first mapping relationship table, wherein the first mapping relationship table includes a mapping relationship between the hardware efficiency value and the hardware anomaly point.

[0031] The software efficiency value is obtained and software anomalies are determined based on a pre-set second mapping table, wherein the second mapping table includes a mapping relationship between the software efficiency value and the software anomaly.

[0032] In some embodiments, the step of determining anomalies based on the hardware efficiency values ​​and software efficiency values ​​and periodically generating health reports to optimize the health of the storage system further includes:

[0033] The storage system periodically acquires the hardware efficiency value, software efficiency value, and anomalies to generate a health report, wherein the health report includes at least one of the hardware efficiency value, software efficiency value, and anomalies.

[0034] The storage system triggers corresponding repair operations for the anomalies in the health report in order to optimize the health of the storage system.

[0035] If the hardware efficiency value and / or the software efficiency value in the health report are lower than the warning value, the storage system also generates a health warning to remind the user.

[0036] Secondly, this application provides a health optimization system, characterized in that the system includes:

[0037] The data acquisition module is used to monitor multiple hardware components in the storage system and obtain the current hardware efficiency acquisition values ​​of the multiple hardware components.

[0038] The data acquisition module is also used to monitor the software running on the storage system and obtain the current software service status of the software.

[0039] An efficiency calculation module is used to determine a hardware efficiency value based on the current hardware efficiency acquisition value, and to determine a software efficiency value based on the current software service status.

[0040] The health optimization module is used to identify anomalies and generate a health report based on the hardware efficiency value and software efficiency value in order to optimize the health of the storage system.

[0041] Thirdly, this application provides an electronic device, the electronic device comprising:

[0042] One or more processors;

[0043] and a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the following operations:

[0044] Monitor multiple hardware components in the storage system and obtain the current hardware efficiency values ​​of the multiple hardware components;

[0045] Monitor the software running on the storage system and obtain the current software service status of the software;

[0046] The hardware efficiency value is determined based on the current hardware efficiency acquisition value, and the software efficiency value is determined based on the current software service status.

[0047] Based on the hardware efficiency values ​​and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system.

[0048] Fourthly, this application also provides a computer-readable storage medium storing a computer program that causes a computer to perform the following operations:

[0049] Monitor multiple hardware components in the storage system and obtain the current hardware efficiency values ​​of the multiple hardware components;

[0050] Monitor the software running on the storage system and obtain the current software service status of the software;

[0051] The hardware efficiency value is determined based on the current hardware efficiency acquisition value, and the software efficiency value is determined based on the current software service status.

[0052] Based on the hardware efficiency values ​​and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system.

[0053] The beneficial effects achieved by this application are as follows:

[0054] This application provides a method for optimizing the health of a storage system, including monitoring multiple hardware components in the storage system and obtaining current hardware efficiency values ​​for the multiple hardware components; monitoring software running on the storage system and obtaining the current software service status of the software; determining hardware efficiency values ​​based on the current hardware efficiency values ​​and software efficiency values ​​based on the current software service status; identifying anomalies and generating a health report based on the hardware and software efficiency values ​​to optimize the health of the storage system. By periodically monitoring key hardware and software and outputting corresponding hardware and software efficiency values, the health of the storage system is reflected and promptly synchronized to customers, improving the timeliness of health notifications, avoiding business interruptions due to anomalies, and improving system operating efficiency. Furthermore, it stabilizes the health of the storage system in an optimal state. In addition, by automatically analyzing and resolving anomalies, the storage system can automatically perform corresponding optimizations, reducing human intervention, thereby improving the robustness of the storage system and ultimately enhancing product competitiveness. Furthermore, by generating health reports on storage system health and optimized operations, customers can more easily grasp the health status of their storage systems in a concise and clear manner; this changes the traditional way of obtaining storage system health status and enhances product competitiveness. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:

[0056] Figure 1 This is an overall flowchart of the storage system health optimization provided in the embodiments of this application;

[0057] Figure 2 This is a flowchart illustrating the hardware efficiency value determination process provided in an embodiment of this application.

[0058] Figure 3 This is a flowchart of the software efficiency value determination provided in the embodiments of this application;

[0059] Figure 4 This is a schematic diagram of anomaly analysis and optimization provided in the embodiments of this application;

[0060] Figure 5 This is a schematic diagram of the storage system health optimization method provided in the embodiments of this application;

[0061] Figure 6 This is a health optimization system architecture diagram provided in the embodiments of this application;

[0062] Figure 7This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] It should be understood that, in the description of this application, unless the context explicitly requires it, the words "comprising," "including," and similar terms throughout the specification and claims should be interpreted as encompassing rather than being exclusive or exhaustive; that is, meaning "including but not limited to."

[0065] It should also be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0066] It should be noted that the terms "S1," "S2," etc., are used only for descriptive purposes and do not specifically refer to the order or sequence, nor are they intended to limit this application. They are merely for the convenience of describing the method of this application and should not be construed as indicating the sequential order of the steps. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0067] Example 1

[0068] This application provides a method for optimizing the health of a storage system, specifically, as shown in the embodiments below. Figure 1 As shown, the process of applying this method to achieve automatic optimization of the storage system includes:

[0069] S1. Obtain the hardware configuration and software information of the storage system.

[0070] Specifically, the operating status of a storage system is typically affected by both software information and hardware configuration. Hardware configuration includes key components such as CPU, memory, hard drive, and chassis temperature. Therefore, in this embodiment, the system monitors these key hardware components in real time and periodically collects current hardware efficiency values ​​to reflect the hardware configuration. These hardware efficiency values ​​include current CPU efficiency (i.e., current CPU processing power), current memory efficiency (i.e., current memory usage), current hard drive efficiency (i.e., current hard drive health and load), and current chassis temperature (i.e., current chassis temperature reading). The system reflects the software service status of the storage system by monitoring software information such as the current service running status, current log status, and current thread running status.

[0071] S2. Based on the obtained hardware configuration, further determine the hardware efficiency value of the storage system.

[0072] like Figure 2 The flowchart shown illustrates how, after periodically acquiring current hardware efficiency values ​​for multiple hardware components of the storage system, the system determines its current hardware efficiency value by comparing these values ​​with pre-defined theoretical hardware efficiency values. The theoretical hardware efficiency values ​​are pre-defined optimal values ​​derived from customer business needs, including theoretical CPU efficiency, theoretical memory efficiency, theoretical hard disk efficiency, and theoretical chassis temperature. Specifically, the theoretical CPU efficiency value represents the CPU processing power under ideal conditions, the theoretical memory efficiency value represents memory usage under ideal conditions, the theoretical hard disk efficiency value represents hard disk health and load under ideal conditions, and the theoretical chassis temperature efficiency value represents temperature detection under ideal conditions.

[0073] Specifically, determining the hardware efficiency value by comparing the currently acquired hardware efficiency sampling value with the theoretical hardware efficiency value includes: if any acquired current hardware efficiency sampling value is less than the corresponding theoretical hardware efficiency value, then the hardware efficiency value is determined to be the first hardware efficiency value, in which case the utilization rate of each hardware component is low, and therefore the corresponding efficiency is high; if any acquired current hardware efficiency sampling value is greater than or equal to the corresponding theoretical hardware efficiency value, then the hardware efficiency value is determined to be the second hardware efficiency value; if the acquired current chassis temperature sampling value is greater than the theoretical chassis temperature value, and the differences between the current CPU efficiency sampling value, the current hard disk efficiency sampling value, and the current memory efficiency sampling value and their corresponding theoretical CPU efficiency values, hard disk efficiency values, and memory efficiency values ​​are all less than a first preset threshold, then the hardware efficiency value is determined to be the third hardware efficiency value; if the acquired current chassis temperature sampling value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold .... If the sum of the differences between the theoretical values ​​of CPU efficiency and memory efficiency is greater than or equal to a second preset threshold, then the hardware efficiency value is determined to be the fourth hardware efficiency value. If the current chassis temperature is less than the theoretical chassis temperature and the difference is less than a first preset threshold, and the sum of the differences between the current CPU efficiency, current hard disk efficiency, and current memory efficiency values ​​and their corresponding theoretical values ​​is less than a second preset threshold, then the hardware efficiency value is determined to be the fifth hardware efficiency value. The first hardware efficiency value is any percentage greater than 90%, the second hardware efficiency value is any percentage less than 90%, the third hardware efficiency value is any percentage between 30% and 50%, the fourth hardware efficiency value is any percentage between 50% and 70%, and the fifth hardware efficiency value is any percentage between 70% and 90%. Furthermore, the first and second preset thresholds are set by staff according to the actual scenario; preferably, the first preset threshold can be set to 5%, and the second preset threshold can be set to 50%.

[0074] S3. Based on the obtained software information, further determine the software efficiency value of the storage system.

[0075] like Figure 3The flowchart shown illustrates the process of obtaining the current service running status, current log status, and current thread running status of the software running on the storage system. Based on these obtained service running status, log status, and thread running status (i.e., thread busyness), a software efficiency value is determined. If the obtained current service running status is normal, the current log status is normal, and there are idle threads, then the software efficiency value is determined to be a first software efficiency value. If the obtained current service running status is normal, the current log status is normal, and there are no idle threads, then the software efficiency value is determined to be a second software efficiency value. If the obtained current service running status is normal but the current log status is abnormal, then the software efficiency value is determined to be a third software efficiency value. If an abnormal service running status is detected, then the software efficiency value is determined to be a fourth software efficiency value. The first software efficiency value is any percentage greater than 90%, the second software efficiency value is any percentage between 50% and 90%, the third software efficiency value is any percentage between 30% and 50%, and the fourth software efficiency value is any percentage less than 30%.

[0076] It is understandable that steps S2 and S3 have no specific order. Steps S2 and S3 can be executed simultaneously, or step S2 can be executed first and then step S3, or step S3 can be executed first and then step S2.

[0077] S4. Analyze the hardware efficiency values ​​and software efficiency values ​​determined above to identify outliers.

[0078] The storage system can identify anomalies by analyzing hardware and software efficiency values. These anomalies include both hardware and software anomalies. Specifically, this can be achieved using a pre-defined first mapping table and a second mapping table. The first mapping table maps hardware efficiency values ​​to hardware anomalies, and the second mapping table maps software efficiency values ​​to software anomalies. For hardware anomalies, if the hardware efficiency value is the first or fifth hardware efficiency value (i.e., greater than 70%), the number of hardware anomalies is 0. If the hardware efficiency value is the second hardware efficiency value, the number of hardware anomalies is 4, representing CPU anomalies, memory anomalies, hard disk anomalies, and rack environment anomalies. If the hardware efficiency value is the third hardware efficiency value, the number of hardware anomalies is 1, representing a rack environment anomaly. If the hardware efficiency value is the fourth hardware efficiency value, the number of hardware anomalies is 3, representing CPU anomalies, memory anomalies, and hard disk anomalies. Specifically, for software anomalies, if the software efficiency value is the first software efficiency value, the number of software anomalies is 0; if the software efficiency value is the second software efficiency value, the number of software anomalies is 1, which is a thread anomaly; if the software efficiency value is the third software efficiency value, the number of software anomalies is 2, which are thread anomalies and log anomalies respectively; if the software efficiency value is the fourth software efficiency value, the number of software anomalies is 3, which are thread anomalies, log anomalies, and service operation anomalies respectively.

[0079] S5. Based on the identified anomalies, hardware efficiency values, and software efficiency values, a health report is generated; the storage system is then optimized based on the health report.

[0080] The storage system periodically acquires hardware efficiency values, software efficiency values, and corresponding anomalies, generating a health report. Based on the anomalies in this health report, the storage system automatically triggers corresponding repair operations, such as... Figure 4 Specifically, if a CPU anomaly exists, the number of idle / expired tasks will be reduced based on CPU processing capacity; if a content anomaly exists, the memory usage objects and causes will be analyzed, and release improvement prompts will be generated to remind users to release memory; if a hard disk anomaly exists, capacity will be evenly distributed to other disks or expansion prompts will be generated to remind users to release hard disk capacity; if a rack environment anomaly exists, the fan speed / output will be increased through software control to reduce the ambient temperature or prompts will be generated to remind users to take software control measures; if a service operation anomaly exists, the service will be restarted through software or external blocking points will be reduced; if a log anomaly exists, the anomaly log will be analyzed and an anomaly report will be generated; if a thread anomaly exists, useless threads will be released immediately.

[0081] Furthermore, the storage system can periodically output the health status to the user and provide the aforementioned repair operations as system suggestions. Further, if the hardware efficiency value and / or software efficiency value in the health report is lower than the warning value, the storage system also generates a health warning to remind the user. This health warning includes at least anomalies; the warning value is set by the user, and this embodiment does not limit its setting, but preferably it can be set to 30%.

[0082] This application proposes a storage system health optimization method. By monitoring key hardware and software and outputting corresponding hardware and software efficiency values, the method reflects the health of the storage system and synchronizes these values ​​with customers in a timely manner. This avoids business interruptions caused by anomalies and improves system operating efficiency. Furthermore, it stabilizes the storage system's health at an optimal level. Additionally, by automatically analyzing and resolving anomalies, it reduces human intervention, thereby improving the robustness of the storage system and ultimately enhancing product competitiveness.

[0083] Example 2

[0084] Corresponding to Embodiment 1 above, this application also provides a method for optimizing the health of a storage system, such as... Figure 5 As shown, the details are as follows:

[0085] 5100. Monitor multiple hardware components in the storage system and obtain the current hardware efficiency acquisition values ​​of the multiple hardware components;

[0086] 5200. Monitor the software running on the storage system and obtain the current software service status of the software;

[0087] 5300. Determine the hardware efficiency value based on the current hardware efficiency acquisition value, and determine the software efficiency value based on the current software service status;

[0088] Preferably, determining the hardware efficiency value based on the current hardware efficiency acquisition value includes:

[0089] 5310. Compare the current hardware efficiency acquisition values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple hardware components set in advance, and determine the hardware efficiency value.

[0090] The hardware components include a CPU, memory, hard drive, and chassis. The current hardware efficiency acquisition values ​​include the corresponding current CPU efficiency acquisition values, current memory efficiency acquisition values, current hard drive efficiency acquisition values, and current chassis temperature acquisition values. The theoretical hardware efficiency values ​​include the corresponding theoretical CPU efficiency values, theoretical memory efficiency values, theoretical hard drive efficiency values, and theoretical chassis temperature values.

[0091] Specifically, the theoretical CPU efficiency value represents the CPU's processing power under ideal conditions; the theoretical memory efficiency value represents the memory usage under ideal conditions; the theoretical hard disk efficiency value represents the hard disk's health and load under ideal conditions; and the theoretical chassis temperature efficiency value represents the temperature reading under ideal conditions. Among the current hardware efficiency data collected, the current CPU efficiency value reflects the current CPU processing power, the current memory efficiency value reflects the current memory usage, the current hard disk efficiency value reflects the current hard disk's health and load, and the current chassis temperature value reflects the current chassis temperature reading.

[0092] Preferably, the step of comparing the current hardware efficiency acquisition values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple hardware components pre-set to determine the hardware efficiency value includes:

[0093] 5311. If the current hardware efficiency values ​​of the multiple hardware components are all less than the corresponding theoretical hardware efficiency values, then the hardware efficiency value of the storage system is determined to be the first hardware efficiency value.

[0094] 5312. If the current hardware efficiency values ​​of the multiple hardware components are all greater than or equal to the corresponding theoretical hardware efficiency values, then the hardware efficiency value is determined to be the second hardware efficiency value.

[0095] The first hardware efficiency value is greater than the second hardware efficiency value.

[0096] Preferably, the step of comparing the current hardware efficiency acquisition values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple hardware components pre-set to determine the hardware efficiency value further includes:

[0097] 5313. If the current chassis temperature acquisition value is greater than the theoretical chassis temperature value, and the difference between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than the first preset threshold, then the hardware efficiency value is determined to be the third hardware efficiency value.

[0098] 5314. If the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold, and the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is greater than or equal to the second preset threshold, then the hardware efficiency value is determined to be the fourth hardware efficiency value.

[0099] 5315. If the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold, and the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than the second preset threshold, then the hardware efficiency value is determined to be the fifth hardware efficiency value.

[0100] Among them, the third hardware efficiency value is greater than the second hardware efficiency value, the fourth hardware efficiency value is greater than the third hardware efficiency value, the fifth hardware efficiency value is greater than the fourth hardware efficiency value, and the first hardware efficiency value is greater than the fifth hardware efficiency value.

[0101] Preferably, the current software service status includes the current service running status, the current log status, and the current thread running status of the software. Determining the software efficiency value based on the current software service status includes:

[0102] 5320. If the current service running status is normal, the current log status is normal, and there are idle threads, then the software efficiency value is determined to be the first software efficiency value.

[0103] 5330. If the current service running status is normal, the current log status is normal, and there are no idle threads, then the software efficiency value is determined to be the second software efficiency value.

[0104] 5340. If the current service running status is normal but the current log status is abnormal, then the software efficiency value is determined to be the third software efficiency value.

[0105] 5350. If the detected service operation status is abnormal, then the software efficiency value is determined to be the fourth software efficiency value.

[0106] Among them, the first software efficiency value is greater than the second software efficiency value, the second software efficiency value is greater than the third software efficiency value, and the third software efficiency value is greater than the fourth software efficiency value.

[0107] It is understandable that the steps for confirming hardware efficiency values ​​and software efficiency values ​​are not sequential; they can be performed simultaneously, or the steps for confirming hardware efficiency values ​​can be performed first and then the steps for confirming software efficiency values, or vice versa.

[0108] 5400. Based on the hardware efficiency value and software efficiency value, identify anomalies and generate a health report to optimize the health of the storage system.

[0109] Preferably, the anomalies include hardware anomalies and software anomalies, and determining the anomalies based on the hardware efficiency values ​​and software efficiency values ​​includes:

[0110] 5410. Obtain the hardware efficiency value and determine the hardware anomaly point based on the pre-set first mapping relationship table, wherein the first mapping relationship table includes the mapping relationship between the hardware efficiency value and the hardware anomaly point;

[0111] 5420. Obtain the software efficiency value and determine the software anomaly point based on the pre-set second mapping relationship table, wherein the second mapping relationship table includes the mapping relationship between the software efficiency value and the software anomaly point.

[0112] Specifically, the first mapping table represents the mapping relationship between hardware efficiency values ​​and hardware anomalies, and the second mapping table represents the mapping relationship between software efficiency values ​​and software anomalies. For hardware anomalies, if the hardware efficiency value is the first or fifth hardware efficiency value, the number of hardware anomalies is 0; if the hardware efficiency value is the second hardware efficiency value, the number of hardware anomalies is 4, including CPU anomalies, memory anomalies, hard disk anomalies, and rack environment anomalies; if the hardware efficiency value is the third hardware efficiency value, the number of hardware anomalies is 1, and it is a rack environment anomaly; if the hardware efficiency value is the fourth hardware efficiency value, the number of hardware anomalies is 3, including CPU anomalies, memory anomalies, and hard disk anomalies. For software anomalies, if the software efficiency value is the first software efficiency value, the number of software anomalies is 0; if the software efficiency value is the second software efficiency value, the number of software anomalies is 1, and it is a thread anomaly; if the software efficiency value is the third software efficiency value, the number of software anomalies is 2, including thread anomalies and log anomalies; if the software efficiency value is the fourth software efficiency value, the number of software anomalies is 3, including thread anomalies, log anomalies, and service operation anomalies.

[0113] Preferably, the step of determining anomalies and periodically generating health reports based on the hardware and software efficiency values ​​to optimize the health of the storage system further includes:

[0114] 5430. The storage system periodically acquires the hardware efficiency value, software efficiency value, and anomalies to generate a health report, wherein the health report includes at least one of the hardware efficiency value, software efficiency value, and anomalies;

[0115] 5440. The storage system triggers corresponding repair operations for the anomalies in the health report to optimize the health of the storage system;

[0116] 5450. If the hardware efficiency value and / or the software efficiency value in the health report are lower than the warning value, the storage system also generates a health warning to remind the user.

[0117] Example 3

[0118] Corresponding to Embodiments 1 and 2 above, this application also provides a health optimization system applicable to storage systems, such as... Figure 6 As shown, the system includes:

[0119] The data acquisition module 610 is used to monitor multiple hardware components in the storage system and obtain the current hardware efficiency acquisition value of the multiple hardware components.

[0120] The data acquisition module 610 is also used to monitor the software running on the storage system and obtain the current software service status of the software.

[0121] The efficiency calculation module 620 is used to determine the hardware efficiency value based on the current hardware efficiency acquisition value, and to determine the software efficiency value based on the current software service status.

[0122] The health optimization module 630 is used to determine anomalies and generate a health report based on the hardware efficiency value and software efficiency value in order to optimize the health of the storage system.

[0123] In some embodiments, the efficiency calculation module 620 is further configured to compare the current hardware efficiency acquisition values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple hardware components set in advance, and determine the hardware efficiency value.

[0124] The hardware components include a CPU, memory, hard drive, and chassis. The current hardware efficiency acquisition values ​​include the corresponding current CPU efficiency acquisition values, current memory efficiency acquisition values, current hard drive efficiency acquisition values, and current chassis temperature acquisition values. The theoretical hardware efficiency values ​​include the corresponding theoretical CPU efficiency values, theoretical memory efficiency values, theoretical hard drive efficiency values, and theoretical chassis temperature values.

[0125] In some embodiments, the efficiency calculation module 620 is further configured to determine the hardware efficiency value of the storage system as a first hardware efficiency value when all the current hardware efficiency values ​​of the multiple hardware components are less than the corresponding theoretical hardware efficiency values; the efficiency calculation module 620 is further configured to determine the hardware efficiency value as a second hardware efficiency value when all the current hardware efficiency values ​​of the multiple hardware components are greater than or equal to the corresponding theoretical hardware efficiency values; wherein, the first hardware efficiency value is greater than the second hardware efficiency value.

[0126] In some embodiments, the efficiency calculation module 620 is further configured to determine the hardware efficiency value as a third hardware efficiency value when the acquired current chassis temperature is greater than the theoretical chassis temperature, and the differences between the current CPU efficiency value, the current hard disk efficiency value, and the current memory efficiency value and their corresponding theoretical CPU efficiency values, theoretical hard disk efficiency values, and theoretical memory efficiency values ​​are all less than a first preset threshold; the efficiency calculation module 620 is further configured to determine the hardware efficiency value as a third hardware efficiency value when the acquired current chassis temperature is less than the theoretical chassis temperature and the difference is less than the first preset threshold, and the differences between the current CPU efficiency value, the current hard disk efficiency value, and the current memory efficiency value and their corresponding theoretical CPU efficiency values, theoretical hard disk efficiency values, and theoretical memory efficiency values ​​are all less than a first preset threshold. When the sum of the differences between the values ​​is greater than or equal to a second preset threshold, the hardware efficiency value is determined to be a fourth hardware efficiency value; the efficiency calculation module 620 is further configured to determine the hardware efficiency value to be a fifth hardware efficiency value when the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than a first preset threshold, and when the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than a second preset threshold; wherein, the third hardware efficiency value is greater than the second hardware efficiency value, the fourth hardware efficiency value is greater than the third hardware efficiency value, the fifth hardware efficiency value is greater than the fourth hardware efficiency value, and the first hardware efficiency value is greater than the fifth hardware efficiency value.

[0127] In some embodiments, the efficiency calculation module 620 is further configured to determine the software efficiency value as a first software efficiency value when the current service running status is normal, the current log status is normal, and there are idle threads; the efficiency calculation module 620 is further configured to determine the software efficiency value as a second software efficiency value when the current service running status is normal, the current log status is normal, and there are no idle threads; the efficiency calculation module 620 is further configured to determine the software efficiency value as a third software efficiency value when the current service running status is normal but the current log status is abnormal; the efficiency calculation module 620 is further configured to determine the software efficiency value as a fourth software efficiency value when an abnormal service running status is detected; wherein, the first software efficiency value is greater than the second software efficiency value, the second software efficiency value is greater than the third software efficiency value, and the third software efficiency value is greater than the fourth software efficiency value.

[0128] In some embodiments, the health optimization module 630 is further configured to acquire the hardware efficiency value and determine hardware anomalies based on a pre-set first mapping table, wherein the first mapping table includes a mapping relationship between hardware efficiency values ​​and hardware anomalies; the health optimization module 630 is further configured to acquire the software efficiency value and determine software anomalies based on a pre-set second mapping table, wherein the second mapping table includes a mapping relationship between software efficiency values ​​and software anomalies.

[0129] In some embodiments, the health optimization module 630 is further configured to periodically acquire the hardware efficiency value, software efficiency value, and anomalies to generate a health report, wherein the health report includes at least one of the hardware efficiency value, software efficiency value, and anomalies; the health optimization module 630 is further configured to trigger corresponding repair operations for the anomalies in the health report to optimize the health of the storage system; if the hardware efficiency value and / or the software efficiency value in the health report is lower than a warning value, the health optimization module 630 is further configured to generate a health warning to remind the user.

[0130] Example 4

[0131] Corresponding to all the above embodiments, this application provides an electronic device, including:

[0132] One or more processors; and memory associated with the one or more processors, the memory storing program instructions that, when read and executed by the one or more processors, perform the following operations:

[0133] Monitor multiple hardware components in the storage system and obtain the current hardware efficiency values ​​of the multiple hardware components;

[0134] Monitor the software running on the storage system and obtain the current software service status of the software;

[0135] The hardware efficiency value is determined based on the current hardware efficiency acquisition value, and the software efficiency value is determined based on the current software service status.

[0136] Based on the hardware efficiency values ​​and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system.

[0137] in, Figure 7An exemplary architecture of an electronic device is shown, which may include a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, and a memory 720. The processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, and memory 720 can communicate with each other via a bus 730.

[0138] The processor 710 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs and implement the technical solution provided in this application.

[0139] The memory 720 can be implemented in the form of ROM (Read-Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 720 can store the operating system 721 for controlling the execution of the electronic device 700, and the basic input / output system (BIOS) 722 for controlling the low-level operations of the electronic device 700. Additionally, it can store a web browser 723, a data storage management system 724, and an icon font processing system 725, etc. The aforementioned icon font processing system 725 can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when the technical solution provided in this application is implemented through software or firmware, the relevant program code is stored in the memory 720 and is called and executed by the processor 710.

[0140] Input / output interface 713 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0141] Network interface 714 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0142] Bus 730 includes a pathway for transmitting information between various components of the device, such as processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, and memory 720.

[0143] In addition, the electronic device 700 can also obtain information on specific claim conditions from the virtual resource object claim condition information database for condition judgment, etc.

[0144] It should be noted that although the above-described device only shows the processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, memory 720, bus 730, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the solution of this application, and does not necessarily include all the components shown in the figures.

[0145] Example 5

[0146] Corresponding to all the above embodiments, this application also provides a computer-readable storage medium, characterized in that it stores a computer program that causes a computer to perform the following operations:

[0147] Monitor multiple hardware components in the storage system and obtain the current hardware efficiency values ​​of the multiple hardware components;

[0148] Monitor the software running on the storage system and obtain the current software service status of the software;

[0149] The hardware efficiency value is determined based on the current hardware efficiency acquisition value, and the software efficiency value is determined based on the current software service status.

[0150] Based on the hardware efficiency values ​​and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system.

[0151] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a cloud server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0152] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0153] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A storage system health optimization method, comprising: The method includes: Monitor multiple hardware components in the storage system and obtain the current hardware efficiency values ​​of the multiple hardware components; Monitor the software running on the storage system and obtain the current software service status of the software; The hardware efficiency value is determined based on the current hardware efficiency acquisition value, and the software efficiency value is determined based on the current software service status. Determining the hardware efficiency value based on the current hardware efficiency acquisition value includes: The current hardware efficiency values ​​of multiple hardware components are compared with the corresponding theoretical hardware efficiency values ​​of multiple hardware components set in advance to determine the hardware efficiency value. The current hardware efficiency values ​​include the corresponding current CPU efficiency values, current memory efficiency values, current hard disk efficiency values, and current chassis temperature values. The theoretical hardware efficiency values ​​include the corresponding theoretical CPU efficiency values, theoretical memory efficiency values, theoretical hard disk efficiency values, and theoretical chassis temperature values. If the current chassis temperature acquisition value is greater than the theoretical chassis temperature value, and the difference between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than the first preset threshold, then the hardware efficiency value is determined to be the third hardware efficiency value. If the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold, and the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is greater than or equal to the second preset threshold, then the hardware efficiency value is determined to be the fourth hardware efficiency value. Determining the software efficiency value based on the current software service status includes: If the current service running status is normal, the current log status is normal, and there are idle threads, then the software efficiency value is determined to be the first software efficiency value. If the current service running status is normal, the current log status is normal, and there are no idle threads, then the software efficiency value is determined to be the second software efficiency value. Based on the hardware and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system, including: The hardware efficiency value is obtained and the hardware anomaly point is determined based on the pre-set first mapping relationship table. The first mapping relationship table includes the mapping relationship between hardware efficiency value and hardware anomaly point. If the hardware efficiency value is the third hardware efficiency value, it is an abnormality of the rack environment; if the hardware efficiency value is the fourth hardware efficiency value, it is an abnormality of CPU, memory, and hard disk. The software efficiency value is obtained and the software anomaly is determined based on a pre-set second mapping table. The second mapping table includes a mapping relationship between the software efficiency value and the software anomaly. If the software efficiency value is a first software efficiency value, the software anomaly is 0; if the software efficiency value is a second software efficiency value, it is a thread anomaly. The storage system periodically acquires the hardware efficiency value, software efficiency value, and anomalies to generate a health report. The health report includes at least one of the hardware efficiency value, software efficiency value, and anomalies, and the anomalies include both hardware anomalies and software anomalies. The storage system triggers corresponding repair operations for the anomalies in the health report to optimize the health of the storage system.

2. The method of claim 1, wherein, The process of comparing the current hardware efficiency values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple pre-set hardware components to determine the hardware efficiency value includes: If the current hardware efficiency values ​​of the multiple hardware components are all less than the corresponding theoretical hardware efficiency values, then the hardware efficiency value of the storage system is determined to be the first hardware efficiency value. If the current hardware efficiency acquisition values ​​of the multiple hardware components are all greater than or equal to the corresponding theoretical hardware efficiency values, then the hardware efficiency value is determined to be the second hardware efficiency value. The first hardware efficiency value is greater than the second hardware efficiency value.

3. The method of claim 2, wherein, The step of comparing the current hardware efficiency acquisition values ​​of multiple hardware components with the corresponding theoretical hardware efficiency values ​​of multiple pre-set hardware components to determine the hardware efficiency value further includes: If the current chassis temperature acquisition value is less than the theoretical chassis temperature value and the difference is less than the first preset threshold, and the sum of the differences between the current CPU efficiency acquisition value, the current hard disk efficiency acquisition value, and the current memory efficiency acquisition value and the corresponding theoretical CPU efficiency value, theoretical hard disk efficiency value, and theoretical memory efficiency value is less than the second preset threshold, then the hardware efficiency value is determined to be the fifth hardware efficiency value. The third hardware efficiency value is greater than the second hardware efficiency value, the fourth hardware efficiency value is greater than the third hardware efficiency value, the fifth hardware efficiency value is greater than the fourth hardware efficiency value, and the first hardware efficiency value is greater than the fifth hardware efficiency value.

4. The method of claim 1, wherein, The step of determining the software efficiency value based on the current software service status also includes: If the current service running status is normal but the current log status is abnormal, then the software efficiency value is determined to be the third software efficiency value. If the detected service operation status is abnormal, the software efficiency value is determined to be the fourth software efficiency value. The first software efficiency value is greater than the second software efficiency value, the second software efficiency value is greater than the third software efficiency value, and the third software efficiency value is greater than the fourth software efficiency value.

5. The method according to claim 2 or 3, characterized in that, The step of obtaining the hardware efficiency value and determining hardware anomalies based on a pre-set first mapping table further includes: If the hardware efficiency value is the first or fifth hardware efficiency value, then there are 0 hardware anomalies; if the hardware efficiency value is the second hardware efficiency value, then there are CPU anomalies, memory anomalies, hard disk anomalies, and rack environment anomalies.

6. The method of claim 4, wherein, The step of obtaining the software efficiency value and determining software anomalies based on a pre-set second mapping table further includes: If the software efficiency value is the third software efficiency value, then it indicates thread exceptions and log exceptions; if the software efficiency value is the fourth software efficiency value, then it indicates thread exceptions, log exceptions, and service operation exceptions.

7. The method of claim 1, wherein, Based on the aforementioned hardware and software efficiency values, anomalies are identified and a health report is generated to optimize the health of the storage system. This also includes: If the hardware efficiency value and / or the software efficiency value in the health report are lower than the warning value, the storage system also generates a health warning to remind the user.

8. A health-optimization system, characterized by, For implementing the storage system health optimization method as described in claim 1, the system comprises: The data acquisition module is used to monitor multiple hardware components in the storage system and obtain the current hardware efficiency acquisition values ​​of the multiple hardware components. The data acquisition module is also used to monitor the software running on the storage system and obtain the current software service status of the software. An efficiency calculation module is used to determine a hardware efficiency value based on the current hardware efficiency acquisition value, and to determine a software efficiency value based on the current software service status. The health optimization module is used to identify anomalies and generate a health report based on the hardware efficiency value and software efficiency value in order to optimize the health of the storage system.

9. An electronic device, comprising: The electronic device includes: One or more processors; And a memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that causes the computer to perform the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • An insurance core system monitoring platform

    CN109903175A

  • Communication terminal abnormal power consumption monitoring method and system, terminal equipment and storage medium

    CN111722993A