Method, device, computer device and storage medium for improving system availability
By acquiring and optimizing volume performance in a multi-controller array storage system, and dynamically managing the number of volumes, CPU core policies, and cache space, the problem of unbalanced read and write performance is solved, and the availability and stability of the system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2023-03-09
- Publication Date
- 2026-05-29
AI Technical Summary
Current multi-controller array storage systems suffer from uneven read and write performance in high-end storage, leading to decreased cluster performance and affecting the availability of the entire system. They also lack unified volume number, CPU core strategy, and volume cache allocation management.
By acquiring the theoretical and actual performance of volumes in each business model, the number of volumes, CPU core strategy, and volume cache space are dynamically adjusted, and the strategy is optimized in real time to maintain performance consistency. The system is managed using a data acquisition module, a monitoring module, and a processing module.
Without compromising reliability, the system's stability and availability are improved, ensuring the stability of cluster performance and the efficient operation of the entire system.
Smart Images

Figure CN116774921B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to methods, apparatus, computer devices, and storage media for improving system availability. Background Technology
[0002] Currently, multi-controller array storage, especially high-end storage, requires high read and write performance to support front-end services. Many factors influence read and write performance, but there is generally no unified management of volume quantity, CPU core strategy, and volume cache allocation. When front-end services are saturated, a particular module can easily become a significant performance bottleneck, causing a decline in cluster performance and consequently affecting the availability of the entire system. Summary of the Invention
[0003] In order to solve the technical problems existing in the prior art, the present invention provides a method, apparatus, computer equipment and storage medium for improving system availability.
[0004] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0005] In a first aspect, in one embodiment of the present invention, a method for improving system availability is provided, the method comprising the following steps:
[0006] S10. Obtain the theoretical performance of the volume in each preset business model;
[0007] S20. Real-time acquisition of the actual performance of volumes in the currently running business model;
[0008] S30. Based on the theoretical performance data and actual performance of the volume, optimize the volume strategy.
[0009] As a further aspect of the present invention, before step S10, which involves obtaining the theoretical performance of the volume in each preset business model, the following steps are also included:
[0010] When in the first mode, when creating a volume, it is determined whether there is a setting instruction for the number of volumes. If so, the number of volumes, CPU core policy and volume cache space size are set according to the setting instruction and business model.
[0011] If not, the matching number of volumes, CPU core policy, and volume cache size are selected from the embedded configuration library based on the business model.
[0012] As a further aspect of the present invention, S30 involves optimizing the volume based on its theoretical performance and actual performance, including in a first performance mode.
[0013] If the difference between the actual performance and the theoretical performance of the roll exceeds the threshold percentage B1 within the duration A1, an alarm will be issued.
[0014] If the actual performance of the volume is less than its theoretical performance, it is necessary to simultaneously report whether to switch to the second performance mode.
[0015] As a further aspect of the present invention, before step S10, which involves obtaining the theoretical performance of the volume in each preset business model, the following steps are also included:
[0016] If the second performance mode is confirmed before carrying the business, when creating the volume, it is determined whether there is a setting instruction for setting the number of volumes. If so, the number of volumes, CPU core policy and volume cache space size are set according to the setting instruction and business model.
[0017] If not, select the matching number of volumes, CPU core policy, and volume cache space size from the embedded configuration library based on the business model;
[0018] Based on the business model, a simulation test is conducted under the current configuration for a duration of D1. If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm is reported. If the actual performance of the volume is greater than the theoretical performance of the volume, no action is taken.
[0019] As a further aspect of the present invention, in step S30, based on the theoretical performance and actual performance of the volume, strategy optimization is performed on the volume, including in the second performance mode.
[0020] If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm will be reported. If the actual performance of the volume is greater than the theoretical performance of the volume, and the duration continues to exceed C1, the theoretical data under the current configuration in the embedded database will be modified.
[0021] If the actual performance of the roll is less than its theoretical performance, then adjust the portion with the larger percentage difference based on the percentage difference between the actual and theoretical performance of the roll until the alarm is cleared.
[0022] As a further aspect of the present invention, in step S30, based on the theoretical performance and actual performance of the volume, strategy optimization is performed on the volume, including in the second performance mode.
[0023] If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm is reported. If the actual performance of the volume is greater than the theoretical performance of the volume, and the duration continues to exceed C1, the theoretical data under the current configuration in the embedded database is modified. If the actual performance of the volume is less than the theoretical performance of the volume, only the part with the larger difference percentage is adjusted according to the difference percentage between the actual performance and the theoretical performance of the volume until the alarm is cleared, and a practice log is reported at the same time.
[0024] As a further aspect of the present invention, the theoretical performance of the volume is the theoretical volume cache utilization and the theoretical CPU core utilization.
[0025] Secondly, in another embodiment provided by the present invention, an apparatus for improving device availability is provided, the apparatus comprising: a data acquisition module, a monitoring module, and a processing module;
[0026] The data acquisition module is used to acquire the theoretical performance of the volume in each preset business model.
[0027] The monitoring module is used to obtain the actual performance of the volume in the currently running business model in real time.
[0028] The processing module is used to optimize the volume strategy based on the volume's theoretical performance data and actual performance.
[0029] Thirdly, in another embodiment provided by the present invention, a computer device is provided, including an LCA management module, a CPU module, a cache module, and a disk array module, wherein the CPU module, the cache module, and the disk array module are all communicatively connected to the LCA management module, and the LCA management module implements the steps of the method for improving system availability when loading and executing the computer program.
[0030] Fourthly, in another embodiment of the present invention, a storage medium is provided storing a computer program that, when loaded and executed by a processor, implements the steps of the method for improving system availability.
[0031] The technical solution provided by this invention has the following beneficial effects:
[0032] The present invention provides a method, apparatus, computer device, and storage medium for improving system availability. These include: obtaining the theoretical performance of volumes in each preset business model; obtaining the actual performance of volumes in the currently running business model in real time; and optimizing volume strategies based on the theoretical and actual performance data. The present invention dynamically manages the CPU module, cache module, and disk array module in the system, particularly optimizing the number of volumes, CPU core strategies, and volume cache allocation in high-performance mode. This ensures stable cluster performance without affecting reliability, thereby improving the availability of the entire system.
[0033] These or other aspects of the invention will become more apparent from the following description of embodiments. It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0035] Figure 1 A flowchart illustrating a method for improving system availability according to an embodiment of the present invention;
[0036] Figure 2 This is a structural block diagram of an apparatus for improving system availability according to an embodiment of the present invention;
[0037] Figure 3 This is a structural block diagram of a computer system according to an embodiment of the present invention.
[0038] In the diagram: Data acquisition module-100, monitoring module-200, processing module-300. Detailed Implementation
[0039] Various embodiments and / or forms are described below with reference to the accompanying drawings. In the following description, numerous specific details are disclosed for illustrative purposes to provide a general understanding of more than one form. However, those skilled in the art will understand that these forms can be implemented without specific details. Specific examples of more than one form will be described in detail in the following description and drawings. However, these forms are merely illustrative and may utilize a portion of the principles and methods of various forms; the descriptions are intended to encompass all forms and their equivalents. Specifically, the terms "embodiment," "example," "form," "illustration," etc., as used in this specification can be interpreted as meaning that any form or design described may be better or more advantageous than other forms or designs.
[0040] Furthermore, various forms and characteristics can be embodied in systems that include more than one device, terminal, server, equipment, component, and / or module. It should be understood and recognized that various systems may include additional devices, terminals, servers, equipment, components, and / or modules, and / or may not include all of the multiple devices, terminals, servers, equipment, components, and modules shown in the figure.
[0041] The terms "computer program," "component," "module," and "system" used in this specification are used interchangeably, and "computer" refers to related entities, hardware, firmware, software, combinations of software and hardware, or the execution of software. For example, a component can be a process executing on a processor, a processor, an object, a thread of execution, a program, and / or a computer, but is not limited thereto. For example, it can be an application program executing on a computer device and / or all components of the computing device. More than one component can be installed within a processor and / or a thread of execution. A component can be localized within a single computer. A component can also be distributed between two or more computers.
[0042] Furthermore, these components can be executed by various computer-readable media constructed to internally store various data. These components, for example, can communicate locally and / or remotely based on signals having more than one data packet (e.g., data emitted by a component interacting with other components on a local system or a distributed system, and data transmitted to other systems via networks such as the Internet).
[0043] Hereinafter, regardless of the symbols used in the drawings, the same or similar constituent elements will be assigned the same symbols, and repeated descriptions of these elements will be omitted. Furthermore, when describing the embodiments disclosed in this specification, detailed descriptions of well-known technologies will be omitted if it is determined that such detailed descriptions would obscure the essence of the invention. Moreover, the accompanying drawings are only for easier understanding of the embodiments disclosed in this specification, and the technical concepts disclosed in this specification are not limited to the drawings.
[0044] The terminology used in this specification is for illustrative purposes and not for limiting the invention. Unless otherwise specified, the singular includes the plural. The use of “comprises” and / or “comprising” in this specification does not exclude the presence or addition of more than one other constituent element in addition to the mentioned constituent elements.
[0045] The terms "first," "second," etc., can be used to describe various elements or components, but the elements or components are not limited to those terms. The terms are used to distinguish one element or component from others. Therefore, the first element or component mentioned below can also be a second element or component within the technical concept of this invention.
[0046] Unless otherwise defined, all terms used in this specification (including technical and scientific terms) are to be understood in the sense commonly understood by one of ordinary skill in the art to which this invention pertains. Furthermore, terms defined in commonly used dictionaries should not be interpreted ideally or excessively unless specifically defined otherwise.
[0047] Furthermore, the term "or" does not mean exclusive "or" but inclusive "or". That is, unless otherwise specific or contextually ambiguous, "X uses A or B" implies one of the natural connotations. That is, "X uses A or B" can be any of the above when X uses A or B; X uses B or X uses both A and B. And it should be understood that the term "and / or" as used in this specification refers to all possible combinations of more than one of the related items listed.
[0048] In addition, the terms “information” and “data” used in this specification are generally used interchangeably.
[0049] The suffixes “module” and “section” used in the following description of the constituent elements are merely assigned or used interchangeably for the convenience of writing the specification, and they do not have any distinguishing meaning or function in themselves.
[0050] Specifically, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0051] Please see Figure 1 , Figure 1 This is a flowchart of a method for improving system availability provided by an embodiment of the present invention, such as... Figure 1 As shown, the method for improving system availability includes steps S10 to S30. This method is applied to a computer system that includes a first performance mode and a second performance mode, wherein the first performance mode is a normal performance mode and the second performance mode is a high-performance mode.
[0052] S10. Obtain the theoretical performance of the volume in each preset business model.
[0053] In embodiments of the present invention, the theoretical performance of a volume is defined as its theoretical volume cache utilization and theoretical CPU core utilization. The business model includes large and small data blocks, random / sequential operations, read / write operations, etc.
[0054] Before step S10, which involves obtaining the theoretical performance of the volume in each preset business model, the following steps are also included:
[0055] When in the first mode, when creating a volume, it is determined whether there is a setting instruction for the number of volumes. If so, the number of volumes, CPU core policy and volume cache space size are set according to the setting instruction and business model.
[0056] If not, the matching number of volumes, CPU core policy, and volume cache size are selected from the embedded configuration library based on the business model. The CPU core policy is mainly a core-binding policy.
[0057] The embedded configuration library stores the default configuration determined by best practices.
[0058] In an embodiment of the present invention, before step S10, which involves obtaining the theoretical performance of a volume in each preset business model, the method further includes:
[0059] If the second performance mode is confirmed before carrying the business, when creating the volume, it is determined whether there is a setting instruction for setting the number of volumes. If so, the number of volumes, CPU core policy and volume cache space size are set according to the setting instruction and business model.
[0060] If not, the matching number of volumes, CPU core policy, and volume cache size are selected from the embedded configuration library based on the business model. The CPU core policy is mainly a core-binding policy.
[0061] Based on the business model, a simulation test is conducted under the current configuration for a duration of D1. If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm is reported. If the actual performance of the volume is greater than the theoretical performance of the volume, no action is taken.
[0062] S20: Obtain the actual performance of the volume in the currently running business model in real time.
[0063] In embodiments of the present invention, the actual performance of a volume is the actual volume cache utilization and the actual CPU core utilization.
[0064] S30. Based on the theoretical performance data and actual performance of the volume, optimize the volume strategy.
[0065] In an embodiment of the present invention, step S30 involves optimizing the volume based on its theoretical and actual performance, including in a first performance mode.
[0066] If the difference between the actual performance and the theoretical performance of the roll exceeds the threshold percentage B1 within the duration A1, an alarm will be issued.
[0067] If the actual performance of the volume is less than its theoretical performance, it is necessary to simultaneously report whether to switch to the second performance mode.
[0068] S30, based on the theoretical and actual performance of the volume, performs strategy optimization on the volume, including in the second performance mode.
[0069] If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm will be reported. If the actual performance of the volume is greater than the theoretical performance of the volume, and the duration continues to exceed C1, the theoretical data under the current configuration in the embedded database will be modified.
[0070] If the actual performance of the roll is less than its theoretical performance, then adjust the portion with the larger percentage difference based on the percentage difference between the actual and theoretical performance of the roll until the alarm is cleared.
[0071] S30, based on the theoretical and actual performance of the volume, performs strategy optimization on the volume, including in the second performance mode.
[0072] If the difference between the actual performance and the theoretical performance of a volume exceeds a threshold percentage B2 within a duration A2, an alarm is reported. If the actual performance of a volume is greater than its theoretical performance, and the duration continues to exceed C1, the theoretical data under the current configuration in the embedded database is modified. If the actual performance of a volume is less than its theoretical performance, only the portion with the larger percentage difference is adjusted according to the percentage difference between the actual and theoretical performance of the volume until the alarm is cleared. At the same time, a practice log is reported to recommend that the volume creation and other operations be completed in high-performance mode before new services are launched.
[0073] Among them, A1, A2, B1, B2, C1 and D1 are all system preset parameters, which can be adjusted in the system or through serial port modules.
[0074] This invention dynamically manages the CPU module, cache module, and disk array module in the system. In particular, it optimizes the number of volumes, CPU core strategy, and volume cache allocation in high-performance mode, ensuring the stability of cluster performance without affecting reliability, thereby improving the availability of the entire system.
[0075] It should be understood that although the above description follows a certain order, these steps are not necessarily executed in that order. Unless otherwise expressly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, some steps in this embodiment may include multiple steps or multiple stages, which are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least a portion of the steps or stages in other steps.
[0076] In one embodiment, see Figure 2 As shown, an apparatus for improving system availability is also provided in an embodiment of the present invention, the apparatus including a data acquisition module 100, a monitoring module 200 and a processing module 300.
[0077] The data acquisition module 100 is used to acquire the theoretical performance of the volume in each preset business model.
[0078] The monitoring module 200 is used to obtain the actual performance of the volume in the currently running business model in real time.
[0079] The processing module 300 is used to optimize the volume strategy based on the volume's theoretical performance data and actual performance.
[0080] In one embodiment, see Figure 3 As shown, an embodiment of the present invention also provides a computer device, including an LCA management module, a CPU module, a cache module and a disk array module, wherein the CPU module, the cache module and the disk array module are all communicatively connected to the LCA management module.
[0081] The LCA management module is used to execute the method for improving system availability. When executing instructions, the LCA module implements the steps in the above method embodiments:
[0082] S10. Obtain the theoretical performance of the volume in each preset business model;
[0083] S20. Real-time acquisition of the actual performance of volumes in the currently running business model;
[0084] S30. Based on the theoretical performance data and actual performance of the volume, optimize the volume strategy.
[0085] The CPU module is the control module for the array storage, carrying the storage system software. It is managed by LCA (especially core binding actions) and sends real-time CPU core utilization information to the LCA management module.
[0086] The cache module is located on the board and generally includes memory and high-speed cache within the component. It is managed by LCA and sends information such as volume cache usage rate to the LCA management module in real time.
[0087] The disk array module carries volume data and is managed by LCA.
[0088] The indicator module is located on the board and is directly controlled by the serial port module. It indicates the real-time status of the current LCA management module.
[0089] The wireless module can convert the serial port module signal into a wireless signal such as WIFI, allowing external devices to exchange information with the LCA management module without the need for a physical serial cable.
[0090] The serial port module allows for information exchange, parameter presets, and enabling of related functions between the external system and the LCA management module.
[0091] The LCA management module is located on the board and is primarily used with programmable logic devices such as ARM. This module can dynamically manage the CPU module, cache module, and disk array module in high-performance mode and normal mode respectively (LCA is used to refer to volume combination affinity in the following text). The LCA management module monitors the performance data, volume cache utilization, and CPU core utilization of each volume in the storage system in real time under different front-end service models (large and small data blocks, random / sequential, read / write, etc.).
[0092] For example, when creating a volume in normal performance mode, if the customer does not specify the number of volumes, the LCA management module selects the matching number of volumes, CPU core policy, and volume cache space size from the embedded configuration library based on the customer's selected business model. If the customer specifies the number of volumes, the LCA management module selects the matching CPU core policy (mainly core binding policy), volume cache space size, etc., from the embedded configuration library based on the customer's selected business model and number of volumes. The embedded configuration library is the default configuration determined by LCA according to storage best practices. In actual operation of normal mode, if the difference between the actual performance and the theoretical performance of the volume exceeds a threshold percentage of 25% within a 20-minute period, an alarm is reported (the following assumes that the business model has not changed significantly at this time; if the model has changed, it is not within the constraints of this document). If the actual performance of the volume is greater than the theoretical performance of the volume, and the duration continues to exceed 50 minutes, the theoretical performance of the volume under the current configuration in the embedded database is modified. If the actual performance of the volume is less than the theoretical performance of the volume, it is necessary to simultaneously report whether to switch to high-performance mode. In high-performance mode, there are two scenarios: 1. If the high-performance mode is confirmed before undertaking the business, when creating a volume, if the customer does not specify the number of volumes, the LCA management module selects the matching number of volumes, CPU core policy (mainly core binding policy), and volume cache space size from the embedded configuration library based on the customer's selected business model. If the customer specifies the number of volumes, the LCA management module selects the matching CPU core policy (mainly core binding policy), volume cache space size, etc., from the embedded configuration library based on the customer's selected business model and number of volumes. Before the business runs, LCA will conduct a 30-minute simulation test based on the business model and the current configuration. During this period, if the difference between the actual performance and the theoretical performance of the volume exceeds a threshold percentage of 15% within 10 minutes, an alarm will be reported. If the actual performance of the volume is greater than the theoretical performance, no action will be taken. If the actual performance of the volume is less than the theoretical performance, LCA will adjust the volume cache space size, CPU core policy, and number of volumes (if not specified by the customer) according to the difference between the actual and theoretical data, such as volume cache utilization and CPU core utilization, and re-perform the simulation test until it passes. In actual operation of high-performance mode, if the difference between the actual performance and the theoretical performance of a volume exceeds a threshold percentage of 15% within a 10-minute period, an alarm will be reported. If the actual performance of a volume is greater than its theoretical performance, and this difference persists for more than 50 minutes, the theoretical data under the current configuration in the embedded database will be modified. If the actual performance of a volume is less than its theoretical performance, then according to LCA, only the portion with the larger percentage difference between the actual and theoretical data, such as volume cache utilization and CPU core utilization, will be adjusted until the alarm is cleared. 2. If the service being carried is confirmed to be in high-performance mode (generally switched from normal mode), then the number of volumes, etc., is already determined.If the difference between the actual performance and the theoretical performance of a volume exceeds a threshold percentage of 15% within a 10-minute period, an alarm will be reported. If the actual performance of a volume is greater than its theoretical performance and the duration continues to exceed 50 minutes, the theoretical data under the current configuration in the embedded database will be modified. If the actual performance of a volume is less than its theoretical performance, the LCA will adjust only the portion with the larger percentage difference based on the actual and theoretical data, such as volume cache utilization and CPU core utilization, until the alarm is cleared. At the same time, a practice log will be reported, and customers will be advised to complete volume creation and other operations in high-performance mode before starting new services.
[0093] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0094] The communication interface is used for communication between the aforementioned terminal and other devices.
[0095] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0096] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0097] The computer equipment includes user equipment and network equipment. The user equipment includes, but is not limited to, computers, smartphones, and PDAs; the network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing, which is a type of distributed computing consisting of a super virtual computer composed of a group of loosely coupled computers. The computer equipment can operate independently to implement the present invention, or it can connect to a network and interact with other computer equipment within the network to implement the present invention. The network in which the computer equipment is located includes, but is not limited to, the Internet, wide area networks (WANs), metropolitan area networks (MANs), local area networks (LANs), and VPN networks.
[0098] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0099] In one embodiment of the present invention, a storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps in the above method embodiments:
[0100] S10. Obtain the theoretical performance of the volume in each preset business model;
[0101] S20. Real-time acquisition of the actual performance of volumes in the currently running business model;
[0102] S30. Based on the theoretical performance data and actual performance of the volume, optimize the volume strategy.
[0103] This invention dynamically manages the CPU module, cache module, and disk array module in the system. In particular, it optimizes the number of volumes, CPU core strategy, and volume cache allocation in high-performance mode, ensuring the stability of cluster performance without affecting reliability, thereby improving the availability of the entire system.
[0104] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Furthermore, any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory.
[0105] It should be understood that, as used herein, the singular form "a" is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations of one or more of the associatedly listed items. The embodiment numbers disclosed above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0106] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the invention (including the claims) is limited to these examples. Within the framework of the invention, technical features of the above embodiments or different embodiments can be combined, and many other variations of different aspects of the invention exist, which are not provided in the details for the sake of brevity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.
Claims
1. A method for improving system availability, characterized in that, The method includes: S10. Obtain the theoretical performance of the volume in each preset business model; S20. Real-time acquisition of the actual performance of volumes in the currently running business model; S30. Based on the theoretical performance data and actual performance of the volume, optimize the volume strategy. Before step S10, obtaining the theoretical performance of the volume in each preset business model, when in the first performance mode, when creating a volume, it is determined whether there is a setting instruction for setting the number of volumes. If there is, the number of volumes, CPU core policy and volume cache space size are set according to the setting instruction and business model. If not, the matching number of volumes, CPU core policy and volume cache space size are selected from the embedded configuration library according to the business model. S30: Based on the volume's theoretical performance and actual performance, perform strategy optimization on the volume, including in the first performance mode; if the difference between the volume's actual performance and theoretical performance exceeds a threshold percentage B1 within a duration A1, then issue an alarm; if the volume's actual performance is less than its theoretical performance, then it is necessary to simultaneously report whether to switch to the second performance mode. Before step S10, which obtains the theoretical performance of the volume in each preset business model, if the second performance mode is confirmed before carrying the business, when creating the volume, it is determined whether there is a setting instruction for setting the number of volumes. If there is, the number of volumes, CPU core policy, and volume cache space size are set according to the setting instruction and the business model. If not, the matching number of volumes, CPU core policy, and volume cache space size are selected from the embedded configuration library according to the business model. A simulation test is performed under the current configuration according to the business model for a duration of D1. If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm is reported. If the actual performance of the volume is greater than the theoretical performance of the volume, no action is taken. S30 involves optimizing the volume strategy based on its theoretical and actual performance. This includes, in the second performance mode, reporting an alarm if the difference between the volume's actual and theoretical performance exceeds a threshold percentage B2 within a duration A2; modifying the theoretical data in the embedded database under the current configuration if the volume's actual performance is greater than its theoretical performance and the duration continues to exceed C1; and adjusting the portion with the larger difference percentage based on the difference between the volume's actual and theoretical performance until the alarm is cleared.
2. The method for improving system availability as described in claim 1, characterized in that, S30: Based on the volume's theoretical performance and actual performance, optimize the volume strategy. In the second performance mode, if the volume's actual performance is less than its theoretical performance, adjust only the portion with the larger percentage difference between the volume's actual and theoretical performance until the alarm is cleared, and simultaneously report a practice log.
3. The method for improving system availability as described in claim 1, characterized in that, The theoretical performance of a volume consists of its theoretical volume cache utilization and theoretical CPU core utilization.
4. An apparatus for improving system availability, characterized in that, The device includes: a data acquisition module, a monitoring module, and a processing module; The data acquisition module is used to acquire the theoretical performance of the volume in each preset business model; The monitoring module is used to obtain the actual performance of the volume in the currently running business model in real time; The processing module is used to optimize the volume strategy based on the volume's theoretical performance data and actual performance. And modules for the following functions: When in the first performance mode, when creating a volume, determine whether there is a setting instruction for setting the number of volumes. If there is, set the number of volumes, CPU core policy and volume cache space size according to the setting instruction and business model. If not, select the matching number of volumes, CPU core policy and volume cache space size from the embedded configuration library according to the business model. If the second performance mode is confirmed before carrying the business, when creating the volume, it is determined whether there is a setting instruction for setting the number of volumes. If there is, the number of volumes, CPU core policy and volume cache space size are set according to the setting instruction and business model. If not, the matching number of volumes, CPU core policy and volume cache space size are selected from the embedded configuration library according to the business model. A simulation test with a duration of D1 is performed under the current configuration according to the business model. If the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm is reported. If the actual performance of the volume is greater than the theoretical performance of the volume, no action is taken. The processing module is also used for: In the first performance mode; if the difference between the actual performance of the volume and the theoretical performance of the volume exceeds the threshold percentage B1 within the duration A1, an alarm is issued; if the actual performance of the volume is less than the theoretical performance of the volume, it is necessary to report whether to switch to the second performance mode at the same time. In the second performance mode, if the difference between the actual performance and the theoretical performance of the volume exceeds the threshold percentage B2 within the duration A2, an alarm is reported; if the actual performance of the volume is greater than the theoretical performance of the volume, and the duration continues to exceed C1, the theoretical data under the current configuration in the embedded database is modified; if the actual performance of the volume is less than the theoretical performance of the volume, the portion with the larger difference percentage is adjusted according to the difference percentage between the actual performance and the theoretical performance of the volume until the alarm is cleared.
5. A computer device, comprising an LCA management module, a CPU module, a cache module, and a disk array module, wherein, The CPU module, cache module, and disk array module are all communicatively connected to the LCA management module. When the LCA management module loads and executes the computer program, it implements the steps of the method for improving system availability as described in any one of claims 1-3.
6. A storage medium storing a computer program that, when loaded and executed by a processor, implements the steps of the method for improving system availability as described in any one of claims 1-3.