Server availability scheme combination determination method and electronic device
By constructing a server availability assessment model and a pruning optimization algorithm based on partial order relations, the problem that existing technologies cannot decompose high availability targets into component-level indicators is solved, thus achieving optimization of server availability scheme combinations and the design with the lowest cost.
Patent Information
- Application Number
- CN202511375142.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing server availability specifications cannot be directly mapped to implementable circuit, microarchitecture, or firmware parameters, which makes it impossible to decompose high availability goals into component-level metrics during the system-level design phase, and the reliability databases of various server, storage, and memory vendors are isolated from each other.
By constructing a server availability assessment model, availability scheme combinations are generated based on availability parameters. Partial order relation and pruning optimization algorithm are used to select availability scheme combinations that meet the target constraints and preset cost rules, and decompose them into the indicators of each component.
This approach enables the decomposition of high availability goals into component-level metrics during the early design phase of servers, improving the selection efficiency of pruning optimization algorithms and ensuring the optimization and lowest cost of server availability solution combinations.
Smart Images

Figure CN120872505B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method for determining a combination of server availability schemes and an electronic device. Background Technology
[0002] With the increasing computing power of cloud computing and artificial intelligence, server system availability has become a core indicator determining business continuity and customer satisfaction. However, most current server availability specifications only provide framework-level guidance and cannot be directly mapped to implementable circuit, microarchitecture, or firmware parameters. Furthermore, the reliability databases of various server, storage, and memory vendors are isolated from each other, making it impossible to decompose the goal of high availability into component-level indicators corresponding to various availability solutions during the system-level design phase.
[0003] Therefore, there is an urgent need for a quantifiable method to achieve the goal of high server availability by combining server availability schemes, so as to achieve precise allocation of component parameters in the early design stage of the server. Summary of the Invention
[0004] This application provides a method and electronic device for determining server availability scheme combinations. It can improve the screening efficiency of pruning optimization algorithms by using the partial order relationship between schemes, and can decompose the goal of high availability into a combination of technologies that achieve the indicators of each component.
[0005] This application provides a method for determining a combination of server availability schemes, including:
[0006] A server availability assessment model is constructed based on the availability parameters of each server component in the target server. The server availability assessment model is a model for evaluating the availability indicators of the target server based on availability parameters. Availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component.
[0007] Several availability scheme combinations are generated based on several availability schemes corresponding to the target server, and target constraints are constructed based on several availability scheme combinations, several cost data corresponding to several availability schemes, server availability evaluation models, and preset availability indicators.
[0008] A partial order relationship is determined for several availability schemes. Based on the partial order relationship, a preset pruning optimization algorithm is used to select a target availability scheme combination that meets the target constraints and preset cost rules from several availability scheme combinations. The target availability scheme combination includes the component parameter indicators of each target server component corresponding to each target availability scheme.
[0009] This application also provides a device for determining a combination of server availability schemes, including:
[0010] The model building module is used to build a server availability assessment model based on the availability parameters of each server component in the target server. The server availability assessment model is a model that evaluates the availability indicators of the target server based on availability parameters. Availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component.
[0011] The constraint construction module is used to generate several availability scheme combinations based on several availability schemes corresponding to the target server, and to construct target constraints based on several availability scheme combinations, several cost data corresponding to several availability schemes, server availability evaluation models, and preset availability indicators.
[0012] The scheme combination screening module is used to determine the partial order relationship corresponding to several availability schemes, and select the target availability scheme combination that meets the target constraints and preset cost rules from several availability scheme combinations based on the partial order relationship using a preset pruning optimization algorithm; the target availability scheme combination includes the component parameter indicators of each target server component corresponding to each target availability scheme.
[0013] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described server availability scheme combination determination methods when executing the computer program.
[0014] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described methods for determining a combination of server availability schemes.
[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for determining a combination of server availability schemes.
[0016] In this application, a server availability assessment model can be constructed based on the availability parameters of each server component in the target server. The server availability assessment model is a model for evaluating the availability indicators of the target server based on availability parameters. The availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component. Several availability scheme combinations are generated based on several availability schemes corresponding to the target server, and target constraints are constructed based on several availability scheme combinations, several cost data corresponding to several availability schemes, the server availability assessment model, and preset availability indicators. The partial order relationship corresponding to several availability schemes is determined, and the target availability scheme combination is determined based on the target constraints and the partial order relationship using a preset pruning optimization algorithm. The target availability scheme combination is an availability scheme combination that satisfies the target constraints and the preset cost rules. The server availability indicators corresponding to the target availability scheme combination are decomposed into component parameter indicators corresponding to each server component, and the design and development scheme of the target server is determined based on the corresponding decomposition results.
[0017] Therefore, the method of this application requires constructing a server availability assessment model based on the availability parameters of each server component in the target server. Then, constraints are constructed based on several availability schemes corresponding to the target server, cost data corresponding to these schemes, the availability assessment model, and preset availability indicators. A preset pruning optimization algorithm, based on the constraints and the partial order relationship between availability schemes, determines the combination of availability schemes that satisfy the constraints and preset cost rules. Finally, the combination of availability schemes is decomposed into component parameter indicators corresponding to each server component, and the design and development scheme of the target server is determined based on the decomposition results. In this way, the application objects of server availability technologies can be divided, the partial order relationship of availability technologies can be defined, and an optimization pruning algorithm combining server availability characteristics can be implemented to solve for the combination of availability technologies that satisfies the availability objective and has the lowest cost. Thus, the goal of high availability is decomposed into a combination of technologies that achieve the indicators of each component. Attached Figure Description
[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for determining a combination of server availability schemes provided in this application embodiment;
[0020] Figure 2 A flowchart illustrating a specific method for determining a combination of server availability schemes, as provided in this application embodiment;
[0021] Figure 3 This application provides a schematic diagram of a server architecture hierarchy.
[0022] Figure 4 This is a schematic diagram of a server availability scheme combination determination device provided in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0024] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0025] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Current server availability specifications mostly provide framework-level guidance, which cannot be directly mapped to implementable circuit, microarchitecture, or firmware parameters. Furthermore, the reliability databases of various server, storage, and memory vendors are isolated from each other, making it impossible to decompose high availability goals into verifiable component-level metrics during the system-level design phase.
[0027] To overcome the aforementioned technical problems, this application discloses a method and electronic device for determining server availability scheme combinations. It can improve the screening efficiency of pruning optimization algorithms through the partial order relationship between schemes, and can decompose the goal of high availability into a combination of technologies that achieve the indicators of each component.
[0028] Embodiments of this application provide a method for determining a combination of server availability schemes. The method is described in detail below, in conjunction with its execution flow. The method includes:
[0029] Step S11: Construct a server availability assessment model based on the availability parameters of each server component in the target server; the server availability assessment model is a model for evaluating the availability indicators of the target server based on availability parameters; availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component.
[0030] In this embodiment, to achieve availability index decomposition, it is first necessary to establish an evaluation model for the server availability structure to quantitatively describe the server's availability. Specifically, it is necessary to determine the Mean Time Between Failures (MTBF) and Mean Time To Repair (MTTR) for each server component in the target server. Then, based on the MTBF and MTTR, the overall MTBF and MTTR of the target server are determined. It should be noted that for any component, module, subsystem, or entire server, the set {Item, (Re, Ma)} represents the hardware structure and its availability parameters, where Item is a string representing the hardware structure; Re is a numeric value representing the component's reliability index, which in this embodiment is the MTBF of the server component; and Ma is a numeric value representing the component's maintainability index, which in this embodiment is the MTTR of the server component. Furthermore, through... Represents a set of availability models, where each s i Both are strings representing typical availability models, such as serial, parallel, voting, and side-connection. `r` represents the number of all typical availability models. And each typical availability model `s`... i This corresponds to an explicit functional relation f(s) i This function provides formulas for calculating the reliability and maintainability indices of the modules composed of this availability model, derived from the reliability and maintainability indices of all component parts. In other words, the availability parameters of the entire machine can be calculated based on the availability parameters of each component using these formulas, for example, through f(s) i This converts the mean time between failures (MTBF) of each server component into the mean time between failures (MTBF) of the entire server. It should be noted that server components refer to all parts of the server, including but not limited to memory, hard drives, motherboard, network interface card (NIC), fans, power supply, RAID (Redundant Arrays of Independent Disks) card, and CPU (Central Processing Unit).
[0031] For any server, the available set This represents the availability structure of the server, where Item is a string representing the module, subsystem, or entire machine; s is a string representing the typical availability model corresponding to the top layer of the hierarchical model of the module, subsystem, or entire machine. This is a string type, representing all the components that make up the server after decomposition at the top level of the hierarchical model. Therefore, based on the above, the availability parameter of the Item (Re) can be determined. I Ma I ) can be obtained through the functional relationship f(s) i The specific relationship is determined as follows:
[0032] ;
[0033] Through iterative calculation (Re) I Ma I The overall reliability and maintainability indices of the server can be derived from the reliability and maintainability indices of all its components. Therefore, based on the mean time between failures (MTBF) and mean time to repair (MTB), the MTBF and MTBF of the target server can be determined. Then, a server availability assessment model can be constructed based on the MTBF and MTBF of the server.
[0034] The server availability assessment model is defined as the ratio between the server's mean time between failures (MTBF) and the target time / value, where the target time / value is the sum of the MTBF and the server's mean time to repair. Therefore, the server availability assessment model can be expressed as:
[0035] ;
[0036] Where A is the server availability metric, MTBF 服务器 MTTR is the mean time between failures (MTTR) of the entire server. 服务器 This represents the average repair time for the entire server.
[0037] In this way, the baseline availability metrics of servers can be quantified by building a server availability assessment model.
[0038] Step S12: Generate several availability scheme combinations based on several availability schemes corresponding to the target server, and construct target constraints based on several availability scheme combinations, several cost data corresponding to several availability schemes, server availability evaluation model and preset availability indicators.
[0039] In this embodiment, several availability scheme combinations can be generated based on several availability schemes corresponding to the target server. Then, target constraints are constructed based on these availability scheme combinations, cost data corresponding to the availability schemes, an availability evaluation model, and preset availability indicators. Specifically, it is necessary to determine several availability schemes corresponding to the target server, and to determine the design cost data and production cost data corresponding to each availability scheme. Then, a cost indicator function is constructed based on the design cost data and production cost data. Specifically, the ratio of the design cost data to the preset number of servers can be used as the first function term of the cost indicator function, and the production cost data can be used as the second function term. The sum of the first and second function terms is the cost indicator function. For example, the target server corresponds to an availability scheme A, and the corresponding design cost of availability scheme A is... d and production cost p Cost d To develop and design this availability technology so that it can be applied to production servers at a one-time total cost, Cost p This represents the increased unit cost after applying this availability technology to a server compared to before. To amortize the design cost across each server, a numeric parameter N is defined as representing the projected total sales volume of a server product during its primary sales cycle. Therefore, the cost indicator function can be expressed as:
[0040] ;
[0041] Furthermore, it is necessary to determine the availability impact of several availability schemes and construct a cost-benefit model for the target server based on the availability impact and cost indicator function. It should be noted that the application of availability schemes may have a certain availability impact on the server. For example, the use of a certain availability technology (Tech) may directly change the reliability or maintainability indicators of a component, module, or subsystem (Item), i.e., Tech:(Re, Ma) → (Re', Ma'). Therefore, it is necessary to construct a cost-benefit model for the target server based on the aforementioned availability impact and cost indicator function.
[0042] The next step involves generating several availability scheme combinations from several availability schemes, and then calculating the costs of these combinations using a cost-benefit model. For example, if the target server corresponds to availability scheme A and availability scheme B, and {0, 1} represents whether an availability scheme is applied (1 indicates application, 0 indicates no application), then four combinations can be generated. That is, the server's availability scheme Com can be represented as {0, 0}, {0, 1}, {1, 0}, and {1, 1}. The cost of each combination needs to be determined using the cost indicator function in the cost-benefit model to obtain the costs of the various combinations.
[0043] Ultimately, target constraints need to be constructed using several combinations of availability solutions, preset availability metrics, availability assessment models, and the costs of these combinations. It's important to note that the preset availability metrics are the expected availability metrics achievable after applying the combined solutions. For example, if the current server availability metric is 99.30% and the preset availability metric is 99.90%, then the expected combined solutions will increase the availability metric from 99.30% to 99.90%. Further, the target constraints are: the selected availability solution combinations belong to several availability solution combinations; the availability metric of the selected availability solution combinations is not less than the preset availability metric; the cost of the selected availability solution combinations belongs to any of the costs of several combined solutions; and the availability metric of the selected availability solution combinations is obtained through the server availability assessment model. For example, assuming the server's target availability metric is 99.95%, let Com = (b1, b2, b3, b4) ∈ (0, 1)^4 represent whether availability technologies Tech1 to Tech4 are applied, and Re... Com and Ma Com Let represent the overall server reliability index and maintainability index after applying the corresponding availability technology combination, respectively. Then the target constraint is:
[0044] ;
[0045] Since the reliability index and maintainability index in this embodiment represent the mean time between failures (MTBF) and mean time to repair (MTTR) respectively, the above target constraint can also be expressed as:
[0046] ;
[0047] In this way, by constructing server constraints, the most suitable combination of solutions can be selected, thereby improving the effectiveness of determining the final server availability solution combination.
[0048] Step S13: Determine the partial order relationship corresponding to several availability schemes, and select the target availability scheme combination that meets the target constraints and preset cost rules from several availability scheme combinations based on the partial order relationship using a preset pruning optimization algorithm; the target availability scheme combination includes the component parameter indicators of each target server component corresponding to each target availability scheme.
[0049] In this embodiment, it is first necessary to determine the partial order relationship of several availability schemes. Specifically, it is necessary to determine the partial order relationship between several availability schemes based on the component type of the server component corresponding to the several availability schemes, the availability parameters of the server component corresponding to the several availability schemes, and the scheme cost corresponding to the several availability schemes. It should be noted that for any component, module, subsystem, or system Item contained in the server system, the availability technology partial order relationship on the Item is defined. For availability technology Tech1 that applies to Item1 and availability technology Tech2 that applies to Item2, it is said that there is a partial order relationship on the Item between Tech1 and Tech2: Tech1≥Tech2. And the partial order relationship can be expressed in three aspects: the first aspect is: the first component type corresponding to the first availability scheme in the several availability schemes is the same as the second component type corresponding to the second availability scheme, that is, Item1 and Item2 are both Items; or one of Item1 and Item2 is an Item, and the other is a component or submodule of an Item; or Item1 and Item2 are both components or submodules of Items. The second aspect is that the cost of the first availability plan is no greater than the cost of the second availability plan, meaning the cost index satisfies the relationship Cost(Tech1)≤Cost(Tech2). The third aspect is that the first availability parameter corresponding to the first availability plan and the second availability parameter corresponding to the second availability plan satisfy a preset availability parameter relationship. This preset relationship is that the first mean time between failures (MTBF) corresponding to the first availability plan is no less than the second MTBF corresponding to the second availability plan, and the first average repair time corresponding to the first availability plan is no greater than the second average repair time corresponding to the second availability plan. In other words, for the availability model {Item, {Re, Ma}} of an item, availability model {Item, {Re1, Ma1}} is calculated using availability technology Tech1, and availability model {Item, {Re2, Ma2}} is calculated using availability technology Tech2. The availability index satisfies the relationship: MTBF Re1≥Re2 and average repair time Ma1≤Ma2.
[0050] It should be noted that the partial order relation characterizes the cost of availability techniques and the resulting benefits of improved availability. Availability techniques with a larger partial order in the relation have lower costs and provide greater improvements to reliability metrics (higher mean time between failures, higher reliability) and maintainability metrics (lower mean time to repair, higher maintainability). Therefore, in the optimal availability metric decomposition scheme, if an availability technique with a larger partial order is not applied, then an availability technique with a smaller partial order will also not be applied.
[0051] Therefore, a partial order relation can be used to filter several availability scheme combinations to obtain a number of filtered availability scheme combinations. For example, if the target server corresponds to 4 availability schemes, and there exists a partial order relation Tech4≥Tech3, then the four availability scheme combinations (0,0,1,0), (0,1,1,0), (1,0,1,0), and (1,1,1,0) can be excluded. The partial order relation Tech4≥Tech3 can be seen from the third and fourth positions of these four combinations. Then, a preset pruning optimization algorithm is used to traverse the several filtered availability scheme combinations based on the target constraints to determine several undetermined availability scheme combinations that satisfy the target constraints. Finally, the target availability scheme combination is selected from these undetermined availability scheme combinations based on a preset cost rule. It should be noted that the preset cost rule selects the undetermined availability scheme combination with the minimum cost among the several undetermined availability scheme combinations as the target availability scheme combination. It should be noted that pruning optimization requires that when the availability technique of the larger partial order relation is not applied, the availability technique of the smaller partial order relation cannot be applied either. Therefore, for a partial order relation of length... The availability technology of partial order relation chains, the pruning optimization strategy of this invention reduces the search space size of the subset from an exponential level. Reduced to linear level In this way, by filtering the combination of usability solutions through partial order relations, the processing efficiency of the pruning optimization algorithm can be effectively improved.
[0052] Furthermore, the server availability metrics corresponding to the final selected target availability scheme combination can be decomposed into component parameter metrics for each server component. Specifically, the server availability metrics corresponding to the target availability scheme combination can be determined, and then decomposed according to the server component type corresponding to the server availability metrics to obtain the component parameter metrics for each server component. For example, the final target availability scheme combination might be Scheme A: memory mirroring technology, replacing 8 memory modules with 8 groups of two memory modules in parallel, and Scheme B: dual power redundancy technology, replacing a single power supply with two power supplies in parallel. This allows for the decomposition of server availability metrics, and ultimately, the server development scheme can be set based on the decomposition results; for example, the final development scheme might be a combination of Scheme A and Scheme B. Therefore, the overall server availability metrics can be decomposed into component parameter metrics for each server component. Then, by executing the target availability schemes included in the target availability scheme combination, the server components corresponding to each scheme can achieve their respective component parameter metrics. Finally, through the combined action of each server component under its respective component parameter metrics, the server availability can reach the preset availability metrics in the target constraints. For example, if the preset availability metric is that the overall availability of the server reaches 99.95%, and the current availability metric of the target server obtained through the server availability assessment model is 99.90%, then by executing the above-mentioned schemes A and B in the target availability scheme combination, the availability metric can be increased from 99.90% to 99.95%. Therefore, the availability target can be decomposed into a combination of availability schemes to achieve the availability metrics of each component.
[0053] In this embodiment, a server availability assessment model can be constructed based on the availability parameters of each server component in the target server. The server availability assessment model is a model that evaluates the availability index of the target server based on the availability parameters. The availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component. Several availability scheme combinations are generated based on several availability schemes corresponding to the target server, and target constraints are constructed based on several availability scheme combinations, several cost data corresponding to several availability schemes, the server availability assessment model, and preset availability indexes. The partial order relationship corresponding to several availability schemes is determined, and the target availability scheme combination is determined based on the target constraints and the partial order relationship using a preset pruning optimization algorithm. The target availability scheme combination is an availability scheme combination that satisfies the target constraints and the preset cost rules. The server availability index corresponding to the target availability scheme combination is decomposed into component parameter indexes corresponding to each server component, and the design and development scheme of the target server is determined based on the corresponding decomposition results. Therefore, the method in this embodiment requires constructing a server availability assessment model using the availability parameters of each server component in the target server. Then, constraints are constructed based on several availability schemes corresponding to the target server, cost data corresponding to each scheme, the availability assessment model, and preset availability indicators. A preset pruning optimization algorithm, based on the constraints and the partial order relationship between availability schemes, determines the combination of availability schemes that satisfy the constraints and preset cost rules. Finally, the combination of availability schemes is decomposed into component parameter indicators corresponding to each server component, and the design and development scheme of the target server is determined based on the decomposition results. In this way, the application objects of server availability technologies can be divided, partial order relationships of availability technologies can be defined, and an optimization pruning algorithm combining server availability characteristics can be implemented. Furthermore, the partial order relationship between schemes improves the screening efficiency of the pruning optimization algorithm, thereby quickly selecting the server availability scheme combination that meets the requirements from several server availability schemes. Moreover, the goal of high availability can be decomposed into a combination of availability schemes that achieve the indicators of each component.
[0054] As a preferred embodiment, such as Figure 2The diagram illustrates the implementation process of determining a server availability scheme combination. First, a server availability assessment model needs to be constructed to calculate the server's baseline availability index. For the type of server to be designed, reliability and maintainability models, reliability diagrams, Markov models, and other methods are used to construct an availability assessment model based on the server structure, serving as a baseline model for which availability technologies have not yet been applied. The mean time between failures (MTBF) and mean time to repair (MTBR) parameters of each component included in the baseline model are obtained and input into the baseline model to calculate the server's baseline availability index. The server's baseline availability index represents the server's current availability index. Next, a candidate set of availability technologies needs to be defined, and an availability technology cost-benefit model needs to be established. All available availability technologies are obtained and defined as the candidate set. The cost of applying each availability technology in the candidate set to the design is evaluated. The changes in the baseline model before and after applying each availability technology to the server design are analyzed to clarify the impact of availability technologies on the baseline model's availability. A cost-benefit model is established by combining the cost and availability impact of each availability technology. Furthermore, availability constraints and a cost objective function need to be defined based on the availability objectives. Using the application or non-application of each availability technology as the independent variable and the preset availability objectives as the constraints, a boundary inequality for the actual availability of the server is defined. Based on the cost and usage status of each availability technology, a cost objective function for the server availability index decomposition scheme is defined. Combining the aforementioned independent variable ranges, constraints, and optimization function, a combinatorial optimization mathematical model for the availability index decomposition problem is established. Finally, a pruning optimization algorithm is used to solve for the optimal combination of availability technologies. Each availability technology in the candidate set has two states: applied or not applied. All application states of the candidate set constitute the search space. Based on the application objects of availability technologies, a partial order relationship between availability technologies is defined. A pruning optimization strategy search algorithm is applied to the search space to achieve an optimized search combining server availability characteristics. The solution that satisfies the boundary inequalities corresponding to the constraints and minimizes the cost objective function is found, outputting the optimal combination of availability technologies that satisfies the availability objective.
[0055] As a preferred embodiment, such as Figure 3 The diagram shows a hierarchical structure of the server. Figure 3 Taking a server as an example, this paper illustrates the method for determining server availability scheme combinations in this application using a practical scenario. Here, it is assumed that the availability metric for the entire server is Re. 整机 =MTBF 整机 ≈1.05×10^5h, Ma 整机 =MTTR 整机If the latency is approximately 104 hours, then based on the server availability assessment model in step S11, the baseline availability of the server can be estimated to be approximately 99.90%. Assume there are four available solutions for this server: Tech1: Memory mirroring technology, changing 8 memory modules to 8 groups of two memory modules connected in parallel. Tech2: Dual power supply redundancy technology, changing a single power supply to two power supplies connected in parallel. Tech3: A memory error correction technology that can improve the memory's MTBF. Tech4: A memory voltage control technology that can improve the memory's MTBF. Assume the main sales cycle for this server is one year, and the estimated sales volume is 10,000 units. Assuming the cost and availability impact of the above availability technologies are as follows: For Tech1, its R&D cost is 5 million, and the price of each memory stick is 1,000, then according to the cost index function in step S12, its solution cost can be determined to be 8,500; For Tech2, its R&D cost is 500,000, and the price of each power supply is 500, then its solution cost can be obtained as 550; For Tech3, its R&D cost is 2 million, then its solution cost is 200; For Tech4, its R&D cost is 1.5 million, then its solution cost is 150.
[0056] Assuming the target availability metric for the server is 99.95%, let Com = (b_1, b_2, b_3, b_4) ∈ (0, 1). 4 Indicates whether usability technologies Tech1 through Tech4 are applied, Re Com and Ma Com Let represent the overall server reliability index and maintainability index after applying the corresponding availability technology combination, respectively. Then the target constraints are as follows:
[0057] .
[0058] Next, we need to determine the partial order relationship between availability technologies. Assuming there exists a partial order relationship Tech4≥Tech3, we can eliminate the four cases (0,0,1,0), (0,1,1,0), (1,0,1,0), and (1,1,1,0), leaving a remaining search range of 12. Assuming that after traversal, when Com=(1,0,0,0), Cost(Com)=8500, the server availability Ava Com ≈99.98%; when Com=(0,0,0,1), Cost(Com)=150, server availability Ava Com≈99.96%. The availability solution that achieves the target availability metric with the lowest cost is Com=(0,0,0,1), also known as Tech4. Therefore, the Tech4 solution can be decomposed to break down the solution that achieves a server availability metric of 99.95% into the memory voltage control technology corresponding to Tech4, and a server design and development scheme can be generated based on the memory voltage control technology corresponding to Tech4.
[0059] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0060] See Figure 4 As shown, embodiments of this application also provide a server availability scheme combination determination apparatus, including:
[0061] Model building module 11 is used to build a server availability assessment model based on the availability parameters of each server component in the target server; the server availability assessment model is a model that evaluates the availability indicators of the target server based on availability parameters; availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component;
[0062] The constraint construction module 12 is used to generate several availability scheme combinations based on several availability schemes corresponding to the target server, and to construct target constraints based on several availability scheme combinations, several cost data corresponding to several availability schemes, server availability evaluation model and preset availability indicators.
[0063] The scheme combination screening module 13 is used to determine the partial order relationship corresponding to several availability schemes, and to select the target availability scheme combination that meets the target constraints and preset cost rules from several availability scheme combinations based on the partial order relationship using a preset pruning optimization algorithm; the target availability scheme combination includes the component parameter indicators of each target server component corresponding to each target availability scheme.
[0064] In some embodiments, the model building module 11 may specifically include:
[0065] The first availability parameter determination unit is used to determine the mean time between failures and the mean time to repair for each server component in the target server.
[0066] The second availability parameter determination unit is used to determine the mean time between failures (MTBF) and mean time to repair (MTBR) of the target server based on the MTBF and MTBR of the server.
[0067] The model building unit is used to build a server availability assessment model based on the server's mean time between failures (MTBF) and mean time to repair (MTBR).
[0068] The server availability assessment model is the ratio between the server's mean time between failures (MTBF) and the target time and value; the target time and value are the sum of the server's MTBF and the server's average maintenance time.
[0069] In some embodiments, the constraint construction module 12 may specifically include:
[0070] The cost data determination submodule is used to determine several availability schemes corresponding to the target server, and to determine the design cost data and production cost data corresponding to each availability scheme.
[0071] The function construction submodule is used to construct cost indicator functions based on design cost data and production cost data.
[0072] The revenue model construction submodule is used to determine the availability impact of several availability schemes, and to construct the cost-benefit model for the target server based on the availability impact and cost indicator function.
[0073] The cost calculation submodule is used to generate several availability scheme combinations from several availability schemes, and to calculate the cost of several combination schemes corresponding to several availability scheme combinations through a cost-benefit model.
[0074] The constraint construction submodule is used to construct target constraints based on several combinations of availability solutions, preset availability indicators, availability assessment models, and the costs of several combinations of solutions.
[0075] The target constraints are: the selected availability scheme combination belongs to several availability scheme combinations; the availability index of the selected availability scheme combination is not less than the preset availability index; and the combination cost of the selected availability scheme combination belongs to any one of the costs of several combination schemes. The availability index of the selected availability scheme combination is obtained through a server availability evaluation model.
[0076] In some embodiments, the function construction submodule may specifically include:
[0077] The function term determination unit is used to take the ratio of design cost data to the preset number of servers as the first function term of the cost index function, and the production cost data as the second function term of the cost index function.
[0078] The function building unit is used to use the sum of the first function term and the second function term as the cost index function.
[0079] In some embodiments, the scheme combination screening module 13 may specifically include:
[0080] The partial order relationship determination unit is used to determine the partial order relationship between several availability schemes based on the component type of the server component corresponding to the several availability schemes, the availability parameters of the server component corresponding to the several availability schemes, and the scheme cost corresponding to the several availability schemes.
[0081] In some embodiments, the partial order relationship is that the first component type corresponding to the first availability scheme in a plurality of availability schemes is the same as the second component type corresponding to the second availability scheme; and, the cost of the first scheme corresponding to the first availability scheme is not greater than the cost of the second scheme corresponding to the second availability scheme; and, the first availability parameter corresponding to the first availability scheme and the second availability parameter corresponding to the second availability scheme satisfy a preset availability parameter relationship.
[0082] In some embodiments, the preset availability parameter relationship is that the first mean time between failures (MTBF) corresponding to the first availability scheme is not less than the second mean time between failures (MTBF) corresponding to the second availability scheme, and the first mean time to repair (MTBR) corresponding to the first availability scheme is not greater than the second mean time to repair (MTBR) corresponding to the second availability scheme.
[0083] In some embodiments, the scheme combination screening module 13 may specifically include:
[0084] The first screening unit is used to screen several combinations of availability schemes through a partial order relation to obtain several filtered combinations of availability schemes.
[0085] The second screening unit is used to traverse several screening availability scheme combinations based on target constraints using a preset pruning optimization algorithm, so as to determine several undetermined availability scheme combinations that meet the target constraints among the several screening availability scheme combinations.
[0086] The third filtering unit is used to filter out the target availability solution combination from several undetermined availability solution combinations based on preset cost rules;
[0087] The preset cost rule is to take the combination of available solutions with the minimum cost among several combinations of available solutions as the target available solution combination.
[0088] In some embodiments, the server availability scheme combination determination apparatus may further include:
[0089] The availability metric determination unit is used to determine the server availability metrics corresponding to the target availability scheme combination.
[0090] The indicator decomposition unit is used to decompose the server availability indicator according to the server component type corresponding to the server availability indicator, so as to obtain the component parameter indicator corresponding to each server component.
[0091] For a description of the features in the embodiment corresponding to the server availability scheme combination determination device, please refer to the relevant description in the embodiment corresponding to the server availability scheme combination determination method, which will not be repeated here.
[0092] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the server availability scheme combination determination method.
[0093] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the server availability scheme combination determination method when it runs.
[0094] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0095] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the server availability scheme combination determination method.
[0096] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the server availability scheme combination determination method.
[0097] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0098] The foregoing has provided a detailed description of a server availability scheme combination determination method and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for determining a combination of server availability schemes, characterized in that, include: A server availability assessment model is constructed based on the availability parameters of each server component in the target server; the server availability assessment model is a model for evaluating the availability indicators of the target server according to the availability parameters; the availability parameters include the mean time between failures (MTBF) and mean time to repair (MTBR) of each server component. Several availability scheme combinations are generated based on several availability schemes corresponding to the target server, and target constraints are constructed based on the several availability scheme combinations, several cost data corresponding to the several availability schemes, the server availability evaluation model, and preset availability indicators. Determine the partial order relationship corresponding to the plurality of availability schemes, and select a target availability scheme combination that satisfies the target constraint and the preset cost rule from the plurality of availability scheme combinations based on the partial order relationship using a preset pruning optimization algorithm; The target availability scheme combination includes the component parameter indicators of each target server component corresponding to each target availability scheme; The server availability assessment model is the ratio between the server's mean time between failures (MTBF) and the target time and value; the target time and value are the sum of the server's MTBF and the server's average maintenance time. Wherein, determining the partial order relation corresponding to the plurality of availability schemes includes: The partial order relationship between the availability schemes is determined based on the component type of the server component corresponding to the availability scheme, the availability parameters of the server component corresponding to the availability scheme, and the scheme cost corresponding to the availability scheme. Wherein, the partial order relationship means that the first component type corresponding to the first availability scheme in the plurality of availability schemes is the same as the second component type corresponding to the second availability scheme; And, the cost of the first solution corresponding to the first availability solution is not greater than the cost of the second solution corresponding to the second availability solution; And, the first availability parameter corresponding to the first availability scheme and the second availability parameter corresponding to the second availability scheme satisfy a preset availability parameter relationship; The preset availability parameter relationship is that the first mean time between failures (MTBF) corresponding to the first availability scheme is not less than the second mean time between failures (MTBF) corresponding to the second availability scheme, and the first mean time to repair (MTBR) corresponding to the first availability scheme is not greater than the second mean time to repair (MTBR) corresponding to the second availability scheme.
2. The method for determining server availability scheme combinations according to claim 1, characterized in that, The method of constructing a server availability assessment model based on the availability parameters of each server component in the target server includes: Determine the mean time between failures (MTBF) and mean time to repair (MTB) for each server component in the target server. Based on the mean time between failures (MTBF) and the mean time to repair (MTB), the mean time between failures (MTBF) and the mean time to repair (MTB) of the target server are determined. A server availability assessment model is constructed based on the mean time between failures (MTBF) and mean time to repair (MTBR) of the server.
3. The method for determining server availability scheme combinations according to claim 1, characterized in that, The process of generating several availability scheme combinations based on several availability schemes corresponding to the target server, and constructing target constraints based on the several availability scheme combinations, several cost data corresponding to the several availability schemes, the server availability evaluation model, and preset availability indicators, includes: Determine several availability schemes corresponding to the target server, and determine the design cost data and production cost data corresponding to each of the several availability schemes; Construct a cost index function based on the design cost data and the production cost data; Determine the availability impact of the aforementioned availability schemes, and construct a cost-benefit model for the target server based on the availability impact and the cost indicator function; Several availability scheme combinations are generated from the aforementioned availability schemes, and the cost of several combination schemes corresponding to the aforementioned availability scheme combinations is calculated using the cost-benefit model. Target constraints are constructed based on the combination of several availability schemes, preset availability indicators, the server availability evaluation model, and the cost of the combination schemes. The target constraints are that the selected availability scheme combination belongs to the plurality of availability scheme combinations, the availability index of the selected availability scheme combination is not less than the preset availability index, and the combination scheme cost of the selected availability scheme combination belongs to any one of the plurality of combination scheme costs; wherein, the availability index of the selected availability scheme combination is obtained through the server availability evaluation model.
4. The method for determining server availability scheme combinations according to claim 3, characterized in that, The step of constructing a cost index function based on the design cost data and the production cost data includes: The ratio of the design cost data to the preset number of servers is used as the first function term of the cost index function, and the production cost data is used as the second function term of the cost index function. The sum of the first function term and the second function term is used as the cost index function.
5. The method for determining a combination of server availability schemes according to claim 1, characterized in that, The step of selecting a target availability solution combination that satisfies the target constraints and preset cost rules from the plurality of availability solution combinations based on the partial order relationship using a preset pruning optimization algorithm includes: The partial order relation is used to filter the several availability scheme combinations to obtain several filtered availability scheme combinations. The preset pruning optimization algorithm is used to traverse the several filtered availability scheme combinations based on the target constraints to determine several undetermined availability scheme combinations that satisfy the target constraints. Based on preset cost rules, a target availability scheme combination is selected from the plurality of availability scheme combinations to be determined; The preset cost rule is to take the combination of available solutions with the minimum cost among the combinations of available solutions to be determined as the target available solution combination.
6. The method for determining a combination of server availability schemes according to any one of claims 1 to 5, characterized in that, After determining the partial order relation corresponding to the plurality of availability schemes, and using a preset pruning optimization algorithm to select a target availability scheme combination that satisfies the target constraint and preset cost rule from the plurality of availability scheme combinations based on the partial order relation, the method further includes: Determine the server availability metrics corresponding to the target availability scheme combination; The server availability index is decomposed according to the server component type corresponding to the server availability index to obtain the component parameter index corresponding to each server component.
7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the server availability scheme combination determination method as described in any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Cloud service availability assessment method and system
CN106571969A
Fault processing method and device and server
CN111414268A