Server hardware quality evaluation method and device, computer equipment, readable storage medium and program product
By acquiring out-of-band hardware, change, and fault information of servers, calculating alarm, replacement, and downtime rates, and combining this with the server's years of service to assess server quality, the limitations of traditional assessment methods are overcome, enabling multi-dimensional quality assessment and risk prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional server quality assessment methods are unable to fully reflect the complex performance bottlenecks and potential risks in actual operation, and are limited to a single dimension of assessment.
By acquiring out-of-band hardware information, change information, and fault information of the server, the alarm generation rate, component replacement rate, and downtime rate are calculated. Combined with the average service life and comprehensive quality assessment formula, the hardware quality of the server is comprehensively evaluated.
It enables comprehensive evaluation of server quality from multiple dimensions, reflecting hardware reliability, failure frequency, and maintenance costs, providing quantitative hardware quality performance and risk assessment, and optimizing hardware selection and operation and maintenance strategies.
Smart Images

Figure CN121833433A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of server hardware quality evaluation, and particularly relates to a server hardware quality evaluation method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] With the rapid development of cloud computing, big data and distributed computing technology, servers as the core infrastructure of data centers undertake core tasks such as massive data processing and high-concurrency request response in enterprise-level applications, Internet services and key business scenarios.
[0003] However, the traditional server quality evaluation method is often limited to a single dimension (such as processor utilization or network delay), and it is difficult to fully reflect the complex performance bottlenecks and potential risks in actual operation. SUMMARY
[0004] Therefore, it is necessary to provide a server hardware quality evaluation method, device, computer equipment, computer readable storage medium and computer program product capable of comprehensively evaluating server quality in view of the above technical problems.
[0005] In a first aspect, the present application provides a server hardware quality evaluation method, comprising:
[0006] obtaining server out-of-band hardware information, server out-of-band change information and server failure information of a plurality of servers;
[0007] determining the alarm generation rate, the spare part replacement rate and the downtime rate of the plurality of servers based on the server out-of-band hardware information, the server out-of-band change information and the server failure information;
[0008] obtaining the average service life of the plurality of servers, and determining the comprehensive quality score of the plurality of servers based on the average service life, the alarm generation rate, the spare part replacement rate, the downtime rate and a comprehensive quality evaluation formula.
[0009] In one of the embodiments, the alarm generation rate, the spare part replacement rate and the downtime rate of the plurality of servers are determined based on the server out-of-band hardware information, the server out-of-band change information and the server failure information, comprising:
[0010] Based on out-of-band hardware information, determine the number of servers and alarms generated for various types of servers; based on the number of servers and alarms generated, determine the alarm generation rate for various types of servers; based on out-of-band change information, determine the number of servers and parts replaced for various types of servers; based on the number of servers and parts replaced, determine the parts replacement rate for various types of servers; based on server fault information, determine the number of servers and the number of downtimes for various types of servers; based on the number of servers and the number of downtimes, determine the downtime rate for various types of servers.
[0011] In one embodiment, the method further includes:
[0012] Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of parts replaced and a preset time span, a parts replacement constant is determined; based on the parts replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of outages and a preset time span, an outage constant is determined; based on the outage constant and a reliability function, a third reliability value for various servers is determined.
[0013] In one embodiment, based on out-of-band hardware information of the server, the number of servers and the number of alarms generated are determined, including:
[0014] Based on the out-of-band hardware information of the servers, the number of servers and the initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server commissioning and change periods, and the number of alarms generated for various types of servers is obtained.
[0015] In one embodiment, based on out-of-band server change information, the number of servers and the number of parts to be replaced for various types of servers are determined, including:
[0016] Based on out-of-band server change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server commissioning period and change period, thus obtaining the component replacement quantity for various types of servers.
[0017] In one embodiment, based on server failure information, the number of servers and the number of downtime servers are determined, including:
[0018] Based on server fault information, the number of servers and initial fault data for various types of servers are determined; the initial fault data is cleaned to remove fault data generated during the server commissioning and change periods, thus obtaining the number of downtime servers for various types of servers.
[0019] Secondly, this application also provides a server hardware quality assessment device, comprising:
[0020] The acquisition module is used to acquire out-of-band hardware information, out-of-band change information, and fault information of various servers.
[0021] The determination module is used to determine the alarm generation rate, component replacement rate, and downtime rate of various servers based on out-of-band hardware information, out-of-band change information, and server fault information.
[0022] The evaluation module is used to obtain the average service life of various servers. Based on the average service life, alarm generation rate, parts replacement rate, downtime rate, and comprehensive quality evaluation formula, the comprehensive quality score of various servers is determined.
[0023] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0024] Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers;
[0025] Based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined.
[0026] The average service life of various servers is obtained. Based on the average service life, alarm generation rate, component replacement rate, downtime rate, and comprehensive quality assessment formula, the comprehensive quality score of various servers is determined.
[0027] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0028] Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers;
[0029] Based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined.
[0030] The average service life of various servers is obtained. Based on the average service life, alarm generation rate, component replacement rate, downtime rate, and comprehensive quality assessment formula, the comprehensive quality score of various servers is determined.
[0031] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0032] Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers;
[0033] Based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined.
[0034] The average service life of various servers is obtained. Based on the average service life, alarm generation rate, component replacement rate, downtime rate, and comprehensive quality assessment formula, the comprehensive quality score of various servers is determined.
[0035] The aforementioned server hardware quality assessment methods, apparatus, computer equipment, computer-readable storage media, and computer program products acquire out-of-band hardware information, out-of-band change information, and server fault information for various servers. Based on this information, they determine the alarm generation rate, component replacement rate, and downtime rate for various servers. They also acquire the average service life of various servers and, based on the average service life, alarm generation rate, component replacement rate, downtime rate, and a comprehensive quality assessment formula, determine the overall quality score for various servers. This application comprehensively assesses the quality of various servers by considering average service life, alarm generation rate, component replacement rate, and downtime rate, enabling a comprehensive evaluation of server quality from multiple dimensions. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart illustrating a server hardware quality assessment method in one embodiment;
[0038] Figure 2 This is a detailed flowchart of a server hardware quality assessment method in one embodiment;
[0039] Figure 3 This is a structural block diagram of a server hardware quality assessment device in one embodiment;
[0040] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0042] In one embodiment, such as Figure 1 As shown, a server hardware quality assessment method is provided. This embodiment illustrates the method by applying it to a terminal. It is understood that this method can also be applied to servers, and further to systems including both terminals and servers, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0043] Step 102: Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers.
[0044] Optionally, various servers of different brands and models can be used; information can be obtained through either the IPMP (Intelligent Platform Management Interface Protocol) or Redfish protocol. IPMP is a standard protocol for out-of-band server management, allowing users to remotely monitor and manage server hardware independently of the server's operating system. It enables hardware status monitoring (such as temperature and voltage), fault alarms, power control, and log collection. Even when the server is powered off or experiencing operating system failure, hardware information can still be accessed through this protocol, providing fundamental support for remote server operation and maintenance and fault diagnosis. Redfish is a modern server management protocol used to obtain server hardware configuration (such as processor, memory, and disk information), monitor hardware health, manage firmware updates, and collect hardware logs and alarm information. It supports large-scale, automated server hardware management and is suitable for efficient operation and maintenance scenarios in large-scale data centers.
[0045] Step 104: Based on server out-of-band hardware information, server out-of-band change information, and server fault information, determine the alarm generation rate, component replacement rate, and downtime rate of various servers.
[0046] Optional out-of-band hardware information includes the statistical period (time span, e.g., once a month), server brand, server model, number of servers, and initial alarm data. Out-of-band server change information includes the statistical period, server brand, server model, number of servers, and initial component replacement data. Server fault information includes the statistical period, server brand, server model, number of servers, and initial fault data.
[0047] Step 106: Obtain the average service life of various servers. Based on the average service life, alarm generation rate, parts replacement rate, downtime rate, and comprehensive quality assessment formula, determine the comprehensive quality score of various servers.
[0048] Optionally, the comprehensive quality assessment formula can be:
[0049]
[0050] in, It is a comprehensive quality score. It is the average years of service factor. It is the average number of years of work experience. It is an alarm-generating factor. It is the alarm generation rate. It is a parts replacement factor. It is the parts replacement rate. It is a downtime factor. It refers to downtime rate. A comprehensive quality score can be used to assess the quality and performance of each server.
[0051] The aforementioned server hardware quality assessment method acquires out-of-band hardware information, out-of-band change information, and server fault information for various servers. Based on this information, it determines the alarm generation rate, component replacement rate, and downtime rate for each server type. It also acquires the average service life of the servers and, based on this average service life, alarm generation rate, component replacement rate, downtime rate, and a comprehensive quality assessment formula, determines the overall quality score for each server type. This application comprehensively assesses the quality of various servers by considering average service life, alarm generation rate, component replacement rate, and downtime rate, enabling a comprehensive evaluation of server quality from multiple dimensions.
[0052] In one exemplary embodiment, based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined, including:
[0053] Based on out-of-band hardware information, determine the number of servers and alarms generated for various types of servers; based on the number of servers and alarms generated, determine the alarm generation rate for various types of servers; based on out-of-band change information, determine the number of servers and parts replaced for various types of servers; based on the number of servers and parts replaced, determine the parts replacement rate for various types of servers; based on server fault information, determine the number of servers and the number of downtimes for various types of servers; based on the number of servers and the number of downtimes, determine the downtime rate for various types of servers.
[0054] For example, for each type of server, the number of servers and the number of alarms generated for that type are obtained from the out-of-band hardware information. The number of alarms generated is divided by the number of servers to obtain the alarm generation rate for each type of server. The number of alarms generated for all types of servers is divided by the total number of servers of all types to obtain the alarm generation standard value. Servers of the corresponding type with an alarm generation rate higher than the alarm generation standard value and a number greater than 500 are classified as servers with poor alarm generation rate. Furthermore, for each type of server, the number of servers and the number of parts replaced are obtained from the out-of-band change information. The number of parts replaced is divided by the number of servers to obtain the parts replacement rate for each type of server. The number of parts replaced for all types of servers is divided by the total number of servers of all types to obtain the parts replacement standard value. Servers of the corresponding type with a parts replacement rate higher than the parts replacement standard value and a number greater than 500 are classified as servers with poor parts replacement rate. Finally, for each type of server, obtain the number of servers and the number of downtimes for that type from the server failure information; divide the number of downtimes by the number of servers to obtain the downtime rate for each type of server; divide the number of downtimes for all types of servers by the total number of servers for all types to obtain the downtime standard value; and classify the corresponding type of servers with downtime rates higher than the downtime standard value and a number of servers greater than 500 as servers with poor downtime quality.
[0055] In this embodiment, by determining the alarm generation rate, component replacement rate, and downtime rate of each type of server, a quantitative evaluation model can be constructed from three core dimensions: hardware reliability, failure frequency, and maintenance cost, to comprehensively reflect the hardware quality performance of the server in actual operation.
[0056] In one exemplary embodiment, the method further includes:
[0057] Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of parts replaced and a preset time span, a parts replacement constant is determined; based on the parts replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of outages and a preset time span, an outage constant is determined; based on the outage constant and a reliability function, a third reliability value for various servers is determined.
[0058] For example, the reliability function is:
[0059]
[0060] in, It is a reliability value (e.g., the first, second, or third reliability value). It is a preset time span. It is a constant (e.g., alarm occurrence constant, parts replacement constant, or system crash constant); Yes, the calculation formula is as follows:
[0061]
[0062] in, This represents the quantity (e.g., number of alarms, number of parts replaced, or number of downtimes). The above formula is used to calculate the first, second, and third reliability values for various servers. A higher reliability value indicates a less reliable server. Based on the reliability values, some unreliable servers can be removed, for example, 20%, before calculating the overall quality score for the remaining servers.
[0063] In this embodiment, by calculating the first, second, and third reliability values of various servers, the hardware stability and risk resistance capabilities of the servers can be quantified from multiple dimensions. These three values work together to form a comprehensive reliability assessment system, which not only allows for horizontal comparison of hardware quality differences between different server models, but also provides quantitative evidence for predicting potential server failure risks, optimizing hardware selection strategies, and developing targeted operation and maintenance solutions, thereby improving the overall reliability management efficiency of large-scale server clusters.
[0064] In one exemplary embodiment, the number of servers and the number of alarms generated are determined based on out-of-band server hardware information, including:
[0065] Based on the out-of-band hardware information of the servers, the number of servers and the initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server commissioning and change periods, and the number of alarms generated for various types of servers is obtained.
[0066] For example, suppose initial data for three types of servers in a data center is collected through out-of-band hardware information: Model A: 500 units deployed, 120 initial alarm data (30 debugging alarms generated during the production phase and 15 alarms caused by hardware upgrades during the change phase); Model B: 800 units deployed, 200 initial alarm data (50 initialization alarms during the production phase and 20 configuration adjustment alarms during the change phase); Model C: 600 units deployed, 80 initial alarm data (10 self-test alarms during the production phase and 5 firmware update alarms during the change phase). After cleaning the initial alarm data: the number of alarms generated for model A = 120 - 30 - 15 = 75; the number of alarms generated for model B = 200 - 50 - 20 = 130; and the number of alarms generated for model C = 80 - 10 - 5 = 65. This determines the number of servers and valid alarms generated for model A (500 units, 75 alarms), model B (800 units, 130 alarms), and model C (600 units, 65 alarms), providing foundational data for subsequent calculations of alarm generation rate and first reliability value.
[0067] In this embodiment, by performing targeted cleaning on the initial alarm data (removing non-fault alarms during the production and change periods), effective alarm information reflecting the actual faults or abnormal states of the server hardware can be accurately extracted, avoiding interference data caused by normal operations such as debugging and upgrading from affecting the evaluation results.
[0068] In one exemplary embodiment, based on out-of-band server change information, the number of servers and the number of parts to be replaced for various types of servers are determined, including:
[0069] Based on out-of-band server change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server commissioning period and change period, thus obtaining the component replacement quantity for various types of servers.
[0070] For example, suppose a data center collects initial component replacement data for three server models through out-of-band change information: Model A: 500 units deployed, 90 initial component replacements (20 during production phase for hardware compatibility testing, 30 during phase 20 for upgrades due to bulk expansion); Model B: 800 units deployed, 160 initial component replacements (15 during production phase for installation errors, 45 during phase 20 for architecture adjustments); Model C: 600 units deployed, 70 initial component replacements (8 during production phase for equipment debugging, 12 during phase 20 for firmware adaptation). After cleaning the initial data: Component replacement count for Model A = 90 - 20 - 30 = 40 times; Component replacement count for Model B = 160 - 15 - 45 = 100 times; Component replacement count for Model C = 70 - 8 - 12 = 50 times. This determines the number of servers and the number of effective parts replaced for Model A (500 units, 40 times), Model B (800 units, 100 times), and Model C (600 units, 50 times), providing basic data for subsequent calculations of parts replacement rate and second reliability value.
[0071] In this embodiment, by collecting and cleaning out-of-band change information of the server, the number of various servers and the effective number of parts to be replaced are determined. This provides accurate data support for subsequent parts replacement rate calculation and second reliability value assessment, eliminates interference from non-faulty replacements during the production and change periods, and ensures that the assessment of the reliability of server hardware components is more in line with the actual operating scenario.
[0072] In one exemplary embodiment, based on server failure information, the number of servers and the number of downtime servers for various types of servers are determined, including:
[0073] Based on server fault information, the number of servers and initial fault data for various types of servers are determined; the initial fault data is cleaned to remove fault data generated during the server commissioning and change periods, thus obtaining the number of downtime servers for various types of servers.
[0074] For example, suppose a data center collects initial fault data for three server models through server fault information: Model A: 500 units deployed, 60 initial fault records (25 due to system initialization restarts during production and 10 due to hardware configuration adjustments during the change period); Model B: 800 units deployed, 100 initial fault records (30 due to cluster testing during production and 20 due to network reconstruction during the change period); Model C: 600 units deployed, 40 initial fault records (5 due to power supply debugging restarts during production and 5 due to software upgrades during the change period). After cleaning the initial data: the number of downtime records for Model A = 60 - 25 - 10 = 25; the number of downtime records for Model B = 100 - 30 - 20 = 50; the number of downtime records for Model C = 40 - 5 - 5 = 30. This determines the number of servers and the number of effective downtimes for Model A (500 units, 25 lines), Model B (800 units, 50 lines), and Model C (600 units, 30 lines), providing basic data for subsequent calculations of downtime rate and third reliability value.
[0075] In this embodiment, by filtering out the number of valid downtimes based on server fault information, it is possible to avoid misjudging downtimes caused by normal operations such as commissioning and configuration adjustments during the change period as fault downtimes. This allows the subsequently calculated downtime rate and third-party reliability value to more accurately reflect the risk of business interruption caused by server hardware failures, providing an objective basis for server hardware quality assessment.
[0076] In one exemplary embodiment, such as Figure 2As shown, a server hardware quality assessment method includes: acquiring out-of-band hardware information, out-of-band change information, and server fault information for various servers. For each type of server, based on the out-of-band hardware information, the number of servers and initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server's production and change periods, resulting in the alarm generation count for various types of servers. The alarm generation count is divided by the number of servers to obtain the alarm generation rate for each type of server; the alarm generation count for all types of servers is divided by the total number of servers of all types to obtain the alarm generation standard value; servers of the corresponding type with an alarm generation rate higher than the alarm generation standard value and a number greater than 500 are classified as servers with poor alarm generation rate. Furthermore, for each type of server, based on the out-of-band change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server's production and change periods, resulting in the component replacement count for various types of servers. The replacement rate for each server type is calculated by dividing the number of replaced parts by the number of servers. The replacement rate for all server types is then divided by the total number of servers of all types to obtain the standard replacement value. Server types with replacement rates exceeding the standard value and a number greater than 500 are classified as servers with poor replacement rate quality. Finally, for each server type, based on server fault information, the number of servers and initial fault data for each type are determined. This initial fault data is then cleaned to remove fault data generated during server deployment and change periods, yielding the number of downtimes for each type of server. The downtime rate for each server type is then divided by the number of servers. The downtime rate for all server types is then divided by the total number of servers of all types to obtain the standard downtime value. Server types with downtime rates exceeding the standard value and a number greater than 500 are classified as servers with poor downtime rate quality. Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of component replacements and a preset time span, a component replacement constant is determined; based on the component replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of server outages and a preset time span, an outage constant is determined; and based on the outage constant and a reliability function, a third reliability value for various servers is determined. The reliability function is:
[0077]
[0078] in, It is a reliability value (e.g., the first, second, or third reliability value). It is a preset time span. It is a constant (e.g., alarm occurrence constant, parts replacement constant, or system crash constant); Yes, the calculation formula is as follows:
[0079]
[0080] in, Representing quantities (e.g., number of alarms, number of parts replaced, or number of downtimes). The first, second, and third reliability values for various servers are calculated using the formula above. A higher reliability value indicates a less reliable server. Based on these reliability values, a subset of unreliable servers can be removed, for example, 20%, before calculating the overall quality score for the remaining servers. The average service life of various servers is obtained. Based on the average service life, alarm occurrence rate, parts replacement rate, downtime rate, and the overall quality assessment formula, the overall quality score for various servers is determined. The overall quality assessment formula can be:
[0081]
[0082] in, It is a comprehensive quality score. It is the average years of service factor. It is the average number of years of work experience. It is an alarm-generating factor. It is the alarm generation rate. It is a parts replacement factor. It is the parts replacement rate. It is a downtime factor. It refers to downtime rate. A comprehensive quality score can be used to assess the quality and performance of each server.
[0083] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0084] In one exemplary embodiment, such as Figure 3 As shown, a server hardware quality assessment device is provided, including: an acquisition module 301, a determination module 302, and an assessment module 303, wherein:
[0085] The acquisition module is used to acquire out-of-band hardware information, out-of-band change information, and fault information of various servers.
[0086] The determination module is used to determine the alarm generation rate, component replacement rate, and downtime rate of various servers based on out-of-band hardware information, out-of-band change information, and server fault information.
[0087] The evaluation module is used to obtain the average service life of various servers. Based on the average service life, alarm generation rate, parts replacement rate, downtime rate, and comprehensive quality evaluation formula, the comprehensive quality score of various servers is determined.
[0088] In one exemplary embodiment, the determining module is further configured to:
[0089] Based on out-of-band hardware information, determine the number of servers and alarms generated for various types of servers; based on the number of servers and alarms generated, determine the alarm generation rate for various types of servers; based on out-of-band change information, determine the number of servers and parts replaced for various types of servers; based on the number of servers and parts replaced, determine the parts replacement rate for various types of servers; based on server fault information, determine the number of servers and the number of downtimes for various types of servers; based on the number of servers and the number of downtimes, determine the downtime rate for various types of servers.
[0090] In one exemplary embodiment, the evaluation module is further configured to:
[0091] Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of parts replaced and a preset time span, a parts replacement constant is determined; based on the parts replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of outages and a preset time span, an outage constant is determined; based on the outage constant and a reliability function, a third reliability value for various servers is determined.
[0092] In one exemplary embodiment, the determining module is further configured to:
[0093] Based on the out-of-band hardware information of the servers, the number of servers and the initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server commissioning and change periods, and the number of alarms generated for various types of servers is obtained.
[0094] In one exemplary embodiment, the determining module is further configured to:
[0095] Based on out-of-band server change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server commissioning period and change period, thus obtaining the component replacement quantity for various types of servers.
[0096] In one exemplary embodiment, the determining module is further configured to:
[0097] Based on server fault information, the number of servers and initial fault data for various types of servers are determined; the initial fault data is cleaned to remove fault data generated during the server commissioning and change periods, thus obtaining the number of downtime servers for various types of servers.
[0098] Each module in the aforementioned server hardware quality assessment device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0099] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores comprehensive quality scores. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a server hardware quality evaluation method.
[0100] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0101] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0102] Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers;
[0103] Based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined.
[0104] The average service life of various servers is obtained. Based on the average service life, alarm generation rate, component replacement rate, downtime rate, and comprehensive quality assessment formula, the comprehensive quality score of various servers is determined.
[0105] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0106] Based on out-of-band hardware information, determine the number of servers and alarms generated for various types of servers; based on the number of servers and alarms generated, determine the alarm generation rate for various types of servers; based on out-of-band change information, determine the number of servers and parts replaced for various types of servers; based on the number of servers and parts replaced, determine the parts replacement rate for various types of servers; based on server fault information, determine the number of servers and the number of downtimes for various types of servers; based on the number of servers and the number of downtimes, determine the downtime rate for various types of servers.
[0107] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0108] Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of parts replaced and a preset time span, a parts replacement constant is determined; based on the parts replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of outages and a preset time span, an outage constant is determined; based on the outage constant and a reliability function, a third reliability value for various servers is determined.
[0109] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0110] Based on the out-of-band hardware information of the servers, the number of servers and the initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server commissioning and change periods, and the number of alarms generated for various types of servers is obtained.
[0111] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0112] Based on out-of-band server change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server commissioning period and change period, thus obtaining the component replacement quantity for various types of servers.
[0113] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0114] Based on server fault information, the number of servers and initial fault data for various types of servers are determined; the initial fault data is cleaned to remove fault data generated during the server commissioning and change periods, thus obtaining the number of downtime servers for various types of servers.
[0115] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0116] Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers;
[0117] Based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined.
[0118] The average service life of various servers is obtained. Based on the average service life, alarm generation rate, component replacement rate, downtime rate, and comprehensive quality assessment formula, the comprehensive quality score of various servers is determined.
[0119] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0120] Based on out-of-band hardware information, determine the number of servers and alarms generated for various types of servers; based on the number of servers and alarms generated, determine the alarm generation rate for various types of servers; based on out-of-band change information, determine the number of servers and parts replaced for various types of servers; based on the number of servers and parts replaced, determine the parts replacement rate for various types of servers; based on server fault information, determine the number of servers and the number of downtimes for various types of servers; based on the number of servers and the number of downtimes, determine the downtime rate for various types of servers.
[0121] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0122] Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of parts replaced and a preset time span, a parts replacement constant is determined; based on the parts replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of outages and a preset time span, an outage constant is determined; based on the outage constant and a reliability function, a third reliability value for various servers is determined.
[0123] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0124] Based on the out-of-band hardware information of the servers, the number of servers and the initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server commissioning and change periods, and the number of alarms generated for various types of servers is obtained.
[0125] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0126] Based on out-of-band server change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server commissioning period and change period, thus obtaining the component replacement quantity for various types of servers.
[0127] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0128] Based on server fault information, the number of servers and initial fault data for various types of servers are determined; the initial fault data is cleaned to remove fault data generated during the server commissioning and change periods, thus obtaining the number of downtime servers for various types of servers.
[0129] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0130] Obtain out-of-band hardware information, out-of-band change information, and server fault information for various servers;
[0131] Based on server out-of-band hardware information, server out-of-band change information, and server fault information, the alarm generation rate, component replacement rate, and downtime rate of various servers are determined.
[0132] The average service life of various servers is obtained. Based on the average service life, alarm generation rate, component replacement rate, downtime rate, and comprehensive quality assessment formula, the comprehensive quality score of various servers is determined.
[0133] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0134] Based on out-of-band hardware information, determine the number of servers and alarms generated for various types of servers; based on the number of servers and alarms generated, determine the alarm generation rate for various types of servers; based on out-of-band change information, determine the number of servers and parts replaced for various types of servers; based on the number of servers and parts replaced, determine the parts replacement rate for various types of servers; based on server fault information, determine the number of servers and the number of downtimes for various types of servers; based on the number of servers and the number of downtimes, determine the downtime rate for various types of servers.
[0135] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0136] Based on the number of alarm occurrences and a preset time span, an alarm occurrence constant is determined; based on the alarm occurrence constant and a reliability function, a first reliability value for various servers is determined; based on the number of parts replaced and a preset time span, a parts replacement constant is determined; based on the parts replacement constant and a reliability function, a second reliability value for various servers is determined; based on the number of outages and a preset time span, an outage constant is determined; based on the outage constant and a reliability function, a third reliability value for various servers is determined.
[0137] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0138] Based on the out-of-band hardware information of the servers, the number of servers and the initial alarm data for various types of servers are determined; the initial alarm data is cleaned to remove alarm data generated during the server commissioning and change periods, and the number of alarms generated for various types of servers is obtained.
[0139] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0140] Based on out-of-band server change information, the number of servers and initial component replacement data for various types of servers are determined; the initial component replacement data is cleaned to remove component replacement data generated during the server commissioning period and change period, thus obtaining the component replacement quantity for various types of servers.
[0141] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0142] Based on server fault information, the number of servers and initial fault data for various types of servers are determined; the initial fault data is cleaned to remove fault data generated during the server commissioning and change periods, thus obtaining the number of downtime servers for various types of servers.
[0143] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0145] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A server hardware quality evaluation method, characterized by, The method comprises: obtaining server out-of-band hardware information, server out-of-band change information and server failure information of a plurality of servers; determining alarm generation rates, accessory replacement rates and downtime rates of the plurality of servers based on the server out-of-band hardware information, the server out-of-band change information and the server failure information; obtaining average service life of the plurality of servers, and determining comprehensive quality scores of the plurality of servers based on the average service life, the alarm generation rates, the accessory replacement rates, the downtime rates and a comprehensive quality evaluation formula.
2. The method of claim 1, wherein, The determination of the alarm generation rates, the accessory replacement rates and the downtime rates of the plurality of servers based on the server out-of-band hardware information, the server out-of-band change information and the server failure information comprises: determining server quantities and alarm generation quantities of the plurality of servers based on the server out-of-band hardware information, and determining the alarm generation rates of the plurality of servers based on the server quantities and the alarm generation quantities; determining server quantities and accessory replacement quantities of the plurality of servers based on the server out-of-band change information, and determining the accessory replacement rates of the plurality of servers based on the server quantities and the accessory replacement quantities; determining server quantities and downtime quantities of the plurality of servers based on the server failure information, and determining the downtime rates of the plurality of servers based on the server quantities and the downtime quantities.
3. The method of claim 2, wherein, The method further comprises: determining alarm occurrence constants based on the alarm occurrence quantities and a preset time span, and determining first reliability values of the plurality of servers based on the alarm occurrence constants and a reliability function; determining accessory replacement constants based on the accessory replacement quantities and the preset time span, and determining second reliability values of the plurality of servers based on the accessory replacement constants and the reliability function; determining downtime constants based on the downtime quantities and the preset time span, and determining third reliability values of the plurality of servers based on the downtime constants and the reliability function.
4. The method of claim 2, wherein, The determination of the server quantities and the alarm generation quantities of the plurality of servers based on the server out-of-band hardware information comprises: determining server quantities and initial alarm data of the plurality of servers based on the server out-of-band hardware information; cleaning the initial alarm data to remove alarm data generated during a server commissioning period and a server change period, and obtaining alarm generation quantities of the plurality of servers.
5. The method of claim 2, wherein, The determination of the server quantities and the accessory replacement quantities of the plurality of servers based on the server out-of-band change information comprises: determining server quantities and initial accessory replacement data of the plurality of servers based on the server out-of-band change information; cleaning the initial accessory replacement data to remove accessory replacement data generated during a server commissioning period and a server change period, and obtaining accessory replacement quantities of the plurality of servers.
6. The method of claim 2, wherein, The determination of the server quantities and the downtime quantities of the plurality of servers based on the server failure information comprises: determining server quantities and initial failure data of the plurality of servers based on the server failure information; cleaning the initial failure data to remove failure data generated during a server commissioning period and a server change period, and obtaining downtime quantities of the plurality of servers.
7. A server hardware quality evaluation apparatus, characterized by comprising: The device comprises: An acquisition module is configured to acquire server out-of-band hardware information, server out-of-band change information and server failure information of a plurality of servers; A determination module is configured to determine an alarm generation rate, a spare part replacement rate and a downtime rate of the plurality of servers based on the server out-of-band hardware information, the server out-of-band change information and the server failure information; An evaluation module is configured to acquire an average service life of the plurality of servers, and determine a comprehensive quality score of the plurality of servers based on the average service life, the alarm generation rate, the spare part replacement rate, the downtime rate and a comprehensive quality evaluation formula.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.