Arbitration methods, systems, and electronic devices

By calculating the health score and business availability score of the controller, the arbitration device selects the controller with the highest score as the target, which solves the problem of low business continuity in the existing technology and improves the stability and reliability of the system.

CN120950428BActive Publication Date: 2026-01-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511468691.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-01-27
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

In existing technologies, when arbitration mechanisms are implemented through third-party arbitration media, the winning node may be unable to undertake business, resulting in low business continuity. Furthermore, the failure to quantify the health status of the controller may lead to business interruption.

Method used

By acquiring hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data of the controller, a health score and a service availability score are calculated. Arbitration is then conducted based on these scores, and the arbitration device selects the controller with the highest score as the target controller.

Benefits of technology

It enables a comprehensive evaluation of controller performance and service availability, avoiding interruptions caused by nodes being unable to handle services, and improving system stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950428B_ABST
    Figure CN120950428B_ABST
Patent Text Reader

Abstract

The application discloses an arbitration method, system and electronic device, and relates to the technical field of computers, and comprises the following steps: a controller calculates a health score of the controller based on hardware index data, software index data and service capability data, calculates a service availability score of the controller based on host path connection data, host path reachability data and controller service type data; the score of the controller is calculated based on the health score and the service availability score; the score is sent to an arbitration device through an arbitration request; and the arbitration device selects the controller corresponding to the maximum score from the multiple scores to successfully preempt. The technical problem of low service continuity is solved, the technical effect of guaranteeing service continuity, improving system stability and reliability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly to an arbitration method, system, and electronic device. Background Technology

[0002] In related technologies, arbitration mechanisms are crucial for ensuring data consistency and high availability in storage systems. Currently, arbitration mechanisms are typically implemented through a third-party arbitration medium. For example, a third-party arbitration mechanism usually uses a dedicated storage device as the arbitration medium, with nodes competing for write locks to determine control. While this approach achieves arbitration, if the winning node is unable to handle the business, it will impact business continuity. Summary of the Invention

[0003] This application provides an arbitration method, system, and electronic device. Its main purpose is to address the problem of low business continuity.

[0004] According to a first aspect of this application, an arbitration method is provided, applied to a controller, the method comprising:

[0005] Acquire the controller's hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data;

[0006] The health score of the controller is determined based on the hardware indicator data, software indicator data, and service capability data; the service availability score of the controller is determined based on the host path connection data, host path reachability data, and controller service type data.

[0007] The controller's score is determined based on the health score and the business availability score;

[0008] The score is sent to the arbitration device via an arbitration request, so that the arbitration device can arbitrate based on the controller's score.

[0009] According to a second aspect of this application, an arbitration method is provided, applied to an arbitration device, the method comprising:

[0010] Receive scores from multiple controllers;

[0011] Select the controller with the highest score from among multiple scores as the target controller;

[0012] The target controller is identified as the controller that successfully preempted the hostage.

[0013] According to a third aspect of this application, an arbitration apparatus is provided, comprising:

[0014] The data acquisition module is used to acquire the hardware indicator data, software indicator data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller.

[0015] The first score calculation module is used to determine the health score of the controller based on the hardware indicator data, software indicator data, and service capability data, and to determine the service availability score of the controller based on the host path connection data, host path reachability data, and controller service type data.

[0016] The second score calculation module is used to determine the score of the controller based on the health score and the business availability score;

[0017] An arbitration module is used to send the score to an arbitration device via an arbitration request, so that the arbitration device can arbitrate based on the controller's score.

[0018] According to a fourth aspect of this application, an arbitration apparatus is provided, comprising:

[0019] The score receiving module is used to receive scores sent by multiple controllers;

[0020] The selection module is used to select the controller corresponding to the maximum score among multiple scores, and use it as the target controller.

[0021] The controller arbitration module is used to determine the target controller as the controller that has successfully preempted the controller.

[0022] According to a fifth aspect of this application, an electronic device is provided, comprising:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0026] According to a sixth aspect of this application, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0027] According to a seventh aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0028] In the embodiments of this application, hardware performance data, software performance data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller are acquired. Based on the hardware performance data, software performance data, and service capability data, a health score of the controller is determined; based on the host path connection data, host path reachability data, and controller service type data, a service availability score of the controller is determined; based on the health score and the service availability score, a score of the controller is determined; the score is sent to an arbitration device through an arbitration request; the arbitration device selects the controller with the highest score from multiple scores as the target controller; and the target controller is determined as the controller that successfully preempted the arbitration. Thus, during arbitration, the controller can calculate its own score based on its own hardware performance data, software performance data, service capability data, and other performance data, as well as host path connection data, host path reachability data, and controller service type data, as well as other service availability data, so that the arbitration device can conduct controller arbitration based on the score. In this way, the performance and service availability of the controller can be comprehensively evaluated through a score system, and dynamic arbitration can be carried out based on the score. This not only avoids situations where the winning node is unable to handle the service, resulting in service interruption, but also improves system stability and reliability. Attached Figure Description

[0029] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A schematic diagram illustrating the principle of a third-party arbitration mechanism provided in related technologies;

[0031] Figure 2 A flowchart illustrating an arbitration method applied to a controller, provided as an embodiment of this application;

[0032] Figure 3 A schematic flowchart illustrating an arbitration method applied to an arbitration device, provided as an embodiment of this application;

[0033] Figure 4 A flowchart illustrating an arbitration method provided in an embodiment of this application;

[0034] Figure 5 This is a schematic diagram of an arbitration device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0036] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0037] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] The specific application environment architecture or specific hardware architecture on which the execution of the arbitration method depends is described here.

[0039] As the background technology shows, in third-party arbitration schemes, control is primarily determined by nodes competing for write locks during the arbitration decision-making process. For example, see [link to relevant documentation]. Figure 1 Taking a dual-active storage scenario as an example, each site has two controllers (Controller A and Controller B), and there is a connection between Site 1 and Site 2. When the connection between Site 1 and Site 2 is lost at some point, Site 1 and Site 2 will each select a controller to compete for arbitration. As an example, the arbitration process can be as follows: both the controllers selected by Site 1 and Site 2 send arbitration requests to the arbitration device. The controller that sends its request first will have priority in writing information to the arbitration device and will successfully preempt arbitration; the controller that sends its request later will fail to preempt arbitration. In this process, the arbitration decision is not decoupled from business requirements. For example, the winning node may be disconnected from the host and unable to handle business, thus compromising business continuity. Furthermore, this arbitration strategy does not quantify the health status of the controllers. If the winning node itself is faulty, a failure may be triggered after successful arbitration, leading to business interruption.

[0040] Based on this, this application provides an arbitration method. When a controller sends an arbitration request to an arbitration device, it needs to calculate its own score and then include the score in the arbitration request and send it to the arbitration device. The arbitration device compares the scores of each controller and successfully arbitrates the controller with the highest score.

[0041] The arbitration methods, systems, and electronic devices of embodiments of this application are described below with reference to the accompanying drawings.

[0042] Figure 2 This is a flowchart illustrating an arbitration method provided in an embodiment of this application, which can be applied to a controller. Figure 2 As shown, the method includes the following steps:

[0043] Step 201: Obtain the controller's hardware metrics data, software metrics data, service capability data, host path connection data, host path reachability data, and controller service type data.

[0044] During arbitration, the controller can obtain its own performance data, such as hardware metrics, software metrics, and service capability data, as well as its own service availability data, such as host path connection data, host path reachability data, and controller service type data. Understandably, multiple controllers typically participate in the arbitration process in a storage system. Each participating controller can execute this process; that is, each controller can obtain its own hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data.

[0045] Step 202: Determine the controller's health score based on hardware indicator data, software indicator data, and service capability data; determine the controller's service availability score based on host path connection data, host path reachability data, and controller service type data.

[0046] After acquiring hardware metric data, software metric data, service capability data, host path connection data, host path reachability data, and controller service type data, the controller's own score can be calculated based on this data. For example, for a given controller, it can calculate its own health score based on its acquired hardware metric data, software metric data, and service capability data. It can also calculate its own service availability score based on its acquired host path connection data, host path reachability data, and controller service type data. Understandably, the calculation of the health score and the service availability score can be performed sequentially or in parallel.

[0047] Step 203: Determine the controller's score based on the health score and business availability score.

[0048] After calculating its own health score and service availability score, the controller can determine its overall score based on these two scores. For example, the health score and service availability score can be aggregated, such as by calculating their sum, as the controller's overall score. Understandably, other methods can also be used to determine the controller's score based on the health score and service availability score, such as weighted average, product, taking the maximum value, or taking the minimum value.

[0049] Step 204: Send the score to the arbitration device through the arbitration request so that the arbitration device can arbitrate based on the score of the controller.

[0050] After determining its own score, the controller can generate an arbitration request based on that score. This arbitration request carries the controller's score, and is then sent to the arbitration device. The arbitration device receives the score from the controllers and can then perform arbitration processing based on each controller's score. For example, among all the received controller scores, the controller with the highest score can be selected, and the controller with the highest score can be determined to have successfully preempted the arbitration.

[0051] In the embodiments of this application, hardware performance data, software performance data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller are acquired. Based on the hardware performance data, software performance data, and service capability data, a health score of the controller is determined; based on the host path connection data, host path reachability data, and controller service type data, a service availability score of the controller is determined; based on the health score and the service availability score, a score of the controller is determined; the score is sent to an arbitration device through an arbitration request; the arbitration device selects the controller with the highest score from multiple scores as the target controller; and the target controller is determined as the controller that successfully preempted the arbitration. Thus, during arbitration, the controller can calculate its own score based on its own hardware performance data, software performance data, service capability data, and other performance data, as well as host path connection data, host path reachability data, and controller service type data, as well as other service availability data, so that the arbitration device can conduct controller arbitration based on the score. In this way, the performance and service availability of the controller can be comprehensively evaluated through a score system, and dynamic arbitration can be carried out based on the score. This not only avoids situations where the winning node is unable to handle the service, resulting in service interruption, but also improves system stability and reliability.

[0052] It should be noted that the embodiments of this application may include multiple steps. For ease of description, these steps are numbered, but these numbers are not a limitation on the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of this application do not limit this.

[0053] Furthermore, the hardware performance data includes at least one of the following: CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data; the software performance data includes process stability data and resource efficiency data; and the service capability data includes historical failure data, successful arbitration data, and latency data.

[0054] The controller's health score is determined based on hardware metrics, software metrics, and service capability data, including:

[0055] Based on data from the central processing unit, memory, motherboard, system disk, network communication capabilities, and chassis, the hardware performance score of the controller is calculated.

[0056] Based on process stability data and resource efficiency data, calculate the software metric score of the controller;

[0057] Based on historical fault data, successful arbitration data, and latency data, calculate the controller's service capability score;

[0058] The controller's health score is calculated based on hardware metric scores, software metric scores, and service capability scores.

[0059] The hardware metrics data can include CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data. Specifically, when acquiring the controller's hardware metrics data, we can acquire CPU data (e.g., instruction retry rate, number of errors recorded in the MCE (Machine Check Exception) log), memory data (e.g., memory throttling event count, number of errors recorded in the MCE log), motherboard data (e.g., Advanced Error Reporting (AER) log count, such as the AER log count of lspci-vvvAER), system disk data (e.g., number of system startup log errors, number of SMART (Self-Monitoring, Analysis, and Reporting Technology) detection errors), network communication capability data (e.g., optical module received optical power, CRC error count), and chassis data (e.g., voltage, chassis temperature, fan speed).

[0060] Software metrics data can include process stability data and resource efficiency data, used to indicate the health status of the controller's processes and the efficiency of resource processing. That is, when obtaining software metrics data for the controller, one can obtain process stability data and resource efficiency data. Service capability data includes historical fault data, arbitration success data, and latency data. That is, when obtaining service capability data, one can obtain historical fault data (e.g., the number of times the controller failed within a preset time period (e.g., the last year, month, or day), arbitration success data (e.g., the number of times the controller successfully preempted within a preset time period (e.g., the last hour),) and latency data (e.g., the controller's average I / O (Input / Output) latency).

[0061] Accordingly, when determining the controller's health score based on hardware, software, and service capability data, the following calculations can be performed: Hardware score: based on CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data; Software score: based on process stability and resource efficiency data; Service capability score: based on historical fault data, successful arbitration data, and latency data. Then, the controller's overall health score can be calculated by summing these scores. This multi-dimensional and comprehensive approach to calculating the controller's health score allows for a more accurate reflection of the control status, providing more accurate data for subsequent arbitration and further improving the accuracy and reliability of the arbitration results.

[0062] Furthermore, based on data from the central processing unit, memory, motherboard, system disk, network communication capabilities, and chassis, the controller's hardware performance score is calculated, including:

[0063] For the i-th hardware indicator data among the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, determine whether the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data; where i∈[1,N], and N is the total number of hardware indicator data;

[0064] If the i-th hardware indicator data falls within the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data will be determined as the first preset score.

[0065] If the i-th hardware indicator data does not belong to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data will be determined as the second preset score.

[0066] The hardware metric score of the controller is calculated based on the score corresponding to each hardware metric data.

[0067] In calculating the controller's hardware performance score based on CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, it can be determined whether each hardware performance indicator falls within its corresponding preset range. If it does, a first preset score is assigned; otherwise, a second preset score is assigned. For example, for any hardware performance indicator among CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, such as the i-th hardware performance indicator, the corresponding data range can be obtained. This range can be set according to actual needs. If the i-th hardware performance indicator falls within its corresponding data range, the score corresponding to the i-th hardware performance indicator can be assigned the first preset score; conversely, if the i-th hardware performance indicator does not fall within its corresponding data range, the score corresponding to the i-th hardware performance indicator can be assigned the second preset score. As a specific example, see Table 1. Table 1 shows a scoring method with a first preset score of 5 points. Understandably, the second preset score can be 0 points, which can be understood as the specific value of the indicator falling within the set value range.

[0068] Table 1

[0069]

[0070] After obtaining the score corresponding to each hardware indicator data point, the controller's hardware indicator score can be calculated based on that score. For example, the sum of the scores for each hardware indicator data point can be calculated as the controller's hardware indicator score. This allows for the calculation of the controller's hardware indicator score, improving the accuracy of the calculation results and providing more accurate data for subsequent arbitration.

[0071] Furthermore, process stability data includes core process survival rate and the number of process restarts within the first preset period; resource efficiency data includes CPU utilization, soft interrupt ratio, and file descriptor reclamation rate.

[0072] Based on process stability data and resource efficiency data, the controller's software metric score is calculated, including:

[0073] For the j-th software indicator data among the core process survival rate, process restart count within the first preset period, CPU utilization, soft interrupt ratio, and file descriptor reclamation rate, determine whether the j-th software indicator data belongs to the data range corresponding to the j-th software indicator data; where j∈[1,M], and M is the total number of software indicator data;

[0074] If the j-th software indicator data falls within the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data will be determined as the third preset score.

[0075] If the j-th software indicator data does not belong to the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data will be determined as the fourth preset score.

[0076] The controller's software indicator score is calculated based on the score corresponding to each software indicator data.

[0077] The process stability data can include the core process survival rate and the number of process restarts within a preset period. The core process survival rate can include, for example, the survival rate of core processes detected by systemd (System Daemon) and the number of restarts within the first preset period (e.g., one week) in / proc / <pid> / oom_score process crash restart count, / proc / <pid> / oom_score, also known as the "process OOM (Out of Memory) score file," is a part of the / proc filesystem in Linux systems used to provide information about a specific process (manufactured by...). <pid>Information (identifier). Resource efficiency data can include CPU utilization, soft interrupt percentage, file descriptor reclamation rate, etc.

[0078] Accordingly, when calculating the controller's software metric score based on process stability data and resource efficiency data, for the j-th software metric data among core process survival rate, process restart count within a preset period, CPU utilization, soft interrupt ratio, and file descriptor reclamation rate, the data range corresponding to the j-th software metric data can be obtained to determine whether the j-th software metric data belongs to the data range corresponding to the j-th software metric data. If the j-th software metric data belongs to the data range corresponding to the j-th software metric data, the score corresponding to the j-th software metric data can be assigned a third preset score, such as 5 points; conversely, if the j-th software metric data does not belong to the data range corresponding to the j-th software metric data, the score corresponding to the j-th software metric data can be assigned a fourth preset score, such as 0 points. Afterwards, the controller's software metric score can be calculated using the scores corresponding to each software metric data, for example, by summing. As an example, see Table 2, which shows one scoring method with a third preset score of 5 points. It is understood that the fourth preset score can be 0 points.

[0079] Table 2

[0080]

[0081] Furthermore, based on historical fault data, successful arbitration data, and latency data, the controller's service capability score is calculated, including:

[0082] The number of controller failures within the second preset period is determined based on historical fault data.

[0083] The fault prediction score of the controller is calculated based on the number of faults, the preset reduction in score for each fault, and the preset total fault score.

[0084] The number of successful preemptions by the controller within a preset time period is determined based on the successful arbitration data;

[0085] The controller calculates the preemption bonus value based on the number of successful preemptions, the preset bonus for a single preemption, and the preset maximum preemption score.

[0086] The delay score of the controller is determined based on the preset delay range to which the delay data belongs;

[0087] The service capability score of the controller is calculated based on the fault prediction score, the number of successful preemption attempts, and the preemption bonus score.

[0088] When calculating the controller's service capability score based on historical fault data, successful arbitration data, and latency data, the number of faults in the controller within a second preset period can be determined based on historical fault data, such as the number of faults in a day, a month, or a year. Then, the preset reduction in points for each fault (preset reduction in points per fault) and the preset maximum value of the fault score (preset total fault score) can be obtained. Based on the number of faults, the preset reduction in points per fault, and the preset total fault score, the controller's fault prediction score is calculated. For an example, see Table 3 for a specific example of the fault prediction score. The number of successful arbitrations can also be used to determine the number of times the controller successfully preempts a data point within a preset time period. For example, the preset time period could be the most recent hour, meaning the number of times the controller successfully preempts a data point within the most recent hour. See Table 3 for a specific example. Furthermore, the preset latency range to which the latency data belongs can be determined. The specific range of the preset latency range can be set according to actual needs. For example, the preset latency range can be set to: less than or equal to 1ms, greater than 1ms and less than or equal to 10ms, etc. Based on the preset delay range to which the delay data belongs, the controller's delay score is assigned. See Table 3 for a specific example. Then, the controller's service capability score can be calculated using the fault prediction score, the number of successful preemption attempts, and the preemption bonus score. For example, the fault prediction score, the number of successful preemption attempts, and the preemption bonus score can be summarized to obtain the controller's service capability score.

[0089] Table 3

[0090]

[0091] Furthermore, the host path connection data includes the number of paths connecting the controller to the host and the total number of paths connecting all controllers to the host; the host path reachability data includes the number of valid volumes on the controller path and the total number of volumes configured on all controllers; and the controller service type data includes service type data and the weight score corresponding to each service type.

[0092] The service availability score of the controller is determined based on host path connection data, host path reachability data, and controller service type data, including:

[0093] The host path connection status score of the controller is calculated based on the number of paths connecting the controller to the host, the total number of paths connecting all controllers to the host, and the first decision weight score corresponding to the host path connection data.

[0094] The system is based on the number of volumes with valid controller paths, the total number of volumes configured on all controllers, the second decision weight score corresponding to the host path reachability data, and the host path reachability score based on the controller.

[0095] The controller's business weight score is calculated based on business type data, the weight corresponding to each business type, and the third decision weight score corresponding to the business type data; wherein, the weight corresponding to each business type is set based on the priority of each business type;

[0096] The controller's service availability score is calculated based on the host path connection status score, host path reachability score, and service weight score.

[0097] The host path connection data can include the number of paths connecting the controller to the host and the total number of paths connecting all controllers to the host. Based on this, when determining the service availability score of a controller based on host path connection data, host path reachability data, and controller service type data, the controller's host path connection status score can be calculated based on the number of paths connecting the controller to the host, the total number of paths connecting all controllers to the host, and the first decision weight score corresponding to the host path connection data. For example, the number of port paths connecting the controller and the host via a switch can be summed with the number of paths directly connecting the controller to the host to obtain the number of paths connecting the controller to the host; the sum of the number of paths connecting all controllers in the cluster to the host can be obtained as the total number of paths connecting all controllers to the host; the decision weight score corresponding to the host path connection data, i.e., the first decision weight score, can be obtained by multiplying a preset total decision score (e.g., 100 points) with the first decision weight corresponding to the host path connection data (e.g., 40%). For example, based on the number of paths connecting the controller to the host, the total number of paths connecting all controllers to the host, and the first decision weight score corresponding to the host path connection data, the host path connection status score of the controller can be calculated as follows: Final score = Number of host paths for this controller / Total number of paths for all hosts in the cluster configuration * Decision weight score. The final score is the host path connection status score. As a specific example, assuming a controller has 6 host paths, the total number of paths for all hosts in the cluster configuration is 10, and the first decision weight is 40%, then the host path connection status score of this controller = 6 / 10 * 40 = 24.

[0098] Host path reachability data can include the number of volumes with valid controller paths and the total number of volumes configured on all controllers. Correspondingly, based on the number of volumes with valid controller paths, the total number of volumes configured on all controllers, and the second decision weight score corresponding to the host path reachability data, and based on the controller's host path reachability score, after the host path connection status check is completed, the list of hosts mounting the volume can be checked according to the host->controller access path of each service volume. SCSI can then be executed through the host agent. The `TEST_UNIT_READY` command ("Test Device Ready" command) determines the validity of the path from the host to the volume. The decision weight score corresponding to the host path reachability data, also known as the second decision weight score, can be obtained by multiplying the preset total decision score (e.g., 100 points) by the second decision weight corresponding to the host path reachability data (e.g., 35%). As an example, the host path reachability score of a controller can be calculated as: Final Score = Number of volumes with valid paths on this controller / Total number of volumes in the cluster configuration * Decision Weight Score. Here, the final score is the host path reachability score. For a specific example, assuming a controller has 10 volumes with valid paths and a total of 20 volumes in the cluster configuration, then the host path reachability score of this controller = 10 / 20 * 35 = 17.5.

[0099] Controller service type data can include service type data and the weight score corresponding to each service type. Correspondingly, when determining the controller's service availability score based on host path connection data, host path reachability data, and controller service type data, the controller's service weight score can be calculated based on the service type data, the weight score corresponding to each service type, and the third decision weight score corresponding to the service type data. Here, the service type data can be the service type of the business to be handled, such as core database, virtualization platform, backup service, test environment, etc.; the decision weight score corresponding to the service type data, i.e., the third decision weight score, can be obtained by multiplying a preset total decision score (e.g., 100 points) by the third decision weight of the service type data (e.g., 25%); the weight corresponding to each service type can be pre-set, for example, the service weights corresponding to core database, virtualization platform, backup service, and test environment can be set to 40%, 30%, 20%, and 10%, respectively. For example, considering that different surviving controllers handle different host services, the importance of different services varies significantly for the user; for example, the importance of production services such as user databases is usually much greater than that of production services such as backup services. Therefore, a score needs to be calculated based on the types of business handled by the controller for score calculation. The final score is equal to the sum of the scores for each type of business handled by the controller. The score for each business type = the business weight corresponding to that business type * the decision weight score. As a specific example, see Table 4, which provides an exemplary table of priorities and business weights for different business types. Assuming a controller deploys core database business and test environment business and the third decision weight is 25%, then the controller's business weight score = 25 * (40% + 10%) = 12.5.

[0100] Understandably, the weight corresponding to each business type can be set based on the priority of each business type. For example, the higher the priority of a business type, the higher its weight.

[0101] Table 4

[0102]

[0103] Furthermore, acquire the controller's hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data, including:

[0104] When the connection between the first node and the second node to which the controller belongs is broken, the controller's hardware indicator data, software indicator data, service capability data, host path connection data, host path reachability data, and controller service type data are obtained; wherein, the second node is a node other than the first node.

[0105] Since arbitration is typically conducted when nodes are disconnected, it is necessary to determine whether the nodes are disconnected before executing the arbitration method provided in this embodiment. Only if all nodes of the controller (i.e., the first node) are disconnected from other nodes in the cluster (i.e., the second node) will the controller's hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data be acquired. Subsequent arbitration processing will then be performed based on these metrics.

[0106] Furthermore, the controller's score is determined based on the health score and business availability score, including at least one of the following:

[0107] The health score and business availability score are aggregated to obtain the controller's score;

[0108] The maximum value between the health score and the business availability score is selected as the controller's score;

[0109] The minimum value between the health score and the business availability score is selected as the controller's score.

[0110] When determining the controller's score based on the health score and service availability score, methods such as summarizing, selecting the maximum or minimum value can be used. For example, the health score and service availability score can be summed, such as calculating the sum of the health score and service availability score, as the controller's score. Alternatively, the health score and service availability score can be compared, and the maximum or minimum value between the two can be selected as the controller's score.

[0111] Based on the same inventive concept, an arbitration method applied to arbitration equipment is also provided, see [link to relevant documentation]. Figure 3 This can include the following processing:

[0112] Step 301: Receive scores sent by multiple controllers.

[0113] In this system, the controller can send its own score to the arbitration device via an arbitration request, so that the arbitration device can receive the controller's score. Understandably, in a cluster, there are usually two or more nodes; therefore, the arbitration device can receive scores from multiple controllers, for example, two or more controllers.

[0114] Step 302: Select the controller with the highest score from among the multiple scores as the target controller.

[0115] The arbitration device receives scores from multiple controllers, compares these scores, and selects the highest score. The controller corresponding to this highest score is then identified as the target controller.

[0116] Step 303: The target controller is identified as the controller that successfully preempted the hostage.

[0117] Once the target controller is determined, it can be designated as the controller that successfully preempts the service, and the controller that successfully preempts the service will continue to handle the business.

[0118] In the embodiments of this application, scores are received from multiple controllers; the controller with the highest score is selected as the target controller; and the target controller is determined as the controller that successfully preempts the service. In this way, the controller with the highest score is selected for successful preemption, which means that the controller with better health and service availability is selected for successful preemption. This can better avoid service interruptions and improve service continuity.

[0119] Furthermore, to make the arbitration method provided in the embodiments of this application clearer, it will be described below with reference to the following specific examples. See Figure 4 In the embodiments of this application, the arbitration preemption process is optimized: when a controller sends a request to the arbitration device, it can calculate its own score, and then include this score in the arbitration request and send it to the arbitration device. The arbitration device compares the scores of the two controllers, and the controller with the higher score successfully preempts the arbitration. For example, the process of the controller calculating its own score may include the following aspects:

[0120] 1. Dynamic health score assessment.

[0121] At the start of the arbitration evaluation, all controllers participating in the arbitration competition are initially assigned a score of 0, and scores are determined based on the dimensions mentioned below. The overall scoring algorithm includes: ① adding points based on different detection items; ② some items have degradation acceleration factors: for example, in the fault prediction section below, faults with more recent dates have higher weights; ③ critical faults are vetoed: for example, in the process stability check section below, critical processes that have experienced faults receive no points.

[0122] Dimension 1: Hardware Indicator Judgment

[0123] Hardware performance evaluation primarily involves testing the CPU, memory, motherboard, system disk, and other components that make up the storage controller, as well as the controller's network communication capabilities and the chassis that maintains its operation. Specific testing items are shown in the table below:

[0124]

[0125] Dimension Two: Software Indicator Judgment

[0126] When determining software metrics, the main focus is on detecting the process health status and resource efficiency of the storage controller system. The specific detection content and scoring are shown in the table below:

[0127]

[0128] Dimension 3: Service Capability Indicator Assessment

[0129] When determining service capability metrics, the main criteria are the number of historical failures of the storage controller, whether it has successfully preempted arbitration, and the controller's IO latency. These criteria are used to evaluate the controller's ability to provide services. The specific test content and scoring are shown in the table below:

[0130]

[0131] 2. Business availability score assessment.

[0132] Business availability assessment can be based on a "business availability priority" principle to restructure the arbitration process, focusing on the connection status between the controller and the host, as well as the business data access capabilities. This embodiment achieves deep coupling between arbitration decisions and business status through host path connectivity verification → host path reachability detection → business weight voting.

[0133] As an example, when making a decision, the total score can be set to 100 points, and then the scores can be allocated according to the decision weights (see the table below). The final score is equal to the sum of the scores of each detection item.

[0134]

[0135] Among these, host path connection status: Each controller can check the port paths connecting itself and hosts via switches, as well as the paths directly connecting to hosts, to determine whether any host paths were lost during arbitration from a physical path perspective. Then, a score can be calculated based on the host path connection retention status. For example: Final score = Number of host paths for this controller / Total number of paths for all hosts in the cluster configuration * Decision weight score. For example: If a controller has 6 host paths and the total number of paths for all hosts in the cluster configuration is 10, then the score = 6 / 10 * 40 = 24.

[0136] Host path reachability: After checking the host path connectivity status, the list of hosts mounting the volume can be checked based on the host -> controller access path for each service volume. Then, the SCSI TEST_UNIT_READY command is executed through the host agent to determine the validity of the path from the host to the volume. For example, the final score = number of volumes with valid paths on the controller / total number of volumes in the cluster configuration * decision weight score. For instance, if a controller has 10 volumes with valid paths and the cluster configuration has 20 volumes, then the score = 10 / 20 * 35 = 17.5.

[0137] Business Weight Voting: After arbitration, due to the differences in the host services handled by the surviving controllers, the importance of different services varies significantly for the customer. For example, the importance of production services such as databases is far greater than that of backup services. Therefore, a score is calculated based on the types of services handled by the controller, which is used for controller score calculation. The final score is equal to the sum of the scores for each business type handled by the controller. The score for each business type = the business weight corresponding to that business type * the decision weight score. For example, if a controller deploys core database services and a test environment, then the score = 25 * (40% + 10%) = 12.5. As a concrete example, the priority and business weight of each business type are shown in the table below:

[0138]

[0139] The controller's score, or comprehensive score, can be obtained by summing the health score and the business availability score.

[0140] In summary, the embodiments of this application can achieve dynamic weighted arbitration of the controller by evaluating the controller's health score and service availability score. This allows the controller to score its own health and service carrying capacity for subsequent competition arbitration. It is understood that the specific implementation and technical effects of each step in this embodiment are similar to those in the above-described method embodiments, and will not be repeated here.

[0141] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0142] Based on the same inventive concept, embodiments of this application also provide an arbitration system, including multiple controllers and arbitration devices, wherein:

[0143] The controller is used for:

[0144] Acquire hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data of the controller;

[0145] The health score of the controller is determined based on the hardware indicator data, software indicator data, and service capability data; the service availability score of the controller is determined based on the host path connection data, host path reachability data, and controller service type data.

[0146] The controller's score is determined based on the health score and the business availability score;

[0147] The score is sent to the arbitration device through an arbitration request;

[0148] The arbitration device is used for:

[0149] Receive scores sent by multiple controllers;

[0150] Select the controller with the highest score from among multiple scores as the target controller;

[0151] The target controller is identified as the controller that successfully preempted the hostage.

[0152] Furthermore, the hardware performance data includes at least one of the following: CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data; the software performance data includes process stability data and resource efficiency data; and the service capability data includes historical fault data, successful arbitration data, and latency data.

[0153] The controller is used for:

[0154] Based on the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, the hardware performance score of the controller is calculated.

[0155] Based on the process stability data and resource efficiency data, the software performance score of the controller is calculated.

[0156] Based on the historical fault data, successful arbitration data, and latency data, the service capability score of the controller is calculated.

[0157] The health score of the controller is calculated based on the hardware indicator score, the software indicator score, and the service capability score.

[0158] Furthermore, the controller is configured to:

[0159] For the i-th hardware indicator data among the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, determine whether the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data; where i∈[1,N], and N is the total number of hardware indicator data;

[0160] If the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data is determined as the first preset score.

[0161] If the i-th hardware indicator data does not belong to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data is determined as the second preset score.

[0162] The hardware indicator score of the controller is calculated based on the score corresponding to each of the aforementioned hardware indicator data.

[0163] Furthermore, the process stability data includes the core process survival rate and the number of process restarts within a first preset period; the resource efficiency data includes the CPU utilization rate, the proportion of soft interrupts, and the file descriptor reclamation rate.

[0164] The controller is used for:

[0165] For the j-th software indicator data among the core process survival rate, the number of process restarts within the first preset period, the CPU utilization rate, the soft interrupt ratio, and the file descriptor reclamation rate, determine whether the j-th software indicator data belongs to the data range corresponding to the j-th software indicator data; where j∈[1,M], and M is the total number of software indicator data;

[0166] If the j-th software indicator data falls within the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data is determined as the third preset score.

[0167] If the j-th software indicator data does not belong to the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data is determined as the fourth preset score.

[0168] The software indicator score of the controller is calculated based on the score corresponding to each of the software indicator data.

[0169] Furthermore, the controller is configured to:

[0170] The number of failures of the controller within a second preset period is determined based on the historical fault data.

[0171] Based on the number of faults, the preset reduction in points per fault, and the preset total fault score, the fault prediction score of the controller is calculated.

[0172] The number of successful preemptions by the controller within a preset time period is determined based on the successful arbitration data.

[0173] Based on the number of successful preemptions, the preset single preemption bonus, and the preset maximum preemption score, the preemption bonus value of the controller is calculated;

[0174] Based on the preset delay range to which the delay data belongs, the delay score of the controller is determined;

[0175] The service capability score of the controller is calculated based on the fault prediction score, the number of successful preemption attempts, and the preemption bonus score.

[0176] Furthermore, the host path connection data includes the number of paths connecting the controller to the host and the total number of paths connecting all controllers to the host; the host path reachability data includes the number of valid volumes for the controller path and the total number of volumes configured on all controllers; and the controller service type data includes service type data and the weight score corresponding to each service type.

[0177] The controller is used for:

[0178] The host path connection status score of the controller is calculated based on the number of paths between the controller and the host, the total number of paths between all controllers and the host, and the first decision weight score corresponding to the host path connection data.

[0179] Based on the number of valid volumes on the controller path, the total number of volumes configured on all controllers, the second decision weight score corresponding to the host path reachability data, and the host path reachability score of the controller;

[0180] Based on the business type data, the weight corresponding to each business type, and the third decision weight score corresponding to the business type data, the business weight score of the controller is calculated; wherein, the weight corresponding to each business type is set based on the priority of each business type;

[0181] Based on the host path connection status score, the host path reachability score, and the service weight score, the service availability score of the controller is calculated.

[0182] Furthermore, the controller is configured to:

[0183] When the connection between the first node and the second node to which the controller belongs is broken, the hardware indicator data, software indicator data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller are obtained; wherein, the second node is a node other than the first node.

[0184] Furthermore, the controller is configured to:

[0185] The controller's score is determined based on health score and business availability score, including at least one of the following:

[0186] The health score and business availability score are aggregated to obtain the controller's score;

[0187] The maximum value between the health score and the business availability score is selected as the controller's score;

[0188] The minimum value between the health score and the business availability score is selected as the controller's score.

[0189] In this embodiment, the specific implementation and technical effects of the controller and arbitration device in performing each step are similar to those in the above embodiments, and will not be repeated here.

[0190] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0191] Embodiments of this application also provide an arbitration device.

[0192] For example, Figure 5 This is a schematic diagram of an arbitration device provided in an embodiment of this application. The arbitration device 500 includes:

[0193] The data acquisition module 510 is used to acquire the hardware indicator data, software indicator data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller.

[0194] The first score calculation module 520 is used to determine the health score of the controller based on the hardware indicator data, software indicator data, and service capability data, and to determine the service availability score of the controller based on the host path connection data, host path reachability data, and controller service type data.

[0195] The second score calculation module 530 is used to determine the score of the controller based on the health score and the business availability score;

[0196] Arbitration module 540 is used to send the score to the arbitration device via an arbitration request, so that the arbitration device can arbitrate based on the score of the controller.

[0197] Furthermore, the hardware performance data includes at least one of the following: CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data; the software performance data includes process stability data and resource efficiency data; and the service capability data includes historical fault data, successful arbitration data, and latency data.

[0198] The first score calculation module 520 is used for:

[0199] Based on the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, the hardware performance score of the controller is calculated.

[0200] Based on the process stability data and resource efficiency data, the software performance score of the controller is calculated.

[0201] Based on the historical fault data, successful arbitration data, and latency data, the service capability score of the controller is calculated.

[0202] The health score of the controller is calculated based on the hardware indicator score, the software indicator score, and the service capability score.

[0203] Furthermore, the first score calculation module 520 is used for:

[0204] For the i-th hardware indicator data among the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, determine whether the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data; where i∈[1,N], and N is the total number of hardware indicator data;

[0205] If the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data is determined as the first preset score.

[0206] If the i-th hardware indicator data does not belong to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data is determined as the second preset score.

[0207] The hardware indicator score of the controller is calculated based on the score corresponding to each of the aforementioned hardware indicator data.

[0208] Furthermore, the process stability data includes the core process survival rate and the number of process restarts within a first preset period; the resource efficiency data includes the CPU utilization rate, the proportion of soft interrupts, and the file descriptor reclamation rate.

[0209] The first score calculation module 520 is used for:

[0210] For the j-th software indicator data among the core process survival rate, the number of process restarts within the first preset period, the CPU utilization rate, the soft interrupt ratio, and the file descriptor reclamation rate, determine whether the j-th software indicator data belongs to the data range corresponding to the j-th software indicator data; where j∈[1,M], and M is the total number of software indicator data;

[0211] If the j-th software indicator data falls within the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data is determined as the third preset score.

[0212] If the j-th software indicator data does not belong to the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data is determined as the fourth preset score.

[0213] The software indicator score of the controller is calculated based on the score corresponding to each of the software indicator data.

[0214] Furthermore, the first score calculation module 520 is used for:

[0215] The number of failures of the controller within a second preset period is determined based on the historical fault data.

[0216] Based on the number of faults, the preset reduction in points per fault, and the preset total fault score, the fault prediction score of the controller is calculated.

[0217] The number of successful preemptions by the controller within a preset time period is determined based on the successful arbitration data.

[0218] Based on the number of successful preemptions, the preset single preemption bonus, and the preset maximum preemption score, the preemption bonus value of the controller is calculated;

[0219] Based on the preset delay range to which the delay data belongs, the delay score of the controller is determined;

[0220] The service capability score of the controller is calculated based on the fault prediction score, the number of successful preemption attempts, and the preemption bonus score.

[0221] Furthermore, the host path connection data includes the number of paths connecting the controller to the host and the total number of paths connecting all controllers to the host; the host path reachability data includes the number of valid volumes for the controller path and the total number of volumes configured on all controllers; and the controller service type data includes service type data and the weight score corresponding to each service type.

[0222] The first score calculation module 520 is used for:

[0223] The host path connection status score of the controller is calculated based on the number of paths between the controller and the host, the total number of paths between all controllers and the host, and the first decision weight score corresponding to the host path connection data.

[0224] Based on the number of valid volumes on the controller path, the total number of volumes configured on all controllers, the second decision weight score corresponding to the host path reachability data, and the host path reachability score of the controller;

[0225] Based on the business type data, the weight corresponding to each business type, and the third decision weight score corresponding to the business type data, the business weight score of the controller is calculated; wherein, the weight corresponding to each business type is set based on the priority of each business type;

[0226] Based on the host path connection status score, the host path reachability score, and the service weight score, the service availability score of the controller is calculated.

[0227] Furthermore, the data acquisition module 510 is used for:

[0228] When the connection between the first node and the second node to which the controller belongs is broken, the hardware indicator data, software indicator data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller are obtained; wherein, the second node is a node other than the first node.

[0229] Furthermore, the second score calculation module 530 is configured to perform at least one of the following:

[0230] The health score and business availability score are aggregated to obtain the controller's score;

[0231] The maximum value between the health score and the business availability score is selected as the controller's score;

[0232] The minimum value between the health score and the business availability score is selected as the controller's score.

[0233] It should be noted that the description of the features in the embodiment corresponding to the device can be found in the relevant description of the embodiment corresponding to the method, and will not be repeated here.

[0234] Embodiments of this application also provide an arbitration apparatus, which may include:

[0235] The score receiving module is used to receive scores sent by multiple controllers;

[0236] The selection module is used to select the controller corresponding to the maximum score among multiple scores, and use it as the target controller.

[0237] The controller arbitration module is used to determine the target controller as the controller that has successfully preempted the controller.

[0238] For a description of the features in the embodiment corresponding to the arbitration device, please refer to the relevant description of the embodiment corresponding to the arbitration method, which will not be repeated here.

[0239] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described arbitration method embodiments.

[0240] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described arbitration method embodiments when it is run.

[0241] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0242] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described arbitration method embodiments.

[0243] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described arbitration method embodiments.

[0244] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0245] The above provides a detailed description of an arbitration method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.< / pid> < / pid> < / pid>

Claims

1. An arbitration method, characterized in that, Applied to a controller, the method includes: The system acquires hardware metrics, software metrics, service capability data, host path connection data, host path reachability data, and controller service type data for the controller. The host path connection data includes the number of paths connecting the controller to the host and the total number of paths connecting all controllers to the host. The host path reachability data includes the number of valid volumes on the controller path and the total number of volumes configured on all controllers. The controller service type data includes service type data and the weight score corresponding to each service type. The controller's health score is determined based on the hardware metric data, software metric data, and service capability data; the controller's host path connection status score is calculated based on the number of paths, the total number of paths, and the first decision weight score corresponding to the host path connection data; the controller's host path reachability score is calculated based on the number of volumes, the total number of volumes, and the second decision weight score corresponding to the host path reachability data; the controller's service weight score is calculated based on the service type data, the weight corresponding to each service type, and the third decision weight score corresponding to the service type data; and the controller's service availability score is determined based on the host path connection status score, the host path reachability score, and the service weight score; wherein, the weight corresponding to each service type is set based on the priority of each service type. The controller's score is determined based on the health score and the business availability score; The score is sent to the arbitration device via an arbitration request, so that the arbitration device can arbitrate based on the controller's score.

2. The arbitration method according to claim 1, characterized in that, The hardware performance data includes at least one of the following: CPU data, memory data, motherboard data, system disk data, network communication capability data, and chassis data; the software performance data includes process stability data and resource efficiency data. The service capability data includes historical fault data, successful arbitration data, and latency data. The process of determining the health score of the controller based on the hardware indicator data, software indicator data, and service capability data includes: Based on the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, the hardware performance score of the controller is calculated. Based on the process stability data and resource efficiency data, the software performance score of the controller is calculated. Based on the historical fault data, successful arbitration data, and latency data, the service capability score of the controller is calculated. The health score of the controller is calculated based on the hardware indicator score, the software indicator score, and the service capability score.

3. The arbitration method according to claim 2, characterized in that, The calculation of the controller's hardware performance score based on the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data includes: For the i-th hardware indicator data among the central processing unit data, memory data, motherboard data, system disk data, network communication capability data, and chassis data, determine whether the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data; where i∈[1,N], and N is the total number of hardware indicator data; If the i-th hardware indicator data belongs to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data is determined as the first preset score. If the i-th hardware indicator data does not belong to the data range corresponding to the i-th hardware indicator data, the score corresponding to the i-th hardware indicator data is determined as the second preset score. The hardware indicator score of the controller is calculated based on the score corresponding to each of the aforementioned hardware indicator data.

4. The arbitration method according to claim 2, characterized in that, The process stability data includes the core process survival rate and the number of process restarts within the first preset period; the resource efficiency data includes the CPU utilization rate, the proportion of soft interrupts, and the file descriptor reclamation rate. The calculation of the controller's software metric score based on the process stability data and resource efficiency data includes: For the j-th software indicator data among the core process survival rate, the number of process restarts within the first preset period, the CPU utilization rate, the soft interrupt ratio, and the file descriptor reclamation rate, determine whether the j-th software indicator data belongs to the data range corresponding to the j-th software indicator data; where j∈[1,M], and M is the total number of software indicator data; If the j-th software indicator data falls within the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data is determined as the third preset score. If the j-th software indicator data does not belong to the data range corresponding to the j-th software indicator data, the score corresponding to the j-th software indicator data is determined as the fourth preset score. The software indicator score of the controller is calculated based on the score corresponding to each of the software indicator data.

5. The arbitration method according to claim 2, characterized in that, The calculation of the controller's service capability score based on the historical fault data, successful arbitration data, and latency data includes: The number of failures of the controller within a second preset period is determined based on the historical fault data. Based on the number of faults, the preset reduction in points per fault, and the preset total fault score, the fault prediction score of the controller is calculated. The number of successful preemptions by the controller within a preset time period is determined based on the successful arbitration data. Based on the number of successful preemptions, the preset single preemption bonus, and the preset maximum preemption score, the preemption bonus value of the controller is calculated; Based on the preset delay range to which the delay data belongs, the delay score of the controller is determined; The service capability score of the controller is calculated based on the fault prediction score, the number of successful preemption attempts, and the preemption bonus score.

6. The arbitration method according to claim 1, characterized in that, The acquisition of the controller's hardware metrics data, software metrics data, service capability data, host path connection data, host path reachability data, and controller service type data includes: When the connection between the first node and the second node to which the controller belongs is broken, the hardware indicator data, software indicator data, service capability data, host path connection data, host path reachability data, and controller service type data of the controller are obtained; wherein, the second node is a node other than the first node.

7. The arbitration method according to claim 1, characterized in that, The determination of the controller's score based on the health score and the business availability score includes at least one of the following: The health score and the service availability score are aggregated to obtain the controller's score; The maximum value between the health score and the business availability score is selected as the score of the controller; The minimum value between the health score and the business availability score is selected as the score for the controller.

8. An arbitration method, characterized in that, Applied to arbitration equipment, the method is characterized by comprising: Receive scores from multiple controllers, determined using the method described in any one of claims 1 to 7; Select the controller with the highest score from among multiple scores as the target controller; The target controller is identified as the controller that successfully preempted the controller.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the arbitration method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Dual-active cluster arbitration method and device, computer equipment and storage medium

    CN117499210A

  • Storage cluster arbitration method and device, electronic equipment and storage medium

    CN120508456A