Data center high-availability automatic inspection method and device
By conducting domain division and failure and explosion scope analysis on the operation and maintenance objects of the data center infrastructure, a high availability capability view is generated, and the problem of lack of evaluation and management of the data center's high availability capability is solved, efficient and comprehensive inspection and management are achieved, and the business continuity of the system is improved.
Patent Information
- Application Number
- CN202510545486.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-29
Smart Images

Figure CN120560931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information systems, and in particular to a method and device for high-availability automated inspection of a data center. Background Art
[0002] This section is intended to provide a background or context for embodiments of the present invention. No description herein is admitted to be prior art by virtue of its inclusion in this section.
[0003] High availability is the ability of a system to provide fault-free, uninterrupted services. The design of high-availability architecture can be discussed from two levels: application and infrastructure. At the application level, high availability can be achieved through service degradation, current limiting, and circuit breaking, while at the infrastructure level, it focuses on resource redundancy, fault isolation, and other means.
[0004] With the continued adoption of distributed and cloud-native technologies, IT (Information Technology) system complexity is growing exponentially, leading to increased failure frequency and longer resolution times. This poses a significant challenge to high system availability in the financial industry, which demands stringent business continuity. Data centers prioritize high availability as a key architectural design factor and deployment principle when providing various infrastructure resources and operations services for systems. However, from design to operational state, non-standard deployments, often caused by implementation errors, lack of audits, and resource constraints, inevitably lead to a lack of high availability capabilities.
[0005] Existing system high availability checking methods have the following shortcomings:
[0006] 1. Lack of high availability evaluation standards and management basis:
[0007] Data center infrastructure resources are diverse and wide-ranging. Different products offer varying levels of high availability capabilities, and their corresponding high-availability deployment architectures also vary. Data centers often publish various product specifications in the form of product operations manuals, lacking a holistic high-availability perspective.
[0008] 2. Lack of problem discovery methods and normalized management mechanisms:
[0009] Supporting the normal operation of a system requires the joint guarantee of various layers of infrastructure. Problems at any node will have an impact, and information systems have very strict requirements for business continuity. Once an interruption occurs, it may cause unpredictable losses. With existing technologies, it is obviously not advisable to wait until a real production event occurs before locating and rectifying the problem. Summary of the Invention
[0010] To address the problems of existing technologies, this paper proposes a method and device for automated data center high-availability inspection. This method can improve the comprehensiveness and integrity of data center high-availability management, avoid inspection omissions, effectively reduce data center failures and losses caused by high-availability failures, and enhance system business continuity capabilities.
[0011] An embodiment of the present invention provides a data center high availability automated inspection method, comprising:
[0012] Obtain multiple infrastructure operation and maintenance objects of the data center;
[0013] Using feature matching algorithms, the infrastructure operation and maintenance objects are divided into domains to generate domain division results; the domains include computing, basic software, storage, network, security, or computer room environment, or any combination thereof;
[0014] Determine the failure explosion range of each infrastructure operation and maintenance object based on historical data from the data center;
[0015] Based on the fault explosion range and the preset high availability capability tiering model, multiple infrastructure operation and maintenance objects are stratified for high availability capabilities and the stratification results are generated.
[0016] Based on the domain division and stratification results, high availability modeling is performed on multiple infrastructure operation and maintenance objects to generate and display high availability capability views;
[0017] According to the preset high availability inspection model, high availability inspection is performed on the infrastructure operation and maintenance objects, and the high availability inspection results are generated and displayed; the preset high availability inspection model is established based on the high availability capability view.
[0018] An embodiment of the present invention further provides a data center high-availability automated inspection device, comprising:
[0019] The acquisition module is used to obtain multiple infrastructure operation and maintenance objects of the data center;
[0020] The domain division module is used to divide the infrastructure operation and maintenance objects into domains using a feature matching algorithm and generate domain division results; the domains include computing domain, basic software domain, storage domain, network domain, security domain, or computer room environment domain, or any combination thereof;
[0021] The fault explosion range determination module is used to determine the fault explosion range of each infrastructure operation and maintenance object based on the historical data of the data center;
[0022] A stratification result generation module is used to perform high availability stratification on multiple infrastructure operation and maintenance objects based on the fault explosion range and the preset high availability stratification model, and generate stratification results;
[0023] The high-availability capability view generation module is used to perform high-availability modeling on multiple infrastructure operation and maintenance objects based on the domain division and layering results, and to generate and display high-availability capability views;
[0024] The high availability inspection result generation module is used to perform high availability inspection on infrastructure operation and maintenance objects according to a preset high availability inspection model, and generate and display high availability inspection results; the preset high availability inspection model is established according to the high availability capability view.
[0025] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a method for automated inspection of high availability of a data center when executing the computer program.
[0026] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer program implements a method for automated inspection of high availability of a data center.
[0027] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a method for automated inspection of high availability of a data center.
[0028] Compared with the technical solutions for system high availability inspection in the prior art, the embodiments of the present invention obtain multiple infrastructure operation and maintenance objects of the data center; use a feature matching algorithm to divide the infrastructure operation and maintenance objects into domains and generate domain division results; the domains include one or any combination of computing domain, basic software domain, storage domain, network domain, security domain or computer room environment domain; determine the failure explosion range of each infrastructure operation and maintenance object based on the historical data of the data center; perform high availability capability stratification on multiple infrastructure operation and maintenance objects based on the failure explosion range and a preset high availability capability stratification model to generate stratification results; perform high availability modeling on multiple infrastructure operation and maintenance objects based on the domain division results and stratification results, generate and display a high availability capability view; perform high availability inspection on the infrastructure operation and maintenance objects based on a preset high availability inspection model, generate and display high availability inspection results; the preset high availability inspection model is established based on the high availability capability view, which can improve the comprehensiveness and integrity of the high availability management of the data center, avoid inspection omissions, effectively reduce data center failures and losses caused by high availability failures, and enhance the system business continuity capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 This is a flow chart of a data center high availability automated inspection method according to an embodiment of the present invention;
[0031] Figure 2 This is a flowchart of a specific example of a method for automated inspection of high availability of a data center according to an embodiment of the present invention;
[0032] Figure 3 This is a flowchart of a specific example of a method for automated inspection of high availability of a data center according to an embodiment of the present invention;
[0033] Figure 4 is a schematic diagram of a data center high-availability automated inspection device according to an embodiment of the present invention;
[0034] Figure 5 Schematic diagram of the computer device structure according to an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0036] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.
[0037] First, let’s introduce the professional terms involved in this article:
[0038] The High Availability Model (HAM) is a model established for the high availability capabilities of operation and maintenance objects, including its high availability deployment requirements, high availability failure scenarios, and high availability recovery capabilities.
[0039] A TOR (Top of Rack Switch) is a network device widely used in data centers, typically deployed atop server cabinets. In this context, it refers to a group of access switches that provide high availability.
[0040] AZ (Availability Region) is a physical area with independent wind, fire, water, electricity, and network services within the same region, which isolates faults between availability zones.
[0041] Data center infrastructure resources are diverse and wide-ranging. Different products can provide different high-availability capabilities, and their corresponding high-availability deployment architectures are also different. In order to manage them scientifically and completely, an overall view of high-availability capabilities is required.
[0042] Supporting the normal operation of a system requires the coordinated efforts of various layers of infrastructure. Problems at any node can have an impact. Financial information systems, in particular, have stringent requirements for business continuity. A disruption can lead to unforeseen losses. Waiting until a real production incident occurs to identify and rectify the problem is simply not advisable. We need proactive risk detection methods and regularized management mechanisms to effectively prevent high-availability production incidents.
[0043] In response to the above situation, an embodiment of the present invention proposes an automated inspection method for high availability of data centers. First, the high availability capabilities of data center infrastructure are modeled, and an overall view of the high availability capabilities of infrastructure is formed through domain division and high availability capability level division, key elements such as fault scenarios and fault recovery capabilities are clarified, and a high availability management system is established; second, the high availability architecture standard is converted into an inspection model, and by comparing whether the operating state and the design state of the high availability deployment architecture are consistent, it is determined whether and where high availability risks exist in the system, and regular inspection services can be provided in an online, automated, and one-click manner at a predetermined frequency, so that risk warnings can be issued at the first time to guide operation and maintenance personnel to respond in a timely manner.
[0044] The principles and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0045] Figure 1 This is a flow chart of a data center high availability automated inspection method according to an embodiment of the present invention. Figure 1 As shown, the method may include:
[0046] Step 101: Acquire multiple infrastructure operation and maintenance objects of a data center;
[0047] Step 102: Using a feature matching algorithm, the infrastructure operation and maintenance objects are divided into domains to generate domain division results; the domains include one or any combination of the computing domain, basic software domain, storage domain, network domain, security domain, or computer room environment domain;
[0048] Step 103: Determine the failure explosion range of each infrastructure operation and maintenance object based on the historical data of the data center;
[0049] Step 104: Perform high availability stratification on multiple infrastructure operation and maintenance objects based on the fault explosion range and a preset high availability stratification model, and generate a stratification result.
[0050] Step 105: Based on the domain division results and the layering results, high availability modeling is performed on multiple infrastructure operation and maintenance objects to generate and display a high availability capability view;
[0051] Step 106: Perform a high availability check on the infrastructure operation and maintenance object according to a preset high availability check model, and generate and display high availability check results; the preset high availability check model is established according to the high availability capability view.
[0052] The embodiments of the present invention overcome the shortcomings of the above-mentioned existing technologies through the above-mentioned steps, and provide a more comprehensive, efficient and accurate data center high-availability automated inspection method, which can improve the comprehensiveness and integrity of data center high-availability management, avoid inspection omissions, effectively reduce data center failures and losses caused by high-availability failures, and enhance the system's business continuity capabilities.
[0053] Figure 2 FIG. 1 is a flow chart of a specific example of a method for automated inspection of high availability of a data center according to an embodiment of the present invention. Figure 2 As shown, the data center high availability automated inspection method of an embodiment of the present invention may specifically include: dividing the infrastructure operation and maintenance objects into fields, and dividing the high availability capabilities into levels according to the size of the fault explosion radius, and standardizing the high availability deployment architecture and high availability capability indicators of each operation and maintenance object by establishing a high availability model, and finally forming an overall view of the high availability capability.
[0054] First, in step 101, multiple infrastructure operation and maintenance objects of a data center are obtained. The data center may be a financial data center, and all infrastructure operation and maintenance objects of the financial data center may be included in the high availability capability analysis scope.
[0055] Next, in order to improve the accuracy of domain division, in one embodiment, a feature matching algorithm is used to perform domain division on the infrastructure operation and maintenance object to generate a domain division result, which may include: using a feature extraction algorithm to perform feature extraction on the historical data of the infrastructure operation and maintenance object to obtain a first feature; using a text feature extraction algorithm to perform feature extraction on a pre-stored domain definition text to obtain a second feature; the domain definition text includes one or any combination of definition texts in the computing field, basic software field, storage field, network field, security field or computer room environment field; generating a first feature descriptor based on the first feature, and generating a second feature descriptor based on the second feature; performing feature matching on the first feature and the second feature based on the first feature descriptor and the second feature descriptor to generate a matching result; and generating a domain division result based on the matching result.
[0056] The entire infrastructure operation and maintenance objects of the financial data center are classified for easy management. In the embodiment of the present invention, they can be divided into 6 areas:
[0057] (1) Computing domain: includes various infrastructure operation and maintenance objects that provide computing services, such as virtual machines and physical machines;
[0058] (2) Basic software field: databases, operating systems, middleware, containers and other infrastructure operation and maintenance objects;
[0059] (3) Storage: This includes various infrastructure operation and maintenance objects that provide storage services, such as object storage and file storage;
[0060] (4) Network domain: includes various infrastructure operation and maintenance objects that provide network services, such as switches, load balancers, network firewalls, etc.
[0061] (5) Security domain: includes various infrastructure operation and maintenance objects that provide security services, such as encryption and decryption equipment, key machines, application firewalls, etc.
[0062] (6) Computer room environment: includes various infrastructure operation and maintenance objects that provide the basic operating environment in the computer room, such as power supplies, cabinets, etc.
[0063] In one embodiment, in order to provide a data basis for high-availability capability stratification, historical data may include the operating data of the data center before and after the individual failure of each infrastructure operation and maintenance object; determining the failure explosion range of each infrastructure operation and maintenance object based on the historical data of the data center may include: determining the type of data affected by the individual failure of each infrastructure operation and maintenance object based on the operating data of the data center before and after the individual failure of each infrastructure operation and maintenance object; determining the failure explosion range of each infrastructure operation and maintenance object based on the type of data affected by the individual failure of each infrastructure operation and maintenance object.
[0064] In one embodiment, high availability stratification is performed on multiple infrastructure operation and maintenance objects based on the fault explosion range and a preset high availability stratification model to generate stratification results, which may include: inputting the fault explosion range of each infrastructure operation and maintenance object into the preset high availability stratification model, and outputting the high availability stratification level of each infrastructure operation and maintenance object; the preset high availability stratification model includes multiple high availability stratification levels, multiple fault explosion ranges, and a mapping relationship between multiple high availability stratification levels and multiple fault explosion ranges; and determining the output high availability stratification level of each infrastructure operation and maintenance object as the stratification result.
[0065] High availability capabilities are divided into different levels according to the scope of failure explosion. Each infrastructure operation and maintenance object evaluates the high availability capability level that can be met based on its own design and actual capabilities. The embodiment of the present invention divides high availability capabilities into 8 levels:
[0066] (1) Component level: The high availability design of each component in the device, including network cards, network ports, etc., avoids the high availability risk of service failure caused by a single component failure.
[0067] (2) Device level: High availability design of physical or logical devices to avoid the risk of service failure caused by failure of a single device or a single logical host.
[0068] (3) Cabinet level: Equipment deployment should consider cross-cabinet high availability design to avoid the risk of high availability deployment failure due to cabinet-level failures such as cabinet power supply, resulting in the inability to provide services.
[0069] (4) TOR level: Devices are deployed and connected to different TORs (The Onion Router) to avoid the risk of service failure caused by physical or logical failures in the same group of switches.
[0070] (5) Computer room level: Equipment deployment takes into account high availability at the computer room module level to avoid the risk of failure of high availability deployment due to failure of a single computer room module, resulting in the inability to provide services.
[0071] (6) Building level / AZ (Availability Zone) level: Equipment deployment should consider high availability across buildings / AZs to avoid the risk of high availability deployment failure and service inability due to failure of a single building / AZ.
[0072] (7) Campus level: Equipment deployment takes into account high availability across campuses in the same city. When a campus-level failure occurs, another campus in the same city can take over the business and provide external services.
[0073] (8) Regional level: Equipment deployment takes into account the high availability of remote campuses. When a regional failure occurs, the remote campus can take over the business and provide external services.
[0074] The domain division and stratification results obtained after analyzing the infrastructure operation and maintenance objects of switches and TiDB databases are shown in Table 1 below.
[0075] Table 1
[0076]
[0077] In one embodiment, the data center high availability automated inspection method may also include: determining the fault recovery time, availability rate and failure business impact rate of each infrastructure operation and maintenance object based on the operation data of the data center before and after the individual failure of each infrastructure operation and maintenance object; the failure business impact rate represents the impact degree of the failure of the infrastructure operation and maintenance object.
[0078] By establishing an indicator system, the high availability capability of each infrastructure operation and maintenance object is digitally evaluated. By comparing the differences between the design state and the operating state, it is not only possible to intuitively monitor whether the high availability operating status of each infrastructure operation and maintenance object meets expectations, but also to reversely optimize and correct the high availability design. The embodiment of the present invention proposes three key indicators for evaluating high availability capability, namely, fault recovery time, availability rate, and fault business impact rate:
[0079] (1) Fault recovery time = fault recovery time - fault occurrence time;
[0080] (2) Availability = (Current Planned Service Time - Current Outage Time) / Current Planned Service Time × 100%. "Current" can be based on quarterly time. Current Outage Time = Fault Recovery Time.
[0081] (3) Fault business impact rate = number of business transactions affected during the fault period / expected number of business transactions during the fault period × 100%.
[0082] In one embodiment, based on the domain division results and stratification results, high availability modeling is performed on multiple infrastructure operation and maintenance objects, and a high availability capability view is generated and displayed. This may include: based on the domain division results, stratification results, fault recovery time, availability rate and fault business impact rate, high availability modeling is performed on multiple infrastructure operation and maintenance objects, and a high availability capability view is generated and displayed.
[0083] In one embodiment, the high availability capability view may further include: a high availability deployment description corresponding to the hierarchical results; and a preset high availability check model established specifically according to the high availability deployment description.
[0084] High availability modeling is performed on all infrastructure operation and maintenance objects, including operation and maintenance objects, domains, high availability capabilities, high availability deployment instructions, and high availability capability indicators. This creates a holistic high availability capability view, which serves as an important basis for high availability capability management of financial data center infrastructure. The high availability capability view template is shown in Table 2.
[0085] Table 2
[0086]
[0087] In one embodiment, the preset high availability inspection model may include a preset inspection script; according to the preset high availability inspection model, a high availability inspection is performed on the infrastructure operation and maintenance object, and high availability inspection results are generated and displayed, which may include: through the preset inspection script, a high availability inspection is performed on the instance information of the infrastructure operation and maintenance object, and high availability inspection results are generated and displayed.
[0088] In one embodiment, according to a preset high availability inspection model, a high availability inspection is performed on the infrastructure operation and maintenance object, and the high availability inspection results are generated and displayed, which may include: performing a high availability inspection on the instance information of the infrastructure operation and maintenance object according to one or any combination of the following algorithms to generate a high availability inspection result: comparing the instance information of the numerical type of the infrastructure operation and maintenance object and the associated operation and maintenance objects that have an associated relationship with the infrastructure operation and maintenance object with a preset threshold, and determining the high availability inspection result based on the comparison result; performing regular matching on the instance information of the character type of the infrastructure operation and maintenance object and the associated operation and maintenance objects that have an associated relationship with the infrastructure operation and maintenance object, and determining the high availability inspection result based on the regular matching result.
[0089] Based on the overall view of high availability capabilities, a high-availability automated inspection mechanism is established, which supports both regular inspections of all operation and maintenance objects at a fixed frequency and manual triggering of inspection actions for individual operation and maintenance objects, meeting inspection needs in different scenarios.
[0090] Figure 3 FIG. 1 is a flow chart of a specific example of a method for automated inspection of high availability of a data center according to an embodiment of the present invention. Figure 3 As shown, the data center high availability automated inspection method may specifically include:
[0091] 1. Build a high availability check object:
[0092] The high availability inspection object is formed by matching the inspection data and the high availability inspection model.
[0093] Check the data level: First, the financial data center needs to establish a data warehouse to manage the instance information (basic information of the instance) of each infrastructure operation and maintenance object, and build a high availability group on this basis. The high availability group is the minimum instance set of the infrastructure operation and maintenance object that can provide a certain high availability capability, and it is also the smallest unit for performing high availability automated inspections.
[0094] High availability inspection model level: Based on the high availability deployment instructions in the overall high availability capability view, a high availability inspection model is formed.
[0095] 2. High availability check execution:
[0096] Input high availability check object, supporting multiple types of algorithm implementations:
[0097] (1) Supports comparison of the number of instances and replicas of numerical types with preset thresholds;
[0098] (2) Support regular matching of character type configuration fields;
[0099] (3) Supports searching for numerical values and character type judgments of associated operation and maintenance objects that have an associated relationship with infrastructure operation and maintenance objects;
[0100] (4) Support more flexible rule judgment through scripts.
[0101] 3. High availability check result display:
[0102] High availability check results can be displayed from two perspectives: application system and infrastructure operation and maintenance object. In the application system dimension, with the system as the primary key, the high availability check results for all instances of the infrastructure operation and maintenance objects that support the system are displayed. In the infrastructure operation and maintenance object dimension, with the infrastructure operation and maintenance object as the primary key, the high availability check results for all instances of the infrastructure operation and maintenance object are displayed, along with the application systems associated with each instance.
[0103] The premise for implementing the above design is to establish an association between the application system and various infrastructure operation and maintenance objects.
[0104] 4. Establish a high availability problem management ledger:
[0105] The results of automated high-availability inspections should effectively track their corrective actions, responsibilities, results, and subsequent optimization recommendations. Therefore, a high-availability issue management ledger should be established. For each high-availability issue, an online record should be generated. This not only serves as a record of the rectification of existing issues, but also allows for problem-oriented, reverse optimization of high-availability designs, forming a virtuous cycle and establishing a normalized management mechanism. An example template for the high-availability issue management ledger is shown in Table 3.
[0106] Table 3
[0107]
[0108] The embodiment of the present invention proposes a method for modeling the high availability capability of data center infrastructure, abstracts the high availability capability of infrastructure into 8 levels, and intuitively reflects the high availability capability of infrastructure; proposes a high availability capability evaluation index to further digitally evaluate the high availability capability of infrastructure operation and maintenance objects; proposes a high availability automated inspection method, constructs high availability inspection objects based on the overall view of high availability capability, and meets the inspection requirements of both application and infrastructure dimensions; proposes a high availability normalization management mechanism, through the establishment of a high availability problem management ledger, problem-oriented, reversely optimizes high availability design, and promotes the formation of a virtuous circle of high availability capability construction.
[0109] Although the above solutions have improved and optimized the core technical solutions in some aspects, they are generally centered around the core purpose of the present invention, that is, to achieve efficient and accurate inspection and maintenance of cloud database connectivity. They are all reasonable improvements and extensions based on the present invention and can also achieve the purpose of the invention.
[0110] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and drawings, this does not require or imply that these operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0111] After introducing the method of the exemplary embodiment of the present invention, next, reference is made to Figure 3 A high-availability automated inspection device for a data center according to an exemplary embodiment of the present invention is introduced.
[0112] The implementation of the data center high-availability automated inspection device can be referenced to the implementation of the above-mentioned method, and any repetitions will not be repeated. The terms "module" or "unit" used below may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0113] Based on the same inventive concept, the present invention also proposes a high-availability automated inspection device for a data center. Figure 4 Schematic diagram of a data center high-availability automated inspection device according to an embodiment of the present invention. Figure 4 As shown, the device includes:
[0114] An acquisition module 401 is used to acquire multiple infrastructure operation and maintenance objects of a data center;
[0115] Domain division module 402 is used to divide infrastructure operation and maintenance objects into domains using a feature matching algorithm and generate domain division results; the domains include one or any combination of computing, basic software, storage, network, security, or computer room environment.
[0116] The fault explosion range determination module 403 is used to determine the fault explosion range of each infrastructure operation and maintenance object based on the historical data of the data center;
[0117] A stratification result generation module 404 is configured to perform high availability stratification on multiple infrastructure operation and maintenance objects based on the fault explosion range and a preset high availability stratification model, and generate a stratification result;
[0118] A high availability capability view generation module 405 is used to perform high availability modeling on multiple infrastructure operation and maintenance objects based on the domain division results and the layering results, and to generate and display a high availability capability view;
[0119] The high availability inspection result generation module 406 is used to perform high availability inspection on the infrastructure operation and maintenance objects according to a preset high availability inspection model, and generate and display high availability inspection results; the preset high availability inspection model is established according to the high availability capability view.
[0120] In one embodiment, the domain division module 402 is specifically configured to:
[0121] Using a feature extraction algorithm, extract features from the historical data of the infrastructure operation and maintenance object to obtain the first feature;
[0122] Using a text feature extraction algorithm, extracting features from pre-stored domain definition text to obtain a second feature; the domain definition text includes one or any combination of definition texts in the computing field, basic software field, storage field, network field, security field, or computer room environment field;
[0123] generating a first feature descriptor according to the first feature, and generating a second feature descriptor according to the second feature;
[0124] Performing feature matching on the first feature and the second feature according to the first feature descriptor and the second feature descriptor to generate a matching result;
[0125] Based on the matching results, domain division results are generated.
[0126] In one embodiment, the historical data includes the operational data of the data center before and after the individual failure of each infrastructure operation and maintenance object;
[0127] The fault explosion range determination module 403 is specifically configured to:
[0128] Determine the type of data affected by the failure of each infrastructure operation and maintenance object based on the data center's operating data before and after the failure of each infrastructure operation and maintenance object.
[0129] Determine the failure explosion range of each infrastructure operation and maintenance object based on the type of data affected by the individual failure of each infrastructure operation and maintenance object.
[0130] In one embodiment, the hierarchical result generation module 404 is specifically configured to:
[0131] Input the failure explosion range of each infrastructure operation and maintenance object into a preset high availability capability layering model, and output the high availability capability layer of each infrastructure operation and maintenance object; the preset high availability capability layering model includes multiple high availability capability layers, multiple failure explosion ranges, and mapping relationships between multiple high availability capability layers and multiple failure explosion ranges;
[0132] The high availability capability level of each output infrastructure operation and maintenance object is determined as a tiered result.
[0133] In one embodiment, the data center high availability automated inspection device further includes: an availability determination module, configured to:
[0134] Based on the operating data of the data center before and after the failure of each infrastructure operation and maintenance object, the failure recovery time, availability rate and failure business impact rate of each infrastructure operation and maintenance object are determined; the failure business impact rate represents the impact of the failure of the infrastructure operation and maintenance object.
[0135] In one embodiment, the high availability capability view generation module 405 is specifically configured to:
[0136] Based on the domain division results, stratification results, fault recovery time, availability rate, and fault business impact rate, high availability modeling is performed on multiple infrastructure operation and maintenance objects to generate and display high availability capability views.
[0137] In one embodiment, the high availability capability view further includes: a high availability deployment description corresponding to the hierarchical results; and a preset high availability check model is specifically established according to the high availability deployment description.
[0138] In one embodiment, the preset high availability inspection model includes a preset inspection script;
[0139] The high availability check result generation module 406 is specifically configured to:
[0140] Through preset inspection scripts, high availability inspections are performed on the instance information of infrastructure operation and maintenance objects, and high availability inspection results are generated and displayed.
[0141] In one embodiment, the high availability check result generation module 406 is specifically configured to:
[0142] Perform a high availability check on the instance information of the infrastructure operation and maintenance object according to one or any combination of the following algorithms to generate a high availability check result:
[0143] Compare instance information of numerical types of infrastructure operation and maintenance objects and associated operation and maintenance objects associated with the infrastructure operation and maintenance objects with preset thresholds, and determine a high availability check result based on the comparison result;
[0144] Regular matching is performed on the character type instance information of the infrastructure operation and maintenance object and the associated operation and maintenance object that has an associated relationship with the infrastructure operation and maintenance object, and the high availability check result is determined based on the regular matching result.
[0145] It should be noted that while the detailed description above mentions several modules of the data center high-availability automated inspection device, this division is merely exemplary and not mandatory. In practice, according to embodiments of the present invention, the features and functions of two or more modules described above may be embodied in a single module. Conversely, the features and functions of a single module described above may be further divided and embodied by multiple modules.
[0146] Based on the above invention concept, Figure 5 As shown, the present invention also proposes a computer device 500, including a memory 510, a processor 520 and a computer program 530 stored in the memory 510 and executable on the processor 520, wherein the processor 520 implements the aforementioned data center high availability automated inspection method when executing the computer program 530.
[0147] Based on the aforementioned inventive concept, the present invention proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the aforementioned data center high-availability automated inspection method is implemented.
[0148] Based on the aforementioned inventive concept, the present invention proposes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a method for automated inspection of high availability of a data center.
[0149] Compared with the technical solutions for system high availability inspection in the prior art, the embodiments of the present invention obtain multiple infrastructure operation and maintenance objects of the data center; use a feature matching algorithm to divide the infrastructure operation and maintenance objects into domains and generate domain division results; the domains include one or any combination of computing domain, basic software domain, storage domain, network domain, security domain or computer room environment domain; determine the failure explosion range of each infrastructure operation and maintenance object based on the historical data of the data center; perform high availability capability stratification on multiple infrastructure operation and maintenance objects based on the failure explosion range and a preset high availability capability stratification model to generate stratification results; perform high availability modeling on multiple infrastructure operation and maintenance objects based on the domain division results and stratification results, generate and display a high availability capability view; perform high availability inspection on the infrastructure operation and maintenance objects based on a preset high availability inspection model, generate and display high availability inspection results; the preset high availability inspection model is established based on the high availability capability view, which can improve the comprehensiveness and integrity of the high availability management of the data center, avoid inspection omissions, effectively reduce data center failures and losses caused by high availability failures, and enhance the system business continuity capability.
[0150] Compared with the prior art, the embodiments of the present invention have the following advantages:
[0151] 1. High availability management provides a clear and comprehensive perspective. High availability capability modeling is performed across the entire data center infrastructure, focusing on the high availability design of each operation and maintenance object from a unified, holistic perspective. This facilitates management and avoids omissions.
[0152] 2. Proactively discover problems. Establishing a mechanism for proactively discovering problems through automated inspections can effectively reduce system failures caused by high availability failures and improve the system's business continuity capabilities.
[0153] The acquisition, storage, use, and processing of data in the technical solution of this application comply with relevant laws and regulations.
[0154] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0155] The present invention is described with reference to flowcharts and / or block diagrams of methods and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0156] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0158] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A data center high availability automated inspection method, characterized in that: include: Obtain multiple infrastructure operation and maintenance objects of the data center; Using a feature matching algorithm, the infrastructure operation and maintenance object is divided into domains to generate domain division results; the domains include one or any combination of computing, basic software, storage, network, security, or computer room environment; Determine the failure explosion range of each infrastructure operation and maintenance object based on historical data from the data center; Based on the failure explosion range and the preset high availability capability stratification model, multiple infrastructure operation and maintenance objects are stratified for high availability capabilities to generate stratification results; Based on the domain division and stratification results, high availability modeling is performed on multiple infrastructure operation and maintenance objects to generate and display high availability capability views; According to a preset high availability inspection model, a high availability inspection is performed on the infrastructure operation and maintenance object, and high availability inspection results are generated and displayed; the preset high availability inspection model is established according to the high availability capability view.
2. The method according to claim 1, characterized in that Using a feature matching algorithm, the infrastructure operation and maintenance objects are divided into domains, and domain division results are generated, including: Using a feature extraction algorithm, extract features from the historical data of the infrastructure operation and maintenance object to obtain the first feature; Using a text feature extraction algorithm, extracting features from pre-stored domain definition text to obtain a second feature; the domain definition text includes one or any combination of definition texts in the computing field, basic software field, storage field, network field, security field, or computer room environment field; generating a first feature descriptor according to the first feature, and generating a second feature descriptor according to the second feature; Performing feature matching on the first feature and the second feature according to the first feature descriptor and the second feature descriptor to generate a matching result; Based on the matching results, domain division results are generated.
3. The method according to claim 1, characterized in that The historical data includes the operation data of the data center before and after the failure of each infrastructure operation and maintenance object; Based on historical data from the data center, determine the scope of failure explosion for each infrastructure operation and maintenance object, including: Determine the type of data affected by the failure of each infrastructure operation and maintenance object based on the data center's operating data before and after the failure of each infrastructure operation and maintenance object. Determine the failure explosion range of each infrastructure operation and maintenance object based on the type of data affected by the individual failure of each infrastructure operation and maintenance object.
4. The method according to claim 3, characterized in that Based on the failure explosion range and the preset high availability capability stratification model, multiple infrastructure operation and maintenance objects are stratified for high availability capabilities, and stratification results are generated, including: Inputting the failure explosion range of each infrastructure operation and maintenance object into a preset high availability capability hierarchical model, and outputting the high availability capability level of each infrastructure operation and maintenance object; the preset high availability capability hierarchical model includes multiple high availability capability levels, multiple failure explosion ranges, and mapping relationships between the multiple high availability capability levels and the multiple failure explosion ranges; The high availability capability level of each output infrastructure operation and maintenance object is determined as a tiered result.
5. The method according to claim 3, characterized in that Also includes: Based on the operating data of the data center before and after the failure of each infrastructure operation and maintenance object, the failure recovery time, availability rate and failure business impact rate of each infrastructure operation and maintenance object are determined; the failure business impact rate represents the impact of the failure of the infrastructure operation and maintenance object.
6. The method according to claim 5, characterized in that Based on the domain division and stratification results, high availability modeling is performed on multiple infrastructure operation and maintenance objects to generate and display high availability capability views, including: Based on the field division results, stratification results, fault recovery time, availability rate and fault business impact rate, high availability modeling is performed on multiple infrastructure operation and maintenance objects to generate and display a high availability capability view.
7. The method according to claim 1, characterized in that The high availability capability view also includes: a high availability deployment description corresponding to the layered result; and the preset high availability inspection model is specifically established according to the high availability deployment description.
8. The method according to claim 1, characterized in that The preset high availability inspection model includes a preset inspection script; Perform high availability checks on infrastructure operation and maintenance objects based on the preset high availability check model, and generate and display high availability check results, including: Through preset inspection scripts, high availability inspections are performed on the instance information of infrastructure operation and maintenance objects, and high availability inspection results are generated and displayed.
9. The method according to claim 8, characterized in that Perform high availability checks on infrastructure operation and maintenance objects based on the preset high availability check model, and generate and display high availability check results, including: Perform a high availability check on the instance information of the infrastructure operation and maintenance object according to one or any combination of the following algorithms to generate a high availability check result: Compare instance information of numerical types of infrastructure operation and maintenance objects and associated operation and maintenance objects associated with the infrastructure operation and maintenance objects with preset thresholds, and determine a high availability check result based on the comparison result; Regular matching is performed on the character type instance information of the infrastructure operation and maintenance object and the associated operation and maintenance object that has an associated relationship with the infrastructure operation and maintenance object, and the high availability check result is determined based on the regular matching result.
10. A data center high-availability automated inspection device, characterized in that: include: The acquisition module is used to obtain multiple infrastructure operation and maintenance objects of the data center; A domain division module is used to divide the infrastructure operation and maintenance objects into domains using a feature matching algorithm and generate domain division results; the domains include one or any combination of the computing domain, basic software domain, storage domain, network domain, security domain, or computer room environment domain; The fault explosion range determination module is used to determine the fault explosion range of each infrastructure operation and maintenance object based on the historical data of the data center; A stratification result generation module is used to perform high availability stratification on multiple infrastructure operation and maintenance objects according to the fault explosion range and a preset high availability stratification model, and generate a stratification result; A high availability capability view generation module is used to perform high availability modeling on multiple infrastructure operation and maintenance objects based on the domain division results and layering results, and to generate and display a high availability capability view; The high availability inspection result generation module is used to perform high availability inspection on infrastructure operation and maintenance objects according to a preset high availability inspection model, and generate and display high availability inspection results; the preset high availability inspection model is established according to the high availability capability view.
11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.