Computing power resource intelligent operation and maintenance fault self-healing method and device of intelligent computing center
By real-time monitoring and correlation analysis of multi-dimensional operation indicators in the management platform of the Intelligent Computing Center, fault handling instructions are generated and task scripts are allocated, the problem of inefficient fault positioning in the existing technology is solved, the independent management of computing power resources and fault self-healing is realized, and the system stability and operation and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202510622203.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-19
AI Technical Summary
The operation and maintenance system of the existing intelligent computing center lacks the ability to analyze the correlation of multi-dimensional indicators, resulting in low fault positioning efficiency, making it difficult to quickly generate repair strategies in complex fault scenarios, and is unable to achieve independent management of computing power resources and self-healing of faults.
By monitoring multi-dimensional operation indicators in real time in the management platform, performing correlation analysis, generating fault processing instructions, calling task script information, and assigning target tasks to associated computing nodes to repair faults, and using the monitoring module of the intelligent computing center and preset security channels for task allocation and repair.
It realizes rapid detection and precise positioning of faults, improves the response timeliness of fault repair, reduces the complexity of manual operation, optimizes resource utilization, realizes the independent management of computing resources and self-healing of faults, and improves system stability and operation and maintenance efficiency.
Smart Images

Figure CN120508426A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical fields of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a method and device for intelligent operation and maintenance fault self-healing of computing power resources in an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.
[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] As intelligent computing centers evolve towards large-scale distributed architectures, the abnormal response mechanism for computing resources faces severe challenges. Existing operation and maintenance systems mostly adopt a passive monitoring and manual intervention mode. After triggering alarms through preset thresholds, they rely on operation and maintenance personnel to troubleshoot faults one by one, resulting in low fault location efficiency and long recovery cycles. In cross-domain resource scheduling scenarios, traditional methods lack the ability to correlate and analyze multi-dimensional indicators (such as CPU and GPU utilization, network latency, etc.), making it difficult to quickly generate corresponding repair strategies in complex fault scenarios. Therefore, since the emergence of intelligent computing centers, how to deal with sudden hardware failures or software anomalies and achieve autonomous management of computing resources and self-healing of faults has been an urgent problem to be solved. Summary of the Invention
[0008] The present invention provides a method and device for intelligent operation and maintenance of computing power resources in an intelligent computing center, which can fully cope with sudden hardware failures or software anomalies and realize autonomous management and fault self-healing of computing power resources.
[0009] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0010] In a first aspect, the present invention provides a method for intelligent operation and maintenance fault self-healing of computing resources in an intelligent computing center, the method being applied to a management platform in the intelligent computing center, the method comprising:
[0011] Step S1: determining a fault to be repaired based on the operating status information of the intelligent computing center and generating a fault handling instruction;
[0012] Step S2: Based on the fault handling instruction, calling task script information containing the fault recovery logic of the fault to be repaired;
[0013] Step S3: Based on a task query request sent by at least one computing power node deployed in the intelligent computing center, assign a target task corresponding to the task query request to a target node, wherein the target task carries the task script information, and the target node is configured to repair the fault to be repaired according to the task script information;
[0014] Among them, the target node is the computing power node associated with the fault to be repaired among the at least one computing power node, and the task query request is a request for the computing power node to periodically call the task to be executed to the management platform.
[0015] In one embodiment, step S1 includes:
[0016] Step S11: Based on the monitoring module deployed in the intelligent computing center, multi-dimensional operating indicators of the intelligent computing center are collected, and the multi-dimensional operating indicators include at least one of central processing unit (CPU) utilization, GPU utilization, memory usage, disk space occupancy, network latency, and service process status;
[0017] Step S12: performing correlation analysis on the multi-dimensional operating indicators, and determining the fault to be repaired in the intelligent computing center based on a preset fault judgment rule;
[0018] Step S13: matching a predefined fault handling strategy according to the fault type of the fault to be repaired, and generating a fault handling instruction including a target repair action.
[0019] In one embodiment, step S2 includes:
[0020] Step S21: Based on the target repair action in the fault handling instruction, a basic repair script corresponding to the fault type to be repaired is retrieved from a script library to obtain the task script information;
[0021] Step S22: adjusting the execution parameters of the basic repair script according to at least one of the occurrence node information, fault severity, and associated operation indicators of the fault to be repaired, to obtain updated task script information;
[0022] The execution parameters include at least one of an execution account, a timeout period, a number of retries, and a concurrency level, and the task script information is packaged into an executable file format and is attached with a digital signature.
[0023] In one embodiment, step S3 includes:
[0024] Step S31: receiving a task query request periodically sent by the at least one computing node, wherein the task query request carries first identification information and operation status information of the computing node initiating the task query request;
[0025] Step S32: determining the target node based on a matching relationship between the second identification information of the node where the fault to be repaired occurs and the first identification information;
[0026] Step S33: Filtering out a target task corresponding to the task query request of the target node from the task queue to be assigned, wherein the target task includes the task script information;
[0027] Step S34: Allocate the target task to the target node.
[0028] In one embodiment, step S34 includes:
[0029] Step S341: Encapsulate the target task into an encrypted response message containing a digital signature;
[0030] Step S342: Sending the encrypted response message to the target node via a preset secure channel in the intelligent computing center; wherein the preset secure channel is established based on the Transport Layer Security (TLS) protocol and uses a dynamic session key for data transmission;
[0031] After step S34, the method further includes:
[0032] Step S35: Record the allocation status and expected execution time of the target task, and generate an execution tracking identifier associated with the target task. The execution tracking identifier is used to match the execution result returned by the target node.
[0033] In one embodiment, after step S3, the method further includes:
[0034] Step S4: Based on the execution result returned by the target node, a fault repair effect evaluation model is constructed, where the fault repair effect evaluation model is constructed based on multi-dimensional indicators, and the multi-dimensional indicators include at least one of the CPU utilization recovery rate, GPU utilization recovery rate, memory leak suppression rate, network delay improvement value, and service availability improvement of the target node after repair;
[0035] Step S5: when the result output by the fault repair effect evaluation model shows that the fault repair fails, instructing the target node and / or the associated nodes of the target node to perform a secondary repair operation according to the task script information;
[0036] The associated node is a node that is connected to the target node in at least one of physical architecture, network topology, storage system or business logic.
[0037] In a second aspect, the present invention provides a computing resource intelligent operation and maintenance fault self-healing device for an intelligent computing center, which is applied to a management platform of the intelligent computing center, and the device includes:
[0038] A first generating module is used to determine the fault to be repaired according to the operating status information of the intelligent computing center and generate a fault handling instruction;
[0039] A script calling module is used to call task script information containing fault recovery logic of the fault to be repaired based on the fault handling instruction;
[0040] A task allocation module is used to allocate a target task corresponding to a task query request to a target node based on a task query request sent by at least one computing power node deployed in the intelligent computing center;
[0041] Among them, the target node is the computing power node associated with the fault to be repaired among the at least one computing power node, the task query request is a request for the computing power node to periodically call the task to be executed to the management platform, and the target task carries the task script information, which is executed by the target node to repair the fault to be repaired.
[0042] In the third aspect, the present invention provides a server comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, the steps of the intelligent operation and maintenance fault self-healing method of computing power resources of the intelligent computing center as described in the first aspect above are implemented.
[0043] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for intelligent operation and maintenance fault self-healing of computing power resources of an intelligent computing center as described in the first aspect above are implemented.
[0044] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the method for intelligent operation and maintenance fault self-healing of computing power resources in an intelligent computing center as described in the first aspect above.
[0045] In the present invention, a method for intelligent operation and maintenance fault self-healing of computing power resources running in an intelligent computing center is provided. The management platform in the intelligent computing center determines the fault to be repaired based on the operating status information of the intelligent computing center, and generates a fault handling instruction accordingly. Then, based on the fault handling instruction, task script information containing the fault recovery logic of the fault to be repaired is generated. After receiving a task query request sent by multiple computing power nodes, a corresponding target task is assigned to the computing power node associated with the fault to be repaired, and the target task carries task script information so that the target node repairs the fault to be repaired according to the task script. Thus, the embodiment of the present invention accurately locates the fault by real-time monitoring of the operating status, combines automated task script generation with intelligent node scheduling, realizes closed-loop management of fault repair, and improves response timeliness. At the same time, the complexity of manual operation is reduced through the structured encapsulation of the task script, and computing power nodes are allocated according to fault correlation to optimize resource utilization, reduce redundant calculations, realize autonomous management of computing power resources and fault self-healing, and improve computing power resource operation and maintenance efficiency and system stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0047] Figure 1 A flowchart of a method for intelligent operation and maintenance of computing resources in an intelligent computing center for fault self-healing in the present invention;
[0048] Figure 2 for Figure 1 Schematic diagram of the architecture of the Intelligent Computing Center;
[0049] Figure 3 This is a structural diagram of a computing power resource intelligent operation and maintenance fault self-healing device in an intelligent computing center in the present invention;
[0050] Figure 4The figure is a schematic structural diagram of an electronic device in the present invention. DETAILED DESCRIPTION
[0051] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0052] First, the technical terms involved in the present invention are briefly explained below.
[0053] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0054] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .
[0055] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0056] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.
[0057] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.
[0058] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0059] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.
[0060] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0061] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.
[0062] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.
[0063] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0064] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".
[0065] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0066] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0067] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0068] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, computing resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.
[0069] For details, see Figure 1 , Figure 1 This is a flow chart of a method for intelligent operation and maintenance of computing resources in an intelligent computing center, provided by an embodiment of the present invention. The method is applied to a management platform in the intelligent computing center and specifically includes the following steps:
[0070] Step S1: Determine the fault to be repaired based on the operating status information of the intelligent computing center and generate a fault handling instruction.
[0071] In the above steps, by collecting the operating status data of the intelligent computing center in real time, using the rule engine or artificial intelligence model to analyze abnormal patterns, identify the fault type and the scope of the fault impact, and generate targeted fault handling instructions, such as restarting services, migrating tasks, etc.
[0072] The above-mentioned operating status may include hardware status data (such as CPU utilization, GPU utilization, memory usage, disk I / O, temperature, fan speed, etc.), software status data (such as process status, service logs, number of network connections, error code statistics, etc.), and business indicator data (such as task processing delay, throughput, success rate, resource quota usage, etc.). In addition, the data sources of the above-mentioned operating status information also include logging systems, API interfaces, and third-party tools, etc., which are not specifically limited in this application.
[0073] It should be noted that the fault handling instructions generated above can be matched and generated according to different fault types. For example, when the hardware sends a fault, it marks the node for repair, and an instruction to trigger the spare parts replacement process can be generated. Alternatively, if the software service is abnormal, instructions such as "restart service", "reload configuration", and "roll back version" can be generated. When resources are overloaded, load balancing strategies can be triggered, such as migrating tasks to other nodes. Among them, the fault handling instructions can include fault location information (such as node identification (Identifier, ID), faulty components, etc.) and target repair actions (such as restarting the GPU service of node 1, etc.).
[0074] In this way, the above steps achieve rapid detection and precise location of faults through multi-dimensional data collection and intelligent analysis, avoiding the lag and misjudgment of manual inspections, and solidifying the fault handling logic into executable instructions, reducing the operation and maintenance personnel's reliance on experience and ensuring consistency in the handling of similar faults.
[0075] It is worth mentioning that the intelligent computing center in the embodiment of the present application may include a management platform and multiple computing nodes. Each computing node is connected to the management platform. The computing node asks the management platform which tasks need to be executed through a direct service (server) interface with the management platform. If the management node needs to assign a task to a computing node, it will return the task script information to the computing node so that the computing node can create (fork) a process locally through the task script information and then report the result to the management platform. Figure 2As shown in the example, the management platform can be the ibex-server deployed in the Nightingale system, and the computing nodes can be the ibex-agant deployed on the Categraf side of the intelligent computing center. To simplify deployment, the management platform and computing nodes can be combined into a binary system, or ibex. Different roles can be activated by passing different parameters, reducing the complexity of manual operations and improving the flexibility and stability of the intelligent computing center.
[0076] Step S2: Based on the fault handling instruction, call the task script information containing the fault recovery logic of the fault to be repaired.
[0077] It should be noted that, based on the fault handling instructions, a preset fault recovery logic template (such as a script library or knowledge graph) can be called to dynamically generate a task script containing repair steps, parameters, and dependencies. For example, if the fault is a node downtime, the script may include operations such as restarting the service and checking dependencies.
[0078] In a specific embodiment, key parameters such as node ID, operation type, and target service can be extracted by parsing fault handling instructions. Based on the fault type, the corresponding script template library is called, such as predefined script templates for Shell, Python, and Ansible Playbook. For example, when a service restart is required, the preset template service_restart.sh can be called, and the parameters "node IP = 192.168.1.100, service name = GPU_Service" can be filled in. If a configuration update is required, the template config_rollback.py can be called, specifying the target configuration file path and version number.
[0079] In this way, the above steps can standardize the repair process, reduce human configuration errors, and improve the automation and consistency of fault repair.
[0080] Step S3: Based on a task query request sent by at least one computing power node deployed in the intelligent computing center, assign a target task corresponding to the task query request to a target node, wherein the target task carries the task script information, and the target node is configured to repair the fault to be repaired according to the task script information;
[0081] Among them, the target node is the computing power node associated with the fault to be repaired among the at least one computing power node, and the task query request is a request for the computing power node to periodically call the task to be executed to the management platform.
[0082] In a specific embodiment of the present invention, computing nodes periodically send task query requests to the management platform. The management platform selects the optimal target node based on fault relevance (e.g., node load, fault type, and geographic location) and sends the task script to that node for execution. For example, a network fault repair task can be assigned to a low-load node in the same region.
[0083] Specifically, each computing power node can send a task query request to the management platform through a periodic heartbeat mechanism. The task query request can carry node status information, including CPU, GPU, memory load, current task list, supported fault type labels, etc. The management platform can maintain a node capability mapping table to record the types of faults that each node can handle, for example, node A supports GPU service repair, and node B supports network configuration repair. After receiving the query request, the target node can be filtered according to the following strategies: relevance priority, that is, priority is given to the node where the fault occurred (such as local service restart) or the same cluster node (reducing cross-network communication); load balancing, that is, selecting nodes with low current load and processing capabilities; fault tolerance strategy, that is, avoiding nodes that have frequently failed recently. Subsequently, the task script information can be encapsulated into the target task and sent to the target node through a remote procedure call (RPC) or a message queue.
[0084] Furthermore, after receiving the task, the target node can parse the script information and execute the repair logic, while recording the execution log (for example, including the start time, key step output, error information, etc.). After the execution is completed, the result (success / failure) is fed back to the management platform. If it fails, step S1 is triggered to re-detect the fault or invoke the fallback strategy.
[0085] As a result, the intelligent computing center in the embodiment of the present application has achieved a leap from "manual operation and maintenance" to "intelligent self-healing", significantly improving the stability and operation and maintenance efficiency of the system. It is especially suitable for scenarios such as AI training and big data processing that have extremely high reliability requirements. It can optimize resource utilization, shorten repair delays, and ensure the efficiency and fault tolerance of task allocation.
[0086] In one embodiment, step S1 includes:
[0087] Step S11: Based on the monitoring module deployed in the intelligent computing center, multi-dimensional operating indicators of the intelligent computing center are collected, and the multi-dimensional operating indicators include at least one of central processing unit (CPU) utilization, GPU utilization, memory usage, disk space occupancy, network latency, and service process status;
[0088] Step S12: performing correlation analysis on the multi-dimensional operating indicators, and determining the fault to be repaired in the intelligent computing center based on a preset fault judgment rule;
[0089] Step S13: matching a predefined fault handling strategy according to the fault type of the fault to be repaired, and generating a fault handling instruction including a target repair action.
[0090] In some specific embodiments, a monitoring module deployed in an intelligent computing center can be used to collect multi-dimensional operating indicators in real time, and the collected multi-dimensional indicators can be correlated and analyzed, and faults can be identified in combination with preset fault judgment rules. Subsequently, the collected multi-dimensional indicators are correlated and analyzed, and faults can be identified in combination with preset fault judgment rules (such as threshold alarms, timing anomaly detection, logical association rules, etc.). Finally, a predefined processing strategy library can be matched according to the fault type (such as hardware failure, network interruption, service anomaly) to generate specific repair instructions, which may include restarting services, migrating tasks, expanding resources, etc. For example, if it is determined that the network delay is too high, the instructions may include switching backup links or optimizing routing strategies.
[0091] Specifically, in step S11 above, the monitoring module deployed in the intelligent computing center collects multi-dimensional operating indicators such as CPU utilization, memory usage, disk space occupancy, network latency, and service process status in real time. These indicators are collected using various technical means, such as monitoring tools such as Prometheus and Zabbix. Subsequently, the collected operating indicators are transmitted to the management platform and stored in a time series database or distributed log system for subsequent analysis and processing.
[0092] In the above step S12, the management platform can clean and standardize the collected multi-dimensional operating indicators. Since the data formats collected by different monitoring sources may be different, these format differences can be eliminated through data preprocessing, and abnormal data can be corrected to ensure the quality of the data. In addition, association rule mining, time series analysis and other technologies can be used to conduct in-depth analysis of the pre-processed multi-dimensional operating indicators. For example, when it is found that the CPU usage rate continues to increase while the memory usage rate is also increasing, and the service process status is abnormal, it can be determined that there is a correlation between these indicators. In addition, the management platform can also compare the results of the correlation analysis with the preset fault judgment rules to make a fault judgment. Since the preset rules are based on historical fault data and expert experience, the faults to be repaired in the intelligent computing center can be determined through comparison.
[0093] In step S13 above, the management platform can classify the fault to be repaired into different types, such as hardware fault, software fault, network fault, etc., based on the nature of the fault. Subsequently, the management platform can match the corresponding handling strategy from a predefined fault handling strategy library based on the fault type. For example, if the fault is determined to be a software service process crash, the corresponding restart service strategy will be matched. Finally, based on the matched handling strategy, fault handling instructions containing the target repair actions can be generated. These instructions clearly specify the specific operations required to repair the fault, such as restarting the service and adjusting resource allocation.
[0094] In this way, the embodiments of the present application can achieve automated fault detection and diagnosis through multi-dimensional data collection, correlation analysis, and intelligent strategy matching. Through correlation analysis and preset rules, faults can be accurately located. This avoids the false alarms that may be caused by judging based on only a single indicator, improves the accuracy of fault detection, and further achieves automated fault detection and diagnosis through multi-dimensional data collection, correlation analysis, and intelligent strategy matching. This not only improves the timeliness of fault discovery, but also lays a solid foundation for subsequent automated repairs, allowing the entire intelligent operation and maintenance fault self-healing system to operate more efficiently and reliably.
[0095] In one embodiment, step S2 includes:
[0096] Step S21: Based on the target repair action in the fault handling instruction, a basic repair script corresponding to the fault type to be repaired is retrieved from a script library to obtain the task script information;
[0097] Step S22: adjusting the execution parameters of the basic repair script according to at least one of the occurrence node information, fault severity, and associated operation indicators of the fault to be repaired, to obtain updated task script information;
[0098] The execution parameters include at least one of an execution account, a timeout period, a number of retries, and a concurrency level, and the task script information is packaged into an executable file format and is attached with a digital signature.
[0099] In some embodiments, the template script that matches the corresponding fault type can be called from the pre-stored script library according to the specific repair action in the fault handling instruction. For example, if the fault is a "storage node failure", the "data migration script" or "storage volume mount script" is called. In the application, the basic script can be quickly located and loaded through the mapping relationship between the fault type and the script library (such as a database index or rule engine). Therefore, in the embodiment of the present application, the standardized repair process is used to reduce duplicate development and improve the script reuse rate.
[0100] Furthermore, the management platform can also dynamically adjust the parameters of the basic repair script based on one or more of the fault node information (such as IP address, resource type), severity (such as impact range, priority) and related indicators (such as CPU, GPU / memory usage) to obtain updated task script information. For example, high-priority faults increase the number of retries, and low-resource nodes reduce concurrency. The script is ultimately encapsulated as an executable file (such as a Python script or a Docker image) and a digital signature (such as SHA-256) is attached to ensure integrity. In the application, dynamic variables can be injected through a parameter template engine (such as Jinja2), and an encryption algorithm can be called to generate a signature. In this way, the execution efficiency can be optimized through parameter adaptation, and the digital signature can ensure the credibility of the script and prevent malicious tampering.
[0101] In one embodiment, step S3 includes:
[0102] Step S31: receiving a task query request periodically sent by the at least one computing node, wherein the task query request carries first identification information and operation status information of the computing node initiating the task query request;
[0103] Step S32: determining the target node based on a matching relationship between the second identification information of the node where the fault to be repaired occurs and the first identification information;
[0104] Step S33: Filtering out a target task corresponding to the task query request of the target node from the task queue to be assigned, wherein the target task includes the task script information;
[0105] Step S34: Allocate the target task to the target node.
[0106] In some embodiments, the management platform can continuously monitor the network port and receive task query requests periodically sent from each computing power node in the intelligent computing center. In the application, the first identification information (such as the ID and IP address of the computing power node) and the operating status information (such as CPU utilization, GPU utilization, memory usage, current load, etc.) of the computing power node that initiated the request are extracted from the received task query request. By receiving the periodic requests and status information of the nodes, the management platform can grasp the operating status of each computing power node in real time, provide accurate data support for subsequent task allocation, and ensure that tasks can be allocated to appropriate nodes.
[0107] Subsequently, the second identification information of the node where the fault to be repaired occurs can be compared with the first identification information of each computing power node. Among them, if the first identification information of a computing power node matches the second identification information of the node where the fault occurs, then the node is preferentially determined as the target node. If there is no match, the node's load condition, resource status, processing power and other factors will be comprehensively considered to select the most suitable node from other nodes as the target node. Therefore, step S32 in the embodiment of the present application ensures that the fault can be assigned to the node that is most suitable for repairing it. Giving priority to repairing the node where the fault occurs itself can reduce the overhead of data transmission and coordination and improve repair efficiency. In the case of node mismatch, taking other factors into consideration can make full use of the resources of the entire computing center and ensure the overall performance of the system.
[0108] It should be noted that the management platform can maintain a task queue to be assigned, which has wherein stored all pending tasks. Therefore, in the embodiment of the present application, according to the task query request of target node and the capability label of this node (such as the fault type supported, processing capability etc.), from the task queue to be assigned, filter out the target task corresponding thereto. When screening, the task script information carried by the target task can be guaranteed to match the processing capability of the target node. As can be seen, by screening the task of matching, it is possible to ensure that the target node possesses the ability to handle this task, avoid the repair failure or inefficient problem caused by task and node capability mismatch, and improve the success rate of fault repair.
[0109] Furthermore, the embodiment of the present application can encapsulate the screened target tasks, including task script information, execution parameters, timeout settings, etc. The encapsulated target tasks are then sent to the target node through a secure and reliable communication protocol (such as Hypertext Transfer Protocol Secure (HTTPS), gRPC). This step achieves reliable task allocation, ensuring that the target node can accurately receive task information and execute it in a timely manner. The secure communication protocol ensures the integrity and security of task data during transmission, providing a guarantee for the smooth progress of fault repair.
[0110] It can be seen that the above-mentioned embodiments of the present application achieve efficient distribution of fault repair tasks through node status perception, intelligent matching and reliable allocation, fully utilize distributed computing resources, improve the parallelism and success rate of fault processing, and also enhance the elasticity and reliability of the entire intelligent computing center.
[0111] In one embodiment, step S34 includes:
[0112] Step S341: Encapsulate the target task into an encrypted response message containing a digital signature;
[0113] Step S342: Sending the encrypted response message to the target node via a preset secure channel in the intelligent computing center; wherein the preset secure channel is established based on the Transport Layer Security (TLS) protocol and uses a dynamic session key for data transmission;
[0114] After step S34, the method further includes:
[0115] Step S35: Record the allocation status and expected execution time of the target task, and generate an execution tracking identifier associated with the target task. The execution tracking identifier is used to match the execution result returned by the target node.
[0116] In some embodiments, the key information (such as task script, execution parameters, priority) of the target task can be encapsulated into a structured message (such as JSON or Protocol Buffers format) by the management platform. In addition, the private key of the management platform can be used to perform a hash operation on the message content to generate a digital signature, thereby ensuring message integrity and source credibility. Subsequently, a symmetric encryption algorithm can be adopted to encrypt the message after encapsulation. Thus, the target task in the embodiment of the present application can ensure that the target task is not maliciously tampered with during transmission by carrying out a digital front, and the recipient (target node) can confirm that the source of the message is legal by the public key verification signature of the management platform, and can also prevent sensitive information (such as repairing script logic) from being eavesdropped during transmission by an encryption mechanism.
[0117] Furthermore, the management platform and the target node can establish an encrypted communication channel based on the pre-configured Transport Layer Security (TLS) certificate to verify the identities of both parties. Among them, the TLS handshake protocol (such as the ECDHE algorithm) can be used to dynamically generate temporary session keys, which are renegotiated each time the communication is made to avoid long-term exposure of the key. In this way, an encrypted response message is sent through the established TLS channel, and end-to-end encryption is performed using the session key. Therefore, the embodiment of the present application ensures the authenticity of the identities of the communicating parties through the TLS protocol, prevents man-in-the-middle attacks, and the session key is dynamically generated and replaced regularly, reducing the risk of the key being cracked.
[0118] It is also worth mentioning that in step S35, the status of the target task allocation can also be recorded, and the allocation status of the target task (such as "allocated" or "transmitting") and the expected execution time (calculated based on task priority and node load) can be written into a distributed log system (such as Elasticsearch) or a relational database for recording. In addition, a globally unique execution tracking identifier can be generated for each target task and stored in association with the task metadata (such as fault ID, node ID). After the task is executed, the execution result (success / failure) returned by the target node will carry this tracking identifier, and the management platform can quickly locate the corresponding task record through the identifier.
[0119] In this way, the above embodiment achieves full-link traceability by tracking the entire process of allocation, execution, and result feedback of tasks associated with the identification. By recording the expected execution time, it helps to monitor the timeliness of task processing and ensure that key fault repairs comply with the service level agreement. In addition, by automatically matching the execution results through tracking identification, subsequent processes (such as repair verification and alarm cancellation) are triggered to achieve a complete closed loop of fault self-healing. The embodiment of the present application ensures the security of the task allocation process, which is particularly suitable for sensitive environments. It also improves the transparency and automation of operation and maintenance through the tracking mechanism, allowing system administrators to monitor the progress of fault repair in real time and intervene quickly when anomalies occur, thereby jointly enhancing the reliability, security and manageability of the intelligent computing center.
[0120] In one embodiment, after step S3, the method further includes:
[0121] Step S4: Based on the execution result returned by the target node, a fault repair effect evaluation model is constructed, where the fault repair effect evaluation model is constructed based on multi-dimensional indicators, and the multi-dimensional indicators include at least one of the CPU utilization recovery rate, GPU utilization recovery rate, memory leak suppression rate, network delay improvement value, and service availability improvement of the target node after repair;
[0122] Step S5: when the result output by the fault repair effect evaluation model shows that the fault repair fails, instructing the target node and / or the associated nodes of the target node to perform a secondary repair operation according to the task script information;
[0123] The associated node is a node that is connected to the target node in at least one of physical architecture, network topology, storage system or business logic.
[0124] In some embodiments, an evaluation model is constructed by monitoring the multi-dimensional indicators of the target node after repair (such as CPU utilization recovery rate, GPU utilization recovery rate, memory leak suppression rate, etc.). For example, if the CPU utilization is restored to more than 90% of the normal threshold after repair, the repair is determined to be effective. In the specific implementation process, the monitoring system can be used to collect the index data after repair in real time, and the repair effect can be evaluated in combination with a preset threshold or a machine learning model (such as a random forest). In this way, the repair effect can be quantified by the fault repair effect evaluation model to provide data support for subsequent optimization and avoid blind repair. When the output result shows that the fault repair has failed, one or more of the target node and the associated nodes of the target node can perform a secondary repair operation according to the task script information to improve the success rate of the repair.
[0125] In another embodiment, the execution parameters of the task script information can also be adjusted to generate optimized task script information. In addition, the optimized task script information is pushed to the target node and / or the associated node of the target node for performing a secondary repair operation. Specifically, the execution parameters of the task script can be dynamically adjusted according to the evaluation results (for example, increasing the number of retries, optimizing concurrency), and an optimized script can be generated. For example, if the network delay improvement value is insufficient, the task scheduling strategy can be adjusted to reduce the network load. In the application, the evaluation results can be fed back to the script generation module through a feedback mechanism, and a parameter tuning algorithm (for example, a genetic algorithm) can be used to generate a better configuration. Thus, the embodiment of the present application improves repair efficiency and adaptability, reduces waste of resources, and enhances system robustness.
[0126] Furthermore, the management platform can push the optimized script to the target node and its associated nodes (such as other nodes in the same cluster or dependent service nodes) to ensure the coordination of the repair operation. For example, if a storage node failure is repaired, the configuration of its backup node needs to be updated synchronously. Therefore, the embodiment of the present application pushes the script to the associated nodes through a distributed task distribution mechanism (such as Kafka or Redis Pub / Sub), ensuring the consistency of the data script, realizing global collaborative repair, avoiding system inconsistencies caused by local repairs, and improving overall stability.
[0127] See Figure 3 , Figure 3 This is a structural diagram of a computing power resource intelligent operation and maintenance fault self-healing device for an intelligent computing center provided by an embodiment of the present invention, which is applied to the management platform of the intelligent computing center. The computing power resource intelligent operation and maintenance fault self-healing device 20 includes:
[0128] A first generating module 21 is configured to determine a fault to be repaired based on the operating status information of the intelligent computing center and generate a fault handling instruction;
[0129] A script calling module 22 is configured to call task script information containing fault recovery logic for the fault to be repaired based on the fault handling instruction;
[0130] A task assignment module 23 is configured to assign a target task corresponding to a task query request sent by at least one computing power node deployed in the intelligent computing center to a target node based on the task query request;
[0131] Among them, the target node is the computing power node associated with the fault to be repaired among the at least one computing power node, the task query request is a request for the computing power node to periodically call the task to be executed to the management platform, and the target task carries the task script information, which is executed by the target node to repair the fault to be repaired.
[0132] In one embodiment, the first generating module 21 is configured to:
[0133] Based on a monitoring module deployed in the intelligent computing center, multi-dimensional operating indicators of the intelligent computing center are collected, wherein the multi-dimensional operating indicators include at least one of central processing unit (CPU) utilization, GPU utilization, memory usage, disk space occupancy, network latency, and service process status;
[0134] Performing correlation analysis on the multi-dimensional operating indicators and determining the fault to be repaired in the intelligent computing center based on preset fault judgment rules;
[0135] A predefined fault handling strategy is matched according to the fault type of the fault to be repaired, and a fault handling instruction including a target repair action is generated.
[0136] In one embodiment, the script calling module 22 includes:
[0137] A script calling unit is used to call a basic repair script corresponding to the fault type of the fault to be repaired from a script library based on the target repair action in the fault handling instruction to obtain the task script information;
[0138] A script updating unit, configured to adjust the execution parameters of the basic repair script according to at least one of the occurrence node information of the fault to be repaired, the fault severity, and the associated operation index, to obtain updated task script information;
[0139] The execution parameters include at least one of an execution account, a timeout period, a number of retries, and a concurrency level, and the task script information is packaged into an executable file format and is attached with a digital signature.
[0140] In one embodiment, the task allocation module 23 includes:
[0141] A first receiving unit is configured to receive a task query request periodically sent by the at least one computing power node, wherein the task query request carries first identification information and operation status information of the computing power node that initiates the task query request;
[0142] A first determining unit is configured to determine the target node based on a matching relationship between the second identification information of the node where the fault to be repaired occurs and the first identification information;
[0143] A task screening unit, configured to screen out a target task corresponding to the task query request of the target node from a queue of tasks to be assigned, wherein the target task includes the task script information;
[0144] A task allocation unit is used to allocate the target task to the target node.
[0145] In one embodiment, the task allocation unit is configured to:
[0146] Encapsulating the target task into an encrypted response message containing a digital signature;
[0147] Sending the encrypted response message to the target node through a preset secure channel in the intelligent computing center; wherein the preset secure channel is established based on the Transport Layer Security (TLS) protocol and uses a dynamic session key for data transmission;
[0148] After the task allocation unit, the computing power resource intelligent operation and maintenance fault self-healing device of the intelligent computing center is also used to:
[0149] The allocation status and expected execution time of the target task are recorded, and an execution tracking identifier associated with the target task is generated, where the execution tracking identifier is used to match the execution result returned by the target node.
[0150] In one embodiment, after the task allocation module, the computing power resource intelligent operation and maintenance fault self-healing device 20 of the intelligent computing center is further used to:
[0151] Based on the execution result returned by the target node, a fault repair effect evaluation model is constructed. The fault repair effect evaluation model is constructed based on multi-dimensional indicators, and the multi-dimensional indicators include at least one of the CPU utilization recovery rate, GPU utilization recovery rate, memory leak suppression rate, network delay improvement value and service availability improvement of the target node after repair.
[0152] Adjusting the execution parameters of the task script information according to the output of the fault repair effect evaluation model to generate optimized task script information;
[0153] Pushing the optimized task script information to the target node and / or an associated node of the target node for performing a secondary repair operation;
[0154] The associated node is a node that is connected to the target node in at least one of physical architecture, network topology, storage system or business logic.
[0155] It should be noted that the intelligent operation and maintenance fault self-healing device 20 of the computing power resource of the intelligent computing center in the embodiment of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.
[0156] The intelligent operation and maintenance fault self-healing device 20 for computing power resources of an intelligent computing center provided in an embodiment of the present invention is capable of realizing the various processes of each embodiment of the intelligent operation and maintenance fault self-healing method for computing power resources of the above-mentioned intelligent computing center. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.
[0157] The embodiment of the present invention further provides an electronic device, see Figure 4 , Figure 4 The electronic device includes a memory 31, a processor 32, and a program or instruction stored in the memory 31. When the program or instruction is executed by the processor 32, the program or instruction can be realized. Figure 1 Any steps in the corresponding embodiment of the intelligent operation and maintenance fault self-recovery method for computing power resources of the intelligent computing center and the same beneficial effects are achieved will not be repeated here.
[0158] The processor 32 may be a CPU, an ASIC, an FPGA or a GPU.
[0159] Those skilled in the art will understand that all or part of the steps of the embodiment of the intelligent operation and maintenance fault self-healing method for computing power resources of the above-mentioned intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.
[0160] The embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, which can realize the above-mentioned Figure 1 Any step in the corresponding embodiment of the method for intelligent operation and maintenance of computing power resources in an intelligent computing center can achieve the same technical effect. To avoid repetition, it is not repeated here. The storage medium is such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0161] The present application also provides a computer program product including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the embodiment of the intelligent operation and maintenance fault self-healing method of the computing power resources of the intelligent computing center shown can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0162] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0164] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for intelligent operation and maintenance of computing resources in an intelligent computing center, characterized in that: The method is applied to a management platform in an intelligent computing center, and includes: Step S1: determining a fault to be repaired based on the operating status information of the intelligent computing center and generating a fault handling instruction; Step S2: Based on the fault handling instruction, calling task script information containing the fault recovery logic of the fault to be repaired; Step S3: Based on a task query request sent by at least one computing power node deployed in the intelligent computing center, assign a target task corresponding to the task query request to a target node, wherein the target task carries the task script information, and the target node is configured to repair the fault to be repaired according to the task script information; Among them, the target node is the computing power node associated with the fault to be repaired among the at least one computing power node, and the task query request is a request for the computing power node to periodically call the task to be executed to the management platform.
2. The method according to claim 1, characterized in that The step S1 comprises: Step S11: Based on the monitoring module deployed in the intelligent computing center, multi-dimensional operating indicators of the intelligent computing center are collected, and the multi-dimensional operating indicators include at least one of central processing unit (CPU) utilization, GPU utilization, memory usage, disk space occupancy, network latency, and service process status; Step S12: performing correlation analysis on the multi-dimensional operating indicators, and determining the fault to be repaired in the intelligent computing center based on a preset fault judgment rule; Step S13: matching a predefined fault handling strategy according to the fault type of the fault to be repaired, and generating a fault handling instruction including a target repair action.
3. The method according to claim 2, characterized in that The step S2 comprises: Step S21: Based on the target repair action in the fault handling instruction, a basic repair script corresponding to the fault type to be repaired is called from a script library to obtain the task script information; Step S22: adjusting the execution parameters of the basic repair script according to at least one of the occurrence node information, fault severity, and associated operation indicators of the fault to be repaired, to obtain updated task script information; The execution parameters include at least one of an execution account, a timeout period, a number of retries, and a concurrency level, and the task script information is packaged into an executable file format and is attached with a digital signature.
4. The method according to claim 1, wherein The step S3 comprises: Step S31: receiving a task query request periodically sent by the at least one computing node, wherein the task query request carries first identification information and operation status information of the computing node initiating the task query request; Step S32: determining the target node based on a matching relationship between the second identification information of the node where the fault to be repaired occurs and the first identification information; Step S33: Filtering out a target task corresponding to the task query request of the target node from the task queue to be assigned, wherein the target task includes the task script information; Step S34: Allocate the target task to the target node.
5. The method according to claim 4, characterized in that The step S34 includes: Step S341: Encapsulate the target task into an encrypted response message containing a digital signature; Step S342: Sending the encrypted response message to the target node via a preset secure channel in the intelligent computing center; wherein the preset secure channel is established based on the Transport Layer Security (TLS) protocol and uses a dynamic session key for data transmission; After step S34, the method further includes: Step S35: Record the allocation status and expected execution time of the target task, and generate an execution tracking identifier associated with the target task. The execution tracking identifier is used to match the execution result returned by the target node.
6. The method according to any one of claims 1 to 5, characterized in that After step S3, the method further includes: Step S4: Based on the execution result returned by the target node, a fault repair effect evaluation model is constructed, where the fault repair effect evaluation model is constructed based on multi-dimensional indicators, and the multi-dimensional indicators include at least one of the CPU utilization recovery rate, GPU utilization recovery rate, memory leak suppression rate, network delay improvement value, and service availability improvement of the target node after repair; Step S5: when the result output by the fault repair effect evaluation model shows that the fault repair has failed, instructing the target node and / or the associated nodes of the target node to perform a secondary repair operation according to the task script information; The associated node is a node that is connected to the target node in at least one of physical architecture, network topology, storage system or business logic.
7. A computing power resource intelligent operation and maintenance fault self-healing device for an intelligent computing center, characterized in that: A management platform for an intelligent computing center, comprising: A first generating module is used to determine the fault to be repaired according to the operating status information of the intelligent computing center and generate a fault handling instruction; A script calling module is used to call task script information containing fault recovery logic of the fault to be repaired based on the fault handling instruction; A task allocation module is used to allocate a target task corresponding to a task query request to a target node based on a task query request sent by at least one computing power node deployed in the intelligent computing center; Among them, the target node is the computing power node associated with the fault to be repaired among the at least one computing power node, the task query request is a request for the computing power node to periodically call the task to be executed to the management platform, and the target task carries the task script information, which is executed by the target node to repair the fault to be repaired.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for intelligent operation and maintenance fault self-healing of computing power resources of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for intelligent operation and maintenance fault self-healing of computing power resources in an intelligent computing center according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method for intelligent operation and maintenance fault self-healing of computing power resources of an intelligent computing center as described in any one of claims 1 to 6.