Industrial computing network intelligent operation and maintenance method and system, electronic equipment and storage medium

By collaborating with cloud-side intelligent agents and computing agents, unified perception and automated fault location of heterogeneous computing resources are achieved, solving the problems of resource dispersion and reliance on human experience in traditional operation and maintenance methods. This improves fault location efficiency and the level of operation and maintenance automation, ensuring the stability and high availability of industrial production.

CN121940261APending Publication Date: 2026-04-28INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
Filing Date
2025-12-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional operation and maintenance methods are difficult to achieve unified management and visual monitoring of heterogeneous and distributed computing resources. Fault location relies on manual experience and is too time-consuming. The degree of automation in operation and maintenance is low, which cannot meet the high availability requirements of industrial production.

Method used

By deploying computing agents on cloud-side intelligent agents and heterogeneous computing power nodes, operational events are acquired for root cause analysis, operational strategies are matched and distributed to target computing agents for execution, and closed-loop updates are performed using artificial intelligence models to achieve automated fault location and self-healing recovery.

Benefits of technology

It achieves unified perception of heterogeneous and distributed computing power nodes, quickly and accurately locates the root cause of faults, significantly shortens fault handling time, improves the level of operation and maintenance automation, and ensures the stability and high availability of industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940261A_ABST
    Figure CN121940261A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent operation and maintenance method and system for an industrial computing network, electronic equipment and a storage medium, and the method comprises the steps: deploying computing power agents at heterogeneous computing power nodes to obtain operation and maintenance events in a unified manner, and carrying out the centralized root cause analysis of the operation and maintenance events through a cloud end, so as to match an operation and maintenance strategy, the method comprises the following steps: acquiring an operation and maintenance strategy library, issuing the operation and maintenance strategy to a corresponding computing power agent for execution and feeding back a result, and finally performing closed-loop updating on the operation and maintenance strategy library and an artificial intelligence model based on a complete event case, thereby realizing unified perception of heterogeneous dispersed computing power nodes through the computing power agent, and solving the problem that unified management and monitoring are difficult due to resource dispersion; through automatic root cause analysis and strategy matching, a fault root cause can be rapidly and accurately positioned and an operation and maintenance scheme can be generated so that serious dependence on artificial experience can be eliminated and fault processing time can be substantially shortened. Through continuous optimization of a closed-loop learning mechanism, the operation and maintenance automation level is greatly improved, and high availability and stability of an industrial computing network are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent operation and maintenance technology, and in particular to an intelligent operation and maintenance method, system, electronic device and storage medium for industrial computing networks. Background Technology

[0002] With the development of the Industrial Internet, industrial computing networks are typically composed of widely distributed heterogeneous computing power nodes (such as factory all-in-one machines, edge clouds, and central clouds).

[0003] In this environment, traditional operation and maintenance methods face enormous challenges. First, the heterogeneous and dispersed nature of resources makes it difficult to achieve unified management and visualized monitoring. Second, application performance problems may originate from computing, networks, platforms, or the application itself, making fault localization extremely difficult and heavily reliant on human experience, resulting in excessively long average repair times. Finally, the low level of automation in operation and maintenance fails to meet the stringent high availability requirements of industrial production. Summary of the Invention

[0004] This invention provides an intelligent operation and maintenance method, system, electronic device, and storage medium for industrial computing networks. It addresses the shortcomings of existing technologies that struggle to achieve unified management and real-time visual monitoring of heterogeneous and dispersed computing resources, leading to low fault location efficiency and heavy reliance on manual experience. Furthermore, the existing operation and maintenance systems suffer from insufficient automation, failing to meet the stringent requirements of high availability and rapid response in industrial production scenarios.

[0005] This invention provides an intelligent operation and maintenance method for an industrial computing network, wherein the industrial computing network includes a cloud-side intelligent agent and multiple computing agents deployed on heterogeneous computing power nodes. The method is applied to the cloud-side intelligent agent and includes the following steps: Obtain the operation and maintenance events reported by the computing power agent; Root cause analysis is performed on the operation and maintenance events to obtain the root causes of the failures, and operation and maintenance strategies are obtained by matching the root causes of the failures from the operation and maintenance strategy library. The operation and maintenance strategy is distributed to the target computing power agent related to the root cause of the failure, and the execution results are received from the target computing power agent. Based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, an event case is determined, and based on the event case, the operation and maintenance strategy in the operation and maintenance strategy library is updated and / or the artificial intelligence model used for the root cause analysis is updated.

[0006] According to the present invention, an intelligent operation and maintenance method for industrial computing networks includes performing root cause analysis on the operation and maintenance events to obtain the root causes of the faults, comprising: The operation and maintenance events are subjected to fault identification to obtain fault events; Root cause analysis is performed based on the fault event to obtain the root cause of the fault.

[0007] According to the present invention, an intelligent operation and maintenance method for industrial computing networks includes performing root cause analysis on the fault event to obtain the root cause of the fault, comprising: Obtain the dependency topology relationships related to the fault event; The root cause of the failure is obtained by performing index correlation analysis on multiple resource indicators in the dependent topology relationship and / or performing time-series correlation analysis on multiple abnormal indicators in the dependent topology relationship.

[0008] According to the present invention, an intelligent operation and maintenance method for industrial computing networks is provided, wherein the dependency topology relationship includes a dependency topology graph; The step of performing correlation analysis on multiple resource indicators in the dependent topology includes: Query the configuration management database to obtain the dependency graph of the resource; the dependency graph includes multiple dependency nodes and dependency edges between the multiple dependency nodes; Along the dependency edges in the dependency graph, check in turn whether there are any abnormal metrics for the dependency nodes connected to the dependency edges.

[0009] According to the present invention, an intelligent operation and maintenance method for industrial computing networks includes updating the operation and maintenance strategies in the operation and maintenance strategy library and / or updating the artificial intelligence model used for root cause analysis based on the event cases, comprising: Adjust the weight of the operation and maintenance strategy in the event case within the operation and maintenance strategy library; and / or, The event cases are used as new training samples, and the artificial intelligence model is updated based on the new training samples.

[0010] According to the present invention, an intelligent operation and maintenance method for industrial computing networks includes, wherein distributing the operation and maintenance strategy to target computing power agents related to the root cause of the fault includes: The operation and maintenance strategy is constrained and verified, and the operation and maintenance strategy that passes the constraint verification is distributed to the target computing power agent related to the root cause of the fault. The constraint verification includes at least one of the following: Check whether the remaining resources of the computing node corresponding to the target computing power agent meet the requirements of the operation and maintenance strategy; Based on the monitoring data obtained from the target computing power agent, assess whether the current business load on the computing power node allows the operation in the operation and maintenance strategy to be executed.

[0011] According to the present invention, an intelligent operation and maintenance method for industrial computing networks, the step of distributing the operation and maintenance strategy that has passed the constraint verification to the target computing power agent related to the root cause of the fault further includes: Assess the execution risk level of the operation and maintenance strategy that has passed the constraint verification, and obtain the current business criticality level of the business unit corresponding to the target computing power agent; If the execution risk level is low risk, then the operation and maintenance strategy will be issued. If the execution risk level is high and the business criticality level is high, then an execution instruction is generated, and in response to the approval operation of the execution instruction, the operation and maintenance strategy is issued. If the execution risk level is medium risk, the operation and maintenance strategy will be added to the automatic execution queue of the preset off-peak maintenance window.

[0012] This invention also provides an industrial computing network intelligent operation and maintenance system, comprising the following units: The acquisition unit is used to acquire the operation and maintenance events reported by the computing power agent; The root cause analysis unit is used to perform root cause analysis on the operation and maintenance event, obtain the root cause of the failure, and obtain the operation and maintenance strategy from the operation and maintenance strategy library based on the root cause of the failure. The distribution unit is used to distribute the operation and maintenance strategy to the target computing power agent related to the root cause of the fault, and to receive the execution results fed back by the target computing power agent; The determining unit is configured to determine event cases based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, and update the operation and maintenance strategy in the operation and maintenance strategy library and / or update the artificial intelligence model used for the root cause analysis based on the event cases.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the industrial computing network intelligent operation and maintenance method described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the industrial computing network intelligent operation and maintenance method as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the industrial network intelligent operation and maintenance method described above.

[0016] The present invention provides an intelligent operation and maintenance method, system, electronic device, and storage medium for industrial computing networks, which acquires operation and maintenance events reported by computing power agents; performs root cause analysis on the operation and maintenance events to obtain the root causes of the faults, and matches operation and maintenance strategies from the operation and maintenance strategy library based on the root causes of the faults; distributes the operation and maintenance strategies to the target computing power agents related to the root causes of the faults, and receives the execution results fed back by the target computing power agents; determines event cases based on the operation and maintenance events, root causes of the faults, operation and maintenance strategies, and execution results, and updates the operation and maintenance strategies in the operation and maintenance strategy library and / or updates the artificial intelligence model used for root cause analysis based on the event cases. This invention deploys computing power agents on heterogeneous computing power nodes to uniformly acquire operation and maintenance events. The cloud performs centralized root cause analysis on these events to match operation and maintenance strategies. These strategies are then distributed to the corresponding computing power agents for execution, and the results are fed back. Finally, the operation and maintenance strategy library and artificial intelligence model are updated in a closed loop based on complete event cases. Thus, unified awareness of heterogeneous and distributed computing power nodes is achieved through computing power agents, solving the problem of difficult unified management and monitoring due to resource dispersion. Furthermore, through automated root cause analysis and strategy matching, the root cause of faults can be quickly and accurately located and operation and maintenance solutions can be generated, reducing reliance on manual experience and significantly shortening fault handling time. Moreover, continuous optimization through a closed-loop learning mechanism enables automatic fault discovery, accurate fault location, and execution, greatly improving the level of operation and maintenance automation and ensuring the stability of industrial computing networks. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts of the intelligent operation and maintenance method for industrial computing networks provided by the present invention.

[0019] Figure 2 This is the second flowchart of the intelligent operation and maintenance method for industrial computing networks provided by this invention.

[0020] Figure 3 This is a schematic diagram of the structure of the industrial computing network intelligent operation and maintenance system provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] This invention provides an intelligent operation and maintenance method for industrial computing networks. In this field, operation and maintenance in industrial computing network environments face numerous challenges, including difficulties in unified management due to heterogeneous and dispersed resources, time-consuming reliance on manual experience for fault root cause localization, and low levels of automation failing to meet high availability requirements. The method proposed in this embodiment aims to construct an intelligent operation and maintenance system with cloud-edge-device collaboration and a complete closed-loop capability of "perception-analysis-decision-execution-learning," thereby achieving automatic and accurate fault localization and rapid self-healing recovery, ensuring the continuity and stability of industrial production. In a specific application scenario, the industrial computing network includes a cloud-side intelligent agent and multiple computing agents deployed on heterogeneous computing power nodes. The method is applied to the cloud-side intelligent agent, which acts as the system's sole decision-making center, communicating and collaborating with the computing agents deployed on each computing power node. Figure 1 This is one of the flowcharts illustrating the intelligent operation and maintenance method for industrial computing networks provided by this invention, such as... Figure 1 As shown, the method includes steps 110, 120, 130 and 140.

[0024] Step 110: Obtain the operation and maintenance events reported by the computing power agent.

[0025] Specifically, in this embodiment, the computing power agent can be understood as a lightweight software program or module deployed on different computing power nodes in an industrial computing network. The computing power nodes may include end-side nodes, edge-side nodes, and cloud-side nodes, etc. This embodiment of the invention does not specifically limit this.

[0026] Here, "end-side node" refers to computing devices located close to the industrial production site, such as industrial all-in-one machines and gateway devices deployed in industrial sites. "Edge-side node" refers to nodes deployed in industrial parks, factories, or regional data centers that provide stronger computing power support for local businesses, such as edge servers and edge cloud nodes. "Cloud-side node" refers to a larger-scale computing resource pool located in a central data center or public cloud, such as a server cluster within a data center; this embodiment of the invention does not specifically limit this.

[0027] One of the core responsibilities of the computing power agent is intelligent sensing, which involves proactively collecting various monitoring data from the computing power node it resides on. This monitoring data covers multiple layers, from the underlying hardware to the upper-level applications, to ensure comprehensive sensing. Here, the monitoring data may include infrastructure layer data, platform and middleware layer data, and application and business layer data, etc., and this embodiment of the invention does not specifically limit these.

[0028] Here, infrastructure layer data may include CPU (Central Processing Unit) utilization, memory usage, and swap usage of computing resources; disk I / O (Input / Output) read / write speed and utilization of storage resources; and bandwidth utilization, TCP (Transmission Control Protocol) connection count, packet loss rate and error rate, and network latency of network resources. This embodiment of the invention does not specifically limit these data.

[0029] Here, the platform and middleware layer data may include Pod status in the container environment, number of restarts, number of database connections (such as MySQL), number of slow queries, message backlog in message queues (such as Kafka), etc., and this embodiment of the invention does not specifically limit these.

[0030] Here, application and business layer data may include service interface response time, throughput, error rate, application output business logs, and, in a microservice architecture, call chain data used to trace the call path of a single request.

[0031] To address the data heterogeneity issues across different computing power nodes and monitoring items, the computing power agent can standardize the collected monitoring data to ensure consistent format and semantics. Standardization methods can include: aligning all data timestamps to the same time zone (e.g., UTC+8) and precision (e.g., milliseconds); converting data in different units to standard units for metric normalization, such as standardizing memory to GB and network bandwidth to Mbps; structuring unstructured text logs into key-value pairs (e.g., JSON format); and adding rich contextual tags to each data entry, such as asset ID, business unit, and application name.

[0032] After standardization, the computing power agent encapsulates this data into one or more operation and maintenance events and reports them. Here, an operation and maintenance event is a structured data object containing information such as timestamps, standardized indicator data, and context labels. An operation and maintenance event clearly describes an observable phenomenon that occurs at a specific time and on a specific computing power node.

[0033] The reporting method can be flexible. For example, the computing agent can be configured to proactively push events when a certain indicator exceeds a preset threshold or when a critical error log is captured, or it can passively poll and report in response to periodic query requests from the cloud-side intelligent agent. By receiving these operation and maintenance events, the cloud-side intelligent agent achieves comprehensive and real-time perception of the entire industrial computing network's operational status.

[0034] Step 120: Perform root cause analysis on the operation and maintenance event to obtain the root cause of the failure, and obtain the operation and maintenance strategy from the operation and maintenance strategy library based on the root cause of the failure.

[0035] Specifically, after acquiring an operational event, root cause analysis can be performed to identify the root cause of the failure. Root cause analysis is a diagnostic process aimed at deeply mining and locating the root cause of anomalies or failures from a massive amount of superficial operational events. Within the general scope of this step, the specific implementation of root cause analysis can be diverse. For example, a cloud-side intelligent agent can utilize an AI-based analysis model. This model, by learning from massive amounts of historical operational data and cases, can identify complex relationships between different operational events, thereby inferring the root cause of the failure. Alternatively, a rule engine based on expert knowledge can be used to reason according to a pre-defined chain of logical rules. For example, when receiving an operational event indicating a "surge in application programming interface response time for application A," the root cause analysis process will comprehensively analyze other events related to application A, potentially concluding that the root cause is "a bottleneck in disk I / O on server D, where database C resides."

[0036] After identifying the root cause of a failure, the system needs to decide how to resolve the issue. To this end, the system maintains an operations and maintenance (O&M) policy repository. This repository can be understood as a knowledge base or database, containing pre-defined automated handling solutions for various known root causes of failures—the O&M policies. O&M policies can be stored in various formats; for example, a common format is an "IF-THEN" rule template, such as: IF The root cause of the failure is a disk I / O bottleneck THEN Execute a strategy to expand disk capacity or migrate part of the load.

[0037] Then, operational strategies can be obtained from the operational strategy library based on the root cause of the failure. Here, direct matching can be performed in the operational strategy library based on keywords in the root cause of the failure, or a more complex semantic matching algorithm can be used to understand the deeper meaning of the root cause of the failure and match strategies with similar intent in the operational strategy library. When multiple matching strategies exist, they can also be weighted and sorted according to metadata such as historical success rate, execution cost, and risk level of the operational strategies to select the optimal one or a group of strategies.

[0038] Step 130: Distribute the operation and maintenance strategy to the target computing power agent related to the root cause of the fault, and receive the execution result fed back by the target computing power agent.

[0039] Specifically, after obtaining the operation and maintenance strategy, the operation and maintenance strategy can be distributed to the target computing power agent related to the root cause of the failure, and the execution results fed back by the target computing power agent can be received.

[0040] The cloud-side intelligent agent first instantiates the matched operation and maintenance policy into a specific, target-device-oriented sequence of instructions. For example, if the matched operation and maintenance policy is disk expansion, the generated instruction could be: "Execute the command on host 192.168.1.10: lvresize -L +50G / dev / mapper / vg01-lv_data". Subsequently, the cloud-side intelligent agent, through a secure and reliable communication channel, such as a channel encrypted with TLS (Transport Layer Security), precisely sends the instruction packet containing this instruction sequence to the target computing power agent directly related to the root cause of the fault. For example, if the aforementioned "disk I / O bottleneck" root cause is located on server D, then the instruction will only be sent to the computing power agent deployed on server D.

[0041] Upon receiving an instruction, the target computing power agent executes it locally. To ensure security, the agent first verifies the legality and permissions of the instruction to prevent the execution of malicious or erroneous commands. After successful verification, the agent invokes local system commands or APIs (Application Programming Interfaces) on its node to execute the instruction. During execution, the agent monitors the instruction's execution status, such as success rate, execution time, and resource consumption. If a timeout or failure occurs, the agent proactively terminates the process and records the exception. After execution, the agent encapsulates the results into a structured feedback report and asynchronously sends it back to the cloud-side agent. This feedback report should at least include the execution status, error messages, and changes in relevant monitoring metrics after execution, such as changes in disk utilization and I / O latency after disk expansion, to allow the cloud-side agent to assess the effectiveness of the measures.

[0042] Step 140: Based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, determine the event case, and based on the event case, update the operation and maintenance strategy in the operation and maintenance strategy library and / or update the artificial intelligence model used for the root cause analysis.

[0043] Specifically, event cases are identified based on operational events, root causes of failures, operational strategies, and execution results. Each event case comprehensively records the event's background, root cause of the failure, the executed operational strategies, and the strategies' effects. Then, based on the event cases, the operational strategies in the operational strategy library are updated, and / or the artificial intelligence model used for root cause analysis is updated.

[0044] For example, the system dynamically adjusts the weight and priority of each strategy in the operation and maintenance strategy library based on the execution results of event case feedback. If the operation and maintenance strategy is executed successfully and achieves the expected results, its weight will be increased accordingly to enhance the probability of its selection in similar scenarios in the future; if the execution results do not meet expectations, its weight will be reduced, and revision suggestions can be issued to operation and maintenance experts to promote the continuous improvement of the strategy.

[0045] For example, if the root cause analysis process uses an artificial intelligence model, newly generated event cases, especially the accurate correspondence between operational events and root causes of failures, will be added to the AI ​​model's training set as high-quality training samples. The system can use these new samples for incremental or full training of the AI ​​model based on a preset period or the amount of accumulated samples, thereby continuously optimizing the AI ​​model's judgment capabilities. This enables it to identify new and rare failure modes, improving the accuracy and timeliness of root cause localization.

[0046] Through the above steps, the industrial computing network intelligent operation and maintenance method constructs a complete closed loop including perception, analysis, decision-making, execution and learning, realizing continuous optimization of the operation and maintenance process and improving the system's autonomous evolution capability.

[0047] The method provided in this embodiment of the invention obtains operation and maintenance events reported by computing power agents; performs root cause analysis on the operation and maintenance events to obtain the root cause of the failure, and matches the operation and maintenance strategy from the operation and maintenance strategy library based on the root cause of the failure; distributes the operation and maintenance strategy to the target computing power agent related to the root cause of the failure, and receives the execution result fed back by the target computing power agent; determines the event case based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, and updates the operation and maintenance strategy in the operation and maintenance strategy library and / or updates the artificial intelligence model used for root cause analysis based on the event case. This invention deploys computing power agents on heterogeneous computing power nodes to uniformly acquire operation and maintenance events. The cloud performs centralized root cause analysis on these events to match operation and maintenance strategies. These strategies are then distributed to the corresponding computing power agents for execution, and the results are fed back. Finally, the operation and maintenance strategy library and artificial intelligence model are updated in a closed loop based on complete event cases. Thus, unified awareness of heterogeneous and distributed computing power nodes is achieved through computing power agents, solving the problem of difficult unified management and monitoring due to resource dispersion. Furthermore, through automated root cause analysis and strategy matching, the root cause of faults can be quickly and accurately located and operation and maintenance solutions can be generated, reducing reliance on manual experience and significantly shortening fault handling time. Moreover, continuous optimization through a closed-loop learning mechanism enables automatic fault discovery, accurate fault location, and execution, greatly improving the level of operation and maintenance automation and ensuring the stability of industrial computing networks.

[0048] Based on the above embodiments, step 120, which involves performing root cause analysis on the maintenance event to obtain the root cause of the failure, includes: Step 1201: Perform fault identification on the operation and maintenance events to obtain fault events; Step 1202: Perform root cause analysis based on the fault event to obtain the root cause of the fault.

[0049] Specifically, the operational events reported by the computing power agent are massive and raw, containing both a large amount of normal operational status information and potentially genuine fault signals. Performing complex root cause analysis directly on all events would consume significant computing resources and be inefficient. Therefore, fault identification can be performed on these operational events to determine the fault events.

[0050] Here, fault identification may include rule matching, anomaly detection, and event aggregation, etc., and the embodiments of the present invention do not specifically limit this.

[0051] Rule matching involves matching reported operational events against a predefined rule base. For example, operations personnel can pre-set a rule: if the CPU utilization metric exceeds 90% for 5 consecutive minutes, a CPU overload failure event will be generated.

[0052] Anomaly detection refers to the use of machine learning algorithms, such as statistical methods based on the 3-sigma principle or unsupervised learning algorithms like Isolation Forest, to analyze metric data in operational events. These algorithms learn historical data patterns of metrics and can automatically identify abnormal fluctuations that deviate from their normal behavior patterns, even if the fluctuation does not reach a fixed, preset threshold.

[0053] Here, event aggregation refers to combining multiple highly correlated original events from the same source, such as the same computing node or with the same context label, into a higher-level, semantically clear composite fault event within a short period of time. For example, if a server reports multiple original events such as insufficient memory, critical process crash, and unresponsive service port within one minute, event aggregation can merge them into a composite fault event indicating that server XXX is suspected of being down.

[0054] Then, root cause analysis is performed based on the failure events to obtain the root cause of the failure.

[0055] The method provided in this invention avoids unnecessary computation and manpower costs on a large number of irrelevant or low-priority operation and maintenance events by introducing a fault identification preprocessing step. Then, it performs root cause analysis based on the fault events to obtain the root cause of the fault, which improves the efficiency and accuracy of the root cause analysis process and enables the system to respond to real faults more quickly.

[0056] Based on the above embodiments, step 1202 includes: Step 1202-1: Obtain the dependency topology related to the fault event; Step 1202-2: Perform index correlation analysis on multiple resource indicators in the dependent topology relationship, and / or perform time-series correlation analysis on multiple abnormal indicators in the dependent topology relationship to obtain the root cause of the failure.

[0057] Specifically, the first step is to identify the dependency topology related to the failure event. When a failure event occurs, such as a spike in the API response time of application A, it is often just a symptom; the root cause may be hidden in other related system components. Therefore, to accurately pinpoint the root cause, it is first necessary to clarify the components involved in the failure event and their inter-component dependency topology.

[0058] Here, dependency topology is used to reflect the calling, dependency, and hosting relationships between various resources such as applications, services, middleware, and infrastructure. For example, an agent can query the Configuration Management Database (CMDB) or obtain the dependency topology of application A through an automatic discovery tool: application A runs on container P, container P depends on service B, service B depends on database C, and database C ultimately runs on physical server D.

[0059] Then, we can perform correlation analysis on multiple resource indicators in the dependency topology, and / or perform time-series correlation analysis on multiple abnormal indicators in the dependency topology to obtain the root cause of the failure.

[0060] Here, the correlation analysis focuses on analyzing the cross-level and cross-component correlations of multiple resource metrics across different nodes in a dependency topology within the same time window. Its basic assumption is that multiple anomalies caused by the same root cause are strongly correlated in time. For example, when analyzing a failure event such as a spike in application A's response time, the agent will examine the key performance indicators of each node along the dependency chain "Application A -> Container P -> Service B -> Database C -> Server D" at the moment the failure occurred. If the analysis finds that at the same moment application A's response time spiked, server D's disk I / O utilization also reached a peak of 100%, while the metrics of other components in the dependency chain remained normal, then the system can preliminarily infer that the root cause of the failure is likely the disk I / O bottleneck of server D.

[0061] Here, temporal correlation analysis focuses on the order in which anomalous indicators occur. Its basic assumption is that the anomalous indicators of the root cause of a fault will appear before the anomalous indicators of the apparent faults they trigger. For example, system analysis reveals that the network outgoing bandwidth of server D reaches saturation at 10:00:01 AM, and subsequently, at 10:00:05 AM, the interface error rate of application A, which relies on the database on this server, begins to rise sharply. This clear temporal sequence provides strong evidence that the network bandwidth saturation of server D is the root cause of the interface errors in application A.

[0062] The method provided in this invention introduces dependent topological relationships and combines index correlation analysis and time-series correlation analysis, which greatly narrows the scope of fault investigation and uses the inherent temporal and spatial correlation of data to infer causal relationships, significantly improving the accuracy and efficiency of root cause localization.

[0063] Based on the above embodiments, the dependency topology relationship includes a dependency topology graph; Step 1202-2, which involves performing correlation analysis on multiple resource indicators in the dependent topology, includes: Steps 1202-21: Query the configuration management database to obtain the dependency graph of the resource; the dependency graph includes multiple dependency nodes and dependency edges between the multiple dependency nodes; Steps 1202-22: Along the dependency edges in the dependency graph, check in turn whether there are any abnormal indicators in the dependency nodes connected to the dependency edges.

[0064] Specifically, in this embodiment, the dependency topology relationship includes a dependency topology graph. A dependency topology graph is a graphical data structure composed of multiple dependent nodes and dependency edges connecting the nodes. A dependent node can represent any resource entity in an industrial computing network, such as an application service, a database instance, a container, or a physical server. Dependency edges represent the dependency relationships between nodes, such as calling, connecting to, or running. This embodiment of the invention does not specifically limit the specific relationships between nodes, and dependency edges can be directed to explicitly indicate the direction of the dependency.

[0065] Accordingly, the configuration management database can be queried to obtain the resource dependency graph. The configuration management database statically stores the configuration information of enterprise IT resources and their interrelationships, serving as the foundational data source for building the dependency graph. Furthermore, dynamic service discovery tools can be used to obtain more real-time dependency topology relationships, supplementing and correcting the information in the CMDB.

[0066] After obtaining the dependency graph, the next step is to check for any abnormal metrics among the dependent nodes connected to those edges. This process can be understood as a traversal search on the dependency graph. For example, for a failure event where the API response time of application A spikes, the starting point of the analysis is the node representing application A. The AI ​​will check downstream dependent nodes one by one along the dependency edges originating from application A. For instance, it will first check the service B node that application A depends on, querying and analyzing various performance metrics of service B at the time of the failure, such as response time and error rate. If the metrics of service B are normal, it will continue along the dependency edge of service B to check the next node, database C, analyzing its connection count, slow queries, and other metrics. This process continues layer by layer until the end of the dependency chain is reached, such as server D, or a clear and explainable abnormal metric is found on a certain node.

[0067] The method provided in this embodiment of the invention concretizes the abstract dependency topology into a traversable dependency graph and defines a clear operation path for checking along the dependency graph, ensuring the logic and coverage of root cause analysis and making it easy to implement in a programmatic manner.

[0068] Based on the above embodiments, step 140, which involves updating the operation and maintenance policies in the operation and maintenance policy library and / or updating the artificial intelligence model used for root cause analysis based on the event cases, includes: Step 141, adjust the weight of the operation and maintenance strategy in the event case in the operation and maintenance strategy library; and / or, Step 142: Use the event cases as new training samples, and update the artificial intelligence model based on the new training samples.

[0069] Specifically, the weight of the operation and maintenance strategy in the event case in the operation and maintenance strategy library can be adjusted, and / or the event case can be used as a new training sample, and the artificial intelligence model can be updated based on the new training sample.

[0070] In this embodiment, each operation and maintenance (O&M) strategy in the O&M strategy library can be associated with a weight value. This weight value can represent the priority, historical success rate, recommendation index, etc., of the O&M strategy. This embodiment does not specifically limit this. Once an event case is identified, the system evaluates the effectiveness of the O&M strategy executed in that event case. If the execution is successful and the fault is effectively resolved, the system can increase the weight value of the O&M strategy. This means that the O&M strategy has proven effective in practice, and it will be more likely to be recommended and executed when the same or similar root causes of faults are matched again in the future. If the execution fails, or the problem is not improved after execution, the system can decrease the weight value of the O&M strategy, or even mark it as pending review, prompting O&M experts to analyze and correct it.

[0071] Here, event cases are used as new training samples, and the AI ​​model is updated based on these new training samples. This operation aims to optimize the AI ​​model. A complete event case, especially the verified mapping relationship between the operation and maintenance event and the root cause of the failure, and the mapping relationship between the root cause of the failure and the effective operation and maintenance strategy, serves as supervised training data for the AI ​​model. The system can add this event case to a dedicated training sample dataset. After accumulating a certain number of new samples, or according to a preset period, the system can trigger a retraining or incremental learning process for the AI ​​model. By learning from these new cases, the AI ​​model can continuously recognize and master new failure modes and new causal relationships, thereby continuously improving its accuracy and generalization ability in future root cause analysis.

[0072] Based on the above embodiments, step 130, which involves distributing the operation and maintenance strategy to the target computing power agent related to the root cause of the fault, includes: Step 131: Perform constraint verification on the operation and maintenance strategy, and distribute the operation and maintenance strategy that passes the constraint verification to the target computing power agent related to the root cause of the fault. The constraint verification includes at least one of the following: Check whether the remaining resources of the computing node corresponding to the target computing power agent meet the requirements of the operation and maintenance strategy; Based on the monitoring data obtained from the target computing power agent, assess whether the current business load on the computing power node allows the operation in the operation and maintenance strategy to be executed.

[0073] Specifically, in this embodiment, before issuing the operation and maintenance strategy, the operation and maintenance strategy will be constrained and verified. Only operation and maintenance strategies that pass the constraint verification can be issued to the target computing power agent.

[0074] Here, constraint verification may include at least one of the following checks: Check whether the remaining resources of the computing node corresponding to the target computing power agent meet the requirements of the operation and maintenance strategy; Based on monitoring data obtained from the target computing power agent, assess whether the current business load on the computing power node allows the execution of operations in the operation and maintenance strategy.

[0075] This involves checking whether the remaining resources of the computing node corresponding to the target computing power agent meet the requirements of the operation and maintenance policy. For example, if the operation and maintenance policy to be executed is to expand the memory of the container instance of application A to 8GB, the constraint verification step will first query the current available memory of the physical machine or virtual machine where the container is located through the computing power agent. If the available memory is less than the required increment, the verification will fail, the operation will be blocked, and a message may be displayed indicating that the target node resources are insufficient. Similarly, if the operation and maintenance policy is to expand the disk by 50GB, it will check whether the corresponding storage volume group has sufficient remaining space.

[0076] The process involves assessing whether the current business load on the computing power node allows the execution of operations within the operational strategy, based on monitoring data obtained from the target computing power agent. Since some operational operations may temporarily interrupt services or impact performance, such as application restarts or database master-slave switching, the constraint verification step evaluates whether the current execution time is appropriate. For example, by analyzing real-time QPS (Queries Per Second) and transaction volume monitoring data of the business applications associated with the computing power node, the current business load is assessed. If the current period is a peak business period, an operational strategy requiring a service restart might be deemed to have excessively high business load and therefore not allowed to execute, resulting in a failed verification.

[0077] The method provided in this invention performs constraint verification on the operation and maintenance strategy, and distributes the operation and maintenance strategy that passes the constraint verification to the target computing power agent related to the root cause of the fault. This significantly improves the security and reliability of automated operation and maintenance, avoids serious consequences such as business interruption or system crash that may be caused by blindly executing disposal instructions when conditions are not met, and improves the stability of intelligent operation and maintenance of industrial computing networks.

[0078] Based on the above embodiments, step 131 further includes the following prior steps: Step 131-1: Assess the execution risk level of the operation and maintenance strategy that has passed the constraint verification, and obtain the current business criticality level of the business unit corresponding to the target computing power agent; Step 131-2: If the execution risk level is low risk, then issue the operation and maintenance strategy; Step 131-3: If the execution risk level is high risk and the business criticality level is high, then generate an execution instruction and issue the operation and maintenance strategy in response to the approval operation of the execution instruction. Step 131-4: If the execution risk level is medium risk, then the operation and maintenance strategy is added to the automatic execution queue of the preset business off-peak maintenance window.

[0079] Specifically, first, assess the execution risk level of the operation and maintenance strategy that has passed the constraint verification, and obtain the current business criticality level of the business unit corresponding to the target computing power agent.

[0080] The risk level can be categorized as low, medium, or high. For example, read-only or benign operations such as querying logs and adding caches can be rated as low risk; operations that may cause brief service fluctuations, such as restarting stateless applications and switching traffic, can be rated as medium risk; while operations such as restarting the database, formatting the disk, and changing core network configurations may be rated as high risk.

[0081] The business criticality level reflects the criticality level of a business unit. For example, the criticality level of a core production line quality inspection system might be high, while the criticality level of a back-end data reporting and analysis system might be low.

[0082] If the risk level is low, the operation and maintenance strategy will be issued. For low-risk operations, the system considers their impact negligible, so after passing the constraint verification, it will be automatically issued and executed to pursue the highest repair efficiency.

[0083] If the risk level is high and the business criticality level is high, a pending instruction is generated. In response to the approval of this pending instruction, an operational strategy is issued. For high-risk operations that are about to impact critical business operations, the system adopts the most cautious approach. It does not execute automatically; instead, it transforms the generated instruction into a pending state and initiates an approval process, such as notifying the relevant operations manager via a work order system or instant messaging tools. Only after obtaining explicit approval from a human expert will the instruction be officially issued for execution.

[0084] If the risk level is medium, the maintenance strategy will be added to the automatic execution queue of the preset low-peak maintenance window. For operations with moderate risk that may have some impact on business, in order to balance repair timeliness and business stability, the system adopts a delayed execution strategy. It will automatically add the maintenance strategy to a preset maintenance window with the lowest business volume, such as the automatic execution queue from 2:00 AM to 4:00 AM, and then automatically execute it when the time window arrives.

[0085] The method provided in this invention introduces a hierarchical management and control strategy based on risk and business criticality, thereby maximizing operational efficiency while ensuring the stability and continuity of core businesses, and thus improving the accuracy and reliability of intelligent operation and maintenance of industrial computing networks.

[0086] Based on any of the above embodiments Figure 2 This is the second flowchart of the intelligent operation and maintenance method for industrial computing networks provided by this invention, as shown below. Figure 2 As shown, the process begins with the edge and endpoint agents of the system performing intelligent perception and event reporting. These agents process the collected data to form standardized events, which are then reported to the cloud-side intelligent agent. Subsequently, the cloud-side intelligent agent enters the intelligent analysis and decision-making stage. This stage can be further subdivided into event identification, boundary and location to determine the root cause of the fault, and finally, decision generation to output an executable instruction sequence. After decision-making, the process enters the precise execution and real-time feedback stage. In this stage, the cloud-side intelligent agent distributes the executable instruction sequence to the relevant computing power agents for execution and receives the execution results. Finally, the entire process enters the closed-loop learning and evolution stage. Based on the full-cycle information of this event, the system updates and optimizes the operation and maintenance knowledge base, AI model module, and operation and maintenance strategy and handling solution module, thus forming a continuously self-evolving intelligent operation and maintenance closed loop. The optimized knowledge is then used to guide future analysis and decision-making processes.

[0087] The intelligent operation and maintenance system for industrial computing networks provided by this invention is described below. The intelligent operation and maintenance system for industrial computing networks described below can be referred to in correspondence with the intelligent operation and maintenance method for industrial computing networks described above.

[0088] Based on any of the above embodiments, the present invention provides an intelligent operation and maintenance system for industrial computing networks. Figure 3 This is a schematic diagram of the structure of the industrial computing network intelligent operation and maintenance system provided by the present invention, as shown below. Figure 3 As shown, the system includes: Acquisition unit 310 is used to acquire the operation and maintenance events reported by the computing power agent; The root cause analysis unit 320 is used to perform root cause analysis on the operation and maintenance event, obtain the root cause of the failure, and obtain the operation and maintenance strategy from the operation and maintenance strategy library based on the root cause of the failure. The distribution unit 330 is used to distribute the operation and maintenance strategy to the target computing power agent related to the root cause of the fault, and to receive the execution result fed back by the target computing power agent; The determining unit 340 is used to determine event cases based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, and update the operation and maintenance strategy in the operation and maintenance strategy library and / or update the artificial intelligence model used for the root cause analysis based on the event cases.

[0089] The system provided in this embodiment of the invention acquires operation and maintenance events reported by computing power agents; performs root cause analysis on the operation and maintenance events to obtain the root causes of the faults, and matches operation and maintenance strategies from the operation and maintenance strategy library based on the root causes of the faults; distributes the operation and maintenance strategies to the target computing power agents related to the root causes of the faults, and receives the execution results fed back by the target computing power agents; determines event cases based on the operation and maintenance events, root causes of the faults, operation and maintenance strategies, and execution results, and updates the operation and maintenance strategies in the operation and maintenance strategy library and / or updates the artificial intelligence model used for root cause analysis based on the event cases. This invention deploys computing power agents on heterogeneous computing power nodes to uniformly acquire operation and maintenance events. The cloud performs centralized root cause analysis on these events to match operation and maintenance strategies. These strategies are then distributed to the corresponding computing power agents for execution, and the results are fed back. Finally, the operation and maintenance strategy library and artificial intelligence model are updated in a closed loop based on complete event cases. Thus, unified awareness of heterogeneous and distributed computing power nodes is achieved through computing power agents, solving the problem of difficult unified management and monitoring due to resource dispersion. Furthermore, through automated root cause analysis and strategy matching, the root cause of faults can be quickly and accurately located and operation and maintenance solutions can be generated, reducing reliance on manual experience and significantly shortening fault handling time. Moreover, continuous optimization through a closed-loop learning mechanism enables automatic fault discovery, accurate fault location, and execution, greatly improving the level of operation and maintenance automation and ensuring the stability of industrial computing networks.

[0090] Based on any of the above embodiments, the root cause analysis unit 320 specifically includes: The fault identification unit is used to identify faults in the operation and maintenance events to obtain fault events. An analysis unit is used to perform root cause analysis based on the fault event to obtain the root cause of the fault.

[0091] Based on any of the above embodiments, the analysis unit specifically includes: Obtain dependency topology unit, used to obtain dependency topology relationships related to the fault event; The fault root cause determination unit is used to perform index correlation analysis on multiple resource indicators in the dependent topology relationship, and / or to perform time-series correlation analysis on multiple abnormal indicators in the dependent topology relationship to obtain the fault root cause.

[0092] Based on any of the above embodiments, the dependency topology relationship includes a dependency topology graph; The fault root cause determination unit is specifically used for: Query the configuration management database to obtain the dependency graph of the resource; the dependency graph includes multiple dependency nodes and dependency edges between the multiple dependency nodes; Along the dependency edges in the dependency graph, check in turn whether there are any abnormal metrics for the dependency nodes connected to the dependency edges.

[0093] Based on any of the above embodiments, the determining unit 340 is specifically used for: Adjust the weight of the operation and maintenance strategy in the event case within the operation and maintenance strategy library; and / or, The event cases are used as new training samples, and the artificial intelligence model is updated based on the new training samples.

[0094] Based on any of the above embodiments, the sending unit 330 specifically includes: The constraint verification unit is used to perform constraint verification on the operation and maintenance strategy, and to send the operation and maintenance strategy that passes the constraint verification to the target computing power agent related to the root cause of the fault. The constraint verification includes at least one of the following: Check whether the remaining resources of the computing node corresponding to the target computing power agent meet the requirements of the operation and maintenance strategy; Based on the monitoring data obtained from the target computing power agent, assess whether the current business load on the computing power node allows the operation in the operation and maintenance strategy to be executed.

[0095] Based on any of the above embodiments, an evaluation unit is further included, wherein the evaluation unit is specifically used for: Assess the execution risk level of the operation and maintenance strategy that has passed the constraint verification, and obtain the current business criticality level of the business unit corresponding to the target computing power agent; If the execution risk level is low risk, then the operation and maintenance strategy will be issued. If the execution risk level is high and the business criticality level is high, then an execution instruction is generated, and in response to the approval operation of the execution instruction, the operation and maintenance strategy is issued. If the execution risk level is medium risk, the operation and maintenance strategy will be added to the automatic execution queue of the preset off-peak maintenance window.

[0096] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an intelligent operation and maintenance method for the industrial computing network. The method includes: acquiring operation and maintenance events reported by the computing power agents; performing root cause analysis on the operation and maintenance events to obtain the root causes of the faults, and matching operation and maintenance strategies from the operation and maintenance strategy library based on the root causes of the faults; distributing the operation and maintenance strategies to target computing power agents related to the root causes of the faults, and receiving the execution results fed back by the target computing power agents; determining event cases based on the operation and maintenance events, the root causes of the faults, the operation and maintenance strategies, and the execution results, and updating the operation and maintenance strategies in the operation and maintenance strategy library and / or updating the artificial intelligence model used for the root cause analysis based on the event cases.

[0097] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0098] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the intelligent operation and maintenance method for industrial computing networks provided by the above methods. The method includes: acquiring operation and maintenance events reported by the computing power agent; performing root cause analysis on the operation and maintenance events to obtain the root cause of the fault, and matching the operation and maintenance strategy from the operation and maintenance strategy library based on the root cause of the fault; distributing the operation and maintenance strategy to the target computing power agent related to the root cause of the fault, and receiving the execution result fed back by the target computing power agent; determining event cases based on the operation and maintenance events, the root cause of the fault, the operation and maintenance strategy, and the execution result, and updating the operation and maintenance strategy in the operation and maintenance strategy library and / or updating the artificial intelligence model used for the root cause analysis based on the event cases.

[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program is implemented to perform the intelligent operation and maintenance method for industrial computing networks provided by the methods described above. The method includes: acquiring operation and maintenance events reported by the computing power agents; performing root cause analysis on the operation and maintenance events to obtain the root causes of the faults, and matching operation and maintenance strategies from an operation and maintenance strategy library based on the root causes of the faults; distributing the operation and maintenance strategies to target computing power agents related to the root causes of the faults, and receiving execution results fed back by the target computing power agents; determining event cases based on the operation and maintenance events, the root causes of the faults, the operation and maintenance strategies, and the execution results, and updating the operation and maintenance strategies in the operation and maintenance strategy library and / or updating the artificial intelligence model used for the root cause analysis based on the event cases.

[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent operation and maintenance method for industrial computing networks, characterized in that, The industrial computing network includes a cloud-side intelligent agent and multiple computing agents deployed on heterogeneous computing power nodes. The method is applied to the cloud-side intelligent agent, and the method includes: Obtain the operation and maintenance events reported by the computing power agent; Root cause analysis is performed on the operation and maintenance events to obtain the root causes of the failures, and operation and maintenance strategies are obtained by matching the root causes of the failures from the operation and maintenance strategy library. The operation and maintenance strategy is distributed to the target computing power agent related to the root cause of the failure, and the execution results are received from the target computing power agent. Based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, an event case is determined, and based on the event case, the operation and maintenance strategy in the operation and maintenance strategy library is updated and / or the artificial intelligence model used for the root cause analysis is updated.

2. The intelligent operation and maintenance method for industrial computing networks according to claim 1, characterized in that, The root cause analysis of the maintenance event to obtain the root cause of the failure includes: The operation and maintenance events are subjected to fault identification to obtain fault events; Root cause analysis is performed based on the fault event to obtain the root cause of the fault.

3. The intelligent operation and maintenance method for industrial computing networks according to claim 2, characterized in that, The root cause analysis of the fault event to obtain the root cause of the fault includes: Obtain the dependency topology relationships related to the fault event; The root cause of the failure is obtained by performing index correlation analysis on multiple resource indicators in the dependent topology relationship and / or performing time-series correlation analysis on multiple abnormal indicators in the dependent topology relationship.

4. The intelligent operation and maintenance method for industrial computing networks according to claim 3, characterized in that, The dependency topology relationship includes a dependency topology graph; The step of performing correlation analysis on multiple resource indicators in the dependent topology includes: Query the configuration management database to obtain the dependency graph of the resource; the dependency graph includes multiple dependency nodes and dependency edges between the multiple dependency nodes; Along the dependency edges in the dependency graph, check in turn whether there are any abnormal metrics for the dependency nodes connected to the dependency edges.

5. The industrial computing network intelligent operation and maintenance method according to any one of claims 1 to 4, characterized in that, The step of updating the operation and maintenance policies in the operation and maintenance policy library and / or updating the artificial intelligence model used for the root cause analysis based on the event cases includes: Adjust the weight of the operation and maintenance strategy in the event case within the operation and maintenance strategy library; and / or, The event cases are used as new training samples, and the artificial intelligence model is updated based on the new training samples.

6. The industrial computing network intelligent operation and maintenance method according to any one of claims 1 to 4, characterized in that, The step of distributing the operation and maintenance strategy to the target computing power agent related to the root cause of the failure includes: The operation and maintenance strategy is constrained and verified, and the operation and maintenance strategy that passes the constraint verification is distributed to the target computing power agent related to the root cause of the fault. The constraint verification includes at least one of the following: Check whether the remaining resources of the computing node corresponding to the target computing power agent meet the requirements of the operation and maintenance strategy; Based on the monitoring data obtained from the target computing power agent, assess whether the current business load on the computing power node allows the operation in the operation and maintenance strategy to be executed.

7. The industrial computing network intelligent operation and maintenance method according to claim 6, characterized in that, The step of distributing the operation and maintenance strategy that has passed the constraint verification to the target computing power agent related to the root cause of the failure also includes the following before: Assess the execution risk level of the operation and maintenance strategy that has passed the constraint verification, and obtain the current business criticality level of the business unit corresponding to the target computing power agent; If the execution risk level is low risk, then the operation and maintenance strategy will be issued. If the execution risk level is high and the business criticality level is high, then an execution instruction is generated, and in response to the approval operation of the execution instruction, the operation and maintenance strategy is issued. If the execution risk level is medium risk, the operation and maintenance strategy will be added to the automatic execution queue of the preset off-peak maintenance window.

8. An industrial computing network intelligent operation and maintenance system, characterized in that, include: The acquisition unit is used to acquire the operation and maintenance events reported by the computing power agent; The root cause analysis unit is used to perform root cause analysis on the operation and maintenance event, obtain the root cause of the failure, and obtain the operation and maintenance strategy from the operation and maintenance strategy library based on the root cause of the failure. The distribution unit is used to distribute the operation and maintenance strategy to the target computing power agent related to the root cause of the fault, and to receive the execution results fed back by the target computing power agent; The determining unit is configured to determine event cases based on the operation and maintenance event, the root cause of the failure, the operation and maintenance strategy, and the execution result, and update the operation and maintenance strategy in the operation and maintenance strategy library and / or update the artificial intelligence model used for the root cause analysis based on the event cases.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the intelligent operation and maintenance method for industrial computing networks as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the intelligent operation and maintenance method for industrial computing networks as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent computing cloud platform providing computing power based on multi-agent intelligent operation and maintenance method and device

    CN122363986A