Operation and maintenance prediction and self-repairing method, system and equipment and storage medium

By fusing multi-source heterogeneous data and constructing a dynamic knowledge graph using deep graph neural networks, combined with reinforcement learning algorithms, a panoramic digital mapping and autonomous repair of IT operations and maintenance systems have been achieved. This solves the problems of single data and passive response in traditional operations and maintenance models, improves the accuracy of fault prediction and self-repair efficiency, and reduces operations and maintenance costs.

CN120973566APending Publication Date: 2025-11-18YANCHI ZHONGYING CHUANGNENG NEW ENERGY CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511061182.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional IT operations and maintenance models rely on manual monitoring and rule-based alarm systems, resulting in single data dimensions, passive fault response, lack of predictive and autonomous decision-making capabilities, and inability to effectively cope with complex multi-system and multi-component interconnected faults. Furthermore, existing AIOps solutions have failed to form an automated closed loop from prediction to self-healing.

Method used

By employing multi-source heterogeneous data fusion, deep graph neural networks, and reinforcement learning algorithms, a dynamic knowledge graph is constructed to achieve fault prediction and autonomous repair. This includes data acquisition, parsing, format conversion, dynamic graph construction, fault node identification, repair strategy generation, and execution, forming an automated closed-loop optimization process.

Benefits of technology

It achieves a panoramic digital mapping of the operation and maintenance system, improves the accuracy of fault prediction by more than 40%, reduces the fault handling time by 70%, stabilizes the self-repair success rate at more than 95%, and reduces the operation and maintenance manpower cost by 60%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973566A_ABST
    Figure CN120973566A_ABST
Patent Text Reader

Abstract

The invention discloses an operation and maintenance prediction and self-repairing method, system and device and a storage medium, and the method comprises the following steps: collecting structured, semi-structured and non-structured data, and carrying out the analysis and format conversion of the collected data, and obtaining the preprocessed data; based on the preprocessed data, constructing a dynamic knowledge graph containing entities, relationships, attributes and timestamps; a depth map neural network model is applied to analyze the dynamic knowledge graph, node features are aggregated, key fault nodes are identified, and potential faults are predicted; when the potential fault is predicted, a reinforcement learning algorithm is applied, and a repair strategy is generated according to the operation and maintenance cost, the service influence and the repair duration; and executing the restoration strategy, and feeding back an execution result and performance index data in the restoration process to update the dynamic knowledge graph and depth graph neural network model. Panoramic digital mapping of the operation and maintenance system is realized, and the fault prediction accuracy and the self-repairing success rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of computer operation and maintenance, and relates to an operation and maintenance prediction and self-repairing method, system, device and storage medium. BACKGROUND

[0002] With the rapid development of information technology, especially the wide application of technologies such as cloud computing, microservices and containerization, the IT (information technology) system architecture of modern enterprises is increasingly complex and dynamic. Servers, virtual machines, containers, network devices, application programs and their calling relationships constitute a huge and ever-changing ecosystem, which brings unprecedented challenges to the stable operation and efficient operation and maintenance of the system.

[0003] The traditional IT operation mode mainly relies on manual monitoring and rule-based alarm systems. The operation and maintenance personnel set static thresholds through monitoring tools (such as Zabbix, Nagios, etc.), and when a certain performance indicator (such as CPU usage, memory occupancy) exceeds the preset value, an alarm is triggered. This mode has the following significant defects: Firstly, the data dimension is single and cannot form a global view. Traditional monitoring mainly focuses on isolated, structured performance indicator data, while ignoring a large amount of semi-structured log data and unstructured data (such as configuration change records, work order texts, etc.) that contain rich context information. This fragmentation of data makes it difficult for operation and maintenance personnel to understand the real running state of the system and the complex dependency relationship between components from a global perspective, and it is difficult to effectively discover the correlation faults that span multiple systems and components.

[0004] Secondly, the fault response is passive and the processing efficiency is low. The alarm mechanism based on static thresholds belongs to "after-the-fact response", that is, the problem can only be discovered after it has occurred and produced an impact. In addition, when the system has complex faults, the alarm storm (a large number of alarms emerging at the same time) often makes the operation and maintenance personnel at a loss, and a large amount of time is needed for manual troubleshooting, correlation analysis and root cause positioning, not only the processing efficiency is low, but also the business may be interrupted for a long time due to human error, causing huge economic losses.

[0005] Thirdly, it lacks prediction ability and autonomous decision-making ability. Most existing technologies are limited to monitoring and alarming the current or historical state, and cannot effectively predict the future running trend of the system, and cannot give early warning of potential fault risks. After the fault occurs, the repair means is highly dependent on the personal experience and knowledge base of the operation and maintenance personnel, and lacks intelligent decision support. For unknown or complex fault scenarios, multiple experts often need to be consulted, and the repair process is long and uncertain.

[0006] To solve the above problems, the industry began to try to introduce artificial intelligence technology, namely AIOps (AI for ITOperations). However, the existing AIOps scheme still has limitations. Some schemes only apply machine learning algorithms to a single scenario, such as log anomaly detection or time series data prediction, and fail to effectively fuse multi-source heterogeneous data; other schemes attempt to build system topology, but are mostly static or slow to update, and cannot adapt to the dynamic changes in system architecture in a cloud-native environment. More importantly, most schemes still require human intervention for repair decisions and execution after fault prediction, and fail to form an automated closed loop from prediction to self-healing. SUMMARY

[0007] The purpose of the present application is to overcome the above-mentioned shortcomings of the prior art, and to provide an operation and maintenance prediction and self-repair method, system, device and storage medium, which realizes panoramic digital mapping of the operation and maintenance system and improves the fault prediction accuracy and self-repair success rate.

[0008] To achieve the above-mentioned purpose, the present application adopts the following technical solutions: An operation and maintenance prediction and self-repair method, comprising the following processes: Collecting structured, semi-structured and unstructured data, and analyzing and converting the collected data to obtain preprocessed data; Based on the preprocessed data, a dynamic knowledge graph containing entities, relationships, attributes and timestamps is constructed; A deep graph neural network model is applied to analyze the dynamic knowledge graph, aggregate node features, identify key fault nodes, and predict potential faults; When a potential fault is predicted, a reinforcement learning algorithm is applied to generate a repair strategy based on operation and maintenance costs, business impact and repair duration; The repair strategy is executed, and the execution results and performance indicator data during the repair process are fed back to update the dynamic knowledge graph and the deep graph neural network model.

[0009] Preferably, after collecting the data, the data is also subjected to quality evaluation by entropy calculation and outlier detection algorithms to filter out noise data.

[0010] Preferably, the process of constructing the dynamic knowledge graph is: applying a dynamic Bayesian network algorithm to automatically update the topology structure of the graph according to real-time data changes; and applying a time series model to predict the future trend of entity attribute changes and update the node state in advance.

[0011] Preferably, the process of applying the deep graph neural network model is: applying a graph aggregation algorithm at the bottom to aggregate node features and learn local neighborhood information of the nodes; and applying a graph attention network mechanism at the top to dynamically allocate node attention weights according to fault propagation characteristics.

[0012] Preferably, after generating the repair strategy, the application of the generative adversarial network is also included to verify the repair strategy, and the strategy with the highest success rate and the smallest resource consumption is selected.

[0013] Preferably, the process of executing the repair strategy is: generating an automated tool script, remotely executing a configuration modification operation; and deploying a monitoring probe to collect performance index changes in real time during the repair process.

[0014] Preferably, the process of feeding back the execution result and the performance index data is: taking the execution result and the performance index data as new samples, and applying a transfer learning technique to update the model parameters of the deep graph neural network model.

[0015] An operation and maintenance prediction and self-repair system, comprising: A multi-source heterogeneous data fusion and collection module for collecting structured, semi-structured and unstructured data, and parsing and format converting the collected data to obtain preprocessed data; An adaptive dynamic knowledge graph construction module for constructing a dynamic knowledge graph containing entities, relationships, attributes and timestamps based on the preprocessed data; A deep graph neural network prediction analysis module for applying a deep graph neural network model to analyze the dynamic knowledge graph, aggregate node features, identify key fault nodes, and predict potential faults; An intelligent decision-making and risk assessment module for generating a repair strategy based on operation and maintenance costs, business impact and repair duration when a potential fault is predicted by applying a reinforcement learning algorithm; An automated closed-loop execution and intelligent optimization module for executing the repair strategy and feeding back the execution result and performance index data during the repair process to update the dynamic knowledge graph and the deep graph neural network model.

[0016] A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the operation and maintenance prediction and self-repair method when executing the computer program.

[0017] A computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the steps of the operation and maintenance prediction and self-repair method.

[0018] Compared with the prior art, the present application has the following beneficial effects: 1. Super-dimensional data insight: Break through the limitations of traditional single-dimensional data, fuse structured, semi-structured and unstructured data, construct a dynamic knowledge graph containing millions of nodes and relationships, realize panoramic digital mapping of the operation and maintenance system, and improve the fault prediction accuracy by more than 40%.

[0019] 2. Intelligent dynamic decision-making: Deep graph neural network and reinforcement learning algorithm cooperate to realize the leap from "passive response" to "active prevention" of fault prediction, which can provide 2-hour early warning of potential faults and generate a globally optimal repair strategy, reducing the average fault handling time by 70%.

[0020] 3. Autonomous evolution capability: Through data feedback driven model iteration, the system continuously learns new fault patterns and repair experience during operation and maintenance, and after more than 100,000 training iterations, the self-repair success rate is stably maintained at more than 95%.

[0021] 4. Significant cost-effectiveness: The fully automated operation and maintenance process greatly reduces manual intervention, reducing operation and maintenance labor costs by 60%; precise fault prediction and efficient repair avoid business interruption losses, and it is estimated that the enterprise can save millions of operation and maintenance costs per year. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a flow chart of the operation and maintenance prediction and self-repair method of the embodiment of the present application; Figure 2 is a link process schematic diagram of the operation and maintenance prediction and self-repair system of the embodiment of the present application. DETAILED DESCRIPTION

[0023] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] As shown in Figure 1 , the operation and maintenance prediction and self-repair method described in the embodiment includes the following processes: (I) Collect structured, semi-structured and unstructured data, and parse and format the collected data to obtain preprocessed data.

[0026] After collecting data, the quality of the data is evaluated by entropy calculation and outlier detection algorithm to filter out noise data.

[0027] (II) Based on the preprocessed data, a dynamic knowledge graph containing entities, relationships, attributes and timestamps is constructed.

[0028] The process of constructing a dynamic knowledge graph is: applying a dynamic Bayesian network algorithm to automatically update the topology of the graph according to real-time data changes; and applying a time series model to predict the future trend of entity attributes and update the node state in advance.

[0029] (III) Apply a deep graph neural network model to analyze the dynamic knowledge graph, aggregate node features, identify key fault nodes, and predict potential faults.

[0030] The process of applying a deep graph neural network model is: applying a graph aggregation algorithm at the bottom to aggregate node features and learn local neighborhood information of nodes; and applying a graph attention network mechanism at the top to dynamically allocate node attention weights according to fault propagation characteristics.

[0031] (IV) When potential faults are predicted, apply reinforcement learning algorithm to generate repair strategies according to operation and maintenance cost, business impact and repair duration.

[0032] After generating the repair strategy, apply the generative adversarial network to verify the repair strategy and select the strategy with the highest success rate and the smallest resource consumption.

[0033] (V) Execute the repair strategy and feed back the execution results and performance indicator data during the repair process to update the dynamic knowledge graph and the deep graph neural network model.

[0034] The process of executing the repair strategy is: generating an automated tool script to remotely execute configuration modification operations; and deploying monitoring probes to collect real-time performance indicator changes during the repair process.

[0035] The process of feeding back the execution results and performance indicator data is: taking the execution results and performance indicator data as new samples and applying transfer learning technology to update the model parameters of the deep graph neural network model.

[0036] The specific process of the above method is: (I) System deployment In the enterprise data center, the multi-source heterogeneous data fusion and collection module is deployed in the form of a Docker container on each server node and network device, and is managed and scheduled uniformly through a Kubernetes cluster. The adaptive dynamic knowledge graph construction module, the deep graph neural network prediction analysis module, and the intelligent decision-making and risk assessment module are deployed on a high-performance GPU server cluster, and use a distributed storage architecture (Ceph) to store graph data and model parameters. The automated closed-loop execution and intelligent optimization module is deployed on an operation and maintenance control center server and interacts with each business system through an API interface.

[0037] In this embodiment, the logical architecture of the system is closely combined with the physical architecture. Logically, each module is designed in a way of high cohesion and low coupling, and communicates through a high-performance remote procedure call framework such as gRPC to ensure the efficiency and stability of the interaction between modules. Physically, the containerized deployment of the collection agent ensures lightweight resource occupation and rapid scaling; the core computing modules are deployed centrally on a GPU cluster to fully utilize the advantages of GPU in parallel computing and accelerate the training and inference process of graph neural networks; and the selection of a distributed storage architecture is to cope with the massive storage demand of knowledge graph data and model parameters as they grow over time, and to ensure the high availability and fault tolerance of data.

[0038] (II) Data collection and processing Taking a server cluster of an e-commerce platform as an example, the collection module collects hardware indicators such as server CPU usage, memory occupation, and disk I / O every 500 milliseconds; application program logs are collected in real time using Flume, and key error information is extracted using regular expressions; network traffic data is captured using Wireshark, and API call records are obtained by analyzing HTTP and TCP protocol packets. After data collection, real-time cleaning and conversion are performed using the Flink streaming computing framework to remove duplicate data and invalid fields, and the data is transmitted to the graph construction module after being unified in format.

[0039] In the specific processing flow, in order to ensure the quality and availability of data, the preprocessing step also includes more refined operations. First, the data streams from different sources are strictly aligned according to the timestamp, for example, unified to a second-level time window, to ensure the accuracy of the causal relationship in subsequent analysis. Second, the collected structured time series indicators are normalized, for example, using Min-Max Scaling to map their numerical values to the [0, 1] interval, to eliminate the dimensional differences between different indicators and facilitate uniform processing by the model. Finally, for unstructured text data such as logs, after extracting key fields using regular expressions, word embedding models such as Word2Vec or BERT are further used to convert text information such as error descriptions into high-dimensional numerical vectors, thereby integrating unstructured information into subsequent mathematical models.

[0040] (Three) Graph construction and update After the graph construction module receives the data, it first extracts entities (such as server names, application IDs) from the log text through named entity recognition (NER) technology, and determines the relationship between entities (such as "call" "dependence") using dependency syntax analysis. When detecting that a micro-service application has added API interface call relationship, through graph database transaction operation, new edges are created in the graph, and the attribute information of related nodes is updated. At the same time, based on the time series data prediction model, the prediction value of the node attributes such as server load and network bandwidth is automatically updated every hour to ensure that the graph reflects the system running state in real time.

[0041] To further enhance the expression ability and real-time performance of the graph, the construction and update mechanism of the graph is deepened in this embodiment. In the construction phase, a clear graph schema is defined, which specifies the core entity types (such as host, service, database, API interface) and relationship types (such as deployed, dependent on, call), providing a unified framework for data fusion. In the update mechanism, a combination of event-driven and periodic refreshing is adopted. Event-driven update refers to the real-time change of the graph topology structure when the CI / CD system releases new services or network configuration changes. Periodic refreshing refers to the system aggregating time series data within a time window every minute, calculating the dynamic attributes of the nodes (such as "average request delay in the past five minutes"), and updating them to the graph. This hybrid update mechanism ensures that the graph can respond to system changes in real time and fully reflect the long-term running state of the system.

[0042] (Four) Fault prediction and analysis The prediction and analysis module performs full analysis on the dynamic knowledge graph every morning, using the GNN model to mine the potential association between nodes. For example, when it is found that the CPU usage of a database server node is continuously rising, and multiple application server nodes connected to the node show a trend of increasing request delay, the fault propagation simulation engine predicts that the database server may have a performance bottleneck in the next 4 hours. Combined with the Transformer time series prediction module, the CPU usage is predicted for the next 6 hours, and if the prediction value exceeds 80%, the warning mechanism is triggered.

[0043] In a specific prediction analysis process, the training of the GNN model is a key link. The system takes the graph snapshot in the historical data as the input feature, and takes whether the node actually fails in a specific time window (e.g. 4 hours) in the future as the label (0 or 1) to generate training samples. After the model training is completed, in the prediction, the GNN outputs an "abnormal score" in the interval [0, 1] for each node in the graph. When the abnormal score of a node exceeds a preset threshold (e.g. 0.7), the system will mark it as a high-risk node. Subsequently, the Transformer time series prediction module will predict the future trend of the key performance indicators (KPI) of the high-risk node. Only when the predicted trend will also break through the dynamically set health threshold, the system will formally trigger the failure warning. In addition, after the warning is triggered, the system can identify the upstream nodes that contribute most to the warning node by backtracking the attention weights of the graph attention network (GAT) in the GNN model, so as to locate the most possible failure propagation path and potential root cause.

[0044] (V) Repair decision and execution When the prediction module issues a failure warning, the decision module first determines whether there is a matching simple repair rule through the rule engine. If not, it starts the reinforcement learning algorithm to search for the optimal repair scheme in the strategy space. For example, for the database performance bottleneck problem, candidate strategies such as "increase the size of the database connection pool" and "migrate part of the business to the standby database" are generated, and the strategy with the highest success rate and the smallest resource consumption is selected through GAN simulation verification. After receiving the repair instruction, the execution module automatically generates an Ansible playbook script, remotely logs into the database server to perform configuration modification operations, and monitors the performance indicator changes in real time during the repair process. After the repair is completed, the execution result and performance data are fed back to the system for model optimization and graph update.

[0045] When starting the reinforcement learning algorithm to make decisions, the system formalizes the definition of the problem: the state of the current warning node and its adjacent subgraph (including the performance indicators of each node, the anomaly score, etc.) is defined as the "state" of the environment; a combination of a series of atomic operations (such as restarting services, expanding instances, rolling back configurations, etc.) is defined as the "action space"; and a comprehensive "reward function" is designed, which positively encourages business recovery while negatively punishing resource consumption and repair time. Through interaction with the environment, Q-learning and other algorithms learn a strategy that maximizes cumulative rewards. After the repair is complete, the complete record of this incident, including fault characteristics, repair strategies, execution results, and performance change data, is packaged into a structured feedback sample, which is used to incrementally update the knowledge graph and as new training data to fine-tune the GNN prediction model and the reinforcement learning decision-making model, forming a complete, automated closed-loop optimization process.

[0046] The following is an apparatus embodiment of the present application, which can be used to perform the method embodiments of the present application. For details not covered in the apparatus embodiment, please refer to the method embodiments of the present application.

[0047] As shown in Figure 2 In another embodiment of the present application, a system for operation and maintenance prediction and self-repair is provided, which can be used to implement the operation and maintenance prediction and self-repair method described above. Specifically, the system for operation and maintenance prediction and self-repair includes a multi-source heterogeneous data fusion and collection module, an adaptive dynamic knowledge graph construction module, a deep graph neural network prediction and analysis module, an intelligent decision-making and risk assessment module, and an automated closed-loop execution and intelligent optimization module.

[0048] The multi-source heterogeneous data fusion and collection module is used to collect structured, semi-structured and unstructured data, and to analyze and format the collected data to obtain preprocessed data.

[0049] The multi-source heterogeneous data fusion and collection module uses edge computing and distributed collection architecture, and deploys lightweight collection agents on server, network equipment, storage terminal and other data source nodes. For structured log data (such as operating system event logs), semi-structured network traffic data (JSON format API call records) and unstructured user behavior data (text format operation logs), regular expressions and natural language processing techniques are used for unified analysis and format conversion. At the same time, a data quality evaluation mechanism is set up to filter noise data in real time through entropy calculation and outlier detection algorithms, ensuring the accuracy and integrity of the collected data.

[0050] The adaptive dynamic knowledge graph construction module constructs a dynamic knowledge graph containing entities, relationships, attributes, and timestamps based on pre-processed data.

[0051] The adaptive dynamic knowledge graph construction module constructs a multi-dimensional dynamic graph based on the Neo4j graph database, innovatively adopting a "entity-relation-attribute-timestamp" quadruple model. The dynamic Bayesian network algorithm is introduced to automatically update the graph topology according to real-time data changes: when changes in application program call relationships are detected, the Dijkstra shortest path algorithm is used to recalculate the impact path and adjust the weight of the edge; the LSTM time series model is used to predict the future trend of entity attributes (such as server load) and update the node state in advance. At the same time, a semantic understanding engine is constructed to parse fault descriptions in log texts based on the BERT pre-training model, automatically extract new entities and relationships, and realize the autonomous evolution of the graph.

[0052] The deep graph neural network prediction analysis module is used to analyze the dynamic knowledge graph using a deep graph neural network model, aggregate node features, identify key fault nodes, and predict potential faults.

[0053] A double-layer GNN architecture is constructed, with the bottom layer using the GraphSAGE algorithm for node feature aggregation to learn the local neighborhood information of nodes; the upper layer uses the GAT (graph attention network) mechanism to dynamically allocate node attention weights according to the fault propagation characteristics, accurately identifying key fault nodes. Combined with the Transformer time series prediction module, multi-step prediction of performance indicators such as CPU usage and network latency is performed, and a prediction error compensation model is constructed through historical data training to improve the prediction accuracy to more than 98%. In addition, a fault propagation simulation engine is designed based on the Monte Carlo simulation algorithm to deduce the fault diffusion path in a virtual environment and evaluate the potential impact range.

[0054] The intelligent decision-making and risk assessment module is used to generate a repair strategy based on operation and maintenance costs, business impact, and repair duration when a potential fault is predicted using a reinforcement learning algorithm.

[0055] A hierarchical decision tree model is established, with the bottom layer based on a rule engine to quickly handle high-frequency simple faults (such as automatically cleaning temporary files when disk space is insufficient); the upper layer uses a reinforcement learning algorithm to optimize the operation and maintenance costs, business impact, and repair duration, and searches for the optimal repair solution in the strategy space through the Q-learning algorithm. An adversarial generative network (GAN) is introduced to verify the repair strategy, with the generator simulating fault scenarios and the discriminator evaluating the effectiveness of the strategy to improve the reliability of the strategy through adversarial training. At the same time, a risk assessment matrix is constructed to quantitatively score the repair strategy from the dimensions of business continuity, data security, and resource consumption, ensuring safe and reliable decision-making.

[0056] The automatic closed-loop execution and intelligent optimization module is used for executing the repair strategy and feeding back execution results and performance index data in the repair process to update the dynamic knowledge graph and the deep graph neural network model.

[0057] The Ansible automation tool is used to implement cross-platform execution of the repair strategy, and supports script automatic distribution, service start and stop, configuration modification and other operations. Real-time monitoring probes are deployed on the execution nodes, and the combination of Prometheus and Grafana is used to collect performance index changes in the repair process in real time, the Kalman filtering algorithm is used to smooth the data, and abnormal fluctuations are found in time. After the repair is completed, the execution results and performance change data are fed back to the graph construction module and the prediction analysis module as new samples, and the model parameters are quickly updated using the transfer learning technology to continuously optimize the system operation and maintenance capability.

[0058] In another embodiment of the present application, a terminal device is provided, which includes a processor and a memory, the memory being used to store a computer program, the computer program including program instructions, and the processor being used to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions; the processor in the embodiments of the present application can be used for operation of the operation and maintenance prediction and self-repair method, including: collecting structured, semi-structured and unstructured data, and analyzing and format-converting the collected data to obtain preprocessed data; based on the preprocessed data, a dynamic knowledge graph containing entities, relationships, attributes and timestamps is constructed; a deep graph neural network model is applied to analyze the dynamic knowledge graph, aggregate node features, identify key fault nodes, and predict potential faults; when the potential faults are predicted, a reinforcement learning algorithm is applied to generate a repair strategy according to operation and maintenance costs, business impacts and repair time lengths; the repair strategy is executed, and execution results and performance index data in the repair process are fed back to update the dynamic knowledge graph and the deep graph neural network model.

[0059] In another embodiment, the present application also provides a computer readable storage medium (Memory) which is a memory device in the terminal device, used for storing programs and data. It can be understood that the computer readable storage medium herein can include the built-in storage medium in the terminal device, and of course can also include the expansion storage medium supported by the terminal device. The computer readable storage medium provides a storage space which stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can include any entity or device, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory) capable of carrying the computer program code.

[0060] The one or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the operation and maintenance prediction and self-repair method in the above embodiments. The one or more instructions in the computer readable storage medium are loaded and executed by the processor to perform the following steps: collecting structured, semi-structured and unstructured data, and parsing and format converting the collected data to obtain preprocessed data; based on the preprocessed data, constructing a dynamic knowledge graph containing entities, relationships, attributes and timestamps; applying a deep graph neural network model to analyze the dynamic knowledge graph, aggregate node features, identify key fault nodes, and predict potential faults; when predicting potential faults, applying a reinforcement learning algorithm to generate a repair strategy according to operation and maintenance costs, business impact and repair duration; executing the repair strategy and feeding back the execution result and performance index data in the repair process to update the dynamic knowledge graph and the deep graph neural network model.

[0061] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer usable program code.

[0062] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0063] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0064] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0065] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0066] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0067] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0068] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or also can be distributed to multiple units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0069] The above description is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can also be made. These improvements and refinements should be considered as the protection scope of the present application.

[0070] It should be understood that the above description is for illustration only and is not intended to be limiting. Many embodiments and many applications other than those described herein will be readily apparent to those skilled in the art from this description. The scope of the application should therefore not be determined with reference to the above description, but instead should be determined with reference to the appended claims along with their full scope of equivalents. For purposes of completeness, all articles and references, including patents and patent documents, are incorporated herein by reference in their entirety. The disclosure of any aspect of the subject matter disclosed herein that is omitted from any particular claim is hereby incorporated into that claim as if fully set forth in the claim.

Claims

1. A method for predictive maintenance and self-repair, characterized in that, Includes the following processes: Collect structured, semi-structured, and unstructured data, and parse and convert the collected data to obtain preprocessed data; Based on the preprocessed data, a dynamic knowledge graph containing entities, relationships, attributes, and timestamps is constructed. A deep graph neural network model is applied to analyze dynamic knowledge graphs, aggregate node features, identify key fault nodes, and predict potential faults. When a potential failure is predicted, a reinforcement learning algorithm is applied to generate a repair strategy based on the operation and maintenance cost, business impact and repair time. The system executes the repair strategy and feeds back the execution results and performance metrics data during the repair process to update the dynamic knowledge graph and deep graph neural network model.

2. The operation and maintenance prediction and self-repair method according to claim 1, characterized in that, After data collection, the data quality is assessed using entropy calculation and outlier detection algorithms to filter out noisy data.

3. The operation and maintenance prediction and self-repair method according to claim 1, characterized in that, The process of constructing a dynamic knowledge graph involves: applying a dynamic Bayesian network algorithm to automatically update the graph's topology based on real-time data changes; and applying a time series model to predict future trends in entity attributes and update node states in advance.

4. The operation and maintenance prediction and self-repair method according to claim 1, characterized in that, The process of applying the deep graph neural network model is as follows: the bottom layer uses a graph aggregation algorithm to aggregate node features and learn the local neighborhood information of the nodes; the upper layer uses a graph attention network mechanism to dynamically allocate node attention weights according to the fault propagation characteristics.

5. The operation and maintenance prediction and self-repair method according to claim 1, characterized in that, After generating the repair strategy, an adversarial generative network is applied to verify the repair strategy, and the strategy with the highest success rate and the lowest resource consumption is selected.

6. The operation and maintenance prediction and self-repair method according to claim 1, characterized in that, The process of implementing the remediation strategy involves: generating automated tool scripts to remotely execute configuration modification operations; and deploying monitoring probes to collect performance metric changes in real time during the remediation process.

7. The operation and maintenance prediction and self-repair method according to claim 1, characterized in that, The process of providing feedback on execution results and performance metrics data is as follows: using the execution results and performance metrics data as new samples, and applying transfer learning techniques to update the model parameters of the deep graph neural network model.

8. A predictive and self-healing system for operation and maintenance, characterized in that, include: The multi-source heterogeneous data fusion acquisition module is used to acquire structured, semi-structured and unstructured data, and to parse and convert the acquired data to obtain pre-processed data. The adaptive dynamic knowledge graph construction module builds a dynamic knowledge graph containing entities, relationships, attributes, and timestamps based on preprocessed data. The Deep Graph Neural Network Prediction and Analysis Module is used to analyze dynamic knowledge graphs using deep graph neural network models, aggregate node features, identify key fault nodes, and predict potential faults. The intelligent decision-making and risk assessment module is used to generate a repair strategy based on operation and maintenance costs, business impact and repair time when a potential failure is predicted, by applying reinforcement learning algorithms. The automated closed-loop execution and intelligent optimization module is used to execute the repair strategy and feed back the execution results and performance index data during the repair process to update the dynamic knowledge graph and deep graph neural network model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the operation and maintenance prediction and self-repair method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the operation and maintenance prediction and self-repair method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Storage system fault repairing method and device, electronic equipment, medium and product

    CN121478539A

  • Fault processing method and device and electronic equipment

    CN121690964A