Intelligent operation and maintenance methods, systems, equipment and media based on deep learning
Through intelligent operation and maintenance methods based on deep learning, system failures are predicted and automatic repairs are triggered, and the problems of passive response and lack of self-repair in traditional operation and maintenance methods are solved, and efficient and intelligent operation and maintenance management is achieved.
Patent Information
- Application Number
- CN202411593641.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-08
AI Technical Summary
Traditional operation and maintenance methods cannot predict system failures in advance, resulting in passive responses, lack of intelligent analysis and self-repair capabilities, and increase the time for troubleshooting and resolution.
Using an intelligent operation and maintenance method based on deep learning, we train a fault prediction model, predict the failure probability, and trigger automatic operation and maintenance repair processing when the probability is greater than the preset threshold.
It realizes early prediction and early warning of system failures, introduces automatic repair functions, reduces manual intervention and fault handling time, and improves operation and maintenance efficiency and system stability.
Smart Images

Figure CN119168626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of operation and maintenance technology, and in particular to an intelligent operation and maintenance method, system, equipment and medium based on deep learning. Background Art
[0002] Operation and maintenance refers to the maintenance of established network hardware and software in large organizations. In essence, it is the operation and maintenance of networks, servers, and services at all stages of their life cycle to achieve an acceptable state in terms of cost, stability, and efficiency.
[0003] At present, with the popularization of Internet applications and the development of technologies such as cloud computing and big data, the scale and complexity of information systems are increasing. Traditional operation and maintenance methods mainly rely on manual monitoring and post-processing. When a system fails, it can only be learned through the alarm system. This passive response method is prone to business interruption and reduced user experience. Summary of the invention
[0004] The purpose of the present invention is to provide an intelligent operation and maintenance method, system, device and medium based on deep learning to solve the problem in related technologies that operation and maintenance failures can only be handled passively after the failure occurs.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] First aspect: The present invention provides an intelligent operation and maintenance method based on deep learning, including: obtaining first historical operation and maintenance data, and preprocessing the first historical operation and maintenance data to obtain second historical operation and maintenance data; inputting the second historical operation and maintenance data into a preset deep learning-based fault prediction model for training, and bringing it to the trained fault prediction model; obtaining first real-time operation and maintenance data, and preprocessing the first real-time operation and maintenance data to obtain second real-time operation and maintenance data; inputting the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of failure; if the probability of failure is greater than a preset fault threshold, triggering automatic operation and maintenance repair processing.
[0007] Preferably, the method further comprises: if the automatic operation and maintenance repair process fails, sending a warning message to the operation and maintenance management personnel.
[0008] Preferably, the method further comprises: receiving voice instructions from the operation and maintenance manager, and performing operation and maintenance fault analysis and location according to the voice instructions; and generating a recommended operation and maintenance solution according to the results of the fault analysis and location.
[0009] Preferably, the historical operation and maintenance data includes: system log data, monitoring data and performance indicator data.
[0010] Preferably, the preprocessing includes: data cleaning, data standardization and data format unification.
[0011] Preferably, the fault prediction model based on deep learning adopts the following calculation method:
[0012] a. Input data definition:
[0013] Historical operation and maintenance data matrix X: X∈R N×T×F represents the value of N operation and maintenance indicator data in the past T time windows, where F is the feature dimension; X n,t,f represents the fth eigenvalue of the nth operation and maintenance indicator data in the tth time window;
[0014] Fault label vector Y: Y∈0,1 T , indicating whether a failure will occur at some point in the future;
[0015] b. Multi-mode data fusion:
[0016]
[0017] in, is the weight coefficient of the mth mode; is the scoring function; W m is the weight matrix; b m is the bias term; tanh is the activation function; H t is the fused hidden state vector; is the data of the mth mode at time t;
[0018] c. Failure probability prediction:
[0019] P t =softmax(W fc ·H t +b fc )
[0020]
[0021] Among them, P t is the fault probability distribution vector; softmax is the normalized exponential function; W fc is the weight matrix; b fc is the bias term; is the predicted fault label.
[0022] In a second aspect, the present invention further provides an intelligent operation and maintenance system based on deep learning, comprising:
[0023] A first acquisition module, used to acquire first historical operation and maintenance data, and pre-process the first historical operation and maintenance data to obtain second historical operation and maintenance data;
[0024] A training module, used for inputting the second historical operation and maintenance data into a preset deep learning-based fault prediction model for training, and bringing it to the trained fault prediction model;
[0025] A second acquisition module is used to acquire first real-time operation and maintenance data, and pre-process the first real-time operation and maintenance data to obtain second real-time operation and maintenance data;
[0026] A fault prediction module, used for inputting the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of a fault;
[0027] The automatic operation and maintenance module is used to trigger automatic operation and maintenance repair processing if the probability of the failure is greater than a preset failure threshold.
[0028] Preferably, the intelligent operation and maintenance system further comprises: an early warning module, configured to send an early warning message to an operation and maintenance management personnel if the automatic operation and maintenance repair process fails.
[0029] Preferably, the intelligent operation and maintenance system further includes: a fault processing module, which is used to receive voice instructions from the operation and maintenance manager, and perform operation and maintenance fault analysis and location according to the voice instructions; and generate a recommended operation and maintenance solution based on the results of the fault analysis and location.
[0030] Aspect 3: The present invention also provides a computer electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above-mentioned deep learning-based intelligent operation and maintenance methods are implemented.
[0031] Aspect 4: The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any one of the deep learning-based intelligent operation and maintenance methods described above are implemented.
[0032] The present invention provides an intelligent operation and maintenance method, system, device and medium based on deep learning. The method obtains first historical operation and maintenance data, and pre-processes the first historical operation and maintenance data to obtain second historical operation and maintenance data; inputs the second historical operation and maintenance data into a preset fault prediction model based on deep learning for training, and takes it into the trained fault prediction model; obtains first real-time operation and maintenance data, and pre-processes the first real-time operation and maintenance data to obtain second real-time operation and maintenance data; inputs the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of failure; if the probability of failure is greater than the preset fault threshold, automatic operation and maintenance repair processing is triggered. Compared with the related art, it has the following effects: the present application uses a deep learning prediction model to perform real-time analysis and prediction of the real-time operation data of the system, so as to predict and warn possible faults in advance, and introduces an automatic repair function. When potential risks are detected, it can also autonomously try to repair the problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of an intelligent operation and maintenance method based on deep learning according to an embodiment of the present invention;
[0034] Figure 2 It is a structural diagram of an intelligent operation and maintenance system based on deep learning according to an embodiment of the present invention;
[0035] Figure 3 It is a structural schematic diagram of a computer electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. In contrast, when an element is referred to as being "directly on" another element, there is no intermediate element. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0038] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0039] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0040] The terms used in one or more embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present invention. The singular forms "a", "said" and "the" used in one or more embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of the template are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0042] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present invention, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when...".
[0043] At present, the traditional operation and maintenance methods have the following problems:
[0044] 1. Passive response: The existing alarm system cannot predict potential system failures in advance and can only issue alarms after a failure occurs, which lacks initiative.
[0045] 2. Lack of intelligent analysis and self-repair capabilities: Most operation and maintenance systems lack intelligent analysis capabilities and automatic repair functions. They are unable to predict and prevent problems before they occur in the system, nor can they automatically restore services after problems occur.
[0046] 3. Difficulty in obtaining information: When faced with complex faults, operation and maintenance personnel often need to query a large amount of documents and information, making it difficult to find effective solutions in a timely manner, which increases the time for troubleshooting and resolution.
[0047] The above problems have brought great inconvenience to the management system of operation and maintenance personnel.
[0048] Please refer to Figure 1 , an intelligent operation and maintenance method based on deep learning provided by an embodiment of the present invention is applied to an intelligent operation and maintenance system, comprising the following steps:
[0049] S10: Acquire first historical operation and maintenance data, and preprocess the first historical operation and maintenance data to obtain second historical operation and maintenance data.
[0050] Specifically, the first historical operation and maintenance data may be collected from multiple data sources such as system logs, monitoring data, performance indicators, etc., and a preprocessing operation may be performed on the first historical operation and maintenance data to obtain the second historical operation and maintenance data.
[0051] It is understandable that the preprocessing of the first historical operation and maintenance data is for the convenience of subsequent use of the data. In some optional embodiments, the preprocessing includes: data cleaning, data standardization and data format unification.
[0052] S20. Input the second historical operation and maintenance data into a preset deep learning-based fault prediction model for training, and bring it to the trained fault prediction model.
[0053] Specifically, after the second historical operation and maintenance data is obtained, the second historical operation and maintenance data needs to be input into a pre-built fault prediction model for training to obtain a trained fault prediction model.
[0054] It should be noted that, in this embodiment, a fault prediction model based on a deep learning algorithm is used. For example, the fault prediction model can be constructed by using an attention mechanism, a convolutional neural network, a recurrent neural network or other algorithms.
[0055] S30: Acquire first real-time operation and maintenance data, and preprocess the first real-time operation and maintenance data to obtain second real-time operation and maintenance data.
[0056] Specifically, the first real-time operation and maintenance data may be collected from multiple data sources such as system logs, monitoring data, performance indicators, etc., and a preprocessing operation may be performed on the first real-time operation and maintenance data to obtain the second real-time operation and maintenance data.
[0057] It is understandable that the preprocessing of the first real-time operation and maintenance data is for the convenience of subsequent use of the data. In some optional embodiments, the preprocessing includes: data cleaning, data standardization and data format unification.
[0058] S40: Input the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of a fault occurring.
[0059] Specifically, after the second real-time operation and maintenance data is acquired, the second real-time operation and maintenance data is input into the fault prediction model for prediction, thereby obtaining the probability of a system failure.
[0060] S50: If the probability of the failure occurring is greater than a preset failure threshold, automatic operation and maintenance repair processing is triggered.
[0061] It is understandable that when the predicted probability of failure is greater than the pre-set failure threshold, the intelligent operation and maintenance system will trigger automatic operation and maintenance repair processing, such as restarting the service, clearing the cache, reconfiguring network parameters, etc.
[0062] It should be noted that the fault threshold can be set according to actual conditions. For example, for certain specific usage scenarios, the system requires that no faults occur as much as possible. In this case, the fault threshold can be set lower.
[0063] The present invention provides an intelligent operation and maintenance method based on deep learning, which obtains first historical operation and maintenance data, and pre-processes the first historical operation and maintenance data to obtain second historical operation and maintenance data; inputs the second historical operation and maintenance data into a preset fault prediction model based on deep learning for training, and takes it into the trained fault prediction model; obtains first real-time operation and maintenance data, and pre-processes the first real-time operation and maintenance data to obtain second real-time operation and maintenance data; inputs the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of failure; if the probability of failure is greater than the preset fault threshold, automatic operation and maintenance repair processing is triggered. Compared with the related art, it has the following effects: the present application uses a deep learning prediction model to perform real-time analysis and prediction of the real-time operation data of the system, so as to predict and warn possible faults in advance, and introduces an automatic repair function. When potential risks are detected, it can also autonomously try to repair the problem.
[0064] In some optional embodiments, the method further includes: S60, if the automatic operation and maintenance repair process fails, sending a warning message to the operation and maintenance management personnel.
[0065] Specifically, if the automatic operation and maintenance repair process fails, the corresponding warning information will be given to the operation and maintenance management personnel for processing.
[0066] It should be noted that the warning information includes: the type of fault and the attempted repair plan processed by automatic operation and maintenance.
[0067] In some optional embodiments, the method further comprises:
[0068] S70: Receive the voice command of the operation and maintenance manager, and perform operation and maintenance fault analysis and location according to the voice command.
[0069] It is understandable that after the operation and maintenance management personnel receive the early warning information, they can interact with the voice assistant of the intelligent operation and maintenance system through natural language. The voice assistant will mine the fault-related information from the historical data and monitoring logs based on the questions raised by the operation and maintenance personnel, and quickly locate the source of the fault.
[0070] S80: Generate a recommended operation and maintenance solution based on the results of the fault analysis and location.
[0071] Specifically, the intelligent operation and maintenance system will generate recommended operation and maintenance solutions based on the results of fault analysis and location. For example, based on the large model knowledge base and combined with the current system status, it can provide targeted solutions, such as parameter adjustment suggestions and system optimization strategies.
[0072] It should be noted that in order to enable operation and maintenance personnel to quickly troubleshoot, the intelligent operation and maintenance system will also generate operation instructions based on the fault type and historical processing experience to help operation and maintenance personnel quickly troubleshoot.
[0073] In some optional instances, the historical operation and maintenance data includes: system log data, monitoring data, and performance indicator data.
[0074] In some optional examples, the preprocessing includes: data cleaning, data standardization, and data format unification.
[0075] In some optional embodiments, the fault prediction model based on deep learning adopts the following calculation method:
[0076] a. Input data definition:
[0077] Historical operation and maintenance data matrix X: X∈R N×T×F represents the value of N operation and maintenance indicator data in the past T time windows, where F is the feature dimension; X n,t,f represents the fth eigenvalue of the nth operation and maintenance indicator data in the tth time window;
[0078] Fault label vector Y: Y∈0,1T , indicating whether a failure will occur at some point in the future;
[0079] b. Multi-mode data fusion:
[0080] In order to fuse monitoring data of different modes (such as time series, logs, keyword frequencies, etc.), we introduce the attention mechanism to perform weighted fusion of multimodal data:
[0081]
[0082] in, is the weight coefficient of the mth mode; it represents the weight coefficient of the mth mode at time t; each mode may represent a different data source. For example, performance data mode (CPU, memory, etc.), network data mode (bandwidth, packet loss, etc.), log mode (log keyword frequency), It is used to measure the importance of data of different modes to fault prediction at the current time point. A high weight means that the mode is more important at the current moment.
[0083] is a scoring function that measures the importance of the mth mode. The scoring function measures the impact of each mode (such as performance data, network data, and log data) on the system state. For example, if the network traffic surges, then the score of the network mode Probably higher due to the increased importance of network data for system failure prediction.
[0084] W m is the weight matrix, which is a parameter learned by the system specifically for each modality (data source). It helps the system adjust the influence of different modalities in scoring according to their characteristics. This matrix can be understood as a "weight adjuster for data sources" and is responsible for scoring according to the characteristics of different modalities.
[0085] b m is the bias term, which is a small numerical adjustment used to fine-tune the calculation results. The role of the bias term is usually to make the model more flexible to adapt to different data situations.
[0086] tanh tanh is an activation function (hyperbolic tangent function), which controls the range of the calculation result between -1 and 1; the activation function can help the model handle nonlinear relationships and improve the performance of the model. When the calculation result is very close to 0, it means that the influence of the current mode at this moment is relatively small. When the result is close to 1 (or -1), it means that the mode is very important at this moment and the system needs to pay close attention.
[0087] H tIt is the fused hidden state vector, which represents the current overall health status of the system and contains information of all modal data (such as performance data, network data, log data, etc.).
[0088] It is the data of the mth mode at time t; it is a "hidden state" generated after being processed by some deep learning models (such as LSTM or Transformer). It contains the main features and information of the mode. For example, performance data such as CPU utilization, memory usage, and disk I / O will be analyzed by the model to generate a hidden vector representing the current performance state.
[0089] c. Failure probability prediction:
[0090] The fused feature vector is input into the fully connected layer, and finally the probability of future failure is obtained through the softmax function.
[0091] P t =softmax(W fc *H t +b fc )
[0092]
[0093] Among them, P t is the fault probability distribution vector; softmax is a normalized exponential function used to convert the output of the model into a probability distribution; its output is a number between 0 and 1, and the sum of all possible outputs is always 1;
[0094] W fc and b fc They are the weight matrix and bias term of the fully connected layer, which are used to transform the hidden state H t Converted into a specific output result, namely the probability of predicted failure. is the predicted fault label. arg max: This function means "select the value with the largest probability". In simple terms, it will t Choose the outcome (1 or 0) with the highest probability of occurrence.
[0095] In order to better understand this application, the application of this application in the operation and maintenance of e-commerce platform cloud services is used as an example for explanation:
[0096] A large e-commerce platform is hosted in a hybrid cloud environment, which includes some self-built computer rooms and public cloud resources. The system supports multiple microservice architectures, covering core business modules such as user login, order processing, and product management. During the "618" promotion, the platform faced a huge access traffic peak, and the system performance pressure rose sharply. In the past, during similar promotions, the platform often crashed due to server overload, database bottlenecks, or network congestion, resulting in business interruptions and economic losses.
[0097] In order to solve these problems, the platform operation and maintenance team introduced deep learning intelligent operation and maintenance methods and applied them to the intelligent operation and maintenance system.
[0098] Method implementation process:
[0099] 1. Data Collection and Analysis
[0100] The system collects the following multimodal data from the e-commerce platform in real time:
[0101] Server performance indicators: CPU, memory, disk IO, network traffic, etc.
[0102] Database status: query latency, number of connections, transaction processing time, etc.
[0103] Application logs: error logs in microservices, abnormal request information, timeout warnings, etc.
[0104] Network traffic: IP source of user requests, request frequency, response time, etc.
[0105] All data are standardized through the system's data acquisition module and stored in a distributed data warehouse to provide support for subsequent analysis.
[0106] 2. Failure prediction
[0107] The fault prediction module in the system is based on the LSTM time series prediction model, combined with data from historical promotions, to predict potential faults in advance. By comparing current monitoring data with historical data, the model detects the following abnormal trends:
[0108] The database query latency increases exponentially with the increase in user requests and is expected to reach a bottleneck within the next 30 minutes.
[0109] The CPU usage of a certain microservice (order processing service) continues to soar, and it is expected to trigger a server overload failure within 10 minutes.
[0110] The response speed of user login service has increased compared to before, and it is expected to increase to 3 seconds within 30 minutes.
[0111] Predictive models identify these potential bottlenecks in good time and issue early warnings before expected failures occur.
[0112] 3. Self-healing operation
[0113] When the system detected that the order processing service was about to be overloaded, the self-healing module immediately initiated the following self-healing operations:
[0114] Automatically scale the number of instances of the order processing service from 10 to 15 to spread the CPU load.
[0115] Reallocate database connection pool resources to ensure that more resources are allocated to the order processing module.
[0116] Clean up historical cache data to free up memory and avoid memory leaks that can cause system crashes.
[0117] All of the above operations were completed without human intervention. The platform's order processing service successfully avoided overload problems and the system operated stably.
[0118] 4. Natural language interaction
[0119] The intelligent operation and maintenance system found that the response speed of the login service has decreased, but it cannot be solved by self-repair, so it sent an early warning to the operation and maintenance personnel. After receiving the early warning, the operation and maintenance personnel asked the system questions through the natural language interaction module:
[0120] Question: "Why is the response time of the login service slow? What is the solution?"
[0121] The AI assistant uses natural language understanding technology to parse the question and find relevant information from the system's monitoring data:
[0122] Cause analysis: The CPU usage of the login service is normal, but the overall login service becomes slower due to the increased response time of a third-party authentication interface.
[0123] The AI assistant then provided the following suggestions to the operation and maintenance personnel:
[0124] Solution: It is recommended to increase local cache to reduce dependence on third-party authentication interfaces, or enable backup authentication services during peak traffic periods.
[0125] The operation and maintenance personnel optimized the configuration of the login service according to the solution provided by the AI assistant, and successfully improved the response speed of the service.
[0126] See also Figure 2 The present invention also provides an intelligent operation and maintenance system 200 based on deep learning, comprising:
[0127] A first acquisition module 201 is used to acquire first historical operation and maintenance data, and pre-process the first historical operation and maintenance data to obtain second historical operation and maintenance data;
[0128] A training module 202 is used to input the second historical operation and maintenance data into a preset deep learning-based fault prediction model for training, and bring it to the trained fault prediction model;
[0129] The second acquisition module 203 is used to acquire first real-time operation and maintenance data, and pre-process the first real-time operation and maintenance data to obtain second real-time operation and maintenance data;
[0130] A fault prediction module 204 is used to input the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of a fault;
[0131] The automatic operation and maintenance module 205 is used to trigger automatic operation and maintenance repair processing if the probability of the failure is greater than a preset failure threshold.
[0132] In some optional embodiments, the intelligent operation and maintenance system 200 further includes: an early warning module 206, which is used to send an early warning message to the operation and maintenance management personnel if the automatic operation and maintenance repair process fails.
[0133] In some optional embodiments, the system 200 also includes: a fault processing module 207, which is used to receive voice instructions from the operation and maintenance management personnel, and perform operation and maintenance fault analysis and location according to the voice instructions; and generate a recommended operation and maintenance solution based on the results of the fault analysis and location.
[0134] See also Figure 3 An embodiment of the present invention further provides a computer electronic device 400, comprising a memory 403 and a processor 402, wherein the memory 403 stores a computer program, and when the processor executes the computer program, the steps of any of the above-mentioned deep learning-based intelligent operation and maintenance methods are implemented.
[0135] Specifically, the electronic device 400 includes: a transceiver 401, a bus interface and a processor 402, wherein the processor 402 is used to obtain first historical operation and maintenance data, and pre-process the first historical operation and maintenance data to obtain second historical operation and maintenance data; input the second historical operation and maintenance data into a preset deep learning-based fault prediction model for training, and bring it to the trained fault prediction model; obtain first real-time operation and maintenance data, and pre-process the first real-time operation and maintenance data to obtain second real-time operation and maintenance data; input the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of a fault; if the probability of a fault is greater than a preset fault threshold, automatic operation and maintenance repair processing is triggered.
[0136] In the embodiment of the present invention, the electronic device 400 further includes a memory 403. Figure 3 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 402 and memory represented by memory 403. The bus architecture may also link various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface. The transceiver 401 may be a plurality of components, i.e., including a transmitter and a receiver, providing a unit for communicating with various other devices on a transmission medium. The processor 402 is responsible for managing the bus architecture and general processing, and the memory 403 may store data used by the processor 402 when performing operations.
[0137] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned deep learning-based intelligent operation and maintenance methods are implemented.
[0138] In this embodiment, the computer readable storage medium may be a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include but is not limited to: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0139] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not as limiting, and thus other examples of the exemplary embodiments may have different values.
[0140] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0141] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or the flow diagram, and the combination of boxes in the structure diagram and / or the flow diagram, can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0142] In addition, the functional modules or units in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0143] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a terminal device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application.
[0144] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. An intelligent operation and maintenance method based on deep learning, characterized in that: include: Acquire first historical operation and maintenance data, and preprocess the first historical operation and maintenance data to obtain second historical operation and maintenance data; The second historical operation and maintenance data is input into a preset fault prediction model based on deep learning for training, and brought into the trained fault prediction model, wherein the fault prediction model based on deep learning adopts the following calculation method: a. Input data definition: Historical operation and maintenance data matrix X: X∈R N×T×F represents the value of N operation and maintenance indicator data in the past T time windows, where F is the feature dimension; X n,t,f represents the fth eigenvalue of the nth operation and maintenance indicator data in the tth time window; Fault label vector Y: Y∈0,1 T , indicating whether a failure will occur at some point in the future; b. Multi-mode data fusion: in, is the weight coefficient of the mth mode; is the scoring function; W m is the weight matrix; b m is the bias term; tanh is the activation function; H t is the fused hidden state vector; is the data of the mth mode at time t; c. Failure probability prediction: P t =softmax(W fc ·H t +b fc ) Among them, P t is the fault probability distribution vector; softmax is the normalized exponential function; W fc is the weight matrix; b fc is the bias term; is the predicted fault label; Acquire first real-time operation and maintenance data, and preprocess the first real-time operation and maintenance data to obtain second real-time operation and maintenance data; Inputting the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of a fault occurring; If the probability of the failure occurring is greater than a preset failure threshold, automatic operation and maintenance repair processing is triggered.
2. The intelligent operation and maintenance method based on deep learning according to claim 1, characterized in that: The method further comprises: If the automatic operation and maintenance repair process fails, an early warning message is sent to the operation and maintenance management personnel.
3. The intelligent operation and maintenance method based on deep learning according to claim 2, characterized in that: The method further comprises: Receive voice instructions from the operation and maintenance manager, and perform operation and maintenance fault analysis and location according to the voice instructions; Based on the results of the fault analysis and location, a recommended operation and maintenance solution is generated.
4. The intelligent operation and maintenance method based on deep learning according to claim 1, characterized in that: The historical operation and maintenance data includes: system log data, monitoring data and performance indicator data.
5. The intelligent operation and maintenance method based on deep learning according to claim 1, characterized in that: The preprocessing includes: data cleaning, data standardization and data format unification.
6. An intelligent operation and maintenance system based on deep learning, characterized in that: include: A first acquisition module, used to acquire first historical operation and maintenance data, and pre-process the first historical operation and maintenance data to obtain second historical operation and maintenance data; A training module, used for inputting the second historical operation and maintenance data into a preset deep learning-based fault prediction model for training, and bringing it to the trained fault prediction model; A second acquisition module is used to acquire first real-time operation and maintenance data, and pre-process the first real-time operation and maintenance data to obtain second real-time operation and maintenance data; A fault prediction module, used for inputting the second real-time operation and maintenance data into the trained fault prediction model to predict the probability of a fault; The automatic operation and maintenance module is used to trigger automatic operation and maintenance repair processing if the probability of the failure is greater than a preset failure threshold.
7. The deep learning-based intelligent operation and maintenance system according to claim 6, characterized in that: The intelligent operation and maintenance system also includes: The early warning module is used to send an early warning message to the operation and maintenance management personnel if the automatic operation and maintenance repair process fails.
8. A computer electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the deep learning-based intelligent operation and maintenance method described in any one of claims 1 to 5 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the deep learning-based intelligent operation and maintenance method described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Database intelligent operation and maintenance system and method based on machine learning
CN116701652A
Space-time fusion wind turbine generator fault prediction method based on SCADA data
CN118194222A
Power grid health assessment and analysis method based on multiple modes
CN118657404A
Power transmission line operation and maintenance method and system based on fault location and prediction
CN118710234A