Operation and maintenance fault positioning method and device based on large model, equipment and storage medium
Through the operation and maintenance fault location method based on large models, system operation and maintenance data is collected and analyzed in real time and fault trees are built, which solves the problems of fault prediction and positioning difficulties in traditional operation and maintenance management methods, and achieves fast and accurate fault location and efficient fault handling.
Patent Information
- Application Number
- CN202510678067.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Traditional operation and maintenance management methods rely on manual experience and simple monitoring tools, resulting in insufficient fault prediction capabilities and difficulty in positioning the root cause of faults.
The operation and maintenance fault location method based on the big model is adopted to collect system operation and maintenance data in real time, perform preprocessing and feature extraction, analyze data characteristics through preset large models, and build a fault tree to determine the cause of the fault.
Quickly and accurately locate the root cause of the fault, reduce troubleshooting time, improve fault handling efficiency, and ensure the rapid recovery and normal operation of the system.
Smart Images

Figure CN120196528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to an operation and maintenance fault location method, device, equipment, and storage medium based on a large model. Background Art
[0002] In today's digital age, the scale and complexity of various information systems, network systems, and industrial control systems are constantly increasing. These systems are usually composed of a large number of hardware devices, software components, and network connections, and their operating states are affected by various factors, such as hardware aging, software vulnerabilities, network fluctuations, human operation errors, etc. Once a system fails, it may lead to serious consequences such as business interruption, data loss, and service quality degradation, causing huge losses to enterprises and users.
[0003] Traditional operation and maintenance management methods mainly rely on manual experience and simple monitoring tools. Operation and maintenance personnel monitor key indicators of the system, such as processor usage rate, memory occupancy, network traffic, etc., by setting some fixed thresholds. When these indicators exceed the thresholds, the system will send an alarm to notify the operation and maintenance personnel. However, this method has problems such as insufficient fault prediction ability due to rigid threshold setting and difficulty in locating the root cause of faults due to complex alarm information. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an operation and maintenance fault location method, device, equipment, and storage medium based on a large model, which can quickly collect various information related to the fault when the fault occurs, accurately locate the root cause of the fault, reduce the fault troubleshooting time, and improve the fault handling efficiency. The specific solutions are as follows: In the first aspect, the present application discloses an operation and maintenance fault location method based on a large model, including: If a system fault is detected, real-time collection of current system operation and maintenance data is performed, and the current system operation and maintenance data is preprocessed to obtain target preprocessed system operation and maintenance data; Extract the data features corresponding to the target preprocessed system operation and maintenance data, and input the data features into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; Based on the fault tree analysis method, construct a fault tree corresponding to the fault analysis result to determine the fault cause of the system fault through the fault tree, and analyze the fault cause to complete fault location.
[0005] Optionally, the step of if a system fault is detected, real-time collection of current system operation and maintenance data is performed, and the current system operation and maintenance data is preprocessed to obtain target preprocessed system operation and maintenance data includes: If a system failure is detected, the current system operation and maintenance data of the system locally is collected in real time, and data annotation and data classification are performed on the current system operation and maintenance data to obtain the first preprocessed system operation and maintenance data; Identify the invalid data, duplicate data, and abnormal data in the first preprocessed system operation and maintenance data, and remove the invalid data, duplicate data, and abnormal data from the first preprocessed system operation and maintenance data to obtain the second preprocessed system operation and maintenance data; Perform normalization processing on the second preprocessed system operation and maintenance data to convert the second preprocessed system operation and maintenance data into a preset data format to obtain the target preprocessed system operation and maintenance data.
[0006] Optionally, the real-time collection of the current system operation and maintenance data of the system locally, and the data annotation and data classification of the current system operation and maintenance data to obtain the first preprocessed system operation and maintenance data, includes: Real-time collect the current system logs, alarm information, performance metric data, network traffic data, and hardware status data of the system; Add a timestamp to the system logs and sort the system logs based on the timestamp to obtain the sorted system logs; Generate an alarm event chain based on the alarm information, classify the performance metric data based on the metric type, classify the hardware status data based on the component type, and classify the network traffic data based on the traffic type.
[0007] Optionally, the analysis of the data features by the preset large model to obtain a fault analysis result includes: Extract information from the data features through a preset large model, and match the extracted target information with preset fault cases to determine the target fault cases in the preset fault cases that match the target information; Construct a fault causality graph based on the alarm event chain in the data features; Compare the performance metric data with historical performance metric data to determine abnormal changes in performance metrics; Analyze the network traffic data to determine abnormal network behaviors existing in the network traffic data; Compare the hardware status data with a preset hardware data threshold to determine abnormal hardware statuses in the hardware status data; Use the fault causality graph, the abnormal changes in performance metrics, the abnormal network behaviors, the abnormal hardware statuses, and the target fault cases as the fault analysis result.
[0008] Optionally, constructing a fault tree corresponding to the fault analysis result based on the fault tree analysis method to determine the fault cause of the system fault through the fault tree, including: Identifying a fault top event, fault intermediate events, and fault bottom events corresponding to the system fault based on the fault analysis result, and constructing a fault tree corresponding to the fault analysis result based on the fault top event, the fault intermediate events, and the fault bottom events; Determining, through the fault tree, a target fault bottom event with the highest influence probability on the fault top event among the fault bottom events, and using the target fault bottom event as the fault cause of the system fault.
[0009] Optionally, analyzing the fault cause to complete fault location, including: Analyzing the fault cause to determine the fault occurrence time, the faulty component, and the fault abnormal data corresponding to the fault cause, so as to complete fault location.
[0010] Optionally, the operation and maintenance fault location method based on a large model further includes: Constructing a fault prediction model based on a pre-trained model through an ensemble learning method; Collecting current operation and maintenance data of the target system in real time based on a preset time interval, and inputting the operation and maintenance data of the target system into the fault prediction model to perform real-time system fault prediction through the fault prediction model, so as to obtain a real-time fault prediction result; Adjusting system parameters based on the fault prediction result to prevent system faults.
[0011] In a second aspect, the present application discloses an operation and maintenance fault location device based on a large model, including: A data preprocessing module, configured to, if a system fault is detected, collect current system operation and maintenance data in real time, and preprocess the current system operation and maintenance data to obtain target preprocessed system operation and maintenance data; A fault analysis module, configured to extract data features corresponding to the target preprocessed system operation and maintenance data, and input the data features into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; A fault location module, configured to construct a fault tree corresponding to the fault analysis result based on the fault tree analysis method to determine the fault cause of the system fault through the fault tree, and analyze the fault cause to complete fault location.
[0012] In a third aspect, the present application discloses an electronic device, including: A memory, configured to store a computer program; A processor for executing the computer program to implement the operation and maintenance fault location method based on a large model as described above.
[0013] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the operation and maintenance fault location method based on a large model as described above.
[0014] In the present application, if a system fault is detected, current system operation and maintenance data is collected in real time, and the current system operation and maintenance data is preprocessed to obtain target preprocessed system operation and maintenance data; data features corresponding to the target preprocessed system operation and maintenance data are extracted, and the data features are input into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; a fault tree corresponding to the fault analysis result is constructed based on the fault tree analysis method to determine the fault cause of the system fault through the fault tree, and the fault cause is analyzed to complete fault location. It can be seen that through the method of the present application, it is necessary to collect system operation and maintenance data in real time after detecting a system fault, then preprocess the system operation and maintenance data, and extract the operation and maintenance features of the preprocessed data. After extracting the features, it is necessary to analyze the data features through a preset large model and construct a fault tree corresponding to the obtained analysis result to determine the system fault cause through the finally obtained fault tree. In this way, when a fault occurs, various information related to the fault can be quickly collected, the root cause of the fault can be accurately located, the fault troubleshooting time can be reduced, the fault handling efficiency can be improved, and the rapid recovery and normal operation of the system can be guaranteed. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0016] Figure 1 It is a flowchart of an operation and maintenance fault location method based on a large model disclosed in the present application; Figure 2 It is a timing diagram of an operation and maintenance fault location method based on a large model disclosed in the present application; Figure 3 It is a schematic structural diagram of an operation and maintenance fault location device based on a large model disclosed in the present application; Figure 4 It is a structural diagram of an electronic device disclosed in the present application. Detailed Embodiments
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] In the prior art, traditional operation and maintenance management methods mainly rely on manual experience and simple monitoring tools. For example, thresholds are set manually to detect key system indicators and generate alarms. However, this method has problems such as insufficient fault prediction ability due to rigid threshold setting and difficulty in locating the root cause of faults due to complex alarm information.
[0019] To overcome the above technical problems, the present application discloses an operation and maintenance fault location method, device, equipment and storage medium based on a large model. When a fault occurs, it can quickly collect various information related to the fault, accurately locate the root cause of the fault, reduce the fault troubleshooting time, and improve the fault handling efficiency.
[0020] See Figure 1 As shown, the embodiments of the present invention disclose an operation and maintenance fault location method based on a large model, including: Step S11: If a system fault is detected, collect the current system operation and maintenance data in real time, and preprocess the current system operation and maintenance data to obtain target preprocessed system operation and maintenance data.
[0021] In this embodiment, as Figure 2As shown, if a system failure is detected, it is necessary to collect the current system operation and maintenance data in real time and perform corresponding preprocessing on the system operation and maintenance data. Specifically, if a system failure is detected, the current system operation and maintenance data of the local system is collected in real time, and the current system operation and maintenance data is labeled and classified to obtain the first preprocessed system operation and maintenance data, wherein it is necessary to collect the current system log, alarm information, performance indicator data, network traffic data and hardware status data of the system in real time, such as server, application, database and other logs. After the log collection is completed, it is necessary to add a timestamp to the log and sort the system log based on the timestamp to obtain the sorted system log; the alarm information needs to be collected from the hardware , network, and application performance monitoring systems, and the alarm information needs to be integrated and classified, duplicates need to be removed, and the alarm information needs to be associated to form an alarm event chain; for performance indicator data, such as processor utilization, memory occupancy, etc., the performance indicator data needs to be classified according to the indicator type, and then the historical data and change trends need to be displayed in the form of a chart; for hardware status data, the hardware status data needs to be classified according to the component type corresponding to the hardware, such as temperature, fan speed, power status, etc.; for traffic data, the network traffic data needs to be obtained through a network traffic collection tool, and the network traffic data needs to be classified according to the traffic type, and the traffic direction, size, and data packet characteristics can be analyzed.
[0022] Furthermore, the first pre-processed system operation and maintenance data needs to be further processed. Specifically, invalid data, duplicate data, and abnormal data in the first pre-processed system operation and maintenance data need to be determined, and the invalid data, duplicate data, and abnormal data need to be removed from the first pre-processed system operation and maintenance data to obtain the second pre-processed system operation and maintenance data. In this way, by removing invalid data, duplicate data, and abnormal data, noise interference in the data can be effectively reduced, thereby more truly reflecting the data characteristics.
[0023] In the next step, the obtained second pre-processed system operation and maintenance data needs to be normalized to convert the second pre-processed system operation and maintenance data into a preset data format, thereby obtaining the target pre-processed system operation and maintenance data. It should be noted that the data is normalized and converted into a unified data format, that is, JSON (JavaScript Object Notation) format. It should be noted that the data is normalized and converted into a unified data format to ensure data consistency and avoid analysis errors caused by data inconsistencies.
[0024] Step S12: Extract the data features corresponding to the target pre - processed system operation and maintenance data, and input the data features into a preset large - model to analyze the data features through the preset large - model to obtain a fault analysis result.
[0025] In this embodiment, as Figure 2 shown, it is necessary to extract the features of the target pre - processed system operation and maintenance data, and analyze the extracted data features through a large - model to obtain the corresponding fault analysis result. Specifically, it is necessary to extract information from the data features through a preset large - model, and match the extracted target information with preset fault cases to determine the target fault case that matches the target information in the preset fault cases. It should be noted that after obtaining the target pre - processed system operation and maintenance data, it is necessary to extract the data features. The time - domain features include the mean, variance, maximum and minimum values, median of performance indicators, as well as the occurrence frequency and time interval of specific events in the log; the frequency - domain features obtain the frequency components of the data through Fourier transform, etc., such as the main frequency and harmonics, reflecting the periodicity and stability of the system operation; the trend features use time - series analysis methods such as moving average and exponential smoothing to extract the rising, falling or stable trends of performance indicators to detect system performance anomalies in advance; the text features are for text data such as system logs. Then, the extracted data features are input into the preset large - model. And the large - model, based on the analysis of operation and maintenance domain knowledge and data, determines which features are important for fault prediction, avoiding the influence of feature redundancy on the model performance, and the large - model will match the fault cases in the operation and maintenance domain with the feature information to determine the corresponding target fault case.
[0026] Furthermore, it is necessary to construct a fault causality graph based on the alarm event chain in the data features to determine which alarms trigger other alarms and the root cause of the fault; it is also necessary to compare the performance indicator data with the historical performance indicator data to determine the abnormal change of the performance indicators, and the fault correlation can be judged according to the normal and fault modes of the system, and the processes and reasons leading to the abnormal performance indicators can be analyzed; the network traffic data can also be analyzed to determine the abnormal network behaviors in the network traffic data, and judge whether there are fault factors such as network attacks, congestion or abnormal network behaviors of applications; the hardware status data can be compared with the preset hardware data threshold to determine the abnormal hardware status in the hardware status data, such as judging the reason for the server temperature being too high and its impact on other components of the system. Finally, the fault causality graph, the abnormal change of performance indicators, the abnormal network behaviors, the abnormal hardware status and the target fault case are used as the fault analysis result.
[0027] Step S13: Construct a fault tree corresponding to the fault analysis result based on the fault tree analysis method to determine the fault cause of the system fault through the fault tree, and analyze the fault cause to complete the fault location.
[0028] In this embodiment, it is necessary to determine the cause of the system failure through the constructed fault tree and complete the corresponding fault location. Specifically, it is necessary to identify the fault top event, fault intermediate event, and fault bottom event corresponding to the system failure based on the fault analysis results, and construct a fault tree corresponding to the fault analysis results based on the fault top event, fault intermediate event, and fault bottom event. That is, based on the analysis results of the large model, the fault tree analysis method is used to construct the fault tree. Taking the fault phenomenon as the top event, according to the fault causes and causal relationships analyzed by the large model, the intermediate events and bottom events are gradually constructed. For example, if the system service is unavailable, the intermediate events may be network failure, server hardware failure, or application error. Further expanding, the network failure may be bottom events such as network device failure, line interruption, or configuration error.
[0029] Then, it is necessary to determine the target fault bottom event with the highest influence probability on the fault top event among the fault bottom events through the fault tree, and use the target fault bottom event as the cause of the system failure. Specifically, it is necessary to use fault tree analysis software to input parameters such as the fault tree structure, event logical relationship, and occurrence probability of the bottom event, calculate the contribution degree of the bottom event to the occurrence of the top event, determine the target fault bottom event with the highest influence probability on the fault top event among the fault bottom events, and use the target fault bottom event as the cause of the system failure. Further, it is necessary to analyze the fault cause to determine the fault occurrence time, fault occurrence component, and fault abnormal data corresponding to the fault cause to complete the fault location. It should be noted that after completing the fault location, a fault report can be generated, including the fault occurrence time, location, phenomenon, root cause analysis process, fault tree structure, and solution suggestions. The report details the large model analysis idea and the fault tree construction process, providing a clear basis for fault handling for the operation and maintenance personnel. Then, specific operation steps and suggestions are provided according to the fault root cause. For example, for a network device port failure, it is recommended to replace the port, check the line connection, reconfigure the parameters, etc., and provide relevant technical documents and reference material links. Finally, the fault analysis report needs to be provided to the operation and maintenance personnel. In this way, when a fault occurs, various information related to the fault can be quickly collected, including system logs, alarm information, performance indicators, etc., and with the help of the knowledge reasoning and semantic understanding capabilities of the large model, the root cause of the fault can be quickly and accurately located, reducing the fault troubleshooting time, improving the fault handling efficiency, and ensuring the rapid recovery and normal operation of the system.
[0030] Further explanations are needed. A fault prediction model can be constructed to predict the system through the prediction model. Specifically, a fault prediction model can be constructed based on a pre-trained model by using a learning method. It should be noted that the pre-trained model can be selected according to requirements. For example, algorithms such as Support Vector Machine (SVM), Random Forest (RF), Long Short-Term Memory (LSTM), and Convolutional Neural Networks (CNN) can be used, and appropriate algorithms or combinations can be selected according to the system characteristics and data features. Moreover, after the fault prediction model is constructed, the current operation and maintenance data of the target system can be collected in real time based on a preset time interval, and the operation and maintenance data of the target system can be input into the fault prediction model to perform real-time system fault prediction through the fault prediction model and obtain real-time fault prediction results. Additionally, the preset large model can combine the fault prediction results with operation and maintenance cases and data to understand the model performance under different parameter settings, and provide optimal parameter suggestions for the fault prediction model. For example, it can provide adjustment suggestions for the kernel function parameters and penalty factors of the SVM (Support Vector Machine) model, and optimize the model architecture. For example, according to the system data complexity and fault patterns, it can suggest adjusting the number of network layers, the number of neurons, etc. of the LSTM or CNN model to improve the model's prediction ability. In this way, fault prediction can be performed in real time, and parameter adjustment can be carried out through the large model, effectively improving the efficiency and accuracy of fault prediction.
[0031] In this embodiment, if a system failure is detected, the current system operation and maintenance data is collected in real time, and the current system operation and maintenance data is preprocessed to obtain the target preprocessed system operation and maintenance data; the data features corresponding to the target preprocessed system operation and maintenance data are extracted, and the data features are input into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; a fault tree corresponding to the fault analysis result is constructed based on the fault tree analysis method to determine the fault cause of the system failure through the fault tree, and the fault cause is analyzed to complete fault location. It can be seen that through the method of this application, it is necessary to collect the system operation and maintenance data in real time after detecting the system failure, then preprocess the system operation and maintenance data, and extract the operation and maintenance features of the preprocessed data. After extracting the features, it is necessary to analyze the data features through a preset large model and construct a fault tree corresponding to the obtained analysis result to determine the system fault cause through the finally obtained fault tree. In this way, on the one hand, when a fault occurs, various information related to the fault can be quickly collected to accurately locate the root cause of the fault, reduce the fault troubleshooting time, improve the fault handling efficiency, and ensure the rapid recovery and normal operation of the system; on the other hand, through automated fault prediction and root cause location, the dependence on manual operation and maintenance is reduced, the work intensity and professional requirements of operation and maintenance personnel are reduced, thereby reducing the operation and maintenance costs of the enterprise. At the same time, the availability and stability of the system are improved, and the business losses caused by faults are reduced; on the other hand, through real-time collection and analysis of multi-source heterogeneous data generated during the system operation process, using the constructed prediction model, potential fault hazards in the system can be discovered in advance, accurate fault prediction can be achieved, sufficient time is provided for operation and maintenance personnel for fault prevention and handling, and the probability and impact degree of fault occurrence are reduced.
[0032] See Figure 3 As shown, an operation and maintenance fault location device based on a large model according to an embodiment of the present invention includes: A data preprocessing module 11, configured to collect the current system operation and maintenance data in real time if a system failure is detected, and preprocess the current system operation and maintenance data to obtain the target preprocessed system operation and maintenance data; A fault analysis module 12, configured to extract the data features corresponding to the target preprocessed system operation and maintenance data, and input the data features into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; A fault location module 13, configured to construct a fault tree corresponding to the fault analysis result based on the fault tree analysis method to determine the fault cause of the system failure through the fault tree, and analyze the fault cause to complete fault location.
[0033] In this embodiment, if a system failure is detected, the current system operation and maintenance data is collected in real time, and the current system operation and maintenance data is preprocessed to obtain the target preprocessed system operation and maintenance data; the data features corresponding to the target preprocessed system operation and maintenance data are extracted, and the data features are input into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; a fault tree corresponding to the fault analysis result is constructed based on the fault tree analysis method to determine the fault cause of the system failure through the fault tree, and the fault cause is analyzed to complete fault location. It can be seen that through the method of this application, it is necessary to collect the system operation and maintenance data in real time after detecting the system failure, then preprocess the system operation and maintenance data, and extract the operation and maintenance features of the preprocessed data. After the features are extracted, it is necessary to analyze the data features through a preset large model and construct a fault tree corresponding to the obtained analysis result to determine the system fault cause through the finally obtained fault tree. In this way, when a fault occurs, various information related to the fault can be quickly collected to accurately locate the root cause of the fault, reduce the fault troubleshooting time, improve the fault handling efficiency, and ensure the rapid recovery and normal operation of the system.
[0034] In some embodiments, the data preprocessing module 11 may specifically include: The first preprocessing sub-module is configured to, if a system failure is detected, collect the current system operation and maintenance data of the system locally in real time, and perform data annotation and data classification on the current system operation and maintenance data to obtain the first preprocessed system operation and maintenance data; The second preprocessing sub-module is configured to remove the invalid data, the duplicate data, and the abnormal data from the first preprocessed system operation and maintenance data to obtain the second preprocessed system operation and maintenance data; The data conversion sub-module is configured to perform normalization processing on the second preprocessed system operation and maintenance data to convert the second preprocessed system operation and maintenance data into a preset data format to obtain the target preprocessed system operation and maintenance data.
[0035] In some embodiments, the first preprocessing sub-module may specifically include: The data collection unit is configured to collect the system log, alarm information, performance index data, network traffic data, and hardware status data of the system in real time; The log sorting unit is configured to add a timestamp to the system log and sort the system log based on the timestamp to obtain the sorted system log; A data classification unit for generating an alarm event chain based on the alarm information, classifying the performance metric data based on metric types, classifying the hardware status data based on component types, and classifying the network traffic data based on traffic types.
[0036] In some embodiments, the fault analysis module 12 may specifically include: An information matching unit for extracting information from the data features through a preset large model and matching the extracted target information with preset fault cases to determine a target fault case in the preset fault cases that matches the target information; A relationship graph construction unit for constructing a fault causality graph based on the alarm event chain in the data features; A first data comparison unit for comparing the performance metric data with historical performance metric data to determine abnormal changes in the performance metrics; A second data comparison unit, a data analysis unit, for analyzing the network traffic data to determine abnormal network behaviors in the network traffic data; A third data comparison unit for comparing the hardware status data with a preset hardware data threshold to determine abnormal hardware statuses in the hardware status data; A data definition unit for taking the fault causality graph, the abnormal changes in the performance metrics, the abnormal network behaviors, the abnormal hardware statuses, and the target fault case as fault analysis results.
[0037] In some embodiments, the fault location module 13 may specifically include: A fault tree construction unit for identifying a fault top event, fault intermediate events, and fault bottom events corresponding to the system fault based on the fault analysis results and constructing a fault tree corresponding to the fault analysis results based on the fault top event, the fault intermediate events, and the fault bottom events; A fault cause analysis unit for determining, through the fault tree, a target fault bottom event with the highest influence probability on the fault top event among the fault bottom events and taking the target fault bottom event as the fault cause of the system fault.
[0038] In some embodiments, the fault location module 13 may specifically include: A fault location unit for analyzing the fault cause to determine the fault occurrence time, the fault occurrence component, and the fault abnormal data corresponding to the fault cause to complete fault location.
[0039] In some embodiments, the operation and maintenance fault location device based on a large model may further include: A model construction unit for constructing a fault prediction model based on a pre-trained model by an ensemble learning method; A fault real-time prediction unit for collecting current operation and maintenance data of a target system in real time based on a preset time interval, and inputting the operation and maintenance data of the target system into the fault prediction model to perform real-time system fault prediction through the fault prediction model and obtain a real-time fault prediction result; A parameter adjustment unit for adjusting system parameters based on the fault prediction result to prevent system faults.
[0040] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 4 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of the present application.
[0041] Figure 4 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the operation and maintenance fault location method based on a large model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0042] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0043] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0044] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the operation and maintenance fault location method based on the large model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.
[0045] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the operation and maintenance fault location method based on the large model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0046] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for related parts.
[0047] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0048] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0049] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0050] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for locating operation and maintenance faults based on a large model, characterized in that Including: If a system failure is detected, the current system operation and maintenance data is collected in real time, and the current system operation and maintenance data is preprocessed to obtain the target preprocessed system operation and maintenance data; Extract the data features corresponding to the target preprocessed system operation and maintenance data, and input the data features into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; Construct a fault tree corresponding to the fault analysis result based on the fault tree analysis method, to determine the fault cause of the system failure through the fault tree, and analyze the fault cause to complete fault location.
2. The operation and maintenance fault location method based on a large model according to claim 1, characterized in that The step of if a system failure is detected, the current system operation and maintenance data is collected in real time, and the current system operation and maintenance data is preprocessed to obtain the target preprocessed system operation and maintenance data, includes: If a system failure is detected, the current system operation and maintenance data on the system local is collected in real time, and the current system operation and maintenance data is data labeled and data classified to obtain the first preprocessed system operation and maintenance data; Determine the invalid data, duplicate data and abnormal data in the first preprocessed system operation and maintenance data, and remove the invalid data, the duplicate data and the abnormal data from the first preprocessed system operation and maintenance data to obtain the second preprocessed system operation and maintenance data; Perform normalization processing on the second preprocessed system operation and maintenance data to convert the second preprocessed system operation and maintenance data into a preset data format to obtain the target preprocessed system operation and maintenance data.
3. The operation and maintenance fault location method based on a large model according to claim 2, wherein The step of collecting the current system operation and maintenance data on the system local in real time, and data labeling and data classifying the current system operation and maintenance data to obtain the first preprocessed system operation and maintenance data, includes: Collect the current system log, alarm information, performance index data, network traffic data and hardware status data of the system in real time; Add a timestamp to the system log, and sort the system log based on the timestamp to obtain the sorted system log; Generate an alarm event chain based on the alarm information, classify the performance index data based on the index type, classify the hardware status data based on the component type, and classify the network traffic data based on the traffic type.
4. The operation and maintenance fault location method based on a large model according to claim 3, wherein The step of analyzing the data features through the preset large model to obtain a fault analysis result, includes: Extract information from the data features through a preset large model, and match the extracted target information with preset fault cases to determine the target fault case in the preset fault cases that matches the target information; Construct a fault causal relationship diagram based on the alarm event chain in the data features; Compare the performance index data with the historical performance index data to determine abnormal changes in the performance index; Analyze the network traffic data to determine abnormal network behaviors existing in the network traffic data; Compare the hardware status data with a preset hardware data threshold to determine abnormal hardware status in the hardware status data; Take the fault causal relationship diagram, the abnormal change of performance indicators, the abnormal network behavior, the abnormal hardware state, and the target fault case as the fault analysis results.
5. The operation and maintenance fault location method based on a large model according to claim 1, characterized in that, Construct a fault tree corresponding to the fault analysis results based on the fault tree analysis method to determine the fault cause of the system fault through the fault tree, including: Identify the fault top event, fault intermediate event, and fault bottom event corresponding to the system fault based on the fault analysis results, and construct a fault tree corresponding to the fault analysis results based on the fault top event, the fault intermediate event, and the fault bottom event; Determine the target fault bottom event with the highest influence probability on the fault top event among the fault bottom events through the fault tree, and take the target fault bottom event as the fault cause of the system fault.
6. The operation and maintenance fault location method based on a large model according to claim 1, wherein, Analyze the fault cause to complete fault location, including: Analyze the fault cause to determine the fault occurrence time, fault occurrence component, and fault abnormal data corresponding to the fault cause, so as to complete fault location.
7. The operation and maintenance fault location method based on a large model according to claim 1, characterized in that, It also includes: Construct a fault prediction model based on a pre-trained model through an ensemble learning method; Real-time collect the current target system operation and maintenance data at preset time intervals, and input the target system operation and maintenance data into the fault prediction model to perform real-time system fault prediction through the fault prediction model to obtain real-time fault prediction results; Adjust system parameters based on the fault prediction results to prevent system faults.
8. An operation and maintenance fault location device based on a large model, characterized in that, Include: A data preprocessing module, configured to, if a system fault is detected, collect the current system operation and maintenance data in real time, and preprocess the current system operation and maintenance data to obtain target preprocessed system operation and maintenance data; A fault analysis module, configured to extract the data features corresponding to the target preprocessed system operation and maintenance data, and input the data features into a preset large model to analyze the data features through the preset large model to obtain a fault analysis result; A fault location module, configured to construct a fault tree corresponding to the fault analysis results based on the fault tree analysis method to determine the fault cause of the system fault through the fault tree, and analyze the fault cause to complete fault location.
9. An electronic device, characterized in that, Include: A memory for storing computer programs; A processor for executing the computer programs to implement the large model-based operation and maintenance fault location method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing computer programs, wherein the computer programs, when executed by a processor, implement the large model-based operation and maintenance fault location method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Methods and apparatus for performing system fault diagnosis
CA2002222A1
Operation and maintenance system and method
CN110659173A
Method and device for processing system fault, equipment and storage medium
CN117170925A
System-level fault tree aided modeling method based on large model
CN118052281A
Internet equipment fault diagnosis method and system
CN118827342A
Cited By
Multi-mode Internet of Things field computing AI gateway
CN121356944A
Energy storage system anomaly analysis method and device, electronic equipment, storage medium and product
CN121388879A