A service operation and maintenance method, system, device, medium and product
By preprocessing and intelligently diagnosing multi-source operational data of IT systems, the problem of error-prone fault diagnosis in traditional manual operation and maintenance is solved, enabling real-time monitoring and efficient fault handling, and reducing the impact of faults.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE COMM CORP TIANJIN
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-19
Smart Images

Figure CN122247885A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operation and maintenance technology, and in particular to a business operation and maintenance method, system, equipment, medium and product. Background Technology
[0002] Information Technology (IT) operations and maintenance are primarily used to ensure the continuous and stable operation of an enterprise's IT systems and networks. Its main goal is to ensure the availability, security, and efficiency of the IT environment through standardized processes and technical means, thereby supporting business continuity.
[0003] In current mainstream production IT operations, fault handling primarily relies on manual methods. When an IT system or operating environment malfunctions, technical personnel use their experience and expertise to troubleshoot, determine the cause, and manually implement solutions. However, this approach is flawed. Firstly, individual differences in experience and expertise among technical personnel can lead to incorrect fault diagnosis due to human error. This not only reduces efficiency but can also create new problems and disrupt the overall IT system's operation. Secondly, this method struggles to identify potential faults early, resulting in delayed resolution, a wider impact, and significantly reduced timeliness and effectiveness of fault handling. Summary of the Invention
[0004] This invention provides a business operation and maintenance method, system, equipment, medium and product to achieve real-time and comprehensive monitoring of IT systems, accurate and intelligent fault diagnosis, timely early warning of potential risks, improve operation and maintenance efficiency and accuracy, and reduce the risk of fault impact.
[0005] In a first aspect, embodiments of this disclosure provide a service operation and maintenance method, including: Collect multi-source operational data of the business environment at a preset sampling frequency, and preprocess the collected multi-source operational data to form an operation and maintenance dataset; The operation and maintenance dataset is input into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results. Based on the model monitoring results, fault diagnosis is performed to obtain the fault diagnosis results. When the fault diagnosis result indicates that a fault exists, an operation and maintenance fault alarm is issued and operation and maintenance fault handling is performed.
[0006] Secondly, embodiments of this disclosure provide a business operation and maintenance system, including: The data acquisition module is used to collect multi-source operational data of the business environment at a preset sampling frequency, and to preprocess the collected multi-source operational data to form an operation and maintenance dataset. The operation and maintenance monitoring module is used to input the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results, and to perform fault diagnosis based on the model monitoring results to obtain the fault diagnosis results. The fault handling module is used to issue an operation and maintenance fault alarm and perform operation and maintenance fault handling when the fault diagnosis result indicates that a fault exists.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: At least one processor; and A memory that is communicatively connected to at least one processor; wherein, The memory stores a computer program that can be executed by at least one processor, such that the at least one processor can execute a business operation and maintenance method provided in the first aspect embodiment described above.
[0008] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer instructions that, when executed by a processor, implement a business operation and maintenance method provided in the first aspect of the embodiments described above.
[0009] Fifthly, this disclosure provides a computer program product, which includes a computer program that, when executed by a processor, implements a business operation and maintenance method provided in the first aspect of the embodiment.
[0010] The technical solution of this invention collects multi-source operational data of the business environment at a preset sampling frequency, preprocesses the collected multi-source operational data to form an operation and maintenance dataset, inputs the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain model monitoring results, performs fault diagnosis based on the model monitoring results to obtain fault diagnosis results, and issues an operation and maintenance fault alarm and executes operation and maintenance fault handling when the fault diagnosis results indicate that a fault exists. These technical features enable real-time comprehensive monitoring of IT systems, accurate and intelligent fault diagnosis, timely early warning of potential problems, improved operation and maintenance efficiency and accuracy, and reduced risk of fault impact.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a business operation and maintenance method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a business operation and maintenance system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] In one embodiment, Figure 1 This is a flowchart of a business operation and maintenance method provided by an embodiment of the present invention. This embodiment can be applied to the situation of performing operation and maintenance monitoring and fault handling of IT business. The method can be executed by a business operation and maintenance system, which can be implemented in hardware and / or software.
[0017] For example, a business operation and maintenance system includes a data acquisition module, an operation and maintenance monitoring module, and a fault handling module; The data acquisition module is used to collect multi-source operational data of the business environment at a preset sampling frequency, and to preprocess the collected multi-source operational data to form an operation and maintenance dataset. The operation and maintenance monitoring module is used to input the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results, and to perform fault diagnosis based on the model monitoring results to obtain the fault diagnosis results. The fault handling module is used to issue an operation and maintenance fault alarm and perform operation and maintenance fault handling when the fault diagnosis result indicates that a fault exists.
[0018] Alternatively, the business operation and maintenance system may include a hardware and software information acquisition module, an operating environment information acquisition module, a data preprocessing module, an AI model training module, an IT operation and maintenance monitoring module, a fault diagnosis module, an alert module, an IT operation and maintenance processing module, and a model optimization module. The hardware and software information acquisition module is used to collect information on the hardware and software during operation in real time, including hardware model, hardware usage time, software version, and software operation log information. The operating environment information acquisition module is used to collect the operating environment information in real time during operation, including data center temperature, humidity, voltage, current, network bandwidth, and network latency. The data preprocessing module is used to clean, transform, and integrate the data collected by the hardware and software information acquisition module and the runtime environment information acquisition module; The AI model training module builds an IT operations and maintenance monitoring model based on a convolutional neural network (CNN) and trains the model using preprocessed data. The IT operations and maintenance monitoring module receives hardware and software information and operating environment information processed by the data preprocessing module, and performs real-time analysis and processing of this information based on the trained IT operations and maintenance monitoring model to determine whether the IT system and operating environment are in a normal state. When the IT operations and maintenance monitoring module detects an anomaly, the fault diagnosis module performs fault diagnosis based on the model output to determine the cause of the fault. Once the IT operations and maintenance monitoring module detects an anomaly in the information or the fault diagnosis module determines the cause of the fault, the alert module will receive the alert information from the corresponding module, issue an alarm, and transmit the alert information to relevant personnel for timely handling. When a fault occurs during IT operation, the IT operations and maintenance module can quickly retrieve the corresponding fault handling operation plan from the fault handling operation database based on the fault diagnosis module and provide it to operations and maintenance personnel or automatically execute some handling operations, thereby improving fault handling efficiency. The model optimization module is used to continuously optimize and update the IT operations and maintenance monitoring model.
[0019] like Figure 1 As shown, the method includes: S101. Collect multi-source operational data of the business environment according to the preset sampling frequency, and preprocess the collected multi-source operational data to form an operation and maintenance dataset.
[0020] In this embodiment, the preset sampling frequency can be understood as a differentiated collection frequency set for different types of operation and maintenance data, which is a collection rule adapted to the characteristics of software, hardware, and operating environment data. Multi-source operation data can be understood as a full-dimensional data set covering software and hardware operation data and operating environment data in the IT business environment. Among them, software and hardware operation data includes hardware model, hardware usage time, software version, and software operation log information, while operating environment data includes environmental parameters such as data center temperature, humidity, voltage, and current, and network performance indicators such as network bandwidth and network latency. The operation and maintenance dataset can be understood as a unified, standardized dataset that can be directly used for model analysis after preprocessing and merging according to the collection timestamp.
[0021] Specifically, the software and hardware information acquisition module and the operating environment information acquisition module are first activated. They synchronously collect multi-source raw operating data from the business environment according to differentiated preset sampling frequencies, so as to achieve comprehensive and real-time data capture of the IT system's software and hardware status and operating environment. The collected raw multi-source data is then transmitted to the data preprocessing module. The software and hardware operating data and the operating environment data are first cleaned to remove noise and outliers. Then, the data is standardized to convert data of different magnitudes to the same magnitude. Finally, the two types of processed data are merged according to the collection timestamp to form a unified operation and maintenance dataset.
[0022] S102. Input the operation and maintenance dataset into the pre-trained operation and maintenance monitoring model to obtain the model monitoring results. Based on the model monitoring results, perform fault diagnosis to obtain the fault diagnosis results.
[0023] In this embodiment, the operation and maintenance monitoring model is an artificial intelligence model built on a convolutional neural network. Using historical fault data as the training set, it is trained, tested, evaluated, and optimized to possess IT system status analysis and fault diagnosis capabilities. The model monitoring results can be understood as the IT business operation status assessment results output by the model after real-time analysis of the input operation and maintenance dataset, categorized into normal and abnormal states. The fault diagnosis results can be understood as the conclusions output after fault diagnosis, including three types of results: no fault, fault causes corresponding to specific fault subclasses, and general fault causes corresponding to fault parent classes.
[0024] Specifically, the operation and maintenance dataset is input into the pre-trained operation and maintenance monitoring model in real time. The model performs intelligent analysis and processing on the dataset and outputs corresponding business operation status monitoring results. If the model monitoring results indicate that the business operation status is normal, the fault diagnosis result is directly determined as no fault. If the model monitoring results indicate that the business operation status is abnormal, abnormal data feature keyword groups are extracted from the operation and maintenance dataset. These keyword groups are then matched with fault feature keyword groups in the fault database using weighted similarity. If a corresponding fault subclass can be matched, the specific fault cause is determined based on the association between the subclass and the fault cause database. If no fault subclass can be matched, the process backtracks to the corresponding fault parent class to determine its general fault cause, ultimately forming the corresponding fault diagnosis result.
[0025] S103. When the fault diagnosis result indicates that a fault exists, issue an operation and maintenance fault alarm and perform operation and maintenance fault handling.
[0026] In this embodiment, when the fault diagnosis result indicates the existence of a current fault, the system first triggers a matching maintenance fault alarm type based on the fault level corresponding to the fault diagnosis result. Simultaneously, it generates a warning message containing the specific fault cause (or a general fault cause) and accurately pushes this message to relevant maintenance personnel for timely fault notification. Then, based on the fault cause determined by the fault diagnosis result, the system retrieves the corresponding fault handling operation plan from the fault handling operation database. If the fault is a first-level fault, the fault handling operation is automatically executed directly according to the plan. If the fault is a second-level fault, the handling plan is pushed to the maintenance personnel as a professional reference for their fault handling.
[0027] This invention provides a business operation and maintenance method, comprising collecting multi-source operational data of the business environment at a preset sampling frequency, preprocessing the collected multi-source operational data to form an operation and maintenance dataset; inputting the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain model monitoring results; performing fault diagnosis based on the model monitoring results to obtain fault diagnosis results; and issuing an operation and maintenance fault alarm and executing operation and maintenance fault handling when the fault diagnosis results indicate that a fault exists. This technical solution enables real-time and comprehensive monitoring of the IT system, accurate and intelligent fault diagnosis, timely warning of potential risks, improved operation and maintenance efficiency and accuracy, and reduced risk of fault impact.
[0028] As a first optional embodiment of this example, multi-source operational data of the business environment is collected at a preset sampling frequency, and the collected multi-source operational data is preprocessed to form an operation and maintenance dataset, including: S1011. Collect multi-source operational data of the business environment at differentiated preset sampling frequencies. The multi-source operational data includes hardware and software operational data and operational environment data. The hardware and software operational data includes hardware model, hardware usage time, software version, and software operation log information. The operational environment data includes data center environmental parameters and network performance indicators. Data center environmental parameters include temperature, humidity, voltage, and current. Network performance indicators include network bandwidth and network latency.
[0029] In this embodiment, hardware and software operation data can be understood as data reflecting the basic attributes of IT system hardware and the operational status of software, including four core categories: hardware model, hardware usage time, software version, and software operation log information. Operational environment data can be understood as quantitative indicator data of the physical data center and network environment supporting the operation of the IT system, divided into two categories: data center environment parameters and network performance indicators. Data center environment parameters are indicators characterizing the physical operating environment of the data center, including temperature, humidity, voltage, and current; network performance indicators are indicators reflecting the network operation and transmission status, including network bandwidth and network latency.
[0030] Specifically, the hardware and software information acquisition module and the runtime environment information acquisition module work together to complete the differentiated acquisition of multi-source data. The hardware and software information acquisition module includes a sensor data acquisition subunit and a program implantation monitoring subunit. The sensor data acquisition subunit collects hardware operation data such as hardware model and hardware usage time through various sensors installed in the hardware devices, with a sampling frequency of, for example, 1Hz. The program implantation monitoring subunit records software operation data such as software version and software operation log information through a monitoring program embedded in the software, with a sampling frequency of, for example, 0.5Hz. The runtime environment information acquisition module includes an environmental sensor unit and a network monitoring unit. The environmental sensor unit is deployed in various areas of the computer room to collect computer room environmental parameters such as temperature, humidity, voltage, and current. The temperature measurement range can be, for example, 0-50℃ with an accuracy of ±0.5℃, and the humidity measurement range can be, for example, 20%-90%RH with an accuracy of ±3%RH. The network monitoring unit monitors network performance indicators such as network bandwidth utilization and network latency in real time through network probes. The network latency measurement accuracy can be, for example, ±1ms. Ultimately, this achieves real-time, multi-dimensional, and differentiated data acquisition of hardware and software, computer room environment, and network status in the business environment.
[0031] S1012. Perform data cleaning and standardization on the software and hardware operation data and the operation environment data respectively. Merge the cleaned and standardized software and hardware operation data and the operation environment data according to the collection timestamp to form an operation and maintenance dataset.
[0032] In this embodiment, preprocessing operations are performed through a data preprocessing module. This module includes a data cleaning subunit, a data standardization subunit, and a data fusion subunit. First, the hardware and software operation data and operating environment data collected by S1011 are classified and processed independently. The data cleaning subunit uses the Z-score method for outlier detection, using the formula... Calculate the Z-value of the data, where, The original data values, The mean of the data. For the standard deviation of the data, when Values identified as outliers are removed, eliminating noise and invalid values from both data types. The subsequent data standardization sub-unit employs the min-max standardization method, using the formula... Data normalization is achieved by transforming two types of data of different magnitudes and dimensions to the same magnitude. The original data values, The minimum value of the data. For the maximum value of the data, This is the standardized data. After cleaning and standardization, the data fusion subunit uses the collection timestamp of each data point as a matching benchmark to integrate and fuse the processed hardware and software operation data and operation environment data, breaking down the heterogeneity of multi-source data and ultimately constructing a unified multi-dimensional operation and maintenance dataset.
[0033] As a second optional embodiment of this example, the operation and maintenance dataset is input into a pre-trained operation and maintenance monitoring model to obtain model monitoring results. Based on the model monitoring results, fault diagnosis is performed to obtain fault diagnosis results, including: S1021. Input the operation and maintenance dataset into the pre-trained operation and maintenance monitoring model to obtain the model monitoring results output by the operation and maintenance monitoring model.
[0034] In this embodiment, the real-time data receiving subunit within the IT operations and maintenance monitoring module receives the pre-processed operations and maintenance dataset via a message queue and inputs the dataset into the pre-trained operations and maintenance monitoring model in real time. The status analysis subunit within the IT operations and maintenance monitoring module performs intelligent analysis and processing on the operations and maintenance dataset, and, in conjunction with the normal operating characteristics of the IT system trained in the model, outputs the corresponding business operation status evaluation result, i.e., the model monitoring result, providing a basis for status determination for subsequent fault diagnosis.
[0035] S1022. If the model monitoring results indicate that the business operation status is normal, then the fault diagnosis result is determined to be no fault.
[0036] In this embodiment, "normal business operation status" means that, after model analysis, all indicators of the IT system's hardware, software, and operating environment meet reasonable operating conditions. The fault diagnosis result is the final judgment on whether a fault exists in the IT system and its cause, including three categories: no fault, specific fault cause, and general fault cause. "No fault" is the basic type of fault diagnosis result, indicating that the IT system currently has no potential faults or actual faults.
[0037] Specifically, the system performs status identification on the model monitoring results output by the operation and maintenance monitoring model. If the model monitoring results clearly indicate that the business operation status of the IT system software and hardware operation, data center environment, network status and other dimensions are all in the normal range and meet the normal operation conditions of model training, then the fault diagnosis module will directly make a judgment and determine the fault diagnosis result as no fault, without the need to perform subsequent abnormal fault investigation operations.
[0038] S1023. If the model monitoring results indicate that the business operation status is abnormal, extract the abnormal data feature keyword groups from the operation and maintenance dataset, match the abnormal data feature keyword groups with the fault feature keyword groups in the fault database, and determine the cause of the business fault based on the matching results and the fault cause database, thus forming a fault diagnosis result.
[0039] In this embodiment, "abnormal business operation status" refers to the deviation of one or more operational indicators of the IT system from reasonable operating conditions, as discovered by model analysis. The abnormal data feature keyword set can be understood as a set of keywords extracted from the operations and maintenance dataset that characterizes abnormal business operation features. The fault database can be understood as a fault feature library constructed according to a multi-level classification architecture of parent-child categories, storing feature keyword sets corresponding to various types of faults. The fault cause database is a database associated with the fault database, storing the fault causes corresponding to each fault parent / child category.
[0040] Specifically, if the system identifies an anomaly in the business operation status as indicated by the model monitoring results, the fault diagnosis module will perform fault investigation. First, it extracts abnormal data feature keyword groups that reflect the core characteristics of the anomaly from the operation and maintenance dataset. Then, the fault feature matching subunit of the fault diagnosis module matches these keyword groups with various fault feature keyword groups pre-stored in the fault database. Subsequently, it combines the fault occurrence cause database, which is deeply associated with the fault database, to infer and determine the specific fault cause that led to the abnormal business operation based on the feature matching results, and finally forms the corresponding fault diagnosis result.
[0041] Furthermore, the abnormal data feature keyword groups are matched with the fault feature keyword groups in the fault database. Based on the matching results and the fault occurrence cause database, the cause of the business fault is determined, and a fault diagnosis result is formed, including: a1. Calculate the weight of the abnormal data feature keyword group, and perform similarity matching between the weight of the abnormal data feature keyword group and the weight of the fault feature keyword group in the fault database.
[0042] In this embodiment, the weight is the quantified value of the abnormal data feature keywords in the corresponding abnormal category, calculated using the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, which reflects the feature importance of the keywords. The TF-IDF algorithm is an algorithm used to calculate the feature keyword weights, which determines the keyword weights by multiplying the term frequency and the inverse document frequency.
[0043] Specifically, the fault diagnosis module uses the TF-IDF algorithm to calculate the weight of each keyword in the abnormal data feature keyword group in turn, forming the weight distribution corresponding to the anomaly. Then, this weight distribution is compared with the weight distribution of fault feature keyword groups corresponding to all fault parent and child classes in the fault database one by one. The degree of matching between the abnormal data and each fault feature is determined by quantitative calculation, providing a matching basis for subsequent fault cause determination.
[0044] b1. If a corresponding fault subclass can be matched, the cause of the business fault corresponding to the fault subclass is determined based on the association between the fault subclass and the fault cause database, and a fault diagnosis result is formed.
[0045] In this embodiment, a fault subclass can be understood as a further subdivision of the same parent fault class in the fault database based on differences in fault causes and processing operations. Each fault subclass corresponds to a unique fault cause. The association between the fault subclass and the fault cause database is a mapping relationship between the two databases, used to achieve a one-to-one correspondence between fault categories and fault causes. The business fault cause corresponding to a fault subclass is a specific fault formation cause exclusive to that subclass, representing a precise fault diagnosis conclusion.
[0046] Specifically, if the abnormal data feature keyword group can accurately match a certain fault subclass in the fault database after weighted similarity matching, the cause reasoning subunit of the fault diagnosis module will directly retrieve the unique and specific fault cause corresponding to the fault subclass based on the preset association mapping relationship between the fault subclass and the fault cause database, determine it as the cause of the business fault, and form a fault diagnosis result accordingly.
[0047] c1. If the corresponding fault subclass cannot be matched, backtrack to the corresponding fault parent class. Based on the association between the fault parent class and the fault cause database, determine the general business fault cause corresponding to the fault parent class and form a fault diagnosis result.
[0048] In this embodiment, the fault parent class can be understood as a basic fault type categorized based on the external symptoms of the fault. A fault parent class can contain multiple fault subclasses, corresponding to common fault causes. The common business fault cause is the cause corresponding to the fault parent class, which can cover the commonalities of faults in all subclasses under that parent class, and serves as a fallback conclusion for fault diagnosis.
[0049] Specifically, if, after weighted similarity matching, the abnormal data feature keyword group cannot accurately match any fault subclass in the fault database, the fault diagnosis module triggers a downgrade processing rule, backtracking the matching scope to the fault parent class corresponding to the anomaly. Then, based on the association between the fault parent class and the fault cause database, it retrieves the general business fault cause corresponding to the fault parent class, identifies this general cause as the business fault cause, and forms a fallback fault diagnosis result to ensure the effectiveness of the fault diagnosis.
[0050] As a third optional embodiment of this example, when the fault diagnosis result indicates that a fault currently exists, an operation and maintenance fault alarm is issued, and operation and maintenance fault handling is performed, including: S1031. Based on the fault level of the fault diagnosis result, trigger the corresponding type of operation and maintenance fault alarm, generate warning information containing the fault cause, and push the warning information to relevant operation and maintenance personnel. The operation and maintenance fault alarm shall include at least sound alarm, light alarm, and SMS alarm.
[0051] In this embodiment, the fault level is a fault category classified according to the degree of abnormality, scope of impact, and difficulty of handling the IT system fault, and serves as the basis for determining different alarm types. Operational fault alarms are multi-type alerts issued by the warning module after detecting a fault, including three basic alarm types: sound, light, and SMS. Warning messages are fault notifications containing specific / general causes of the fault, serving as a reference for operations and maintenance personnel in handling the fault; the relevant operations and maintenance personnel are the professional technicians responsible for troubleshooting and handling IT system faults, and are the recipients of the warning messages.
[0052] Specifically, the system's alert module includes an alarm triggering subunit and an information push subunit. First, the alarm triggering subunit triggers a matching type of operation and maintenance fault alarm based on the fault level corresponding to the fault diagnosis result, triggering at least one or more of the following types of alarms: sound, light, and SMS, to achieve immediate fault notification. At the same time, it generates an alert message containing the specific cause of the business fault (the fault subclass corresponds to the precise cause / the fault parent class corresponds to the general cause). Then, the information push subunit accurately pushes the alert message to the relevant operation and maintenance personnel via email, instant messaging tools, etc., to ensure that the operation and maintenance personnel grasp the core fault information as soon as possible and prepare for subsequent fault handling.
[0053] S1032. Based on the cause of the business failure determined by the fault diagnosis results, retrieve the corresponding fault handling operation plan from the fault handling operation database. If the failure is a first-level failure, execute the fault handling according to the fault handling operation plan. If the failure is a second-level failure, push the fault handling operation plan to the relevant operation and maintenance personnel as a reference for fault handling.
[0054] In this embodiment, the fault handling operation database is a pre-built database that stores corresponding handling schemes for each fault parent / child class, with multiple handling schemes preset for each type of fault. The fault handling operation scheme is a standardized fault handling process and operation method formulated for specific fault causes / general fault causes. Level 1 faults are those with low handling difficulty and simple operation procedures, supporting automatic system execution of handling operations; Level 2 faults are those with high handling difficulty and require professional human judgment, requiring maintenance personnel to manually handle them by referring to the scheme.
[0055] Specifically, the system's IT operations and maintenance module includes a solution retrieval subunit and an automatic execution subunit. First, the solution retrieval subunit, based on the cause of the business fault determined by the fault diagnosis results, accurately retrieves the corresponding fault handling operation solutions with a high success rate from the fault handling operation database. If the fault is determined to be a first-level fault, the automatic execution subunit directly executes the fault handling operation according to the retrieved high-success-rate solution, achieving unmanned and rapid handling of simple faults. If the fault is determined to be a second-level fault, the retrieved high-success-rate solution is pushed to relevant operations and maintenance personnel as an operational reference for their fault investigation and handling work, improving the efficiency and accuracy of manual fault handling.
[0056] As a fourth optional embodiment of this example, the method further includes: a2. Select samples from historical fault data to construct a training set. Divide the historical fault data in the training set into fault parent classes based on different external fault phenomena. If the fault causes or fault handling operations corresponding to the same fault parent class are different, divide the fault parent class into multiple fault subclasses. Construct a fault database with a hierarchical structure of parent class-subclass based on the hierarchical historical fault data. At the same time, construct the corresponding fault cause database and fault handling operation database.
[0057] In this embodiment, the training set is a sample set selected from historical fault data for training the operation and maintenance monitoring model, with at least 100 samples selected for each type of fault. The external manifestations of a fault are the abnormal behaviors and status characteristics that are intuitively presented when an IT system and operating environment malfunction. These can be directly captured through multi-source operation and maintenance data such as hardware and software operation data, data center environment parameters, and network performance indicators, covering all dimensions of IT operation and maintenance, including hardware and software, data center environment, and network transmission. The fault parent class is the basic fault category based on the external manifestations of the fault and is the upper-level classification of fault grading. Fault subclasses are further subdivisions of the fault parent class, based on the differences in the causes or handling operations of faults within the same parent class. The fault database, fault cause database, and fault handling operation database are related databases built according to a multi-level parent-subclass architecture, storing fault characteristics, root causes, and standardized fault handling schemes, respectively, serving as a knowledge base for model training and fault diagnosis.
[0058] Specifically, the training set construction subunit of the AI model training module first selects samples from historical fault data to construct the model training set, ensuring that the number of samples for each type of fault is no less than 100. Then, based on the external phenomena of the faults, the historical fault data in the training set is initially classified into different fault parent classes. Each fault parent class is further analyzed. If it is found that there are different causes of faults or different fault handling operations under the same parent class, the parent class is further subdivided into multiple fault subclasses. Finally, based on the historical fault data that has completed the parent-subclass hierarchical classification, a hierarchical fault database, as well as a related fault cause database and fault handling operation database, are constructed to provide knowledge data support for subsequent model training.
[0059] b2. Construct an operation and maintenance monitoring model based on convolutional neural networks, extract fault feature keyword groups from historical fault data in the fault database, calculate the weight of the fault feature keyword groups, and associate the fault feature keyword groups with the fault causes in the fault cause database and the fault handling operations in the fault handling operation database through the weight.
[0060] In this embodiment, a Convolutional Neural Network (CNN) is the basic network architecture for building the operation and maintenance monitoring model. It includes an input layer, convolutional layers, pooling layers, fully connected layers, and an output layer. The convolutional layers use 3×3 convolutional kernels, and the pooling layers use 2×2 max pooling. The fault feature keyword group is a set of keywords extracted from historical fault data in the fault database that can characterize the core features of various faults. The weight is a quantified value of the feature keywords calculated using the TF-IDF algorithm, reflecting the importance of the keywords in the corresponding fault category. The TF-IDF algorithm calculates the keyword weight by multiplying the term frequency (TF) and the inverse document frequency (IDF), serving as the basis for associating fault features with their causes and handling solutions.
[0061] Specifically, the model structure design subunit and the model training subunit of the AI model training module work together. First, the model structure design subunit builds an operation and maintenance monitoring model based on a convolutional neural network, configuring a core network layer with 3×3 convolutional kernels and 2×2 max pooling. Then, the model training subunit extracts fault feature keyword groups corresponding to various fault parent and child classes from historical fault data in the fault database. Subsequently, the TF-IDF algorithm is used to calculate the weight of each feature keyword, using the formula... The quantized weight values are obtained, where, Keywords In the document The word frequency, i.e., keywords In the document The number of times it appears in the document The ratio of the total number of words in the text; Inverse document frequency, , Total number of documents For keywords The document count; finally, by using the weights of feature keyword groups, a mapping relationship is established between various fault features in the fault database and the corresponding fault causes in the fault cause database, and the corresponding handling solutions in the fault handling operation database.
[0062] c2. Input the associated fault data into the operation and maintenance monitoring model for training until the training termination condition is met, and obtain a fully trained operation and maintenance monitoring model.
[0063] In this embodiment, the associated fault data is structured fault data that integrates fault feature keyword groups, corresponding weights, and establishes a mapping relationship with fault causes and fault handling operations. It serves as the input data for model training. The training termination condition can be understood as the condition that the model can accurately complete fault feature identification, cause matching, and handling solution association based on the input fault data, and that the model's fault diagnosis capability reaches a preset standard.
[0064] Specifically, the model training subunit of the AI model training module will complete the structured fault data, which associates features, causes, and solutions, and adapt it to the input format requirements of convolutional neural networks before inputting it into the pre-built operation and maintenance monitoring model for training. The model extracts deep features from the fault data through convolutional and pooling layers, and learns the correlation between fault feature weights and fault causes and handling operations by combining fully connected layers. The training process is continuously iterated until the model can stably and accurately output the corresponding fault causes and handling solutions based on the fault features, meeting the preset training termination conditions, and finally obtaining a fully trained operation and maintenance monitoring model.
[0065] d2. Test and evaluate the fully trained operation and maintenance monitoring model, and continuously iterate and optimize it after the operation and maintenance monitoring model is put into business application.
[0066] In this embodiment, the model optimization module includes a test set construction subunit, a model evaluation subunit, and a model update subunit. The test set construction subunit selects a predetermined number (e.g., 1000) IT fault cases from the internet and constructs a test set containing a test fault database, a test fault cause database, and a test fault handling operation database. Then, the model evaluation subunit uses the test set as input to the operation and maintenance monitoring model for testing, calculating the model's accuracy. The accuracy calculation formula is: .in, True cases refer to the number of failure cases correctly predicted by the model. True negatives are the number of normal cases correctly predicted by the model. False positives are the number of normal cases that the model incorrectly predicts as faults. False negatives are the number of fault cases that the model incorrectly predicted as normal. When the accuracy is 99.5% or higher, the IT operations monitoring model is considered sufficiently accurate. If the accuracy is below 99.5%, the model is retrained using these 1000 cases as the training set, and then tested again until the accuracy reaches 99.5% or higher. The model update subunit continuously collects new fault case data generated during operations and maintenance, and combines this with the evaluation results of actual model operation to continuously adjust the model's parameters and optimize its structure, ensuring the model continuously adapts to the dynamically changing IT operations and maintenance environment and always guarantees the accuracy and practicality of fault diagnosis.
[0067] This invention discloses an IT operations and maintenance method and system based on an AI-powered large-scale model, belonging to the field of AI-driven IT operations and maintenance. It aims to solve the technical problems of traditional manual IT operations and maintenance, such as error-prone fault diagnosis, low processing efficiency, lack of real-time and comprehensive monitoring, and high probability of fault occurrence. The system is meticulously composed of nine modules: hardware and software information acquisition, operating environment information acquisition, data preprocessing, AI model training, IT operations and maintenance monitoring, fault diagnosis, alerts, IT operations and maintenance processing, and model optimization. Each module has refined sub-units equipped with dedicated technical parameters and processing algorithms, forming a complete intelligent operations and maintenance system. The corresponding operations and maintenance method sequentially includes five steps: data collection and preprocessing, database establishment, model training, model testing and optimization, and real-time monitoring and fault handling, achieving intelligent management and control of the entire process from data collection to fault handling.
[0068] First, a multi-dimensional data fusion monitoring method for hardware and software environments integrates hardware sensor data sampled at 1Hz and software operation logs sampled at 0.5Hz with environmental parameters such as room temperature and humidity, and performance indicators such as network latency, using a unified timestamp to construct a global multi-dimensional operation and maintenance data stream, accurately identifying complex fault hazards caused by multiple intertwined factors. Second, a TF-IDF weighted feature keyword diagnosis method for IT operation and maintenance logs calculates the weight of fault feature keywords using the TF-IDF algorithm, and performs similarity matching between the weight distribution of real-time abnormal logs and the weighted feature set of the fault knowledge base, achieving quantitative and accurate diagnosis of fault root causes and reducing reliance on expert experience. Third, a multi-level fault classification and degradation processing method based on detailed causes and operations constructs a parent-child hierarchical fault knowledge base according to the differences in fault causes and processing operations, with a sample size of at least 100 faults per category. When a subcategory cannot be accurately matched, it automatically backtracks to the parent category and executes a general high-success-rate processing solution, taking into account both accurate handling of known faults and effective response to unknown faults.
[0069] Compared to traditional manual operation and maintenance methods, this invention relies on real-time comprehensive collection and standardized processing of multi-source data to provide a high-quality data foundation for AI models. The operation and maintenance monitoring model built on convolutional neural networks has been tested and verified with 1,000 fault cases, achieving an accuracy rate of 99.5% or higher, and can continuously adapt to dynamic operation and maintenance environments. The IT operation and maintenance monitoring, fault diagnosis, alert, and handling modules form a closed-loop process of fault discovery, diagnosis, alarm, and handling. It can not only intelligently analyze and diagnose faults, quickly retrieve and automatically execute handling solutions as needed, improving the efficiency and accuracy of fault judgment and handling, but also monitor the status of IT system hardware and software and operating environment in real time, promptly discover potential fault hazards and intervene in advance, reducing the probability of fault occurrence and the impact on business operations, thus forming an efficient, intelligent, and robust IT operation and maintenance solution.
[0070] In one embodiment, Figure 2This is a schematic diagram of the structure of a business operation and maintenance system provided in an embodiment of the present invention. For example... Figure 2 As shown, the system includes: The data acquisition module 21 is used to collect multi-source operational data of the business environment at a preset sampling frequency, and to preprocess the collected multi-source operational data to form an operation and maintenance dataset. The operation and maintenance monitoring module 22 is used to input the operation and maintenance dataset into the pre-trained operation and maintenance monitoring model to obtain the model monitoring results, and perform fault diagnosis based on the model monitoring results to obtain the fault diagnosis results. The fault handling module 23 is used to issue an operation and maintenance fault alarm and perform operation and maintenance fault handling when the fault diagnosis result indicates that a fault exists.
[0071] The business operation and maintenance system adopted in this technical solution enables real-time and comprehensive monitoring of IT systems, accurate and intelligent fault diagnosis, timely early warning of potential risks, improved operation and maintenance efficiency and accuracy, and reduced risk of failure impact.
[0072] Optionally, the data acquisition module 21 is specifically used for: Multi-source operational data of the business environment is collected at a differentiated preset sampling frequency. The multi-source operational data includes hardware and software operational data and operational environment data. The hardware and software operational data includes hardware model, hardware usage time, software version and software operation log information. The operational environment data includes data center environmental parameters and network performance indicators. The data center environmental parameters include temperature, humidity, voltage and current. The network performance indicators include network bandwidth and network latency. The software and hardware operation data and the operation environment data are cleaned and standardized respectively. The cleaned and standardized software and hardware operation data and the operation environment data are then merged according to the collection timestamp to form an operation and maintenance dataset.
[0073] Optionally, the operation and maintenance monitoring module 22 includes: The model monitoring unit is used to input the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results output by the operation and maintenance monitoring model. The first diagnostic unit is used to determine that the fault diagnosis result is no fault if the model monitoring result indicates that the business operation status is normal. The second diagnostic unit is used to extract abnormal data feature keyword groups from the operation and maintenance dataset if the model monitoring results indicate that the business operation status is abnormal, match the abnormal data feature keyword groups with the fault feature keyword groups in the fault database, and determine the cause of the business fault based on the matching results and the fault cause database, thereby forming a fault diagnosis result.
[0074] Optionally, the second diagnostic unit is specifically used for: Calculate the weight of the abnormal data feature keyword group, and perform similarity matching between the weight of the abnormal data feature keyword group and the weight of the fault feature keyword group in the fault database; If a corresponding fault subclass can be matched, the cause of the business fault corresponding to the fault subclass is determined based on the association between the fault subclass and the fault cause database, and a fault diagnosis result is formed. If the corresponding fault subclass cannot be matched, the fault parent class is traced back. Based on the association between the fault parent class and the fault cause database, the general business fault cause corresponding to the fault parent class is determined, and a fault diagnosis result is formed.
[0075] Optionally, the fault handling module 23 is specifically used for: Based on the fault level of the fault diagnosis result, the corresponding type of operation and maintenance fault alarm is triggered, a warning message containing the cause of the fault is generated, and the warning message is pushed to the relevant operation and maintenance personnel. The operation and maintenance fault alarm includes at least sound alarm, light alarm, and SMS alarm. Based on the cause of the business failure determined by the fault diagnosis results, the corresponding fault handling operation plan is retrieved from the fault handling operation database. If the failure is a first-level failure, the fault handling is performed according to the fault handling operation plan. If the failure is a second-level failure, the fault handling operation plan is pushed to the relevant operation and maintenance personnel as an operation reference for fault handling.
[0076] Optionally, the apparatus further includes a model training unit, specifically used for: Samples are selected from historical fault data to construct a training set. Based on different external fault phenomena, the historical fault data in the training set are divided into fault parent classes. If the fault causes or fault handling operations corresponding to the same fault parent class are different, multiple fault subclasses are divided under the fault parent class. A fault database with a parent-subclass hierarchical architecture is constructed based on the hierarchical historical fault data. At the same time, a corresponding fault cause database and fault handling operation database are constructed. An operation and maintenance monitoring model is constructed based on a convolutional neural network. Fault feature keyword groups are extracted from historical fault data in the fault database, the weight of the fault feature keyword groups is calculated, and the fault feature keyword groups are associated with the fault causes in the fault occurrence cause database and the fault handling operations in the fault handling operation database through the weight. The associated fault data is input into the operation and maintenance monitoring model for training until the training end condition is met, resulting in a fully trained operation and maintenance monitoring model. The fully trained operation and maintenance monitoring model is tested and evaluated, and iterative optimization is continuously carried out after the operation and maintenance monitoring model is put into business application.
[0077] The business operation and maintenance system provided in the embodiments of the present invention can execute the business operation and maintenance method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0078] In one embodiment, Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 3 The diagram illustrates a schematic representation of an electronic device 10 that can be used to implement embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0079] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0080] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0081] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as business operations methods.
[0082] In some embodiments, the business operation and maintenance method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the business operation and maintenance method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the business operation and maintenance method by any other suitable means (e.g., by means of firmware).
[0083] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0084] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0085] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0086] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0087] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0088] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0089] This invention also provides a computer program product, including a computer program that, when executed by a processor, can implement the business operation and maintenance method provided in any embodiment of this application.
[0090] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0091] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0092] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A business operation and maintenance method, characterized in that, include: Collect multi-source operational data of the business environment at a preset sampling frequency, and preprocess the collected multi-source operational data to form an operation and maintenance dataset; The operation and maintenance dataset is input into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results. Based on the model monitoring results, fault diagnosis is performed to obtain the fault diagnosis results. When the fault diagnosis result indicates that a fault exists, an operation and maintenance fault alarm is issued and operation and maintenance fault handling is performed.
2. The method according to claim 1, characterized in that, The process involves collecting multi-source operational data of the business environment at a preset sampling frequency, and preprocessing the collected multi-source operational data to form an operation and maintenance dataset, including: Multi-source operational data of the business environment is collected at a differentiated preset sampling frequency. The multi-source operational data includes hardware and software operational data and operational environment data. The hardware and software operational data includes hardware model, hardware usage time, software version and software operation log information. The operational environment data includes data center environmental parameters and network performance indicators. The data center environmental parameters include temperature, humidity, voltage and current. The network performance indicators include network bandwidth and network latency. The software and hardware operation data and the operation environment data are cleaned and standardized respectively. The cleaned and standardized software and hardware operation data and the operation environment data are then merged according to the collection timestamp to form an operation and maintenance dataset.
3. The method according to claim 1, characterized in that, The step involves inputting the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain model monitoring results, and then performing fault diagnosis based on the model monitoring results to obtain fault diagnosis results, including: The operation and maintenance dataset is input into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results output by the operation and maintenance monitoring model. If the model monitoring results indicate that the business operation status is normal, then the fault diagnosis result is determined to be no fault. If the model monitoring results indicate an abnormal business operation status, extract the abnormal data feature keyword groups from the operation and maintenance dataset, match the abnormal data feature keyword groups with the fault feature keyword groups in the fault database, and determine the cause of the business fault based on the matching results and the fault cause database, thus forming a fault diagnosis result.
4. The method according to claim 3, characterized in that, The step of matching the abnormal data feature keyword groups with the fault feature keyword groups in the fault database, and determining the cause of the business fault based on the matching results and the fault occurrence cause database, and forming a fault diagnosis result, includes: Calculate the weight of the abnormal data feature keyword group, and perform similarity matching between the weight of the abnormal data feature keyword group and the weight of the fault feature keyword group in the fault database; If a corresponding fault subclass can be matched, the cause of the business fault corresponding to the fault subclass is determined based on the association between the fault subclass and the fault cause database, and a fault diagnosis result is formed. If the corresponding fault subclass cannot be matched, the fault parent class is traced back. Based on the association between the fault parent class and the fault cause database, the general business fault cause corresponding to the fault parent class is determined, and a fault diagnosis result is formed.
5. The method according to claim 1, characterized in that, When the fault diagnosis result indicates that a fault exists, an operation and maintenance fault alarm is issued, and operation and maintenance fault handling is performed, including: Based on the fault level of the fault diagnosis result, the corresponding type of operation and maintenance fault alarm is triggered, a warning message containing the cause of the fault is generated, and the warning message is pushed to the relevant operation and maintenance personnel. The operation and maintenance fault alarm includes at least sound alarm, light alarm, and SMS alarm. Based on the cause of the business failure determined by the fault diagnosis results, the corresponding fault handling operation plan is retrieved from the fault handling operation database. If the failure is a first-level failure, the fault handling is performed according to the fault handling operation plan. If the failure is a second-level failure, the fault handling operation plan is pushed to the relevant operation and maintenance personnel as an operation reference for fault handling.
6. The method according to claim 1, characterized in that, Also includes: Samples are selected from historical fault data to construct a training set. Based on different external fault phenomena, the historical fault data in the training set are divided into fault parent classes. If the fault causes or fault handling operations corresponding to the same fault parent class are different, multiple fault subclasses are divided under the fault parent class. A fault database with a parent-subclass hierarchical architecture is constructed based on the hierarchical historical fault data. At the same time, a corresponding fault cause database and fault handling operation database are constructed. An operation and maintenance monitoring model is constructed based on a convolutional neural network. Fault feature keyword groups are extracted from historical fault data in the fault database, the weight of the fault feature keyword groups is calculated, and the fault feature keyword groups are associated with the fault causes in the fault occurrence cause database and the fault handling operations in the fault handling operation database through the weight. The associated fault data is input into the operation and maintenance monitoring model for training until the training end condition is met, resulting in a fully trained operation and maintenance monitoring model. The fully trained operation and maintenance monitoring model is tested and evaluated, and iterative optimization is continuously carried out after the operation and maintenance monitoring model is put into business application.
7. A business operation and maintenance system, characterized in that, include: The data acquisition module is used to collect multi-source operational data of the business environment at a preset sampling frequency, and to preprocess the collected multi-source operational data to form an operation and maintenance dataset. The operation and maintenance monitoring module is used to input the operation and maintenance dataset into a pre-trained operation and maintenance monitoring model to obtain the model monitoring results, and to perform fault diagnosis based on the model monitoring results to obtain the fault diagnosis results. The fault handling module is used to issue an operation and maintenance fault alarm and perform operation and maintenance fault handling when the fault diagnosis result indicates that a fault exists.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a business operation and maintenance method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute a business operation and maintenance method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements a business operation and maintenance method according to any one of claims 1-6.