Fault detection method and device of service system, storage medium and electronic equipment
By acquiring and analyzing log data from business systems and utilizing the matching relationship between log features and a fault log feature database, the system automatically detects and locates fault causes, generates repair information, and solves the problem of low fault monitoring efficiency in existing technologies, achieving efficient fault handling and accurate fault location.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, fault monitoring of business systems is inefficient, unable to accurately identify abnormal patterns, resulting in high false alarm and false negative rates, and failing to quickly locate the root cause of faults, leading to low processing efficiency.
By acquiring the log data set of the business system and using the matching relationship between log features and the fault log feature library, the system can automatically detect the operating mode, generate location information and repair information, and send them to the maintenance account to achieve automatic location of fault causes and instruction for repair operations.
It improves the accuracy and efficiency of fault location and processing, reduces the false alarm rate, and enables rapid response and efficient fault repair.
Smart Images

Figure CN121841963A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more specifically, to a fault detection method and apparatus for a business system, a storage medium, and an electronic device. Background Technology
[0002] In existing technologies, fault monitoring of business systems typically relies on real-time data analysis platforms. These solutions generally collect log data generated during the operation of the business system and compare each log entry with preset abnormal logs. When an abnormal log is detected, an alarm is generated to notify maintenance personnel of potential anomalies. However, the above methods have significant limitations: Firstly, the alarm information generated by these methods is often not accurate or comprehensive enough. When faced with complex system faults, it is difficult to accurately identify specific abnormal patterns, resulting in low monitoring accuracy and a high false alarm and false negative rate in practical applications. Secondly, these methods can only achieve anomaly alerts; the alarm information only indicates that an anomaly has occurred and cannot pinpoint the root cause of the fault. The root cause still needs to be manually investigated and analyzed, resulting in low fault handling efficiency and difficulty in quickly responding to system anomalies.
[0003] There is currently no effective solution to the problem of low efficiency in fault monitoring of business systems in related technologies. Summary of the Invention
[0004] The main objective of this application is to provide a fault detection method and apparatus, storage medium and electronic device for business systems, in order to solve the problem of low efficiency in fault monitoring of business systems in related technologies.
[0005] To achieve the above objectives, according to one aspect of this application, a fault detection method for a business system is provided. The method includes:
[0006] During the operation of the target business in the business system, the log data set generated by one or more functional subsystems included in the business system is obtained. The business system executes the corresponding function through one or more functional subsystems to run the target business, and the log data set includes the log data generated by each functional subsystem.
[0007] The current operating mode of the business system is detected by matching the log features of the log data set with the fault log features in the fault log feature library. The fault log features are the log features of the log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode.
[0008] If the operating mode is detected to be the target fault operating mode corresponding to the target fault log characteristics, target location information and target repair information are generated according to the target fault operating mode. The target location information is used to locate the reason why the business system is in the target fault operating mode, and the target repair information is used to indicate the repair operation required to switch the business system from the target fault operating mode to the normal operating mode.
[0009] Send target location information and target repair information to the maintenance account of the business system.
[0010] To achieve the above objectives, according to another aspect of this application, a fault detection device for a business system is provided. The device includes:
[0011] The acquisition module is used to acquire a set of log data generated by one or more functional subsystems included in the business system during the execution of the target business in the business system. The business system executes the corresponding functions through one or more functional subsystems to run the target business, and the log data set includes the log data generated by each functional subsystem.
[0012] The detection module is used to detect the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library. The fault log features are the log features of the log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode.
[0013] The generation module is used to generate target location information and target repair information based on the target fault operation mode when the detected operation mode belongs to the target fault operation mode corresponding to the target fault log characteristics. The target location information is used to locate the reason why the business system is in the target fault operation mode, and the target repair information is used to indicate the repair operation required to switch the business system from the target fault operation mode to the normal operation mode.
[0014] The sending module is used to send target location information and target repair information to the maintenance account of the business system.
[0015] In this embodiment, by acquiring the log data set generated by one or more functional subsystems included in the business system, and detecting the current operating mode of the business system based on the matching relationship between the log characteristics of the log data set and the fault log characteristics in the fault log characteristic library, and then, when the operating mode is detected to be the target fault operating mode corresponding to the target fault log characteristic, target location information and target repair information are generated according to the target fault operating mode, and the target location information and target repair information are sent to the maintenance account of the business system. This achieves the purpose of automatically locating the reason why the business system is in the target fault operating mode, and instructing the repair operation required to switch from the target fault operating mode to the normal operating mode. This achieves the technical effect of improving the accuracy of fault location and the efficiency of fault handling, and solves the technical problem of low fault detection efficiency caused by the existing technology that can only alarm but cannot locate the root cause. Attached Figure Description
[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0017] Figure 1 A hardware structure block diagram of a computer terminal for implementing a fault detection method for a business system is shown.
[0018] Figure 2 This is a flowchart of a fault detection method for a business system provided according to an embodiment of this application;
[0019] Figure 3 This is a schematic diagram of a log data set provided according to an embodiment of this application;
[0020] Figure 4 This is a flowchart of business system fault detection and repair provided according to the embodiments of this application;
[0021] Figure 5 This is a schematic diagram of a fault detection device for a business system provided according to an embodiment of this application;
[0022] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0026] Example 1
[0027] According to an embodiment of this application, a method embodiment for fault detection of a business system is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0028] The method embodiment provided in Embodiment 1 of this application can be executed in a mobile terminal, computer terminal or similar computing device. Figure 1A hardware structure block diagram of a computer terminal (or mobile device) for implementing a fault detection method for a business system is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0029] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the fault detection method of the business system in this embodiment of the application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the fault detection method of the business system described above. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0031] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0032] The display may be, for example, a touchscreen LCD display that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0033] Under the aforementioned operating environment, this application provides the following: Figure 2 The fault detection method for the business system shown. Figure 2 This is a flowchart of a fault detection method for a business system according to Embodiment 1 of this application.
[0034] Step S101: During the operation of the target business in the business system, obtain the log data set generated by one or more functional subsystems included in the business system. The business system executes the corresponding function through one or more functional subsystems to run the target business, and the log data set includes the log data generated by each functional subsystem.
[0035] Step S102: Detect the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library. Here, the fault log features are the log features of log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode.
[0036] Step S103: If it is detected that the operating mode belongs to the target fault operating mode corresponding to the target fault log feature, target location information and target repair information are generated according to the target fault operating mode. The target location information is used to locate the reason why the business system is in the target fault operating mode, and the target repair information is used to indicate the repair operation required to switch the business system from the target fault operating mode to the normal operating mode.
[0037] Step S104: Send the target location information and target repair information to the maintenance account of the business system.
[0038] By acquiring the log data set generated by one or more functional subsystems included in the business system, and detecting the current operating mode of the business system based on the matching relationship between the log characteristics of the log data set and the fault log characteristics in the fault log characteristic library, and then generating target location information and target repair information according to the target fault operating mode when the detected operating mode belongs to the target fault operating mode corresponding to the target fault log characteristics, and sending the target location information and target repair information to the maintenance account of the business system, the purpose of automatically locating the cause of the business system being in the target fault operating mode is achieved, and indicating the repair operation required to switch from the target fault operating mode to the normal operating mode is indicated. This achieves the technical effect of improving the accuracy of fault location and the efficiency of fault handling, and solves the technical problem of low fault detection efficiency caused by existing technologies that can only issue alarms but cannot locate the root cause.
[0039] Optionally, in this embodiment, the business system may refer to, but is not limited to, a software system with complete business processing capabilities, such as a financial transaction system, an e-commerce platform, a streaming media service, or an enterprise resource planning system.
[0040] Optionally, in this embodiment, the functional subsystem may be, but is not limited to, a modular component that constitutes the business system and undertakes some of the functional modules of the business system, such as user service, order service, and payment service in a microservice architecture, or log module, authentication module, and database interface module in a monolithic application.
[0041] Optionally, in this embodiment of the application, the target business may be, but is not limited to, a specific business operation initiated by the user on the business system, which is a core task of the business system, such as user login, order submission, payment execution, data query, batch data processing, etc.
[0042] Optionally, in the embodiments of this application, Figure 3 This is a schematic diagram of a log data set provided according to an embodiment of this application, such as... Figure 3 As shown, the log data set can be, but is not limited to, a summary of various log information generated by each functional subsystem during operation, such as: business logs, application logs, application transaction tracking logs, inter-application interaction logs, application start / stop logs (system logs), batch logs, database logs, security logs, performance logs, audit logs, etc. This log data is collected, transmitted, and uniformly stored in a data lake or a dedicated log aggregation platform (such as a platform built on Kafka, ELK Stack, or Splunk) after its generation, for unified analysis and querying. This log data can exist in, but is not limited to, file format, database record format, or data stream format.
[0043] In the embodiments provided in step S101 above, log data sets can be obtained, but are not limited to, through proactive pulling by a log broker. For example, a log collection broker such as Filebeat can be deployed on the server to monitor changes in local log files in real time and fetch the changed local log files to a unified data lake or log center. Log data sets can also be obtained, but are not limited to, through proactive pushing by a log framework. For example, integrating a log framework such as Log4j can asynchronously push logs to a message queue (such as Kafka) or repository. Furthermore, this can also be achieved through database replication, API polling, and other methods.
[0044] Optionally, after obtaining the log data set in step S101 and before performing the detection in step S102, this method may further include data cleaning of the log data set. Data cleaning may include, but is not limited to: removing duplicate values: for example, by comparing key fields of the log data (such as the link identifier of the log record, the key business primary key, timestamp, etc.), filtering out log records with completely identical content or duplicate core information, and deleting or merging them to avoid duplicate data affecting subsequent analysis results. Data standardization and normalization: for example, for log data with different formats generated by different functional subsystems (such as timestamp formats such as "YYYY-MM-DD HH:MM:SS" and "MM / DD / YYYY HH:mm", and log levels such as "ERROR", "Err", "Error", etc.), converting them into a unified standard format to ensure the consistency and accuracy of the log data and provide a unified data foundation for subsequent feature extraction. Data denoising: For example, filtering rules can be defined based on regular expressions to remove data in log data that does not meet the needs of business analysis (such as debug logs generated in the test environment, meaningless empty log entries, and invalid data with malformed formats), and retain valid log data.
[0045] Optionally, in this embodiment, log features may be, but are not limited to, key information extracted from the original log data that can reflect the core information of the log or the system's operating status, such as: specific error keywords ("Timeout", "Exception", "Connection Refused"), the frequency of log events, the combination pattern of specific log sequences, or log vectors extracted by a deep learning model.
[0046] Optionally, in this embodiment, the fault log feature library may be, but is not limited to, a pre-built database (such as MySQL or Redis) that stores log features corresponding to known fault modes. For example, the fault log feature library may record the features corresponding to the "database connection pool exhaustion" fault mode as follows: "ConnectionTimeout" appears frequently in the logs of the "Accounting Core Subsystem", and "TNS-12519" error appears in the logs of the "Database Cluster Subsystem".
[0047] Optionally, in this embodiment, the operating mode may be, but is not limited to, the overall operating status of the business system within the current time period, which can reflect whether the system is normal and what kind of fault exists, such as normal operating mode, database connection pool exhaustion mode, payment interface delay mode, batch task failure mode, etc.
[0048] In the embodiment provided in step S102 above, the current operating mode of the business system can be detected by, but is not limited to, matching using a rule engine. For example, a predefined rule can be set, such as "If the frequency of the keyword 'Connection Timeout' in the logs of the 'Accounting Core Subsystem' exceeds 100 times per minute within a 5-minute time window, and the error keyword 'TNS-12519' appears in the logs of the 'Database Cluster Subsystem,' then it is determined that the 'Database Connection Pool Exhaustion' fault log feature is matched." The log features in the log data set are scanned in real time, and when the predefined rule is met, the corresponding fault operating mode is determined as the current operating mode of the business system. The current operating mode of the business system can also be detected by, but is not limited to, vector similarity matching. For example, a deep learning model (such as LSTM or Transformer model) can be used to convert the log data set into log feature vectors; the cosine similarity between the log feature vector and each fault log feature vector stored in the fault log feature library is calculated; if the similarity between a certain fault log feature vector and the real-time log feature vector is higher than a preset threshold (such as 0.9), then the fault operating mode corresponding to the fault log feature is determined as the current operating mode of the business system.
[0049] Optionally, in the embodiments of this application, the target location information may be, but is not limited to, information used to clarify the root cause and location of the business system being in the target fault operation mode, such as structured data (JSON format: {"fault module": "accounting core subsystem (instance core-app-01)", "root cause of fault": "database connection pool core-db-pool is full (Active: 200 / 200)"}), descriptive text ("the root cause of the fault is located in the 'database connection module' of the 'payment subsystem', and the specific exception is 'connection pool full'"), etc.
[0050] Optionally, in this embodiment, the target repair information may be, but is not limited to, an operation guide used to instruct operations personnel to switch the business system from the target fault operation mode to the normal operation mode, such as an ordered list of steps ("1. Immediately restart instances core-app-01, 02, 03; 2. Evaluate and increase the maximum capacity of the connection pool to 300"), the path to an executable automated operations script (" / opt / ops / scripts / restart_app.sh"), and a description of the operation instructions ("Log in to the database management platform and execute the command 'ALTER SYSTEM SET processes=300 SCOPE=SPFILE' to adjust the number of connections"), etc.
[0051] In the embodiment provided in step S103 above, target location information and target repair information can be generated through knowledge base mapping lookup, but are not limited to this method. For example, a fault knowledge base is constructed, with "target fault operation mode" as the key and the corresponding "target location information" and "target repair information" as values. For instance, the key is "database connection pool exhaustion mode," and the value is "Location information: the database connection pool of the accounting core subsystem is full; Repair information: restart the relevant instance and increase the connection pool capacity." When a target fault operation mode is detected, the corresponding key is retrieved from the fault knowledge base to directly obtain the associated target location information and target repair information. Target location information and target repair information can also be dynamically generated through a large language model, but are not limited to this method. For example, the detected "target fault operation mode" and the corresponding abnormal log fragments are used as prompt words and input into a large language model (such as GPT, Llama, etc.) that has been fine-tuned with knowledge from the operations and maintenance domain. Based on the trained operations and maintenance knowledge, the model generates target location information and target repair information in natural language format in real time.
[0052] Optionally, in this application embodiment, the maintenance account of the business system may be, but is not limited to, the receiving channel used by the personnel or team responsible for monitoring and maintaining the operating status of the system, such as: one or more email addresses, group or robot accounts of instant messaging tools (such as Slack, Microsoft Teams), API interfaces of ticketing systems (such as JIRA, ServiceNow), or accounts of SMS alarm platforms, etc.
[0053] In the embodiment provided in step S104 above, target location information and target repair information can be sent via email or SMS gateway, but are not limited to these methods. For example, the Simple Mail Transfer Protocol (SMT) service interface can be called to format the target location information and target repair information into standardized email content (including fault level, fault time, location details, and repair steps) and send it to a preset maintenance team email list; or the API interface provided by the SMS service provider can be called to send a brief fault information to the mobile phone of the on-duty personnel. Target location information and target repair information can also be sent via Webhook interface to push to instant messaging tools or maintenance platforms, but are not limited to these methods.
[0054] Optionally, in this embodiment, the above solution can be explained using, but is not limited to, a "unified payment platform" for processing payment transactions as an example:
[0055] Step 1: Collect logs from each subsystem using Filebeat and clean and standardize them.
[0056] Step 2: The analysis module detected 1,500 "Connection Timeout" logs from the "Accounting Core Subsystem" and 20 "TNS-12519" logs from the "Database Cluster Subsystem". The log features of "High-Frequency Connection Timeout + TNS-12519 Error" were extracted and matched with the fault log feature library. It was found that the similarity with the fault log feature of "Database Connection Pool Exhaustion" reached 0.95 (higher than the preset threshold of 0.9), and it was determined to be "Database Connection Pool Exhaustion".
[0057] Step 3: The system uses the fault knowledge base to map and search, and generates target location information based on the "database connection pool exhausted" pattern: "The database connection pool core-db-pool of the accounting core subsystem (instances core-app-01, 02, 03) is full (Active: 200 / 200)", and generates target repair information: "1. Immediately restart instances core-app-01, 02, 03; 2. Evaluate and increase the maximum capacity of the connection pool to 300".
[0058] Step 4: The system combines the target location information and target repair information into a P0-level alarm, pushes it to the operations and maintenance team channel through the Slack Webhook interface, and simultaneously calls the SMS API to send alarm notifications to on-duty personnel to ensure that operations and maintenance personnel can handle the fault in a timely manner.
[0059] Optionally, in the fault detection method for a business system provided in this application embodiment, detecting the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library includes: extracting features from the log data included in the log data set to obtain log entity features, wherein the log entity features are used to characterize multiple log segments included in the log data and the correlation between the multiple log segments, and the log features include log entity features; and detecting the current operating mode of the business system based on the degree of matching between the log entity features and the fault log features in the fault log feature library.
[0060] Optionally, in the embodiments of this application, the log fragment may be, but is not limited to, a structured or unstructured data fragment with independent information units split from the original log data, such as timestamp ("2025-11-04 09:00:00"), log level (Error, Warn), specific error code ("TNS-12519"), exception stack information ("java.sql.SQLException: Connection refused"), etc.
[0061] Optionally, in this embodiment, the association may be, but is not limited to, the relationship between multiple log segments at the temporal, logical, or contextual level, such as temporal association, contextual association, causal association, etc.
[0062] Optionally, in the embodiments of this application, the log entity features may be, but are not limited to, feature data used to comprehensively characterize multiple log segments and their relationships, such as graph vectors constructed by graph neural networks (nodes are log segments, edges are relationships, and the overall vector of the graph represents log entity features), time-series vectors extracted by LSTM models (containing the time-series relationships of log segments), etc.
[0063] Optionally, in the embodiments of this application, the matching degree may be, but is not limited to, an indicator used to quantify the similarity between real-time log entity features and fault log features in the fault log feature library, such as cosine similarity (value range 0-1, the higher the value, the higher the similarity), Euclidean distance (the smaller the value, the higher the similarity), edit distance (the smaller the value, the more similar the structure), etc.
[0064] Optionally, in this embodiment, feature extraction can be performed through graph construction and embedding, but is not limited to. For example, log fragments (such as service name "Accounting Core Subsystem", error code "ConnectionTimeout", IP address "192.168.1.1") in the log data set can be used as nodes of a graph, and relationships such as "association with the same link ID", "time interval less than 5 seconds", and "causal dependency" can be used as edges of the graph to construct a heterogeneous graph. Graph embedding algorithms (such as Node2Vec and GraphSAGE) can be used to learn the vector representations of nodes and edges in the graph, and finally, a graph-level vector representing the entire graph structure can be generated as the log entity feature. Feature extraction can also be performed through temporal model extraction of contextual features, but is not limited to. For example, the log data set can be grouped according to time order and session ID to form a log sequence of "user login → order submission → payment request → payment failure". The sequence can be input into a Transformer model, and the model can capture the temporal relationship between different log fragments (such as "payment request" and "payment failure") through a self-attention mechanism, and output the [CLS] label vector of the sequence as the log entity feature.
[0065] Optionally, in this embodiment, the current operating mode of the business system can be detected by, but is not limited to, using a vector similarity threshold. For example, the fault log feature library stores standard log entity feature vectors (graph embedding vectors) corresponding to the "database connection pool exhaustion" mode. The system calculates the cosine similarity between the real-time extracted log entity feature vector and the standard vector. If the similarity is 0.92 (higher than the preset threshold of 0.9), the current operating mode is determined to be the "database connection pool exhaustion" fault mode. Alternatively, but not limited to, the current operating mode of the business system can be detected by probability prediction using a classifier model. For example, a pre-trained multilayer perceptron classifier is used as input for real-time log entity features. The classifier outputs the probability that the system belongs to "normal operation mode," "database connection pool exhaustion mode," "payment interface delay mode," etc. If the probability of "database connection pool exhaustion mode" is 0.95 (the highest probability), the current operating mode is determined to be this fault mode.
[0066] Based on the above, feature extraction is performed on log data to obtain log entity features that characterize multiple log segments and their relationships. The operating mode is detected based on the degree of matching between these log entity features and the fault log feature library. This enables the system to understand the context and logical relationships between log events, thereby more accurately identifying complex, cross-subsystem linked fault modes. This avoids the shortcomings of related technologies that rely solely on isolated log events, cannot accurately identify complex faults, and have insufficiently accurate and comprehensive alarm information. The result is a significant improvement in fault detection accuracy and a reduction in false alarm rate.
[0067] Optionally, in the fault detection method for a business system provided in this application embodiment, the current operating mode of the business system is detected based on the degree of matching between log entity features and fault log features in the fault log feature library. This includes: calculating the feature similarity between the log entity features and each fault log feature in the fault log feature library to obtain corresponding fault log features and feature similarities; selecting reference fault log features whose corresponding feature similarity is greater than or equal to a similarity threshold from the corresponding fault log features and feature similarities; and determining the reference fault operating mode corresponding to the reference fault log features as the current operating mode of the business system.
[0068] Optionally, in this embodiment, feature similarity may be, but is not limited to, an indicator used to quantify the similarity between log entity features and individual fault log features in the fault log feature library, such as normalized values like 0.92 obtained through vector cosine similarity, 0.88 obtained through Pearson correlation coefficient, and 0.95 obtained through Jaccard similarity coefficient. Feature similarity may be calculated using, but is not limited to, vector cosine similarity or Jaccard similarity coefficient.
[0069] Optionally, in this embodiment, the similarity threshold may be, but is not limited to, a pre-set threshold (such as 0.9 or 0.95) used to determine whether a strong match exists. This threshold is configurable and is used to balance the sensitivity (recall, i.e., no false negatives) and accuracy (precision, i.e., no false positives) of detection. Setting a high threshold (such as 0.95) will reduce false positives, but may miss some atypical faults; setting a low threshold (such as 0.8) can detect more problems, but false positives will increase.
[0070] Optionally, in this embodiment, the reference fault log feature may refer to one or more fault log features whose corresponding feature similarity meets a similarity threshold (greater than or equal to). The reference fault operation mode is the fault mode represented by these reference fault log features. In some scenarios, multiple reference fault log features may be selected. In this case, the mode corresponding to the feature with the highest feature similarity can be determined as the current operation mode, or all matching modes can be used as candidates for the current operation mode. This solution does not limit this.
[0071] Optionally, in this embodiment, reference fault log features can be filtered using a fixed threshold, but not limited to this. For example, a preset similarity threshold of 0.9 can be used to iterate through all "fault log feature - feature similarity" correspondences, and fault log features with a feature similarity ≥ 0.9 can be selected and identified as reference fault log features. Alternatively, reference fault log features can be filtered using a dynamic threshold, but not limited to this. For example, the threshold can be dynamically adjusted based on the current system load (the threshold is lowered to 0.85 under high load to reduce false negatives, and raised to 0.95 under low load to reduce false positives). If the current load is high, fault log features with a feature similarity ≥ 0.85 can be selected as reference fault log features.
[0072] Optionally, in this embodiment, the current operating mode of the business system can be determined by prioritizing the highest similarity. For example, the selected reference fault log features include "database connection pool exhaustion" (similarity 0.94) and "database lock conflict" (similarity 0.91). The reference fault operating mode corresponding to "database connection pool exhaustion" with the highest feature similarity is selected as the current operating mode of the business system. Alternatively, the current operating mode of the business system can be determined by merging multiple modes. For example, if the reference fault operating modes corresponding to the selected reference fault log features are related (such as a causal relationship between "database connection pool exhaustion" and "application process abnormality"), then multiple modes are merged into "application process abnormality caused by database connection pool exhaustion" as the current operating mode.
[0073] Based on the above, by calculating the feature similarity between log entity features and fault log features, and filtering out reference fault log features that are greater than or equal to the similarity threshold, the current operating mode is determined. This makes the fault decision-making mechanism clear, quantifiable, and configurable, thereby ensuring that an alarm is triggered only when the real-time log features are highly similar to known fault features. This avoids the shortcomings of related technologies that rely on fuzzy matching and subjective judgment, resulting in inaccurate alarm information, and achieves the technical effect of making the detection results more reliable.
[0074] Optionally, in the fault detection method for a business system provided in this application embodiment, detecting the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library includes: obtaining a log analysis model, wherein the log analysis model is trained using log samples labeled with operating mode tags; inputting the log data included in the log data set into the log analysis model to obtain the operating mode output by the log analysis model, wherein the log analysis model includes an input layer, a feature extraction layer, a classification layer and an output layer, the input layer is used to transmit log data to the feature extraction layer, the feature extraction layer is used to extract features from the log data to obtain log entity features, and transmit the log entity features to the classification layer, the classification layer is used to classify the log entity features according to the fault log features in the fault log feature library to obtain the current operating mode of the business system, and outputs the operating mode through the output layer.
[0075] Optionally, in the embodiments of this application, the log analysis model may refer to, but is not limited to, a machine learning or deep learning model trained with labeled samples for detecting the operating mode of a business system, such as a classification model built based on convolutional neural networks (CNN), LSTM, or Transformer (such as BERT).
[0076] Optionally, in this embodiment, the log samples labeled with the running mode may refer to, but are not limited to, historical log data with explicit running mode labels, which are the basic data for model training, such as log sequences labeled as "normal mode", log sets labeled as "database connection pool exhaustion mode", and log fragment groups labeled as "payment interface timeout mode".
[0077] Optionally, in this embodiment, the training process of the log analysis model may include, but is not limited to, the following: Sample preparation stage: Collect historical log data, manually or semi-automatically label the operating mode (e.g., "normal mode" "database connection pool exhausted mode"), and construct a training sample set. Model construction stage: Build a model architecture including an input layer, a feature extraction layer, a classification layer, and an output layer. The input layer uses a word embedding module, the feature extraction layer uses a Transformer Encoder, the classification layer uses a fully connected network, and the output layer uses a Softmax function. Training and optimization stage: Train the model using labeled samples, optimize parameters through backpropagation, and enable the model to learn the mapping relationship between log entity features and operating modes (where the classification layer weights implicitly fuse fault log feature library information); during training, use reinforcement learning algorithms (e.g., group relative policy optimization) to fine-tune the model and improve the ability to identify complex relationships. Online update stage: After the model is deployed, receive feedback from maintenance personnel on the prediction results (e.g., "false positives" "false negatives"), incorporate the corrected samples into the training set, perform incremental training periodically, and update the model parameters to adapt to new fault modes.
[0078] Optionally, in this embodiment, the input layer may be, but is not limited to, the module in the log analysis model responsible for receiving and preprocessing the raw log data, providing standardized input for subsequent feature extraction, such as converting log text into word vectors, converting structured log fields into numerical vectors, and converting log sequences into token ID sequences.
[0079] Optionally, in this embodiment, the feature extraction layer may be, but is not limited to, the module in the log analysis model responsible for extracting key features from the preprocessed log data, capable of generating log entity features that represent the core information of the log, such as an LSTM-based temporal feature extraction layer (outputting a vector containing log temporal relationships) or a Transformer-based self-attention layer (outputting a vector containing log context associations), etc.
[0080] Optionally, in this embodiment, the classification layer may be, but is not limited to, the module in the log analysis model responsible for pattern classification based on log entity features and fault log feature library, capable of outputting the operating mode of the business system, such as a neural network structure containing fully connected layers (implicitly storing fault log feature library information through weight parameters), a classifier based on support vector machines, etc.
[0081] Optionally, in this embodiment of the application, the log analysis model can be obtained by loading an offline pre-trained model or by deploying an online training model, and the log data set can be input into the log analysis model.
[0082] Optionally, in this embodiment, the operating mode can be obtained through end-to-end batch inference, but is not limited to. For example, log data sets are input into the model input layer in batches, and the input layer converts the log text into word vectors after word segmentation; the feature extraction layer (Transformer's Encoder) extracts log entity features (including the association between "Connection Timeout" and "TNS-12519") through a self-attention mechanism; the classification layer classifies the features by combining fault log feature library information (implicitly stored in weight parameters); the output layer outputs a probability distribution ("database connection pool exhaustion mode" probability 0.96) to obtain the operating mode. The operating mode can also be obtained through real-time streaming data inference, but is not limited to. For example, log data is input into the model input layer in real-time in streaming form, and the input layer processes the logs line by line and generates real-time feature vectors; the feature extraction layer dynamically updates the time-series features to capture real-time changes in the log sequence; the classification layer performs classification every 30 seconds based on the latest features; the output layer outputs the current operating mode in real-time (e.g., switching from "normal mode" to "database connection pool exhaustion mode").
[0083] Optionally, in this embodiment, to improve fault response efficiency and model accuracy, the method may further include setting up a multi-layered feedback mechanism and supporting interactive operations through a graphical user interface. For example, after receiving target location information and target repair information, maintenance personnel can query detailed alarm information, log context, and solution documents through the graphical user interface; at the same time, maintenance personnel can provide feedback on the detection results (such as clicking "Confirm Fault" to mark an actual fault, and clicking "False Alarm" to mark a non-fault situation). This feedback information will trigger online incremental updates of the log analysis model (incorporating new fault logs into the training data) or parameter tuning (adjusting model feature weights), thereby responding to abnormal system behavior more quickly and accurately through the multi-layered feedback mechanism.
[0084] Based on the above, a log analysis model was trained using log samples labeled with operating mode tags. End-to-end detection was achieved through the model's feature extraction layer and classification layer. This enabled the model to automatically learn and extract complex, abstract, and non-linear log patterns from massive logs, thereby capturing fault associations that are difficult to define manually. This avoids the shortcomings of related technologies that rely on manually defined rules and cannot accurately identify complex or unknown faults. The result is a technical effect of improving the intelligence level of fault detection and the ability to generalize to unknown faults.
[0085] Optionally, in the fault detection method for a business system provided in this application embodiment, generating target location information and target repair information based on the target fault operation mode includes: finding a target log segment from the log data set whose correlation with the target fault operation mode is greater than a preset correlation, wherein the target log segment records the operation status of the corresponding functional subsystem when the business system is in the target fault operation mode; and generating target location information and target repair information based on the target log segment.
[0086] Optionally, in this embodiment, the preset correlation degree can be, but is not limited to, a quantitative threshold value used to determine the correlation between log fragments and target fault operation modes. It can be configured according to the fault detection accuracy requirements, such as 0.8 set by attention weight, 0.75 set by mutual information calculation, or 0.85 set by feature importance score.
[0087] Optionally, in this embodiment, the target log may be, but is not limited to, log records in the log data set that are associated with the target fault operation mode, such as the "Accounting Core Subsystem Connection Timeout Log" and the "Database Cluster Subsystem TNS Error Log" corresponding to the "Database Connection Pool Exhaustion" fault.
[0088] Optionally, in the embodiments of this application, the target log fragment may be, but is not limited to, key information units extracted from the target log that record the fault operation status of the functional subsystem, such as error code ("TNS-12519"), connection pool name ("main-pool"), transaction ID ("TID:abc_123"), exception description ("Failed to getconnection"), etc. in the log.
[0089] Optionally, in the embodiments of this application, the operating status of the functional subsystem may be, but is not limited to, the functional subsystem status information related to the fault carried by the target log fragment, such as the subsystem's resource usage (connection pool full), error type (timeout error), business interaction status (transaction link interruption), etc.
[0090] Optionally, in this embodiment, target log segments can be found using, but is not limited to, association rule matching. For example, historical fault data can be analyzed in advance using the Apriori algorithm to establish association rules for "target fault operation mode - target log segment" (e.g., "database connection pool exhaustion mode" associated with "Connection Timeout error code" and "connection pool name field"). When a target fault operation mode is detected, target logs containing the associated fields are matched in the log data set, and information such as error codes and connection pool names are extracted from the target logs as target log segments. If the degree of association between the segment and the fault mode (rule matching score) is ≥ a preset association degree of 0.8, it is determined to be a valid target log segment. Target log segments can also be found using, but is not limited to, model attribution analysis. For example, after detecting the target fault operation mode using a log analysis model with an attention mechanism, the contribution (degree of association) of each log segment to the fault judgment is calculated; log segments with a contribution degree ≥ a preset association degree of 0.85 are selected as target log segments.
[0091] Optionally, in this embodiment, target location information and target repair information can be generated by filling in templates, but are not limited to. For example, a location information template ("Fault-related transaction ID: {TID}, Fault module: {subsystem name}, Core error: {error code}, Resource anomaly: {connection pool status}") and a repair information template ("1. Check the {connection pool name} configuration of {subsystem name}; 2. Restart {subsystem instance} to release the connection; 3. Verify the problem corresponding to {error code} (such as database listener status)") can be pre-built; the TID ("abc_123"), subsystem name ("Accounting core subsystem"), error code ("TNS-12519"), and connection pool name ("main-pool") can be extracted from the target log fragment and filled into the template to generate target location information and target repair information. Target location information and target repair information can also be dynamically generated through a generative model, but are not limited to. For example, the target log fragment is input into the T5 generation model, which has been fine-tuned with operational knowledge; based on the fault details in the fragment and combined with historical repair experience, the model dynamically generates target location information and target repair information in natural language format.
[0092] Based on the above, after detecting the target fault operation mode, the system searches for target log segments in the log data set that are more correlated with the mode than a preset correlation level. Based on these target log segments, it generates location and repair information, enabling the system to uncover the specific log segments that led to the fault diagnosis. This achieves an automated closed loop from abstract fault detection to specific fault location, avoiding the shortcomings of related technologies that can only achieve abnormal alarms and require manual investigation and analysis of fault root causes. It achieves the technical effect of generating specific, accurate, and highly operable location and repair information, and greatly shortening the fault analysis time.
[0093] Optionally, in the fault detection method for a business system provided in this application embodiment, finding a target log segment from the log data set whose correlation with the target fault operation mode is greater than a preset correlation degree includes: when the target fault operation mode is a data fault operation mode, finding a data log and a data log segment from the log data set, wherein the data fault operation mode is the operation mode in which the business system is when the database system in one or more functional subsystems fails, the database system is used to provide business data for the business system to run the target business, the data log is the log generated when the database system fails, and the data log segment records the database fault parameters of the database system; when the target fault operation mode is an application fault operation mode, finding an application log and an application log segment from the log data set, wherein the application fault operation mode is the operation mode in which the business system is when the application system in one or more functional subsystems fails, the business system runs the target business through the application system, the application log is the log generated when the application system fails, and the application log segment records the application fault parameters of the application system.
[0094] Optionally, in this embodiment, the data failure operation mode may refer to, but is not limited to, the business system operation state caused by a database system failure in the functional subsystem, such as "database tablespace full mode", "database deadlock mode", "Redis cache breakdown mode", "master-slave synchronization failure mode", etc.
[0095] Optionally, in this embodiment, the database system may be, but is not limited to, a functional subsystem that provides data storage, query, and transaction processing support for the business system to run target business operations, such as relational databases (Oracle, MySQL), NoSQL databases (MongoDB), distributed caches (Redis), etc.
[0096] Optionally, in this embodiment, the data log may be, but is not limited to, log records generated when the database system fails, such as Oracle alarm logs, MySQL error logs, Redis slow query logs, etc.
[0097] Optionally, in this embodiment, the data log fragment may be, but is not limited to, information units extracted from the data log that record key parameters of database failure, such as error code (“ORA-01653”), tablespace name (“DATA_TS”), deadlock transaction ID (“TX-0001-0002”), cache hit rate (“hit_rate=10%”), etc.
[0098] Optionally, in this embodiment, the application fault operation mode may refer to the business system operation state caused by the application system failure in the functional subsystem, with the root cause of the failure located in the business logic layer, such as "null pointer exception mode", "payment interface timeout mode", "authentication service unreachable mode", "503 service error mode", etc.
[0099] Optionally, in this embodiment, the application system may be, but is not limited to, a functional subsystem that carries business logic and directly runs the target business, such as order service, payment service, user authentication service, product query service, etc.
[0100] Optionally, in this embodiment, the application logs may be, but are not limited to, log records generated when the application system fails, reflecting the abnormal operation of the application, such as application error logs output by Log4j, transaction exception logs of the Spring framework, microservice call chain logs, etc.
[0101] Optionally, in this embodiment, the application log fragment may be, but is not limited to, information units extracted from the application log that record key parameters of the application failure, such as exception type ("NullPointerException"), error code ("503Service Unavailable"), exception location ("PaymentService.java:250"), timeout ("timeout=30s"), etc.
[0102] Optionally, in this embodiment, application logs and application log fragments can be found through log type classification and keyword filtering, but not limited to. For example, when a target failure mode (such as "database deadlock") is detected, the system first determines that the mode belongs to the "data failure mode". Then, the system searches only the subset of "data logs" (such as MySQL Error Log) in the "log data set" to find data log fragments containing "data failure parameters" (such as "deadlock"). Similarly, if the mode is "application failure mode" (such as "null pointer exception"), then the system searches only the subset of "application logs" for "application failure parameters" (such as "NullPointerException"). Application logs and application log fragments can also be found by targeted search based on log source, but not limited to. For example, the system maintains a mapping relationship, indicating which "functional subsystems" (such as the database system itself or database interface modules) are usually associated with the "data failure mode", and which "functional subsystems" (such as order service, payment service) are usually associated with the "application failure mode". Once the target fault operation mode is detected, the system uses this mapping relationship to prioritize or only search for logs generated by these specific functional subsystems from the log data set, and extracts "data log fragments" or "application log fragments" containing the corresponding fault parameters.
[0103] Based on the above, by searching for corresponding data log segments or application log segments in the log data set according to the type of the target fault operation mode, the system can pre-distinguish the type of fault mode, thereby significantly reducing the search space for the target log segment. This avoids the inefficiency of manual troubleshooting in massive logs, which is a drawback of related technologies. It achieves the technical effects of improving the efficiency and accuracy of fault location and reducing the consumption of computing resources.
[0104] Optionally, in the fault detection method of the business system provided in this application embodiment, generating target location information and target repair information according to the target fault operation mode includes: obtaining fault operation modes and location information with corresponding relationships; filtering out the target location information corresponding to the target fault operation mode from the fault operation modes and location information with corresponding relationships; obtaining location information and repair information with corresponding relationships; and filtering out the target repair information corresponding to the target location information from the location information and repair information with corresponding relationships.
[0105] Optionally, in this embodiment, the corresponding fault operation mode and location information may refer to, but are not limited to, a pre-built data set used to store the mapping relationship between fault modes and location information, such as a hash table (key is "database connection pool exhaustion mode", value is "connection pool configuration error and database listener failure"), a relational database table (fields include "fault_mode" and "location_info"), a JSON configuration file (database connection pool exhaustion mode and database listener failure), etc.
[0106] Optionally, in this embodiment, the location information and repair information with corresponding relationships may refer to, but are not limited to, a pre-built data set used to store the mapping relationship between location information and repair information, such as a dictionary (key is "connection pool configuration error", value is "restart instance + adjust connection pool parameters"), a database table (fields include "location_info" and "remedy_info"), an XML configuration file (with nested tags storing the mapping relationship), etc.
[0107] Optionally, in this embodiment, the target location information can be filtered out by searching a memory hash table, but is not limited to this method. For example, the system loads the corresponding fault operation modes and location information into a memory hash table, where the key is the fault operation mode (e.g., "database connection pool exhausted mode" or "payment interface timeout mode"), and the value is the corresponding location information (e.g., "accounting core subsystem connection pool full, database listener abnormal" or "payment gateway response timeout"). When the target fault operation mode is detected as "database connection pool exhausted mode", the system directly queries the hash table using this mode as the key to filter out the corresponding "accounting core subsystem connection pool full, database listener abnormal" as the target location information. Alternatively, the target location information can be filtered out by querying a database index, but is not limited to this method. For example, the correspondence between fault operation modes and location information is stored in the "fault_location" table of a MySQL database, and an index is created on the "fault_mode" field. When the target fault operation mode is detected, the SQL query "SELECT location_info FROM fault_location WHERE fault_mode = 'target fault operation mode'" is executed to quickly locate and filter the target location information using the index.
[0108] Optionally, in this embodiment, target repair information can be filtered using a dictionary-based query method, but is not limited to. For example, the system maintains a nested dictionary where the outer key is the fault operation mode, the inner key is the location information, and the value is the repair information. After obtaining the target location information "the core accounting subsystem connection pool is full, and the database listener is abnormal," the dictionary is queried using this location information as the inner key to filter out the corresponding "1. Restart the core accounting subsystem instance; 2. Check the database listener status; 3. Adjust the maximum capacity of the connection pool" as the target repair information. Target repair information can also be filtered using a relational database query method, but is not limited to. For example, the correspondence between location information and repair information is stored in the "location_remedy" table and associated with the "fault_location" table through the "location_info" field. After obtaining the target location information, the SQL query "SELECT remedy_info FROM location_remedy WHERE location_info = 'target location information'" is executed to filter out the corresponding target repair information.
[0109] The above content achieves the following technical effects: by obtaining the correspondence between fault operation mode and location information, as well as the correspondence between location information and repair information, target location information and target repair information can be filtered out. This enables the provision of pre-programmed standardized guidance for known faults with standard processing procedures, thus providing an extremely fast, stable, and predictable information generation mechanism. This avoids the shortcomings of related technologies that rely on manual analysis, resulting in low fault handling efficiency and inconsistent results. It also achieves the technical effects of reducing the computational complexity of real-time analysis and ensuring high consistency and reliability of solutions.
[0110] Optionally, in this embodiment of the application, to better illustrate the implementation of the above method, a "user order payment" business scenario of an e-commerce platform is used as an example to explain the implementation process of the above procedure in detail. Figure 4 This is a flowchart of business system fault detection and repair provided according to an embodiment of this application, such as... Figure 4As shown, the method includes: a user submits an order and initiates a payment operation on an e-commerce platform (user accesses / operates on the application system). The platform's payment service subsystem, accounting subsystem, and other functional modules automatically generate log information at code embedding points, including payment request time, transaction serial number, interface response status, database operation records, etc. (embedding point-generated log information). This log information is transmitted in real-time to a log analysis model, which matches and analyzes the log features against fault features such as "payment timeout" and "accounting processing anomaly" in the fault log feature library (log analysis model). If the log features do not trigger the preset alarm threshold (No), the model continuously maintains and optimizes its ability to analyze payment scenario logs based on the normal payment process data recorded at that embedding point. If the feature "payment interface response timeouts exceed 10 times within 5 minutes" is detected (Yes), an automatic alarm is triggered, and an alarm notification is sent to system maintenance personnel via email, clearly informing them of the problem and solution: "The payment subsystem's interaction with the third-party gateway timed out; it is recommended to temporarily switch to a backup payment gateway and check the network link." The operations and maintenance personnel log into the log analysis-based alarm system, query the detailed log context, related transaction records, and repair steps for the payment failure, perform the backup gateway switchover operation, and complete the problem resolution.
[0111] Optionally, in the fault detection method for a business system provided in this application embodiment, after sending the target location information and target repair information to the maintenance account of the business system, the method further includes: obtaining repair log data generated by the business system after executing the repair operation indicated by the target repair information; extracting result fragments from the repair log data, wherein the result fragments record the repair results; determining that the repair operation was successful when the repair result indicates that the business system switched from the target fault operation mode to the normal operation mode after executing the repair operation, and storing the log features of the corresponding log data set and the target fault operation mode in the fault log feature library; determining that the repair operation was unsuccessful when the repair result indicates that the business system did not switch from the target fault operation mode to the normal operation mode after executing the repair operation, updating the target location information and target repair information according to the repair log data to obtain the target location information and target repair information; and sending the target location information and target repair information to the maintenance account of the business system.
[0112] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0113] Example 2
[0114] This application also provides a fault detection device for a business system. It should be noted that the fault detection device for a business system in this application can be used to execute the fault detection method for a business system provided in this application. The following describes the fault detection device for a business system provided in this application.
[0115] According to an embodiment of this application, an apparatus for implementing the fault detection method of the above-described business system is also provided. Figure 5 This is a schematic diagram of a fault detection device for a business system provided according to an embodiment of this application, such as... Figure 5 As shown, the device includes:
[0116] The acquisition module 502 is used to acquire a set of log data generated by one or more functional subsystems included in the business system during the operation of the target business in the business system. The business system executes corresponding functions through one or more functional subsystems to run the target business, and the log data set includes log data generated by each functional subsystem.
[0117] The detection module 504 is used to detect the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library. The fault log features are the log features of the log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode.
[0118] The generation module 506 is used to generate target location information and target repair information according to the target fault operation mode when the operation mode is detected to be the target fault operation mode corresponding to the target fault log feature. The target location information is used to locate the reason why the business system is in the target fault operation mode, and the target repair information is used to indicate the repair operation required to switch the business system from the target fault operation mode to the normal operation mode.
[0119] The sending module 508 is used to send target location information and target repair information to the maintenance account of the business system.
[0120] The fault detection device for a business system provided in this application acquires a set of log data generated by one or more functional subsystems included in the business system, and detects the current operating mode of the business system based on the matching relationship between the log characteristics of the log data set and the fault log characteristics in the fault log characteristic library. Then, when the operating mode is detected to be a target fault operating mode corresponding to the target fault log characteristics, target location information and target repair information are generated according to the target fault operating mode, and the target location information and target repair information are sent to the maintenance account of the business system. This achieves the purpose of automatically locating the cause of the business system being in the target fault operating mode and instructing the repair operation required to switch from the target fault operating mode to the normal operating mode. This achieves the technical effect of improving the accuracy of fault location and the efficiency of fault handling, and solves the technical problem of low fault detection efficiency caused by the existing technology that can only alarm but cannot locate the root cause.
[0121] Optionally, in the fault detection device for a business system provided in this application embodiment, the detection module includes: a feature extraction unit, used to extract features from the log data included in the log data set to obtain log entity features, wherein the log entity features are used to characterize multiple log segments included in the log data and the correlation between the multiple log segments, and the log features include log entity features; and a detection unit, used to detect the current operating mode of the business system based on the degree of matching between the log entity features and the fault log features in the fault log feature library.
[0122] Optionally, in the fault detection device for the business system provided in this application embodiment, the detection unit is further configured to: calculate the feature similarity between the log entity feature and each fault log feature in the multiple fault log features in the fault log feature library, to obtain corresponding fault log features and feature similarities; filter out reference fault log features whose corresponding feature similarity is greater than or equal to a similarity threshold from the corresponding fault log features and feature similarities; and determine the reference fault operation mode corresponding to the reference fault log feature as the current operation mode of the business system.
[0123] Optionally, in the fault detection device for a business system provided in this application embodiment, the detection module includes: a first acquisition unit, used to acquire a log analysis model, wherein the log analysis model is trained using log samples labeled with operating mode tags; an input unit, used to input log data included in the log data set into the log analysis model to obtain the operating mode output by the log analysis model, wherein the log analysis model includes an input layer, a feature extraction layer, a classification layer and an output layer, the input layer is used to transmit log data to the feature extraction layer, the feature extraction layer is used to extract features from the log data to obtain log entity features, and transmit the log entity features to the classification layer, the classification layer is used to classify the log entity features according to the fault log features in the fault log feature library to obtain the current operating mode of the business system, and output the operating mode through the output layer.
[0124] Optionally, in the fault detection device for a business system provided in this application embodiment, the generation module includes: a search unit, used to search for a target log segment from the log data set that has a correlation degree greater than a preset correlation degree with the target fault operation mode, wherein the target log segment records the operation status of the functional subsystem corresponding to the business system when it is in the target fault operation mode; and a generation unit, used to generate target location information and target repair information based on the target log segment.
[0125] Optionally, in the fault detection device for the business system provided in this application embodiment, the searching unit is further configured to: when the target fault operation mode is a data fault operation mode, search for data logs and data log fragments from the log data set, wherein the data fault operation mode is the operation mode in which the business system is when the database system in one or more functional subsystems fails, the database system is used to provide business data for the business system to run the target business, the data log is the log generated when the database system fails, and the data log fragment records the database fault parameters of the database system; when the target fault operation mode is an application fault operation mode, search for application logs and application log fragments from the log data set, wherein the application fault operation mode is the operation mode in which the business system is when the application system in one or more functional subsystems fails, the business system runs the target business through the application system, the application log is the log generated when the application system fails, and the application log fragment records the application fault parameters of the application system.
[0126] Optionally, in the fault detection device for the business system provided in this application embodiment, the generation module further includes: a second acquisition unit, used to acquire fault operation modes and location information with corresponding relationships; a first filtering unit, used to filter out target location information corresponding to the target fault operation mode from the fault operation modes and location information with corresponding relationships; a third acquisition unit, used to acquire location information and repair information with corresponding relationships; and a second filtering unit, used to filter out target repair information corresponding to the target location information from the location information and repair information with corresponding relationships.
[0127] It should be noted that the acquisition module 502, detection module 504, generation module 506, and sending module 508 mentioned above correspond to steps S101 to S104 in Embodiment 1. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. The above modules or units may be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules may also be part of a device and run in the computer terminal 10 provided in Embodiment 1.
[0128] Example 3
[0129] Embodiments of this application may provide an electronic device. Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) Processor 602, memory 604, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0130] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0131] The processor can invoke information and applications stored in memory via a transmission device to perform the following steps: During the operation of the target business function in the business system, acquire a set of log data generated by one or more functional subsystems included in the business system, wherein the business system executes corresponding functions through one or more functional subsystems to run the target business, and the log data set includes log data generated by each functional subsystem; detect the current operating mode of the business system based on the matching relationship between the log characteristics of the log data set and the fault log characteristics in the fault log characteristic library, wherein the fault log characteristics are the log characteristics possessed by the log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode; if the detected operating mode belongs to the target fault operating mode corresponding to the target fault log characteristics, generate target location information and target repair information based on the target fault operating mode, wherein the target location information is used to locate the reason why the business system is in the target fault operating mode, and the target repair information is used to indicate the repair operations required to switch the business system from the target fault operating mode to the normal operating mode; send the target location information and target repair information to the maintenance account of the business system.
[0132] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: extracting features from the log data included in the log data set to obtain log entity features, wherein the log entity features are used to characterize the multiple log segments included in the log data and the correlation between the multiple log segments, and the log features include log entity features; detecting the current operating mode of the business system based on the degree of matching between the log entity features and the fault log features in the fault log feature library.
[0133] The processor can also invoke information and applications stored in the memory via the transmission device to perform the following steps: calculate the feature similarity between the log entity feature and each fault log feature in the fault log feature library to obtain corresponding fault log features and feature similarities; select reference fault log features whose corresponding feature similarity is greater than or equal to the similarity threshold from the corresponding fault log features and feature similarities; and determine the reference fault operation mode corresponding to the reference fault log feature as the current operation mode of the business system.
[0134] The processor can also invoke information and applications stored in the memory via the transmission device to perform the following steps: obtaining a log analysis model, wherein the log analysis model is trained using log samples labeled with operating mode tags; inputting the log data included in the log dataset into the log analysis model to obtain the operating mode output by the log analysis model, wherein the log analysis model includes an input layer, a feature extraction layer, a classification layer, and an output layer, the input layer is used to transmit log data to the feature extraction layer, the feature extraction layer is used to extract features from the log data to obtain log entity features, and transmit the log entity features to the classification layer, the classification layer is used to classify the log entity features according to the fault log features in the fault log feature library to obtain the current operating mode of the business system, and output the operating mode through the output layer.
[0135] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: find the target log segment from the log data set that has a correlation degree greater than the target fault operation mode, wherein the target log segment records the operation status of the corresponding functional subsystem when the business system is in the target fault operation mode; generate target location information and target repair information based on the target log segment.
[0136] The processor can also invoke information and applications stored in memory via a transmission device to perform the following steps: When the target fault operating mode is a data fault operating mode, retrieve data logs and data log fragments from the log data set. The data fault operating mode is the operating mode of the business system when the database system in one or more functional subsystems fails. The database system provides business data for the business system to run its target business. The data logs are logs generated when the database system fails, and the data log fragments record the database system's database fault parameters. When the target fault operating mode is an application fault operating mode, retrieve application logs and application log fragments from the log data set. The application fault operating mode is the operating mode of the business system when the application system in one or more functional subsystems fails. The business system runs its target business through the application system. The application logs are logs generated when the application system fails, and the application log fragments record the application system's application fault parameters.
[0137] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain the corresponding fault operation mode and location information; filter out the target location information corresponding to the target fault operation mode from the corresponding fault operation mode and location information; obtain the corresponding location information and repair information; filter out the target repair information corresponding to the target location information from the corresponding location information and repair information.
[0138] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic device. For example, the electronic device may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.
[0139] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0140] Example 4
[0141] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the fault detection method of the business system provided in Embodiment 1.
[0142] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0143] This application also provides a computer program product, which, when executed on a data processing device, is suitable for performing the fault detection method steps of the above-described business system.
[0144] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above embodiments of this application, the descriptions of each embodiment have their own emphasis, and for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0145] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0147] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0148] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory, random access memory, portable hard drives, magnetic disks, or optical disks.
[0149] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A fault detection method for a business system, characterized in that, include: During the operation of the target business in the business system, a set of log data generated by one or more functional subsystems included in the business system is obtained. The business system executes corresponding functions through one or more of the functional subsystems to run the target business, and the set of log data includes log data generated by each of the functional subsystems. The current operating mode of the business system is detected based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library. The fault log features are the log features of log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode. If the operating mode is detected to be a target fault operating mode corresponding to the target fault log feature, target location information and target repair information are generated according to the target fault operating mode. The target location information is used to locate the reason why the business system is in the target fault operating mode, and the target repair information is used to indicate the repair operation required to switch the business system from the target fault operating mode to the normal operating mode. Send the target location information and target repair information to the maintenance account of the business system.
2. The method according to claim 1, characterized in that, The step of detecting the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library includes: Feature extraction is performed on the log data included in the log data set to obtain log entity features, wherein the log entity features are used to characterize multiple log segments included in the log data and the correlation between the multiple log segments, and the log features include the log entity features; The current operating mode of the business system is detected based on the degree of matching between the log entity characteristics and the fault log characteristics in the fault log characteristic library.
3. The method according to claim 2, characterized in that, The step of detecting the current operating mode of the business system based on the degree of matching between the log entity characteristics and the fault log characteristics in the fault log characteristic library includes: Calculate the feature similarity between the log entity feature and each of the multiple fault log features in the fault log feature library to obtain the fault log features and the feature similarity with a corresponding relationship; From the corresponding fault log features and the feature similarity, filter out the corresponding reference fault log features whose feature similarity is greater than or equal to the similarity threshold; The reference fault operation mode corresponding to the reference fault log feature is determined as the current operation mode of the business system.
4. The method according to claim 1, characterized in that, The step of detecting the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library includes: Obtain a log analysis model, wherein the log analysis model is obtained by training log samples labeled with running mode tags; The log data included in the log data set is input into the log analysis model to obtain the operating mode output by the log analysis model. The log analysis model includes an input layer, a feature extraction layer, a classification layer, and an output layer. The input layer is used to transmit the log data to the feature extraction layer. The feature extraction layer is used to extract features from the log data to obtain log entity features, and then transmits the log entity features to the classification layer. The classification layer is used to classify the log entity features according to the fault log features in the fault log feature library to obtain the current operating mode of the business system, and then outputs the operating mode through the output layer.
5. The method according to claim 1, characterized in that, The step of generating target location information and target repair information based on the target fault operation mode includes: Find the target log segment from the log data set that has a correlation greater than a preset correlation with the target fault operation mode. The target log segment records the operation status of the functional subsystem corresponding to the business system when the business system is in the target fault operation mode. The target location information and the target repair information are generated based on the target log fragment.
6. The method according to claim 5, characterized in that, The step of finding the target log segment from the log data set whose correlation with the target fault operation mode is greater than a preset correlation degree includes: When the target fault operation mode is a data fault operation mode, data logs and data log fragments are retrieved from the log data set. The data fault operation mode is the operation mode of the business system when the database system in one or more of the functional subsystems fails. The database system is used to provide business data for the business system to run the target business. The data log is the log generated when the database system fails. The data log fragment records the database fault parameters of the database system. When the target fault operation mode is the application fault operation mode, the application log and application log fragment are retrieved from the log data set. The application fault operation mode is the operation mode of the business system when the application system in one or more of the functional subsystems fails. The business system runs the target business through the application system. The application log is the log generated when the application system fails. The application log fragment records the application fault parameters of the application system.
7. The method according to claim 1, characterized in that, The step of generating target location information and target repair information based on the target fault operation mode includes: Obtain corresponding fault operation modes and location information; Filter out the target location information corresponding to the target fault operation mode from the corresponding fault operation modes and location information; Obtain corresponding location and repair information; The target repair information corresponding to the target location information is selected from the corresponding location information and repair information.
8. A fault detection device for a business system, characterized in that, include: The acquisition module is used to acquire a set of log data generated by one or more functional subsystems included in the business system during the operation of the target business in the business system. The business system executes corresponding functions through one or more of the functional subsystems to run the target business, and the log data set includes log data generated by each of the functional subsystems. The detection module is used to detect the current operating mode of the business system based on the matching relationship between the log features of the log data set and the fault log features in the fault log feature library. The fault log features are the log features of log data generated by one or more functional subsystems when the business system is in the corresponding fault operating mode. The generation module is used to generate target location information and target repair information according to the target fault operation mode when the operation mode is detected to belong to the target fault operation mode corresponding to the target fault log feature. The target location information is used to locate the reason why the business system is in the target fault operation mode, and the target repair information is used to indicate the repair operation to be performed to switch the business system from the target fault operation mode to the normal operation mode. The sending module is used to send the target location information and target repair information to the maintenance account of the business system.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the steps of the fault detection method of the business system according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the steps of the fault detection method for the business system according to any one of claims 1 to 7.
11. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the fault detection method for the business system according to any one of claims 1 to 7.