Operation and maintenance management method and device, electronic equipment and storage medium
By implementing end-to-end anomaly monitoring and correlation analysis, the system addresses the challenges of identifying business-level anomalies and the low platform independence within the operations and maintenance system, thereby improving operational efficiency.
Patent Information
- Application Number
- CN202511194183.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-19
AI Technical Summary
The existing operation and maintenance system is unable to proactively identify anomalies at the business level, and the system status monitoring and work order management platform are independent, resulting in low operation and maintenance efficiency.
By acquiring operational data from business systems and maintenance work orders from user feedback, we can conduct end-to-end anomaly monitoring, correlate and analyze anomaly monitoring results with work orders, and achieve linkage between monitoring data and work orders to improve operational efficiency.
It enables proactive identification and rapid repair of business-level anomalies, improving operational efficiency by 30% to 50%.
Smart Images

Figure CN121166418A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and particularly relates to an operation and maintenance management method and device, electronic equipment and a storage medium. BACKGROUND
[0002] The current operation and maintenance system can monitor system-level exceptions (such as server failure and interface timeout) at the technical level, but it is difficult to actively identify business-level exceptions (such as functional logic errors). These business exceptions usually need to be known by the technical party after user feedback. The specific process is as follows: after the user finds the problem, the user needs to submit a work order to the operation and maintenance on-duty personnel; if the operation and maintenance on-duty personnel cannot solve the problem, the operation and maintenance on-duty personnel manually assigns the problem to an internal technical personnel, and the multi-link forwarding is easy to cause deviation in the transmission of information, thereby affecting the problem processing efficiency.
[0003] In addition, the centralized management of the current operation and maintenance work order is responsible for a dedicated work order management platform, and the centralized monitoring of the system state is responsible for another dedicated monitoring platform, the two platforms are independent of each other, and the data correlation between the two platforms is low. In the operation and maintenance management process, the technical personnel needs to separately manage the work order of the work order management platform and the monitoring data of the monitoring platform, which further reduces the overall operation and maintenance efficiency. SUMMARY
[0004] The main purpose of the present application is to provide an operation and maintenance management method, device, electronic equipment and storage medium, which aims to improve the overall operation and maintenance efficiency of the system.
[0005] To achieve the above purpose, the present application provides an operation and maintenance management method, comprising:
[0006] obtaining to-be-processed information, wherein the to-be-processed information comprises running data of each business system and an operation and maintenance work order fed back by a user;
[0007] performing full-link exception monitoring on the running data to obtain an exception monitoring result;
[0008] if there is an operation and maintenance work order associated with an exception event in the exception monitoring result, performing correlation analysis on the exception monitoring result and the operation and maintenance work order associated with the exception event to obtain an exception analysis result;
[0009] repairing the exception event according to the exception analysis result.
[0010] In an embodiment, the running data comprises transaction data; and the performing full-link exception monitoring on the running data to obtain an exception monitoring result comprises:
[0011] performing analysis on the transaction data to obtain a business process state of each transaction;
[0012] abnormal monitoring is performed on the business process state of each transaction to obtain the abnormal monitoring result; and / or,
[0013] the transaction data is analyzed to obtain the calling result of the calling interface associated with each transaction;
[0014] abnormal monitoring is performed on the calling result of the calling interface associated with each transaction to obtain the abnormal monitoring result.
[0015] In an embodiment, the running data includes system log information and business log information, and the full-link abnormal monitoring on the running data to obtain an abnormal monitoring result includes:
[0016] the system log information and the business log information are preprocessed;
[0017] the preprocessed system log information and the preprocessed business log information are associatedly analyzed to obtain an associated analysis result;
[0018] the associated analysis result is used to determine the abnormal monitoring result.
[0019] In an embodiment, the business log information is obtained by:
[0020] log collection rules associated with each business system are determined, wherein the log collection rules include at least one of collection rules corresponding to a system granularity, a function module granularity, and a method function granularity;
[0021] the business log information of each business system is collected according to at least one of the collection rules corresponding to the system granularity, the function module granularity, and the method function granularity.
[0022] In an embodiment, if there is no associated operation and maintenance work order, the operation and maintenance work order is repaired by:
[0023] the operation and maintenance work order is checked;
[0024] if the checking is successful, the operation and maintenance work order is distributed to a corresponding technical processing personnel for repair processing according to a preset task approval process.
[0025] In an embodiment, if there is no associated operation and maintenance work order, the abnormal monitoring result is repaired by:
[0026] the abnormal event in the abnormal monitoring result is analyzed according to the running data to obtain an abnormal root cause analysis result;
[0027] the abnormal event is repaired according to the abnormal root cause analysis result.
[0028] In an embodiment, after the repairing processing of the abnormal event according to the abnormal analysis result, the method further comprises:
[0029] determining a repairing scheme associated with the abnormal event in the operation and maintenance work order and the abnormal event in the abnormal monitoring result;
[0030] storing the abnormal event and the repairing scheme associated with the abnormal event into a preset knowledge base;
[0031] if the abnormal event is monitored again, matching the monitored abnormal event with the preset knowledge base;
[0032] if the matching is successful, performing repairing processing by using the matched repairing scheme associated with the abnormal event.
[0033] In addition, to achieve the above-mentioned purpose, the present application further provides an operation and maintenance management device, which comprises:
[0034] an acquisition module, configured to acquire to-be-processed information, wherein the to-be-processed information comprises running data of each business system and operation and maintenance work orders fed back by users;
[0035] a monitoring module, configured to perform full-link abnormal monitoring on the running data to obtain an abnormal monitoring result;
[0036] an analysis module, configured to, if there is an operation and maintenance work order associated with an abnormal event in the abnormal monitoring result, perform correlation analysis on the abnormal monitoring result and the operation and maintenance work order associated with the abnormal event to obtain an abnormal analysis result;
[0037] a repairing module, configured to perform repairing processing on the abnormal event according to the abnormal analysis result.
[0038] In addition, to achieve the above-mentioned purpose, the present application further provides an electronic device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the operation and maintenance management method as described above.
[0039] In addition, to achieve the above-mentioned purpose, the present application further provides a storage medium, which is a computer readable storage medium, and a computer program is stored on the storage medium, wherein the computer program is executed by a processor to implement the steps of the operation and maintenance management method as described above.
[0040] In addition, to achieve the above-mentioned purpose, the present application further provides a computer program product, which comprises a computer program, wherein the computer program is executed by a processor to implement the steps of the operation and maintenance management method as described above.
[0041] The application provides an operation and maintenance management method and device, electronic equipment and a storage medium. The operation and maintenance management method comprises the following steps: obtaining to-be-processed information, wherein the to-be-processed information comprises operation data of each business system and user feedback operation and maintenance work orders; performing full-link abnormality monitoring on the operation data to obtain an abnormality monitoring result; if there is an operation and maintenance work order associated with an abnormal event in the abnormality monitoring result, performing correlation analysis on the abnormality monitoring result and the operation and maintenance work order associated with the abnormal event to obtain an abnormality analysis result; and performing repair processing on the abnormal event according to the abnormality analysis result. The application performs full-link abnormality monitoring on the operation data of each business system, and further, when there is an operation and maintenance work order associated with an abnormal event in the abnormality monitoring result, the abnormality monitoring result and the user feedback operation and maintenance work order are associated and repaired, the linkage between monitoring data and work order information is realized, and the efficiency of overall operation and maintenance is effectively improved. BRIEF DESCRIPTION OF DRAWINGS
[0042] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, serve to explain the principles of the application.
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced here. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0044] Figure 1 A flowchart provided by the operation and maintenance management method embodiment one of the application;
[0045] Figure 2 A flowchart provided by the operation and maintenance management method embodiment two of the application;
[0046] Figure 3 A flowchart provided by the operation and maintenance management method embodiment three of the application;
[0047] Figure 4 A flowchart provided by the operation and maintenance management method embodiment four of the application;
[0048] Figure 5 A task approval flowchart provided by an embodiment of the application;
[0049] Figure 6 A flowchart provided by the operation and maintenance management method embodiment six of the application;
[0050] Figure 7 A module structure diagram of the operation and maintenance management device of the embodiment of the application;
[0051] Figure 8 A device structure schematic diagram of a hardware running environment involved in the operation and maintenance management method in the embodiments of the present application.
[0052] The object implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0053] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.
[0054] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings and specific embodiments of the specification.
[0055] It should be noted that the execution subject of the present embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a big data service platform, an operation and maintenance management system, etc. capable of realizing the above functions. The present embodiment and the following embodiments will be described below taking an operation and maintenance management platform as an example.
[0056] The current operation and maintenance system can easily monitor system-level exceptions (such as server failure, interface timeout) through technical means at the technical level, but it is difficult to actively identify business-level exceptions (such as transaction process jam, function logic error). These business exceptions usually need to be fed back by users before the technical party can know. The specific process is as follows: after the user finds the problem, the user needs to submit a work order to the operation and maintenance on-duty personnel; if the operation and maintenance on-duty personnel cannot solve the problem, the operation and maintenance on-duty personnel manually assigns the problem to an internal technical personnel, and the multi-link forwarding is easy to cause deviation in the transmission of information, affecting the efficiency of problem handling. In addition, the centralized management of the current operation and maintenance work order (handled by a dedicated work order management platform) and the centralized monitoring of the system state (handled by another dedicated monitoring platform) are independent of each other, the data correlation between the two core platforms is low, and the technical personnel needs to separately manage the work order of the work order management platform and the monitoring data of the monitoring platform for operation and maintenance, further reducing the overall operation and maintenance efficiency.
[0057] Based on the above problems, the present embodiment provides an operation and maintenance management method, which will be described below with reference to Figure 1 , Figure 1 A flowchart provided by the operation and maintenance management method embodiment one of the present application. In the present embodiment, the operation and maintenance management method comprises the following steps:
[0058] Step S11, obtaining to-be-processed information, wherein the to-be-processed information comprises running data of each business system and operation and maintenance work orders fed back by users;
[0059] It should be noted that the operation data includes transaction data, system log information and business log information, the business log information refers to the log recorded by the system when performing a specific business process, which is directly related to business operation, business data, business state, and the system log information is the log recording the system running state, technical level operation, bottom layer resource and component behavior.
[0060] It should be noted that the transaction data includes user information, transaction counterpart information, transaction amount, transaction time, transaction state, transaction belonging business and the like. The system log information includes log generation time, log level, log source information (function module name and code position lamp), log operation information and the like. The business log information includes business process link, execution state, business data content (specific data related to business), business start time, final result of business execution and the like.
[0061] It should be noted that the operation and maintenance work order includes work order number, work order creation time, work order creator, problem detailed description, related screenshot, processing state and the like. Alternatively, the user can log in the system to fill in the operation and maintenance work order when finding abnormality in use.
[0062] It should be noted that the monitored operation data can be visually displayed, and the large screen data display technology is bix platform in the industry platform, which integrates some visual charts of table, picture, line statistics, histogram statistics and the like to display data characteristics. The platform presents large screen data visualization effect through data interface API and chart dragging development, and the technical personnel can intuitively understand the operation of each business system according to the displayed data.
[0063] It should be noted that the operation and maintenance management platform of the embodiment retains the original system operation and maintenance platform and system monitoring platform, and adds a connection bridge between the system operation and maintenance platform and the system monitoring platform. When the system fails, the bridge gets abnormal data through API interface and KAFKA listening, and alarms to operation and maintenance personnel, technical person in charge and business personnel through the robot, so as to realize the integrated operation and maintenance system of system monitoring and work order management.
[0064] Step S12, performing full-link abnormality monitoring on the operation data to obtain an abnormality monitoring result;
[0065] It needs to be explained that the core transaction system is centered, and the business systems related to the upstream and downstream are linked from the transaction receiving, transaction sending, transaction processing to the transaction clearing nodes. The core transaction system monitors the transfer process status of each transaction. If a problem occurs in a node, it can be monitored and warned in real time through the whole link monitoring system. In an embodiment, the transaction data is analyzed to record the business process status of each transaction in real time, and the business process status of each transaction is monitored abnormally. In addition, the transaction data is analyzed to obtain the calling result of the API calling of each transaction, and the calling result of the calling interface associated with each transaction is monitored abnormally to record the health status of each transaction node. The abnormal monitoring result is obtained. Through the business health status and the flow health status of each transaction, the accuracy of abnormal monitoring is comprehensively improved.
[0066] In an embodiment, according to different application scenarios, a unified business log collection tool is set among system clusters, and then the business log information is collected by using the business log collection tool. Optionally, the collection rules of system level, class level, method level and the like granularity can be pre-configured, and further, the collection of key business data is realized in a low-intrusive manner, so as to obtain the business log information. Through asynchronous log information consumption, the collected system log information and business log information are associated and analyzed to identify and warn different abnormal scenarios, such as abnormal alarm, timeout warning and the like.
[0067] In addition, the request amount, response time, success rate and response rate of each application are monitored, and monitoring alarm and link positioning are automatically performed for these indicators to quickly find and locate problems.
[0068] Step S13, if there is a maintenance work order associated with the abnormal event in the abnormal monitoring result, the abnormal monitoring result and the maintenance work order associated with the abnormal event are associated and analyzed to obtain an abnormal analysis result.
[0069] Step S14, according to the abnormal analysis result, the abnormal event is repaired.
[0070] In this embodiment, in each maintenance work order, it is judged whether there is a maintenance work order associated with the abnormal event in the abnormal monitoring result. Optionally, it is judged whether the abnormal monitoring result and the maintenance work order point to the same problem, the same business module or the same technical root cause, so as to avoid repeated dispatching of the same problem.
[0071] For example, the monitoring platform finds that the abnormal monitoring result is a payment interface timeout. The system will first query the collected operation and maintenance work order: whether anyone has submitted a work order because of payment problems. Optionally, the feature set of the abnormal event: interface ID = PAY001, root cause ID = ERR_PAY_API_TIMEOUT, business module = payment business-payment interface, response time 1.2s timeout, occurrence time = 2024-10-01 14;00. Find 1 work order: work order ID = WO123, title = payment interface PAY001 timeout troubleshooting, status = in processing, creation time = 2024-10-01 13;50.
[0072] If the judgment result is that there is an associated operation and maintenance work order, the monitoring abnormal result or the operation and maintenance work order will not be handled separately, but the information of the two will be combined for analysis, and then the abnormal event will be repaired. Optionally, the detail information of the abnormal event in the abnormal monitoring result and the detail information in the operation and maintenance work order associated with the abnormal event are analyzed to obtain an abnormal analysis result. Optionally, the business module, response time, business detail information, and other information associated with the abnormal event in the abnormal monitoring result and the abnormal module, abnormal detail description, and other information in the operation and maintenance work order are analyzed to accurately troubleshoot the problem root cause and improve the accuracy and efficiency of problem repair; further, according to the abnormal analysis result, the abnormal event is repaired. Optionally, the abnormal analysis result is sent to a technical personnel to help the technical personnel to quickly locate and repair (such as expansion failure, suggesting checking whether the server is online).
[0073] In a feasible embodiment, if there is an associated operation and maintenance work order, it is determined whether the processing state of the operation and maintenance work order is a processing completed state. If it is a processing completed state, the repair scheme of the operation and maintenance work order that has completed processing and the abnormal event in the abnormal monitoring result are pushed to a technical personnel to assist the technical personnel in repairing the problem. In addition, if the processing state of the operation and maintenance work order is a processing state or a waiting for allocation state, the operation and maintenance work order and the information of the abnormal monitoring result are combined for analysis. The detail information of the abnormal event in the abnormal monitoring result and the detail information in the operation and maintenance work order associated with the abnormal event are analyzed. Further, according to the abnormal analysis result, the abnormal event is repaired to realize the linkage of monitoring data and work order information.
[0074] In an embodiment, if there is no associated operation and maintenance work order, the operation and maintenance work order and the abnormal event of the abnormal monitoring result need to be handled separately. Optionally, for the operation and maintenance work order, the operation and maintenance work order is allocated to a corresponding technical processing personnel for repair processing according to a pre-set task approval process. For the abnormal event in the abnormal monitoring result, according to the running data, the abnormal event in the abnormal monitoring result is attributed analyzed to repair the abnormal event according to the attribution analysis result.
[0075] In addition, the user can pre-save a draft of the operation and maintenance work order, and the draft is edited and formally submitted to the operation and maintenance personnel processing system for operation and maintenance application. The operation and maintenance personnel can assign a technical processing personnel for the operation and maintenance work order of the matter, and the technical processing personnel can refuse the operation and maintenance work order of the user application if the system operation is normal. After the technical personnel receives the operation and maintenance work order, the processing is performed, and after the processing is completed, the operation and maintenance work order can be submitted to the user for acceptance. If there is a problem that cannot be solved in a short time, the technical processing personnel can apply for leaving the system defect. In addition, the operation and maintenance management platform in the embodiment provides a unified entry platform connecting platform users, system operation and maintenance personnel and technical personnel, which not only meets the needs of the user side and the business side to quickly find the technical side to solve the system problem, but also meets the needs of the operation and maintenance personnel and the technical personnel to quickly solve the problem through the operation and maintenance knowledge base. Before the operation and maintenance closed-loop system is not used, the team personnel need to communicate and confirm in the business platform, the search engine platform, the user feedback platform and the monitoring and early warning platform, which is time-consuming and labor-intensive. Through data analysis and comparison, the operation and maintenance management platform in the embodiment improves the operation and maintenance efficiency by 30% to 50%.
[0076] The embodiment further associates and repairs the abnormal monitoring result and the operation and maintenance work order fed back by the user when there is an operation and maintenance work order associated with the abnormal event in the abnormal monitoring result, realizes the linkage of the monitoring data and the work order information, and effectively improves the overall operation and maintenance efficiency.
[0077] In a feasible implementation manner, referring to Figure 2 , Figure 2 The flowchart provided by the operation and maintenance method embodiment two of the present application. The running data is subjected to full-link abnormal monitoring to obtain an abnormal monitoring result, including:
[0078] Step S21, analyzing the transaction data to obtain the business process state of each transaction;
[0079] Step S22, performing abnormal monitoring on the business process state of each transaction to obtain the abnormal monitoring result.
[0080] It should be noted that the business process state refers to judging whether the business process progress of each transaction meets the expected state from the perspective of business rules, for example, the transaction state flows through the initiation of transfer, bank verification, successful transfer, and normal fund arrival. The abnormal situation includes state stagnation, for example, the bank verification is not entered 1 hour after the initiation of transfer (timeout).
[0081] In this embodiment, the transaction data is analyzed to analyze the business process state of each transaction. Optionally, the relevant information of the entire business process is aggregated according to the transaction ID from the original transaction data to determine which node of the preset business process the transaction is currently in. Further, the preset business monitoring rule is used to compare the actual business process state of the current node to identify deviations from the rule. Optionally, the business monitoring rule includes process sequence rules, time threshold rules, state legality rules, and the like. As a specific example, the process sequence rule is: to be paid, payment success, to be shipped, and the entire business process sequence cannot be skipped, for example, from the to-be-paid state directly to the shipped state. The time threshold rule is: the payment time threshold of the to-be-paid state is set to 30 minutes, and more than 30 minutes proves that the delivery process is interrupted. The state legality rule is: the business state can only be to-be-paid, payment success, payment failure, to-be-shipped, and shipped, and cannot be unknown or error code. The business process state of each transaction is monitored for abnormalities using the preset business monitoring rule, thereby obtaining the abnormal monitoring result of the related business process.
[0082] In this embodiment, by monitoring the business process state of each transaction for abnormalities, the business-level abnormalities can be quickly located, rather than only looking at technical-level error reports. The business-level abnormalities are directly related to the actual experience of users, which facilitates the business party to quickly intervene and comprehensively improves the effect and efficiency of platform operation and management.
[0083] In a feasible implementation manner, refer to Figure 3 , Figure 3 The flowchart provided for the third embodiment of the operation and management method of the present application. The running data is monitored for full-link abnormalities to obtain an abnormal monitoring result, including:
[0084] Step S31, analyzing the transaction data to obtain the calling result of the calling interface associated with each transaction;
[0085] Step S32, monitoring the calling result of the calling interface associated with each transaction for abnormalities to obtain the abnormal monitoring result.
[0086] It should be noted that the calling result includes the calling state, response code (for example, 200 indicates success, 503 indicates service unavailable, and 404 indicates that the interface does not exist), response time, and calling timestamp.
[0087] In the embodiment, all interfaces relied on by each transaction are analyzed from transaction data, and the calling results of the interfaces are analyzed, and then the calling results of the calling interfaces associated with each transaction are monitored for exceptions. Understandably, the calling exception state and the exception type of each transaction are determined according to the calling state and the response code in the calling results, and further, the abnormal monitoring result corresponding to each transaction is obtained according to the calling exception state and the exception type of each transaction.
[0088] As a specific example, the result exception is that the calling result of the payment interface (PAY_001) fails, the response code is 504 (gateway timeout), and the exception point is that the core interface calling fails.
[0089] The timing exception is that the result synchronization interface (SYNC_001) is not called (the process is interrupted due to the failure of the payment interface). The exception point is that the key interface is not called.
[0090] The response time exception is that the response time of the payment interface exceeds the preset time threshold. The exception point is that the response of the core interface is timed out.
[0091] In the embodiment, the calling results of the calling interfaces associated with each transaction are analyzed, and then the calling results of the calling interfaces associated with each transaction are monitored for exceptions, so that the technical root cause of the financial transaction exception can be directly located, instead of only staying on the surface of the business state exception, and accurate basis is provided for the technical team to troubleshoot faults (such as interface service downtime and network delay), and the effect of platform operation and maintenance management is comprehensively improved.
[0092] In a feasible implementation manner, referring to Figure 4 , Figure 4 The flowchart provided for the fourth embodiment of the operation and maintenance method of the application. The running data is monitored for full-link exceptions to obtain an exception monitoring result, including:
[0093] In step S41, the system log information and the business log information are preprocessed.
[0094] In step S42, the preprocessed system log information and business log information are associated and analyzed to obtain an association analysis result.
[0095] In step S43, the exception monitoring result is determined according to the association analysis result.
[0096] It should be noted that the system log information is a log recording the running state of the system itself, the technical level operation, and the behavior of the underlying resources and components, including system startup log, server log, network device log, operating system log and the like.
[0097] It should be noted that the business log information refers to the log directly related to business operation, business data, business state, etc. recorded by the system when performing a specific business process. The business log information includes business process link, execution state, business data content (specific data related to business), business start time, final result of business execution, etc. so that the business system can be monitored abnormally according to the business log information.
[0098] In an embodiment, the business log information is acquired, including:
[0099] Step A11, determining the log collection rule associated with each business system, wherein the log collection rule includes at least one of the collection rules corresponding to the system granularity, the function module granularity and the method function granularity;
[0100] Step A12, collecting the business log information of each business system according to at least one of the collection rules corresponding to the system granularity, the function module granularity and the method function granularity.
[0101] The log collection rule associated with each business system is set in advance according to different business scenarios. Optionally, the log collection rule includes collection rules corresponding to the system granularity, the function module granularity and the method function granularity, etc. It should be noted that the system granularity is configured with a unified basic log format (such as timestamp, log level, cluster node ID) for the whole system cluster to ensure that the whole engineering log is consistent. The function module granularity is configured with whether to collect the log of all methods under the class (for example, PaymentService class of payment engineering, SettlementService class of settlement engineering) to avoid the redundancy of the log of other classes. The method function granularity is configured with the business data (such as payment amount, order number, user ID) that needs to be collected for the key business method (for example, doPay() payment method and queryBalance() balance query method of PaymentService) to realize the collection of key data actually needed, without the collection of other redundant information, improve the efficiency of log information collection, and further improve the efficiency of operation and maintenance. In this embodiment, the business log information of each business system is collected according to the collection rules of the system granularity, the function module granularity and the method function granularity, etc.
[0102] In an embodiment, after collecting the system log information and the business log information, the system log information and the business log information are not directly analyzed in the business thread, but are sent to a message queue, for example, Kafka, and are asynchronously consumed by an independent log analysis service, which not only avoids the delay of business interface due to log processing, but also can cope with the log peak in the high concurrency scenario.
[0103] In the embodiment, the system log information and the business log information are preprocessed, optionally, invalid or duplicate log records are removed, log of different sources is formatted into a unified format, and subsequent processing is facilitated. In addition, the system log information and the business log information can be time-aligned to ensure that the time stamp formats of all logs are consistent, and the system log information and the business log information are aligned by the time stamp, so that the system log information and the business log information are analyzed in chronological order. Optionally, the system log information and the business log information are matched according to the time stamp, for example, system errors and user operations occurring in the same time range are found. Further, algorithms such as association rule mining and clustering analysis are used to discover potential relationships between the system log information and the business information. In an embodiment, the logs can also be associated according to context information, for example, all operation logs of a certain user ID are associated with corresponding error logs in the system log. In other embodiments, the system log information and the business information can also be associated according to a specific event, for example, the business log of a certain transaction failure is associated with the related error log in the system.
[0104] Further, the preprocessed system log information and business log information are analyzed, optionally, feature information in the system log information and feature information in the business log information are extracted, and then the feature information in the system log information and the feature information in the business log information are integrated, and then the integrated feature information is analyzed to obtain an association analysis result. By analyzing the system log information and the business log information, deep abnormalities that cannot be discovered by single log can be identified, and the accuracy of abnormal monitoring can be improved.
[0105] In a feasible implementation, for example, the system log includes: interface information: called core interface (PAY_API_001), interface state (timeout, success, failure), response time (1.2s), error code (504); resource state: CPU (85%) of the server where the interface is located; time information: interface calling time (14;00;05). Business log information: business operation type (initiate payment), operator (USER_456), operation time (14;00;03); state information: order state change (to be paid→payment failure), failure description (payment timeout, please retry); business data: payment amount (299 yuan), payment method (WeChat payment). The integrated feature information includes interface ID; PAY_API_001, interface state; timeout, response time; 1.2s, business feature; order state; to be paid→payment failure, payment amount; 299 yuan, failure description; payment timeout, association time; 2024-10-0114;00;05.
[0106] Further, the association analysis result is monitored for abnormality, and optionally, an abnormality detection model pre-constructed is used to automatically identify abnormal behavior, to obtain the abnormality monitoring result, wherein the abnormality detection model is obtained by iterative training according to historical system log information, historical business log information, and pre-labeled abnormal event labels. In other embodiments, a log information related index threshold value can also be set, and an alarm is triggered when the business data exceeds the index threshold value. For example, the number of error logs exceeds a certain threshold. Define patterns of abnormal behavior, such as the frequent occurrence of certain error logs or a specific sequence of logs.
[0107] The embodiment can directly associate technical problems in the system log with user operations in the business log by performing association analysis on the system log information and the business log information, quickly locate the root cause of the problem, and reduce the time to troubleshoot the problem. In addition, by performing association analysis on the log information, deep abnormalities that cannot be discovered by a single log can be identified, thereby faster problem resolution and reduced impact on business operation.
[0108] In a feasible implementation, if there is no associated operation and maintenance work order, the operation and maintenance work order is repaired, including:
[0109] Step S51, the operation and maintenance work order is checked;
[0110] Step S52, if the check is successful, the operation and maintenance work order is distributed to the corresponding technical processing personnel for repair processing according to a preset task approval process;
[0111] In the embodiment, the basic information of the operation and maintenance work order is checked for integrity, business compliance, and redundancy, etc. For example, the operation and maintenance work order includes a description of the specific abnormal function, and it is determined whether the abnormal problem described in the operation and maintenance work order is consistent with the business rules, for example, the payment interface timeout needs to be associated with the specific order range to avoid ambiguous descriptions. Redundancy checking refers to comparing existing processing or solved operation and maintenance work orders, and if there is a work order with the same root cause or the same module, it is determined as a duplicate work order. If the check fails, a prompt message is generated and returned to the user. If the check is successful, the operation and maintenance work order is distributed to the corresponding technical processing personnel for repair processing according to a preset task approval process; the task approval process can be specifically referred to Figure 5 , Figure 5 The task approval process provided by an embodiment of the present application can optionally determine the priority of all unassigned operation and maintenance work orders during the distribution process, and distribute each unassigned operation and maintenance work order according to the priority of each unassigned operation and maintenance work order.
[0112] Further, during the repairing process, the processing state of the operation and maintenance order in the task approval process is tracked, such as in processing, processing completed, and the like, and the processing state is fed back to the user.
[0113] Further, the user can record or collect the operation and maintenance order, and the next time the system failure is encountered, it can be checked whether similar problems have occurred before, and a solution is obtained therefrom to quickly respond to the production accident of the system.
[0114] The embodiment constructs an operation and maintenance closed-loop system of a user end, a system operation and maintenance end, and a system technology end, the user logs in the platform, a newly created operation and maintenance order of a problem is sent to an operation and maintenance management personnel, and the operation and maintenance management personnel distributes the operation and maintenance order of the system to a technical personnel, thereby realizing quick response of the problem.
[0115] In a feasible implementation manner, if there is no associated operation and maintenance order, the abnormal monitoring result is repaired, including:
[0116] In step S61, according to the running data, the abnormal event in the abnormal monitoring result is analyzed to obtain an abnormal root cause analysis result.
[0117] In step S62, according to the abnormal root cause analysis result, the abnormal event is repaired.
[0118] In an embodiment, according to the running data, time series data before and after the abnormal event is analyzed to find out the abnormal point, and in addition, detailed information of the abnormal event is also analyzed. According to the detailed information of the abnormal event, possible root causes, and influence range, a detailed abnormal root cause analysis result is generated. In other embodiments, an analysis model can also be trained according to the running data and the pre-labeled attribution result of the abnormal event in advance, and then the analysis model is used to attribute analyze the abnormal event in the abnormal monitoring result to obtain the abnormal root cause analysis result.
[0119] Further, according to the abnormal root cause analysis result, the abnormal event is repaired. Optionally, the abnormal root cause analysis result is pushed to the technical personnel to assist the technical personnel in repairing the abnormal event.
[0120] The embodiment analyzes the abnormal event in the abnormal monitoring result to obtain an abnormal root cause analysis result, and then according to the abnormal root cause analysis result, the abnormal event is repaired, thereby ensuring that the abnormal event is timely and effectively processed, and the stability and reliability of the system are improved.
[0121] In a feasible implementation manner, with reference to Figure 6 , Figure 6A flowchart provided by the sixth embodiment of the operation and maintenance management method of the present application. After the repair processing of the abnormal event according to the abnormal analysis result, further comprising:
[0122] Step S71, determining the repair scheme associated with the abnormal event in the operation and maintenance work order and the abnormal event in the abnormal monitoring result;
[0123] Step S72, storing the abnormal event and the repair scheme associated with the abnormal event in the preset knowledge base;
[0124] In an embodiment, after processing any abnormal event, the repair scheme associated with the abnormal event is recorded, and then the repair scheme associated with the abnormal event in the operation and maintenance work order is stored in the preset knowledge base, and the repair scheme associated with the abnormal event in the abnormal monitoring result is stored in the preset knowledge base. For the convenience of subsequent rapid matching, the knowledge base needs to use the storage structure of index, detail information of abnormal event and repair scheme.
[0125] Step S73, if the abnormal event is monitored again, the monitored abnormal event and the preset knowledge base are matched;
[0126] Step S74, if the matching is successful, the repair processing is performed using the matched repair scheme associated with the abnormal event.
[0127] In an embodiment, when the monitoring system detects the abnormal event again, the system first extracts the core features of the abnormal event, and then queries and matches in the preset knowledge base to find the most similar repair scheme. If the matching is successful, the repair processing is performed using the matched repair scheme associated with the abnormal event. In order to further improve the accuracy of operation and maintenance, after the repair scheme is queried, the abnormal event and the repair scheme are pushed to the technical personnel again, and after manual verification, the system performs repair processing using the matched repair scheme associated with the abnormal event;
[0128] The present embodiment stores the abnormal event and the repair scheme associated with the abnormal event in the preset knowledge base, so that the operation and maintenance personnel and the technical personnel can directly reuse the scheme for similar subsequent abnormalities, avoid repeated problem investigation, and effectively improve the problem repair efficiency.
[0129] It should be noted that the examples in the figure are only used to understand the present application and do not constitute a limitation on the operation and maintenance management method of the present application. More simple transformations based on this technical concept are within the scope of protection of the present application.
[0130] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0131] The application also provides an operation and maintenance management device, please refer to Figure 7 , Figure 7 The application also provides an operation and maintenance management device, please refer to
[0132] The acquisition module 81 is used for acquiring to-be-processed information, wherein the to-be-processed information comprises running data of each business system and user feedback operation and maintenance work orders;
[0133] The monitoring module 82 is used for performing full-link abnormality monitoring on the running data to obtain an abnormality monitoring result;
[0134] The analysis module 83 is used for, if there is an operation and maintenance work order associated with an abnormal event in the abnormality monitoring result, performing correlation analysis on the abnormality monitoring result and the operation and maintenance work order associated with the abnormal event to obtain an abnormality analysis result;
[0135] The repair module 84 is used for performing repair processing on the abnormal event according to the abnormality analysis result.
[0136] The monitoring module 82 is further used for:
[0137] performing analysis on the transaction data to obtain a business process state of each transaction;
[0138] performing abnormality monitoring on the business process state of each transaction to obtain the abnormality monitoring result; and / or,
[0139] performing analysis on the transaction data to obtain a calling result of a calling interface associated with each transaction;
[0140] performing abnormality monitoring on the calling result of the calling interface associated with each transaction to obtain the abnormality monitoring result.
[0141] The monitoring module 82 is further used for:
[0142] performing preprocessing on the system log information and the business log information;
[0143] performing correlation analysis on the preprocessed system log information and business log information to obtain a correlation analysis result;
[0144] determining the abnormality monitoring result according to the correlation analysis result.
[0145] The acquisition module 81 is further used for:
[0146] determining log collection rules associated with each business system, wherein the log collection rules comprise at least one of collection rules corresponding to system granularity, function module granularity and method function granularity;
[0147] According to at least one of the acquisition rules corresponding to the system granularity, the function module granularity and the method function granularity, the business log information of each business system is acquired.
[0148] The operation and maintenance management device further comprises:
[0149] The verification module is configured to verify the operation and maintenance order.
[0150] The distribution module is configured to, if the verification is successful, distribute the operation and maintenance order to a corresponding technical processing personnel for repair processing according to a preset task approval process.
[0151] The operation and maintenance management device further comprises:
[0152] The abnormal root cause analysis module is configured to analyze the abnormal events in the abnormal monitoring result according to the running data to obtain an abnormal root cause analysis result.
[0153] The repair processing module is configured to perform repair processing on the abnormal events according to the abnormal root cause analysis result.
[0154] The operation and maintenance management device further comprises:
[0155] The determination module is configured to determine the abnormal events in the operation and maintenance order and a repair scheme associated with the abnormal events in the abnormal monitoring result.
[0156] The storage module is configured to store the abnormal events and the repair scheme associated with the abnormal events in a preset knowledge base.
[0157] The matching module is configured to, if an abnormal event is monitored again, match the monitored abnormal event with the preset knowledge base.
[0158] The repair processing module is configured to, if the matching is successful, perform repair processing by using the matched repair scheme associated with the abnormal event.
[0159] The operation and maintenance management device provided in the present application adopts the operation and maintenance management method in the above embodiments and can solve the technical problems in the background art. Compared with the prior art, the operation and maintenance management device provided in the present application has the same beneficial effects as the operation and maintenance management method provided in the above embodiments, and other technical features in the operation and maintenance management device are the same as the features disclosed in the above embodiments, which will not be described herein.
[0160] The electronic device provided in the present application comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the operation and maintenance management method in the above embodiment one.
[0161] Reference will be made to the following description Figure 8 , which shows a structural schematic diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0162] As Figure 8 shown, the electronic device can include a processing device 1001 (such as a central processor, a graphics processor, or the like) that can perform various appropriate actions and processes according to programs stored in a read-only memory 1002 or loaded from a storage device 1003 into a random access memory 1004. Various programs and data required for operation of the electronic device are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; the storage device 1003 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the electronic device with various systems is shown in the figure, it should be understood that all the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.
[0163] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.
[0164] The electronic device provided by the present application adopts the operation and maintenance management method in the above-mentioned embodiments, and can solve the technical problems in the background art. Compared with the prior art, the electronic device provided by the present application has the same beneficial effects as the operation and maintenance management method provided by the above-mentioned embodiments, and other technical features in the electronic device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.
[0165] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0166] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0167] The present application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e. computer program) for executing the operation and maintenance management method in the above-mentioned embodiments.
[0168] The computer readable storage medium provided in the application may be, for example, a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system or device, or any combination of the above. More specific examples of the computer readable storage medium may include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment, the computer readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system or device. The program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to an electric wire, an optical cable, an RF (Radio Frequency), and the like, or any suitable combination of the above.
[0169] The computer readable storage medium described above may be contained in an electronic device, or may exist separately without being assembled into an electronic device.
[0170] The computer readable storage medium described above carries one or more programs, which, when executed by an electronic device, cause the electronic device to:
[0171] Obtain to-be-processed information, wherein the to-be-processed information includes running data of each service system and user feedback operation and maintenance work orders;
[0172] Perform full-link abnormality monitoring on the running data to obtain an abnormality monitoring result;
[0173] If there is an operation and maintenance work order associated with an abnormal event in the abnormality monitoring result, the abnormality monitoring result and the operation and maintenance work order associated with the abnormal event are associated for analysis to obtain an abnormality analysis result;
[0174] According to the abnormality analysis result, the abnormal event is repaired.
[0175] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0176] The flow diagrams and the block diagrams in the drawings are meant as methodological and functional description of implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0177] The modules involved in the embodiments of the present application can be implemented in software or hardware. In some cases, the names of the modules do not constitute a limitation on the modules themselves.
[0178] The readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the above-mentioned operation and maintenance management method, and can solve the technical problems in the background art. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the operation and maintenance management method provided by the above-mentioned embodiments, which will not be described here.
[0179] The embodiment of the present application provides a computer program product, comprising a computer program, which realizes the steps of the operation and maintenance management method when executed by a processor.
[0180] The computer program product provided by the present application can solve the technical problems in the background art. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiment of the present application are the same as those of the operation and maintenance management method provided by the above-mentioned embodiment, and are not described here.
[0181] The above-mentioned is only part of the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation, direct / indirect application in other related technical fields made by using the content of the present application specification and drawings, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An operation and maintenance management method, characterized by, The method comprises the following steps: acquiring to-be-processed information, wherein the to-be-processed information comprises operation data of each business system and user feedback operation and maintenance work orders; performing full-link abnormality monitoring on the operation data to obtain abnormality monitoring results; if there is an operation and maintenance work order associated with an abnormal event in the abnormality monitoring results, performing correlation analysis on the abnormality monitoring results and the operation and maintenance work order associated with the abnormal event to obtain abnormality analysis results; performing repair processing on the abnormal event according to the abnormality analysis results.
2. The operation and maintenance management method according to claim 1, wherein The operation data comprises transaction data; the step of performing full-link abnormality monitoring on the operation data to obtain abnormality monitoring results comprises the following steps: performing analysis on the transaction data to obtain a business process state of each transaction; performing abnormality monitoring on the business process state of each transaction to obtain the abnormality monitoring results; and / or performing analysis on the transaction data to obtain a calling result of a calling interface associated with each transaction; performing abnormality monitoring on the calling result of the calling interface associated with each transaction to obtain the abnormality monitoring results. 3.The operation and maintenance management method of claim 1, wherein, The operation data comprises system log information and business log information; the step of performing full-link abnormality monitoring on the operation data to obtain abnormality monitoring results comprises the following steps: performing preprocessing on the system log information and the business log information; performing correlation analysis on the preprocessed system log information and business log information to obtain correlation analysis results; determining the abnormality monitoring results according to the correlation analysis results.
4. The operation and maintenance management method according to claim 1, wherein The operation data comprises business log information; the step of acquiring business log information comprises the following steps: determining log collection rules associated with each business system, wherein the log collection rules comprise at least one of collection rules corresponding to a system granularity, a function module granularity and a method function granularity; collecting business log information of each business system according to at least one of the collection rules corresponding to the system granularity, the function module granularity and the method function granularity. 5.The operation and maintenance management method of claim 1, wherein, If there is no associated operation and maintenance work order, performing repair processing on the operation and maintenance work order comprises the following steps: verifying the operation and maintenance work order; if the verification is successful, distributing the operation and maintenance work order to corresponding technical processing personnel for repair processing according to a preset task approval process.
6. The operation and maintenance management method according to claim 1, wherein If there is no associated operation and maintenance work order, performing repair processing on the abnormality monitoring results comprises the following steps: performing analysis on an abnormal event in the abnormality monitoring results according to the operation data to obtain abnormality root cause analysis results; performing repair processing on the abnormal event according to the abnormality root cause analysis results.
7. The operation and maintenance method according to any one of claims 1 to 6, wherein After performing repair processing on the abnormal event according to the abnormality analysis results, the method further comprises the following steps: determining a repair scheme associated with the abnormal event in the operation and maintenance work order and the abnormal event in the abnormality monitoring results; storing the abnormal event and the repair scheme associated with the abnormal event in a preset knowledge base; if an abnormal event is monitored again, matching the monitored abnormal event with the preset knowledge base; if the matching is successful, performing repair processing by using the repair scheme associated with the matched abnormal event.
8. An operation and maintenance management apparatus characterized by comprising: The method comprises the following steps: An acquisition module is configured to acquire to-be-processed information, wherein the to-be-processed information comprises operation data of each service system and user feedback operation and maintenance work orders; A monitoring module is configured to perform full-link abnormality monitoring on the operation data to obtain abnormality monitoring results; An analysis module is configured to, if there is an operation and maintenance work order associated with an abnormal event in the abnormality monitoring results, perform correlation analysis on the abnormality monitoring results and the operation and maintenance work order associated with the abnormal event to obtain abnormality analysis results; A repair module is configured to perform repair processing on the abnormal event according to the abnormality analysis results.
9. An electronic device, comprising: The electronic device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the operation and maintenance management method according to any one of claims 1 to 7.
10. A storage medium, characterized by The storage medium is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the operation and maintenance management method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for operating an internal combustion engine
WO2024100113A1