Operation and maintenance monitoring management method and device, electronic equipment and medium
By using anomaly detection models in the financial system to process performance and monitoring data, the inefficiency and misjudgment of traditional operation and maintenance monitoring methods are solved, and more efficient operation and maintenance and system stability are achieved.
Patent Information
- Application Number
- CN202510091934.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The traditional operation and maintenance monitoring method has low rates, which are prone to omissions and misjudgments, and cannot meet the complex business environment and technical challenges of modern financial institutions.
By obtaining the performance data and monitoring data of the financial system and inputting it into the pre-trained exception detection model, an exception alarm message is obtained, and the abnormal location and cause of the exception are determined based on the exception alarm message and monitoring data.
It realizes timely discovering and solving potential risks, preventing system failures, and improving operation and maintenance efficiency and system stability.
Smart Images

Figure CN120013676A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to an operation and maintenance monitoring management method, device, electronic equipment and medium. Background Art
[0002] In today's era of rapid digital development, financial institutions are facing increasingly complex business environments and technical challenges. With the continuous expansion and innovation of financial services, the scale and complexity of information technology systems are also growing rapidly. Key applications such as financial institutions' core business systems, channel systems, and risk management systems need to maintain stable operation to ensure that customers can conduct financial transactions at any time.
[0003] The network architecture of financial institutions is becoming increasingly large, covering connections between multiple data centers, branches and external institutions, and some businesses have been deployed in the cloud, which greatly increases the difficulty of monitoring and management. Traditional operation and maintenance monitoring methods have low efficiency, are prone to omissions and misjudgments, and can no longer meet the needs of modern financial institutions. Summary of the invention
[0004] In view of this, an object of the present invention is to provide an operation and maintenance monitoring management method, device, electronic device and medium to improve the operation and maintenance efficiency and system stability.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention provides an operation and maintenance monitoring management method, comprising: obtaining performance data and monitoring data of a financial system; inputting the performance data and monitoring data into a pre-trained anomaly detection model to obtain an abnormal alarm message; and determining the abnormal location and cause of the abnormality based on the abnormal alarm message and monitoring data.
[0007] Optionally, the performance data and monitoring data are input into a pre-trained anomaly detection model to obtain an abnormal alarm message, including: using pattern recognition to identify the data types of the performance data and monitoring data to obtain time series periodic data; inputting the time series periodic data into a pre-trained anomaly detection model to obtain an abnormal alarm message; wherein the anomaly detection model is trained using a machine learning model, including a single-indicator anomaly detection model and a multi-indicator anomaly detection model.
[0008] Optionally, the monitoring data also includes log data; the above method also includes: parsing the log data based on a log parsing algorithm to obtain a log pattern of the log data; performing log anomaly detection based on the log pattern of the log data to obtain an abnormal alarm message, and performing root cause analysis on the abnormal alarm message to determine the cause of the abnormality.
[0009] Optionally, log anomaly detection is performed based on the log pattern of the log data, including: comparing the log pattern of the log data with a pre-set normal log pattern to determine whether the log data is an abnormal pattern log; determining whether the log data is a quantity abnormality log based on the relationship between the log quantities of different log patterns; performing log sequence anomaly detection based on a pre-trained deep learning model to determine whether the log data is a sequence abnormality log.
[0010] Optionally, a root cause analysis is performed on the abnormal alarm message to determine the cause of the abnormality, including: obtaining time series data from the log of the abnormal alarm message, and constructing a directed acyclic graph based on the time series data and a causal relationship algorithm; and determining the cause of the abnormality of the abnormal alarm message based on the directed acyclic graph.
[0011] Optionally, the monitoring data also includes business indicator data; the above method also includes: obtaining historical data of the financial system, and extracting business processes and data processes from the historical data to build a business prediction model; predicting the business indicator data based on the business prediction model to obtain prediction results; determining abnormal alarm messages based on the prediction results of the business indicator data.
[0012] Optionally, it also includes: compressing the abnormal alarm message into an alarm according to the alarm source type of the abnormal alarm message and a preset compression rule; and merging the alarm into an abnormal alarm event according to a preset merging rule.
[0013] In the second aspect, the present invention provides an operation and maintenance monitoring management device, including: a data acquisition module, used to acquire performance data and monitoring data of a financial system; an anomaly detection module, used to input the performance data and monitoring data into a pre-trained anomaly detection model to obtain an abnormal alarm message; an anomaly analysis module, used to determine the abnormal location and cause of the abnormality based on the abnormal alarm message and monitoring data.
[0014] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any one of the methods provided in the first aspect above.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any one of the methods provided in the first aspect are executed.
[0016] The present invention brings the following beneficial effects:
[0017] The above-mentioned operation and maintenance monitoring management method, device, electronic device and medium provided by the present invention first obtain the performance data and monitoring data of the financial system; then input the performance data and monitoring data into a pre-trained anomaly detection model to obtain an abnormal alarm message; finally, determine the abnormal location and abnormal cause based on the abnormal alarm message and monitoring data. The above-mentioned method can obtain an abnormal alarm message based on the performance data and monitoring data of the system through a pre-trained anomaly detection model, and further determine the abnormal location and abnormal cause, so as to timely discover and solve potential risks, prevent system failures, and improve operation and maintenance efficiency and system stability.
[0018] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the preferred embodiments are specifically mentioned below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A flowchart of an operation and maintenance monitoring management method provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of an alarm suppression mechanism provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of the structure of an operation and maintenance monitoring management device provided by an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] At present, the traditional operation and maintenance monitoring method has a low efficiency, is prone to omissions and misjudgments, and can no longer meet the needs of modern financial institutions. Based on this, an operation and maintenance monitoring management method, device, electronic device and medium provided by the embodiment of the present invention can improve the operation and maintenance efficiency and system stability.
[0027] To facilitate understanding of this embodiment, firstly, a method for operation and maintenance monitoring management disclosed in an embodiment of the present invention is introduced in detail. The method can be executed by an electronic device, such as a smart phone, a computer, a tablet computer, etc. Figure 1 The flowchart of an operation and maintenance monitoring management method shown in FIG. 1 illustrates that the method mainly includes the following steps S101 to S103:
[0028] Step S101: Acquire performance data and monitoring data of the financial system.
[0029] In one implementation, a unified collection and data processing module can be used to uniformly collect, clean, convert, and send various types of performance data and monitoring data, while providing unified and standardized scheduling and control of data collection tasks and behaviors, and quickly accessing various types of production data within financial institutions.
[0030] The data collection methods in this embodiment include: Agent mode and non-Agent mode; Agent mode includes plug-in collection, script collection, log collection, process collection, APM probe, etc.; non-Agent mode includes general protocol collection, Web dialing, API interface, etc. The collection frequency is divided into three types: second level, minute level, and random level. The commonly used collection frequency in the embodiment of the present invention is minute level. The collection transmission can be classified according to the transmission initiation or the transmission link. According to the transmission initiation classification, there are active collection Pull (pull) and passive reception Push (push). According to the transmission link classification, there are direct connection mode and Proxy transmission. Among them, Proxy transmission can not only solve the problem of cross-network transmission of monitoring data, but also alleviate the bottleneck of data transmission caused by too many monitoring nodes, and use Proxy to achieve data diversion.
[0031] In the specific implementation, the monitoring data in this embodiment includes at least three types: indicator data, log data, and tracking data. Indicator data is a numerical monitoring item, which is mainly identified by dimensions. Log data is character data, which is mainly used to find some keyword information for monitoring. Tracking data feedback is the process of tracking the data flow of a link, and observes whether the time-consuming performance in the process is normal.
[0032] In the embodiment of the present invention, the following three storage methods can also be used for data storage:
[0033] (1) Relational databases. For example: MySQL, MSSQL, DB2.
[0034] (2) Time series database. Good at storing and calculating indicator data; for example, InfluxDB, OpenTSDB (based on Hbase), Prometheus, etc.
[0035] (3) Full-text search database. Mainly used for log-type storage, such as Elasticsearch.
[0036] Step S102: Input the performance data and monitoring data into a pre-trained anomaly detection model to obtain an anomaly alarm message.
[0037] In one implementation, a machine learning algorithm can be used to implement anomaly detection of performance data, monitoring data, etc. First, based on historical data, a machine learning algorithm is used to train an anomaly detection model, and then the trained anomaly detection model is used to perform anomaly detection on the acquired performance data and monitoring data to obtain anomaly alarm messages.
[0038] Step S103: Determine the abnormal location and abnormal cause based on the abnormal alarm message and monitoring data.
[0039] In one implementation, after the abnormality alarm message is obtained, the abnormality can be located and the root cause analyzed according to the monitoring data and the abnormality alarm message to determine the abnormality location and the abnormality cause.
[0040] The above-mentioned operation and maintenance monitoring management method provided by the present invention can obtain abnormal alarm messages based on the system's performance data and monitoring data through a pre-trained abnormality detection model, and further determine the abnormal location and cause of the abnormality, so as to timely discover and resolve potential risks, prevent system failures, and improve operation and maintenance efficiency and system stability.
[0041] In one embodiment, for the aforementioned step S102, that is, when the performance data and the monitoring data are input into a pre-trained anomaly detection model to obtain an abnormal alarm message, the following method can be used but is not limited to: first, the data types of the performance data and the monitoring data are identified by using pattern recognition to obtain time series periodic data; then the time series periodic data is input into a pre-trained anomaly detection model to obtain an abnormal alarm message; wherein the anomaly detection model is trained using a machine learning model, including a single-indicator anomaly detection model and a multi-indicator anomaly detection model.
[0042] In specific implementation, the anomaly detection model used in the embodiment of the present invention can learn the historical data baseline based on historical data, and further predict the future data baseline, and perform data anomaly detection according to the future data baseline. Specifically, it can be trained in the following way:
[0043] (1) Data collection and preprocessing
[0044] Collect relevant data of historical time periods and use segmentation methods to remove statistical outliers. Then preprocess the collected data, including:
[0045] Data cleaning. For example, log data cleaning. Since log data is unstructured data with low information density, it is necessary to extract useful data from it through data cleaning.
[0046] Data calculation. Since many raw performance data cannot be directly used to determine whether the data is abnormal, it is necessary to calculate it. For example, if the collected data is the total disk volume and disk usage, if you want to detect the disk usage rate, perform four arithmetic operations on the existing data to obtain the disk usage rate.
[0047] Data enrichment. That is, data is tagged with tags, such as host and computer room tags, to facilitate aggregate calculations.
[0048] Indicator derivation: that is, obtaining new data through calculation based on existing data.
[0049] (2) Model establishment
[0050] Specifically, the model is constructed as follows: y(t) = g(t) + s(t) + h(t) + ∈
[0051] Among them, g(t) is the trend function, which is used to analyze non-periodic changes in time series. s(t) represents periodic changes, such as a week or a year. h(t) represents the impact of an accidental day or a few days such as holidays. ∈ is the error term, which represents the impact of errors not considered in this model.
[0052] (3) Data fitting and prediction
[0053] When anomalies occur in historical data, there are often some anomalies that cannot be detected by the above statistical methods, and these anomalies cannot be removed in the preprocessing stage. In existing methods, when there are many anomalies in historical data that cannot be removed by statistical methods (a common scenario in actual data), the robustness of the baseline obtained by model fitting will be greatly affected, resulting in unstable and inaccurate predicted baselines, which cannot meet the expected results.
[0054] Based on this, in this embodiment, when solving the optimal solution of the baseline, an improved method is adopted, which can greatly reduce the impact of outliers on baseline prediction and enhance the robustness of model learning to outliers.
[0055] (4) Anomaly detection algorithm design
[0056] Through the collected data, machine learning algorithms are used to detect single-indicator anomalies of performance data, monitoring data, and other data. At the same time, according to different business scenarios, after selecting the corresponding indicator data, machine learning algorithms can also be used to perform multi-indicator correlation analysis to achieve multi-indicator anomaly detection for application system clusters and various business scenarios. For periodic data, the platform can automatically set the dynamic threshold of the single indicator curve to help operation and maintenance personnel improve operation and maintenance efficiency.
[0057] The anomaly detection model provided by the embodiment of the present invention can perform anomaly detection on business performance indicator data, such as transaction volume, response rate, response time, success rate, and other data with fixed time intervals, time series regularity or periodicity, as well as indicator data that can reflect the health of the business system, identify abnormal changes in business indicator trends, discover problem risks early, diagnose and repair them, and shorten fault discovery and recovery time.
[0058] In one implementation, the anomaly detection process is as follows:
[0059] 1) Collect relevant data (performance data and monitoring data) in historical time periods, and access the data through the big data platform.
[0060] 2) Data preprocessing: Detect and remove noise data and irrelevant data in the data set through data cleaning, data integration, data transformation, data reduction, etc. This includes processing leaky data and removing blank data, thereby improving data quality.
[0061] 3) Data pattern recognition: Different data types (time series periodic data) use different algorithm models. Therefore, first use pattern recognition to effectively diagnose the data type, then use the time series clustering model to analyze different time series, and further classify different categories of time series.
[0062] 4) Perform anomaly detection on different data: For different data patterns, time series decomposition is used for periodic data, and periodic statistical test algorithms are used for binary data.
[0063] The single-indicator anomaly detection algorithm uses machine learning to summarize the changing rules of the periodicity and stability of time series data based on data characteristics such as distance, density, and frequency. At the same time, under the premise of ensuring the robustness of the algorithm based on the 3σ rule, the collected data is input into the anomaly detection model for comparative testing to determine the abnormal alarm message.
[0064] In the embodiment of the present invention, the anomaly detection model can use fixed rules and machine learning algorithms to identify anomalies. Fixed algorithms are more common algorithms, such as static thresholds, year-on-year and month-on-month changes, and custom rules, while machine learning mainly includes dynamic baselines, burr detection, indicator prediction, multi-indicator correlation detection and other algorithms. Whether it is a fixed rule or machine learning, there will be corresponding judgment rules, that is, the common <, >, > = and and / or combination judgments.
[0065] In one embodiment, the monitoring data further includes log data; and the above method further includes:
[0066] Step (1) parses the log data based on a log parsing algorithm to obtain a log pattern of the log data.
[0067] In specific implementation, log parsing is an important foundation for all subsequent log analysis technologies. Log parsing technology can extract the inherent patterns of log text from source code or log information. On this basis, anomaly detection and root cause analysis can reduce data processing pressure and improve analysis accuracy.
[0068] The log text consists of a constant part and a parameter part. The log parsing algorithm can extract the constant part in the log as the log pattern. For example, in the log 'Received block blk_-562725280853087685of size67108864from / 10.251.91.84', the constant part is 'Received', 'block', 'of', 'size', 'from', and the parameter part is 'blk_-562725280853087685', '67108864', ' / 10.251.91.84'. Therefore, the log pattern of the log is 'Received block*of size*from*', where '*' represents the position of the parameter part.
[0069] At present, the existing log parsing algorithms are mainly divided into two categories: source code-based log parsing technology and data mining-based log parsing technology. Source code-based log parsing: Log events are uniquely associated with log statements in the source code. Therefore, automatic log parsing can be performed based on the relevant statements in the source code that print logs. First, static program analysis is used to extract the log template in the source code, and then regular expressions are automatically generated based on the log template to match the corresponding log message.
[0070] Log parsing technology based on data mining can be divided into three categories: clustering algorithm, heuristic algorithm, and frequent pattern mining algorithm. Clustering algorithm divides logs into different categories by calculating the similarity between logs; heuristic algorithm divides logs based on prior knowledge, token position, log length, etc.; frequent pattern mining algorithm counts high-frequency words or high-frequency word pairs in logs to obtain log patterns.
[0071] Due to the semi-structured data characteristics of logs, log parsing needs to consider the structured and unstructured (text) properties of logs. For the structured part, multiple algorithm model sets can be used for algorithm optimization; for unstructured logs, natural language processing technology can be used to construct word vectors for unstructured logs, further enhancing the ability of log parsing.
[0072] Step (2): Perform log anomaly detection based on the log pattern of the log data to obtain an abnormal alarm message, and perform root cause analysis on the abnormal alarm message to determine the cause of the abnormality.
[0073] In specific implementation, log anomaly detection is performed based on the log mode of log data, including:
[0074] (1) Compare the log mode of the log data with the pre-set normal log mode to determine whether the log data is an abnormal mode log.
[0075] In the event of abnormal machine login or system failure, the system will generate abnormal logs. These abnormal logs are often submerged in a large number of logs. If they cannot be detected immediately, they will seriously affect the stability of the system. Log abnormal pattern detection can detect logs that are different from the normal pattern in historical logs and online streaming logs. Specifically, the log pattern of the obtained log data can be matched with the historical normal log pattern to determine whether it is an abnormal pattern log.
[0076] (2) Determine whether the log data is an abnormal log based on the log quantity relationship of different log modes.
[0077] Quantitative relationships that consistently hold in system logs under different inputs and workloads are considered program invariants. These linear relationships can capture normal program execution behavior. If new logs break certain invariants, it can be assumed that an anomaly has occurred during system execution.
[0078] Log statistical anomaly detection is used to detect the abnormality of the quantitative relationship between log patterns, and for logs with workflows, detect execution anomalies therein. For logs with workflows, there is a constant quantitative relationship between the number of logs generated by each execution node in the process, that is, the program invariant. The program invariant is a linear relationship, which is always maintained during the operation of the system even under different inputs and different workloads. For example, when the system is executing normally, the number of 'Openfile' logs is equal to 'Close file'. When this quantitative relationship is destroyed, it means that an abnormality has occurred in the file operation. Based on this, in the embodiment of the present invention, it is possible to detect whether it is a log with abnormal quantity based on the number of logs in different log patterns.
[0079] (3) Perform log sequence anomaly detection based on the pre-trained deep learning model to determine whether the log data is a sequence anomaly log.
[0080] Business processes usually have a logical order, so logs are also printed in a certain order. When a process exception occurs, a large number of out-of-order logs will be generated, disrupting the normal execution path.
[0081] In addition to text attributes, logs also have sequence attributes. In the embodiments of the present invention, the sequence attributes of logs can be used in combination with deep learning models to perform log sequence anomaly detection. The algorithm transforms the log sequence anomaly detection problem into a multi-classification problem, outputs probability distribution, implements anomaly detection through prediction, and identifies execution anomalies in program logic flow.
[0082] Furthermore, when performing root cause analysis on abnormal alarm messages and determining the cause of the abnormality, the following method can be used: first, obtain time series data from the log of the abnormal alarm message, and construct a directed acyclic graph based on the time series data and the causal relationship algorithm; then determine the abnormal cause of the abnormal alarm message based on the directed acyclic graph.
[0083] In specific implementation, anomaly detection is part of the log analysis process. After an anomaly occurs, the operation and maintenance personnel need to understand what caused the system failure, so further root cause (causal relationship) analysis is required.
[0084] Causality is a partial order relationship different from correlation (correlation is usually quantified by correlation coefficient). Two events are positively correlated and do not necessarily have causal relationship. Using correlation as causal relationship will produce a large number of false positives. Checking the timestamps of two events helps to determine causality, but due to NTP time synchronization errors, jitters, and network failures, the timestamps of system logs are not completely reliable for determining causality. Therefore, it is necessary to determine the causal relationship between events without timestamps. In this embodiment, time series data is first extracted from log data according to the preprocessing method, and then the causal relationship algorithm is used to output a directed acyclic graph. The network in the figure can reflect the causal relationship between events, and then assist in analyzing the root cause of the anomaly. In addition, combining system resource usage data with error logs can accurately detect anomalies in large distributed systems.
[0085] In order to analyze the root cause of the fault more accurately, in this embodiment, logs and indicator data from different sources can be combined to perform root cause analysis and locate the root cause of the abnormality from multiple perspectives.
[0086] The present invention also provides a method for performing log analysis and model training based on machine learning, deep learning and natural language processing, and then performing log anomaly detection, as shown below:
[0087] (1) Machine Learning
[0088] Machine learning algorithms are used to analyze logs, which are mainly divided into two methods: classification and clustering.
[0089] Classification algorithms are based on two assumptions: (1) data has labels; (2) normal and abnormal instances are separable in feature space. In log analysis, decision trees and SVMs are the most commonly used classification algorithms. After extracting features from logs, the logs are classified, model parameters are trained, and then used to detect abnormal logs in the system. The accuracy of the classification algorithm is highly dependent on the quality of the labels.
[0090] Since in actual situations, most log data is unlabeled, classification methods are relatively limited, while clustering methods do not have this limitation and are more widely used. Clustering methods first need to calculate the distance between two log texts, and then aggregate similar log texts to obtain several categories. Logs belonging to the same category are considered to have the same pattern, thereby identifying patterns in massive log texts and detecting abnormal logs that are far away from all categories. Through clustering algorithms, log pattern discovery is automatically realized, and a large amount of log original text is converted into a small number of log patterns, greatly reducing the time for manual screening. Machine learning can automatically learn system behavior, assist in fault diagnosis, and has strong interpretability.
[0091] (2) Deep Learning
[0092] In addition to text attributes, logs also have sequence attributes. By using the sequence attributes of logs and combining them with deep learning models, we can mine the contextual information in the log sequence, feedback the abnormalities of system execution, and make it easier to understand the cause of the failure from the perspective of system behavior, thus providing assistance for failure recovery.
[0093] In addition, a generative adversarial network is used to analyze the logs, which mainly consists of two parts: the generator and the discriminator. The generator attempts to capture the data distribution of the real training dataset and synthesize reasonable instances (i.e., normal and abnormal data), while the discriminator aims to distinguish synthetic data from datasets built using real data and synthetic data. Finally, the fully trained generator will detect whether the upcoming logs are normal or abnormal based on the latest events, thereby generating anomaly alerts and effectively helping administrators diagnose workflows. The imbalance problem between normal and abnormal instances can also be alleviated by generating abnormal data.
[0094] (3) Natural speech processing technology
[0095] Logs are a special type of semi-structured text that has some properties of natural language. Therefore, natural language processing (NLP) technology can be used to analyze logs, convert logs into semantic vectors, mine the semantic information in the logs, and combine deep learning models to detect log sequence anomalies, thereby improving detection accuracy.
[0096] In the embodiment of the present invention, machine learning algorithms can also be used to perform intelligent prediction of indicator data according to different strategies. For important data indicators of business performance, such as transaction volume, response rate, response time, success rate and other data, algorithm indicator anomaly detection is performed, and a business indicator anomaly detection and root cause location algorithm engine is constructed. The algorithms implemented include variational autoencoder, progressive gradient regression tree, differential exponential sliding average, extreme value theory, periodic median detection, LightGBM, Monte Carlo search tree, etc. Identify abnormal changes in business indicator trends, discover problem risks early, and shorten fault discovery and recovery time. When business indicators fluctuate abnormally or show signs of degradation, the fault root cause location function is automatically triggered. From a large number of transaction details in the abnormal time period of the faulty business system, anomaly detection is performed after statistics in multiple attribute dimensions, and the candidate root cause set is sorted according to the indicator conversion rate and inclusion relationship, and the abnormal root cause is finally determined. Based on this, the above method provided by an embodiment of the present invention also includes: first, obtaining historical data of the financial system, and extracting business processes and data flows from the historical data to construct a business prediction model; then, predicting the business indicator data based on the business prediction model to obtain prediction results; finally, determining abnormal alarm messages based on the prediction results of the business indicator data.
[0097] In the specific implementation, by continuously collecting historical data and manually analyzing the historical data, we sort out the complete business process and data process to establish a business prediction model, which includes a business model and a prediction algorithm model. Among them, historical data mainly includes business data, performance evaluation data, system architecture information and system monitoring data.
[0098] a. Business data mainly refers to statistics and detailed data related to business operations.
[0099] Such as overall transaction volume, concurrent transaction volume, transaction details, etc., which are mainly business data related to system operation complexity and contain time information. At the same time, business development planning data also has a key impact on the correction of prediction models.
[0100] b. Performance evaluation data mainly refers to the performance and monitoring results data of the system that has been executed.
[0101] Such as maximum transaction capacity (TPS), system response time, maximum concurrency, high-water mark user experience, exceptions and errors, etc. System architecture information, including system architecture, network topology, etc.
[0102] c. System monitoring data includes network resource utilization, server load monitoring, third-party service monitoring and other related data.
[0103] The focus is on monitoring and counting the capacity of the resource pool (mainly including: computing resources, storage resources, and number of containers), and performing trend analysis based on statistical indicators. Monitor and count the capacity of applications and business systems, and perform trend analysis based on statistical indicators.
[0104] Furthermore, through machine learning predictions combined with expert consultation and analysis, we can predict the expected growth of business and systems.
[0105] Specifically, according to the business forecast model, select the appropriate forecast analysis algorithm, import historical data to execute the machine learning algorithm. Based on the execution results of the machine learning algorithm, combined with non-systematic data such as business planning, policy changes and other information, manually correct the business forecast results to complete the business and system capacity forecast report. After the system evaluation and optimization complete the stage goals, it is necessary to continue to collect data, and then regularly correct the business and system capacity forecast results, and perform system capacity evaluation regression as needed.
[0106] Correspondingly, business volume and system capacity forecasts, as data collection continues, also need to be continuously or regularly executed to continuously revise the business volume growth forecast and system capacity demand growth forecast.
[0107] Develop and implement system stress testing and monitoring plans based on system capacity forecast results.
[0108] Machine learning is combined with expert consultation to analyze system stress testing and monitoring data to complete system expansion and optimization recommendations.
[0109] Develop and implement system expansion and optimization plans based on business development, cost estimates and other factors.
[0110] After the system expansion and optimization is completed, perform system stress testing to complete regression verification.
[0111] After the system optimization goal is achieved, a continuous system monitoring and data collection mechanism is established for subsequent iterative optimization of the process.
[0112] In the embodiments of the present invention, a self-developed prediction algorithm can be used to perform single indicator prediction and multi-indicator correlation prediction.
[0113] a. Single indicator prediction: Based on the historical data of a single indicator, predict the indicator changes in the future. For example, if the overall disk space occupancy is predicted to reach 80% within 30 days, expand the capacity; if it is less than 10% within 30 days, recycle it.
[0114] In one embodiment, the method further includes: compressing the abnormal alarm message into an alarm according to the alarm source type of the abnormal alarm message according to a preset compression rule; and merging the alarm into an abnormal alarm event according to a preset merging rule.
[0115] In specific implementation, the default merge rules and the merge rules customized by the operation and maintenance personnel can be uniformly displayed and managed, and frequent item set mining, intelligent alarm and other algorithms can be used for compression. Under certain conditions, the same type of alarms can be compressed and merged into one alarm information to reduce the number of alarm data presentations. In the embodiment of the present invention, custom merge rules can be added, edited, enabled / suspended, deleted and viewed, and default merge rules can be viewed, edited, enabled / suspended.
[0116] When the same alarm occurs repeatedly, it will be sent to the real-time event library repeatedly. When the event library receives these repeated alarms, it needs to judge these alarms, compress the same alarms into one, and only update the number of repetitions, the last occurrence time and the alarm description. When storing, it is also stored as one alarm.
[0117] Specifically, the alarm suppression mechanism is divided into two steps: alarm compression and alarm merging. Figure 2 As shown, the alarm compression rule is a background preset rule, which will match the compression rule according to the alarm source type, and compress the alarm message into an alarm that meets the compression rule. Then, the alarm is merged into an alarm event according to the merging rule.
[0118] On the one hand, the alarm compression link can compress conditions such as business system, host IP, alarm level, alarm keyword, suppression cycle and event number. On the other hand, the rules of the alarm merging link are divided into default merging rules and custom merging rules. After successfully accessing the alarm source, the default merging rules will be automatically created for each alarm source. Custom merging rules need to be customized by operation and maintenance personnel.
[0119] Furthermore, in the embodiment of the present invention, the enrichment of alarms can also prepare for the subsequent analysis of alarm events, and use auxiliary information to determine how to handle, analyze and notify. Alarm enrichment is generally achieved through rules, linkage of CMDB, knowledge base, operation history records and other data sources to enrich alarm fields and related information.
[0120] There are three approaches to alarm convergence: suppression, shielding, and aggregation. Suppression means suppressing the same problem to avoid repeated alarms. Common suppression schemes include anti-shake suppression, dependency suppression, time suppression, combined condition suppression, high availability suppression, etc. Shielding means shielding predictable situations, such as changing maintenance periods and fixed periodic tasks. Aggregation is to merge similar or identical alarms. For example, if the business access volume increases, the CPU, memory, disk IO, network IO and other performance of the host carrying the business will increase. Aggregating these performance indicators together will make it easier to analyze and process alarms.
[0121] In the embodiment of the present invention, after receiving the alarm, an alarm notification is also required, and the person is notified through some conventional notification channels, and the staff can be triggered through WeChat, SMS, and email. Push to a third-party system through an API to facilitate subsequent event processing. In addition, it is also necessary to support custom channel expansion (for example, the IM system in the enterprise can be accessed by itself).
[0122] The above method provided by the embodiment of the present invention has the following beneficial effects compared with the prior art:
[0123] (1) Enhanced system stability: Real-time monitoring and intelligent analysis can promptly detect potential risks, effectively prevent system failures, and improve the overall stability and reliability of the system. It realizes the automated execution and management of operation and maintenance tasks, reducing the workload of operation and maintenance personnel and the risk of human error.
[0124] (2) Improved operation and maintenance efficiency: Intelligent algorithms are used to realize automated monitoring, early warning, alarm, and troubleshooting, allowing operation and maintenance personnel to more easily grasp the overall situation of the system and improve work efficiency. Operation and maintenance observability and process management significantly reduce the workload of operation and maintenance personnel and improve the response speed and efficiency of operation and maintenance.
[0125] (3) Optimize resource allocation: Through intelligent analysis and prediction of system resource requirements, resources can be rationally allocated and optimally utilized, thereby reducing operating costs.
[0126] (4) Improve decision accuracy: By introducing intelligent algorithms, the automatic analysis and processing of operation and maintenance data is realized, which significantly improves the efficiency and accuracy of operation and maintenance. At the same time, it provides a rich visualization and reporting function, providing operation and maintenance personnel with intuitive and comprehensive system status information, allowing them to understand the overall situation and key indicators of the system more quickly, which helps to formulate more scientific operation and maintenance strategies and decisions.
[0127] For the operation and maintenance monitoring management method provided in the above embodiment, the present invention also provides an operation and maintenance monitoring management device, see Figure 3 The schematic diagram of the structure of an operation and maintenance monitoring and management device shown in the figure shows that the device mainly includes the following parts:
[0128] Data acquisition module 301, used to acquire performance data and monitoring data of the financial system;
[0129] Anomaly detection module 302, used to input performance data and monitoring data into a pre-trained anomaly detection model to obtain anomaly alarm messages;
[0130] The abnormality analysis module 303 is used to determine the abnormality location and abnormality cause based on the abnormality alarm message and monitoring data.
[0131] The above-mentioned operation and maintenance monitoring and management device provided by the present invention can obtain abnormal alarm messages based on the system's performance data and monitoring data through a pre-trained abnormality detection model, and further determine the abnormal location and cause of the abnormality, so as to timely discover and resolve potential risks, prevent system failures, and improve operation and maintenance efficiency and system stability.
[0132] In one embodiment, the above device further comprises:
[0133] Log management and analysis module: This module can uniformly collect, process, store, query and analyze all logs of the financial system (business systems, cloud resources, servers, network equipment, security equipment, databases, middleware, etc.), including log retrieval, log monitoring and alarm, log pattern recognition and anomaly detection, log desensitization, log correlation analysis, full-link tracking, log visualization analysis, etc., and supports log viewing and alarm according to different user permissions.
[0134] Basic monitoring module: This module can realize comprehensive and in-depth monitoring of the operating system, database, middleware, hardware network equipment and other resources of important business systems within financial institutions, ensuring the continuous and stable operation of the network and IT business systems.
[0135] Three-dimensional monitoring module: This module can integrate existing monitoring tools (basic monitoring, NPM, BPC, cloud platform monitoring, etc.) within financial institutions, build unified monitoring capabilities, and provide system health assessment and display. Through multi-dimensional monitoring and analysis of indicators, alarm events, logs, business call relationships, resource dependencies, etc., it enriches monitoring and fault analysis paths, ensures continuous and stable business operation, and improves operation and maintenance efficiency.
[0136] In an embodiment of the present invention, based on the native functions of a Zabbix open source monitoring system, and in combination with Shell scripts and database SQL statements, the availability and health of the core business systems of financial institutions are monitored and displayed, helping users to timely understand the key indicators of the business system and facilitating the location and handling of problems.
[0137] Unified event management module: This module can realize unified access and processing of alarm messages and data indicators from various monitoring systems, filter, notify, respond, handle, classify, track and multi-dimensionally analyze alarm events, and use a variety of algorithms to achieve convergence, noise reduction, anomaly detection and root cause analysis of alarm events, and realize global control of the entire life cycle of problem events, thereby freeing operation and maintenance personnel from massive repeated alarms and having more energy to ensure and handle production events.
[0138] Operation and maintenance process management module: This module can build reasonable event handling specifications, clarify the process direction from aspects such as alarm / event management and duty management, and provide a comprehensive, closed-loop service model for operation and maintenance work.
[0139] Intelligent algorithm management module: This module can provide unified algorithm management, scenario-based algorithm configuration and other functions based on the mainstream TensorFlow framework and artificial intelligence algorithms. It supports unified access to intelligent algorithms, intelligent data analysis, model experimental training and tuning, and supports the release and application of generic algorithms. It has high availability and high concurrency performance, and provides powerful algorithm capabilities for upper-level businesses and products.
[0140] Report management module: This module can provide a variety of analysis models based on different business scenarios to meet users' different analysis goals. It can generate various reports according to various conditions to meet various needs such as application system management, alarm statistics, business statistics, and technology risk audit assessment.
[0141] Visual display module: Based on the characteristics of adaptive management, this module can provide multi-level display, integrate and analyze IT and business application related data from different perspectives, build a visual analysis command center, and improve data application and data decision-making capabilities.
[0142] The device uses the performance data collected by the Zabbix monitoring system, the event data of the centralized event management platform, and the data of the business system as the data source, deeply customizes the dynamic large-screen display, organically integrates the operation and maintenance monitoring data, event data, and business data, and displays them dynamically and in real time.
[0143] It should be noted that the implementation principle and technical effects of the device provided in the embodiment of the present invention are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned method embodiment.
[0144] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above implementation methods.
[0145] Figure 4 A structural diagram of an electronic device provided in an embodiment of the present invention, the electronic device 100 includes: a processor 40, a memory 41, a bus 42 and a communication interface 43, wherein the processor 40, the communication interface 43 and the memory 41 are connected via the bus 42; the processor 40 is used to execute an executable module stored in the memory 41, such as a computer program.
[0146] The memory 41 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 43 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0147] The bus 42 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0148] Among them, the memory 41 is used to store programs, and the processor 40 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 40 or implemented by the processor 40.
[0149] The processor 40 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 40. The above processor 40 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present invention can be directly embodied as a hardware decoding processor to execute, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 41, and the processor 40 reads the information in the memory 41 and completes the steps of the above method in combination with its hardware.
[0150] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be referred to the previous method embodiments, which will not be repeated here.
[0151] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0152] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. An operation and maintenance monitoring management method, characterized in that: include: Obtain performance and monitoring data for financial systems; Inputting the performance data and the monitoring data into a pre-trained anomaly detection model to obtain an anomaly alarm message; The abnormal location and the abnormal cause are determined based on the abnormal alarm message and the monitoring data.
2. The method according to claim 1, characterized in that The performance data and the monitoring data are input into a pre-trained anomaly detection model to obtain an anomaly alarm message, including: Using pattern recognition to identify the data types of the performance data and the monitoring data to obtain time series periodic data; The time series periodic data is input into a pre-trained anomaly detection model to obtain an abnormal alarm message; wherein the anomaly detection model is trained using a machine learning model, including a single-indicator anomaly detection model and a multi-indicator anomaly detection model.
3. The method according to claim 1, characterized in that: The monitoring data also includes log data; the method also includes: Parsing the log data based on a log parsing algorithm to obtain a log mode of the log data; Perform log anomaly detection based on the log mode of the log data to obtain an abnormal alarm message, and perform root cause analysis on the abnormal alarm message to determine the cause of the abnormality.
4. The method according to claim 3, characterized in that Performing log anomaly detection based on the log mode of the log data includes: Compare the log mode of the log data with a preset normal log mode to determine whether the log data is an abnormal mode log; Determining whether the log data is a log with abnormal quantity based on the relationship between the log quantities of different log modes; Log sequence anomaly detection is performed based on a pre-trained deep learning model to determine whether the log data is a sequence anomaly log.
5. The method according to claim 3, characterized in that: Perform root cause analysis on the abnormal alarm message to determine the cause of the abnormality, including: Acquire time series data from the log of the abnormal alarm message, and construct a directed acyclic graph based on the time series data and a causal relationship algorithm; The abnormal cause of the abnormal alarm message is determined based on the directed acyclic graph.
6. The method according to claim 1, characterized in that The monitoring data also includes business indicator data; the method also includes: Acquire historical data of the financial system, extract business processes and data processes from the historical data, and build a business prediction model; Predicting the business indicator data based on the business prediction model to obtain a prediction result; An abnormal alarm message is determined based on the prediction result of the business indicator data.
7. The method according to claim 1, characterized in that Also includes: According to the alarm source type of the abnormal alarm message, and according to a preset compression rule, compressing the abnormal alarm message into an alarm; The alarms are merged into abnormal alarm events according to preset merging rules.
8. An operation and maintenance monitoring and management device, characterized in that: include: A data acquisition module, used to obtain performance data and monitoring data of the financial system; An anomaly detection module, used for inputting the performance data and the monitoring data into a pre-trained anomaly detection model to obtain an anomaly alarm message; The abnormality analysis module is used to determine the abnormality location and abnormality cause based on the abnormality alarm message and the monitoring data.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.
Citation Information
Cited By
Network equipment early warning method, electronic equipment, storage medium and program product
CN121690976A