Log anomaly detection method and device, storage medium and electronic equipment
By converting bank system log data into a time series matrix and using a long-term extreme learning machine model for event prediction, the problems of false positives and missed negatives in traditional methods are solved, and efficient and low-resource log anomaly detection is achieved, which is suitable for real-time security monitoring in the financial industry.
Patent Information
- Application Number
- CN202511023942.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional log anomaly detection methods are difficult to avoid false positives or omissions when faced with the dynamic changes in business requests in banking systems. Existing deep learning solutions rely on massive amounts of labeled data and high computing resources, with long training cycles and high deployment costs, and cannot meet the financial industry's needs for real-time and efficiency.
The original log data of the banking system is converted into a time series matrix, and dynamic weighting and dimensionality reduction processing are performed. Event prediction is performed using a long-term extreme learning machine model. By comparing the predicted results with the actual results, it is determined whether there are any anomalies in the log data.
It achieves efficient log anomaly detection, reduces dependence on massive labeled data and high computing resources, improves the accuracy and timeliness of detection, and meets the financial industry's needs for security and stability.
Smart Images

Figure CN120804995A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log data processing, and in particular to a log anomaly detection method, device, storage medium and electronic equipment. Background Art
[0002] In the field of financial technology, log anomaly detection in banking systems is becoming increasingly important as a key component in ensuring stable business operations and risk prevention. With the widespread adoption of distributed architectures and the rapid expansion of high-frequency trading scenarios, the volume of bank log data is increasing exponentially. Large banking systems can generate over 50GB of logs per hour, with daily log volumes reaching tens of millions. Furthermore, call chain structures are becoming increasingly complex, and time series characteristics are exhibiting significant dynamic fluctuations.
[0003] However, traditional methods based on fixed thresholds or rules struggle to cope with the dynamic nature of banking system business requests. For example, in scenarios like sudden increases in transaction volume or abnormal cross-system calls, static thresholds can easily lead to false positives or false negatives. Time series-based detection techniques place strict demands on data stability, while bank logs show significant fluctuations in call duration and business response times, leading to inadequate extraction of time-series correlation features. Existing deep learning solutions rely on massive amounts of labeled data and high computing power, resulting in lengthy training cycles and high deployment costs.
[0004] Therefore, how to achieve efficient log anomaly detection has become a technical problem that technical personnel in this field urgently need to solve. Summary of the Invention
[0005] In view of the above problems, the present invention provides a log anomaly detection method, device, storage medium, and electronic device that overcome the above problems or at least partially solve the above problems. The technical solution is as follows:
[0006] A log anomaly detection method, comprising:
[0007] Obtaining original log data of a designated bank system, wherein the original log data includes log events automatically generated during operation of the designated bank system;
[0008] Converting the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data;
[0009] Performing dynamic weighting and dimensionality reduction processing on the time series matrix to obtain log feature information;
[0010] inputting the log feature information into a long-time extreme learning machine model constructed in advance to obtain an event prediction result output by the long-time extreme learning machine model, wherein the event prediction result is a log event result generated by the specified bank system in a current time window predicted by the long-time extreme learning machine model;
[0011] obtaining an event real result of the specified bank system, wherein the event real result is a log event result actually generated by the specified bank system in the current time window;
[0012] using the event prediction result and the event real result to determine whether there is abnormal log data in the original log data, and generating an abnormality determination result.
[0013] Optionally, after the step of using the event prediction result and the event real result to determine whether there is abnormal log data in the original log data, and generating an abnormality determination result, the method further comprises:
[0014] updating model parameters of the long-time extreme learning machine model using a plurality of abnormality determination results in a first time window to obtain an updated long-time extreme learning machine model.
[0015] Optionally, the step of converting the original log data into a time series matrix according to event types and event frequencies of each log event in the original log data comprises:
[0016] clustering the original log data into event templates according to event types by using a log clustering algorithm;
[0017] separating static parameters and dynamic parameters in the event templates to generate an event encoding vector;
[0018] based on the event encoding vector, counting event frequencies according to a second time window to construct a time series matrix.
[0019] Optionally, the step of performing dynamic weighting and dimension reduction processing on the time series matrix to obtain log feature information comprises:
[0020] generating a dynamic weighted feature vector using the time series matrix and a preset time decay weight;
[0021] performing dimension reduction compression on the dynamic weighted feature vector by using a principal component analysis algorithm to obtain log feature information.
[0022] Optionally, the step of inputting the log feature information into a long-time extreme learning machine model constructed in advance to obtain an event prediction result output by the long-time extreme learning machine model comprises:
[0023] inputting the log feature information into a long-time extreme learning machine model constructed in advance, so that the log feature information is transmitted to a dynamic subnet and a static subnet of the long-time extreme learning machine model, wherein the dynamic subnet comprises a long-term memory unit; obtaining long-term dynamic information output by the long-term memory unit and static information output by the static subnet; and generating an event prediction result and outputting the event prediction result by using the long-term dynamic information and the static information.
[0024] Optionally, the log event result comprises a log event frequency distribution result, and the determining whether the original log data comprises abnormal log data by using the event prediction result and the event real result, and generating an abnormality determination result comprises:
[0025] determining a first confidence interval based on a residual distribution of the log event frequency distribution result in the event prediction result and the event real result;
[0026] calculating a KL divergence of the log event frequency distribution result in the event prediction result and the event real result under the current time window;
[0027] generating the abnormality determination result that the original log data comprises the abnormal log data in a case where the KL divergence exceeds the first confidence interval.
[0028] Optionally, the log event result comprises a log event sequence result, and the determining whether the original log data comprises abnormal log data by using the event prediction result and the event real result, and generating an abnormality determination result comprises:
[0029] determining a second confidence interval based on a residual distribution of the log event sequence result in the event prediction result and the event real result;
[0030] calculating a matching degree of the log event sequence result in the event real result and a specified historical event sequence result under the current time window;
[0031] generating the abnormality determination result that the original log data comprises the abnormal log data in a case where the matching degree exceeds the second confidence interval.
[0032] A log abnormality detection device comprises an original log data obtaining unit, a time series matrix conversion unit, a log feature information obtaining unit, an event prediction result obtaining unit, an event real result obtaining unit and an abnormality determination result generating unit,
[0033] the original log data obtaining unit is configured to obtain original log data of a specified bank system, wherein the original log data comprises log events automatically generated by the specified bank system during running;
[0034] The time series matrix conversion unit is configured to convert the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data.
[0035] The log feature information obtaining unit is configured to perform dynamic weighting and dimension reduction processing on the time series matrix to obtain log feature information.
[0036] The event prediction result obtaining unit is configured to input the log feature information into a long-time extreme learning machine model constructed in advance to obtain an event prediction result output by the long-time extreme learning machine model, wherein the event prediction result is a log event result predicted by the long-time extreme learning machine model for the specified bank system in a current time window.
[0037] The event real result obtaining unit is configured to obtain an event real result of the specified bank system, wherein the event real result is a log event result actually generated by the specified bank system in the current time window.
[0038] The abnormality determination result generating unit is configured to determine whether there is abnormal log data in the original log data by using the event prediction result and the event real result, and generate an abnormality determination result.
[0039] A computer readable storage medium having a program stored thereon, the program being executed by a processor to implement the log abnormality detection method.
[0040] An electronic device, the electronic device comprising at least one processor, and at least one memory connected with the processor through a bus; wherein the processor, the memory complete mutual communication through the bus; the processor is used to call the program instruction in the memory, to execute the log abnormality detection method.
[0041] By the above technical scheme, the log anomaly detection method, device, storage medium and electronic equipment provided by the application obtain original log data of a specified bank system, wherein the original log data includes log events automatically generated by the specified bank system in a running process; the original log data is converted into a time series matrix according to the event type and event frequency of each log event in the original log data; the time series matrix is dynamically weighted and dimensionally reduced to obtain log feature information; the log feature information is input into a long-time extreme learning machine model constructed in advance to obtain an event prediction result output by the long-time extreme learning machine model, wherein the event prediction result is a log event result predicted by the long-time extreme learning machine model for the specified bank system in a current time window; an event real result of the specified bank system is obtained, wherein the event real result is a log event result actually generated by the specified bank system in the current time window; whether there is abnormal log data in the original log data is judged by using the event prediction result and the event real result, and an abnormality judgment result is generated. The original log data is converted into a time series matrix and dynamically weighted and dimensionally reduced to extract features, the long-time extreme learning machine model is used for event prediction, and efficient log anomaly detection capability can be provided, thereby meeting the needs of the financial industry for safety and stability.
[0042] The above description is only a summary of the technical scheme of the application, in order to more clearly understand the technical means of the application, the content of the specification can be implemented, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0043] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not meant to limit the present application. Furthermore, the same reference numerals are used throughout the several drawings to designate the same or similar parts. In the drawings:
[0044] Figure 1 A flowchart of an embodiment of the log anomaly detection method provided by the application is shown;
[0045] Figure 2 A flowchart of a first specific embodiment of the log anomaly detection method provided by the application is shown;
[0046] Figure 3 A flowchart of a second specific embodiment of the log anomaly detection method provided by the application is shown;
[0047] Figure 4 A structural diagram of the log anomaly detection device provided by the application is shown;
[0048] Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0049] Exemplary embodiments of the present application will be described in detail with reference to the drawings. Although exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present application can be more thoroughly understood, and so that the scope of the present application can be completely conveyed to those skilled in the art.
[0050] In the field of financial technology, log anomaly detection of bank systems is a key link to ensure business continuity and prevent risks. With the widespread application of distributed architecture and the rapid development of high-frequency trading scenarios, the scale of bank log data has increased dramatically, leading to unprecedented challenges in log processing and analysis. Specifically, modern bank systems can generate more than 50 GB of log data per hour, with single-day log volume reaching tens of millions, which makes traditional log processing methods seem inadequate.
[0051] The complexity of bank logs not only lies in the surge in data volume, but also includes the complexity of call links and the dynamic changes of time series features. Traditional log anomaly detection methods rely on fixed thresholds or rules, which show obvious shortcomings when facing dynamic changes in business requests of bank systems, especially in the case of sudden increase in transaction volume or abnormal cross-system calls. Static thresholds often lead to false positives or false negatives, affecting the timely response to real anomalies.
[0052] At the same time, time series-based detection techniques usually require data to be stationary, while the call duration and business response time of bank logs have significant fluctuations, making traditional methods insufficient in extracting time sequence correlation features. In addition, although existing deep learning solutions can improve detection accuracy, they rely on massive labeled data and high computing resources, with long training cycles and high deployment costs. Studies have shown that when processing tens of millions of logs, the model iteration of traditional neural networks may take more than 48 hours, which contradicts the strict real-time requirements of bank businesses.
[0053] Furthermore, the poor interpretability of deep learning models makes it difficult for operations personnel to trace the root causes of anomalies. In banking systems, abnormal log patterns often exhibit long-term dependencies across time dimensions, such as cascading failures caused by distributed transaction timeouts. Traditional clustering methods and shallow learning models often struggle to model these long-term features. While anomaly detection based on text clustering can reduce reliance on regular expressions, it lacks sensitivity to low-frequency anomalies and cannot effectively quantify pattern shifts within a time window.
[0054] Based on this, an embodiment of the present invention provides a log anomaly detection method. First, the original log data of the specified bank system is obtained and converted into a time series matrix according to the event type and frequency. Subsequently, the matrix is dynamically weighted and dimensionally reduced to extract log feature information. These features are input into the long-term extreme learning machine model to obtain the event prediction results under the current time window. At the same time, the actual event results of the system in the same time window are obtained. By comparing the predicted results with the actual results, it is determined whether there are abnormal logs in the original log data, thereby generating an anomaly judgment result. It can be seen that this solution uses the long-term extreme learning machine model for event prediction, which makes it more efficient in resource consumption and avoids dependence on massive labeled data and high computing power resources, thereby achieving efficient log anomaly detection with low resource consumption.
[0055] like Figure 1 As shown, a flow chart of an implementation of a log anomaly detection method provided by an embodiment of the present invention is provided. The method may include:
[0056] S100: Obtain original log data of a designated bank system, wherein the original log data includes log events automatically generated during operation of the designated bank system.
[0057] The designated bank system may be a specific information management system for processing the bank's daily operations, transaction records, and user information, and may automatically generate various log events generated during operation.
[0058] Among them, raw log data refers to the records automatically generated by the designated banking system during operation. It can contain a variety of information such as system activities, user operations and error messages. It is usually stored in text form and contains information such as timestamp, event type, event source and event details.
[0059] Among them, a log event refers to a single record item in the original log data, which specifically describes a specific operation or status occurring in a specified banking system at a certain moment, including information such as timestamp, event type, event source, and event details.
[0060] It can be understood that the log management module in the designated bank system automatically records and stores various log events generated during system operation, thereby periodically recording each operation of the user, the response of the system, and potential error information. The original log data is usually stored in an unstructured format. The embodiment of the present application can call the corresponding interface of the designated bank system or directly read the log file, thereby extracting the original log data of the designated bank system, thereby providing basic data support for subsequent analysis and anomaly detection.
[0061] S110, converting the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data.
[0062] The event type refers to the classification of different operations or states recorded in the original log data. For example, the event type can include user login, transaction execution, system error, and warning notification. Each event type has its specific meaning and influence, facilitating identification and classification.
[0063] The event frequency refers to the number of times a specific event type occurs within a certain time window. The event frequency can be used to analyze the regularity and abnormal fluctuations of system activity, helping to discover potential problems or security threats.
[0064] The time series matrix is a multi-dimensional matrix formed by taking time as the independent variable and event type and its corresponding frequency as the dependent variable, which is a structured data representation form. Each row in the time series matrix can represent a time point, and each column can represent a specific event type and its occurrence frequency.
[0065] Specifically, the embodiment of the present application can first extract the event type and timestamp of each log event from the original log data. Then, according to the specified time window (such as every hour, every day, etc.), the events are grouped and counted, and the frequency of each event type within the time window is calculated. Then, these frequencies are arranged into a matrix structure, with rows representing time windows and columns representing different event types. Finally, each element in the matrix is the number of occurrences of the corresponding event type within the time window, forming a complete time series matrix.
[0066] S120, dynamically weighting and dimensionality reduction processing the time series matrix to obtain log feature information.
[0067] The log feature information refers to the numerical values or indicators extracted from the time series matrix that can represent the important features of the log data. The log feature information can include event frequency, event type change trend, and anomaly detection indicators, etc., aiming to reveal the behavior pattern and potential anomalies of the system.
[0068] Dynamic weighting is a method of weighting various features in a time series matrix according to specific rules or conditions. Dynamic weighting can help highlight the importance of certain log features within a specific time period, thereby more effectively capturing system changes.
[0069] Dimensionality reduction involves using compression algorithms to reduce the number of features in a time series matrix, retaining the most important information while removing redundant or irrelevant data. Dimensionality reduction helps improve the model's ability to detect log anomalies.
[0070] Specifically, embodiments of the present invention assign weights to each event type based on a predetermined weighting rule. These weights are then applied to the various data types in the time series matrix to generate a weighted time series matrix. Dimensionality reduction techniques are then used to process the weighted time series matrix to extract the most representative log feature information. This dimensionality reduction process effectively reduces the number of features while retaining key information, thereby providing concise and effective input data for subsequent long-term extreme learning machine model training and anomaly detection.
[0071] S130: Input the log feature information into a pre-built long-term extreme learning machine model to obtain an event prediction result output by the long-term extreme learning machine model, wherein the event prediction result is a result of the long-term extreme learning machine model predicting a log event generated by the specified bank system in the current time window.
[0072] The Long-Term Extreme Learning Machine (L-ELM) model is an improved Extreme Learning Machine (ELM) algorithm that specializes in processing time series data and features fast training and good generalization capabilities. By introducing a long-term memory mechanism, the L-ELM can capture long-term dependencies in time series data, thereby improving the accuracy of event predictions.
[0073] Among them, the event prediction result refers to the prediction output generated by the long-term extreme learning machine model based on the input log feature information, which reflects the long-term extreme learning machine model's expectations of future log events in the current time window, including the frequency and sequence of possible log events.
[0074] Log event results refer to all log events generated by a specified banking system within a specific time window, including specific operation records and related information. Log event results include log event frequency distribution results and / or log event sequence results.
[0075] The log event frequency distribution result is a statistical result of the occurrence frequency of various log events in a specific time period, and can reveal the running state and normal mode of the system and be used for anomaly detection.
[0076] The log event sequence result is a log event record arranged in chronological order, provides the time relationship and change trend between events, helps to analyze the dynamic characteristics of log data in depth, and is used for anomaly detection.
[0077] Specifically, the log feature information is input into the long-time extreme learning machine model constructed in advance, and the input log feature information is processed by using the long-time memory mechanism of the neural network structure of the long-time extreme learning machine model, so as to generate the event prediction result.
[0078] S140, obtaining an event real result of the specified bank system, wherein the event real result is a log event result actually generated by the specified bank system in the current time window.
[0079] The event real result refers to the log event record actually generated by the specified bank system in the current time window, which is used to evaluate the accuracy of the long-time extreme learning machine model prediction and provides real benchmark data, including the actual occurrence frequency and sequence of the log event.
[0080] Specifically, the log event record generated in the current time window can be extracted from the actually running specified bank system. These log event records include all related system operations and are analyzed against the prediction result of the model.
[0081] S150, using the event prediction result and the event real result to determine whether there is abnormal log data in the original log data, and generating an abnormality determination result.
[0082] The abnormal log data refers to the log data that deviates significantly from the normal running mode. The abnormal log data can indicate system failure, abnormal activity or security event.
[0083] The abnormality determination result is a judgment on whether the log data is abnormal based on the comparison between the event prediction result and the event real result. The abnormality determination result can be used to guide subsequent system monitoring and maintenance.
[0084] Specifically, the embodiment of the present application compares the event prediction result output by the model with the actually obtained event real result, analyzes the difference between the two. If the deviation between the event prediction result and the event real result exceeds the preset confidence interval, it can be determined that there is abnormal log data in the original log data. The generated abnormality determination result will provide an important reference for subsequent system maintenance and security monitoring. At the same time, the construction of the confidence interval can further increase the reliability of the abnormality detection, and ensure the accuracy of the abnormality determination result.
[0085] The present application provides a log abnormality detection method, which comprises: obtaining original log data of a specified bank system, wherein the original log data comprises log events automatically generated by the specified bank system during operation; converting the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data; performing dynamic weighting and dimensionality reduction processing on the time series matrix to obtain log feature information; inputting the log feature information into a pre-constructed long-time extreme learning machine model to obtain an event prediction result output by the long-time extreme learning machine model, wherein the event prediction result is a log event result predicted by the long-time extreme learning machine model for the specified bank system under a current time window; obtaining an event real result of the specified bank system, wherein the event real result is a log event result actually generated by the specified bank system under the current time window; using the event prediction result and the event real result to determine whether there is abnormal log data in the original log data, and generating an abnormality determination result. The present application can provide efficient log abnormality detection capability by converting the original log data into a time series matrix and dynamically weighting and dimensionality reduction to extract features, and using a long-time extreme learning machine model for event prediction, thereby meeting the needs of the financial industry for security and stability.
[0086] Optionally, based on Figure 1 As shown in the method, Figure 2 As shown in the method,
[0087] S200, using a plurality of abnormality determination results in the first time window to update the model parameters of the long-time extreme learning machine model, and obtaining an updated long-time extreme learning machine model.
[0088] The model parameters refer to the values learned by the long-time extreme learning machine model during training, which determine the behavior and output of the model. The model parameters can include the weights, biases and other adjustable hyperparameters of the long-time extreme learning machine model. These parameters can be optimized through training samples, so that the model can more accurately predict and classify.
[0089] Specifically, embodiments of the present invention can collect multiple anomaly determination results within a first time window. These anomaly determination results are then used to adjust the relevant parameters of the long-term extreme learning machine model through algorithms such as backpropagation to better adapt it to the current log data characteristics and anomaly patterns. During the update process, a specific learning rate is used to control the amplitude of parameter adjustments, and regularization methods may be applied to prevent overfitting. Ultimately, after multiple iterative optimizations, an updated long-term extreme learning machine model is obtained, enabling more accurate log event prediction and anomaly detection in subsequent time windows.
[0090] Furthermore, to improve the performance of long-term extreme learning machine models in anomaly detection, a dynamic update strategy can be adopted. For example, a fixed time window (such as 24 hours) is set, and only the latest log data is retained within this window. Expired data is automatically discarded. To ensure the timeliness of training data, a ring buffer is used to store log data within the window, effectively preventing old data from interfering with the model. After obtaining multiple anomaly determination results within the first time window, the recursive least squares (RLS) method is used based on new batches of data to dynamically adjust the weights from the hidden layer to the output layer. This incremental parameter update method eliminates the need to retrain the entire model, thereby improving training efficiency. Furthermore, the error distribution of the model prediction within the window can be calculated and compared with the historical distribution. If the difference between the two is too large, it is determined that concept drift has occurred. If drift is detected, the weights from the hidden layer to the output layer are fully updated based on the data in the sliding window, ensuring that the model always adapts to the current data characteristics and improving the accuracy and reliability of anomaly detection.
[0091] The embodiments of the present invention improve the LELM model's adaptability to new log formats and anomaly patterns by allowing it to dynamically adjust based on anomalies encountered during actual operation. Secondly, the updated LELM model can more accurately reflect the current operating status and potential anomalies of the banking system, enhancing the accuracy and timeliness of anomaly detection. Finally, by continuously updating the model parameters, online learning can be achieved, enabling the system to maintain efficient monitoring capabilities over extended periods of use, promptly responding to new security threats and anomalies, and ensuring the stability and security of the banking system.
[0092] Optional, based on Figure 1 The method shown, such as Figure 3 As shown, a flowchart of a second specific implementation of the log anomaly detection method provided by an embodiment of the present invention is shown, and step S110 may include:
[0093] S300, adopting a log clustering algorithm to cluster the original log data into event templates according to event types.
[0094] The event template refers to a unified log record format formed by analyzing the original log data through the log clustering algorithm. The event template captures the common characteristics in the log, so that similar log events can be classified into the same type, facilitating subsequent processing and analysis.
[0095] The log clustering algorithm is used to automatically group a large amount of log data according to its characteristics or similarity, forming several event templates. Through clustering, the collection of similar log events can be identified, facilitating subsequent processing and analysis. Optionally, the log clustering algorithm provided by the embodiment of the present application can be Drain algorithm.
[0096] Specifically, the embodiment of the present application can input the original log data into the Drain algorithm, analyze the characteristics in the log by the Drain algorithm, and cluster the log according to the event type. After clustering, similar log events will be classified into the same event template, thereby forming a standardized log format. In the process of converting the original log data of the specified bank system into a time series matrix, the embodiment of the present application classifies the log events by using the log clustering algorithm, which can effectively identify event templates with similar characteristics, not only reducing the redundancy of data, but also making subsequent analysis more efficient and accurate.
[0097] S310, separating static parameters and dynamic parameters in the event template to generate an event encoding vector.
[0098] The static parameter refers to the attribute that remains unchanged when the event occurs. For example: error type and event category.
[0099] The dynamic parameter refers to the attribute that may change when the event occurs. For example: timestamp, IP address and user ID. Dynamic parameters provide context information about the event, reflecting the timeliness and environment of the event.
[0100] The event encoding vector refers to the vector representation generated by encoding the static parameters and dynamic parameters in the event template. The event encoding vector converts log information into numerical form, so that it can be input as a feature into a machine learning model, thereby realizing the standardized processing of log data.
[0101] Specifically, the embodiment of the present application can parse the generated event template, identify and separate static parameters and dynamic parameters. The static parameters (such as error types) are fixed, while the dynamic parameters (such as timestamps, IP addresses) are variable. After encoding the static parameters and the dynamic parameters, an event encoding vector is combined. The event encoding vector represents the characteristics of the log event in numerical form, facilitating subsequent model input. The embodiment of the present application separates the static parameters and the dynamic parameters in the event template, and can generate structured event encoding vectors, which provide a standardized input format for subsequent statistical analysis.
[0102] S320, based on the event encoding vector, the event frequency is counted according to the second time window, and a time series matrix is constructed.
[0103] Specifically, the embodiment of the present application can use the generated event encoding vector to count the frequency of events according to a set time window (such as 5 minutes). By analyzing the event encoding vector in each time window, the statistical results are organized into a time series matrix, which reflects the trend and pattern of events over time, for subsequent model input.
[0104] The embodiment of the present application counts the frequency of each event according to the set time window based on the event encoding vector, and the time series matrix constructed can clearly present the trend of log events over time. This structured data representation provides high-quality feature input for the subsequent long-time extreme learning machine model, significantly improves the accuracy and efficiency of event prediction, and thus more effectively supports the detection of abnormal logs and the security guarantee of the system.
[0105] Optionally, in the above Figure 1 Based on one or more embodiments, in another optional embodiment provided by the embodiment of the present application, step S120 can include:
[0106] A dynamic weighted feature vector is generated using the time series matrix and a preset time decay weight. Principal component analysis algorithm is used to reduce dimension compression of the dynamic weighted feature vector, and log feature information is obtained.
[0107] The preset time decay weight refers to a weight value calculated based on the difference (i.e. time interval Δt) between the event occurrence time and the current time when analyzing the time series matrix, so that the recent events have more significant influence on the current state, and the interference of obsolete data is reduced. The preset time decay weight can be calculated by an exponential decay function, and the formula is , wherein is an adjustable decay coefficient. This weight determines the contribution degree of historical events to the current analysis.
[0108] wherein the dynamically weighted feature vector is a new feature representation generated by multiplying the event encoding vector in the time series matrix with the time decay weight. The specific generation process is to perform weighted calculation on each element (such as a log template ID and an error code) with the corresponding time decay weight, so as to obtain a feature vector that comprehensively considers the time factor.
[0109] wherein the principal component analysis algorithm (PCA) projects the data into a new coordinate system through linear transformation, thereby reducing the dimension of the data while retaining the main features of the data. The principal component analysis algorithm is used to extract principal components by calculating the covariance matrix of the data, select the first few principal components as features, finally compress the high-dimensional feature vector, reduce the computational complexity, and improve the efficiency and accuracy of the model.
[0110] Specifically, the embodiment of the present application can calculate the interval Δt of the timestamp of each log event and the current time, and obtain the corresponding time decay weight through an exponential decay function . Next, each original event encoding vector (such as a log template ID and an error code) is multiplied element by element with the calculated time decay weight , thereby generating a dynamically weighted feature vector : so that recent events occupy a greater proportion in the feature vector, and the influence of the current state is improved. Finally, since the feature vector generated by dynamic weighting usually has a high dimension, the principal component analysis algorithm is used to reduce the dimension of the features, extract the most representative principal components, compress the feature dimension and reduce the redundant information, and finally the obtained log feature information is transmitted to the long-time extreme learning machine model for subsequent anomaly detection and analysis.
[0111] The embodiment of the present application strengthens the influence of recent log events on the overall features by generating a dynamically weighted feature vector, so that the data is more consistent with the current system state, and the interference of obsolete data is suppressed. Then, the principal component analysis algorithm is used to reduce the dimension of the dynamically weighted feature vector, effectively extract the main information in the data, and reduce the redundancy and complexity caused by the dimension, thereby not only improving the expression ability of the log features, but also providing high-quality input features for the subsequent long-time extreme learning machine model, so that the accuracy of event prediction is significantly improved, thereby more effectively supporting the detection of abnormal logs.
[0112] Optionally, on the basis of one or more embodiments of the above Figure 1 , another optional embodiment provided by the embodiment of the present application can include that the step S130 can include:
[0113] The log feature information is input into a pre-constructed long-time extreme learning machine model, so that the log feature information is transmitted to a dynamic subnet and a static subnet of the long-time extreme learning machine model, wherein the dynamic subnet comprises a long-term memory unit; long-term dynamic information output by the long-term memory unit and static information output by the static subnet are obtained; and the long-term dynamic information and the static information are used to generate an event prediction result and output.
[0114] The dynamic subnet is a component of the long-time extreme learning machine model, mainly responsible for extracting time-related dynamic information from input data. The dynamic subnet usually adopts a variant of LSTM (Long short-term memory), which has long-term memory units and short-term memory units, and can capture short-term and long-term dependencies in time series data.
[0115] The static subnet is a component of the long-time extreme learning machine model, mainly responsible for extracting static information from input data. Unlike the dynamic subnet, the static subnet ignores time factors and focuses on analyzing the overall characteristics of the data, ensuring that the extracted information does not overlap with the dynamic information.
[0116] The long-term memory unit is a key component in the dynamic subnet, responsible for storing and processing long-term state information of the data set.
[0117] The long-term dynamic information is a feature output by the long-term memory unit, reflecting important change patterns of the data over time.
[0118] The static information is a feature output by the static subnet, representing the fixed attributes of the input data, and is not affected by time changes.
[0119] The long-time extreme learning machine model provided by the embodiment of the application introduces a long-time memory mechanism, enhances the ability of the traditional extreme learning machine (ELM), and forms a framework with dynamic learning characteristics. The L-ELM model comprises two main subnets: a dynamic subnet and a static subnet. The dynamic subnet is responsible for extracting dynamic information from input log feature information, and especially adopts a variant of a long short-term memory network to better capture long-term and short-term changes in time series. In the dynamic subnet, the long-term memory unit is specially used to store overall state information that changes over time, and the short-term memory unit focuses on input features at the current time point to respond to immediate changes. During training of the long-time extreme learning machine model, only the output of the long-term memory unit is selected to avoid information overlap with the static subnet, thereby preventing redundant information from affecting the performance of the model. The dynamic subnet and the static subnet are independent of each other, ensuring that their information extraction processes are not disturbed.
[0120] In the specific implementation process, the log feature information is first input into the long-term extreme learning machine model and normalized by Z-score to eliminate the dimensional differences between the features. Then, the long-term extreme learning machine model is initialized. During the initialization phase, the dynamic subnet will use historical log data for training and calculate the loss value of each training iteration. :
[0121] ,
[0122] in, is the iteration number of the current training, is the expected total number of iterations, is the target value, The training values are then used. The backpropagation algorithm is then used to adjust the dynamic subnet weights and biases, continuously training the dynamic subnet parameters until the set training conditions (such as the number of training iterations or the minimum error condition) are met, allowing the dynamic subnet to extract long-term dynamic information. The trained dynamic subnet information (hidden layer output) is then extracted, and the hidden layer information of the static subnet is calculated. Finally, the generalized inverse is calculated to obtain the hidden-to-output layer weights of the long-term extreme learning machine model, completing the initialization of the long-term extreme learning machine model. Subsequently, the model parameters are automatically updated through recursive least squares (as well as drift detection and retraining triggers).
[0123] The long-term extreme learning machine model provided by the embodiment of the present invention can generate and output event prediction results by combining long-term dynamic information from long-term memory units and static information output by static subnets, supporting subsequent anomaly detection and analysis, thereby ensuring that the long-term extreme learning machine model fully considers the dual effects of time changes and static characteristics when processing log data, thereby improving the accuracy and reliability of event prediction results.
[0124] Optional, in the above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, step S150 may include:
[0125] Based on the residual distribution of the frequency distribution results of log events in the event prediction results and the actual event results, the first confidence interval is determined; the KL divergence of the frequency distribution results of log events in the event prediction results and the actual event results in the current time window is calculated; when the KL divergence exceeds the first confidence interval, an anomaly determination result is generated indicating that abnormal log data exists in the original log data.
[0126] Specifically, the embodiment of the present application can calculate the residual distribution between the event prediction result and the log event frequency distribution result in the event real result, take 95% as the first confidence interval to dynamically adapt to business fluctuations, and avoid false positives caused by static thresholds. Subsequently, frequency level anomaly detection is performed in the current time window, and the KL (Kullback-Leibler) divergence of the log event frequency distribution in the event prediction result and the event real result is calculated to quantify the difference between the two distributions. If the calculated KL divergence exceeds the previously determined first confidence interval, it means that the current event frequency distribution deviates significantly from the model prediction, so that it is judged that there may be abnormal logs in the original log data. Finally, the corresponding abnormal judgment result is generated to further support subsequent abnormal analysis and processing.
[0127] It can be understood that, in the case where the KL divergence does not exceed the first confidence interval, the embodiment of the present application can directly generate the abnormal judgment result that there is no abnormal log data in the original log data, or can generate the final abnormal judgment result in cooperation with the result generated by other abnormal judgment means.
[0128] The embodiment of the present application calculates the KL divergence between the event prediction result and the event real result in the current time window, which helps to quantify the difference between the two distributions. When the KL divergence exceeds the first confidence interval, it indicates that there is a significant deviation between the prediction and the actual, which can effectively indicate that there may be abnormal log data in the original log data. Finally, according to this judgment, the generated abnormal judgment result not only improves the monitoring ability of the system, but also provides strong support for timely identification and processing of potential abnormalities, ensuring the stability and security of the banking system.
[0129] Optionally, in the above Figure 1 Based on one or more embodiments corresponding to the above, in another optional embodiment provided by the embodiment of the present application, the step S150 can include:
[0130] Based on the residual distribution of the log event sequence result in the event prediction result and the event real result, a second confidence interval is determined; the matching degree of the log event sequence result in the event real result and the specified historical event sequence result in the current time window is calculated; and in the case where the matching degree exceeds the second confidence interval, an abnormal judgment result that there is abnormal log data in the original log data is generated.
[0131] Specifically, the embodiment of the present application can calculate the residual distribution between the event prediction result and the log event sequence result in the event real result, take 95% as the second confidence interval to dynamically adapt to business fluctuations, and avoid false positives caused by static thresholds. Subsequently, the log event sequence result in the event real result under the current time window is compared with the specified historical event sequence result, sequence-level similarity analysis is performed using the dynamic time warping (DTW) algorithm to calculate the matching degree between them, thereby effectively processing the differences caused by time delay or deformation in time series data. The calculated matching degree is compared with the previously determined second confidence interval. If the matching degree exceeds the second confidence interval, it indicates that there is a significant difference between the current log event sequence and the historical normal sequence, thereby indicating that there may be abnormal logs in the original log data. Finally, an abnormality determination is generated based on this result to help identify and handle system abnormal conditions in a timely manner.
[0132] It can be understood that, in the case where the matching degree does not exceed the second confidence interval, the embodiment of the present application can directly generate an abnormality determination result that there is no abnormal log data in the original log data, or can generate a final abnormality determination result in cooperation with the result generated by other abnormality determination means.
[0133] The embodiment of the present application compares the real result of the current event sequence with the pre-set historical normal event sequence, calculates the matching degree between them, and if the matching degree exceeds the second confidence interval, it can be determined that the current log event sequence deviates significantly from the historical normal sequence, thereby indicating that there are abnormal logs in the original log data. Finally, according to this determination, the generated abnormality determination result not only improves the accuracy of abnormality detection, but also enhances the adaptability of the system to dynamic changes, ensuring the stability and security of system operation.
[0134] Although the operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing can be advantageous.
[0135] It should be understood that each of the steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present application is not limited in this respect.
[0136] Corresponding to the above-mentioned method embodiments, the embodiment of the present application also provides a log abnormality detection device, the structure of which is as follows Figure 4As shown, the log anomaly detection device can include: an original log data obtaining unit 10, a time series matrix conversion unit 20, a log feature information obtaining unit 30, an event prediction result obtaining unit 40, an event real result obtaining unit 50, and an abnormality determination result generating unit 60.
[0137] The original log data obtaining unit 10 is configured to obtain original log data of a specified bank system, wherein the original log data includes log events automatically generated by the specified bank system during operation.
[0138] The time series matrix conversion unit 20 is configured to convert the original log data into a time series matrix according to event types and event frequencies of each log event in the original log data.
[0139] The log feature information obtaining unit 30 is configured to perform dynamic weighting and dimensionality reduction processing on the time series matrix to obtain log feature information.
[0140] The event prediction result obtaining unit 40 is configured to input the log feature information into a pre-constructed long-time extreme learning machine model to obtain an event prediction result output by the long-time extreme learning machine model, wherein the event prediction result is a log event result predicted by the long-time extreme learning machine model for the specified bank system in a current time window.
[0141] The event real result obtaining unit 50 is configured to obtain an event real result of the specified bank system, wherein the event real result is a log event result actually generated by the specified bank system in the current time window.
[0142] The abnormality determination result generating unit 60 is configured to determine whether there is abnormal log data in the original log data by using the event prediction result and the event real result, and generate an abnormality determination result.
[0143] Optionally, the log anomaly detection device can further include a model parameter updating unit.
[0144] The model parameter updating unit is configured to update model parameters of the long-time extreme learning machine model by using a plurality of abnormality determination results in the first time window after the abnormality determination result generating unit 60 determines whether there is abnormal log data in the original log data by using the event prediction result and the event real result, and generates an abnormality determination result, to obtain an updated long-time extreme learning machine model.
[0145] Optionally, the time series matrix conversion unit 20 can be specifically configured to cluster the original log data into event templates according to event types by using a log clustering algorithm; separate static parameters and dynamic parameters in the event templates to generate event encoding vectors; and construct the time series matrix according to a second time window by counting event frequencies based on the event encoding vectors.
[0146] Optionally, the log feature information obtaining unit 30 can be specifically configured to generate a dynamic weighted feature vector by using the time series matrix and the preset time decay weight; and perform dimension reduction compression on the dynamic weighted feature vector by using a principal component analysis algorithm to obtain the log feature information.
[0147] Optionally, the event prediction result obtaining unit 40 can be specifically configured to input the log feature information into a pre-constructed long-time extreme learning machine model, so that the log feature information is transmitted into a dynamic subnetwork and a static subnetwork of the long-time extreme learning machine model, wherein the dynamic subnetwork comprises a long-term memory unit; obtain long-term dynamic information output by the long-term memory unit and static information output by the static subnetwork; and generate and output an event prediction result by using the long-term dynamic information and the static information.
[0148] Optionally, the log event result comprises a log event frequency distribution result, and the abnormality determination result generating unit 60 can be specifically configured to determine a first confidence interval based on a residual distribution of the log event frequency distribution result in the event prediction result and an event real result; calculate a KL divergence of the log event frequency distribution result in the event prediction result and the event real result under a current time window; and generate an abnormality determination result that there is abnormal log data in the original log data in a case where the KL divergence exceeds the first confidence interval.
[0149] Optionally, the log event result comprises a log event sequence result, and the abnormality determination result generating unit 60 can be specifically configured to determine a second confidence interval based on a residual distribution of the log event sequence result in the event prediction result and the event real result; calculate a matching degree of the log event sequence result in the event real result and a specified historical event sequence result under the current time window; and generate an abnormality determination result that there is abnormal log data in the original log data in a case where the matching degree exceeds the second confidence interval.
[0150] The log anomaly detection device provided by the application comprises: an original log data obtaining unit configured to obtain original log data of a specified banking system, wherein the original log data comprises log events automatically generated by the specified banking system during operation; a time series matrix converting unit configured to convert the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data; a log feature information obtaining unit configured to obtain log feature information by performing dynamic weighting and dimension reduction processing on the time series matrix; an event prediction result obtaining unit configured to input the log feature information into a pre-constructed long-time extreme learning machine model to obtain an event prediction result output by the long-time extreme learning machine model, wherein the event prediction result is a log event result predicted by the long-time extreme learning machine model for the specified banking system in a current time window; an event real result obtaining unit configured to obtain an event real result of the specified banking system, wherein the event real result is a log event result actually generated by the specified banking system in the current time window; and an anomaly determination result generating unit configured to determine whether there is abnormal log data in the original log data by using the event prediction result and the event real result, and generate an anomaly determination result. The original log data is converted into a time series matrix and features are extracted by dynamic weighting and dimension reduction, and the long-time extreme learning machine model is used for event prediction, so that efficient log anomaly detection capability can be provided to meet the demand of the financial industry for safety and stability.
[0151] As to the device in the above-mentioned embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and thus will not be described in detail here.
[0152] The log anomaly detection device comprises a processor and a memory, and the original log data obtaining unit 10, the time series matrix converting unit 20, the log feature information obtaining unit 30, the event prediction result obtaining unit 40, the event real result obtaining unit 50 and the anomaly determination result generating unit 60 are all stored in the memory as program units, and the corresponding functions are realized by the processor executing the above-mentioned program units stored in the memory.
[0153] The processor comprises a core, and the core retrieves the corresponding program units from the memory. One or more than one core can be set, and the core parameters are adjusted to obtain the original log data of the specified banking system, and the original log data is converted into a time series matrix according to the event type and frequency, and then the matrix is dynamically weighted and dimensionally reduced to extract the log feature information. The features are input into the long-time extreme learning machine model to obtain the event prediction result in the current time window. At the same time, the real event result of the specified banking system in the same time window is obtained. By comparing the prediction result with the real result, it is determined whether there is abnormal log in the original log data, and an anomaly determination result is generated.
[0154] The embodiment of the application provides a computer readable storage medium, which stores a program, and the program is executed by a processor to realize the log anomaly detection method.
[0155] An embodiment of the present invention provides a processor, which is used to run a program, wherein the log anomaly detection method is executed when the program is running.
[0156] like Figure 5 As shown, an embodiment of the present invention provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003. The processor 1001 and the memory 1002 communicate with each other via the bus 1003. The processor 1001 is configured to call program instructions in the memory 1002 to execute the above-described log anomaly detection method. The electronic device herein may be a server, a PC, a PAD, a mobile phone, etc.
[0157] The present invention also provides a computer program product, which, when executed on an electronic device, is suitable for executing a program that initializes the steps of the log anomaly detection method.
[0158] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to produce a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0159] In a typical configuration, an electronic device includes one or more processors (CPUs), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, and the like.
[0160] Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip. Memory is an example of a computer-readable medium.
[0161] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0162] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.
[0163] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scene, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0164] In the description of the present application, it should be understood that if the terms "upper", "lower", "front", "rear", "left" and "right" and the like indicate the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the position or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0165] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; it is not intended to exclude myriad other embodiments of the present application that other present or future technologies can provide. It must be noted that as used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. The terms "includes" and "comprising," as well as derivatives thereof, mean that various embodiments include, but are not limited to, the listed material or step or steps. The terms "sub- steps" of any method, "steps" of any method, "elements" of any system, and "components" of any system are interchangeable and do not require a particular order of execution. The terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term "about" when used before a numerical designation, has its ordinary meaning in the art, for example, it can refer to an approximation within 1, 2, 3, 4, 5%, 10% or another suitable non-zero amount of the numerical designation. The term "consisting essentially of" when used in this specification, specifies the presence of stated features, integers, steps, operations, elements, and / or components as well as those that do not materially affect the method, system, and / or result. The term "consisting of" when used in this specification, specifies the presence of stated features, integers, steps, operations, elements, and / or components but not others. The term "comprising" when used in this specification, means "including, but not limited to."
[0166] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer readable program code.
[0167] The foregoing is merely illustrative of the embodiments of the present application and the present application should not be limited to the described embodiments. Numerous other embodiments can be devised by those skilled in the art without departing from the spirit and principles of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the scope of the present application.
Claims
1. A log anomaly detection method, characterized in that: include: Obtaining original log data of a designated bank system, wherein the original log data includes log events automatically generated during operation of the designated bank system; Converting the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data; Performing dynamic weighting and dimensionality reduction processing on the time series matrix to obtain log feature information; Inputting the log feature information into a pre-built long-term extreme learning machine model to obtain an event prediction result output by the long-term extreme learning machine model, wherein the event prediction result is a result predicted by the long-term extreme learning machine model for log events generated by the designated bank system in the current time window; Obtaining a true event result of the designated bank system, wherein the true event result is a log event result actually generated by the designated bank system in the current time window; The event prediction result and the actual event result are used to determine whether there is abnormal log data in the original log data, and generate an abnormality determination result.
2. The method according to claim 1, characterized in that After determining whether there is abnormal log data in the original log data by using the event prediction result and the actual event result and generating an abnormality determination result, the method further includes: The model parameters of the long-term extreme learning machine model are updated using the multiple abnormality determination results in the first time window to obtain an updated long-term extreme learning machine model.
3. The method according to claim 1, characterized in that The converting the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data includes: Clustering the original log data into event templates according to event types using a log clustering algorithm; Separating static parameters and dynamic parameters from the event template to generate an event coding vector; Based on the event coding vector, the event frequency is counted according to the second time window to construct a time series matrix.
4. The method according to claim 1, wherein The dynamically weighting and dimensionality reduction processing of the time series matrix to obtain log feature information includes: Generating a dynamic weighted feature vector using the time series matrix and preset time decay weights; A principal component analysis algorithm is used to perform dimensionality reduction compression on the dynamic weighted feature vector to obtain log feature information.
5. The method according to claim 1, wherein Inputting the log feature information into a pre-built long-term extreme learning machine model to obtain an event prediction result output by the long-term extreme learning machine model includes: The log feature information is input into a pre-built long-term extreme learning machine model so that the log feature information is transmitted to the dynamic subnet and static subnet of the long-term extreme learning machine model, wherein the dynamic subnet includes a long-term memory unit; long-term dynamic information output by the long-term memory unit and static information output by the static subnet are obtained; and event prediction results are generated and output using the long-term dynamic information and the static information.
6. The method according to claim 1, wherein The log event result includes a log event frequency distribution result. The use of the event prediction result and the actual event result to determine whether there is abnormal log data in the original log data and generate an abnormality determination result includes: Determining a first confidence interval based on a residual distribution of a frequency distribution result of a log event in the event prediction result and the actual event result; Calculate the KL divergence of the frequency distribution of the event in the event prediction result and the actual event result in the current time window; When the KL divergence exceeds the first confidence interval, the abnormality determination result indicating that the abnormal log data exists in the original log data is generated.
7. The method according to claim 1, characterized in that The log event result includes a log event sequence result. The use of the event prediction result and the actual event result to determine whether there is abnormal log data in the original log data and generate an abnormality determination result includes: Determining a second confidence interval based on a residual distribution of a log event sequence result in the event prediction result and the actual event result; Calculating the matching degree between the log event sequence result in the actual result of the event in the current time window and the specified historical event sequence result; When the degree of matching exceeds the second confidence interval, the abnormality determination result indicating that the abnormal log data exists in the original log data is generated.
8. A log anomaly detection device, characterized in that: include: Original log data acquisition unit, time series matrix conversion unit, log feature information acquisition unit, event prediction result acquisition unit, event real result acquisition unit and abnormal judgment result generation unit, The original log data obtaining unit is configured to obtain original log data of a designated bank system, wherein the original log data includes log events automatically generated during operation of the designated bank system; The time series matrix conversion unit is used to convert the original log data into a time series matrix according to the event type and event frequency of each log event in the original log data; The log feature information obtaining unit is used to perform dynamic weighting and dimensionality reduction processing on the time series matrix to obtain log feature information; The event prediction result obtaining unit is configured to input the log feature information into a pre-built long-term extreme learning machine model to obtain an event prediction result output by the long-term extreme learning machine model, wherein the event prediction result is a result predicted by the long-term extreme learning machine model regarding a log event generated by the designated bank system in a current time window; The event real result obtaining unit is configured to obtain the event real result of the designated bank system, wherein the event real result is the log event result actually generated by the designated bank system in the current time window; The abnormality determination result generating unit is used to use the event prediction result and the actual event result to determine whether there is abnormal log data in the original log data, and generate an abnormality determination result.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the log anomaly detection method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The electronic device includes at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the log anomaly detection method according to any one of claims 1 to 7.