Distributed server cluster log processing method and device
By building a distributed server cluster log acquisition network and intelligent log analysis model, the problems of insufficient load balancing, data transmission efficiency and intelligent analysis in the log processing system are solved, efficient log acquisition and intelligent analysis are realized, and system performance and practicality are improved.
Patent Information
- Application Number
- CN202510601646.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing log processing system lacks a distributed load balancing mechanism, which is difficult to meet the log collection needs of large-scale server clusters, has low data transmission efficiency, has failed to achieve incremental acquisition and real-time compression, and lacks intelligent abnormal detection and correlation analysis capabilities.
Build a distributed server cluster log acquisition network, dynamically allocate acquisition tasks through the load balancing scheduling center, design incremental data acquisition channels and websocket long connection pools, combine real-time compression coding and duplicate data detection mechanisms, and build an intelligent log analysis model, including anomaly detection, pattern recognition and event association analysis models, providing visual display and data export functions.
It significantly improves the performance and practicality of the log processing system, realizes efficient log collection and analysis, reduces network bandwidth usage and storage space consumption, and improves the accuracy of abnormal detection and the efficiency of event correlation analysis.
Smart Images

Figure CN120123184B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a distributed server cluster log processing method and device. Background Art
[0002] Existing log processing methods have obvious shortcomings. Traditional systems often use a single server processing model and lack a distributed load balancing mechanism, making it difficult to meet the log collection needs of large-scale server clusters.
[0003] Furthermore, existing technologies face bottlenecks in data transmission efficiency. Most systems use full data transmission, failing to implement incremental data collection and real-time compression, resulting in high network bandwidth usage and significant transmission delays.
[0004] Existing systems have technical shortcomings in log analysis. They lack intelligent anomaly detection and correlation analysis capabilities, making it difficult to quickly identify abnormal patterns and discover event correlations within massive log volumes. Addressing these issues is crucial for improving the performance and availability of log processing systems. Summary of the Invention
[0005] In response to the problems in the existing technology, the present application provides a distributed server cluster log processing method and device, which can effectively solve the shortcomings of traditional technologies in distributed collection, data transmission and intelligent analysis, and significantly improve the performance and practicality of the log processing system.
[0006] In order to solve at least one of the above problems, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides a distributed server cluster log processing method, comprising:
[0008] Establish a distributed server cluster log collection network, create a server cluster configuration table on the management terminal, the server cluster configuration table contains the network address, port number, access credentials and server load threshold of each server to be monitored, establish a load balancing scheduling center based on the server cluster configuration table, the load balancing scheduling center periodically collects system load information of each server to be monitored, dynamically allocates log collection tasks based on the system load information, and generates a task scheduling queue according to preset task priorities;
[0009] Constructing an incremental log data collection channel, establishing a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, recording the latest read position of each log file, performing incremental data collection according to the latest read position, performing real-time compression encoding on the collected log data, detecting duplicate data based on a sliding time window, removing the duplicate data, and transmitting the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0010] An intelligent log analysis model is constructed, and deep learning methods are used to train historical log data. An analysis engine including an anomaly detection model, a pattern recognition model, and an event correlation analysis model is constructed. The restored log data is input into the analysis engine. The anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information. The abnormal event warning information, log classification and labeling results, and correlation analysis results are visualized in the management terminal interface. When a user triggers a download operation, the analysis result data within the specified time range is packaged, compressed, and downloaded to local storage.
[0011] Furthermore, the method further includes: entering server cluster information in a configuration interface of a management terminal, saving the server cluster information in a database to generate a server cluster configuration table, wherein the server cluster configuration table includes a server identifier, a network address, a port number, an access credential, a load threshold, and a priority field; reading the network address and port number in the server cluster configuration table to establish a TCP connection, verifying the validity of the access credential through the SSH protocol, and marking the server node as available after the verification is passed;
[0012] Build a load monitoring agent program and deploy it to each server to be monitored. The load monitoring agent program collects CPU usage, memory occupancy, disk IO and network bandwidth data, calculates the server's comprehensive load score, compares the comprehensive load score with the load threshold in the server cluster configuration table, and issues a warning signal to the server that exceeds the load threshold. Adjust the collection task priority of the server according to the warning signal.
[0013] Furthermore, the method further includes: creating a load balancing scheduling center process, reading the load threshold and priority parameters in the server cluster configuration table, establishing a system load information collection timer task, wherein the system load information collection timer task obtains CPU usage, memory occupancy, disk IO and network bandwidth data from the load monitoring agent program of each monitored server at a preset time interval, and stores the system load information in a load status cache table;
[0014] Based on the system load information in the load status cache table, the available resource capacity of each server is calculated, the task volume of the log collection task is evaluated, the task volume size is matched with the available resource capacity, and a task allocation plan is generated. According to the task allocation plan and the priority parameters in the server cluster configuration table, a task scheduling queue is constructed, and the task scheduling queue is distributed to the load monitoring agent program of each server to be monitored.
[0015] Furthermore, the method further includes: obtaining the network address and port number of the server to be monitored based on the task scheduling queue, establishing a persistent connection between the management terminal and the server to be monitored through the websocket protocol, storing the persistent connection in a connection pool according to the server identifier, creating a connection pool manager, wherein the connection pool manager monitors the connection status, automatically reestablishes disconnected connections, and maintains the availability of the connection pool;
[0016] Create a log file reading position record table, which contains the log file path, file size, last read timestamp and reading position offset fields. Before reading the log file each time, obtain the offset of the last read position from the log file reading position record table, and use the offset as the starting position of the new round of reading. After the reading is completed, update the timestamp and offset in the log file reading position record table.
[0017] Furthermore, the method further includes: obtaining an offset from the log file reading position record table, reading newly added log data using the offset as a file pointer position, performing row segmentation processing on the read log data, calculating a hash value for each row of data, obtaining a hash value of historical data from a log data cache according to a preset time window range, comparing the hash value of the current data with the hash value of the historical data, removing duplicate data, and writing the deduplication result into a compression buffer;
[0018] Execute the LZ4 compression algorithm on the log data in the compression buffer to generate a compressed data block, add a data identification header to the compressed data block, the data identification header includes the data block size, compression algorithm type and checksum information, send the compressed data block to the management terminal through the websocket long connection, select the corresponding decompression algorithm on the management terminal according to the data identification header to decompress the compressed data, and write the decompressed data into the log storage area.
[0019] Furthermore, the method further includes: reading historical log data from a log storage area, performing text preprocessing on the historical log data, extracting a timestamp, event type, operation object, status code, execution result, error code, user identifier, and operation instruction in the log to form a feature set, converting the feature set into a vector representation, using an LSTM neural network to build an anomaly detection model and a pattern recognition model, using a graph neural network to build an event association analysis model, using the feature vector to train the anomaly detection model, pattern recognition model, and event association analysis model, and saving the trained model to a model library;
[0020] The anomaly detection model, pattern recognition model and event association analysis model are loaded from the model library to build a log analysis engine. The restored log data is converted into a feature vector and input into the log analysis engine. The anomaly detection model calculates the anomaly score of the feature vector. The pattern recognition model performs multi-classification prediction on the log content and outputs a category label. The event association analysis model mines the temporal association and causal relationship between log events based on the graph structure, and generates abnormal event warning information based on the anomaly score, category label and association relationship.
[0021] Furthermore, the system further includes: creating a management terminal visualization panel, dividing the abnormal event warning information into three levels: high, medium, and low according to the abnormality level and displaying it in the warning information area; displaying the distribution of each category in the form of a pie chart; displaying the correlation analysis result through a force-directed graph to show the correlation strength between log event nodes; creating a time range selector in the visualization panel, and updating the visualization data in real time according to the start and end time of the time range selector;
[0022] Listen for the user's download button click event in the visualization panel, obtain the start and end time parameters in the time range selector, query the abnormal event warning information, log classification and annotation results and correlation analysis results within the time range from the database, convert the query results into JSON format, execute the ZIP compression algorithm on the JSON format data to generate a compressed file, and save the compressed file to the local storage path specified by the user.
[0023] In a second aspect, the present application provides a distributed server cluster log processing device, comprising:
[0024] A dynamic allocation module is used to establish a distributed server cluster log collection network, create a server cluster configuration table on the management terminal, and the server cluster configuration table contains the network address, port number, access credentials and server load threshold of each server to be monitored. A load balancing scheduling center is established based on the server cluster configuration table. The load balancing scheduling center periodically collects system load information of each server to be monitored, dynamically allocates log collection tasks based on the system load information, and generates a task scheduling queue according to the preset task priority.
[0025] A log processing module is used to build an incremental log data collection channel, establish a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, record the latest read position of each log file, perform incremental data collection according to the latest read position, perform real-time compression encoding on the collected log data, detect duplicate data based on a sliding time window, remove the duplicate data, and then transmit the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0026] The log analysis module is used to build an intelligent log analysis model, use deep learning methods to train historical log data, and build an analysis engine that includes an anomaly detection model, a pattern recognition model, and an event correlation analysis model. The restored log data is input into the analysis engine. The anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and annotates the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information. The abnormal event warning information, log classification and annotation results, and correlation analysis results are visualized in the management terminal interface. When a user triggers a download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
[0027] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the distributed server cluster log processing method when executing the program.
[0028] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the distributed server cluster log processing method when executed by a processor.
[0029] In a fifth aspect, the present application provides a computer program product, including a computer program / instruction, which implements the steps of the distributed server cluster log processing method when executed by a processor.
[0030] It can be seen from the above technical solution that the present application provides a distributed server cluster log processing method and device, which constructs a distributed server cluster log collection network and dynamically allocates collection tasks through a load balancing scheduling center. Design an incremental data collection channel, use the websocket long connection pool to achieve efficient transmission, and combine real-time compression encoding and duplicate data detection mechanism to optimize transmission efficiency. Construct an intelligent log analysis model, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, it realizes anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis, and provides visual display and data export functions. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0032] Figure 1 This is a flow chart of a distributed server cluster log processing method in an embodiment of the present application;
[0033] Figure 2 This is a structural diagram of a distributed server cluster log processing device in an embodiment of the present application;
[0034] Figure 3 Schematic diagram of the structure of the electronic device in the embodiment of the present application.
[0035] Reference numerals:
[0036] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION
[0037] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.
[0039] Taking into account the problems existing in the prior art, the present application provides a distributed server cluster log processing method and device, which constructs a distributed server cluster log collection network and dynamically allocates collection tasks through a load balancing scheduling center. Design an incremental data collection channel, use the websocket long connection pool to achieve efficient transmission, and combine real-time compression encoding and duplicate data detection mechanism to optimize transmission efficiency. Construct an intelligent log analysis model, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, it realizes anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis, and provides visual display and data export functions. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system.
[0040] In order to effectively solve the shortcomings of traditional technologies in distributed collection, data transmission and intelligent analysis, and significantly improve the performance and practicality of the log processing system, this application provides an embodiment of a distributed server cluster log processing method, see Figure 1 The distributed server cluster log processing method specifically includes the following contents:
[0041] Step S101: Establishing a distributed server cluster log collection network, creating a server cluster configuration table on a management terminal, the server cluster configuration table including the network address, port number, access credentials, and server load threshold of each server to be monitored, establishing a load balancing scheduling center based on the server cluster configuration table, the load balancing scheduling center periodically collecting system load information of each server to be monitored, dynamically allocating log collection tasks based on the system load information, and generating a task scheduling queue according to a preset task priority;
[0042] Optionally, this embodiment achieves efficient log data collection and processing by building a distributed server cluster log collection network. On the management terminal, a highly scalable configuration interface is designed. This modular interface includes three main functional modules: a server information entry area, a connection status display area, and a load monitoring area. The configuration interface uses the React framework to achieve a responsive layout, ensuring good display effects on different terminal devices.
[0043] This embodiment utilizes a step-by-step verification mechanism when entering server information. First, the format of the entered network address is verified, supporting IPv4, IPv6, and domain names. Next, the port number is verified for validity. The system defaults to port 8080, but supports custom port configuration. Finally, access credentials are encrypted and stored, using an asymmetric encryption algorithm to encrypt sensitive information such as passwords. The ciphertext is stored in the database to ensure security. This step-by-step verification mechanism effectively avoids connection failures caused by configuration errors.
[0044] The server cluster configuration table created in this embodiment is stored in a relational database. The table structure includes fields such as server ID, network address, port number, access credentials, load threshold, and priority. The server ID serves as the primary key, and a UUID generation algorithm is used to ensure uniqueness. The load threshold field is used to store threshold parameters for four dimensions: CPU usage, memory occupancy, disk IO, and network bandwidth. The priority field is represented by an integer; the smaller the value, the higher the priority. The database uses a master-slave replication architecture to provide data backup and failover capabilities.
[0045] This embodiment utilizes an innovative design for the load balancing scheduling center. The scheduling center runs as an independent microservice, using the Spring Cloud framework for service registration and discovery. The scheduling center utilizes ZooKeeper for distributed coordination. If the master node fails, slave nodes can quickly take over scheduling tasks, ensuring high service availability. The scheduling center also implements dynamic configuration refresh, allowing changes to the server cluster configuration to take effect without requiring a restart.
[0046] This embodiment designs a precise system load collection mechanism. A load collection agent is deployed on each monitored server. This agent, developed in the Go language, features low resource usage and high concurrency processing capabilities. The agent uses system calls to obtain CPU usage, memory usage, disk I / O, and network bandwidth data. The sampling interval is dynamically adjustable and defaults to 30 seconds. The collected data is locally aggregated and sent to the dispatch center via the gRPC protocol, ensuring efficient and reliable data transmission.
[0047] This embodiment implements an intelligent task allocation algorithm. After the dispatch center receives the load data, it first calculates the comprehensive load score of each server. The calculation formula is:
[0048] Score = w1 CPU + w2 Memory + w3 DiskIO + w4 Network, where w1 to w4 are weight coefficients that can be adjusted based on actual business scenarios. When the score exceeds the preset threshold, a load alarm is triggered, and the system automatically lowers the priority of the collection task on that server and transfers some tasks to servers with lower loads.
[0049] This embodiment builds an efficient task scheduling queue. The queue is implemented using a priority queue data structure, supporting dynamic task insertion and adjustment. Each task contains information such as server identification, collection path, and collection interval. Queue processing utilizes a multi-threaded model, dynamically adjusting the thread pool size based on server performance. A batch processing mechanism is also implemented, allowing tasks of the same priority to be combined for execution, improving processing efficiency.
[0050] Through the above technical innovations, this embodiment effectively solves the key problems of log collection in a distributed environment: complex configuration management, unreasonable load balancing, low task scheduling efficiency, etc. In practical applications, this solution can support the log collection needs of large-scale server clusters, and ensure the continuous and stable operation of collection tasks through dynamic load balancing and intelligent task scheduling. It is particularly suitable for microservice architectures and containerized deployment environments, and significantly improves the reliability and efficiency of log collection. The systematic and innovative nature of this solution enables it to adapt to the application needs of enterprises of different sizes, and through flexible configuration management and intelligent task scheduling, it achieves a comprehensive improvement in the log collection system.
[0051] Step S102: constructing an incremental log data collection channel, establishing a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, recording the latest read position of each log file, performing incremental data collection according to the latest read position, performing real-time compression encoding on the collected log data, detecting duplicate data based on a sliding time window, removing the duplicate data, and transmitting the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0052] Optionally, this embodiment achieves efficient distributed log collection by establishing an incremental log data collection channel. Based on the Node.js event-driven model, the management terminal creates a WebSocket server instance and listens for connection requests from each monitored server. A handshake protocol is used for authentication during connection establishment. Once authenticated, the connection object is stored in a connection pool. A heartbeat detection mechanism is also activated, periodically sending ping frames to ensure connection activity.
[0053] This embodiment employs a layered design for connection pool management. Connection pools are grouped by server cluster, with each group maintaining independent connection counters and status markers. Core connection pool parameters include the maximum number of connections, idle timeout, and reconnection interval, which can be dynamically adjusted through the configuration file. When a disconnect is detected, the connection pool manager automatically triggers a reconnection mechanism to ensure continued connection availability.
[0054] This embodiment implements a precise file location tracking mechanism. A location record table is created in Redis, using a hash structure to store metadata for each log file. The key is generated by combining the server ID and the file path, while the value contains fields such as the file size, the last read timestamp, and the read position offset. The collection program queries the location record table before each read to obtain the last read position, implementing incremental collection and avoiding reprocessing of already collected data.
[0055] This embodiment designs an efficient data compression scheme. It uses the LZ4 compression algorithm for real-time compression of log data, which boasts extremely fast compression speeds and high compression ratios. The compression process uses a streaming mode with a 64KB block size. Each block contains independent header information, supporting random access and parallel decompression. It also implements dynamic adjustment of the compression level, automatically selecting the optimal compression parameters based on CPU load.
[0056] This embodiment establishes a reliable duplicate data detection mechanism. A sliding time window method is used, and the window size is set to 5 minutes by default, which can be adjusted according to actual needs. Within the window range, a Bloom filter is used to record the hash value of the processed log line. When a new log line arrives, its hash value is first calculated and the Bloom filter is queried. If a duplicate is found, it is directly discarded. The Bloom filter uses multiple hash functions to improve accuracy and controls memory usage by regularly clearing expired data.
[0057] This embodiment implements an efficient data transmission protocol. A custom application layer protocol is encapsulated within the WebSocket frame, consisting of a 4-byte magic number, a 1-byte version number, a 4-byte data length, a 2-byte checksum, and a variable-length data body. The sender packages the compressed data according to the protocol format, and the receiver parses it using a protocol parser. The transmission process supports fragmented transmission, with the size of a single fragment limited to 1MB to prevent network impact from large data block transmission.
[0058] This embodiment designs a complete data restoration process. After receiving compressed data, the management terminal first performs protocol parsing and data integrity verification. Once verification is successful, the corresponding decompression algorithm is selected based on the compressed block header information for decompression. The decompression process uses multi-threaded parallel processing, with each thread responsible for decompressing a compressed block. The decompressed data is written to the local cache, awaiting subsequent analysis and processing.
[0059] Through the above technological innovations, this embodiment effectively solves the key problems in distributed log collection: low data transmission efficiency, duplicate data interference, large decompression and restoration delays, etc. In practical applications, this solution can support the real-time log collection needs of large-scale server clusters, and significantly reduces network bandwidth occupancy and storage space consumption through incremental collection and efficient compression transmission. It is particularly suitable for microservice architecture environments, and through reliable connection pool management and data deduplication mechanisms, it ensures the integrity and accuracy of log data. The systematic and innovative nature of this solution enables it to adapt to the application needs of enterprises of different sizes, and through technological innovation and optimized design, it achieves a comprehensive improvement in log collection efficiency.
[0060] Step S103: Construct an intelligent log analysis model, use deep learning methods to train historical log data, build an analysis engine including an anomaly detection model, a pattern recognition model and an event correlation analysis model, input the restored log data into the analysis engine, the anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information, and visualize the abnormal event warning information, log classification and labeling results and correlation analysis results in the management terminal interface. When it is detected that the user triggers the download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
[0061] Optionally, this embodiment uses deep learning methods to build an intelligent log analysis engine. First, historical log data is preprocessed, and key fields in the logs, including timestamps, event types, operation objects, and status codes, are extracted using regular expressions. The Word2Vec model is used to convert log text into a dense vector representation. During training, the window size is set to 5 and the vector dimension is set to 128 to fully capture the semantic relationships between words.
[0062] This embodiment uses a bidirectional LSTM network structure in the anomaly detection model. The input layer receives a sequence of log feature vectors. The LSTM layer contains 128 hidden units and learns the long-term dependencies of the log sequence through a gating mechanism. The model training uses a contrastive learning strategy, using normal log sequences as positive samples and artificially constructed anomalous sequences as negative samples. The loss function uses a contrastive loss, which reduces the distance between normal samples and increases the distance between anomalous samples. The anomaly score is calculated by calculating the average distance between the test sample and the normal sample set.
[0063] This embodiment designs an innovative pattern recognition model. It employs an attention-enhanced CNN network. The convolutional layer uses multiple convolution kernels of different sizes to extract features in parallel, capturing text patterns at different scales. The attention mechanism calculates weights on the feature map to highlight important pattern information. The classification layer uses a softmax function to output the probability distribution of each category. The model can recognize log categories including system errors, security warnings, performance anomalies, and configuration changes, and supports dynamic expansion of new categories.
[0064] This embodiment implements a complex event correlation analysis model. It constructs an event correlation graph based on a graph neural network, where nodes represent log events and edges represent the relationships between events. Each node's features include information such as event type, occurrence time, and impact range. Edge weights are calculated based on temporal correlation and causal relationships. The model uses a graph attention network (GAT) structure, using a multi-head attention mechanism to learn correlation patterns between nodes, enabling tracking and analysis of complex event chains.
[0065] This embodiment establishes an adaptive early warning mechanism. It comprehensively considers the anomaly detection model's score, the category probability of pattern recognition, and the impact range of event associations to calculate the event's risk level. Risk assessment uses a weighted scoring method, with the weight coefficient dynamically adjusted based on historical early warning results. When the risk level exceeds the threshold, an early warning message is generated, including a description of the anomaly, possible causes, and recommended actions. This warning message is pushed to the management terminal via a message queue to ensure a timely response.
[0066] This embodiment designs an intuitive visualization interface. The front-end application is built using the React framework, and the ECharts library is used for data visualization. Abnormal warning information is displayed in a timeline format, with different risk levels indicated by different colors. Log classification results are displayed using a pie chart to show the distribution ratio of each category. Event correlation analysis results are displayed using a force-directed graph, where the size of the node indicates the importance of the event and the thickness of the edge indicates the strength of the correlation. The visualization component supports interactive operations such as time range selection, node expansion, and path tracing.
[0067] This embodiment implements flexible data export functionality. Users can select a time range and export content type on the interface. The system then queries the analysis results for the corresponding time period and converts the data into JSON format. For large amounts of data, sharding is used to prevent memory overflows. Data compression uses the GZIP algorithm to effectively reduce the size of exported files. The export process supports resumable downloads, ensuring reliable downloads of large files.
[0068] Through the above technological innovations, this embodiment effectively solves key problems in log analysis: low anomaly detection accuracy, insufficient event correlation analysis, and delayed early warning responses. In practical applications, this solution can accurately identify various types of anomalies and provide increasingly accurate analysis results through continuous optimization of deep learning models. It is particularly suitable for operation and maintenance monitoring of large-scale distributed systems. Through multi-dimensional intelligent analysis, it significantly improves the efficiency of problem location and fault prevention. The systematic and innovative nature of this solution enables it to adapt to different types of log analysis needs. Through continuous learning and optimization of the model, it achieves a comprehensive improvement in log analysis capabilities.
[0069] As can be seen from the above description, the distributed server cluster log processing method provided in the embodiment of the present application can dynamically allocate collection tasks through the load balancing scheduling center by building a distributed server cluster log collection network. Design an incremental data collection channel, use the websocket long connection pool to achieve efficient transmission, and combine real-time compression coding and duplicate data detection mechanism to optimize transmission efficiency. Construct an intelligent log analysis model, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, it realizes anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis, and provides visual display and data export functions. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system.
[0070] In one embodiment of the distributed server cluster log processing method of the present application, the following contents may also be specifically included:
[0071] Step S201: Enter the server cluster information in the configuration interface of the management terminal, save the server cluster information in the database to generate a server cluster configuration table, the server cluster configuration table includes the server identification, network address, port number, access credentials, load threshold and priority fields, read the network address and port number in the server cluster configuration table to establish a TCP connection, verify the validity of the access credentials through the SSH protocol, and mark the server node as available after the verification is passed;
[0072] Step S202: Build a load monitoring agent program and deploy the load monitoring agent program to each server to be monitored. The load monitoring agent program collects CPU usage, memory occupancy, disk IO and network bandwidth data, calculates the server's comprehensive load score, compares the comprehensive load score with the load threshold in the server cluster configuration table, and issues a warning signal to the server that exceeds the load threshold, and adjusts the collection task priority of the server according to the warning signal.
[0073] Optionally, this embodiment adopts a distributed architecture concept in the design of the management terminal configuration interface. This interface is implemented using the Vue.js framework and employs a component-based design pattern, decoupling functions such as the server information entry form, status monitoring panel, and load threshold configuration into independent components. Form validation utilizes the async-validator library, enabling real-time validation of fields such as network addresses and port numbers to prevent the submission of invalid data.
[0074] This embodiment implements a secure data storage mechanism. Server cluster information is first normalized, with special characters escaped to prevent SQL injection attacks. The database uses PostgreSQL, and a server configuration table is created with an auto-incrementing primary key. Unique index constraints are also established for the network address and port number. Access credentials are one-way encrypted using the bcrypt algorithm before storage to ensure the security of sensitive information. The database connection pool uses HikariCP, and access performance is improved through connection pool configuration optimization.
[0075] This embodiment establishes a reliable connection verification mechanism. The TCP connection is established in non-blocking mode, and the connection timeout is set to 5 seconds. After the connection is successful, SSH verification is implemented through the JSch library, supporting both password and key authentication methods. A retry mechanism is used during the verification process, with a maximum of 3 retries, and the interval between each retry increases exponentially. After verification, the server node status is recorded in Redis, using a hash structure for storage. The key is the server identifier, and the value contains the status mark and the last update time.
[0076] This embodiment designs an efficient load monitoring agent. The agent is developed in the Go language and uses goroutines for concurrent data collection. CPU usage is obtained by parsing the / proc / stat file, calculating the proportion of CPU time in user mode and system mode. Memory usage is obtained by using the / proc / meminfo file to obtain physical memory and swap space usage. Disk I / O read and write rates are obtained using / proc / diskstats, and network bandwidth statistics are obtained using / proc / net / dev.
[0077] This embodiment implements an innovative load calculation model. The comprehensive load score adopts a weighted calculation method, and the calculation formula is:
[0078] Score = w1 CPU + w2 Memory + w3 DiskIO + w4 Network, where w1 through w4 are weight coefficients, with initial values of 0.4, 0.3, 0.2, and 0.1, respectively. Weight coefficients can be dynamically adjusted through the configuration file to adapt to the resource characteristics of different servers. Sampled data is smoothed using a sliding window with a window size of 5 minutes to reduce the impact of instantaneous fluctuations.
[0079] This embodiment establishes an intelligent early warning mechanism. Load threshold detection uses a multi-level threshold design, divided into warning and critical levels. When the load score exceeds the warning threshold, the system issues a warning-level alert; when it exceeds the critical threshold, a critical-level alert is issued. Alert information is pushed to the management terminal via a message queue (RabbitMQ) to ensure real-time delivery. Furthermore, alert information is persistently stored in a time series database (InfluxDB) for subsequent trend analysis.
[0080] This embodiment implements a dynamic task adjustment strategy. Upon receiving a warning signal, the scheduling system first suspends allocating new collection tasks to overloaded servers. For ongoing tasks, the system evaluates their priority and migrates low-priority tasks to less-loaded servers. This task migration process utilizes a smooth handover strategy to ensure continuous data collection. The server's available resource capacity is also updated for subsequent task allocation decisions.
[0081] Through the above technological innovations, this embodiment effectively solves the key problems in server cluster management: complex configuration management, inaccurate load monitoring, unreasonable task scheduling, etc. In actual applications, this solution can adapt to the dynamic management needs of large-scale server clusters, and ensure the stable operation of the log collection system through precise load monitoring and intelligent task scheduling. It is particularly suitable for microservice architecture environments, and significantly improves the availability and reliability of the system through distributed monitoring and dynamic adjustment. The systematic and innovative nature of this solution enables it to support the application needs of enterprises of different sizes, and through continuous optimization and improvement, it achieves a comprehensive improvement in the efficiency of server cluster management.
[0082] In one embodiment of the distributed server cluster log processing method of the present application, the following contents may also be specifically included:
[0083] Step S301: Create a load balancing scheduling center process, read the load threshold and priority parameters in the server cluster configuration table, establish a system load information collection timer task, and obtain CPU usage, memory occupancy, disk I / O, and network bandwidth data from the load monitoring agent of each monitored server at a preset time interval, and store the system load information in a load status cache table;
[0084] Step S302: Calculate the available resource capacity of each server based on the system load information in the load status cache table, evaluate the task volume of the log collection task, match the task volume with the available resource capacity, generate a task allocation plan, build a task scheduling queue based on the task allocation plan and the priority parameters in the server cluster configuration table, and distribute the task scheduling queue to the load monitoring agent program of each server to be monitored.
[0085] Optionally, this embodiment adopts a microservices architecture for the design of the load balancing scheduling center. The scheduling center process is implemented using the Spring Cloud framework and implements service discovery and registration through the service registry (Eureka). When the process starts, it first obtains the server cluster configuration from the configuration center. A hot configuration update mechanism is used to support dynamic configuration refresh, avoiding service interruptions caused by restarting services.
[0086] This embodiment implements a precise load information collection mechanism. Scheduled tasks are scheduled using the Quartz framework and configured with CronTrigger expressions to implement flexible scheduling strategies. The collection interval defaults to 30 seconds and can be adjusted dynamically based on actual needs. The collection process is asynchronous, using CompletableFuture to parallelize data acquisition requests from multiple servers, significantly improving collection efficiency.
[0087] This embodiment designs an efficient cache management strategy. The load status cache table is deployed using a Redis cluster, using a hash structure to store server load information. The key-value design uses the "serverid:timestamp" format to ensure data timeliness. A multi-level caching mechanism is also implemented, storing hot data in a local cache (Caffeine) to reduce network request overhead. Cache eviction uses a LRU algorithm and sets a reasonable expiration time to prevent cached data from expiring.
[0088] This embodiment builds an innovative resource capacity calculation model. Available resource capacity is comprehensively evaluated through multi-dimensional indicators, and the calculation formula is:
[0089] Capacity = min(1-CPURatio, 1-MemRatio, 1-IOWeight, 1-NetWeight) BaseCapacity represents the server's basic processing capacity. Each metric is normalized to ensure dimensional consistency. A fluctuation coefficient is also introduced, using exponential smoothing to predict short-term resource usage trends.
[0090] This embodiment implements an intelligent task load evaluation mechanism. The evaluation of log collection tasks takes into account multiple dimensions: log file size, update frequency, parsing complexity, etc. The evaluation results are converted into standardized resource requirements, and the calculation formula is:
[0091] TaskLoad = FileSize UpdateFreq ComplexityFactor, where ComplexityFactor is dynamically determined based on the log format type. Task evaluation results are cached locally and updated regularly to avoid repeated calculations.
[0092] This embodiment designs a precise task matching algorithm. Using an improved Best Fit algorithm, tasks are sorted in descending order of resource requirements, prioritizing allocation to servers with the closest remaining resources. The matching process considers server location and network latency, prioritizing servers in the same region. When resource contention arises, arbitration is performed based on task priority to ensure resource availability for critical tasks.
[0093] This example builds a reliable task scheduling queue. The queue is implemented as a priority queue, supporting dynamic adjustment of task priorities. The queue is stored in Redis using a Sorted Set structure. The score value is generated by combining the task priority and creation time, ensuring a first-in-first-out order for tasks of the same priority. The queue supports batch operations, improving task dispatch efficiency.
[0094] This embodiment implements an efficient task distribution mechanism. The distribution process uses a push-pull model. While the dispatch center actively pushes tasks to the agent, the agent periodically pulls task updates. Task distribution utilizes a resumable transmission mechanism to ensure reliable transmission even under network fluctuations. It also provides real-time feedback on task execution status and supports automatic retry of abnormal tasks.
[0095] Through the above technological innovations, this embodiment effectively solves the key problems in distributed log collection: inaccurate load balancing, unreasonable task allocation, and low scheduling efficiency. In practical applications, this solution can support the dynamic scheduling needs of large-scale server clusters, and significantly improves the overall performance of the system through accurate load assessment and intelligent task allocation. It is particularly suitable for microservice architecture environments, and through flexible scheduling strategies and reliable task distribution, it ensures the real-time and integrity of log collection. The systematic and innovative nature of this solution enables it to adapt to the application needs of enterprises of different sizes, and through continuous optimization and improvement, it achieves a comprehensive improvement in the efficiency of the scheduling system.
[0096] In one embodiment of the distributed server cluster log processing method of the present application, the following contents may also be specifically included:
[0097] Step S401: obtaining the network address and port number of the server to be monitored based on the task scheduling queue, establishing a persistent connection between the management terminal and the server to be monitored via the websocket protocol, storing the persistent connection in a connection pool according to the server identifier, creating a connection pool manager, and monitoring the connection status, automatically reestablishing disconnected connections, and maintaining the availability of the connection pool;
[0098] Step S402: Create a log file reading position record table, which contains the log file path, file size, last read timestamp and read position offset fields. Before reading the log file each time, obtain the offset of the last read position from the log file reading position record table, and use the offset as the starting position of a new round of reading. After the reading is completed, update the timestamp and offset in the log file reading position record table.
[0099] Optionally, this embodiment uses a hierarchical handshake mechanism during the WebSocket persistent connection establishment process. First, a WebSocket connection is established using an HTTP Upgrade request, with the request header containing the Upgrade field and the Sec-WebSocket-Key field. After verifying the request's legitimacy, the server returns a response containing the Sec-WebSocket-Accept field, completing the protocol upgrade. Identity authentication is performed immediately after the connection is established, using the JSON Web Token (JWT) mechanism to ensure connection security.
[0100] This embodiment implements an innovative connection pool management strategy. The connection pool uses a multi-level grouping structure, grouping by server cluster, geographic region, and business type. Each group maintains an independent connection counter and status monitor. Core connection pool parameters include the maximum number of connections, the minimum number of idle connections, and the connection lifetime, which can be dynamically adjusted through the configuration center. A connection reuse mechanism is also implemented, prioritizing existing connections for requests to the same target server.
[0101] This embodiment designs a reliable connection status monitoring mechanism. This monitoring uses a heartbeat detection scheme, sending a ping frame every 30 seconds. If no pong response is received three times in a row, the connection is considered disconnected. The heartbeat interval can be dynamically adjusted based on network quality, shortening it appropriately in poor network conditions. Connection quality assessment is also implemented, measuring round-trip time (RTT) and packet loss rate to score connections, prioritizing reestablishment of low-scoring connections.
[0102] This embodiment establishes an efficient reconnection mechanism. The reconnection strategy uses an exponential backoff algorithm, with an initial reconnection interval of 1 second. This interval doubles after each failure, with a maximum interval limit of 60 seconds. The connection context, including authentication status and session data, is saved during the reconnection process. Upon successful reconnection, the previous session state can be quickly restored. Concurrent reconnection control is also implemented to prevent excessive server pressure from initiating too many reconnection requests simultaneously.
[0103] This embodiment implements a precise file location record mechanism. The location record table uses a distributed storage solution and MongoDB. The document structure includes fields such as _id (a unique identifier generated by the server ID and file path), filePath (the absolute path of the log file), fileSize (the current file size), lastReadTime (the last read timestamp), and offset (the read position offset). A composite index is also established to optimize query performance.
[0104] This embodiment designs an innovative location update strategy. It uses an optimistic locking mechanism to handle concurrent updates, controlling the atomicity of update operations through the version number field. Update operations use atomic update commands to ensure data consistency in high-concurrency scenarios. It also implements an incremental update mechanism that only updates changed fields, reducing data transmission. To prevent data loss, update operations use a write concern mechanism to ensure that data is written to a sufficient number of replicas.
[0105] This embodiment establishes a reliable file reading mechanism. The reading process uses a buffer with a default buffer size of 8KB, which can be dynamically adjusted based on the file size. The file pointer is positioned using a seek operation, supporting reading from a specified offset. File status detection is also implemented to verify whether the file has been modified or deleted before reading, preventing invalid data from being read. For appended log files, the tail -f mode is used to continuously monitor file changes.
[0106] This embodiment implements a complete exception handling process. When a file read exception occurs, detailed error information is recorded, including the file status and system error code. A retry mechanism is used for temporary errors (such as file locks). For permanent errors (such as file deletions), the upper-layer application is promptly notified and the location record status is updated. Data consistency verification is also implemented, using a checksum mechanism to verify the integrity of the read data.
[0107] Through the above technical innovations, this embodiment effectively solves the key problems in distributed log collection: complex connection management, inaccurate location records, unreliable file reading, etc. In actual applications, this solution can support the real-time log collection needs of large-scale server clusters, and significantly improves the reliability and efficiency of data collection through reliable connection management and precise location records. It is particularly suitable for microservice architecture environments, and ensures the integrity and continuity of log data through flexible connection pool management and reliable file reading mechanisms. The systematic and innovative nature of this solution enables it to adapt to the application needs of enterprises of different sizes, and through continuous optimization and improvement, it achieves a comprehensive improvement in the log collection system.
[0108] In one embodiment of the distributed server cluster log processing method of the present application, the following contents may also be specifically included:
[0109] Step S501: Obtain an offset from the log file reading position record table, use the offset as the file pointer position to read new log data, perform line segmentation processing on the read log data, calculate the hash value of each line of data, obtain the hash value of historical data from the log data cache according to a preset time window range, compare the hash value of the current data with the hash value of the historical data, remove duplicate data, and write the deduplication result into the compression buffer;
[0110] Step S502: Execute the LZ4 compression algorithm on the log data in the compression buffer to generate a compressed data block, add a data identification header to the compressed data block, the data identification header includes the data block size, compression algorithm type and checksum information, send the compressed data block to the management terminal via the websocket long connection, select the corresponding decompression algorithm according to the data identification header, decompress the compressed data at the management terminal, and write the decompressed data into the log storage area.
[0111] Optionally, this embodiment implements an efficient incremental read mechanism based on file pointer position. First, the log file is mapped to memory using mmap technology to avoid frequent disk I / O operations. The read process employs a dual-buffer design: one buffer for data reading and the other for data processing. Buffer swapping enables continuous data processing. The buffer size is dynamically adjusted based on the frequency of file updates, with the buffer capacity appropriately increased in high-frequency write scenarios.
[0112] This embodiment implements a precise line segmentation algorithm. Taking into account the differences in line ending characters across operating systems, it supports three line ending methods: LF (\n), CR (\r), and CRLF (\r\n). The segmentation process utilizes zero-copy technology, avoiding data duplication through character pointer operations. It also supports multiple character sets, correctly processing log files encoded in formats such as UTF-8 and GBK, ensuring accurate line segmentation.
[0113] This embodiment employs an innovative hash calculation method. To improve computational efficiency, the MurmurHash3 algorithm is employed, which features high speed and low collision rate. Timestamp information is incorporated into the hash calculation process, serving as a seed value for the calculation, enhancing the uniqueness of the hash value. Furthermore, parallel hash calculations are implemented, allowing multiple threads to simultaneously process different data blocks, significantly improving processing speed.
[0114] This embodiment establishes a reliable duplicate data detection mechanism. The time window uses a sliding window design with a configurable window size, defaulting to 5 minutes. Hash values within the window are stored using a Bloom filter, which reduces the false positive rate through the use of multiple hash functions. The Bloom filter parameters are dynamically calculated based on the expected data volume, maintaining an optimal ratio between the bitmap size and the number of hash functions. Data outside the time window is automatically cleared to prevent continued memory usage growth.
[0115] This embodiment implements an efficient compression strategy. The LZ4 compression algorithm was chosen based on its fast compression and decompression characteristics, making it particularly suitable for text data with a high degree of repetitive patterns, such as log data. The compression process uses a streaming mode, with a data block size of 64KB, which strikes a balance between compression ratio and processing latency. Dynamic adjustment of the compression level is also implemented, selecting different compression levels based on CPU load.
[0116] This embodiment designs a complete data identification header structure. The header includes fields such as a magic number (for rapid data block identification), version number, compression algorithm type, original data size, compressed size, timestamp, and checksum. The checksum uses the CRC32 algorithm to verify the compressed data and ensure the integrity of data transmission. The header is designed to support backward compatibility, facilitating future expansion of new compression algorithms or the addition of new metadata fields.
[0117] This embodiment establishes a reliable data transmission mechanism. WebSocket data frames use binary mode, avoiding the encoding overhead of text mode. Flow control is implemented during transmission, using a sliding window mechanism to control the sending rate and prevent buffer overflows on the receiving end. Data frames are also transmitted in fragments, with the size of each fragment limited to 1MB, ensuring smooth transmission.
[0118] This embodiment implements an efficient decompression and restoration process. After receiving the data, the management terminal first verifies the integrity and correctness of the data identification header. The corresponding decompression method is selected based on the compression algorithm type. LZ4 decompression uses streaming processing, supporting the parallel decompression of multiple data blocks. The decompressed data undergoes format verification before being written to the log storage area to ensure data validity. The storage area adopts a sharding storage strategy, sharding the data by time range to facilitate subsequent query and analysis.
[0119] Through the above technological innovations, this embodiment effectively solves the key problems in log collection: data duplication, low transmission efficiency, large decompression delay, etc. In practical applications, this solution can support the real-time log collection needs of large-scale server clusters, and significantly reduces network bandwidth occupancy and storage space consumption through efficient deduplication compression and reliable transmission mechanisms. It is particularly suitable for high-concurrency microservice environments, and through streaming processing and parallel optimization, it ensures the real-time and integrity of log collection. The systematic and innovative nature of this solution enables it to adapt to the application needs of enterprises of different sizes, and through continuous optimization and improvement, it achieves a comprehensive improvement in the log collection system.
[0120] In one embodiment of the distributed server cluster log processing method of the present application, the following contents may also be specifically included:
[0121] Step S601: Read historical log data from the log storage area, perform text preprocessing on the historical log data, extract the timestamp, event type, operation object, status code, execution result, error code, user identifier, and operation instruction in the log to form a feature set, convert the feature set into a vector representation, use an LSTM neural network to build an anomaly detection model and a pattern recognition model, use a graph neural network to build an event association analysis model, use the feature vectors to train the anomaly detection model, pattern recognition model, and event association analysis model, and save the trained model to the model library;
[0122] Step S602: Load the anomaly detection model, pattern recognition model and event association analysis model from the model library to build a log analysis engine, convert the restored log data into a feature vector and input it into the log analysis engine, the anomaly detection model calculates the anomaly score of the feature vector, the pattern recognition model performs multi-classification prediction on the log content and outputs a category label, and the event association analysis model mines the temporal association and causal relationship between log events based on the graph structure, and generates abnormal event warning information based on the anomaly score, category label and association relationship.
[0123] Optionally, this embodiment first implements an efficient log text preprocessing process. A regular expression engine is used for structured parsing, establishing a matching rule base for log templates of different formats. The parsing process considers various timestamp formats and supports automatic recognition and conversion of standard formats such as ISO8601 and Unix timestamps. For unstructured log content, a sliding window and word frequency statistics method are used to extract key information to ensure the integrity of feature extraction.
[0124] This embodiment employs an innovative feature engineering solution. Time features are parsed through timestamp analysis to obtain specific time granularity information, including periodic features such as hours, weeks, and months. Event types and operation objects are converted into dense vectors using word embedding technology, and a 300-dimensional vector representation is obtained using Word2Vec training. Status codes and error codes are converted using one-hot encoding, while retaining the original numerical information, making it easier to capture correlations between error codes.
[0125] This example constructs a deep anomaly detection model. The LSTM network adopts a bidirectional architecture, consisting of two LSTM layers, each with 128 hidden units. The network input is a time series feature sequence, with each time step containing a vector representation of all features. Training utilizes a contrastive learning strategy, using normal sequences as positive samples and artificially constructed anomalous sequences as negative samples. The loss function combines contrastive loss and reconstruction loss to improve the model's sensitivity to anomalous patterns.
[0126] This embodiment implements a precise pattern recognition model. Based on an LSTM-based sequence classification network, an attention mechanism is added above the LSTM layer to adaptively focus on important time steps and feature dimensions. Attention weights are calculated using a softmax function, enabling dynamic adjustment of feature importance. The model output layer uses a multi-label classification architecture, supporting simultaneous recognition of multiple category labels for a single log entry, improving classification flexibility.
[0127] This embodiment designs an innovative graph neural network model. The nodes of the event correlation graph represent log events, and the edges represent the relationships between events. Node features contain all event attribute information, and edge weights are calculated based on time intervals and feature similarity. The graph neural network uses a graph attention network (GAT) structure, using a multi-head attention mechanism to learn complex dependencies between nodes, effectively capturing causal relationships within event chains.
[0128] This example implements an efficient model training mechanism. Training data is organized using a sliding window approach, with the window size adaptively adjusted based on temporal correlation. The batch size is set to 64, and the Adam optimizer is used for parameter updates. The learning rate adopts a warm-up strategy, gradually increasing to 0.001 at the beginning of training and then decaying using cosine annealing. An early stopping mechanism is also implemented, terminating training if validation set performance does not improve for five consecutive epochs.
[0129] This embodiment builds a complete model deployment framework. The model library utilizes a distributed storage architecture, supporting model version management and rapid rollback. Model loading utilizes a lazy loading strategy, dynamically loading required models based on actual needs. The inference process uses the ONNX format for model optimization, improving inference efficiency through operator fusion and quantization optimization. A hot update mechanism is also implemented, supporting online updates without impacting services.
[0130] This embodiment designs a reliable anomaly warning mechanism. Anomaly scores are mapped to the [0, 1] interval through probability distribution transformation, and dynamic thresholds are set using expert rules. Category prediction results are compared with historical statistical distributions to identify changes in anomaly category distributions. Event correlation analysis results are used to construct a fault propagation diagram and predict potential fault spread paths. Warning information includes anomaly description, impact scope, and handling suggestions, supporting multi-level alerting strategies.
[0131] Through the above technological innovations, this embodiment effectively solves key problems in intelligent log analysis: incomplete feature extraction, inaccurate anomaly detection, insufficient correlation analysis, etc. In practical applications, this solution can accurately identify system anomalies and provide increasingly accurate analysis results through continuous optimization of deep learning models. It is particularly suitable for operation and maintenance monitoring of large-scale distributed systems. Through multi-dimensional intelligent analysis, it significantly improves the efficiency of problem location and fault prevention. The systematic and innovative nature of this solution enables it to adapt to different types of log analysis needs. Through continuous learning and optimization of the model, it achieves a comprehensive improvement in log analysis capabilities.
[0132] In one embodiment of the distributed server cluster log processing method of the present application, the following contents may also be specifically included:
[0133] Step S701: Create a management terminal visualization panel, divide the abnormal event warning information into three levels: high, medium, and low according to the abnormality level, and display it in the warning information area. The log classification and annotation results are displayed in the form of a pie chart to show the proportion distribution of each category. The association analysis results are displayed in the form of a force-directed graph to show the association strength between log event nodes. A time range selector is created in the visualization panel, and the visualization data is updated in real time according to the start and end time of the time range selector.
[0134] Step S702: Listen for the user's click event of the download button in the visualization panel, obtain the start and end time parameters in the time range selector, query the abnormal event warning information, log classification and annotation results and correlation analysis results within the time range from the database, convert the query results into JSON format, execute the ZIP compression algorithm on the JSON format data to generate a compressed file, and save the compressed file to the local storage path specified by the user.
[0135] Optionally, this embodiment adopts a modular architecture for the design of the visualization panel. The front-end application is built on the React framework, and Redux is used for state management, enabling data flow and state synchronization between various visualization components. The panel layout adopts a responsive design, using CSS Grid and Flexbox to achieve adaptive layout, ensuring good display effects on different screen sizes.
[0136] This embodiment implements a precise warning information display mechanism. Anomaly levels are classified based on a multi-dimensional scoring system, including impact scope, duration, and business importance. High-level alarms are highlighted in red with a flashing animation, while medium-level alarms are highlighted in orange and low-level alarms in yellow. The warning information area uses a card-based layout. Each alarm card contains information such as time, type, description, and handling suggestions, and can be expanded to view detailed information.
[0137] This example designs a dynamic visualization of classification results. The pie chart, implemented using the ECharts library, supports multi-level category display, with the inner ring displaying the main categories and the outer ring showing the subcategory distribution. The chart interactivity supports click-to-filter and drill-down analysis of categories, with detailed statistical information displayed upon hovering. The color scheme is professionally designed to ensure good readability and aesthetics. The chart also implements automatic scaling to adapt to display areas of varying sizes.
[0138] This example demonstrates innovative visualization of association analysis. The force-directed graph is implemented using the D3.js library. Node size indicates the importance of an event, while color indicates the event type. Edge thickness and color indicate the strength of the association, while arrows indicate the temporal relationship between events. The graph layout utilizes a force-directed algorithm, automatically adjusting node positions to avoid overlap. Graph zooming, panning, and node dragging are also supported, enhancing the interactive experience.
[0139] This embodiment implements a reliable time range selector. The selector uses a dual-slider design, supporting time range selection with precision down to the second. Time selection supports shortcuts such as the last hour, today, and this week. Changes in the selector trigger real-time data updates, and the update process uses anti-shake processing to avoid frequent data requests. Automatic time zone recognition and conversion are also implemented, ensuring the correct time is displayed in different regions.
[0140] This embodiment designs an efficient data update mechanism. It uses a persistent WebSocket connection to implement server push. When new analysis results are available, the server proactively pushes data to the front-end. Data updates utilize an incremental update strategy, transmitting only the changed parts, reducing network traffic. A data caching mechanism is also implemented, caching data for commonly used time ranges in the browser to improve query response speed.
[0141] This example implements a complete data export function. It monitors the download button's click event and uses throttling to avoid repeated clicks. Data queries use paging to avoid memory overflows caused by fetching too much data at once. Data cleansing is performed during JSON format conversion, removing unnecessary fields and optimizing the data structure. Compression is implemented using Web Workers to avoid blocking the main thread and improve the user experience.
[0142] This embodiment implements a reliable file saving mechanism. It uses the native File System Access API to save files, allowing users to select the save path. Large files are saved using streaming, writing fragments through Blob objects to avoid excessive memory usage. It also implements save progress notifications and error retry mechanisms, improving file saving reliability. For browsers that don't support the new API, a compatible download solution is provided.
[0143] Through the above technical innovations, this embodiment effectively solves the key problems in log analysis visualization: non-intuitive display effects, poor interactive experience, inconvenient data export, etc. In actual applications, this solution can intuitively display the system operation status, and significantly improve the work efficiency of operation and maintenance personnel through rich visualization components and smooth interactive experience. It is particularly suitable for monitoring scenarios of large-scale distributed systems, and meets analysis needs at different levels through multi-dimensional data display and flexible data export. The systematic and innovative nature of this solution enables it to adapt to different types of visualization needs, and through continuous optimization and improvement, it achieves a comprehensive improvement of the log analysis system.
[0144] In order to effectively solve the shortcomings of traditional technologies in distributed collection, data transmission and intelligent analysis, and significantly improve the performance and practicality of the log processing system, the present application provides an embodiment of a distributed server cluster log processing device for implementing all or part of the content of the distributed server cluster log processing method, see Figure 2 The distributed server cluster log processing device specifically includes the following contents:
[0145] A dynamic allocation module 10 is configured to establish a distributed server cluster log collection network, create a server cluster configuration table on a management terminal, the server cluster configuration table including the network address, port number, access credentials, and server load threshold of each server to be monitored, establish a load balancing scheduling center based on the server cluster configuration table, and periodically collect system load information of each server to be monitored, dynamically allocate log collection tasks based on the system load information, and generate a task scheduling queue according to a preset task priority.
[0146] The log processing module 20 is used to build an incremental log data collection channel, establish a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, record the latest read position of each log file, perform incremental data collection according to the latest read position, perform real-time compression encoding on the collected log data, detect duplicate data based on a sliding time window, remove the duplicate data, and then transmit the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0147] The log analysis module 30 is used to build an intelligent log analysis model, use deep learning methods to train historical log data, build an analysis engine including an anomaly detection model, a pattern recognition model and an event association analysis model, and input the restored log data into the analysis engine. The anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event association analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information. The abnormal event warning information, log classification and labeling results and correlation analysis results are visualized in the management terminal interface. When it is detected that the user triggers the download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
[0148] From the above description, it can be seen that the distributed server cluster log processing device provided in the embodiment of the present application can dynamically allocate collection tasks through the load balancing scheduling center by building a distributed server cluster log collection network. Design an incremental data collection channel, use the websocket long connection pool to achieve efficient transmission, and combine real-time compression coding and duplicate data detection mechanism to optimize transmission efficiency. Construct an intelligent log analysis model, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, it realizes anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis, and provides visual display and data export functions. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system.
[0149] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies in distributed data collection, data transmission, and intelligent analysis, and significantly improve the performance and practicality of log processing systems, the present application provides an embodiment of an electronic device for implementing all or part of the content of the distributed server cluster log processing method. The electronic device specifically includes the following content:
[0150] A processor, memory, communication interface, and bus; wherein the processor, memory, and communication interface communicate with each other via the bus; the communication interface is used to implement information transmission between the distributed server cluster log processing device and related devices such as the core business system, user terminals, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited to this. In this embodiment, the logic controller can be implemented with reference to the embodiments of the distributed server cluster log processing method and the embodiments of the distributed server cluster log processing device in the embodiments, the contents of which are incorporated herein and repeated parts are not repeated.
[0151] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0152] In practical applications, portions of the distributed server cluster log processing method can be executed on the electronic device side as described above, or all operations can be performed on the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.
[0153] The aforementioned client device may include a communication module (i.e., a communication unit) capable of establishing a communication connection with a remote server to facilitate data transmission with the server. The server may include a server at the task scheduling center or, in other implementation scenarios, a server on an intermediate platform, such as a server on a third-party server platform that is communicatively linked to the task scheduling center server. The server may comprise a single computer device, a server cluster consisting of multiple servers, or a distributed server configuration.
[0154] Figure 3 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 3 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0155] In one embodiment, the distributed server cluster log processing method function may be integrated into the central processing unit 9100. The central processing unit 9100 may be configured to perform the following control:
[0156] Step S101: Establishing a distributed server cluster log collection network, creating a server cluster configuration table on a management terminal, the server cluster configuration table including the network address, port number, access credentials, and server load threshold of each server to be monitored, establishing a load balancing scheduling center based on the server cluster configuration table, the load balancing scheduling center periodically collecting system load information of each server to be monitored, dynamically allocating log collection tasks based on the system load information, and generating a task scheduling queue according to a preset task priority;
[0157] Step S102: constructing an incremental log data collection channel, establishing a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, recording the latest read position of each log file, performing incremental data collection according to the latest read position, performing real-time compression encoding on the collected log data, detecting duplicate data based on a sliding time window, removing the duplicate data, and transmitting the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0158] Step S103: Construct an intelligent log analysis model, use deep learning methods to train historical log data, build an analysis engine including an anomaly detection model, a pattern recognition model and an event correlation analysis model, input the restored log data into the analysis engine, the anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information, and visualize the abnormal event warning information, log classification and labeling results and correlation analysis results in the management terminal interface. When it is detected that the user triggers the download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
[0159] As can be seen from the above description, the electronic device provided in the embodiment of the present application constructs a distributed server cluster log collection network and dynamically allocates collection tasks through a load balancing scheduling center. An incremental data collection channel is designed, and a websocket long connection pool is used to achieve efficient transmission. The transmission efficiency is optimized by combining real-time compression coding and duplicate data detection mechanisms. An intelligent log analysis model is constructed, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis are implemented, and visual display and data export functions are provided. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system.
[0160] In another embodiment, the distributed server cluster log processing device can be configured separately from the central processor 9100. For example, the distributed server cluster log processing device can be configured as a chip connected to the central processor 9100, and the distributed server cluster log processing method function can be implemented through the control of the central processor.
[0161] like Figure 3As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 3 In addition, the electronic device 9600 may also include all components shown in Figure 3 For components not shown, reference may be made to the prior art.
[0162] like Figure 3 As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.
[0163] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.
[0164] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.
[0165] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), or SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is capable of storing additional data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs, or processes used by the central processing unit 9100 to execute operations of the electronic device 9600.
[0166] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, images, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0167] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.
[0168] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless local area network modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130, providing audio output via the speaker 9131 and receiving audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.
[0169] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the distributed server cluster log processing method in the above-mentioned embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the distributed server cluster log processing method in the above-mentioned embodiment, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:
[0170] Step S101: Establishing a distributed server cluster log collection network, creating a server cluster configuration table on a management terminal, the server cluster configuration table including the network address, port number, access credentials, and server load threshold of each server to be monitored, establishing a load balancing scheduling center based on the server cluster configuration table, the load balancing scheduling center periodically collecting system load information of each server to be monitored, dynamically allocating log collection tasks based on the system load information, and generating a task scheduling queue according to a preset task priority;
[0171] Step S102: constructing an incremental log data collection channel, establishing a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, recording the latest read position of each log file, performing incremental data collection according to the latest read position, performing real-time compression encoding on the collected log data, detecting duplicate data based on a sliding time window, removing the duplicate data, and transmitting the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0172] Step S103: Construct an intelligent log analysis model, use deep learning methods to train historical log data, build an analysis engine including an anomaly detection model, a pattern recognition model and an event correlation analysis model, input the restored log data into the analysis engine, the anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information, and visualize the abnormal event warning information, log classification and labeling results and correlation analysis results in the management terminal interface. When it is detected that the user triggers the download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
[0173] As can be seen from the above description, the computer-readable storage medium provided in the embodiment of the present application constructs a distributed server cluster log collection network and dynamically allocates collection tasks through a load balancing scheduling center. An incremental data collection channel is designed, and a websocket long connection pool is used to achieve efficient transmission. The transmission efficiency is optimized by combining real-time compression coding and duplicate data detection mechanisms. An intelligent log analysis model is constructed, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis are implemented, and visual display and data export functions are provided. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system.
[0174] The embodiments of the present application also provide a computer program product capable of implementing all steps of the distributed server cluster log processing method in the above-mentioned embodiment, where the execution subject is a server or a client. When the computer program / instructions are executed by a processor, the steps of the distributed server cluster log processing method are implemented. For example, the computer program / instructions implement the following steps:
[0175] Step S101: Establishing a distributed server cluster log collection network, creating a server cluster configuration table on a management terminal, the server cluster configuration table including the network address, port number, access credentials, and server load threshold of each server to be monitored, establishing a load balancing scheduling center based on the server cluster configuration table, the load balancing scheduling center periodically collecting system load information of each server to be monitored, dynamically allocating log collection tasks based on the system load information, and generating a task scheduling queue according to a preset task priority;
[0176] Step S102: constructing an incremental log data collection channel, establishing a websocket persistent connection pool between the management terminal and each monitored server based on the task scheduling queue, recording the latest read position of each log file, performing incremental data collection according to the latest read position, performing real-time compression encoding on the collected log data, detecting duplicate data based on a sliding time window, removing the duplicate data, and transmitting the compressed encoded data to the management terminal, where the received compressed encoded data is decompressed and restored;
[0177] Step S103: Construct an intelligent log analysis model, use deep learning methods to train historical log data, build an analysis engine including an anomaly detection model, a pattern recognition model and an event correlation analysis model, input the restored log data into the analysis engine, the anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information, and visualize the abnormal event warning information, log classification and labeling results and correlation analysis results in the management terminal interface. When it is detected that the user triggers the download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
[0178] As can be seen from the above description, the computer program product provided by the embodiment of the present application constructs a distributed server cluster log collection network and dynamically allocates collection tasks through a load balancing scheduling center. An incremental data collection channel is designed, and a websocket long connection pool is used to achieve efficient transmission. The transmission efficiency is optimized by combining real-time compression coding and duplicate data detection mechanisms. An intelligent log analysis model is constructed, which includes three sub-models: anomaly detection, pattern recognition, and event correlation analysis. Based on deep learning methods, anomaly score calculation, log classification and labeling, and multi-dimensional correlation analysis are implemented, and visual display and data export functions are provided. This method effectively solves the shortcomings of traditional technologies in distributed collection, data transmission, and intelligent analysis, and significantly improves the performance and practicality of the log processing system.
[0179] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatuses, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0180] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0181] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0183] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A distributed server cluster log processing method, characterized in that: The method comprises: Establish a distributed server cluster log collection network, create a server cluster configuration table on the management terminal, the server cluster configuration table contains the network address, port number, access credentials and server load threshold of each server to be monitored, establish a load balancing scheduling center based on the server cluster configuration table, the load balancing scheduling center periodically collects system load information of each server to be monitored, dynamically allocates log collection tasks based on the system load information, and generates a task scheduling queue according to preset task priorities; Construct an incremental log data collection channel, establish a websocket long connection pool between the management terminal and each monitored server based on the task scheduling queue, create a log file reading position record table, the log file reading position record table includes the log file path, file size, last read timestamp and reading position offset fields, obtain the offset of the last read position from the log file reading position record table before each log file reading, use the offset as the starting position of a new round of reading, update the timestamp and offset in the log file reading position record table after the reading is completed, obtain the offset from the log file reading position record table, use the offset as the file pointer position to read the newly added log data, perform real-time compression encoding on the collected log data, detect duplicate data based on a sliding time window, transmit the compressed encoded data to the management terminal after removing the duplicate data, and decompress and restore the received compressed encoded data at the management terminal; An intelligent log analysis model is constructed, and deep learning methods are used to train historical log data. An analysis engine including an anomaly detection model, a pattern recognition model, and an event correlation analysis model is constructed. The restored log data is input into the analysis engine. The anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and labels the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information. The abnormal event warning information, log classification and labeling results, and correlation analysis results are visualized in the management terminal interface. When a user triggers a download operation, the analysis result data within the specified time range is packaged, compressed, and downloaded to local storage.
2. The distributed server cluster log processing method according to claim 1, characterized in that: The distributed server cluster log collection network is established, and a server cluster configuration table is created on the management terminal. The server cluster configuration table contains the network address, port number, access credentials and server load threshold of each server to be monitored, including: Enter the server cluster information in the configuration interface of the management terminal, save the server cluster information to the database to generate a server cluster configuration table, which contains the server identifier, network address, port number, access credentials, load threshold and priority fields. Read the network address and port number in the server cluster configuration table to establish a TCP connection, verify the validity of the access credentials through the SSH protocol, and mark the server node as available after verification. Build a load monitoring agent program and deploy it to each server to be monitored. The load monitoring agent program collects CPU usage, memory occupancy, disk IO and network bandwidth data, calculates the server's comprehensive load score, compares the comprehensive load score with the load threshold in the server cluster configuration table, and issues a warning signal to the server that exceeds the load threshold. Adjust the collection task priority of the server according to the warning signal.
3. The distributed server cluster log processing method according to claim 1, characterized in that: The load balancing scheduling center is established based on the server cluster configuration table, and the load balancing scheduling center periodically collects system load information of each server to be monitored, dynamically allocates log collection tasks according to the system load information, and generates a task scheduling queue according to a preset task priority, including: Create a load balancing scheduling center process, read the load threshold and priority parameters in the server cluster configuration table, establish a system load information collection timer task, and obtain CPU usage, memory occupancy, disk I / O, and network bandwidth data from the load monitoring agent program of each monitored server at a preset time interval, and store the system load information in a load status cache table; Based on the system load information in the load status cache table, the available resource capacity of each server is calculated, the task volume of the log collection task is evaluated, the task volume size is matched with the available resource capacity, and a task allocation plan is generated. According to the task allocation plan and the priority parameters in the server cluster configuration table, a task scheduling queue is constructed, and the task scheduling queue is distributed to the load monitoring agent program of each server to be monitored.
4. The distributed server cluster log processing method according to claim 1, characterized in that: The step of constructing an incremental log data collection channel and establishing a websocket persistent connection pool between the management terminal and each server to be monitored based on the task scheduling queue includes: Based on the task scheduling queue, the network address and port number of the server to be monitored are obtained, a long connection is established between the management terminal and the server to be monitored through the websocket protocol, the long connection is stored in the connection pool according to the server identifier, and a connection pool manager is created. The connection pool manager monitors the connection status, automatically reestablishes the disconnected connection, and maintains the availability of the connection pool.
5. The distributed server cluster log processing method according to claim 4, characterized in that: The method includes performing real-time compression encoding on the collected log data, detecting duplicate data based on a sliding time window, removing the duplicate data, and transmitting the compressed encoded data to a management terminal, and decompressing and restoring the received compressed encoded data at the management terminal, including: The read log data is split into rows, the hash value of each row of data is calculated, the hash value of the historical data is obtained from the log data cache according to the preset time window range, the hash value of the current data is compared with the hash value of the historical data, and the deduplication result is written to the compression buffer after removing duplicate data; Execute the LZ4 compression algorithm on the log data in the compression buffer to generate a compressed data block, add a data identification header to the compressed data block, the data identification header includes the data block size, compression algorithm type and checksum information, send the compressed data block to the management terminal through the websocket long connection, select the corresponding decompression algorithm on the management terminal according to the data identification header to decompress the compressed data, and write the decompressed data into the log storage area.
6. The distributed server cluster log processing method according to claim 1, characterized in that: The intelligent log analysis model is constructed, and a deep learning method is used to train historical log data. An analysis engine including an anomaly detection model, a pattern recognition model, and an event correlation analysis model is constructed. The restored log data is input into the analysis engine. The anomaly detection model calculates anomaly scores based on log feature vectors, the pattern recognition model classifies and labels log content, and the event correlation analysis model mines the correlation relationships between multi-dimensional logs to generate abnormal event warning information, including: Read historical log data from the log storage area, perform text preprocessing on the historical log data, extract the timestamp, event type, operation object, status code, execution result, error code, user identifier, and operation instruction in the log to form a feature set, convert the feature set into a vector representation, use an LSTM neural network to build an anomaly detection model and a pattern recognition model, use a graph neural network to build an event association analysis model, use the feature vector to train the anomaly detection model, pattern recognition model, and event association analysis model, and save the trained model to a model library; The anomaly detection model, pattern recognition model and event association analysis model are loaded from the model library to build a log analysis engine. The restored log data is converted into a feature vector and input into the log analysis engine. The anomaly detection model calculates the anomaly score of the feature vector. The pattern recognition model performs multi-classification prediction on the log content and outputs a category label. The event association analysis model mines the temporal association and causal relationship between log events based on the graph structure, and generates abnormal event warning information based on the anomaly score, category label and association relationship.
7. The distributed server cluster log processing method according to claim 1, characterized in that: The abnormal event warning information, log classification and annotation results, and correlation analysis results are visually displayed in the management terminal interface. When a user triggers a download operation, the analysis result data within a specified time range is packaged and compressed and downloaded to the local storage, including: Create a management terminal visualization panel, classify the abnormal event warning information into three levels: high, medium, and low according to the abnormality level, and display it in the warning information area. The log classification and annotation results are displayed in the form of a pie chart to show the distribution of each category. The association analysis results are displayed through a force-directed graph to show the association strength between log event nodes. Create a time range selector in the visualization panel, and update the visualization data in real time according to the start and end times of the time range selector; Listen for the user's download button click event in the visualization panel, obtain the start and end time parameters in the time range selector, query the abnormal event warning information, log classification and annotation results and correlation analysis results within the time range from the database, convert the query results into JSON format, execute the ZIP compression algorithm on the JSON format data to generate a compressed file, and save the compressed file to the local storage path specified by the user.
8. A distributed server cluster log processing device, characterized in that: The device comprises: A dynamic allocation module is used to establish a distributed server cluster log collection network, create a server cluster configuration table on the management terminal, and the server cluster configuration table contains the network address, port number, access credentials and server load threshold of each server to be monitored. A load balancing scheduling center is established based on the server cluster configuration table. The load balancing scheduling center periodically collects system load information of each server to be monitored, dynamically allocates log collection tasks based on the system load information, and generates a task scheduling queue according to the preset task priority. A log processing module is used to build an incremental log data collection channel, establish a websocket long connection pool between the management terminal and each monitored server based on the task scheduling queue, create a log file reading position record table, the log file reading position record table includes the log file path, file size, last read timestamp and reading position offset fields, obtain the offset of the last read position from the log file reading position record table before each log file reading, use the offset as the starting position of a new round of reading, update the timestamp and offset in the log file reading position record table after the reading is completed, obtain the offset from the log file reading position record table, use the offset as the file pointer position to read the newly added log data, perform real-time compression encoding on the collected log data, detect duplicate data based on a sliding time window, transmit the compressed encoded data to the management terminal after removing the duplicate data, and decompress and restore the received compressed encoded data at the management terminal; The log analysis module is used to build an intelligent log analysis model, use deep learning methods to train historical log data, and build an analysis engine that includes an anomaly detection model, a pattern recognition model, and an event correlation analysis model. The restored log data is input into the analysis engine. The anomaly detection model calculates the anomaly score based on the log feature vector, the pattern recognition model classifies and annotates the log content, and the event correlation analysis model mines the correlation relationship between multi-dimensional logs to generate abnormal event warning information. The abnormal event warning information, log classification and annotation results, and correlation analysis results are visualized in the management terminal interface. When a user triggers a download operation, the analysis result data within the specified time range is packaged and compressed and downloaded to local storage.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the distributed server cluster log processing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the distributed server cluster log processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Real-time log collection and analysis method on basis of B2B (Business to Business) platform
CN105824744A
Time-limited automatic processing method for multi-source heterogeneous mass data
CN111124679A