Log analysis method and system for optimizing K-means clustering based on pigeon bionic algorithm
By optimizing K-means clustering using a pigeon-inspired biomimetic algorithm, the problems of unstable clustering quality and low computational efficiency in system log parsing are solved, achieving efficient and accurate log parsing results.
Patent Information
- Application Number
- CN202511242818.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies suffer from unstable clustering quality, high parameter sensitivity, and low computational efficiency in system log parsing. Furthermore, the selection of initial cluster centers has a significant impact on the results, making it difficult to determine the optimal number of clusters.
A method based on pigeon biomimetic algorithm to optimize K-means clustering is adopted. By initializing the pigeon population, the initial cluster center selection is optimized, and the fitness function is used to search for the cluster point with minimum loss. Combined with the K-means clustering algorithm for iterative optimization, a log template is finally generated.
It improves the stability and accuracy of clustering, enhances computational efficiency, avoids getting trapped in local optima, and achieves efficient and accurate log parsing.
Smart Images

Figure CN120929600A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer system log analysis technology, and in particular to a log parsing method and system based on pigeon biomimetic algorithm to optimize K-means clustering. Background Technology
[0002] System logs are an important source for checking the system status in software systems. The runtime status reports and error messages contained in system logs are widely used in system operation and maintenance. Currently, the mainstream method for parsing system logs is based on machine learning and utilizes clustering methods.
[0003] Mainstream methods for parsing system logs often utilize clusterers to analyze and divide the logs. This method has the following problems: 1) It is sensitive to the value of K, requiring a pre-set value for K, making it difficult to determine the optimal number of clusters; 2) It is sensitive to outliers and noise points, which may affect the clustering results of other points; 3) The selection of initial cluster centers has a significant impact on the results, and multiple attempts may be needed to obtain better clustering results. Summary of the Invention
[0004] In view of the above-mentioned defects in the prior art, the present invention aims to solve the technical problems of unstable clustering quality, high parameter sensitivity and low computational efficiency in traditional log parsing methods, and provide an efficient and accurate log parsing scheme.
[0005] To achieve the above objectives, this invention provides a log parsing method based on the pigeon-inspired bionic algorithm to optimize K-means clustering, comprising the following steps: The S100 uses a log collection tool to collect system logs; S200 preprocesses log data, including cleaning invalid entries and standardizing log formats; The S300 extracts log features and converts text information into numerical feature vectors; The S400 algorithm uses the pigeon optimization algorithm to optimize the initial cluster center selection for K-means clustering, including: S410 initializes the pigeon population, with each individual pigeon representing a set of candidate cluster centers. Data points to be classified are then randomly sampled as candidate points for clustering. S420 calculates the fitness value, and the fitness function is defined as the clustering loss function. The search algorithm is used to search for the cluster point with the minimum loss. S430 updates the positions based on the pigeon's role classification: discoverer, joiner, and watcher, and uses these cluster points as the initial cluster points for the clustering algorithm. S440 iterative optimization continues until the convergence condition is met, and the final cluster points are obtained using the partitioned clustering (K-means clustering) algorithm. S500 performs K-means clustering using optimized cluster centers; S600 generates log templates based on clustering results.
[0006] Furthermore, the formula for updating the discoverer's location in step S430 is as follows:
[0007] Where t represents the current iteration number, j = 1, 2, 3, ..., d is a natural number, is a constant, represents the maximum number of iterations (the number of iterations during each clustering, a convergence condition; if the number of iterations is less than iter.max, the convergence condition is met and iteration stops), represents the position information of the i-th pigeon in the j-th dimension, a is a random number, represents the warning value, ST represents the safety value, Q is a random number following a normal distribution, and L represents a matrix of type n, where each element in the matrix is 1.
[0008] Furthermore, the formula for updating the joiner's position in step S430 is:
[0009] Where is the best position currently occupied by the discoverer, represents the worst position globally, A represents a matrix where each element is randomly assigned a value of 1 or -1, and when i>n / 2, it indicates that the i-th joiner with a lower fitness value has not obtained food and is in a state of search completion. At this time, it needs to fly to other places to forage for food to obtain more energy; where X_{worst} is the worst position globally, and i is the pigeon index number.
[0010] Furthermore, the formula for updating the vigilant's position in step S430 is as follows:
[0011] Where X_{best} is the current global best position, K is a random number in the interval [-1, 1], rand is a random number in the interval [0, 1], and ε is a small constant to avoid division by zero.
[0012] Furthermore, the feature extraction in step S300 employs the TFIDF vectorization method.
[0013] Furthermore, the convergence condition in step S400 is: Reaching the preset maximum number of iterations T_max, or The fitness value changes less than the threshold δ after M consecutive iterations.
[0014] The present invention also provides a log parsing system, comprising: Log collection module: Used to collect system logs; Preprocessing module: used for cleaning and standardizing log data; Feature extraction module: used to convert log text into feature vectors; Optimize the clustering module: Use the pigeon-inspired algorithm to optimize the initial K-means centers; Template generation module: Generates log templates based on clustering results.
[0015] The present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes the program to implement the method described above.
[0016] This invention introduces a pigeon-inspired bionic algorithm into clustering calculations to avoid the original clustering method getting stuck in local optima during log information analysis. The advantages of this technical solution are fast convergence speed and high accuracy, thus improving the efficiency of analysis.
[0017] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of a preferred embodiment of the present invention. Detailed Implementation
[0019] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0020] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.
[0021] Algorithm explanation and terminology: The main idea of clustering algorithms is to group similar instances together by comparing their similarity. This is similar to the purpose of log parsing, which groups similar log messages (log messages generated from the same message template) together and extracts the common parts as message templates. Furthermore, since clustering algorithms involve a large amount of similarity calculation, we process the log data in the preceding step to reduce time complexity.
[0022] Analysis of a large amount of log data revealed that logs generated from the same message template are usually quite similar, which is related to the fact that they all have the same constant markers.
[0023] The log parsing algorithm based on clustering achieves log parsing through four steps. Preprocessing counts the message length of each message and replaces obvious marker variables in the message with simple regular expressions. Messages of the same length are grouped together. Within the same group, a clustering algorithm is used to group similar messages in the same group into a category. Since messages of the same category may still be generated by different message templates after clustering, the same category is further grouped after clustering so that each group corresponds to an independent print statement. Message templates / event types are generated in each group.
[0024] The principle of partitioning clustering algorithm: Partitioning clustering is an unsupervised classification algorithm used with an infinite number of datasets.
[0025] The task of this algorithm is to cluster the dataset into k clusters. The loss function to be minimized is:
[0026] in It is the cluster center point:
[0027] Finding the optimal solution to the above problem requires traversing all possible cluster partitions. The standard partitioning clustering algorithm uses a greedy strategy to find an approximate solution. The specific steps are: (1) Extract k sample points from the sample to serve as the center points of each cluster. (2) Calculate the distance between all sample points and each cluster center, and then put the sample points into the nearest cluster. (3) Recalculate the cluster center based on the existing sample points in the cluster. (4) Repeat steps two and three. The clustering results obtained by partitioning clustering algorithms heavily depend on the selection of initial cluster centers. If the initial cluster centers are not selected well, the algorithm will get trapped in local optima. Therefore, the selection of initial cluster centers is very important. So, it is necessary to optimize the initial cluster centers to improve the problem of easily getting trapped in local optima.
[0028] Clustering in optimization algorithms: The principle behind optimizing the search algorithm is based on the process of pigeons foraging for food: The nouns included: Pigeon: Refers to the probe factor in the algorithm, used to represent the basic existing elements in the matrix; Joiner: Actually, it's a state of a pigeon, used for moving element factors within a matrix; Eagle: refers to pigeons (element factors) in a monitoring matrix that meet certain conditions, and they are removed when the conditions are met. Predator: Same as eagle; Discoverer: The pigeon that found food (finds element factors that meet the clustering criteria); Watchdog: A pigeon that detects whether an eagle is nearby (a certain number of element factors in the matrix are used to determine whether the position meets the condition for deleting an element); (1) The discoverer usually has a high flying ability and is responsible for searching for areas with abundant food throughout the population, providing foraging areas and directions for all participants. In the model establishment, the level of energy reserves depends on the fitness value of the individual pigeon.
[0029] (2) Once a pigeon spots an eagle, it begins to send out alarm signals. When the alarm threshold exceeds the safe threshold, the spotter will lead the other pigeons to other safe areas to forage.
[0030] The roles of discoverers and joiners are dynamic. Every pigeon can become a discoverer as long as a better food source is found, but the proportion of discoverers and joiners in the overall population remains constant. In other words, if one pigeon becomes a discoverer, another pigeon will inevitably become a joiner.
[0031] (3) The lower the energy of the newcomers, the worse their foraging position in the population. Some hungry newcomers are more likely to fly to other places to forage for more energy.
[0032] (4) During foraging, participants are always able to find the discoverer who provides the best food, and then obtain food from the best food or forage around the discoverer. At the same time, some participants may constantly monitor the discoverer in order to increase their predation rate and compete for food resources.
[0033] (5) When they realize danger, pigeons on the edge of the flock will quickly move to a safe area to get a better position, while pigeons in the middle of the flock will move around randomly to get closer to other pigeons.
[0034] One embodiment of the present invention provides a log parsing method based on the pigeon biomimetic algorithm to optimize K-means clustering, comprising the following steps: The S100 uses a log collection tool to collect system logs; S200 preprocesses log data, including cleaning invalid entries and standardizing log formats; The S300 extracts log features and converts text information into numerical feature vectors; The S400 algorithm uses the pigeon optimization algorithm to optimize the initial cluster center selection for K-means clustering, including: S410 initializes the pigeon population, with each individual pigeon representing a set of candidate cluster centers. Data points to be classified are then randomly sampled as candidate points for clustering. S420 calculates the fitness value, and the fitness function is defined as the clustering loss function. The search algorithm is used to search for the cluster point with the minimum loss. S430 updates the positions based on the pigeon's role classification: discoverer, joiner, and watcher, and uses these cluster points as the initial cluster points for the clustering algorithm. S440 iterative optimization continues until the convergence condition is met, and the final cluster points are obtained using the partitioned clustering (K-means clustering) algorithm. S500 performs K-means clustering using optimized cluster centers; S600 generates log templates based on clustering results.
[0035] Preferably, the formula for updating the discoverer's location in step S430 is:
[0036] Where t represents the current iteration number, j = 1, 2, 3, ..., d is a natural number, a constant, and represents the maximum number of iterations (the number of iterations per clustering session; a convergence condition is met when the number of iterations is less than iter.max, at which point iteration stops). represents the position information of the i-th pigeon in the j-th dimension. a is a random number. represents the warning value, and ST represents the safety value. Q is a random number following a normal distribution. L represents a matrix of type , where each element is 1.
[0037] Preferably, the formula for updating the joiner's position in step S430 is:
[0038] Where is the best position currently occupied by the discoverer, and represents the worst position globally. A represents a matrix , where each element is randomly assigned a value of 1 or -1, and . When i > n / 2, this indicates that the i-th joiner with a lower fitness value has not obtained food and is in a state of search completion. At this time, it needs to fly to other places to forage for food to obtain more energy; where X_{worst} is the worst position globally, and i is the pigeon index number.
[0039] Preferably, the formula for updating the vigilant's position in step S430 is:
[0040] Where X_{best} is the current global best position, K is a random number in the interval [-1, 1], rand is a random number in the interval [0, 1], and ε is a small constant to avoid division by zero.
[0041] Preferably, the feature extraction in step S300 uses the TFIDF vectorization method.
[0042] Preferably, the convergence condition in step S400 is: Reaching the preset maximum number of iterations T_max, or The fitness value changes less than the threshold δ after M consecutive iterations.
[0043] Preferably, the present invention also provides an embodiment of a log parsing system including the above methods, comprising: Log collection module: Used to collect system logs; Preprocessing module: used for cleaning and standardizing log data; Feature extraction module: used to convert log text into feature vectors; Optimize the clustering module: Use the pigeon-inspired algorithm to optimize the initial K-means centers; Template generation module: Generates log templates based on clustering results.
[0044] The present invention preferably provides an electronic device, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes to implement the above methods and systems.
[0045] (1) Initialize parameters, including the number of clusters and the number of iterations; (2) Optimize clustering using the pigeon-inspired algorithm: Initialize pigeon positions, including setting the number of clusters, population size (maximum and minimum boundaries of the set), and maximum number of iterations; Both the Discoverer and the Watcher are types of pigeons; see the glossary for details. The default ratio of Discoverers to Watchers is 20%, and the default safety threshold is 80% (factors that do not move when the calculation condition is below 0.8 will be removed). Each pigeon represents a cluster center; (3) Calculate pigeon fitness Fitness is calculated using the inverse of WCSS, where WCSS minimizes the sum of squares within the cluster. (4) Update pigeon location Cluster labels are added to the original data to evaluate the silhouette coefficient (which measures the similarity (density) of samples within the same cluster and the separation between samples from different clusters). Analyze the clustering results and update the positions of the pigeons (candidate cluster centers); (5) Check the termination conditions; The sample is selected cyclically. When a pigeon is among the top 20% of discoverers, it moves closer to the optimal solution; otherwise, it joins the current solution and randomly perturbs it. Boundary handling is performed to ensure that the center point is within the data range. (6) Output the best cluster center, i.e. run the final K-means with the optimal solution.
[0046] Then output the best cluster centers; (7) Application of cluster centers This invention introduces a pigeon-inspired bionic algorithm into clustering calculations to avoid the original clustering method getting stuck in local optima during log information analysis. The advantages of this technical solution are fast convergence speed and high accuracy, thus improving the efficiency of analysis.
[0047] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A log parsing method based on pigeon-inspired bionic algorithm to optimize K-means clustering, characterized in that, Includes the following steps: The S100 uses a log collection tool to collect system logs; S200 preprocesses log data, including cleaning invalid entries and standardizing log formats; The S300 extracts log features and converts text information into numerical feature vectors; The S400 algorithm applies the pigeon optimization algorithm to optimize the initial cluster center selection for K-means clustering, including: S410 initializes the pigeon population, with each individual pigeon representing a set of candidate cluster centers. Randomly sampled data points to be classified are used as candidate points for clustering. S420 calculates the fitness value, and the fitness function is defined as the clustering loss function. The search algorithm is used to search for the cluster point with the minimum loss. S430 updates the positions based on the pigeon's role classification: discoverer, joiner, and watcher. These cluster points are used as the initial cluster points for the clustering algorithm. S440 iterative optimization continues until the convergence condition is met, and the final cluster points are obtained using the partitioned clustering (K-means clustering) algorithm. S500 performs K-means clustering using optimized cluster centers; S600 generates log templates based on clustering results.
2. The method according to claim 1, characterized in that, The formula for updating the discoverer's location in step S430 is: Where t represents the current iteration number, j = 1, 2, 3, ..., d is a natural number, is a constant, represents the maximum number of iterations (the number of iterations during each clustering, a convergence condition; if the number of iterations is less than iter.max, the convergence condition is met and iteration stops), represents the position information of the i-th pigeon in the j-th dimension, a is a random number representing the warning value, ST represents the safety value, Q is a random number following a normal distribution, and L represents a matrix where each element in the matrix is 1.
3. The method according to claim 1, characterized in that, The formula for updating the joiner's position in step S430 is: Where is the best position currently occupied by the discoverer, and represents the worst position globally. A represents a matrix where each element is randomly assigned a value of 1 or -1. When i > n / 2, it indicates that the i-th joiner with a lower fitness value has not obtained food and is in a state of completed search. At this time, it needs to fly to other places to forage for food to obtain more energy. Where X_{worst} is the worst position globally, and i is the pigeon index number.
4. The method according to claim 1, characterized in that, The formula for updating the vigilant's position in step S430 is as follows: Where X_{best} is the current global best position, K is a random number in the interval [-1, 1], rand is a random number in the interval [0, 1], and ε is a small constant to avoid division by zero.
5. The method according to claim 1, characterized in that, The feature extraction in step S300 uses the TFIDF vectorization method.
6. The method according to claim 1, characterized in that, The convergence condition in step S400 is: The fitness value changes less than the threshold δ after reaching the preset maximum number of iterations T_max, or after M consecutive iterations.
7. A log parsing system, characterized in that, include: Log collection module: Used to collect system logs; Preprocessing module: used for cleaning and standardizing log data; Feature extraction module: used to convert log text into feature vectors; Optimize the clustering module: Use the pigeon-inspired algorithm to optimize the initial K-means centers; Template generation module: Generates log templates based on clustering results.
8. An electronic device comprising a processor, a memory, and a computer program stored in the memory, characterized in that, When the processor executes the program, it implements the method of any one of claims 1-6.