Log classification method and computing equipment
By calculating the text similarity between logs and log samples by computing the computing device, log classification is automated, and the problem of high complexity in log type recognition in the prior art is solved, and efficient and accurate log classification is achieved.
Patent Information
- Application Number
- CN202510460231.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-26
AI Technical Summary
In the existing message subscription system, log classification requires manual definition of log identifiers, which has certain technical thresholds, and it is difficult to accurately identify logs of different log types for the same producer.
The text similarity between the log to be classified and the log sample of the sender is calculated through the computing device, and the log type is determined using the bag-of-word model and similarity threshold, which automates the log classification process.
It improves the efficiency and accuracy of log classification, simplifies the operation process, reduces the complexity of manually defining log identifiers, and ensures the accuracy of log type recognition.
Smart Images

Figure CN120541223A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information processing technology, and in particular to a log classification method and computing device. Background Art
[0002] Producers in the message subscription system publish logs to the message cluster through topics, and consumers subscribe to logs through topics. To facilitate consumer subscriptions, the message subscription system normalizes logs from different producers. However, logs of different log types from the same producer may be published to the same topic, so logs need to be categorized before normalization.
[0003] Currently, the message subscription system achieves log classification by manually entering log identifiers for different log types from different producers. However, log identifiers require manual definition, which has certain technical barriers. Summary of the Invention
[0004] The embodiments of the present application provide a log classification method and a computing device that can accurately identify the type of log.
[0005] To achieve the above technical objectives, this application adopts the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a log classification method, which is applied to a computing device, and the method includes: in response to receiving a log to be classified, querying a log list corresponding to the sender of the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender; and determining the type of the log to be classified based on the similarity between the log to be classified and the log sample.
[0007] It is understandable that the computing device calculates the text similarity between the log to be classified and the log sample corresponding to the log type of the sender of the log to be classified, and then determines the log type of the log to be classified based on the text similarity, and achieves the effect of accurate log classification by comparing the text similarity between the log and the log sample of its sender. In related solutions, it is necessary to manually pre-construct log identifiers of different log types to achieve the purpose of log classification. Log identifiers need to be manually defined, which has a certain technical threshold. Therefore, the computing device identifies the type of log by the text similarity between the log and the log sample, which is more convenient and easy to operate than identifying the type of log based on manually constructed log identifiers in related solutions.
[0008] In a possible implementation, the method further includes:
[0009] According to the type of the log to be classified, the log to be classified is sent to the corresponding log processor, and the log processor is used to process the log to be classified. It can be understood that after determining the type of the log to be classified, the log to be classified is sent to different log processors for normalization processing according to the log type to improve the efficiency of log normalization. In a possible implementation, the above-mentioned determination of the type of the log to be classified based on the similarity between the log to be classified and the log sample includes: using a bag-of-words model to determine the first text vector corresponding to the log to be classified, wherein the bag-of-words model is constructed based on the log fields included in the historical log data; based on the similarity between the first text vector and the second text vector corresponding to the log sample, determining the type of the log to be classified, wherein the second text vector is determined in advance based on the bag-of-words model.
[0010] It's understandable that the computing device compares the log fields in the log to words in text, pre-constructs a bag-of-words model, and uses this to pre-determine the second text vector corresponding to each log sample from each sender. When performing log classification tasks, the pre-stored second text vector corresponding to each log sample from the sender to be classified is directly called, eliminating the need to determine the text vector corresponding to each log sample in real time, thereby improving log classification efficiency.
[0011] In another possible implementation, the above-mentioned determination of the type of the log to be classified based on the similarity between the first text vector and the second text vector corresponding to the log sample includes: determining the log sample corresponding to the second text vector having the greatest similarity to the first text vector as the target log sample; and determining the type of the log to be classified based on the target log sample.
[0012] It is understandable that the sender may have two logs of different log types that contain very similar log fields. In this case, the classification and identification of the two types of logs are very easy to confuse. Based on this, the computing device uses the log type corresponding to the log sample that is most similar to the text vector of the log to be classified as the most likely log type of the log to be classified, eliminating the mutual interference in the classification and identification of logs with similar log fields, and improving the accuracy of log classification. In related schemes, the two types of logs with very similar log fields are distinguished by increasing the complexity of the log identifier, which further increases the difficulty of defining the log identifier. Compared with related schemes, this implementation method achieves accurate classification on the basis of simple operation.
[0013] In another possible implementation, the above-mentioned determination of the type of the log to be classified based on the target log sample includes: in response to determining that the similarity between the log to be classified and the target log sample is greater than a target similarity threshold, determining that the log type of the log to be classified is the log type corresponding to the target log sample, wherein the target similarity threshold is a similarity threshold bound to the sender.
[0014] It is understandable that the computing device sets a similarity threshold for each sender. When classifying the sender's logs, only the log types corresponding to the log samples in the log list corresponding to the sender, whose text similarity with the log to be classified is greater than the similarity threshold set by the sender, can be regarded as the log types of the log to be classified, thereby further improving the accuracy of log classification.
[0015] In another possible implementation, the above method also includes: in response to determining the log type of the log to be classified, determining the number of fields that match the log to be classified and the target log sample; determining the classification accuracy score of the log to be classified based on the proportion of the number to the total number of fields in the log to be classified, wherein the classification accuracy score is proportional to the proportion; and optimizing the target similarity threshold based on the classification accuracy score.
[0016] It is understandable that the computing device performs a classification accuracy score on each log to be classified, and uses the classification accuracy score of each sender's log to be classified to optimize the similarity threshold of each sender, thereby improving the log classification accuracy by continuously optimizing the similarity threshold of each sender.
[0017] In another possible implementation, the above-mentioned optimization of the target similarity threshold based on the classification accuracy score includes: determining the average classification accuracy score of the logs corresponding to the sender within the optimization period in which the logs to be classified are located; in response to the average value being lower than a preset score, raising the target similarity threshold.
[0018] It is understandable that for each sender, when the average classification accuracy score of the logs to be classified by the computing device within the optimization period is lower than the preset score, that is, when the classification accuracy is low, the target similarity threshold is raised to reduce the probability of log classification errors.
[0019] In another possible implementation, the method further includes: in response to the log type of the log to be classified not being determined, storing the log to be classified in an unidentified log list corresponding to the sender.
[0020] Understandably, log messages are generated very quickly. Within a short period of time, a large number of log messages from different senders are published to the message subscription system. Some of these logs have difficult-to-identify log types. The computing device stores these logs in the sender's unidentified list. Each sender's unidentified list is subsequently analyzed to refine the log list corresponding to each sender. This continuous improvement helps maintain the long-term accuracy of log classification.
[0021] In another possible implementation, the above method also includes: dividing the logs in the unidentified log list into different clusters based on text vector similarity matching; extracting a predetermined number of logs to be classified from the clusters; obtaining the actual log type of the extracted logs to be classified and a standard log sample corresponding to the actual log type; and updating the log list based on the actual log type and the standard log sample.
[0022] It is understandable that most of the logs in the sender's unidentified log list belong to a limited number of log types, and the reasons why logs of the same log category are difficult to be identified in a short period of time are likely to be the same. Therefore, the computing device divides the logs in the sender's unidentified log list into different clusters based on text vector similarity matching, selects one or several logs to be classified from each cluster, and uses the actual log types of the selected logs to be classified and the standard log samples corresponding to the actual log types to update the sender's log list, thereby improving the efficiency of updating the log list.
[0023] In another possible implementation, updating the log list based on the real log type and the standard log sample includes: in response to determining that the real log type does not exist in the log list, updating the real log type and the standard log sample to the log list.
[0024] It is understandable that one reason why the log type of a log to be classified is difficult to identify is that the sender has added a log of a new log type but has not synchronized it with the computing device. Therefore, when updating the log list, the computing device first determines whether the actual log type of the selected log to be classified is recorded in the log list. If not, the actual log type and its corresponding standard log sample are updated to the log list to eliminate the possibility of log type identification failure due to the log list not recording the newly added log type and its corresponding log sample.
[0025] In another possible implementation, the above-mentioned updating of the log list based on the real log type and the standard log sample includes: in response to determining that the real log type exists in the log list, and the log sample corresponding to the real log type is inconsistent with the standard log sample, updating the log sample corresponding to the real log type to the standard log sample.
[0026] It is understandable that one reason why the log type of the log to be classified is difficult to identify is that the sender has updated the log sample corresponding to the log type but has not synchronized it with the computing device. Therefore, after the computing device determines that the actual log type of the selected log to be classified is recorded in the log list, it determines whether the log sample corresponding to the actual log type in the log list is inconsistent with the standard log sample corresponding to the actual log type. If so, the log sample corresponding to the actual log type is updated to the standard log sample, eliminating the possibility of log type identification failure due to the log list not recording the latest log sample corresponding to the log type.
[0027] In a second aspect, the present application provides a computing device, which includes modules applied to the method of the first aspect or any possible design method of the first aspect.
[0028] In a third aspect, the present application provides a computing device comprising a memory and a processor. The memory and the processor are coupled; the memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the computing device performs the log classification method according to the first aspect and any possible implementation thereof.
[0029] In a fourth aspect, the present application provides a computing device comprising a processor, wherein the processor executes the log classification method of the first aspect and any possible implementation thereof.
[0030] Illustratively, the computing device may be a server, a tablet computer, a desktop, a laptop, a notebook computer, a netbook, and the like.
[0031] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer instructions, wherein when the computer instructions are executed on a computing device, the computing device executes the log classification method of the first aspect and any possible implementation thereof.
[0032] In a sixth aspect, the present application provides a computer program product comprising computer instructions, wherein when the computer instructions are executed on a computing device, the computing device is caused to execute the log classification method of the first aspect and any possible implementation thereof.
[0033] For the specific descriptions of the second to sixth aspects and their various implementations in this application, reference can be made to the detailed descriptions in the first aspect and its various implementations; and for the beneficial effects of the second to sixth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be repeated here.
[0034] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A schematic diagram of an implementation environment involved in a log classification method provided in an embodiment of the present application;
[0036] Figure 2 A schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application;
[0037] Figure 3 A flow chart of a log classification method provided in an embodiment of the present application;
[0038] Figure 4 A flow chart of another log classification method provided in an embodiment of the present application;
[0039] Figure 5 A flow chart of another log classification method provided in an embodiment of the present application;
[0040] Figure 6 A flow chart of another log classification method provided in an embodiment of the present application;
[0041] Figure 7 A schematic diagram of the structure of a log classification device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] In the following, to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean different.
[0043] Meanwhile, in the description of this application, unless otherwise specified, "plurality" means two or more than two. Words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.
[0044] For ease of understanding, the following briefly introduces the relevant terms involved in the embodiments of this application:
[0045] (1) Message subscription system (e.g., Apache Kafka): It is a distributed event streaming system used to build real-time data pipelines and streaming applications. It can efficiently process large amounts of data streams and supports multiple data processing models, such as message queues and publish-subscribe models. Message subscription systems are widely used in many fields, such as log aggregation, website activity tracking, metric collection, and real-time data processing.
[0046] (2) Topic: In a message subscription system, a Topic is used to represent the category or channel of a message. Simply put, a Topic is the name of a message queue. Producers publish messages to the message subscription system using a specified Topic name, while consumers subscribe to and consume messages from the Topic.
[0047] (3) Apache NiFi: is a flow-based programming and data integration tool that supports powerful and configurable data flow automation.
[0048] (4) Log: A file that records the time, operations, and activities that occur in a computer system, including errors, problems, or current operating information. Logs can be used to monitor and understand system operation, debug problems, or conduct audits.
[0049] (5) Log normalization: converting logs from different senders into a unified format.
[0050] (6) Cosine similarity algorithm: Cosine similarity is a method for calculating the similarity between two vectors in a multidimensional space. It is commonly used in fields such as text mining, information retrieval, and recommendation systems to evaluate the similarity between documents or to compare the similarity between different feature vectors. The range of cosine similarity is [-1, 1], where 1 indicates exact similarity, -1 indicates exact opposite, and 0 indicates no correlation.
[0051] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0052] To facilitate consumer subscriptions, the message subscription system normalizes different logs from different producers. However, logs of different log types from the same producer may be published through the same topic. Therefore, the log type of the log must be identified before normalization.
[0053] In the embodiment of the present application, the text similarity between the log to be classified and the log sample corresponding to the log type of the sender of the log to be classified is calculated, and the log type of the log to be classified is determined based on the text similarity. By comparing the text similarity between the log and the log sample of its sender, the effect of accurate log classification is achieved. The operation is simple and easy to implement.
[0054] For further information, please refer to Figure 1 , which shows a schematic diagram of the implementation environment involved in a log classification method provided in an embodiment of the present application. Figure 1 As shown, the implementation environment may include: a producer 100, a first server 110, a terminal 120, a second server 130 and a consumer 140;
[0055] A producer 100 refers to a sending device set up at a client (i.e., log sender) of the message subscription system. Each sending device publishes the logs generated during its operation to the message subscription system through the topic of the message subscription system. A topic may receive different types of logs from the same sending device. The types of logs include but are not limited to general alarm logs, email network management logs, terminal summary logs, database logs, certificate logs, and audit logs.
[0056] The first server 110 can be a server or server cluster running a message subscription system. It sends the logs to be classified in the message subscription system to the terminal 120 to determine the log type corresponding to the log through the terminal 120; then normalizes the log according to the log type, and returns the normalized log to the message subscription system for subscription by multiple consumers 140.
[0057] The terminal 120 can determine the sending device identifier of the log to be classified in response to receiving the log to be classified, and send the sending device identifier to the second server 130, so that the second server 130 queries the sender corresponding to the sending device identifier, and queries the log list corresponding to the sender; it can also receive the sender and its corresponding log list fed back by the second server 130, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender; and determine the type of the log to be classified based on the similarity between the log to be classified and the log sample.
[0058] The second server 130 can provide a query service server, which can query the sender and its corresponding log list according to the device identifier sent by the terminal 120, and feed back the sender and its corresponding log list to the terminal 120.
[0059] Illustratively, the terminal 120 in the embodiment of the present application may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a notebook computer, a netbook computer, etc. The embodiment of the present application does not impose any special restrictions on the specific form of the terminal 120.
[0060] The first server 110 and the second server 130 can be independent physical servers, or a server cluster or distributed file system composed of multiple physical servers, or at least one of the cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data or artificial intelligence platforms. The embodiments of the present application are not limited to this.
[0061] It should be noted that the log classification method provided in the embodiment of the present application can be applied to the terminal 120, the second server 130, or the terminal 120 and the second server 130. The terminal 120 and the second server 130 can be collectively referred to as electronic devices.
[0062] Figure 2 This is a hardware structure diagram of a computing device provided in an embodiment of the present application. Figure 2 , Figure 2 The computing device shown may include: a processor 201 , a memory 202 , a communication interface 203 , and a bus 204 . The processor 201 , the memory 202 , and the communication interface 203 may be connected via the bus 204 .
[0063] The processor 201 is the control center of the computing device, and can be a general-purpose central processing unit such as a CPU, or other general-purpose processors, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0064] As an example, the processor 201 may include one or more CPUs, such as Figure 2 CPU 0 and CPU 1 are shown in Figure 1.
[0065] The memory 202 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0066] In one possible implementation, memory 202 may exist independently of processor 201. Memory 202 may be connected to processor 201 via bus 204 and used to store data, instructions, or program code. When processor 201 calls and executes the instructions or program code stored in memory 202, the log classification method provided in the embodiments of the present application can be implemented.
[0067] In another possible implementation, the memory 202 may also be integrated with the processor 201 .
[0068] The communication interface 203 is used to connect the computing device to other devices via a communication network, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 203 may include a receiving unit for receiving data and a sending unit for sending data.
[0069] The bus 204 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of presentation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0070] It should be pointed out that Figure 2 The structure shown in the figure does not constitute a limitation on the computing device, except Figure 2 In addition to the components shown, the computing device may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0071] Logs of the same log type published by multiple producers in a message subscription system may have different log formats. To standardize management, logs of the same log type but different log formats need to be normalized. Because the same topic may receive logs of different log types from the same producer, it is difficult to distinguish log types based on the topic. Therefore, log classification methods are needed to complete log classification.
[0072] Figure 3 This is a flow chart of a log classification method provided in an embodiment of the present application. Figure 3, applied to a computing device, the message subscription system may be run by a server or a server cluster, and the method includes the following S301-S302:
[0073] S301: In response to receiving a log to be classified, query a log list corresponding to the sender of the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender.
[0074] One or more sending devices are set at a client (i.e., log sender) of the message subscription system. The log types and log formats corresponding to the logs generated by one or more sending devices of a log sender during operation are basically the same. Therefore, with one sender as a unit, a log list corresponding to each sender is pre-established. The log list lists log samples corresponding to each log type that exists for the sender.
[0075] In one possible implementation, the sender corresponding to the device ID of each sending device can be pre-stored. Upon receiving a log to be classified, the sender of the log to be classified can be queried based on the device ID of the sending device, and then the log list corresponding to the sender can be retrieved. The device ID is an important component of device management, used to uniquely identify each device. It can include basic information such as the device name, model, and installation address.
[0076] Alternatively, in another possible implementation, when a sending device publishes a log to a topic in a message subscription system, the log carries the corresponding sender identifier. In response to receiving a log to be categorized, the sender identifier is parsed from the log to be categorized, and a list of logs corresponding to the sender identifier is queried. Each sender is pre-assigned an identifier that uniquely identifies it, and the identifier is communicated to each sender so that the sender can indicate the source of the log when publishing it.
[0077] S302: Determine the type of the log to be classified based on the similarity between the log to be classified and the log sample.
[0078] The similarity between the log to be classified and the log sample may be the text vector similarity between the log to be classified and the log sample. Before determining the type of the log to be classified based on the similarity between the log to be classified and the log sample, it is necessary to determine the text vectors corresponding to the log to be classified and the log sample.
[0079] In some possible implementations, the determining of the text vectors corresponding to the log to be classified and the log sample includes: determining the text vectors corresponding to the log to be classified and the log sample by constructing a bag-of-words model.
[0080] The process of constructing a bag-of-words model includes: obtaining a certain number of historical logs from a message subscription system; extracting log fields from the historical logs; using all the log fields that have appeared to construct a field set, wherein each log field has a unique number; initializing a bag-of-words vector, wherein the length of the bag-of-words vector is the number of fields in the field set; generating a bag-of-words model based on the field set and the bag-of-words vector, wherein when determining the text vector of each log, the bag-of-words model adds 1 (or the number of times the log field appears in the log) to the corresponding number position in the bag-of-words vector for the log field that appears in the log, thereby generating a text vector corresponding to the log.
[0081] Exemplarily, a bag-of-words model is constructed using two historical logs, where historical log one is "I like watching movies" and historical log two is "I love watching movies very much". Log field extraction is performed on historical log one to obtain "I", "like" and "watch movies"; log field extraction is performed on historical log two to obtain "I", "very", "love" and "watch movies". All the log fields that have appeared are used to construct a field set of "I", "like", "very", "love" and "watch movies", and the numbers correspond to 1, 2, 3, 4, and 5 respectively; a bag-of-words vector with the same length as the number of fields in the field set is established, and a bag-of-words model is constructed based on the field set and the bag-of-words vector. When determining the text vector of each log, the bag-of-words model adds 1 to the corresponding number position in the bag-of-words vector for the log field that appears in the log (or the number of times the log field appears in the log). The bag-of-words model is used to determine that the text vector corresponding to historical log one is [1,1,0,0,1], and the text vector corresponding to historical log two is [1,0,1,1,1].
[0082] Alternatively, the determining of the text vectors corresponding to the log to be classified and the log sample includes: determining the text vectors corresponding to the log to be classified and the log sample by using a method such as one-hot encoding or word embedding.
[0083] Among them, the principle of one-hot encoding is to construct a field set from a certain number of historical logs. Each log field in the field set has a corresponding number, and each log field is assigned a unique binary vector, in which only one position is 1 and the other positions are 0. The position of this 1 corresponds to the number of the log field in the field set. By determining the one-hot encoding of each log field appearing in the log, the text vector corresponding to the log is determined. The principle of word embedding is based on mapping each log field into a high-dimensional real vector space, and generating the text vector corresponding to the log by learning the co-occurrence relationship between log fields. Word embedding methods include: Word2Vec, GloVe, FastText, etc. One-hot encoding and word embedding are both existing text vector identification methods, and will not be elaborated on in detail here.
[0084] In some possible implementations, determining the type of the log to be classified based on the similarity between the log to be classified and the log sample includes: determining a first text vector corresponding to the log to be classified, and a second text vector corresponding to the log sample; determining the type of the log to be classified based on the similarity between the first text vector and the second text vector, and a similarity threshold corresponding to the sender of the log to be classified.
[0085] The similarity between the first text vector and the second text vector may be cosine similarity.
[0086] Cosine similarity is a method that measures the cosine value of the angle between two vectors and is used to evaluate the similarity of texts in the vector space model. The formula for calculating cosine similarity is:
[0087]
[0088] Where A·B is the dot product of vector A and vector B, ‖A‖ and ‖B‖ are the modulos of vector A and vector B, respectively.
[0089] For example, the text vector 1 corresponding to the history log 1 is [1, 1, 0, 0, 1], and the text vector 2 corresponding to the history log 2 is [1, 0, 1, 1, 1]. The dot product of the text vector 1 and the text vector 2 is 1×1+1×0+0×1+0×1+1×1=2, and the modulus of the text vector 1 is The modulus of text vector 2 is The cosine similarity between text vector 1 and text vector 2 is
[0090] Alternatively, the similarity between the first text vector and the second text vector may be Jacard similarity.
[0091] Alternatively, the similarity between the first text vector and the second text vector may be edit distance similarity.
[0092] The above-mentioned calculation methods of Jacard similarity and edit distance similarity are prior arts and will not be described in detail here.
[0093] The above-mentioned process of determining the type of the log to be classified based on the similarity between the first text vector and the second text vector and the similarity threshold corresponding to the sender of the log to be classified includes:
[0094] Log samples are selected in order from top to bottom in the log list until the log sample meets a first preset condition, where the first preset condition is that the similarity between the second text vector corresponding to the log sample and the first text vector is greater than or equal to a similarity threshold corresponding to the sender of the log to be classified; and the log type corresponding to the log sample that meets the first preset condition is determined to be the log type of the log to be classified.
[0095] It is important to note that if the similarity between the log to be classified and a log sample is greater than or equal to the similarity threshold corresponding to the sender of the log to be classified, it means that the log to be classified is highly likely to belong to the same log type as the log sample. Therefore, the log type of the log to be classified can be determined by searching for log samples of the same log type as the log to be classified in the order of log samples from top to bottom in the log list. This eliminates the need to calculate the text similarity between each log sample in the log list and the log to be classified, thus reducing the amount of similarity calculations.
[0096] In some possible implementations, after determining the log type of the log to be classified, the log to be classified may be sent to a corresponding log processor according to the log type of the log to be classified. The log processor is used to process the log to be classified.
[0097] As an example: The method provided in the embodiment of the present application can be applied to a NiFi cluster. The NiFi cluster may include multiple log processors, each of which performs normalization processing in a different way. After determining the log type of the log to be classified, the NiFi cluster distributes the log to be classified and its corresponding log type to the corresponding log processor for log processing according to the log type.
[0098] The technical solution provided in the embodiment of the present application calculates the text similarity between the log to be classified and the log samples corresponding to all log types of the sender of the log to be classified, and then determines the log type of the log to be classified based on the text similarity. By comparing the text similarity between the log and the log samples of its sender, the effect of accurate log classification is achieved. Compared with the solution in the related solution that identifies the type of log based on manually defined log identifiers, this solution avoids the impact of the difficulty in defining log identifiers on log type identification and is simpler to operate.
[0099] Figure 4 This is a flow chart of another log classification method provided in the embodiment of the present application. Figure 4 , applied to a computing device, the method includes: S401-S405:
[0100] S401: In response to receiving a log to be classified, query a log list corresponding to the sender of the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender.
[0101] S402 : Determine a first text vector corresponding to the log to be classified using a bag-of-words model, wherein the bag-of-words model is constructed based on log fields included in historical log data.
[0102] S403 : Calculate the similarity between the first text vector and the second text vector corresponding to each log sample in the log list, wherein the second text vector is determined in advance based on a bag-of-words model.
[0103] The above steps S401-S403 are as follows Figure 3 Steps S301 and S302 of the illustrated embodiment have been described in detail and will not be repeated here.
[0104] In this embodiment, the second text vector corresponding to each log sample of each sender is pre-stored. When executing a log classification task, the pre-stored second text vector corresponding to each log sample of the sender to be classified is directly called, eliminating the need to determine the text vector of each log sample of the sender to be classified in real time, thereby improving the efficiency of log classification.
[0105] S404: Determine the log sample corresponding to the second text vector having the greatest similarity to the first text vector as the target log sample.
[0106] In this embodiment, the sender may have two logs of different log types that contain very similar log fields. In this case, the classification and identification of the two types of logs are easily confused. Taking this situation into consideration, the log type corresponding to the log sample that is most similar to the text vector of the log to be classified is used as the most likely log type of the log to be classified. This eliminates the interference of similar log fields in log classification and identification, and improves the accuracy of log classification. In related solutions, the two types of logs with very similar log fields can only be distinguished by increasing the complexity of the log identifier, which further increases the difficulty of defining the log identifier. Compared with related solutions, this embodiment achieves accurate classification while being easy to operate.
[0107] S405 . In response to determining that the similarity between the log to be classified and the target log sample is greater than a target similarity threshold, determine the log type of the log to be classified as the log type corresponding to the target log sample, wherein the target similarity threshold is a similarity threshold bound to the sender.
[0108] In this embodiment, a similarity threshold is set for each sender. For this sender, only when the similarity between the log to be classified and the target log sample is greater than the corresponding similarity threshold, that is, when the classification accuracy is high, can the log type corresponding to the target log sample be regarded as the log type of the log to be classified, thereby further improving the log classification accuracy.
[0109] The technical solution provided by the embodiments of the present application determines the selection of a log sample from the log list of the sender of the log to be classified that meets a second preset condition, wherein the second preset condition is: the highest similarity to the log to be classified, and the similarity to the log to be classified is greater than or equal to the similarity threshold of the sender of the log to be classified. In other words, the second preset condition is used to find the log sample most likely to belong to the same log type as the log to be classified, and the log type of the log sample is then determined as the log type of the log to be classified, thereby improving log classification accuracy.
[0110] Figure 5 This is a flow chart of another log classification method provided in the embodiment of the present application. Figure 5 , applied to a computing device, the method includes: S501-S506:
[0111] S501: In response to receiving a log to be classified, querying a log list corresponding to the sender of the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender.
[0112] S502: Determine the type of the log to be classified according to the similarity between the log to be classified and the log sample, and the similarity threshold corresponding to the sender of the log to be classified.
[0113] The above steps S501-S502 are as follows Figure 3 Steps S301 and S302 of the illustrated embodiment have been described in detail and will not be repeated here.
[0114] S503: In response to determining the log type of the log to be classified, determine the number of fields that match the log to be classified and the target log sample.
[0115] In some possible implementations, the number of fields that match the log to be classified and the target log sample may be the number of fields that appear simultaneously in the log to be classified and the target log sample.
[0116] S504: Determine a classification accuracy score of the log to be classified based on the proportion of the number in the total number of fields of the log to be classified, wherein the classification accuracy score is proportional to the proportion.
[0117] In some possible implementations, determining the classification accuracy score of the log to be classified based on the proportion of the number of fields in the total number of fields in the log to be classified includes:
[0118] Based on the proportion and a pre-set classification accuracy score calculation formula, the classification accuracy score of the log to be classified is determined, wherein the classification accuracy score calculation formula can be any function that can ensure that the proportion is proportional to the classification accuracy score, such as a linear function.
[0119] For example, the classification accuracy score calculation formula can be:
[0120]
[0121] Among them, Score is the classification accuracy score, Score min is the lower limit of the classification accuracy score, Score max is the upper limit of the classification accuracy score, M is the number of fields that match the target log sample and the log to be classified, and N is the number of fields in the target log sample.
[0122] S505: In response to the end time of the optimization period of the log to be classified, determining an average value of the classification accuracy scores of the logs corresponding to the sender of the log to be classified within the optimization period.
[0123] In some possible implementations, a division standard for the optimization period may be preset, for example, a certain period of time is an optimization period, or an optimization period is when the number of logs of an identified log type reaches a certain value.
[0124] In this embodiment, at the end of each optimization cycle, the similarity threshold corresponding to each sender is optimized based on the average classification accuracy score of the log corresponding to each sender within the optimization cycle. By continuously adjusting the similarity threshold corresponding to each sender, the accuracy of log classification is maintained for a long time.
[0125] S506: In response to the average value being lower than a preset score, increase the similarity threshold corresponding to the sender of the log to be classified.
[0126] In this embodiment, for each sender, when the average classification accuracy score of the logs to be classified within the optimization period is lower than the preset score, that is, when the classification accuracy is low, the target similarity threshold is increased to reduce the probability of log classification errors.
[0127] It can be understood that the log classification method of this embodiment can initially set the similarity threshold corresponding to each sender based on expert experience. Among them, when setting the similarity threshold, the experts will consider whether each sender log belongs to structured data such as JSON, XML, or unstructured data such as plain text. If the sender log belongs to structured data, then its log type is easier to identify, and a relatively low similarity threshold setting can also ensure the accuracy of log type identification. If the sender log belongs to unstructured data, then its log type is difficult to identify, and a relatively high similarity threshold setting can ensure the accuracy of log type identification.
[0128] The above-mentioned log classification method is then used to determine the log type corresponding to each log in the first data set, wherein the first data set includes multiple historical logs and their corresponding real log types and log senders; based on the determined log type and the real log type, the classification accuracy score of the log is determined; for each sender, the similarity threshold is optimized based on the average classification accuracy score of the log within the optimization period, and the above operation is repeated until the classification accuracy of the log reaches the expected level; the similarity threshold at this time is used as the similarity threshold of the sender.
[0129] The technical solution provided in the embodiment of the present application adjusts the sender similarity threshold for each sender based on whether the average classification accuracy score of the logs to be classified is lower than a preset score in each optimization cycle. By continuously adjusting the sender similarity threshold, the accuracy of log classification is ensured.
[0130] Figure 6 This is a flow chart of another log classification method provided in the embodiment of the present application. Figure 6 , applied to a computing device, the method includes: S601-S610:
[0131] S601: In response to receiving a log to be classified, querying a log list corresponding to the sender of the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender.
[0132] S602: Determine the type of the log to be classified according to the similarity between the log to be classified and the log sample, and the similarity threshold corresponding to the sender of the log to be classified.
[0133] The above steps S601-S602 are as follows Figure 3 Steps S301 and S302 of the illustrated embodiment have been described in detail and will not be repeated here.
[0134] S603: In response to the log type of the log to be classified not being determined, the log to be classified is stored in an unidentified log list corresponding to the sender.
[0135] In some possible implementations, in response to the log type of the log to be classified not being determined, the above steps may be repeated to determine the log type of the log to be classified. In response to the number of days in which the log type of the log to be classified has not been identified reaching a first preset number, the log to be classified is stored in the unidentified log list corresponding to the sender. This can avoid situations where the log type of the log to be classified cannot be identified due to a temporary error.
[0136] It is understandable that log messages are generated very quickly. In a short period of time, a large number of log messages from different senders will be published to the message subscription system. The log types of some of these logs are difficult to identify. Logs with unidentified log types are stored in the unidentified list corresponding to their senders. The unidentified list of each sender will be analyzed later to improve the log list corresponding to each sender. By continuously improving the log list corresponding to each sender, the long-term accuracy of log classification can be maintained.
[0137] S604: Based on text vector similarity matching, the logs in the unrecognized log list are divided into different clusters.
[0138] In some possible implementations, when the number of logs in the unidentified log list reaches a second number threshold, or the time interval from the most recent clearing time of the unidentified log list reaches a second preset time interval, the logs in the unidentified log list are divided into different clusters based on text vector similarity matching.
[0139] In some possible implementations, logs in the unidentified log list are divided into different clusters based on text vector similarity matching, including: based on the text vector similarity corresponding to the logs in the unidentified log list, using a clustering algorithm such as the K-Means algorithm, the K-Medoids algorithm, the DBSCAN algorithm, the OPTICS algorithm, etc., to divide the logs in the unidentified log list into different clusters.
[0140] Among them, K-Means algorithm, K-Medoids algorithm, DBSCAN algorithm, OPTICS algorithm and other clustering algorithms are existing clustering algorithms and will not be described in detail here.
[0141] It is understandable that most of the logs in the sender's unidentified log list belong to a limited number of log types, and the reasons why logs of the same log category are difficult to be identified in a short period of time are likely to be the same. Therefore, the logs in the sender's unidentified log list are divided into different clusters, and the sender's log list is updated according to the clusters, thereby improving the efficiency of log list updates.
[0142] S605: Extract a predetermined number of logs to be classified from the cluster.
[0143] In some possible implementations, the predetermined number may be one or more. The smaller the number of extractions, the fewer operations are required to optimize the log list. The larger the number of extractions, the better the effect of optimizing the log list.
[0144] S606: Acquire the actual log type of the extracted log to be classified and a standard log sample corresponding to the actual log type.
[0145] S607: Determine whether there is a real log type in the log list. In response to determining that there is no real log type in the log list, execute step S608; in response to determining that there is a real log type in the log list, execute step S609.
[0146] S608: Update the real log type and standard log sample to the log list.
[0147] It is understandable that one possible reason why the log type of the log to be classified is difficult to identify is that the sender has added a new log type but has not synchronized it to the execution entity.
[0148] Therefore, when updating the log list, first determine whether the actual log type of the selected log to be classified is recorded in the log list. If not, update the actual log type and its corresponding standard log sample to the log list to eliminate the possibility of log type recognition failure due to the log list not recording the newly added log type and its corresponding log sample, thereby improving the richness of the identifiable log types.
[0149] S609: Determine whether the log sample corresponding to the actual log type in the log list is consistent with the standard log sample. In response to the log sample corresponding to the actual log type in the log list being inconsistent with the standard log sample, execute step S610; in response to the log sample corresponding to the actual log type in the log list being consistent with the standard log sample, the execution ends.
[0150] S610: Update the log sample corresponding to the real log type in the log list to a standard log sample.
[0151] It is understandable that one possible reason why the log type of the log to be classified is difficult to identify is that the sender updated the log sample corresponding to the log type but did not synchronize it with the execution entity. Therefore, after confirming that the actual log type of the selected log to be classified is recorded in the log list, determine whether the log sample corresponding to the actual log type in the log list is consistent with the standard log sample corresponding to the actual log type. If not, update the log sample corresponding to the actual log type to the standard log sample to eliminate the possibility of log type identification failure due to the log list not recording the latest log sample corresponding to the log type.
[0152] In some possible implementations, after the log list is updated using the extracted logs to be classified, the unidentified log list is cleared.
[0153] The technical solution provided by the embodiment of the present application utilizes the unidentified log list of each sender to improve the log list of each sender, so that the log list adapts to changes in the sender, thereby ensuring the accuracy of log classification.
[0154] The above mainly introduces the scheme of the embodiment of the present application from the perspective of method. It is understandable that the computing device (such as server) in the embodiment of the present application includes a hardware structure and / or software module corresponding to the execution of each function in order to realize the above functions. Those skilled in the art should easily appreciate that, in conjunction with the units and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in a way that hardware or computer software drives hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present application.
[0155] Figure 7 This is a structural diagram of a log classification device provided in an embodiment of the present application. Figure 7 The log classification device is applied to a computing device and includes a query module 701 and a determination module 702.
[0156] A query module 701 is configured to query a log list corresponding to a sender of a log to be classified in response to receiving the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender;
[0157] The determination module 702 is configured to determine the type of the log to be classified based on the similarity between the log to be classified and the log sample.
[0158] The technical solution provided in the embodiment of the present application calculates the text similarity between the log to be classified and the log samples corresponding to all log types of the sender of the log to be classified, and then determines the log type of the log to be classified based on the text similarity. By comparing the text similarity between the log and the log samples of its sender, the effect of accurate log classification is achieved. Compared with the solution in the related solution that identifies the type of log based on manually defined log identifiers, this solution avoids the impact of the difficulty in defining log identifiers on log type identification and is simpler to operate.
[0159] In some possible implementations, the device further includes a sending module, which is configured to send the log to be classified to a corresponding log processor according to the type of the log to be classified, and the log processor is configured to process the log to be classified.
[0160] In some possible implementations, the determining module 702 includes: a first determining unit and a second determining unit, wherein:
[0161] A first determining unit is configured to determine a first text vector corresponding to the log to be classified by using a bag-of-words model, wherein the bag-of-words model is constructed based on log fields included in the historical log data;
[0162] The second determining unit is configured to determine the type of the log to be classified based on a similarity between the first text vector and a second text vector corresponding to the log sample, wherein the second text vector is pre-determined based on a bag-of-words model.
[0163] In some possible implementations, the first determining unit includes: a first determining subunit and a second determining subunit; wherein:
[0164] A first determining subunit is configured to determine the log sample corresponding to the second text vector having the greatest similarity to the first text vector as the target log sample;
[0165] The second determining subunit is configured to determine the type of the log to be classified based on the target log sample.
[0166] In some possible implementations, the second determination subunit is specifically used to: in response to determining that the similarity between the log to be classified and the target log sample is greater than a target similarity threshold, determine that the log type of the log to be classified is the log type corresponding to the target log sample, wherein the target similarity threshold is a similarity threshold bound to the sender.
[0167] In some possible implementations, the log classification device further includes an optimization module, which includes: a third determination unit, a fourth determination unit, and an optimization unit; wherein:
[0168] a third determining unit, configured to determine, in response to determining the log type of the log to be classified, the number of fields matching the log to be classified and the target log sample;
[0169] a fourth determining unit, configured to determine a classification accuracy score of the log to be classified based on a proportion of the number of fields in the total number of the log to be classified, wherein the classification accuracy score is proportional to the proportion;
[0170] The optimization unit is used to optimize the target similarity threshold based on the classification accuracy score.
[0171] In some possible implementations, the optimization unit includes: a third determining subunit and an upward adjustment unit, wherein:
[0172] A third determining subunit is configured to determine an average value of the classification accuracy scores of the logs corresponding to the sender within the optimization period in which the logs to be classified are located;
[0173] The upward adjustment subunit is configured to increase the target similarity threshold in response to the average value being lower than a preset score.
[0174] In some possible implementations, the log classification device further includes a storage module, which is specifically configured to: in response to an undetermined log type of the log to be classified, store the log to be classified in an unidentified log list corresponding to the sender.
[0175] In some possible implementations, the log classification device further includes an update module, which includes a clustering unit, an extraction unit, an acquisition unit, and an update unit; wherein:
[0176] The clustering unit is used to divide the logs in the unidentified log list into different clusters based on text vector similarity matching;
[0177] An extraction unit, configured to extract a predetermined number of logs to be classified from the cluster;
[0178] An acquisition unit, configured to acquire the actual log type of the extracted log to be classified and a standard log sample corresponding to the actual log type;
[0179] The update unit is used to update the log list based on real log types and standard log samples.
[0180] In some possible implementations, the updating unit is specifically configured to:
[0181] In response to determining that the real log type does not exist in the log list, the real log type and the standard log sample are updated to the log list.
[0182] In some possible implementations, the updating unit is further configured to:
[0183] In response to determining that a real log type exists in the log list and the log sample corresponding to the real log type is inconsistent with the standard log sample, the log sample corresponding to the real log type is updated to the standard log sample.
[0184] The present application also provides a computing device including a processor and a memory, wherein the processor and the memory are coupled, wherein the memory is used to store computer program instructions, and the processor is used to call the computer program instructions in the memory to execute the log classification method shown in the above embodiment.
[0185] An embodiment of the present application further provides a computer-readable storage medium storing computer program instructions, which are used to enable a computing device to execute the log classification method shown in the above embodiment.
[0186] An embodiment of the present application further provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computing device, the computing device executes the log classification method shown in the above embodiment.
[0187] The computing devices, computer-readable storage media, or computer program products provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0188] Through the description of the above implementation methods, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device (such as a computing device) can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices (such as computing devices) and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0189] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices (such as computing devices) and methods can be implemented in other ways. For example, the device (such as computing device) embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0190] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0192] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0193] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A log classification method, characterized in that: Applied to a computing device, the method includes: In response to receiving a log to be classified, querying a log list corresponding to a sender of the log to be classified, wherein the log list pre-stores a log sample corresponding to at least one log type of the sender; The type of the log to be classified is determined according to the similarity between the log to be classified and the log sample.
2. The method according to claim 1, characterized in that The method further comprises: According to the type of the log to be classified, the log to be classified is sent to a corresponding log processor, and the log processor is used to process the log to be classified.
3. The method according to claim 1, characterized in that The determining the type of the log to be classified according to the similarity between the log to be classified and the log sample includes: Determine a first text vector corresponding to the log to be classified using a bag-of-words model, wherein the bag-of-words model is constructed based on log fields included in historical log data; The type of the log to be classified is determined based on a similarity between the first text vector and a second text vector corresponding to the log sample, wherein the second text vector is determined in advance based on the bag-of-words model.
4. The method according to claim 3, characterized in that The determining the type of the log to be classified based on the similarity between the first text vector and the second text vector corresponding to the log sample includes: Determine the log sample corresponding to the second text vector having the greatest similarity to the first text vector as the target log sample; Based on the target log sample, the type of the log to be classified is determined.
5. The method according to claim 4, characterized in that The determining the type of the log to be classified based on the target log sample includes: In response to determining that the similarity between the log to be classified and the target log sample is greater than a target similarity threshold, determining the log type of the log to be classified to be the log type corresponding to the target log sample, wherein the target similarity threshold is a similarity threshold bound to the sender.
6. The method according to claim 5, characterized in that The method further comprises: In response to determining the log type of the log to be classified, determining the number of fields that match between the log to be classified and the target log sample; Determining a classification accuracy score of the log to be classified based on a proportion of the number to the total number of fields in the log to be classified, wherein the classification accuracy score is proportional to the proportion; Based on the classification accuracy score, the target similarity threshold is optimized.
7. The method according to claim 6, characterized in that Optimizing the target similarity threshold based on the classification accuracy score includes: Determine an average value of classification accuracy scores of logs corresponding to the sender within the optimization period of the log to be classified; In response to the average being lower than a preset score, the target similarity threshold is increased.
8. The method according to any one of claims 1 to 7, characterized in that The method further includes: in response to not determining the log type of the log to be classified, storing the log to be classified in an unidentified log list corresponding to the sender; The method further comprises: Based on text vector similarity matching, the logs in the unidentified log list are divided into different clusters; Extracting a predetermined number of logs to be classified from the cluster; Obtaining the real log type of the extracted log to be classified and a standard log sample corresponding to the real log type; The log list is updated based on the real log type and the standard log sample.
9. The method according to claim 8, characterized in that The updating of the log list based on the real log type and the standard log sample includes: In response to determining that the real log type does not exist in the log list, updating the real log type and the standard log sample to the log list; Or, in response to determining that the real log type exists in the log list and the log sample corresponding to the real log type is inconsistent with the standard log sample, the log sample corresponding to the real log type is updated to the standard log sample.
10. A computing device, characterized in that The method comprises a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor calls the program instructions to execute the method according to any one of claims 1 to 9.