Log collection method and device, equipment and medium

By pre-constructing the log feature library to screen and collect the collection logs, the high cost problems caused by a large number of logs are solved, and dynamic log collection and cost savings are achieved.

CN120066899APending Publication Date: 2025-05-30BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197404.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

As business development occurs, business applications output a large number of logs, resulting in high cost of storing, processing and analysis of logs.

Method used

By obtaining the target application's to be collected logs and the pre-built log feature library, we can determine whether the to be collected logs belong to the log category based on comparative calculations, and collect the logs according to the preset ratio or conduct full collection.

Benefits of technology

Dynamic collection of logs is realized, reducing the number of transmissions during log processing, and saving the cost of log processing, storage and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066899A_ABST
    Figure CN120066899A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a log collection method and device, equipment and a medium, and relates to the technical field of data processing, and the method comprises the following steps: obtaining a to-be-collected log corresponding to a target application program and a pre-constructed log feature library; performing comparison calculation based on the to-be-collected log and a log feature library, and judging whether the to-be-collected log belongs to a log category or not; when the to-be-collected logs belong to the log category, determining a target collection log from the to-be-collected logs according to a preset target proportion value to perform log collection; wherein the target proportion value is less than 1; and when the to-be-collected log does not belong to the log category, carrying out full-amount log collection on the to-be-collected log. According to the embodiment of the invention, the to-be-collected log is compared with the log feature library to judge that the to-be-collected log is the abnormal log, and the to-be-collected log is judged to be the normal log and is collected according to a certain proportion, so that the analysis volume of the normal log is reduced, the dynamic collection of the log is realized, and the log collection and analysis cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a method, apparatus, device, and medium for log collection. Background Art

[0002] Currently, various business application programs, such as video services, audio services, etc., will output printed logs to record important information of key links for scenarios such as monitoring and alarming, auditing, and troubleshooting based on the logs.

[0003] Generally, the logs are first saved to the server where the application program runs, and then collected by a dedicated log platform for subsequent processing and analysis. However, with the development of the business, a large number of services have output a very large number of logs, bringing relatively high costs for log storage, processing, and analysis. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for log collection.

[0005] An embodiment of the present disclosure provides a method for log collection. The method includes: obtaining the logs to be collected corresponding to a target application program and a pre-constructed log feature library; wherein the log feature library includes at least one log category of the target application program, and each log category includes at least one log feature field; performing a comparison calculation based on the logs to be collected and the log feature library to determine whether the logs to be collected belong to the log category; when the logs to be collected belong to the log category, determining target collected logs from the logs to be collected according to a preset target ratio value for log collection; wherein the target ratio value is less than 1; when the logs to be collected do not belong to the log category, performing full-volume log collection on the logs to be collected.

[0006] Optionally, the performing a comparison calculation based on the logs to be collected and the log feature library to determine whether the logs to be collected belong to the log category includes: obtaining the log fields corresponding to the logs to be collected; performing a match based on the log fields and at least one log feature field included in the log category to obtain the number of field matches; when the number of field matches is greater than or equal to a preset match quantity threshold, determining that the logs to be collected belong to the log category, and when the number of field matches is less than the match quantity threshold, determining that the logs to be collected do not belong to the log category.

[0007] Optionally, the comparison calculation based on the log to be collected and the log feature library to determine whether the log to be collected belongs to the log category includes: inputting the log to be collected into a pre-trained comparison model; wherein, the comparison model has the log feature library, the comparison model obtains the log vector corresponding to the log to be collected, and obtains the log feature vectors of at least one log feature field corresponding to the log category; calculating the vector similarity based on the log vector and the log feature vectors, and when the vector similarity is greater than or equal to a preset vector similarity threshold, determining that the log to be collected belongs to the log category, and when the vector similarity is less than the vector similarity threshold, determining that the log to be collected does not belong to the log category.

[0008] Optionally, the method further includes: obtaining a plurality of historical logs of the target application in a historical time period; obtaining string variables in each historical log based on regular matching, and replacing the string variables in each historical log with a preset target identifier to obtain a plurality of preprocessed logs; performing word segmentation processing on each of the preprocessed logs, obtaining a word segmentation set and a log length corresponding to each of the preprocessed logs, and obtaining a log matching word corresponding to each of the preprocessed logs; comparing the log lengths and the log matching words of any two of the preprocessed logs, and when the log lengths between any two of the preprocessed logs are the same and the log matching words between any two of the preprocessed logs are the same, regarding any two of the preprocessed logs as two logs to be clustered; calculating the similarity based on the word segmentation sets corresponding to at least the two logs to be clustered, and when the similarity is greater than a preset similarity threshold, performing clustering processing on at least the two logs to be clustered to obtain clustered logs; obtaining the logs to be processed for which there are no logs to be clustered, and constructing the log feature library based on the log lengths, log matching words, and word segmentation sets corresponding to the logs to be processed and the clustered logs.

[0009] Optionally, the method further includes: obtaining a plurality of historical logs of the target application in a historical time period; merging at least two historical logs whose string lengths of the longest common substring obtained from the plurality of historical logs are greater than or equal to a preset length threshold to obtain merged logs; wherein, the length threshold is determined based on the log lengths of the at least two historical logs; updating the non-common substrings in the merged logs to the target identifier to obtain updated logs; constructing the log feature library based on the historical logs that have not been merged in the plurality of historical logs and the updated logs.

[0010] Optionally, the method further includes: obtaining usage information of the target application; determining an update log of the target application based on the usage information; and updating the log feature library with the update log.

[0011] An embodiment of the present disclosure also provides a log collection device, which includes: an acquisition module, configured to acquire logs to be collected corresponding to a target application program and a pre-constructed log feature library; wherein, the log feature library includes at least one log category of the target application program, and each log category includes at least one log feature field; a comparison module, configured to perform a comparison calculation based on the logs to be collected and the log feature library to determine whether the logs to be collected belong to the log category; a determination module, configured to, when the logs to be collected belong to the log category, determine target collected logs from the logs to be collected according to a preset target ratio value for log collection, and when the logs to be collected do not belong to the log category, perform full-volume log collection on the logs to be collected; wherein, the target ratio value is less than 1.

[0012] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the log collection method provided by the embodiment of the present disclosure.

[0013] An embodiment of the present disclosure also provides a computer-readable storage medium, which stores a computer program, and the computer program is used to execute the log collection method provided by the embodiment of the present disclosure.

[0014] An embodiment of the present disclosure also provides a computer program product, including a computer program, wherein the computer program is executed by a processor to implement the log collection method provided by the embodiment of the present application.

[0015] The above technical solution provided by the embodiment of the present disclosure screens and collects the logs to be collected of the application program through the pre-constructed log feature library of the application program, that is, pre-constructs the log feature library of the application program, including one or more log categories of the target application program. When generating the logs to be collected, compare the logs to be collected with the log feature library corresponding to the application program, so that when the logs to be collected belong to the log category of the target application program, partial collection or non-collection can be selected, and when the logs to be collected do not belong to the log category of the target application program, full collection can be selected and other collection strategies can be used to realize dynamic log collection, thereby greatly reducing the number of logs transmitted in the entire log processing process and greatly saving the costs of log processing, storage, and analysis.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0017] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0018] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can also be obtained based on these accompanying drawings without creative efforts.

[0019] Figure 1 A schematic diagram of a log collection process provided for the related art;

[0020] Figure 2 A schematic diagram of a log collection method provided for an embodiment of the present disclosure;

[0021] Figure 3 A schematic diagram of a log collection method provided for an embodiment of the present disclosure;

[0022] Figure 4 A schematic diagram of a log collection process provided for an embodiment of the present disclosure;

[0023] Figure 5 A schematic diagram of the structure of a log collection device provided for an embodiment of the present disclosure;

[0024] Figure 6 A schematic diagram of the structure of an electronic device provided for an embodiment of the present disclosure. Detailed implementation manners

[0025] In order to be able to more clearly understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0026] Many specific details are set forth in the following description to facilitate a thorough understanding of the present disclosure, but the present disclosure may be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.

[0027] In the existing log processing scenario, business application programs will set breakpoints at key points, such as video application programs setting breakpoints at key points such as click play and like. Specifically, it is selected and set according to business needs, so as to output corresponding log information to a specified location, usually a certain file on the server where the application program is located, and then collected by the log platform for analysis.

[0028] For ease of understanding, reference may be made toFigure 1 A schematic diagram of a log collection process provided by the related art shown. In the log platform in the related art, it includes a collection module, which is responsible for collecting the logs generated by business application programs and transmitting them to the unified log center at the back end; generally referred to as a log collection agent, which is deployed on the server where the business application program is located; a processing module, which is used to receive the logs delivered by the collection module, process them according to business rules, and store them in the log storage system; an analysis module, which is used to read and analyze the logs in the log storage system for scenarios such as monitoring and alarming, auditing, and troubleshooting.

[0029] It can be understood that with the continuous development of the business, business application programs will output a large number of logs, bringing relatively high costs for log storage, processing, and analysis; among them, most are duplicate logs and useless logs, and only a small part is used for auditing or occasional troubleshooting. Each time all logs are collected for subsequent processing and analysis, there are relatively large numbers of logs transmitted during the whole process, and the costs of the log processing, storage, and analysis links are relatively high.

[0030] In view of the above problems, an embodiment of the present disclosure proposes a log collection method. By obtaining the logs to be collected corresponding to the target application program and a pre-built log feature library, and performing comparison calculations based on the logs to be collected and the log feature library, it is determined whether the logs to be collected belong to the log category. When the logs to be collected belong to the log category, the target collection logs are determined from the logs to be collected according to a preset target ratio value for log collection; where the target ratio value is less than 1. When the logs to be collected do not belong to the log category, full-volume log collection is performed on the logs to be collected. Thus, by comparing the logs to be collected with the log feature library, when it is determined that the logs to be collected are abnormal logs, full-volume collection is performed, and when it is determined that the logs to be collected are normal logs, collection is performed according to a certain ratio, so as to reduce the analysis volume of normal logs, realize dynamic log collection, and thus reduce the costs of log collection and analysis. The following is a detailed explanation:

[0031] Figure 2 A schematic diagram of the process of a log collection method provided by an embodiment of the present disclosure. This method can be applied to an electronic device including a central processing unit, and such an electronic device can be, for example, a computer, a mobile phone, a tablet computer, a television, a server, etc., without limitation here. As Figure 2 shown, this method mainly includes the following steps S202 to step S206:

[0032] Step S202, obtaining the logs to be collected corresponding to the target application program and a pre-built log feature library; where the log feature library includes at least one log category of the target application program, and each log category includes at least one log feature field.

[0033] In the embodiments of the present disclosure, the target application program may be any business application program, such as a video application program, an audio application program, etc. One or more target application programs are specifically determined according to the actual application scenario.

[0034] In the embodiments of the present disclosure, the logs to be collected refer to one or more logs generated during the running of the target application program; obtaining the logs to be collected corresponding to the target application program may be obtaining the logs generated during the running of the target application program within a preset time period (such as one minute) as the logs to be collected; it may also be obtaining a batch (such as 10) of logs generated during the running of the target application program as the logs to be collected; specifically, it is selected and set according to the actual application scenario.

[0035] Specifically, during the log collection process, the logs to be collected are obtained from a specified file or location on the server where the logs generated during the running of the target application program are stored; among them, there may be multiple logs to be collected.

[0036] In the embodiments of the present disclosure, the log feature library includes at least one log category of the target application program, and each log category includes at least one log feature field. The log category refers to the category of the log template. The fact that the log category includes at least one log feature field can be understood as the log template including one or more log feature fields. The log template is a standard format used to record the events that occur during the running of the target application program and the running state of the program. The log template may include log feature fields such as time, log level, program method, log source, and specific content.

[0037] In the embodiments of the present disclosure, the log feature libraries of each application program are pre-constructed. As an example, multiple historical logs of the application program within a historical time period may be obtained, and the multiple historical logs are clustered through a preset log clustering algorithm to obtain the log feature library of the application program; among them, the log feature library includes one or more log categories, and each log category has a corresponding log template, that is, one or more log feature fields form the log template.

[0038] Step 204, perform a comparison calculation based on the logs to be collected and the log feature library, and determine whether the logs to be collected belong to the log category.

[0039] In the embodiments of the present disclosure, there are many ways to determine whether the log to be collected belongs to a log category by performing comparison calculations based on the log to be collected and the log feature library. As an example, obtain the log fields corresponding to the log to be collected, and match them based on at least one log feature field included in the log fields and the log category, and obtain the number of field matches. When the number of field matches is greater than or equal to the preset matching quantity threshold, it is determined that the log to be collected belongs to the log category. When the number of field matches is less than the matching quantity threshold, it is determined that the log to be collected does not belong to the log category.

[0040] As another example, input the log to be collected into a pre-trained comparison model; wherein, the comparison model has a log feature library, the comparison model obtains the log vector corresponding to the log to be collected, and obtains the log feature vectors of at least one log feature field corresponding to the log category; calculate the vector similarity based on the log vector and the log feature vectors. When the vector similarity is greater than or equal to the preset vector similarity threshold, it is determined that the log to be collected belongs to the log category. When the vector similarity is less than the vector similarity threshold, it is determined that the log to be collected does not belong to the log category.

[0041] It can be understood that the log feature library includes one or more log categories corresponding to the application program. By determining whether the log to be collected belongs to the existing log categories of the application program to determine the collection strategy, there are many ways to determine whether the log to be collected belongs to the log category corresponding to the application program. For example, obtain one or more log feature fields corresponding to each log category and calculate with the log to be collected to determine whether there is a log category that matches the log to be collected.

[0042] Specifically, calculate the similarity between the set of log fields of the log to be collected and the set of log feature fields of each log category to obtain the matching similarity, and compare the matching similarity with the preset similarity threshold. When the matching similarity is greater than or equal to the similarity threshold, it is determined that the log to be collected belongs to the log category. When the matching similarity is less than the similarity threshold, it is determined that the log to be collected does not belong to the log category.

[0043] Step S206, when the log to be collected belongs to the log category, determine the target collection log from the log to be collected according to the preset target ratio value for log collection; wherein, the target ratio value is less than 1.

[0044] Step S208, when the log to be collected does not belong to the log category, perform full-volume log collection on the log to be collected.

[0045] In the embodiments of the present disclosure, the target collection log can be all or part of the collection log, or there can be no collection log.

[0046] It can be understood that the log to be collected belongs to the log category corresponding to the application program, indicating that the log to be collected corresponds to the log template of the application program, and indicating that the log to be collected is a normal log. Usually, it is not necessary to collect and further analyze it to improve the processing efficiency. When the log to be collected belongs to the log category corresponding to the application program, the probability that the log to be collected is an abnormal log is relatively high. Therefore, full-scale collection is required for analysis to further improve security.

[0047] Specifically, there are many ways to determine the log collection strategy after determining whether the log to be collected belongs to the log category. As an example, determine the log category to which the log to be collected belongs, and determine the target collection log from the log to be collected according to the preset target ratio value for log collection. Usually, the target ratio value is less than 1 to save the log collection cost and improve the collection efficiency; as another example, if it is determined that the log to be collected does not belong to the log category, then full-scale log collection is performed on the log to be collected.

[0048] In summary, the log collection scheme of the embodiments of the present disclosure obtains the log to be collected corresponding to the target application program and the pre-constructed log feature library, and performs comparison calculations based on the log to be collected and the log feature library to determine whether the log to be collected belongs to the log category. When the log to be collected belongs to the log category, determine the target collection log from the log to be collected according to the preset target ratio value for log collection; wherein, the target ratio value is less than 1. When the log to be collected does not belong to the log category, full-scale log collection is performed on the log to be collected. Thus, by comparing the log to be collected with the log feature library, full-scale collection is performed when it is determined that the log to be collected is an abnormal log, and collection is performed at a certain ratio when it is determined that the log to be collected is a normal log, so as to reduce the analysis volume of normal logs, realize dynamic collection of logs, and thus reduce the log collection and analysis costs.

[0049] In some embodiments, based on the comparison calculation between the log to be collected and the log feature library to determine whether the log to be collected belongs to the log category, it includes: obtaining the log fields corresponding to the log to be collected; based on the log fields and at least one log feature field included in the log category for matching, obtaining the number of field matches; when the number of field matches is greater than or equal to the preset matching quantity threshold, determine that the log to be collected belongs to the log category, and when the number of field matches is less than the matching quantity threshold, determine that the log to be collected does not belong to the log category.

[0050] Specifically, each log to be collected is usually a text. By performing word segmentation on the log to be collected, one or more corresponding log fields can be obtained, and the set of log feature fields corresponding to each log category can be acquired. The log fields are matched with the log feature fields in the set of log feature fields to obtain the number of field matches. Thus, when the number of field matches is greater than or equal to the preset matching quantity threshold, it is determined that the log to be collected belongs to the log category. When the number of field matches is less than the matching quantity threshold, it is determined that the log to be collected does not belong to the log category. Among them, the matching quantity threshold can be selected and set according to actual applications.

[0051] In the above embodiment, the method of keyword field matching is used to determine whether the log to be collected belongs to the log category, thereby improving the judgment efficiency and further improving the subsequent log collection efficiency.

[0052] In some embodiments, based on the comparison calculation between the log to be collected and the log feature library, it is determined whether the log to be collected belongs to the log category, including: inputting the log to be collected into a pre-trained comparison model. Among them, the comparison model has a log feature library. The comparison model obtains the log vector corresponding to the log to be collected and the log feature vectors of at least one log feature field corresponding to the log category. Based on the log vector and the log feature vectors, the vector similarity is calculated. When the vector similarity is greater than or equal to the preset vector similarity threshold, it is determined that the log to be collected belongs to the log category. When the vector similarity is less than the vector similarity threshold, it is determined that the log to be collected does not belong to the log category.

[0053] In the embodiments of the present disclosure, the comparison model is pre-trained. Specifically, the log samples corresponding to the target application program are obtained and input into the comparison model to be trained, such as a convolutional neural network. Thus, the log sample vectors corresponding to the log samples and the log feature vectors of the log feature library are calculated to determine whether the log samples belong to the log category. Then, the result estimated by the comparison model is compared with the preset annotation result, that is, whether the actual log category to which the log sample belongs is the same as the log category judged by the comparison model, to adjust the model parameters of the comparison model to obtain the trained comparison model.

[0054] Specifically, multiple log samples can be obtained and divided into log training samples and log validation samples. Each log sample is labeled as belonging to a log category in the log feature library or not belonging to the log category in the log feature library. The log training samples are input into a preset contrast model to be trained to obtain prediction results, and the loss value is determined by comparing the prediction results with the labeled results. The model parameters of the contrast model to be trained are continuously adjusted to obtain a trained contrast model. Further, the log validation samples are input into the trained contrast model for processing to output judgment results, so as to verify the accuracy of the trained contrast model according to the judgment results and the labeled results, thereby further improving the accuracy of the judgment results of the contrast model.

[0055] It can also be understood that the contrast model can also be trained based on a large language model. By using the log feature library as the guiding language of the large language model to input log samples for training, a contrast model is obtained, further improving the subsequent processing efficiency and effect.

[0056] In the above embodiment, by inputting the log to be collected into the pre-trained contrast model, it is possible to quickly and accurately obtain whether the log to be collected belongs to the log category in the log feature library, thereby improving the judgment efficiency and further improving the subsequent log collection efficiency.

[0057] To further improve the log collection efficiency, a corresponding log feature library is pre-constructed for each application program. In some embodiments, the log collection method provided by the embodiments of the present disclosure further includes: obtaining multiple historical logs of the target application program in a historical time period, obtaining variables in each historical log based on regular matching, and updating the variables in each historical log to target identifiers to obtain multiple preprocessed logs; performing word segmentation processing on each preprocessed log to obtain the word segmentation set and log length corresponding to each preprocessed log, and obtaining the log matching words corresponding to each preprocessed log; comparing the log lengths and log matching words of any two of the preprocessed logs. When the log lengths between any two preprocessed logs are the same and the log matching words between any two preprocessed logs are the same, any two preprocessed logs are used as two logs to be clustered. Based on the word segmentation sets corresponding to at least two logs to be clustered, a similarity calculation is performed. When the similarity is greater than a preset similarity threshold, at least two logs to be clustered are clustered to obtain clustered logs; based on the matching results, the logs to be processed for which there are no logs to be clustered are determined, and a log feature library is constructed based on the log lengths, log matching words, and word segmentation sets corresponding to the logs to be processed and the clustered logs.

[0058] Among them, the historical time period can be selected and set according to specific scenarios, such as one day, two days, etc. All historical logs generated during the operation of the target application program in the historical time period, that is, multiple historical logs, are used to construct the log feature library corresponding to the target application program.

[0059] Among them, regular matching means presetting a rule string, and then matching each historical log with the preset rule string to determine the matching string in each historical log as a string variable, and replacing the string variable in each historical log with a preset target identifier. That is to say, the string variable is replaced by the target identifier, and the target identifier can be selected and set as needed. For example, using "*" as the target identifier.

[0060] Specifically, when replacing the string variable in each historical log with a preset target identifier, multiple preprocessed logs are obtained. That is, the preprocessed log is a log in which the string variable is replaced by the target identifier. Then, each preprocessed log is segmented through a segmentation tool or algorithm, so that multiple segments corresponding to each preprocessed log can be obtained as a segmentation set, and the number of segments of multiple segments is counted as the log length. In addition, the prefix word of each preprocessed log is obtained as the log matching word corresponding to each preprocessed log.

[0061] Furthermore, for all preprocessed logs, any preprocessed log is matched with other preprocessed logs. First, the log matching words of the preprocessed log are matched. If there are the same log matching words, then the log lengths of the preprocessed logs are further matched. If the log lengths are the same, these preprocessed logs are determined as the logs to be clustered. Further, the similarity of the segmentation sets corresponding to the logs to be clustered is calculated, and all clustered logs with a similarity greater than the preset similarity threshold are clustered to obtain clustered logs. Among them, the similarity threshold can be set according to actual application needs. That is to say, all logs to be clustered are merged to obtain a clustered log.

[0062] It can also be understood that there are unmatched preprocessed logs among all preprocessed logs, that is, preprocessed logs with no same log matching words, or preprocessed logs with the same matching words but different log lengths, or preprocessed logs with the same matching words, the same log lengths but a similarity less than the similarity threshold. The logs to be processed without logs to be clustered, and a log feature library is constructed based on the log lengths, log matching words and segmentation sets corresponding to the logs to be processed and the clustered logs. That is to say, the log templates of the logs to be processed and the log templates corresponding to the clustered logs are determined, and all log templates are combined to obtain a log feature library. Among them, the lengths of each template in the log template can be determined based on the log length, and the log template includes log feature fields such as time, log level, program method, log source and specific content, which can be determined through the log matching words and the segmentation set.

[0063] In some embodiments, the log collection method provided by the embodiments of the present disclosure further includes: obtaining a plurality of historical logs of a target application within a historical time period; obtaining at least two historical logs corresponding to a string length of the longest common substring greater than or equal to a preset length threshold from the plurality of historical logs for merging to obtain a merged log; wherein, determining the length threshold based on the log lengths of the at least two historical logs; updating the non-common substrings in the merged log to target identifiers to obtain an updated log; constructing a log feature library based on the historical logs that are not merged among the plurality of historical logs and the updated log.

[0064] Specifically, compare a plurality of historical logs generated during the operation of the target application within the historical time period, obtain the longest common substring between the historical logs, and calculate the string length of the longest common substring. Merge at least two historical logs whose string length of the longest common substring is greater than or equal to the length threshold, that is, merge all the strings corresponding to the at least two historical logs and only keep one of the same strings to obtain a merged log; wherein, in order to improve the merging effect, determine the length threshold based on the log lengths of the at least two historical logs, such as 0.5 to 1 times the log length corresponding to any one of the at least two historical logs, or directly use the longest log length among the at least two historical logs as the length threshold.

[0065] It can be understood that a string of the longest common string greater than or equal to the length threshold indicates that the corresponding two or more historical logs are similar logs and can share a log template. Therefore, the merged log can be analyzed to obtain a log template composed of one or more log strings; in addition, there may still be historical logs that are not merged among the plurality of historical logs, that is, this historical log has no similar logs and needs to be analyzed separately to determine one or more log strings as the log template of this historical log, so as to construct a log feature library of the target application from multiple log templates.

[0066] In the above solution, the log feature library corresponding to the application can be flexibly constructed by analyzing a plurality of historical logs of the application, so that when the application runs and generates logs in the future, the effective logs can be directly screened and collected based on the log feature library, thereby ensuring the effectiveness of log collection and improving the log collection efficiency.

[0067] It can be understood that after the log feature library is constructed, as time goes by, the usage time and usage content of the application may change, and the log feature library needs to be updated in a timely manner to improve the accuracy of subsequent comparison, thereby improving the accuracy of collection.

[0068] In some embodiments, the method further includes: obtaining usage information of a target application; determining an update log of the target application based on the usage information; and obtaining the update log to update a log feature library.

[0069] In the embodiments of the present disclosure, the usage information includes information such as the usage time and usage location of the target application. Thus, it is determined that the number of generated logs is relatively large. Therefore, the update log of the target application is re-obtained, that is, the newly generated log is used to construct a new log feature library in the foregoing manner, and the log template in the log feature library can also be updated, such as adding or deleting, in the foregoing manner.

[0070] Thereby, the log feature library can be updated in a timely manner based on the specific usage information of the application, further ensuring the subsequent log collection effect and meeting the log collection requirements.

[0071] To further improve the log collection effect, the log to be collected is compared with a pre-constructed log feature library. The comparison method can be flexibly selected to further meet the log collection requirements. In some embodiments, a similarity calculation is performed on the log field set of the log to be collected and the word segmentation set of the matching log to obtain a matching similarity; based on the comparison between the matching similarity and a preset similarity threshold, when the matching similarity is greater than or equal to the similarity threshold, it is determined that the log to be collected belongs to the log category in the log feature library, and when the matching similarity is less than the similarity threshold, it is determined that the log to be collected does not belong to the log category in the log feature library.

[0072] Specifically, each log to be collected is segmented to obtain multiple segments of each log to be collected, and a log field set of the log to be collected is obtained.

[0073] It can be understood that each log template is stored in the log feature library, and each log template corresponds to a log category including log features fields such as log length, log matching words, and word segmentation set. Thus, a similarity calculation is performed on the log field set of the log to be collected and the log feature field set to obtain a matching similarity.

[0074] Further, the matching similarity is compared with a preset similarity threshold. When the matching similarity is greater than or equal to the similarity threshold, it is determined that the log to be collected belongs to the log category in the log feature library, and when the matching similarity is less than the similarity threshold, it is determined that the log to be collected does not belong to the log category in the log feature library; among them, the similarity threshold can be determined according to the actual application.

[0075] Figure 3 The flowchart of a log collection method provided by the embodiments of the present disclosure mainly includes the following steps S302 to step S308:

[0076] Step S302: Obtain the logs to be collected corresponding to the target application and the pre-built log feature library. The log feature library includes at least one log category of the target application, and each log category includes at least one log feature field.

[0077] Step S304: Input the logs to be collected into a pre-trained comparison model. The comparison model has a log feature library. The comparison model obtains the log vector corresponding to the logs to be collected and obtains the log feature vectors of at least one log feature field corresponding to the log category.

[0078] Step S306: Calculate the vector similarity based on the log vector and the log feature vectors. When the vector similarity is greater than or equal to the preset vector similarity threshold, it is determined that the logs to be collected belong to the log category. When the vector similarity is less than the vector similarity threshold, it is determined that the logs to be collected do not belong to the log category.

[0079] Step S308: When the logs to be collected belong to the log category, determine the target collected logs from the logs to be collected according to the preset target proportion value for log collection. When the logs to be collected do not belong to the log category, perform full-volume log collection on the logs to be collected.

[0080] Specifically, the log platform constructs a log feature library of the application based on a preset clustering algorithm. During the log collection process, each log is read and compared with the log feature library. If it belongs to the log category in the log feature library, it can be not collected or sampled at a certain proportion. If it does not belong to the log category in the log feature library, it may belong to abnormal logs and can be 100% collected.

[0081] Exemplarily, as Figure 4 shown, in Step 1, the log collection module collects normal logs for a period of time, such as one day's logs. These logs are transmitted to the log feature library construction module, and a log feature library belonging to the application is constructed based on a preset clustering algorithm for the logs collected for a period of time. The log feature library includes log categories, and each log category corresponds to a log template, including one or more log feature fields. In Step 2, the log feature library is then distributed to each machine (such as a server, etc.) where the application is located. In Step 3, the log collection module reads a batch of logs each time, which can be one or multiple logs within a short period of time, such as one minute's logs, and compares them with the log feature library. In Step 4, if the collected logs are consistent with the log category in the log feature library, it means that this batch of logs belongs to the normal range and can be not collected or sampled at a certain proportion. If they are inconsistent with the log category in the log feature library, it means that this batch of logs may be abnormal logs generated by a service failure and are usually 100% collected in order to check and locate each detailed information or improve the accuracy of monitoring and alarm.

[0082] Specifically, after the above screening and judgment, the logs to be collected are determined and then transmitted to the backend log processing module. It can also be understood that the business form may change over time. For example, for a video application, if the number of people watching videos increases during a certain period, more log volume may be generated and the log template may change. At this time, the log feature library can be rebuilt and updated, that is, repeat steps 1 and 2 to further ensure the log collection effect.

[0083] In summary, the log collection method provided by the embodiments of the present disclosure realizes dynamic log collection based on the log feature library, rather than traditional 100% collection. It can achieve dynamic log collection, greatly reduce the number of logs transmitted during the whole process, and greatly save the costs of log processing, storage, and analysis.

[0084] Corresponding to the foregoing log collection method, the embodiments of the present disclosure further provide a log collection device. Figure 5 The following is a schematic structural diagram of a log collection device provided by the embodiments of the present disclosure. The device can be implemented by software and / or hardware. The device includes:

[0085] An acquisition module 502, configured to acquire the logs to be collected corresponding to the target application program and a pre-built log feature library; wherein, the log feature library includes at least one log category of the target application program, and each log category includes at least one log feature field;

[0086] A comparison module 504, configured to perform comparison calculation based on the logs to be collected and the log feature library, and determine whether the logs to be collected belong to the log category;

[0087] A determination module 506, configured to determine target collection logs from the logs to be collected according to a preset target ratio value for log collection when the logs to be collected belong to the log category, and perform full-volume log collection on the logs to be collected when the logs to be collected do not belong to the log category; wherein, the target ratio value is less than 1.

[0088] The above device provided by the embodiments of the present disclosure screens and collects the logs to be collected of the application program through the pre-built log feature library of the application program. That is to say, the log feature library of the application program is pre-built to represent the log categories corresponding to the application program. When the logs to be collected are generated, the logs to be collected are compared with the set of log feature fields of the log categories corresponding to the application program. Thus, when the logs to be collected belong to a log category, partial collection or non-collection can be selected, and when the logs to be collected do not belong to a log category, full collection and other collection strategies can be selected to achieve dynamic log collection, thereby greatly reducing the number of logs transmitted during the whole log processing process and greatly saving the costs of log processing, storage, and analysis.

[0089] In some embodiments, the comparison module 504 is specifically configured to: obtain the log fields corresponding to the log to be collected; match based on the log fields and at least one log feature field included in the log category to obtain the number of field matches; when the number of field matches is greater than or equal to a preset match quantity threshold, determine that the log to be collected belongs to the log category, and when the number of field matches is less than the match quantity threshold, determine that the log to be collected does not belong to the log category.

[0090] In some embodiments, the comparison module 504 is specifically configured to: input the log to be collected into a pre-trained comparison model; wherein, the comparison model has the log feature library, the comparison model obtains the log vector corresponding to the log to be collected, and obtains the log feature vectors of at least one log feature field corresponding to the log category; calculate the vector similarity based on the log vector and the log feature vectors, when the vector similarity is greater than or equal to a preset vector similarity threshold, determine that the log to be collected belongs to the log category, and when the vector similarity is less than the vector similarity threshold, determine that the log to be collected does not belong to the log category.

[0091] In some embodiments, the device further includes: a first log acquisition module, configured to acquire a plurality of historical logs of the target application within a historical time period; a matching and updating module, configured to obtain string variables in each historical log based on regular matching, and replace the string variables in each historical log with preset target identifiers to obtain a plurality of preprocessed logs; a word segmentation and acquisition module, configured to perform word segmentation processing on each of the preprocessed logs to obtain a word segmentation set and a log length corresponding to each of the preprocessed logs, and obtain a log matching word corresponding to each of the preprocessed logs; a matching and calculation module, configured to compare the log lengths and the log matching words of any two of the preprocessed logs, and when the log lengths between any two of the preprocessed logs are the same and the log matching words between any two of the preprocessed logs are the same, use any two of the preprocessed logs as two logs to be clustered, calculate the similarity based on the word segmentation sets corresponding to at least the two logs to be clustered, and when the similarity is greater than a preset similarity threshold, perform clustering processing on at least the two logs to be clustered to obtain clustered logs;; a first construction module, configured to obtain a log to be processed for which there are no logs to be clustered, and construct the log feature library based on the log lengths, log matching words, and word segmentation sets corresponding to the log to be processed and the clustered logs.

[0092] In some embodiments, the device further includes: a second log acquisition module, configured to acquire a plurality of historical logs of the target application within a historical time period; an acquisition and merging module, configured to acquire at least two historical logs corresponding to a string length of a maximum common substring greater than or equal to a preset length threshold from the plurality of historical logs for merging to obtain a merged log; wherein the length threshold is determined based on the log lengths of the at least two historical logs; an update module, configured to update non-common substrings in the merged log to the target identifier to obtain an updated log; a second construction module, configured to construct the log feature library based on the historical logs that are not merged in the plurality of historical logs and the updated log.

[0093] In some embodiments, the device further includes: an acquisition and update module, configured to acquire usage information of the target application, determine an update log of the target application based on the usage information, and acquire the update log to perform an update process on the log feature library.

[0094] The log acquisition device provided by the embodiments of the present disclosure can execute the log acquisition method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0095] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device embodiments described above can refer to the corresponding process in the method embodiments, which will not be elaborated herein.

[0096] Embodiments of the present disclosure provide an electronic device, which includes: a storage device storing a computer program thereon; a processing device, configured to execute the computer program in the storage device to implement the steps of any method in the present disclosure.

[0097] Next, with reference to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0098] As Figure 6As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0099] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 6 an electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0100] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0101] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run on a processor, cause the processor to execute the image processing methods provided by the embodiments of the present disclosure. The computer program products may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0102] In addition, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and the computer program instructions, when run on a processor, cause the processor to execute the log collection method provided by the embodiments of the present disclosure.

[0103] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0104] Embodiments of the present disclosure also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the log collection method in the embodiments of the present disclosure.

[0105] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0106] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosed technical solution based on the prompt message.

[0107] As an optional but non-limiting implementation manner, when responding to receiving an active request from a user, the manner of sending a prompt message to the user may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0108] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manners of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0109] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0110] The above are only specific implementation manners of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A log collection method, characterized in that: The method comprises: Obtaining the logs to be collected and the pre-built log feature library corresponding to the target application; wherein the log feature library includes at least one log category of the target application, and each of the log categories includes at least one log feature field; Performing a comparative calculation based on the log to be collected and the log feature library, determining whether the log to be collected belongs to the log category; When the log to be collected belongs to the log category, determining a target collection log from the log to be collected according to a preset target ratio value to perform log collection; wherein the target ratio value is less than 1; When the log to be collected does not belong to the log category, full log collection is performed on the log to be collected.

2. The method according to claim 1, characterized in that The comparing and calculating based on the log to be collected and the log feature library to determine whether the log to be collected belongs to the log category includes: Obtain the log field corresponding to the log to be collected; Matching the log field with at least one log feature field included in the log category to obtain the number of field matches; When the number of field matches is greater than or equal to a preset matching number threshold, it is determined that the log to be collected belongs to the log category; when the number of field matches is less than the matching number threshold, it is determined that the log to be collected does not belong to the log category.

3. The method according to claim 1, characterized in that The comparing and calculating based on the log to be collected and the log feature library to determine whether the log to be collected belongs to the log category includes: Input the log to be collected into a pre-trained comparison model; wherein the comparison model has the log feature library, and the comparison model obtains the log vector corresponding to the log to be collected, and obtains the log feature vector of at least one log feature field corresponding to the log category; The vector similarity is calculated based on the log vector and the log feature vector. When the vector similarity is greater than or equal to a preset vector similarity threshold, it is determined that the log to be collected belongs to the log category. When the vector similarity is less than the vector similarity threshold, it is determined that the log to be collected does not belong to the log category.

4. The method according to claim 1, characterized in that: The method further comprises: Obtaining multiple historical logs of the target application within a historical time period; Obtaining a string variable in each historical log based on regular expression matching, and replacing the string variable in each historical log with a preset target identifier to obtain multiple pre-processed logs; Perform word segmentation processing on each of the preprocessed logs, obtain a word segmentation set and a log length corresponding to each of the preprocessed logs, and obtain a log matching word corresponding to each of the preprocessed logs; Comparing the log lengths and the log matching words of any two of the pre-processed logs, and taking any two of the pre-processed logs as two logs to be clustered when the log lengths of any two of the pre-processed logs are the same and the log matching words of any two of the pre-processed logs are the same; Performing similarity calculation based on the word segmentation sets corresponding to at least the two logs to be clustered, and when the similarity is greater than a preset similarity threshold, clustering the at least two logs to be clustered to obtain clustered logs; Obtain a to-be-processed log for which no to-be-clustered log exists, and construct the log feature library based on the log length, log matching words, and word segmentation set corresponding to the to-be-processed log and the clustered log.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining multiple historical logs of the target application within a historical time period; At least two historical logs corresponding to the maximum common substring whose string length is greater than or equal to a preset length threshold are obtained from the multiple historical logs and merged to obtain a merged log; wherein the length threshold is determined based on the log length of the at least two historical logs; Updating the non-common substring in the merge log to the target identifier to obtain an update log; The log feature library is constructed based on the unmerged historical logs among the plurality of historical logs and the update log.

6. The method according to claim 1, characterized in that The method further comprises: Obtaining usage information of the target application; determining an update log of the target application based on the usage information; The update log is obtained to update the log feature library.

7. A log collection device, characterized in that: The device comprises: An acquisition module, used to acquire the logs to be collected and the pre-built log feature library corresponding to the target application; wherein the log feature library includes at least one log category of the target application, and each of the log categories includes at least one log feature field; A comparison module, used for performing comparison calculation based on the log to be collected and the log feature library to determine whether the log to be collected belongs to the log category; A determination module is used to determine a target collection log from the logs to be collected according to a preset target ratio value for log collection when the logs to be collected belong to the log category, and to perform full log collection on the logs to be collected when the logs to be collected do not belong to the log category; wherein the target ratio value is less than 1.

8. An electronic device, characterized in that: The electronic device comprises: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the log collection method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the log collection method described in any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises a computer program, wherein the computer program is executed by a processor to perform the log collection method described in any one of claims 1 to 6.