Log Data Testing Method, Device, System, Electronic Device, and Storage Medium
Automated log testing using white-list data generated from pre-set log data and clustering algorithms addresses inefficiencies in human-based log observation, enhancing accuracy and system stability.
Patent Information
- Application Number
- CN201811519316.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-12-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2038-12-12
AI Technical Summary
In the prior art, log inspection when the module is online relies on manual observation, resulting in low efficiency and accuracy, which may cause log misses and affect system stability.
By generating whitelist data associated with the test mode, using the clustering algorithm to classify and analyze the log data frequently, automatically test whether the target log data is normal and avoid manual omissions.
It improves the accuracy and efficiency of log testing, ensures that each log data is tested, enhances system stability, and reduces the number of rollbacks after going online.
Smart Images

Figure CN111309585B_ABST
Abstract
Description
Background Art
[0002] Currently, when testing products launched online, end-to-end tests such as performance testing and functional testing are mainly carried out. However, there may be potential risks in the logs when the modules are launched online.
[0003] In related technologies, generally, the logs are opened manually for observation, and whether the logs are abnormal completely depends on the cognition of the testers. In the case of a large number of logs, it is impossible to pay attention to each log, which may cause the problem of log omission; the log inspection completely relies on the cognition of the testers, so the efficiency and accuracy of log inspection may be low, thereby reducing the system stability.
[0004] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a log data testing method, device, system, electronic device, and storage medium, so as to at least to some extent overcome the problem of quickly and accurately performing log testing caused by the limitations and defects of related technologies.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0007] According to one aspect of the present disclosure, there is provided a log data testing method, including: generating whitelist data associated with a test mode according to preset log data corresponding to the test mode; testing target log data associated with the test mode through the whitelist data to determine whether the target log data is normal.
[0008] In an exemplary embodiment of the present disclosure, generating whitelist data associated with a test mode according to preset log data corresponding to the test mode includes: preprocessing preset log data belonging to a first tag; determining the occurrence frequency of each type of log data in the preprocessed partial log data; determining the whitelist data according to the occurrence frequency.
[0009] In an exemplary embodiment of the present disclosure, determining the occurrence frequency of each type of log data in the preprocessed partial log data includes: classifying the preprocessed partial log data through a clustering algorithm to obtain multiple types of log data; determining the occurrence frequency of each type of log data according to the maximum value of the occurrence times of each type of log data.
[0010] In an exemplary embodiment of the present disclosure, the method further includes: testing preset log data other than the partial log data that belongs to the first tag and / or belongs to the second tag by using the whitelist data to determine whether to update the whitelist data.
[0011] In an exemplary embodiment of the present disclosure, testing preset log data other than the partial log data that belongs to the first tag and / or belongs to the second tag by using the whitelist data to determine whether to update the whitelist data includes: if the preset log data other than the partial log data that belongs to the first tag and belongs to the second tag all pass the test of the whitelist data, then the whitelist data is not updated; if the preset log data other than the partial log data that belongs to the first tag or belongs to the second tag fails the test of the whitelist data, then the clustering algorithm is adjusted to update the whitelist data.
[0012] In an exemplary embodiment of the present disclosure, the test mode includes regression testing and online testing. Generating whitelist data associated with the test mode according to the preset log data corresponding to the test mode includes: adopting an offline method to generate whitelist data corresponding to the regression testing according to historical log data; or adopting a real-time method to generate whitelist data corresponding to the online testing according to online log data of a stable version.
[0013] In an exemplary embodiment of the present disclosure, testing the target log data associated with the test mode by using the whitelist data to determine whether the target log data is normal includes: clustering the target log data associated with the test mode by using a clustering algorithm to obtain multiple categories of target log data; testing each category of target log data by using the whitelist data to determine whether each category of target log data is normal.
[0014] In an exemplary embodiment of the present disclosure, testing each category of target log data by using the whitelist data to determine whether each category of target log data is normal includes: if the occurrence frequency of each category of target log data meets the whitelist data, then it is determined that each category of target log data is normal; if the occurrence frequency of each category of target log data does not meet the whitelist data, then a reminder message is generated and the reminder message is sent to the client.
[0015] In an exemplary embodiment of the present disclosure, if the test mode is online testing, the target log data is online log data of an online version.
[0016] According to one aspect of the present disclosure, there is provided a log data testing device, including: a whitelist data generation module configured to generate whitelist data associated with a test mode according to preset log data corresponding to the test mode; a log testing module configured to test target log data associated with the test mode by using the whitelist data to determine whether the target log data is normal.
[0017] According to one aspect of the present disclosure, there is provided a log data testing system, including: a server configured to generate whitelist data associated with a test mode according to preset log data corresponding to the test mode, test target log data of the test mode according to the whitelist data, and send a test result to a client; a client configured to receive the test result returned by the server and display the test result.
[0018] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the log data testing method according to any one of the above via executing the executable instructions.
[0019] According to one aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, where the computer program, when executed by a processor, implements the log data testing method according to any one of the above.
[0020] In a log data testing method, device, system, electronic device, and computer-readable storage medium provided in an exemplary embodiment of the present disclosure, on the one hand, automatic testing is performed on target log data associated with a test mode by using whitelist data generated according to preset log data corresponding to the test mode, and vertical comparison of log data can be achieved. Compared with manual testing, the accuracy and efficiency of testing log data are improved; on the other hand, since the target log data can be automatically tested by using the whitelist data, each log data can be tested, so that the omission problem caused by manual testing can be avoided, and the system stability is improved.
[0021] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 Schematically shows a schematic diagram of a system architecture for implementing a log data testing method in an exemplary embodiment of the present disclosure.
[0024] Figure 2 Schematically shows a flowchart of a log data testing method in an exemplary embodiment of the present disclosure.
[0025] Figure 3 Schematically shows a flowchart of generating whitelist data in an exemplary embodiment of the present disclosure.
[0026] Figure 4 Schematically shows a specific flowchart of generating whitelist data in an exemplary embodiment of the present disclosure.
[0027] Figure 5 Schematically shows a flowchart of updating whitelist data in an exemplary embodiment of the present disclosure.
[0028] Figure 6 Schematically shows a flowchart of generating whitelist data in an offline manner in an exemplary embodiment of the present disclosure.
[0029] Figure 7 Schematically shows a flowchart of generating whitelist data in a real-time manner in an exemplary embodiment of the present disclosure.
[0030] Figure 8 Schematically shows a flowchart of analyzing target log data in an exemplary embodiment of the present disclosure.
[0031] Figure 9 Schematically shows a block diagram of a log data testing device in an exemplary embodiment of the present disclosure.
[0032] Figure 10 Schematically shows a block diagram of a log data testing system in an exemplary embodiment of the present disclosure.
[0033] Figure 11 Schematically shows a block diagram of an electronic device in an exemplary embodiment of the present disclosure.
[0034] Figure 12 Schematically shows a program product in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0036] In addition, the drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0037] In this example embodiment, a system architecture for implementing log data testing is first provided. Referring to Figure 1 as shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0038] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send request instructions, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as picture processing applications, shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0039] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0040] Server 105 may be a server that provides various services. For example, it may be a background management server (only for example) that supports shopping websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - only for example) to the terminal devices.
[0041] It should be noted that the log data testing method provided in the embodiments of the present application is generally executed by server 105. Correspondingly, the log data testing device is generally set in terminal device 101.
[0042] Based on Figure 1 the system architecture in, in this exemplary embodiment, a log data testing method is provided, which can be applied to scenarios where log data during module go-live or product go-live is processed to determine whether the log data is abnormal. Next, with reference to Figure 2 shown in, the log data testing method in this exemplary embodiment will be specifically described.
[0043] In step S210, white list data associated with the test mode is generated according to the preset log data corresponding to the test mode.
[0044] In this exemplary embodiment, the test mode may include two modes: regression testing and online testing. Regression testing refers to checking whether the original functions of the software remain intact after modification. Regression testing does not reduce the test requirements for new functions and features of the system. The regression test package should include tests for new functions and features. Regression testing does not occur only when requirements change. Regression testing can occur in any part of the software life cycle, from unit testing, functional testing, integration testing, to even release testing. The test mode may also include online testing, which refers to the testing of the newly released version or the newly released product to improve the quality of the product through online testing.
[0045] For regression testing, the preset log data refers to the historical log data after historical regression testing, that is, the log data used for regression testing statistically after each regression testing. For online testing, the preset log data refers to the online log data, generally the log data of the stable version of the code in the online environment. The preset log data can generally be represented in the form of a LOG set, and the LOG set may include user information, test information, etc.
[0046] Whitelist data refers to the data used to test target log data. Through the whitelist data, it can be determined whether the target log data meets the requirements. In this exemplary embodiment, the method of generating whitelist data for different test modes is also different to cope with different test scenarios. For example, in regression testing, the whitelist data corresponding to the regression testing is generated in an offline manner. In the offline manner, mainly the historical log data is crawled. During the testing process, by comparing with the historical log data set, all suspicious points are found from the historical log data and reported to reduce the probability of code problems. Another example is that in online testing, the whitelist data corresponding to the online testing is generated in a real-time manner. In the real-time manner, mainly the cluster logs of the online stable version are crawled in real time to generate the whitelist data. The whitelist data generated by any method can include the frequency of occurrence of certain logs to perform subsequent log analysis based on the recorded frequency.
[0047] Specifically, the whitelist data associated with the test mode can be generated by learning the preset log data corresponding to the test mode. That is to say, by learning the historical log data of the regression testing, the whitelist data corresponding to the regression testing is generated in an offline manner. By learning the online log data of the online testing, the whitelist data corresponding to the online testing is generated in a real-time manner. Generating the whitelist data associated with different test modes through different methods and different preset log data can make the obtained whitelist data more accurate and more in line with the test mode, so as to achieve accurate log testing.
[0048] Reference Figure 3 As shown, the specific steps of generating the whitelist data associated with the test mode by analyzing the preset log data corresponding to the test mode include step S301 and step S302. Among them:
[0049] In step S301, the first label or the second label corresponding to the preset log data is obtained.
[0050] In this step, the first label and the second label can be manually marked or marked by an algorithm, and no special limitation is imposed here. The first label can be a label indicating that the preset log data passes the test. For example, it can be represented by identifiers such as PASS or the number 1. The second label can be a label indicating that the preset log data fails the test. For example, it can be represented by identifiers such as FAIL or the number 0. The specific manifestation form is not specially limited here, as long as it can represent passing the test and failing the test respectively. Taking the process of generating whitelist data in regression testing as an example for illustration. In the daily regression testing in the initial stage, historical log data will be automatically collected and stored. According to the confirmation by personnel and the final result of going online (i.e., whether to go online), the collected historical log data is detected and labeled, and the preset log data of the regression test is divided into log data with two labels, PASS and FAIL.
[0051] In step S302, by analyzing some of the preset log data belonging to the first label, the whitelist data is generated.
[0052] In this exemplary embodiment, some of the preset log data belonging to the first label refers to the PASS log data set. Here, some of the log data can be used as training samples. Some of the log data can be more than half of the preset log data. The more the data volume, the more accurate the obtained whitelist data. Analyzing some of the preset log data belonging to the first label can be understood as training or learning on a set of historical log data with some PASS. Referring to Figure 4 As shown in, the specific process of step S302 includes steps S401 to S403. Among them:
[0053] In step S401, the preset log data belonging to the first label is preprocessed.
[0054] In this exemplary embodiment, the preprocessing includes, but is not limited to, operations such as word segmentation, data filtering, and statistics. Specifically, word segmentation can be performed on the set of historical log data with the label PASS, removing meaningless words or some specific words, and leaving general and meaningful sentences. For example, removing user information in the LOG. Frequency statistics of word segmentation and sentence splitting can also be performed, and the error sentences ranked in the top N positions in terms of the number of occurrences are analyzed. N can be set according to actual needs. At the same time, known non-serious errors or normal log data, as well as sentences with a frequency not exceeding a certain threshold, can be ignored. By performing preprocessing, the influence of meaningless words or specific words can be avoided, and log data analysis can be performed more accurately. At the same time, due to the preprocessing process, the interference of irrelevant or less influential preset log data can be avoided, reducing the amount of data to be analyzed and improving the efficiency of log data analysis.
[0055] In step S402, determine the occurrence frequency of each type of log data in the preprocessed partial log data.
[0056] In this exemplary embodiment, the preprocessed partial log data can be classified to obtain multiple types of log data, such as the first type, the second type, the third type, and so on. Specifically, after preprocessing the historical log data of PASS, the preprocessed partial log data can be classified by a clustering algorithm. The clustering algorithm can include, for example, the K-means clustering algorithm, mean shift clustering, density-based spatial clustering of applications with noise (DBSCAN), hierarchical clustering, and so on. By the clustering algorithm, log statements with the same characteristics can be aggregated into one class, thus obtaining multiple types of log data. Through the clustering algorithm, the preprocessed partial log data can be accurately classified to obtain an accurate classification result, that is, multiple types of log data.
[0057] After classifying the partial log data, the occurrence frequency of each type of log data can be determined according to the occurrence times of each type of log statement in each regression test. Specifically, to improve accuracy, the maximum value of the occurrence times of each type of log data can be used as the occurrence frequency of each type of log data. For example, if the occurrence times of the first type of log data are 10 times, 20 times, and 30 times respectively, then its occurrence frequency is 30 times.
[0058] In step S403, determine the whitelist data according to the occurrence frequency.
[0059] In this exemplary embodiment, if this type of log data appears in each regression test, the occurrence frequency determined by the maximum occurrence times of this type of log data is used as the frequency in the whitelist data, and the whitelist data for this type of log data in the regression test stage is generated. For example, if the occurrence times of the first type of log data are 10 times, 20 times, 30 times, etc., then for the first type of log data, the whitelist data in the regression test stage includes the first type of log data 30 times. In the same way, the whitelist data in the regression test stage can be generated according to the maximum occurrence times of all types of log data, so as to display existing error statements and so on in the form of whitelist data.
[0060] After generating the whitelist data corresponding to the regression test, since it is not known whether the whitelist data is accurate, it is necessary to test the generated whitelist data to check the accuracy rate of whether the whitelist data test log data is incorrect. Specifically, when testing the whitelist data, the preset log data other than some log data belonging to the first label and / or belonging to the second label can be tested by the whitelist data to determine whether to update the whitelist data. That is to say, although the test data is part of the preset log data, it cannot be the same as the training data. By determining whether the whitelist data needs to be updated, the accuracy and timeliness of the whitelist data can be improved, thus avoiding the problem of system instability caused by missing the test of a certain type of log.
[0061] Reference Figure 5 As shown in the reference, the specific steps of testing the preset log data other than some log data belonging to the first label and / or belonging to the second label by the whitelist data to determine whether to update the whitelist data include step S501 and step S502.
[0062] In step S501, if the preset log data other than some log data belonging to the first label and belonging to the second label all pass the test by the whitelist data, the whitelist data is not updated.
[0063] In this step, when testing the whitelist data, the preset log data other than some log data used in steps S401 to S403 can be used as the test data. That is to say, the remaining log data belonging to the first label and all the preset log data belonging to the second label can be used as the test data to test the generated whitelist data. Specifically, the historical log data with all FAIL and some PASS historical log data are used as the test data. If the label of log data 1 is PASS, the result after testing by the whitelist data is still PASS. If the label of log data 2 is FAIL, the result after testing by the whitelist data is still FAIL. By the same method, if the whitelist data tests all PASS historical log data and the result is still PASS, and the whitelist data tests all FAIL historical log data and the result is still FAIL, it can be considered that the whitelist data is judged correctly and there is no need to adjust the whitelist data.
[0064] In step S502, if the preset log data other than part of the log data belonging to the first tag or the second tag fails the whitelist data test, the clustering algorithm is adjusted to update the whitelist data. Specifically, the historical log data with all FAIL and part of the historical log data with PASS are used as test data. If the label of log data 1 is PASS and the result after testing with the whitelist data is still PASS. If the label of log data 2 is FAIL and the result after testing with the whitelist data is PASS. If the result obtained by testing all historical log data with the whitelist data is different from the originally determined label PASS or FAIL, it can be considered that the whitelist data determination is incorrect and the whitelist data needs to be adjusted.
[0065] Since the whitelist is determined by clustering the preset log data and then obtaining the occurrence frequency of each type of log data, when adjusting the whitelist data, the parameters of the clustering algorithm need to be adjusted so as to continue through Figure 4 and Figure 5 The steps in are trained and tested to obtain accurate whitelist data. In addition, more historical log data can be collected for training to obtain more accurate whitelist data. In this exemplary embodiment, the remaining historical log data of the first tag and all historical log data of the second tag are analyzed and verified through the whitelist data. If the label obtained from the analysis result is different from the initially calibrated label, it is considered that there is a problem with the whitelist data determination result, and the parameters of the clustering algorithm are adjusted or more logs are collected to continue strengthening the learning of historical log data, upgrading and updating the whitelist data until the label tested by the whitelist data is consistent with the initially calibrated label.
[0066] Figure 6 The application process of the whitelist data generated in the offline mode is shown in, specifically including:
[0067] In step S601, historical log data LOG is obtained.
[0068] In step S602, the historical log data LOG is tested to determine the label of the historical log data according to whether the test passes. If not, the label of the historical log data is determined to be FAIL through step S603.
[0069] In step S604, if the historical log data LOG passes the test, the label of the historical log data is determined to be PASS.
[0070] In step S605, learning is performed on the historical log data with the label PASS.
[0071] In step S606, the whitelist data for regression testing is obtained by learning the historical log data with the label PASS.
[0072] In step S607, the log data in the regression testing is verified by the whitelist data.
[0073] Through the offline method and the historical log data, the whitelist data for regression testing can be generated, and the whitelist data that better conforms to the regression testing mode can be obtained.
[0074] For online testing, the whitelist data for online testing can be generated through the online log data and the real-time method. The whitelist data generated by the real-time method is mainly used for discriminating the log data during the online process to ensure that no new unknown errors will occur during the online process. It should be noted that the specific process of generating the whitelist data by the real-time method is the same as that of the offline method, that is Figure 3 and Figure 4 the steps in. Only the preset log data in the real-time method refers to the online log data instead of the historical log data. Therefore, the PASS LOG belonging to the first label used to generate the whitelist data also comes from the online log data.
[0075] Figure 7 The application process of the whitelist data generated by the real-time method is shown in
[0076] In step S701, the FileBeat tool is used to obtain the online log data LOG, and the label of the online log data is determined to be FAIL or PASS according to whether it is online.
[0077] In step S702, the online log data LOG with the label PASS is divided into two topics, the stable version and the online version, and stored in the Kafka cluster. Among them, FileBeat is a log file shipper tool. After installing the client on the server, Filebeat will monitor the log directory or the specified log file and track and read the content of these files.
[0078] In step S703, the Kafka cluster is consumed in real time, and the online log data of the two topics is analyzed to determine whether there are abnormal unknown errors. Among them, Kafka is a high-throughput distributed publish-subscribe messaging system used for producing and consuming data.
[0079] In step S704, the stable version LOG is obtained from the Kafka cluster.
[0080] In step S705, the whitelist data generated from the stable version of the online log data is obtained.
[0081] In step S706, the whitelist data is iterated in real time, the online log data of the online version machine is inspected, and it is determined whether there are any problems with the new version during the online process.
[0082] Specifically, after installing the client on the server, the log file shipping tool FileBeat can be used to monitor the log directory or the specified log file, and track and read the content of these files. Further, the online log data can be stored in the high-throughput distributed publish-subscribe messaging system Kafka for data production and consumption, and can be stored as a stable version and an online version respectively. By default, the code in the online environment can be considered the stable version. That is to say, in online testing, the whitelist data is generated based on the online log data of the stable version, so that the generated whitelist data better meets the requirements of online testing and improves the accuracy. In addition, the whitelist data for online testing can also be verified and updated until the labels tested by the whitelist data are consistent with the initially calibrated labels, so as to improve the accuracy rate of testing the log data through the whitelist data.
[0083] In addition to generating whitelist data, a preset rule can also be set. The preset rule refers to a rule set for specific types of log data. For example, whitelist data cannot be set for logical errors. For timeout errors, the number of occurrences that can be ignored is set to avoid large-area alarms in log analysis. In addition, the ratio control can be set to determine how much higher than the whitelist frequency to trigger an alarm. By setting the preset rule, the whitelist data can be reduced and the whitelist data can be made more accurate.
[0084] Next, in step S220, the target log data associated with the test mode is tested using the whitelist data to determine whether the target log data is normal.
[0085] In this exemplary embodiment, the target log data refers to the log data to be finally tested, and the target log data for regression testing or online testing may be different. For example, the target log data for regression testing can be the log data within a preset time, and the target log data for online testing can be the online log data of the online version stored in the distributed publish-subscribe messaging system Kafka. According to the process of testing with the whitelist data, mainly the machines after the new version is launched are compared with the machines without the new version launched in terms of logs, so as to ensure that the launched code version has no impact on the online business.
[0086] Specifically, the steps of testing the target log data associated with the test mode by using the whitelist data to determine whether the target log data is normal include: First step, clustering the target log data associated with the test mode by using a clustering algorithm to obtain multiple categories of target log data. Among them, the clustering algorithm can be the same as the clustering algorithm in step S402 above. The target log data of regression testing or online testing can be clustered respectively by using the clustering algorithm, and the specific clustering process can also be the same, which will not be elaborated here.
[0087] Second step, testing each category of target log data by using the whitelist data to determine whether each category of target log data is normal. Specifically, if the occurrence frequency of each category of target log data meets the whitelist data, it is determined that each category of target log data is normal. For example, for regression testing, after regression testing is completed, obtain the online log data within a preset time, classify the online log data by using a clustering algorithm, and then check whether the occurrence frequency of each category of online log data greatly exceeds the occurrence frequency in the whitelist data. If it does not exceed the occurrence frequency in the whitelist data, it is considered that each category of online log data within the preset time is normal.
[0088] If it exceeds the occurrence frequency in the whitelist data, that is, if the frequency is abnormal or a new type of problem appears, it is considered that the online log data within the preset time in the regression testing is abnormal.
[0089] If the test mode is online testing, then use the whitelist data generated from the online log data of the stable version to test each category of target log data (online log data) of the online version. The specific testing process is the same as the method of regression testing, which will not be elaborated here. Determine whether there is a problem in the target log data through the whitelist data of the online testing, and prevent the release version or online when there is a problem with the target log data to ensure the code quality of the online version. During the online process, continuously compare the online version and the stable version logs horizontally to ensure that the online code will not cause new problems.
[0090] It should be added that if it is determined that the occurrence frequency of the target log data does not meet the occurrence frequency in the whitelist data for regression testing or online testing, it is considered that the target log data is abnormal. When it is determined that the target log data is abnormal, a reminder message can be generated and sent to the client for display. The manifestation form of the reminder message can be an email, a message, etc. In this way, after detecting that the target log data is abnormal, the online process will be stopped at this time and analyzed. Send the reminder message of the target log data indicating the abnormality to the client of the developer or tester in the form of an email, which can facilitate the timely solution and troubleshooting of the abnormality, thereby improving the system stability.
[0091] In addition, during the whole process, the display template can be configured according to specific business modules, priority definitions can be made, and rules of different levels can be configured to achieve hierarchical analysis of target log data. By displaying target log data of different levels according to specific display rules, it is convenient to conduct troubleshooting, thus improving efficiency.
[0092] Figure 8 A schematic overall flowchart for analyzing log data is shown, specifically including the following steps:
[0093] In step S801, the test starts, which refers to regression testing or online testing here.
[0094] In step S802, the test is completed. If the test passes, go to step S803. If the test fails, go to step S809.
[0095] In step S803, preset log data LOG is collected.
[0096] In step S804, the preset log data LOG is processed, specifically including preprocessing, tagging, etc.
[0097] In step S805, it is judged whether to train or analyze the preset log data LOG.
[0098] In step S806, the preset log data LOG is trained to generate whitelist data.
[0099] In step S807, the preset log data LOG is analyzed and judged according to the whitelist data and preset rules.
[0100] In step S808, customized error reporting is performed according to the module.
[0101] In step S809, it is judged whether the log data generated by the whitelist data test passes.
[0102] In step S810, if it passes, online difference verification is performed on the log data. Mainly, the logs of the machines after the new version is launched are compared with those of the machines without the new version launched to ensure that the launched code version has no impact on the online business.
[0103] Through the method in this exemplary embodiment, the whitelist data is generated based on the training data with the label PASS in the preset log data corresponding to the test mode, and another part of the preset log data with labels is used as the test data to test the generated whitelist data, which can make the generated whitelist data more accurate. Further, the target log data associated with the test mode is automatically tested through the whitelist data, and the vertical comparison of log data can be realized. Compared with manual testing, the accuracy and efficiency of testing log data are improved. In addition, since the target log data can be automatically tested through the whitelist data and each log data can be tested, the omission problem caused by manual testing can be avoided, and the system stability is improved. The technical solution of the present invention does not have the problem of missing or misreading compared with manually checking the logs. At the same time, problems other than module testing can also be found through the logs, such as some problems of upstream and downstream teams. During the online process, the stability of the online is also ensured, the number of rollbacks caused by problems after going online is reduced, the stability of the online system is ensured, and potential problems can be found.
[0104] The present disclosure also provides a log data testing device. Refer to Figure 9 As shown, the log data testing device may include:
[0105] A whitelist data generation module 901, configured to generate whitelist data associated with the test mode according to the preset log data corresponding to the test mode;
[0106] A log testing module 902, configured to test the target log data associated with the test mode through the whitelist data to determine whether the target log data is normal.
[0107] In an exemplary embodiment of the present disclosure,
[0108] In an exemplary embodiment of the present disclosure, the whitelist data generation module includes: a preprocessing module, configured to preprocess the preset log data belonging to the first label; a frequency determination module, configured to determine the occurrence frequency of each type of log data in a part of the preprocessed log data; a whitelist generation control module, configured to determine the whitelist data according to the occurrence frequency.
[0109] In an exemplary embodiment of the present disclosure, the frequency determination module includes: a log classification module, configured to classify a part of the preprocessed log data through a clustering algorithm to obtain multiple types of log data; a frequency determination control module, configured to determine the occurrence frequency of each type of log data according to the maximum value of the occurrence times of each type of log data.
[0110] In an exemplary embodiment of the present disclosure, the apparatus further includes: a whitelist update module, configured to test preset log data other than the partial log data that belongs to the first tag and / or belongs to the second tag by using the whitelist data, so as to determine whether to update the whitelist data.
[0111] In an exemplary embodiment of the present disclosure, the whitelist update module includes: a first determination module, configured not to update the whitelist data if the preset log data other than the partial log data that belongs to the first tag and belongs to the second tag both pass the test of the whitelist data; a second determination module, configured to adjust the clustering algorithm to update the whitelist data if the preset log data other than the partial log data that belongs to the first tag or belongs to the second tag fails the test of the whitelist data.
[0112] In an exemplary embodiment of the present disclosure, the test mode includes regression testing and online testing, and the whitelist data generation module includes: an offline generation module, configured to generate, in an offline manner, whitelist data corresponding to the regression testing according to historical log data; or a real-time generation module, configured to generate, in a real-time manner, whitelist data corresponding to the online testing according to online log data of a stable version.
[0113] In an exemplary embodiment of the present disclosure, the log testing module includes: a target log clustering module, configured to cluster the target log data associated with the test mode by using a clustering algorithm to obtain multiple types of target log data; a test control module, configured to test each type of target log data by using the whitelist data to determine whether each type of target log data is normal.
[0114] In an exemplary embodiment of the present disclosure, the test control module includes: a first test module, configured to determine that each type of target log data is normal if the occurrence frequency of each type of target log data meets the whitelist data; a second test module, configured to generate a reminder message and send the reminder message to the client if the occurrence frequency of each type of target log data does not meet the whitelist data.
[0115] In an exemplary embodiment of the present disclosure, if the test mode is online testing, the target log data is online log data of an online version.
[0116] It should be noted that the specific details of each functional module in the above log data testing apparatus have been described in detail in the corresponding log data testing method, and thus will not be elaborated herein.
[0117] In this exemplary embodiment, a log data testing system 1000 is further provided. Refer to Figure 10As shown in the figure, the system includes: a server 1010, which is used to generate whitelist data associated with a test mode according to preset log data corresponding to the test mode, test target log data of the test mode according to the whitelist data, and send the test result to the client. Among them, the server may include:
[0118] An offline generation module 1011, which is used to generate whitelist data for regression testing in an offline manner according to historical log data in regression testing.
[0119] A real-time generation module 1012, which is used to generate whitelist data for online testing in a real-time manner according to online log data of a stable version in online testing.
[0120] A first log testing module 1013, which is used to test target log data within a preset time period according to the whitelist data corresponding to regression testing.
[0121] A second log testing module 1014, which is used to test online log data of the online version according to the whitelist data corresponding to online testing.
[0122] A client 1020, which is used to receive the test result returned by the server and display the test result. Specifically, according to specific business modules, a display template can be configured, priority definitions can be made, different-level rule configurations can be completed, and hierarchical analysis of logs can be achieved. Specifically, different-level log logs can be displayed according to specific display rules. For example, abnormal ordinary characters are displayed first, then abnormal new characters, then abnormal basic service characters, and finally critical errors, etc., to facilitate developers to conduct troubleshooting and analysis and process abnormal logs in a timely manner.
[0123] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units for embodiment.
[0124] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0125] Those skilled in the art can easily understand from the description of the above embodiments that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0126] In an exemplary embodiment of the present disclosure, there is also provided an electronic device capable of implementing the above method.
[0127] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0128] The following refers to Figure 11 to describe the electronic device 1100 according to this embodiment of the present invention. Figure 11 The electronic device 1100 shown is only an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present invention.
[0129] As Figure 11 shown, the electronic device 1100 is presented in the form of a general-purpose computing device. The components of the electronic device 1100 may include, but are not limited to: at least one of the above-mentioned processing units 1110, at least one of the above-mentioned storage units 1120, and a bus 1130 connecting different system components (including the storage unit 1120 and the processing unit 1110).
[0130] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 1110, so that the processing unit 1110 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 1110 can execute the steps as Figure 2 shown.
[0131] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 11201 and / or a cache storage unit 11202, and may further include a read-only storage unit (ROM) 11203.
[0132] The storage unit 1120 may also include a program / utility 11204 having a set (at least one) of program modules 11205. Such program modules 11205 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0133] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0134] The display unit 1140 may be a display with display functions to display the processing results obtained by the processing unit 1110 executing the method in this exemplary embodiment through the display. The display includes, but is not limited to, a liquid crystal display or other displays.
[0135] The electronic device 1100 may also communicate with one or more external devices 1300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 1100, and / or communicate with any device that enables the electronic device 1100 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 1150. Moreover, the electronic device 1100 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1160. As shown in the figure, the network adapter 1160 communicates with other modules of the electronic device 1100 through the bus 1130. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0136] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having stored thereon a program product capable of implementing the above-described method of this specification. In some possible implementation manners, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0137] Reference Figure 12As shown, a program product 1200 for implementing the above method according to an embodiment of the present invention is described. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited to this. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus.
[0138] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0139] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or apparatus.
[0140] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0141] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0142] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.
[0143] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are to be considered as exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A method for testing log data, characterized in that, Including: Generating whitelist data associated with a test mode according to preset log data corresponding to the test mode, including: obtaining a first tag or a second tag corresponding to the preset log data; generating the whitelist data by analyzing some of the log data in the preset log data belonging to the first tag; Testing target log data associated with the test mode by means of the whitelist data to determine whether the target log data is normal; Among them, generating the whitelist data includes: preprocessing the preset log data belonging to the first tag; classifying some of the preprocessed log data by means of a clustering algorithm to obtain multiple categories of log data; determining the occurrence frequency of each category of log data according to the maximum value of the occurrence times of each category of log data; determining the whitelist data according to the occurrence frequency; The method further includes: Testing preset log data other than the some of the log data that belongs to the first tag and / or belongs to the second tag by means of the whitelist data to determine whether to update the whitelist data.
2. The log data testing method according to claim 1, wherein Testing preset log data other than the some of the log data that belongs to the first tag and / or belongs to the second tag by means of the whitelist data to determine whether to update the whitelist data includes: If the preset log data other than the some of the log data that belongs to the first tag and belongs to the second tag all pass the test of the whitelist data, the whitelist data is not updated; If the preset log data other than the some of the log data that belongs to the first tag or belongs to the second tag fails the test of the whitelist data, the clustering algorithm is adjusted to update the whitelist data.
3. The log data testing method according to claim 1, characterized in that The test mode includes regression testing and online testing. Generating whitelist data associated with the test mode according to preset log data corresponding to the test mode includes: Adopting an offline method to generate whitelist data corresponding to the regression testing according to historical log data; or Adopting a real-time method to generate whitelist data corresponding to the online testing according to online log data of a stable version.
4. The log data testing method according to claim 1, wherein Testing target log data associated with the test mode by means of the whitelist data to determine whether the target log data is normal includes: Clustering the target log data associated with the test mode by means of a clustering algorithm to obtain multiple categories of target log data; Testing each category of target log data by means of the whitelist data to determine whether each category of target log data is normal.
5. The log data testing method according to claim 4, wherein Testing each category of target log data by means of the whitelist data to determine whether each category of target log data is normal includes: If the occurrence frequency of each category of target log data meets the whitelist data, it is determined that each category of target log data is normal; If the occurrence frequency of each category of the target log data does not meet the whitelist data, a reminder message is generated and the reminder message is sent to the client.
6. The log data testing method according to claim 4 or 5, characterized in that, If the test mode is online testing, the target log data is online log data of a released version.
7. A log data testing device, characterized in that Including: A whitelist data generation module, configured to generate whitelist data associated with a test mode according to preset log data corresponding to the test mode, including: obtaining a first tag or a second tag corresponding to the preset log data; generating the whitelist data by analyzing part of the log data in the preset log data belonging to the first tag; A log testing module, configured to test target log data associated with the test mode by using the whitelist data to determine whether the target log data is normal; Wherein, generating the whitelist data includes: preprocessing the preset log data belonging to the first tag; classifying part of the preprocessed log data through a clustering algorithm to obtain multiple classes of log data; determining the occurrence frequency of each class of log data according to the maximum value of the occurrence times of each class of log data; determining the whitelist data according to the occurrence frequency; The device is further configured to test preset log data other than the part of the log data belonging to the first tag and / or belonging to the second tag by using the whitelist data to determine whether to update the whitelist data.
8. A log data testing system, characterized in that, Including: A server, configured to generate whitelist data associated with a test mode according to preset log data corresponding to the test mode, test target log data of the test mode by using the whitelist data, and send a test result to a client; wherein, generating the whitelist data associated with the test mode includes: obtaining a first tag or a second tag corresponding to the preset log data; generating the whitelist data by analyzing part of the log data in the preset log data belonging to the first tag; A client, configured to receive the test result returned by the server and display the test result; Wherein, generating the whitelist data includes: preprocessing the preset log data belonging to the first tag; classifying part of the preprocessed log data through a clustering algorithm to obtain multiple classes of log data; determining the occurrence frequency of each class of log data according to the maximum value of the occurrence times of each class of log data; determining the whitelist data according to the occurrence frequency; The server is further configured to test preset log data other than the part of the log data belonging to the first tag and / or belonging to the second tag by using the whitelist data to determine whether to update the whitelist data.
9. An electronic device, characterized in that, Including: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the log data testing method according to any one of claims 1-6 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the log data testing method according to any one of claims 1-6.
Citation Information
Patent Citations
Real-time online log detection method and system
CN103514398A
Log testing method and device
CN106294151A