A method for analyzing and locating problems in failure cases based on decision trees

By using a decision tree-based method in automated testing, samples preprocessing, feature extraction and information gain calculation are performed on failed cases, and the problem of analyzing and positioning failure cases is solved, which is unable to effectively locate and solve failed cases in the existing technology, and efficient and accurate analysis and positioning are achieved.

CN114218094BActive Publication Date: 2025-06-17CHINA CITIC BANK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111501968.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-06-17
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

In the analysis of failure cases caused by the existing technology during automated testing, it can only provide preliminary statistical results or classification results, and cannot further locate the root cause of the problem or provide solutions, resulting in a large amount of manual analysis costs.

Method used

Through a decision tree-based method, sample preprocessing, feature extraction and information gain calculation are performed on failed cases, and a decision tree is generated to analyze and locate the problems of failed cases.

Benefits of technology

It has realized the classification and precise positioning of automated problems of failed cases, significantly improving analysis efficiency and accuracy, and reducing manual analysis costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218094B_ABST
    Figure CN114218094B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for analyzing and locating problems of failure cases based on decision trees, which relates to the field of information technology. The method includes: performing sample preprocessing on a first sample set to obtain a first preprocessing result; obtaining a p-th feature list of the first preprocessing result; calculating the feature gains of each feature in the p-th feature list to obtain a p-th information gain set; selecting the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and deleting the p-th feature from the p-th feature list; generating a decision tree according to the corresponding features of the n sample subsets; and analyzing and locating problems of failure cases based on the decision tree. It solves the technical problem that the existing technical solutions only give preliminary statistical results or preliminary classification results, without further locating the root cause of the problem or giving a solution, so a large amount of manual analysis costs are still required.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for analyzing and locating problems of failure cases based on a decision tree. Background Art

[0002] During the automated testing process, with the high-frequency and efficient execution of automated tests, a large number of execution failure cases will be generated in a short time, which brings a large amount of manual analysis and location costs to us. Currently, the widespread solution to reduce labor costs is to assume that there already exists a set of error keywords for classification, and based on keyword matching, a preliminary classification result of the failure cases is given. The existing solutions only give preliminary statistical results or preliminary classification results.

[0003] However, in the process of implementing the technical solution of the present invention in the embodiments of the present application, the inventors of the present application found that the above technology has at least the following technical problems:

[0004] The existing technical solutions only give preliminary statistical results or preliminary classification results, without further locating the root cause of the problem or giving a solution, so there is still a problem of a large amount of manual analysis costs. Summary of the Invention

[0005] Embodiments of the present application provide a method for analyzing and locating problems of failure cases based on a decision tree, which solves the technical problem that the existing technical solutions only give preliminary statistical results or preliminary classification results, without further locating the root cause of the problem or giving a solution, so there is still a large amount of manual analysis costs. By constructing a decision tree technology for predictive classification based on failure cases, it effectively realizes the automated problem classification of failure cases. In addition, an automated location technology for various problems is provided, which realizes the effective and accurate location of failure cases, and effectively improves the analysis efficiency and accuracy of failure cases.

[0006] In view of the above problems, the present invention is proposed to provide a method that overcomes the above problems or at least partially solves the above problems.

[0007] In a first aspect, an embodiment of the present application provides a method for analyzing and locating failure case problems based on a decision tree. The method includes: obtaining a first sample set; performing sample preprocessing on the first sample set to obtain a first preprocessing result, where the first preprocessing result includes n sample subsets, and n is a positive integer greater than 1; obtaining the p-th sample subset of the first preprocessing result, performing feature extraction on the p-th sample subset to obtain a p-th feature list, where p is a positive integer less than or equal to n; calculating the information gain of each feature in the p-th feature list for the p-th sample subset to obtain a p-th information gain set; selecting the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and deleting the p-th feature from the p-th feature list; repeating the above steps until corresponding features are obtained for all n sample subsets, then stopping the repetition, and generating a decision tree according to the corresponding features of the n sample subsets; analyzing and locating failure case problems based on the decision tree.

[0008] In another aspect, the present application further provides a system for analyzing and locating failure case problems based on a decision tree. The system includes: a first obtaining unit for obtaining a first sample set; a second obtaining unit for performing sample preprocessing on the first sample set to obtain a first preprocessing result, where the first preprocessing result includes n sample subsets, and n is a positive integer greater than 1; a third obtaining unit for obtaining the p-th sample subset of the first preprocessing result, performing feature extraction on the p-th sample subset to obtain a p-th feature list, where p is a positive integer less than or equal to n; a fourth obtaining unit for calculating the information gain of each feature in the p-th feature list for the p-th sample subset to obtain a p-th information gain set; a first selecting unit for selecting the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and deleting the p-th feature from the p-th feature list; a first generating unit for repeating the operations of the third obtaining unit to the first selecting unit until corresponding features are obtained for all n sample subsets, then stopping the repetition, and generating a decision tree according to the corresponding features of the n sample subsets; a first analyzing unit for analyzing and locating failure case problems based on the decision tree.

[0009] In a third aspect, an embodiment of the present invention provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor. The transceiver, the memory, and the processor are connected through the bus. When the computer program is executed by the processor, it implements the steps in the method for controlling and outputting data described in any one of the above.

[0010] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for controlling and outputting data described in any one of the above.

[0011] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0012] By performing sample preprocessing on the first sample set to obtain a first preprocessing result; obtaining the p-th sample subset of the first preprocessing result, performing feature extraction on the p-th sample subset to obtain a p-th feature list; calculating the information gain of each feature in the p-th feature list for the p-th sample subset to obtain a p-th information gain set; selecting the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and deleting the p-th feature from the p-th feature list; until corresponding features are obtained for all n sample subsets, stopping the repetition of the steps, generating a decision tree based on the corresponding features of the n sample subsets; and analyzing and locating the problems of failure cases based on the decision tree. Thus, through the technology of constructing a decision tree for predictive classification based on failure cases, the effective automatic problem classification of failure cases is realized. In addition, an automatic positioning technology for various problems is provided, realizing the effective and accurate positioning of failure cases, and effectively improving the analysis efficiency and accuracy of failure cases.

[0013] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the following specific embodiments of the present application are specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a flowchart of a method for analyzing and locating failure case problems based on a decision tree according to an embodiment of the present application;

[0015] Figure 2 It is a flowchart of sample preprocessing of the first sample set in a method for analyzing and locating failure case problems based on a decision tree according to an embodiment of the present application;

[0016] Figure 3 This is a schematic flowchart for defect location by analyzing the location result in a method for analyzing and locating failure case problems based on a decision tree according to an embodiment of the present application;

[0017] Figure 4 This is a schematic flowchart for providing a reference solution for the analysis and location result in a method for analyzing and locating failure case problems based on a decision tree according to an embodiment of the present application;

[0018] Figure 5 This is a schematic structural diagram of a system for analyzing and locating failure case problems based on a decision tree according to an embodiment of the present application;

[0019] Figure 6 This is a schematic structural diagram of an electronic device for a method of executing control output data provided by an embodiment of the present application.

[0020] Explanation of reference numerals: First acquisition unit 11, second acquisition unit 12, third acquisition unit 13, fourth acquisition unit 14, first selection unit 15, first generation unit 16, first analysis unit 17, bus 1110, processor 1120, transceiver 1130, bus interface 1140, memory 1150, operating system 1151, application program 1152, and user interface 1160. Detailed implementation manners

[0021] In the description of the embodiments of the present invention, those skilled in the art should know that the embodiments of the present invention can be implemented as a method, a device, an electronic device, and a computer-readable storage medium. Therefore, the embodiments of the present invention can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), and a combination of hardware and software. In addition, in some embodiments, the embodiments of the present invention can also be implemented in the form of a computer program product in one or more computer-readable storage media, and the computer-readable storage media contains computer program code.

[0022] The above computer-readable storage media can be any combination of one or more computer-readable storage media. Computer-readable storage media include: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media include: portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, flash memories, optical fibers, compact disc read-only memories, optical storage devices, magnetic storage devices, or any combination of the above. In the embodiments of the present invention, the computer-readable storage media can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or component.

[0023] Application overview

[0024] Embodiments of the present invention describe the provided methods, apparatuses, and electronic devices through flowcharts and / or block diagrams.

[0025] It should be understood that each block of the flowchart and / or block diagram, as well as the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine. These computer-readable program instructions, when executed by a computer or other programmable data processing device, produce a device that implements the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0026] These computer-readable program instructions can also be stored in a computer-readable storage medium that can cause a computer or other programmable data processing device to work in a specific manner. In this way, the instructions stored in the computer-readable storage medium produce an instruction device product that includes instructions for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0027] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device, to produce a computer-implemented process. Thus, the instructions executed on the computer or other programmable data processing device can provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0028] The embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0029] Example 1

[0030] As Figure 1 shown, an embodiment of the present application provides a method for analyzing and locating failure case problems based on a decision tree. Among them, the method includes:

[0031] Step S100: Obtain a first sample set;

[0032] Specifically, during the automated testing process, with the high-frequency and efficient execution of automated testing, a large number of execution failure cases will be generated in a short time. The first sample set is a set of error messages for the execution failure cases in the automated testing, providing a case data basis for subsequent analysis of failure case problems.

[0033] Step S200: Perform sample preprocessing on the first sample set to obtain a first preprocessing result, where the first preprocessing result includes n sample subsets, where n is a positive integer greater than 1;

[0034] As Figure 2 shown, furthermore, wherein, the step S200 further includes:

[0035] Step S210: Obtain the first failure problem classification criterion;

[0036] Step S220: Classify the first sample set based on the first failure problem classification criterion to obtain a first classification result;

[0037] Step S230: Analyze and mark the reasons for failure in the first classification result to obtain the first preprocessing result.

[0038] Specifically, perform sample preprocessing on the first sample set. Based on the first failure problem classification criterion, that is, according to the different coping measures of development and testing personnel, classify the failure problems of the first sample set to obtain a first classification result. The first classification result includes system defect class, test environment class, and case data class. Analyze and mark the reasons for failure in the first classification result to obtain the analyzed first preprocessing result. Among them, the first preprocessing result includes n sample subsets, where n is a positive integer greater than 1. Preprocess the samples to make subsequent feature extraction more accurate.

[0039] Step S300: Obtain the p-th sample subset of the first preprocessing result, and perform feature extraction on the p-th sample subset to obtain a p-th feature list, where p is a positive integer less than or equal to n;

[0040] Step S400: Calculate the information gain of each feature in the p-th feature list for the p-th sample subset to obtain a p-th information gain set;

[0041] Specifically, the p-th sample subset is a preprocessed sample subset included in the n sample subsets. Perform feature extraction on the p-th sample subset, segment the error message of the failure case, and extract feature words by removing duplicates and stop words to obtain the corresponding p-th feature list, where p is a positive integer less than or equal to n. According to the information gain, select the classification features of the decision tree, and calculate the information gain of each feature in the p-th feature list for the p-th sample subset to obtain the p-th information gain set corresponding to the gain calculation.

[0042] Step S500: Select the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and delete the p-th feature from the p-th feature list;

[0043] Step S600: Repeat steps S300 to S500 until corresponding features are obtained for all the n sample subsets, then stop repeating the steps, and generate a decision tree based on the corresponding features of the n sample subsets.

[0044] Specifically, compare the information gains of each feature, select the feature with the largest information gain as the optimal feature for the current classification, and then delete it from the feature list. That is, select the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and delete the p-th feature from the p-th feature list, indicating that this feature has been used. Then recursively select the feature with the largest information gain for each sample subset until corresponding features are obtained for all the n sample subsets, then stop repeating steps S300 to S500, and generate a decision tree for problem classification based on the corresponding features of the n sample subsets.

[0045] Step S700: Analyze and locate the problems of failure cases based on the decision tree.

[0046] Specifically, analyze and locate the problems of failure cases based on the decision tree. When a case fails, use the error message of the failure case as the input of the decision tree, and classify the problem for this failure according to the final result hit by executing the decision tree. By using the decision tree technology for predicting and classifying based on failure cases, it effectively realizes automatic problem classification for failure cases, greatly reduces the human troubleshooting cost, and effectively improves the analysis efficiency and accuracy of failure cases.

[0047] As Figure 3 shown, furthermore, embodiment S230 of the present application further includes:

[0048] Step S231: Locate the defects according to the first classification result to obtain a first defect location set.

[0049] Step S232: Construct a mapping relationship based on the first defect location set and the first classification result to obtain a first mapping relationship construction result.

[0050] Step S233: After analyzing and locating the problems of failure cases based on the decision tree, locate the defects by analyzing the location result and the first mapping relationship construction result.

[0051] Specifically, defect localization is performed according to the first classification result to obtain a corresponding first defect localization set. The first classification result includes system defect category, test environment category, and case data category. Based on the first defect localization set and the first classification result, a mapping relationship is constructed to obtain a first mapping relationship construction result. After analyzing and locating the problem of the failure case based on the decision tree, defect localization is performed by analyzing the localization result and the first mapping relationship construction result. For example, for the system defect category, a mapping between code functions and test cases is established through a code coverage statistics tool, and then during the execution of the test case, which functions in the program are executed is counted, so as to locate the program and code corresponding to the test case, and finally the mapping relationship from the test case to the function can be obtained. When there is a failure case of the system defect category, according to the mapping relationship between the test case and the code, the list of functions covered by this failure case is searched, and the defective program can be located; for the test environment category, according to the pre-set transaction-system mapping relationship topology diagram, the system link passed by the transaction and the upstream and downstream call relationships between systems can be viewed. When a certain transaction is located as an environment problem, the link of this transaction is obtained from the mapping diagram according to the system, and starting from the most downstream system T of the transaction link, the error log is successfully retrieved, then the abnormal environment of system T can be located; if there is no serial number log in system T, then system T-1 is retrieved and located... until the problem system is successfully located; for the case data category, the problem field is located. After analyzing and locating the problem of the failure case based on the decision tree, defect localization is performed by analyzing the localization result and the first mapping relationship construction result, which realizes the automated analysis and localization of the failure case and effectively improves the analysis efficiency of the failure case.

[0052] As Figure 4 shown, furthermore, the embodiment of the present application further includes:

[0053] Step S810: Construct a solution library for failure cases based on the case execution status;

[0054] Step S820: Obtain the analysis and localization result of the decision tree;

[0055] Step S830: Provide a reference solution for the analysis and localization result based on the solution library for failure cases.

[0056] Specifically, by constructing a failure case solution library based on the case execution status, solutions are automatically recommended to cases with the same error message. First, based on the transformation of the case execution status, the data configuration, transaction code, and error message of the same case executed successfully and the last execution failed are recorded to construct a failure case solution library. Based on the failure case solution library, reference solutions are provided for the analysis and positioning results of the decision tree. For example, when there is a failure case with data problems, using the transaction code and error message as keys, records with the same transaction code and error message as the failure case are retrieved from the solution library, and the data configuration and modified fields of the historical successful cases are recommended to the testers, providing ideas for the testers to solve the problems. Solutions are automatically recommended to cases with the same error message, realizing the automated analysis and positioning of failure cases and effectively improving the analysis efficiency of failure cases.

[0057] Furthermore, the embodiments of the present application further include:

[0058] Step S910: Obtain the first information entropy of the p-th sample subset;

[0059] Step S920: Obtain the second information entropy of the p-th sample subset after removing the p-th feature;

[0060] Step S930: Obtain the information gain of the p-th feature with respect to the p-th sample subset based on the first information entropy and the second information entropy.

[0061] Furthermore, the information gain of the p-th feature with respect to the p-th sample subset is calculated by the formula, and the calculation formula is as follows:

[0062] g(D,p) = H(D) - H(D|p)

[0063] where g(D,p) is the information gain, H(D) is the first information entropy, and H(D|p) is the second information entropy.

[0064] Specifically, calculate the information gain of each feature in the p-th feature list for the p-th sample subset. Specifically, denote the set of unclassified failure cases as the data set D, and its information entropy is the first information entropy Y1 of the p-th sample subset. After removing the failure cases containing the p-th feature, the information entropy of this set drops to the second information entropy Y2 of the p-th sample subset. Based on the first information entropy and the second information entropy, obtain the information gain of the p-th feature for the p-th sample subset. Then, the information gain of feature P for the data set D is Y1 - Y2. Denote the information gain of feature P for the data set D as g(D, p), and the calculation formula is as follows: g(D, p) = H(D) - H(D|p), where g(D, p) is the information gain, H(D) is the first information entropy, and H(D|p) is the second information entropy. For the calculation of H(D), according to the formula of information entropy: It can be obtained that:

[0065] H(D) = -(probability of system defect * log2 probability of system defect + probability of environmental problem * log2 probability of environmental problem + probability of data problem * log2 probability of data problem);

[0066] For H(D|p), H(D|p) is the empirical conditional entropy of D under the condition of feature p. The conditional entropy H(Y|X) represents the uncertainty of random variable Y under the condition of known random variable X, and its calculation formula is: H(Y|X) = ∑ j=1 p i H(Y|X = x i ).

[0067] In the present invention, use feature A1 to represent "the error message of the failure case contains feature word a1". Feature A1 divides the data D into two subsets D1 and D2, and D1 and D2 respectively represent the sample subsets in D where the error message contains feature word a1 and does not contain feature word a1. Then H(D|A1) = probability of containing feature word a1 * H(D1) + probability of not containing feature word a1 * H(D2), g(D, A1) = H(D) - probability of containing feature word a1 * H(D1) + probability of not containing feature word a1 * H(D2), where H(D), H(D1), and H(D2) are calculated according to the formula of entropy.

[0068] Use feature A2 to represent "the error message of the failure case contains feature word a2". D1 and D2 respectively represent the sample subsets in D where the error message contains feature word a2 and does not contain feature word a2. Then g(D, A2) = H(D) - probability of containing feature word a2 * H(D1) + probability of not containing feature word a2 * H(D2). According to the information gain, select the classification features of the decision tree to make the construction of the decision tree for prediction classification more accurate and improve the accuracy of case problem classification.

[0069] Furthermore, the embodiments of the present application further include:

[0070] Step S1010: Obtain the first predetermined decision tree generation condition;

[0071] Step S1020: Determine whether the decision tree generated in step S600 meets the first predetermined decision tree generation condition;

[0072] Step S1030: When the decision tree generated in step S600 meets the first predetermined decision tree generation condition, complete the construction of the decision tree.

[0073] Specifically, there are two first predetermined decision tree generation conditions: one is that the divided data all belong to one class, and the other is that all features have been used. In the second end case, the divided data may not belong to one category, and the classification of this sub-data set needs to be determined according to the majority voting criterion. Determine whether the decision tree generated in step S600 meets the first predetermined decision tree generation condition. When the decision tree generated in step S600 meets the first predetermined decision tree generation condition, that is, the end condition for generating the decision tree is met, and the construction of the decision tree is completed, effectively realizing the technical effect of automatically classifying problems for failure cases and then improving the analysis efficiency.

[0074] In summary, the method for analyzing and locating failure case problems based on a decision tree provided by the embodiments of the present application has the following technical effects:

[0075] By performing sample preprocessing on the first sample set to obtain the first preprocessing result; obtaining the p-th sample subset of the first preprocessing result, performing feature extraction on the p-th sample subset to obtain the p-th feature list; calculating the information gain of each feature in the p-th feature list for the p-th sample subset to obtain the p-th information gain set; selecting the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and deleting the p-th feature from the p-th feature list; until corresponding features are obtained for all n sample subsets, stop repeating the steps, generate a decision tree according to the corresponding features of the n sample subsets; and perform analysis and location of failure case problems based on the decision tree. Furthermore, by using the technology of constructing a predictive classification decision tree based on failure cases, it is effectively realized to automatically classify problems for failure cases. In addition, an automatic location technology for various problems is provided, realizing the effective and accurate location of failure cases, and effectively improving the analysis efficiency and accuracy of failure cases.

[0076] Example 2

[0077] Based on the same inventive concept as the method for analyzing and locating failure case problems based on a decision tree in the foregoing embodiments, the present invention further provides a system for analyzing and locating failure case problems based on a decision tree, as Figure 5 shown, the system includes:

[0078] A first acquisition unit 11, which is used to acquire a first sample set;

[0079] A second acquisition unit 12, which is used to perform sample preprocessing on the first sample set to obtain a first preprocessing result, wherein the first preprocessing result includes n sample subsets, where n is a positive integer greater than 1;

[0080] A third acquisition unit 13, which is used to acquire the p-th sample subset of the first preprocessing result, perform feature extraction on the p-th sample subset to obtain a p-th feature list, where p is a positive integer less than or equal to n;

[0081] A fourth acquisition unit 14, which is used to calculate the information gain of each feature in the p-th feature list for the p-th sample subset to obtain a p-th information gain set;

[0082] A first selection unit 15, which is used to select the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and delete the p-th feature from the p-th feature list;

[0083] A first generation unit 16, which is used to repeat the third acquisition unit to the first selection unit until corresponding features are obtained for all n sample subsets, stop repeating the steps, and generate a decision tree according to the corresponding features of the n sample subsets;

[0084] A first analysis unit 17, which is used to analyze and locate failure case problems based on the decision tree.

[0085] Further, the system further includes:

[0086] A fifth acquisition unit, which is used to acquire a first failure problem classification criterion;

[0087] A sixth acquisition unit, which is used to classify the first sample set based on the first failure problem classification criterion to obtain a first classification result;

[0088] A seventh acquisition unit, which is used to analyze and mark the failure reasons for the first classification result to obtain the first preprocessing result.

[0089] Further, the system further includes:

[0090] An eighth acquisition unit, configured to perform defect localization according to the first classification result to obtain a first defect localization set;

[0091] A ninth acquisition unit, configured to construct a mapping relationship based on the first defect localization set and the first classification result to obtain a first mapping relationship construction result;

[0092] A first localization unit, configured to perform defect localization by analyzing the localization result and the first mapping relationship construction result after analyzing and localizing the failure case problem based on the decision tree.

[0093] Further, the system further includes:

[0094] A first construction unit, configured to construct a failure case solution library based on the case execution status;

[0095] A tenth acquisition unit, configured to acquire the analysis and localization result of the decision tree;

[0096] A first reference unit, configured to provide a reference solution for the analysis and localization result based on the failure case solution library.

[0097] Further, the system further includes:

[0098] An eleventh acquisition unit, configured to acquire the first information entropy of the p-th sample subset;

[0099] A twelfth acquisition unit, configured to acquire the second information entropy of the p-th sample subset after the p-th feature is removed;

[0100] A first gain unit, configured to obtain the information gain of the p-th feature for the p-th sample subset based on the first information entropy and the second information entropy.

[0101] Further, the system further includes:

[0102] A thirteenth acquisition unit, configured to acquire a first predetermined decision tree generation condition;

[0103] A first judgment unit, configured to judge whether the decision tree generated by the first generation unit satisfies the first predetermined decision tree generation condition;

[0104] A second construction unit, which is used to complete the construction of the decision tree when the decision tree generated by the first generation unit meets the first predetermined decision tree generation condition.

[0105] The foregoing Figure 1 All the various change methods and specific examples of the method for analyzing and locating failure case problems based on a decision tree in the first embodiment also apply to the system for analyzing and locating failure case problems based on a decision tree in this embodiment. Through the foregoing detailed description of the method for analyzing and locating failure case problems based on a decision tree, those skilled in the art can clearly know the implementation method of the system for analyzing and locating failure case problems based on a decision tree in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated herein.

[0106] In addition, an embodiment of the present invention further provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are respectively connected through the bus. When the computer program is executed by the processor, it implements each process of the method embodiment for controlling the output data, and can achieve the same technical effect. To avoid repetition, it will not be described herein again.

[0107] Exemplary electronic device

[0108] Specifically, referring to Figure 6 As shown, an embodiment of the present invention further provides an electronic device, which includes a bus 1110, a processor 1120, a transceiver 1130, a bus interface 1140, a memory 1150, and a user interface 1160.

[0109] In an embodiment of the present invention, the electronic device further includes: a computer program stored in the memory 1150 and executable on the processor 1120. When the computer program is executed by the processor 1120, it implements each process of the method embodiment for controlling the output data.

[0110] The transceiver 1130 is used to receive and send data under the control of the processor 1120.

[0111] In an embodiment of the present invention, the bus architecture (represented by the bus 1110), the bus 1110 may include any number of interconnected buses and bridges. The bus 1110 connects various circuits including one or more processors represented by the processor 1120 and the memory represented by the memory 1150 together.

[0112] Bus 1110 represents one or more of any of several types of bus structures, including a memory bus and memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include: Industry Standard Architecture bus, Micro Channel Architecture bus, Extended bus, Video Electronics Standards Association, Peripheral Component Interconnect bus.

[0113] Processor 1120 may be an integrated circuit chip having signal processing capabilities. In implementation, the steps of the above method embodiments may be completed by the integrated logic circuit in the processor or instructions in software form. The above-mentioned processor includes: general-purpose processor, central processor, network processor, digital signal processor, application specific integrated circuit, field programmable gate array, complex programmable logic device, programmable logic array, micro control unit or other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. For example, the processor may be a single-core processor or a multi-core processor, and the processor may be integrated on a single chip or located on multiple different chips.

[0114] Processor 1120 may be a microprocessor or any conventional processor. The method steps disclosed in combination with the embodiments of the present invention may be directly executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a readable storage medium well known in the art such as random access memory, flash memory, read only memory, programmable read only memory, erasable programmable read only memory, register, etc. The readable storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0115] Bus 1110 may also connect together various other circuits, such as peripheral devices, voltage regulators, or power management circuits, etc. The bus interface 1140 provides an interface between the bus 1110 and the transceiver 1130, which are all well known in the art. Therefore, the embodiments of the present invention will not be further described herein.

[0116] The transceiver 1130 may be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. For example: the transceiver 1130 receives external data from other devices, and the transceiver 1130 is used to send the data processed by the processor 1120 to other devices. Depending on the nature of the computer device, a user interface 1160 may also be provided, such as: touch screen, physical keyboard, display, mouse, speaker, microphone, trackball, joystick, stylus.

[0117] It should be understood that in the embodiments of the present invention, the memory 1150 may further include a memory remotely disposed relative to the processor 1120, and these remotely disposed memories can be connected to the server through a network. One or more parts of the above network may be an ad hoc network, an intranet, an extranet, a virtual private network, a local area network, a wireless local area network, a wide area network, a wireless wide area network, a metropolitan area network, the Internet, a public switched telephone network, a plain old telephone service network, a cellular telephone network, a wireless network, a Wi-Fi network, and a combination of two or more of the above networks. For example, the cellular telephone network and the wireless network may be a global mobile communication device, a code division multiple access device, a worldwide interoperability for microwave access device, a general packet radio service device, a wideband code division multiple access device, a long term evolution device, an LTE frequency division duplex device, an LTE time division duplex device, an advanced long term evolution device, a universal mobile telecommunications system device, an enhanced mobile broadband device, a massive machine type communication device, an ultra-reliable low latency communication device, etc.

[0118] It should be understood that the memory 1150 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory includes: read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, or flash memory.

[0119] The volatile memory includes: random access memory, which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as: static random access memory, dynamic random access memory, synchronous dynamic random access memory, double data rate synchronous dynamic random access memory, enhanced synchronous dynamic random access memory, synchronous link dynamic random access memory, and direct memory bus random access memory. The memory 1150 of the electronic device described in the embodiments of the present invention includes but is not limited to the above and any other suitable types of memory.

[0120] In the embodiments of the present invention, the memory 1150 stores the following elements of the operating system 1151 and the application program 1152: executable modules, data structures, or subsets thereof, or extended sets thereof.

[0121] Specifically, the operating system 1151 includes various device programs, such as: a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 1152 includes various application programs, such as: a media player, a browser, for implementing various application services. The program for implementing the method of the embodiments of the present invention may be included in the application program 1152. The application program 1152 includes: applets, objects, components, logics, data structures, and other computer device executable instructions for performing specific tasks or implementing specific abstract data types.

[0122] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the method embodiment for controlling the output data is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0123] As mentioned above, the above is only the specific implementation manner of the embodiment of the present invention, but the protection scope of the embodiment of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiment of the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the embodiment of the present invention. Therefore, the protection scope of the embodiment of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for analyzing and locating problems in failure cases based on decision trees, wherein, The method includes: Step S100: Obtain a first sample set; Step S200: Perform sample preprocessing on the first sample set to obtain a first preprocessing result, where the first preprocessing result includes n sample subsets, and n is a positive integer greater than 1; Step S300: Obtain the p-th sample subset of the first preprocessing result, perform feature extraction on the p-th sample subset to obtain a p-th feature list, where p is a positive integer less than or equal to n; Step S400: Calculate the information gain of each feature in the p-th feature list for the p-th sample subset to obtain a p-th information gain set; Step S500: Select the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and delete the p-th feature from the p-th feature list; Step S600: Repeat steps S300 to S500 until corresponding features are obtained for all the n sample subsets, then stop repeating the steps, and generate a decision tree based on the corresponding features of the n sample subsets; Step S700: Analyze and locate the problems of failure cases based on the decision tree; Among them, the method further includes: Obtain the first information entropy of the p-th sample subset; Obtain the second information entropy of the p-th sample subset after removing the p-th feature; Obtain the information gain of the p-th feature for the p-th sample subset based on the first information entropy and the second information entropy; Calculate the information gain of the p-th feature for the p-th sample subset through a formula, and the calculation formula is as follows: g(D,p) = H(D) - H(D|p) where g(D,p) is the information gain, H(D) is the first information entropy, and H(D|p) is the second information entropy; Calculating the information gain of each feature in the p-th feature list for the p-th sample subset specifically means recording the set of unclassified failure cases as dataset D, and its information entropy is the first information entropy Y1 of the p-th sample subset; After removing the failure cases containing the p-th feature, the information entropy of the set drops to the second information entropy Y2 of the p-th sample subset. Based on the first information entropy and the second information entropy, obtain the information gain of the p-th feature for the p-th sample subset. Then the information gain of feature P for dataset D is Y1 - Y2; Among them, for the calculation of H(D), according to the formula of information entropy: H(D) = -(probability of system defect * log2 probability of system defect + probability of environmental problem * log2 probability of environmental problem + probability of data problem * log2 probability of data problem); H(D|p) is the empirical conditional entropy of D given the condition of feature p. The conditional entropy H(Y|X) represents the uncertainty of random variable Y under the condition that random variable X is known, and its calculation formula is: Specifically, use feature A1 to represent "the error message of the failure case contains feature word a1". Feature A1 can divide data D into two subsets D1 and D2. D1 and D2 respectively represent the sample subsets in D whose error messages contain feature word a1 and do not contain feature word a1. Then H(D|A1) = Probability of containing feature word a1 * H(D1) + Probability of not containing feature word a1 * H(D2) g(D,A1) = H(D) - Probability of containing feature word a1 * H(D1) + Probability of not containing feature word a1 * H(D2) H(D), H(D1) and H(D2) are calculated according to the entropy calculation formula; Let feature A2 represent "the error message of the failure case contains feature word a2", and D1 and D2 respectively represent the sample subsets of D where the error message contains feature word a2 and does not contain feature word a2 g(D,A2) = H(D) - Probability of containing feature word a2 * H(D1) + Probability of not containing feature word a2 * H(D2).

2. The method according to claim 1, wherein, The step S200 further includes: Obtain the first failure problem classification criterion; Classify the first sample set based on the first failure problem classification criterion to obtain the first classification result; Perform failure cause analysis and marking on the first classification result to obtain the first preprocessing result.

3. The method according to claim 2, wherein, The method further includes: Locate defects according to the first classification result to obtain the first defect location set; Construct a mapping relationship based on the first defect location set and the first classification result to obtain the first mapping relationship construction result; After analyzing and locating the failure case problem based on the decision tree, locate the defect by analyzing the location result and the first mapping relationship construction result.

4. The method according to claim 1, wherein The method further includes: Construct a failure case solution library based on the case execution status; Obtain the analysis and location result of the decision tree; Provide a reference solution for the analysis and location result based on the failure case solution library.

5. The method according to claim 1, wherein The method further includes: Obtain the first predetermined decision tree generation condition; Judge whether the decision tree generated in step S600 meets the first predetermined decision tree generation condition; When the decision tree generated in step S600 meets the first predetermined decision tree generation condition, complete the construction of the decision tree.

6. A system for analyzing and locating problems in failure cases based on a decision tree, wherein The system includes: The first obtaining unit, which is used to obtain the first sample set; The second obtaining unit, which is used to perform sample preprocessing on the first sample set to obtain the first preprocessing result, where the first preprocessing result includes n sample subsets, where n is a positive integer greater than 1; The third obtaining unit, which is used to obtain the p-th sample subset of the first preprocessing result, perform feature extraction on the p-th sample subset to obtain the p-th feature list, where p is a positive integer less than or equal to n; The fourth obtaining unit, which is used to calculate the information gain of each feature in the p-th feature list for the p-th sample subset to obtain the p-th information gain set; The first selection unit, which is used to select the p-th feature with the largest information gain in the p-th information gain set as the p-th feature of the p-th sample subset, and delete the p-th feature from the p-th feature list; A first generation unit, which is configured to repeat the third acquisition unit to the first selection unit until corresponding features of all the n sample subsets are obtained, then stop repeating the steps, and generate a decision tree according to the corresponding features of the n sample subsets; A first analysis unit, which is configured to analyze and locate problems of failure cases based on the decision tree Wherein, it further includes: Obtain the first information entropy of the p-th sample subset; Obtain the second information entropy of the p-th sample subset after removing the p-th feature; Obtain the information gain of the p-th feature with respect to the p-th sample subset based on the first information entropy and the second information entropy; Calculate the information gain of the p-th feature with respect to the p-th sample subset through a formula, and the calculation formula is as follows: g(D,p) = H(D) - H(D|p) Wherein, g(D,p) is the information gain, H(D) is the first information entropy, and H(D|p) is the second information entropy; Calculate the information gain of each feature in the p-th feature list with respect to the p-th sample subset. Specifically, the set of unclassified failure cases is denoted as dataset D, and its information entropy is the first information entropy Y1 of the p-th sample subset; After removing the failure cases containing the p-th feature, the information entropy of the set drops to the second information entropy Y2 of the p-th sample subset. Obtain the information gain of the p-th feature with respect to the p-th sample subset based on the first information entropy and the second information entropy. Then the information gain of feature P with respect to dataset D is Y1 - Y2; Wherein, for the calculation of H(D), according to the formula of information entropy: H(D) = -(probability of system defect * log2 probability of system defect + probability of environmental problem * log2 probability of environmental problem + probability of data problem * log2 probability of data problem); H(D|p) is the empirical conditional entropy of D under the condition of feature p. The conditional entropy H(Y|X) represents the uncertainty of random variable Y under the condition of known random variable X, and its calculation formula is: Specifically, let feature A1 represent "the error message of the failure case contains feature word a1". Feature A1 divides data D into two subsets D1 and D2. D1 and D2 respectively represent the sample subsets in D whose error messages contain feature word a1 and do not contain feature word a1. Then H(D|A1) = probability of containing feature word a1 * H(D1) + probability of not containing feature word a1 * H(D2) g(D,A1) = H(D) - probability of containing feature word a1 * H(D1) + probability of not containing feature word a1 * H(D2) H(D), H(D1), and H(D2) are calculated according to the formula of entropy; Let feature A2 represent "the error message of the failure case contains feature word a2". D1 and D2 respectively represent the sample subsets in D whose error messages contain feature word a2 and do not contain feature word a2. g(D,A2) = H(D) - probability of containing feature word a2 * H(D1) + probability of not containing feature word a2 * H(D2).

7. An electronic device for failure case problem analysis and localization based on a decision tree, comprising a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected through the bus, and is characterized in that, When the computer program is executed by the processor, it implements the steps in the method described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps in the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Fault data processing method and device, computer equipment and storage medium

    CN111338836A

  • Method for automatically testing, positioning and repairing a fault and storage medium

    CN113656323A