Root cause mining method and device, electronic equipment, storage medium and program product

By constructing a multi-level dataset and calculating the generalized latent score, initial leaf nodes are screened and clustered, and false positive nodes are pruned. This solves the problem of difficulty in quickly and accurately locating root causes in existing technologies, and enables rapid and accurate root cause location in low-volume experiments, thereby improving the efficiency of product or strategy iteration.

CN116383277BActive Publication Date: 2026-03-03BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310370164.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-03-03
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

During product or strategy iteration, existing technologies struggle to quickly and accurately pinpoint the root cause of abnormal changes in experimental metrics, impacting the efficiency of product or strategy iteration.

Method used

By constructing a multi-level dataset, calculating generalized latent scores, filtering and clustering initial leaf nodes, pruning and filtering false positive nodes, and using difference values ​​and influence parameters to accurately locate root causes.

Benefits of technology

It enables the rapid and accurate identification of the root causes of abnormal changes in experimental indicators in low-volume experiments, improving the efficiency and accuracy of product or strategy iteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116383277B_ABST
    Figure CN116383277B_ABST
Patent Text Reader

Abstract

The disclosure provides a root cause mining method and device, electronic equipment, storage medium and program product, relates to the technical field of data processing, and particularly relates to the technical field of intelligent search and big data. The specific implementation scheme is: obtaining experimental group values and control group values corresponding to each dimension value combination under a preset dimension combination, taking each dimension value combination and the experimental group values and the control group values corresponding to the dimension value combination as an initial leaf node; for each category of initial leaf node, constructing a first level to a first preset number of levels of data set based on the initial leaf node of the category; for each category, according to the first level to the first preset number of levels of data set of the category, the generalized potential score of each dimension value combination of the first level to the first preset number of levels is calculated in turn, and the root cause causing the abnormal change of the experimental group value is mined from the dimension value combination satisfying the root cause condition. In this way, the root cause can be quickly and accurately located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and more particularly to the field of intelligent search and big data technology. Background Technology

[0002] In small-scale experiments during product or strategy iteration, when significant abnormal changes occur in experimental metrics, it's crucial to pinpoint the root cause of these changes. This allows for the identification of the underlying causes leading to these abnormal changes, and adjustments to the product or strategy can be made based on these findings. This approach enables the rapid development of iterative plans for products or strategies, driving product improvement through small-scale experiments. Summary of the Invention

[0003] This disclosure provides a root cause analysis method, apparatus, electronic device, storage medium, and program product, specifically including:

[0004] In a first aspect, embodiments of this disclosure provide a root cause analysis method, including:

[0005] Obtain the experimental group value and control group value corresponding to each dimension value combination under the preset dimension combination, and take each dimension value combination and the corresponding experimental group value and control group value as an initial leaf node;

[0006] For each category's initial leaf node, a dataset from the first level to the preset number of levels is constructed based on the initial leaf node of that category. The dataset at the Nth level includes potential leaf nodes corresponding to the combinations of each dimension value under the combination of N dimensions. The potential leaf node corresponding to a combination of dimension values ​​includes: the experimental group value and the control group value corresponding to the combination of the dimension value and each dimension value of the other single dimension respectively.

[0007] For each category, based on the dataset from the first level to the preset number of levels for that category, the generalized latent score of each dimension value combination from the first level to the preset number of levels is calculated sequentially. From the dimension value combinations where the generalized latent score satisfies the root cause condition, the root cause that leads to the abnormal change in the experimental group value is extracted.

[0008] Secondly, embodiments of this disclosure provide a root cause analysis device, the device comprising:

[0009] The acquisition module is used to acquire the experimental group value and control group value corresponding to each dimension value combination under the preset dimension combination, and to take each dimension value combination and the corresponding experimental group value and control group value as an initial leaf node.

[0010] The construction module is used to construct datasets from the first level to the preset number of levels based on the initial leaf nodes of each category. The dataset of the Nth level includes potential leaf nodes corresponding to the combinations of each dimension value under the combination of N dimensions. The potential leaf node corresponding to a combination of dimension values ​​includes: the experimental group value and the control group value corresponding to the combination of the dimension value and each dimension value of the other single dimension respectively.

[0011] The calculation module is used to calculate the generalized latent score of each dimension value combination from the first level to the preset number of levels of the dataset for each category, and to mine the root causes that cause abnormal changes in the experimental group values ​​from the dimension value combinations that satisfy the root cause condition from the generalized latent score.

[0012] Thirdly, embodiments of this disclosure provide an electronic device, including:

[0013] At least one processor; and

[0014] A memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0016] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described in the first aspect above.

[0017] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0019] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0020] Figure 1 This is a flowchart of a root cause analysis method provided in an embodiment of this disclosure;

[0021] Figure 2 This is a flowchart of another root cause discovery method provided in this embodiment of the disclosure;

[0022] Figure 3 This is a flowchart of yet another root cause analysis method provided in this disclosure embodiment;

[0023] Figure 4 This is an exemplary flowchart of a root cause analysis method provided in an embodiment of this disclosure;

[0024] Figure 5 This is an exemplary schematic diagram showing the comparison results between a root cause mining method of this disclosure and a traditional root cause mining algorithm provided in this embodiment;

[0025] Figure 6 This is an exemplary schematic diagram showing the comparison results between another root cause mining method provided in this disclosure and a traditional root cause mining algorithm.

[0026] Figure 7 This is a schematic diagram of the structure of a root cause discovery device provided in an embodiment of this disclosure;

[0027] Figure 8 This is a block diagram of an electronic device used to implement the root cause analysis method of the present disclosure. Detailed Implementation

[0028] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0029] This disclosure can be applied to low-volume experiments during product or strategy iteration processes, in scenarios where experimental metrics exhibit significant abnormal changes. Taking product iteration as an example, experiments can be conducted on the iterated product to obtain experimental and control group values ​​generated during the experiment. Then, a significance test is performed on each experimental group value. If there are significant abnormal changes in the experimental group values, root cause analysis is required to locate the root cause of the abnormal changes.

[0030] For example, in a small-scale experiment on a search engine, if a significance test reveals a significant abnormal change in the experimental group's page views (PV) for repeated searches, then it is necessary to investigate the root cause of this significant abnormal change in PV for repeated searches.

[0031] To accurately pinpoint the root cause of abnormal changes, embodiments of this disclosure provide a root cause analysis method, which can be applied to electronic devices, such as... Figure 1As shown, the method includes:

[0032] S101. Obtain the experimental group value and control group value corresponding to each dimension value combination under the preset dimension combination, and take each dimension value combination and the corresponding experimental group value and control group value as an initial leaf node.

[0033] The preset dimension combinations include dimensions that are suspected of causing significant abnormal changes in the experimental group values.

[0034] After identifying the experimental group values ​​that show significant abnormal changes, the dimensions that may cause significant abnormal changes in the experimental group values ​​can be manually compiled. Then, the electronic device will compile the enumerated values ​​of all the compiled dimensions into a cross-tabulation to obtain the combination of the most granular dimension values. Each combination of dimension values ​​includes one dimension value from each of the compiled dimensions.

[0035] For a given dimension, the enumerated value of that dimension refers to all the enumerated values ​​of that dimension.

[0036] For example, the dimension values ​​for the operating system dimension can include Android and iOS systems, and the dimension values ​​for the network type dimension can include WiFi networks, 4G networks, 5G networks, etc.

[0037] The electronic device can then obtain the experimental group values ​​and control group values ​​corresponding to each combination of dimensional values, resulting in a two-dimensional table, which is used as the basic data.

[0038] Each row in this two-dimensional table includes a combination of dimension values ​​and the corresponding experimental and control group values. Accordingly, each row in this two-dimensional table, excluding the header, can be used as an initial leaf node. As an example, this two-dimensional table is shown in Table 1:

[0039] Table 1

[0040]

[0041] Table 1 illustrates an example of a preset dimension combination, which includes six dimensions: search type, network type, browser, search category, login status, and operating system. Starting from the second row of Table 1, the first six columns of each row represent a combination of dimension values, and the last two columns are the experimental and control group values ​​corresponding to that dimension value combination. Table 1 provides an example of the experimental and control group values ​​for the repeated search PV for each dimension value combination.

[0042] It should be noted that Table 1 is only an example for easy understanding and does not fully show all combinations of dimension values. The amount of basic data in the actual implementation is not limited to this.

[0043] S102. For each category's initial leaf node, construct a dataset from the first level to the preset number of levels based on the initial leaf node of that category.

[0044] The dataset at level N includes potential leaf nodes corresponding to the combinations of values ​​of each dimension under the combination of N dimensions. The potential leaf node corresponding to a combination of dimension values ​​includes the experimental group value and the control group value corresponding to the combination of the dimension value with each dimension value of the other single dimension.

[0045] The preset quantity is the number of dimensions included in the preset dimension combination minus 1, and the value of N ranges from 1 to the preset quantity.

[0046] As an example, if the preset dimension combination includes 3 dimensions, then a dataset for the first level to the second level needs to be constructed.

[0047] For the first level, the dataset includes potential leaf nodes corresponding to each dimension value under one dimension. This means that potential leaf nodes need to be constructed separately for each dimension value. When constructing a potential leaf node corresponding to a dimension value, this dimension value can be combined with each dimension value of the other single dimensions, and the experimental group value and control group value for each newly obtained combination of dimension values ​​can be calculated.

[0048] For example, if the three dimensions are search type, whether logged in, and operating system, the dimension values ​​for the search type dimension include sug, se, and inp; the dimension values ​​for the whether logged in dimension include 1 and 0; and the dimension values ​​for the operating system dimension include android and iOS.

[0049] For example, for the dimension value "sug" in the search type dimension, combining "sug" with the values ​​of each dimension in the login / non-login dimension yields the following combinations: "sug+1" and "sug+0". Combining "sug" with the values ​​of each dimension in the operating system dimension yields the following combinations: "sug+android" and "sug+iOS". Therefore, the potential leaf nodes corresponding to the dimension value "sug" in the search type dimension include: "sug+1" and its corresponding experimental and control group values; "sug+0" and its corresponding experimental and control group values; "sug+android" and its corresponding experimental and control group values; and "sug+iOS" and its corresponding experimental and control group values.

[0050] For the dimension value 'se' in the search type dimension, 'se' is combined with the values ​​of each dimension in the login status dimension, resulting in the following combination: 'se+0' and 'se+1'. Similarly, 'se' is combined with the values ​​of each dimension in the operating system dimension, resulting in the following combination: 'se+android' and 'se+iOS'. Therefore, the potential leaf nodes corresponding to the dimension value 'se' in the search type dimension include: 'se+0' and its corresponding experimental and control group values; 'se+1' and its corresponding experimental and control group values; 'se+android' and its corresponding experimental and control group values; and 'se+iOS' and its corresponding experimental and control group values.

[0051] Similarly, for the inp value in the retrieval type dimension and for the dimension values ​​in other dimensions, multiple potential leaf nodes corresponding to a dimension value can be obtained in the same way as described above, which will not be elaborated here.

[0052] For the second level, the dataset includes potential leaf nodes corresponding to combinations of values ​​from each of the two dimensions. That is, it is necessary to construct potential leaf nodes for each combination of values ​​from each of the two dimensions separately. When constructing a potential leaf node corresponding to a combination of dimensional values, this combination can be combined with each value from each of the other individual dimensions, and the experimental and control group values ​​for each newly obtained combination of dimensional values ​​can be calculated.

[0053] For example, combining the dimension value "sug" from the search type dimension and the dimension value "1" from the login status dimension yields the dimension value combination "sug+1". This "sug+1" can then be combined with the values ​​of each dimension from the operating system dimension, resulting in dimension value combinations such as "sug+1+android" and "sug+1+iOS". Therefore, the potential leaf nodes corresponding to the dimension value combination "sug+1" include: "sug+1+android" along with its corresponding experimental and control group values, and "sug+1+iOS" along with its corresponding experimental and control group values.

[0054] The method for constructing potential leaf nodes for other combinations of dimensional values ​​is the same, and will not be listed here.

[0055] S103. For each category, based on the dataset from the first level to the preset number of levels of that category, sequentially traverse and calculate the generalized latent score of each dimension value combination from the first level to the preset number of levels. From the dimension value combinations where the generalized latent score satisfies the root cause condition, mine the root cause that leads to the abnormal change in the experimental group value.

[0056] For each combination of dimension values, the generalized latent score of that combination can be calculated from all the latent leaf nodes corresponding to that combination of dimension values.

[0057] Based on the example in S102 above, for the first level, the generalized latent scores of the dimension values ​​sug, se, inp, 1, 0, android and iOS need to be calculated respectively.

[0058] Taking the dimension value sug as an example, the generalized latent score of the dimension value sug can be calculated based on all the potential leaf nodes corresponding to the dimension value sug listed above.

[0059] Using the above method, the experimental group values ​​and control group values ​​corresponding to each dimension value combination under the preset dimension combinations are obtained, thus obtaining multiple initial leaf nodes. Then, the electronic device constructs initial leaf nodes based on each category, constructing datasets from the first level to the preset number of levels. The Nth level dataset includes potential leaf nodes corresponding to each dimension value combination under N dimension combinations. Since a potential leaf node corresponding to a dimension value combination includes the experimental group values ​​and control group values ​​corresponding to each dimension value combination of that dimension value combination and each dimension value of other individual dimensions, it is equivalent to constructing potential leaf nodes using one dimension value combination plus one dimension, avoiding the problem of excessively sparse experimental group and control group values ​​corresponding to potential leaf nodes. Subsequently, using the datasets from the first level to the preset number of levels of that category, the generalized latent scores of each dimension value combination obtained through traversal calculation are more stable and accurate. Therefore, from the dimension value combinations whose generalized latent scores satisfy the root cause condition, the root cause causing abnormal changes in experimental group values ​​can be more accurately located.

[0060] In some embodiments of this disclosure, prior to S102 described above, the initial leaf nodes need to be screened and clustered, such as... Figure 2 As shown, the method includes S201-S206.

[0061] S201 is the same as S101, and S205-S206 are the same as S102-S103 above. Please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0062] S202. Calculate the difference value for each initial leaf node, which represents the difference between the experimental group value and the control group value of the initial leaf node.

[0063] Specifically, based on the experimental group values ​​and control group values ​​included in each initial leaf node, the difference value for each initial leaf node is calculated. The difference value can be represented by the offset score. ds1) represents the initial leaf node's ds1 calculation formula as follows:

[0064]

[0065] in, This is the experimental group value for the initial leaf node. This is the control group value for the initial leaf node.

[0066] For example, if the experimental group value in the initial leaf node 1 is 1111 and the control group value is 2323, then according to the above formula, the ds1 of the initial leaf node 1 is calculated to be 2. (1111-2323) / (1111+2323) is approximately equal to -0.7.

[0067] S203. If the experimental group value shows an abnormal upward trend, delete the initial leaf node whose difference value is less than the mean difference value; or, if the experimental group value shows an abnormal downward trend, delete the initial leaf node whose difference value is greater than the mean difference value.

[0068] The mean difference value is obtained by summing the difference values ​​of all initial leaf nodes and then taking the average.

[0069] Understandably, if the experimental group values ​​show an abnormal upward trend, initial leaf nodes with a difference value less than the mean difference value will not cause the abnormal increase in experimental group values. The possibility of finding the root cause from these initial leaf nodes is small. Therefore, initial leaf nodes with a difference value less than the mean difference value are deleted. Similarly, if the experimental group values ​​show an abnormal downward trend, initial leaf nodes with a difference value greater than the mean difference value will not cause the abnormal decrease in experimental group values. The possibility of finding the root cause from these initial leaf nodes is small. Therefore, initial leaf nodes with a difference value greater than the mean difference value are deleted.

[0070] For example, if the experimental group value is the failure rate and the experimental group value shows an abnormal upward trend, the initial leaf node with a difference value greater than the mean difference value is the initial leaf node with a high failure rate. Therefore, it is easier to locate the root cause in the initial leaf node with a high failure rate. So, the initial leaf node with a difference value less than the mean is excluded.

[0071] For example, if the experimental group value is the pass rate and the experimental group value shows an abnormal downward trend, the initial leaf node with a difference value less than the mean difference value is the initial leaf node with a low pass rate. Therefore, it is easier to locate the root cause in the initial leaf node with a low pass rate. So, the initial leaf node with a difference value greater than the mean difference value is excluded.

[0072] S204. Based on the difference values ​​of the remaining initial leaf nodes, cluster each initial leaf node.

[0073] Specifically, one-dimensional k-means clustering can be used to perform clustering calculations based on the difference values ​​of the remaining initial leaf nodes, thereby dividing the remaining initial leaf nodes into multiple categories, each category including multiple initial leaf nodes.

[0074] By employing the above method, the difference values ​​of each initial leaf node are used to filter out initial leaf nodes that may contain the root cause, narrowing the screening scope and reducing the number of initial leaf nodes that need to be processed subsequently, thus reducing computational load. Based on the difference values ​​of the remaining initial leaf nodes, they are clustered to locate the root cause within each category. This facilitates rapid and accurate root cause location.

[0075] In some embodiments of this disclosure, such as Figure 3 As shown, in S103 above, based on the dataset from the first level to the preset number of levels of the category, the generalized latent score of each dimension value combination from the first level to the preset number of levels is calculated sequentially. From the dimension value combinations whose generalized latent scores satisfy the root cause condition, the root causes causing abnormal changes in the experimental group values ​​are mined. Specifically, this may include the following steps:

[0076] S1031. Starting from the first level, for each combination of dimension values ​​included in each level of the category, calculate the generalized latent score of the combination of dimension values ​​based on the latent leaf node corresponding to the combination of dimension values ​​in the dataset of that level.

[0077] The generalized latent score is used to characterize the probability that the combination of dimension values ​​corresponding to the generalized latent score is the root cause. The specific method for calculating the generalized latent score of the combination of dimension values ​​corresponding to the dataset will be described in detail in subsequent embodiments.

[0078] S1032. Whenever the generalized latent score of the calculated combination of dimension values ​​is greater than the first preset threshold, the combination of dimension values ​​is added to the candidate root cause set, and the potential leaf nodes under the combination of dimension values ​​are pruned.

[0079] Understandably, if the generalized latent score of a combination of dimension values ​​is greater than the first preset threshold, it indicates that the combination of dimension values ​​may be a root cause, and therefore, the combination of dimension values ​​is added to the candidate root cause set. To avoid the potential leaf nodes under this dimension value combination being traversed again when calculating other dimension value combinations later, the potential leaf nodes under this dimension value combination can be pruned. This can optimize performance. As an example, the first preset threshold can be 0.85.

[0080] Among them, the potential leaf nodes under a combination of dimension values ​​refer to the potential leaf nodes that include that combination of dimension values ​​and are more granular. For example, if a combination of dimension values ​​is: retrieval type sug + network type WiFi, a potential leaf node is: retrieval type sug + network type WiFi + browser oppo, as well as the corresponding experimental group value and control group value.

[0081] S1033. After completing the traversal of the category, extract the root causes that cause abnormal changes in the experimental group values ​​from the candidate root cause set of the category.

[0082] Using the above method, the generalized latent score is used to measure the probability that the corresponding dimension value combination is the root cause. Therefore, when the generalized latent score of a dimension value combination is greater than a first preset threshold, it indicates that the dimension value combination is more likely to be the root cause. Consequently, the potential leaf nodes under this dimension value combination are pruned, avoiding unnecessary calculations and improving the efficiency of root cause discovery. Furthermore, other dimension value combinations unrelated to this one will continue to be traversed and calculated at a deeper level, accurately discovering other dimension value combinations that may be root causes. This avoids the problem of finer-grained dimension value combinations in subsequent levels not being traversed, thus enabling efficient and accurate root cause discovery.

[0083] In this embodiment of the disclosure, S1033 can be specifically implemented as steps A to C.

[0084] Step A: For each combination of dimension values ​​included in the candidate root cause set of this category, determine the proportion of leaf nodes of the same type for that dimension value combination. The proportion of leaf nodes of the same type is the ratio between the number of potential leaf nodes corresponding to that dimension value combination in this category and the total number of potential leaf nodes corresponding to that dimension value combination in all categories.

[0085] Specifically, for each combination of dimension values, the proportion of leaf nodes of the same type ( The formula for calculating ) is:

[0086]

[0087] in, This represents the number of potential leaf nodes corresponding to this combination of dimension values ​​within this category. This represents the total number of potential leaf nodes corresponding to the combination of values ​​for this dimension across all categories.

[0088] For example, if the dimension value combination is retrieval type 1 and network type WiFi, and this dimension value combination belongs to category A, and the number of potential leaf nodes corresponding to this dimension value combination in category A is 70, while the number of potential leaf nodes corresponding to this dimension value combination across all categories is 100, then this dimension value combination... It is 0.7.

[0089] Step B: If the proportion of leaf nodes of the same type is less than the second preset threshold, then the combination of dimension values ​​is deleted from the candidate root cause set.

[0090] Understandably, if the proportion of leaf nodes of the same type is less than the second preset threshold, it indicates that only a small portion of the potential leaf nodes of that dimension value combination belong to that category. This may be due to a small sample size, leading to inaccurate generalized potential scores for that dimension value combination, which mistakenly adds it to the candidate root cause set. Therefore, the probability that this dimension value combination is a root cause of that category is low, and it can be removed from the candidate root cause set. As an example, the second preset threshold can be 0.5.

[0091] Step C: From the remaining combination of dimension values ​​in the candidate root cause set for this category, select the combination of dimension values ​​that has the greatest impact on the overall experimental group value, and use it as the root cause for this category.

[0092] The impact of dimensional value combinations on the overall experimental group value can be characterized by an influence parameter. Based on the influence parameter, the root cause of the category can be determined. Specifically, this involves calculating the influence parameter of each dimensional value combination from the remaining dimensional value combinations in the candidate root cause set of the category, and taking the dimensional value combination with the largest influence parameter as the root cause of the category.

[0093] The influence parameter is the difference between the first and second magnitudes of change. The first magnitude of change is the difference between the overall experimental group value and the overall control group value. The second magnitude of change is the difference between the overall experimental group value and the overall control group value of the remaining dimension value combinations after removing that dimension value combination. Specifically, the influence parameter can be the value of each dimension value combination. , The calculation method will be described in subsequent embodiments.

[0094] Using the above method, for each dimension value combination included in each level of the category, if the proportion of similar leaf nodes is less than a preset threshold, then the dimension value combination is considered a false positive and is deleted. This avoids inaccurate calculation results of the generalized latent score of dimension value combinations due to a small sample size, thus preventing the erroneous addition of false positive dimension value combinations to the candidate root cause set. This further narrows the screening scope of root causes, facilitating rapid root cause location. Simultaneously, deleting false positive dimension value combinations from the candidate root cause set improves the accuracy of root cause discovery. Furthermore, for the remaining dimension value combinations in the candidate root cause set of this category, the influence parameter can be used to characterize the impact of the dimension value combination corresponding to the influence parameter on the overall experimental group value. Therefore, among the remaining dimension value combinations, the dimension value combination with the largest influence parameter is most likely to be the root cause of this category. Thus, the dimension value combination with the largest influence parameter is taken as the root cause of this category. In this way, the root cause can be located quickly and accurately.

[0095] In some embodiments of this disclosure, the generalized latent score of the dimension value combination is calculated based on the latent leaf node corresponding to the dimension value combination in the dataset at that level, including:

[0096] When the experimental group values ​​are proportional indicators, the generalized potential score (gps) for this combination of dimension values ​​is calculated using the following formula:

[0097]

[0098] As an example, as shown in Table 2, the experimental group value and control group value corresponding to each row of dimension value combination in Table 2 are both proportional indicators. Therefore, the GPS of each dimension value combination included in Table 2 can be calculated based on the above GPS calculation formula and Table 2.

[0099] Table 2

[0100]

[0101] Table 2 and Table 1 show the same preset combination of dimensions. The difference is that the experimental group value and control group value corresponding to each combination of dimension values ​​in Table 1 are proportional indicators.

[0102] It should be noted that Table 2 is only an example for easy understanding and does not fully show all combinations of dimension values. The amount of basic data in the actual implementation is not limited to this.

[0103] When the experimental group values ​​are non-proportional indicators, the generalized latent score for this dimension value combination is calculated using the following formula:

[0104]

[0105] in, This refers to the experimental group values ​​of the potential leaf nodes corresponding to this combination of dimension values ​​in the dataset at this level. = , This refers to the control group values ​​for the potential leaf nodes corresponding to this combination of dimension values ​​in the dataset at this level. This represents the sum of the experimental group values ​​for each potential leaf node corresponding to the combination of values ​​for this dimension in the dataset at this level. This is the sum of the control group values ​​for each potential leaf node corresponding to this dimension value combination in the dataset at this level. This refers to taking the weighted average of the values ​​calculated for each potential leaf node within the parentheses.

[0106] When the experimental group value is a proportional indicator... = , The overall experimental set value is the sum of the experimental set values ​​of all potential leaf nodes in the dataset at this level. This is the overall control group value for all potential leaf nodes in the dataset at this level. The overall control group value is the sum of the control group values ​​for all potential leaf nodes in the dataset at this level. This is the overall experimental set value of the potential leaf nodes in the dataset at this level, excluding the combination of dimension values. The overall experimental set value is the sum of the experimental set values ​​of all potential leaf nodes in the dataset at this level, excluding the combination of dimension values. This is the overall control group value for all potential leaf nodes in the dataset at this level, excluding the combination of values ​​for this dimension. The overall control group value is the sum of the control group values ​​for all potential leaf nodes in the dataset at this level, excluding the combination of values ​​for this dimension.

[0107] When the experimental group values ​​are non-proportional indicators... .

[0108] Using the above method, this embodiment of the disclosure improves the GPS calculation formula. On the one hand, it... The arithmetic mean of each node has been changed to a weighted average of all potential leaf nodes, avoiding the low accuracy of GPS calculation results caused by excessively sparse experimental and control group values ​​of potential leaf nodes. On the other hand, when the experimental group values ​​are non-proportional indicators, this embodiment changes the absolute difference in the proportional indicator GPS calculation formula to a relative difference. This unifies the dimensions of the experimental group values ​​for non-proportional indicators, improving the stability of the GPS calculation results. Furthermore, for both GPS calculation formulas, the sum of the differences between the experimental and control group values ​​for all dimension value combinations other than the specified dimension value combination has been removed from the numerator. This reduces the value of the numerator, making GPS more sensitive to multiple root causes. Consequently, root causes can be located more accurately.

[0109] The following combination Figure 4 The complete process of the root cause discovery method provided in the embodiments of this disclosure is described below, such as... Figure 4 As shown, the method includes:

[0110] S401. Based on a preset combination of dimensions, acquire data and generate a two-dimensional table.

[0111] The preset dimension combination consists of dimensions corresponding to experimental indicators that show significant abnormal changes, which are manually compiled. The preset dimension combination includes multiple dimension values. The electronic device combines all dimension values ​​under the preset dimension combination into a cross-tabulation. Each row of the cross-tabulation is a combination of the most granular dimension values.

[0112] Obtain the experimental group values ​​and control group values ​​corresponding to each combination of dimension values, and generate a two-dimensional table.

[0113] S402. Take each row of the two-dimensional table as an initial leaf node, and calculate the difference value for each initial leaf node.

[0114] Specifically, the formula for calculating ds1 is described in the above embodiments.

[0115] Taking the initial leaf node corresponding to the last row in Table 1 as an example, assuming that the experimental group value with significant abnormal changes is the repeated search PV, then the ds1 of this initial leaf node is 2. (2-23) / (2+23)=-1.68, so the ds1 of the initial leaf node is -1.68.

[0116] S403. If the experimental group values ​​show an abnormal upward trend, delete the initial leaf nodes whose difference values ​​are less than the mean difference values.

[0117] S404. If the experimental group values ​​show an abnormal downward trend, delete the initial leaf nodes whose difference values ​​are greater than the average difference values.

[0118] S405. The remaining initial leaf nodes fall into the candidate dataset.

[0119] S406. Cluster the initial leaf nodes in the candidate dataset.

[0120] S407. For each category, construct a dataset from the first level to the preset number of levels for the initial leaf nodes in that category.

[0121] The method for constructing the dataset has been described in the above embodiments and will not be repeated here.

[0122] S408. Based on the dataset from the first level to the preset number of levels, sequentially traverse and calculate the GPS of each dimension value combination from the first level to the preset number of levels.

[0123] Once a GPS value is calculated, S410 is executed.

[0124] S409. Determine if the GPS value is greater than 0.85.

[0125] If yes, execute S410; otherwise, execute S408.

[0126] S410. Prune the potential leaf nodes corresponding to the dimension value combinations with GPS greater than 0.85, and continue to traverse other dimension value combinations.

[0127] S411. After traversal, store the dimension value combinations with GPS greater than 0.85 in the candidate root cause set, and filter out the dimension value combinations in the candidate root cause set.

[0128] Taking Table 1 as an example, the dimension value combination search type "sug+browser oppo" is a dimension value combination in the candidate root cause set. The number of potential leaf nodes of the dimension value combination search type "sug+browser oppo" in category B of the candidate root cause set is 70. The total number of potential leaf nodes in the dataset corresponding to the dimension value combination search type "sug+browser oppo" is 100. Therefore, the proportion of leaf nodes of the same type corresponding to the dimension value combination search type "sug+browser oppo" is 0.7. This proportion of leaf nodes of the same type indicates that the dimension value combination search type "sug+browser oppo" is likely to be a root cause of category B. Therefore, the dimension value combination search type "sug+browser oppo" can be retained in the candidate root cause set.

[0129] Conversely, the dimension value combination search type sug+network type 4 and the aforementioned dimension value combination search type sug+browser oppo belong to the same candidate root cause set. The number of potential leaf nodes of dimension value combination search type sug+network type 4 in category B corresponding to this candidate root cause set is 20, and the total number of potential leaf nodes corresponding to dimension value combination search type sug+network type 4 is 100. Therefore, the proportion of leaf nodes of the same type corresponding to dimension value combination search type sug+network type 4 is 0.2. This proportion of leaf nodes of the same type indicates that the probability of dimension value combination search type sug+network type 4 being a root cause of category B is low. Therefore, dimension value combination search type sug+network type 4 can be deleted from this candidate root cause set.

[0130] S412. Based on the influence parameters corresponding to the combination of dimension values ​​in the candidate root cause set, sort the combination of dimension values ​​in the candidate root cause set in reverse order, and take the first one as the root cause of that type.

[0131] Specifically, the combination of dimension values ​​with the largest influence parameter is determined from the GPS corresponding to each combination of dimension values, and the combination of dimension values ​​corresponding to the largest influence parameter is output as the root cause.

[0132] The influence parameter is the same as the parameter in the two GPS calculation formulas mentioned above. .

[0133] S413. Have all classes been traversed?

[0134] If yes, execute S414; otherwise, execute S407.

[0135] S414. Output the root cause of all classes.

[0136] Using the above method, initial leaf nodes are selected based on differences, reducing the number of initial leaf nodes to be processed and thus reducing computational load. For each category of initial leaf nodes, a dataset from the first level to a predetermined number of levels is constructed. The GPS of each dimension value combination is calculated by iterating through the potential leaf nodes corresponding to each dimension value combination. This improves the stability of the GPS. Furthermore, the above GPS calculation method is applicable to finer-grained dimension value combinations, thus it is also suitable for cases of dimension explosion. After iteration, dimension value combinations with GPS greater than 0.85 are stored in a candidate root cause set. False positive dimension value combinations in the candidate root cause set are filtered out using the proportion of similar leaf nodes, further narrowing the selection range of root causes. Finally, the largest dimension value combination is determined from the remaining dimension value combinations in the candidate root cause set. The corresponding combination of dimension values ​​is then output as the root cause. In this way, the root cause can be accurately located.

[0137] Based on the above embodiments, in a low-traffic experimental scenario, the GPS results output using the GPS calculation formula provided in this disclosure embodiment are compared with the GPS results output by the traditional squeeze method, where the squeeze method is a root cause analysis algorithm. Figure 5 As shown, Figure 5 The x-axis represents the relative difference, and the y-axis represents GPS. The relative difference characterizes the impact of a specified root cause on the overall change in the experimental group values. The relative difference is calculated as: (Overall experimental group values ​​under the specified dataset - Overall control group values ​​under the specified dataset) / Overall control group values ​​under the specified dataset. The specified dataset is the dataset used in the low-flow experiment.

[0138] Figure 5 In the experimental group, the values ​​are non-proportional indicators. Using the GPS calculation formula provided in this embodiment, 100 GPS results with relative differences between 0% and 5% are calculated. The above 100 GPS results are connected into a curve, which is the GPS result curve obtained by using the GPS calculation formula provided in this embodiment. Similarly, the GPS result curve corresponding to the traditional squeeze method can be obtained.

[0139] Figure 5 The curve above is the GPS result curve obtained using the GPS calculation formula provided in the embodiments of this disclosure. Figure 5 The curve below shows the GPS result obtained using the traditional squeeze method. It can be seen that, when the experimental group value is a non-proportional indicator, after the relative difference exceeds 1.14%, the GPS result obtained using the GPS calculation formula provided in this embodiment can stably reach above 0.85. Figure 5The GPS result curve corresponding to the traditional squeeze method has never reached 0.85.

[0140] Similarly, such as Figure 6 As shown, Figure 6 The x-axis represents the relative difference, and the y-axis represents GPS. Figure 6 In the figure, the experimental group value is a proportional indicator. The GPS result curve obtained using the GPS calculation formula provided in this embodiment is as follows: Figure 6 The curve in the upper middle section is the GPS result curve calculated using the traditional squeeze method. Figure 6 The lower curve, after the relative difference is greater than 0.96%, the GPS result obtained by using the GPS calculation formula provided in this embodiment can stably reach above 0.85, while the GPS result curve corresponding to the traditional squeeze method always fluctuates around 0.

[0141] Therefore, the GPS calculation formula provided in this embodiment is more sensitive, and the root cause mining method provided in this embodiment can effectively mine root causes.

[0142] It should be noted that, Figure 5 and Figure 6 This is merely an example for verification; in practical applications, the GPS calculated in the root cause mining method provided in this disclosure is not limited to this.

[0143] Based on the same concept, embodiments of this disclosure provide a root cause analysis device, such as... Figure 7 As shown, the device includes:

[0144] The acquisition module 701 is used to acquire the experimental group value and control group value corresponding to each dimension value combination under the preset dimension combination, and to take each dimension value combination and the corresponding experimental group value and control group value as an initial leaf node.

[0145] The construction module 702 is used to construct datasets from the first level to the preset number of levels based on the initial leaf nodes of each category. The dataset of the Nth level includes potential leaf nodes corresponding to the combinations of dimension values ​​under the combination of N dimensions. The potential leaf node corresponding to a combination of dimension values ​​includes: the experimental group value and the control group value corresponding to the combination of dimension values ​​and each dimension value of the other single dimension respectively.

[0146] The calculation module 703 is used to calculate the generalized latent score of each dimension value combination from the first level to the preset number of levels of the dataset for each category, and to mine the root cause that causes the abnormal change of the experimental group value from the dimension value combination that satisfies the root cause condition from the generalized latent score.

[0147] Optionally, the device also includes a deletion module and a clustering module:

[0148] The calculation module 703 is also used to calculate the difference value of each initial leaf node, which represents the difference between the experimental group value and the control group value of the initial leaf node.

[0149] The deletion module is used to delete the initial leaf nodes whose difference value is less than the mean difference value if the experimental group value shows an abnormal upward trend; or to delete the initial leaf nodes whose difference value is greater than the mean difference value if the experimental group value shows an abnormal downward trend.

[0150] The clustering module is used to cluster the initial leaf nodes based on the difference values ​​of the remaining initial leaf nodes.

[0151] Optional, the calculation module 703 is specifically used for:

[0152] Starting from the first level, for each combination of dimension values ​​included in each level of the category, the generalized latent score of the combination of dimension values ​​is calculated based on the latent leaf nodes corresponding to the combination of dimension values ​​in the dataset of that level.

[0153] Whenever the generalized latent score of the calculated combination of dimension values ​​is greater than the first preset threshold, the combination of dimension values ​​is added to the candidate root cause set, and the potential leaf nodes under the combination of dimension values ​​are pruned.

[0154] After completing the traversal of the category, the root causes that lead to the abnormal changes in the experimental group values ​​are extracted from the candidate root cause set of the category.

[0155] Optional, the calculation module 703 is specifically used for:

[0156] For each combination of dimension values ​​included in the candidate root cause set of this category, determine the proportion of leaf nodes of the same type for that dimension value combination. The proportion of leaf nodes of the same type is the ratio between the number of potential leaf nodes corresponding to that dimension value combination in this category and the total number of potential leaf nodes corresponding to that dimension value combination in all categories.

[0157] If the proportion of leaf nodes of the same type is less than the second preset threshold, then the combination of that dimension value will be deleted from the candidate root cause set.

[0158] From the remaining combination of dimension values ​​in the candidate root cause set for that category, select the combination of dimension values ​​that has the greatest impact on the overall experimental group value, and use it as the root cause for that category.

[0159] Optional, the calculation module 703 is specifically used for:

[0160] From the remaining combination of dimension values ​​in the candidate root cause set of this category, calculate the influence parameter of each combination of dimension values. The influence parameter is the difference between the first change magnitude and the second change magnitude. The first change magnitude is the difference between the overall experimental group value and the overall control group value. The second change magnitude is the difference between the overall experimental group value and the overall control group value of the remaining combination of dimension values ​​after removing this combination of dimension values.

[0161] The combination of dimension values ​​with the highest influence parameter is taken as the root cause of the category.

[0162] Optional, the calculation module 703 is specifically used for:

[0163] When the experimental group values ​​are proportional indicators, the generalized latent score for this combination of dimension values ​​is calculated using the following formula:

[0164]

[0165] Alternatively, if the experimental group values ​​are non-proportional indicators, the generalized latent score for this combination of dimension values ​​can be calculated using the following formula:

[0166]

[0167] in, This refers to the experimental group values ​​of the potential leaf nodes corresponding to this combination of dimension values ​​in the dataset at this level. = , This refers to the control group values ​​for the potential leaf nodes corresponding to this combination of dimension values ​​in the dataset at this level. This represents the sum of the experimental group values ​​for each potential leaf node corresponding to the combination of values ​​for this dimension in the dataset at this level. This is the sum of the control group values ​​for each potential leaf node corresponding to this dimension value combination in the dataset at this level. This refers to taking the weighted average of the values ​​calculated for each potential leaf node within the parentheses.

[0168] When the experimental group value is a proportional indicator... = , This represents the overall experimental group value for all potential leaf nodes in the dataset at this level. This represents the overall control group value for all potential leaf nodes in the dataset at this level. This represents the overall experimental group values ​​of potential leaf nodes in the dataset at this level, excluding the combination of values ​​for this dimension. This represents the overall control group value for potential leaf nodes in the dataset at this level, excluding the combination of values ​​for this dimension.

[0169] When the experimental group values ​​are non-proportional indicators... .

[0170] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0171] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0172] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0173] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0174] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as root cause mining methods. For example, in some embodiments, the root cause mining method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the root cause mining method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the root cause mining method by any other suitable means (e.g., by means of firmware).

[0175] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0176] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0179] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0180] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0181] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0182] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A root cause mining method, comprising: obtaining experimental group values and control group values corresponding to each dimension value combination under a preset dimension combination, taking each dimension value combination and the experimental group values and control group values corresponding to the dimension value combination as an initial leaf node; for each category of initial leaf node, constructing a first layer to a preset number of layers of data sets based on the initial leaf node of the category, wherein the Nth layer of data sets comprises each dimension value combination under a combination of N dimensions, and the potential leaf node corresponding to each dimension value combination comprises experimental group values and control group values corresponding to each dimension value combination of a single dimension combined with the dimension value combination; the preset number is the number of dimensions included in the preset dimension combination minus 1, and the value range of N is 1 to the preset number; for each category, according to the first layer to the preset number of layers of data sets of the category, iteratively calculating the generalized potential score of each dimension value combination of the first layer to the preset number of layers, and mining the root cause causing the abnormal change of the experimental group values from the dimension value combination whose generalized potential score meets the root cause condition. 2.The method of claim 1, before the for each category of initial leaf node, constructing a first layer to a preset number of layers of data sets based on the initial leaf node of the category, the method further comprises: calculating the difference value of each initial leaf node, the difference value being used to represent the difference between the experimental group values and the control group values of the initial leaf node; if the experimental group values show an abnormal upward trend, deleting the initial leaf node whose difference value is less than the average difference value; or if the experimental group values show an abnormal downward trend, deleting the initial leaf node whose difference value is greater than the average difference value; clustering the remaining initial leaf nodes based on the difference values of the initial leaf nodes.

3. The method of claim 1, wherein, The according to the first layer to the preset number of layers of data sets of the category, iteratively calculating the generalized potential score of each dimension value combination of the first layer to the preset number of layers, and mining the root cause causing the abnormal change of the experimental group values from the dimension value combination whose generalized potential score meets the root cause condition, comprises: starting from the first layer, for each dimension value combination included in each layer of the category, calculating the generalized potential score of the dimension value combination according to the potential leaf node corresponding to the dimension value combination in the data set of the layer; whenever the calculated generalized potential score of the dimension value combination is greater than a first preset threshold, adding the dimension value combination to a candidate root cause set, and pruning the potential leaf nodes subordinate to the dimension value combination; after completing the traversal of the category, mining the root cause causing the abnormal change of the experimental group values from the candidate root cause set of the category.

4. The method of claim 3, wherein, The mining the root cause causing the abnormal change of the experimental group values from the candidate root cause set of the category, comprises: for each dimension value combination included in the candidate root cause set of the category, determining the proportion of the same type leaf node of the dimension value combination, the proportion of the same type leaf node being the ratio between the number of potential leaf nodes corresponding to the dimension value combination in the category and the total amount of potential leaf nodes corresponding to the dimension value combination in all categories. if the proportion of the same type of leaf nodes is less than a second preset threshold, the dimension value combination is deleted from the set of candidate root causes; from the dimension value combinations remaining in the set of candidate root causes of the category, a dimension value combination having the greatest impact on the overall experimental group value is selected as the root cause of the category.

5. The method of claim 4, wherein, The method further includes: from the dimension value combinations remaining in the set of candidate root causes of the category, an influence parameter of each dimension value combination is calculated, the influence parameter being a difference between a first change amplitude and a second change amplitude, the first change amplitude being a difference between the overall experimental group value and the overall control group value, the second change amplitude being a difference between the overall experimental group value and the overall control group value of the remaining dimension value combinations after the dimension value combination is excluded; a dimension value combination having the greatest influence parameter is selected as the root cause of the category.

6. The method according to any one of claims 3-5, wherein, The method further includes: in a case where the experimental group value is a proportional index, the generalized latent score of the dimension value combination is calculated according to the following formula: ; or, in a case where the experimental group value is a non-proportional index, the generalized latent score of the dimension value combination is calculated according to the following formula: ; wherein, is the experimental group value of the potential leaf node corresponding to the dimension value combination in the data set of the level, = is the control group value of the potential leaf node corresponding to the dimension value combination in the data set of the level, , is the control group value of the potential leaf node corresponding to the dimension value combination in the data set of the level, is the sum of the experimental group values of the potential leaf nodes corresponding to the dimension value combination in the data set of the level, is the sum of the control group values of the potential leaf nodes corresponding to the dimension value combination in the data set of the level; is a weighted average of the values calculated for each potential leaf node in the parentheses. In the case of the experimental group value being a ratio-type indicator, = the experimental group value for the dimension value combination at the level of the data set, = the control group value for the dimension value combination at the level of the data set, = the overall experimental group value for all potential leaf nodes in the data set at the level, = the overall control group value for all potential leaf nodes in the data set at the level, = the overall experimental group value for potential leaf nodes in the data set at the level other than the dimension value combination, = the overall control group value for potential leaf nodes in the data set at the level other than the dimension value combination. In the case of the experimental group value is a non-proportional index, .

7. A root cause mining device, the device comprising: an acquisition module configured to acquire experimental group values and control group values corresponding to each dimension value combination under a preset dimension combination, and to take each dimension value combination and the experimental group values and control group values corresponding to the dimension value combination as an initial leaf node; a construction module configured to, for each category of initial leaf node, construct a first level to an Nth level data set based on the initial leaf node of the category, wherein the Nth level data set includes potential leaf nodes corresponding to each dimension value combination under a combination of N dimensions, a potential leaf node corresponding to a dimension value combination including experimental group values and control group values corresponding to the dimension value combination and each dimension value combination of another single dimension; the preset number being a number of dimensions included in the preset dimension combination minus 1, and the value of N being in a range of 1 to the preset number; a calculation module configured to, for each category, sequentially traverse and calculate generalized latent scores of each dimension value combination of the first level to the Nth level data set of the category, and to mine a root cause causing abnormal change of the experimental group value from among dimension value combinations whose generalized latent scores satisfy a root cause condition.

8. The device of claim 7, further comprising a deletion module and a clustering module: the calculation module is further configured to calculate a difference value of each initial leaf node, the difference value being used to represent a difference between the experimental group value and the control group value of the initial leaf node; the deletion module is configured to, if the experimental group value has an abnormal upward trend, delete the initial leaf node having a difference value less than a mean difference value; or, if the experimental group value has an abnormal downward trend, delete the initial leaf node having a difference value greater than the mean difference value. The clustering module is configured to cluster the initial leaf nodes based on the difference values of the remaining initial leaf nodes.

9. The apparatus of claim 7, wherein, The computing module is specifically configured to: starting from the first level, for each dimension value combination included in each level of the category, calculate a generalized latent score of the dimension value combination according to the latent leaf nodes corresponding to the dimension value combination in the data set of the level; whenever the calculated generalized latent score of the dimension value combination is greater than a first preset threshold, add the dimension value combination to the candidate root cause set, and prune the latent leaf nodes subordinate to the dimension value combination; after the traversal of the category is completed, mine the root cause causing the abnormal change of the experimental group value from the candidate root cause set of the category.

10. The apparatus of claim 9, wherein, The computing module is specifically configured to: for each dimension value combination included in the candidate root cause set of the category, determine a same-category leaf node proportion of the dimension value combination, the same-category leaf node proportion being a ratio between the number of the latent leaf nodes corresponding to the dimension value combination in the category and the total number of the latent leaf nodes corresponding to the dimension value combination in all categories; if the same-category leaf node proportion is less than a second preset threshold, delete the dimension value combination from the candidate root cause set; from the remaining dimension value combinations in the candidate root cause set of the category, select a dimension value combination having the greatest impact on the overall experimental group value as the root cause of the category.

11. The apparatus of claim 10, wherein, The computing module is specifically configured to: from the remaining dimension value combinations in the candidate root cause set of the category, calculate an influence parameter of each dimension value combination, the influence parameter being a difference between a first change amplitude and a second change amplitude, the first change amplitude being a difference between the overall experimental group value and the overall control group value, and the second change amplitude being a difference between the overall experimental group value and the overall control group value of the remaining dimension value combinations after the dimension value combination is excluded; select a dimension value combination having the greatest influence parameter as the root cause of the category.

12. The apparatus of any one of claims 9-11, wherein, The computing module is specifically configured to: in a case where the experimental group value is a proportional index, calculate the generalized latent score of the dimension value combination according to the following formula: ; or, in a case where the experimental group value is a non-proportional index, calculate the generalized latent score of the dimension value combination according to the following formula: ; wherein, is the experimental group value of the potential leaf node corresponding to the dimension value combination in the data set of the level, , is the control group value of the potential leaf node corresponding to the dimension value combination in the data set of the level, is the sum of the experimental group values of the potential leaf nodes corresponding to the dimension value combination in the data set of the level, is the sum of the control group values of the potential leaf nodes corresponding to the dimension value combination in the data set of the level; is a weighted average of the values calculated for each potential leaf node in the parentheses.​ In the case of the experimental group value being a ratio-type indicator, = , is the overall experimental group value for all potential leaf nodes in the data set at this level, is the overall control group value for all potential leaf nodes in the data set at this level, is the overall experimental group value for potential leaf nodes in the data set at this level other than the dimension value combination, is the overall control group value for potential leaf nodes in the data set at this level other than the dimension value combination; In the case of the experimental group value is a non-proportional index, . 13.An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-6. 15.A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Operation and maintenance system fault positioning method and system based on Monte Carlo tree search

    CN112187554A

  • Fault root cause positioning method and device, equipment and storage medium

    CN115730774A