Approximate match
Approximate matching techniques with specified thresholds address inefficiencies in precise matching, reducing computational and memory demands while maintaining accuracy in large datasets.
Patent Information
- Application Number
- CN202080011129.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-31
- Filing Date
- 2020-01-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-01-22
AI Technical Summary
In the process of big data environment, the precise matching method leads to an exponential growth of potential matching, occupying memory and networking resources, and the complexity of the precise matching scheme increases when the symptom set is not completely correct or incompletely informed.
The approximate matching method is used to specify the matching function and the distance threshold function, allowing approximate matching between rows and masks, reducing the number of matched rows, and using binary, ternary, or quaternary notation to optimize storage and computing efficiency.
Effectively reduces the storage and computing requirements of matching operations and improves processing efficiency, especially when the symptom set is not fully informed, it can efficiently identify potential failure scenarios and root causes.
Smart Images

Figure CN113366465B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 799,613, filed on January 31, 2019, entitled "APPROXIMATE MATCHING", which is hereby incorporated by reference in its entirety for all purposes. Background of the Invention
[0003] The matching of row values against a table can be used in various fields such as root cause analysis, rule - based systems, and database queries using indexes. In terms of processing resources, using exact matching for such value matching can be efficient, but as the data size increases, using exact matching may suffer from an exponential growth of potential matches, which slows down processing and may also consume too much memory and / or networking resources. Brief Description of the Drawings
[0004] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.
[0005] Figure 1 is a functional diagram of a programmed computer / server system for approximate matching according to some embodiments.
[0006] Figure 2 is an example of a network created using element types for illustrating a partial - matching scenario.
[0007] Figure 3 is a block diagram of a power example illustrating additional partial - matching examples.
[0008] Figure 4A 、 Figure 4B and Figure 4C are flowcharts illustrating embodiments of a process for approximate matching. Detailed Description
[0009] The present invention may be implemented in many ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations or any other form that the present invention may take may be referred to as techniques. Generally, the order of the steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise stated, components such as a processor or a memory described as being configured to perform a task may be implemented as a general component temporarily configured to perform the task at a given time or a special component manufactured to perform the task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0010] The following, together with the accompanying drawings which illustrate the principles of the present invention, Figure 1 provides a detailed description of one or more embodiments of the present invention. The present invention is described in connection with such embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses numerous alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description in order to provide a thorough understanding of the present invention. These details are provided for purposes of example, and the present invention may be practiced without some or all of these specific details in accordance with the claims. For clarity, technical material known in the technical field related to the present invention has not been described in detail so as not to obscure the present invention unnecessarily.
[0011] Parametric matching is disclosed. When a value has one of multiple states, an exact match of a row's value against a table is efficient and useful, and where the multiple states include an "irrelevant" state, e.g., if the multiple states are two states, such as a binary representation having a "true" state and an "irrelevant" state, if the multiple states are three states, such as a ternary representation having a "true" state, a "false" state, and an "irrelevant" state, if the multiple states are four states, such as a quaternary representation having a "true" state, a "false" state, a "don't know true" state, and an "irrelevant" state, and so on. Useful applications of such exact matching include rule-based engines with sub-conditions, e.g., in the field of automated root cause analysis, where the table includes a root cause table of symptoms and the rows include actual failure scenarios.
[0012] Various techniques address the problem that, in some situations, the exact set of symptoms is not known or is not entirely correct. As one technique, a "partial match" operation provides multiple matches corresponding to additional rows for partial matches as well as rows for exact matches. The technique can be refined where some of the multiple matches can be suppressed as being subsumed by other matches.
[0013] The table matching mechanism can be generalized as an "index table" where columns are considered to correspond to sub-conditions rather than symptoms. Specific examples are given for "sub-conditions" and "masks" as generalizations of symptoms and actual fault scenarios respectively. Another specific example of a "sub-condition table" is used as a generalization of a root cause table. Any person of ordinary skill in the art will appreciate that the disclosed content applies to the general case of sub-conditions and masks without being limited to the specific examples of root cause analysis, rule-based systems, and index database queries.
[0014] Approximate matching for identifying problems is disclosed. For example, in a medical diagnosis application for identifying different types of diseases and disorders based on patient symptoms, there may be K different sub-conditions or features associated with a certain disease, but given that a symptom may be incorrect or missing, matching K - 1 features may be sufficient. Attempting to handle this situation using exact matching may require introducing K additional rows, one for each feature, where the i-th row sets the i-th feature entry to "irrelevant". Further, when the mask appropriately sets all K sub-conditions (e.g., with low overlap), the matching produces all K rows, or K + 1 rows if there are also rows corresponding to all K sub-conditions set to their actual values. If the application accepts matches with up to two mismatched entries, thus requiring K 2 additional rows, or accepts matches with three mismatched entries, thus requiring K 3 additional rows, these solutions are even more complex. Thus, using exact matching may suffer from an exponential growth of potential matches. If the application needs to dynamically select a distance or other threshold for the match, exact matching also becomes more challenging.
[0015] Approximate matching for such situations and other applications where not all features and / or sub-conditions are necessarily known are disclosed. As cited herein, "approximate matching" is at least partially based on the result of a specified matching function applied between sub-condition rows and masks. Examples of specified matching functions are distance threshold functions and intensity threshold functions.
[0016] Binary, ternary, and quaternary representations. In one embodiment, values have three or more states including values for "true", "false", and "irrelevant". In one embodiment, use ternarysub - condition values such that the sub - condition is represented as a "known" bit, indicating known or unknown by a bit value of true or false respectively, and a second "value" bit indicating true or false, which is interpreted as such only when the "known" bit is set to true. Using two bits to represent a ternary value quaternary is represented herein as [a, b], where a is whether the state is known (0 = unknown, 1 = known), and b is the value associated with the state (0 = false, 1 = true). Using this convention, the allowed interpretation of [0, 1] is that the associated sub - condition "is not known to be true". Compare [0, 0] which may correspond to an unknown state with [0, 1] which may be interpreted as "not known to be true" (an alternative convention could be that [0, 0] is "not known to be false"). Completing these possibilities, [1, 0] can correspond to "false", and [1, 1] can correspond to "true". Note that the [0, 1] sub - condition can match an input that is false or unknown, and unlike [0, 0], [0, 1] does not match any input that is true. Thus, [0, 1] may not necessarily be considered the same as [0, 0], and / or may not be allowed in an internal and / or consistent representation.
[0017] In one embodiment, the value has two states that allow different matching behaviors for different columns in a table. For a value with binary states, in many applications, an object has symptoms that are otherwise "irrelevant" to the client. For example, an object representing isOverheating (i.e., whether the unit is overheating). A "true" indication for isOverheating is a symptom of a particular problem, but there may be many root causes that are not relevant to whether the unit is overheating and are thus "irrelevant". Recognizing this, an alternative approach is that each entry has two values corresponding to "true" and "irrelevant" instead of three values. The corresponding matching operator output table is:
[0018] .
[0019] This reduces the number of bits required per entry from two bits in a ternary representation to one bit in a binary representation, so the memory overhead is only half.
[0020] There are cases where it is necessary to match on entities that all have symptom S1 and also on separate rules / root causes that do not have the symptom, such as when S1 is false. For such cases, an additional symptom S2 corresponding to the negation of S1 can be explicitly introduced. For example, symptom S1 could be lossOfPower and S2 could be hasPower. Then, the rows that require lossOfPower can have the corresponding table entry with symptom S1 set to "true" and S2 set to "irrelevant". Conversely, the rows that require hasPower can have the corresponding table entry with S2 set to "true" and the S1 entry set to "irrelevant". Thus, if for example 10% of the symptoms need to be negated, the 1-bit approach will still be spatially cheaper than the 2-bit approach, that is, an extra 10% of space rather than 100%.
[0021] In one embodiment, the combination where both S1 and S2 are true can correspond to "unknown". That is, if the input sets both S1 and S2, the symptom is considered "unknown". In an application, representing the state as "unknown" may be more useful than an interpretation of "don't know true" for a fourth value. There are no changes or extensions to the above matching for this "unknown" support to take effect. Both the S1 and S2 entries in the row are specified as true, and the input also specifies these entries as true, so it matches "unknown".
[0022] There are also symptoms that are actually discrete enumerations of values, such as very cold, cold, cool, normal, warm, hot, and very hot. In a normal binary representation, these seven values can be represented with three bits. However, with the above "true" / "irrelevant" binary approach, the affirmatives and negatives will require six symptoms as separate symptoms for each bit. Specifically, the above 7 enumerations can be represented with 3 bits using a conventional boolean representation, for example, 001 can indicate "very cold". However, since the true / irrelevant embodiment does not explicitly handle false, another 3 bits or symptoms are needed to specify false, that is, 0. Thus, the 001 encoding for "very cold" will be indicated by six symptoms [X, X, 1, 1, 1, X], where "X" indicates "irrelevant", and the first 3 entries correspond to true or 1 values, and the last 3 entries correspond to false or 0 entries. Alternatively, seven separate symptoms for a "one-hot" representation of these seven values can be used, where the "one-hot" representation is referred to herein as a group of bits, where the only legal combinations of values are those with a single "true" bit and all others not "true".
[0023] In one embodiment, a specific set of columns is used, and the symptoms are used to represent an enumeration of such multiple symptom values. In this case, the specific columns are marked as matching "true" / "false", rather than "true" / "not relevant", to represent the enumeration in three columns rather than six or seven columns. The "not relevant" value is still needed so that rows that are not relevant to the enumerated symptom values can indicate this. Therefore, it is allowed to specify the column value width and reserve a value to indicate "not relevant", for example, all zeros. In this case, a logical column width of three bits can be specified, so [0, 0, 0] can correspond to "not relevant", and the combinations of the remaining seven values respectively represent seven different symptom states. Therefore, this enumeration can be represented in three bits, but still allows specifying "not relevant" in rows that are not relevant to the enumerated symptom.
[0024] Therefore, this additional extension allows the table to specify the logical column width and treat the all-zero case as "not relevant". Then, the columns corresponding to isOverheating or S1, S2 can be marked as one-bit wide, while the columns corresponding to the above enumeration will be marked as a logical column with three bits.
[0025] In one embodiment, if the next column is also part of a logical column / entry that spans multiple columns, an efficient representation of the column width is to have one bit set for each column. Equivalently, this bit can indicate the continuation of the previous column. Therefore, the columns corresponding to the enumeration can be marked as [1, 1, 0], where the last bit indicates no continuation beyond the third column, to indicate a three-bit column that forms a logical column / entry.
[0026] Generally, the table can specify the column width and the matching behavior of the logical columns. As an example of this more general case, the table can mark the logical column as five bits wide, and the matching behavior is a test for "greater than". That is, if the input value in these five bits is a binary value greater than the value in the table entry, the input can match the entry, except that if it is zero, the table entry is considered "not relevant", or if all are ones, the table entry is considered "unknown". As another example of specifying the matching behavior, the matching behavior can be specified as the logical OR of the inputs on these columns. By having a bit in the logical column indicating whether the symptom is relevant to this row and excluding it from the OR, the all-zero value in the table entry can be considered "not relevant".
[0027] In one embodiment, hardware instructions such as "ANDN" or logical AND NOT instructions in Intel and AMD instruction sets can be used for binary representation. Hardware instructions generally execute faster than their software counterparts and are more programmatically efficient. The ANDN instruction performs a bitwise logical AND on the inverted first operand and the second operand, for example, inBlock and tableBlock the two operations mismatchThe results are as follows:
[0028]
[0029] The corresponding output table:
[0030]
[0031] And thus, if mismatch is match the complement of ,
[0032] .
[0033] Figure 1 is a functional diagram of a programmed computer / server system for approximate matching according to some embodiments. As shown, Figure 1 provides a functional diagram of a general-purpose computer system programmed for approximate matching according to some embodiments. As will be clear, other computer system architectures and configurations can be used for approximate matching.
[0034] Computer system 100, which includes various subsystems as described below, includes at least one microprocessor subsystem, also referred to as a processor or central processing unit ("CPU") (102). For example, the processor (102) can be implemented by a single-chip processor or by multiple cores and / or processors. In some embodiments, the processor (102) is a general-purpose digital processor that controls the operation of computer system 100. By using instructions retrieved from the memory (110), the processor (102) controls the reception and manipulation of input data, as well as the output and display of data on output devices such as a display and a graphics processing unit (GPU) (118).
[0035] The processor (102) is bi-directionally coupled to the memory (110), which may include: a first main storage device, typically a random access memory (“RAM”); and a second main storage area, typically a read-only memory (“ROM”). As is well known in the art, the main storage device can be used as a general storage area and as a scratch-pad memory, and can also be used to store input data and processed data. In addition to other data and instructions for the processes operating on the processor (102), the main storage device can also store data and programming instructions in the form of data objects and text objects. Also as is well known in the art, the main storage device typically includes basic operation instructions, program code, data, and objects used by the processor (102) to perform its functions, such as programmed instructions. For example, the main storage device (110) can include any suitable computer-readable storage medium as described below, depending on whether the data access needs to be bi-directional or unidirectional. For example, the processor (102) can also retrieve and store frequently needed data directly and very quickly in a cache memory (not shown). The processor (102) can also include a co-processor (not shown) as a supplementary processing component to assist the processor and / or the memory (110).
[0036] The removable mass storage device (112) provides additional data storage capacity for the computer system 100 and is bi-directionally (read / write) or unidirectionally (read-only) coupled to the processor (102). For example, the storage device (112) can also include computer-readable media, such as flash memory, portable mass storage devices, holographic storage devices, magnetic devices, magneto-optical devices, optical devices, and other storage devices. The fixed mass storage device (120) can also provide additional data storage capacity, for example. An example of the mass storage device (120) is an eMMC or a micro SD device. In one embodiment, the mass storage device (120) is a solid-state drive connected via a bus (114). The mass storage devices (112), (120) generally store additional programming instructions, data, etc. that are not normally in active use by the processor (102). It will be appreciated that the information retained within the mass storage devices (112), (120) can be incorporated in a standard manner, if desired, as part of the main storage device (110) (e.g., RAM), as virtual memory.
[0037] In addition to providing access to the processor (102) of the storage subsystem, the bus (114) can also be used to provide access to other subsystems and devices. As shown, these can include, as needed, a display monitor (118), a communication interface (116), a touch (or physical) keyboard (104), and one or more auxiliary input / output devices (106), including an audio interface, a sound card, a microphone, an audio port, an audio recording device, an audio card, speakers, a touch (or pointing) device, and / or other subsystems. In addition to a touchscreen and / or a capacitive touch interface, the auxiliary device (106) can be a mouse, a stylus, a trackball, or a tablet device, and is useful for interacting with a graphical user interface.
[0038] The communication interface (116) allows the processor (102) to be coupled to another computer, a computer network, or a telecommunications network by using a network connection as shown. For example, through the communication interface (116), the processor (102) can receive information, such as data objects or program instructions, from another network, or output information to another network during the execution of method / process steps. Information, typically represented as a sequence of instructions to be executed on the processor, can be received from and output to another network. Interface cards or similar devices, as well as appropriate software implemented by execution / implementation, for example, on the processor (102), can be used to connect the computer system 100 to an external network and transfer data according to standard protocols. For example, various process embodiments disclosed herein can be executed on the processor (102), or can be executed in combination with a remote processor sharing a portion of the processing, across a network such as the Internet, an intranetwork, or a local area network. Throughout this specification, "network" refers to any interconnection between computer components, including the Internet, Bluetooth, WiFi, 3G, 4G, 4GLTE, GSM, Ethernet, TCP / IP, intranet, local area network ("LAN"), home area network ("HAN"), serial connection, parallel connection, wide area network ("WAN"), Fibre Channel, PCI / PCI-X, AGP, VLbus, Fast PCI, ExpressCard, InfiniBand, ACCESS.bus, wireless LAN, HomePNA, fibre, G.hn, infrared network, satellite network, microwave network, cellular network, virtual private network ("VPN"), universal serial bus ("USB"), FireWire, Serial ATA, 1-Wire, UNI / O, or any form that connects homogeneous, heterogeneous systems, and / or groups of systems together. Additional mass storage devices not shown can also be connected to the processor (102) through the communication interface (116).
[0039] Auxiliary I / O device interfaces (not shown) may be used in conjunction with computer system 100. Auxiliary I / O device interfaces may include both general and customized interfaces that allow the processor (102) to send data and, more typically, receive data from other devices such as microphones, touch-sensitive displays, transducer readers, tape readers, voice or handwriting recognizers, biometric readers, cameras, portable mass storage devices, and other computers.
[0040] In addition, various embodiments disclosed herein further relate to computer storage products having computer-readable media that include program code for performing various computer-implemented operations. A computer-readable medium is any data storage device that can store data that can be thereafter read by a computer system. Examples of computer-readable media include, but are not limited to, all of the above-mentioned media: flash media such as NAND flash, eMMC, SD, compact flash; magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD-ROM disks, etc.; magneto-optical media such as optical disks; and specially configured hardware devices such as application specific integrated circuits (“ASICs”), programmable logic devices (“PLDs”), and ROM and RAM devices. Examples of program code include both machine code, such as produced, for example, by a compiler, or files containing higher-level code, such as scripts that can be executed by using an interpreter.
[0041] Figure 1 The computer / server system shown in is merely an example of a computer system suitable for use with the various embodiments disclosed herein. Other computer systems suitable for such use may include additional or fewer subsystems. Additionally, bus (114) is illustrative of any interconnect scheme for linking subsystems. Other computer architectures with different subsystem configurations may also be utilized.
[0042] Partial match example. Figure 2 is a network example for creating a diagram of a partial match scenario using an element type. In Figure 2 , two switches - switch 0 / SW0 (202) and switch 1 / SW1 (242) - each are instances of the switch types described earlier. Each of switches SW0 (202) and SW1 (242) is shown to have one component interface that is I14-a / eth14 (204) and I3-a / eth3 (244) respectively, and power sensors for the switches that are SW0 (206) and SW1 (246) respectively. Figure 2 Power sensors for the network interfaces (204, 244) may also be not shown in - if they are separately powered.
[0043] In a real network, each switch can typically have multiple interfaces, and potentially many more switches. The link (222) between two connected interfaces is modeled as two component one-way links (222a and 222b), each being an instance of a single link. This level of detail allows for the specification of implied directionality. It also allows for the modeling of faults, such as one direction of the link failing while the other continues to operate.
[0044] This simple model illustrates how rule conditions can be specified in an element type such as a link (222) where there may be no telemetry on the link (222) as there are no sensors on the link (222). However, the associated sub-conditions may imply sub-conditions in intermediate elements such as one-way links and then to the connecting interfaces to their parent switches where they imply observable sub-conditions.
[0045] For example, a symptom in the current symptoms for the network is that a computer network switch reports "loss of signal" on a particular interface among its interfaces. The actual fault scenario of "loss of signal" reported by SW0 (202) on interface I14-a (204) can match a fault scenario FS1 corresponding to a link fault in link a (222) between switch SW0 (202) and switch SW1 (242). However, the same symptom can also match a fault scenario FS2 where the two interfaces (204, 244) at both ends of the link have failed simultaneously. It can also match a fault scenario FS3 corresponding to a link fault in link a (222), but without considering secondary symptoms such as symptoms corresponding to power loss at SW0 sensor (206) and SW1 sensor (246) known to be false. Thus, in this example, the set of matches consists of FS1, FS2, and FS3. The tabular expression for this situation is:
[0046]
[0047] A base scenario is subsumed by its associated derived scenarios of potential fault scenarios. In one embodiment, the attributes of potential fault scenarios indicate when one potential fault scenario FSa is subsumed by another potential fault scenario FSb. That is, whenever FSb matches, FSa will also match. As referred to herein, FSa is basic scenario and FSb is derived scenario . In the case where both FSa and FSb match, the refinement of the set of matches is to remove FSa from the set of matches before transforming the fault scenario into its associated root cause.
[0048] To illustrate this situation, continue Figure 2For the network example, the matching refinement step will identify that FS3 is subsumed by FS1 because FS3 only needs to match a subset of the symptoms required by FS1.
[0049] 。
[0050] Another simple example of the base scenario subsumed by a derived scenario is a medical example:
[0051] • The potential failure scenario FSm shows a 90% confidence level for the root cause of influenza given the symptoms of high body temperature and pain; and
[0052] • The potential failure scenario FSn shows a 90% confidence level for the root cause of influenza given the symptoms of high body temperature, pain, and headache.
[0053]
[0054] Therefore, in the case of an actual failure scenario including the symptoms of high body temperature, pain, and headache, FSm is identified as the base scenario subsumed by the derived scenario FSn, and thus the root cause of influenza with a 90% confidence level is output.
[0055] 。
[0056] Combination of output probabilities. In one embodiment, the refinement can identify that two potential failure scenarios present in the match set are actually two different symptom sets of the same root cause and may in fact both be true, and thus output the potential root cause that may have an associated probability that is a combination of the probabilities of the two potential failure scenarios. For example, FSn can be the following potential failure scenario: it shows a 90% confidence level for the root cause of influenza given the symptoms of high body temperature, pain, and headache, and FSp can be the following potential failure scenario: it shows a 5% confidence level for the root cause of influenza given the symptoms of runny nose and earache.
[0057]
[0058] A patient with the symptoms of high body temperature, pain, headache, runny nose, and earache can be identified as having a combined associated probability that is a combination of a 90% confidence level and a 5% confidence level. In one embodiment, the confidence levels can be linearly summed.
[0059]
[0060] Alternative explanations. In one embodiment, the attributes of the potential failure scenario indicate when one potential failure scenario FSc is of another potential failure scenario FSdsubstitution possibility 。Thus, when both FSc and FSd are present in the match set, the refinement indicates them as part of a subset of the potential root causes of the actual failure scenario, rather than indicating the two matches as two separate possible failures and / or indicating the two matches as part of different root cause groups. In an embodiment, the property of indicating a potential root cause as an alternative can be calculated by comparing the symptoms of two potential root causes. It is an alternative and has a subset of the symptoms of other potential root causes, and it is not the basic root cause of the symptom; it is an alternative. substitute
[0061] For example, using the Figure 2 network example, the refinement will indicate FS1 and FS2 as alternatives to each other, given that both scenarios correspond to a common set or subset of symptoms.
[0062] 。
[0063] Another simple example of an alternative explanation is a medical example:
[0064] • The potential failure scenario FSn shows a 90% confidence level of a flu root cause given the symptoms of high body temperature, pain, and headache; and
[0065] • The potential failure scenario FSq shows a 3% confidence level of a hay fever root cause given the symptoms of high body temperature, pain, and headache;
[0066]
[0067] Thus, in the case of an actual failure scenario including the symptoms of high body temperature, pain, and headache, FSq is considered an alternative explanation to FSn.
[0068]
[0069] In one embodiment, another property of a potential failure scenario is the probability of that failure scenario relative to its associated alternative failure scenario. For illustration, using the Figure 2 network example, the probability of FS1 can be 0.95, and the probability of FS2 as an alternative to FS1 can be assigned 0.05. Then, the match set refinement can rank the associated root causes according to the probabilities associated with each alternative. Thus, in the Figure 2 network example, the refined set of root causes can be:
[0070] [RC1: 0.95, RC2: 0.05]
[0071] Where RC1 corresponds to the root cause associated with the failure scenario FS1, and RC2 corresponds to the root cause associated with the failure scenario FS2. This refinement eliminates the third entry because FS3 is subsumed by FS1.
[0072]
[0073] Associating probabilities with potential failure scenarios may be more feasible than input probability methods because each failure scenario represents a situation where a top-level failure requires remediation. Thus, the operational data can indicate the frequency of a given root cause compared to the frequency of alternatives (i.e., alternatives with the same symptoms) occurring. For example, returning to Figure 2 the network example of a if a broken link
[0074] (222) is the actual root cause observed 95 out of 100 times for the associated symptoms, and only 5 out of 100 times are the cases where two interfaces (204, 244) actually fail simultaneously, the recorded operational data provides a basis for weighting and ranking these two alternative root causes using these probabilities. a Thus, initially considering the output as a detected broken link
[0075] (222), the remediation action will resolve the actual root cause failure immediately most of the time and will only require an alternative failure remediation action 5% of the time. In some cases, such as when two interfaces (204, 244) fail simultaneously, the user can estimate the probability-based average time for interface repair, the frequency of individual interface failures, and the number of interfaces, thereby further qualifying the likelihood that the two interfaces that fail within the same recovery window are actually at either end of the link. Note that it is possible, although unlikely, that the link has failed and the two interfaces have also failed. That is, the alternative root causes may not be mutually exclusive. In such a case, remediation actions for both failures are required.
[0076] For example, if the potential failure scenario designates symptom Si as an oven temperature above 100 degrees Celsius, the actual failure scenario should include that symptom reported as above 100 degrees Celsius.
[0077] This matching is in contrast to input probability methods used, for example, in traditional ARCA techniques, where, given the uncertainty about the sensor as captured by the correlation probability, there may be some probability that the symptom is true even if the sensor is not reporting this. In one embodiment, the generation of the matching set is performed by a ternary matching mechanism having a ternary RCT representation as described herein.
[0078] The set of incomplete fault scenario matches can include multiple members, even in the case of matching a single actual fault, partly because the set of potential fault scenarios should cover cases where some telemetry is missing or incorrect. For example, FS3 in the above network example is provided such that there is some match even if the telemetry of the associated symptom is incomplete or incorrect. That is, it would be unacceptable to diagnose a link a fault in link (222) just because one (202) or the other of the switches (242) is unable to report to the interface regarding power (206, 246).
[0079] In general, matching can be implemented efficiently and simultaneously match multiple independent root causes, as described above in the application of the ternary fault scenario representation. A drawback of matching is that it fails to match when any of the specified symptoms in the potential fault scenario corresponding to the actual fault scenario do not match the symptoms determined from the telemetry. This can occur even when a human assessment of the symptoms may quickly arrive at a conclusion as to what the root cause is.
[0080] Partial matching. Figure 3 is a block diagram of a power example that illustrates another example of partial matching. In this power example, switch SW0 (302) is fully coupled via interfaces and links to 24 other switches SW1 (342) and SW2 (362) through SW15 (392). As shown previously in Figure 2 each switch, such as switch SW0 (302), includes a power sensor (302z) and one or more interfaces I1-a (302a), I1-b (302b),..., I1-x (302x) corresponding respectively to links a (322a), b (322b),..., x (322x).
[0081] If the power supply to computer network switch SW0 (302) including SW0 power sensor (302z) fails, then each interface that the switch is expected to be connected to via a link will detect a loss of signal. However, if the switch in question is connected via links to 24 separate interfaces I2-a (342a), I3-b (362b), …, I25-x (392x), but only 23 of these interfaces are reporting a loss of signal and the 24th interface I25-x (392x) is missing from the telemetry, then the matching will fail to match the potential failure scenario that specifies all 24 separate interfaces with the symptom of power loss, even though this can be reasonably determined based on the symptoms that the switch has failed, and furthermore, if the SW0 power sensor (302z) of the switch reports a power loss, then the switch has failed due to power loss.
[0082] As shown herein, it is important to utilize this matching ability to match multiple failure scenarios simultaneously in order to compensate for this shortcoming. In particular, in addition to potential failure scenarios corresponding to all symptoms, there are potential failure scenarios that are specified to correspond to partial matches for the same root cause. The extension of the associated attributes of the potential failure scenarios allows for the refinement of the matching set to reduce the number of potential root causes of the actual output.
[0083] In particular, when a match with a complete potential failure scenario occurs, the potential failure scenarios of partial matches corresponding to the same root cause are eliminated and / or subsumed. Similarly, when the root cause exists only due to an effective partial match, the probability attribute associated with the potential failure scenario allows the output to effectively indicate a lower confidence in that root cause in the output.
[0084] In one embodiment, approximate match For situations where not all features (e.g., sub-conditions) need to be known. Thus, approximate matching can be used in combination with partial matching.
[0085] In one embodiment, approximate matching is provided by specifying a distance threshold parameter, and if a row is within the distance threshold according to some distance metric defined between the row and a mask, then the row is output as a match. As will be described in further detail below, processing additional matches to reduce and organize the matches for the efficiency of explanation can be improved by approximate matching, which is partly by treating an approximate match at distance D as a fundamental root cause relative to a match at distance D-1.
[0086] Partial Match Potential Failure Scenarios (PMPFS). PMPFS are referred to herein as potential failure scenarios that are added to effectively handle partial matches with the matching mechanism. Various techniques define PMPFS.
[0087] PMPFS ignoring one symptom First, for each fully potential failure scenario of the root cause, for each symptom, there may be a PMPFS that ignores one of the symptoms. For example, using the Figure 3 power example, for each adjacent interface, there may be a PMPFS that ignores the interface as a symptom or, alternatively, labels the symptom as "irrelevant". For example, the PMPFS may ignore I25-x (392x) as "irrelevant", and thus where I2-a (342a), I3-b (362b) …… I24-w ( Figure 3 not shown) reports a loss of signal, the system may conclude that switch SW0 (302) has failed.
[0088] Furthermore, it may be possible to provide PMPFS for a subset of the symptoms of the fully potential failure scenario. An example is to create PMPFS for both I24-w and I25-x (392x) as "irrelevant". However, this may lead to an impractical number of PMPFS in a system with realistic complexity. For example, in a switch example with 32 directly adjacent switches, there are essentially 2 to the 32nd power or roughly 4 billion possible subsets. Here, approximate matching solves the problem of too many PMPFS. In other words, partial matching can be regarded as adding additional rows that are less complete, while approximate matching can parameterize the matching by relaxing the matching criteria and thus can match rows that do not exactly match the mask or the actual complete set of symptoms at least partially.
[0089] PMPFS excluding value range One way to effectively support partial matching while avoiding an exponential growth in the number of PMPFS is to allow a potential failure scenario to specify a given symptom as excluding a certain value or range of values. Values that are contradictory to the associated failure as the root cause are typically used. In the Figure 3 power example, the PMPFS can be specified to require the lossOfSignal symptom to be true or unknown. Then, a match occurs as long as no adjacent switch claims to have received a signal from the switch that is allegedly losing power. That is, if the symptom is unknown for some of the adjacent switches (such as the unknown I25-x (392x)), the match still occurs.
[0090] In one embodiment, the representation of PMPFS allows for specifying exclusion-based matching in range specifications, not just inclusion. For example, a two-bit representation of a value can use a value of "unknown but true" (e.g., [0,1]), which is otherwise used to label "not known to be true". Generally, there are traditional techniques for data representation that can be used to efficiently encode additional information corresponding to exclusion as well as inclusion.
[0091] limiting the range of PMPFS Another way to effectively support partial matching while avoiding exponential growth in the number of PMPFSs is to limit the scope of PMPFSs and their symptoms and correspondingly reduce the probabilities associated with them. In Figure 3 the power example, the following PMPFS can be generated: The PMPFS matches on the current power fault sensor (302z) for switch SW0 (302) and specifies "irrelevant" that is valid for telemetry of adjacent switches (342, 362... 392). If the power sensor (302z) reports a power fault but there is conflicting information from one or more adjacent switches (such as "unknown" for I25-x (392x), which may be incorrect or outdated), then the PMPFS matches.
[0092] On the other hand, if the above PMPFS for the same switch SW0 (302) matches an exclusion-based match, the lower-probability match is filtered out by the refinement step. In general, the generation of PMPFSs can be scoped based on relationships with other elements, the types of other elements, specific characteristics of these elements, and other attributes.
[0093] defining aggregated symptoms Another way to effectively support partial matching while avoiding exponential growth in the number of PMPFSs is to define aggregated symptoms set based on telemetry across multiple sensor inputs. In Figure 3 the power example, an aggregated symptom can be defined that corresponds to more than a certain threshold K among adjacent switches SW1 (342), SW2 (362)... SW15 (392) that have lost signals from a given switch SW0 (302). Then, the PMPFS for switch power loss can specify this aggregated symptom such that if most of the direct neighbors of the switch have lost signals from it, the switch is considered to have had a power failure. Specifically, the benefit of incorporating information from its direct neighbors is that it helps to eliminate the ambiguity of the situation where the current sensor on the switch has failed rather than the power itself.
[0094] In one embodiment, the matching operation accepts zero or more parameters that indicate criteria for approximate matching on an acceptance row. Approximate row matching is the result of comparing a given row with a specified mask, where the row meets certain matching criteria relative to the mask, even though not every pair of entries can be strictly matched. A match operator is defined between two subconditions s0 and s1 to return true
[0095] ;
[0096] If each sub - condition entry in s0 is either "irrelevant" or matches the value in the corresponding entry in s1. Note that the matching operation is non - commutative; match ( a , b ) may not necessarily be equal to match ( b , a ).
[0097] The approximate matching operation on a table is to perform a row - wise approximate match between the mask and each row of the table and output the results of those rows that meet the approximate matching criteria. In one embodiment, a specified matching function is applied between the mask and a row indicating whether the specified row approximately matches the mask. The approximate match of the mask against the table is performed by calling the matching function on each row of the table with the specified mask and outputting those rows in the result set that match the approximate matching criteria.
[0098] In one embodiment, approximate matching is provided by: specifying distance a threshold parameter and, based on some distance metric defined between the row and the mask, outputting them as matches if the rows are within the distance threshold. For example, one distance metric would be
[0099]
[0100] where N is the number of entries in the rows where the mask is true and the entry is false, or vice versa, P is the number of entries where the entry is "not known to be true" and the mask is true, and R is the number of rows where the entry is known and the mask entry is unknown. The constants preceding these parameters can be specified in registers or variables and thus vary dynamically depending on application requirements. Many other distance metrics and associated parameters can be implemented. In one embodiment, the approximate matching operation incrementally calculates the distance as it matches each entry in the row with the mask and determines a non - match once the calculated distance exceeds the distance threshold. In one embodiment, the distance metric is specified as a relative percentage rather than an absolute value.
[0101] This approximate matching contrasts with the use of distance in traditional binary matching for root - cause analysis, where, for example, after comparing with each row, the row with the shortest Hamming distance to the mask across all rows is selected. In particular, as described herein, instead, there is a distance threshold applied per row. Another difference is that, as described herein, the distance metric is flexible and not limited to a particular distance metric such as the Hamming distance.
[0102] In one embodiment, any distance metric can be provided for approximate matching and be better suited for a particular application. A maximum distance parameter can be specified so that only those rows at a distance at most as large as the maximum distance are returned in the matching result. For example, if there is no exact match with a given mask having a distance of at most one, the application can query for matches with a distance of at most two to see if any rows now match.
[0103] As a special case of the distance metric, each mismatch adds one to the distance, so the distance is only the number of mismatches allowed for that row before the row is declared a mismatch. For example, in the above example of a medical diagnosis application, a threshold of one would allow a match of K-1 sub-conditions by accepting a row as a match if the row has mismatched entries in only zero or one of the entries relative to the mask.
[0104] In some implementations, a distance threshold will be provided to the matching operation. When starting the matching operation for the next row against the mask, the matching operation will initialize a distance variable to zero. Then, it will continue to match the mask against the row entry by entry. In the case of a mismatch of an entry, the process will increment the distance variable by the cost of the mismatch. Then the variable will be compared against the distance threshold. If the variable is now greater than the distance threshold, it will declare the row a mismatch and continue with the next row in the table. Otherwise, it will continue with the entry-by-entry matching in that row. In one embodiment, matching the mask against the row is block-by-block instead of entry-by-entry, where the block structure is a fixed-length (e.g., 32-bit or 64-bit wide) representation of multiple / consecutive values for compact memory representation.
[0105] In one embodiment, different distance parameters are used for different types of mismatches. In particular, there are an unknown distance threshold parameter and a known distance threshold parameter, where the former applies when the row entry contains a known value and the corresponding mask entry contains an unknown entry, while the latter applies when there is a mismatch due to different known values in the corresponding entries in the row entry and the mask. A separate parameter can handle "don't know as true" values, or it can be handled under an existing threshold.
[0106] In one embodiment, approximate matching includes specifying intensity a threshold parameter and calculating, for each row, a strength associated with the match between that row and the mask according to some strength function. If the match strength of a row is equal to or greater than the specified strength threshold, that row is considered to match the mask. In one embodiment, the strength threshold is specified as a relative percentage rather than an absolute value.
[0107] As with distance, a variety of strength functions are possible. In a particular embodiment, the weight of a match is simply one, and thus the strength of a match corresponds to the minimum number of entries that need to be matched in order to declare a row a match. For example, in the above example of a medical diagnosis application, if a row matches at least three entries, a strength of three will accept that row as a match. For example, if it is observed that a patient has these three symptoms, the approximate match result set for the row will indicate the possible diseases to consider. Additionally, if the match result returns multiple rows, the sub-condition differences between these rows can indicate which sub-conditions to test for will most effectively distinguish between these matching rows in order to narrow the diagnosis.
[0108] For example, if the match result rows differ on sub-conditions that can be determined by a blood test, the doctor can arrange a blood test for the patient to efficiently refine the diagnosis. In the case of having the results of the blood test, the approximate match can be run again with the (one or more) additional sub-conditions and a greater strength value specified in the mask, thereby producing a new match result set that provides a more accurate or refined diagnosis. If necessary, this sequence of actions can be repeated to further refine the diagnosis. In one embodiment, the strength may not be set to be greater than the number of entries set in the mask, because then it is not possible to match that many entries for the mask, at least when each match has a weight of one.
[0109] In one embodiment, different threshold parameters are used for the strength of different types of match / mismatch, similar to different distance thresholds. For example, there may be an exact match strength threshold that requires a row to match exactly on at least so many entries. An alternative is to have a factor in the strength calculation for the exact match case that is set high enough such that other mismatches generally cannot compensate for the lack of an exact match. For example, if the factor for an exact match is 1000 and other mismatches have a very small strength, setting the strength threshold to 4000 means the mask must have at least 4 exact matches.
[0110] In one embodiment, weights are associated with table entries. When a match operation is invoked, for each row to be matched, the match initializes a current strength variable to zero. Then, for each entry in the row that matches the mask, the current strength variable is updated as a function of the weight associated with that entry. The resulting weight is then compared to a strength threshold parameter. If the current strength is greater than or equal to the value of the strength threshold parameter, the row is considered a match. Otherwise, the match continues to match against the row entries.
[0111] In one embodiment, different weights are associated with matching entries based on the nature of the match. For example, matching a row entry that specifies an "irrelevant" value may have a different weight compared to matching a row entry because of a known value that is true or false. In one embodiment, the parameters of the matching operation indicate the weight of an "irrelevant" entry match. A weight of zero effectively means that the "irrelevant" match has no impact on the weight calculated.
[0112] In one embodiment, a potentially negative weight is associated with a mismatch of an entry. For example, when the mismatch corresponds to false in a row entry and unknown in a mask entry, one specified weight can be used, and when the mismatch corresponds to false in a row entry and true in the corresponding mask entry, another specified weight can be used.
[0113] In one embodiment, the weights associated with the entries in a table are specified by a weight matrix corresponding to column and row groups in the table. For example, the weight matrix can be initialized such that the weight matrix [i, j] entry includes the weight of the entry within the i-th row range and the j-th column range.
[0114] In one embodiment, the parameters of the matching operation indicate that if the parameter is true, then a specified / known entry in the mask that matches an "irrelevant" entry in the row increments the current strength variable, and otherwise the current strength variable remains unchanged. To effectively disable the strength mechanism, the matching operation will be called with the parameter set to true and the strength parameter set to the number of specified / known entries in the mask.
[0115] In one embodiment, both a distance threshold and a strength threshold parameter are supported. The implementation has two separate counters and performs the processing specified above for mismatches as well as for matches. Used together, the matching operation can return all / approximate matches that have a strength above a specified threshold and below a specified distance from the mask. These distance and strength threshold parameters can specify percentages rather than absolute values. For example, the percentages can be specified relative to the non-"irrelevant" entries in the row being matched.
[0116] In an application, when the intensity threshold parameter is set, the distance threshold parameter is set to a larger value, so it is effectively disabled to prevent row matching due to mismatches. For example, in a medical diagnosis application, the number of additional sub-conditions associated with a given row is not important; what is desired is to match to indicate which rows are to be considered as potential diagnoses to determine additional tests to narrow down the diagnosis. In this case, entry mismatches may occur because the mask entries are unknown, as these tests have not yet been performed. When using the distance threshold parameter, from the perspective of the intensity parameter, the matching operation can be configured to consider mismatches below the mismatch threshold as matches. Thus, if the intensity threshold is set to K, a mask with K set entries and a mismatch threshold of three can match as expected, because in the calculation of intensity, at most three mismatched entries will be counted as matches.
[0117] In one embodiment, different distance thresholds are associated with different subsets of table entries. The matching operation then maintains a distance variable for each different subset and considers mismatches within each subset as matches until the distance threshold for the subset is reached or a row match is determined. In one embodiment, a distance threshold matrix stores these different thresholds, where the rows of the matrix correspond to subsets of the rows in the table and each column corresponds to a subset of the columns in the table. Then, when a mismatch occurs on an entry in a row, the matching operation increments the distance variable associated with that entry and compares the result against the threshold stored in the entry of the distance threshold matrix; if the count is greater than the threshold, a mismatch is declared. In one embodiment, the subsets are determined by index ranges on columns and indices.
[0118] To illustrate the utility of different thresholds, consider again the medical diagnosis application. The sub-condition table can be organized such that the least common or most detailed sub-conditions / symptoms appear in the first half of the columns. It can also be organized such that the best-understood diseases and disorders appear in the first few rows of the table. The distance threshold matrix can then specify lower distance thresholds for common sub-conditions associated with rows corresponding to well-understood diseases, and higher distance thresholds for less common sub-conditions associated with less well-understood diseases, as well as appropriate thresholds for other combinations. A similar approach can be used to support multiple intensity thresholds associated with different subsets of the entries in the table.
[0119] In one embodiment, the columns in a table are sorted to facilitate efficient matching. For example, using the above example of a medical diagnosis application, the least common sub-conditions can be placed in the initial range of columns. The "least common" characteristic can be determined by examining all rows and counting the number of rows in which each sub-condition appears. Then, a distance threshold can be associated with the columns of that range, without the need for a more complex mapping from entries to distance thresholds. This sorting also means that when row matching proceeds sequentially through the rows, the matching operation for a row frequently terminates early, resulting in greater efficiency than when the most common sub-conditions appear first.
[0120] As referred to herein, sparse matrix is any matrix that has more zero entries than non-zero entries in the matrix. Approximate matching can be implemented using sparse matrix-vector multiplication (SPVM) techniques, where the multiplication operation is replaced with a matching calculation. This can be: a vector multiplication operation applied to each pair of vectors / rows being matched; or an entry-level multiplication operator. In the latter case, the addition of the results is replaced by an appropriate distance or intensity calculation. The resulting vector indicates those that match the entries within a threshold. Clearly, exact matching can also be achieved using this method with a suitable setting of the threshold. As an example of applying the above technique, for a GPGPU (general-purpose computing on graphics processing units (GPU)) implementation, the same representation of a table as the matrix and a mask as the vector used for SPMV GPU implementation can be used, except that the kernel is replaced with a matching calculation instead of an entry or vector multiplication.
[0121] In various embodiments, multiple approximate matching functions can be provided, and multiple functions can be applied to each row simultaneously, such as the distance and intensity calculations mentioned earlier, and different approximate matching functions can be used depending on the row.
[0122] The performance impact on matching may be minimal because these calculations do not cause significant additional access to memory, unless there are a large number of entry-specific thresholds - which is not expected to be required. Even for a hardware implementation, the additional implementation complexity may be minimal because it is mainly a matter of accumulator registers and logic for each threshold, which are used to increment the accumulator and compare against the accumulator when performing the matching.
[0123] In addition, various column and row sortings can be applied to minimize the cost of matching, which includes grouping of rows such that certain matching operations only have to be performed on a subset of the table rows.
[0124] Traditionally, other unrelated work on "approximate matching" has focused on approximation stringMatching, where it is a matter of matching an input string of values / characters with a given internal string. Usually there is a table of such internal strings for matching, each of which can be considered as a row. In contrast to what is disclosed, the values / characters are not identified with columns, so there is no concept of "corresponding" entries between the two strings being matched. In the disclosed application, each value in the mask and row is instead associated with a column with semantic meaning in the form of a sub-condition, and a pairwise comparison is performed between the entries in the mask and row in the corresponding column.
[0125] Figure 4A , Figure 4B and Figure 4C is a flow chart illustrating an embodiment of a process for approximate matching. In one embodiment, Figure 1 System implementation Figure 4A , Figure 4B and Figure 4C For example, the process for approximate matching can be used in the medical diagnosis application mentioned above as a practical application.
[0126] exist Figure 4A In step (402), a first sub-condition set is obtained. In one embodiment, the first sub-condition set includes sub-conditions to be matched. Figure 4A In step (404), an approximate match of the first sub-condition set against the plurality of sub-condition sets is performed. The sub-conditions may be obtained for users and / or administrators of actual applications (e.g., nurses, doctors, and / or scientists associated with root cause analysis). In one embodiment, the matching described herein is a matching of database rows.
[0127] In one embodiment, Figure 4B The process used to achieve Figure 4A Steps (404): Figure 4B In step (422), a second sub-condition set among multiple sub-condition sets is accessed, wherein the representation of the sub-conditions in the first sub-condition set and / or the second sub-condition set includes a value having one of multiple states, and wherein the multiple states include a "don't care" state. In one embodiment, the second sub-condition set includes a rule corresponding to a row in a rule table, or a root cause corresponding to a row in a root cause table. In one embodiment, the value is represented by a single bit, and / or an ANDN hardware instruction is used to match two values. In one embodiment, the value is represented by two bits. In one embodiment, the multiple states further include a "true" state. In one embodiment, the second sub-condition set is among the multiple sub-condition sets being evaluated.
[0128] exist Figure 4B In step (424), the first sub-condition set approximately matches the second sub-condition set.
[0129] In one embodiment, Figure 4C the process is Figure 4B part of step (424): In Figure 4C step (442), a first set of sub-conditions is compared against a second set of sub-conditions. In one embodiment, the comparison includes an exact comparison. In one embodiment, comparing the first set of sub-conditions against the second set of sub-conditions includes comparing entry by entry, row by row, and / or includes comparing block by block.
[0130] In Figure 4C step (444), based on the result of the match, the match criteria are determined to be satisfied. In one embodiment, the match criteria can be satisfied in the case of an exact match such that an exact match will result in an approximate match outcome. In a step not shown, the approximate match is repeated for other rules among multiple rules. In one embodiment, an approximate match of a mask to a row is indicated as the result of a specified matching function.
[0131] In one embodiment, an approximate match is found when a row is below a distance threshold from the mask. In one embodiment, an approximate match is found when a row has at most N mismatched entries. In one embodiment, an approximate match is found when a row has a match strength equal to or greater than a specified strength threshold. In one embodiment, the approximate match is based on a parameter depending on the column and / or row of the entry being matched. In one embodiment, the columns and / or rows are organized to allow a simple mapping from the parameter to the mapped entry. In one embodiment, the approximate match includes applying a dedicated SPVM that calculates a metric corresponding to the distance or strength using a multiplication operator.
[0132] In Figure 4A step (406), information indicating that the second set of sub-conditions is at least an approximate match of the first set of sub-conditions is output.
[0133] Although some detailed descriptions have been made of the foregoing embodiments for the purpose of clear understanding, the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative rather than restrictive.
Claims
1. A method for approximate matching, comprising: Receiving a table of symptoms for root cause analysis of a monitored system, wherein the table is a sparse matrix; Receiving a query related to root cause analysis of a monitored system, including obtaining a set of masked sub-conditions associated with a current set of symptoms in the monitored system, wherein the representation of each sub-condition includes a value having one of three or more states, and wherein the three or more states include an "irrelevant" state, and wherein the obtained set of masked sub-conditions is associated with the query; Performing approximate matching of the obtained set of masked sub-conditions against a plurality of sets of sub-conditions, wherein the plurality of sets of sub-conditions are represented by rows in the table, including: Accessing a row sub-condition set among the plurality of sets of sub-conditions, wherein the row sub-condition set is associated with a row in the table; Wherein the row sub-condition set includes a root cause; Determining a matching function value at least in part based on such steps: determining a matching operator value for each specific sub-condition of the row sub-condition set of the obtained masked set and its corresponding masked sub-condition at least in part based on a non-commutative matching operation; Wherein determining the matching operator value includes: using a non-commutative matching operation to match a given specific sub-condition as "irrelevant" or match it to its corresponding masked sub-condition in a first case; Wherein determining the matching operator value includes: applying a sparse matrix-vector multiplication SpMV and a multiplication operator to calculate the matching operator value; Comparing the matching function value with a matching threshold; and In a second case where the matching function value is within the matching threshold, outputting information indicating that the row sub-condition set approximately matches the obtained set of masked sub-conditions; and Responding to the query at least in part with a root cause associated with the approximate matching as a possible root cause failure of the monitored system.
2. The method according to claim 1, wherein the set of masked sub-conditions includes sub-conditions to be matched.
3. The method according to claim 1, wherein the row sub-condition set includes a rule corresponding to a specified row in a rule table, or a root cause corresponding to a given row in a root cause table.
4. The method according to claim 1, wherein the state is represented by a single bit.
5. The method according to claim 1, wherein, The state is represented by a single bit, and an ANDN hardware instruction is used to match two values.
6. The method according to claim 1, wherein the state is represented by two bits.
7. The method according to claim 1, wherein comparing the set of masked sub-conditions against the row sub-condition set includes comparing entry by entry and / or includes comparing block by block.
8. The method according to claim 1, wherein the matching criterion can be satisfied in the case of an exact match, such that the exact match will produce a result of an approximate match.
9. The method according to claim 1, wherein the three or more states further include a "true" state.
10. The method according to claim 1, wherein the approximate match of the set of masked sub-conditions and the row sub-condition set is indicated as the result of a specified matching function.
11. The method according to claim 1, wherein an approximate match is determined based on a row sub-condition set below a distance threshold from the set of masked sub-conditions.
12. The method according to claim 11, wherein the distance threshold is a percentage.
13. The method according to claim 1, wherein an approximate match is determined based on at most N mismatched entries, where N is a natural number.
14. The method according to claim 1, wherein an approximate match is determined based on a set of masked sub-conditions having a match strength with a set of masked sub-conditions equal to or greater than a specified strength threshold.
15. The method according to claim 14, wherein the specified strength threshold is a percentage.
16. The method according to claim 1, wherein the approximate match is based on parameters depending on columns and rows of the entry to be matched.
17. The method according to claim 1, wherein The columns and / or rows of the table are organized to allow a simple mapping from the parameters to the mapped entries.
18. The method according to claim 1, wherein the SpMV multiplication includes an implementation using general-purpose computing on graphics processing units GPGPU.
19. A system for approximate matching, comprising: a processor configured to: receive a table of symptoms for root cause analysis of a monitored system, wherein the table is a sparse matrix; receive a query related to root cause analysis of the monitored system, including obtaining a set of masked sub-conditions associated with a current set of symptoms in the monitored system, wherein the representation of each sub-condition includes a value having one of three or more states, and wherein the three or more states include an "irrelevant" state, and wherein the obtained set of masked sub-conditions is associated with the query; perform an approximate match of the obtained set of masked sub-conditions against a plurality of sets of sub-conditions, wherein the plurality of sets of sub-conditions are represented by rows in the table, including: access a row sub-condition set among the plurality of sets of sub-conditions, wherein the row sub-condition set is associated with a row in the table; wherein the row sub-condition set includes a root cause; determine a match function value at least in part based on the step of: determining a match operator value for each specific sub-condition of the row sub-condition set of the obtained masked set and its corresponding masked sub-condition at least in part based on a non-commutative matching operation; wherein determining the match operator value includes: using a non-commutative matching operation to match a given specific sub-condition as "irrelevant" or match it with its corresponding masked sub-condition in a first case; wherein determining the match operator value includes: applying a sparse matrix-vector multiplication SpMV and a multiplication operator to calculate the match operator value; compare the match function value with a match threshold; and in a second case where the match function value is within the match threshold, output information indicating an approximate match of the row sub-condition set with the obtained set of masked sub-conditions; and respond to the query at least in part with a root cause associated with the approximate match as a possible root cause failure of the monitored system.
20. A computer program product embodied in a non-transitory computer-readable storage medium and including computer instructions for: receiving a table of symptoms for root cause analysis of a monitored system, wherein the table is a sparse matrix; Receives a query related to root cause analysis of a monitored system, including obtaining a set of masked sub-conditions associated with a current set of symptoms in the monitored system, where the representation of each sub-condition includes a value having one of three or more states, and where the three or more states include an "irrelevant" state, and where the obtained set of masked sub-conditions is associated with the query; Performs approximate matching of the obtained set of masked sub-conditions against a plurality of sets of sub-conditions, where the plurality of sets of sub-conditions are represented by rows in a table, including: Accesses a row sub-condition set among the plurality of sets of sub-conditions, where the row sub-condition set is associated with a row in the table; where the row sub-condition set includes a root cause; Determines a matching function value at least in part based on such steps: determining a matching operator value for each specific sub-condition of the row sub-condition set of the obtained set of masks and its corresponding masked sub-condition at least in part based on a non-commutative matching operation; where determining the matching operator value includes: using a non-commutative matching operation to match a given specific sub-condition as "irrelevant" or to its corresponding masked sub-condition in a first case; where determining the matching operator value includes: applying a sparse matrix-vector multiplication SpMV and a multiplication operator to calculate the matching operator value; Compares the matching function value with a matching threshold; and In a second case where the matching function value is within the matching threshold, outputs information indicating that the row sub-condition set approximately matches the obtained set of masked sub-conditions; and Responds to the query at least in part with the root cause associated with the approximate match as a possible root cause failure of the monitored system.
Citation Information
Patent Citations
Methods and systems for predicting electromagnetic scattering
US20050071097A1
Dynamic bitmap processing, identification and reusability
US20050154710A1
Answer Support System and Method
US20110161274A1
System and method for fingerprinting datasets
US20130259211A1
Time-based method of human-computer interaction for controlling storage and retrieval of multimedia information
US5907320A