Identify the causes of anomalies observed in integrated circuit chips
By integrating monitoring circuit devices in the SoC, the reasons for abnormal characteristics are identified by using the difference in feature probability distribution, the problem that monitoring tools in the SoC are difficult to directly access the system bus, and the monitoring capabilities of system security vulnerabilities and defects are improved.
Patent Information
- Application Number
- CN202080083699.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-03
- Filing Date
- 2020-11-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-11-26
AI Technical Summary
In a system-on-chip (SoC), monitoring tools have difficulty directly accessing the system bus, resulting in reduced ability of external monitoring tools to monitor security vulnerabilities, defects, and security concerns of the system within the time frame required by the industry.
By implementing a monitoring circuit device on an integrated circuit (IC) chip, the system circuit device is monitored by measuring the system circuit device characteristics in a series of windows, and the cause of abnormal features is identified by identifying the difference in the feature probability distribution in the candidate window set.
A more detailed analysis of abnormal features measured by the circuit device of the SoC system is realized, and the causes of abnormal features can be identified, thereby improving the monitoring ability of system security vulnerabilities and defects.
Smart Images

Figure CN114761928B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the analysis of characteristics measured from system circuitry in a system on chip (SoC) or a multi-chip module (MCM). Background Art
[0002] In the past, embedded systems with multiple core devices (processors, memory, etc.) were integrated onto a printed circuit board (PCB) and connected via buses on the PCB. The business in the embedded system is delivered via these buses. This arrangement is convenient for monitoring core devices, because monitoring tools such as oscilloscopes and logic analyzers can be attached to the buses of the PCB to allow direct access to the core devices.
[0003] The market demand for smaller products coupled with advances in semiconductor technology has led to the development of System-on-Chip (SoC) devices. In an SoC, multiple core devices of an embedded system are integrated onto a single chip. In an SoC, the traffic in the embedded system is delivered via an internal bus, so monitoring tools can no longer connect directly to the system bus. The resulting reduction in access coupled with the increased volume of data moving around the chip (due to the integration of multiple processing cores and higher internal clock frequencies resulting from advances in SoC technology) reduces the ability of external monitoring tools to monitor the system for security vulnerabilities, defects, and security concerns within the timeframes required by the industry. In addition, when multiple core devices are embedded onto the same single chip, each individual core device behaves differently than it would in isolation due to its interaction with other core devices and real-time events such as triggers and alarms.
[0004] Therefore, the development of SoC devices requires associated developments in monitoring technology so that some monitoring functions are integrated into the SoC. It is now known that for monitoring circuitry within the SoC to track the output of the processor executing the program on the core device (such as the CPU). The trace data is usually exported for off-chip analysis.
[0005] It would be desirable to be able to perform more detailed analysis of the data collected by the on-chip monitoring circuitry, particularly to investigate anomalies in the data. Summary of the invention
[0006] According to a first aspect, a method for identifying the cause of an abnormal feature measured from a system circuit device on an integrated circuit (IC) chip is provided, the IC chip comprising a system circuit device and a monitoring circuit device, the monitoring circuit device being used to monitor the system circuit device by measuring the feature of the system circuit device in each window of a series of windows, the method comprising: (i) identifying a candidate window set from a window set preceding an abnormal window including the abnormal feature, so as to search for the cause of the abnormal feature in the candidate window set; (ii) for each measured feature of the system circuit device: (a) calculating a first feature probability distribution of the measured feature for the candidate window set; (b) calculating a second feature probability distribution of the measured feature for a window not in the candidate window set; (c) comparing the first and second feature probability distributions; and (d) if the first and second feature probability distributions differ by more than a threshold, identifying the measured feature within the time range of the candidate window set as the cause of the abnormal feature; (iii) iterating steps (i) and (ii) for other candidate window sets from the window set preceding the abnormal window; and (iv) outputting a signal indicating the measured feature identified as the cause of the abnormal feature in steps (ii)(d).
[0007] Step (ii)(c) may include determining a difference measure between the first feature probability distribution and the second feature probability distribution; and step (ii)(d) may include identifying that the measured feature within the time range of the candidate window set is the cause of the abnormal feature if the difference measure is greater than the threshold.
[0008] The difference measure may be scaled by the percentile of the difference over time between the first and second feature probability distributions of the iteration.
[0009] The set of windows preceding the anomaly window can be bounded by (i) the anomaly window and (ii) the distal earlier windows.
[0010] Step (ii)(b) may include calculating a second feature probability distribution of the measured features for a set of windows between the set of candidate windows and the abnormal windows.
[0011] The candidate window set may include fewer than 10 windows.
[0012] The candidate window set may include only one window.
[0013] The first and second feature probability distributions may be calculated in steps (ii)(a) and (ii)(b) by fitting a Gaussian model to the measured features of the identified window.
[0014] The method may also include identifying a measured feature affected by the abnormal feature, wherein the affected measured feature is located in a window after the abnormal window, the method comprising: (v) identifying a subsequent candidate window set from the window set after the abnormal window, so as to search for the influence of the abnormal feature in the subsequent candidate window set; (vi) for each of the measured features of the system circuit device: (a) calculating a third feature probability distribution of the measured feature for the subsequent candidate window set; (b) calculating a fourth feature probability distribution of the measured feature for the subsequent window not in the subsequent candidate window set; (c) comparing the third and fourth feature probability distributions; and (d) if the third and fourth feature probability distributions differ by more than another threshold, identifying the measured feature within the time range of the subsequent candidate window set as being affected by the abnormal feature; and (vii) iterating steps (v) and (vi) for other subsequent candidate window sets of the window set after the abnormal window; and (viii) outputting a signal indicating the measured feature identified as being affected by the abnormal feature in step (vi)(d).
[0015] Step (vi)(c) may include determining other difference measures between the third feature probability distribution and the fourth feature probability distribution; and step (vi)(d) may include identifying that the measured features within the time range of the subsequent candidate window set are affected by abnormal features if the other difference measures are greater than other thresholds.
[0016] Other difference measures may be the time-scaled difference between the third and fourth feature probability distributions.
[0017] The set of windows following the anomaly window can be bounded by (i) the anomaly window and (ii) the distal late window.
[0018] Step (vi)(b) may include calculating a fourth feature probability distribution of the measured features for the set of windows between the subsequent candidate windows and the abnormal windows.
[0019] The subsequent candidate window set may include fewer than 10 windows.
[0020] The subsequent candidate window set may include only one window.
[0021] The third and fourth feature probability distributions may be calculated in steps (vi)(a) and (vi)(b) by fitting a Gaussian model to the measured features of the identified windows.
[0022] Measured characteristics may include those derived from trace data generated by monitoring circuitry from data output by components of the system circuitry.
[0023] Measured characteristics may include those characteristics derived from matching events identified by the monitoring circuitry from data input to or output from components of the system circuitry.
[0024] The measured characteristic may include a characteristic derived from a counter of the monitoring circuitry, the counter being configured to count each time a particular item is observed from a component of the system circuitry. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The present invention will now be described by way of example with reference to the accompanying drawings.
[0026] Figure 1 is a schematic diagram of an exemplary integrated circuit chip device;
[0027] Figure 2 is a schematic diagram of an exemplary monitoring network and system circuit arrangement on an integrated circuit chip device;
[0028] Figure 3 is a flow chart of a method for identifying a cause of an abnormal characteristic measured from a system circuit device;
[0029] Figure 4 The time window in which the characteristic sequence was measured is shown;
[0030] Figure 5 is a diagram that describes the features that are the cause of subsequent abnormal features;
[0031] Figure 6 is a flow chart of a method of identifying a subsequent measured feature affected by an abnormal feature measured from a system circuit device;
[0032] Figure 7 The time window over which the characteristic sequence was measured is shown.
[0033] Figure 8 is a diagram that describes the features that are the cause of the subsequent abnormal features and the features that are affected by the abnormal features;
[0034] Figure 9a , Figure 9b , Fig.9c and Figure 9d It is a graph describing the features that are the causes of subsequent abnormal features and the features affected by the abnormal features for candidate window sets of different lengths. DETAILED DESCRIPTION
[0035] The following disclosure describes a monitoring architecture suitable for implementation on an integrated circuit chip. The integrated circuit chip may be a SoC or a multi-chip module (MCM).
[0036] Figure 1 and Figure 2Schematic diagrams of exemplary system architectures and components within the system architecture. These figures present the structure in the form of functional blocks. Some functional blocks for performing functions known in the art are omitted in these figures. Figure 3 and Figure 6 The present invention is a flowchart that illustrates a method of analyzing statistical data measured by a monitoring circuit device. Each flowchart describes an order in which the method of the flowchart can be implemented. However, the flowcharts are not intended to limit the described methods to be implemented in the described order. The steps of the methods can be performed in an alternative order to the order described in the flowcharts.
[0037] Figure 1 The general structure of an exemplary monitoring network of SoC 100 is shown. Monitoring circuitry 101 is arranged to monitor system circuitry 102. For example, to detect improper operation of core devices related to safety or security concerns.
[0038] Figure 2 An exemplary system circuit arrangement is shown, which includes core devices 201, 202 connected via a SoC interconnect 203. Core devices 201a, 201b, 201c are master devices. Core devices 202a, 202b, 202c are slave devices. Any number of core devices may be suitably integrated into the system circuit arrangement, such as Figure 2 The master devices and slave devices are numbered 1, 2, ..., N. The SoC interconnect 203 forms the communication backbone of the SoC, through which the master devices and slave devices communicate with each other. These communications are bidirectional.
[0039] A master device is a device that initiates a service, such as a read / write request in a network. Examples of master devices are processors, such as DSPs (digital signal processors), video processors, application processors, CPUs (central processor units), and GPUs (graphics processor units). Any programmable processor can be a master device. Other examples of master devices are devices with DMA (direct memory access) capabilities, such as traditional DMA for moving data from one location to another, autonomous coprocessors with DMA capabilities (such as encryption engines), and peripheral devices with DMA capabilities (such as Ethernet controllers).
[0040] A slave device is a device that responds to commands from a master device. Examples of slave devices are on-chip memory, memory controllers for off-chip memory (such as DRAM), and peripheral units.
[0041] The topology of SoC interconnect 203 depends on the SoC. For example, it may include any one or combination of the following network types for routing communications around system circuitry: a bus network, a ring network, a tree network, or a mesh network.
[0042] The monitoring circuit arrangement 101 comprises monitoring units 204 a , 204 b which are connected to a communicator 206 via a monitoring interconnect 205 .
[0043] Any number of monitoring units can be integrated in the monitoring circuit device. Each monitoring unit is connected to the communication link between the master device and the slave device. This connection can be between the master device and the SoC interconnect, such as at the interface between the master device and the SoC interconnect. The connection can be between the SoC interconnect and the slave device, such as at the interface between the slave device and the SoC interconnect. Each monitoring unit can be connected to a separate communication link. Alternatively, one or more monitoring units of the monitoring circuit device 101 can be connected to multiple communication links. The monitoring unit 204 monitors the operation of the core device by monitoring the communication on the monitored communication link. Optionally, the monitoring units may also be able to manipulate the operation of the core devices they monitor.
[0044] The communicator 206 may be an interface for communicating with an off-chip entity. For example, the monitoring circuit device 101 may communicate with an off-chip analyzer via the communicator 206. Additionally or alternatively, the communicator 206 may be configured to communicate with other on-chip entities. For example, the monitoring circuit device 101 may communicate with an on-chip analyzer via the communicator 206. Figure 2 One communicator 206 is shown, but any number of communicators may be integrated onto the SoC. The communicator implemented is selected based on the type of connection to be made. Exemplary communicators include: JTAG, parallel trace input / output, and Aurora-based high-speed serial interfaces; as well as reuse of system interfaces such as USB, Ethernet, RS232, PCIe, and CAN.
[0045] The topology of monitoring interconnect 205 may include any one or combination of the following network types for transmitting communications around monitoring circuitry: a bus network, a ring network, a tree network, or a mesh network. The communication link between monitoring unit 204 and communicator 206 is bidirectional.
[0046] As mentioned above, Figure 2 The monitoring unit 204 monitors the communication between the master device 201 and the slave device 202. The monitoring units can collect statistical data from the communications they monitor. The collected statistical data and the time window for completing the statistical collection are configurable. For example, each monitoring unit 204 can receive a configuration command from an on-chip or off-chip analyzer, which commands the monitoring unit to monitor specific communication parameters. The analyzer can also specify the length of the time window for monitoring parameters. The analyzer can also specify when the collected data is reported to the analyzer. Typically, the analyzer requires regular reporting of the collected data.
[0047] Therefore, monitoring unit 204 can be configured to monitor the communication of its connected components (whether master device 201 or slave device 202) over a series of monitoring time windows. The length of each monitoring time window can be specified by the above-mentioned analyzer. The monitoring time windows can be non-overlapping. For example, the monitoring time windows can be continuous. Alternatively, the monitoring time windows can be overlapping.
[0048] Examples of data that may be generated by a monitoring unit observing one or more components of a system circuit arrangement include:
[0049] - Trace data. The trace data generated may be a copy of the data observed by the monitoring unit. For example, a copy of a sequence of instructions executed by a CPU, or a set of transactions on a bus.
[0050] - Matching data. The monitoring unit may be configured to monitor the system circuit device for the occurrence of a specific event. When the specific event is identified, the monitoring unit generates matching data. The monitoring unit may immediately output the matching data to the analyzer.
[0051] - Counter data. The monitoring unit may include one or more counters. Each counter is configured to count the occurrence of a specific event. The count value of the counter may be periodically output to the analyzer.
[0052] The raw data generated by the monitoring unit is suitably converted into a set of measured features for each of a series of time windows. Each measured feature has a value for each window.
[0053] Examples of measured characteristics include:
[0054] - Aggregate bandwidth captured from the bus. This can be split into the aggregate bandwidth for read operations, and a separate aggregate bandwidth for write operations.
[0055] - Maximum latency, minimum latency, and / or average latency of read operations captured from the bus.
[0056] -The number of address match events. In other words, the number of accesses to the selected memory area.
[0057] - Based on software execution traces, in each individual thread: (i) the cumulative time spent in the thread; and / or (ii) the minimum, maximum, and / or average thread interval time; and / or (iii) the number of thread scheduling events, optionally specifying which thread to take over from.
[0058] - Based on software execution trace, the number of interrupts, and / or the minimum, maximum, and / or average time spent in interrupt handling.
[0059] - Based on the CPU instruction trace, the number of instructions executed, optionally grouped by instruction categories which may include branches.
[0060] The conversion of raw data to measured features may be performed by any method known in the art. Such conversion may be performed by the on-chip monitoring circuit device 101. Alternatively, the conversion may be performed by an analyzer, which may be on-chip or off-chip. In generating the measured features, data obtained from resources other than the monitoring unit 204 may be used in conjunction with the raw data generated by the monitoring unit. The time window for aggregating the measured features may have a length between 1 ms and 1000 ms. The time window for aggregating the measured features may have a length between 10 ms and 100 ms.
[0061] The measured features in the series of time windows may then be input into an anomaly detection method to identify whether any of the measured features is abnormal. The anomaly detection method may be performed by an analyzer. However, alternatively, the anomaly detection method may be performed by the monitoring circuit device 101.
[0062] In a first example, anomaly detection is performed using a model trained from a known good sequence. In this example, the model captures the behavior of a series of time windows whose measured features are known not to be abnormal. Building the model includes constructing a feature distribution for each feature. For example, a kernel density estimator (KDE) can be used to build the distribution. KDE starts with a flat zero line and adds a small Gaussian kernel to the value of each feature from each window in a series of time windows. Each feature value contributes the same amount to the distribution. The final value can then be scaled. The result is a feature distribution that indicates the likelihood that a particular value of the feature represents normal behavior. Therefore, the model includes a set of feature distributions that represent normal behavior for these features.
[0063] A subsequent sequence can then be compared to the model. The subsequent sequence includes a series of time windows whose measured features are unknown to be abnormal or normal. The subsequent sequence can be compared to the model by comparing a single window of the subsequent sequence to the model. In this case, the value of the model feature distribution corresponding to the value of the feature in the single window is determined. If the value of the distribution indicates that there is a low probability that the feature value is normal behavior, then the feature is determined to be abnormal in the single window. For example, if the distribution value is below a threshold, then the feature is determined to be abnormal in the single window. The threshold can be different for different features.
[0064] The abnormal features are output as electrical signals to the user (e.g., as visual signals on a screen). If two or more abnormal features are identified, then these features can be ranked in the output signal. The abnormal features can be ranked in the order of their values below the threshold, with the abnormal feature farthest below the threshold being ranked first and the abnormal feature closest below the threshold being ranked last.
[0065] Subsequent sequences can be compared to the model by first constructing a feature distribution for each feature of the subsequent sequence. For example, each feature distribution can be generated using KDE, as described above with respect to model generation. The difference between the model feature distribution and the subsequent sequence feature distribution is then obtained for each feature. If the average difference between the two feature distributions for a feature is greater than a threshold, then the feature in the subsequent sequence is determined to contain an anomaly.
[0066] The abnormal features are output as electrical signals to the user (e.g., as visual signals on a screen). If two or more abnormal features are determined, then the features can be ranked in the output signal. The abnormal features can be ranked in order of their average difference between the model feature distribution and the subsequent sequence feature distribution, with the abnormal feature with the largest average difference ranked first and the abnormal feature with the smallest average difference ranked last.
[0067] In a second example, anomaly detection is performed without utilizing the behavior of a series of time windows whose measured features are known to be non-abnormal. Anomaly detection is performed on a sequence of time windows containing a series of measured features. These measured features are not known to be abnormal or normal. The example includes constructing a feature distribution for each feature of the sequence. This can be performed using KDE as described above for the first example. The lowest value in the feature distribution of each feature is identified as a potential anomaly. These potential abnormal features are output to the user as electrical signals (e.g., as visual signals on a screen). The user can reject the identified features as not abnormal, or accept the identified features as abnormal. The user can also manually mark other features as abnormal.
[0068] The output abnormal features can be grouped into abnormal windows. The abnormal windows can be sorted in order of their likelihood across all features, with the abnormal window with the lowest likelihood of representing normal behavior across all features being ranked first and the abnormal window with the highest likelihood of representing normal behavior across all features being ranked last.
[0069] Suitably, multiple iterations are performed on the selected anomaly detection method, each iteration using a different time window length. For example, a time window length range from 10ms to 100ms can be utilized in the iteration. This can enable anomalies caused by temporary properties that are more easily observed within a specific window length to be identified. In each iteration, the time windows may be non-overlapping. For example, the time windows may be continuous. Alternatively, the time windows may be overlapping.
[0070] Now, refer to Figure 3 A method is described for identifying the causes of anomalous features in component activity on a SoC. These anomalous features have been identified using an anomaly detection method (one of the methods described above). Figure 3 The method described is performed on a processor. Suitably, the processor is located at the analyzer. The analyzer may be on the SoC 100. Alternatively, the analyzer may be an off-chip component. Alternatively, the processor may also be located in the monitoring circuit device 101 of the SoC.
[0071] The processor receives as input the sequence of measured features. The processor also receives as input one or more time windows, which are identified as having at least one abnormal feature therein. Optionally, the abnormal features themselves may also be identified. The processor uses these inputs to search for possible causes of the abnormal features in time windows preceding the abnormal window.
[0072] In step 301, the processor selects a candidate window set j and searches for the cause of the abnormal feature in the candidate window set. For each abnormal window, the processor selects one or more windows to add to the candidate window set j. For each abnormal window, the processor selects a window to be added to the candidate window set j from the window set before the abnormal window in the sequence of measured features and the abnormal window. Figure 4 An example sequence of measured features is shown. An abnormal window is labeled 401. For the abnormal window, a window is selected from the windows labeled 402 to be added by the processor to the candidate window set j. The window set 402 is defined by: (i) an abnormal window 404, and (ii) a far early window 405. The far early window is the earliest window in which the cause of the abnormal feature is to be searched. For ease of illustration, only 10 windows are shown before the abnormal window 401. In practice, the window set 402 may include up to 1000 windows. For example, the window set 402 may include 100 windows. The window Cj added to the candidate window set j for the abnormal window 401 is labeled 403. The length of the window Cj added to the candidate window set j is configurable. The length of the window Cj may be only a single window. Alternatively, the length of the window Cj may include two or more windows. The length of the window Cj may include up to 10 windows. In Figure 4In the example of FIG. 1 , the length of window Cj is shown to include three windows: windows 4, 5, and 6. The processor selects window Cj to add to the candidate window set j for each abnormal window in the measured feature sequence. For example, in this iteration, the processor can add three windows to the candidate window set j for each abnormal window, which are consecutive windows that are 4 to 6 windows back from the abnormal window.
[0073] After step 301, the processor goes to step 302. In step 302, for the measured feature i, the processor calculates the first feature probability distribution PD1 of the measured feature i for the candidate window set j.
[0074] At step 303, for each measured feature i, the processor calculates a second feature probability distribution PD2 for the measured feature i for windows in the sequence but not in the candidate window set j. The second feature probability distribution PD2 may be calculated for a window set including all windows 402 not in the candidate window set j.
[0075] Steps 302 and 303 may be performed simultaneously. Alternatively, Figure 3 As shown, step 302 may precede step 303. Alternatively, step 303 may precede step 302.
[0076] The first and second feature probability distributions can be calculated by the processor using the above-mentioned KDE method to the identified window of the measured feature sequence. Alternatively, the first and second feature probability distributions can be calculated by the processor by fitting a Gaussian mixture model to the identified window of the measured feature sequence. Before identifying other abnormal abnormal windows therein, there may be only a small amount of windows. The distribution generated by the Gaussian mixture model is simpler than the KDE model and is more effective for fewer data points, so it may be more preferred here. In addition, different models known in the art can be used to generate the first and second feature probability distributions.
[0077] After calculating the first and second feature probability distributions in steps 302 and 303, the processor compares the two distributions in step 304. A large difference between the distributions indicates that the feature is the cause or contributor to the anomaly observed in the anomaly window. Therefore, the processor determines whether the difference between the first and second feature probability distributions exceeds a threshold value Vt. In step 304, if the first and second feature probability distributions PD1 and PD2 differ by more than the threshold value Vt, the processor proceeds to step 305, in which the processor identifies feature i in candidate window set j as the cause of the abnormal feature in the anomaly window. If, in step 304, the difference between the first and second feature probability distributions PD1 and PD2 is less than the threshold value Vt, the processor does not identify feature i in candidate window set j as the cause of the abnormal feature in the anomaly window.
[0078] In order to evaluate whether the difference between the first and second feature probability distributions exceeds a threshold, the processor can determine a difference metric between the two probability distributions. The difference metric is a single value. The single value can represent the average difference between the probability distributions. In other words, the average difference between the number of features observed at each feature value in the two distributions. Alternatively, the single value can represent the total difference between the probability distributions. In other words, the total difference between the number of features observed at each feature value in the two distributions. The difference metric can be calculated by any method known in the art. Then, in step 304, the difference metric |PD1-PD2| is compared with a threshold value Vt.
[0079] The processor then proceeds to step 306. In step 306, the processor determines whether there are Figure 3 The method of the present invention is applied to other measured features of the candidate window set j that have not been applied. If there are still measured features, the processor proceeds to step 307, in which the next measured feature is selected. The processor then repeats steps 302 to 306 for the next measured feature for the candidate window set j. In step 306, if the processor determines that there are no other measured features, it proceeds to step 308.
[0080] In step 308, the processor determines whether there are Figure 3 The method has not been applied to the candidate window set. The next candidate window set j+1 may overlap with the candidate window set j. For example, for the next candidate window set j+1, for each abnormal window, the processor may select one or more different windows to add to the candidate window set j+1 compared to the window selected to be added to the candidate window set j. As with the candidate window set j, for each abnormal window, the window added to the candidate window set j+1 is selected from the abnormal window and the window set that precedes the abnormal window in the measured feature sequence. Figure 4 , the window Cj+1 added by the processor to the candidate window set j+1 for the abnormal window 401 is marked as 408, which only includes windows 5, 6 and 7. In this example, windows 5 and 6 of the candidate window set j+1 overlap with the candidate window set j. Alternatively, the next candidate window set j+1 may not overlap with the candidate window set j, but be adjacent to it. For example, referring to Figure 4 , the window Cj+1 added by the processor to the candidate window set j+1 for the abnormal window 401 is marked as 409, which includes windows 7, 8 and 9. In this example, the candidate window set j+1 and the candidate window set j selected for the abnormal window 401 have no overlapping windows. Once the iteration reaches the far early window 405, there are no other candidate window sets.
[0081] If it is determined at step 308 that there are more candidate window sets, the processor proceeds to step 309, where the next candidate window set is selected. The processor then repeats steps 302 to 308 for the next candidate window set. If it is determined at step 308 that there are no more candidate window sets, the processor proceeds to step 310, where the identified cause (if any) of the abnormal feature of the abnormal window is output.
[0082] At step 310, the cause of the abnormal characteristic may be output to the user as an electrical signal (e.g., as a visual signal on the analyzer screen). For example, the cause of the abnormal characteristic may be output to the screen as Figure 5 The diagram shown in . Figure 5 The diagram of FIG. 1 shows a plurality of graphs. Each graph plots a scaled difference measure between a first and a second feature distribution of a measured feature (on the y-axis) versus the number of windows backed off from the anomaly window (on the x-axis). The number of windows backed off in time can be used as the window of the candidate window set that is closest to the anomaly window. For example, see Figure 4 , the scaled difference measure between the first and second feature probability distributions for candidate window set j is marked on the x-axis at three windows back in time.
[0083] The difference measures between different measured features are not consistent. For example, the accumulation time may be consistently more variable than the memory saturation. Figure 5 The difference measures of all measured features are plotted on a graph, and the difference measures are scaled so that they are comparable. Thus, the difference measures provide a relative measure for different measured features. For example, the difference measures can be scaled by the percentile of their differences over time. For example, the difference measures can be scaled by the 50th percentile.
[0084] By plotting the scaled difference metric over time offsets from the anomaly window, the measured feature that is the cause of the anomaly feature is readily apparent to the user. A large scaled difference for a measured feature over a particular number of windows back in time indicates a high probability that the cause occurred in that measured feature over that number of windows back in time.
[0085] Figure 5 The diagram shows the execution of software thread switching data Figure 3 The Gaussian mixture model (GMM) is used to generate the first and second feature probability distributions respectively. The candidate window set has a single window. The relative GMM distribution difference is used as the difference measure. The figure shows that the rt threads that are temporally close to the abnormal window are the cause of the abnormal feature. This is reflected in the fact that the measured difference for the maximum and minimum rt times is significantly higher than the measured difference for other features in the window range of 0 to 1.
[0086] With reference Figure 3The method corresponding to the method described can also be applied to the window after the abnormal window in the measured feature sequence. Figure 6 , and can be used to identify subsequent measured features that are affected by the abnormal feature in the abnormal window. Figure 6 The method is to execute Figure 3 The method is executed on the same processor.
[0087] and Figure 3 Same, for Figure 6 The method comprises a processor receiving as input a sequence of measured features and one or more time windows identified as having at least one abnormal feature therein. The processor uses these inputs to search for subsequent features affected by the abnormal feature in a time window following the abnormal window.
[0088] In step 601, the processor selects a subsequent candidate window set k in which measured features affected by abnormal features are searched. For each abnormal window, the processor selects one or more windows to add to the subsequent candidate window set k. For each abnormal window, the windows added to the subsequent candidate window set k are selected from the abnormal window and the window set that is located after the abnormal window in the measured feature sequence. Figure 7 An example sequence of measured features is shown. An anomaly window is labeled 401. For the anomaly window, the window added by the processor to the subsequent candidate window set j is selected from the windows labeled 701. The window set 701 is defined by: (i) an anomaly window 702, and (ii) a distal late window 703. The distal late window 703 is the latest window in which the influence of the anomaly feature is to be searched. For ease of illustration, only 10 windows are shown after the anomaly window 401. In practice, the window set 701 may include up to 1000 windows. For example, the window set 701 may include 100 windows. The window Ck added to the subsequent candidate window set k for the anomaly window 401 is labeled 704. The length of the window Ck added to the subsequent candidate window set k is configurable. The length of the window Ck may be a single window. Alternatively, the length of the window Ck may include two or more windows. The length of the window Ck may include up to 10 windows. In Figure 7 In the example of FIG. 1 , the length of window Ck is shown to include three windows: windows 5, 6, and 7. The processor selects window Ck to add to the subsequent candidate window set k for each abnormal window in the measured feature sequence. For example, in this iteration, the processor can add three windows to the candidate window set k for each abnormal window, which are consecutive windows that are 5 to 7 windows ahead of the abnormal window.
[0089] After step 601, the processor goes to step 602. In step 602, for the measured feature I, the processor calculates a third feature probability distribution PD3 of the measured feature I for the subsequent candidate window set k.
[0090] At step 603, for each measured feature I, the processor calculates a fourth feature probability distribution PD4 of the measured feature I for windows in the sequence but not in the subsequent candidate window set k. The fourth feature probability distribution PD4 may be calculated for a window set 704 including all windows 701 not in the subsequent candidate window set k.
[0091] Steps 602 and 603 may be performed simultaneously. Alternatively, Figure 3 As shown, step 602 may precede step 603. Alternatively, step 603 may precede step 602.
[0092] The third and fourth feature probability distributions may be calculated by the processor using any of the methods described above with respect to the first and second feature probability distributions.
[0093] After calculating the third and fourth feature probability distributions in steps 602 and 603, the processor compares the two distributions in step 604. A large difference between the distributions indicates that the feature is affected by the anomaly observed in the anomaly window. Therefore, the processor determines whether the difference between the third and fourth feature probability distributions exceeds a threshold value Vt'. In step 604, if the third and fourth feature probability distributions PD3 and PD4 differ by more than the threshold value Vt', the processor proceeds to step 605, where the processor identifies the feature I in the subsequent candidate window set k as being affected by the abnormal feature in the anomaly window. In step 604, if the third and fourth feature probability distributions PD3 and PD4 differ by less than the threshold value Vt', the processor does not identify the feature I in the subsequent candidate window set k as being affected by the abnormal feature in the anomaly window.
[0094] To evaluate whether the difference between the third and fourth feature probability distributions exceeds a threshold, the processor may determine a difference metric between the two probability distributions. The difference metric may be as described above with reference to Figure 3 The first and second characteristic probability distributions are calculated.
[0095] The processor then proceeds to step 606. In step 606, the processor determines whether there are Figure 6 The method of is applied to any measured features that have not been applied to the subsequent candidate window set k. If there are more measured features, the processor proceeds to step 607, in which the next measured feature is selected. The processor then repeats steps 602 to 606 for the next measured feature for the subsequent candidate window set k. If, at step 606, the processor determines that there are no more measured features, then the processor proceeds to step 608.
[0096] In step 608, the processor determines whether there are Figure 6 The subsequent candidate window set to which the method has not been applied. The next subsequent candidate window set k+1 may overlap with the subsequent candidate window set k. For example, for the next candidate window set k+1, for each abnormal window, the processor may select one or more different windows to add to the subsequent candidate window set k+1 compared to the window selected to be added to the subsequent candidate window set k. As with the subsequent candidate window set k, for each abnormal window, the window added to the subsequent candidate window set k+1 is selected from the abnormal window and the set of windows located after the abnormal window in the measured feature sequence. Figure 7 , the window Ck+1 added by the processor to the subsequent candidate window set k+1 for the abnormal window 401 is marked as 707, which only includes windows 6, 7 and 8. In this example, windows 6 and 7 of the subsequent candidate window set k+1 overlap with the subsequent candidate window set k. Alternatively, the subsequent candidate window set k+1 may not overlap with the subsequent candidate window set k, but be adjacent to it. For example, referring to Figure 7 , the window Ck+1 added by the processor to the subsequent candidate window set k+1 for the abnormal window 401 is marked as 708, which includes windows 8, 9 and 10. In this example, there are no overlapping windows between the subsequent candidate window set k+1 and the subsequent candidate window set k selected for the abnormal window 401. Once the iteration reaches the far late window 703, there are no other subsequent candidate window sets.
[0097] If it is determined at step 608 that there are more subsequent candidate window sets, the processor proceeds to step 609, where the next subsequent candidate window set is selected. The processor then repeats steps 602 to 608 for the next subsequent candidate window set. If it is determined at step 608 that there are no more subsequent candidate window sets, the processor proceeds to step 610, where the measured features identified as being affected by the abnormal features of the abnormal window are output.
[0098] At step 610, the affected measured characteristic may be output as an electrical signal to a user (e.g., as a visual signal on an analyzer screen). For example, a signal corresponding to Figure 5 A corresponding graph showing the scaled difference measure between the third and fourth feature distributions for the measured features (on the y-axis) versus the number of time windows forward in time from the anomaly window (on the x-axis). The number of windows forward in time can be used as the window closest to the anomaly window in the set of subsequent candidate windows. For example, see Figure 7 , the scaled difference measure between the third and fourth feature probability distributions for the subsequent candidate window set k is marked four windows forward in time on the x-axis. The difference measure can be referred to as Figure 3 and Figure 5Scaling is done in the same way as described.
[0099] Figure 8 is a graph showing the performance of a reference Figure 3 and Figure 6 The results of the two methods described are the measured characteristic sequences and the Figure 5 The measured characteristic sequence of the graph of is the same. These methods can be performed separately, as described above. Alternatively, these methods can be performed together as a single method, with Figure 3 The candidate window set in includes two windows before and after the abnormal window. Figure 8 In the diagram of , the feature type of the measured feature i is the same as that of the measured feature I. The length of the candidate window set j is the same as the length of the subsequent candidate window set k. Figure 5 As with , the feature probability distributions are all generated using a Gaussian mixture model, and the (subsequent) candidate window set has a single window. Figure 5 Same, Figure 8 The graph indicates that the RT threads close to the abnormal window time are the cause of the abnormal feature, and indicates that the RT threads after the abnormal window are also affected by the abnormal feature. This is reflected in the fact that the measured difference of the minimum RT time is significantly larger than the measured difference of other features in the -1 to 1 window range.
[0100] Figure 5 and Figure 8 All are referenced by the processor execution Figure 3 and Figure 6 The described methods are generated when using a (subsequent) candidate window set with only a single window. The number of windows in the (subsequent) candidate window set can be greater than one. This may make these methods less sensitive to the exact point in time when the error occurs. The implemented generation is repeated for (subsequent) candidate window sets with lengths of 2, 3, 5 and 10 windows. Figure 8 The results are shown in Figure 9a , Figure 9b , Fig.9c and Figure 9d These figures show that for the (subsequent) candidate window set of length 2 windows, the rt thread can be easily identified as the cause and the affected feature. For the (subsequent) candidate window sets of length 3 and 5 windows, the rt thread can be identified as the cause and the affected feature. However, in Figure 9d When the length of the (subsequent) candidate window set is 10 windows, the feature probability distribution of the (subsequent) candidate window set is too similar to the feature probability distribution of windows outside the (subsequent) candidate window set, and the rt thread cannot be identified as the cause and affected features.
[0101] Monitoring circuit devices on IC chips, such as Figure 2As shown, a huge amount of monitoring data can be generated. The methods described herein provide a method for analyzing such data to identify the cause of abnormal features measured from the system circuit device, as well as subsequent features affected by the abnormal features. These methods can be implemented in real time while the system circuit device continues to perform its functions. Alternatively, these methods can be implemented offline at a later time.
[0102] Anomaly detection can be applied to a wide range of fields such as financial, commercial, enterprise, industrial and engineering markets. Exemplary uses of the methods described herein are for security monitoring, such as fraud detection or intrusion detection, safety monitoring, preventive maintenance and performance monitoring of industrial equipment (such as sensors).
[0103] Figure 1 and Figure 2 Each component of the SoC shown in FIG. 1 may be implemented with dedicated hardware. Alternatively, Figure 1 and Figure 2 Each component of the SoC shown in FIG. 1 may be implemented by software. Some components may be implemented by software, while other components may be implemented by dedicated hardware.
[0104] The described SoC is suitably integrated into a computing-based device. The computing-based device may be an electronic device. Suitably, the computing-based device includes one or more processors for processing computer-executable instructions to control the operation of the device, thereby implementing the methods described herein. Any computer-readable medium, such as a memory, may be used to provide computer-executable instructions. The methods described herein may be executed by software in a machine-readable form on a tangible storage medium. Software may be provided on a computing-based device to implement the methods described herein.
[0105] The above description describes that the system circuit device and the monitoring circuit device are included on the same SoC. In an alternative embodiment, the system circuit device and the monitoring circuit device are included on two or more integrated circuit chips of an MCM. In an MCM, the integrated circuit chips are usually stacked together or located adjacently on an intermediate substrate. Some system circuit devices may be located on one integrated circuit chip, and other system circuit devices are located on different integrated circuit chips of the MCM. Similarly, the monitoring circuit device can be distributed on more than one integrated circuit chip of the MCM. Therefore, the methods and devices described above in the context of SoC are also applicable to the context of MCM.
[0106] The applicant hereby separately discloses each independent feature described herein and any combination of two or more such features, provided that such features or combinations can be made by those skilled in the art based on the whole specification according to the common general knowledge, regardless of whether such features or combinations of features solve any of the problems disclosed herein, and without limiting the scope of the claims. The applicant states that various aspects of the present invention may be constituted by any such individual feature or combination of features. In view of the above description, it is obvious to those skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. A method for identifying the cause of an abnormal characteristic measured from a system circuit device on an integrated circuit IC chip, the IC chip comprising the system circuit device and a monitoring circuit device, the monitoring circuit device being used to monitor the system circuit device by measuring the characteristic of the system circuit device in each window of a series of windows, the method include: (i) identifying a candidate window set from a window set preceding the abnormal window including the abnormal feature, so as to search for the cause of the abnormal feature in the candidate window set; (ii) for each of the measured characteristics of the system circuitry: (a) calculating, for the candidate window set, a first feature probability distribution of the measured feature; (b) calculating a second feature probability distribution of the measured feature for a window that is not in the candidate window set; (c) comparing the first feature probability distribution and the second feature probability distribution; as well as (d) if the first feature probability distribution and the second feature probability distribution differ by more than a threshold, identifying the measured feature within the time range of the candidate window set as the cause of the abnormal feature; (iii) iterating steps (i) and (ii) for a plurality of other candidate window sets from the window set preceding the abnormal window; and (iv) outputting a signal indicative of those measured features identified in step (ii)(d) as being the cause of the abnormal features.
2. The method according to claim 1, in: Step (ii)(c) comprises determining a difference measure between the first feature probability distribution and the second feature probability distribution; as well as Step (ii)(d) includes identifying that the measured feature within the time range of the candidate window set is the cause of the abnormal feature if the difference measure is greater than the threshold.
3. The method according to claim 2, in, The difference metric is scaled by the percentile of the difference over time between the first feature probability distribution and the second feature probability distribution of the iteration.
4. The method according to any one of claims 1 to 3, in, The set of windows preceding the anomaly window is bounded by (i) the anomaly window and (ii) a distal early window.
5. The method according to claim 4, in, Step (ii)(b) comprises, for a window set between the candidate window set and the abnormal window, calculating the second feature probability distribution of the measured feature.
6. The method according to any one of claims 1 to 3, in, The candidate window set includes less than 10 windows.
7. The method according to claim 6, in, The candidate window set includes only one window.
8. The method according to any one of claims 1 to 3, in, The first feature probability distribution and the second feature probability distribution are calculated in steps (ii)(a) and (ii)(b) by fitting a Gaussian model to the measured features of the identified window.
9. The method according to any one of claims 1 to 3, further comprising: identifying a measured feature affected by the abnormal feature, wherein the affected measured feature is located in a window after the abnormal window, wherein the method include: (v) identifying a subsequent candidate window set from a window set following the abnormal window, so as to search for the influence of the abnormal feature in the subsequent candidate window set; (vi) for each of the measured characteristics of the system circuitry: (a) calculating a third feature probability distribution of the measured feature for the subsequent candidate window set; (b) calculating a fourth feature probability distribution of the measured feature for a subsequent window that is not in the subsequent candidate window set; (c) comparing the third feature probability distribution and the fourth feature probability distribution; as well as (d) if the third feature probability distribution and the fourth feature probability distribution differ by more than another threshold, identifying the measured feature within the time range of the subsequent candidate window set as being affected by the abnormal feature; as well as (vii) iterating steps (v) and (vi) for a plurality of other subsequent candidate window sets of the window set following the abnormal window; and (viii) outputting a signal indicative of those measured features identified in step (vi)(d) as being affected by the anomalous feature.
10. The method according to claim 9, in: Step (vi)(c) includes determining another difference measure between the third feature probability distribution and the fourth feature probability distribution; as well as Step (vi)(d) includes identifying that the measured features within the time range of the subsequent candidate window set are affected by the abnormal feature if the other difference measure is greater than the other threshold.
11. The method according to claim 10, in, The other difference measure is a time-scaled difference between the third feature probability distribution and the fourth feature probability distribution.
12. The method according to claim 9, in, The set of windows following the anomaly window is bounded by (i) the anomaly window and (ii) a distal late window.
13. The method according to claim 12, in, Step (vi)(b) includes, for a window set between the subsequent candidate window set and the abnormal window, calculating the fourth feature probability distribution of the measured feature.
14. The method according to claim 9, in, The subsequent candidate window set includes less than 10 windows.
15. The method according to claim 14, in, The subsequent candidate window set includes only one window.
16. The method according to claim 9, in, The third feature probability distribution and the fourth feature probability distribution are calculated in steps (vi)(a) and (b) by fitting a Gaussian model to the measured features of the identified window.
17. The method according to any one of claims 1 to 3, in, The measured characteristics include those derived from trace data generated by the monitoring circuitry from data output by a plurality of components of the system circuitry.
18. The method according to any one of claims 1 to 3, in, The measured characteristics include those characteristics derived from matching events identified by the monitoring circuitry from data input to or output from the plurality of components of the system circuitry.
19. The method according to any one of claims 1 to 3, in, The measured characteristics include characteristics derived from a plurality of counters of the monitoring circuitry, the counters being configured to count each time a particular item is observed from a plurality of components of the system circuitry.
Citation Information
Patent Citations
System and method for a multi view learning approach to anomaly detection and root cause analysis
CN108028776A
Abnormality detection method and device based on log graph modeling
CN108920947A