A causality matrix creation process for severely unbalanced and evolving environments

The innovative causality matrix creation process addresses the challenges of data variability and noise in unbalanced environments by employing banned intervals, temporal spacing, minimum occurrences, and predictor subdivision, resulting in improved accuracy and robustness of causality relationships identified by machine learning models.

WO2025120515A1PCT designated stage expired Publication Date: 2025-06-12ALTICE LABS SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/062176
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing causality matrix creation processes are susceptible to data variability and noise, particularly in severely unbalanced and evolving environments, leading to false learning and inability to discern complete hierarchical correlations between events.

Method used

The process involves banned intervals based on the predictor under analysis, homogeneous temporal spacing between negative windows, a minimum number of occurrences in negative windows, and subdivision of predictors based on their temporal delta to the event prompting the window creation, ensuring robustness and accuracy in causality matrix creation.

Benefits of technology

This approach enhances the quality of the dataset created, improves the extraction of causality relationships by machine learning models, and minimizes false positives and correlations, effectively supporting severely unbalanced and evolving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024062176_12062025_PF_FP_ABST
    Figure IB2024062176_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention describes a new causality matrix creation process for analyzing correlations between events in any severely unbalanced deterministic environment, thus allowing the creation of a robust dataset for machine learning models. The proposed innovative process seamlessly integrates multiple functionalities to enhance its resilience in addressing the inherent challenges of real-world environments, ensuring optimal extraction of correlations while mitigating the risk of false cor-relations. Furthermore, specific optimizations are detailed to streamline the process, not only diminishing its complexity and execution time but also minimizing hardware requirements. These enhancements render the solution scalable to accommodate diverse sizes of real environments, a critical attribute in the context of big data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A CAUSALITY MATRIX CREATION PROCESS FOR SEVERELY UNBALANCED AND EVOLVING ENVIRONMENTS

[0002] FIELD OF THE INVENTION

[0003] The present invention is enclosed within the field of data mining in the context of causality matrix creation for Root Cause Analysis in any deterministic environment. Particularly, the present invention relates to a dataset creation process that will be used as input for a machine learning model to identify events correlation, thus being able to detect root causes. Additionally, the proposed process is capable of supporting extreme class imbalance by a unique selection process.

[0004] PRIOR ART

[0005] Completely deterministic event analysis environments are quite common, but in a real-world scenario, their scale can be enormous, making it difficult to detect these correlations and impossible for humans to detect their root causes.

[0006] An early solution to this problem was the creation of Root Cause Analysis tools. These operate on a set of predefined rules, based on human knowledge, to create causal relationships between multiple events in the environment. However, due to the increase of network complexity and volume of fault events, these tools often fall short, unable to discern the complete hierarchy of correlations. Their limitations are further exacerbated by the lack of specificity and completeness in the rules integrated into them. Due to the current scale, the present solution to this problem is the application of machine learning techniques, which thrive on data dimensionality. However, their application is not straightforward, and data mining techniques need to be applied to convert raw events into a useful dataset on which ML models can learn and extract correlations. In this regard, an exceptional approach is the creation and usage of a causality matrix, which discards most of the events' information and operates merely based on their causality, as the aim is to understand the impact between events to create a hierarchy tree.

[0007] Although this technique aids a lot with the mission of identifying correlations, it has one major shortcoming: it is very susceptible to data variability and noise. In real-world scenarios, this is almost always the case, where an imbalance of classes can be found in the order of thousands and the root causes can be both multiple and non-atomic.

[0008] This makes it imperative to create a process to generate causality matrices that support scenarios that are not only severely unbalanced but also disparate, thus managing to keep up with the evolution of environments over time.

[0009] PROBLEM TO BE SOLVED

[0010] When applying machine learning techniques to identify correlations between events and their respective root cause in any deterministic environment, which is a mandatory requirement for its success, a prior data mining and analytics process is necessary to convert all the information from the system's events into a useful dataset for an ML model learn upon.

[0011] One of the state-of-the-art techniques is the usage of causality matrices. Thus, instead of processing all the information from each single event, the analysis is carried out by the mere causality of events. So, initially, almost all the event's attributes are discarded, preserving only the minimum to be able to characterize it as a unique point of failure and then create a sliding window with the count of single events that have occurred previously .

[0012] However, this is not only a cumbersome process, due to the volume of data to be processed in any modern environment, but it is also very susceptible to false learning since the entropy of the systems is enormous and class imbalance can be extreme. It is, therefore, necessary to evolve the process of causality matrix creation to overcome these problems, allowing an ML model to effectively succeed in identifying the hierarchical relationships between events and be able to determine the real root cause, minimizing any scenario of false positives or false correlations.

[0013] SUMMARY OF THE INVENTION

[0014] The innovative process of creating causality matrices for severely unbalanced deterministic environments includes several distinct concepts: (i) banned intervals depending on the predictor under analysis (ii) homogeneous temporal spacing between negative windows (iii) implementation of a minimum number of occurrences in the negative windows based on the average number of occurrences on the positive windows. (iv) sub-division of the predictors based on their temporal delta to the event prompting the window creation. The combination of these four concepts has a major impact on the quality of the dataset created, thus improving the causality relationships extracted by the ML models.

[0015] Another concept introduced in the invention is mandatory time spacing. Since the environment can, and usually is, severely unbalanced, when analyzing a predictor, it may only occur several times over weeks if not months. Thus, without any kind of control, it would be possible to find the same number of negative windows in hours / days. Although the number of entries would be balanced, AT would be severely unbalanced, thus showing no different behavior over time or even any evolution of the environment. With this strategy, it is guaranteed that the selected negative events are chosen equally spaced over time, thus making it possible to better introduce the behavior of the environment over time under analysis .

[0016] In addition, a minimum number of occurrences in the negative windows is also forced. After calculating the positive windows, which must be done before the negative windows to calculate the intervals to be banned, the average number of events per window is calculated and then this value is used to find the negative windows to be used. The idea of this obligation is to approximate the behavior of the windows, preventing the model from being able to determine the polarity of the window by the inherent quantity of events (since the event under analysis could only occur at a time of high volume of events and the negative windows could be found at quieter times in the environment, thus allowing the two to be easily distinguished, without the need to identify causal relationships .

[0017] Finally, the last innovative factor of the invention is that the predictors are subdivided by time distance from the event that triggered the creation of the time window. In other words, instead of a single origin being just one predictor and all the events from that origin being counted in the same column of the dataset, several (at least 2) predictors are created instead. In this way, you can separate the impact that one event has on another, not only by their occurrence but also by their proximity, thus better understanding their relationship and causality.

[0018] In sum, the invention hereby described comprises a new process for creating causality matrices for root cause analysis in any deterministic environment, including not only new features to support severely unbalanced environments without biasing model learning but also implementation optimizations for the same, reducing matrix creation time and hardware requirements.

[0019] DESCRIPTION OF FIGURES

[0020] FIG. 1 is a flowchart showing the entire creation process and the order in which the functionalities are applied to create the causality matrix in a severely unbalanced and evolving environment.

[0021] FIG. 2A is a tabular diagram illustrating the result of calculating positive windows in the context of generating banned intervals. It represents the two main steps, showcasing the total intervals with overlaps and their subsequent merging into unique banned intervals .

[0022] FIG. 2B is a diagram showing the use of a number line with banned intervals to optimize the calculation of viable negative intervals through set theory.

[0023] FIG. 3 is a flowchart demonstrating the behavior of the enforcement of homogeneous temporal spacing between negative windows.

[0024] FIG. 4A is a tabular representation illustrating the application of the base process for the creation of a causality matrix, using the Cartesian product of two attributes as unique identifiers for a single point of failure.

[0025] FIG. 4B is a tabular representation showcasing the application of the functionality of subdividing predictors based on their AT regarding the window originating event.

[0026] DETAILED DESCRIPTION

[0027] The present disclosure relates to a new process for creating causality matrices to study hierarchical and causal relationships between events and is agnostic of the environment in which it is applied, provided that it is deterministic. In addition, this new approach was created to support severe class imbalance, thus allowing for the complete analysis of real environments, minimizing any type of false correlation, such as Post Hoc, Ergo Propter Hoc scenarios. In addition to the new process, this document also includes some optimizations used in the new process to reduce computing time, which has a significant impact on big data scenarios, and the hardware requirements needed for it. As aforementioned, this new process is made up of four main features: (i) event intervals depending on the predictor being analyzed to prevent predecessor events from one positive window from being used in other windows. In this step, there are some options as to which intervals should be banned, depending on the use case in question. This step is essential in "protecting" the root cause, increasing its impact on model learning. (ii) applying a homogeneous time distance between negative windows, based on an index of hours calculated a priori. This step is used to allow the network's behavior to be introduced via the negative windows, avoiding false correlations derived from a negative analysis on a small substrate of data, (iii) creating a minimum number of occurrences requirement for the use of a negative window. This step ensures greater proximity of behavior between a positive and negative window, thus avoiding false correlations derived from the frequency inherent in the window under analysis, (iv) application of a temporal division within the same predictor, based on the AT to the event originating the window, thus allowing for a greater specificity and perception, aiding analysis in environments with multiple and / or non-atomic root causes.

[0028] Regarding the banned intervals, this step is essential since the causality matrix discards all the event's attributes, except the bare minimum to allow for single failure source characterization. Thus, each time window (a row of the causality matrix) consists of the number of occurrences of each single source in the predefined interval before the event under analysis . So, initially, an issue that arises is, that by using contiguous positive and negative windows, the root cause is devalued, as it appears in both positive and negative scenarios (and two contiguous windows can be almost identical) . When learning the data, the machine learning models are wrongly influenced to consider that the true root cause has no impact and either discover false correlation scenarios or simply fail the learning phase, thus not reaching the minimum success threshold for extracting correlations.

[0029] With the introduction of banned windows, this issue is severely mitigated or even solved, as specific events used in a window can never appear in a negative window, thus valuing their impact and facilitating the process of detecting which events have a causality relationship (increasing the contrast between positive and negative windows) . In this respect, two strategies have emerged for banned windows :

[0030] - ban for positive and negative: In this scenario, once a positive window has been calculated, the events are completely banned and cannot be reused in any other data window. In this way, the next window to be calculated has to check that its AT does not collide with that of the previously calculated window, thus ensuring that there is no overlap. In this case, set theory is applied to speed up the calculation, severely reducing the complexity and thus the execution time of the task.

[0031] - ban only for negatives : In this scenario, all the positive windows can be calculated without any concern about their possible overlapping of events. Once they have been completed, it is calculated which event intervals have been used and only negative windows are calculated in the sub-segments to be used, thus ensuring that the intersection of events between the negative and positive windows is null, hence: {positive window events} fl {negative window events} = 0. Set theory is also applied for a huge performance gain, which is essential in a big data scenario.

[0032] In both strategies, positive windows must be calculated before negative windows and the dataset of events must be temporally ordered and synchronized by the same clock (usually that of the event capture server) . Even though, both solve the issue at hand and achieve positive results, thus minimizing false correlation scenarios such as Post Hoc Ergo Propter Hoc, the second strategy is considered to be the most optimized. By banning the used events only for the negative windows, it was possible to return slightly better results with a smaller number of banned intervals, allowing all positive events to be used, thus guaranteeing maximum learning on the same data sample.

[0033] The invention further comprises a step relating to the temporal spacing between the windows used, thus ensuring that the negative inputs are temporally homogeneously separated, creating a better representation of the behavior of the environment over time. Since the environment under analysis can be severely unbalanced, i.e. the distribution of the occurrence of events is not homogeneous across the various single points of failure, there can be frequency variations in the range of tens of thousands of times. Thus, over the entire period of the base event data, certain points of failure may only have occurred several times over weeks or even months, causing a random selection of negative entries (to create a correct balance) to possibly be close in time, occurring within a few hours or days. Such can lead to a scenario of false correlations, as the passive behavior of the environment can be so distinct over the entire interval that, as the negative windows are very close, they are easy to distinguish from the positive ones, without any need to find the correct root cause.

[0034] To efficiently solve this issue, a minimum time spacing between negative windows is implemented. For that, initially, all the indices of the events at the start of each hour (in small datasets minutes can be used instead of hours) are calculated, making it easy to navigate through the immensity of the data. Next, the total number of events is divided by the number of positive occurrences of the point of failure event under analysis, thus obtaining the perfect spacing to be used for the causality matrix. Once the correct spacing has been obtained for the matrix in question, the indices calculated a priori are used, allowing only a negative window to be used whose end of the time window is greater than the index. This feature ensures that all negative windows are evenly spaced across all the data and correspond to the negative behavior of the environment throughout its evolution, thus minimizing false correlations.

[0035] Another innovative factor implemented was the sub-division of each predictor, i.e. each single failure point, by its AT to the event originating the time window. Since the main objective is to study causality between events and root cause detection, the time distance between events is an essential factor, as the same event may only have an impact if it occurs at a close interval to another (non-atomic root cause scenario) . It was therefore limiting to only consider the number of occurrences of each predictor in the predefined interval prior to the event under analysis. To efficiently overcome this obstacle, three sub-divisions (average success value, but configurable for each predictor and / or for each environment) were created, so for predictor X, there would then be three predictors :

[0036] - X < 30s

[0037] - X > 30s A X < 60s :

[0038] - X > 60s

[0039] As the equations imply, these divisions make it possible to identify whether the event that occurred was less than 30 seconds, between 30 and 60 seconds, or more than 60 seconds from the event that originated the time window. In this way, it is possible to determine the impact that one event has on another with the degree of specificity of the moment in which it occurred, thus facilitating the process of detecting causality in non-atomic root cause scenarios. The values used for the divisions are relative to a 5- minute window scenario. If the window is longer, it is generally sufficient to multiply the values represented by the order of magnitude of the increase, thus maintaining the advantage of the temporal division (since if the intervals are too large or poorly distributed, the gain inherent in this step is lost) .

[0040] The last feature encompassed in this innovation is the incorporation of a minimum event occurrence requirement for negative windows. This ensures a balanced event frequency akin to positive windows, a strategic choice crafted to address severely unbalanced environments. In such scenarios, certain predictors may manifest only sporadically over an extended period, necessitating adaptability to shifts in environmental behavior. An integral facet of this innovation lies in recognizing the significant impact of event frequency on the learning process. Some periods are marked by a dense concentration of events, while others exhibit a notable scarcity. Aligning with the preceding features, the imposed threshold obligation compels the behavior of negative windows to mirror that of the positive ones. Despite the complexity it introduces to machine learning models, this approach minimizes false correlations, contributing to successful learning outcomes.

[0041] To avoid the undue influence of sparsely populated moments, where the target truth value could be inaccurately predicted based solely on the density or frequency of events, a minimum requirement is enforced for negative windows. This requirement is an additional criterion to be met alongside the existing functionalities. The calculation of this threshold is meticulously tailored to each matrix, derived from the average occurrence of events in positive windows, with an optional adjustment based on the standard deviation (depending on the predictor under analysis. This methodical approach ensures that only negative windows meeting the calculated metric are eligible for use, effectively mitigating false correlations and enhancing the overall robustness of the learning process.

[0042] The integration of these four key functionalities enhances the resilience of the causality matrix, destined for input into a machine learning model. This strengthened matrix not only supports severely unbalanced environments but also adapts to variations in environmental behavior, encompassing shifts in event frequency and the evolving nature of the environment, including changes in points of failure. Thus, minimizing false correlations while maximizing the extraction of meaningful causal knowledge. This comprehensive approach ensures that the resulting matrix serves as a robust and reliable foundation for subsequent machine-learning applications .

[0043] DESCRIPTION OF THE EMBODIMENTS

[0044] Fig. 1 is a flowchart representing the entire process of creating a causality matrix in a severely unbalanced and evolving environment, displaying the order in which all the features of this innovation are applied.

[0045] Initially, it is necessary to identify all the positive events that will be used for the matrix 100, since it is the positive windows that determine the behavior of the negative ones. Next, the number of positive occurrences is checked 101 to evaluate if the minimum conditions for training / creat ion of the dataset have been met. The extent of this value fluctuates depending on the specific use case at hand. Nevertheless, it remains crucial to establish a minimum threshold to gauge the effectiveness of the models and the extraction of meaningful correlations. If the threshold is not reached, the behavior depends on the scenario under analysis: (i) if it is real-time 102, it returns to the first step, waiting for more events to arrive until the threshold is reached; (ii) if the analysis is based on a complete dataset, the process comes to an end and it is considered that the analysis encounters a deficiency in the number of events, hindering the learning stage and extraction of meaningful correlations.

[0046] Once the threshold has been reached, the dataset predictors can then be calculated. In this setting, the functionality of temporal sub-division of the predictors 103 is applied, i.e. instead of merely calculating all the single failure points, they are also sub-divided according to their AT to the event causing the window, thus allowing a more specific analysis of their impact. Subsequently, it is possible to calculate the positive windows 104 through an iterative process, capable of being optimized from 0(V2) to 0(JV) through a temporal ordering, which allows a min-max passage between windows. After determining the positive windows, the next step involves calculating the negative windows, which encompass a majority of the features associated with the innovation.

[0047] Initially, the average number of events occurring in the positive windows is calculated 105, to understand what frequency baseline should be required of the negative windows, avoiding any kind of frequency bias, which could occur due to the inhomogeneous behavior of the environment under analysis. Following this, the banned intervals are calculated 106, i.e. which event intervals were used in the positive windows and therefore cannot be studied in the negative windows. This ensures that root causes are protected, since without banned intervals, with contiguous windows, root causes can be devalued.

[0048] Once the analysis of the positive windows for the creation of the negative windows is complete, a time check is then carried out 107. This checks whether the number of positive occurrences is less than the number of hours (or other configured degree of specificity) of the total dataset. If not, no restrictions are placed on the iteration of the indexes, as the extensive frequency necessitates the utilization of the entire dataset to fulfill the other functionalities requirements. If indeed that is the case, the negative windows need to be properly spaced out to better demonstrate the behavior of the environment over time. Such is accomplished by checking whether the time mapping, which has the index of the first event of each hour (or other configured degree of specificity) , has already been created and / or updated 109. If not, it has to then be created / updated 110. This can be used to create the causality matrix for any event in the environment, as long as no further events occur, in which case this basic mapping can be kept and will have to be updated as new events arrive. Once the mapping has been created and updated, it is then possible to calculate which indices will be iterated over 111, i.e. for each base index a negative window will be found that meets all the requirements of the functionalities and then skips to the next index to use, thus forcing a homogeneous time gap. However, for the window calculation, two final checks are made: if the index window under analysis has as many events as 75% of the average of the positive windows, and if the events present in the window are not in any banned interval 112. If these 2 requirements are not met, the index's window fails and the next index is analyzed, which is always directly adjacent, except for specific optimizations regarding the banned intervals, allowing skipping certain indices known a priori to be invalid. After meeting both of these thresholds, the negative window is created 113 and the balance of positive and negative indices is checked 114. If successful, the positive and negative windows are then concatenated *115*, and the process is finished, with the causality matrix capable of withstanding severely unbalanced and evolving environments having been created. Otherwise, the next index is analyzed, which can be the directly adjacent one in the 108 scenario or the index of the next hour, thus forcing a time gap, in the 111 scenario.

[0049] Fig. 2A is a tabular representation of the functionality for applying banned intervals from positive windows, thus preventing their subsequent use in negative windows. Each window, composed with the history of occurrences before its originating event, has its root cause. However, without the application of banned intervals, the root cause would be devalued as it could appear in both a positive and negative window if they were contiguous. The banned intervals effectively mitigate this scenario. This functionality consists of two main steps. Initially, a calculation is made of all the intervals to be banned, thus identifying all the intervals that were used in the 200 positive windows. However, since no limitation is imposed between positive windows, interval overlap scenarios may occur 201. Then, since there is no point in reanalyzing the same interval several times just because it belongs to multiple windows, the second step is applied. In this way, a merge process is carried out between the intervals, thus creating a list of unique intervals 202 that cannot be used in the negative windows. This significantly reduces the number of comparisons and iterations required, thus minimizing complexity and therefore execution time.

[0050] Fig. 2B is a diagram showcasing the application of a number line to represent the entire dataset and the intervals banned by the positive windows, to help calculate the viable intervals, via set theory, for the creation of negative windows. As shown in the diagram in Fig. 2A, once the positive windows have been calculated, the events are precluded from appearing in negative windows, however, even after calculating the merged intervals it can be a very heavy and slow process if not properly optimized. Therefore, an efficient mechanism for calculating the negative intervals and iterating over them was developed. Initially, having calculated unique positive intervals 204, it is possible to identify three possible scenarios:

[0051] - Initial iterations, prior to a banned interval 203: in this scenario, it suffices to iterate through the windows ensuring that the id of the event originating in the window is lower than the first id (the one with the lowest index) of the first banned interval. When a window is found that meets this and all other requirements of the innovation, it is unnecessary to iterate to the next id, as it's possible to skip to the event whose window last id is greater than the originator id of the window just calculated .

[0052] - Intermediate iterations, having a lower and upper banned interval 205: in this scenario, the viable interval for a negative window is more restricted, as there are upper and lower constraints. It is therefore necessary to preserve the highest index value of the lower banned interval (min) and the lowest index value of the upper banned interval (max) . At each iteration through the 205 viable intervals, the following checks are made: (i) the event index value is greater than min and less than max (ii) the index value of the first event in the window under analysis is greater than min (iii) the event window does not overlap with a previously calculated negative in the same viable interval. If condition (i) fails, it means that the event under analysis is in a banned interval and can be skipped directly to the next highest index in the upper banned interval. If condition (ii) fails, it means that although the event under analysis is not in the banned interval, its window has an overlap with the banned interval and can be skipped to the index whose window mapping is greater than min. If condition (iii) fails, it means that although there was no collision with the positive banned events, there was an overlap with the last negative window, so it is possible to skip to the index whose window mapping is greater than the index of the last event that caused a negative window.

[0053] - Final iterations, after the last banned interval 206: in this scenario, it suffices to iterate through the windows ensuring that the first id of the window of the event under analysis is greater than the highest id of the lower banned interval (min) . If the event under analysis has an overlap with the banned interval, it's possible to skip to the event whose window mapping is higher than min.

[0054] Fig. 3 is a flowchart showing the enforcement behavior of the homogeneous temporal spacing between the negative windows, to better convey the passive and evolutionary behavior of the environment. As previously mentioned, an environment is not necessarily constant over time and can therefore have variations in frequency and single points of failure, so it is necessary, for extraction of better correlations, to transpose this behavior into the negative windows as well (making false correlations based on frequency impossible) .

[0055] Initially, the positive windows are calculated 300 based on the occurrences of the point of failure events under analysis, and it is mandatory to calculate these first for all the reasons aforementioned in the banned windows functionality. After calculating the positive entries, a mapping containing the first index of each time change is used to guarantee efficient spacing of the negative entries. This allows to efficiently skip through the initial dataset without having to iterate continuously. At the start of the negative window calculation, a check is performed to verify that the mapping is up to date 301. This verification is imperative, especially when studying the environment in real-time, where events unfold dynamically as a continuous stream of data. Therefore, it is crucial to ensure that the mapping is consistently updated with new time indices to accurately reflect the evolving nature of the events. If the validation fails, the mapping will be recalculated 302. This process involves iterating through the events, starting from the last index present in the mapping (thus ensuring the minimum number of iterations possible) .

[0056] Once the mapping has been updated, the number of positive occurrences can then be checked 303. This step is used to identify how many negative windows need to be calculated to create a proper balance in the causality matrix, thus not forming any bias in the subsequent learning of the ML models. Next, it is checked whether the number of positive occurrences is greater than the number of indices calculated in the mapping 304. This verification is necessary for assessing the frequency of the failure point under analysis. In the scenario where it is greater, it is inferred that it is so frequent that it is not necessary to force the even time spacing, as the entire dataset will be used regardless, so a normal iteration process can be applied 305. This scenario occurs in less than 1% of single fault points, but due to its high frequency, it can represent 30+% of all events (thus demonstrating once again the degree of severe unbalance present in real environments) . In the scenario where the number of positive events is lesser than the number of indices, enforcing time-spacing becomes necessary. Next, the number of indices is divided by the number of positive occurrences 306, to calculate the minimum duration of spacing required between negative windows. Following this, the corresponding indices are extracted from the mapping 307, and the iteration through the list of indices is performed to generate a single negative window that meets the specified criteria outlined in the other aspects of the innovation.

[0057] Finally, once all the negative windows have been calculated, the positive and negative windows can be concatenated 308, thus generating the balanced and temporally spaced causality matrix. It is worth noting that all the windows are calculated using sparse matrices, as they represent the occurrence counts of all predictors, resulting in a matrix predominantly composed of zero, thus reducing hardware requirements by several orders of magnitude.

[0058] Fig. 4A is a tabular representation of the state-of-the-art base process for constructing a causality matrix, on which, in this innovation, new functionalities have been developed to support severely unbalanced and evolving environments. From a macro perspective, the process consists of three phases. Initially, almost all the information from the event is discarded 401, preserving only the minimum to identify the point of failure. The resulting Cartesian product (if more than one attribute is required 400 402) will be used as a predictor, so the predictors in the dataset are all the unique single points of failure 404 405 406 407 in the environment under analysis. Next, all the causality windows 408 are created by counting all the occurrences of events for each failure point. This leads to the creation of an intermediate dataset equal in size to the number of events. In addition, a temporary array 409 of equal size is created which contains, for each window, the point of failure that caused the window. Finally, the positive and negative entries are selected to create a balanced dataset and a Boolean comparison is made between the temporary array of failure points and the failure point under analysis, thus creating the target of the dataset 410. This array is an optimization that allows the causality matrix to be calculated only once for all the data and then, for the analysis of each failure point, a simple Boolean comparison with the array is adequate, thus avoiding the need for a comprehensive recalculation of the entire matrix .

[0059] Fig. 4B is a tabular representation of the predictors' temporal sub-division functionality based on their AT to the windoworiginating event. An inherent challenge associated with the causality matrix approach lies in its limited capacity to discern non-atomic root causes— situations wherein multiple events must occur concurrently to trigger the analyzed point of failure. As these may have to occur in a specific order or with certain temporal proximity, simply counting fault occurrences is not sufficient for this detection. Therefore, it is imperative to separate the predictors based on the time of their occurrence. For this purpose, several intervals can be created, where these three showed the best results (for the 5-minute window scenario) : (i) the event occurred less than 30 seconds from the window originator event (ii) the event occurred between 30 and 60 seconds (iii) the event occurred after 60 seconds. Thus, from each predictor, three new ones appear, i.e. from 403 to 411 412 413, from 404 to 413 414 415, from 405 to 416 417 418, and from 406 to 419 420 421. With the new predictors calculated, it is possible to create every window with the occurrences count, where, assuming a gap of 25 seconds between the events in 403, the windows in 422 would be created. A notable difference between 407 and 422 is in the fluctuation of column counts throughout the windows, as over time, events' counts converge on the later (third) column of their predictor (thus speeding up the process of the temporal evolution of the environment, aiding the process of identifying non-atomic root causes) .

Claims

Claims1. A method for creating a causality matrix from any severely unbalanced, evolving and deterministic environment, enabling the full utilization of the available dataset while maintaining the integrity of machine learning model training, he method comprising : a) identifying positive events associated with a point of failure (100) ; b) determining a minimum number of occurrences of the positive events needed to trigger causality matrix creation (101) ; c) calculating a time window for each positive event (300) , wherein the time window includes a set of events preceding the positive event; d) calculating negative time windows, each containing events occurring during a time period before and separate from the time windows of positive events and wherein the minimum number of occurrences of events in each negative time window is calculated based on the average number of positive events in the positive time windows; e) enforcing a minimum time separation between negative time windows ; f) subdividing predictors based on their temporal distance (AT) from the event triggering the creation of a time window; g) generating a causality matrix comprising the positive and negative time windows .

2. The method of claim 1, wherein step (d) further comprises calculating the average number of events per positive time window and establishing a minimum number of occurrences for negative time windows based on the calculated average (105) . Root cause accuracy is further enhanced by generating bannedintervals during positive window calculations, preventing the influence of network noise on negative window calculations (202) .

3. The method of claim 1, wherein step (e) further comprises determining a minimum time spacing between negative windows based on the frequency of positive events (304) .

4. The method of claim 1, wherein step (f) further comprises dividing each predictor into sub-predictors based on their temporal distance (AT) from the triggering event using multiple time thresholds (400) .

5. The method of claim 1, incorporates optimizations based on temporal ordering that enable the use of a min-max iterative approach, thereby reducing the complexity of positive window creation from 0 (N2) to 0 (N) .

6. The method of claim 1, wherein the causality matrix is created using sparse matrix techniques to improve efficiency and severely reduce storage requirements.