Causal Model Construction via Data Clustering for Network Abnormality Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for constructing causal models in communication network systems face challenges with rule-based methods requiring expert knowledge and data-driven methods struggling with insufficient abnormal case examples, especially with varying types of observed data and frequent topology changes.
Innovation Solution
A model construction apparatus that collects observed data, clusters it, determines representative values, and constructs a causal model using both rule-based and data-driven methods, combining them to estimate abnormality locations or causes effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a rule-based method is used to construct a causal model, then the model can be built using expert knowledge, but it becomes difficult to establish rules for the relationship between abnormalities and each piece of various types of observed data one by one
Solution Approach 1:
The patent segments the observed data into multiple clusters based on data types (e.g., flow data, telemetry data, sensor data). For each cluster, a representative value is determined, and rules are established between abnormalities and cluster representatives rather than individual data pieces. This segmentation reduces the complexity of rule establishment while maintaining model construction feasibility through expert knowledge.
2Adaptability or versatility
If a data-driven method is used to construct a causal model, then the model can be built from observed data, but it becomes difficult to collect abnormal case examples when the types of observed data are various
Solution Approach 1:
The patent segments various types of observed data into multiple clusters, where each cluster contains data of a specific type. By determining a representative value for each cluster, the system reduces the dimensionality of the data space. This segmentation enables the collection of sufficient abnormal case examples for each cluster separately, rather than requiring comprehensive examples for all possible combinations of various data types.
Solution Approach 2:
The patent introduces representative values as intermediaries between the actual observed data and the causal model. These representative values serve as mediators that capture the essential characteristics of each data cluster, allowing the model to learn from a reduced set of examples while still being applicable to the full range of observed data types.
3Measurement precision
If various types of observed data are used to construct a causal model, then the estimation of abnormality location or cause can be achieved with finer granularity, but the complexity of model construction increases
Solution Approach 1:
The patent segments various types of observed data into multiple clusters based on data types, with each cluster representing a specific category (e.g., flow data, telemetry data, sensor data). A representative value is determined for each cluster, which maintains the fine-grained information from the original data types while reducing the overall complexity of model construction by working with clustered representatives rather than individual data pieces.
Data Source
AI summary
A model construction apparatus according to an embodiment includes a processor and a memory storing program instructions that cause the processor to receive pieces of observed data from a communication network system that is a target for estimation of a location or a cause of an abnormality; divide the received pieces of observed data into a plurality of clusters according to the types of information represented by the respective pieces of observed data; determine, for each location or each cause of an abnormality, a representative value as representative observed data for each of the plurality of clusters; and construct, using the representative observed data, a first causal model for estimating the location or the cause of the abnormality from the pieces of observed data based on a rule-based method.


