Policy Data Source Mining via Hierarchical Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for creating a list of policy data sources in enterprises are inefficient due to heterogeneity, scalability issues, and the inability to guarantee discreteness, making it difficult to apply policies effectively, especially in large IT systems with billions of information objects.
Innovation Solution
A method that cleanses information objects, determines sorting criteria, and uses hierarchical clustering with an Euclidean distance measure to create discrete policy target groups, providing human-understandable names and descriptions, thereby overcoming the limitations of manual and database-oriented approaches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual collection of policy data sources using office tools, meta-data from crawling IT repositories, and interviews with employees is used, then flexibility and adaptability are maintained, but productivity and scalability deteriorate significantly when dealing with billions of information objects
Solution Approach 1:
The system enables self-service by automatically harvesting meta-data from IT repositories and performing clustering without requiring manual interviews with employees or manual collection using office tools, thus maintaining adaptability while dramatically improving productivity
Solution Approach 2:
The patent replaces manual mechanical processes (office tools, interviews) with automated computational processes including meta-data crawling, data cleansing, and hierarchical clustering algorithms, eliminating the scalability bottleneck while preserving the ability to handle diverse policy data sources
2Productivity
If data warehouse type querying of indexed meta-data collections is used, then productivity improves through automated processing, but the ability to guarantee discreteness of policy data sources deteriorates
Solution Approach 1:
The system incorporates feedback mechanisms where clustering results are evaluated against discreteness criteria, and the process iterates to ensure that policy data sources are pairwise discrete. The human-understandable names and descriptions provide feedback for validation of proper discretization
Solution Approach 2:
The patent introduces hierarchical clustering as an intermediary process between raw meta-data and final policy data sources. This intermediary step with human-understandable naming ensures that the automated process produces discrete, non-overlapping policy data sources while maintaining productivity
3Adaptability or versatility
If traditional methods are used to handle extreme heterogeneity of policy data sources, then comprehensive coverage is achieved, but device complexity and processing requirements increase dramatically
Solution Approach 1:
The patent handles heterogeneity by transforming diverse policy data sources into a unified meta-data representation with standardized parameters. The hierarchical clustering operates on these standardized parameters, allowing comprehensive coverage of heterogeneous sources without proportionally increasing system complexity
Solution Approach 2:
The system segments the complex heterogeneous data landscape into hierarchical clusters with human-understandable names and descriptions. This segmentation organizes the diversity into manageable groups, reducing the effective complexity while maintaining comprehensive coverage of all policy data sources
Data Source
AI summary
A method and system determines discrete policy target groups for information objects stored in an enterprise IT system. The method and system provide cleansed information about information objects stored on the enterprise IT system. Criteria for sorting the information objects is determined. Initial sorting of the information objects is carried out, resulting in an initial set of clusters. The information objects are clustered into discrete policy target groups based on the information about the information objects and the initial set of clusters, and human-understandable names and definite descriptions for policy target groups are computed.


