TIE* Algorithm Discovers Multiple Markov Boundaries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for discovering Markov boundaries in high-dimensional data are either computationally inefficient or fail to identify all maximally accurate and non-redundant predictive models, leading to incomplete sets and reproducibility issues in bioinformatics and biotechnology applications.
Innovation Solution
The TIE* method systematically identifies all Markov boundaries and maximally accurate, non-redundant predictive models by using a generative approach that iteratively generates and verifies subsets of variables, ensuring admissibility criteria are met, thus overcoming the limitations of existing methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing methods for discovering Markov boundaries are used, then computational efficiency is improved, but the completeness of identified predictive models deteriorates
Solution Approach 1:
The method segments the search space of predictive models by systematically exploring different variable subsets through a structured algorithm that divides the complex search into manageable steps, ensuring all Markov boundaries are identified without computational redundancy
Solution Approach 2:
The method performs preliminary actions by pre-processing the data to identify candidate variable sets before applying the main discovery algorithm, ensuring that all potential Markov boundaries are captured in the final result set
2Productivity
If existing methods for discovering Markov boundaries are used, then computational efficiency is improved, but the accuracy of predictive models deteriorates
Solution Approach 1:
The method incorporates feedback mechanisms where the identified Markov boundaries are validated against the original data distribution to ensure accuracy, and the algorithm adjusts its search based on validation results to maintain high predictive accuracy
Solution Approach 2:
The method changes parameters such as the set of candidate variables and conditioning sets dynamically during the search process, allowing the algorithm to adapt to different data structures and identify accurate predictive models across diverse datasets
3Loss of information
If multiple predictive models are identified, then understanding of underlying mechanisms is improved, but device complexity increases
Solution Approach 1:
The method segments the complex task of identifying multiple predictive models into distinct algorithmic steps, making the system more manageable and easier to implement while maintaining comprehensive model discovery
Solution Approach 2:
The method creates a universal algorithm that can handle various data types and structures through a single unified approach, reducing the need for multiple specialized tools and simplifying the overall system architecture
Data Source
AI summary
Methods for discovery of a Markov boundary from data constitute one of the most important recent developments in pattern recognition and applied statistics, primarily because they offer a principled solution to the variable/feature selection problem and give insight about local causal structure. Even though there is always a single Markov boundary of the response variable in faithful distributions, distributions with violations of the intersection property of probability theory may have multiple Markov boundaries. Such distributions are abundant in practical data-analytic applications, and there are several reasons why it is important to discover all Markov boundaries from such data. The present invention is a novel computer implemented generative method (termed TIE*) that can discover all Markov boundaries from a data sample drawn from a distribution. TIE* can be instantiated to discover all and only Markov boundaries independent of data distribution. TIE* has been tested with simulated and re-simulated data and then applied to (a) identify the set of maximally accurate and non-redundant molecular signatures and to (b) discover Markov boundaries in datasets from several application domains including but not limited to: biology, medicine, economics, ecology, digit recognition, text categorization, and computational biology.


