Sparse PCA Feature Filter for IoT Data Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-dimensional data from IoT and mobile computing devices often includes significant irrelevant data, leading to misleading outcomes in machine learning and data mining applications, and increases computational processing, storage, and communication requirements, making them unsuitable for mobile devices.
Innovation Solution
A filtering mechanism using sparse principle component analysis (PCA) to identify and select significant features from a dataset, grouping right-singular vectors into clusters and determining medoids to identify relevant features, which are then utilized by data-analytics applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all available data from sensors and monitoring technologies is used, then machine learning systems gain greater insight into monitored events, but the computational processing, storage, and communication requirements substantially increase
Solution Approach 1:
The patent extracts and removes irrelevant data and noise from the dataset, keeping only the significant features that contribute to machine learning outcomes. This is achieved through filtering mechanisms that identify and eliminate redundant information, thereby reducing computational requirements while preserving essential insights.
Solution Approach 2:
The patent changes the parameter of data dimensionality by transforming high-dimensional data into a lower-dimensional representation that retains the most important information. This dimensionality reduction approach maintains information quality while substantially decreasing computational processing, storage, and communication requirements.
2Loss of information
If all available data from sensors and monitoring technologies is used, then machine learning systems gain greater insight into monitored events, but storage and communication requirements increase
Solution Approach 1:
The patent extracts and removes irrelevant data and noise from the dataset, keeping only the significant features that contribute to machine learning outcomes. This is achieved through filtering mechanisms that identify and eliminate redundant information, thereby reducing computational requirements while preserving essential insights.
Solution Approach 2:
The patent changes the parameter of data dimensionality by transforming high-dimensional data into a lower-dimensional representation that retains the most important information. This dimensionality reduction approach maintains information quality while substantially decreasing computational processing, storage, and communication requirements.
3Loss of information
If high-dimensional data is processed, then more comprehensive analysis is achieved, but timely results become difficult or nearly impossible to provide
Solution Approach 1:
The patent extracts and removes irrelevant data and noise from the dataset, keeping only the significant features that contribute to machine learning outcomes. This is achieved through filtering mechanisms that identify and eliminate redundant information, thereby reducing computational requirements while preserving essential insights.
Solution Approach 2:
The patent changes the parameter of data dimensionality by transforming high-dimensional data into a lower-dimensional representation that retains the most important information. This dimensionality reduction approach maintains information quality while substantially decreasing computational processing, storage, and communication requirements.
Data Source
AI summary
A data-analytics application may be optimized for implementation on a computing device for conserving computing or providing timely results, such as a prediction, recommendation, inference, or diagnosis about a monitored system, process, event, or a user, for example. A feature filter or classifier is generated and incorporated into or used by the application to provide the optimization. The feature filter and classifier are generated based on a set of significant features, determined using a data condensation and summarization process, from a high-dimensional set of available features characterizing the target. For example, a process that includes utilizing combined sparse principal component analysis with sparse singular value decomposition and applying k-medoids clustering may determine the significant features. Insignificant features may be filtered out or not used, as information represented by the insignificant features is expressed by the significant features.


