Automatic Hyper-Local Feature Selection for Model Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Workbench software platforms lack an efficient method for users to automatically select relevant hyper-local data sources and features for model generation, as they are restricted from copying or exporting native hyper-local data, hindering the development of accurate models.
Innovation Solution
A computer-implemented method and system that generates a feature profile relation graph based on client data profiles, hyper-local feature importances, and use-case profiles, allowing for the automatic determination and ranking of hyper-local features for new models, while preserving data confidentiality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users manually select hyper-local data sources and features for model generation, then model accuracy can be improved, but user effort and time consumption increase significantly
Solution Approach 1:
The system performs automatic feature selection and model generation without requiring manual user intervention. The server autonomously selects relevant hyper-local data sources and features based on the user's description, then automatically generates and trains the predictive model, eliminating the need for users to manually configure complex parameters.
Solution Approach 2:
The server acts as an intermediary between the user and the hyper-local data sources. It receives the user's natural language description, automatically interprets the requirements, selects appropriate features from multiple data sources, and generates the model, thereby mediating the complex interaction between user intent and data source selection.
2Reliability
If users are restricted from copying or exporting native hyper-local data, then data confidentiality and integrity are maintained, but the ability to perform external analysis and model generation is hindered
Solution Approach 1:
The server serves as a secure intermediary that allows users to access and utilize hyper-local data for model generation without exposing the raw data. The system processes data requests, performs feature selection, and generates models within the platform's secure environment, preventing data export while enabling analytical capabilities.
Solution Approach 2:
Instead of copying or exporting raw hyper-local data, the system creates and exports simplified representations such as feature importance scores, selected feature lists, and model predictions. These derived artifacts contain the essential analytical value while preserving the confidentiality of the underlying sensitive data.
3Adaptability or versatility
If multiple hyper-local data sources are integrated into the platform, then the variety and quality of available features increase, but the complexity of data source selection and feature engineering increases
Solution Approach 1:
The system automatically performs feature selection and engineering tasks without requiring user expertise in data science. The server analyzes the user's description, automatically identifies relevant features from multiple hyper-local data sources, and handles the complex process of feature engineering, making the platform accessible to non-experts.
Solution Approach 2:
The system provides feedback to users in the form of feature importance scores and explanations for why certain features were selected. This feedback mechanism helps users understand the system's decisions, validate the selected features, and iteratively refine their model requests, reducing the perceived complexity through transparency.
Data Source
AI summary
Methods, systems and computer program products for providing automatic determination of recommended hyper-local data sources and features for use in modeling is provided. Responsive to training each model of a plurality of models, aspects include receiving client data, a use-case description and a selection of hyper-local data sources, generating a client data profile, determining feature importance and generating a use-case profile. Aspects also include generating a feature profile relation graph including client data profile nodes, hyper-local feature nodes and a use-case profile nodes, wherein each hyper-local feature node is associated with one or more client data profile nodes and user-case profile nodes by a respective edge having an associated edge weight. Responsive to receiving a new client data set and a new use-case description, aspects also include determining one or more hyper-local features as suggested hyper-local features for use in building a new model.


