Dynamic Data Subset Selection with Velocity-Based Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Individuals face challenges in determining which subsets of information are optimally suited for a particular task, leading to inefficient information allocation and user experience issues.
Innovation Solution
A computing platform generates data silos within a distributed ledger system, identifies relevant data silos using a machine learning model, assigns data velocity values (DVVs) to these silos, and trains the model based on efficiency to select optimal data subsets for specific use cases, refining the model with updated DVVs and performance data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all available information is processed to complete a task, then the completeness of information is improved, but the efficiency of information allocation deteriorates
Solution Approach 1:
The patent segments the large volume of available information into organized data silos grouped by use cases. The machine learning model then segments the selection process by evaluating only relevant subsets rather than processing all information, thus maintaining completeness while improving efficiency.
Solution Approach 2:
The patent introduces data velocity values as a new parameter to characterize and evaluate data silos. By changing the evaluation parameter from brute-force completeness checking to velocity-based relevance scoring, the system efficiently identifies appropriate information subsets without processing all available data.
2Productivity
If a machine learning model is trained to select optimal data subsets, then the efficiency of information allocation is improved, but the complexity of the system increases
Solution Approach 1:
The patent introduces data velocity values as an intermediary metric that simplifies the machine learning model's task. Instead of directly optimizing complex information allocation decisions, the model works with the intermediate DVV parameter, which captures data characteristics and enables more efficient subset selection.
Solution Approach 2:
The system performs self-training by automatically generating data velocity values from monitoring information and using these to train the machine learning model. This self-service approach reduces external configuration complexity while improving information allocation efficiency.
3Measurement precision
If data velocity values are monitored and used to train the model, then the accuracy of data subset selection is improved, but the amount of processing required increases
Solution Approach 1:
The patent applies partial action by monitoring and processing only the specific metrics needed to calculate data velocity values, rather than analyzing all possible data characteristics. This selective monitoring approach improves selection accuracy while controlling processing resource consumption.
Data Source
AI summary
Aspects of the disclosure relate to an information reduction platform. The information reduction platform may generate data silos within a distributed ledger system. The information reduction platform may receive a data use case request from a client device corresponding to a first user. The information reduction platform may identify relevant data silos corresponding to the data use case request. The information reduction platform may direct the distributed ledger system to grant the client device access to the relevant data silos. The information reduction platform may monitor the efficiency of the relevant data silos. The information reduction platform may generate data velocity values (DVVs) for each relevant data silo. The information reduction platform may train a machine learning model to select a subset of relevant data silos for a particular data use case. The information reduction platform may create an iterative feedback loop to update the first machine learning model.


