Intelligent Model Training with Guidance-Based Data Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional machine-learning models face challenges in efficiently processing large text inputs due to increased computing cost and complexity, often leading to inaccurate predictions and resource wastage, while existing data reduction tools fail to align with human understanding and regulatory guidelines.
Innovation Solution
The use of extrinsic guidance data, such as clinical guidance documents, to filter and highlight relevant portions of training data, thereby training a machine-learning model to focus on material pertinent to decision-making, reducing computing complexity and enhancing accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If large text input is used for training machine-learning models, then the model may capture more information, but computing cost and complexity increase significantly
Solution Approach 1:
The patent extracts only the relevant portions of training data using guidance data to identify and filter out unnecessary information. This extraction approach maintains the essential information needed for accurate predictions while removing redundant content that contributes to computing complexity.
Solution Approach 2:
The training data is segmented into relevant and irrelevant portions based on guidance data. By dividing the large text input into manageable segments and processing only the relevant ones, the system reduces computing complexity while preserving critical information.
2Loss of information
If large text input is used for training machine-learning models, then the model may capture more information, but computing cost increases
Solution Approach 1:
The system extracts only the essential information from large text inputs using guidance data to filter out redundant content. This extraction process reduces the volume of data that needs to be processed, thereby lowering computing cost while maintaining information quality.
Solution Approach 2:
Instead of processing all available training data, the system applies partial action by selectively processing only the relevant portions identified through guidance data. This approach avoids the excessive computing cost associated with processing entire large datasets while still capturing necessary information.
3Device complexity
If conventional data reduction techniques are used, then input size is reduced, but information loss occurs and alignment with human understanding is compromised
Solution Approach 1:
Guidance data serves as an intermediary between the raw training data and the model processing. This intermediary provides domain-specific knowledge that guides the filtering process, ensuring that only truly relevant information is retained while maintaining alignment with human understanding and decision-making processes.
Solution Approach 2:
The system changes the parameter of data selection from random or uniform sampling to guidance-based selective sampling. By altering how training data is chosen and filtered, the system reduces input size while preserving information quality through domain-expert-guided selection.
4Device complexity
If conventional data reduction techniques are used, then input size is reduced, but computing cost increases due to trade-off between accuracy and computational cost
Solution Approach 1:
Guidance data acts as an intermediary that enables efficient filtering of training data without requiring complex processing. This intermediary approach reduces input size to the extent that computing cost is lowered, breaking the conventional trade-off between accuracy and computational cost.
Data Source
AI summary
Systems and methods are described for training and/or using a machine-learning model. A first set of textual data is received. Using a trained machine-learning model that is applied to the first set, a classification of the first set is generated. The trained machine-learning model has been trained based on a subset of textual data that resulted from filtering a set of training textual data. The filtering of the set of training textual data to generate the subset of textual data is based on a comparison between the training textual data and a second set of textual data.


