Predictive Modeling for Drug Portfolio Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The lack of large, annotated datasets hinders the performance of supervised machine learning algorithms in predicting consumer preferences for prescription drugs, limiting the ability to accurately predict product adoption and inclusion in a portfolio, especially for new products.
Innovation Solution
The development of computer-implemented systems and methods that utilize consumer preference data and drug feature data to train predictive models, incorporating natural language processing and feature selection techniques to generate machine learning classifiers for predicting consumer preferences and recommending substitution or new drugs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised machine learning algorithms are used to predict consumer preferences, then prediction capability is improved, but the lack of large annotated datasets causes performance degradation and overfitting
Solution Approach 1:
The system performs preliminary actions by collecting and annotating transaction data, product attribute data, and consumer preference data before training the machine learning model. This includes gathering historical transaction records, extracting product features using NLP, and preparing labeled datasets in advance to ensure sufficient training data is available when model training begins, thereby preventing data scarcity issues
Solution Approach 2:
The patent introduces an intermediary data processing layer that includes natural language processing components and feature extraction mechanisms. This intermediary layer transforms raw transaction data and product descriptions into structured, annotated features that serve as mediators between the available raw data and the machine learning model, effectively amplifying the utility of limited training data
2Measurement precision
If data annotation is performed to create training datasets, then model training quality is improved, but the process becomes expensive and time consuming
Solution Approach 1:
The system implements self-service annotation by automatically extracting product attributes from product descriptions using natural language processing, automatically labeling transaction data with purchase outcomes, and generating training datasets without human intervention. This automated self-annotation process maintains high data quality while eliminating the time and cost associated with manual expert annotation
Solution Approach 2:
The patent replaces the mechanical process of manual data annotation with automated computational processes including NLP-based feature extraction, automated data labeling algorithms, and machine learning preprocessing pipelines. This substitution transforms the annotation process from a labor-intensive manual operation to an automated computational task, dramatically reducing time and cost while maintaining annotation quality
3Loss of information
If only in-house transaction data is used, then data privacy is maintained, but insight into underlying factors driving patient preference is limited
Solution Approach 1:
The system merges multiple data sources including in-house transaction data, publicly available product attribute data, and external consumer feedback data into a unified training dataset. This combination allows the model to learn from diverse information sources while maintaining data privacy through anonymization and aggregation techniques, thereby gaining comprehensive insights into patient preference drivers without compromising confidentiality
4Productivity
If machine learning models are trained to predict consumer preferences, then portfolio decision-making is improved, but the models suffer from overfitting due to insufficient training data
Solution Approach 1:
The system applies partial action by using regularization techniques and feature selection methods that focus on the most important predictors while excluding redundant features. This selective approach prevents the model from attempting to learn from all available features, thereby reducing overfitting risk while maintaining predictive accuracy for portfolio decision-making applications
Data Source
AI summary
Methods, systems, and apparatuses for predicting prescription drug products or substitution drug products for inclusion in a company's portfolio.


