Machine Learning Model for Clinical Trial Data Quality Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Clinical trial designers face challenges in designing visit schedules and content to collect sufficient data without overburdening participants, leading to issues with data quality due to confusing survey questions and participant confusion.
Innovation Solution
A computer-implemented method using trained machine learning models to evaluate data quality in clinical trials by analyzing query-related information from participants, incorporating parameters such as meta-factors, participant characteristics, and survey data points to predict a data quality score and provide suggestions for improvement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If more data points are collected and collected more frequently, then data quality improves, but participant burden increases leading to adverse effects on recruitment and retention
Solution Approach 1:
The system dynamically adjusts data collection parameters (frequency, number of data points) based on machine learning predictions of data quality and real-time monitoring of participant engagement metrics, transforming static survey designs into adaptive ones that optimize both data quality and participant burden
Solution Approach 2:
The system implements continuous feedback loops where participant queries and responses are monitored in real-time, fed into machine learning models that predict data quality outcomes, and used to adjust subsequent data collection strategies during the clinical trial
2Ease of operation
If fewer data points are collected or collected less frequently, then participant burden decreases, but data quality deteriorates
Solution Approach 1:
The system uses machine learning models to identify optimal parameter settings for data collection that maintain adequate data quality while minimizing participant burden, dynamically adjusting the balance between these competing objectives
Solution Approach 2:
The system replaces manual trial design adjustments with automated machine learning-based optimization, where algorithms analyze participant responses and queries to automatically determine optimal data collection frequency and intensity
3Productivity
If confusing survey questions are used, then data collection efficiency improves, but participant confusion increases leading to inappropriate responses
Solution Approach 1:
The system monitors participant queries about survey questions in real-time and uses this feedback to identify confusing questions, which are then flagged for revision or clarification to maintain both efficiency and response accuracy
Solution Approach 2:
The system uses machine learning models trained on historical query data to predict which survey questions are likely to cause confusion before they are administered, allowing pre-emptive refinement of question wording or structure
Data Source
AI summary
A method, computing platform, and computer program product are provided for evaluating data quality during a clinical trial. A computing platform receives, for a clinical trial, study design information including a set of parameters and corresponding parameter values related to data quality of the clinical trial. During the clinical trial, the computing platform receives query-related information associated with queries from at least some of a plurality of participants of the clinical trial. The computing platform applies the study design information and the query-related information to at least one trained machine learning model to calculate a predicted data quality score indicating data quality for the clinical trial. At least one suggestion for improving the data quality is determined and the predicted data quality score and the at least one suggestion for improving the data quality are output.


