Machine Learning Model for Clinical Trial Data Quality Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinical trial designers face challenges in designing visit schedules and content to collect sufficient data without overburdening participants, leading to issues with data quality due to confusing survey questions and participant confusion.

Innovation Solution

A computer-implemented method using trained machine learning models to evaluate data quality in clinical trials by analyzing query-related information from participants, incorporating parameters such as meta-factors, participant characteristics, and survey data points to predict a data quality score and provide suggestions for improvement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If more data points are collected and collected more frequently, then data quality improves, but participant burden increases leading to adverse effects on recruitment and retention

Engineering Contradiction:
Improvedata qualityVSAvoidparticipant burden
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system dynamically adjusts data collection parameters (frequency, number of data points) based on machine learning predictions of data quality and real-time monitoring of participant engagement metrics, transforming static survey designs into adaptive ones that optimize both data quality and participant burden

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements continuous feedback loops where participant queries and responses are monitored in real-time, fed into machine learning models that predict data quality outcomes, and used to adjust subsequent data collection strategies during the clinical trial

Inventive Principle:
Principle #23Feedback

2Ease of operation

If fewer data points are collected or collected less frequently, then participant burden decreases, but data quality deteriorates

Engineering Contradiction:
Improveparticipant burdenVSAvoiddata quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system uses machine learning models to identify optimal parameter settings for data collection that maintain adequate data quality while minimizing participant burden, dynamically adjusting the balance between these competing objectives

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system replaces manual trial design adjustments with automated machine learning-based optimization, where algorithms analyze participant responses and queries to automatically determine optimal data collection frequency and intensity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If confusing survey questions are used, then data collection efficiency improves, but participant confusion increases leading to inappropriate responses

Engineering Contradiction:
Improvedata collection efficiencyVSAvoidresponse accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system monitors participant queries about survey questions in real-time and uses this feedback to identify confusing questions, which are then flagged for revision or clarification to maintain both efficiency and response accuracy

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system uses machine learning models trained on historical query data to predict which survey questions are likely to cause confusion before they are administered, allowing pre-emptive refinement of question wording or structure

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11651243B2Using machine learning to evaluate data quality during a clinical trial based on participant queries
Publication Date: 2023.05.16 MERATIVE US LP
  • US11651243B2 patent drawing
  • US11651243B2 patent drawing
  • US11651243B2 patent drawing

AI summary

A method, computing platform, and computer program product are provided for evaluating data quality during a clinical trial. A computing platform receives, for a clinical trial, study design information including a set of parameters and corresponding parameter values related to data quality of the clinical trial. During the clinical trial, the computing platform receives query-related information associated with queries from at least some of a plurality of participants of the clinical trial. The computing platform applies the study design information and the query-related information to at least one trained machine learning model to calculate a predicted data quality score indicating data quality for the clinical trial. At least one suggestion for improving the data quality is determined and the predicted data quality score and the at least one suggestion for improving the data quality are output.