Database Reduction for Clinical Trial Record Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinical trial databases become overly complex and inefficient due to duplicate and inconsistent entries of clinical trial investigators, leading to manual review challenges and errors, especially when data is sourced from multiple locations and merged over time.

Innovation Solution

A system utilizing computer algorithms for data processing, including Damerau-Levenshtein scoring and machine-learning models, to standardize and match descriptors, produce record scores, and selectively write records to a database, thereby reducing database size and improving data quality and processing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual review of database entries is performed to identify and remove duplicates, then data quality can be maintained, but the process becomes time-consuming and error-prone

Engineering Contradiction:
Improvedata qualityVSAvoidreview time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical review processes with automated computer algorithms that perform data standardization, matching, and duplicate identification. The system uses automated scoring mechanisms to evaluate database records and identify duplicates, eliminating the need for time-consuming manual review while maintaining or improving data quality through consistent algorithmic application.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The database system performs self-service by automatically identifying and flagging duplicate records through algorithmic matching and scoring. The system autonomously evaluates data quality metrics, generates match scores, and highlights potential duplicates without requiring external manual intervention, thereby maintaining data quality while reducing time investment.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If multiple data sources are merged to accumulate investigator information over time, then database comprehensiveness increases, but database complexity and duplicate entries increase

Engineering Contradiction:
Improvedatabase comprehensivenessVSAvoiddatabase complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent extracts and isolates specific descriptor elements from multiple data sources that are relevant for identifying duplicates (such as name components, geographic information, specialty). By extracting only the critical matching elements rather than processing entire records, the system maintains comprehensiveness while reducing the complexity of comparison operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments database records into standardized descriptor components (name parts, location, specialty, etc.) that can be independently matched and scored. This segmentation allows the system to handle multiple data sources systematically by comparing individual descriptor elements rather than entire complex records, thereby managing database complexity while accumulating comprehensive investigator information.

Inventive Principle:
Principle #1Segmentation

3Productivity

If database entries are standardized and matched using algorithms, then processing speed improves, but computational resources increase

Engineering Contradiction:
Improveprocessing speedVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial matching by focusing computational resources on key descriptor elements that are most indicative of duplicates (such as name components and geographic location) rather than analyzing every aspect of each record. The scoring system prioritizes certain descriptors over others, performing detailed algorithmic matching only on the most discriminative features, thereby improving processing speed while moderating computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11899692B2Database reduction based on geographically clustered data to provide record selection for clinical trials
Publication Date: 2024.02.13 LABORATORY CORPORATION OF AMERICA HOLDINGS INC
  • US11899692B2 patent drawing
  • US11899692B2 patent drawing
  • US11899692B2 patent drawing

AI summary

Aspects and features relate to computationally reducing the size or complexity of a database in order to improve the speed and efficiency with which such a database is processed by a computing system in order to identify investigators for clinical trials. In some aspects, a processing device performs operations including identifying data sources for geographically clustered data containing corresponding descriptors for database records. The operations further include formatting the corresponding descriptors to produce standardized, corresponding descriptors, and matching each standardized, corresponding descriptor to produce a record score for the descriptor. The record scores can be combined to produce an overall score for each database record and the database record can be selected and written to the data store based on the overall score.