Database Reduction for Clinical Trial Record Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Clinical trial databases become overly complex and inefficient due to duplicate and inconsistent entries of clinical trial investigators, leading to manual review challenges and errors, especially when data is sourced from multiple locations and merged over time.
Innovation Solution
A system utilizing computer algorithms for data processing, including Damerau-Levenshtein scoring and machine-learning models, to standardize and match descriptors, produce record scores, and selectively write records to a database, thereby reducing database size and improving data quality and processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual review of database entries is performed to identify and remove duplicates, then data quality can be maintained, but the process becomes time-consuming and error-prone
Solution Approach 1:
The patent replaces manual mechanical review processes with automated computer algorithms that perform data standardization, matching, and duplicate identification. The system uses automated scoring mechanisms to evaluate database records and identify duplicates, eliminating the need for time-consuming manual review while maintaining or improving data quality through consistent algorithmic application.
Solution Approach 2:
The database system performs self-service by automatically identifying and flagging duplicate records through algorithmic matching and scoring. The system autonomously evaluates data quality metrics, generates match scores, and highlights potential duplicates without requiring external manual intervention, thereby maintaining data quality while reducing time investment.
2Quantity of substance
If multiple data sources are merged to accumulate investigator information over time, then database comprehensiveness increases, but database complexity and duplicate entries increase
Solution Approach 1:
The patent extracts and isolates specific descriptor elements from multiple data sources that are relevant for identifying duplicates (such as name components, geographic information, specialty). By extracting only the critical matching elements rather than processing entire records, the system maintains comprehensiveness while reducing the complexity of comparison operations.
Solution Approach 2:
The system segments database records into standardized descriptor components (name parts, location, specialty, etc.) that can be independently matched and scored. This segmentation allows the system to handle multiple data sources systematically by comparing individual descriptor elements rather than entire complex records, thereby managing database complexity while accumulating comprehensive investigator information.
3Productivity
If database entries are standardized and matched using algorithms, then processing speed improves, but computational resources increase
Solution Approach 1:
The patent applies partial matching by focusing computational resources on key descriptor elements that are most indicative of duplicates (such as name components and geographic location) rather than analyzing every aspect of each record. The scoring system prioritizes certain descriptors over others, performing detailed algorithmic matching only on the most discriminative features, thereby improving processing speed while moderating computational resource consumption.
Data Source
AI summary
Aspects and features relate to computationally reducing the size or complexity of a database in order to improve the speed and efficiency with which such a database is processed by a computing system in order to identify investigators for clinical trials. In some aspects, a processing device performs operations including identifying data sources for geographically clustered data containing corresponding descriptors for database records. The operations further include formatting the corresponding descriptors to produce standardized, corresponding descriptors, and matching each standardized, corresponding descriptor to produce a record score for the descriptor. The record scores can be combined to produce an overall score for each database record and the database record can be selected and written to the data store based on the overall score.


