Schema Search for Clinical Trial Data Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Clinical trial data sets from different trials often have varying formats, making it difficult for researchers to identify and compare corresponding data due to differences in labels, data types, variable names, domains, applications, or protocols, requiring manual searches or limited analysis from individual studies.
Innovation Solution
A system and method that searches a schema to identify and visualize corresponding columns with different labels and data types by receiving a search query, determining results based on machine learning, and generating graphical representations of data distributions to normalize and filter data, enabling comparison across multiple clinical trials.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual search methods are used to identify corresponding columns across different data sets, then researchers can potentially find matching data, but the process requires significant time and effort
Solution Approach 1:
The patent replaces manual mechanical search processes with an automated machine learning system that uses natural language processing and algorithmic matching to identify corresponding columns across different data sets, dramatically reducing the time required while maintaining or improving accuracy
Solution Approach 2:
The patent introduces an intermediary machine learning model that acts as a mediator between different data sets with varying formats, translating and matching columns based on learned patterns rather than requiring direct manual comparison between all possible column pairs
2Adaptability or versatility
If data from multiple clinical trials with different formats is analyzed, then more comprehensive results can be obtained, but the complexity of handling different labels, data types, and protocols increases
Solution Approach 1:
The patent creates a universal machine learning-based data matching system that can handle multiple data formats, labels, and protocols through a single platform, eliminating the need for separate processing systems for each data type while managing complexity through automated algorithms
Solution Approach 2:
The patent transforms the complexity of handling diverse data formats by changing the approach from rule-based manual mapping to parameter-based machine learning models that automatically adapt to different data characteristics, reducing the perceived complexity for users
3Reliability
If columns with different labels and data types are normalized to enable comparison, then data compatibility improves, but additional processing steps are required
Solution Approach 1:
The patent performs preliminary normalization and standardization of data columns during the initial machine learning matching process rather than as a separate subsequent step, integrating the compatibility adjustment into the core matching algorithm to reduce the total number of discrete processing steps
Data Source
AI summary
Systems and methods are provided for searching a schema to identify and visualize corresponding data. The system may be configured to receive a search query and search the schema for columns that correspond to one another. Individual ones of the corresponding columns may have different labels or different data types. Results for the search query may be determined. The results may include a first set of columns that satisfy the search query including a first column and a second column that correspond to each other but have one or more different labels or different data types. The system may be configured to generate a graphical representation of a data distribution for the first column responsive to receiving a selection of the first column.


