Schema Search for Clinical Trial Data Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinical trial data sets from different trials often have varying formats, making it difficult for researchers to identify and compare corresponding data due to differences in labels, data types, variable names, domains, applications, or protocols, requiring manual searches or limited analysis from individual studies.

Innovation Solution

A system and method that searches a schema to identify and visualize corresponding columns with different labels and data types by receiving a search query, determining results based on machine learning, and generating graphical representations of data distributions to normalize and filter data, enabling comparison across multiple clinical trials.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual search methods are used to identify corresponding columns across different data sets, then researchers can potentially find matching data, but the process requires significant time and effort

Engineering Contradiction:
Improveaccuracy of identifying corresponding dataVSAvoidtime required for manual search
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical search processes with an automated machine learning system that uses natural language processing and algorithmic matching to identify corresponding columns across different data sets, dramatically reducing the time required while maintaining or improving accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary machine learning model that acts as a mediator between different data sets with varying formats, translating and matching columns based on learned patterns rather than requiring direct manual comparison between all possible column pairs

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data from multiple clinical trials with different formats is analyzed, then more comprehensive results can be obtained, but the complexity of handling different labels, data types, and protocols increases

Engineering Contradiction:
Improveability to handle multiple data formatsVSAvoidcomplexity of data processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal machine learning-based data matching system that can handle multiple data formats, labels, and protocols through a single platform, eliminating the need for separate processing systems for each data type while managing complexity through automated algorithms

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms the complexity of handling diverse data formats by changing the approach from rule-based manual mapping to parameter-based machine learning models that automatically adapt to different data characteristics, reducing the perceived complexity for users

Inventive Principle:
Principle #35Parameter changes

3Reliability

If columns with different labels and data types are normalized to enable comparison, then data compatibility improves, but additional processing steps are required

Engineering Contradiction:
Improvecompatibility of corresponding dataVSAvoidnumber of processing steps
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary normalization and standardization of data columns during the initial machine learning matching process rather than as a separate subsequent step, integrating the compatibility adjustment into the core matching algorithm to reduce the total number of discrete processing steps

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11494454B1Systems and methods for searching a schema to identify and visualize corresponding data
Publication Date: 2022.11.08 PALANTIR TECHNOLOGIES INC
  • US11494454B1 patent drawing
  • US11494454B1 patent drawing
  • US11494454B1 patent drawing

AI summary

Systems and methods are provided for searching a schema to identify and visualize corresponding data. The system may be configured to receive a search query and search the schema for columns that correspond to one another. Individual ones of the corresponding columns may have different labels or different data types. Results for the search query may be determined. The results may include a first set of columns that satisfy the search query including a first column and a second column that correspond to each other but have one or more different labels or different data types. The system may be configured to generate a graphical representation of a data distribution for the first column responsive to receiving a selection of the first column.