Clinical Trial Data Abstraction from Unstructured Documents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Clinical trial protocol documents are typically unstructured, making them difficult to analyze and reuse, leading to inefficiencies in generating new protocol documents and manual searches for specific data.
Innovation Solution
A computer-implemented method converts unstructured clinical trial documents into a structured format using a data transformation system, identifying template-mapped, high-interest, and non-template-mapped content, generating a dictionary data structure with key-value pairs in JSON format for easy analysis and reuse.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If unstructured clinical trial documents are used, then the documents can be created and stored easily, but they cannot be easily analyzed or reused
Solution Approach 1:
The patent introduces an intermediary processing system that acts as a mediator between unstructured document creation and structured data analysis. The system includes components for format conversion (e.g., XML to HTML), template matching, and data extraction that transform unstructured documents into structured formats with dictionary data structures, enabling both easy creation and easy analysis of clinical trial documents
Solution Approach 2:
The patent applies parameter changes by transforming the structural parameters of documents from unstructured to structured formats. This involves changing the organization parameters through template mapping, data extraction, and serialization to JSON format, thereby improving analyzability while preserving the ease of initial document creation
2Adaptability or versatility
If unstructured protocol documents are used, then document creation is flexible, but manual searches for specific data are required
Solution Approach 1:
The patent applies preliminary action by performing automated data extraction, template matching, and structuring operations during the document processing phase rather than requiring manual searches later. The system pre-structures documents with dictionary data structures and JSON serialization, making specific data immediately searchable without manual intervention
Solution Approach 2:
The patent replaces the mechanical manual search process with an automated computer-implemented system that uses template matching, pattern recognition, and structured data queries to locate specific clinical trial data, eliminating the need for manual searching while preserving document creation flexibility
3Reliability
If manual processing of clinical trial documents is used, then data accuracy can be maintained, but productivity is reduced
Solution Approach 1:
The patent implements self-service by designing an automated system that performs data extraction, validation, and structuring operations independently without requiring manual review. The template-matching algorithm and dictionary data structure generation process automatically ensure data accuracy while dramatically improving processing productivity
Solution Approach 2:
The patent incorporates feedback mechanisms where the automated processing system validates extracted data against predefined templates and schemas, providing immediate feedback on data quality and accuracy. This automated feedback loop maintains reliability while enabling high-speed processing that would be impossible with manual methods
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure relates to techniques for abstraction of unstructured clinical trial health data. Particularly, aspects are directed to obtaining, by a data transformation system, an unstructured document that includes various types of content for a given event. The data transformation system converts the unstructured document from an original format into a standardized format. The data transformation system processes the unstructured document in the standardized format to identify template-mapped content data and processes the template-mapped content data to identify high-interest content data. The data transformation system generates a dictionary data structure comprising key -value pairs for each of the template-mapped content data and the high-interest content data. The data transformation system generates a structured document based on the dictionary data structure for each of the template-mapped content data and the high- interest content data. Generating the structured document comprises serializing the key-value pairs to data objects in an open standard file format.