Clinical Trial Data Abstraction from Unstructured Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinical trial protocol documents are typically unstructured, making them difficult to analyze and reuse, leading to inefficiencies in generating new protocol documents and manual searches for specific data.

Innovation Solution

A computer-implemented method converts unstructured clinical trial documents into a structured format using a data transformation system, identifying template-mapped, high-interest, and non-template-mapped content, generating a dictionary data structure with key-value pairs in JSON format for easy analysis and reuse.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If unstructured clinical trial documents are used, then the documents can be created and stored easily, but they cannot be easily analyzed or reused

Engineering Contradiction:
Improveease of document creationVSAvoidease of data analysis
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The patent introduces an intermediary processing system that acts as a mediator between unstructured document creation and structured data analysis. The system includes components for format conversion (e.g., XML to HTML), template matching, and data extraction that transform unstructured documents into structured formats with dictionary data structures, enabling both easy creation and easy analysis of clinical trial documents

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by transforming the structural parameters of documents from unstructured to structured formats. This involves changing the organization parameters through template mapping, data extraction, and serialization to JSON format, thereby improving analyzability while preserving the ease of initial document creation

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If unstructured protocol documents are used, then document creation is flexible, but manual searches for specific data are required

Engineering Contradiction:
Improvedocument creation flexibilityVSAvoidtime for manual data search
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing automated data extraction, template matching, and structuring operations during the document processing phase rather than requiring manual searches later. The system pre-structures documents with dictionary data structures and JSON serialization, making specific data immediately searchable without manual intervention

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical manual search process with an automated computer-implemented system that uses template matching, pattern recognition, and structured data queries to locate specific clinical trial data, eliminating the need for manual searching while preserving document creation flexibility

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If manual processing of clinical trial documents is used, then data accuracy can be maintained, but productivity is reduced

Engineering Contradiction:
Improvedata accuracyVSAvoiddocument processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements self-service by designing an automated system that performs data extraction, validation, and structuring operations independently without requiring manual review. The template-matching algorithm and dictionary data structure generation process automatically ensure data accuracy while dramatically improving processing productivity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the automated processing system validates extracted data against predefined templates and schemas, providing immediate feedback on data quality and accuracy. This automated feedback loop maintains reliability while enabling high-speed processing that would be impossible with manual methods

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4348444B1Techniques for abstraction of unstructured clinical trial health data
Publication Date: 2025.10.29 GENENTECH INC
  • EP4348444B1 patent drawingFigure 1
  • EP4348444B1 patent drawingFigure 2
  • EP4348444B1 patent drawingFigure 3

AI summary

The present disclosure relates to techniques for abstraction of unstructured clinical trial health data. Particularly, aspects are directed to obtaining, by a data transformation system, an unstructured document that includes various types of content for a given event. The data transformation system converts the unstructured document from an original format into a standardized format. The data transformation system processes the unstructured document in the standardized format to identify template-mapped content data and processes the template-mapped content data to identify high-interest content data. The data transformation system generates a dictionary data structure comprising key -value pairs for each of the template-mapped content data and the high-interest content data. The data transformation system generates a structured document based on the dictionary data structure for each of the template-mapped content data and the high- interest content data. Generating the structured document comprises serializing the key-value pairs to data objects in an open standard file format.