Electronic binder system (EBINDER) for processing source data into EDC systems
The eBinder system automates clinical trial data conversion and secure transfer, addressing inefficiencies and security gaps in existing methods, enhancing data quality and compliance.
Patent Information
- Application Number
- JP2025511787
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-23
- Filing Date
- 2023-08-23
- Publication Date
- 2025-09-09
AI Technical Summary
Current clinical trial data processing methods are inefficient, prone to transcription errors, and fail to meet FDA requirements for electronic data capture, while manually masking PII is labor-intensive and existing security measures are inadequate for protecting sensitive data.
An eBinder system that automates the conversion of clinical trial source data into machine-readable format using OCR and NLP, encrypts and masks PII, and securely transfers data to EDC systems, ensuring compliance with regulatory standards.
Enhances data quality, reduces errors, and improves efficiency by eliminating manual transcription and ensuring secure, compliant data processing.
Smart Images

Figure 2025529899000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Application No. 63 / 400,158, filed August 23, 2022. The contents and disclosures of the aforementioned application are incorporated herein by reference in their entireties. Throughout this application, various publications are cited. The disclosures of these publications in their entireties are incorporated herein by reference in order to more fully describe the state of the art to which this invention pertains.
[0002] The present invention relates to a method and system for automatically and seamlessly processing source data in a clinical trial into a clinical trial's Electronic Data Capture (EDC) system. [Background technology]
[0003] Electronic data capture (EDC) systems are commonly used in clinical trials to collect, manage, and store clinical trial data. Traditionally, study sites collect patient data on site-specific source data capture forms (SDCFs) according to study protocols and store them in physical patient binders. Clinical research coordinators (CRCs) then manually enter data from the SDCFs into corresponding case report forms (CRFs) in a clinical trial data management system, typically called electronic data capture. This manual data entry process at the site level can lead to transcription errors and is highly inefficient. As a key quality control process in clinical trials, clinical research organizations (CROs) send clinical trial monitors, also known as clinical research associates (CRAs), to study sites to perform manual source data verification (SDV). SDV is a tedious, time-consuming, and expensive process.
[0004] In 2013, the U.S. Food and Drug Administration (FDA) issued a guidance document to industry on electronic source data in clinical trials [1], in which the agency listed requirements for collecting source data electronically and transmitting it in an electronic CRF (eCRF). These requirements included: Eliminating unnecessary duplication of data; -reducing the possibility of transcription errors; Encouraging subject entry of source data during visits, when appropriate; Eliminate transcription of source data before entering it into the eCRF; and Facilitating remote monitoring of data; Facilitating real-time access for data review; · Facilitate the collection of accurate and complete data.
[0005] There is no comprehensive method or system that can meet FDA requirements. TransCelerate BioPharma Inc., a nonprofit organization dedicated to guiding biopharmaceutical companies, has published two papers [2, 3] to encourage its member companies to begin developing methods for optimizing the use of electronically sourced data in clinical trials. The papers acknowledge that "data collection methods and technologies are not being utilized to their full potential, and transcription between electronic systems remains the standard."
[0006] Developing a system to automatically process source data into a remotely accessible EDC, thereby eliminating transcription errors and in-person SDV, would address a current unmet need in the field. Addressing such a need could significantly improve data quality and clinical trial efficiency, resulting in better clinical trials and enormous cost savings.
[0007] Site-level source data may contain personally identifiable information (PII) or personal health information, both of which must be protected in accordance with HIPAA, GDPR, and other regulations. Therefore, care must be taken when transmitting, processing, and storing such data.
[0008] Before processing source data, any PII contained in the source must be de-identified or desensitized. This step, called "masking," can be performed manually by field personnel such as CRCs or study coordinators (SCs). Masking often involves remediating sensitive data. However, manually masking source data typically requires a significant amount of work. Therefore, a means of automatically masking PII would be of great benefit to the art. Such masking must protect the PII from unauthorized access while still allowing SCs and CRAs to view the unmasked source data for verification purposes. Sensitive data may be transmitted or retrieved via secure processes such as HTTPS (Secure Communications HTTP Protocol for the Internet) or SFTP (Secure FTP Server Process). However, protecting sensitive data in transit is not enough. Sensitive data at rest on the server, or anywhere else, must also be protected. HTTPS or SFTP alone do not solve this security problem. Therefore, additional security measures are required.
[0009] The submitted source files are often in image format and may contain more information than required by the CRF. Challenges remain: how to convert images into machine-readable data, how to make the machine-readable data machine-understandable, and how to accurately input them into the EDC in order to automatically process the source data into the EDC.
[0010] Source data and source files are interchangeable in this application. Summary of the Invention
[0011] The present invention provides a method and system for transferring and directly processing clinical trial source data into a clinical trial database (EDC), including an eBinder system in electronic communication with a web application for receiving source data from trial sites, an encryption module for encrypting the uploaded source data, a PII masking module for automatically masking patient identifiable information (PII), an eBinder database for storing the data, an Optical Character Recognition (OCR) module for converting images (source files) into machine-readable plain text, and a Prompt Engineering (PE, i.e., in-context learning) module for a Generative Pre-Trained Transformer (GPT) to convert the plain text into machine-understandable data and populate corresponding data fields in a Case Report Form (CRF). [Brief explanation of the drawings]
[0012] [Figure 1] Figure 1 shows an overview of the eBinder process. [Figure 2] Figure 2 shows the eBinder system. [Figure 3] Figure 3 shows the structure of the eBinder system. [Figure 4] Figure 4 illustrates the patient registration process at the site level. [Figure 5] FIG. 5 illustrates the steps of encrypting and masking source data and files. [Figure 6] Figure 6 illustrates the source data upload, PII masking, and QC process. [Figure 7] Figure 7 shows the process of processing the image into a machine-readable plain text file. [Figure 8] Figure 8 shows the process of formatting machine-readable JSON into a tabular eSource Form (eSF). [Figure 9] Figure 9 shows the NPL / GPT process for adding eSource forms to an eCRF. [Figure 10] Figure 10 shows the Data Manager (DM) QC eSource file. [Figure 11] FIG. 11 shows a CRA performing remote source data verification (SDV). [Figure 12] Figure 12 shows the FDA or regulatory agency accessing eBinder for source data validation. [Figure 13] Figure 13 shows the processing of a JSON file into HTML. [Figure 14] Figure 14 shows a sample masked source file. [Figure 15] Figure 15 shows the transcription from the source file to the HTML tabular e-source form. [Figure 16] Figure 16 shows the transcription of the HTML tabular e-source form into the EDC. DETAILED DESCRIPTION OF THE INVENTION
[0013] The present invention provides a method and system for securely transferring clinical trial source data to electronic data capture (EDC) clinical trial systems. Essentially, the present invention protects private data, such as personally identifiable information (PII), and other sensitive information that needs to be protected when accessed from various devices and transferred between systems. The present invention enables various types of clinical trial EDC systems to use an "eBinder" system to upload clinical trial source data without compromising the security of the PII and other private information contained therein.
[0014] The present invention provides multiple processes, further described below, including, but not limited to, creating an electronic source data management platform called "eBinder" to which research sites can upload source data; automatically masking PII upon upload; encrypting the entire original source data stored in eBinder, which can only be viewed by CRCs and / or CRAs; converting images of the source data into machine-readable data; converting the machine-readable data into machine-understandable data; storing the converted data in a secure database; mapping the converted data to corresponding clinical databases; automatically validating the supplied data against the original source data; and creating access for regulatory inspection or third-party audits.
[0015] In one embodiment, the present invention provides a method for automatically and seamlessly converting clinical trial source data into a key-value structured dataset, which is then processed into an EDC dataset. The method includes the steps of: defining a file structure for storing source data in the eBinder system; extracting source data from multiple sources; encrypting the data using a dual-key encryption algorithm that uses a private key for PII held by the data owner and a shared key for non-PII; masking the PII with a CRC and quality control of the masking with a CRA; converting the masked source file image into machine-readable plain text in JavaScript Object Notation (JSON) format using OCR technology; converting the JSON format data into tabular machine-readable data in HyperText Markup Language (HTML) format using NLP technology; correcting formatting or spelling errors using NLP technology to convert the machine-readable HTML data into machine-understandable data; adding selected source data into the correct data filed in the CRF of the EDC using AI technology; displaying the source files and the converted data side-by-side to enable remote SDV; and providing a platform for verifying submitted data against the source files for regulatory review or audit purposes.
[0016] In one embodiment, the present invention provides a method for automatically and seamlessly converting clinical trial source data into a key-value structured dataset and processing it into an EDC dataset, the method including the steps of: defining a file structure for storing the source data in an eBinder system, extracting the source data from multiple sources, encrypting the source data using an encryption algorithm, masking PII and verifying the masking, converting source file images into machine-readable plain text using OCR technology, correcting any formatting or spelling errors in the machine-readable plain text using NPL, converting the machine-readable plain text into tabular machine-readable text using NPL, converting the tabular machine-readable plain text into machine-understandable plain text using NPL, adding eCRFs using the machine-understandable plain text, aligning source data associated with the added eCRFs, and validating the eCRFs against the associated source data.
[0017] In one embodiment, the source data is encrypted using a dual-key encryption algorithm.
[0018] In one embodiment, where the source data is encrypted using a dual-key encryption algorithm, one encryption key allows access to the unmasked data and a second encryption key allows access to the masked data.
[0019] In one embodiment, the machine-readable plain text is in JSON format.
[0020] In one embodiment, the tabular, machine-readable plain text is in HTML format.
[0021] In one embodiment, NPL processing is performed using the GPT process.
[0022] In one embodiment, the eCRF and associated source data are validated using a web-based platform.
[0023] In one embodiment, Figure 1 illustrates an overview of the entire Direct Source Data to EDC (DSDE) process for clinical trials.
[0024] When clinical trial documents are uploaded to the eBinder system, the original documents are encrypted with an advanced encryption algorithm and stored on a Virtual Private Cloud (VPC) server. The encryption key is available only to study administrators. While uploading the original clinical trial documents, the eBinder system also automatically masks all PII data and stores the masked documents on the VPC server. The masked documents are then sent to an Amazon Web Services (AWS) S3 bucket for optical character recognition (OCR) processing. The OCR process extracts plain text data from the masked documents in the S3 bucket and stores them in a cloud database that clinical trial researchers can search and view via the eBinder UI Portal.
[0025] In the present invention, the system includes an eBinder system (Figure 2) in electronic communication with a web application for receiving data from a clinical trial study, an encryption module (Figure 5) for encrypting documents, a PII masking module (Figure 6) for masking the data, an eBinder database that may be hosted as a cloud database for storing the data, an OCR module (Figure 7) for converting source file images into machine-readable plain text in JSON format, an HTML module (Figure 8) for formatting the machine-readable JSON file into a tabulated e-Source Form (eSF) having the same layout as the original source file, and a Natural Language Processing (NLP) / Generative Pre-Trained Transformer (GPT) module (Figure 9) for converting the eSF into machine-understandable data and adding selected source data to corresponding data fields in the eCRF and EDC dataset.
[0026] Once a study protocol is finalized, the eBinder system can be built according to the study schema defined in the protocol. The study schema consists of case report forms (CRFs) and the time points (visits) at which the CRFs will be collected. In one embodiment, the eBinder system includes an eBinder database and an operational module (Figure 1). The eBinder database is for storing source files and data and includes an eBinder folder structure according to the study schema, typically per visit and per form within a visit (Figure 2). Because a site may collect the same type of source data across all visits on one or more pages of the same type of form (e.g., vitals for all visits collected on one or more pages), a per form structure can be used instead. In one embodiment, the operational module includes study-level processes for adding sites, users, and patients. When a new patient is added, the system automatically creates an eBinder folder structure (Figure 3).
[0027] User roles and access controls are created in accordance with Good Clinical Practice (GCP) as defined in FDA guidance. The roles consist of (1) Principal Investigator (PI), (2) Clinical Research Coordinator (CRC), (3) Clinical Research Associate (CRA), and (4) Data Manager (DM). When authorized, only the PI, CRC, and / or CRA have the right to access unmasked PII information (Figure 1).
[0028] When a patient visits a study site, the PI and / or CRC typically manually record the data on site-specific Data Collection Forms (sDCFs) and maintain them in a patient binder (pBinder). The CRC then transcribes the data into corresponding case report forms (CRFs) in the EDC, often not immediately. This manual process at the site level is prone to transcription errors and is highly inefficient.
[0029] In one embodiment, the CRC uploads the source data to the eBinder system. The system automatically masks the PII information on the source file before storing it in the eBinder database. In one embodiment, the private key for masking the PII reports will be held by the research site. A system control key will be used for the non-PII portions of the source file (Figure 5). In one embodiment, the source file can only be reviewed by the CRC and the CRA. The uploaded source file may be distorted (rotated, scaled down, or zoomed in) during scanning, and the CRA will log into the system to verify that the PII information is completely masked (Figure 5).
[0030] Once the CRA completes the quality control (QC) review and saves the source file, the DM can view the masked source file. In one embodiment (FIG. 7), the masked source file is automatically pushed to an OCR processor to convert the image file into a machine-readable plain text file in JSON format.
[0031] In one embodiment (FIG. 8), machine-readable plain text files in JSON format are automatically pushed to an HTML module for conversion to eSourceForm (eSF) format. This conversion not only makes it easier for the DM to compare the machine-readable plain text with its corresponding source file, but also improves the efficiency of downstream processes. Compared to similar data in JSON format, eSF machine-readable plain text contains fewer redundant variables, lowering the likelihood of subsequent NLP / AI errors, reducing the likelihood that the machine-readable plain text will exceed any NLP / AI input limits, and reducing NLP / AI processor time.
[0032] In one embodiment (FIG. 13), the detailed process of the HTML module is illustrated.
[0033] In one embodiment (FIG. 9), the HTML formatted eSF is automatically pushed to an NLP / AI module for conversion into machine-understandable data tables and addition to the eCRF and EDC.
[0034] Once the original source file image is processed into EDC, pre-built logical edit checks will be processed. If there is a data query, the DM will perform a data logical review as shown in one embodiment (Figure 10).
[0035] Once the DM has completed the data review and all queries have been resolved, the CRA can perform a remote SDV as shown in one embodiment (FIG. 11).
[0036] In one embodiment, the present invention provides a computerized system for automatically and seamlessly converting clinical trial source data into multiple key-value structured datasets and processing them into EDC datasets and eCRFs. The system includes an eBinder system in electronic communication with a secure web application for receiving source data from a clinical trial study, an eBinder database for storing the data, an encryption module for encrypting source documents, a PII masking module for masking the source data, an OCR module for converting source file images into machine-readable plain text in JSON format, an HTML module for converting the JSON formatted plain text into tabular machine-readable plain text in HTML format, an NLP module for converting the HTML formatted machine-readable plain text into machine-understandable plain text, an AI module for adding the source data to the eCRF, a SDV module for displaying the source data and associated converted data side-by-side for remote SDV, and a viewing module for remotely accessing the source data for regulatory review or audit purposes.
[0037] In one embodiment, the present invention provides a system for automatically and seamlessly converting clinical trial source data into multiple key-value structured datasets and processing them into an EDC dataset and an eCRF. The system includes a non-transitory computer-readable storage medium storing computer program instructions defined by modules of the computerized system, a non-transitory computer-readable storage medium storing all data in a secure location, a limited-access file server with an encryption module for encryption and decryption, and a web-based interface for displaying masked source data and corresponding converted data side by side. The modules include a masking module for automatically masking PII, an OCR module for converting source file images into machine-readable plain text, a transposition module for converting the machine-readable plain text into tabular machine-readable text, an NLP module for converting the tabular machine-readable plain text into machine-understandable data, and a generation module for generating an EDC dataset and an eCRF from the machine-understandable data.
[0038] In one embodiment, the system's encryption module implements a dual-key encryption algorithm.
[0039] In one embodiment, the encryption module implements a dual-key encryption algorithm, where one encryption key allows access to unmasked data and the other encryption key allows access to masked data.
[0040] In one embodiment, the machine-readable plain text is in JSON format.
[0041] In one embodiment, the tabular, machine-readable plain text is in HTML format.
[0042] In one embodiment, the NPL module utilizes the GPT.
[0043] In one embodiment, blockchain technology is used to authenticate the submitted data and source files.
[0044] Throughout this application, various methods are implemented on a non-transitory computer-readable storage medium. Those skilled in the art will understand that a "non-transitory computer-readable storage medium" may refer to one or more storage media capable of storing instructions for execution by a processor. For example, a "non-transitory computer-readable storage medium" includes a hard drive, a solid-state drive, a random access memory (RAM), and similar media.
[0045] Throughout this application, various publications are cited by author and year of publication. Full citations for the publications are listed below. The disclosures of these publications in their entireties are incorporated herein by reference in order to more fully describe the state of the art to which this invention pertains.
[0046] The invention has been described in an illustrative manner, and it is to be understood that the terminology used is intended to be in the nature of words of description and not of limitation.
[0047] Many modifications and variations of the present invention are possible in light of the above teachings. It is therefore to be understood that, within the scope of the appended claims, the invention may be practiced otherwise than as specifically described. References [1] USFood and Drug Administration.(2013,September).Guidance for Industry:Electronic Source Data in Clinical Investigations.Retrieved from https: / / www.fda.gov / media / 85183 / download [2] Kellar,E.,Bornstein,S.,Caban,A.,Crouthamel,M.,Celingant,C.,McIntire,P.A.,...& Wilson,B.(2017).Optimizing the use of electronic data sources in clinical trials:the technology landscape.Therapeutic Innovation & Regulatory Science,51,551-567. [3] Kellar,E.,Bornstein,S.M.,Caban,A.,Celingant,C.,Crouthamel,M.,Johnson,C.,...& Wilson,B.(2016).Optimizing the use of electronic data sources in clinical trials:the landscape,part 1.Therapeutic Innovation & Regulatory Science,50(6),682-696. [4] TransCelerate Biopharma Inc.(2017).Issues Related to Non-CRF Data Practices.Retrieved from http: / / www.transceleratebiopharmainc.com / wp content / uploads / 2018 / 01 / eSource-Non-CRF-Data-Practices.pdf
Claims
1. 1. A method for automatically and seamlessly converting clinical trial source data into key-value structured data sets and processing the same into electronic data capture data sets (EDCs), the clinical trial source data including site source file images, electronic medical records, and electronic patient-reported outcomes for use in conducting a clinical trial, the method comprising: a. defining a file structure for storing source data in an eBinder system; b. extracting source data from a plurality of sources; c. Encrypting the data using a dual-key encryption algorithm that uses a private key held by the data owner for the Patient Identifiable Information and a shared key for the non-patient identifiable information; d. A Clinical Research Coordinator masks patient identifiable information and a Clinical Research Associate performs quality control of said masking; e. Converting the masked source file image in JavaScript Object Notation format into machine-readable plain text using Optical Character Recognition technology; F. Converting the data in JavaScript Object Notation format into tabular, machine-readable data in HyperText Markup Language format; g. correcting formatting or spelling errors using natural language processing techniques to convert the machine-readable HyperText Markup Language data into machine-understandable data; h. Using artificial intelligence techniques to add the selected source data to the correct data filed in the electronic data collection case report form; i. displaying the source file and the transformed data side-by-side to allow for remote source data verification; j. providing a platform for validating the submitted data against the source files for regulatory review or audit purposes.
2. 1. A method for automatically and seamlessly converting clinical trial source data into key-value structured datasets and processing the same into electronic data collection datasets, the clinical trial source data including site source file images, electronic medical records, and electronic patient-reported outcomes for use in conducting a clinical trial, the method comprising: a. defining a file structure for storing source data in an eBinder system; b. extracting source data from a plurality of sources; c. encrypting the source data using an encryption algorithm; d. masking patient identifiable information; e. verifying the masking of said patient identifiable information; f. converting said source file image into machine-readable plain text using optical character recognition techniques; g. using natural language processing to correct any formatting or spelling errors in said machine-readable plain text; h. converting said plain machine-readable text into tabular machine-readable text using natural language processing; i. converting said tabular machine-readable plain text into machine-understandable plain text using natural language processing; populating an electronic case report form using said machine understandable plain text; k. displaying the added electronic case report form and associated source data side-by-side; l. validating the electronic case report form against the associated source data.
3. The method of claim 2 , wherein the encryption algorithm is a dual-key encryption algorithm.
4. 4. The method of claim 3, wherein one encryption key allows access to unmasked data and a second encryption key allows access to masked data.
5. The method of claim 2 , wherein the machine-readable plain text is in JavaScript Object Notation format.
6. The method of any one of claims 2 to 5, wherein the tabular, machine-readable plain text is in HyperText Markup Language format.
7. The method of any one of claims 2 to 6, wherein the natural language processing is performed using a Generative Pre-Trained Transformer.
8. 8. The method of any one of claims 2 to 7, wherein the validation of the electronic case report form and associated source data is performed using a web-based platform.
9. 1. A computerized system for automatically and seamlessly transforming clinical trial source data into multiple key-value structured datasets and processing the same into electronic data collection datasets and electronic case report forms, wherein the source data includes site source file images, electronic medical records, and electronic patient-reported outcomes, all related to clinical trial subjects, the system comprising: a. an eBinder system in electronic communication with a secure web application for receiving source data from a clinical trial study; b. an eBinder database for storing data; c. an encryption module for encrypting the source document; d. a patient identifiable information masking module for masking the source data; e. an optical character recognition module for converting source file images in JavaScript Object Notation format into machine-readable plain text; f. a HyperText Markup Language module for converting said plain text in JavaScript Object Notation format into tabular, machine-readable plain text in HyperText Markup Language format; g. a natural language processing module for converting said machine-readable plain text in HyperText Markup Language format into machine-understandable plain text; h. An artificial intelligence module for adding source data to an electronic case report form; i. a source data validation module for displaying source data and associated transformed data side-by-side to perform remote source data validation; j. a viewing module for remotely accessing the source data for regulatory review or audit purposes.
10. 1. A computerized system for automatically and seamlessly transforming clinical trial source data into multiple key-value structured datasets and processing the same into electronic data collection datasets and electronic case report forms, wherein the source data includes site source file images, electronic medical records, and electronic patient-reported outcomes, all related to clinical trial subjects, the system comprising: a. a non-transitory computer-readable storage medium storing computer program instructions defined by the modules of said computerized system; b. at least one computing unit coupled to the non-transitory computer-readable storage medium and configured to execute computer program instructions defined by the modules of the computerized system, the modules comprising: i. a cryptographic module for encrypting and decrypting data; ii. a masking module for automatically masking patient identifiable information; iii. An optical character recognition module for converting the source file image into machine-readable plain text; iv. a transposition module for converting said machine-readable plain text into tabular machine-readable text; v. a natural language processing module for converting said tabular, machine-readable plain text into machine-understandable data; vi. a generation module for generating an electronic data collection dataset and an electronic case report form from the machine-understandable data; A computing unit; c. a non-transitory computer-readable storage medium storing all data located on a secure, access-restricted file server; d. A web-based interface for displaying the masked source data and the corresponding transformed data side-by-side.
11. The system of claim 10 , wherein the encryption module implements a dual-key encryption algorithm.
12. 12. The system of claim 11, wherein one encryption key allows access to unmasked data and another encryption key allows access to masked data.
13. 13. The system of claim 10, wherein the machine-readable plain text is in JavaScript Object Notation format.
14. 14. The system of claim 10, wherein the tabular, machine-readable plain text is in HyperText Markup Language format.
15. The system of any one of claims 10 to 14, wherein the natural language processing module utilizes a Generative Pre-Trained Transformer.