System and method for building an electronic data capture (EDC) system
An AI-driven EDC build model automates the generation of clinical trial specifications from research protocol documents, reducing setup time to 1-2 weeks and improving accuracy and scalability for pharmaceutical companies.
Patent Information
- Application Number
- JP2024145051
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-09-22
- Filing Date
- 2024-08-27
- Publication Date
- 2025-12-15
- Estimated Expiration
- 2044-08-27
AI Technical Summary
Designing a new study plan for each clinical trial or modifying an existing one is time-consuming, typically taking 8-10 weeks, delaying the start of clinical trials and extending the time to market for new drugs, and requires significant manual effort with a risk of human error and inconsistencies.
An EDC build model trained to understand textual research protocol documents generates machine-readable specifications for EDC systems, reducing the setup time to 1-2 weeks by automating the process and minimizing human error.
The model significantly reduces setup time, improves accuracy and consistency, and facilitates simultaneous setup of multiple studies, enhancing scalability and data interoperability for large pharmaceutical companies.
Smart Images

Figure 0007785877000001 
Figure 0007785877000002 
Figure 0007785877000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to building an electronic data capture (EDC) system from research protocol documents. [Background technology]
[0002] Electronic data capture (EDC) systems provide the framework for designing and managing clinical trials, collecting data from patients and investigators, and ensuring compliance with various requirements. Summary of the Invention [Problem to be solved by the invention]
[0003] However, designing a new study plan for each clinical trial or modifying (e.g., amending) an existing study plan is time-consuming, typically requiring approximately 8–10 weeks. This delays the start of clinical trials and extends the time to market for new drugs. This process also requires significant manual effort and expertise, which can lead to human error and inconsistencies that can negatively impact the quality of collected data and the reliability of clinical trials. While efforts are underway to standardize and automate parts of the EDC setup process, such as the use of Clinical Data Interchange Standards Consortium (CDISC) standards, these solutions still rely heavily on manual effort.
[0004] Clinical trials tend to have diverse requirements. Each trial is unique and has different requirements. EDC systems must accommodate different study designs, study endpoints, patient populations, and more. Understanding these diverse requirements and configuring an EDC system accordingly requires significant effort. While EDCs have standard components, it's common for studies to require custom modules and functionality. Furthermore, implementing real-time data validation rules specific to each study is critical for data quality but requires careful design and testing to ensure inconsistencies, outliers, and / or missing data are quickly flagged. Furthermore, because traditional EDC study construction processes are largely manual, these approaches lack the scalability needed to efficiently manage multiple studies simultaneously, especially for large pharmaceutical companies conducting multiple clinical trials simultaneously. [Means for solving the problem]
[0005] Disclosed herein is an approach that overcomes the above-mentioned challenges by using an EDC build model that is trained to understand textual research protocol documents that do not have a standardized format or structure, and that generates machine-readable specifications for EDC systems (i.e., specifications for operational databases, architectures, or schemas) from such research protocol documents.
[0006] In one aspect, the disclosed embodiments provide a method, system, and computer-readable medium for building an EDC (Electronic Data Capture) system. The method includes forming a training dataset from a protocol document set and a corresponding EDC build set. The method further includes fine-tuning a pre-trained language model using the protocol document set and the corresponding EDC build set of the training dataset to generate an EDC build model. The protocol document is input into the EDC build model, and an EDC build prediction is generated. A data structure of the EDC system is generated based at least in part on the EDC build prediction.
[0007] Other embodiments may include one or more of the following features, alone or in combination.
[0008] Generating a data structure for the EDC system may include generating one or more electronic case report forms, generating one or more edit checks for checking data entered via the one or more electronic case report forms, and generating one or more folder structures for managing EDC system information. Fine-tuning the pre-trained language model using the set of protocol documents and the EDC build of the corresponding training dataset may include adjusting hyperparameters of the pre-trained language model. The protocol documents may be natural language text documents. The protocol documents may include a tabular evaluation schedule. The EDC build predictions may be in the form of an extensible markup language (XML) document. Forming the training dataset may include dividing the set of protocol documents and the corresponding EDC builds into a training dataset and a test dataset. Forming the training dataset may further include inputting the protocol documents of the test dataset into an EDC build model to generate EDC build predictions, and analyzing the EDC build predictions generated from the protocol documents of the test dataset to determine accuracy of the EDC build model. The method may include fine-tuning the EDC build model based at least in part on the analysis of the EDC build predictions generated from the protocol documents of the testing dataset. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates a system for building an electronic data capture (EDC) system according to disclosed embodiments. [Figure 2] FIG. 1 shows an example of an excerpt from a study protocol document discussing the study's inclusion criteria in natural language text. [Figure 3] FIG. 3 shows a portion of a simulated EDC build in XML specification format generated by the EDC build model based on the excerpt of the research protocol document shown in FIG. 2 . [Figure 4]FIG. 3 shows a portion of a simulated EDC build in XML specification format generated by the EDC build model based on the excerpt of the research protocol document shown in FIG. 2 . [Figure 5] This figure shows the EDC construct specified in the XML in Figure 3 converted into the format of a "Fields" spreadsheet that is part of an alternative spreadsheet specification for the EDC construct. Each row in this spreadsheet defines a specific field in the eCRF form. [Figure 6] 5 is a diagram illustrating the XML-specified EDC Build of FIG. 4 converted into the format of a "Data Dictionary Entries" spreadsheet included in an alternative spreadsheet-specified version of the EDC Build. [Figure 7] This is an example excerpt from a study protocol document that explains the study's screening rejection criteria in natural language. [Figure 8] FIG. 8 shows a portion of a simulated EDC build in XML specification format generated by the EDC build model based on the excerpt from the research protocol document shown in FIG. 7. [Figure 9] FIG. 9 illustrates the XML specification EDC Build of FIG. 8 converted into the form of a "Custom Functions" spreadsheet that is part of an alternative spreadsheet-specific version of the EDC Build. [Figure 10] This figure shows an example of an excerpt from a research protocol document in tabular form, which expresses the evaluation schedule of a trial in natural language. [Figure 11] FIG. 11 illustrates a portion of a simulated EDC build in XML specification format generated by the EDC build model based on the excerpt of the protocol document shown in FIG. 10 . [Figure 12] FIG. 11 illustrates a portion of a simulated EDC build in XML specification format generated by the EDC build model based on the excerpt of the protocol document shown in FIG. 10 . [Figure 13] 12 illustrates the XML-specified EDC construct of FIG. 11 converted into the format of a "folder" spreadsheet that is part of an alternative spreadsheet-specified version of the EDC construct. [Figure 14] FIG. 13 illustrates the XML-specified EDC build of FIG. 12 converted into the format of a "BASE matrix" spreadsheet included in an alternative spreadsheet-specified version of the EDC build. [Figure 15] FIG. 1 illustrates a method for building an electronic data capture (EDC) system. [Figure 16] FIG. 1 illustrates further aspects of a method for building an electronic data capture (EDC) system. [Figure 17] FIG. 10 illustrates a user interface screen showing the generation of edit checks related to adverse event dates using a demonstration EDC build model. [Figure 18] FIG. 10 illustrates a user interface screen for manually entering edit check elements related to adverse event dates for a demonstration EDC build. DETAILED DESCRIPTION OF THE INVENTION
[0010] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the present invention. However, it will be understood by those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as to clearly explain the present invention.
[0011] The disclosed embodiments address technical challenges associated with building electronic data capture (EDC) systems to significantly reduce the time and manual effort required to set up new clinical trials in EDC systems. In some cases, the typical timeline for EDC study builds—8 to 10 weeks—can be reduced to a fraction of that period, e.g., 1 to 2 weeks or less. This allows clinical trials to begin sooner, resulting in an overall reduction in time to market for new drugs. The EDC build model described herein reduces the potential for human error in the EDC study build process and improves the accuracy and consistency of the generated database. Furthermore, the disclosed approach facilitates the simultaneous setup of multiple EDC study builds, providing scalability essential for large pharmaceutical companies running multiple clinical trials simultaneously.
[0012] The disclosed embodiments provide artificial intelligence (AI) models that learn and improve over time, thereby improving their ability to handle increasingly complex clinical study protocols. This adaptability is a significant advantage as clinical trials continue to grow in complexity and evolve. Furthermore, by using AI models, the disclosed approach also offers the potential for more standardized study construction, thereby increasing data interoperability and streamlining the regulatory compliance process. Thus, compared to conventional techniques for conducting EDC study construction, the disclosed embodiments offer significant advances in speed, accuracy, efficiency, and scalability, providing valuable innovation in the field of clinical trials.
[0013] FIG. 1 is a block diagram of a system 100 for building an electronic data capture (EDC) system. The system uses artificial intelligence (AI) on a protocol document and a corresponding EDC build dataset 110. The EDC build corresponds to the protocol document in that it is derived from the protocol document, i.e., through manual effort or using a model trained according to the disclosed embodiments. A study protocol document is a text document, e.g., natural language text, that includes information about the study design, objectives, patient cohort, procedures, etc. (See, e.g., Figures 2, 7, and 10 for example excerpts from a study protocol). The EDC build, on the other hand, is a machine-readable specification and / or schema, such as an extensible markup language (XML) specification (See, e.g., Figures 3, 4, 8, 11, and 12 for example examples of the XML specification portion of an EDC build). The EDC build effectively serves as an operational database that defines the various data structures used in the study, such as electronic case report forms (eCRFs), data fields for the eCRFs, edit checks to validate input data, and folder structures for storing input data. Traditionally, EDC builds are created manually based on research protocols.
[0014] The dataset 110 (and similarly the set of protocol documents 160) is generated by receiving as input a study protocol. The study protocol is typically a text document (e.g., a PDF or Word document) that outlines all aspects of a clinical trial, including the objectives, design, methodology, statistical considerations, and organization. The input document is preprocessed to extract and structure relevant information. Preprocessing may include steps such as tokenization (i.e., dividing text into individual words or phrases), lemmatization (i.e., reducing words to their base forms or roots), and stop-word removal (i.e., filtering out common words that do not contribute to meaning).
[0015] Dataset 110 is split into a training dataset 120 and a testing dataset 130. For example, 80% of dataset 110 may be used for training and 20% may be set aside for testing. An EDC build of the protocol document set and corresponding training dataset 120 is used to fine-tune a pre-trained language model 140 to generate an EDC-built model 150. In an embodiment, pre-trained language model 140 may be a widely available natural language processing (NLP) model, such as Google Research's Pathways Language Model (PaLM) or Meta's Large Language Model Meta AI (LLaMA).
[0016] Fine-tuning a pre-trained model involves further training the model using specific training data, such as the protocol document and corresponding EDC build dataset 110, to adjust the model's parameters to better perform on the specific dataset. During this process, the model learns to associate inputs (e.g., clinical study protocol documents) with outputs (e.g., the corresponding EDC study build). This is done by adjusting internal parameters to minimize the discrepancy between predictions in the training data and the actual outputs. In this embodiment, fine-tuning the pre-trained language model 140 may include, for example, adjusting hyperparameters of the pre-trained language model 140. This approach adapts the model to a specific task or domain with a smaller dataset while leveraging the extensive knowledge captured during pre-training on a large dataset.
[0017] The dashed line between language model 140 and EDC build model 150 in Figure 1 indicates that once language model 140 is fine-tuned, it is used as the runtime model, i.e., EDC build model 150 in this example. EDC build model 150 is trained on a large number of clinical research protocols and corresponding EDC study builds. This training enables EDC build model 150 to make accurate predictions for new, unproven research protocols.
[0018] In runtime operations, protocol documents are obtained from protocol document dataset 160 and input into EDC build model 150 to generate EDC build prediction 170. Protocol document dataset 160 may include, for example, unconfirmed protocols for research that are in the EDC system design phase. In embodiments, EDC build prediction 170 may be manually analyzed to assess how well it meets expectations. Adjustments may then be made to improve the accuracy of EDC build model 150 based on this analysis.
[0019] In embodiments, users interact with system 100 through a user interface that allows them to upload study protocols, initiate processing, and download generated EDC study builds. Users can also provide feedback on the generated output, which can be used to further improve the system's performance. The user interface may include communication elements, such as a chatbot, to explain various aspects of the built study (e.g., how to configure edit checks).
[0020] The EDC build prediction 170 is input to an EDC system architect 180, such as Medidata Architect®. The EDC system architect 180 uses the information in the EDC build prediction 170 to generate data structures for the EDC system 190, such as electronic case report forms (eCRFs), data fields for various forms, edit checks to validate input data, and folder structures for storing input data. In embodiments, the EDC build prediction 170 is an Architecture Loader Specification (ALS) file in extensible markup language (XML) format, which specifies forms, fields, etc. in a specific format that Medidata Architect® can use to generate data structures for the Medidata Rave® EDC system. In embodiments, the EDC system architect 180 can be incorporated as part of the EDC system 190.
[0021] In an embodiment, a study can be designed using tools provided by software in an EDC system, such as Medidata Rave®. Medidata Rave® provides a user interface screen for defining fields, rules, and so forth (see, e.g., Figure 18). The completed study design, or study build, can be downloaded for review and approval. The downloadable study build file, called an Architect Loader Specification (ALS) file, may be in the form of an XML specification (sometimes embodied in an XML file). The XML file can be opened using either a text editor or a spreadsheet application (e.g., Microsoft Excel). When an ALS file is opened in a text editor, it appears in the format shown in Figures 3, 4, 8, 11, and 12. When an ALS file is opened as a spreadsheet, i.e., converted to spreadsheet format, it appears in the format shown in Figures 5, 6, 9, 13, and 14. (In implementation, the spreadsheets shown in these figures may be individual tabs within a spreadsheet "workbook" that, combined, constitute the ALS for a particular study build.) As mentioned above, the spreadsheet format allows the XML files for study builds to be more easily inspected and modified. In either case, the downloaded study build files can be used to create and / or update studies. This allows users to reuse and modify existing study designs without having to manually enter design parameters.
[0022] As described above, data set 110 is divided into training data set 120 and test data set 130. In an embodiment, the protocol documents of test data set 120 can be input to trained EDC build model 150 to generate EDC build predictions 170. EDC build analyzer 175 receives EDC build predictions 170 based on test data set 130 and also receives corresponding protocol documents directly from test data set 130. This allows EDC build analyzer 175 to analyze EDC build predictions 170 generated from the protocol documents of test data set 130 to evaluate the accuracy of EDC build model 150. Results of analysis by EDC build analyzer 175 based on test data set 130 can be used to further fine-tune EDC build model 150.
[0023] As described above, a set of study protocol documents and a corresponding set of EDC builds are used as the dataset 110 for fine-tuning the pre-trained language model 140. The EDC build defines various data structures used in the EDC system 190, such as electronic case report forms (eCRFs), data fields for the eCRFs, edit checks to verify the validity of input data, and folder structures for storing input data. The EDC build may be obtained from a trial management software system, such as Medidata Rave®. In an embodiment, the EDC build may be an Architecture Loader Specification (ALS) file in XML format.
[0024] A study protocol is a text document, e.g., natural language text, that contains information about the overall study design, objectives, patient cohort, procedures, etc. For example, a study protocol might include the following: background and rationale, objectives (e.g., primary, secondary, exploratory), type (e.g., randomized, double-blind, placebo-controlled), number of subjects, duration and phases, eligibility and exclusion criteria, assessments and procedures (e.g., details of medical examinations and laboratory tests, schedule of events or visits, etc.), subject treatment (e.g., investigational drug, dose, mode of administration), drug efficacy and safety monitoring procedures, adverse events and data management and statistical methods, quality control and quality assurance, ethical considerations, publication and data sharing policies, references to relevant scientific literature, and various forms (e.g., informed consent, questionnaires, surveys, etc.).
[0025] Although research protocol documents for different studies generally contain similar information, such documents do not have universally standardized structure and / or content. Furthermore, such documents can vary significantly with regard to text formatting, font, punctuation, etc. (e.g., see excerpts of research protocol documents in Figures 2, 7, and 10). Research protocol documents are not intended to be used solely for study design but to be read by people in various roles within the study, e.g., study designers, regulators, and clinicians.
[0026] In implementation, study protocols used as part of the training dataset can be obtained from a public repository (e.g., ClinicalTrials.gov) or an internal document management system, if the researcher has a repository of such information. Various preprocessing tasks can be performed on the study protocols, such as removing irrelevant information and standardizing the format of the input data.
[0027] Figure 2 shows an example of a study protocol document excerpt that discusses the study inclusion criteria in natural language. This is an example of the type of document that can be used in the protocol document dataset 160 that is input into the EDC Build Model 150 to generate EDC Build Predictions 170 (see Figure 1). Traditionally, such criteria have been manually incorporated into EDC systems, for example, by reviewing and interpreting the protocol document and creating eligibility entry fields and rules in the electronic case report form (eCRF).
[0028] 3 and 4 each show a portion of a simulated EDC build in XML specification format generated by EDC build model 150 based on the excerpt of the study protocol document shown in FIG. 2. The EDC build is derived from the study protocol document, for example, by using the trained EDC build model 150 (see FIG. 1) in accordance with disclosed embodiments. Such an EDC build may be used in the dataset 110 of the protocol document and corresponding EDC build (paired with the corresponding protocol document) and / or used to generate the data structure of the EDC system 190. The XML statements in this example define the fields of an eCRF for entry of subject eligibility data. In an embodiment, the EDC build may be an Architecture Loader Specification (ALS) file in XML format that specifies forms, fields, etc. in a particular format that can be used by Medidata Architect® to generate the data structure of the Medidata Rave® EDC system.
[0029] Figure 5 represents the XML-specified EDC build of Figure 3 converted into the format of a "fields" spreadsheet, which is part of an alternative spreadsheet-specified version of the EDC build. Each row of the sheet defines a specific field in the eCRF form. The spreadsheet format allows the study designer to more easily inspect and modify the various components of the constructed study. In implementation, studies can be designed using Medidata Rave®, which provides user interface screens for defining fields, rules, etc. (See, for example, Figure 18). The completed study design, or study build, can be downloaded for review and approval.
[0030] Downloadable study build files are called Architect Loader Specification (ALS) files and may be in XML specification format. XML files can be opened using either a text editor or a spreadsheet application (e.g., Microsoft Excel). When an ALS file is opened in a text editor, it appears in the format shown in Figures 3, 4, 8, 11, and 12. When an ALS file is opened as a spreadsheet, i.e., converted to spreadsheet format, it appears in the format shown in Figures 5, 6, 9, 13, and 14. As mentioned above, the alternative spreadsheet format makes it easier to review the modified study build XML file. In either case, the downloaded study build file can be used to create and / or update a study. This allows users to reuse and modify existing study designs without manually entering design parameters.
[0031] In this example, a field for the variable "IEYN" is defined on form "IE_1," which is an inclusion / exclusion criteria form. In an embodiment, the variable IEYN has a value of "yes" or "no" ("Y" or "N") indicating whether the subject meets all of the inclusion criteria (in this case, the value is "Y"). The control type on the form is specified as a vertical radio button, allowing the user to click the button to select "Y" or "N" as the value for this variable. A "PreText" string is specified to provide a prompt to the user entering data (e.g., "Does the subject meet all of the inclusion criteria?").
[0032] Figure 6 shows the EDC build specified in the XML in Figure 4 converted into a "Data Dictionary Items" spreadsheet, part of an alternative spreadsheet specification for the EDC build. The spreadsheet format allows study designers to more easily inspect and modify various components of the study build. Each row in the sheet defines an entry in a dictionary of user data strings. The ability to define data dictionary entries allows study designers to standardize components of the study build, such as fields and rules, and speed up the study design process. In this example, the study designer has created a data dictionary containing all inclusion criteria. Physicians entering data through the EDC inclusion / exclusion form are asked, "If a subject is not eligible for this study, which inclusion criteria did they not meet?" These same instructions could potentially be used in other situations by referencing dictionary entries.
[0033] In the illustrated example, the dictionary named "IETESTI" has multiple entries, each associated with a "CodedData" parameter, such as IN02, IN03, and IN04. For example, if this parameter is assigned the code "IN02," the corresponding data string would be: "2. Subject is scheduled for laparoscopic / minimally invasive colorectal surgery." This string could be used as a user prompt in one or more forms of the defined eCRF. Using a dictionary for such strings allows for efficient and consistent reuse of the strings across the defined EDC system.
[0034] Figure 7 shows an example excerpt from a study protocol document that describes in natural language the criteria for a study's screen failures (i.e., screen failures) and the requirements for handling them. Such criteria can be incorporated into the EDC system, for example, by including input fields and rules for eligibility screening in the electronic case report form (eCRF).
[0035] FIG. 8 shows a portion of a simulated EDC build in XML specification format generated by the EDC build model based on the excerpt of the study protocol document shown in FIG. 7. The EDC build is derived from the study protocol document, for example, by using a model trained according to the disclosed embodiments. The XML statements in this portion of the EDC build define a "custom function" for subject status (CF_SUBJECT_STATUS). When building edit checks, i.e., data validation, the EDC system may provide a set of parameters for performing common types of checks, such as copying data from one field to another or asking a physician a question (e.g., "Is this the right temperature?").
[0036] Custom functions allow study designers to configure more specialized and / or complex data validation. For example, a blood cell designer may want to set a subject's status based on a specific set of conditions. To trigger a custom function, the designer configures an edit check with predefined parameters: check, check step, and check action. The check action is the desired action that triggers the custom function code (e.g., CF_SUBJECT_STATUS) in this example. The check and check step define the value and function of the edit check to be performed, respectively. In this particular case, there are multiple edit checks that trigger this custom function (e.g., DS1_SUBJECT_STATUS).
[0037] Figure 9 shows the XML-specified EDC build from Figure 8 converted into a "Custom Functions" spreadsheet format that is part of an alternative spreadsheet-specified EDC build. The spreadsheet format allows study designers to more easily explore and modify various components of the study build. Each row in the sheet defines a custom function, including the function name, source code, and programming language (e.g., C#). As previously mentioned, custom functions are components of a study build, along with folders, forms, fields, edit checks, and so on. Custom functions are easier to integrate into a study build and provide more complex types of edit checks.
[0038] Figure 10 shows an example of an excerpt from a study protocol document in tabular form that presents the study's assessment schedule in natural language. This information includes the assessments that will be performed, such as informed consent, physical examination, and laboratory tests, and the specific time periods during which the assessments will be performed, such as screening, treatment, and follow-up periods.
[0039] 11 and 12 each show a portion of a simulated EDC build in XML specification format generated by the EDC Build Model based on the study protocol document excerpt shown in FIG. 10. The EDC build corresponds to the study protocol document excerpt in that it is derived from the study protocol document, e.g., by using a model trained according to the disclosed embodiments. The XML statements define an eCRF for entry of data regarding the subject's status.
[0040] Figure 13 represents the XML-specified EDC build of Figure 11 converted into a "folder" spreadsheet format that is part of an alternative spreadsheet-specified version of the EDC build. The spreadsheet format allows the study designer to more easily explore and modify the various components of the study build. Each row of the sheet defines a folder associated with a particular period or phase of the study, such as screening, treatment, or follow-up, and appears as a column heading in the table shown in the excerpt of the study protocol document in Figure 10.
[0041] Figure 14 depicts the XML-specified EDC construct in Figure 12 converted into a "BASE matrix" spreadsheet format, which is part of an alternative spreadsheet-specified version of the EDC construct. Each row in the sheet corresponds to an assessment to be performed, e.g., informed consent (IC_1), eligibility / exclusion criteria (IE_1, IE_2), medical history (MH_1), etc., and is shown as a row heading in the table presented in the excerpt from the study protocol document shown in Figure 10.
[0042] 15 illustrates a method 1500 for building an electronic data capture (EDC) system. The method 1500 includes forming a training dataset from a set of protocol documents and corresponding EDC builds (1510). The method 1500 further includes using the set of protocol documents and corresponding EDC builds of the training dataset to fine-tune a pre-trained language model to generate an EDC build model (1520). The method 1500 further includes inputting the protocol documents into the EDC build model to generate an EDC build prediction (1530). The method 1500 further includes generating a data structure for the EDC system based at least in part on the EDC build prediction (1540).
[0043] 16 illustrates a method 1600 that includes a further aspect of the method 1500 for building an electronic data capture (EDC) system, the further aspect relating to testing an EDC build model. The method 1600 includes dividing a set of protocol documents and corresponding EDC builds into a training data set and a test data set (1610). The method 1600 further includes inputting the protocol documents of the test data set into an EDC build model to generate EDC build predictions (1620). The method 1600 further includes analyzing the EDC build predictions generated from the protocol documents of the test data set to determine the accuracy of the EDC build model (1630). The method 1600 further includes fine-tuning the EDC build model based at least in part on the analysis of the EDC build predictions generated from the protocol documents of the test data set (1640).
[0044] FIG. 17 illustrates a user interface screen for generating edit checks related to adverse event dates using a demo EDC build model. The edit check includes a series of check steps and check actions that prompt the user to correct errors in the start and end dates, such as an error where the end date is earlier than the start date. In a disclosed embodiment, the elements of the edit check are generated by processing the corresponding portions of the study protocol document. In the illustrated demonstration, a text string specifying the edit check was entered, and the edit check shown in the figure was generated. In the demonstration, entering the text string and converting it to an edit check took approximately 15 minutes.
[0045] 18 is a user interface screen for manually entering elements of an edit check for adverse event dates in a demo EDC build. The example shown is entering data value elements of the edit check by entering data values and / or selecting various parameters, such as event type. In the demo, manually entering the various data values and parameters of the edit check took over two hours to complete.
[0046] While the implementations presented herein focus on clinical trials (e.g., pharmaceutical research), the disclosed approaches can be used in other fields where structured data collection and management is complex. These solutions can also be used in other applications (and other fields) where building and / or developing them is time-consuming and labor-intensive.
[0047] Aspects of the present invention may be embodied in the form of a system, a computer program product, or a method. Likewise, aspects of the present invention may be embodied in the form of hardware, software, or a combination of both. Aspects of the present invention may also be embodied in the form of a computer program product in the form of computer-readable program code stored on one or more computer-readable mediums.
[0048] The computer-readable medium may be a computer-readable storage medium, which may be, for example, an electronic, optical, magnetic, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof.
[0049] The computer program code in embodiments of the present invention can be written in any suitable programming and / or scripting language. The program code can be executed on a single computer or on multiple computers. The computer can include a processing unit in communication with a computer-usable medium, the computer-usable medium including a set of instructions, and the processing unit can include a machine learning algorithm designed and / or trained to execute the set of instructions.
[0050] The above description is intended to be illustrative of the principles and various embodiments of the present invention. Numerous variations and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. It is intended that the following claims be interpreted to embrace all such variations and modifications.
Claims
1. A system construction method for constructing an EDC system (electronic data collection system), comprising: a training dataset generation step of generating a training dataset from a set of protocol documents and a corresponding set of EDC builds; an EDC build model generation step of fine-tuning a pre-trained language model using the protocol document set and the corresponding EDC build set of the training dataset to generate an EDC build model; an EDC build prediction generation step of inputting one protocol document into the EDC build model to generate an EDC build prediction; a data structure generation step of generating data structures for the EDC system based at least in part on the EDC build predictions; A system construction method including:
2. The system construction method of claim 1 , wherein the data structure generating step includes generating one or more electronic case report forms.
3. 3. The system configuration method of claim 2, wherein the data structure generating step includes generating one or more edit checks for checking data entered via one or more electronic case report forms.
4. 2. The system construction method according to claim 1, wherein the data structure generating step includes the step of generating one or more folder structures for managing EDC system information.
5. The system construction method according to claim 1 , wherein the EDC-built model generation step includes a step of adjusting hyperparameters of the pre-trained language model.
6. 2. The system construction method according to claim 1, wherein the protocol document is a natural language text document.
7. The system construction method according to claim 6, wherein the protocol document includes an evaluation schedule in a tabular format.
8. The system construction method of claim 1 , wherein the EDC build prediction is in the form of an Extensible Markup Language (XML) document.
9. 2. The system construction method of claim 1, wherein the step of forming the training data set includes the step of dividing a set of protocol documents and corresponding EDC builds into a training data set and a test data set.
10. inputting a set of protocol documents of the test dataset into the EDC build model to generate an EDC build prediction; analyzing the EDC build predictions generated from a set of protocol documents of the testing dataset to determine the accuracy of an EDC model; The system construction method according to claim 9, further comprising:
11. 11. The system construction method of claim 10, further comprising fine-tuning the EDC build model based, at least in part, on an analysis of EDC build predictions generated from a set of protocol documents of the test dataset.
12. A system for building an EDC system (electronic data collection system), a computer having one or more processors in communication with a memory; The memory a training dataset generation step of generating a training dataset from a set of protocol documents and a corresponding set of EDC builds; an EDC build model generation step of fine-tuning a pre-trained language model using the protocol document set and corresponding EDC build set of the training dataset to generate an EDC build model; an EDC build prediction generation step of inputting one protocol document into the EDC build model to generate an EDC build prediction; a data structure generation step of generating data structures for the EDC system based at least in part on the EDC build predictions; and a system configuration system that stores instructions for causing the processor to execute the above.
13. The system of claim 12 , wherein the data structure generating step includes generating one or more electronic case report forms.
14. 14. The system of claim 13, wherein the data structure generating step includes generating one or more edit checks for checking data entered via one or more electronic case report forms.
15. 13. The system configuration system according to claim 12, wherein the data structure generating step includes the step of generating one or more folder structures for managing EDC system information.
16. The system construction system of claim 12 , wherein the EDC build prediction is in the form of an Extensible Markup Language (XML) document.
17. 13. The system construction system of claim 12, wherein the step of forming the training data set includes the step of dividing a set of protocol documents and corresponding EDC builds into a training data set and a testing data set.
18. The memory includes: inputting the protocol documentation of the test dataset into the EDC build model to generate an EDC build prediction; analyzing the EDC build predictions generated from the protocol documents of the test dataset to determine the accuracy of the EDC model; 18. The system configuration system according to claim 17, further storing instructions for causing the processor to execute the following:
19. The memory includes:
20. The system configuration system of claim 18, further comprising instructions that cause the processor to perform the step of fine-tuning the EDC build model based, at least in part, on an analysis of EDC build predictions generated from protocol documents of the test dataset.
20. 1. A non-volatile computer-readable medium having stored thereon instructions for causing one or more processors of a computer to perform an EDC system construction method, the EDC system construction method comprising: a training dataset generation step of generating a training dataset from a set of protocol documents and a corresponding set of EDC builds; an EDC build model generation step of fine-tuning a pre-trained language model using the protocol document set and the corresponding EDC build set of the training dataset to generate an EDC build model; an EDC build prediction generation step of inputting one protocol document into the EDC build model to generate an EDC build prediction; a data structure generation step of generating data structures for the EDC system based at least in part on the EDC build predictions; 1. A computer-readable medium comprising:
21. 21. The computer-readable medium of claim 20, wherein the data structure generating step includes generating one or more electronic case report forms.
22. 22. The computer-readable medium of claim 21, wherein the data structure generating step includes generating one or more edit checks for checking data entered via one or more electronic case report forms.
23. 21. The computer-readable medium of claim 20, wherein the data structure generating step includes generating one or more folder structures for managing EDC system information.
24. 21. The computer-readable medium of claim 20, wherein the EDC build prediction is in the form of an Extensible Markup Language (XML) document.
25. 21. The computer-readable medium of claim 20, wherein forming the training data set comprises dividing the set of protocol documents and corresponding EDC builds into a training data set and a testing data set.
26. The EDC system construction method includes: inputting a set of protocol documents of the test dataset into the EDC build model to generate an EDC build prediction; analyzing the EDC build predictions generated from a set of protocol documents of the testing dataset to determine the accuracy of an EDC model; 26. The computer-readable medium of claim 25, further comprising:
27. The EDC system construction method includes:
27. The computer-readable medium of claim 26, further comprising fine-tuning the EDC build model based, at least in part, on an analysis of EDC build predictions generated from a set of protocol documents of the testing dataset.
Citation Information
Patent Citations
Clinical trial / clinical study support system, clinical trial / clinical study support method, electronic medical chart and EDC coordination sub-system, automatic transcription program, and electronic medical chart-recording program
JP2015153137A
Direct data connection for universal collection of health data
US20220351810A1