Data-driven clinical trial full-process automatic management system

Through a data-driven, fully automated clinical trial management system, the entire process from multi-source data to application documents has been automated, solving problems such as time-consuming data standardization and delayed compliance verification. This ensures real-time linkage and compliance between data and documents, thereby improving the efficiency and quality of clinical trials.

CN122117290APending Publication Date: 2026-05-29SUZHOU BLUE POINT RESPIRATORY TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU BLUE POINT RESPIRATORY TECHNOLOGY CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies in clinical trial data processing and management suffer from problems such as time-consuming data standardization, delayed compliance verification, system fragmentation, excessive duplication of documentation work, and lagging risk management, making it impossible to achieve closed-loop automation from raw data collection to final application document generation.

Method used

A data-driven, fully automated clinical trial management system is adopted, including a data integration and standardization module, an intelligent validation and quality control module, an automated statistical analysis module, and an intelligent document generation module. Through intelligent mapping, compliance validation, statistical analysis, and automated document generation, real-time linkage between data and documents and full-process automation are achieved.

Benefits of technology

It has achieved end-to-end automation from multi-source raw data to application-level document generation, reducing manual intervention, ensuring data compliance, traceability, and real-time document updates, solving the problem of data and document disconnect, and improving the overall quality and compliance of clinical trials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117290A_ABST
    Figure CN122117290A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of clinical trial data management, in particular to a kind of clinical trial whole-process automation management system based on data driving, including with data as core driving source, pass through clinical trial whole-process data link, realize the seamless cooperation and automatic linkage of each module, realized from the instruction automatic triggering of the full-link of the generation of the declaration level document of multiple original data access, process automatic advance, result automatic update, problem automatic interception, whole process without manual intervention data transmission and process triggering, let data become the natural power of process advancement, solve the problem that data and document are disjointed, after statistical result update, corresponding clinical summary report often needs artificial rechecking and modification, cannot realize data update driven document automatic update.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of clinical trial data management technology, and more specifically to a data-driven, fully automated management system for the entire clinical trial process. Background Technology

[0002] Clinical trials are the most critical, time-consuming, and costly stage in new drug development. Currently, clinical trial data processing and management involve multiple fragmented stages, including data collection, data standardization, statistical analysis, documentation, and quality control. Traditional operating models heavily rely on manual labor and single-function software tools, leading to numerous problems in the entire clinical trial data processing and management process. For example, data standardization is time-consuming; programmers need to manually write SAS code to map raw data to SDTM and ADaM standards, which is inefficient and error-prone. Validation is delayed; compliance verification often occurs after data conversion, requiring backtracking and modifications if problems are found, resulting in long cycles. Systems are fragmented; data is not shared between pharmacovigilance systems, project management systems, etc., making it difficult for monitoring and data management to grasp all data in real time. Documentation involves repetitive work; although templates exist for protocols, SAP, CSRs, etc., a large amount of data still needs to be manually entered, increasing the risk of inconsistencies between data and results. Existing technologies often focus on a single stage (such as only performing EDC or only performing validation), failing to achieve closed-loop automation from raw data collection to final application document generation. Standardization conversion heavily relies on professional programmers writing code, lacking zero-code or visual configuration solutions. Data and documentation are disconnected; after statistical results are updated, the corresponding clinical summary reports often require manual verification and modification, making it impossible to achieve automatic document updates driven by data updates. Risk management is lagging behind, lacking a centralized monitoring mechanism based on real-time data, making it difficult to detect data fraud or quality risks in a timely manner during trials.

[0003] Therefore, the present invention provides a data-driven automated management system for the entire clinical trial process to solve the above problems. Summary of the Invention

[0004] To address the above issues and overcome the shortcomings of existing technologies, this invention provides a data-driven automated management system for the entire clinical trial process. This system solves the problem of data and documentation being disconnected, and the clinical summary reports often requiring manual verification and modification after statistical results are updated, making it impossible to achieve automatic document updates driven by data updates.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A data-driven, fully automated clinical trial management system includes: a data integration and standardization module, which intelligently maps and processes acquired multivariate clinical trial data to obtain standardized data; an intelligent verification and quality control module, which performs compliance verification and approval processing on the standardized data to obtain an approval dataset; an automated statistical analysis module, which processes the approval dataset based on preset built-in rules to obtain statistical analysis results; an intelligent document generation module, which embeds node-linked analysis of the statistical analysis results based on preset standardized templates to obtain trial data submission results; and a comprehensive project management module, which performs full-dimensional visual monitoring, early warning, and adjustment of each module based on preset modes to update trial data submission results.

[0006] Preferably, the intelligent mapping processing of the acquired multivariate clinical trial data to obtain standardized data includes: parsing the multivariate clinical trial data to obtain a structured raw data pool; determining the mapping relationship between raw variables and standard variables based on the basic trial information and the structured raw data pool, and generating an initial mapping table; performing compliance processing on the initial mapping table to obtain a target mapping table; and mapping the structured raw data pool to standardized data based on the acquired transformation mode and the target mapping table.

[0007] Preferably, the process of performing compliance verification and approval processing on standardized data to obtain an approval dataset includes: acquiring and overlaying institutional verification rules and custom rules corresponding to the test application target to generate a target rule set; performing embedded verification and adjustment on the standardized data of each node in the system based on the target rule set to obtain a compliance dataset; and performing approval processing on the compliance dataset to obtain an approval dataset.

[0008] Preferably, the intelligent verification and quality control module further includes: pre-verification of manual operations during the verification process to intercept abnormal operations.

[0009] Preferably, the step of processing the approval dataset based on preset built-in rules to obtain statistical analysis results includes: extracting requirements from the acquired clinical trial protocol to obtain core requirement information; determining the corresponding statistical model from the built-in algorithm library based on the core requirement information; analyzing the approval dataset based on the statistical model to obtain initial TFLs; personalizing the initial TFLs to obtain customized TFLs; and verifying the compliance and rationality of the customized TFLs to obtain statistical analysis results.

[0010] Preferably, the step of performing embedded node linkage analysis on statistical analysis results based on a preset standardized template to obtain experimental data application results includes: fine-tuning the acquired standardized template to obtain a final review core document template; performing correlation analysis on statistical analysis results based on the final review core document template to generate a structured data pool and a data node association table; filling the structured data pool into the embedded nodes of the final review core document template and generating medical description content to obtain a core document draft; establishing a real-time two-way linkage relationship based on the core document draft and the data node association table to obtain a global data association graph and a core document draft containing dynamic association relationships; verifying the core document draft containing dynamic association relationships and outputting experimental data application results containing the application-level core document draft.

[0011] Preferably, the step of performing correlation analysis on the statistical analysis results based on the lifetime core document template to generate a structured data pool and a data node association table includes: extracting the corresponding structured data from the statistical analysis results based on the unique identifier of the data embedding node in the lifetime core document template, establishing a one-to-one correspondence between the data embedding node and the underlying data source, and generating a structured data pool and a data node association table.

[0012] Preferably, the method of performing full-dimensional visual monitoring, early warning, and adjustment of each module based on a preset mode to update the test data declaration results includes: building a multi-dimensional visual monitoring warehouse; identifying and analyzing real-time data in the visual monitoring warehouse based on a preset risk early warning model to obtain risk early warning results; and performing multi-role collaborative processing on problematic data in the risk early warning results to update the test data declaration results.

[0013] Preferably, the step of identifying and analyzing real-time data in the visual monitoring warehouse based on a preset risk warning model to obtain risk warning results includes: identifying abnormal data patterns in the real-time data based on the preset risk warning model; classifying the abnormal data patterns by risk to obtain risk classification results; triggering in-depth compliance verification on the risk data corresponding to the risk classification results to locate the root cause of the risk, so as to obtain risk warning results and send them to the corresponding management personnel.

[0014] Preferably, the multi-role collaborative processing of problem data in the risk warning results to update the test data declaration results includes: generating a standardized problem list based on the problem data and assigning it to the corresponding management personnel; obtaining the problem processing results of each management personnel on their respective role's dedicated workbench and triggering system verification; and changing the problem status and updating the corresponding test data declaration results after verification.

[0015] The beneficial effects of this invention are as follows: 1. This invention uses data as the core driving force, connects the data links of the entire clinical trial process, and realizes seamless collaboration and automated linkage of various modules. It realizes automatic command triggering, automatic process advancement, automatic result updating, and automatic problem interception throughout the entire chain from multi-source raw data access to application-level document generation. No manual intervention is required for data transmission and process triggering, making data the natural driving force for process advancement. It solves the problem of data and document disconnect, where after the statistical results are updated, the corresponding clinical summary report often needs to be manually re-verified and modified, and the inability to realize automatic document update driven by data updates.

[0016] 2. This invention is driven by data and builds a full-domain data association graph. It establishes a real-time two-way linkage relationship between the original data, standardized datasets, TFLs and data embedding nodes in the document. The data-driven automated flow will trigger the automatic refresh of the document: any data update in any link will be transmitted to the document generation module through automated flow. The corresponding values, charts and descriptions in the document will be automatically updated synchronously without manual intervention.

[0017] 3. This invention achieves operational linkage and functional complementarity through an automated command center, forming a closed-loop quality control system covering the entire clinical trial process: For example, if the project comprehensive management module discovers a risk in the data of a certain center, it can automatically link with the intelligent verification module to conduct in-depth compliance verification of the data of that center; if the intelligent verification module discovers data anomalies, it can automatically link with the project management module to generate problematic data and assign it to the corresponding responsible person; after the TFLs are generated by the statistical module, they can automatically link with the document generation module to populate them, and at the same time link with the intelligent verification module to perform a second verification of the compliance of the TFLs. Thus, it realizes the full-process linkage control from data standardization, compliance verification, statistical analysis, document generation to project management, with no blind spots in quality control, and the overall quality and compliance of clinical trials are comprehensively guaranteed. Attached Figure Description

[0018] Figure 1 This is a schematic block diagram of a data-driven automated management system for the entire clinical trial process according to the present invention. Detailed Implementation

[0019] The following will refer to the attached reference. Figure 1 The various embodiments of the present invention will be described in detail below. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0020] A data-driven, end-to-end automated management system for clinical trials, as shown in the attached document. Figure 1As shown, it includes: a data integration and standardization module, which performs intelligent mapping processing on the acquired multivariate clinical trial data to obtain standardized data; multivariate clinical trial data includes, but is not limited to, heterogeneous raw data (in Excel, CSV, Oracle, HL7, etc. formats) from multiple systems such as EDC, PV, central laboratory, and hospital HIS / LIS; and basic trial information (indications, treatment areas, and regulatory agencies to which the trial is submitted).

[0021] After obtaining the standardized data, the system automatically triggers a data push command, synchronously pushing the SDTM / ADaM dataset package and its accompanying documents to the intelligent verification and quality control module. At the same time, the dataset package is synchronized to the system's unified database for use by the statistical analysis automation module.

[0022] Only after passing compliance verification and result confirmation can the system enter the intelligent verification and quality control module. For example, if data parsing fails, the system cannot enter the mapping configuration; if the mapping verification fails, the system cannot enter the data transformation, thus ensuring the validity of the output results at each step.

[0023] The data integration and standardization module enables unified access and standardized processing of multi-source heterogeneous data, breaking down data silos; it replaces manual SAS programming with zero-code intelligent mapping, shortening the data standardization cycle from weeks to hours; it also stores mapping rules in the enterprise's private library, enabling reuse in experiments within the same domain and continuously improving subsequent processing efficiency; and it generates standardized datasets and supplementary documents that meet regulatory requirements, providing compliant basic data for subsequent verification, statistics, and document generation.

[0024] The intelligent verification and quality control module performs compliance verification and approval processing on standardized data to obtain an approval dataset. This invention targets standardized datasets, full-process operational behaviors, compliance verification, quality risk interception, and end-to-end traceability. Through multi-rule adaptation, real-time verification at all nodes, human error interception, and digital approval traceability, it achieves full lifecycle quality and compliance management of data from access to output, shifting post-event verification to process monitoring.

[0025] Through the intelligent verification and quality control module, this invention realizes full-process, multi-supervisory, real-time compliance verification and quality control from data access to document generation, avoiding secondary errors caused by manual operation, ensuring the compliance, traceability and integrity of data throughout its entire lifecycle, and solving the problem that traditional compliance verification lags behind data transformation and requires retrospective modification after problems are discovered.

[0026] After obtaining the approved dataset, the system automatically triggers two instructions: ① Push the approved and compliant SDTM / ADaM dataset package to the statistical analysis automation module as the basic data for statistical analysis; ② Push all verification reports, traceability reports, and operation logs to the project comprehensive management module to achieve data quality visualization monitoring; At the same time, synchronize the compliant dataset package to the system's unified database for use by the document intelligent generation module.

[0027] The automated statistical analysis module processes the approval dataset based on preset built-in rules to obtain statistical analysis results. After obtaining the statistical analysis results, the system automatically triggers a data push command, pushing the application-level compliance TFLs package, two statistical reports, and a dynamically bound relationship table to the intelligent document generation module as the statistical result basis for core document writing; at the same time, the TFLs generation progress and compliance rate are pushed to the project comprehensive management module to realize the visualization of the statistical analysis process.

[0028] TFLs refer to a standardized outcome data system for clinical trials, consisting of tables, graphs, and lists.

[0029] The automated statistical analysis module targets compliance ADaM datasets, test plans, application-level TFLs, and statistical traceability reports. Through intelligent analysis of test plans, automatic matching of statistical models, automated generation and customization of TFLs, and compliance verification of statistical results, it achieves zero-code statistical analysis and application-level standardization of TFLs, ensuring the accuracy, compliance, and traceability of statistical results.

[0030] Based on a compliant ADaM dataset, it enables zero-code, intelligent, and compliant TFLs generation, providing declaration-level and traceable statistical results for the document intelligent generation module. This solves the problems of traditional statistical analysis requiring manual model selection and coding, which has a high technical threshold; it also solves the problems of incomplete traceability of statistical results, making it impossible to quickly reproduce during regulatory verification; and the problem of low efficiency, requiring the regeneration of all TFLs after data updates.

[0031] The intelligent document generation module, based on preset standardized templates, embeds statistical analysis results into linked nodes for analysis to obtain experimental data reporting results. Through a full-process analysis involving intelligent template loading, automatic extraction of structured data, intelligent document filling and natural language generation, dynamic binding of data and documents, and compliance format verification, it automates the writing of core documents, ensuring real-time consistency, format compliance, and content completeness between document content and underlying data.

[0032] After the content of the document intelligent generation module is completed, the system automatically pushes the core document draft, verification report, and export package to the project comprehensive management module, realizing visual monitoring of document writing progress and compliance, and supporting multi-role online review triggered by the project management module.

[0033] Based on structured data throughout the entire process, it automates the population, intelligent generation, and dynamic updating of core documents such as DMP, SAP, CSR, and Protocol, outputting initial drafts of application-level core documents that meet regulatory requirements. It also solves problems such as the traditional need for manual data extraction and population from multiple stages in core document writing, resulting in a large amount of repetitive work; the need for manual verification and modification of document content after TFLs updates, which can easily lead to data inconsistencies; inconsistent document formats requiring manual formatting; and the potential for version confusion and content conflicts in multi-role collaborative writing.

[0034] The project integrated management module provides comprehensive, visualized monitoring, early warning, and adjustment of each module based on preset modes to update trial data submission results. This module connects the data interfaces of the five major modules, enabling comprehensive visualized monitoring, intelligent RBQM risk warning, multi-role collaborative management, and precise project progress control. It provides a one-stop management platform for clinical trial project leaders, addressing issues such as the inability to monitor progress / quality in real time at each stage of traditional clinical trials, lack of transparency, reliance on manual verification for risk discovery (which is delayed and inaccurate), lack of unified management of multi-role tasks (leading to low collaboration efficiency), lengthy manual tracking of problem handling, and the inability to promptly identify and analyze the causes of project delays.

[0035] This invention uses data as the core driving force, connecting the entire data link of clinical trials to achieve seamless collaboration and automated linkage of various modules. It realizes automatic command triggering, automatic process advancement, automatic result updating, and automatic problem interception throughout the entire chain, from multi-source raw data access to application-level document generation. No manual intervention is required for data transmission and process triggering, making data the natural driving force for process advancement. This solves the problem of data and document disconnect, where after statistical results are updated, the corresponding clinical summary report often needs to be manually re-verified and modified, making it impossible to achieve automatic document updates driven by data updates.

[0036] In one embodiment of the present invention, the intelligent mapping processing of the acquired multivariate clinical trial data to obtain standardized data includes: parsing the multivariate clinical trial data to obtain a structured raw data pool; determining the mapping relationship between raw variables and standard variables based on the basic trial information and the structured raw data pool, and generating an initial mapping table; performing compliance processing on the initial mapping table to obtain a target mapping table; and mapping the structured raw data pool to standardized data according to the acquired transformation mode and the target mapping table.

[0037] Specifically, users select data sources and upload them with one click in the visual interface. The system's multi-source data intelligent access unit automatically triggers the data format parsing engine to complete data type identification (character / numerical / date), preliminary matching of field meanings, preliminary screening of null / abnormal values, and generate a "Data Access Parsing Report" that marks the location and type of abnormal data, while also outputting a structured raw data pool.

[0038] The system's intelligent metadata mapping engine then automatically recommends mapping relationships between original variables and SDTM / ADaM standard variables based on the trial indications / treatment areas in the trial information, generating an initial mapping table. Users can confirm / fine-tune the initial mapping table through a visual drag-and-drop interface, supporting batch mapping, group mapping, and conditional mapping. During the configuration process, the system performs real-time intelligent mapping verification (variable type matching, value range compliance), intercepts invalid mappings, and generates mapping operation logs. After the mapping configuration is completed, users can save the mapping rules with one click, and the rules are automatically stored in the enterprise's private mapping library. Finally, the final audited metadata mapping table, mapping operation logs, and mapping rule version records are output.

[0039] After the user selects the conversion mode (full conversion / incremental conversion, initial conversion to full, subsequent data updates to incremental), the system automatically converts the original data into an SDTM dataset based on the mapping table. Then, based on the association relationships of the SDTM dataset and the ADaM standard, it automatically generates an ADaM analysis dataset. The system monitors the progress in real time during the conversion process, generates conversion logs, and outputs compliant SDTM datasets, compliant ADaM datasets, and data conversion logs.

[0040] The system automatically calls the standard supplementary document generation component and generates Define.xml (including data dictionary, mapping relationship, data traceability), SDRG (data standardization report), and SDTM / ADaM dataset package (including dataset and supplementary documents) with one click, based on the requirements of CDISC and the target regulatory agency. The supplementary documents and datasets are dynamically bound together to obtain the final standardized data.

[0041] Among them, SAS refers to Statistical Analysis System; SDTM refers to Research Data Tabulation Model, which is a standardized presentation format of raw clinical trial data; and ADaM refers to Analytical Data Model, which is an analytical dataset processed based on SDTM, specifically used for statistical analysis and TFL generation.

[0042] Through the data integration and standardization module, this invention enables one-click access to multi-source heterogeneous data and zero-code SDTM / ADaM standardization conversion, providing a compliant, unified, and structured basic dataset for all subsequent modules. This solves the problems of low efficiency and error-proneness in traditional manual SAS code-based data standardization; the need for manual preprocessing due to inconsistent multi-source data formats; compliance risks caused by lagging standard library updates; and repetitive work due to the lack of template support for mapping configuration.

[0043] In one embodiment of the present invention, the compliance verification and approval processing of standardized data to obtain an approval dataset includes: acquiring and overlaying institutional verification rules and custom rules corresponding to the test application target to generate a target rule set; based on the target rule set, performing embedded verification and adjustment on the standardized data of each node in the system to obtain a compliance dataset; and performing approval processing on the compliance dataset to obtain an approval dataset.

[0044] Specifically, it acquires SDTM / ADaM dataset packages pushed from the data integration and standardization module, trial application targets (NMPA / FDA / PMDA / multi-center), and the system's built-in dynamically updated rule base (FDA Validator Rules, NMPA guidance principles, CDISC Validation Checks, PMDA requirements).

[0045] The system's intelligent multi-regulatory rule adaptor automatically loads the corresponding regulatory agency's verification rules based on the test application objectives, supporting parallel rule loading from multiple agencies (such as loading NMPA and FDA simultaneously). If the user has internal quality specifications, custom rules can be overlaid and loaded, generating a target rule set containing official rules and optional custom rules.

[0046] Furthermore, the system's built-in real-time verification engine performs embedded verification on each of the four nodes: data access, mapping configuration, data transformation, and dataset output. Each node triggers corresponding verification rules, such as the mapping configuration node verifying variable matching compliance and the dataset output node verifying data format compliance.

[0047] If a violation is found during the verification process, the system will immediately locate the specific node and field, generate actionable modification suggestions, such as "The value range of field XX in the ADaM dataset exceeds the NMPA requirements, and it is recommended to correct it to [0,1]", and highlight the problematic data in red and send it back to the data integration and standardization module. After the user makes the modification, the verification will be retried.

[0048] After all nodes pass verification, a "Data Full-Process Compliance Verification Report" is generated, which marks the verification results, compliance rate, problem handling records, and verified SDTM / ADaM dataset packages for each node.

[0049] Furthermore, the intelligent verification and quality control module also includes: pre-verification of manual operations during the verification process to intercept abnormal operations; the system performs pre-verification and access control on all manual operations involved in the verification process, preventing unauthorized personnel from modifying core configurations (such as Metadata mapping tables and verification rules); when manually entering / modifying data, the system automatically verifies the data format, value range, and compliance, intercepts abnormal operations, generates manual operation monitoring logs, and outputs the corresponding manual operation monitoring logs and access control records.

[0050] The system automatically creates double bookmarks for verified dataset packages, supplementary documents (such as aCRF, Define.xml), and operation logs for each step, binding them to unique dataset identifiers; it also triggers a multi-role hierarchical signing process, with each role completing an electronic signature in the system.

[0051] The system binds all operational behaviors to data / documents, generating an immutable end-to-end traceability chain, recording the operator, operation time, operation content, and data version, and outputting the completed dataset package, supplementary documents, end-to-end traceability report, and electronic signature record; among which, operational behaviors include data access, mapping, transformation, verification, and signing.

[0052] Through the intelligent verification and quality control module, this invention realizes full-process, multi-supervisory, real-time compliance verification and quality control from data access to document generation, avoiding secondary errors caused by manual operation, ensuring the compliance, traceability and integrity of data throughout its entire lifecycle, and solving the problem that traditional compliance verification lags behind data transformation and requires retrospective modification after problems are discovered.

[0053] In one embodiment of the present invention, the step of processing the approval dataset based on preset built-in rules to obtain statistical analysis results includes: extracting requirements from the acquired clinical trial protocol to obtain core requirement information; determining the corresponding statistical model from the built-in algorithm library based on the core requirement information; analyzing the approval dataset based on the statistical model to obtain initial TFLs; personalizing the initial TFLs to obtain customized TFLs; and verifying the compliance and rationality of the customized TFLs to obtain statistical analysis results.

[0054] Specifically, the system automatically parses the trial protocol uploaded by the user and extracts the core statistical requirements. The core statistical requirements include: trial design type (parallel control / crossover / single-arm trial), primary / secondary endpoint indicators, sample size, subject grouping method, statistical method requirements, subgroup analysis / stratified analysis requirements, etc., and generates a "Trial Protocol Statistical Requirements Extraction Report" for statisticians to confirm. After confirmation, the "Trial Protocol Statistical Requirements Extraction Report" and structured statistical requirement parameters are output.

[0055] Based on statistical requirement parameters and the system's built-in algorithm library (survival analysis, analysis of variance, chi-square test, logistic regression, etc., including models specifically for international multicenter trials), the system automatically recommends the optimal statistical model and marks the applicable basis, calculation formula, and applicable scenarios for the model. Statisticians can confirm / fine-tune the statistical model in the visual interface without writing any code. The model selection results are automatically saved to the statistical analysis plan template, and the system outputs the final statistical model, model configuration parameters, statistical method descriptions, etc.

[0056] The system automatically calculates and analyzes the ADaM dataset based on the final statistical model, generating initial TFLs. The initial TFLs include, but are not limited to, tables: baseline table, efficacy table, safety table; graphs: line graph, bar graph, forest graph; lists: AE / SAE list, subject visit list.

[0057] Users can personalize TFLs through a zero-code visual interface, such as adjusting table styles, graph types, displayed fields, font line spacing, etc. The system previews the effect in real time, and the customization process strictly follows the declaration format requirements of the target regulatory agency.

[0058] If the ADaM dataset is updated, the system triggers an incremental update of TFLs, recalculating only the statistical results corresponding to the changed data and partially refreshing the TFLs without regenerating all content. The final output is the initial customized TFLs and TFLs update log.

[0059] The system verifies the compliance of statistical methods: confirming that the statistical model is consistent with the trial protocol requirements and that the calculation process conforms to statistical standards; the system verifies the reasonableness of statistical results: screening for outliers (such as P-value < 0.0001, rate values ​​exceeding the reasonable range) and logical contradictions (such as inconsistencies between efficacy indicators and group trends). After successful verification, the system generates traceability information for each TFL, annotating the data source (specific fields in the ADaM dataset), statistical methods, calculation formulas, and calculation processes, and generating the "TFL Statistical Result Compliance Verification Report" and the "TFL Traceability Report"; finally, it outputs statistical analysis results including a TFL package (containing tables / graphs / lists) that meets the application-level compliance requirements, the "TFL Statistical Result Compliance Verification Report," the "TFL Traceability Report," and a table showing the dynamic binding relationship between TFLs and the ADaM dataset.

[0060] The automated statistical analysis module targets compliance ADaM datasets, test plans, application-level TFLs, and statistical traceability reports. Through intelligent analysis of test plans, automatic matching of statistical models, automated generation and customization of TFLs, and compliance verification of statistical results, it achieves zero-code statistical analysis and application-level standardization of TFLs, ensuring the accuracy, compliance, and traceability of statistical results.

[0061] Based on a compliant ADaM dataset, it enables zero-code, intelligent, and compliant TFLs generation, providing declaration-level and traceable statistical results for the document intelligent generation module. This solves the problems of traditional statistical analysis requiring manual model selection and coding, which has a high technical threshold; it also solves the problems of incomplete traceability of statistical results, making it impossible to quickly reproduce during regulatory verification; and the problem of low efficiency, requiring the regeneration of all TFLs after data updates.

[0062] In one embodiment of the present invention, the step of performing embedded node linkage analysis on statistical analysis results based on a preset standardized template to obtain experimental data application results includes: fine-tuning the acquired standardized template to obtain a final review core document template; performing correlation analysis on the statistical analysis results based on the final review core document template to generate a structured data pool and a data node association table; filling the structured data pool into the embedded nodes of the final review core document template and generating medical description content to obtain a core document draft; establishing a real-time bidirectional linkage relationship based on the core document draft and the data node association table to obtain a global data association graph and a core document draft containing dynamic association relationships; verifying the core document draft containing dynamic association relationships and outputting experimental data application results containing the application-level core document draft.

[0063] Specifically, the system automatically loads corresponding standardized templates based on the trial application objectives and treatment areas. These templates contain pre-set structured data embedding nodes (each node has a unique identifier linked to the underlying data source). Users can personalize the templates through a zero-code visual interface (e.g., adding or removing content modules, adjusting chapter order, modifying embedding node positions, etc.). After customization, the system automatically saves the template version and outputs a final review core document template containing data embedding nodes and unique identifiers, along with a template customization record.

[0064] Based on the unique identifier of the data embedding node in the template, the system automatically extracts the corresponding structured data from each module, establishes a one-to-one correspondence between the data embedding node and the underlying data source, and generates a "Structured Data Extraction and Association Report" that marks the data source, data version, and association relationship. At the same time, it outputs a structured data pool containing experimental / data / statistical / validation data across all dimensions and a data-node association table.

[0065] The system automatically and accurately populates the structured data into the corresponding embedded nodes of the template. For example, TFLs are automatically embedded into the "Statistical Results" section of the CSR, the enrollment progress is populated into the "Data Acquisition Plan" section of the DMP, and the statistical model is populated into the "Statistical Methods" section of the SAP.

[0066] For descriptive content in documents (such as CSR "Overview of Trial Process" and "Data Quality Assessment" and DMP "Data Management Process"), the system automatically generates standardized medical written language based on the extracted structured data, and supports manual fine-tuning by users.

[0067] Multiple roles (DM / statistics / medical writing) edit documents through online collaborative writing units. Each role only has editing permissions for its corresponding chapter. Editing operations are saved in real time and logs are generated to avoid content conflicts. The final output includes a core draft document containing populated data, TFLs, descriptive content, document editing logs, and multi-role collaboration records.

[0068] The system builds a global data association graph, establishing a real-time bidirectional linkage relationship between "raw data → SDTM / ADaM dataset → TFLs → document embedding nodes" and configuring update trigger rules: when any upstream data is updated, the corresponding downstream content is automatically refreshed, and document update tracing rules are generated at the same time.

[0069] The final output data includes the core document draft, which contains a dynamic relationship graph of documents, update trigger rules, and binding dynamic relationship relationships.

[0070] The system performs content compliance checks: verifying data consistency (no conflicts between documents and underlying data), logical rationality, and compliance with experimental protocols and regulatory requirements. The system also performs format compliance checks: verifying font, line spacing, chapter numbering, citation standards, TFLs layout, etc., and automatically optimizing formatting issues.

[0071] After verification, the system automatically generates a document table of contents, page numbers, and reference index, and supports users to export the document to the required format for application with one click. The exported document has no formatting errors. Finally, the system outputs the test data application results, which include the initial draft of the core application-level document (PDF / Word / XML), the "Document Compliance and Format Verification Report", and the document export package.

[0072] Through the setup method of this embodiment, the present invention, based on full-process structured data, achieves automated filling, intelligent generation, and dynamic updating of core documents such as DMP, SAP, CSR, and Protocol, outputting initial drafts of application-level core documents that meet regulatory requirements. It also solves problems such as the traditional need for manual data extraction and filling from multiple stages in core document writing, resulting in a large amount of repetitive work; the need for manual verification and modification of document content after TFLs updates, which can easily lead to data inconsistencies; inconsistent document formats requiring manual formatting; and the potential for version confusion and content conflicts in multi-role collaborative writing.

[0073] Furthermore, the step of performing correlation analysis on the statistical analysis results based on the lifetime core document template to generate a structured data pool and a data node association table includes: extracting the corresponding structured data from the statistical analysis results based on the unique identifier of the data embedding node in the lifetime core document template, establishing a one-to-one correspondence between the data embedding node and the underlying data source, and generating a structured data pool and a data node association table.

[0074] This invention is driven by data and builds a full-domain data association graph. It establishes a real-time two-way linkage relationship between raw data, standardized datasets, TFLs and data embedding nodes in documents. The data-driven automated flow will trigger the automatic refresh of documents: data updates at any stage will be passed to the document generation module through automated flow, and the corresponding values, charts and descriptions in the document will be automatically updated synchronously without manual intervention.

[0075] In one embodiment of the present invention, the method of performing full-dimensional visual monitoring, early warning and adjustment of each module based on a preset mode to update the test data declaration results includes: building a multi-dimensional visual monitoring warehouse; identifying and analyzing real-time data in the visual monitoring warehouse based on a preset risk early warning model to obtain risk early warning results; and performing multi-role collaborative processing on the problematic data in the risk early warning results to update the test data declaration results.

[0076] Specifically, the system integrates data interfaces with the four main modules, synchronizing all data across the entire process to a unified data platform for data cleaning and structuring, generating standardized management indicators. The system also establishes a multi-dimensional visual monitoring dashboard, divided into five sections: data management, quality control, statistical analysis, document writing, and project progress. Core indicators for each section are displayed in real-time using charts (bar charts, line charts, pie charts, dashboards), supporting multi-dimensional drill-down analysis. For example, clicking on the abnormal event occurrence rate directly drills down to the corresponding center's original data. Customized visual dashboards are created based on user roles, displaying only the core indicators relevant to that role. The integrated output includes the system's overall data platform, a multi-dimensional visual monitoring dashboard for clinical trials, and role-specific visual dashboards.

[0077] The system identifies abnormal data patterns in real-time data based on a pre-set risk warning model; it then classifies these abnormal data patterns into risk levels, obtaining risk classification results; and triggers in-depth compliance verification for the risk data corresponding to the risk classification results to pinpoint the root cause of the risk, thereby obtaining risk warning results and sending them to the relevant management personnel. The system monitors core indicators in real time, identifying abnormal data patterns (such as a data missing rate of 0 for a certain center, an AE reporting rate lower than the industry average by 2 standard deviations, and abnormally fast group entry speed) based on the risk warning model in the system's built-in intelligent RBQM risk warning engine. The system automatically classifies the identified risks into levels and generates customized risk control suggestions for different risk levels, such as triggering on-site monitoring for high risks, remote data verification for medium risks, and continuous system monitoring for low risks. The system automatically sends multi-channel warnings (system messages, emails, SMS) to the relevant responsible persons and pushes risk information to the visual monitoring dashboard, marking the risk location, level, and handling suggestions. After a risk warning is issued, the system automatically activates the intelligent verification and quality control module to trigger in-depth compliance verification of the risk data and quickly locate the root cause of the risk. After the risk is handled, the system automatically verifies the handling results, lifts the warning, and generates a "Risk Warning and Handling Report". At the same time, it outputs risk warning information, risk level classification results, customized risk control suggestions, risk handling logs, etc.

[0078] Secondly, a standardized problem list is generated based on the problem data and assigned to the corresponding managers; the problem handling results of each manager are obtained from their respective role-specific workbench and system verification is triggered; once verification is successful, the problem status is changed and the corresponding experimental data declaration results are updated. The system automatically generates a standardized problem list based on the problem data and automatically assigns it to the corresponding responsible person based on the data ownership (e.g., center data problems are assigned to CRA, data mapping problems are assigned to DM); the responsible person receives the problem list to be done on their role-specific workbench, processes it online and submits the results, the system automatically verifies the processing results, automatically closes the problem mark after successful verification, and returns it for reprocessing if it fails; the system monitors the entire lifecycle of each problem in the problem list in real time and generates a "Problem Handling Statistics Report" which displays the number of problems, processing rate, average processing cycle, and processing efficiency of each center / role. Simultaneously, it outputs a standardized problem list, problem assignment records, problem handling logs, the "Problem Handling Statistics Report," and a multi-role-specific workbench.

[0079] In addition, the system automatically generates a Gantt chart based on the project plan, monitors the progress of each stage in real time, compares the planned progress with the actual progress, and marks lagging stages. The system automatically analyzes the reasons for the delays (such as data collection delays, too many verification issues, and long problem handling cycles) and generates progress catch-up suggestions (such as increasing data collection personnel and prioritizing high-priority issues). Project managers can adjust the project plan in the system, which automatically updates the Gantt chart and pushes the progress adjustment information to the responsible persons at each stage. Simultaneously, it outputs a real-time project Gantt chart, progress comparison report, analysis of the reasons for progress delays, progress catch-up suggestions, and project plan adjustment records for easy recording and review.

[0080] This invention achieves operational linkage and functional complementarity through an automated command center, forming a closed-loop quality control system covering the entire clinical trial process. For example, if the project management module detects a risk in the data of a certain center, it can automatically link with the intelligent verification module to conduct in-depth compliance verification of the data of that center; if the intelligent verification module detects data anomalies, it can automatically link with the project management module to generate problematic data and assign it to the corresponding responsible person; after the statistical module generates TFLs, it can automatically link with the document generation module to populate them, and at the same time link with the intelligent verification module to perform a second verification of the compliance of the TFLs. Thus, it realizes the full-process linkage control from data standardization, compliance verification, statistical analysis, document generation to project management, with no blind spots in quality control, and the overall quality and compliance of clinical trials are comprehensively guaranteed.

[0081] In one embodiment of the present invention, the data-driven automated management system for the entire clinical trial process includes a data integration and standardization module, an intelligent verification and quality control module, an automated statistical analysis module, an intelligent document generation module, and a comprehensive project management module.

[0082] The data integration and standardization module is used to achieve zero-code CDISC generation; the interface unit connects to EDC, PV and central laboratory data.

[0083] The Metadata mapping engine has a built-in SDTM / ADaM standard library (supporting the latest NMPA / FDA / CDISC standards). Users can adjust the mapping relationship between raw variables and standard variables (Metadata Mapping) through a visual interface without writing code.

[0084] Regarding the transformation engine, based on the mapping relationship, the original data is automatically converted into an SDTM dataset, and further, an ADaM dataset is automatically generated based on SDTM.

[0085] The intelligent verification and quality control module is equipped with multi-standard validators and double-blind / double-signature controls. The multi-standard validators incorporate FDA Validator Rules, NMPA guidelines, and CDISC Validation Checks to perform real-time compliance scanning of the generated datasets. The double-blind / double-signature controls support double bookmarking and electronic signature processes for PDF documents (such as aCRF) to ensure the immutability of data traceability.

[0086] The automated statistical analysis module includes an algorithm library and a TFLs generator for automating the generation of TFLs. The algorithm library contains pre-built commonly used clinical statistical models (such as survival analysis and analysis of variance) and randomization / sample size calculation algorithms.

[0087] TFLs Generator: Reads ADaM data and automatically generates tables, figures, and listings that meet the submission requirements based on preset SAP (Statistical Analysis Plan) templates.

[0088] The document intelligent generation module features an intelligent fill engine and dynamic association technology to automate document generation. The intelligent fill engine automatically fills extracted structured data (such as experimental parameters and statistical results) into document templates like DMP, SAP, CSR, and Protocol. The dynamic association technology maintains a dynamic link between the document content and the underlying data; when statistical data is updated, the corresponding values ​​and charts in the document are automatically refreshed.

[0089] The project's comprehensive management module includes a visual monitoring dashboard, a risk warning mechanism, and an online collaboration platform. The visual monitoring dashboard displays key indicators such as participant enrollment progress and adverse event (AE) rates in real-time using charts. The risk warning mechanism monitors abnormal data patterns based on statistical principles (e.g., data from a particular center being excessively perfect) and automatically triggers quality control alerts. The online collaboration platform supports multi-role (DM, Stat, Medical Writer) collaborative editing of documents and problem-solving.

[0090] Through the configuration method described in this embodiment, the present invention achieves cost reduction and efficiency improvement throughout the entire process, connecting the entire chain from EDC data source to CSR report output. By replacing manual programming with "zero-code" configuration, it significantly shortens the clinical trial data processing cycle. Secondly, the system incorporates the latest validation rules from major global regulatory agencies (NMPA / FDA / PMDA), ensuring data compliance and avoiding secondary errors introduced by manual operation through automation. Furthermore, role collaboration and data transparency cover all positions in clinical trials (data management, statistics, monitoring, medical writing), breaking down information silos and enabling collaborative work based on the same set of real-time data.

[0091] In terms of intelligent risk management, the RBQM concept is introduced, shifting quality management from "post-event verification" to "process monitoring," effectively reducing the risk of clinical trial failure. Simultaneously, this system automates the writing of TFLs and various core documents (DMP, SAP, CSR), greatly freeing up the time of highly skilled personnel for more complex scientific analysis.

[0092] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0093] It should be noted that in the description of this invention, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0094] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0095] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0096] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0097] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0098] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0099] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A data-driven, fully automated management system for the entire clinical trial process, characterized in that, include: The data integration and standardization module intelligently maps and processes the acquired multivariate clinical trial data to obtain standardized data; The intelligent verification and quality control module performs compliance verification and approval processing on standardized data to obtain an approval dataset. The automated statistical analysis module processes the approval dataset based on preset built-in rules to obtain statistical analysis results; The document intelligent generation module, based on a preset standardized template, embeds node-linked analysis of statistical analysis results to obtain experimental data reporting results; The project integrated management module provides full-dimensional visual monitoring, early warning, and adjustment of each module based on preset modes to update the test data reporting results.

2. The fully automated clinical trial management system according to claim 1, characterized in that, The process of intelligently mapping the acquired multivariate clinical trial data to obtain standardized data includes: parsing the multivariate clinical trial data to obtain a structured raw data pool; determining the mapping relationship between raw variables and standard variables based on the basic trial information and the structured raw data pool, and generating an initial mapping table; performing compliance processing on the initial mapping table to obtain a target mapping table; and mapping the structured raw data pool to standardized data based on the acquired transformation mode and the target mapping table.

3. The fully automated clinical trial management system according to claim 1, characterized in that, The aforementioned compliance verification and approval processing of standardized data to obtain an approval dataset includes: acquiring and overlaying institutional verification rules and custom rules corresponding to the test application target to generate a target rule set; based on the target rule set, performing embedded verification and adjustment on the standardized data of each node in the system to obtain a compliance dataset; and performing approval processing on the compliance dataset to obtain an approval dataset.

4. The fully automated clinical trial management system according to claim 1, characterized in that, The intelligent verification and quality control module also includes: pre-verification of manual operations during the verification process to intercept abnormal operations.

5. The fully automated clinical trial management system according to claim 1, characterized in that, The aforementioned process for processing the approval dataset based on preset built-in rules to obtain statistical analysis results includes: extracting requirements from the acquired clinical trial protocols to obtain core requirement information; determining the corresponding statistical model from the built-in algorithm library based on the core requirement information; analyzing the approval dataset based on the statistical model to obtain initial TFLs; personalizing the initial TFLs to obtain customized TFLs; and verifying the compliance and rationality of the customized TFLs to obtain statistical analysis results.

6. The fully automated clinical trial management system according to claim 1, characterized in that, The aforementioned method, based on a pre-defined standardized template, performs embedded node linkage analysis on statistical analysis results to obtain experimental data application results. This includes: fine-tuning the acquired standardized template to obtain a final review core document template; performing correlation analysis on the statistical analysis results based on the final review core document template to generate a structured data pool and a data node association table; filling the structured data pool into the embedded nodes of the final review core document template and generating medical description content to obtain a core document draft; establishing a real-time bidirectional linkage relationship based on the core document draft and the data node association table to obtain a global data association graph and a core document draft containing dynamic association relationships; validating the core document draft containing dynamic association relationships and outputting experimental data application results containing the application-level core document draft.

7. The fully automated clinical trial management system according to claim 6, characterized in that, The process of performing correlation analysis on statistical analysis results based on the lifetime core document template to generate a structured data pool and a data node association table includes: extracting corresponding structured data from the statistical analysis results based on the unique identifier of the data embedding node in the lifetime core document template, establishing a one-to-one correspondence between the data embedding node and the underlying data source, and generating a structured data pool and a data node association table.

8. The fully automated clinical trial management system according to claim 1, characterized in that, The aforementioned method for performing comprehensive visual monitoring, early warning, and adjustment of each module based on a preset mode to update the test data reporting results includes: building a multi-dimensional visual monitoring warehouse; identifying and analyzing real-time data in the visual monitoring warehouse based on a preset risk early warning model to obtain risk early warning results; and performing multi-role collaborative processing on problematic data in the risk early warning results to update the test data reporting results.

9. The fully automated clinical trial management system according to claim 8, characterized in that, The method of identifying and analyzing real-time data in the visual monitoring warehouse based on a preset risk warning model to obtain risk warning results includes: identifying abnormal data patterns in real-time data based on the preset risk warning model; classifying the abnormal data patterns by risk to obtain risk classification results; triggering in-depth compliance verification on the risk data corresponding to the risk classification results to locate the root cause of the risk, so as to obtain risk warning results and send them to the corresponding management personnel.

10. The fully automated clinical trial management system according to claim 8, characterized in that, The aforementioned multi-role collaborative processing of problem data in risk warning results to update test data declaration results includes: generating a standardized problem list based on the problem data and assigning it to the corresponding management personnel; obtaining the problem processing results of each management personnel on their respective role's dedicated workbench and triggering system verification; and changing the problem status and updating the corresponding test data declaration results after verification.