A method and management system for identifying deviations from a clinical trial protocol
By establishing data storage location labels and constructing a multi-model deviation identification library in clinical trials, clinical trial data can be detected and corrected. This solves the problem of reduced identification accuracy caused by data source diversity and interaction errors, and improves the accuracy and transparency of data analysis.
Patent Information
- Application Number
- CN202411110749.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-08-14
AI Technical Summary
In existing technologies, the diversity and complexity of data sources lead to inconsistent data quality, and errors occur during data interaction across multiple platforms, resulting in deviations from clinical trial protocols and reduced accuracy in identification.
By acquiring clinical trial data, establishing data storage location labels, distinguishing between usable and data to be repaired, constructing a clinical trial protocol deviation identification model library, and using predictive models, pattern recognition models, log analysis models, and behavior tracking models for data detection and repair, a data repair sequence is generated to ensure the accuracy and consistency of data analysis.
It improved the accuracy and transparency of clinical trial protocol deviation identification, enhanced the reliability of data analysis, enabled timely identification and correction of deviations, and improved trial efficiency.
Smart Images

Figure CN119028601B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing and analysis technology, and in particular to a method and management system for identifying deviations from clinical trial protocols. Background Technology
[0002] With the rapid development of science and technology, digital technology has permeated all walks of life, bringing revolutionary changes to many fields. In the medical field, especially in clinical trials, the application of digital technology is becoming increasingly widespread. Currently, advanced technologies such as cloud computing, big data, and the Internet of Things provide powerful support for data management, monitoring, and analysis in clinical trials, significantly improving the efficiency and accuracy of clinical trials.
[0003] In clinical trials, digital tools are widely used to identify protocol deviations. Specifically, big data technology enables real-time analysis and mining of massive amounts of clinical trial data, allowing for the timely detection of anomalies and deviations. Cloud computing provides powerful computing capabilities and storage space, making data processing and analysis more efficient. Furthermore, the Internet of Things (IoT) can be used to monitor subjects' physiological data and the trial process in real time, ensuring strict adherence to the protocol. The combined use of these digital tools significantly improves the transparency and traceability of clinical trials, facilitating the timely identification and correction of protocol deviations.
[0004] However, despite the enormous potential of digital methods in identifying protocol deviations in clinical trials, several technical challenges remain. First, the diversity and complexity of data sources lead to inconsistent data quality. Clinical trials involve data interaction across multiple systems and platforms; errors at the data source and during multi-platform data interaction can cause data analysis models trained on clinical data to lose accuracy, thus reducing the accuracy of protocol deviation identification. Summary of the Invention
[0005] The purpose of this invention is to provide a method and management system for identifying deviations in clinical trial protocols, in order to solve the problem that errors in the source of data and in the process of data interaction across multiple platforms can lead to the loss of accuracy in data analysis models trained based on clinical data, thereby reducing the accuracy of identifying deviations in clinical trial protocols.
[0006] In a first aspect, the present invention provides a method for identifying deviations from clinical trial protocols, comprising:
[0007] Step S101: Obtain clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data;
[0008] Step S102: Establish corresponding data storage location labels for each data item in the clinical trial data, store each data item in the clinical trial data to the corresponding data storage location, and perform data simulation use in their respective data storage locations. Label the unusable data to obtain data to be repaired, and label the usable data to obtain usable data.
[0009] Step S103: Construct a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model. Establish data detection rules and generate a data inspection sequence from the usable data obtained through the data detection rules.
[0010] Step S104: Set the usable data as the first data group to be detected, set the data to be repaired and the usable data to form the second data group to be detected, repair the data to be repaired to obtain the repaired data, and set the repaired data and the usable data as the third data group to be detected.
[0011] Step S105: Establish data detection rules. The first, second, and third data groups to be tested are processed according to the data detection rules to obtain a data detection order. The corresponding detection models in the data detection order are retrieved from the clinical trial protocol deviation identification model library. The first, second, and third data groups to be tested are then tested according to the data detection order. The detection results are compared, and based on the comparison results, a clinical trial protocol deviation result is generated.
[0012] Furthermore, the method for identifying deviations from clinical trial protocols provided by the present invention includes step S101: acquiring clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data, including:
[0013] Patient information data: Patient medical records are obtained by extracting data features from medical records. Patient electronic data is also collected through an electronic data acquisition system. The patient electronic data and patient medical records are combined to obtain patient information data.
[0014] Clinical trial operational data: Clinical trial operational data is obtained by converting laboratory records, output information from medical devices, and researchers' records into information data.
[0015] Clinical observation data: Clinical observation data consists of information obtained through direct observation and recording by doctors or researchers, as well as information automatically recorded by observation equipment;
[0016] Other external data: Data and information obtained through third-party research institutions and public databases.
[0017] Furthermore, the method for identifying deviations in clinical trial protocols provided by this invention includes step S102: establishing corresponding data storage location labels for each data item in the clinical trial data, storing each data item in the clinical trial data to its corresponding data storage location, simulating data use in its respective data storage location, labeling unusable data to obtain data to be repaired, and labeling usable data to obtain usable data, including:
[0018] Based on the data types in clinical trial data, a data simulation and validation knowledge base is established. The data simulation and validation knowledge base includes clinical trial simulation usage models, patient information data validation models, clinical observation data simulation usage models, and other external data validation models.
[0019] Establish corresponding data storage location labels for patient information data, clinical trial operation data, clinical observation data and other external data in clinical trial data. The data storage location label includes the folder name and the database table name. Store each clinical trial data in the corresponding label location to obtain a data container corresponding to the data storage location label.
[0020] Establish a backup data container corresponding to the data container, and verify the data in the data container through a data simulation verification knowledge base. Data verification includes reading data, performing simple statistical analysis or data query operations.
[0021] If data is lost or damaged during data repair, a backup data container corresponding to the original data container will be retrieved to replace the lost or damaged data.
[0022] Record any problems encountered during the simulation, including reading errors, formatting issues, or data inconsistencies;
[0023] Based on the results of data verification, data that cannot be used normally or has problems will be marked as data to be repaired. Data to be repaired includes data with incorrect format, missing key information, and data corruption.
[0024] Data that has successfully passed data validation and has not been found to have any problems is marked as usable data.
[0025] Furthermore, the method for identifying deviations from clinical trial protocols provided by this invention also includes:
[0026] Create a data record table to be repaired. The data record table includes the file to be repaired, the data to be repaired, and the location information of the data to be repaired.
[0027] Create a usable data record table, which includes the file containing the usable data, the usable record data, and the location information of the usable data.
[0028] Dedicated data storage location labels are set for patient information data, clinical trial operation data, clinical observation data, and other external data. The data storage location labels include folder name, database table name, and corresponding data container identifier.
[0029] Furthermore, the method for identifying deviations from clinical trial protocols provided by this invention includes step S103: constructing a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model; establishing data detection rules; and generating a data inspection sequence from the obtained usable data through the data detection rules, including:
[0030] Obtain the trial objectives, design principles, operational procedures, and core data points of the clinical trial protocol; identify key data indicators for assessing deviations; and establish data quality standards and scope.
[0031] Based on the quality standards and ranges that the data should meet, the data is matched with historical data to obtain the quality standards and ranges of the data and the corresponding historical data. Feature extraction is then performed on the quality standards and ranges of the data and the corresponding historical data to obtain the data deviation patterns and features.
[0032] Identify key data metrics for deviation detection, translate the identified detection logic into server-side executable code, including numerical range verification and time series comparison, set up automated scripts, and perform deviation detection on newly added data at preset time periods or in real time.
[0033] The initially established detection rules are validated using a dataset with known results. The rule parameters are adjusted based on the validation feedback. The validated detection rules are then integrated into the clinical trial protocol identification deviation model library on the server side. During the clinical trial, data changes are monitored in real time on the server side, and the detection rules are adjusted according to the actual situation.
[0034] Furthermore, the method for identifying deviations from clinical trial protocols provided by the present invention includes step S104: setting usable data as a first data group to be tested, setting the data to be repaired and usable data as a second data group to be tested, repairing the data to be repaired to obtain repaired data, and setting the repaired data and usable data as a third data group to be tested, comprising:
[0035] The first group of data to be tested is established. The server selects records marked as usable data from the database based on the data tags.
[0036] The usable data records are integrated into a dataset, the dataset is named as the first data group to be detected, and the first data group to be detected is stored in the specified server location;
[0037] In the second group of data to be tested, the server selects records marked as data to be repaired and usable data based on the data tags.
[0038] The selected data to be repaired is mixed with usable data, and the mixed dataset contains both normal data and data that needs to be repaired.
[0039] Set the mixed dataset as the second data group to be detected, and store the second data group to be detected in the specified server location;
[0040] The third set of data to be tested is established. The server calls the data repair algorithm to repair the data to be repaired. The repair process includes filling in missing values and correcting erroneous data.
[0041] After the repair is completed, the repaired data is verified. Once the verification is successful, the repaired data is integrated with the original usable data to form a new dataset. This newly integrated dataset is set as the third data group to be detected and stored in a specified location on the server.
[0042] Furthermore, the clinical trial protocol deviation identification method provided by the present invention includes step S105: establishing data detection rules, obtaining a data detection order by applying the data detection rules to the first, second, and third data groups to be tested, retrieving the corresponding detection models from the clinical trial protocol deviation identification model library, detecting the first, second, and third data groups to be tested according to the data detection order, comparing the detection results, and generating a clinical trial protocol deviation result based on the comparison results, including:
[0043] Based on the established data detection rules, the first data group to be detected, the second data group to be detected, and the third data group to be detected are detected one by one;
[0044] For the first, second, and third data sets to be tested, the corresponding detection models are retrieved from the deviation model library identified in the clinical trial protocol, and the tests are performed in the predetermined data testing order.
[0045] After the test is completed, the test results for the first, second and third data groups to be tested are collected and recorded. The test results include the identified potential deviations, detailed information on the abnormal data, and the severity of the deviations. By comparison, common deviations and deviations in specific data groups are identified.
[0046] The analysis examines the improvements in deviation identification between the repaired data in the third data set and the original data in the first and second data sets. Based on the comparison results, a clinical trial protocol deviation report is generated.
[0047] The clinical trial protocol deviation report includes the deviations identified in each group's data, including the type of deviation, the cause of the deviation, and the degree of impact of the deviation. It also includes the impact of corrections on the accuracy of deviation identification and outputs the generated deviation report to a designated location.
[0048] Furthermore, the clinical trial protocol deviation identification method provided by this invention analyzes the improvement in deviation identification of the corrected data in the third test data group and the original data in the first and second test data groups, including:
[0049] Compare the number of deviation points in the first, second, and third data groups to be tested. If the number of deviation points in the repaired data in the third data group decreases, the repair operation is effective.
[0050] Compare the repaired data in the third data group to be tested with the original data in the first and second data groups to be tested;
[0051] By comparing the data before and after the repair, if the number of deviations in the repaired data decreases, then the repair effect is positive.
[0052] By comparing the accuracy of deviation identification in the data before and after repair, the improvement of the repair operation on the accuracy of deviation identification can be evaluated. If the accuracy of deviation identification is high after repair, the repair operation reduces the number of deviations and improves the accuracy of deviation identification in the clinical trial protocol.
[0053] Secondly, a clinical trial protocol deviation identification and management system, using the aforementioned clinical trial protocol deviation identification method, includes a server, a backend control terminal, and a user terminal. The server establishes a communication connection with the backend control terminal and the user terminal. The server includes:
[0054] The server acquires clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data.
[0055] The server-side establishes corresponding data storage location tags for each data item in the clinical trial data, stores each data item in the clinical trial data to its corresponding data storage location, and performs data simulation usage in its respective data storage location. Data that cannot be used is tagged to obtain data to be repaired, and data that can be used is tagged to obtain usable data.
[0056] The server-side constructs a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model. Data detection rules are established, and the usable data obtained is processed through these rules to generate a data inspection sequence.
[0057] The server sets the usable data as the first data group to be detected, sets the data to be repaired together with the usable data to form the second data group to be detected, repairs the data to be repaired to obtain the repaired data, and sets the repaired data together with the usable data as the third data group to be detected.
[0058] The server establishes data detection rules, and obtains the data detection order by applying the data detection rules to the first, second, and third data groups to be detected. It then retrieves the corresponding detection models from the clinical trial protocol deviation identification model library and performs the detection on the first, second, and third data groups according to the data detection order. The detection results are compared, and based on the comparison results, a clinical trial protocol deviation result is generated. The server then sends the clinical trial protocol deviation result to the user.
[0059] The beneficial effects of this invention are as follows: This invention provides a method and management system for identifying deviations from clinical trial protocols. By combining multiple models, this invention can effectively solve errors in data source and multi-platform data interaction, ensuring the accuracy of data analysis and improving the precision of deviation identification. Simultaneously, by analyzing clinical trial data and monitoring subject information in real time, it enhances the transparency and traceability of clinical trials, helping to identify and correct deviations in a timely manner, thereby improving trial efficiency. Attached Figure Description
[0060] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 This is a schematic diagram of the clinical trial protocol deviation identification method provided in an embodiment of the present invention. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The technical solutions provided by various embodiments of this invention will be described in detail below with reference to the accompanying drawings.
[0063] Firstly, please refer to Figure 1 The present invention provides a method for identifying deviations from clinical trial protocols, including:
[0064] In step S101: Acquire clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data;
[0065] The server needs to clearly identify the source of the clinical trial data. The data comes from multiple sources, including but not limited to:
[0066] Patient information data primarily originates from the hospital's medical record system, electronic health records (EHRs), or information directly entered by the patient. This includes key information such as the patient's age, gender, medical history, and family medical history.
[0067] Clinical trial operational data is typically generated and recorded automatically by laboratory technicians, researchers, or medical devices, including the dosage of the investigational drug, the time of administration, and laboratory test results.
[0068] Clinical observation data is directly observed and recorded by doctors and researchers based on patients' clinical manifestations and responses, such as changes in symptoms and physical signs.
[0069] Other external data comes from third-party research institutions, public databases, or relevant literature, and is used to assist in analysis and comparison.
[0070] The server needs to establish a stable and secure data interface to obtain data from the aforementioned sources, specifically involving the following aspects:
[0071] Data format standardization ensures that data from different sources can be stored and processed in a unified format.
[0072] Data encryption and secure transmission: Encryption technologies such as SSL / TLS are used to ensure the security of data transmission.
[0073] Data validation and cleaning: Before the data is received, preliminary data validation is performed to remove invalid or erroneous data and ensure the accuracy and integrity of the data.
[0074] Before data is officially stored, the server needs to perform the following preparatory work: Design a reasonable database schema based on the data type and structure, including tables, fields, and indexes. Estimate the data volume and allocate storage space appropriately to ensure data scalability and accessibility. Develop a data backup plan to prevent data loss or corruption and ensure rapid data recovery if necessary.
[0075] The data acquisition process involves the server sending requests to various data sources, requesting the transmission of relevant data. Upon receiving the data, the server performs initial verification and cleaning to ensure its accuracy and validity. The cleaned data is then stored in a pre-designed database for subsequent processing and analysis.
[0076] Log recording and monitoring are also crucial throughout the data acquisition process. The server needs to record the specific time and content of each operation to facilitate quick problem identification and troubleshooting. Real-time monitoring of data transmission speed, stability, and integrity is essential to ensure the smooth operation of the data acquisition process.
[0077] In step S102: establish corresponding data storage location labels for each data item in the clinical trial data, store each data item in the clinical trial data to the corresponding data storage location, and perform data simulation use in their respective data storage locations. Label the unusable data to obtain data to be repaired, and label the usable data to obtain usable data.
[0078] The server first meticulously categorizes the acquired clinical trial data. This includes, but is not limited to, basic patient information, trial operation records, and observational data; each type of data should be clearly defined and labeled.
[0079] For each type of data, the server assigns a specific data storage location label. The label not only indicates the storage path of the data, but also includes the data type and how it is used.
[0080] Using the previously established data storage location tags, various types of data are accurately written to their corresponding storage locations. Throughout this process, data integrity and consistency must be ensured.
[0081] To verify data availability, the server needs to set up a simulated usage environment that closely resembles a real-world scenario. In this simulation, the server attempts to read and use data stored in various locations. The purpose of this step is to detect issues such as formatting errors, missing values, and outliers. Based on the simulation results, the server determines which data is usable and which has problems. Data that is unusable is labeled "to be repaired," while data that is usable is labeled "usable."
[0082] For data marked "to be repaired," the server records its detailed information and the nature of the problem for subsequent repair operations. Simultaneously, this data is temporarily isolated to prevent interference with subsequent analysis. For data marked "usable," the server further processes and optimizes it to improve its efficiency for later use. This data will then be incorporated into subsequent data analysis workflows.
[0083] Throughout the process, the server meticulously records every operation and result to facilitate rapid identification and resolution of any issues. For any anomalies encountered during the simulation, such as data read failures or format errors, the server employs a robust exception handling mechanism to ensure the smooth operation of the entire process.
[0084] In step S103: Construct a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model; establish data detection rules; and generate a data inspection sequence by passing the obtained usable data through the data detection rules.
[0085] Based on the specific needs of clinical trials and historical data analysis, suitable models are selected to construct a deviation identification model library. These models include, but are not limited to, predictive models, pattern recognition models, log analysis models, and behavior tracking models.
[0086] Historical data is used to train predictive models to predict specific indicators or outcomes and identify deviations from expected results.
[0087] Pattern recognition model: A pattern recognition model is trained using machine learning algorithms to discover hidden patterns and anomalies in the data.
[0088] Log analysis model: Analyze system logs and user behavior logs to identify abnormal operations or system errors.
[0089] Behavioral tracking model: Tracks the sequence of user or system behaviors and detects abnormal behavior patterns.
[0090] Model integration and testing: Integrate the trained models into the deviation recognition model library and conduct comprehensive testing to ensure the collaborative work and accuracy of the models.
[0091] Based on clinical trial guidelines, historical deviation cases, and expert knowledge, a comprehensive set of data detection rules was developed. These rules aim to identify anomalies, errors, or unexpected patterns in the data. These rules were then tested on real-world data, and adjusted and optimized based on the test results to improve the accuracy and efficiency of deviation identification.
[0092] The usable data undergoes necessary preprocessing, such as standardization and noise reduction, to ensure data quality and meet model input requirements. The data is then processed using detection rules to identify deviations. Based on the output of these rules, a data inspection sequence is generated. This sequence indicates the data points and related models that require further inspection, repair, or validation.
[0093] In step S104: the usable data is set as the first data group to be detected, the data to be repaired and the usable data are set to form the second data group to be detected, the data to be repaired is repaired to obtain the repaired data, and the repaired data and the usable data are set as the third data group to be detected.
[0094] First group of data to be tested: The server first sets the data previously marked as "usable" as the first group of data to be tested. This data has been confirmed to be accurate and reliable in the previous verification process, so it can be directly used for subsequent testing and analysis.
[0095] Second group of data to be tested: Subsequently, the server mixes the data to be repaired with the usable data to form a second group of data to be tested. The purpose of this group is to observe the impact of the data to be repaired on the overall data quality in subsequent tests and to attempt to identify potential problems with this data in a mixed environment.
[0096] Data Repair Strategy: Based on previously established data inspection rules and issues identified during simulated use, the server checks each piece of data to be repaired. For different types of errors or anomalies (such as format errors, missing values, logical errors, etc.), the server employs corresponding repair strategies, such as data interpolation, format conversion, and logical validation.
[0097] Data repair execution: During the repair process, the server records each repair operation and result to ensure data traceability and repair accuracy. Once the repair is complete, this data will be considered "repaired data".
[0098] Finally, the server merges the repaired data with the original usable data, setting it as the third set of data to be tested. This set of data will be used for subsequent comprehensive testing and verification to ensure the effectiveness of the repair operation and the overall quality of the data.
[0099] Throughout the data grouping and repair process, the server meticulously records the operations and results of each step, including the specific repair methods and data comparisons before and after the repair. This log information is crucial for subsequent data auditing and issue tracking. The server regularly performs quality control checks on the grouping and repair process to ensure that each step conforms to established standards and specifications. This helps in the timely identification and adjustment of problems, thereby continuously improving the accuracy and efficiency of data processing.
[0100] In step S105: Establish data detection rules, obtain the data detection order by passing the data detection rules on the first data group to be detected, the second data group to be detected, and the third data group to be detected. Retrieve the corresponding detection models in the data detection order from the clinical trial protocol identification deviation model library, and detect the first data group to be detected, the second data group to be detected, and the third data group to be detected according to the data detection order. Compare the detection results, and generate the clinical trial protocol deviation results based on the comparison results.
[0101] Based on the specific needs of clinical trials and historical data, the data detection rules should be further refined and improved. Before practical application, the established data detection rules should be validated to ensure their effectiveness and applicability. This can be accomplished by testing on a small-scale dataset.
[0102] Data grouping and testing rules: The first data group to be tested, the second data group to be tested, and the third data group to be tested are respectively tested according to the data testing rules.
[0103] Based on the importance of the data and the urgency of the repair, the detection order of each data group is determined. For example, the repaired data (third data group to be detected) is detected first to verify the repair effect, followed by the original usable data (first data group to be detected), and finally the data group containing the data to be repaired (second data group to be detected).
[0104] Based on the data detection order, the corresponding detection models are retrieved from the clinical trial protocol deviation identification model library. These models include predictive models, pattern recognition models, log analysis models, and behavior tracking models.
[0105] Following the established testing sequence, the retrieved model is used to test each group of data. During the testing process, the integrity and consistency of the data should be ensured, and any potential deviations or anomalies should be promptly identified and recorded.
[0106] The test results of each group of data were compared and analyzed, with particular attention paid to the differences between the restored data and the original data, and the impact of these differences on the clinical trial protocol.
[0107] Based on the comparison results, clinical trial protocol deviation results are generated. These results include information such as the specific deviations identified, the type and extent of the deviations, and their impact on the clinical trial protocol.
[0108] Specifically, the method for identifying deviations from clinical trial protocols provided by this invention includes step S101: acquiring clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data, including:
[0109] Patient information data: Patient medical records are obtained by extracting data features from medical records. Patient electronic data is also collected through an electronic data acquisition system. The patient electronic data and patient medical records are combined to obtain patient information data.
[0110] Clinical trial operational data: Clinical trial operational data is obtained by converting laboratory records, output information from medical devices, and researchers' records into information data.
[0111] Clinical observation data: Clinical observation data consists of information obtained through direct observation and recording by doctors or researchers, as well as information automatically recorded by observation equipment;
[0112] Other external data: Data and information obtained through third-party research institutions and public databases.
[0113] Specifically, the clinical trial protocol deviation identification method provided by this invention includes step S102: establishing corresponding data storage location labels for each data item in the clinical trial data, storing each data item in the clinical trial data to its corresponding data storage location, simulating data use in each data storage location, labeling unusable data to obtain data to be repaired, and labeling usable data to obtain usable data, including:
[0114] Based on the data types in clinical trial data, a data simulation and validation knowledge base is established. The data simulation and validation knowledge base includes clinical trial simulation usage models, patient information data validation models, clinical observation data simulation usage models, and other external data validation models.
[0115] Establish corresponding data storage location labels for patient information data, clinical trial operation data, clinical observation data and other external data in clinical trial data. The data storage location label includes the folder name and the database table name. Store each clinical trial data in the corresponding label location to obtain a data container corresponding to the data storage location label.
[0116] Establish a backup data container corresponding to the data container, and verify the data in the data container through a data simulation verification knowledge base. Data verification includes reading data, performing simple statistical analysis or data query operations.
[0117] If data is lost or damaged during data repair, a backup data container corresponding to the original data container will be retrieved to replace the lost or damaged data.
[0118] Record any problems encountered during the simulation, including reading errors, formatting issues, or data inconsistencies;
[0119] Based on the results of data verification, data that cannot be used normally or has problems will be marked as data to be repaired. Data to be repaired includes data with incorrect format, missing key information, and data corruption.
[0120] Data that has successfully passed data validation and has not been found to have any problems is marked as usable data.
[0121] Specifically, the method for identifying deviations from clinical trial protocols provided by this invention further includes:
[0122] Create a data record table to be repaired. The data record table includes the file to be repaired, the data to be repaired, and the location information of the data to be repaired.
[0123] Create a usable data record table, which includes the file containing the usable data, the usable record data, and the location information of the usable data.
[0124] Dedicated data storage location labels are set for patient information data, clinical trial operation data, clinical observation data, and other external data. The data storage location labels include folder name, database table name, and corresponding data container identifier.
[0125] Specifically, the clinical trial protocol deviation identification method provided by this invention includes step S103: constructing a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model; establishing data detection rules; and generating a data inspection sequence by applying the obtained usable data to the data detection rules, including:
[0126] Obtain the trial objectives, design principles, operational procedures, and core data points of the clinical trial protocol; identify key data indicators for assessing deviations; and establish data quality standards and scope.
[0127] Based on the quality standards and ranges that the data should meet, the data is matched with historical data to obtain the quality standards and ranges of the data and the corresponding historical data. Feature extraction is then performed on the quality standards and ranges of the data and the corresponding historical data to obtain the data deviation patterns and features.
[0128] Identify key data metrics for deviation detection, translate the identified detection logic into server-side executable code, including numerical range verification and time series comparison, set up automated scripts, and perform deviation detection on newly added data at preset time periods or in real time.
[0129] The initially established detection rules are validated using a dataset with known results. The rule parameters are adjusted based on the validation feedback. The validated detection rules are then integrated into the clinical trial protocol identification deviation model library on the server side. During the clinical trial, data changes are monitored in real time on the server side, and the detection rules are adjusted according to the actual situation.
[0130] Specifically, the method for identifying deviations from clinical trial protocols provided by this invention includes step S104: setting usable data as a first data group to be tested, setting data to be repaired and usable data together to form a second data group to be tested, repairing the data to be repaired to obtain repaired data, and setting the repaired data and usable data together to form a third data group to be tested, including:
[0131] The first group of data to be tested is established. The server selects records marked as usable data from the database based on the data tags.
[0132] The usable data records are integrated into a dataset, the dataset is named as the first data group to be detected, and the first data group to be detected is stored in the specified server location;
[0133] In the second group of data to be tested, the server selects records marked as data to be repaired and usable data based on the data tags.
[0134] The selected data to be repaired is mixed with usable data, and the mixed dataset contains both normal data and data that needs to be repaired.
[0135] Set the mixed dataset as the second data group to be detected, and store the second data group to be detected in the specified server location;
[0136] The third set of data to be tested is established. The server calls the data repair algorithm to repair the data to be repaired. The repair process includes filling in missing values and correcting erroneous data.
[0137] After the repair is completed, the repaired data is verified. Once the verification is successful, the repaired data is integrated with the original usable data to form a new dataset. This newly integrated dataset is set as the third data group to be detected and stored in a specified location on the server.
[0138] Specifically, the clinical trial protocol deviation identification method provided by this invention includes step S105: establishing data detection rules, obtaining a data detection order by applying the data detection rules to the first, second, and third data groups to be detected, retrieving the corresponding detection models from the clinical trial protocol deviation identification model library, detecting the first, second, and third data groups to be detected according to the data detection order, comparing the detection results, and generating a clinical trial protocol deviation result based on the comparison results, including:
[0139] Based on the established data detection rules, the first data group to be detected, the second data group to be detected, and the third data group to be detected are detected one by one;
[0140] For the first, second, and third data sets to be tested, the corresponding detection models are retrieved from the deviation model library identified in the clinical trial protocol, and the tests are performed in the predetermined data testing order.
[0141] After the test is completed, the test results for the first, second and third data groups to be tested are collected and recorded. The test results include the identified potential deviations, detailed information on the abnormal data, and the severity of the deviations. By comparison, common deviations and deviations in specific data groups are identified.
[0142] The analysis examines the improvements in deviation identification between the repaired data in the third data set and the original data in the first and second data sets. Based on the comparison results, a clinical trial protocol deviation report is generated.
[0143] The clinical trial protocol deviation report includes the deviations identified in each group's data, including the type of deviation, the cause of the deviation, and the degree of impact of the deviation. It also includes the impact of corrections on the accuracy of deviation identification and outputs the generated deviation report to a designated location.
[0144] Specifically, the clinical trial protocol deviation identification method provided by this invention analyzes the improvement in deviation identification of the corrected data in the third test data group and the original data in the first and second test data groups, including:
[0145] Compare the number of deviation points in the first, second, and third data groups to be tested. If the number of deviation points in the repaired data in the third data group decreases, the repair operation is effective.
[0146] Compare the repaired data in the third data group to be tested with the original data in the first and second data groups to be tested;
[0147] By comparing the data before and after the repair, if the number of deviations in the repaired data decreases, then the repair effect is positive.
[0148] By comparing the accuracy of deviation identification in the data before and after repair, the improvement of the repair operation on the accuracy of deviation identification can be evaluated. If the accuracy of deviation identification is high after repair, the repair operation reduces the number of deviations and improves the accuracy of deviation identification in the clinical trial protocol.
[0149] Secondly, a clinical trial protocol deviation identification and management system, using the aforementioned clinical trial protocol deviation identification method, includes a server, a backend control terminal, and a user terminal. The server establishes a communication connection with the backend control terminal and the user terminal. The server includes:
[0150] The server acquires clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data.
[0151] The server-side establishes corresponding data storage location tags for each data item in the clinical trial data, stores each data item in the clinical trial data to its corresponding data storage location, and performs data simulation usage in its respective data storage location. Data that cannot be used is tagged to obtain data to be repaired, and data that can be used is tagged to obtain usable data.
[0152] The server-side constructs a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model. Data detection rules are established, and the usable data obtained is processed through these rules to generate a data inspection sequence.
[0153] The server sets the usable data as the first data group to be detected, sets the data to be repaired together with the usable data to form the second data group to be detected, repairs the data to be repaired to obtain the repaired data, and sets the repaired data together with the usable data as the third data group to be detected.
[0154] The server establishes data detection rules, and obtains the data detection order by applying the data detection rules to the first, second, and third data groups to be detected. It then retrieves the corresponding detection models from the clinical trial protocol deviation identification model library and performs the detection on the first, second, and third data groups according to the data detection order. The detection results are compared, and based on the comparison results, a clinical trial protocol deviation result is generated. The server then sends the clinical trial protocol deviation result to the user.
[0155] This invention effectively addresses errors that arise during data source and multi-platform data interaction, which can lead to accuracy loss in data analysis models trained on clinical data, thus affecting the accuracy of clinical trial protocol deviation identification. First, this invention establishes data storage location labels to ensure accurate data storage and consistency. Next, through data simulation, usable and problematic data are distinguished and processed. Then, a deviation identification model library containing multiple models is constructed for in-depth detection of clinical data. Furthermore, data detection rules are established, and multiple sets of data are compared to verify data accuracy. Finally, the effectiveness of the repair operations is confirmed through repair effect evaluation, thereby improving the accuracy of clinical trial protocol deviation identification. This invention ensures the accuracy of the data analysis model and the reliability of clinical trial protocol deviation identification.
[0156] The specific implementation and usage of this invention in the context of digital medical diagnosis are as follows:
[0157] Data acquisition and preprocessing begin with the system collecting clinical trial data from various sources, including patient information, clinical trial operation records, and clinical observation data. Next, each data item is assigned a specific data storage location label to ensure structured data storage. During this process, data quality is also verified; unusable data is marked as requiring repair, while usable data is marked as usable.
[0158] In terms of model library construction and data inspection, the system will build a clinical trial protocol deviation identification model library. This library includes various models such as predictive models and pattern recognition models for in-depth data analysis and inspection. Simultaneously, a set of data inspection rules will be established to perform preliminary checks on previously marked usable data and generate a data review sequence.
[0159] In the data grouping and repair phase, the system sets the usable data as the first group to be tested, and mixes the data to be repaired with the usable data to form the second group to be tested. Then, the data to be repaired is repaired, and the repaired data, together with the usable data, forms the third group to be tested.
[0160] Data detection and result generation: Based on the previously established data detection rules, the system performs detailed detection on the three sets of data and compares the results. The system will then generate a clinical trial protocol deviation report based on the comparison results. The report will detail the deviations identified in each set of data, along with their type, cause, and degree of impact.
[0161] Results Output and Application: This deviation report will be output to a designated location for reference by relevant personnel such as researchers, physicians, or regulatory agencies. In digital medical diagnostic scenarios, this report can be used to monitor the progress of clinical trials in real time, promptly identify and correct deviations, thereby ensuring the accuracy and reliability of trial data.
[0162] In terms of system architecture, this invention also proposes a comprehensive clinical trial protocol deviation identification and management system architecture, including a server-side, a backend control terminal, and a user-side. The server-side is responsible for core tasks such as data processing, model building, and result generation, while the user-side provides researchers with intuitive results displays and reports.
[0163] Through the above steps, this invention can effectively identify and manage deviations in clinical trial data in digital medical diagnostic scenarios, thereby significantly improving the accuracy and reliability of trials.
[0164] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0165] The embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention.
Claims
1. A method for identifying deviations from clinical trial protocols, characterized in that, include; Step S101: Obtain clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data; Step S102: Establish corresponding data storage location labels for each data item in the clinical trial data, store each data item in the clinical trial data to the corresponding data storage location, and perform data simulation use in their respective data storage locations. Label the unusable data to obtain data to be repaired, and label the usable data to obtain usable data. Step S103: Construct a clinical trial protocol deviation identification model library. The clinical trial protocol deviation identification model library includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model. The prediction model is used to identify deviations from the expected results, the pattern recognition model is used to identify hidden patterns and anomalies in the data, the log analysis model is used to identify abnormal operations or system errors, and the behavior tracking model is used to detect abnormal behavior patterns. Establish data detection rules, and use the obtained usable data to generate a data inspection sequence through the data detection rules. Step S104: Set the usable data as the first data group to be detected, set the data to be repaired and the usable data to form the second data group to be detected, repair the data to be repaired to obtain the repaired data, and set the repaired data and the usable data as the third data group to be detected, including: The first group of data to be tested is established. The server selects records marked as usable data from the database based on the data tags. The usable data records are integrated into a dataset, the dataset is named as the first data group to be detected, and the first data group to be detected is stored in the specified server location; In the second group of data to be tested, the server selects records marked as data to be repaired and usable data based on the data tags. The selected data to be repaired is mixed with usable data, and the mixed dataset contains both normal data and data that needs to be repaired. Set the mixed dataset as the second data group to be detected, and store the second data group to be detected in the specified server location; The third set of data to be tested is established. The server calls the data repair algorithm to repair the data to be repaired. The repair process includes filling in missing values and correcting erroneous data. After the repair is completed, the repaired data is verified. Once the verification is successful, the repaired data is integrated with the original usable data to form a new dataset. This newly integrated dataset is set as the third data group to be detected and stored in a specified location on the server. Step S105: Establish data detection rules. Using these rules, determine the data detection order for the first, second, and third data groups to be tested. Retrieve the corresponding detection models from the clinical trial protocol deviation identification model library. Perform detection on the first, second, and third data groups according to the data detection order. Then, sequentially call the corresponding models from the model library to perform detection on each group of data. Collect the test results for each group, including potential deviations, detailed information on abnormal data, and the severity of the deviations; By comparing the three sets of results, common / unique deviations were identified, and the improvement of deviation identification by the repaired data was analyzed. Based on the comparison results, deviations from the clinical trial protocol are generated, including: Based on the established data detection rules, the first data group to be detected, the second data group to be detected, and the third data group to be detected are detected one by one. For the first, second, and third data sets to be tested, the corresponding detection models are retrieved from the deviation model library identified in the clinical trial protocol, and the tests are performed in the predetermined data testing order. After the test is completed, the test results for the first, second and third data groups to be tested are collected and recorded. The test results include the identified potential deviations, detailed information on the abnormal data, and the severity of the deviations. By comparison, common deviations and deviations in specific data groups are identified. The analysis examines the improvements in deviation identification between the repaired data in the third data set and the original data in the first and second data sets. Based on the comparison results, a clinical trial protocol deviation report is generated. The clinical trial protocol deviation report includes the deviations identified in each group's data, including the type of deviation, the cause of the deviation, and the degree of impact of the deviation. It also includes the impact of corrections on the accuracy of deviation identification and outputs the generated deviation report to a designated location.
2. The method for identifying deviations from clinical trial protocols as described in claim 1, characterized in that, Step S101: Obtain clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data, including: Patient information data: Patient medical records are obtained by extracting data features from medical records. Patient electronic data is also collected through an electronic data acquisition system. The patient electronic data and patient medical records are combined to obtain patient information data. Clinical trial operational data: Clinical trial operational data is obtained by converting laboratory records, output information from medical devices, and researchers' records into information data. Clinical observation data: Clinical observation data consists of information obtained through direct observation and recording by doctors or researchers, as well as information automatically recorded by observation equipment; Other external data: Data and information obtained through third-party research institutions and public databases.
3. The method for identifying deviations from clinical trial protocols as described in claim 1, characterized in that, Step S102: Establish corresponding data storage location labels for each data item in the clinical trial data, store each data item in the clinical trial data to its corresponding data storage location, and perform data simulation usage in each data storage location. Label unusable data to obtain data to be repaired, and label usable data to obtain usable data, including: Based on the data types in clinical trial data, a data simulation and validation knowledge base is established. The data simulation and validation knowledge base includes clinical trial simulation usage models, patient information data validation models, clinical observation data simulation usage models, and other external data validation models. Establish corresponding data storage location labels for patient information data, clinical trial operation data, clinical observation data and other external data in clinical trial data. The data storage location label includes the folder name and the database table name. Store each clinical trial data in the corresponding label location to obtain a data container corresponding to the data storage location label. Establish a backup data container corresponding to the data container, and verify the data in the data container through a data simulation verification knowledge base. Data verification includes reading data, performing simple statistical analysis or data query operations. If data is lost or damaged during data repair, a backup data container corresponding to the original data container will be retrieved to replace the lost or damaged data. Record any problems encountered during the simulation, including reading errors, formatting issues, or data inconsistencies; Based on the results of data verification, data that cannot be used normally or has problems will be marked as data to be repaired. Data to be repaired includes data with incorrect format, missing key information, and data corruption. Data that has successfully passed data validation and has not been found to have any problems is marked as usable data.
4. The method for identifying deviations from clinical trial protocols as described in claim 1, characterized in that, Also includes: Create a data record table to be repaired. The data record table includes the file to be repaired, the data to be repaired, and the location information of the data to be repaired. Create a usable data record table, which includes the file containing the usable data, the usable record data, and the location information of the usable data. Dedicated data storage location labels are set for patient information data, clinical trial operation data, clinical observation data, and other external data. The data storage location labels include folder name, database table name, and corresponding data container identifier.
5. The method for identifying deviations from clinical trial protocols as described in claim 1, characterized in that, Step S103: Construct a clinical trial protocol deviation identification model library. This library includes prediction models, pattern recognition models, log analysis models, and behavior tracking models. Establish data detection rules and use these rules to generate a data inspection sequence, including: Obtain the trial objectives, design principles, operational procedures, and core data points of the clinical trial protocol; identify key data indicators for assessing deviations; and establish data quality standards and scope. Based on the quality standards and ranges that the data should meet, the data is matched with historical data to obtain the quality standards and ranges of the data and the corresponding historical data. Feature extraction is then performed on the quality standards and ranges of the data and the corresponding historical data to obtain the data deviation patterns and features. Identify key data metrics for deviation detection, translate the identified detection logic into executable code on the server side, including numerical range verification and time series comparison, set up automated scripts, and perform deviation detection on newly added data according to preset time periods or in real time. The initially established detection rules are validated using a dataset with known results. The rule parameters are adjusted based on the validation feedback. The validated detection rules are then integrated into the clinical trial protocol identification deviation model library on the server side. During the clinical trial, data changes are monitored in real time on the server side, and the detection rules are adjusted according to the actual situation.
6. The method for identifying deviations from clinical trial protocols as described in claim 1, characterized in that, The analysis examines the improvements in deviation identification made to the repaired data in the third data group to be detected, as well as the original data in the first and second data groups to be detected, including: Compare the number of deviation points in the first, second, and third data groups to be tested. If the number of deviation points in the repaired data in the third data group is reduced, then the repair operation is effective. Compare the repaired data in the third data group to be tested with the original data in the first and second data groups to be tested; By comparing the data before and after the repair, if the number of deviations in the repaired data decreases, then the repair effect is positive. By comparing the accuracy of deviation identification in the data before and after repair, we can assess the improvement of the repair operation on the accuracy of deviation identification. If the accuracy of deviation identification is high after repair, the repair operation reduces the number of deviations and improves the accuracy of deviation identification in the clinical trial protocol.
7. A clinical trial protocol deviation identification and management system, using the clinical trial protocol deviation identification method as described in any one of claims 1 to 6, characterized in that, It includes a server, a backend control terminal, and a user terminal. The server establishes communication connections with the backend control terminal and the user terminal. The server includes: The server acquires clinical trial data, which includes patient information data, clinical trial operation data, clinical observation data, and other external data. The server-side establishes corresponding data storage location tags for each data item in the clinical trial data, stores each data item in the clinical trial data to its corresponding data storage location, and performs data simulation usage in its respective data storage location. Data that cannot be used is tagged to obtain data to be repaired, and data that can be used is tagged to obtain usable data. The server-side constructs a clinical trial protocol deviation identification model library, which includes a prediction model, a pattern recognition model, a log analysis model, and a behavior tracking model. Data detection rules are established, and the usable data obtained is processed through these rules to generate a data inspection sequence. The server sets the usable data as the first data group to be detected, sets the data to be repaired together with the usable data to form the second data group to be detected, repairs the data to be repaired to obtain the repaired data, and sets the repaired data together with the usable data as the third data group to be detected. The server establishes data detection rules, and obtains the data detection order by applying the data detection rules to the first, second, and third data groups to be detected. It then retrieves the corresponding detection models from the clinical trial protocol deviation identification model library and performs the detection on the first, second, and third data groups according to the data detection order. The detection results are compared, and based on the comparison results, a clinical trial protocol deviation result is generated. The server then sends the clinical trial protocol deviation result to the user.
Citation Information
Patent Citations
Quality control method, device and equipment for medical data detection, medium and program product
CN115691722A
Clinical test quality control method and equipment for scheme deviation semi-quantitative evaluation
CN116864050A
Communication content monitoring method based on Internet of Things
CN118473902A